Skip to main content

LUP Student Papers

LUND UNIVERSITY LIBRARIES

Diffusion Policy for Force- and Vision-Guided Robotic Liquid Pouring Task

Xie, Wenrui LU and Zhang, Jiacheng LU (2026) In LU-CS-EX EDAM01 20252
Department of Computer Science
Abstract
Liquid pouring is a critical, high-precision task in laboratory automation and pharmaceutical manufacturing. Existing imitation learning methods struggle with the high demands for accuracy and robustness in contact-rich pouring tasks. This thesis presents an extension of the diffusion policy framework by integrating force feedback for a dual-arm robot pouring approach. Unlike traditional methods, diffusion-based policy effectively captures the multi-modal action distributions inherent in pouring tasks, where diverse trajectories may lead to successful outcomes. Additionally, a diffusion policy generates temporal action chunks rather than point-wise outputs, which ensures smooth, continuous control suitable for temporally extended pouring... (More)
Liquid pouring is a critical, high-precision task in laboratory automation and pharmaceutical manufacturing. Existing imitation learning methods struggle with the high demands for accuracy and robustness in contact-rich pouring tasks. This thesis presents an extension of the diffusion policy framework by integrating force feedback for a dual-arm robot pouring approach. Unlike traditional methods, diffusion-based policy effectively captures the multi-modal action distributions inherent in pouring tasks, where diverse trajectories may lead to successful outcomes. Additionally, a diffusion policy generates temporal action chunks rather than point-wise outputs, which ensures smooth, continuous control suitable for temporally extended pouring motions. To complement these capabilities with physical awareness, we augment the observation space by directly integrating joint torque signals as proprioceptive inputs together with visual observations. This multimodal observation design allows the policy to condition action generation on visual information, robot joint states, and raw joint torque signals, which provide force-related cues about load changes and physical interaction during pouring. For safety during data collection, rice is used as a proxy medium instead of real liquids. The policy, trained on 150 human demonstrations, was evaluated over 60 trials (20 per configuration) to compare the single-view baseline, multi-view baseline, and force-integrated multi-view policy. A trial is recorded as successful if the robot precisely transfers rice into the target container without missing and both arms return to designated positions. Results show that the force-integrated multi-view policy achieved the highest overall success rate and improved task progression compared with policies without force-feedback input. However, the system is evaluated using rice as a granular proxy medium, and its generalization to real liquid pouring with different fluid properties remains an objective for future work. (Less)
Popular Abstract
Pouring liquids is a common laboratory task, but even a small mistake can waste samples, disrupt experiments, or expose staff to hazardous substances. A useful laboratory robot must therefore do more than move one container above another. It must grasp both containers, coordinate two arms, follow the progress of the pour, and stop at the right moment.

Cameras provide important information about position and alignment, but they do not reveal everything. A gripper can block the view, and a full container may look almost identical to an empty one. This project investigated whether a robot could complement vision with signals already available inside its joints. These joint-torque signals change when the load or physical interaction... (More)
Pouring liquids is a common laboratory task, but even a small mistake can waste samples, disrupt experiments, or expose staff to hazardous substances. A useful laboratory robot must therefore do more than move one container above another. It must grasp both containers, coordinate two arms, follow the progress of the pour, and stop at the right moment.

Cameras provide important information about position and alignment, but they do not reveal everything. A gripper can block the view, and a full container may look almost identical to an empty one. This project investigated whether a robot could complement vision with signals already available inside its joints. These joint-torque signals change when the load or physical interaction changes, giving the robot a simple form of physical awareness without requiring extra force sensors.

The robot learned the task from 150 demonstrations performed by a human operator using virtual-reality controls. The learning method, called a diffusion policy, does not predict only the next isolated movement. Instead, it generates short sequences of coordinated actions, which is useful for a smooth task such as pouring. Rice was used instead of liquid during the experiments. It is safer and still requires controlled flow, accurate alignment, and careful timing, although it does not reproduce liquid effects such as splashing, viscosity, or surface tension.

Three versions of the system were tested 20 times each. A robot using only one overhead camera completed the full task once. Adding two wrist-mounted cameras improved the result to 7 successful trials. When joint-torque signals were also included, the robot reached the pouring stage in 19 of 20 trials and completed the whole task in 9.

One unexpected result was that estimating the exact changing weight of the container from processed force data was too unreliable. The robot's raw joint signals were more useful when supplied directly to the learning system. The approach is not yet fully robust: when the containers collided, unusual torque signals could confuse the policy and cause unexpected motions.

The results show that seeing the task from several angles is important, but vision alone does not always reveal whether a grasp has succeeded or how far pouring has progressed. Combining vision with physical signals could help future robots perform delicate laboratory and manufacturing tasks more safely and consistently. Further work must test larger datasets, different containers, and real liquids before the method can be used in practical laboratory automation. (Less)
Please use this url to cite or link to this publication:
author
Xie, Wenrui LU and Zhang, Jiacheng LU
supervisor
organization
course
EDAM01 20252
year
type
H2 - Master's Degree (Two Years)
subject
keywords
Diffusion Policy, Force Feedback, Robotic Pouring, Laboratory Automation, Imitation Learning
publication/series
LU-CS-EX
report number
2026-53
ISSN
1650-2884
language
English
id
9247448
date added to LUP
2026-08-05 14:10:05
date last changed
2026-08-05 14:10:05
@misc{9247448,
  abstract     = {{Liquid pouring is a critical, high-precision task in laboratory automation and pharmaceutical manufacturing. Existing imitation learning methods struggle with the high demands for accuracy and robustness in contact-rich pouring tasks. This thesis presents an extension of the diffusion policy framework by integrating force feedback for a dual-arm robot pouring approach. Unlike traditional methods, diffusion-based policy effectively captures the multi-modal action distributions inherent in pouring tasks, where diverse trajectories may lead to successful outcomes. Additionally, a diffusion policy generates temporal action chunks rather than point-wise outputs, which ensures smooth, continuous control suitable for temporally extended pouring motions. To complement these capabilities with physical awareness, we augment the observation space by directly integrating joint torque signals as proprioceptive inputs together with visual observations. This multimodal observation design allows the policy to condition action generation on visual information, robot joint states, and raw joint torque signals, which provide force-related cues about load changes and physical interaction during pouring. For safety during data collection, rice is used as a proxy medium instead of real liquids. The policy, trained on 150 human demonstrations, was evaluated over 60 trials (20 per configuration) to compare the single-view baseline, multi-view baseline, and force-integrated multi-view policy. A trial is recorded as successful if the robot precisely transfers rice into the target container without missing and both arms return to designated positions. Results show that the force-integrated multi-view policy achieved the highest overall success rate and improved task progression compared with policies without force-feedback input. However, the system is evaluated using rice as a granular proxy medium, and its generalization to real liquid pouring with different fluid properties remains an objective for future work.}},
  author       = {{Xie, Wenrui and Zhang, Jiacheng}},
  issn         = {{1650-2884}},
  language     = {{eng}},
  note         = {{Student Paper}},
  series       = {{LU-CS-EX}},
  title        = {{Diffusion Policy for Force- and Vision-Guided Robotic Liquid Pouring Task}},
  year         = {{2026}},
}