HybridFlow
A 2-NFE Generative Policy for
Real-Time Robotic Manipulation
Abstract
Generative policies often require multiple denoising steps, which increase inference latency in robotic manipulation. HybridFlow combines a one-step MeanFlow proposal, parameter-free ReNoise interpolation, and an instantaneous-velocity refinement using the same network. This three-stage procedure requires only two function evaluations, with no model distillation or separate refinement network. We evaluate HybridFlow in simulation, across five physical manipulation settings, and as a VLA action expert. On Jetson AGX Thor, action generation takes approximately 19 ms, compared with 152 ms for 16-step DDIM.
Overview
Video chapters
A global view.
A precise finish.
MeanFlow supplies both interval-average and
instantaneous velocities. HybridFlow puts
them to work in the same model.
Interactive illustration, not measured action trajectories. The paper uses α = 0.15; both noise sources work in the controlled ablation.
Real-robot experiments
All physical policies run on Jetson AGX Thor. Select a task and method to inspect recorded behavior.
Intercept a car moving at approximately 0.6 m/s and place it in a basket. Each evaluated baseline achieves 0/63.
Clips illustrate behavior; aggregate results below use the full evaluations. Side-by-side playback shows separate recorded trials.
Quantitative results
Physical task performance
Select a task to compare every evaluated policy.
Action-generation latency
At 0.6 m/s, a car travels 9.1 cm during DDIM-16’s action generation, versus 1.1 cm during HybridFlow’s.
Additional videos Simulation and qualitative tasks · 8 recordings
Additional physical demonstrations are qualitative examples, separate from the five evaluated settings.