Simulation Rollouts
We showcase how CycleVLA corrects failures arising from the policy's own execution in simulation.
Failure: Misgrasp
❌
Put BBQ sauce in basket
🔄 → ✅
Failure: Wrong task
❌
Open the stove
🔄 → ✅
Failure: Wrong object
❌
Put soup and sauce in basket
🔄 → ✅
Failure: Misgrasp
❌
Put milk in basket
🔄 → ✅
Failure: Wrong bowl
❌
Place bowl on cabinet onto plate
🔄 → ✅
Failure: Wrong task
❌
Open the middle drawer
🔄 → ✅
Real-Robot Rollouts: Natural Failure
We showcase how CycleVLA corrects failures arising from the policy's own execution on an AgileX PiPER arm. Speed 3×
Teapot hanging
❌ Misgrasp
🔄 → ✅
Fruit sorting
❌ Premature movement
🔄 → ✅
Cookware packing
❌ Premature movement
🔄 → ✅
Real-Robot Rollouts: Human-Injected Errors
We stress-test CycleVLA by manually injecting three progressively more challenging error types: Distractor Introduction, Target Displacement, and Target Substitution. Speed 3×
Target Displacement: We perturb the position and orientation of the placement target during interaction.
❌ Perturb mug holder location
🔄 → ✅
❌ Perturb grape and plate locations
🔄 → ✅
❌ Perturb corn can and milk locations
🔄 → ✅
Target Substitution: We replace the target object with an unrelated one and move the original target elsewhere.
❌ Substitute with pink teapot
🔄 → ✅
❌ Substitute with eggplant and sausage
🔄 → ✅
❌ Substitute with pepper and mayonnaise
🔄 → ✅
Distractor Introduction: The target object remains unchanged, but we place a visually similar distractor nearby.
❌ Distract with pink teapot and pink cup
🔄 → ✅
❌ Distract with eggplant and sausage
🔄 → ✅
❌ Distract with pepper and mayonnaise
🔄 → ✅
BibTeX
@article{ma2026cyclevla,
title={CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding},
author={Ma, Chenyang and Lu, Kai and Yang, Guangyu and Liu, Jiuming and Xu, Shitong and Byrne, Bill and Havoutis, Ioannis and Trigoni, Niki and Markham, Andrew},
journal={arXiv preprint arXiv:2601.02295},
year={2026}
}