Skip to content
Desmond Zee

dice-rl-wam

DICE-RL post-training of a video-action world model on LIBERO-10

2026-09

Post-trained LingBot-VA, a video-action world model, on LIBERO-10 with DICE-RL: supervised fine-tuning on 300 demonstrations, then residual reinforcement learning against the frozen prior with an ensemble of ten critics. The fine-tuned prior reached 69% under a 200-episode protocol; residual training moved individual tasks strongly in both directions without changing the mean.

Model
LingBot-VA, ~5B video-action world model
Fine-tuning
300 demonstrations, 8× H100
Reinforcement learning
Residual MLP actor, ensemble of ten critics
Result
Prior 69.0%; residual 63.0% mid-run, 69.5% at 100k steps