SO-101 Pi0.5 ACP Round 1 Rollouts
This dataset contains 120 complete SO-101 episodes used for the first Pi0.5 Advantage-Conditioned Policy (ACP) training round.
Summary
- Episodes: 120
- Frames: 137,200
- Cameras: front and hand-eye
- Autonomous successes: 51
- Successful intervention recoveries: 33
- Failures: 36
Episodes are recorded from the initial state through the terminal outcome. Observations, policy actions, executed actions, intervention state, task text, and episode-level outcomes are retained for value learning and ACP label generation.
Method
Pi0.5 ACP adds a trajectory-value and offline-advantage loop to Pi0.5 training. A value model learns from successful and failed complete episodes. Per-frame n-step advantages are converted to a binary conditioning signal used during Pi0.5 fine-tuning with indicator dropout.
Source code and reproducible workflow: BurningDawn8888/lerobot-pi05-acp.
This is an experimental extension built on Hugging Face LeRobot, not an official Pi0.5 feature.
Related resources
- Value model: ted88168/colorlogo_value_round1
- Fixed-matrix evaluation: ted88168/rollout_colorlogo_stage6_eval
- Round 2 training pool: ted88168/rollout_colorlogo_rl_round2_combined
Limitations
The dataset represents one SO-101 setup, workspace, lighting arrangement, camera configuration, object family, and task distribution. It is intended for robotics research and does not establish deployment safety.
- Downloads last month
- 59