SO-101 Pi0.5 ACP R2 Targeted Rollouts v1
This public dataset contains 84 reviewed SO-101 rollout episodes collected for the second round of the Pi0.5 Advantage-Conditioned Policy (ACP) reinforcement-learning loop.
Dataset summary
- Episodes: 84
- Frames: 101,893
- FPS: 30
- Cameras:
observation.images.frontandobservation.images.handeye - Robot state/action dimensions: 6
- Results: 43 success, 36 not picked, 2 picked but not placed, and 3 wrong color
Every episode includes robot observations, actions, timestamps, one front-camera video, and one hand-eye-camera video. source_manifest.csv maps each merged episode to its fixed-matrix task, operator-reviewed result, run ID, and original source dataset.
The 84 tasks target failure conditions identified during fixed-matrix comparison of a frozen Pi0.5 baseline and Pi0.5 ACP Round 1. The source trajectories were merged without re-encoding or modifying episode content.
Pi0.5 ACP enhancement
Pi0.5 ACP combines trajectory-value learning, per-frame n-step advantage inference, binary advantage conditioning, indicator dropout, and iterative collection from failed evaluation conditions. The implementation is open source at BurningDawn8888/lerobot-pi05-acp.
This is an experimental extension built on Hugging Face LeRobot, not an official Pi0.5 feature.
Related resources
- Source code: BurningDawn8888/lerobot-pi05-acp
- Stage 6 evaluation: ted88168/rollout_colorlogo_stage6_eval
- Combined Round 2 training pool: ted88168/rollout_colorlogo_rl_round2_combined
Safety and limitations
The dataset reflects one physical robot and workspace. Models trained from it require independent calibration, camera, action-range, reset-pose, and emergency-stop validation before real-robot use.
- Downloads last month
- 115