Axis Robotics Turns Robot Data Into a Compounding Training Engine

Axis Robotics released a new dataset for robot manipulation. Axis Sim Dataset V1 packages more than 50,000 human-teleoperated simulation trajectories across 207 manipulation tasks and more than 60,000 scene variants. The release targets Franka arm manipulation and ranks among the largest open-source simulation datasets in that area.
The dataset has already drawn more than 160,000 downloads on Hugging Face, giving Axis a useful test of demand beyond the announcement itself. Its main value, however, is not the download count; it is the measured effect of adding the data to robot policy training.
Continual pretraining on V1 lifted π0.5 from 83.9% to 88.8% success on LIBERO-Plus. Performance improved as the pretraining data increased from 25% to 100% of the dataset, with the largest gains appearing under camera, sensor-noise, and layout perturbations.
More Data, With More Places To Fail
That result points to the problem Axis is trying to solve: a robot policy must handle changes in its surroundings, sensors, and camera views rather than succeed only inside one carefully arranged demonstration. Simulation gives the company a way to produce broader task coverage and test those conditions before relying on real-robot demonstrations.
Axis describes V1 as one output of a larger data engine built around a hybrid strategy across four data lines: simulation, egocentric data collection, loco-manipulation, and human-gated DAgger post-training. Across those lines, the engine includes more than 4.7 million trajectories across 13 embodiments, more than 200,000 hours of egocentric data, more than 500 hours of loco-manipulation data, and more than 500 hours of human-gated DAgger data.
The company records every task and trajectory on-chain on Base for provenance. That makes the dataset’s history part of the system rather than a detail left to a spreadsheet, which is a sensible choice when training data becomes the product’s central asset.
Axis Robotics co-founder Chris Feng framed the strategy in blunt terms: “The future of Physical AI isn’t a static dataset you download once. It’s an engine that keeps producing the data the model needs next. Scale gets you broad coverage. Diversity keeps the noise unbiased. The closed loop turns every failure into progress. That’s what compounds.”
V2 Raises The Scale
Axis has already started work on V2, which will scale to 1.2 million trajectories across 1,200 tasks. The jump from V1’s 207 tasks to 1,200 tasks is not a small refresh; it is a move toward a training pipeline designed to keep expanding its coverage.
The company is also using the engine for partner-specific work. Axis collaborated with Booster Robotics to build a digital twin and collected more than 42,000 simulation episodes for the Booster workspace.
Those episodes produced a clear gap in the resulting model prior. With 30 real-robot demonstrations, Booster’s model prior reached 87.5% success, compared with 37.5% for an out-of-the-box π0.5.
The result does not erase the role of physical robots, but it shows how simulation data and a small number of real-robot demonstrations can work together. The digital twin supplies a large training base, while the demonstrations connect that prior to the target workspace.
Axis was founded by researchers from UC Berkeley, CMU, Georgia Tech, and SJTU. Its research is advised by Jiachen Li, Assistant Professor at Georgia Tech.
Axis’s pitch is straightforward: robot intelligence improves when data collection becomes a continuous loop instead of a one-time dataset release. V1 provides the first large public demonstration of that approach, V2 sets a much larger target, and the performance numbers suggest that diversity matters most when the world refuses to stay still.
Based on




