We publish open benchmarks so the field can measure the thing that matters: whether a policy or agent works outside the lab. Results are posted as they are produced, never before.
Status is shown honestly. A benchmark gets a leaderboard only after trials are run.
Open-source VLA policies evaluated on household tasks captured in real kitchens and laundry rooms, not simulation.
| Tasks | Folding, loading, sorting, wiping |
| Policies under test | π0, OpenVLA, GR00T, Octo |
| Metric | Success rate, 3 trials per task |
Bimanual assembly and packaging tasks from real production lines, scored by the line workers who trained on them.
Join as a partner lab →Computer-use agents on real e-commerce seller operations: listing edits, ad bids, returns and review handling.
Join as a partner lab →Short write-ups from the capture floor and the eval bench. First notes publish with the first benchmark results.
Three questions. We reply within one business day with a sample set and a quote.