Research

Benchmarks on real-world tasks

We publish open benchmarks so the field can measure the thing that matters: whether a policy or agent works outside the lab. Results are posted as they are produced, never before.

Featured benchmarks

Status is shown honestly. A benchmark gets a leaderboard only after trials are run.

In progress · results Q4 2026

Realset Household Manipulation Bench

Open-source VLA policies evaluated on household tasks captured in real kitchens and laundry rooms, not simulation.

TasksFolding, loading, sorting, wiping
Policies under testπ0, OpenVLA, GR00T, Octo
MetricSuccess rate, 3 trials per task
Request early access →
Planned

Light Assembly Bench

Bimanual assembly and packaging tasks from real production lines, scored by the line workers who trained on them.

Join as a partner lab →
Planned

Commerce Ops Agent Bench

Computer-use agents on real e-commerce seller operations: listing edits, ad bids, returns and review handling.

Join as a partner lab →

Research notes

Short write-ups from the capture floor and the eval bench. First notes publish with the first benchmark results.

What 200 hours of egocentric kitchen video taught us about task segmentation

Drafting

De-identification at capture: what it costs and what it does to policy performance

Drafting

Sim-to-real gap on household tasks: benchmark results, first release

Planned

Tell us what your model needs

Three questions. We reply within one business day with a sample set and a quote.

Looking for capture or annotation work? Go to the Experts page instead.

Thanks, you're all set.

Your mail client opened with the request pre-filled. Send it and we'll reply within one business day.