The International Top 4
Who left verifiable operating records on live lines
The four, ranked by evidence strength
Mid-2026 verifiable deployment evidence ranks: Agility (65,000+ hours, nine customer sites, GXO/Toyota commercial contracts) > Figure (1,250 h / 90k parts / 99% accuracy at one BMW plant, independently confirmed by BMW) > Boston Dynamics (production electric Atlas delivered to Hyundai and DeepMind, line work from 2028) > Tesla (Optimus line being installed; Musk himself said it is 'not in usage in our factories in a material way'). The ranking measures public operating data, not technical merit.

| Vendor / model | Form / DoF | Payload | Verified operating data | Training-data route |
|---|---|---|---|---|
| Figure 02 to 03 | Bipedal humanoid | ~20-25 kg (02, public) | BMW Spartanburg 1,250+ h, 90k+ parts, 99%+ accuracy, 84 s cycle | Real pilots plus spoken-instruction interaction data |
| Agility Digit | Bipedal with ostrich-leg hybrid | ~16 kg | 100k+ totes at GXO, 65k+ hours, Toyota Canada commercial contract | Continuous fleet feedback from RaaS sites |
| Boston Dynamics Atlas (production electric) | Bipedal humanoid, 56 DoF | 50 kg | 2026 deliveries to Hyundai Metaplant and DeepMind; parts sequencing from 2028 | Joint field-task data with Hyundai / DeepMind |
| Tesla Optimus Gen 3 | Bipedal humanoid | No audited public figure | Fremont line being installed; Musk: not yet used materially in factories | Human motion capture plus imitation learning on Dojo |
Training and test data: three schools
Figure runs a speech-to-action closed loop: the robot hears an instruction, plans a sequence, confirms verbally, executes - pilot data feeds Helix. Agility runs RaaS fleet feedback: every tote path and failure recovery at commercial sites enters the training set. Tesla runs human motion-capture distillation: workers in mocap suits provide demonstrations distilled into neural policies - the cheapest data route but the weakest real-machine validation.

On testing, the 2026 consensus is 'three looks': look at operating hours not demo videos, at placement/grasp accuracy not single-point success rates, at exception recovery (jams, slips, power-resume) not smooth-run performance. BMW's Center of Competence for Physical AI in Production institutionalizes exactly this rubric.