Home Artificial Intelligence Dyna Robotics Trains DYNA-2 on a Million Hours of Human Video, No Robot Data – Unite.AI

Dyna Robotics Trains DYNA-2 on a Million Hours of Human Video, No Robot Data – Unite.AI

by admin
Dyna Robotics Trains DYNA-2 on a Million Hours of Human Video, No Robot Data – Unite.AI

The claim matters because of where the field’s constraint sits. Generalist robot policies have been trained largely on teleoperated action data: humans physically guiding robots through tasks, hour by hour. That data is expensive and slow to collect, and it has capped how far robot foundation models can scale. DYNA-2’s bet is that the physical intuition robots need can be learned from video of humans doing things, then transferred to robot hardware.

“By building a World-Action Model that imagines how the physical world moves before taking action, we give robots spatial reasoning and contact physics that traditional vision-language models simply lack.”

How DYNA-2 learns from video without seeing a robot

The architecture is what Dyna calls a World-Action Model, built on video generation rather than the vision-language-action adaptations that have dominated the field. During pre-training, the model runs a dual objective (predicting the next frame and the next action) over the human video corpus. The company says this gives the model a working model of contact physics and spatial reasoning that transfers across embodiments: stationary robot arms, humanoid prototypes, and dexterous five-fingered hands, despite none of those robots appearing in pre-training.

Adapting the result to a specific platform takes hours of local fine-tuning rather than weeks of data collection. In one case Dyna cites, 13 minutes of data was enough to teach a pair of five-fingered robot hands to twist open a bottle cap. Aggregated over 15 benchmark tasks, policies pre-trained on more human data consistently outperformed those trained on less, the scaling curve the company is pointing to as the law.

The clearest comparison is against Dyna’s own prior model. DYNA-1, the vision-language-action model running in the company’s commercial deployments, faced DYNA-2 in head-to-head physical evaluations under matched training steps and datasets. On dexterous tasks like chopping food and clearing workspaces, Dyna says DYNA-2 recovered from physical disturbances without human intervention, where the VLA baseline failed and needed manual reset. A video co-training algorithm lifted instruction-following scores by 133% on tasks requiring distinct motions in response to user commands.

DYNA-2 by the numbers

  • 1 million+ hours of egocentric human video in pre-training — described by the company as roughly 170 years of continuous waking experience
  • 20% → 80–90% task success rates on high-precision manufacturing tasks, from pre-training scale alone
  • 1.55x more tasks completed than DYNA-1 in real-world head-to-head evaluations
  • 87% vs. 46% pass rates at a customer deployment, DYNA-2 against DYNA-1
  • 13 minutes of data to train five-fingered hands to open a bottle cap
  • 133% improvement on instruction-following tasks from the video co-training algorithm
  • 10 million hours: the training-data scale Dyna says the approach opens a path toward

Dyna’s deployments were already the test bed

DYNA-2 arrives on top of an unusually concrete commercial base for a company this young. Dyna’s robots, running DYNA-1, are deployed in production at hotels, restaurants, laundromats, and gyms. The company’s own materials describe DYNA-1 folding more than 40 shirts per hour continuously and running sixteen hours a day at customer sites, with a 99%-plus success rate over 24-hour non-stop operation.

That deployment footprint is also the data flywheel behind the research.

The company’s stated path from here is scale: if the human-video scaling curve holds, Dyna says, training on 10 million hours becomes a matter of collecting video rather than building fleets of teleoperation rigs. The 1-million-hour result is the evidence the curve exists. Whether it holds at 10x is the open engineering question — and it’s now the one Dyna has staked its roadmap on.

Source Link

Related Posts

Leave a Comment