Generalist - GEN-1: Scaling Embodied Foundation Models to Mastery
Generalist Team
Introduction
GEN-1, our latest milestone in scaling robot learning. We believe it to be the first general-purpose AI model that crosses a new performance threshold: mastery of simple physical tasks. It improves average success rates to 99% on tasks where previous models achieve 64%, completes tasks roughly 3x faster than state of the art, and requires only 1 hour of robot data for each of these results. GEN-1 unlocks commercial viability across a broad range of applications—and while it cannot solve all tasks today, it is a significant step towards our mission of creating generalist intelligence for the physical world.
Scaling the Pretraining Era of Embodied Intelligence
Previously, with GEN-0, we showed for the first time that scaling laws exist in robotics. Importantly, it demonstrated that it was possible to scale up robot learning in a generalized way – every zero-shot task we tracked would simultaneously improve. However, its performance was not sufficient to be used in commercial settings. Now with GEN-1, through further scaling of data and compute, and accelerated by algorithmic advances, we are starting to see some tasks cross the level of performance needed to be deployed in economically useful settings.
This parallels what has underpinned progress in large language models (LLMs) as they have been scaled over the past 8 years. GPT-2 showed a scalable path for multitask learning but struggled to be deployed in economically valuable or useful software products. Scaling the model to GPT-3 showed the scaling laws held, new capabilities emerged, and the model became economically viable for certain tasks, such as copywriting for ads. Similarly, GEN-1 can begin to master simple tasks, but the more important concept supported by scaling is that we can expect each new generation of model to result in a new set of increasingly complex tasks that can be mastered.
Notably, this progression also validates the data engine behind these models. Previous general models in robotics that surpass 90% success have depended on enormous teleoperation datasets that are expensive and difficult to scale. Instead, for GEN-0 and GEN-1 the base foundation model is trained without any robot data—it instead uses data from low-cost wearable devices on humans doing millions of activities, and provides an existence proof that this pretraining can lead to high levels of mastery without requiring large teleoperation or simulation datasets.
Capabilities
Reliability
GEN-1 can perform several tasks at high levels of reliability over long durations without intervention. We show here 6 tasks: kitting auto parts for more than an hour, folding t-shirts 86 times in a row, servicing robot vacuums over 200 times in a row, packing blocks over 1,800 times in a row, folding boxes over 200 times in a row, and packing phones over 100 times in a row.
Tasks Performance:
- Kitting Auto Parts: 1 hour without intervention
- T-shirt Folding: 86 times in a row without intervention
- Servicing Robot Vacuum: 200+ times in a row without intervention
- Packing Blocks: 1,800 times in a row without intervention
- Folding Boxes: 200 times in a row without intervention
- Packing Phones: 100 times in a row without intervention
Speed
On two challenging dexterous tasks, GEN-1 enables task completion speeds at roughly ~3x the state of the art. Importantly, GEN-1 can improve task completion speeds to be faster than demonstrations and can react to new object physics at those speeds accordingly. GEN-1 can assemble a box in 12.1 seconds – this is 2.8x faster than prior SOTA.
Improvisational Intelligence
We see a notable shift in how these models respond creatively to unexpected scenarios. In a long-horizon automotive kitting example, if a washer is bumped so far that it’s no longer held properly, the robot can either set it back down to regrasp it or decide to use its other hand to enable bimanual in-hand regrasping.
For physical tasks, improvisation is essential to thrive in unstructured environments.
Limitations
GEN-1 is not without limitations. For instance, while we have shown several dexterous tasks at 99%+ success rates, not all tasks that we have attempted are able to hit these rates. Nevertheless, we expect the next generation of models to unlock a broader range of more complex tasks that can be mastered.
Looking Ahead
Building GEN-1 was not easy—we redesigned our distributed training infrastructure to support petabytes of physical interaction data. We spent months improving training stability and honing post-training techniques. Nevertheless, we believe these advances will lay the groundwork for future research as we continue to scale our data engine into the next phase of capabilities.
Citation
Please cite this work as:
Generalist Team, "GEN-1: Scaling Embodied Foundation Models to Mastery", Generalist AI Blog, Apr 2026.
Or use the BibTeX citation:
@article{generalist2026gen1,
author = {Generalist Team},
title = {GEN-1: Scaling Embodied Foundation Models to Mastery},
journal = {Generalist AI Blog},
year = {2026},
note = {https://generalistai.com/blog/gen-1},
}