Physical AI Begins Where Prediction Ends

Physical AI Begins Where Prediction Ends

One helpful way to look at the current robotics boom is to focus less on what happens when we put AI into machines, and more on how things change when an AI model must deal with the results of its own actions.

A language model can give a wrong answer and simply try again. But if a robot makes a mistake, it might already have dropped a glass, hit a shelf, or put itself in a tricky position for its next move. This difference is key to what sets Physical AI apart from the generative AI systems that have led the tech industry in recent years. Generative AI mostly produces information, while physical systems actually change their environment, which the model then has to sense again almost immediately. Because of this, we can’t just judge intelligence by whether a model understands a scene or gives a good answer. Instead, we have to look at whether perception, reasoning, and action stay in sync as the world changes around—and because of—the machine.

Jensen Huang’s statement at CES 2026 that the “ChatGPT moment for robotics is here” captures the excitement in the industry, but isn’t a perfect technical comparison. ChatGPT showed how a general-purpose model could change how we use digital information when enough data, computing power, and model capability come together. Robotics is now seeing similar benefits, but moving from models that just represent the world to machines that act in it brings new challenges that can’t be solved just by making models bigger. NVIDIA describes Physical AI as systems that understand, reason about, and plan actions in the real world. Its robotics stack now covers models, simulation, data generation, evaluation, and edge computing, instead of treating intelligence as just a single model. In robotics, the model is only one part of a much bigger process. [1]

Once a prediction becomes an action, timing becomes part of intelligence

Think about what happens when a robot is told to clear a table. Today’s multimodal models are getting better at understanding the meaning of that instruction. They can recognize objects they haven’t seen before, tell a drinking glass from a decoration, and guess that dishes should go in a cabinet or dishwasher. But these skills don’t tell the robot how tightly to grip a glass, whether it can safely reach for it, how to react if the glass moves, or how to change its motion if something gets in the way. What seems like a simple task to a person actually involves many types of intelligence, from understanding the goal to making precise movements and constantly checking if reality still matches the plan.

That difference in timing helps explain why Vision-Language-Action models have become a key research area. In 2023, RT-2 showed that a vision-language model trained on huge amounts of internet data could be adapted for robotic control by turning actions into tokens. This was important because it proved that robots could use knowledge learned outside of robotics, instead of having to learn everything through their own costly physical experience. Moving from understanding to action, however, revealed another problem: the part of the model that reasons about tasks doesn’t need to work as fast as the part that controls movement. Google DeepMind’s original RT-2 work showed the value of linking these areas, and newer systems are now being designed around the idea that thinking and moving don’t have to happen at the same speed. [2]

Figure’s Helix is a good example of this design. Its vision-language part works at 7 to 9 times per second, while a smaller visuomotor policy controls movement at 200 times per second. This lets the robot’s intent change slowly, while its body reacts quickly to what its sensors pick up. NVIDIA’s original GR00T N1 used a similar split, with a vision-language ‘System 2’ for reasoning and a diffusion-transformer ‘System 1’ for continuous actions. Google DeepMind is also moving VLA models to run directly on robots with Gemini Robotics On-Device 2, so they don’t have to rely on network connections. These systems differ in important ways, and the field hasn’t settled on a final design. Still, they point toward the same requirement: physical intelligence needs an architecture that matches the timing of the machine, without forcing every component to run at the same pace. [3][4][5]

Robotics cannot rely on the Internet’s data advantage alone

Another big difference between generative AI and robotics shows up during training. Language and vision models had a huge advantage because people had already created massive amounts of text, photos, videos, and other digital content before today’s large-scale models were trained. This meant that models could be trained on data that was already available. Robotics doesn’t have this kind of archive. For robots, useful data includes not just what the environment looks like, but also what actions the machine took, its state, how the scene changed, and other signals that only matter when something physical is being controlled. Getting more of this data usually means a robot, a person, or a simulator has to spend real time making it.

Open X-Embodiment shows both how far the field has come and how big the challenge still is. Google DeepMind and 33 academic labs combined data from 22 types of robots, creating a dataset with over a million episodes, more than 500 skills, and 150,000 tasks. This was enough to prove that knowledge can be shared across different datasets and robot types, but it’s still very different from just collecting another billion documents from the web. Today, model development often mixes different types of experience instead of relying on one source for all training. For example, in July 2026, NVIDIA’s GR00T 1.7 was pretrained on about 32,000 hours of real demonstrations and human data, plus around 8,000 hours of simulated rollouts. The trend is toward carefully mixing data sources so each one fills in gaps the others can’t. [6][7]

A central challenge in Physical AI, then, is finding ways to create useful experience at a cost that allows steady improvement. Real robots give the best evidence about what happens when sensors, motors, materials, and unpredictable environments interact. But collecting this data is hard to scale up and always uses hardware, people, and time. Human video shows a lot about physical behavior, but it doesn’t include the robot’s actions. Teleoperation helps, though it brings back the cost of human demonstrations. The problem becomes as much about data engineering as model training, with the added twist that a sample’s value depends on whether it contains the right state and action information to teach a robot how to respond.

Simulation is becoming a data factory, not a substitute for reality

Simulation changes the economics because physical time is no longer the main limit. A real robot can only do one thing at a time, but many simulated robots can run in parallel, fail over and over without breaking anything, and try situations that would be too costly, risky, or boring in real life. This makes simulation especially useful for reinforcement learning and for creating new versions of tasks based on a small set of real demonstrations. There’s a catch: the policy is learning from a model of reality, not reality itself. A 2026 review in the Annual Review of Control, Robotics, and Autonomous Systems explains that the simplifications in simulation create gaps that make it hard to transfer what’s learned to real robots, even as methods like domain randomization, real-to-sim transfer, and mixed training keep getting better. [8]

Simulation is most useful as a way to explore more situations at a lower cost, with real-world data revealing where the simulator stops being accurate. Domain randomization helps by changing things like how objects look or where they are, so a policy doesn’t rely on a perfect synthetic world. Randomization can reduce some of these gaps, but it cannot compensate for every modeling error or missing physical effect. This is especially clear in tasks that involve a lot of contact: a scene might look realistic, but small differences in friction, flexibility, or how forces move can make a simulated grasp work in the computer but fail in real life.

World foundation models add another layer by providing a second way to create artificial experience. Traditional robotics simulators use clear rules about shapes and physics to predict how things change, while generative world models learn patterns from lots of real-world data and can make up realistic variations or possible futures. NVIDIA’s Cosmos 3 brings together vision reasoning, world generation, and action prediction in one foundation model for physical AI. NVIDIA now uses world models alongside regular simulation and synthetic data, rather than treating them as replacements for physics engines. The two approaches solve different parts of the problem: explicit simulation gives control, repeatability, and access to details that video can’t always show, while generative models can create a wider range of visual and behavioral situations that would be hard to design by hand. [9]

A more likely future is a training loop where real demonstrations, simulation, and generative models work together and learn from each other. A small amount of real-world experience can set the foundation for the system’s behavior. Simulation can then build on that by covering more objects, layouts, and situations. World models can add even more variety or help predict how a scene might change. Testing on real hardware shows where the synthetic data was off, creating new data for the next round of training. Seen this way, ‘sim-to-real’ becomes an ongoing process where different sources of experience contribute something the others can’t provide.

The important breakthrough will be the loop, not the demonstration

This changes how we should view the impressive robotics demos we see online. A humanoid folding clothes, a robot handling a new object, or one following a natural-language command can show real progress in generalization. But these demos don’t tell us how often the robot fails after many tries, how it knows when it’s in an unfamiliar situation, or if it can recover before a small mistake turns into a big one. These questions matter even more as models become more general, since a robot that can do many tasks faces more situations where its training doesn’t fully prepare it. Reliability in Physical AI depends on more than picking the right action in a test. It also depends on the systems around the policy—knowing when to trust an action, watching what happens, and deciding what to do when things go wrong.

Here the comparison with ChatGPT becomes useful again, though maybe for a different reason than people first thought. The breakthrough in generative AI didn’t come from one trick. It came from building a stack that was powerful, scalable, and easy enough for millions to use for tasks the creators never listed out. Robotics is now building its own stack: multimodal models give broad understanding, VLA architectures link that understanding to actions, edge hardware allows fast decisions on the robot, real-world datasets ground the policies, and simulation plus world models create much more experience than hardware alone could. The big question is no longer whether these parts can work in isolation. Connecting them into a development and operating cycle that is affordable, flexible, and reliable enough to move beyond staged demos is the harder problem.

Physical AI begins where prediction stops being the end of the process. When a model’s output actually moves a machine, every action affects what the robot sees next, every mistake changes how recovery works, and every success becomes new evidence for improvement. Companies and research groups may build better models, but the real long-term advantage will likely come from building better learning loops around them—gathering experience, making useful synthetic data, training and testing policies, spotting failures, and feeding real-world results back into the system. AI has gotten very good at representing the physical world; Physical AI will show if those representations still work when the world pushes back.

References

  1. NVIDIA. (2026, January 5). NVIDIA Releases New Physical AI Models as Global Partners Unveil Next-Generation Robots. NVIDIA Newsroom.
  2. Chebotar, Y., & Yu, T. (2023, July 28). RT-2: New Model Translates Vision and Language Into Action. Google DeepMind.
  3. Figure. (2025, February 20). Helix: A Vision-Language-Action Model for Generalist Humanoid Control.
  4. 1. Vadrevu, K. M., & Omotuyi, O. (2025, March 18). Accelerate Generalist Humanoid Robot Development With NVIDIA Isaac GR00T N1. NVIDIA Technical Blog.
  5. Google DeepMind. (2026, July 30). Gemini Robotics On-Device 2. Model Card.
  6. Vuong, Q., & Sanketi, P. (2023, October 3). Scaling Up Learning Across Many Different Robot Types. Google DeepMind.
  7. Llontop, E., & Neel, B. (2026, July 7). Develop Humanoid Robot Policies End-to-End With NVIDIA Isaac GR00T. NVIDIA Technical Blog.
  8. Aljalbout, E., Xing, J., Romero, A., Akinola, I., Garrett, C. R., Heiden, E., Gupta, A., Hermans, T., Narang, Y., Fox, D., Scaramuzza, D., & Ramos, F. (2026). The Reality Gap in Robotics: Challenges, Solutions, and Best Practices. Annual Review of Control, Robotics, and Autonomous Systems, 9, 403–432.
  9. NVIDIA. (2026, May 31). NVIDIA Launches Cosmos 3, the Open Frontier Foundation Model for Physical AI. NVIDIA Newsroom.