Physical AI Is Moving From Demos to Dependable Work

Physical AI is leaving the demo floor. Companies are building systems that must perceive the world, make decisions, and complete work without turning every mistake into a production incident.
Standard Bots claims to be “America’s largest AI-native industrial robot manufacturer” and has raised $200 million at a $1 billion valuation in a Series C round led by General Catalyst and RoboStrategy. Its customers include NASA, Amazon, and Lockheed Martin, while its industrial robot arms target machine tending, welding, and assembly.
The company uses a shared base model that customers adapt through demonstrations and fine-tuning, alongside a range of models including a zero-shot perception system for machine tending. The largest model has parameters in the low billions, and the backbone for machine tending is trained on more than a billion images.
That scale is less important than the way Standard Bots narrows the job. The company focuses on short-horizon tasks because production demands predictable cycle times and reliability, not a robot that can deliver a philosophical monologue before missing the assembly step.
Standard Bots runs inference locally rather than in the cloud, and uses simulation where possible. Some production tasks remain difficult to reproduce in simulators, so deployments also capture failure signals and human corrections; customers decide how those signals and corrections are used.
Leif Jentoft, head of AI at Standard Bots, put the emphasis on training quality: “We believe data quality matters far more than raw volume.” He also said, “Despite claims elsewhere, no model today is truly hardware-agnostic, and co-optimizing low-level control and higher-level functions is a major advantage for both performance and iteration speed.”
Safety Becomes the Physical AI Bottleneck
Nvidia launched Halos in 2025 with hardware and software tools for safety in autonomous vehicles, then expanded the architecture to more physical AI technologies with Nvidia Halos for Robotics in 2026. The supplied timeline places the announcement in April 2026 and identifies the robotics expansion in June 2026.
The system includes the Halos operating system, Nvidia Holoscan Sensor Bridge, simulation and inspection tools, and the Nvidia IGX Thor computing module. Agility Robotics uses Halos in its Digit humanoid robot to work safely around people, while Boston Dynamics has been an early partner focused on Spot, Stretch, and Atlas.
KION Group and LG also work with Nvidia Halos. Nvidia has invested in Agility Robotics and Figure AI, and partnered with Unitree, giving the safety architecture a growing list of hardware environments to support.
Amit Goel, Nvidia’s head of robotics ecosystem and edge computing, described the problem plainly: “Now the AI models are getting capable, the robot hardware is getting capable, and a thing we thought is going to be the next bottleneck is safety.” He added, “Because giving flexibility can come at the cost of losing some control over the stack,” and said, “Now the robot is literally unchained—the safety goes with the robot wherever it goes.”
Jensen Huang, Nvidia’s CEO, described the physical AI business as already driving nearly $10 billion in annual revenue during an appearance on the All-In Podcast in March 2026. That figure puts safety architecture beside model capability as a commercial requirement, not a compliance footnote.
Causal Reasoning and First-Person Data
Aether AI, founded in 2026 by UC San Diego assistant professor Biwei Huang, is pursuing a different route to reliable action. Its CRIS-0 system maintains an explicit causal state and reasons before acting, allowing it to respond to changes instead of blindly repeating a plan.
In a coffee bean task, CRIS-0 recovered from disruptions such as a coffee bag being moved mid-task and replanned in an average of two seconds. In a microwave scenario, it stopped or adjusted within an average of 0.2 seconds after detecting a safety risk, while it chose and moved the intended object with a success rate above 90% in a personalized pick-and-place task.
Aether AI’s causal world model, CausalWM, ranked first on TriWorldBench and first in the robot track of PAI-Bench as of September 2026. Its agent framework, RSIAgent, surpassed frontier closed-source models on OSWorld 2.0 and Agents’ Last Exam, and Aether AI plans to extend this reasoning pattern beyond robotics into forecasting and scientific discovery.
Video is supplying another piece of the puzzle. TwelveLabs released Pegasus 1.6 on October 6, 2026, with support for footage recorded from the perspective of the person or machine performing a task; its Pegasus models are designed for native temporal reasoning and end-to-end video understanding.
Pegasus 1.6 can process approximately 17 years of first-person footage in roughly 18 hours without failure. TwelveLabs prices it at $1.75 per video hour, $3 per million image-input tokens, and $7.50 per million output tokens.
Jae Lee, CEO of TwelveLabs, said, “I think we’re still in the go go go train mode.” The broader direction is clear: reliable physical AI will depend on local inference, task-specific control, safety systems, causal reasoning, and useful data from real deployments—not just larger models waiting for a body.
Based on
- Building AI for Reliable Execution: Lessons From Industrial Robotics — latent.space
- Nvidia’s big bet on physical AI aims for safer robotaxis, humanoid robots – Ars Technica — arstechnica.com
- Aether AI builds causal intelligence for the real world | TechCrunch — techcrunch.com
- TwelveLabs debuts Pegasus 1.6 to improve robotics training data from first-person video | VentureBeat — venturebeat.com




