Baidu Apollo Go robotaxi operating on public roads in China. The expansion of commercial autonomous-driving fleets turned China into one of the industry's most important real-world testing and deployment environments. Source: Baidu
While Silicon Valley and Detroit dominated much of the early Western narrative, China built a parallel autonomous-driving ecosystem through companies such as Pony.ai, WeRide, Baidu Apollo, AutoX, DeepRoute.ai, and Momenta. China also possessed structural conditions that could support large-scale development and deployment, including enormous urban populations, dense mobility ecosystems, extensive digital infrastructure, a large domestic automotive market, and a rapidly expanding electric-vehicle industry. The electric vehicle served as an increasingly suitable platform because modern EVs relied heavily on electronic control systems, powerful onboard computing, software updates, and digital architectures. By developing electrification and autonomy together, the vehicle underwent a broader transformation into a connected, software-defined, electric, and eventually autonomous platform.
From Laboratory to Fleet
Autonomous driving increasingly became not only a laboratory problem, but a fleet problem. A commercial fleet must demonstrate that it works repeatedly, creating a continuous loop: drive collect data identify failures train or improve the system validate deploy collect more data. The more vehicles operate, the more situations they encounter, creating more opportunities to improve and making autonomous driving increasingly dependent on artificial intelligence.
Then Came Another AI Revolution Just as autonomous driving was learning how difficult the physical world was, OpenAI released ChatGPT in late 2022. The significance for autonomous driving was conceptual: large foundation models demonstrated that a single model could acquire broad capabilities from enormous quantities of data across tasks that had traditionally been handled by more specialized systems. As these models expanded from language into images, video, audio, and multimodal reasoning, an obvious question emerged: could the same idea be applied to the physical world?
From Specialized Models to More Unified Architectures
NVIDIA's Lincoln MKV research vehicle used in the DAVE-2 experiment, which demonstrated end-to-end learning by mapping camera images directly to steering commands in 2016. Source: NVIDIA.
The traditional autonomous vehicle architecture was a chain where sensors collected information, perception systems identified objects, tracking systems followed them, prediction models estimated actions, planning software determined decisions, and control systems executed them. While this modular architecture was understandable and testable component by component, it created boundaries between systems. One emerging approach makes the architecture more unified through end-to-end neural networks, sometimes combined with deterministic safety layers. However, safety-critical autonomous driving cannot operate on capability alone; it requires predictable behavior, validation, fail-safe mechanisms, traceability, and proven handling of rare situations.
The World Model
The next step involves world models, where an autonomous vehicle does not merely recognize the current state of the world, but models how that world may behave and evolve. A world model can allow AI systems to reason about possible futures—such as predicting pedestrian trajectories, turning cars, or blocked lanes—and evaluate potential outcomes before acting. This connects autonomous driving to what is increasingly described as physical AI: artificial intelligence designed to perceive, reason about, and act in the physical world.