Google DeepMind announced in May 2026 that it is integrating Street View data into Project Genie, its general-purpose world model, allowing users to simulate and interact with real-world streets. The feature launched during the Google I/O 2026 developer conference and is rolling out to select Google AI Ultra subscribers in the United States, with global access expected over the following weeks.
The integration allows users to load a real-world location and then modify environmental conditions — changing the weather, adjusting the time of year, or exploring scenarios like heavy snowfall on a specific city block. Google has spent 20 years collecting Street View imagery via camera-equipped cars and individuals carrying tracker backpacks, amassing more than 280 billion images across 110 countries and seven continents.
Key figures behind the project include Jack Parker-Holder, a research scientist on DeepMind’s open-endedness team, Diego Rivas, a product manager at DeepMind, and Jonathan Herbert, director of Google Maps. Parker-Holder highlighted a robotics use case: Genie could simulate rare sunny conditions in London so that robots deployed there are not caught off guard by unusual lighting. Herbert noted that while Genie cannot yet produce a faithful reconstruction of a street, the model demonstrates meaningful spatial continuity — correctly remembering a simulated environment when a user turns 360 degrees.
Genie 3, the underlying model, is already being used to help power a Waymo simulator for training its self-driving vehicles on rare events. The Street View addition could extend that capability to new cities and different agent perspectives beyond a car’s point of view, including those of humans and robots.
The technology has notable limitations. Simulations currently render at video game quality rather than photorealistic detail, and the models are not yet physics-aware — demonstrated by a simulated figure running through cacti without interaction. Parker-Holder estimated that world model accuracy is roughly six to twelve months behind video generation models. Rivas described the feature as still experimental, with accuracy improvements ongoing.
Source: TechCrunch