AI & Technology
World Labs Atlas: Generating Interactive 3D Worlds and Spatial Intelligence from Any Input

What Is World Labs Atlas?
Atlas is an omni world model for spatial intelligence introduced by Fei-Fei Li's World Labs.
While traditional AI models generate flat 2D images, video clips, or text sequences, Atlas is built to generate, reconstruct, and simulate interactive 3D worlds. Atlas takes any text prompt, single image, or video recording and expands it into an immersive, spatially consistent digital environment.
The model can render any camera angle on demand, enabling users to rotate and move around freely inside generated scenes in 3D.
How to Access Atlas
World Labs is rolling out early access to researchers, creators, and developers working on 3D worlds, robotics, gaming, simulation, and VFX.
You can apply for early access through the official form: Request early access to World Labs Atlas.
What Is Spatial Intelligence?
Spatial intelligence is the ability to understand the geometric, physical, and relational structure of 3D spaces—how objects are positioned, how they occlude each other, how perspective shifts when a camera moves, and how scenes evolve over time.
Pioneered by AI pioneer Dr. Fei-Fei Li and the World Labs team, spatial intelligence bridges the gap between digital generative AI and the physical three-dimensional world.
Core Capabilities of Atlas
Atlas natively operates on text, images, video, and 3D representations through a multimodal autoregressive diffusion transformer architecture:
- Camera-Controlled Generation: Generates photorealistic views and long videos (up to 1 minute at 1440p) with pixel-perfect 3D camera trajectory control from as little as a single reference photo.
- Spatial Reconstruction: Reconstructs real-world scenes from sparse photos (1 to dozens of images), generating both novel camera perspectives and explicit 3D outputs (point clouds and 3D Gaussian splats).
- Space-Time Simulation: Reframes video captures from impossible angles without complex capture studios, enabling Real-to-Sim pipelines for robotics navigation and manipulation.
- Omni Multimodal Synthesis: Follows rich text prompts to create 360-degree panoramas, consistent scene expansions, and navigable visual environments.
How Atlas Works: Multimodal Diffusion Transformer
Atlas combines all inputs into a unified spatial context where images, video frames, and geometry are grounded at 3D coordinates in space.
Instead of treating frames as isolated 2D pixel grids, Atlas maintains persistent geometric grounding. When moving the camera or generating unexplored areas of a room or landscape, Atlas extrapolates what lies beyond while preserving total 3D consistency.
Explicit 3D Outputs: 3D Gaussian Splats & Point Clouds
While many video models output only flat pixels, Atlas generates explicit 3D geometry. It predicts depth, point clouds, and 3D Gaussian splats, allowing generated environments to be exported directly into game engines, robotics simulators, VFX workflows, and interactive web viewers like Marble.
Impact on Robotics, Gaming, and Spatial Computing
Atlas unlocks massive breakthroughs across industries:
- Robotics & Real-to-Sim: Robots can be trained inside high-fidelity simulated 3D environments generated directly from casual smartphone video captures.
- Game Development & VFX: Worldbuilders can design entire virtual scenes with interactive camera control instead of manual 3D modeling from scratch.
- Spatial Computing & XR: Instant generation of navigable 3D spaces from text descriptions or single photos for headsets and immersive browsers.
Key Takeaway
World Labs Atlas marks a major leap in AI—moving from flat 2D content generation to genuine 3D spatial intelligence and interactive world simulation. By turning simple text, images, and videos into persistent navigable worlds, Atlas provides the foundation for the next generation of creative tools, simulation engines, and embodied AI.