World Labs introduced Atlas, a new world model the company says can generate, reconstruct, and simulate scenes across text, images, video, and 3D inputs. Techmeme described the announcement as the unveiling of a multimodal world model from Fei-Fei Li’s World Labs, with image and video frame generation, camera control, and 3D reconstruction capabilities. According to World Labs’ announcement, Atlas is an “omni” model pretrained from scratch to operate natively across text, images, video, and 3D. The company describes it as a multimodal autoregressive diffusion transformer: inputs are combined into a shared spatial context, and the model generates what comes next while trying to remain consistent with the scene geometry it has already seen. The core product claim is camera control. World Labs says Atlas can take one or more reference images and generate new views at user-specified camera positions and angles. Instead of relying only on text prompts such as “pan left” or “zoom in,” the company says Atlas uses precise camera geometry as a native input type, giving users more direct control over framing and motion. World Labs says the model can extrapolate beyond visible pixels. In the company’s examples, Atlas generates a complete scene from a single input image, using the visible content plus its broader learned world knowledge to infer unseen portions of the environment. The examples cited in the announcement include imagined backsides of objects and adjacent scene elements that were not present in the original image. The company frames Atlas around “spatial context.” World Labs compares the model’s workflow to a large language model encoding inputs into context before generating outputs, but says Atlas differs because each image is grounded at a 3D position in space. That spatial grounding is what the company says enables consistent scene extension and more deliberate creative control. World Labs also says Atlas can combine unrelated reference images by placing them in a shared 3D context, then generate a world that interpolates between them. The company characterizes this as a way to manage transitions across spaces, including imagined doorways, hallways, and other connective structures between references. For video, World Labs says Atlas can generate long sequences by combining camera movement with spatial-context management. The announcement cites an example of a one-minute video at 1440p resolution generated from a small number of reference images with a hand-designed camera path. The company also says Atlas reconstructs real-world spaces from one or more input images and does not require special capture equipment, though the provided source text does not include independent validation of that claim. Atlas is not being framed as a standalone one-off demo. World Labs says the model will power future versions of Marble and other company products. The announcement also says Atlas is built to scale, with performance improving as training compute increases, and that World Labs expects that scaling trend to continue. Who benefits: Creative users and teams building visual environments could benefit if Atlas makes camera paths, scene extension, and 3D reconstruction more controllable. World Labs also stands to strengthen Marble and other future products with a model it says was built around spatial context from the start. Who's exposed: Existing image and video generation tools may face pressure if users begin prioritizing precise camera control and 3D consistency over prompt-only generation. The degree of exposure is still unclear because the provided items do not include benchmarks, pricing, access details, or third-party tests.