What I Learned About World Models at a San Francisco Hackathon

Written by

in

I spent a full day, 8 a.m. to 8 p.m., at the Spatial Intelligence and Generative 3D Hackathon at Founders, Inc. in San Francisco. Here is what I built and what a world model is, so the next person exploring this space has a clear starting point.

The Problem I Wanted to Solve

Anyone who has worked with AI video generation runs into the same wall. Every new shot means a new generation, and every new generation means a slightly different set, a slightly different actor, a slightly different light. Shot one and shot two end up looking like two different movies. You can get one beautiful frame, but stitching those frames into a coherent scene is the hard part.

A world model flips that order. Instead of generating a new scene for every shot, you build the environment once, lock it in place, and then film inside it. The set stays fixed. The camera moves. The light stays consistent. That is the whole idea, and it is the reason I picked this track for the hackathon.

Watch the Full Build

I filmed the entire day for my vlog, from the walk into Founders, Inc. and losing my phone in a Waymo, to picking a track, watching the other builds come together, and the final presentation at 8 p.m. It covers the hackathon floor, the track selection process, the Star Wars build step by step:

What I Built

I started with a reference image, a twin sunset scene inspired by Star Wars, generated with Higgsfield. From there I sent it to World Labs, using their Marble world model, and built a full 3D version of that environment. It took a few iterations to get the lighting and geometry to match the reference, but once it clicked, I had a persistent 3D space I could move a camera through freely.

That is the part that stood out to me most. Marble is not producing a single flat image. It is producing a space with depth, one you can walk around, revisit, and edit without starting over. I could change the sky, rerun the render, and everything else in the scene held steady.

Once the world existed, I brought in my own footage and integrated it into that environment, so the final result placed me inside the scene I had built rather than inside something generated fresh for that one shot. I used Tripo AI for a 3D double of myself in the final pass, and leaned on Convex on the backend side to keep the project moving during the build.

Why This Matters for Video Production

Spatial intelligence is the term researchers use for this category of AI. It marks the difference between a model that understands what a scene looks like and a model that understands what a scene is, including its shape, its depth, and how a camera should move through it.

For creators, that distinction is practical, not academic. A persistent world means:

  • Consistent characters and settings across multiple shots
  • The ability to change one element, like lighting or a background object, without regenerating the entire scene
  • Camera movement that behaves like it would on a real set, because the geometry is there

I watched other builders use the same idea in different directions. One hacker built a full simulator starting with a humanoid robot exploring a Jurassic World style environment. Another, Liz, rebuilt an entire Harry Potter scene in 3D that people could walk through and take photos inside, and her project took first place that night. Different goals, same underlying tool.

Where This Fits in My Work

I have been thinking about world models for a while now, long enough that I put together a full breakdown of the space in my book. If you want the deeper version of everything here, including the tools, the workflows, and where this technology goes next for creators, you can find it here:

World Models by Roan Weigert

I will keep building with these tools and sharing what works, alongside the room and community at Founders, Inc. that made this build possible. This hackathon confirmed something I already suspected: the next upgrade in video production is not a better single image generator. It is a model that remembers the world you built and lets you keep working inside it.