ChristianSafka ↩ The garage
← Blog

Real-Time Neural Rendering for World Models

In my last post I made the case for world models that predict the next world state instead of the next frame.  But if the world isn't stored as pixels, something needs to turn it back into pixels, at least 30 times per second.

Before we get to our experiments, let's review neural rendering:

Traditional computer graphics might represent the world as triangles and do:  scene geometry + materials + lights + camera --> pixels.  Rasterization, ray tracing, shadow calculations, etc. are explicit algorithms.

Neural rendering means that at least one of those pieces is learned by a neural network.  The question becomes "what information does the scene representation explicitly contain, vs implicitly in a neural network?"

Neural Radiance Field (NeRF) takes 3D position and view direction as input and gives you a density and color.  To render a pixel, you shoot a ray and sample many points along it, query the network, and composite the samples.  FPS is constrained by all those neural network calls. 

NeRF example 

3D Gaussian Splatting (3DGS) instead stores millions of little ellipsoids with 3D position, size, orientation, opacity, and appearance learned with gradient descent.  Now during inference we can just project those splats to the image plane based on camera viewpoint, instead of making a bunch of neural network calls.  By moving to this learned explicit representation, we're able to render 3DGS at 100+ FPS.

The more we explicitly define the world representation, the easier the job for the renderer and the faster it can run

World state vs rendering capability

When it comes to building world models, this is where you have to decide what your goals are.  If you want a fast real-time, potentially multiplayer experience, it's unlikely you'll want to spin up a B300 per player to host a large renderer.  We'll start from "fast, real-time" and get closer to the Pareto frontier of this balance as we go.

Experiment #1

I decided to first explore how a small neural renderer could add realism to a gaussian splat scene.  In the image below, we have a town square 3DGS scene in which I inserted a few additional objects.  They look very 'pasted' in the splat render (left frame).   The middle frame is a one-step image network that turns the splat render into a finished frame.  The right frame is AlayaRenderer, which was the larger, slower teacher for our small student network.  AlayaRenderer takes game-engine style buffers and renders a final image, and was trained on AAA games.

World unseen in training.  Left is 3DGS splat input, middle is the render from our student, and right is AlayaRenderer.

The student learned to relight new scenes plausibly, but the output on unseen worlds was more blurry.  When scaling up training the model tended to lean toward copying the 3DGS render.  This may be due to the AlayaRenderer straying quite far in style from the splat render and reshuffling its details every few seconds. 

NVIDIA is launching DLSS 5 this fall, a neural renderer that adds photorealism to video game rendered frames at about 8ms per 4K frame on an RTX 5090.  It uses a very explicit prior world state as fully rendered frames are fed as input, but it shows this kind of renderer is ready for wide-adoption. 


Experiment #2

Here I switched gears to see how far we could get on the other end of the spectrum:  Compact world, smart renderer.  

The renderer is a 22B video model with our LoRA that conditions the generation on given object IDs, depth, and a short text label per object.  

TOP: depth input, BOTTOM: renderer output

I also tried giving this smart renderer our 3DGS render.  The rendered frame is then able to follow geometry and appearance accurately while allowing for object additions and edits at will:

TOP: our 3dgs render, BOTTOM: renderer output with couch label edited to "blue couch" + new object with label "wall clock"

The catch is that on an H100 it takes about 1 minute to generate 4 seconds of video, ~0.6s per frame.  Let's continue and see how much we can get the best of both worlds.


Experiment #3

I wondered if we could bake the photorealistic appearance of the big renderer output to our gaussian splats asynchronously.  The large video model paints 80 clips of camera trajectories in a 3DGS splat world.  The stored world swaps out splat appearance for a 16-dimensional latent that is learned such that the world would reproduce those 80 clips.

The small renderer is a 2M-parameter network that takes those numbers, depth, and viewing direction, and outputs pixels at 7ms per frame.  

TOP: VIDEO MODEL, held-out camera path.  Bottom: real-time renderer reading from the UPDATED SPLAT

There are a multitude of problems with this approach that are obvious in hindsight.  The video model paints fresh details each time, and minor differences would get averaged in our 16-dim latent, resulting in blurry output when read by our tiny renderer.  Plus, our renderer would not be able to make any moving objects or new objects look realistic in real-time.

Neural Harmonic Textures (NVIDIA, ECCV 2026) is a much better engineered version of our gaussian splat reader.  It stores learned features around gaussians and decodes them once per pixel, achieving the quality of a million gaussians with roughly a third as many.  230 FPS on an RTX 4090.

Different GPUs and resolutions, so read it as orders of magnitude


Learnings so far

  • Tiny renderers need a lot of explicit world input to be useful
  • Large renderers are powerful, slow, and any non-determinism hurts when trying to create multi-view data
  • Gaussian splats are messy to work with as an explicit intermediate, especially if you're "lifting" the 3D splats from video
  • From my previous post, we also got the hint that latent vectors without spatial structure are likely not capable of  learning dynamic world state transitions

Next steps

By decreasing how explicit the prior needs to be and slightly increasing the responsibility of the renderer, my intuition is we can retain real-time rendering with an H100 and achieve a more adaptable world with physics and appearance learned from internet-scale video.

Toward this Pareto frontier, one idea is to represent the world as a set of latent vectors with 3D position.  In addition to explicit 3D position, we can learn which vectors matter for each pixel for a mid-size renderer.

More in the upcoming blog post!