All articles
AI

From GTA 6 to Tesla: What World Models Could Actually Do

Could an AI get you past a GTA 6 mission—or help your car anticipate trouble? A look at world models, Tesla's simulator and the real-world tests that matter before we trust the predictions.

An intact blue coffee cup sits on a desk while a dark frame behind it shows the cup broken and coffee spilled, imagining a failure before it happens.
Let the failure happen in imagination first. A conceptual illustration for this article on world models.
Also available in中文Español

When GTA 6 comes out, could I hand an AI the mission I’ve been failing for an hour and go make coffee?

Here’s a less forgiving version of that question. On my Monday commute, a truck blocks the view of an intersection. Could my car anticipate someone stepping out from behind it and ease off the accelerator, instead of waiting until the last second to slam on the brakes?

Those are two useful ways into world models. The idea is to learn a practical relationship: given what I can observe now, what might happen if I take this action? Recent work on Odyssey-3 and RoboJEPA moves that idea forward. Whether we should trust the resulting decisions is a harder question.

Could an AI get me past a GTA 6 mission?

First, the calendar matters. As of October 10, 2026, Rockstar lists November 19, 2026 as the release date for GTA VI. The gaming assistant I’m describing here is a possibility, not a feature you can use in GTA 6 today.

Picture a car chase in an open-world game. You keep losing it at the same corner. The walkthrough says to slow down earlier and time the turn, which is helpful right up until you have to do it. An AI that could read the screen, judge speed, steer and try another route might actually help you finish the mission. Perhaps it could take over just that stretch.

There are research prototypes pointing in this direction. Google DeepMind’s SIMA 2, introduced in November 2025, follows user instructions in 3D games by looking at the screen and using virtual keyboard and mouse inputs. Researchers also placed it in new environments generated by the Genie 3 world model to test navigation, instruction following and action.

In that pairing, SIMA 2 is the player; Genie 3 supplies a world to practice in. Building something that looks and responds like a game doesn’t automatically produce an AI that can play it. Finding a target in a new environment is also a long way from finishing GTA 6. The team identifies long tasks, limited memory and precise control as continuing problems.

There’s another detail that gets lost in the excitement: a game already has an engine. It can calculate what happens when you turn or hit a wall. An AI doesn’t necessarily need a separate learned world model to play an existing game. If training inside the actual game is feasible, its own rules are a good place to start. Generated worlds could add more practice situations, but the skills still have to work when you return to the game you care about.

I’d happily take a teammate that handles repetitive travel or helps with the section I’m stuck on. I’d keep the story and the big decisions for myself. That’s a smaller ambition than handing over the entire game, and a feature I can actually imagine wanting.

What could this do for Tesla’s vision-based driving?

Think of the situations an experienced driver would rather never encounter again: a pedestrian appearing from behind a truck, lane markings lost in reflections on a rainy night, a nearby car suddenly cutting across.

A world model offers a way to revisit a situation and vary the driving decisions. What happens if the car slows sooner? What if it keeps going? Repeatedly covering those cases in training and testing could eventually help the person sitting behind the wheel.

Tesla has publicly described work along these lines. In an ICCV 2025 technical talk, Tesla AI lead Ashok Elluswamy demonstrated a neural world simulator that generates future views from multiple cameras, conditioned on past states and driving actions. The driving policy—the system choosing what to do—acts, the simulator produces the next observations, and the loop continues.

The demonstrations included revisiting situations that had caused problems to see how a newer policy would respond, and adding a sudden cut-in to test the system’s reaction. These were Tesla’s own demonstrations. They show how the company uses world models in development; any resulting safety improvement still needs evidence from actual driving.

Much of the benefit could come during training and testing, then reach a customer’s car through a better driving policy. The car doesn’t have to render a little movie of the future every time it turns the wheel. How prediction is used on board depends on the architecture and the available computing power.

Nor can imagination turn an unseen pedestrian into a known fact. A model can’t establish whether someone is behind that truck just by inventing a scene. What I’d want is a system that considers several possibilities and leaves room for uncertainty. From the passenger seat, useful progress might feel like earlier, smoother slowing—and fewer unnecessary hard stops.

Whether Tesla’s vision-based approach makes that progress has to be judged on difficult cases like these. Its consumer FSD (Supervised) still requires active driver supervision. Having a world model doesn’t make the car autonomous.

What do the latest results actually show?

A gaming companion, a driving simulator and a robot planner have different jobs. Their results aren’t interchangeable. Two recent releases make that distinction worth keeping in view.

Odyssey-3: interactive worlds, with conditions attached to the score

In its October 8 announcement, Odyssey opened a research preview of Odyssey-3 Flash. Users can move through generated environments, change their viewpoint and introduce events, then watch the environment respond. That is progress in interactive world generation. It isn’t a demonstration of an AI playing GTA 6.

The Pro version scored 66.1 on Physics-IQ Verified, a video-continuation test that compares generated continuations of real physical experiments with what actually happened. That result used prompt enhancement and best-of-8: generate eight candidates, then select the best. With prompt enhancement but without best-of-8, Pro scored 63.4. The official chart says the standard scores average four runs; the best-of-8 result comes from one run.

Those details belong next to the number. A score of 66.1 isn’t a driving or robot-task success rate, and it isn’t the score of the publicly available Flash preview.

RoboJEPA: better predictions still leave hard physical tasks

RoboJEPA, released by Meta FAIR and collaborators on October 7, predicts future visual features in a compressed representation without rendering every frame. The researchers scaled the predictor from 22 million to 8 billion parameters, using data from 23 public datasets spanning 12 robot platforms. With that data collection and the visual encoder held fixed, more training compute reduced prediction error and improved performance on the planning tasks tested.

Here is how RoboJEPA-8B performed on the physical robot:

Real-world robot taskSuccess rate
Grasp an object67%
Lift and hold it until the episode ends50%
Pick it up and place it at the target27%
Source: RoboJEPA, Table 3. Appendix G.2 specifies 30 episodes per model per task, scored by people. Each task has its own completion criterion; these are not consecutive pass rates through one sequence.

The paper also says a multistep planning call with the 8-billion-parameter model still takes several seconds on a GPU cluster. It isn’t real-time yet. A robot bringing coffee across your living room needs to react in time as well as choose the right action.

You can retry a game. Who pays for a real mistake?

Suppose a model learns that a cup won’t fall even when the gripper barely holds it. A robot might choose a shortcut right over your laptop. It rehearsed the move; everything looked fine. That’s precisely the problem: if a bad action keeps succeeding in simulation, optimization can make the machine more eager to choose it.

The theoretical paper Imperfect World Models are Exploitable studies a crucial kind of mistake: policy A performs better than B in reality, but the model ranks B above A. Under specific mathematical assumptions, the authors investigate when that reversal can be avoided. It has something in common with the reward-hacking problem I discussed in Karotte, though here the error is in predicting what an action will cause.

That’s why I’d first give a world model a role in evaluation. Let it help developers select driving or robot policies worth testing, then check those judgments on real hardware. Screening doesn’t remove model errors. It gives us a chance to catch them before a recommendation becomes a real action.

There is evidence for this narrower use. In research published in February, Runway evaluated eight robot policies using 1,450 simulated rollouts and more than 16,000 human ratings. Across those eight policies, simulated and real-world scores had a Pearson correlation of 0.95. This was a company-reported study of tabletop tasks with a single robot arm. It supports agreement in the relative performance of those policies. It does not mean a 95% success rate or establish anything about driving safety.

A model that helps grade candidates could be useful. But first, someone has to grade the grader.

Start with the coffee cup

Imagine a team comparing three versions of a robot-arm controller. Each has to put a cup on a tray. The team calibrates its evaluation method on real-world records from older versions, then holds back the three new controllers, new cup positions and different fill levels for testing.

Have the model score and rank the candidates first. Record its choices before revealing the real outcomes. Then run repeated physical trials from matched starting conditions on a protected test bench. This first round won’t save much testing money. It has to establish whether the screening works before anyone budgets around the savings.

Track drops, spills, completion and elapsed time separately. Being fast shouldn’t let a controller make up for breaking things.

Once the model starts screening candidates, randomly audit some of the ones it rejects, too. Testing only its favorites won’t reveal how many good options it ruled out. The moment a model decides what deserves a test, it also starts shaping the evidence used to judge it. Skip this check and the screening process can become a machine for confirming its own choices.

A new gripper, cup material or operating environment calls for another check. And do the full cost calculation: are the physical tests saved worth more than the inference, human review and rework caused by wrong judgments? If not, this extra step hasn’t earned its place.

Proposed world-model trial: start with real observations, simulate candidate actions, then check recommendations and randomly sampled rejected candidates on real hardware. Correct the model or narrow its use when predictions fail.
A trial design I’d start with. Physical checks cover recommended candidates and a random sample of rejected ones.

Keep the failures that don’t make the demo reel

One detail from Tesla’s talk stuck with me. A world simulator can learn from poor driving, too: it needs to know what bad actions lead to. Teaching a car to drive well and teaching a model the consequences of driving badly don’t call for exactly the same data.

That sent me back to RoboJEPA’s data table. In the collection the researchers gathered and verified, AgiBot World, associated with Chinese robotics company AgiBot, accounted for 8,036 of 15,022 video hours, or about 53.5%. Multiple camera views can count as separate video hours. Measured by action duration, its share was 2,734 of 6,692 hours, or about 40.9%. Neither share tells us the training sampling weights or the dataset’s causal contribution to model performance. They do show that real interaction data from China is a substantial input to this research.

I’d also want to preserve the less flattering records: the missed grasp, the bad contact, the recovery that did or didn’t work. Aligning actions with sensor logs and making failures available for review and authorized reuse is unglamorous work. It may teach a model more than another carefully chosen successful demo.

Back to those two everyday wishes. In a game, I’d like an AI to save me a few retries. On the road, I’d like dangerous situations to stay in well-tested simulations whenever possible. Those are benefits I’d notice. World models need to earn their place in daily life that way.

Sources checked through October 10, 2026. This analysis draws on papers, official research reports and a technical talk; I have not independently repeated the robot experiments. The GTA 6 assistant, driving scenarios and coffee-cup trial are illustrative uses or proposals. The cover is a conceptual illustration.

Leave a thought

Your email address will not be published.