Games are becoming a community model benchmark
New models are increasingly being judged by the games they can generate. Game generation was always hard, but doing it well in one shot combines long-horizon coding, visual judgment, spatial reasoning, tool use, and product taste. In a way, AI-generated games have become organic distribution mechanisms for model providers:
While these public demos keep getting better, they're not actually a great judge of capability. Each demo has a different prompt, harness, model (even reasoning effort makes a big difference), token budget, time budget, iteration counts, tool calls, you name it.
At Instaplay, our mission is to make AI the most powerful medium for fun in the world. To do that, we started out by testing various model and harness combinations and seeing what users liked. As more and more models and harnesses were released, we figured the right approach was to let our users decide which models we should use. We believe more and more companies at the app layer (especially consumer AI) will do the same.
Here's what we found when asking our users which models they preferred:
One-shot generation
The first public FunBench season is live. We wrote 24 game requests, covering everything from circuit racers to cozy farming sims, and had every model build each one in a single attempt, with the same agent setup across the board. The finished games then went head to head on Instaplay: a player gets two games built from the same request, plays both, and picks the one they like better, without ever knowing which model made which. As of August 6, 2026, players have cast a little over 3,000 of these votes, and the live leaderboard updates with every new one.
Which games players prefer is only part of the picture. Because we ran every generation ourselves, we also know how often each model actually finished a working game and what each game cost to make, and the voting sessions tell us how long people really played:
- Claude Opus 5 makes the games players like best. It finished every one of its 22 attempts and leads the preference ranking by a wide margin. It is also one of the most expensive ways to get there, at $7.11 in model costs per game.
- GPT-5.6 Sol is the best deal in the field: third place with players, a perfect 23 for 23 on finishing games, and only $1.34 per game.
- Grok 4.5 never failed either, finishing all 27 of its attempts at just $0.86 per game. Players rated its games about average.
- Kimi K3 placed fourth with players, but it only managed to finish 14 of its 24 attempts, the worst completion record on the board.
- Claude Fable 5 came second with players and held their attention longest, at about 30 seconds of actual play per matchup. The rest of the field lands between 9 and 26 seconds.
Two things we learned matter beyond this particular leaderboard. First, when generated games break, they break in a handful of repeatable ways, and the most common one is strange: a model builds a world with perfectly consistent physics, then wires the player's controls in backwards, so steering or aiming is inverted the moment a human takes over. The AI opponents in the same game often drive perfectly. Second, the kind of game matters more than the model for how long people play; racing games simply hold attention longer than farming sims. That is why FunBench only compares games built from the same request. Averaging across genres would mostly reward whichever model happened to draw the most arcade-style prompts.