Every model receives the same game brief, same constraints, and same anti-drift guardrails.
02
Clean workspace
No memory, previous builds, project context, or hidden examples. The run starts from zero.
03
One shot
The model plans, builds, tests, and delivers without iterative human steering.
04
Playable result
The output is judged on whether it loads, plays, communicates its rules, and feels worth trying.
Score method
Public scores expose four use-case pillars. Gameplay blends design, fun, and onboarding.
Adherence is prompt-following fidelity. Engineering combines cross-platform support, implementation quality, and QA.
Art combines visual art, UI/UX, and audio. Final score = Gameplay 45, Adherence 15, Engineering 25, Art 15.
What this measures
Builder capability under a real artifact constraint
Game Bench asks models for a complete game: input handling, state,
rendering, tutorial clarity, play feel, failure states, and enough polish
that a human can actually evaluate it. The idea is to test how models working
at tasks closer to real life usage, not just coding exercises.
Reproducibility
Full frozen prompt
Read the exact prompt every model received: same game brief, same constraints,
same one-shot mandate, same mobile and playtest requirements.