Free tools Windows power users keep installed
One-click scans. No signup required.
Yes—Super Mario has become a real AI evaluation environment, but not a universal intelligence test. A March 2025 experiment from UC San Diego’s Hao AI Lab placed large language and vision-language models inside an emulated version of Super Mario Bros. using the GamingAgent framework. The models interpreted screenshots, generated actions, and tried to keep Mario alive while responding to a changing game state.
Claude 3.7 was reported as the strongest performer in that specific comparison. Claude 3.5 followed, while Gemini 1.5 Pro, GPT-4o, and OpenAI’s o1 struggled more. The more important finding was not simply which model finished furthest. Mario exposed a weakness that ordinary question-and-answer benchmarks often miss: an AI can reason well in the abstract and still fail when perception, timing, action, and latency must work together.
What the Super Mario AI test actually was
The experiment was not performed on a Nintendo Switch, and it was not a human-style player holding a controller. The system used an emulator connected to GamingAgent. Models received screenshots and high-level instructions, then generated Python-code inputs that controlled Mario in the game environment.
The basic loop looked like this:
- The emulator produced the current game frame.
- The model received the frame and an instruction or prompt.
- It selected the next action, such as moving, stopping, or jumping.
- The framework converted the response into executable input.
- Mario moved, the game state changed, and another frame was captured.
- The process repeated until the episode ended or reached its limit.
That makes the result a measurement of a model-agent system, not just the underlying model. Prompt wording, screenshot resolution, frame sampling, action duration, memory, code execution, retries, emulator settings, and API latency can all affect the outcome.
Recommended Free Tools
#1 Best Overall
- Find wonder in the Flower Kingdom in the next side-scrolling Super Mario adventure
- Collect Wonder Flowers for surprising, game-changing effects like pipes coming alive, an enemy stampede, and much, much more
- Choose from the largest cast of characters in a side-scrolling Mario game, including Mario, Luigi, Peach, Daisy and other favorites
- Ease into the action with four different-colored Yoshis and Nabbit who can’t take damage
- Discover new power-ups like Elephant Fruit, which transforms Mario and friends into an elephant that can swing its trunk and spray water
TechCrunch’s March 3, 2025 report also emphasized that the game was not quite identical to the original 1985 release. The safest description is therefore “an emulated version of the NES-era game.” The later public project lists Super Mario Bros. 1985 as a Retro environment, but its ROM, emulator configuration, prompts, action timing, episode length, and scoring protocol should not automatically be assumed to match every earlier test.
Which AI performed best?
In the original reported comparison, Anthropic’s Claude 3.7 performed best, with Claude 3.5 next. Google Gemini 1.5 Pro and OpenAI GPT-4o reportedly struggled, while OpenAI’s reasoning-oriented o1 also performed worse than expected in the real-time setting.
Those results should be read as a historical finding from one experiment—not as a current ranking of AI models. Model versions, APIs, prompts, harnesses, and availability change. The public GamingAgent and LMGame Bench repository now lists support for newer systems, including Claude 4 models, OpenAI o3 and o4-mini, Gemini 2.5 models, Grok 3 Mini, DeepSeek, and Qwen3.
The striking lesson is that a model’s reputation for reasoning does not guarantee success in a fast interactive environment. A model that spends several seconds deliberating may produce a better explanation but miss the jump it needed to make immediately.
Why Mario is useful as an AI benchmark
A platform game looks simple because its controls are simple. That simplicity is precisely what makes it useful for controlled testing.
- Visual grounding: The agent must interpret pixels, identify Mario, recognize platforms and enemies, and estimate what is about to happen.
- Sequential decisions: Every action changes the next state. A poorly timed movement can make the following decision impossible.
- Timing pressure: A correct action delivered too late is still a failure.
- Limited controls: Running, stopping, jumping, and directional movement create a relatively clear action interface.
- Longer-horizon behavior: The agent must survive a sequence of obstacles rather than answer one isolated question.
- Visible failure: Falling into a pit or colliding with an enemy produces an outcome that is easy to inspect.
- Repeatability: An emulator can provide repeatable starting states, logs, and scoring.
Mario therefore tests a combination of perception, short-term planning, feedback, and control. It shares that structure with more serious agent problems, but it is much more constrained than driving, robotics, physical manipulation, or open-ended autonomous work.
Why “more reasoning” can hurt
The reported weakness of o1 in this setting illustrates an important distinction between deliberative reasoning and real-time control. In a static problem, extra reasoning time may improve the answer. In Mario, the world continues moving while the model is thinking.
Performance can depend on:
- How quickly the model responds to each screenshot.
- How many frames are grouped into one action.
- Whether the action is sent to the emulator immediately.
- Whether the model is reacting to a current frame or an obsolete one.
- How much context and planning the harness preserves between turns.
This does not show that reasoning models are generally poor at games. It shows that this particular task rewards a different capability: producing sufficiently good actions quickly and repeatedly. A slower model may be more capable at offline analysis and still be worse at a rapid perception-action loop.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteGamingAgent is part of the result
It is tempting to summarize the experiment as “Claude played Mario better than GPT-4o.” That wording hides an important variable: the framework connecting the model to the game.
The project describes two related uses for its software:
- Direct evaluation of models in standardized game environments.
- A customized gaming harness intended to improve performance through more agentic workflows.
A harness may provide memory, reflection, heuristics, structured prompts, action tools, or other assistance. That can make an agent more effective, but it also means the result measures the model together with the wrapper. Comparing models fairly requires stating whether the harness was enabled and keeping the rest of the protocol constant.
The project subsequently expanded into LMGame Bench and Gaming Agent, described in its repository as an ICLR 2026 project and officially released in June 2025. It supports multiple games, model comparisons, computer-use gaming agents, notebooks, Colab-based reproduction, and custom game integration. Its documented harness modes are true, false, or both.
Rank #2
- Save the Flower Kingdom in the latest side-scrolling Super Mario adventure
- Turn classic Mario side-scrolling gameplay on its head with the addition of Wonder Flowers.
Super Mario is one game in a broader benchmark suite
The public repository lists environments including:
- Sokoban
- Tetris
- 2048
- Candy Crush
- Pokémon Red
- Super Mario Bros. 1985
- Ace Attorney
That variety matters because each game stresses different abilities. Sokoban tests planning under irreversible decisions. Tetris combines timing with spatial arrangement. 2048 requires state tracking and heuristic planning. Pokémon Red adds navigation, dialogue, memory, and long-horizon objectives. Ace Attorney emphasizes reading, evidence selection, and structured reasoning.
A suite is more informative than a single title. A model that performs well in Mario but fails at Sokoban may have good visual control without strong spatial planning. One that handles Pokémon but struggles with fast platforming may have useful memory and tool-use skills but poor latency.
What “benchmark” should mean here
A benchmark is not simply an AI playing a game on video. It is a controlled evaluation procedure with a defined environment, observation format, action interface, scoring rule, and repeatable protocol.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsA credible game benchmark should document:
- The exact model name and version.
- The system and user prompts.
- The screenshot resolution, cropping, and frame rate.
- The number of frames or seconds represented by each action.
- The emulator, ROM, input mapping, and initial state.
- The number of episodes and retries.
- The definition of success.
- Average performance, variance, deaths, and progress.
- Latency from observation to action.
- API cost and compute requirements.
- Whether training data or known walkthroughs may have contaminated the result.
Without those details, “the best AI at Mario” is ambiguous. Best could mean furthest horizontal progress, highest score, most levels completed, fewest deaths, fastest completion, or best average across repeated trials.
The main limitations
Classic games may be contaminated
Super Mario Bros. has been documented, streamed, emulated, and discussed for decades. Models may have encountered screenshots, maps, walkthroughs, speedrun knowledge, or code related to the game. A strong result could partly reflect memorization or prior exposure rather than flexible understanding.
A stronger test would include unfamiliar levels, modified layouts, procedurally generated obstacles, or held-out environments.
The harness can dominate the result
Prompt design, memory, reflection, retries, heuristics, code execution, screenshot cadence, and action batching can all change difficulty. A harness-enabled score should not be treated as interchangeable with a raw-model score.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Emulators are not identical
ROM dumps, emulator cores, frame timings, scaling, input mappings, and save-state settings can produce different behavior. Scores from two “Mario” experiments are not directly comparable unless their environments and protocols match.
The action space is narrow
Mario offers far fewer choices than the physical world. The environment is visually legible, mostly deterministic, and instrumentable. Success therefore demonstrates a form of interactive competence, not general intelligence.
API calls are slow and expensive
Screenshot-by-screenshot evaluation can require many frontier-model requests. The repository warns that high-end model evaluation or deployment may incur substantial API and compute costs. Public code does not mean a free experiment.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to reproduce the public project
The repository documents this basic setup:
git clone https://github.com/lmgame-org/GamingAgent.git
cd GamingAgent
conda create -n lmgame python==3.10 -y
conda activate lmgame
pip install -e .
A direct evaluation is documented as:
python3 lmgame-bench/run.py
--model_name {model_name}
--game_names {list_of_games}
--harness_mode false
To run with the gaming harness:
python3 lmgame-bench/run.py
--model_name {model_name}
--game_names {list_of_games}
--harness_mode true
The project supports super_mario_bros among its game names and documents true, false, and both for --harness_mode. Model names and supported providers can change, so readers should use the repository’s current documentation rather than assume that every historical model remains available.
Rank #3
- Step into a world of wonder with Super Mario Bros. Wonder for the Nintendo Switch 2! Experience a fresh twist on classic side-scrolling Mario gameplay, packed with vibrant visuals, new power-ups, and unexpected surprises around every corner.
- Join Mario, Luigi, and friends as they explore imaginative new worlds filled with dynamic challenges and creative level design. The all-new Wonder Effects transform gameplay in exciting ways—pipes come alive, environments shift, and each level delivers a unique adventure.
ROM and emulator requirements
For Retro environments, users must legally obtain compatible game files. The repository documents importing them with Stable Retro:
python3 -m retro.import /path/to/your/ROMs/directory/
Do not download unauthorized ROMs. A legally obtained ROM and compatible emulator setup are prerequisites.
Provider credentials
The repository documents environment variables for several providers:
export OPENAI_API_KEY={YOUR_API_KEY}
export ANTHROPIC_API_KEY={YOUR_API_KEY}
export GEMINI_API_KEY={YOUR_API_KEY}
export XAI_API_KEY={YOUR_API_KEY}
export DEEPSEEK_API_KEY={YOUR_API_KEY}
API pricing, model names, access rules, and availability change frequently. Check each provider’s current documentation before running an evaluation.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
What a better Mario benchmark would report
For results to be meaningful, researchers should publish more than a leaderboard:
- Environment details: ROM identity, emulator version, frame timing, scaling, input mapping, and initial state.
- Model details: exact model version, API settings, context limits, and system prompt.
- Observation and action protocol: screenshot frequency, resolution, action duration, and whether actions are batched.
- Harness configuration: memory, reflection, heuristics, retries, and tool permissions.
- Outcome statistics: average progress, variance, deaths, completion rate, latency, and cost.
- Baselines: novice and expert human performance where possible.
- Generalization: unfamiliar levels or altered layouts rather than only a famous fixed stage.
- Contamination analysis: evidence that success is not merely replaying memorized game knowledge.
| Design choice | Benefit | Risk |
|---|---|---|
| Screenshot input | Tests visual grounding | Sensitive to resolution and frame rate |
| Gaming harness | Can improve planning and reliability | Measures wrapper engineering as well as the model |
| Classic fixed level | Easy to repeat | More vulnerable to memorization |
| Randomized level | Tests generalization | Harder to standardize |
| Short episode | Cheaper and faster | May miss long-horizon failures |
| Long episode | Tests persistence | More expensive and latency-sensitive |
Is Super Mario replacing traditional AI benchmarks?
No. There is no evidence that Mario has replaced standard language, vision, coding, or reasoning benchmarks. It is better understood as one component of a growing family of interactive-agent evaluations.
Its value is diagnostic. It asks questions static tests often cannot:
- Can the model interpret a changing visual scene?
- Can it maintain a useful state across repeated actions?
- Can it recover after mistakes?
- Can it balance planning against latency?
- Does reflection help, or does it make the agent too slow?
- Does competence transfer across different games?
But game performance can also create misleading impressions of progress. Games are artificial, rules are narrow, and success may not transfer to scientific reasoning, social understanding, safety, robotics, physical manipulation, or long-term autonomous work.
The practical takeaway
Super Mario is a useful stress test for real-time interactive agents. It combines visual perception, sequential decision-making, timing, tool use, and feedback in a setting that can be logged and repeated. That makes it more revealing than a static question in some ways.
It is not an IQ test, a replacement for established benchmarks, or proof that a model understands the world. The 2025 result should be read as a snapshot of one model-and-harness comparison, while LMGame Bench represents the more durable direction: testing agents across multiple environments, with explicit protocols and reproducible software.
The most important question is not simply which AI gets furthest in Mario. It is whether a model can act quickly, recover from failure, and generalize its interactive skills to games and environments it has not already seen.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




