PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchClaude 3.7 Sonnet was the top performer in a Super Mario Bros. comparison reported by Hao AI Lab in early 2025. That result came from a custom emulator-and-agent setup, not a standard console speedrun or a universal test of intelligence. Claude 3.5 reportedly followed, while Gemini 1.5 Pro and GPT-4o struggled in the same experiment.
The careful conclusion is narrower: Claude 3.7 appeared particularly well suited to that combination of screenshots, prompts, action-code generation and timing. Later game-agent benchmark results produced a different ranking, so the demonstration does not show that Claude 3.7 was the best game-playing AI—or the best AI generally.
What happened in the Claude 3.7 Mario test?
Hao AI Lab, associated with researchers at the University of California, San Diego, publicized the comparison around late February and early March 2025. Contemporary coverage from TechCrunch and BGR described Claude 3.7 Sonnet as the strongest model in the reported setup, with Claude 3.5 next. Google Gemini 1.5 Pro and OpenAI GPT-4o reportedly had more difficulty. OpenAI o1 was also discussed as an example of a strong reasoning model that did not obviously translate into strong performance in this fast, interactive task.
The comparison was not an esports match, a human-equivalent playthrough or a standardized console benchmark. It was an evaluation of language-and-vision models connected to an emulated game through the GamingAgent framework.
#1 Best Overall
- A variety of playable characters are available, some with unique attributes that affect gameplay and platforming physics.
- Younger and less-experienced players will love playing as Toadette, who is brand new to both games, and Nabbit, who was formerly only playable in New Super Luigi U. Both characters offer extra assistance during play.
- Multiplayer sessions are even more fun, frantic, and exciting thanks to entertaining character interactions. Need a boost? Try jumping off a teammate’s head or getting a teammate to throw you.
- Features a wealth of help features, like a Hints gallery, reference videos, and a Super Guide in New Super Mario Bros. U that can complete levels for you if they’re giving you trouble.
- Three additional modes—Boost Rush, Challenges, and Coin Battle—mix up gameplay and add replayability, while also upping the difficulty for players who want to try something harder. Players can use their Mii characters in these modes.
How the AI controlled Mario
The setup was a closed-loop agent system:
Emulator → screenshot or state → model → action code → emulator
- The 1985 version of Super Mario Bros. ran in an emulator.
- GamingAgent supplied the model with screenshots and basic instructions or observations.
- The model interpreted Mario’s position, platforms, enemies and obstacles.
- It generated control actions, reportedly through Python code, such as moving or jumping.
- The emulator executed those actions and returned a new observation.
- The model repeated the process until it advanced, died or was reset.
This is closer to an LLM or vision-language model operating through an evaluation harness than to a reinforcement-learning bot trained from scratch. It is also different from a person holding a controller: the model’s available actions, observation rate, prompt format and tool interface were determined by the framework.
It was an emulated 1985 game
GamingAgent documents support for Super Mario Bros. 1985. That does not mean the test reproduced every condition of original Nintendo hardware. Emulator timing, frame sampling, ROM, controls, prompt wording and action granularity can all change the outcome. It should not be conflated with a modern Mario release, Super Mario Maker or a browser clone.
Why Super Mario is a useful AI test
The game looks simple, but success requires a repeated interaction loop:
Recommended Free Tools
observe → interpret → decide → act → observe again
Rank #2
- A variety of playable characters are available, some with unique attributes that affect gameplay and platforming physics.
- Younger and less-experienced players will love playing as Toadette, who is brand new to both games, and Nabbit, who was formerly only playable in New Super Luigi U. Both characters offer extra assistance during play.
- Multiplayer sessions are even more fun, frantic, and exciting thanks to entertaining character interactions. Need a boost? Try jumping off a teammate’s head or getting a teammate to throw you!
- Features a wealth of help features, like a Hints gallery, reference videos**, and a Super Guide in New Super Mario Bros. U that can complete levels for you if they’re giving you trouble.
- Three additional modes—Boost Rush, Challenges, and Coin Battle—mix up gameplay and add replayability, while also upping the difficulty for players who want to try something harder. Players can use their Mii characters in these modes!
- Visual scene interpretation: identify Mario, enemies, platforms and gaps.
- Timing: press or release controls at the right moment.
- Distance estimation: judge whether a jump will reach a platform.
- Collision avoidance: distinguish a safe route from an obstacle.
- Short-horizon planning: choose the next action without waiting for a complete level plan.
- Adaptation: change behavior after a missed jump or death.
- Memory: retain useful information about level structure.
- Latency management: act before a moving threat reaches Mario.
A model can describe the correct move and still fail if its response arrives too late, its action syntax is unreliable or its visual estimate is slightly wrong. Those properties are largely invisible in a static question-and-answer benchmark.
Why Claude 3.7 may have had an edge
No independently audited causal analysis establishes one reason for the result. Several factors could have contributed:
Fast, usable action selection
Platform games reward timely, compact decisions. A model that selects a plausible action quickly can outperform one that spends longer producing a more elaborate solution.
Visual-to-action mapping
Claude 3.7 may have been effective at turning a screenshot into a practical instruction: keep moving, brake, jump or wait. Recognizing Mario’s location is not enough; the model must map that recognition to a control sequence.
Output and harness compatibility
Prompt formatting, screenshot handling, allowed tools, action duration and the syntax expected by GamingAgent can favor one model over another. A model that reliably emits executable actions has an advantage even if another model offers deeper written explanations.
Rank #3
- Choose heroic Super Mario characters and power-ups Choose between well-known characters such as Mario, Luigi, Peach, Daisy, Yoshi or Toad. Transform yourself into Elephant Mario with a surprising new power-up and poor opponents with your trunk!
- Share the miracle with friends and Mario fans games with up to three friends to experience the game-changing wonders locally on a Nintendo Switch console as you master the levels as a team and support each other on the way to the goal!
Hybrid reasoning was not automatically the deciding factor
Anthropic introduced Claude 3.7 Sonnet as a hybrid model with standard and extended-thinking modes in its February 2025 announcement. That product description does not prove that more internal reasoning caused the Mario result. In a real-time game, extra deliberation can increase control latency. The likely explanation is a balance of visual interpretation, response speed, output reliability and action policy—not simply “better reasoning.”
Why some reasoning models can lose in a game
The reported contrast with models such as o1 illustrates a speed-versus-deliberation trade-off, but it is only a hypothesis about this setup, not a general rule about reasoning models.
- Longer generation can delay the next control input.
- Token-heavy deliberation may be harmful when an enemy or gap is approaching.
- A game often rewards a short, repeatable policy over a detailed explanation.
- The model can identify the correct move but issue it after the useful timing window.
Performance on mathematics, coding or written reasoning therefore does not directly measure closed-loop motor control.
What “outperformed” means—and what it does not
The original news reports establish a ranking in Hao AI Lab’s comparison, but they do not provide a sufficiently detailed, independently audited score table for the demonstration. The available material does not establish the exact trial count, variance, scoring formula or statistical significance needed to make a broader claim.
Depending on the benchmark, “better” could mean greater distance, more level progress, a higher score, longer survival, more successful jumps, a higher completion rate, or a stronger average across repeated runs. Those measures are not interchangeable.
Rank #4
- Find wonder in the Flower Kingdom in the next side-scrolling Super Mario adventure
- Collect Wonder Flowers for surprising, game-changing effects like pipes coming alive, an enemy stampede, and much, much more
- Choose from the largest cast of characters in a side-scrolling Mario game, including Mario, Luigi, Peach, Daisy and other favorites
- Ease into the action with four different-colored Yoshis and Nabbit who can’t take damage
- Discover new power-ups like Elephant Fruit, which transforms Mario and friends into an elephant that can swing its trunk and spray water
Winning one Mario setup does not prove that Claude 3.7 was the best AI overall, the best game-playing system, or human-level at Super Mario Bros.
Free tools Windows power users keep installed
One-click scans. No signup required.
- It does not establish human-like controller skill.
- It does not predict coding, factuality, research, safety or robotics performance.
- It does not show that every Claude 3.7 configuration beats every competing model.
- It does not transfer automatically to newer Claude, Gemini or OpenAI versions.
Later benchmarks changed the picture
GamingAgent and the LMGame/Orak work expanded the evaluation framework, documented reproducibility tools and separated harness-enabled from non-harness testing. The GamingAgent repository lists support for additional models, including Claude 4, o4-mini, o3, Gemini 2.5 models, DeepSeek and Qwen.
One later Orak table reports the following Super Mario scores:
| Model | Reported score | Reported rank |
|---|---|---|
| Gemini 2.5 Pro | 38.0 ± 14.6 | 1 |
| o3-mini | 34.9 ± 14.6 | 2 |
| GPT-4o | 34.1 ± 14.2 | 3 |
| Claude 3.7 | 31.7 ± 8.2 | 5 |
| DeepSeek-R1 | 28.7 ± 13.2 | 8 |
These values belong to the later benchmark and its evaluation configuration; they must not be merged numerically with Hao AI Lab’s original demonstration. The alternate ordering is the important evidence. Change the harness, prompt, model endpoint, input modality, action timing or scoring method and the leaderboard can change. The benchmark materials are discussed at Orak’s published entry.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Can you reproduce the experiment?
The most direct route is the open-source GamingAgent project. Its documented installation path is:
Best Value
- A Mario game for up to four players, featuring five playable characters; Luigi's first starring role in a platforming adventure, Super Luigi U, is getting the deluxe treatment too and comes packed in
- A single Joy-Con controller is all each player needs; enjoy 164 courses for up to four players anytime, anywhere
- Mario, Luigi and Toad are all here and if that's not enough, Nabbit and Toadette are joining in the fun as well; nabbit doesn't take damage from enemies, which can really come in handy
- Compatible with Nintendo Switch only
- International products have separate terms, are sold from abroad and may differ from local products, including fit, age ratings, and language of product, labeling or instructions.
git clone https://github.com/lmgame-org/GamingAgent.git
cd GamingAgent
conda create -n lmgame python==3.10 -y
conda activate lmgame
pip install -e .
After configuring a supported provider API key, the repository documents commands in this form for harness-enabled and non-harness runs:
python3 lmgame-bench/run.py
--model_name {model_name}
--game_names super_mario_bros
--harness_mode true
python3 lmgame-bench/run.py
--model_name {model_name}
--game_names super_mario_bros
--harness_mode false
These are documented command patterns, not guaranteed turnkey instructions. Check the repository for current model identifiers, configuration files, emulator requirements, ROM handling and provider compatibility before running them. Model endpoints can be deprecated, and high-end evaluations can generate API charges.
Reproduction checklist
- Use a machine that can run the emulator and evaluation software.
- Install the documented Python 3.10 environment.
- Provide your own legally obtained game ROM where required; do not distribute copyrighted ROM files.
- Set provider API keys and confirm that the selected model accepts the required image and tool inputs.
- Fix the prompt, screenshot frequency, action duration and model settings.
- Run repeated trials instead of comparing one best attempt.
- Log latency, actions, retries, deaths, resets, progress and cost per episode.
Why a reproduction may disagree
- Provider interfaces can produce different behavior for the same nominal model.
- Network and API latency can change action timing.
- Prompt or screenshot-description changes invalidate a direct comparison.
- A harness may supply structure or tools unavailable in a non-harness run.
- Randomness and a small number of trials can make a ranking unstable.
- Distance, score, survival and completion are different metrics.
What this teaches us about AI evaluation
The Mario result is valuable because it exposes a capability that static benchmarks often miss: reliable, low-latency interaction with a changing visual environment. It also shows that the evaluated unit is not just the base model. The model, prompt, tools, emulator, observation schedule and scoring rule form one agent system.
A serious comparison should report average progress across many runs, variance, completion rate, action latency, corrective actions, deaths, harness dependence, input modality, cost, reproducibility and sensitivity to prompt wording. A spectacular single run is weaker evidence than consistent performance under a fixed protocol.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
The enduring lesson is not that Claude 3.7 was universally superior. It is that an AI can look excellent or ordinary depending on whether the benchmark rewards deliberation, visual interpretation, precise control, low latency or tool use. Claude 3.7 genuinely won the early Hao AI Lab setup; later results show why that win should remain a narrow, configuration-dependent finding.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




