DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
VGSources
Blog

Why Video Games Still Baffle AI Models

Video games expose the gap between fluent AI reasoning and reliable action. Here is why general-purpose models still fail at perception, memory, planning, and control—and what recent benchmarks show.
Length10 min Posted Quest giverVGSources Team
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A modern AI model may explain Pokémon, write a game clone, or defeat a world-class chess engine—and still get stuck walking into a wall. The reason is that playing an unfamiliar video game is not simply a test of knowledge or verbal reasoning. It is a continuous loop of seeing a changing world, discovering its rules, remembering state, planning ahead, and issuing precisely timed actions.

That distinction explains both the impressive game-playing demonstrations making headlines and the persistent failures reported in recent benchmarks. Specialized systems can dominate particular games. General-purpose language and vision-language models, however, remain unreliable when asked to control unfamiliar games end to end.

As an Amazon Associate I earn from qualifying purchases.

The important distinction: game AI is not one thing

“AI playing a video game” can describe several very different systems.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Specialized game systems are built or trained for a defined environment. Deep Blue, AlphaZero, Atari reinforcement-learning agents, search engines, and game-specific bots can be extraordinarily capable. Their success demonstrates powerful planning and learning, but not that one unchanged system can master arbitrary games.
  • General video-game agents attempt to play multiple unfamiliar games through a common interface. The General Video Game AI framework was designed around this broader challenge.
  • LLM and VLM agents use language or vision-language models to inspect screenshots or symbolic observations, decide what to do, and send keyboard, mouse, controller, or API commands.
  • Tool-assisted demonstrations may add OCR, state extraction, grid overlays, external memory, pathfinding, custom prompts, emulator hooks, automatic retries, or save-state recovery.

So “the model beat the game” is incomplete without describing the setup. Did it receive raw pixels or a prepared map? Did it control individual button presses or high-level actions? Was it trained on the game? Could a human intervene or restore a checkpoint?

#1 Best Overall
Sale
Logitech G213 Prodigy Wired RGB Gaming Keyboard - Black
  • Personalize 5 customizable lighting zones with over 16.8M colors to match your setup or game and synchronize backlit lighting effects with other Logitech G devices using Logitech G Hub
  • G213 Prodigy is a full-sized keyboard designed for gaming and productivity, with a slim body built for gamers of all levels and durable construction to repel liquids, crumbs, and dirt for easy cleanup
  • Each key is tuned to enhance the tactile experience, delivering ultra-quick, responsive feedback while the anti-ghosting gaming matrix is tuned for optimal gaming performance, keeping you in control
  • G213 gaming keyboard features dedicated media controls that can play, pause, and mute music and videos instantly; easily adjust the volume or skip to the next song with the touch of a button
  • Customize lighting, game mode, and macro programming with Logitech G HUB software and stay comfortable during long gaming sessions thanks to an integrated palm rest and adjustable keyboard feet

This is not an argument that scaffolded results are meaningless. A carefully designed agent system can demonstrate genuine progress. It is an argument for measuring the whole system—and for distinguishing a one-off engineered success from general competence.

Five layers at which game-playing agents fail

1. Seeing: recognition is not actionable perception

A model may correctly identify a character, enemy, wall, or item in an image and still be unable to play. Gameplay requires converting a scene into an accurate, changing spatial representation:

  • Which tile is occupied?
  • Which way is the player facing?
  • Is the route open or blocked?
  • Did the previous input take effect?
  • Is the screen changing because of animation, camera movement, or an actual gameplay event?

Small perception errors compound. Losing track of the player after scrolling can invalidate an otherwise sensible plan. Misreading menu focus can turn a healing command into an item discard. Confusing decoration with an interactive object can send an agent into a loop.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Recent game benchmarks identify brittle visual perception as a major obstacle to direct LLM/VLM interaction. The lmgame-Bench work treats perception, memory, and planning as connected evaluation problems rather than assuming that a model’s description of a screenshot proves usable understanding.

2. Understanding: rules are often implicit

Games frequently teach through consequences rather than complete instructions. A player may need to discover that an enemy can be defeated only from above, that an item unlocks a distant door, that an apparently safe surface causes damage, or that an NPC’s dialogue changes after a hidden condition.

Humans test hypotheses rapidly: try an action, observe the result, update the rules, and try again. Models often produce a plausible explanation of the rules without reliably testing whether it is true. They may confidently repeat a failed assumption because the current observation does not immediately contradict their narrative.

3. Remembering: a screenshot is not a world model

Many games require persistent state tracking:

  • inventory, health, money, and party composition;
  • visited locations and blocked routes;
  • NPC conversations and quest conditions;
  • which switches have been activated;
  • the current objective and the observations that explain how to achieve it.

A model can summarize recent events yet still lose the operational state needed for the next decision. It may repeat a route that failed, forget why it entered an area, or treat an old screenshot as current. Long context helps store information, but storage alone does not guarantee that the right state is updated and consulted at the right time.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

4. Planning: local intelligence is not long-horizon success

Games punish decisions that look reasonable in the moment. A wrong turn can waste minutes. A scarce resource spent early can make a later section impossible. A battle choice may have consequences several turns later, and a puzzle may require remembering a clue discovered much earlier.

Rank #2
Sale
AULA F75 Pro Wireless Mechanical Keyboard,75% Hot Swappable Custom Keyboard with Knob,RGB Backlit,Pre-lubed Reaper Switches,Side Printed PBT Keycaps,2.4GHz/USB-C/BT5.0 Mechanical Gaming Keyboards
  • Tri-mode Connection Keyboard: AULA F75 Pro wireless mechanical keyboards work with Bluetooth 5.0, 2.4GHz wireless and USB wired connection, can connect up to five devices at the same time, and easily switch by shortcut keys or side button. F75 Pro computer keyboard is suitable for PC, laptops, tablets, mobile phones, PS, XBOX etc, to meet all the needs of users. In addition, the rechargeable keyboard is equipped with a 4000mAh large-capacity battery, which has long-lasting battery life
  • Hot-swap Custom Keyboard: This custom mechanical keyboard with hot-swappable base supports 3-pin or 5-pin switches replacement. Even keyboard beginners can easily DIY there own keyboards without soldering issue. F75 Pro gaming keyboards equipped with pre-lubricated stabilizers and LEOBOG reaper switches, bring smooth typing feeling and pleasant creamy mechanical sound, provide fast response for exciting game
  • Advanced Structure and PCB Single Key Slotting: This thocky heavy mechanical keyboard features a advanced structure, extended integrated silicone pad, and PCB single key slotting, better optimizes resilience and stability, making the hand feel softer and more elastic. Five layers of filling silencer fills the gap between the PCB, the positioning plate and the shaft,effectively counteracting the cavity noise sound of the shaft hitting the positioning plate, and providing a solid feel
  • 16.8 Million RGB Backlit: F75 Pro light up led keyboard features 16.8 million RGB lighting color. With 16 pre-set lighting effects to add a great atmosphere to the game. And supports 10 cool music rhythm lighting effects with driver. Lighting brightness and speed can be adjusted by the knob or the FN + key combination. You can select the single color effect as wish. And you can turn off the backlight if you do not need it
  • Professional Gaming Keyboard: No matter the outlook, the construction, or the function, F75 Pro mechanical keyboard is definitely a professional gaming keyboard. This 81-key 75% layout compact keyboard can save more desktop space while retaining the necessary arrow keys for gaming. Additionally, with the multi-function knob, you can easily control the backlight and Media. Keys macro programmable, you can customize the function of single key or key combination function through F75 driver to increase the probability of winning the game and improve the work efficiency. N key rollover, and supports WIN key lock to prevent accidental touches in intense games

Language models are good at producing plans in prose. Executing those plans is harder. The agent must notice when the world differs from its expectation, revise the plan, backtrack when necessary, and avoid drifting into repetitive loops. A plan that is correct in theory is worthless if the controller cannot maintain it over hundreds of state transitions.

5. Acting: the interface turns reasoning into control

“Move right” is not always a complete action. The agent may need to know whether to tap or hold a key, how long to wait, whether inputs are buffered, whether an animation temporarily disables control, and whether the character is already moving.

This is a control problem as much as a reasoning problem. Real-time games add frame timing, collision boundaries, camera motion, and narrow windows for action. A single mistimed input can invalidate a good strategy. Commands issued too quickly may be ignored; commands issued too slowly may make an otherwise competent agent inefficient or vulnerable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Feedback is also less clean than in software development. A compiler error or failed test gives a clear intermediate signal. In a game, an action may appear to do nothing, succeed only after a delay, or create a failure that becomes visible several minutes later.

Why models can write games more easily than they can play them

The apparent paradox is straightforward. A model can generate a playable game because conventional mechanics, code patterns, and engine structures are common in its training data. A request for a platformer or maze game maps onto familiar templates, and the result can often be checked for syntax, compilation, or basic execution.

Playing the result demands something different. Good development involves an iterative loop:

  1. implement a mechanic;
  2. play it;
  3. notice whether controls feel responsive;
  4. adjust timing, physics, difficulty, interface, and pacing;
  5. repeat.

A model that cannot reliably control the game cannot fully evaluate the player experience it created. It may produce a technically functional but generic, poorly balanced, or frustrating game. That does not make AI useless for game design or playtesting; it means code generation is not equivalent to embodied evaluation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Chess and Go are important—but unusually clean comparisons

Chess and Go are profoundly difficult games, so their success stories should not be dismissed. But they offer specialized AI an unusually favorable computational interface:

Rank #3
Sale
SteelSeries Apex 3 Gaming Keyboard - Black
  • Ip32 water resistant – Prevents accidental damage from liquid spills
  • 10-zone RGB illumination – Gorgeous color schemes and reactive effects
  • Whisper quiet gaming switches – Nearly silent use for 20 million low friction keypresses
  • Premium magnetic wrist rest – Provides full palm support and comfort
  • Dedicated multimedia controls – Adjust volume and settings on the fly
  • the board is explicit;
  • the state representation is compact;
  • legal actions are structured;
  • the rules are stable and known;
  • the objective is unambiguous;
  • search and evaluation can be optimized repeatedly for the same environment.

Two ordinary video games may differ completely in camera perspective, physics, input device, action timing, hidden state, objective, visual conventions, and failure conditions. AlphaZero’s success in chess and Go demonstrates exceptional game-specific learning and planning. It does not demonstrate that one system can arrive at an unfamiliar game, infer its interface, and play it reliably.

What Pokémon demonstrations really show

Pokémon is a revealing challenge because it combines exploration, menus, battles, inventory and party management, puzzles, long-term progression, and delayed rewards. It tests far more than selecting the strongest move in a fixed board position.

In May 2025, Gemini 2.5 Pro was reported to complete Pokémon Blue. That was a notable demonstration of progress, but not a clean test of an unassisted general model. As IEEE Spectrum reported, the run used custom interaction software and additional processing to make the game state more legible and actionable, and it was far slower and more error-prone than human play.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The fair question is therefore not whether the result was “pure.” It is:

  • What information did the model receive?
  • What actions did it control directly?
  • Were OCR, maps, memory, pathfinding, or other tools available?
  • Could it retry or restore a save state?
  • Was the game known in advance?
  • How many attempts, tokens, actions, and hours were required?
  • Did the same method transfer to a new game?

Completing one familiar title under a carefully engineered harness can be meaningful while still falling short of general game-playing ability.

What current benchmarks reveal

GVGAI-LLM

The GVGAI-LLM benchmark, published on arXiv in August 2025, adapts general video-game evaluation for language models. It uses diverse arcade-style games, compact ASCII representations, rapidly created content, and metrics including meaningful step ratio, step efficiency, and overall score.

Its purpose is to reduce overfitting to famous commercial games and test reasoning about unfamiliar rules. The reported results show persistent spatial and logical errors. Structured prompts and spatial grounding improve performance, but do not resolve the underlying problem.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

lmgame-Bench

lmgame-Bench argues that simply dropping an LLM into a game makes for an unstable evaluation. Performance can be dominated by brittle visual perception, prompt wording, and possible training-data contamination. Its framework covers platformer, puzzle, and narrative games, with perception and memory scaffolds intended to make comparisons more reliable.

Rank #4
Sale
TECKNET Wired Gaming Keyboard, RGB Backlit Keyboard with Metal Panel Design
  • 【Ergonomic Design, Enhanced Typing Experience】Improve your typing experience with our computer keyboard featuring an ergonomic 7-degree input angle and a scientifically designed stepped key layout. The integrated wrist rests maintain a natural hand position, reducing hand fatigue. Constructed with durable ABS plastic keycaps and a robust metal base, this keyboard offers superior tactile feedback and long-lasting durability.
  • 【15-Zone Rainbow Backlit Keyboard】Customize your PC gaming keyboard with 7 illumination modes and 4 brightness levels. Even in low light, easily identify keys for enhanced typing accuracy and efficiency. Choose from 15 RGB color modes to set the perfect ambiance for your typing adventure. After 30 minutes of inactivity, the keyboard will turn off the backlight and enter sleep mode. Press any key or "Fn+PgDn" to wake up the buttons and backlight.
  • 【Whisper Quiet Design】Experience near-silent operation with our whisper-quiet gaming switch, ideal for office environments and gaming setups. The classic volcano switch structure ensures durability and an impressive lifespan of 50 million keystrokes.
  • 【IP32 Spill Resistance】Our quiet gaming keyboard is IP32 spill-resistant, featuring 4 drainage holes in the wrist rest to prevent accidents and keep your game uninterrupted. Cleaning is made easy with the removable key cover.
  • 【25 Anti-Ghost Keys & 12 Multimedia Keys】Enjoy swift and precise responses during games with the RGB gaming keyboard's anti-ghost keys, allowing 25 keys to function simultaneously. Control play, pause, and skip functions directly with the 12 multimedia keys for a seamless gaming experience. (Please note: Multimedia keys are not compatible with Mac)

The project’s broader finding is important: games combine capabilities that are often separated in conventional language benchmarks. The authors also report evidence that game-specific reinforcement learning can transfer to unseen games and external planning tasks. That suggests games can be useful laboratories for generalization—not that every game score is a complete measure of intelligence.

VideoGameBench and strategic-game evaluations

VideoGameBench evaluates vision-language models in real-time interaction with classic games. Reported results indicate that frontier models often make little progress beyond opening sections, especially in less forgiving versions. Exact percentages should be interpreted in the context of the paper’s model versions, game list, interface, and enabled tools rather than treated as permanent rankings.

Google’s GENSTRAT work takes a different approach, evaluating strategic reasoning through thousands of generated games and tens of thousands of matches. Its reported lesson is that models with similar aggregate strength can have different profiles: one may excel at local tactics, another at long-range planning, consistency, or adaptation. A single leaderboard number can hide those differences.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Google and Kaggle’s Game Arena, announced in August 2025, reflects the same movement toward game suites and competitive evaluations rather than one famous title.

Training data can make a game look easier

Popular games leave extensive traces online: walkthroughs, maps, wikis, strategy guides, gameplay videos, speedrunning documentation, source code, and emulator projects. A model may therefore possess prior knowledge of the objective or route before it ever sees the game.

That creates a distinction between:

  • knowing a game’s rules and discovering them;
  • recognizing a screenshot and controlling the underlying world;
  • recalling a walkthrough and adapting a strategy;
  • memorizing a route and generalizing to a new layout.

Newly generated or less-famous games help separate interactive competence from retrieval. GVGAI-LLM uses rapidly created games and levels for this reason, while lmgame-Bench treats contamination and evaluation stability as explicit concerns.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why a larger model is not automatically a better player

Scaling can improve verbal understanding, planning language, tool use, and error explanation. It does not automatically provide high-frequency visual tracking, accurate coordinates, stable world models, real-time motor control, or reliable reward learning.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A larger model may explain the correct move while still pressing the wrong button. It may produce a more elaborate plan without executing it more reliably. It can also be more expensive, slower, or more confidently wrong.

Best Value
GEODMAER 65% Gaming Keyboard, Wired Backlit Mini Keyboard, Ultra-Compact Anti-Ghosting No-Conflict 68 Keys Membrane Gaming Wired Keyboard for PC Laptop Windows Gamer
  • 【65% Compact Design】GEODMAER Wired gaming keyboard compact mini design, save space on the desktop, novel black & silver gray keycap color matching, separate arrow keys, No numpad, both gaming and office, easy to carry size can be easily put into the backpack
  • 【Wired Connection】Gaming Keybaord connects via a detachable Type-C cable to provide a stable, constant connection and ultra-low input latency, and the keyboard's 26 keys no-conflict, with FN+Win lockable win keys to prevent accidental touches
  • 【Strong Working Life】Wired gaming keyboard has more than 10,000,000+ keystrokes lifespan, each key over UV to prevent fading, has 11 media buttons, 65% small size but fully functional, free up desktop space and increase efficiency
  • 【LED Backlit Keyboard】GEODMAER Wired Gaming Keyboard using the new two-color injection molding key caps, characters transparent luminous, in the dark can also clearly see each key, through the light key can be OF/OFF Backlit, FN + light key can switch backlit mode, always bright / breathing mode, FN + ↑ / ↓ adjust the brightness increase / decrease, FN + ← / → adjust the breathing frequency slow / fast
  • 【Ergonomics & Mechanical Feel Keyboard】The ergonomically designed keycap height maintains the comfort for long time use, protects the wrist, and the mechanical feeling brought by the imitation mechanical technology when using it, an excellent mechanical feeling that can be enjoyed without the high price, and also a quiet membrane gaming keyboard

Game-playing performance is a systems property. The model, visual encoder, memory design, controller, emulator, timing loop, observation format, and recovery policy all matter. In practice, the most promising architecture is likely hybrid: a language model handles abstraction, explanation, and flexible goals, while specialized components handle perception, navigation, search, timing, and recovery.

How to judge a game-playing claim

When a demonstration is presented as evidence of AI progress, ask:

  1. Novelty: Was the game newly generated, unfamiliar, or extensively documented?
  2. Observation: Did the agent receive raw pixels, OCR, ASCII, telemetry, or a prepared description?
  3. Action: Did it issue individual button presses or high-level commands?
  4. Tools: Were maps, pathfinding, memory, search, or state tracking delegated?
  5. Training: Was it fine-tuned or reinforcement-trained on this title?
  6. Retries: Could it restore checkpoints or save states?
  7. Intervention: Did a human correct stuck states?
  8. Efficiency: How many actions, tokens, attempts, and hours did completion require?
  9. Transfer: Did the approach work on a different game?
  10. Reproducibility: Can independent researchers run the same protocol?

Completion rate is only one metric. Meaningful progress ratio, efficiency, consistency, recovery from errors, latency, and cost often reveal more about practical competence.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why this matters beyond games

Games are controlled environments: rules are digital, state can sometimes be logged exactly, goals are often explicit, and evaluation can be automated. That makes them valuable testbeds for computer-use agents, simulation training, game testing, robotics research, and autonomous software.

They are also unusually diverse. One game may demand driving, another inventory management, another platforming, dialogue, tactical combat, or physics-based construction. The real world is vastly more complex, but its physical regularities are more consistent across locations than the interfaces and rules of unrelated games.

Success in one simulated environment therefore does not prove real-world general intelligence. Failure across arbitrary games does not prove that an agent cannot operate in the real world. Game evaluations measure a bundle of capabilities under a particular interface and objective.

What could close the gap?

Progress will likely require more than simply increasing a model’s context window or parameter count. Useful directions include:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Persistent world models that track objects, locations, rules, uncertainty, and change.
  • Spatial representations such as maps, coordinates, object relations, and reachability graphs.
  • Active experimentation that lets an agent deliberately test uncertain rules.
  • Reinforcement learning for learning action policies and recovery behaviors.
  • Search and model-predictive control where a forward model is available.
  • Hierarchical control in which a planner chooses goals and a lower-level controller handles timing.
  • External memory for inventory, maps, objectives, and failed-action history.
  • Benchmarks with novel games and standardized observation, action, retry, and cost protocols.

The likely solution is not an LLM operating alone. It is a coordinated system in which general models supply flexible interpretation and specialized modules supply grounded, repeatable control.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More quests from Patch Notes

  1. How to Set Up a RedM (RDR2) Server in 2026: License Key, txAdmin and server.cfg, Step by StepBlog11min
  2. Best RedM (RDR2) Server Hosting in 2026: Comparing Four Hosts on Slots, Memory and PriceBlog10min
  3. How to Host a Mindustry Server in 2026: server-release.jar, Port 6567 and the Commands That MatterBlog7min
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.