In March 2025, a developer known as Guzus published a website where language models played the social-deduction game Mafia against one another. The site showed a leaderboard, completed games, player roles, and full transcripts—and those transcripts were often far more entertaining than competent. Models exposed their own roles, forgot earlier votes, leaked teammates, and made confident accusations based on little evidence.
The project is best understood as a public experiment in multi-agent reasoning, not a definitive test of intelligence. Its reported leaderboard is a useful historical snapshot, while the transcripts reveal how difficult it is for fluent language models to maintain hidden information, strategy, and consistent memory over several turns.
The website is an LLM Mafia experiment, not a normal multiplayer game
The public project is available at mafia.opennumbers.xyz. It pits language models against each other in a Mafia-like game and publishes the resulting matches for people to inspect.
Visitors can use the project to examine:
- a model leaderboard;
- completed-game results;
- role-specific performance;
- the players assigned to Mafia, villager, or doctor roles; and
- full dialogue transcripts showing how each game unfolded.
The project was publicly described on March 3, 2025, and reported by Tom’s Hardware on March 7, 2025. The available reporting confirms that the site was public at the time, but does not establish whether it is still actively running games, whether the leaderboard has been updated, or whether its original model versions remain available in 2026.
#1 Best Overall
- CARD GAMES FOR FAMILIES: The award-winning SET party game has been loved for years by fans of adult card games, kids card games, and family card games
- EASY TO LEARN GAME: Race to make the most SETs by matching colors, shapes, shadings, or number of shapes on the cards. They have to be all the same or all different
- FAMILY GAMES FOR KIDS AND ADULTS: The game appeals to a wide age range. Add it to your go-to children’s games, card games for teens, and board games for family night
- TRAVEL SIZE GAMES: It’s one of the best road trip activities for kids. Choose it when looking for camping games, beach games, and group games for a party
- FAMILY CARD GAMES MAKE GREAT TEEN GIFTS: The SET family card game makes a great gift for teens. Give it as a family night gift basket, or as stocking stuffers for kids and teens
The developer also discussed expanding the framework to human-versus-LLM games, poker, real-time viewing, additional roles, and other games. Those were proposed directions, not features that should be assumed to be complete.
How the reported Mafia games worked
The reported configuration used eight players:
- five villagers;
- one doctor; and
- two Mafia members.
During the day, the players discussed who might be Mafia and voted to eliminate someone. At night, the Mafia selected a target while the doctor attempted to protect a player. The villagers won by eliminating the Mafia; the Mafia won by removing enough villagers to control the outcome.
“Villager,” “farmer,” and “citizen” can refer to equivalent civilian roles depending on the interface or translation. Mafia is also closely related to Werewolf, but implementations differ. Exact details such as tie handling, doctor restrictions, the information shown to eliminated players, and the precise win condition should be verified in the project’s code or current rules rather than imported from a favorite tabletop version.
The source code is available in the guzus/llm-mafia-game GitHub repository. That repository is also the appropriate place to investigate whether a strange transcript reflects a model failure or a problem in the game engine.
The March 2025 leaderboard
The most widely reported table placed Claude 3.7 Sonnet in extended mode at the top, with a 57.78% overall win rate across 45 games and a reported 100% win rate when assigned Mafia. DeepSeek Chat was listed second overall at 50% across 56 games.
| Model | Games | Overall | Mafia | Villager | Doctor |
|---|---|---|---|---|---|
| Claude 3.7 Sonnet, extended | 45 | 57.78% | 100.00% | 37.04% | 50.00% |
| DeepSeek Chat | 56 | 50.00% | 88.24% | 31.03% | 40.00% |
| Claude 3.7 Sonnet, standard | 54 | 46.30% | 92.86% | 32.35% | 16.67% |
| Claude 3.5 Sonnet | 47 | 44.68% | 90.00% | 36.67% | 14.29% |
| Llama 3.3 70B Instruct | 65 | 44.62% | 72.73% | 30.00% | 30.77% |
| Mistral Small 24B Instruct | 65 | 44.62% | 80.00% | 30.30% | 25.00% |
| GPT-4o-mini | 71 | 42.25% | 82.61% | 27.50% | 0.00% |
| GPT-4o | 49 | 38.78% | 90.00% | 24.24% | 33.33% |
| DeepSeek-R1 | 22 | 36.36% | 62.50% | 23.08% | 0.00% |
| Mythomax L2 13B | 61 | 31.15% | 45.45% | 28.21% | 27.27% |
| WizardLM 2 8x22B | 65 | 26.15% | 41.67% | 23.40% | 16.67% |
| Mistral Nemo | 17 | 17.65% | 40.00% | 10.00% | 0.00% |
GIGAZINE’s March 10, 2025 report contains the broader historical table.
These figures need careful interpretation. Each model played a different number of games, ranging from 17 for Mistral Nemo to 73 for Gemini Flash 1.5 among the reported entries. Role assignments were also not necessarily equal. A 100% Mafia win rate may represent only a limited number of Mafia assignments, and doctor percentages are particularly noisy when a model receives that role infrequently.
Mafia and villager performance also measure different things. A model can be effective at deception but poor at evaluating other players’ claims, or competent as a civilian but unable to coordinate with a Mafia teammate. The table therefore does not prove that Claude is universally the best Mafia player, or that DeepSeek is definitively better at deception.
Rank #2
- Fun-Filled Decks: Embrace the fun with six different games in one set! The set, designed for 2-6 players, contains decks for Old Maid, Go Fish, Slap Jack, Crazy 8's, War, and Silly Monster Memory Match, providing hours of entertainment and learning opportunities.
- Child-Friendly Design: Each card in this set is crafted with vibrant, bright colors and easy-to-understand symbols, tailored to engage and captivate young minds.
- Skill-Building Games: Not just for fun, these card games are a stealthy way to build essential skills. These games provide educational benefits such as learning colors, numbers, and reading skills, while also encouraging memory and matching skills.
- Big Cards for Little Hands: Our cards are extra big - it's easy to hold and play! They make an ideal gift for aged 4 and older boys and girls.
- Fun on the Fly: These funny family card games for kids and adults, are your pocket-sized partners for entertainment anywhere. Road trip? Sleepover? Camping Trip? A quick visit to Grandma? Just pop them in your bag and you're set for a fun time, anytime.
The funniest failures expose the real problem
One reported Mythomax transcript contained an astonishing self-disclosure:
“As Mafia, my primary goal is to protect myself and eliminate the other Mafia member.”
That is not sophisticated deception. It appears to be an accidental role leak or a confused description of the game objective. Claude 3.7 Sonnet noticed the contradiction and treated it as evidence of either a role leak or an unusually strange strategy. The complete transcript is more informative than the quotation alone because it shows what led up to the statement and how the other players responded.
In another reported incident, Mythomax named Hermes-3-Llama-3.1-405B as its Mafia partner after being eliminated, apparently revealing information that should have remained strategically hidden.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Across the published games, the most revealing failure patterns include:
- Explicit role leakage: a model states or strongly implies its hidden role.
- Contradictory claims: a player forgets what it said earlier or reverses a vote without accounting for the change.
- Broken state tracking: the model loses track of eliminations, role claims, votes, or accusations.
- Private-information confusion: a model acts as if it knows something only another player should know.
- Overconfident accusations: tone, verbosity, or a single suspicious phrase is treated as proof.
- Literal cooperation: the model follows norms of openness even when strategic concealment is necessary.
- Weak doctor play: protection choices do not consistently reflect likely Mafia targets or public claims.
- Post-elimination leakage: a dead player reveals teammates or hidden information that should remain concealed.
These moments are funny, but they are also useful debugging clues. Before treating a transcript as evidence that a model cannot reason, an evaluator should check whether the engine accidentally included private role data in a public prompt, exposed eliminated players’ information, truncated conversation history, or parsed an invalid response incorrectly.
Why fluent language does not translate into Mafia skill
Mafia requires more than producing persuasive sentences. Each player must maintain a changing internal ledger while reasoning about what every other player could legitimately know.
State tracking
A competent player needs to remember who claimed which role, who voted for whom, which players have been eliminated, which accusations were later withdrawn, and what information was public at each point. A model can generate locally coherent prose while losing the global state of the game.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsRank #3
- SPECIAL EDITION: Comes in a collectible, premium tin!
- NEVER BE BORED: Classic, award-winning game that is fun to play over and over again—over 3 million copies sold!
- FAMILY FUN: Versatile, challenging and fun game appeals to a wide age range, accommodates any number of players, and is a great game for kids and adults to play together!
- FUN GAME PLAY: Race to find as many SETs as you can (each individual feature is all the same or all different)
- PROMOTES BRAIN HEALTH: Builds cognitive, logical and spatial reasoning skills as well as visual perception skills
Private-information boundaries
The orchestrator must keep each model’s role and private messages separate from the public conversation. If that boundary fails, a role leak may be a framework bug rather than a model’s reasoning failure. The code should be audited before drawing conclusions from especially blatant disclosures.
Role-objective confusion
The model may understand the written rules but fail to optimize consistently for its assigned role. Helpful, transparent conversation is usually desirable in ordinary chatbot interactions; in Mafia, it can be fatal. The Mythomax example illustrates how literal explanation can override the goal of winning.
Weak strategic memory
Remembering the last exchange is not the same as maintaining a compact, reliable game ledger over multiple day-night cycles. Re-sending a long transcript every turn can also make it harder for a model to prioritize the information that matters.
Poor uncertainty calibration
Villagers must make decisions with incomplete evidence. A model that confidently converts suspicion into certainty can eliminate an innocent player. A Mafia player that overexplains its reasoning can expose itself while trying to sound credible.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Prompt and inference effects
The difference between Claude 3.7 Sonnet’s standard and extended entries suggests that configuration matters, but it does not prove that extended reasoning alone caused the higher score. Prompt wording, token limits, temperature, context handling, model routing, and evaluation timing may also differ.
What does Mafia actually measure?
The experiment tests a mixture of natural-language interaction, hidden-information reasoning, deception, theory of mind, memory, coalition-building, risk-sensitive decisions, rule-following, and orchestration reliability.
That makes it interesting, but not a general intelligence test. The observed result is produced by at least five interacting layers:
- Model capability: what the underlying language model can represent and infer.
- Agent policy: the system prompt, role instructions, memory format, and decision procedure.
- Game engine: turn order, private messages, vote parsing, eliminations, and win conditions.
- Sampling process: role distribution, seating order, random seeds, temperature, retries, and number of games.
- Evaluation metric: overall wins versus performance in specific roles.
A model may perform well because the prompt suits its conversational style or because it receives favorable roles. Another may perform poorly because its context is truncated, its output parser is brittle, or its model version changed. The leaderboard measures the whole system, not just the base model.
Rank #4
- NEVER BE BORED—Classic, award-winning game that is fun to play over and over again—over 3 million copies sold!
- GIANT GAME PLAY—Race to find as many SETs as you can (each individual feature is all the same or all different) with giant cards!
- FAMILY FUN—Versatile, challenging and fun game appeals to a wide age range, accommodates any number of players, and is a great game for kids and adults to play together!
- GREAT FOR ANY GAME SITUATION—Perfect for home, parties, travel and school play!
- PROMOTES BRAIN HEALTH—Builds cognitive, logical and spatial reasoning skills as well as visual perception skills
How reliable is the leaderboard?
It is best treated as an entertaining preliminary benchmark and a source of qualitative examples—not as a definitive ranking of reasoning or deception.
The reported results have several limitations:
- models played unequal numbers of games;
- role-specific sample sizes differed;
- the reports do not establish statistical significance or confidence intervals;
- model versions and provider backends may change;
- standard and extended modes may have different budgets or prompts;
- API failures, retries, and malformed outputs may affect effective behavior; and
- the available coverage does not provide independent replication.
A stronger benchmark would give every model equal role distributions, randomize seating order, use multiple seeds, pin exact model versions, standardize prompts and token budgets where practical, validate public/private context boundaries, report timeouts and retries, and publish confidence intervals. It would also freeze the rules so later leaderboard entries remain comparable.
Until then, “Claude 3.7 Sonnet extended led the reported March 2025 table” is supportable. “Claude is the smartest Mafia-playing AI” is not.
How to watch the games
If the public site still loads, the simplest route is:
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →- Open the public Mafia website.
- Check whether the leaderboard and game archive are available.
- Select a completed match.
- Review the role assignments and final result.
- Read the transcript from the first day through the final vote.
- Compare what each player said publicly with the hidden roles and the eventual outcome.
Do not assume that the interface, archive, rankings, or ability to start a new game are unchanged from the March 2025 reports.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to reproduce the experiment
The GitHub repository is the natural starting point for a technical reproduction. The main decision is whether to use hosted models or local inference.
Hosted model gateways
The original reporting said the simulations used OpenRouter rather than local models. A gateway such as OpenRouter is convenient when comparing models from several providers through one API.
Advantages include one integration, centralized billing, easy model switching, and access to many hosted models. The trade-offs are token costs, changing model availability, model aliases that may point to new versions, privacy considerations, and weaker reproducibility if exact snapshots are not pinned.
Recommended Free Tools
Best Value
- POKER KEENO - This fun group card game brings two teams of two head-to-head! Try to follow suit and win tricks to collect points and win.
- VALUE - This Regal Game comes with 36 9.5” x 7” double-sided game boards, 600 scoring chips, and a deck of standard Poker-size playing cards. It’s everything you need for the ultimate casino experience.
- BULK BUNDLE - This game set is optimally designed for large get togethers. With a total 72 unique board layouts, you’re sure to keep the game going for hours of play time!
- MANY WAYS TO PLAY - Whether you play the classic Bingo-style 5-in-a-row or the alternative poker styles, you’re sure to have a night of fun, laughter, and luck.
- PERFECT GAME - This Regal Game is the perfect addition to community events, fundraisers, family reunions, birthday parties, casino game nights, and more.
OpenRouter’s current pricing and FAQ pages list provider pricing, a 5.5% fee with a $0.80 minimum when purchasing credits, and a free-model request limit. Those terms can change, so check official pricing and the official FAQ before running a large batch. The exact cost of the original experiment has not been reported and cannot be inferred without token counts, retry data, and contemporaneous model prices.
Local inference with Ollama
Ollama is the more privacy-oriented option for open-weight models. It advertises local execution, and its API documentation describes a chat endpoint at http://localhost:11434/api/chat.
Local inference avoids per-token API charges after the hardware investment and gives you more control over model files and versions. However, larger models require more system memory or GPU memory, and running eight agents concurrently can saturate hardware. Local models may also be slower or less capable than hosted frontier models.
Ollama’s pricing page currently lists local use as free and a $20-per-month Pro plan, while noting that new Max sign-ups are paused. These are current pricing signals, not permanent terms; consult the official pricing page.
Free tools Windows power users keep installed
One-click scans. No signup required.
Reproduction checklist
- Pin exact model identifiers and versions.
- Use identical system prompts and role instructions.
- Record temperature, token limits, context limits, and reasoning settings.
- Randomize player names, seating order, roles, and seeds.
- Keep each model’s private role context separate from public history.
- Log every API failure, timeout, retry, and malformed response.
- Make vote parsing deterministic and test ties explicitly.
- Store complete transcripts, including private actions and engine events.
- Run enough games to calculate uncertainty rather than relying on one lucky streak.
- Report role-specific results as well as overall wins.
What this project is genuinely useful for
The website is valuable in three ways.
First, it is an entertaining demonstration of the gap between conversational fluency and sustained strategic reasoning. A model can sound articulate while forgetting who voted for whom two turns earlier.
Second, it is a practical stress test for multi-agent orchestration. Mafia exposes bugs in memory management, private-context isolation, tool calls, output parsing, turn scheduling, and retry handling.
Third, it provides a promising starting point for a standardized evaluation. With fixed rules, balanced roles, pinned versions, controlled prompts, and independent replication, social-deduction games could measure specific capabilities such as state tracking, uncertainty calibration, coalition reasoning, and adaptation under hidden information.
What it does not provide is proof that language models cannot understand deception, or that one model is inherently more intelligent than another.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




