1. Conclusion: Mahjong AI is not a magic button for huge win rates. It is a machine for stacking a few percentage points
In four-player Mahjong, if all four players are equally strong, the starting point for first place is 25%. That is the first brutal fact.
Even elite humans and strong AIs do not suddenly win half their games. At high levels, strength appears as small long-run differences: a few more firsts, a few fewer lasts, slightly fewer deal-ins, and better choices that change with score, round, seat, hand value, and risk.
High-level Mahjong is therefore less like firing a super move and more like working in a domino factory: keep one safe tile, push one valuable good wait, abandon one cheap bad fight, and choose the correct placement goal in the final hand. Each edge looks tiny. Hundreds of games later, the rating system remembers all of them.
One research caveat matters: there is no solid causal study showing that “using an AI reviewer raises a human player's first-place rate by exactly X points.” The strength of an AI policy and the learning effect of AI coaching are different questions.
2. The cruel 25% baseline of four-player Mahjong
With symmetric players, each rank averages 25%. So “a strong player should win half the time” is the wrong mental model.
Mahjong contains hidden hands, hidden wall tiles, random draws, changing turn order through calls, and three opponents. A good decision can lose; a bad decision can win one hand. Skill therefore appears more reliably in long-run rank distribution and decision quality than in the result of one session.
Saying “I did not get a first today, so my strategy is broken” is like diagnosing climate from this afternoon's weather.
3. Suphx: about a 30% first-place rate is already absurdly strong
Microsoft Research's Suphx is a landmark deep-reinforcement-learning Mahjong system. In the published evaluation, it played 5,760 hanchan in Tenhou's expert room and reached a stable rank of 8.74, compared with 7.46 for the top-human reference group.
Table 5 of the paper reports:
| Metric | Suphx | Top humans | Bakuuchi | NAGA |
|---|---|---|---|---|
| 1st | 29.3% | 28.0% | 28.0% | 25.6% |
| 4th | 18.7% | 20.5% | 22.4% | 21.1% |
| Hand win rate | 22.83% | - | 23.07% | 22.69% |
| Deal-in rate | 10.06% | - | 12.16% | 11.42% |
The important part is not only the 29.3%. Suphx simultaneously had the highest first-place rate in that table, the lowest fourth-place rate, and the lowest deal-in rate.
That breaks the simplistic idea that “more firsts must be purchased with more lasts.” Better decisions can move both ends in the favorable direction.
Suphx also used Global Reward Prediction, connecting local round decisions with the final game reward. That matters because the best move for this hand is not always the best move for the match.
These are Tenhou results, not Mahjong Soul results. Different ranking systems and player pools mean 29.3% should not be copied as a universal target.
4. High ranks are a domino factory
At beginner levels, basic tile efficiency can create a huge advantage. At higher levels, everyone has already collected much of that free money.
The remaining edges are smaller: when to keep a safe tile, when to value a good wait over extra points, how far to push against riichi, whether a one-shanten hand is worth fighting, how dealer status changes risk, and whether second place in South should chase first or protect placement.
One decision looks microscopic. Five hundred repetitions do not.
High-ranked Mahjong is the factory where people who laugh at 0.5% are slowly buried under a very polite pile of expected value.
5. Outsource too much and you can become an “AI operator who does not know the rules”
There is another trap.
“What do I discard?” Ask AI. “Do I call?” Ask AI. “What is my wait?” Ask AI. “What yaku is this?” Apparently the server knows.
Performance can improve before understanding does. You may learn how to operate a Mahjong solver instead of learning Mahjong.
Psychology calls the use of external tools to reduce internal mental processing cognitive offloading. Risko and Gilbert's review describes how offloading reduces cognitive demand and can have cognitive consequences.
A 2025 randomized trial with 120 university students also found lower 45-day knowledge retention in a group using unrestricted ChatGPT study assistance than in a traditional-study group. That was not a Mahjong study, so it does not prove that Mahjong AI causes poor learning. It does support the general warning: outsourcing an answer does not guarantee that you build the internal process that produces it.
It is the gaming equivalent of following a walkthrough to the final boss and then discovering that the exam is on the tutorial.
6. What does “I cannot get first, but I rarely get last” actually mean?
By itself, not enough.
At least four hypotheses fit the same symptom:
- Over-defense: you reject profitable risks and compress into second/third.
- Attack-quality problem: riichi, calls, speed, or hand value are failing to convert opportunities into firsts.
- Placement-state problem: late-game choices do not properly pursue first when the score permits it.
- Variance: nothing important changed; the recent sample simply contains fewer firsts.
So “low fourth-place rate means you are too defensive” is not a research-grade diagnosis.
7. Research logic: placement rates alone cannot identify the cause
Aggregate statistics have an identification problem: different mechanisms can generate the same output.
A record with 20% firsts and 15% lasts might come from excessive folding, unlucky high-value hands, poor final-round conversion, or plain sampling noise.
Suphx itself is evidence against the naive defense-versus-first-place story. It defended extremely well and still had the highest first-place rate in the comparison table.
To diagnose over-defense, we need conditional behavior: what was the hand value, shanten, turn, dealer status, current rank, opponent riichi state, and actual push/fold decision?
8. A bad 100-game stretch can easily be noise
Suppose the true first-place probability were 25%, and—very crudely—we treated each hanchan as an independent Bernoulli trial. The approximate 95% sampling margin is:
| Games | Approx. margin around 25% |
|---|---|
| 100 | ±8.5 percentage points |
| 500 | ±3.8 points |
| 1,000 | ±2.7 points |
This is not a full model of Mahjong Soul. Opponents, ranks, rules and your own strength change, so games are not perfectly identical independent trials. The table is only a scale for how noisy short samples can be.
Mahjong variance has excellent timing: it waits until you say “100 games should be enough” before tapping you on the shoulder.
9. Turn AI back into a teacher: collect reasons, not answers
A better review loop is:
- Choose your move before opening AI.
- Compare your choice with the AI's candidates.
- Ask for the reason in one simple sentence.
- Ask the flip condition: “What change would make the opposite decision correct?”
- Tag the error: efficiency, defense, push/fold, value, call, riichi, or placement.
- Review recurring tags instead of collecting endless new tips.
The most useful hands are usually deal-ins, close push/fold spots, call/no-call decisions, large disagreements with the model, and South/final-hand placement situations.
Mahjong Soul's MAKA can analyze ranked four-player logs and display evaluations and advice. Used after the game, that is ideal for finding repeated decision patterns rather than blindly copying one move.
10. Statistics worth checking
| Metric | What it can reveal |
|---|---|
| 1st/2nd/3rd/4th rates | Where placements are clustering |
| Hand win rate | How often hands convert to wins |
| Deal-in rate | Defensive risk and losses |
| Riichi rate | How often closed tenpai becomes pressure/value |
| Call rate | Speed and open-hand tendencies |
| Average winning value | Quality, not just quantity, of wins |
| Dealer vs non-dealer | Whether dealer value is being captured |
| South/final hand splits | Placement-management quality |
Even better: condition push/fold data on turn, shanten, wait quality, hand value, opponent riichi, dealer status, and current score.
Only then does “I am not getting first lately” become an analyzable problem.
11. Live AI assistance is a different issue: use AI for post-game review
Training with AI and letting an external tool choose moves during ranked play are not the same thing.
Mahjong Soul's official notices list use of external tools and other actions that damage fairness among examples of prohibited conduct and warn that accounts may be suspended.
The clean rule is simple: play the match yourself; analyze the log afterward with permitted review tools such as the official log system and MAKA.
Otherwise you can achieve the rare build known as “high rank, zero internal rules”: Yaku? Ask the AI. Scoring? The server probably knows.
12. Summary: stack the 2% dominoes
Research gives a less glamorous but more useful picture:
- Equal four-player Mahjong starts from a 25% first-place baseline.
- Even Suphx-level performance is around 30% firsts, not 50%.
- Strong policy can produce more firsts and fewer lasts at the same time.
- “Few firsts + few lasts” does not by itself prove over-defense.
- Short samples are noisy.
- Heavy AI outsourcing can separate performance from understanding.
- Ask AI not only “what?”, but “why?” and “what would flip the answer?”
High-level Mahjong is a game of stacking tiny edges.
Someone saves one safe tile. Someone skips one bad call. Someone calculates one final-hand condition correctly.
The dominoes are boring. The rating table remembers every one.
13. References
- Li, J. et al. (2020). Suphx: Mastering Mahjong with Deep Reinforcement Learning. arXiv:2003.13590. https://arxiv.org/abs/2003.13590
- Microsoft Research. Suphx: The World Best Mahjong AI. https://www.microsoft.com/en-us/research/project/suphx-mastering-mahjong-with-deep-reinforcement-learning/
- Microsoft Research Asia (2019). Mahjong AI Suphx. https://www.microsoft.com/en-us/research/articles/mahjong-ai-suphx/
- Mahjong Soul. MAKA AI-assisted game-log analysis. https://mahjongsoul.com/news/254
- Mahjong Soul. Notice on unfair conduct. https://mahjongsoul.com/news/24
- Risko, E. F., & Gilbert, S. J. (2016). Cognitive Offloading. Trends in Cognitive Sciences, 20(9), 676–688. https://doi.org/10.1016/j.tics.2016.07.002
- ChatGPT as a cognitive crutch: Evidence from a randomized controlled trial on knowledge retention (2025). Social Sciences & Humanities Open, 12, 102287. https://doi.org/10.1016/j.ssaho.2025.102287
