When people build a card-game simulator, the first instinct is usually: run more games. One hundred becomes one thousand, one thousand becomes ten thousand, and suddenly a giant CSV makes everything look scientific.
The trap is simple: if the policy is wrong, more simulations only estimate the behavior of a wrong policy with smaller statistical error.
In our case study, a fixed rules engine completed 12,480 games and passed the rules audit well, yet the raw results were extreme: Royal around 79.8% and Witch around 25.6%. That does not mean Royal had an 80% real-world win rate. It means the play policy itself deserved suspicion.
The first rule is therefore simple: before increasing the game count, verify that the CPU is actually playing the card game competently.
1. A simulator is not one AI: separate rules, decisions, and scoring
A useful simulator has at least three layers.
- Rules engine: decides what actions are legal and resolves damage, targets, evolution, and effects.
- Play policy: chooses one action from the legal actions.
- Evaluator: decides whether a deck or policy was actually good.
If these are mixed together, a bad result becomes impossible to diagnose. Is the card effect wrong? Is the AI making bad decisions? Is first-player bias distorting the sample? Is the scoring function rewarding fragile decks?
Even a perfect rules engine can produce nonsense when the policy fires combo pieces too early. A strong policy can still look stronger if it receives more first-player games. A weak evaluator can crown a deck that has a high mean win rate but collapses against one common matchup.
So test legality, decision quality, and evaluation quality separately.
2. Build the rules engine first: physics before intelligence
Before making a clever opponent, build a consistent game world.
The engine must handle deck legality, copy limits, class restrictions, resource points, draw, mulligan, board limits, attacks, Ward-like protection, Storm/Rush-like attack permissions, evolution systems, targeting, generated cards, tokens, special mechanics, and win conditions.
Each card or mechanic should have an explicit support status:
- Full
- Partial
- Unsupported
- Runtime gap
If an unsupported card is present, failing closed is better than printing “57.3% win rate” from a broken rules model. A simulator needs the courage to say, “I cannot evaluate this deck yet.”
3. A play AI needs more than “what is strongest right now?”
Card games are difficult because the best move often depends on the future.
A useful policy should consider board value, hand value, one-to-three-turn planning, lethal probability, the value of holding combo pieces, evolution-resource reservation, likely opponent answers, hand-overflow risk, board-space risk, and alternative win conditions.
A policy that only says “play a playable card, use the biggest body, attack face when possible” can look surprisingly good with straightforward tempo decks. It falls apart with combo, control, and resource-management decks.
That difference is the heart of the “Royal unga-bunga AI” failure later in this article.
4. Human strategy should be a hint, not a prison
Hard-coding every strategy guide as an if-statement creates a brittle bot.
Instead of saying “always keep this card” or “always evolve on turn X,” encode human knowledge as bonuses, penalties, and priorities.
Examples:
- raise the value of a board clear in an aggressive matchup;
- penalize spending a key combo piece too early;
- reward saving evolution resources for a known power turn;
- reduce overcommitting before a likely opponent swing turn;
- increase future value when a two- or three-turn lethal setup exists.
Tag each piece of human knowledge with pack, format, leader, archetype, matchup, first/second, game phase, source date, and confidence. When experts disagree, keep multiple candidate policies instead of flattening them into one fake truth.
The goal is not “obey humans.” The goal is search in a better direction.
5. Fair A/B testing means giving both AIs the same exam
If the old AI bricks and the new AI draws perfectly, the comparison is meaningless.
Use deterministic random seeds so both policies face equivalent random conditions, then reverse first and second player.
seed 001: old AI first / new AI second
seed 001: new AI first / old AI second
seed 002: old AI first / new AI second
seed 002: new AI first / old AI second
...
This pushes the difference toward decision quality instead of luck.
Also record diagnostic mistakes: missed lethal, burned cards, board-slot jams, wasted evolution resources, premature combo use, bad mulligans, and failure to respect the opponent’s counter-turn.
You need to detect the wonderful category of result called: “the AI won, but the play was still stupid.”
6. Why 2,016 games? It is multiplication, not numerology
There is nothing sacred about 2,016.
One planned paired A/B profile was derived from a coverage budget like this:
56 matchup test cells
× 18 random seeds
× 2 paired first/second runs
= 2,016 games
The important idea is to choose the coverage first and let the game count follow.
A practical system should have multiple profiles:
- quick: tens to hundreds of games for smoke tests;
- standard: around two thousand for policy A/B and major matchups;
- full: ten thousand or more for broader environment estimates.
Running the full profile first and discovering afterward that the AI is terrible is an expensive way to manufacture very accurate nonsense.
7. Deck optimization is not brute-forcing every possible 40-card list
The full deck space is too large to enumerate. Start from real competitive or public deck lists, then generate nearby candidates.
Try one-card, two-card, or three-card swaps, count adjustments, role-equivalent replacements, matchup tech cards, and package-level changes.
Do not optimize only mean win rate. A better objective considers average performance, downside risk, first/second stability, and the floor against bad matchups.
A deck with a 62% average but a 20% matchup into a popular opponent may be worse in practice than a 58% deck whose worst major matchup is 45%.
So the output should not say “the absolute strongest deck in the universe.” Say: “the best evaluated list among these candidates under this card snapshot, policy, meta weighting, and simulation budget.”
8. Diagnosis is more useful than a naked win-rate number
A strong user-facing system should explain how to play, not just rank decks.
Given a 40-card list and an opposing leader or archetype, it can return estimated win rate and uncertainty, first/second difference, mulligan keeps and throws, early/mid/late goals, main win condition, cards to hold, removal priority, evolution-resource timing, likely opponent counter-turns, two-to-three-turn lethal routes, common mistakes, alternative lines, and tech candidates.
That turns a simulator into a coach.
9. The Royal 79.8% incident: maybe “play unit, hit face” was simply AI-friendly
The funniest failure was the extreme class split.
A plausible explanation is that simple proactive decks are easier for an immature AI.
Royal can sometimes get away with:
play follower!
win board!
face is open!
hit face!
we won!
Meanwhile a badly designed Witch policy may behave like:
combo piece? playable now! use it!
evolution resource? stronger now! spend it!
hand limit? never heard of it!
three turns from now? future me's problem!
If that is happening, Royal becoming king is not proof of the real metagame. It is evidence that different deck types place different cognitive demands on the policy.
When simulated results disagree sharply with reality, check whether long-horizon decks are disproportionately weak, whether the AI understands hold value, saves resources, manages hand and board space, and predicts opponent responses.
Finding “this AI is only good at unga-bunga tempo” is not a failed experiment. It is a successful diagnosis of the broken layer.
10. Build for environment changes: the final product is a laboratory, not a one-pack script
Store the current rotation as a versioned environment snapshot.
Environment 8
- card snapshot
- legal card pool
- engine version
- policy version
- meta decks
- human knowledge
- simulation results
When the next pack arrives, create the next environment and diff new cards, rotated cards, balance changes, restrictions, new archetypes, engine support, and meta weights.
Recompute only what is affected. If only meta weights changed, reuse the matchup matrix and reweight it. If a new card enters a deck, rerun that deck and relevant opponents. If the engine does not support a new mechanic, block promotion instead of producing fake certainty.
The end-to-end pipeline becomes:
official card data
↓
legal card snapshot
↓
competitive/public seed decks
↓
human knowledge as policy priors
↓
fair CPU-vs-CPU simulation
↓
local 40-card optimization
↓
matchup re-evaluation
↓
win rate + downside + mistake diagnosis
↓
versioned environment stored in GitHub
↓
next pack = update only the diff
The real goal is not a scary-looking win-rate table. It is traceability: which rules, AI policy, card database, seeds, and environment produced this number?
Reproducing one game is more valuable than blindly running ten thousand. Writing the assumptions is better than shouting “best deck.” And the system should make an AI mistake obvious rather than hiding it behind sample size.
If Royal approaches 80% again, ask one question before declaring a new dynasty:
“Does this bot think the whole game is just play a unit and hit face?”


