1. Decide the real goal first: build the game, or build a laboratory that studies winning?
A research simulator does not need to start with card art, animations, voice lines, dragging, and flashy evolution scenes.
If the question is which deck is stronger or which move wins more often, the core loop is enough:
current state
→ enumerate legal actions
→ choose one
→ apply rules
→ next state
Repeat until someone wins.
The product is a numerical laboratory, not necessarily a pretty game client.
If a trustworthy engine already exists, reuse it and build the research layer. If not, build the smallest correct rules engine first.
2. Define inputs and outputs before writing code
Inputs can include card database, 40-card decks, first/second player, random seed, play policy, environment snapshot, and number of games.
Outputs can include winner, turn count, matchup win rate, first/second gap, average and worst matchup, brick rate, key-card-missing penalty, common mistakes, and useful deck changes.
Without this, you may build a huge toy that moves cards but answers no question.
A simulator starts with the research question.
3. Build the card database — data is not behavior
Store structured fields such as ID, name, class, cost, type, stats, rarity, set, rules text, evolved form, related cards, and tags.
JSON or SQLite is enough.
But text saying “Fanfare: deal 4 damage” does not implement targeting, missing targets, damage reduction, death triggers, or effect ordering.
The database department collects manuals.
The engine department makes the world obey them.
An employee directory does not mean anyone has started working.
4. Represent the game as state, not pixels
A machine-readable state may contain turn, active player, leader HP, PP, hand, deck, board, graveyard, Evolution resources, Crests, amulets, counters, and temporary effects.
Visual card position is usually irrelevant.
Attack availability, Ward, Evolution state, and temporary effects matter.
A good model stores the meaning of the world.
5. The rules engine is a state-transition machine
Play cards, attack, evolve, use Act, select targets, end turns, draw, refresh PP, and check victory as state transitions.
The engine does not need to be smart.
It first needs to be correct even when a stupid agent performs a legal action.
Keep the game’s physics separate from the AI’s brain.
6. Implement effects as generic rules plus exceptions
Create reusable operations such as deal damage, heal, draw, buff, summon, destroy, banish, gain Ward, and gain Storm.
Build most cards from these pieces and use explicit overrides for genuinely special cards.
Count unsupported, partial, and runtime-gap cases.
For research, failing closed is better than silently inventing behavior.
7. Generate legal actions carefully
Candidate generation should include playable cards, valid targets, attacks, Evolution, Act, and turn end.
It may also need hold, draw-first, self-clear, board-slot-clearing, and alternate sequences.
A perfect evaluator cannot select a correct action that was never generated.
Before blaming the scoring model, ask:
Was the correct move present in the candidate set?
8. Seeds, replays, and hashes make experiments reproducible
Fix random seeds and record engine commit, card DB hash, deck hash, policy hash, environment hash, schedule hash, and seed hash.
The same conditions should reproduce the same game.
Otherwise “the new AI improved” may actually mean the database, deck, code, or luck changed.
A major research failure is forgetting what was actually compared.
9. The first AI can be simple — build a stable reference policy
A simple baseline is enough:
lethal exists → take it
about to die → defend
play efficient curve
board ahead → pressure face
board behind → trade
Freeze it as a reference policy.
Its purpose is comparison, not perfection.
Later policies can take the same exam.
10. Human strategy should be a prior, not a commandment
Human guides provide mulligans, holding rules, EP/SEP reservation, combo setup, opponent swing turns, hand limits, board space, and future lethal.
But “always hold this” can kill the bot when spending it is necessary to survive.
Encode knowledge as scoring signals such as hold value, future combo value, survival value, and premature-spend penalties.
The guide is not the answer key.
It is a map for the search algorithm.
11. Tree search, beam search, and node budgets let the AI look ahead
With ten actions per step, the tree grows 10, 100, 1,000, 10,000.
Use depth limits, beam width, node budgets, and pruning.
The danger is deleting a setup line that looks weak now but wins two turns later.
Separate immediate, future, survival, setup, lethal, and resource value.
The hard part is not seeing every future.
It is not deleting the important future too early.
12. Hidden hands require belief, POMDP ideas, and Monte Carlo sampling
Giving the policy the opponent’s real hidden hand is cheating.
Use public information to build plausible worlds, such as AoE 30%, Ward 20%, nothing relevant 50%.
This is a belief-state idea.
A hidden-information card game resembles a POMDP.
A practical implementation can sample plausible hands and simulate futures repeatedly using Monte Carlo ideas.
13. Design large simulations as paired A/B experiments
Give both policies the same decks, opponent, first/second orientation, and seed.
Run the old and new policy on the same problem.
Use quick tests first, development next, and an unseen holdout last.
Large samples exist to reduce noise while keeping the comparison fair.
14. Measure more than win rate
Track average win rate, meta-weighted win rate, worst matchup, lower-tail average, matchup variance, first/second gap, brick rate, key-card-missing drop, average game length, and confidence intervals.
CVaR-style metrics examine bad outcomes.
Bootstrap and McNemar-style tests help paired A/B analysis.
Do not search only for the highest average student.
Search for the student who does not fail a required subject.
15. Deck optimization means intelligently shrinking the search space
You cannot exhaustively test every 40-card combination.
Start from tournament and popular decks, then explore one- to three-card swaps, role-equivalent replacements, and matchup techs.
Double Oracle repeatedly adds strong responses.
PSRO treats deck + play policy as the strategy.
No-regret learning reduces long-run regret.
Robust optimization protects the matchup floor.
Deck search is a candidate-compression factory.
16. Record why the AI lost
An action trace can record candidates, scores, selected action, objective, EP/SEP reservations, early combo-piece spend, missed lethal, hand burn, and board lock.
Ablation disables components one at a time.
Counterfactual evaluation compares action A and B from the same public state.
If A merely appears in more losses, that is correlation.
If replacing A with B repeatedly improves outcomes from the same state, the evidence becomes more causal.
17. Treat each expansion as a new environment snapshot
Store environment ID, date, format, legal sets, and hashes for the engine, DB, policy, deck corpus, and meta.
Diff the new DB against the old one.
Identify changed cards, affected decks, and affected matchups.
If one card changed, rerun the affected region. If only metagame weights changed, reuse results and re-aggregate.
Do not rebuild the laboratory every expansion. Make only the affected department work overtime.
18. Keep software architecture small and CLI-first
A practical layout separates upstream, deck, policy, simulation, diagnosis, and statistics, plus config, data, reports, docs, and tests.
Useful CLI commands include status, doctor, db check, env validate, simulate, experiment run/analyze, optimize, diagnose, and context build.
Keep engines, caches, and huge raw traces out of Git.
Commit locks, configs, manifests, checksums, summaries, reproduction commands, and small fixtures.
GitHub becomes the laboratory’s external memory.
19. A practical zero-to-one roadmap
1 read card DB
2 finish one minimal correct match
3 legal actions and targeting
4 deterministic seed/replay/hash
5 simple policy + 100 games
6 expand coverage
7 real deck corpus + matchup matrix
8 human priors
9 tree search + belief
10 paired A/B + holdout
11 deck optimization
12 action diagnosis + expansion diff
Do not begin with one million games, MCTS, reinforcement learning, and global deck optimization all at once.
First make one match correct. Then 100. Then 10,000.
Do not build a machine that is wrong ten thousand times per minute.
20. Conclusion — a card-game simulator is really a research process
The hard part is correctly representing the world, enforcing rules, preserving valid candidate actions, handling hidden information without cheating, comparing policies fairly, measuring weakness beyond average win rate, observing failure causes, and surviving future expansions.
Card DB is memory. Rules engine is physics. Policy is the brain. Search is foresight. Statistics are the report card. Action traces are security cameras. Git is the lab notebook. AI agents are researchers.
The human keeps asking:
“Can I actually trust that number?”
It began with “I want to Storm face.”
Somehow, a card-game AI laboratory appeared.
The final move is still simple.
Remove Ward. Evolve.
Storm face.
