A card-game simulator sounds terrifying if you imagine entering every card by hand: name, cost, stats, text, special behavior, repeat hundreds of times.
The picture changes completely when card data already exists in structured form and a usable battle engine is open source. The database can be loaded in bulk, common effects can reuse existing rule handlers, and only genuinely new mechanics need custom work.
In this case, Beyond Decks already provided card-data integration, deck tools, deterministic seeds, replays, battle simulation, AI logic, and benchmarks. The task was not “build the game again.” It was attach a research layer and a better brain to an existing sandbox.
That is why progress can happen even while the human is away from the desk. The role becomes less “game programmer” and more “director of a very caffeinated laboratory.”
1. Remove the animations and a card-game match becomes hilariously short
A real client spends time on draw animations, dragging cards, voice lines, evolution effects, attack motion, damage popups, and human thinking time.
A simulator mostly does this:
state A
→ enumerate legal actions
→ choose an action
→ update numbers/state
→ state B
Repeat until someone loses.
A match that takes many minutes visually can finish in a fraction of that time internally. No crowd. No voice line. No dramatic evolution cut-in. The CPU may finish several matches while a human player is still saying hello.
The expensive part is not presentation. It is how much future reasoning you ask the AI to perform.
2. The base simulator turned out to be serious, not a toy
The existing simulator already had features such as deterministic seeds, replay inspection, target selection, tactical look-ahead, full-turn planning, a lethal solver, class mechanics, benchmarks, card-coverage audits, and runtime checks.
Its Battle Engine v5 is substantial. The author also describes the Battle Sim as share-ready.
The critical caveat is equally explicit: the AI is not intended to be a perfect tournament oracle. It is an intermediate player.
That is almost ideal for this project. The board, rules, cards, replay system, seeds, and battle machinery are largely present. The main missing layer is human-like decision quality.
Instead of reimplementing hundreds of cards, effort can go into the reasoning layer that sits on top.
3. 12,480 games produced Royal near 79.8%—large samples can amplify policy bias
A previous fixed-engine run completed 12,480 games with good rules coverage, yet the class results were extreme. Royal landed near 79.8% while Witch was around 25.6%.
The wrong conclusion would be, “Twelve thousand games! Royal is objectively unbeatable!”
The better question is: what kind of deck is easy for this AI to pilot?
Royal may look like:
play follower
take board
face is open
hit face
win
while a badly handled combo deck may look like:
combo piece? playable now—use it
evolution resource? value now—spend it
hand limit? never heard of it
three turns later? future me's problem
The danger of large samples is not random noise. It is estimating the wrong policy with extremely small error bars.
4. Human play knowledge was added as a way of thinking, not a lookup table
The next step was to collect actual play knowledge: mulligans, cards to hold, evolution reservation, hand and board-space management, matchup-specific offense/defense switching, opponent swing turns, alternative win conditions, two-to-three-turn lethal planning, and common mistakes.
But hard-coding “always play card X here” creates a brittle guidebook bot.
Instead, human knowledge becomes scoring signals:
hold_value +20
future_combo_value +30
opponent_counter_cost +25
premature_use -40
Human strategy is not the answer key. It is a bias that guides search toward sensible lines.
5. The current run is 4,032 actual games—2,016 is the number of comparison cells
The current A/B design uses 56 scheduled matchup groups, two directions, and 18 deterministic seeds.
56 × 2 × 18 = 2,016 A/B conditions
For every condition, both policies play:
reference policy: 1 game
human-prior policy: 1 game
So the engine processes 4,032 actual games.
The number 2,016 is not sacred. It represents 2,016 matched exam questions given to two different AIs.
Fixed seeds and first/second orientation reduce the chance that the new policy looks better merely because it drew better.
6. Stop before the big run: regressions and smoke tests matter more than they look
The system did not immediately launch all 4,032 games.
First it fixed a regression where finishers were being used too early. Setup cards had to be prioritized under the correct conditions. Then short first/second real-engine matches checked the pinned card database and real decks.
The gates include card coverage, unsupported-card count, rule gaps, data hashes, duplicate IDs, missing references, and corpus-version mismatches.
This is boring and essential.
After learning that a bad policy can happily generate 12,000 games of precise nonsense, the new rule is simple: verify the assumptions before paying for the sample size.
7. Human-prior v1 is still not truly adaptive—the next challenge is changing goals every turn
The current human-prior policy already values holding cards, setup, future combos, and opponent responses. But human play requires another layer: goals must change as the position changes.
The same ramp deck might want:
normal position → ramp
opponent threatens lethal → stop ramping, defend
opponent is exhausted → attack
finisher in hand but setup incomplete → hold
setup complete → release finisher
“Dragon ramps” is not enough.
The AI must recompute what matters most right now every turn.
8. Adaptive AI can switch between ATTACK, SURVIVE, SETUP, RESOURCE, and LETHAL
A practical adaptive policy can maintain dynamic modes:
ATTACK
SURVIVE
SETUP
RESOURCE
LETHAL
Every turn it scores those modes using life totals, board, hand, resources, evolution points, opponent pressure, estimated incoming damage, and combo readiness.
If the player is at 6 life facing 14 damage with no protection, SURVIVE should dominate. Even if a ramp card is technically efficient, removal or protection becomes more valuable.
This is what “adaptive” really means.
Not “never play the finisher before turn six,” but compare the value of using it now with the value of waiting.
9. Look two or three turns ahead—but prune the tree with human knowledge
Full game-tree search explodes quickly.
A practical pipeline is:
50 legal actions
↓
human priors reduce to 20
↓
cheap evaluator reduces to 10
↓
search only the top 10 for 2–3 turns
An action worth +8 immediately but leading to certain death next turn should lose to an action worth +3 now that preserves survival and counterplay.
Human knowledge becomes useful twice: it improves scoring and it reduces the number of future branches worth exploring.
10. Giving the AI the hidden hand would be cheating—model “maybe they have it”
A card-game AI should not simply read the opponent’s hidden hand.
Instead, public information can create weighted possible worlds: leader, archetype, cards already played, remaining deck, hand size, and previous actions.
For example:
opponent has AoE: 30%
opponent does not: 70%
Evaluate the same action across several plausible hands.
That produces behavior closer to human reasoning:
“They might have the clear, so I should respect it—but if I respect everything, I never win.”
11. Order matters—do not upgrade the AI halfway through the current 4,032-game experiment
Adding Adaptive Policy v2 during the current A/B would invalidate the experiment. Early matchup groups and late groups would be using different policies.
The correct order is:
1. freeze human-prior v1
2. finish 4,032 games
3. analyze remaining mistakes
4. build Adaptive Policy v2
5. quick A/B
6. run the ~2,016-condition paired test again
7. run a 12k-class full evaluation
8. optimize 40-card lists
9. generate leader/archetype matchup diagnostics
The most important step is number three. Build v2 from observed failures, not from a pile of features that merely sound intelligent.
12. The destination is not “a CPU that memorized the guide”—it is a CPU allowed to break the guide
The final reasoning stack looks like:
human knowledge → prune sensible candidates
game-tree search → inspect future states
evaluation → compare present and future
opponent-hand model → reason under hidden information
dynamic mode → switch between attack, defense, setup, and lethal
Then the AI can say:
“Ramp normally, but defend if I die.” “Hold the piece normally, but spend everything if lethal exists.” “Develop normally, but keep some resources if the clear is likely.”
It uses the guide, but the board state can overrule the guide.
That is when “adaptive play” starts to look real.
13. This pattern is bigger than one card game: build a winning laboratory instead of rebuilding the game
The same approach transfers to other games.
For Pokémon, reuse an existing simulator and optimize team building, selection, switching, and move choice. For card games, optimize deck construction, mulligans, resource timing, and matchup policy.
The key distinction is:
rebuilding the game and researching how to win the game are different jobs.
If a reliable simulator exists, borrow it.
Let the human decide what looks suspicious, what should be compared, and what failure matters. Let automated coding and simulation handle implementation, regression tests, and thousands of games.
You do not need to rebuild the stadium.
Build the laboratory that figures out how to win inside it.
And if Royal reaches nearly 80% again, check one thing first:
Did the bot just discover that “play unit, hit face” is its favorite subject again?


