1. I only wanted to go face. An AI lab appeared instead.
The original goal was painfully simple: find the strongest current Shadowverse: Worlds Beyond deck. Collect public lists, freeze the rules, let CPUs play thousands of games, and read the answer.
Then reality intervened. Correct rule execution is not strong play. Set 8 was pinned with an 826-card database, engine commit, environment hash, legal decks and deterministic seeds. A historical 12,480-game baseline reached 100% rule coverage with zero unsupported cards and zero known rule gaps, yet the raw bot produced absurd class averages: Sword around 79.8%, Forest 63.7%, Dragon 60.9%, Rune only 25.6%.
Ten thousand repetitions do not convert a mistake into truth. They can merely create a machine that is extremely consistent at being wrong.
2. Human advice made the bot weaker
human-prior-v1 encoded ordinary strategic advice: preserve important resources, respect setup, protect win conditions, and avoid wasting evolution resources. The exact A/B used 2,016 paired units and 4,032 engine games. Baseline was 50.99%; human-prior-v1 was 49.36%, a -1.64 point change.
The bot listened to humans and became worse.
The reason was obvious in hindsight: human advice is conditional. “Hold this card” stops being correct when you die now. “Take the board” does not mean destroy your only combo line. Fixed weights preserved the sentence while murdering the context.
3. Adaptive v2 improved the average and broke Witch
Adaptive v2 switched among LETHAL, SURVIVE, STABILIZE, SETUP, RESOURCE and PRESSURE, with shallow search, hidden-hand approximation and future-value terms. Overall it gained +0.74 points over v1. Runecraft, however, lost -6.25 points and crossed the preregistered -5-point leader floor. Rejected.
T11 later scored +0.30 versus v1 but -2.38 versus reference-v5 and -16.67 for Abysscraft. Rejected again.
The average grade said “pass.” One subject was actively on fire.
4. Root-cause analysis arrested the researchers first
One real implementation bug was found: max_nodes=160 was configured but not actually wired into the planner, and 9,067 decisions exceeded the nominal limit. Connecting it did not repair broad performance.
Then a causal-analysis bug appeared. An intervention labeled “loss → win” had been normalized from a fixed scheduled player. From the perspective of the actor whose action changed, it was actually win → loss. The analysis had arrested the victim.
Actor attribution, direction handling and actor-relative invariants were repaired. The analysis platform became safer. The bot’s performance root cause remained at large.
5. After 31,406 games, “unresolved” became a legitimate conclusion
The Set 8 policy program closed after 20 experiment records and 31,406 confirmed engine games. Among 439 discordant units, 0 had a fully causally explained performance cause, 11 were partially explained, and 428 remained unresolved.
The project deliberately stopped at PAUSED_COMPLETE_NO_PROMOTION. human-prior-v1 stayed active. The planned 8,064-game final unseen holdout remained sealed and unused.
The main achievement was not a superhuman bot. It was a system that refuses to promote an AI merely because it looks better on a convenient slice.
6. Then 52 market decks fought for 46,900 games
The practical question returned: which current decks are strongest, and what should each of the seven leaders play? The market corpus contained 52 decks, 25 archetypes and all seven leaders. Quick screening, broad testing, local 1–3-card optimization, unseen-seed ranking and trace generation brought the market-analysis ledger to 46,900 actual engine games.
The final recommended seven decks were also tested in a direct unseen-seed round robin: 21 unordered pairs, 1,512 games, 72 games per non-diagonal cell, evenly split between reference-v5 and human-prior-v1 and between first and second player. Coverage was 100%; unsupported, rule-gap and hidden-information violations were all zero.
7. The “brain-dead-vs-brain-dead ranking” was born
| Rank | Recommended deck | Direct 7×7 average | Worst matchup |
|---|---|---|---|
| 1 | Buff Forest | 71.99% | Abyss 59.72% |
| 2 | Rally Sword | 65.97% | Forest 31.94% |
| 3 | Midrange Abyss | 54.17% | 40.3% |
| 4 | Artifact Portal | 54.17% | 36.1% |
| 5 | Ramp Dragon | 53.70% | Forest 25.0% |
| 6 | Evolution Haven | 31.71% | Dragon 12.5% |
| 7 | Calgidensula Witch | 18.29% | Forest 1.4% |
The deliberately rude nickname was the brain-dead-vs-brain-dead ranking. Decks where “make a strong play now” tends to remain good later are easy for the current AI. Decks that sacrifice present tempo to buy a stronger future appear much harder.
8. Human tournament data made Witch look like it came from another universe
Dexel Rising Cup #2 on August 8, 2026 featured 64 teams, 192 players and 384 submitted decks. Hand Witch appeared in 116 lists, roughly 30.2% of the entire field, making it the dominant human tournament archetype. AF/Storm Portal followed at 52, Evolution Haven at 34, Rally/Evolution Sword at 29 and Ramp Dragon at 25.
Humans treated Witch as a top metagame deck. The simulator’s Calgidensula Witch: 18.29%, last place. The two archetype labels must not be assumed to represent identical forty-card lists, but the gap between human Witch performance/usage and current AI Witch play is impossible to ignore.
Rally Sword, meanwhile, looked strong both in human competition and in the simulator. A rough but useful practical hypothesis emerges: strong for simple AI + strong for skilled humans may indicate robustness to imperfect play.
9. What does the AI actually understand?
Before changing weights again, the project audited seven layers: rules, state, visibility, immediate value, future value, sequencing, search and planning.
The pattern was blunt. Across 27,810 traced decisions and 118,575 root candidates, only 8,878 candidates — about 7.49% — carried an explicit future-value signal. Rules/state/immediate value were supported 24/24. Sequencing was missing for 23/24 mechanics and planning for 20/24.
The bot is reasonably good at asking “what is best now?” It is much worse at asking “what does this action make possible later?”
10. Witch knows that Spellboost exists; that is not the same as understanding Spellboost
Spellboost rule handling and underlying state existed. Visibility and future value were only partial. Sequence and plan representations were missing.
DRAW → SPELL and SPELL → DRAW can create different futures: a card drawn first may receive the later Spellboost; a card drawn afterward cannot retroactively receive a boost that already happened.
Across 12 synthetic fixtures, DRAW_FIRST_BETTER occurred 4 times, SPELL_FIRST_BETTER 0 times and EQUIVALENT 8 times. Fifty real representative states were classified overall as STATE_DEPENDENT. The correct fix is not “always draw first,” but observe, then re-plan.
11. Dragon: +1 max PP buys future action rights
Ramp can lose tempo now while unlocking expensive cards, multi-card turns, healing, removal and finishers later. The 6→7 transition also crosses the Awakening threshold.
The audit found 12 Dragoncraft MYOPIC_REVERSAL diagnostics: six ramp and six Awakening. Immediate control initially ranked higher, while t+2 diagnostic continuation could reverse the order. Hard-coding “ramp more” would simply create a new flavor of stupidity.
12. Every other class had its own version of the same problem
Forest has combo thresholds, low-cost generation and bounce ordering. Sword has Rally thresholds such as 9→10 and multi-summon interactions. Abyss cares about Shadows, the composition of the reanimation pool and HP as a resource. Haven has Countdown and Act, where not using an ability preserves future choice. Portal’s zero-cost tokens and copy effects create enormous same-turn sequence spaces.
EP, SEP, Extra PP, the nine-card hand limit, five board slots and five leader-area slots are also option-value problems: using a resource now trades against preserving future choices.
13. So we built an actual three-turn closed-loop planner
The next experiment enforced a simple contract: actually advance the engine a minimum of three future turns. Not three actions, not three cards, not depth three. The remainder of the current turn is followed by the opponent turn, the next self turn and the next opponent turn.
The plan was closed-loop, not a fixed three-turn script. Execute one action, observe draws/generation/random effects, then recompute a new three-turn plan. The design resembles receding-horizon / model-predictive control.
Future own draws and the opponent’s true hidden hand were not exposed. Belief sampling was used, and FutureRelevantState carried public Spellboost, Extra PP, Countdown, Act, Rally, Shadows and destroyed-follower-pool information. No fixed optionality, Witch or ramp bonuses were added.
14. Implementation succeeded. Performance collapsed.
This produced the most useful failure yet.
The Ablation stage used 56 actual engine games. The planner evaluated 3,979 / 3,979 root actions, reached the full three-turn horizon for 100% of roots, and recorded zero hidden-information violations. It performed 81,621 observation-triggered replans and changed plans 35,821 times. The implementation suspicion — “maybe it is not really looking three turns ahead” — largely disappeared.
Performance did not. v1-base and widen were -37.5 points versus human-prior-v1; replan, state and robust were -25.0 points. Every candidate violated the preregistered -10-point floor. Correct result: REJECT_ACTIVE_UNCHANGED. Targeted, balanced broad, market 7×7 and final holdout were not run.
We had built a bot that could look farther ahead while becoming more confidently wrong.
15. This does not prove that three turns are enough
Two caveats matter. First, the experiment removed the suspicion that the implementation failed to reach three turns; it did not prove that three turns are a sufficient strategic horizon. Second, the improvement from base -37.5 to replan -25.0 is a directional signal only: each candidate had just eight conditions.
When the old 12 Dragon MYOPIC_REVERSAL diagnostics were replayed with the real three-turn engine search, only 1/12 actually changed ranking toward ramp. The old proxy was therefore not a ground-truth label.
The leading suspects shifted from search reach to the terminal/leaf evaluator that scores the three-turn future, plus the opponent response model used inside the tree.
16. Next question: “You called this future -150. What happens if we actually play it out 100 times?”
The next round should not immediately change evaluator weights. First, freeze each three-turn leaf and decompose every evaluation term from raw feature through normalization, weight and final contribution. Track which term creates extreme values such as -150.
Then restart from the identical leaf and roll the game to termination 32–128 times to estimate an empirical policy-conditioned win rate. Raw scores are not probabilities, so fit a separate score→probability calibration and measure Brier score, log loss and reliability. Separately measure ranking quality with Spearman, Kendall, same-root best-action accuracy, pairwise concordance and decision regret.
Finally cross-play opponent responses using human-prior-v1, reference-v5 and a search-based stress response. That separates four very different failures: correct state but bad scoring; unrealistic opponent responses; sensible features with bad scales/weights; or strategic information that the leaf evaluator simply cannot represent. Only then should horizon 4/5 be reconsidered.
17. Conclusion: the future is visible; the grading rubric is suspicious
The project’s current state fits one sentence:
The AI has moved from “what is strong now?” to “I can simulate three turns ahead,” but we still cannot trust how it grades the future it sees.
The journey began with Sword at 79.8%, continued through human advice making the bot weaker, Adaptive policy breaking Witch, causal analysis confusing actor perspective, a 46,900-game market study creating the “brain-dead ranking,” and a Future Optionality audit revealing missing order and planning.
Then we gave the bot real three-turn foresight.
It got worse.
We only wanted to go face. The next research topic is Brier score, leaf-value calibration and clustered rollout validation. Apparently Shadowverse is graduate school now.
Data appendix
- Fixed Set 8 card DB: 826 cards
- Historical baseline: 12,480 engine games
- Set 8 policy research: 31,406 confirmed actual engine games
- Policy final status:
PAUSED_COMPLETE_NO_PROMOTION - Active policy:
human-prior-v1 - Market corpus: 52 decks / 25 archetypes / 7 leaders
- Market analysis: 46,900 actual engine games
- Direct recommended 7×7: 1,512 games
- Future Optionality audit: 27,810 decisions / 118,575 root candidates / explicit future-value coverage 7.49%
- Mechanics audit: 24 mechanics; sequence MISSING 23, plan MISSING 20
- Dragon MYOPIC_REVERSAL diagnostics: 12 = ramp 6 + Awakening 6
- Three-turn Ablation: 56 actual engine games
- Root action coverage: 3,979 / 3,979
- Three-turn completion: 100%
- Observation replans: 81,621
- Plan changes: 35,821
- Result:
REJECT_ACTIVE_UNCHANGED - Old and new final holdouts:
NOT_RUN
References
- Shadowverse: Worlds Beyond official Deck Portal and class-mechanics pages
- Dexel Rising Cup #2 deck lists and metagame report, Aug. 8, 2026
- Silver & Veness (2010), Monte-Carlo Planning in Large POMDPs (POMCP)
- Cowling, Powley & Whitehouse (2012), Information Set Monte Carlo Tree Search
- Model Predictive Control / Receding Horizon Control literature
- Calibration literature: Brier score, reliability diagrams, grouped/clustered validation
