I Gave a Brain-Dead AI Three-Turn Foresight and It Got Even Worse — What 46,900 Market-Simulation Games and a Separate AI Research Program Taught Us About “Seeing the Future but Scoring It Wrong”

Subject: Shadowverse: Worlds Beyond

The original goal was painfully simple: find the strongest current Shadowverse: Worlds Beyond deck.

How reading tools work

Listen reads the article aloud. Speed read shows phrases in sequence at your chosen pace. Language practice compares available translations. Save keeps a bookmark in this browser; find it in the player’s bookmarks.

Share this article
Advertisement
Advertisement

1. I only wanted to go face. An AI lab appeared instead.

The original goal was painfully simple: find the strongest current Shadowverse: Worlds Beyond deck. Collect public lists, freeze the rules, let CPUs play thousands of games, and read the answer.

Then reality intervened. Correct rule execution is not strong play. Set 8 was pinned with an 826-card database, engine commit, environment hash, legal decks and deterministic seeds. A historical 12,480-game baseline reached 100% rule coverage with zero unsupported cards and zero known rule gaps, yet the raw bot produced absurd class averages: Sword around 79.8%, Forest 63.7%, Dragon 60.9%, Rune only 25.6%.

Ten thousand repetitions do not convert a mistake into truth. They can merely create a machine that is extremely consistent at being wrong.

2. Human advice made the bot weaker

human-prior-v1 encoded ordinary strategic advice: preserve important resources, respect setup, protect win conditions, and avoid wasting evolution resources. The exact A/B used 2,016 paired units and 4,032 engine games. Baseline was 50.99%; human-prior-v1 was 49.36%, a -1.64 point change.

The bot listened to humans and became worse.

The reason was obvious in hindsight: human advice is conditional. “Hold this card” stops being correct when you die now. “Take the board” does not mean destroy your only combo line. Fixed weights preserved the sentence while murdering the context.

3. Adaptive v2 improved the average and broke Witch

Adaptive v2 switched among LETHAL, SURVIVE, STABILIZE, SETUP, RESOURCE and PRESSURE, with shallow search, hidden-hand approximation and future-value terms. Overall it gained +0.74 points over v1. Runecraft, however, lost -6.25 points and crossed the preregistered -5-point leader floor. Rejected.

T11 later scored +0.30 versus v1 but -2.38 versus reference-v5 and -16.67 for Abysscraft. Rejected again.

The average grade said “pass.” One subject was actively on fire.

4. Root-cause analysis arrested the researchers first

One real implementation bug was found: max_nodes=160 was configured but not actually wired into the planner, and 9,067 decisions exceeded the nominal limit. Connecting it did not repair broad performance.

Then a causal-analysis bug appeared. An intervention labeled “loss → win” had been normalized from a fixed scheduled player. From the perspective of the actor whose action changed, it was actually win → loss. The analysis had arrested the victim.

Actor attribution, direction handling and actor-relative invariants were repaired. The analysis platform became safer. The bot’s performance root cause remained at large.

5. After 31,406 games, “unresolved” became a legitimate conclusion

The Set 8 policy program closed after 20 experiment records and 31,406 confirmed engine games. Among 439 discordant units, 0 had a fully causally explained performance cause, 11 were partially explained, and 428 remained unresolved.

The project deliberately stopped at PAUSED_COMPLETE_NO_PROMOTION. human-prior-v1 stayed active. The planned 8,064-game final unseen holdout remained sealed and unused.

The main achievement was not a superhuman bot. It was a system that refuses to promote an AI merely because it looks better on a convenient slice.

6. Then 52 market decks fought for 46,900 games

The practical question returned: which current decks are strongest, and what should each of the seven leaders play? The market corpus contained 52 decks, 25 archetypes and all seven leaders. Quick screening, broad testing, local 1–3-card optimization, unseen-seed ranking and trace generation brought the market-analysis ledger to 46,900 actual engine games.

The final recommended seven decks were also tested in a direct unseen-seed round robin: 21 unordered pairs, 1,512 games, 72 games per non-diagonal cell, evenly split between reference-v5 and human-prior-v1 and between first and second player. Coverage was 100%; unsupported, rule-gap and hidden-information violations were all zero.

7. The “brain-dead-vs-brain-dead ranking” was born

Rank Recommended deck Direct 7×7 average Worst matchup
1 Buff Forest 71.99% Abyss 59.72%
2 Rally Sword 65.97% Forest 31.94%
3 Midrange Abyss 54.17% 40.3%
4 Artifact Portal 54.17% 36.1%
5 Ramp Dragon 53.70% Forest 25.0%
6 Evolution Haven 31.71% Dragon 12.5%
7 Calgidensula Witch 18.29% Forest 1.4%

The deliberately rude nickname was the brain-dead-vs-brain-dead ranking. Decks where “make a strong play now” tends to remain good later are easy for the current AI. Decks that sacrifice present tempo to buy a stronger future appear much harder.

8. Human tournament data made Witch look like it came from another universe

Dexel Rising Cup #2 on August 8, 2026 featured 64 teams, 192 players and 384 submitted decks. Hand Witch appeared in 116 lists, roughly 30.2% of the entire field, making it the dominant human tournament archetype. AF/Storm Portal followed at 52, Evolution Haven at 34, Rally/Evolution Sword at 29 and Ramp Dragon at 25.

Humans treated Witch as a top metagame deck. The simulator’s Calgidensula Witch: 18.29%, last place. The two archetype labels must not be assumed to represent identical forty-card lists, but the gap between human Witch performance/usage and current AI Witch play is impossible to ignore.

Rally Sword, meanwhile, looked strong both in human competition and in the simulator. A rough but useful practical hypothesis emerges: strong for simple AI + strong for skilled humans may indicate robustness to imperfect play.

9. What does the AI actually understand?

Before changing weights again, the project audited seven layers: rules, state, visibility, immediate value, future value, sequencing, search and planning.

The pattern was blunt. Across 27,810 traced decisions and 118,575 root candidates, only 8,878 candidates — about 7.49% — carried an explicit future-value signal. Rules/state/immediate value were supported 24/24. Sequencing was missing for 23/24 mechanics and planning for 20/24.

The bot is reasonably good at asking “what is best now?” It is much worse at asking “what does this action make possible later?”

10. Witch knows that Spellboost exists; that is not the same as understanding Spellboost

Spellboost rule handling and underlying state existed. Visibility and future value were only partial. Sequence and plan representations were missing.

DRAW → SPELL and SPELL → DRAW can create different futures: a card drawn first may receive the later Spellboost; a card drawn afterward cannot retroactively receive a boost that already happened.

Across 12 synthetic fixtures, DRAW_FIRST_BETTER occurred 4 times, SPELL_FIRST_BETTER 0 times and EQUIVALENT 8 times. Fifty real representative states were classified overall as STATE_DEPENDENT. The correct fix is not “always draw first,” but observe, then re-plan.

11. Dragon: +1 max PP buys future action rights

Ramp can lose tempo now while unlocking expensive cards, multi-card turns, healing, removal and finishers later. The 6→7 transition also crosses the Awakening threshold.

The audit found 12 Dragoncraft MYOPIC_REVERSAL diagnostics: six ramp and six Awakening. Immediate control initially ranked higher, while t+2 diagnostic continuation could reverse the order. Hard-coding “ramp more” would simply create a new flavor of stupidity.

12. Every other class had its own version of the same problem

Forest has combo thresholds, low-cost generation and bounce ordering. Sword has Rally thresholds such as 9→10 and multi-summon interactions. Abyss cares about Shadows, the composition of the reanimation pool and HP as a resource. Haven has Countdown and Act, where not using an ability preserves future choice. Portal’s zero-cost tokens and copy effects create enormous same-turn sequence spaces.

EP, SEP, Extra PP, the nine-card hand limit, five board slots and five leader-area slots are also option-value problems: using a resource now trades against preserving future choices.

13. So we built an actual three-turn closed-loop planner

The next experiment enforced a simple contract: actually advance the engine a minimum of three future turns. Not three actions, not three cards, not depth three. The remainder of the current turn is followed by the opponent turn, the next self turn and the next opponent turn.

The plan was closed-loop, not a fixed three-turn script. Execute one action, observe draws/generation/random effects, then recompute a new three-turn plan. The design resembles receding-horizon / model-predictive control.

Future own draws and the opponent’s true hidden hand were not exposed. Belief sampling was used, and FutureRelevantState carried public Spellboost, Extra PP, Countdown, Act, Rally, Shadows and destroyed-follower-pool information. No fixed optionality, Witch or ramp bonuses were added.

14. Implementation succeeded. Performance collapsed.

This produced the most useful failure yet.

The Ablation stage used 56 actual engine games. The planner evaluated 3,979 / 3,979 root actions, reached the full three-turn horizon for 100% of roots, and recorded zero hidden-information violations. It performed 81,621 observation-triggered replans and changed plans 35,821 times. The implementation suspicion — “maybe it is not really looking three turns ahead” — largely disappeared.

Performance did not. v1-base and widen were -37.5 points versus human-prior-v1; replan, state and robust were -25.0 points. Every candidate violated the preregistered -10-point floor. Correct result: REJECT_ACTIVE_UNCHANGED. Targeted, balanced broad, market 7×7 and final holdout were not run.

We had built a bot that could look farther ahead while becoming more confidently wrong.

15. This does not prove that three turns are enough

Two caveats matter. First, the experiment removed the suspicion that the implementation failed to reach three turns; it did not prove that three turns are a sufficient strategic horizon. Second, the improvement from base -37.5 to replan -25.0 is a directional signal only: each candidate had just eight conditions.

When the old 12 Dragon MYOPIC_REVERSAL diagnostics were replayed with the real three-turn engine search, only 1/12 actually changed ranking toward ramp. The old proxy was therefore not a ground-truth label.

The leading suspects shifted from search reach to the terminal/leaf evaluator that scores the three-turn future, plus the opponent response model used inside the tree.

16. Next question: “You called this future -150. What happens if we actually play it out 100 times?”

The next round should not immediately change evaluator weights. First, freeze each three-turn leaf and decompose every evaluation term from raw feature through normalization, weight and final contribution. Track which term creates extreme values such as -150.

Then restart from the identical leaf and roll the game to termination 32–128 times to estimate an empirical policy-conditioned win rate. Raw scores are not probabilities, so fit a separate score→probability calibration and measure Brier score, log loss and reliability. Separately measure ranking quality with Spearman, Kendall, same-root best-action accuracy, pairwise concordance and decision regret.

Finally cross-play opponent responses using human-prior-v1, reference-v5 and a search-based stress response. That separates four very different failures: correct state but bad scoring; unrealistic opponent responses; sensible features with bad scales/weights; or strategic information that the leaf evaluator simply cannot represent. Only then should horizon 4/5 be reconsidered.

17. Conclusion: the future is visible; the grading rubric is suspicious

The project’s current state fits one sentence:

The AI has moved from “what is strong now?” to “I can simulate three turns ahead,” but we still cannot trust how it grades the future it sees.

The journey began with Sword at 79.8%, continued through human advice making the bot weaker, Adaptive policy breaking Witch, causal analysis confusing actor perspective, a 46,900-game market study creating the “brain-dead ranking,” and a Future Optionality audit revealing missing order and planning.

Then we gave the bot real three-turn foresight.

It got worse.

We only wanted to go face. The next research topic is Brier score, leaf-value calibration and clustered rollout validation. Apparently Shadowverse is graduate school now.

Data appendix

  • Fixed Set 8 card DB: 826 cards
  • Historical baseline: 12,480 engine games
  • Set 8 policy research: 31,406 confirmed actual engine games
  • Policy final status: PAUSED_COMPLETE_NO_PROMOTION
  • Active policy: human-prior-v1
  • Market corpus: 52 decks / 25 archetypes / 7 leaders
  • Market analysis: 46,900 actual engine games
  • Direct recommended 7×7: 1,512 games
  • Future Optionality audit: 27,810 decisions / 118,575 root candidates / explicit future-value coverage 7.49%
  • Mechanics audit: 24 mechanics; sequence MISSING 23, plan MISSING 20
  • Dragon MYOPIC_REVERSAL diagnostics: 12 = ramp 6 + Awakening 6
  • Three-turn Ablation: 56 actual engine games
  • Root action coverage: 3,979 / 3,979
  • Three-turn completion: 100%
  • Observation replans: 81,621
  • Plan changes: 35,821
  • Result: REJECT_ACTIVE_UNCHANGED
  • Old and new final holdouts: NOT_RUN

References

  • Shadowverse: Worlds Beyond official Deck Portal and class-mechanics pages
  • Dexel Rising Cup #2 deck lists and metagame report, Aug. 8, 2026
  • Silver & Veness (2010), Monte-Carlo Planning in Large POMDPs (POMCP)
  • Cowling, Powley & Whitehouse (2012), Information Set Monte Carlo Tree Search
  • Model Predictive Control / Receding Horizon Control literature
  • Calibration literature: Brier score, reliability diagrams, grouped/clustered validation

AdFind the work this article discusses

  • Shadowverse: Worlds Beyond

    Search results for Shadowverse: Worlds Beyond, the work this article discusses.

This article contains affiliate links (ads). About advertising As an Amazon Associate I earn from qualifying purchases.

Read this today

Each one answers a question readers of this article tend to ask next.

Browse all articlesMore on Card and board games

Advertisement

One more? Anything fun?

Since you're done reading: a couple of nearby stories and some totally different, fun ones.

  1. Nearby“But I Was Sleepy”The Nap That Leaves You Awake Before a Morning Health Check — How Much Can One Bad Night Shift the Numbers?
  2. Why Does Every One of My Attacks Lose?A Pyra/Mythra Main vs. Kazuya’s Intangibility and the Crocodile’s Belly
  3. Totally different, but funDo Not Become the Tutorial Senior at the Adventurers’ GuildWhy It Can Be Better to Use Polite Language with Juniors
  4. Why Is Malatang So Popular?I Tried It and Found a “Texture Game” Where the Broth Saves Everything
  5. I Want Shaving Out of My LifeHow Many Laser Sessions Does a Beard Really Take?
  6. What Does a Public Health Office Actually Do?From “Animals and COVID” to the City’s Public-Health Final Boss

Find other articles

All articles

Mendoi-chan

Who runs this site

Mendoi-chan

She turns friction at work and in everyday life into clear structure and practical next steps.

Advertisement

Latest articles

  1. 1The Underwear Thief Gets Caught, but the “Breeding Uncle” Is a Hero? The Manga Trims the Most Niche Fetishes While the Web Novel Unlocks “Imaginary Pregnancy”
  2. 2Is Weak Rote Memory the Same as Being Bad at Remembering?
  3. 3The “Young People Welcome” Workplace That Pushes Young People Out
  4. 4Humanity Is Not Winning This — I Treated PRAGMATA Like a Lunar Art Museum and Ended Up Terrified by Industrialized Machine Civilization
  5. 5For People Who Struggle to Do Things Alone: Build a “Solo-Action OS” Instead of Waiting for Someone Else

You may also like

Advertisement