1. I only wanted to go face; an AI lab appeared
The original task was simple: use AI to identify the strongest current deck. Set 8 was pinned with an 826-card database, engine commit, environment hash, legal decks and seeds. A historical 12,480-game baseline reached 100% rule coverage with zero unsupported cards and zero known rule gaps, yet the raw bot produced absurd class results such as Sword around 79.8% and Rune around 25.6%.
Ten thousand repetitions do not turn a mistake into truth. They can create a machine that is extremely consistent at being wrong.
2. Human advice made the bot weaker, and averages hid fires
human-prior-v1 encoded resource preservation and setup advice. Exact A/B used 2,016 paired units / 4,032 engine games: baseline 50.99%, human-prior-v1 49.36%, a -1.64pt change. Adaptive v2 improved the overall average by +0.74pt but Runecraft fell -6.25pt and failed its leader floor; T11 also failed with an Abysscraft -16.67pt result.
Human advice is conditional, while fixed weights can kill context. An average can say “pass” while one class is on fire.
3. Root-cause analysis found bugs in the researchers too
The project found a disconnected max_nodes=160 contract and an actor-attribution error that reversed causal interpretations. Those platform defects were fixed, but broad policy performance did not recover.
The Set 8 policy program stopped after 20 experiment records and 31,406 confirmed engine games at PAUSED_COMPLETE_NO_PROMOTION. Of 439 discordant units, 0 were fully causally explained, 11 partially explained and 428 unresolved. human-prior-v1 remained active and the planned 8,064-game final unseen holdout stayed sealed.
4. Fifty-two market decks fought for 46,900 games and created the “brain-dead ranking”
The market corpus contained 52 decks, 25 archetypes and seven leaders. Market analysis used 46,900 actual engine games. A separate direct recommended 7×7 used 1,512 games. Direct averages were Buff Forest 71.99%, Rally Sword 65.97%, Midrange Abyss 54.17%, Artifact Portal 54.17%, Ramp Dragon 53.70%, Evolution Haven 31.71%, and Calgidensula Witch 18.29%.
These are simulator-direct values, not ladder win rates. Decks where “good now” usually stays good later were easier for the current bot; decks that sacrifice now to buy future power looked harder. Thus the deliberately rude nickname: the brain-dead-vs-brain-dead ranking.
5. Human tournaments put Witch in another universe
At Dexel Rising Cup #2 on August 8, 2026, Hand Witch represented 116 of 384 submitted decks, about 30.2%, while the simulator’s Calgidensula Witch finished last in the direct 7×7 at 18.29%. They must not be treated as identical forty-card lists, but the human-versus-AI Witch gap was too large to ignore.
Rally Sword looked strong both for humans and for the bot, suggesting a practical hypothesis: strategies that survive imperfect decisions may be easier recommendations for non-experts.
6. Future Optionality: the bot knew rules, but barely represented order and plans
Seven leaders and 24 mechanics were audited across RULE, STATE, VISIBILITY, IMMEDIATE VALUE, FUTURE VALUE, SEQUENCE, SEARCH and PLAN. Across 27,810 decisions and 118,575 root candidates, only 8,878 candidates, 7.49%, carried explicit future value. Rules/state/immediate value were supported 24/24, while sequence was missing for 23/24 and plan for 20/24.
Witch exposed draw-order effects; Dragon showed that +1 max PP can unlock future legal actions and thresholds such as Awakening. Dragon had 12 MYOPIC_REVERSAL diagnostics: six ramp and six Awakening. These diagnostics were not promoted to hard-coded bonuses.
7. We built a real three-turn closed-loop planner
The engine was forced to advance at least three future turn boundaries. Only one primitive action was executed before observing draws, random effects and opponent actions, then replanning. Hidden hands and future draw order were not exposed.
The Ablation used 56 actual engine games, evaluated 3,979/3,979 root actions, reached the horizon for 100% of roots, logged 81,621 observation replans and 35,821 plan changes, and had zero hidden-info violations. Implementation worked. Performance fell by -25.0pt to -37.5pt versus human-prior-v1. We built foresight and obtained a bot that could see farther while being wrong farther away.
8. Leaf calibration separated where ranking broke
The system captured 3,979 root-action leaves and 1,009 unique leaves, then produced 54,208 terminal rollout rows. Among 73 first ranking divergences, leaf weighting accounted for 59, chance aggregation 12 and opponent backup 2.
That shifted suspicion from “three turns is simply too short” toward how leaf states are valued. Replaying the 12 old Dragon reversal diagnostics with the real three-turn search changed ranking toward ramp in only 1/12, demonstrating again that a diagnostic proxy is not ground truth.
9. Learned value predicted better and selected the best move worse
Fresh holdout v1.1 covered all seven acting leaders with 727 leaves / 104 roots / 22,768 terminal rollouts. Brier improved from 0.1144 to 0.0972 and pairwise accuracy from 75.7% to 78.9%.
Decision quality collapsed: root argmax 59.6%→39.4%, top-2 empirical-best containment 73.1%→64.4%, top-3 82.7%→76.0%, and mean regret 4.42pt→5.29pt. Forest, Haven and Portal failed leader floors. Status: HOLD_OFFLINE.
Prediction quality and decision quality are not the same objective. A prettier probability model that presses the wrong card stays offline.
10. Reverse the question: reason backward from lethal
The next hypothesis starts from the win state. LETHAL ← required damage ← PP ← cards ← Spellboost/Awakening/Rally thresholds ← current action.
Ramp is not valuable because somebody wrote “ramp +8.” It is valuable when a concrete win route requires eight PP on schedule and skipping ramp makes the route impossible. Draw-first is not a class bonus; it matters when route prerequisites require that ordering. Goal/regression planning proposes the route backward; the real engine verifies it forward.
11. Also reason backward from how we lose
The planner also regresses from SELF_LOSS: opponent required damage, board damage, extra hand damage, Ward-removal prerequisites, PP and public information. Each candidate action is checked for which opponent win routes it cuts.
Safety cannot become the absolute objective. Otherwise we build the Eternal Healing Coward: safe, polite, and dead five turns later. RISK_REQUIRED means that when safe lines have almost no chance to win, a risky non-forced-loss line may be preferred because it preserves more actual win probability.
12. “Forced” requires a declared proof scope
Win labels are separated into PROVEN_WIN_NOW, FORCED_WIN_FULLY_OBSERVABLE_SUBTREE, FORCED_WIN_WITHIN_ENUMERATED_RESPONSE_SET, ROBUST_WIN_BELIEF, and CONDITIONAL_WIN. Loss labels mirror them with proof-scoped forced loss, HIGH_RISK_LOSS_BELIEF, and CONDITIONAL_LOSS.
For proof search, self is OR and opponent is AND. A forced claim within an enumerated response set requires beating every valid response in that set; a certainty claim over chance requires covering all reachable outcomes. Hidden-hand truth and future draw order are forbidden, and strategy fusion is forbidden: the root action cannot secretly depend on an unobserved sampled world. Replanning occurs after observation.
13. Action priority becomes staged gates, not mystery arithmetic
Priority is: (1) PROVEN_WIN_NOW; (2) fully observable forced win; (3) forced win within the declared response set; (4) remove avoidable proof-level forced-loss actions; (5) maximize expected final win probability; (6) for near ties compare near-loss risk; (7) robustness; (8) win-route progress/opponent-route cuts; (9) validated continuation value.
ROBUST_WIN_BELIEF and CONDITIONAL_WIN remain probabilistic. Only actions that are worse in both winning prospects and loss risk may be removed as strictly DOMINATED. The real question is not “what is the board score?” but “what is my chance to win if I press this action?”
14. Three turns becomes a minimum, not a wall
Initial research configuration is H_MIN=3 and selective H_MAX=5. Branches extend only when lethal sequences, forced responses, important thresholds, delayed payoffs or near-tied top candidates remain unresolved at the boundary.
Rollouts start at 32 per candidate; ambiguous leaders of the root receive 64→128→256. Roots are classified as UNIQUE_BEST / PRACTICAL_TIE / UNRESOLVED. Candidate epsilon_decision values 0.02/0.03/0.05 are tested on development data only and frozen before a fresh holdout. Best-arm identification inspires the allocation principle, but the SVWB implementation is not claimed to be Track-and-Stop.
15. New QC: catch confident mistakes before chasing the strongest AI
Primary gates demand zero hidden-info violations, future-draw-order violations, rule gaps, strategy-fusion violations, PROVEN_* false positives, FORCED_* scope mislabels, chance-completeness violations, immediate-lethal misses, avoidable proof-level forced-loss selections and split leakage. Then all-seven-leader coverage, UNIQUE_BEST argmax, mean/p90/p95 regret, leader floors and top-2 best-set coverage are checked.
Brier, log loss, pairwise accuracy, Spearman, Kendall and calibration are Secondary and cannot rescue a Primary failure. Holdout v1.1 is permanently consumed and cannot be used for training or tuning.
We gave the bot foresight and it got weaker. We taught it probabilities and prediction improved while top-one choice got worse. So the next bot reasons backward from winning, backward from losing, and calls something forced only when the proof scope actually supports the word.
Let’s build the strongest AI with AI. At the moment the QC may be stronger than the game bot. That is probably healthy.
