1. 先決定目標——你要做遊戲,還是做研究怎麼贏的實驗室?
研究牌組與操作,不必先做卡圖、動畫、語音與UI。
核心是:
目前狀態
→ 產生合法行動
→ 選擇
→ 套用規則
→ 下一狀態
真正要做的是只靠數字運行卡牌世界的實驗室。
有可靠引擎就重用;沒有就先做最小正確規則。
2. 寫程式前先定義輸入與輸出
輸入:card DB、牌組、對手、先後手、seed、policy、environment、試合數。
輸出:勝敗、turn、matchup勝率、先後手差、worst matchup、brick、key-card drop、錯誤、有效換牌。
模擬器從問題開始。
3. 建card DB——資料不是行為
保存ID、名稱、職業、費用、類型、stats、rarity、set、text、evolution、related cards、tags。
JSON或SQLite即可。
「入場曲:4傷」並沒有自動實作target、減傷、trigger、ordering。
DB收說明書,Rules Engine讓世界照說明書運轉。
4. 把遊戲表示成state
保存turn、active player、HP、PP、hand、deck、board、graveyard、Evolution、Super-Evolution、Crests、amulets、counters、temporary effects。
畫面位置通常不重要,規則意義才重要。
5. Rules Engine = state A → action → state B
出牌、攻擊、進化、Act、target、end turn、draw、PP refresh、win check都是state transition。
Engine不必聰明。
它要在笨bot合法操作時也正確。
6. 卡牌效果=通用規則+例外
建立damage、heal、draw、buff、summon、destroy、banish、Ward、Storm等通用功能。
特殊卡再override。
統計unsupported、partial、runtime gap。
未知時fail closed比亂猜好。
7. 合法行動生成很重要
產生playable cards、targets、attacks、Evolution、Act、end turn。
hold、draw-first、board-slot clearing也可能是正解。
先問:
正確行動有沒有在candidate set裡?
8. seed、replay、hash確保可重現
固定seed並保存engine commit、DB hash、deck hash、policy hash、environment hash、schedule hash、seed hash。
同條件應重現同一局。
否則改善可能只是條件改變。
9. 第一個AI可以很簡單
reference policy:有lethal就結束、快死就防守、照curve、盤面好face、盤面差trade。
目標是穩定baseline。
10. 人類攻略是prior,不是聖旨
利用mulligan、hold、EP/SEP、setup、swing turn、hand cap、board slot、future lethal。
不要硬寫「永遠hold」。
改成hold_value、future_combo、survival、premature_spend等score。
Guide是搜尋地圖。
11. Tree Search、Beam Search、Node Budget
10個候選每層會變10、100、1,000、10,000。
使用depth、beam width、node budget、pruning。
重要setup線不要太早剪掉。
分開immediate、future、survival、setup、lethal、resource value。
12. 隱藏手牌:belief、POMDP、Monte Carlo
policy不能看真實hidden hand。
用公開資訊建可能世界,例如AoE 30%、Ward 20%、無關鍵牌50%。
這是belief。
問題接近POMDP,可以sample多種手牌再反覆rollout。
13. 大量模擬要paired A/B
相同deck、對手、first/second、seed,只換policy。
quick → development → unseen holdout。
目的在公平比較下降低隨機。
14. 不只看勝率
看average、meta-weighted、worst matchup、bottom average、variance、first/second gap、brick、key-card drop、turn、confidence interval。
CVaR看壞結果區域。
Bootstrap與McNemar協助paired A/B。
15. 牌組最佳化是壓縮搜尋空間
不能窮舉所有40張組合。
從真實deck出發,做1~3張swap、role replacement、tech。
Double Oracle增加強counter。
PSRO把deck + policy當策略。
No-Regret降低regret。
Robust Optimization保護matchup floor。
16. 記錄AI為什麼輸
Action Trace保存candidates、scores、choice、objective、reservation、early combo spend、missed lethal、hand burn、board lock。
Ablation一次關一個component。
Counterfactual從同一public state試A與B。
這讓證據更接近因果。
17. 新卡包=新environment snapshot
保存環境、格式、合法卡包與各種hash。
新舊DB做diff,只重算受影響deck/matchup。
只改meta weight就重用舊結果。
不用每個新包重建研究所。
18. 架構小而CLI優先
分離upstream、deck、policy、simulation、diagnosis、statistics,加上config、data、reports、docs、tests。
CLI做status、doctor、db check、validate、simulate、experiment、optimize、diagnose、context。
Git保存manifest與精簡結果。
GitHub是研究所的外部記憶。
19. 從零到一順序
1 card DB
2 一局最小正確對局
3 legal actions/targets
4 seed/replay/hash
5 simple policy + 100 games
6 coverage
7 real decks + matrix
8 human prior
9 search + belief
10 paired A/B + holdout
11 deck optimization
12 diagnosis + expansion diff
不要一開始就百萬局與所有AI理論。
不要做每分鐘高速犯錯一萬次的機器。
20. 結論——模擬器其實是在打造研究流程
DB是記憶,Rules Engine是物理法則,Policy是大腦,Search是預讀,Statistics是成績單,Action Trace是監視器,Git是研究筆記,AI Agents是研究員。
人最後問:
「這個數字真的可信嗎?」
最初只想疾馳打臉。
最後變成卡牌遊戲AI研究所。
最後一手不變:
拆守護,進化,
疾馳打臉。
