1. Chọn mục tiêu trước: làm game hay làm phòng thí nghiệm để tìm cách thắng?
Nếu mục tiêu là nghiên cứu deck và move, không cần bắt đầu bằng hình, animation, voice hay UI.
Cốt lõi:
state hiện tại
→ sinh legal actions
→ chọn action
→ áp rules
→ state mới
Bạn đang xây một phòng thí nghiệm số.
Có engine tốt thì tái dùng; không có thì viết rules tối thiểu đúng.
2. Định nghĩa input và output trước code
Input: card DB, deck, opponent, first/second, seed, policy, environment, số game.
Output: winner, turn, matchup WR, first/second gap, worst matchup, brick, key-card drop, errors, deck changes.
Simulator bắt đầu bằng câu hỏi.
3. Card DB — data không phải behavior
Lưu ID, tên, class, cost, type, stats, rarity, set, text, evolution, related cards, tags.
JSON hoặc SQLite là đủ.
“Fanfare: 4 damage” chưa xử lý target, reduction, trigger, ordering.
DB giữ hướng dẫn; Rules Engine thực thi.
4. Biểu diễn game bằng state
Lưu turn, active player, HP, PP, hand, deck, board, graveyard, Evolution, Super-Evolution, Crests, amulets, counters, temporary effects.
Ý nghĩa rules quan trọng hơn hình ảnh.
5. Rules Engine = state A → action → state B
Play, attack, Evolution, Act, target, end turn, draw, PP refresh, win check đều là transitions.
Engine phải đúng kể cả khi bot ngốc chơi hợp lệ.
6. Card effects = generic + exceptions
Tạo damage, heal, draw, buff, summon, destroy, banish, Ward, Storm.
Card đặc biệt dùng override.
Đếm unsupported, partial, runtime gaps.
7. Legal action generation
Sinh playable cards, targets, attacks, Evolution, Act, end turn và khi cần hold, draw-first, board-slot clearing.
Hỏi:
nước đúng có nằm trong candidate set không?
8. seed, replay, hash
Fix seed và lưu engine commit, DB hash, deck hash, policy hash, environment hash, schedule hash, seed hash.
Cùng điều kiện phải tái tạo cùng game.
9. AI đầu tiên có thể đơn giản
Reference policy: lethal → finish; sắp chết → defend; curve; board tốt → face; board xấu → trade.
Đây là baseline.
10. Human knowledge là prior
Dùng mulligan, hold, EP/SEP, setup, swing turn, hand cap, board slot, future lethal.
Không biến thành luật tuyệt đối.
Dùng score như hold_value, future_combo, survival, premature_spend.
Guide là bản đồ.
11. Tree Search, Beam, Node Budget
10 candidates mỗi layer thành 10, 100, 1.000, 10.000.
Dùng depth, beam width, node budget, pruning.
Đừng cắt setup quan trọng quá sớm.
Tách immediate, future, survival, setup, lethal, resource value.
12. Hidden hand: belief, POMDP, Monte Carlo
Không cho policy đọc hidden hand thật.
Tạo thế giới AoE 30%, Ward 20%, không có 50%.
Đó là belief.
Bài toán gần POMDP; sample hand và rollout nhiều lần.
13. Large simulation = paired A/B
Cùng deck, opponent, first/second, seed; chỉ đổi policy.
Quick → development → unseen holdout.
Mục tiêu là giảm noise với so sánh công bằng.
14. Đo nhiều hơn win rate
Average, meta-weighted, worst matchup, bottom average, variance, first/second gap, brick, key-card drop, turn, confidence interval.
CVaR nhìn phần xấu.
Bootstrap và McNemar giúp paired A/B.
15. Deck optimization thu nhỏ không gian
Không brute-force mọi 40-card combination.
Bắt đầu từ deck thật, thử swap, role replacement, tech.
Double Oracle thêm counter, PSRO dùng deck + policy, No-Regret giảm regret, Robust Optimization giữ floor.
16. Ghi lại tại sao AI thua
Action trace ghi candidate, score, choice, objective, reservation, early combo spend, missed lethal, hand burn, board lock.
Ablation tắt component.
Counterfactual thử A và B từ cùng public state.
17. Set mới = environment snapshot mới
Lưu format, legal sets, hashes.
Diff DB và rerun chỉ deck/matchup bị ảnh hưởng.
Chỉ meta weight đổi thì reuse result.
18. Kiến trúc nhỏ, CLI-first
Tách upstream, deck, policy, simulation, diagnosis, statistics, config, data, reports, docs, tests.
CLI: status, doctor, db check, validate, simulate, experiment, optimize, diagnose, context.
GitHub là external memory.
19. Roadmap zero-to-one
1 card DB
2 một game tối thiểu đúng
3 legal actions/targets
4 seed/replay/hash
5 simple policy + 100 games
6 coverage
7 real decks + matrix
8 human prior
9 search + belief
10 paired A/B + holdout
11 deck optimization
12 diagnosis + expansion diff
Đừng xây máy sai 10.000 lần mỗi phút.
20. Kết luận
DB = trí nhớ. Rules Engine = vật lý. Policy = não. Search = nhìn trước. Statistics = bảng điểm. Action Trace = camera. Git = sổ nghiên cứu. AI Agents = nghiên cứu viên.
Con người hỏi:
“Con số này thật sự đáng tin không?”
Ban đầu chỉ muốn Storm vào mặt.
Cuối cùng thành phòng lab AI card game.
Gỡ Ward, Evolution,
Storm vào leader.
