1. ก่อนอื่นต้องเลือกก่อนว่าจะทำเกมหรือทำห้องทดลองเพื่อหาวิธีชนะ
ถ้าเป้าหมายคือทดสอบเด็คและ move ไม่ต้องเริ่มจากภาพ animation เสียง หรือ UI
แกนจริงคือ:
state ปัจจุบัน
→ สร้าง legal actions
→ เลือก action
→ ใช้ rules
→ state ถัดไป
สิ่งที่สร้างคือ ห้องทดลองเชิงตัวเลข
มี engine ที่ไว้ใจได้ก็ใช้ต่อ ไม่มีก็ทำ rules ขั้นต่ำที่ถูกต้องก่อน
2. กำหนด input และ output ก่อนเขียนโค้ด
Input: card DB, deck, opponent, first/second, seed, policy, environment, จำนวนเกม
Output: winner, turn, matchup WR, first/second gap, worst matchup, brick, key-card drop, mistakes, deck changes
Simulator เริ่มจาก คำถาม
3. Card DB — data ไม่ใช่ behavior
เก็บ ID, ชื่อ, class, cost, type, stats, rarity, set, text, evolution, related cards, tags
JSON หรือ SQLite ก็พอ
“Fanfare: 4 damage” ยังไม่จัดการ target, reduction, trigger, ordering
DB เก็บคู่มือ Rules Engine ทำให้โลกทำตามคู่มือ
4. เก็บเกมเป็น state
เก็บ turn, active player, HP, PP, hand, deck, board, graveyard, Evolution, Super-Evolution, Crests, amulets, counters, temporary effects
ความหมายของ rules สำคัญกว่าหน้าตา
5. Rules Engine = state A → action → state B
Play, attack, Evolution, Act, target, end turn, draw, PP refresh, win check คือ transitions
Engine ต้อง ถูกต้องแม้ bot โง่เล่น legal
6. Card effects = generic + exception
สร้าง damage, heal, draw, buff, summon, destroy, banish, Ward, Storm
การ์ดพิเศษใช้ override
นับ unsupported, partial, runtime gaps
7. Legal action generation
สร้าง playable cards, targets, attacks, Evolution, Act, end turn และถ้าจำเป็น hold, draw-first, board-slot clearing
ถามก่อน:
action ที่ถูกมีอยู่ใน candidate set ไหม?
8. seed, replay, hash
Fix seed และเก็บ engine commit, DB hash, deck hash, policy hash, environment hash, schedule hash, seed hash
เงื่อนไขเดิมต้องสร้างเกมเดิมได้
9. AI รุ่นแรกเรียบง่ายได้
Reference policy: lethal → finish; จะตาย → defend; curve; board ดี → face; board แย่ → trade
นี่คือ baseline
10. Human knowledge เป็น prior
ใช้ mulligan, hold, EP/SEP, setup, swing turn, hand cap, board slot, future lethal
อย่าทำเป็นกฎตายตัว
ใช้ score เช่น hold_value, future_combo, survival, premature_spend
Guide คือ แผนที่
11. Tree Search, Beam, Node Budget
10 candidate ต่อ layer กลายเป็น 10, 100, 1,000, 10,000
ใช้ depth, beam width, node budget, pruning
อย่าตัด setup สำคัญเร็วเกินไป
แยก immediate, future, survival, setup, lethal, resource value
12. Hidden hand: belief, POMDP, Monte Carlo
อย่าให้ policy เห็น hidden hand จริง
สร้างโลก เช่น AoE 30%, Ward 20%, ไม่มีอะไร 50%
นี่คือ belief
ปัญหาใกล้ POMDP; sample hand แล้ว rollout หลายครั้ง
13. Large simulation = paired A/B
deck, opponent, first/second, seed เดียวกัน เปลี่ยน policy เท่านั้น
Quick → development → unseen holdout
เป้าหมายคือ ลด noise ด้วยการเทียบอย่างยุติธรรม
14. วัดมากกว่า win rate
Average, meta-weighted, worst matchup, bottom average, variance, first/second gap, brick, key-card drop, turn, confidence interval
CVaR ดูด้านแย่
Bootstrap และ McNemar ช่วย paired A/B
15. Deck optimization บีบ candidate space
brute-force 40-card combinations ทั้งหมดไม่ได้
เริ่มจาก deck จริง ลอง swap, role replacement, tech
Double Oracle เพิ่ม counter, PSRO ใช้ deck + policy, No-Regret ลด regret, Robust Optimization ป้องกัน floor
16. บันทึกว่าทำไม AI แพ้
Action trace เก็บ candidate, score, choice, objective, reservation, early combo spend, missed lethal, hand burn, board lock
Ablation ปิด component ทีละตัว
Counterfactual ลอง A และ B จาก public state เดียวกัน
17. Set ใหม่ = environment snapshot ใหม่
เก็บ format, legal sets, hashes
diff DB แล้ว rerun เฉพาะ deck/matchup ที่ได้รับผล
เปลี่ยนแค่ meta weight ก็ reuse result
18. Architecture เล็กและ CLI-first
แยก upstream, deck, policy, simulation, diagnosis, statistics, config, data, reports, docs, tests
CLI: status, doctor, db check, validate, simulate, experiment, optimize, diagnose, context
GitHub คือ external memory
19. Roadmap zero-to-one
1 card DB
2 เกมขั้นต่ำ 1 เกมที่ถูกต้อง
3 legal actions/targets
4 seed/replay/hash
5 simple policy + 100 games
6 coverage
7 real decks + matrix
8 human prior
9 search + belief
10 paired A/B + holdout
11 deck optimization
12 diagnosis + expansion diff
อย่าสร้างเครื่องที่ผิด 10,000 ครั้งต่อนาที
20. สรุป
DB = ความจำ. Rules Engine = ฟิสิกส์. Policy = สมอง. Search = มองอนาคต. Statistics = ใบคะแนน. Action Trace = กล้อง. Git = สมุดวิจัย. AI Agents = นักวิจัย.
คนถาม:
“ตัวเลขนี้เชื่อได้จริงไหม?”
เริ่มจากอยาก Storm ใส่หน้า
จบด้วย AI card-game lab
เอา Ward ออก, Evolution,
Storm ใส่หน้า
