pstack
รายงานวิจัยฉบับเต็ม ฉบับภาษาไทย

วิเคราะห์ชุด skills ของ Lauren Tan (@poteto) ลึกถึงระดับไฟล์ต้นฉบับ พร้อมแผนใช้งานจริงบน ZCode — อ่านส่วนสรุปได้ใน 5 นาที เจาะลึกได้ไม่จำกัด

6หัวข้อวิจัย
44skills คู่มือครบทุกตัว
95+แหล่งอ้างอิง
100%validation ผ่านทุกไฟล์
✓ติดตั้งบนเครื่องแล้ว
§

บทสรุปผู้บริหาร — อ่าน 5 นาที

ถ้าอ่านแค่ส่วนเดียว ให้อ่านกล่องนี้

pstack มีประโยชน์จริง — แต่สิ่งที่ควร "ซื้อ" ไม่ใช่ตัวเลข 2,000 PRs/เดือน (self-report ไม่มี audit) หากคือ 3 ชิ้นสถาปัตยกรรม:

  • 🛡️ Verification skill — สอน agent "พิสูจน์" งานบนแอปจริง ไม่ยอมรับแค่ "มัน compile แล้ว"
  • 🗺 Feature map — จดทุก feature เป็นไฟล์เล็กให้ agent เปิดเฉพาะจุด ประหยัด token กว่า wiki
  • 📜 Principles + Playbooks — กติกาที่เรียกดึงทิศกลางงานได้ด้วยชื่อ
1
ติดตั้งแล้ว
pstack@pstack-local บน ZCode นี้ — 44 skills ผ่านตรวจทั้งหมด
2
เริ่มใช้
/setup-pstack ครั้งเดียว → /poteto-mode <โจทย์>
3
ระวัง
token cost จริง ~2x · อย่าเปิด skills จน injection budget ล้น
4
ไม่มีเทียบเท่า
Cloud Agents และ Bugbot ของ Cursor — ส่วนอื่น port ได้หมด
01

กายวิภาคของ pstack

pstack = skills 44 ตัว (port นี้) + 23 playbooks + 21 principles ที่ /poteto-mode คุมทั้งหมด — ติดตั้งข้าม harness ได้เพราะเป็น markdown ล้วน

pstack คือ plugin ชุด Agent Skills แบบ MIT ของ poteto (Lauren Tan, React core team / อดีต Meta, Netflix ปัจจุบันอยู่ Cursor) เป้าหมายคือ 'เขียนโค้ดน้อยลงแต่คุณภาพสูงขึ้น' ไม่ใช่ maximize LOC ประกอบด้วย skill หลัก 27 ตัว + principle skills 24 ตัว + 2 subagents + 1 automation pack โดยมี /poteto-mode เป็น command กลางที่เปิดเป็น 'sticky mode' แล้ว route งานไปยัง playbook ที่เหมาะสมจากทั้งหมด 23 playbooks ติดตั้งได้ทั้งจาก Cursor (/add-plugin pstack) และผ่าน skills CLI ที่รองรับหลาย harness รวมถึง Claude Code / Codex / OpenCode

👤 คุณ (router คนจริง)
พิมพ์ครั้งเดียว:
/poteto-mode <งาน + outcome>
↓
👑 poteto-mode (router)
อ่านโจทย์ → เลือก playbook → เดินตามขั้นแบบ copy verbatim
↓ โหลดกติกา + เลือกแผน + เรียก skills
📜 Principles 21 ตัว
กติกาโหลดทั้งชุด — เรียก redirect กลางงานได้ด้วยชื่อ
📖 Playbooks 23 แบบ
bug-fix · investigation · shipping · autopilot · prototype · …
↓ เรียกใช้ตามที่ playbook กำหนด
🔍 Understand
howwhyteachrecallfigure-it-out
🏛 Design
architectarenaswarmblast-radius
🔨 Build
tddunslopno-commentstechnical-writing
✅ Verify
create-verificationmaintain-verification
↓ ลงมือทำ (subagents ของ ZCode)
🤖 poteto-agent
runner หลัก — งาน code ส่งต่อที่นี่
💀 comment-sicko
agent อ่านอย่างเดียว — ล่า comment ที่ไม่จำเป็น
คุณสั่งจุดเดียว — ที่เหลือ router จัดการเองทั้งชุด · กลางงานดึงทิศกลับได้ด้วยชื่อ principle เช่น "apply laziness-protocol"
ข้อค้นพบสำคัญ
01
โครงสร้างทางการใน repo cursor/plugins/pstack: skills/, agents/, automations/benny/, docs/guide/, assets/, .cursor-plugin/ (manifest) — ยืนยันตรงกับ mirror backnotprop/pstack (~1.1k stars, MIT)
02
รายชื่อ skills เต็ม: poteto-mode, setup-pstack, poteto-help, how, why, recall, blast-radius, architect, arena, swarm, interrogate, automate-me, make-bot-ui, reflect, correct, teach, tdd, benchmark-checklist, no-comments, typescript-best-practices, figure-it-out, show-me-your-work, create-verification-skill, maintain-verification-skill, unslop, bro, technical-writing — บางตัวเป็น 'pairing' เช่น architect ต้องมี arena+how, teach ต้องมี how+why
03
poteto-mode มี 23 playbooks ครบ: investigation, bug-fix, perf-issue, hillclimb, runtime-forensics, trace-forensics, feature, refactoring, prototype, visual-parity, authoring-a-skill, eval, babysit, shipping, autonomous-run, orchestrate, autopilot-full, autopilot-stack, session-pickup, pause-safely, multi-phase-plan, worktree-cleanup, opening-a-pr
04
principle skills 24 ตัวแยกเป็นไฟล์ (principle-laziness-protocol ถึง principle-encode-lessons-in-structure) โดย poteto-mode โหลดทุกตัวเป็นชุดกติกาประจำโหมด
05
มี 2 subagents: poteto-agent (agent ทำงานหลัก) และ comment-sicko (reviewer อ่านอย่างเดียวที่จับ comment ห่วย) และ automation pack 'benny' (bug-triage จาก Slack — ไฟล์ไม่ถูก register เป็น slash skill)
06
การติดตั้ง: Cursor = /add-plugin pstack; หลาย harness = npx skills add backnotprop/pstack (รองรับ Claude Code, Codex, Pi, Cursor, OpenCode); หลังติดตั้งรัน /setup-pstack (เลือก reasoning budget + แผน model) แล้วใช้ /poteto-mode เป็นจุดเข้าหลัก
07
การจัด model: ตั้งค่าเริ่มต้นเป็น panel opus 5.5 / sol / grok — งาน code ส่งให้ grok, งาน judgment ให้ opus (แสดงแนวคิด 'แยก model ตามบทบาท')
08
ของที่ไม่ bundle มาด้วย: /deslop, control-cli, control-ui, /create-skill (อยู่ใน cursor-team-kit หรือมากับ Cursor เอง)
09
ปรัชญาที่น่าสังเกต: ไม่มี dedicated planning skills — README บอก 'the best spec is code' และปรับแต่งต่อได้ด้วย /automate-me (สร้าง mode skill ของตัวเอง) + /setup-pstack (map model เข้าบทบาท)
กลไก · artifacts · quotes แบบเต็ม
⚙️ กลไกการทำงาน (how_it_works)
  • 1) ติดตั้ง plugin → 2) รัน /setup-pstack เพื่อเลือก reasoning budget และ map โมเดลเข้าบทบาท (code vs judgment) → 3) เปิด /poteto-mode เมื่อมีงานจริงจัง
  • /poteto-mode รับ goal + checkable outcome แล้วเลือก playbook ที่ตรง (เช่น bug-fix, perf-issue, autopilot-full) โดยโหลด principle skills ทั้ง 24 ตัวเป็นกติกาพื้นฐานของโหมด — ผู้ใช้เรียกงานด้วยชื่อ principle เพื่อดึง agent กลับเข้าทางกลางงานได้ (เช่น 'laziness-protocol')
  • Playbook แต่ละแบบเป็น workflow จำเป็นขั้น เช่น bug-fix = repro ก่อนแล้วค่อย fix และ verify, autopilot-stack = สร้างคิวงานแบบ autonomy เต็มแล้วส่ง stack ให้คน review ตอนจบ ('You own the stack, never the landing')
  • งานขนานใช้ arena (สร้าง prototype แข่งกัน 2-3 แบบ) และ swarm (ยิง agents หลายตัวเพื่อ sample-size verification / fuzzing) — คู่กับ principle guard-the-context-window ที่บังคับส่งงาน bulk ไป subagent แล้วเก็บแต่ summary
  • ระบบ verification แยกเป็นสอง skill: create-verification-skill (สร้าง skill ตรวจงานเฉพาะ repo ของเรา) และ maintain-verification-skill (รันวันละครั้งเพื่อกัน drift) — รายละเอียดอยู่ใน item verification-skill-and-feature-map
  • agents ที่ใช้ในระบบคือ poteto-agent (ทำงาน) + comment-sicko (อ่านอย่างเดียว เพื่อแยกบทบาทเขียน/ตรวจ)
📎 สิ่งที่จับต้องได้ (concrete_artifacts)
  • Repo ทางการ: https://github.com/cursor/plugins — directory pstack/ (skills 27+24 principle skills, agents/poteto-agent.md, agents/comment-sicko.md, automations/benny/FOR_AGENTS.md, docs/guide/README.md)
  • Mirror repo สำหรับติดตั้งข้าม harness: https://github.com/backnotprop/pstack (~1.1k stars, MIT, MIRROR.md อธิบายที่มา) — ติดตั้งด้วย `npx skills add backnotprop/pstack`
  • Unofficial reader: https://pstack.nishilfaldu.site (guide ที่ /docs, playbooks ที่ /docs/playbooks — เว็บระบุว่าไม่ affiliated)
  • รูปแบบ skill มาตรฐาน: แต่ละ skill เป็น directory มี SKILL.md (เช่น skills/poteto-mode/SKILL.md + skills/poteto-mode/playbooks/*.md)
💬 คำพูดและ claims
  • เป้าหมายของ pstack: 'write less, but higher quality code' (README ของ mirror)
  • จาก autopilot-stack playbook: 'You own the stack, never the landing' — agent มีอิสระเต็มที่จนถึงขั้นเปิด PR แต่การ merge เป็นของคน
  • README ตอบคำถามทำไมไม่มี planning skill: 'the best spec is code'
  • บทความ X ของ poteto (2026-08-31) อ้างว่าชุดนี้ช่วยให้ 'ship 2,000 PRs a month to production with high confidence' (เป็น self-report — ดู risks_limitations)
🎯 นำไปใช้กับ ZCode
⏱ Effort และสิ่งที่ต้องมี

Port ชุด skills markdown อย่างเดียว: ต่ำ (< 1 ชั่วโมง — npx skills add หรือ copy directory) สิ่งที่ต้องมี: พื้นที่ ~/.zcode/skills/ และ ZCode รองรับ SKILL.md (มีอยู่แล้ว) ส่วนการปรับ poteto-mode/arena/swarm ให้ใช้กลไก subagent/workflow ของ ZCode: ประมาณ 1-3 วันทำการ ต้องทดสอบว่า routing + subagent ทำงานได้จริงกับงานของเรา

⚠️ ข้อจำกัดและความเสี่ยง
🔗 แหล่งอ้างอิง (4)
02

Verification skill & Feature map

หัวใจทั้งระบบ: ทำให้ agent พิสูจน์งานบนแอปจริงได้เอง ด้วย skill ที่ commit ลง repo และมีวงจร maintain กันเน่าทุกวัน

หัวใจของ pstack (ชุด skills ส่วนตัวของ Lauren Tan/poteto, MIT license, อยู่ใน github.com/cursor/plugins) คือการทำให้ agent 'พิสูจน์' งานตัวเองบนแอปจริงได้ ผ่าน 2 skills: /create-verification-skill (สัมภาษณ์ repo แล้วสร้าง project-local verification skill ที่ .cursor/skills/verify-<app>/ พร้อม sections Launch, Doctor, Drive, Evidence, Cleanup, Helpers) และ /maintain-verification-skill (loop รักษา feature map ให้ไม่เน่า: source wave ขนานต่อ feature + live pass หนึ่งรอบ + ที่สุดหนึ่ง PR ของการแก้ที่พิสูจน์แล้ว) ส่วน feature map คือ markdown แบบ per-feature ใน references/features/ ที่มีไฟล์ตัวอย่างจริงคือ repo poteto/verification-skill-example (แอปสมมุติ Atlas/Harbor Labs, 34 ไฟล์ feature + README index) โดย README ของ repo ให้เหตุผลที่ไม่ใช้ wiki ชัดเจน: 'A wiki is great for humans. Agents pay for every token they reread.' วงจรการใช้คือ launch → doctor → drive → prove → clean up

วงจรการ verify หนึ่งรอบ — และวงจรดูแลรายวัน
1
launch
เริ่มแอปจริงด้วยคำสั่งที่ skill ระบุ
2
doctor
เช็คก่อนขับว่า instance สด/พร้อม — stale = ไม่ใช่ evidence
3
drive
เดิน user path จริงทุก entry point ที่ diff กระทบ
4
prove
เก็บ evidence ที่ reviewer สันติยอม: side effects ไม่ใช่แค่ pixels
5
clean up
เก็บกวาดเฉพาะสิ่งที่ตัวเองปลุก — evidence ต้องรอด
คู่กัน: /create-verification-skill สร้างครั้งเดียว (ต้องพิสูจน์รันจริง 1 รอบก่อนส่งมอบ) · /maintain-verification-skill กันเน่าทุกวัน
🗂 Codebase
source of truth ที่สมบูรณ์ที่สุด
แต่อ่านทั้งกอง = token ระเบิด
🗺 Feature map
1 ไฟล์ต่อ 1 feature: วิธีถึง · วิธีขับ · Gotchas
"materialized memory"
🎯 Verify จุดเดียว
เปิด 1 ไฟล์ ไม่ใช่ทั้ง wiki
"Agents pay for every token they reread."
⚠ แอปเปลี่ยน → map เน่า = ข้อมูลหลอกที่แย่กว่าไม่มี
🔁 /maintain-verification-skill วันละครั้ง
จบด้วย clean / changed (1 PR) / blocked — ห้ามแก้ product code
ข้อค้นพบสำคัญ
01
ตัวอย่างอ้างอิงหลัก: repo poteto/verification-skill-example — description ว่า 'Example: project-local verification skill + large-app feature map (fictional Atlas / Harbor Labs)' โครงสร้างจริงคือ .cursor/skills/verify-atlas/SKILL.md + control-atlas.mjs (CLI ตั้งใจไม่ใส่ — 'usage only in SKILL.md') + references/features/ มี README.md (index/conventions/sweep order) และไฟล์ per-feature จริง 34 ไฟล์ เช่น sign-in.md, prompt-box.md, toasts.md, multi-surface-journeys.md (lead เดิมบอก ~30 ซึ่งถูกต้องโดยประมาณ นับได้ 34)
02
SKILL.md ของ verify-atlas มี frontmatter: name: verify-atlas, description ระบุการขับแอปผ่าน CDP, และ disable-model-invocation: true; README ระบุวัฏจักร 'launch, doctor, drive, prove, clean up'; คำสั่งของ control-atlas.mjs แบ่ง 6 หมวด: Inspection (info/snapshot/screenshot/components), Navigation, Interaction (send/click/type/press/eval/feature-flag ฯลฯ), Performance (wait-settle ฯลฯ), Streaming (console/network-log), Health & cleanup (doctor/cleanup)
03
Proof bar ใน SKILL.md: 'Do not submit "look, it opens" captures. A proof must exercise the production user path and show the observable result a skeptical reviewer would accept.' และ 'Run doctor first. A video against a stale bundle is not evidence.' และ 'Verify side effects, not just pixels: DOM attributes, clipboard, network/RPC log, file contents, or a reload / switch-away round trip.'
04
ทุกไฟล์ feature ใช้ 4 H2 เดียวกัน: 'Sub-features', 'How to get to it (user POV)', 'Driving it with control-atlas', 'Gotchas' — ตัวอย่าง toasts.md มี gotcha ว่า 'Toasts auto-dismiss. Assert quickly' และ 'Multiple toasts stack; target by name, not by index'; sign-in.md บอกวิธี automation auth ด้วย clone-userdata --force --overwrite แล้ว restart และเตือน selector ที่หลอก (อย่า assert การขาด [data-component="root"])
05
เหตุผลเลือก feature map แทน wiki (จาก README ของ repo) 4 ข้อ: Scoped ('Open one feature file for the change under test, not the whole corpus'), Actionable ('Each file answers the same four questions: what exists, how a user reaches it, how to drive it with the harness, what usually lies'), Sweepable ('features/README.md is a top-to-bottom regression order'), Maintained as code ('Drift gets fixed in the same PR as the UI change, or by a maintain-verification pass')
06
/create-verification-skill (SKILL.md ใน cursor/plugins/pstack) ทำงาน 5 ขั้น: (1) 'Interview the repo, not the user' — วิเคราะห์ Surface/Run/Drive/Observe/Isolate จากโค้ด ถาม user เฉพาะที่สังเกตไม่ได้, (2) Generate skill 6 sections: Launch/Doctor/Drive/Evidence/Cleanup/Helpers, (3) 'Seed the feature map' — features/README.md + per-feature file 'aim for the top 3-5 to start', (4) 'Prove the generated skill before handing it over' — รุนจริง end-to-end หนึ่งรอบ (launch, doctor, drive ONE mapped feature, capture evidence, clean up), (5) Offer /maintain-verification-skill
07
/maintain-verification-skill มี 3 outcome เท่านั้น: clean (ครบและไม่มีอะไรจะ ship — ไม่มี PR), changed (หนึ่ง PR ของการแก้ที่พิสูจน์แล้ว จำกัดแค่ directory ของ verification skill), blocked (บอกว่าอะไรกีดขวาง) — ขั้นตอน: หา skill เป้าหมาย → index hygiene → 'Source wave' (read-only subagent หนึ่งตัวต่อหนึ่ง feature file ยิงขนาน; เด็กห้ามขับแอปและห้ามแก้ไฟล์) → reconcile → live pass (บังคับทำแม้ source ดูสะอาด; 3 invariants: doctor ก่อนขับทุกครั้ง, evidence รอดทุก cleanup, สิ่งที่ drive ปลุกขึ้นมาต้องไม่มีชีวิตยาวกว่า drive นั้น) → triage (doc drift vs harness gap vs product gap — ห้ามแก้ product code) → ship or stop
08
ใน guide บทที่ 6 (docs/guide/06-verify-and-ship.md) เปิดด้วย '"It compiles" is not evidence' และระบุว่า verification คือขั้นที่ช้าที่สุดของ agent work เพราะปกติรอมนุษย์ — ทำให้ agent ทำเองได้ก็เลิกเป็น bottleneck; แนะนำรัน /maintain-verification-skill 'at least once a day, ideally from a scheduled automation' และ 'Treat the verification skill as infrastructure, not a one-off. Commit it'
09
Lauren Tan ประกาศทั้งสอง skill ใน X (2026-07-30) ว่า 'pstack now includes 2 skills i recommend everyone use or copy' อธิบาย feature map ว่า 'a map of all the features in your app and how to get to it and use it from a user's pov. it allows agents to navigate and use the app just like a real user which greatly increases its ability to verify its own work' และเตือนว่า 'this feature map goes out of date very quickly' จึงควรรัน maintain เป็น daily automation ด้วย Cursor cloud agents
10
บทความของ poteto เอง 'The Complete Guide to pstack Pt. 1' (X article, 2026-08-31) ยืนยันว่าเธอ claim ส่ง 2,000 PRs/เดือน, เรียก feature map เป็น 'materialized memory' (codebase คือ memory รูปสมบูรณ์ที่สุด, feature map คือรูปย่อที่ประหยัด token), ใช้ Atlas/verification-skill-example เป็นของจริงในบริษัท (แนบ references เช่น preferences.md), และแนะนำจัด 'oncall rotation' บน verification skill ราวกับ infra จริง
11
แหล่งอื่นยืนยันเชิงภายนอก: Flavio Copes สรุป 5 หน้าที่ของ generated skill (launch/ตรวจสุขภาพ/drive/capture evidence/clean up เฉพาะที่ตัวเองเริ่ม) และว่า 'pstack rejects "the build passed" as complete evidence'; The Neuron บันทึกคำเรียก 'materialized memory' และ 'a cheaper index' ลงแอปที่ยังเป็น source of truth; tenten.dev (สรุปพอดแคสต์กับ Simon Willison) บอกว่า 'The first skill she wrote at Cursor was verification' และทุกแอปที่ Cursor/SpaceXAI มี verification skill ที่ถูก maintain อัตโนมัติ
12
สถาปัตยกรรมเบื้องหลัง: mechanical work ถูกดันลง CLI ที่ skill เรียก (wrap Playwright + Chrome DevTools Protocol) ตาม principle 'Build the Lever' — agent ทุกตัวเรียกเครื่องมือเดียวกัน แทนการเขียน throwaway harness ซ้ำ ๆ; guide สรุปลักษณะ CLI ที่ดี: คำสั่ง composposable, มี --dry-run ของสิ่งทำลาย, rich --help, machine-readable output (JSON)
13
pstack ใช้แนวคิดนี้แบบเป็นระบบ: setup-pstack เสนอสร้าง verification skill ให้ทุกโปรเจกต์ที่ไม่มี (step 'Offer a verification skill (optional)'), principle-prove-it-works บังคับ 'Verify against the real artifact ... not a proxy, self-report, or "it compiles."', และ automation 'benny' ก็มี feature-map.example.md ของตัวเอง (per-feature template: How a user gets there / How the control adapter drives it / Stable selectors / States to exercise) พร้อมกติกา 'Never use generated CSS or StyleX classes, dynamic hashes, child indexes, or brittle DOM position'
กลไก · artifacts · quotes แบบเต็ม
⚙️ กลไกการทำงาน (how_it_works)
  • 1) สร้าง: รัน /create-verification-skill — agent 'interviews the repository' (หา surface ที่ user แตะ, วิธี launch, วิธี drive: existing harness ก่อน ไม่มีก็เลือก browser/CDP สำหรับ web/Electron, PTY/tmux สำหรับ CLI/TUI, plain HTTP สำหรับ service, สิ่งที่ใช้เป็น evidence, และ isolate ได้ไหม) แล้วเขียน .cursor/skills/verify-<app>/ ด้วย 6 sections: Launch (คำสั่งเริ่มจริง + สัญญาณว่าพร้อม), Doctor (read-only check ว่า instance นี้คุ้มที่จะขับไหม), Drive (recipe พร้อม selector/คำสั่งจริงของ repo นั้น), Evidence (มาตรฐาน proof: ต้องเดิน user path จริง, จับ side effects, mocks เฉพาะหลัง production boundary), Cleanup ('Never kill by process name; kill what you started' + evidence ต้องรอดจาก teardown), Helpers (script ต้อง executable และแสดงวิธีเรียกในตัว skill)
  • 2) Seed feature map: สร้าง features/README.md (index + baseline preconditions + driving conventions + proof/skip reporting rules + sweep order) และไฟล์ต่อ feature 3-5 อันแรก — ทุกไฟล์ใช้ 4 H2 ตายตัว: Sub-features / How to get to it (user POV) / Driving it with <harness> / Gotchas เขียนระดับ behavior ไม่ใช่ implementation ('Name only user paths, stable handles, required state, commands, and observable proof')
  • 3) พิสูจน์ generator ก่อนส่งมอบ: รุน skill ที่เพิ่งสร้าง end-to-end หนึ่งรอบ (launch → doctor → drive ONE feature → capture evidence → clean up) แล้วเช็คว่า evidence ยังอยู่ที่ที่ประกาศไว้หลัง cleanup — ถ้า proof พังก็อย่าใช้ output; 'A generated skill that was never executed is a draft, not a deliverable.'
  • 4) ใช้งานรายวัน (วงจร launch → doctor → drive → prove → clean up): agent จับคู่ diff กับ feature file → รัน doctor ให้ instance สด → drive ทุก entry point ที่ถึงได้ และทุก path success/cancel/error/empty/persistence ที่ diff กระทบ → เก็บ evidence (screenshot, DOM assertion, clipboard, network, reload round-trip) → cleanup; ถ้าต้อง sweep กว้าง ให้เดิน features/README.md บนลงล่างแล้วปิดท้ายด้วย multi-surface-journeys.md
  • 5) รักษาให้ตรงกับแอป: รัน /maintain-verification-skill (อย่างน้อยวันละครั้ง ผ่าน scheduled automation) — แก้ index → source wave (subagent read-only ขนานต่อ feature สรุปว่า feature ทำงานอย่างไรจาก source + ธง doc drift + recipe หนึ่งอัน) → reconcile recipe → live pass ขับทุก feature จริง (coordinator เป็นคนขับคนเดียว, serial สำหรับ server/UI หรือ isolated session ต่อ drive สำหรับ CLI สั้น) → triage 3 ทาง (doc drift = แก้ map, harness gap = แก้ harness, product bug = รายงาน ห้ามแปะใน docs) → จบด้วย clean/changed/blocked และอย่างมากหนึ่ง PR ของการแก้ที่พิสูจน์แล้วใน directory ของ skill เอง
  • 6) ขยายผล: /swarm แบ่ง full pass ตาม feature-map entry และ aggregate ผล; verification skill เป็นฐานของ PR loop — open PR ด้วย evidence ใน description, Babysit playbook ดันถึง merge-ready แต่ไม่ merge เอง, Shipping playbook ให้ agent หนึ่งคนต่อ PR พิสูจน์ live โดย 'the agent that judges a change is never the one that wrote it' ก่อน land ทีละ PR จากล่างขึ้น
📎 สิ่งที่จับต้องได้ (concrete_artifacts)
  • Repo ตัวอย่าง: https://github.com/poteto/verification-skill-example — โครงสร้าง: .cursor/skills/verify-atlas/SKILL.md (5,461 bytes), references/features/README.md (6,003 bytes) + 34 ไฟล์ feature (command-palette, extensions-catalog, file-browser, follow-ups, hosted-runtimes, left-rail, live-preview, live-replies, log-pane, multi-surface-journeys, new-session, odds-and-ends, overlays-and-links, preferences, project-brief, prompt-banners, prompt-box, remote-machines, review-requests, runbook, scheduled-jobs, session-menu, session-rows, shared-sessions, shell-pane, sign-in, split-view, toasts, transcript, unattended-runs, web-pane, working-tree, workspace-chrome, workspace-groups); control-atlas.mjs ถูกละไว้ตั้งใจ
  • Skill texts (MIT, Copyright (c) 2026 Lauren Tan): https://github.com/cursor/plugins/blob/main/pstack/skills/create-verification-skill/SKILL.md (5,879 bytes, 5 ขั้น: interview → generate → seed feature map → prove → offer maintenance) พร้อม references/feature-map-example/ (README.md + create-note.md + search.md — แอปสมมุติ Notes ที่ port 4173, ใช้ control-notes harness); https://github.com/cursor/plugins/blob/main/pstack/skills/maintain-verification-skill/SKILL.md (4,922 bytes, 3 outcomes + 6 ขั้น pass)
  • Guide: pstack/docs/guide/06-verify-and-ship.md ใน cursor/plugins ('It compiles' is not evidence / finish condition / benchmark-checklist 7 คำถาม / create-verification-skill / maintain รายวัน / open PR / Babysit / Shipping); เว็บ mirror: https://pstack.nishilfaldu.site (สารบัญบอกว่า 'Prove behavior on the real app, then open a focused PR and drive it to merged.')
  • คำสั่งจริงจาก verify-atlas SKILL.md: node .cursor/skills/verify-atlas/control-atlas.mjs checkout-setup --watch / clone-userdata --force --overwrite / --checkout restart / --checkout info / doctor / cleanup --remove-checkout; CDP port เริ่มต้น 9222; test -f .git && echo checkout ใช้แยก isolated checkout (และคำสั่ง connect/launch จาก checkout ที่ไม่มี --checkout จะถูกปฏิเสธ)
  • ประกาศต้นทาง: ทวีต @poteto 2082874054483255805 (2026-07-30 17:00 UTC) และบทความ 'The Complete Guide to pstack Pt. 1' (X article 2094457600259842065, 2026-08-31) — ข้อความเต็มถูกดึงผ่าน reader และ quote ไว้ใน quotes_and_claims
  • setup-pstack/SKILL.md step 7 'Offer a verification skill (optional)': ถ้าโปรเจกต์ไม่มี verify-* skill หรือ harness จะเสนอหนึ่งครั้งด้วยคำต่ำต้องตายตัว 'want a project-local verification skill, so agents can drive the app the way a user does and prove changes work?' — ข้อความชุดนี้ถูก mirror บน RoutineHub [uncertain: หน้า RoutineHub เองไม่ได้เปิดตรวจโดยตรง ยืนยันจาก snippet ค้นหาว่าเนื้อหาตรงกันและเป็นสำเนา MIT ที่ credit Lauren Tan]
  • benny automation ใน pstack มี feature-map.example.md อีกแบบ: template ต่อ feature (How a user gets there / How the control adapter drives it / Stable selectors / States to exercise — Default, hover, focus-visible, active, disabled, Loading, empty, error ...) และกติกา selector ห้ามใช้ generated CSS/StyleX classes, dynamic hashes, child indexes
💬 คำพูดและ claims
  • poteto/verification-skill-example README: 'A wiki is great for humans. Agents pay for every token they reread. A skill + feature map is: Scoped... Actionable... Sweepable... Maintained as code.' และ 'Nothing here is a real product. "Atlas" / "Harbor Labs" / control-atlas are made up. The shape is what matters.'
  • verify-atlas SKILL.md (Proof bar): 'Do not submit "look, it opens" captures. A proof must exercise the production user path and show the observable result a skeptical reviewer would accept.' / 'Run doctor first. A video against a stale bundle is not evidence.' / 'Verify side effects, not just pixels'
  • create-verification-skill SKILL.md: 'You write the generator's output for the next agent, not for a human: it will be read cold, mid-task, by an agent that has never seen the app.' / 'refusing to double-drive a shared instance beats corrupting the user's session' / 'Never kill by process name; kill what you started.' / 'a cleanup that eats the proof fails this step' / 'A generated skill that was never executed is a draft, not a deliverable.'
  • maintain-verification-skill SKILL.md: 'A feature map rots the moment the app changes.' / 'The unit of rigor is the feature, not every sentence' / 'Never edit product code during a run: a behavior the map describes that the app no longer does is either doc drift (fix the map) or a product regression (report it, don't paper over it in docs).' / 'one read-only subagent per feature file, launched concurrently'
  • Guide บท 06: '"It compiles" is not evidence.' / 'Verification is the slowest step in most agent work, because it's the step that usually waits on a human. Make the agent able to do it, and you stop being the bottleneck.' / 'Treat the verification skill as infrastructure, not a one-off. Commit it, so every person and every agent on the team drives the app the same way.' / 'Apps change and feature maps rot. Run this at least once a day, ideally from a scheduled automation so nobody has to remember.'
  • Lauren Tan ทวีต (2026-07-30): 'i recommend everyone use or copy: /create-verification-skill ... /maintain-verification-skill' ... 'it also creates something i call the feature map: a map of all the features in your app and how to get to it and use it from a user's pov. it allows agents to navigate and use the app just like a real user which greatly increases its ability to verify its own work.' ... 'this feature map goes out of date very quickly. run /maintain-verification-skill as a daily automation with cursor cloud agents.' ... 'having a strong verification skill will end up becoming critical infra for your team'
  • Lauren Tan, 'The Complete Guide to pstack Pt. 1' (2026-08-31): บอกว่า pstack ทำให้เธอ 'ship 2,000 PRs a month to production with high [confidence]' (claim ของเธอเอง), เรียก verification skill ว่า critical infrastructure ไม่ใช่แค่ skill อีกอัน ('Done well, you will 100-1000x your whole team's output'), ให้ feature map ชื่อ 'materialized memory' และย้ำว่า codebase เองคือ memory ที่สมบูรณ์ที่สุด ส่วน feature map คือรูปย่อเพื่อประหยัด token; ตัวอย่างบทที่แนบคือ references/features/preferences.md ซึ่งตรงกับไฟล์จริงใน verification-skill-example
  • Lauren Tan (ผ่าน tenten.dev สรุปพอดแคสต์กับ Simon Willison): เล่าที่มาช่วงต้นเมษายน 2026 ตอนแก้ lag ของ agents window ที่ Cursor เอง ว่าตัวเองเป็น 'the meat proxy between my agent and Chrome DevTools'; 'The first skill she wrote at Cursor was verification'; ทุกแอปที่ Cursor และ SpaceXAI มี verification skill ที่ถูก maintain อัตโนมัติ; และยอมรับว่าถึงจุดนี้ได้ยาก ('very hard to get to this point') ไม่ควรขายว่าติดตั้งแล้วได้ผลทันที
  • Flavio Copes (deep dive): 'pstack rejects "the build passed" as complete evidence' / feature map: 'Each feature records how a user reaches it, how an agent drives it, and what observable state proves it works.' / '"verify it" becomes a repository capability instead of a new conversation every time'
  • The Neuron explainer: feature map คือ 'a set of compact references describing each user-facing feature, how to reach it, how to control it, and any traps an agent should know about' เป็น 'a cheaper index' และแอปยัง 'remain the source of truth'; แนะนำรัน maintain 'at least daily so the agent's model of the product does not slowly diverge from the product itself.'
🎯 นำไปใช้กับ ZCode
  • Port ตรง (แทบไม่ต้องแก้): (a) โครงสร้าง skill — เปลี่ยนจาก .cursor/skills/verify-<app>/ เป็น .zcode/skills/verify-<app>/ โดยคง frontmatter name/description และความหมายของ disable-model-invocation: true (skill ประเภทนี้ให้เรียกเองเมื่อต้อง verify ไม่ให้ model ชวนเรียกเองหลุดบริบท — ใน ZCode ใช้วินัยเดียวกันได้ผ่านการอ้างชื่อ skill ใน prompt เช่น /verify-zcode-app); (b) เอกสาร 6 sections (Launch/Doctor/Drive/Evidence/Cleanup/Helpers) และ 4 H2 ต่อ feature copy ได้ทั้งชุดเพราะเป็น MIT license; (c) feature map แบบ per-feature markdown + README index + sweep order ใช้ได้กับทุก project ในเครื่องนี้ทันที
  • แทน harness: pstack ใช้ control-<app>.mjs ที่ wrap CDP/Playwright — ใน ZCode มี Playwright MCP อยู่แล้ว (browser_navigate/click/snapshot/evaluate/console_messages/network_requests) จึงใช้เป็น 'Driving it with Playwright MCP' แทน 'Driving it with control-atlas' ได้ทันที โดยเขียน feature file ให้อ้างชื่อเครื่องมือ/selector จริง (ARIA role, accessible name, data-component) ตาม gotchas ที่หยิบมาจากตัวอย่าง; กรณี workspace eoffice ในเครื่องนี้ ชุด eo_screen/eo_act/eo_verify ก็คือ control harness ตัวจริงอยู่แล้ว — verification skill จะเป็นตัวเขียนกติกา proof/skip ทับเครื่องมือเหล่านี้
  • ต้องดัดแปลง: (a) /create-verification-skill ปกติรันเป็นครั้งเดียวจบ — ใน ZCode ทำเป็น workflow ที่ใช้ subagents ขนาน (source wave ต่อ feature = subagent read-only หนึ่งตัวต่อไฟล์ ส่งกลับ summary/drift/recipe แล้ว coordinator reconcile + ขับ live ด้วย Playwright ตามแบบ maintain-verification-skill); (b) daily automation ที่ poteto รันบน Cursor cloud agents — ใน ZCode ใช้ hook/loop (เช่น loop-me skill หรือ scheduled bash) เรียก /maintain-verification-skill แทน; (c) ความคิด 'commit skill เป็น infrastructure' ใช้ได้เลยแต่ต้องเลือก repo ที่ agent แก้ได้จริง
  • ใช้ไม่ได้/ไม่จำเป็น: ส่วนที่ผูกกับ Cursor โดยตรง — CDP port 9222 ของ Harbor Labs desktop สมมุติ, checkout-isolated instance ผ่าน --checkout (ZCode บน macOS นี้ทำด้วย git worktree + browser instance แยกได้ถ้าจำเป็น แต่แอปส่วนใหญ่ที่ว่าถึงเป็น web ใช้ Playwright context ใหม่ก็พอ), และ cloud-agent pricing/token ที่เป็นเหตุผลจากฝั่ง Cursor — หลักคิดยังใช้ได้: เปิดเฉพาะไฟล์ feature ที่เกี่ยว เพื่อไม่อ่าน corpus ทั้งกองเข้า context ของ ZCode
  • ข้อดีเชิง practice สำหรับ ZCode โดยเฉพาะ: 'sweepable' README index ตรงกับความต้องการ regression sweep แบบ agent ทำเอง; 'Gotchas' section คือ cache ของบทเรียนที่ agent ทำพลาดซ้ำ (เช่น wait-settle แทน fixed sleep, assert ตาม name ไม่ใช่ index) ซึ่งตรงกับแนว encode-lessons-in-structure ที่ใช้อยู่; และ outcome 3 แบบของ maintain (clean/changed/blocked) ให้ contract ชัดว่าเมื่อไรควรมี PR ที่ agent สร้างเอง
⏱ Effort และสิ่งที่ต้องมี

ประมาณการ: สร้าง verification skill ให้ 1 project ขนาดกลาง ≈ หนึ่ง session (create: สัมภาษณ์ repo + เขียน SKILL.md + seed 3-5 feature files + prove หนึ่งรอบ); ขยาย feature map เต็มแบบ Atlas (30+ ไฟล์) เป็นงานสะสมผ่าน /maintain-verification-skill ที่ค่อย ๆ เพิ่ม ไม่ควรพยายามลงทุนเขียนครั้งเดียวครบ — poteto บอก 'aim for the top 3-5 to start' สิ่งที่ต้องมีก่อน: แอปที่รันได้จริงในเครื่อง (มี dev command), Playwright MCP ติดตั้งพร้อม login ได้, สิทธิ์เขียน .zcode/skills/ ใน project, และถ้าจะ maintain อัตโนมัติต้องมีตัวรันงานตามเวลา (cron/loop)

⚠️ ข้อจำกัดและความเสี่ยง
  • ต้นทุนของ feature map คือการเน่า: 'A feature map rots the moment the app changes' — ถ้าไม่มี maintain loop (อย่างน้อยวันละครั้ง) map กลายเป็นข้อมูลหลอกที่ทำให้ agent verify ผิดทางและมั่นใจผิด ๆ ซึ่งแย่กว่าไม่มี map; ต้นทุน token/เวลาของ live pass ทุกวันไม่ใช่เรื่องเล็ก (poteto เองบอก full autopilot 'quite token intensive')
  • คุณภาพ proof ขึ้นกับวินัยของ map: ถ้า feature file อธิบาย implementation แทน behavior, หรือไม่ list ครบทุก entry point, proof ที่ 'ผ่าน' จะเป็น false confidence — ทั้ง create และ maintain skill มีกติกากันเรื่องนี้ (mocks เฉพาะ production boundary, รายงาน verified-unreachable พร้อม prerequisite จริง) แต่ต้องบังคับใช้
  • ข้อจำกัดเทคนิคจากตัวอย่างเอง: OS file drops, native context menus, system-browser auth ยังเป็น 'manual unless the driver has first-class support'; streaming ไม่ deterministic ต้อง wait-settle; ถ้าแอปต้อง auth จริงกับบริการภายนอก ต้องมี clone-userdata-style seeding ไม่งั้น automation ติดตั้งแต่ต้น
  • ความน่าเชื่อถือของแหล่ง: แหล่งหลัก (verification-skill-example, cursor/plugins/pstack — SKILL.md 2 ไฟล์, guide บท 6, README, benny examples) ถูกอ่านจาก GitHub โดยตรงในงานนี้ = แม่นระดับตัวอักษร; ทวีตและบทความ X ของ poteto เป็นคำพูดเจ้าของเอง (primary แต่เป็น self-report — ตัวเลข 2,000 PRs/เดือน และ 100-1000x เป็น claim ที่ตรวจสอบอิสระไม่ได้); Flavio Copes/The Neuron/tenten.dev เป็น secondary ที่สอดคล้องกันเกี่ยวกับโครงสร้าง (5-6 sections, feature map, materialized memory, daily maintain) — จุดที่เหลื่อมกันเล็กน้อยคือจำนวน sections (guide บอก 5: Launch/Doctor/Drive/Evidence/Cleanup, SKILL.md จริงมี 6 รวม Helpers) และวันที่ที่มา (เมษายน 2026 จาก interview เดียว)
  • ประเด็นการตลาด/ความยากจริง: ทั้ง poteto เอง (ผ่าน tenten.dev) ยอมรับว่าระบบนี้ 'very hard to get to this point' และ skills ควรหดสั้นลงตามความเก่งของ model — อย่าคาดหวัง copy repo มาแล้วได้ผลเท่าที่เห็นในบล็อกทันที
❓ จุดที่ยังไม่ยืนยัน
  • concrete_artifacts
🔗 แหล่งอ้างอิง (17)
03

ผู้เขียนและบทความ

Lauren Tan เล่าว่า verification ทำให้ทีม ship ได้ ~2,000 PRs/เดือน — ตัวเลขเป็น self-report แต่ operating model คือของจริงที่เลียนแบบได้

บทความ 'The Complete Guide to pstack Pt. 1' (2026-08-31) โดย lauren (@poteto) เล่าว่า pstack คือชุด skills ส่วนตัวสำหรับงานวิศวกรรมที่เข้มงวด ให้เธอ ship ~2,000 PRs/เดือนอย่างมั่นใจ ในบทบาท 'gardener and maintainer' ของ Grok @Bot (หลังผ่าน Ember.js, Netflix, Meta/React และ Cursor) แกนบทความ 4 ส่วน: verification คือ infrastructure จริง ๆ, วิธีสร้าง verification skill + Feature Map ('materialized memory'), วิธีใช้งาน (poteto-mode, cloud agents แทน worktrees, swarm), และการลงทุนดูแล (maintain ทุกวัน + oncall rotation) มีภาคต่อ Pt. 2 ที่รายงานตัวเลข ส.ค. = 2,462 PRs และย้ำ parallel agents ผ่าน prototyping playbook

เม.ย. 2026
จุดเริ่มที่ Cursor
เขียน verification skill ตัวแรกแก้ lag ของ agents window
30 ก.ค. 2026
ประกาศ 2 skills
"i recommend everyone use or copy" — create/maintain-verification-skill
31 ส.ค. 2026
Pt. 1 ตีพิมพ์
The Complete Guide to pstack — อ้าง 2,000 PRs/เดือน
10 ก.ย. 2026
Explainer
The Neuron วิเคราะห์ + อ้าง Pt.2 (2,462 PRs ส.ค.)
6 ต.ค. 2026
บนเครื่องเรา
ติดตั้ง pstack@pstack-local สำเร็จ — 44 skills
ข้อค้นพบสำคัญ
01
Intro claim (verbatim จากหลาย mirror): 'my personal set of skills for doing rigorous engineering work. It's allowed me to ship 2,000 PRs a month to production with high confidence'
02
บทบาทผู้เขียนในบทความ: 'Grok @Bot's gardener and maintainer' — รีแฟกเตอร์ codebase ระหว่างที่มันถูก build ไปพร้อมกัน 'with no downtime' และทีม land PRs วันละหลายร้อยตัวบน Grok @Bot
03
โครงสร้างจริงของบทความ 4 ส่วน: (1) 'Verification is all you need' — verification skill คือ critical infrastructure ไม่ใช่ 'แค่ skill' และอ้างว่า '100-1000x your whole team's output'; (2) 'Let's build a verification skill together' — /create-verification-skill, Dr Eggbot, principle 'Build the Lever' (สร้าง CLI แทน markdown สำหรับ agent), agent-friendly CLI design (composability, --dry-run, subcommands, JSON output), เลือก stack ที่ debug ง่าย (CDP, iOS simulator), Feature Map = 'materialized memory'; (3) 'How to use your verification skill' — /poteto-mode pin เป็น Custom Mode ด้วย Opt+Enter, ใช้ cloud agents แทน git worktrees (แนะนำตรงข้ามกับ worktrees), /control-app, /swarm สำหรับ parallel verification/fuzzing, ระบบ auto-reproduce user reports จาก Slack; (4) 'Invest in your verification skill' — /maintain-verification-skill วันละครั้ง และพิจารณาทำ oncall rotation ดูแลมัน
04
ตัวเลขในบทความ: 2,000 PRs/เดือน; PRs วันละหลายร้อยบน Grok @Bot; 100-1000x output; ขนานได้ราว 10 parallel agents (แทนที่จะใช้ worktrees); ปิดท้ายสัญญาว่าภาคหน้าจะคุยเรื่อง 'hundreds of subagents'
05
ลิงก์ในบทความ: pstack plugin (x.ai/bot/plugin/9717366), Dr Eggbot bot, เอกสาร Chrome DevTools Protocol, Cursor cloud agent docs, repo ตัวอย่าง github.com/poteto/verification-skill-example และ SKILL.md ของ create-verification-skill, build-the-lever, swarm
06
Pt. 2 มีจริง (thread 2097732320606507506): quote 'With pstack, we can instead take the measure twice, cut once approach to its limit, using parallel agents. We do this by using the prototyping playbook' และตามรายงานของ The Neuron ภาค 2 แสดงตัวเลข ส.ค. 2026 ปิดที่ '2,462 PRs in production'
07
ประเด็นเสริมที่ The Neuron สรุปเพิ่มจากทั้งซีรีส์: 'I don't believe in planning' แต่ 'Prototyping is planning, but with code'; slogan หน้า marketplace 'if you want to go fast, go deep first'; ทีมใช้ skills ส่วนตัวของเธอราว 10,000 ครั้งในหนึ่งสัปดาห์; คำนิยาม verification = agent ทำงาน โต้ตอบกับ product จริง ตรวจผล จับ failure แล้ว retry จนพิสูจน์ความสำเร็จได้
08
ส่วน guide เต็มใน repo (docs/guide) มี 10 บท กว้างกว่าบทความ (Setup, /poteto-mode, Understand code, Design, Build & clean, Verify & ship, Overnight runs, Principles, Customize, Recipes & pitfalls) — บทความ X เป็นเวอร์ชันเล่าแบบ narrative 4 ส่วน
09
โจทย์ที่บทความตอบตรงกับผู้ใช้ ZCode ที่สุด: 'การ verify บน app จริงต้องเป็น infrastructure ที่ agent เรียกใช้เองได้' — ไม่ใช่การไว้ใจว่าโค้ด compile แล้ว = ถูก
กลไก · artifacts · quotes แบบเต็ม
⚙️ กลไกการทำงาน (how_it_works)
  • เส้นเรื่องของบทความ: ปัญหาของ agent-first dev คือความเชื่อมั่น (trust) → ทางแก้คือ verification loop ที่ปิดเองได้ → สร้างเป็น skill ใน repo → ใช้ CLI ควบคุม app จริง (Build the Lever) → จดทุก feature ลง Feature Map ที่ scope ต่อไฟล์ → ขนานงานด้วย cloud agents + swarm → ดูแลต่อเนื่องด้วย maintain รายวัน
  • การใช้งานประจำวันตามบทความ: เปิด /poteto-mode ค้างไว้เป็นโหมดหลัก ปล่อยงานให้ agent หลายตัวขนาน แล้วคนทำหน้าที่ review landing เท่านั้น ('You own the stack, never the landing' มาจาก playbook ที่เกี่ยวข้อง)
  • ความสัมพันธ์กับ ecosystem: pstack ถูก publish ทั้งใน marketplace ของ Cursor (cursor/plugins) และมีลิงก์ plugin ฝั่ง xAI (x.ai/bot/plugin/9717366) — สะท้อนว่าผู้เขียนย้ายจาก Cursor มา xAI และพาชุด skills ไปใช้ข้าม harness
📎 สิ่งที่จับต้องได้ (concrete_artifacts)
  • บทความ Pt. 1: https://x.com/poteto/article/2094457600259842065 (2026-08-31, author แสดงชื่อ 'lauren') — เข้าถึงตรงต้อง login X; อ่านผ่าน mirror ได้
  • Pt. 2 (thread): https://x.com/poteto/status/2097732320606507506 — ใช้ threadnavigator.com/thread/2097732320606507506 อ่านได้
  • Mirror บทความ: https://www.grokbot.sh/blog/the-complete-guide-to-pstack-pt-1 (หน้าเว็บแสดง attribution 'Roland.W' — ควรยืนยันเนื้อหากับต้นฉบับตอนใช้ quote จริงจัง)
  • รายงานวิเคราะห์: The Neuron explainer (2026-09-10) https://www.theneuron.ai/explainer-articles/pstack-explained-lauren-tans-system-for-trustworthy-ai-agents/
  • ตัวเลข milestone ที่อ้างได้จากซีรีส์: Pt.1 = 2,000 PRs/เดือน; Pt.2 = 2,462 PRs (ส.ค. 2026, อ้างผ่าน The Neuron)
💬 คำพูดและ claims
  • 'It's allowed me to ship 2,000 PRs a month to production with high confidence' (Pt. 1 intro)
  • verification skill ทำให้ได้ '100-1000x your whole team's output' (Pt. 1 — claim เชิง rhetorical ที่ต้องอ่านแบบมีวิจารณญาณ)
  • 'Grok @Bot's gardener and maintainer' (บทบาทตนเองในบทความ)
  • 'With pstack, we can instead take the measure twice, cut once approach to its limit, using parallel agents. We do this by using the prototyping playbook' (Pt. 2)
  • 'I don't believe in planning' / 'Prototyping is planning, but with code' (ซีรีส์ pstack, ผ่านการ quote ของ The Neuron)
  • 'if you want to go fast, go deep first' (slogan หน้า marketplace ของ pstack)
  • บทความเสนอว่า verification skill อาจคุ้มที่จะมี 'oncall rotation' ดูแล — สะท้อนว่าถือเป็น infrastructure ที่พังได้ ไม่ใช่เอกสารตายตัว
🎯 นำไปใช้กับ ZCode
  • บทความให้ blueprint ที่เอามาใช้กับ ZCode ได้โดยไม่ต้องติดตั้ง pstack เลย 3 อย่าง: (1) เขียน verification skill สำหรับ project เราเอง (มีตัวอย่างโครงสร้างใน poteto/verification-skill-example), (2) เปลี่ยนงาน manual ซ้ำ ๆ เป็น agent-friendly CLI (Build the Lever) — ฝั่ง ZCode ใช้ Playwright MCP แทน control-ui/control-app ได้, (3) จด feature ลง Feature Map แบบ per-file markdown ที่ agent เปิดเฉพาะไฟล์ที่เกี่ยว
  • เลนส์ที่ต้องแปล: 'cloud agents แทน worktrees' ของเธอ = ของ ZCode คือ subagents/background agents/dynamic workflows บนเครื่องเรา — การขนานระดับ 10 agents พร้อมกันจะโดนจำกัดด้วย concurrency quota และ token budget ส่วนตัว (ตัวอย่างจริงจากงานวิจัยนี้: ยิง 6 agents พร้อมกันแล้วโดนตัดเหลือ 2)
  • 'I don't believe in planning / prototyping is planning' — เอามาใช้กับ ZCode ได้ตรง: ใช้ skill brainstorming/prototype แล้วปล่อย agent ลอง 2-3 ทาง (arena) เทียบกันจริงดีกว่าเขียน spec ยาว
  • ตัวเลข 2,000 PRs/เดือน ใช้เป็น benchmark ทางอารมณ์ไม่ได้กับ dev คนเดียว — บทเรียนที่ transfer ได้คือสัดส่วนงาน: เธอใช้เวลากับ 'สร้างและดูแล verification infrastructure' มากกว่าเขียน feature เอง
  • อ่านต่อแบบมีหลักฐาน: ตัว guide เต็มใน repo (docs/guide) อ่านได้จาก GitHub โดยไม่ต้อง login X — ใช้แทนบทความสำหรับ implement ได้เลย
⏱ Effort และสิ่งที่ต้องมี

อ่านบทความ/guide ฟรี (mirror + repo) ไม่มีต้นทุนติดตั้งสำหรับ item นี้ สิ่งที่ต้องมีเพื่อลงมือตามบทความ: project ที่ drive ด้วย CLI/CDP ได้ (เว็บ = Playwright MCP ของ ZCode), เวลาสร้าง verification skill ครั้งแรก (ราว ครึ่ง-หนึ่งวัน ตาม item verification-skill-and-feature-map) และงบ token สำหรับการ verify ของจริง (มีรายงานว่าแพงขึ้นจริง)

⚠️ ข้อจำกัดและความเสี่ยง
  • ตัวเลขหลักทั้งหมด (2,000 PRs/เดือน, 2,462 PRs, 100-1000x, 10,000 ครั้ง/สัปดาห์) เป็น self-report จากผู้เขียนเอง — ไม่มีการตรวจสอบอิสระ และความหมายของ 'PR' (ขนาด/ประเภทงาน) ไม่ถูกนิยาม
  • แหล่งขัดแย้งกันเรื่อย: จำนวน skills/principles/playbooks ต่างกันในแต่ละแหล่ง (บล็อกบอก 23/21/22, repo จริงนับได้ 27+24 principles/23 playbooks) — ใช้ repo ทางการเป็น ground truth เสมอ
  • บทบาทองค์กรมีความคลุมเครือ: แหล่งรุ่นเก่าระบุ 'engineer at Cursor', บทความระบุบทบาทที่ xAI (Grok @Bot), README mirror บอก 'ex-Meta, Netflix, Cursor' — ตอนอ้างอิงควรเขียนว่า 'อดีต Cursor ปัจจุบันทำ Grok @Bot' ตามบทความเอง
  • mirror grokbot.sh แสดง attribution 'Roland.W' ที่ไม่ตรงตัวผู้เขียน — น่าจะเป็นเว็บ scrape ไม่เป็นทางการ ใช้ quote สำคัญควร cross-check กับ article บน X ตรง ๆ (ต้อง login)
  • บทความโฆษณา playbook ที่ยังไม่มา ('hundreds of subagents' ในภาคหน้า) — เป็นซีรีส์ต่อเนื่อง อย่าถือว่าระบบที่อ่านคือฉบับสมบูรณ์
  • context ของเธอคือทีมระดับ product ใหญ่ (Grok @Bot) + budget ค่า compute แบบองค์กร — หลายคำแนะนำ (ขนาน 10 agents, oncall rotation) ต้องปรับขนาดก่อนใช้กับงานส่วนตัว
❓ จุดที่ยังไม่ยืนยัน
  • ตัวเลข 2,462 PRs (Pt. 2) อ้างผ่าน The Neuron เท่านั้น — ยังไม่ได้ยืนยันจากตัว thread Pt. 2 โดยตรง
  • เส้นเวลาการย้ายบริษัทของ Lauren Tan (Ember → Netflix → Meta → Cursor → xAI) สรุปจาก README mirror + บทความ — ลำดับบางช่วง (โดยเฉพาะ Netflix) ยังไม่ได้ยืนยันจากแหล่ง primary
  • ความสัมพันธ์เชิงเวลาระหว่างการเป็น 'Cursor cloud agents' ในบทความ กับบทบาทปัจจุบันที่ xAI — บทความอาจเขียนช่วง transition หรือ mirror ตกหล่นรายละเอียด
🔗 แหล่งอ้างอิง (6)
04

การตอบรับจากชุมชน

กระแสบวกชัด (repo หลัก ~10k stars, port ตระกูล Claude Code 20+ repos) แต่มีคำวิจารณ์จริงเรื่อง token cost และพิธีกรรมเกินจำเป็นกับงานเล็ก

การตอบรับ pstack ของ Lauren Tan (poteto) ในชุมชนนักพัฒนาโดยรวมเป็นเชิงบวกอย่างมีนัยสำคัญและวัดผลได้จากตัวเลขจริง: ต้นทาง cursor/plugins มี 10,020 stars (วัด 2026-10-06), port หลักสู่ Claude Code (michael-denyer/pstack-claude) มี 1,504 stars, mirror แบบ standalone (backnotprop/pstack) มี 1,100 stars และเกิด 'port ecosystem' กว่า 20 repos ครอบคลุมแทบทุก harness (Codex, Pi, OpenCode, OMP, Hermes, DeepSeek, ChatGPT plugin, ZCode) ภายใน ~4 เดือน จุดที่ชุมชนชื่นชมมากที่สุดคือ verification skills (agent ต้องพิสูจน์ด้วยหลักฐาน ไม่พูดเองว่าสำเร็จ), การกำหนดขอบเขตด้วย disable-model-invocation และ skill ตัด slop อย่าง unslop ที่ถูกขอ/ดูดไปใช้ใน skill-pack อื่นหลายแห่ง (เช่น obra/superpowers issue #2187) ฝั่งวิจารณ์ที่มีหลักฐานชัดคือ token cost สูง (ทดสอบของ Rob O'Shaughnessy: งาน 30 นาทีกินเวลา ~1 ชั่วโมงบน pstack), ceremony เกินจำเป็นสำหรับงานเล็ก (Flavio Copes), ผู้ใช้จำนวนหนึ่งเลิกใช้หลังสัปดาห์แรกเพราะ friction ไม่เข้ากับสไตล์การเขียนโค้ดของตัวเอง (r/cursor) และตัวเลข '2,000+ PRs/เดือน' ที่เป็นแม่เหล็กความสนใจถูกเตือนว่าเป็น self-reported ที่ 'The raw number is less important than the operating model' ขณะที่ Hacker News แทบไม่มีบทบาท (มีแค่ 2 submissions ต่ำกว่า 3 คะแนน 0 comments) — การอภิปรายจริงเกิดบน Reddit, X และ YouTube แทน

👍 สิ่งที่ชุมชนชื่นชม
  • repo หลัก ~10,020 stars · pstack-claude 1,504 · mirror 1,100 · open-pstack 380
  • port ecosystem 20+ repos — Codex, Pi, Hermes, และ pstack-zcode สำหรับ ZCode โดยเฉพาะ
  • unslop คือของโห่ฮิตสุด — ถูกขอเข้า obra/superpowers + นำเข้า skill packs 4+ ชุด
  • ผู้ทดสอบจริงยืนยัน: จับ hallucination ได้ 3 จุดที่วิธีปกติไม่เห็น
👎 สิ่งที่ถูกวิจารณ์
  • token cost จริง: ~1 ชม. vs baseline 30 นาที — 'burns a lot of tokens'
  • พิธีกรรมเกินงานเล็ก: 'more ceremony than a task needs' (Flavio Copes)
  • คนจำนวนหนึ่งเลิกใช้หลังสัปดาห์แรก — friction ไม่เข้ากับ flow ของแต่ละคน
  • ตัวเลข 2,000+ PRs/เดือน = self-report ไม่มี audit — 'raw number is less important than the operating model'
ข้อค้นพบสำคัญ
01
ตัวเลขการยอมรับบน GitHub (วัดผ่าน GitHub API วันที่ 2026-10-06): cursor/plugins (ต้นทาง) 10,020 stars / 947 forks เปิด 23 ม.ค. 2026; michael-denyer/pstack-claude 1,504 stars / 158 forks (เปิด 26 พ.ค. 2026, push ล่าสุด 6 ต.ค. 2026); backnotprop/pstack (standalone mirror) 1,100 stars / 85 forks (เปิด 27 พ.ค. 2026); ericlitman/open-pstack 380 stars / 62 forks (เปิด 24 ส.ค. 2026); Luks3110/pstack-zcode 8 stars (เปิด 20 ส.ค. 2026); TheOnlyFusionCube/potetos-for-everyone 7 stars (เปิด 8 ก.ย. 2026)
02
GitHub repo search พบ port ecosystem 20+ repos ที่ประกาศชัดว่า derive/port จาก poteto's pstack: uzairansaruzi/p3-stack 111 stars ('pstack reworked for T3 Code'), Aqua-123/pstack-for-codex 80 stars (45 skills + 23 Poteto Mode playbooks), ScriptedAlchemy/pstack-codex 65 stars, dsebban/skills 64 stars (OMP-native), shrimpwtf/oh-my-pstack 32 stars, FetchUpstream/pstack-plugin 16 stars ('Chatgpt plugin for pstack from lauren potato'), kkgogogo17/pi-pstack, Cloeille/pstack (Hermes Agent plugin), irg1008/cstack (46 skills), negoro26/pstack-omp, HustleCoding/pstack-codex, foxytanuki/pstack-opencodex, painhardcore/pstack, iap/pstack-hermes, arjitj2/open-pstack, ajoslin/eng-mode, adjohn/pstack, aa2246740/pstack-dsh (DeepSeek Harness port) — สะท้อนว่า pstack กลายเป็น 'reference architecture' ของ skill-stack ในสาย agent harnesses
03
ระวัง name collision: no-session/pstack (126 stars) ไม่ใช่ของ poteto — README ยืนยันว่าเป็น fork ของ garrytan/gstack ('The solo founder's AI engineering stack. Fork of gstack') เป็นสแตกสำหรับ solo founder ที่ใช้ชื่อ pstack ซ้ำโดยบังเอิญ
04
skill ที่ชุมชนหยิบไปใช้มากที่สุดคือ unslop: obra/superpowers issue #2187 (เปิด 21 ส.ค. 2026 โดย rbjorklin: 'The unslop skill significantly improves the quality of LLM output') ถูกปิด 3 ก.ย. 2026 โดย Jesse Vincent (ผ่าน AI-triage agent 'Claude, Fable 5.1') ด้วยเหตุผลว่า 'unslop is a good skill' แต่ superpowers core 'does not vendor other projects' skills' และแนะทางที่ถูกคือ install จาก cursor/plugins ตรง ๆ (MIT) หรือ wrap เป็น plugin เล็ก; นอกจากนี้ยังมี PR นำ unslop/how/why/teach เข้า skill-pack อื่นอีก 5 แห่ง: hzblj/skills #22, MichaelLeeHobbs/skills #4, jointsome0-lgtm/selfos-skills #128, EpiLogos/ai-kit #110 ('Adopt Matt Pocock + pstack + cursor-team-kit as managed sources; unslop lands in the default set'), MakerCologne/ai-slop-ontology #153 (gap analysis เทียบ slopbeth ของ ehmo/slopkit กับ unslop)
05
Hacker News แทบไม่มีบทบาทต่อการตอบรับ: HN Algolia พบเพียง 2 submissions — 'pstack' (docs README, 1 ก.ย. 2026, 1 point, 0 comments) และ 'The Complete Guide to pstack Pt. 2' (ลิงก์ tweet ของ poteto, 14 ก.ย. 2026, 2 points, 0 comments); แต่มีคอมเมนต์ mention skills ของ pstack ในเธรดอื่น 3 แห่ง — ramoz (5 มิ.ย. 2026, Ask HN: dev stack) ใช้ skill interrogate, m3h (21 ส.ค. 2026) แนะนำ /unslop และชี้ port ที่ michael-denyer/pstack-claude, gosukiwi (22 ส.ค. 2026) ใช้ skill จาก pstack repo ('I use this one and it's ok')
06
Reddit เป็นเวทีหลักและเชิงบวกเป็นส่วนใหญ่: r/cursor 'Killer combo: Cursor Projects + pstack. Huge improvements in code quality and token efficiency' (u/explorioapp, ~27 ก.ย. 2026) เขียนว่า 'I've been using it for weeks and I continue to be amazed. It basically removes all the friction' และเผยว่าให้ coordinator agent ตัดสินใจ merge to main เองเพราะ 'It does a much more careful and thorough job of reviewing the changes than I (or 99.9% of engineers) could ever do'; r/cursor 'Thoughts on Cursor Projects?' มีคนแนะนำว่า 'Teach your projects coordinator bot pstack, invoke poteto-mode, and be prepared to be amazed'; r/ClaudeCode 'Wish me luck' แนะนำมือใหม่ให้ googling 'poteto pstack unslop' ไปใส่เป็น output style ของ Claude Code; r/buildinpublic มีโพสต์เล่าว่าปล่อย Cursor bot fleet ใช้ poteto-mode แบ่งงานเป็น todos + leaf skills ขณะตัวเองตัดหญ้า
07
วิจารณ์เชิงปฏิบัติที่จับต้องได้จาก Reddit: ในเธรด Killer combo เอง u/EzoLabsInc ตั้งข้อสังเกตว่า 'most people I've seen get hyped about pstack seem to stop using it after the first week because the workflow friction doesn't match how they actually code' — เป็นหลักฐานเดียวที่เจอแต่ชัดเจนเรื่อง drop-off; ผู้โพสต์ตอบว่าวิธีใช้ที่ดีคือไม่ต้องเรียก sub-skills เอง ใช้แค่ '/poteto-mode [describe what you want]' ตัวเดียว
08
บล็อกเชิงลึกที่หล่อหลอมภาพรวม: (1) Flavio Copes 'A deep dive into pstack' (อัปเดต 29 ก.ย. 2026) ชมว่า "'verify it' becomes a repository capability instead of a new conversation every time" แต่ติว่า 'That is sometimes more ceremony than a task needs' + 'If they all use frontier models, the tokens add up fast' และเสนอใช้ fstack (ตัวเบา) กับงานเล็ก เก็บ pstack ไว้กับ 'deeper investigation, adversarial review, real runtime proof'; (2) The Neuron (Grant Harvey, 10 ก.ย. 2026) ชม verification-as-infrastructure และ Feature Maps เป็น 'materialized memory' แต่รายงานทดสอบของ Rob O'Shaughnessy (aka Rob Shocks): งานเดียวกันที่ base ใช้ 30 นาที กินเวลา ~1 ชั่วโมงบน pstack และ 'he says plainly that the stack burns a lot of tokens' แม้จะจับ hallucinated claims ได้ 3 จุด — และเตือนเรื่องตัวเลข PR ว่า 'The raw number is less important than the operating model'; (3) Kayvane 'Skill orchestrations, and what I like about pstack' (28 พ.ค. 2026) ชม steps-copied-verbatim ว่า 'I think this is doing a lot of the heavy lifting in practice' รายงานผลใช้จริง 2 วัน ได้ ~12 PRs เป็น diff เล็ก ๆ 'a real velocity enabler without dragging up my cognitive debt' และเสนอแนวคิด continual learning loop: 'I think the next thing is this stuff starting to learn from itself'; (4) tenten.dev (developer.tenten.co, 5 ต.ค. 2026) หัวข้อ 'pstack Lets Agents Merge Their Own PRs Overnight, but Only After Fresh Verifiers Sign Off' ย้ำตัวเลข ~2,500 PRs ในเดือน ส.ค. 2026
09
การอภิปรายบน X/Twitter (อ่านผ่าน search snippets เพราะ login wall): @GorkaCesium วิจารณ์เชิงต้นทุนว่า 'I think most of it is poteto / pstack — verification and agent loops chew tokens fast when you stop watching every step. Less control, more throughput'; @zoomq (ZoomQuiet, นักพัฒนาจีน) โพสต์หลังคุยกับ poteto ว่า 'My hot take from chatting to @poteto is that we should use MORE...' พร้อมความเห็นภาษาจีนว่า 'pstack 是完全等效架构没有必要迁移' (pstack เป็นสถาปัตยกรรมเทียบเท่า ไม่จำเป็นต้อง migrate) — การกล่าวถึงภาษาจีนในชุมชน (juejin/zhihu/v2ex) หาไม่พบในการค้นครั้งนี้ สะท้อนว่า pstack เป็นวงการ English-speaking เป็นหลัก
10
ตัวเลข traction ตอน launch ที่สื่ออ้างไม่ตรงกันเล็กน้อย: Flavio Copes/daily.dev รายงาน skills ถูกใช้ '9,000 times inside Cursor in one week' ส่วน The Neuron รายงาน '10,000 times in one week' — ต่างกันเพราะจุดเวลาที่ quote ต่างกัน [uncertain]
11
pstack-claude คือศูนย์กลางชุมชน port ที่มีชีวิตจริง: มี 214+ issues/PRs, sync ตาม upstream ถึง v0.15.13 (PR #214, 5 ต.ค. 2026), ตอบ demand ของผู้ใช้ผ่าน issue #21 'Plans to update to the latest pstack version?' (22 ส.ค.), มี open issue #105 'could we use grok as a subworker through it's CLI?' (27 ก.ย.) สะท้อนผู้ใช้สนใจ multi-model ข้าม vendor, และมี tools/forks.json ประกาศ policy forks อย่างโปร่งใส (เช่น แก้ architect/blast-radius/bro/create-verification-skill เพื่อให้เข้ากับ Claude Code) — เป็นแบบอย่างการ maintain port ที่ถูกออกแบบให้ upstream กลับมา merge ได้
12
listing ecosystem ครอบคลุมทุก directory หลัก: mcpmarket.com/server/pstack ('Engineering Workflow Skills for AI Coding Assistants') + /tools/skills/poteto-mode-pstack-1 ('Claude Code Skill for Pro Engineers' — listing แสดงเพียง 7 GitHub stars บนหน้า listing) + /tools/skills/pstack-reflect + pstack-model-setup-1/-3; RoutineHub /skill/setup-pstack/ และ /skill/setup-benny/ (automation pack); skills.sh ติดตั้งผ่าน 'npx skills add backnotprop/pstack'; SkillMD ระบุ 'plugin of 41 Agent Skills for Claude Code, Cursor and 60+ agents'; Skill+DEV (skilldev.pro) เสนอ 'Pstack Skills — 23 Playbooks for Any AI Agent' เป็น single-skill port; ClawHub เก็บ mirror ของ skills.sh; drops.scottw.com/pstack-skills-map.html เป็น skill map ที่ชุมชนทำเองอธิบายความสัมพันธ์ของ ~45 skills; รวมถึง listing บน devtools.sh
13
ชื่อ pstack เองกลายเป็น meme เชิงบวกในเชิงเล่าเรียง: youtubesummary.com อธิบายว่า pstack ('Potato snack') ล้อชื่อ gstack ของ Garry Tan ส่วน reddit/สื่ออื่นพรรณนา Lauren Tan ว่า 'I stopped writing code. Now I run quality control on a kitchen of agents' (โพสต์ข้ามหลาย subreddit: r/artificial, r/Agent_AI, r/AgentsOfAI) — ตัว narrative 'engine กลายเป็น reviewer' เป็นสิ่งที่ชุมชนจับไปเล่นซ้ำมากกว่าตัว skill เดี่ยว
14
ไม่พบหลักฐานว่า Lauren Tan แสดงท่าทีต่อ ports อย่างเป็นทางการ (ยกเว้น README เดิมที่เขียนเชิญชวน 'fork it. improve it. make it yours. PRs are welcome!') — ports ทุกตัวให้ credit และใช้ MIT ตามต้นทาง; X ถูก login wall ทำให้ยืนยัน quote tweets ของเจ้าตัวไม่ได้ [uncertain]
กลไก · artifacts · quotes แบบเต็ม
⚙️ กลไกการทำงาน (how_it_works)
  • กลไกการตอบรับของ pstack ทำงานเป็น 3 ชั้นตามลำดับเวลา: ชั้นที่ 1 'story มาก่อนตัวผลิตภัณฑ์' — Lauren Tan เผยแพร่กระบวนการทำงานของตัวเองบน X ก่อน ('How I Use Cursor' 25 พ.ค. 2026, 'The Complete Guide to pstack' Pt.1 31 ส.ค. / Pt.2 14 ก.ย. 2026) พร้อมตัวเลข PRs ที่ merge ด้วย agent (2,462 PRs ในเดือน ส.ค. ในบทบาท 'Grok @Bot gardener') ทำให้ตลาดรับ pstack ในฐานนะ 'operating system ส่วนตัวของ engineer ระดับแนวหน้า' ไม่ใช่แค่ plugin หนึ่ง
  • ชั้นที่ 2 'การติดตั้งบน Cursor รายบุคคล' — ผู้ใช้ทั่วไปเริ่มจากติดตั้ง plugin แล้วใช้แค่ /poteto-mode (router) ตามคำแนะนำของ Flavio Copes ('install, run /setup-pstack, and just use /poteto-mode') และคำยืนยันใน r/cursor ว่า 'All you really need to know/do is install pstack and then use the command /poteto-mode' — จุดนี้ทำให้ newbie-friendly ตามที่ถกกันในเธรด
  • ชั้นที่ 3 'port เป็นระบบนิเวศ' — เพราะ pstack เป็น SKILL.md (plain markdown, MIT) ชุมชนจึง port ได้ 3 แบบ: (a) full-harness port ที่แปลง Cursor primitives เป็น primitives ของ harness ใหม่ (pstack-claude แปลงเป็น marketplace plugin + SessionStart hook + effort-agents แทน reasoning budget; รองรับ Claude Code/Codex/Copilot/Pi/OpenCode/Gemini/Prime Agent), (b) standalone mirror ที่ rewrite เฉพาะส่วน Cursor-specific (backnotprop/pstack: 'Many skills were Cursor-specific. They've been rewritten to work in any harness' + ตาราง dependency ว่า skill ไหนต้องติดตั้งคู่กับอะไร เช่น poteto-mode ต้องมี principle-* ทั้งชุด) และ (c) cherry-pick ทีละ skill (unslop, how, why, teach) เข้า skill-pack ของตัวเอง
  • กลไกการแพร่เชื้อ (adoption vector) ที่ได้ผลที่สุดคือ unslop เพราะเป็น skill เดี่ยวที่ไม่พึ่ง pstack อื่นและแก้ pain point สากล (LLM เขียน prose แบบ slop) — จึงถูก propose เข้า obra/superpowers, import เป็น PR ใน 4 skill-packs และถูกอ้างเป็น output style บน r/ClaudeCode; ขัดกับ poteto-mode ที่ port ทั้งชุดเพราะผูกกับ principle-* 24 ตัว + playbooks
  • กลไกการวิจารณ์ในชุมชนมี 3 เส้นหลัก: (1) ต้นทุน — token burn จาก verification loops (Rob O'Shaughnessy วัดสด ~2 เท่าของ base, @GorkaCesium 'chew tokens fast'), (2) ceremony/fit — workflow หนักเกินงานเล็ก (Flavio 'more ceremony than a task needs', Reddit drop-off หลังสัปดาห์แรก) และ (3) lock-in เชิง ecosystem — skill จำนวนหนึ่งสมมติ GitHub PRs, Graphite (Orchestrate), built-ins ของ Cursor (/create-skill, /babysit, /loop), /deslop + control-cli/control-ui จาก cursor-team-kit ซึ่งเป็นเหตุผลที่ mirrors ต้อง rewrite
  • วงจรป้อนกลับเชิงบวกที่ทำให้ reception โต: บล็อก deep dive (Flavio, The Neuron, daily.dev) → Reddit/X ยกมาใช้ → GitHub stars เพิ่ม → directory เอาไป listing (mcpmarket/RoutineHub/skills.sh/SkillMD/ClawHub) → คนเจอผ่าน search และ port ต่อ — จุดชนวนคือตัวเลข self-reported ของ Lauren (9k-10k ใช้ใน 1 สัปดาห์, 2,000+ PRs/เดือน) ที่ทำหน้าที่เป็น social proof แม้ The Neuron จะเตือนเรื่องการตีความ
📎 สิ่งที่จับต้องได้ (concrete_artifacts)
  • ต้นทาง: github.com/cursor/plugins (10,020 stars / 947 forks, สร้าง 2026-01-23) → ไดเรกทอรี pstack/; README ของ poteto ลงท้ายด้วยคำเชิญ 'fork it. improve it. make it yours. PRs are welcome!'
  • port หลัก: github.com/michael-denyer/pstack-claude — 1,504 stars / 158 forks (สร้าง 2026-05-26, push 2026-10-06, MIT); ติดตั้งด้วย '/plugin marketplace add michael-denyer/pstack-claude' + '/plugin install pstack@pstack-claude'; มี tools/forks.json ประกาศ policy forks; sync ถึง upstream v0.15.13 (PR #214, 2026-10-05); issue #21 'Plans to update to the latest pstack version?' (2026-08-22), open issue #105 'could we use grok as a subworker through it's CLI?' (2026-09-27)
  • mirror standalone: github.com/backnotprop/pstack — 1,110→1,100 stars / 85 forks (สร้าง 2026-05-27, MIT); ติดตั้ง 'npx skills add backnotprop/pstack'; README แบ่ง skills ที่ใช้เดี่ยวได้ (unslop, bro, how, tdd, typescript-best-practices, arena, swarm, interrogate, reflect, show-me-your-work, figure-it-out, automate-me, correct) กับ skills ที่ต้องติดตั้งคู่ (teach→how+why, architect→arena+how, blast-radius→arena+how+why+unslop, poteto-mode→principle-* ทั้งชุด) และหัวข้อ 'What this mirror changes' ยืนยันว่า rewrite ส่วน Cursor-specific แล้ว
  • ports อื่นที่ตรวจสอบแล้ว: ericlitman/open-pstack 380/62 (2026-08-24, 'tracking Cursor pstack upstream', มี UPSTREAM.md); TheOnlyFusionCube/potetos-for-everyone 7/0 (2026-09-08, zero-dependency, วิจารณ์ pstack-claude ว่า pin v0.14.8 เก่ากว่า upstream); uzairansaruzi/p3-stack 111 (2026-10-05, T3 Code); Aqua-123/pstack-for-codex 80; ScriptedAlchemy/pstack-codex 65; dsebban/skills 64; shrimpwtf/oh-my-pstack 32; FetchUpstream/pstack-plugin 16; Luks3110/pstack-zcode 8/0 (2026-08-20, port สำหรับ ZCode — ดู applicability_to_zcode); aa2246740/pstack-dsh 1 (DeepSeek Harness, license NOASSERTION); kkgogogo17/pi-pstack 8; Cloeille/pstack 8; irg1008/cstack 6; negoro26/pstack-omp 5; HustleCoding/pstack-codex 5; adjohn/pstack 4; iap/pstack-hermes 6; arjitj2/open-pstack 6; foxytanuki/pstack-opencodex 7; painhardcore/pstack 6; ajoslin/eng-mode 6
  • GitHub issue สำคัญเรื่อง adoption: obra/superpowers#2187 (2026-08-21, rbjorklin) ขอให้รวม unslop — ปิด 2026-09-03 โดย obra (ผ่าน AI agent 'Claude Fable 5.1' triage) เหตุ 'core is scoped to software-workflow skills and does not vendor other projects' skills'; EpiLogos/ai-kit#110 (2026-08-19) 'Adopt Matt Pocock + pstack + cursor-team-kit as managed sources; unslop lands in the default set'; PRs: hzblj/skills#22 (2026-08-19: how/teach/unslop/why), MichaelLeeHobbs/skills#4 (2026-09-05: unslop fork พร้อม Structure section), jointsome0-lgtm/selfos-skills#128 (2026-08-29: import unslop), MakerCologne/ai-slop-ontology#153 (2026-09-04: gap analysis slopbeth vs unslop)
  • HN artifacts: hn.algolia.com — story 'pstack' (https://news.ycombinator.com/item?id=49523858, 2026-09-01, 1 point, 0 comments), 'The Complete Guide to pstack Pt. 2' (id=49696222, 2026-09-14, 2 points, 0 comments, ลิงก์ไป twitter.com/poteto/status/2097732320606507506); comments: ramoz (2026-06-05, Ask HN dev stack — ใช้ 'interrogate'), m3h (2026-08-21 — แนะนำ /unslop + pstack-claude), gosukiwi (2026-08-22 — 'I use this one and it's ok')
  • Reddit artifacts: r/cursor/comments/1wrmig0 'Killer combo: Cursor Projects + pstack' (u/explorioapp, ~2026-09-27, cross-post ใน r/GrokBot 4 upvotes); r/AI_Agents/comments/1vtskra 'I've ported pstack plugin to ZCode' (u/frnluckhaos, ~2026-08, ลิงก์ github.com/Luks3110/pstack-zcode); r/cursor/comments/1uw5zpn 'AI writes terrible code' (ใช้ pstack แก้ + ยืนยัน skills ใช้ใน Claude Code ได้); r/ClaudeCode/comments/1vtj15v 'Wish me luck' (แนะ 'poteto pstack unslop'); r/cursor/comments/1wow14z 'Does anyone use /goal?' (แนะ pstack เป็นทางเลือกฟรี open-source); r/cursor/comments/1wpefwt 'Thoughts on Cursor Projects?' ('invoke poteto-mode, and be prepared to be amazed'); r/artificial/comments/1vrih0c + r/Agent_AI/1vriigp + r/AgentsOfAI/1vrihkn 'Lauren Tan: I stopped writing code' (2026-08); r/buildinpublic/comments/1w94xmq (bot fleet + poteto-mode); r/opencode/comments/1ws6pvm 'Top 10 OpenCode Skill Repos' (1.4K upvotes, 2026-09-28) — แนวเดียวกับ r/cursor 'Top 10 Cursor Skill Repos' (262 upvotes) ที่ pstack ติดอันดับ
  • listing artifacts: mcpmarket.com/server/pstack; mcpmarket.com/tools/skills/poteto-mode-pstack-1 (มุมมองเชิงขาย: 'transforms the AI's behavior into a senior engineer persona', แสดง 7 GitHub stars); mcpmarket.com/tools/skills/pstack-reflect; mcpmarket.com/tools/skills/pstack-model-setup-1 และ -3; routinehub.co/skill/setup-pstack/ ('Configure which models pstack uses per role'); routinehub.co/skill/setup-benny/ ('Benny ships as a dormant automation pack inside pstack'); skillmd.com/plugins/backnotprop-pstack/pstack ('41 Agent Skills ... 60+ agents'); skilldev.pro/en/skills/pstack-skill/ ('23 Playbooks for Any AI Agent', single-skill port); clawhub.ai/skills-sh/cursor/plugins/setup-pstack และ /skills-sh/backnotprop/pstack/setup-pstack; drops.scottw.com/pstack-skills-map.html (skill map ~45 skills); devtools.sh listing ของ pstack-claude (2026-09-25)
  • blog artifacts: flaviocopes.com/pstack/ (อัปเดต 2026-09-29: '24 workflow skills, 23 engineering principles, 22 task playbooks, 2 specialized subagents', โฆษณา course Solo Lab เปิด 28 ต.ค.); theneuron.ai/explainer-articles/pstack-explained-lauren-tans-system-for-trustworthy-ai-agents/ (Grant Harvey, 2026-09-10, v0.15.1, '23 playbooks, 23 principles, 47 skill entries, 2 subagents', ทดสอบ Rob O'Shaughnessy 13 นาที); kayvane.com/posts/building-a-multi-skill-system (2026-05-28); developer.tenten.co 'pstack Lets Agents Merge Their Own PRs Overnight, but Only After Fresh Verifiers Sign Off' (2026-10-05, 14 min read); daily.dev 'A deep dive into pstack' (2026-08-21, aggregator — permalink หาตรงไม่ได้ [uncertain])
  • X artifacts: twitter.com/poteto — 'How I Use Cursor' (2026-05-25: '/poteto-mode ... maximum signal per line'), 'The Complete Guide to pstack Pt. 1' (2026-08-31) / Pt. 2 (status 2097732320606507506, 2026-09-14); quote จาก @GorkaCesium (token critique) และ @zoomq (ภาษาจีน: 'pstack 是完全等效架构没有必要迁移')
  • YouTube: มีการอ้างว่า 'a ton of articles and youtube videos available' (r/cursor) และ youtubesummary.com มี 'Video Summary - Lauren Tan XAi Grokbot' (2026-09-01) ที่เรียก pstack ว่า 'Potato snack, parodying Gary Stack' — URL เดิม 404 แล้ว [uncertain]
💬 คำพูดและ claims
  • u/EzoLabsInc, r/cursor (2026-09-27): 'most people I've seen get hyped about pstack seem to stop using it after the first week because the workflow friction doesn't match how they actually code.' — วิจารณ์เรื่อง retention/drop-off
  • u/explorioapp, r/cursor (2026-09-27): 'I've been using it for weeks and I continue to be amazed. It basically removes all the friction.' และ 'It does a much more careful and thorough job of reviewing the changes than I (or 99.9% of engineers) could ever do.' — ชื่นชมสุดขั้ว + เปิดให้ agent merge to main เอง
  • r/cursor ('Thoughts on Cursor Projects?'): 'Teach your projects coordinator bot pstack, invoke poteto-mode, and be prepared to be amazed.'
  • Jesse Vincent / obra (ผ่าน AI triage agent, superpowers#2187, 2026-09-03): 'Thanks for the pointer; unslop is a good skill. It does not belong in Superpowers core, though: core is scoped to software-workflow skills and does not vendor other projects' skills ... The right path is what you already did: install it directly from cursor/plugins (it is MIT), or wrap it as a small plugin others can add.'
  • rbjorklin (superpowers#2187, 2026-08-21): 'The unslop skill significantly improves the quality of LLM output.'
  • Flavio Copes (flaviocopes.com/pstack/, 2026-09-29): "'verify it' becomes a repository capability instead of a new conversation every time" — ชม; 'That is sometimes more ceremony than a task needs.' และ 'If they all use frontier models, the tokens add up fast.' — ติ
  • Kayvane (kayvane.com, 2026-05-28): 'I think this is doing a lot of the heavy lifting in practice.' (เรื่อง steps copied verbatim) และ 'a real velocity enabler without dragging up my cognitive debt' (ผลใช้ 2 วัน ~12 PRs); เสนออนาคต: 'I think the next thing is this stuff starting to learn from itself.'
  • The Neuron (Grant Harvey, 2026-09-10) รายงาน Rob O'Shaughnessy (Rob Shocks): งาน 30 นาที base → ~1 ชั่วโมงบน pstack, 'he says plainly that the stack burns a lot of tokens', จับ hallucinated claims ได้ 3 จุด; และเตือนตัวเลข PR: 'The raw number is less important than the operating model.'
  • @GorkaCesium (X): 'I think most of it is poteto / pstack — verification and agent loops chew tokens fast when you stop watching every step. Less control, more throughput.'
  • @zoomq (X, ภาษาจีน): 'pstack 是完全等效架构没有必要迁移' (pstack เป็นสถาปัตยกรรมที่เทียบเท่าอย่างสมบูรณ์ ไม่จำเป็นต้อง migrate) — มุมของผู้ใช้ที่มีสแตกเทียบเท่าอยู่แล้ว
  • gosukiwi (HN, 2026-08-22): 'I use this one and it's ok' (เกี่ยวกับ skill จาก pstack repo) — เสียงกลาง ๆ ไม่สุดโต่ง
  • u/frnluckhaos (r/AI_Agents, ~2026-08): 'pstack is a plugin made by Lauren Tan ... its a great plugin for the overall dev workflow with agents, but its too specific to Cursor capabilities. I've rewritten some skills to fit better to ZCode ... all done with GLM 5.3, and it's being a great experience.' และเหตุผลที่เลือก GLM: 'Cost and quality, I've done a huge amount of work lately with frontier model quality with GLM 5.3 with a 30$ subscription'
  • backnotprop/pstack README: 'Mirror of cursor/plugins/pstack — kept in sync for standalone use. Works in Claude Code, Codex, Pi, and other agents, not only Cursor.' และ 'Many skills were Cursor-specific. They've been rewritten to work in any harness.'
  • Flavio Copes: 'pstack turns those habits into executable workflows' และสรุปท่าที: ใช้ pstack กับ 'deeper investigation, adversarial review, real runtime proof, and autonomous work that must remain auditable' ส่วนงานเล็กใช้ fstack — ปฏิเสธการใช้ทั้งชุดกับทุกงาน
  • ไม่พบ quote ตรงจาก Lauren Tan ตอบโต้/รับรอง ports หรือตอบวิจารณ์ token cost โดยตรง (X ติด login wall) — ข้อความที่ใกล้ที่สุดคือ README เดิม 'fork it. improve it. make it yours. PRs are welcome!' [uncertain]
🎯 นำไปใช้กับ ZCode
port_ตรงที่มีอยู่แล้ว
  • มี port สำหรับ ZCode อยู่จริง: github.com/Luks3110/pstack-zcode (8 stars, สร้าง 2026-08-20, MIT, upstream: cursor/plugins/pstack) — ประกาศบน r/AI_Agents โดย u/frnluckhaos ว่าเขียนใหม่ทั้งหมดด้วย GLM 5.3; แนวทางที่เขาเลือก: 'ZCode runs every subagent on the session model. the multi-model routing of the original becomes routing by subagent type' — ใช้ subagent types (poteto-agent / Explore / code-reviewer / code-architect / general-purpose) แทน per-role models และใช้ Cron automations แทน Cursor automations; repo 'doubles as a local ZCode marketplace' (มี marketplace.json, ติดตั้งผ่าน Settings → Plugin Management → Discover → +)
  • ข้อสังเกตจาก reception ที่ใช้ได้ทันที: ผู้ใช้ทั่วโลกสรุปตรงกันว่าควรใช้แค่ /poteto-mode ตัวเดียวไม่ต้องเรียก sub-skills เอง และควรเลือกติดตั้งเฉพาะ skills ที่ใช้เดี่ยวได้ (unslop, how, why, tdd, teach, blast-radius, bro, correct) — บน ZCode ซึ่งมี skill metadata budget ต่อ turn การเลือก subset แทนทั้ง 51 skills จึงสอดคล้องกับคำเตือนใน zcode-adoption.json และกับ feedback ของชุมชนเรื่อง ceremony
  • unslop คือ skill ที่คุ้มที่สุดที่จะ port ก่อน — ชุมชนพิสูจน์แล้วว่าใช้เดี่ยวได้ทั้งบน Cursor, Claude Code (ผ่าน output style ตาม r/ClaudeCode) และ skill-packs อื่น; บน ZCode copy เข้า ~/.zcode/skills/unslop/ หรือใช้ import/symlink ตาม docs ได้เลย
ต้อง_ดัดแปลง
  • บทเรียนจากวิจารณ์ token cost (Rob O'Shaughnessy ~2x เวลา, @GorkaCesium 'chew tokens fast'): การ port ไป ZCode ควรกำหนด tier — งานเล็กใช้ GLM-5.3-Flash และตัด verification loops, งานหนัก/autonomous ค่อยเปิด pstack เต็ม (ตรงกับที่ Flavio แยก fstack/pstack); ใช้ dynamic-workflows ของ ZCode คุม max_concurrency + subagent_model ต่อ role เพื่อจำกัด burn
  • บทเรียนจาก drop-off หลังสัปดาห์แรก (r/cursor): friction มาจากการที่ workflow บังคับขั้นตอนที่ไม่เข้ากับสไตล์เดิม — การ adapt บน ZCode ควรเปิดให้ 'stated reason to skip' (กลไกที่ Kayvane ชม) ทำงานเต็มที่ และอย่าบังคับ playbooks กับงาน trivial
  • บทเรียนจาก obra/superpowers: หน่วยงาน/ทีมที่มี skill core ของตัวเองไม่ควร vendor skills ของ pstack ทั้งชุด — ทางที่ชุมชนยอมรับคือ install จาก upstream (MIT) หรือ wrap เป็น plugin เล็ก (.zcode-plugin/plugin.json ของ ZCode อ่าน .claude-plugin ได้) — ตรงกับทางเลือก A/C ใน zcode-adoption.json
ใช้ไม่ได้_หรือไม่เกี่ยว
  • ตัวเลข reception (stars, PR counts) ใช้เป็น evidence ของคุณภาพทางเทคนิคไม่ได้ — ต้องทดสอบเองบน ZCode; โดยเฉพาะ pstack-zcode ยังไม่มี review ภายนอก, 8 stars, และ push ครั้งเดียว (สร้างและ push วันเดียวกัน 2026-08-20) จึงอาจไม่รองรับ ZCode รุ่นปัจจุบัน 3.14.x [uncertain]
  • no-session/pstack ที่ 126 stars ไม่เกี่ยวกับ poteto (fork ของ gstack) — ห้ามนำมาอ้างเป็น port
  • ประสบการณ์ Cursor-native (Cloud Agents, Composer routing, Bugbot) ที่ Reddit ชม ใช้เทียบเคียงบน ZCode ตรง ๆ ไม่ได้ เพราะเป็นความสามารถฝั่ง platform ไม่ใช่ของ skills
บทเรียนเชิงกลยุทธ์
  • แนวทางที่ชุมชนยืนยันว่าได้ผลสำหรับ 'prove-it-works': ให้ verification เป็น repository capability (verify-skill ประจำโปรเจกต์) แทนการเชื่อ agent — บน ZCode ใช้ hook Stop (block ได้ ≤3 รอบ) + Playwright MCP เป็นกลไกบังคับ ซึ่งเป็นสิ่งที่ r/cursor ยกเป็นเหตุผลที่ 'An agent that must verify can't lie about its work'
  • ถ้าจะแชร์ pstack-บน-ZCode ให้ทีม ทำตามแบบ pstack-claude: sync ตาม upstream เป็นระยะ, ประกาศ deviations ในไฟล์ประเภท forks.json, ให้ credit poteto + MIT — ชุมชนให้ค่ากับความโปร่งใสนี้ (potetos-for-everyone โจมตี pstack-claude เรื่อง version pin ก็เพราะเรื่องเดียวกันนี้)
⏱ Effort และสิ่งที่ต้องมี

ถ้าจะยกระดับจาก reception สู่การใช้งานบน ZCode: (1) ลอง pstack-zcode ของ Luks3110 ผ่าน local marketplace (~30 นาที รวมทดสอบ) — prereq: ZCode รองรับ plugin marketplace, hooks.enabled ที่ user-level; (2) cherry-pick unslop/how/why/tdd เข้า ~/.zcode/skills/ (~1-2 ชั่วโมง, ตรวจ frontmatter name+description ≤1024 ตัวอักษร); (3) สร้าง plugin ของตัวเองแบบ pstack-claude (SessionStart hook + effort-agents + forks.json) ~2-4 วัน — prereq: เข้าใจ hooks 7 events, subagents, dynamic-workflows และมี repo จริงสำหรับทดสอบ verification loop ทุกอันนี้ต่อยอดข้อสรุปของ zcode-adoption.json โดยตรง

⚠️ ข้อจำกัดและความเสี่ยง
  • ตัวเลขทั้งหมด (stars/forks/issues) วัด ณ 2026-10-06 และเปลี่ยนได้เร็ว — pstack เป็น repo ที่โตเร็วมาก (upstream เพิ่ม skills ระหว่าง 24→47-51 ชิ้นในไม่กี่เดือน ตามที่ zcode-adoption.json ระบุ) การอ้างจำนวนต้องระบุวันที่
  • ตัวเลข PRs (2,000+/เดือน, 2,462 ใน ส.ค. 2026) และ 9k-10k ครั้งใน 1 สัปดาห์ เป็น self-reported ของ Lauren Tan ไม่มี third-party audit; The Neuron เองเตือนว่า 'The raw number is less important than the operating model' — และสื่อรอง (tenten) สะกดนายจ้างผิดเป็น 'SpaceXAI' แสดงว่า chain ของข้อมูลบางส่วนผ่านการ copy ซ้ำ
  • เสียงเชิงบวกใน Reddit มาจากผู้ใช้ที่ยังใช้อยู่ (survivorship bias) — หลักฐานเรื่อง drop-off หลังสัปดาห์แรกมีจากความเห็นเดียว (u/EzoLabsInc) ไม่มีตัวเลข retention จริง; HN ไม่มี thread ใหญ่ทำให้ขาดมุมวิพากษ์เชิงลึกแบบ public
  • X/Twitter ถูก login wall — quote จาก @GorkaCesium/@zoomq มาจาก search snippets ยังไม่ได้อ่าน thread เต็ม; ปริมาณ quote tweets/likes ของประกาศของ poteto วัดไม่ได้ในรอบนี้ [uncertain]
  • การค้นภาษาจีน (juejin/zhihu/v2ex/CSDN) พบการกล่าวถึงเพียงผ่าน X ของ @zoomq — สรุปว่า pstack ยังไม่มี reception ในวงการภาษาจีนอย่างเป็นรูปธรรมอาจผิดพลาดได้ถ้ามีเนื้อหาที่ search engine ไม่ index
  • daily.dev permalink และ youtubesummary.com URL เดิม (404) ทำให้ตรวจสอบย้อนหลังเนื้อหาสองชิ้นนี้ไม่ได้เต็มรูป — ใช้ข้อความจาก search snippets เท่านั้น [uncertain]
  • การประเมิน mcpmarket/RoutineHub ทำได้แค่ระดับ 'มี listing และกรอบขายอย่างไร' — ไม่มีตัวเลข install/downloads สาธารณะบนหน้า listing (Poteto Mode แสดงแค่ 7 GitHub stars) จึงสรุปความนิยมผ่าน channel เหล่านี้ไม่ได้
  • ความน่าเชื่อถือของแหล่ง: บทความ deep dive (Flavio, The Neuron, Kayvane) เป็น first-hand แต่มีแรงจูงใจเชิงพาณิชย์ปะปน (Flavio ขาย course, The Neuron เชื่อม Rob Shocks course waitlist); port repos เป็นแหล่งที่มีผลประโยชน์ในการโปรโมตตัวเอง; ความเห็นที่เป็นกลางที่สุดคือ HN comments และ issue threads บน GitHub
❓ จุดที่ยังไม่ยืนยัน
  • key_findings[8] — จำนวน 9,000 vs 10,000 ครั้งใน 1 สัปดาห์ตอน launch (แหล่งต่างกันให้เลขต่างกัน)
  • key_findings[13] / quotes_and_claims[14] — ท่าทีทางการของ Lauren Tan ต่อ ports และต่อวิจารณ์ (X ติด login wall อ่านไม่ได้)
  • concrete_artifacts[9] — daily.dev permalink ของ 'A deep dive into pstack' (2026-08-21) หา URL ตรงไม่ได้
  • applicability_to_zcode.ใช้ไม่ได้_หรือไม่เกี่ยว[0] — คุณภาพ/ความเข้ากันได้ปัจจุบันของ Luks3110/pstack-zcode กับ ZCode 3.14.x (repo มี 1 commit, ไม่มี review ภายนอก)
  • risks_limitations[3] — ปริมาณ quote tweets/reactions บน X ของประกาศหลักของ poteto
  • risks_limitations[5] — เนื้อหาเต็มของ youtubesummary.com 'Lauren Tan XAi Grokbot' (URL 404) และเนื้อหาเต็มของบทความ daily.dev
  • การกล่าวถึง pstack ในชุมชนภาษาจีน (juejin/zhihu/v2ex) — ค้นแล้วไม่พบ แต่อาจมีนอกเหนือ index ของ search engine
🔗 แหล่งอ้างอิง (43)
05

ของเทียบเคียงในวงการ

ทุกแนวทาง (Anthropic, HumanLayer, CLI design guides) converge ที่จุดเดียวกัน: verification + agent-friendly tools — pstack คือ implementation ที่จัดชุดสำเร็จที่สุด

ภูมิทัศน์ของเทียบเคียงกับ pstack แบ่งเป็น 4 แกน: (1) verification/self-check loops — Anthropic มี guidance ตรงกัน (best practices, verification skills ใน 'How We Use Skills') แต่ pstack ก้าวไกลกว่าด้วยการทำเป็น skill ที่สร้าง/maintain ต่อ repo พร้อม feature map; (2) agent-friendly CLI design — เป็น genre ที่มีตัวอย่างแพร่หลาย (Anthropic 'Writing tools for agents', alexfurrier, Speakeasy, arcjet) โดยหลักการตรงกับที่ pstack ใช้; (3) memory-as-docs — feature map ของ pstack เทียบกับ AGENTS.md/llms.txt โดยต่างที่ feature map แยกต่อ feature เพื่อคุม token; (4) parallelism — Anthropic แนะ git worktrees บนเครื่อง, pstack เลือก cloud agents แทน, ฝั่ง ZCode มี subagents/background agents

ระดับ 3 — ระบบเต็ม (pstack) ← เราอยู่ตรงนี้
สร้าง/maintain verification skill + feature map + swarm ขนาน + principles
ระดับ 2 — วินัย
verification skill แยกไฟล์ + tools ที่ออกแบบให้ agent (Anthropic guidance)
ระดับ 1 — พอตัว
AGENTS.md/CLAUDE.md บอกวิธี test ไว้ในไฟล์เดียว
ยืนยันจากหลายแหล่งว่าหลัก CLI ตรงกันทุกข้อ: --dry-run ก่อนทำลาย · JSON output · idempotent · error ที่บอกทางออก — ไม่ใช่ความเห็นส่วนตัวของ poteto
ข้อค้นพบสำคัญ
01
Verification loops: Anthropic 'Claude Code: Best practices for agentic coding' (เม.ย. 2025) วาง workflow explore-plan-code-commit + ย้ำให้มีการ verify; ต่อมา 'Lessons from Building Claude Code: How We Use Skills' ระบุ verification skills (skill ที่บอกวิธีทดสอบ/ตรวจว่าโค้ดใช้ได้) ว่ามีประโยชน์มากต่อความถูกต้องของ output — แนวคิดเดียวกับ pstack แต่ pstack เพิ่ม create/maintain lifecycle + feature map + 'prove it' บน app จริง
02
Anthropic 'Writing tools for agents' เสนอหัวใจเดียวกับ Build the Lever: tool definition คือ 'contract ระหว่างระบบ deterministic กับ agent ที่ nondeterministic' — ออกแบบให้ model เรียกถูกและปลอดภัยได้ง่าย
03
Agent-friendly CLI design เป็น genre จริงที่มีข้อสรุปตรงกันหลายแหล่ง: alexfurrier.dev (idempotent commands, --dry-run, --yes/--force, resource+verb, machine-readable output), Speakeasy 'Making your CLI agent-friendly' (ก.พ. 2026 — interactive prompts, output formats, skill design), arcjet (predictable commands, machine-readable output), linkly.ai (ทำไม CLI คือ interface ที่ agent-friendly ที่สุด) — ตรงกับ CLI tips ในบทความ pstack ทุกข้อ แสดงว่าเป็น best practice ที่ converged แล้ว ไม่ใช่ความเห็นส่วนตัว
04
Parallelism: แนวทาง Anthropic = git worktrees บนเครื่องเดียว (git worktree add แล้วรันหลาย session ขนาน) — pstack โต้แนวทางนี้ชัดเจน ('cloud agents แทน worktrees') ด้วยเหตุผลลองใช้จริง; ฝั่ง cloud มี Cursor cloud agents, Codex cloud, Claude Code on the web, Devin; มีข้อควรระวังใหม่จาก security report (CVE-2026-55607 Claude Code Git Worktree Confusion — worktree ทำให้ verification steps ถูก bypass ได้) ซึ่งสนับสนุนข้อโต้แย้งของ pstack บางส่วน
05
Memory-as-docs: มาตรฐานทั่วไปคือ AGENTS.md (agents.md spec) / llms.txt / CLAUDE.md — จุดต่างของ feature map คือแยกไฟล์ต่อ feature + เขียนเป็น 'วิธี drive' ไม่ใช่คำอธิบายโค้ด ทำให้ agent เปิดเฉพาะไฟล์ที่เกี่ยว (scoped) และใช้เป็น regression sweep order ได้ — ตอบโจทย์ context budget ที่ wiki ทำไม่ได้ ('A wiki is great for humans. Agents pay for every token they reread.')
06
HumanLayer (12-factor agents, Advanced Context Engineering) วางกรอบเดียวกันว่า agent ดี = วิศวกรรม context + tools ที่ดี — pstack คือการ implement กรอบนี้ในรูป skills สำเร็จรูป
07
การเทียบต้นทุนรวม: ทุกแนวทางยอมรับ verification ที่แน่น = token แพงขึ้น (รายงานผู้ทดสอบ pstack: ~1 ชม. vs baseline 30 นาที) — จุดแลกคือใช้รุนแรงเฉพาะงานที่ 'พลาดแล้วแพง'
กลไก · artifacts · quotes แบบเต็ม
⚙️ กลไกการทำงาน (how_it_works)
  • แกนเทียบ 1 verification: แบบ minimal (best practices = บอก agent วิธี test ใน CLAUDE.md/AGENTS.md) → แบบกลาง (verification skill แยกไฟล์ ตาม Anthropic skills playbook) → แบบ pstack (skill สร้าง skill + maintain รายวัน + feature map + evidence บน app จริงผ่าน CLI) — ความต่างคือระดับ infrastructure และวงจรดูแล
  • แกนเทียบ 2 CLI: หลักทุกแหล่งตรงกัน — composable subcommands, --dry-run ก่อนทำลาย, error message ที่บอกทางออก, --help ละเอียด, JSON output, idempotency; pstack ยกเป็น principle 'Build the Lever' และบังคับผ่าน verification skill ที่ต้องมี CLI ควบคุม app
  • แกนเทียบ 3 memory: AGENTS.md = ข้อตกลงระดับ repo ไฟล์เดียว / llms.txt = index สำหรับ crawler / feature map = หลายไฟล์ต่อ feature ที่อ่านเฉพาะจุด — เลือกตามขนาด app และ token budget
  • แกนเทียบ 4 parallelism: worktrees (ฟรี บนเครื่อง เสี่ยง conflict ทรัพยากร + มี CVE เรื่อง verification bypass) vs cloud agents (แยกสิ่งแวดล้อมจริง แพง ต้องอินเทอร์เน็ต/สิทธิ์) vs subagents ใน session เดียว (เบา แชร์ context แม่ จำกัดด้วย concurrency) — ZCode อยู่แกนสุดท้ายเป็นหลัก
📎 สิ่งที่จับต้องได้ (concrete_artifacts)
  • Anthropic: https://www.anthropic.com/engineering (บทความ 'Claude Code: Best practices for agentic coding' เม.ย. 2025 และ 'Writing tools for agents')
  • Best practices ฉบับสมบูรณ์: https://code.claude.com (official best practices guide — verification approaches, context management)
  • Agent-friendly CLI: https://alexfurrier.dev/blog/2026-03-26-agent-friendly-cli-design ; https://www.speakeasy.com/blog (Making your CLI agent-friendly, ก.พ. 2026) ; arcjet.com 'Designing a CLI for AI agents'
  • HumanLayer: https://github.com/humanlayer/12-factor-agents ; humanlayer.dev/blog/advanced-context-engineering
  • ความเสี่ยง worktrees: CVE-2026-55607 (Claude Code Git Worktree Confusion — verification steps ถูก bypass) รายงานที่ penligent.ai
  • ตัวเทียบของจริงในงานนี้: ZCode มี skills (~/.zcode/skills), hooks, subagents, Playwright MCP, dynamic workflows — ตรงกับแกน verification+parallelism โดยไม่ต้องพึ่ง cloud
💬 คำพูดและ claims
  • 'A wiki is great for humans. Agents pay for every token they reread.' (poteto/verification-skill-example README — ประโยคสรุปฝั่ง feature map)
  • 'Tool definitions are contracts between deterministic systems and nondeterministic agents' (แนวคิดหลักของ Anthropic 'Writing tools for agents' — สรุปผ่านการอ้างอิงใน search)
  • 'Spend more agent compute where being wrong is expensive' (สรุปบทเรียนจากผู้ทดสอบ pstack จริง — The Neuron อ้าง Rob O'Shaughnessy)
  • แนวทาง Anthropic ทางการแนะ worktrees สำหรับขนาน ตรงข้ามกับ pstack ที่เลือก cloud agents — เป็น tradeoff จริงที่ยังไม่มีคำตอบเดียว
🎯 นำไปใช้กับ ZCode
  • ZCode ไม่ต้องเลือกอย่างใดอย่างหนึ่งระหว่างแนวทางเหล่านี้ — เลือก 'แบบกลาง-สูง' ได้ทันที: มี skills format เดียวกับ Anthropic/Claude Code (SKILL.md) จึงอ่าน best practices ของ Anthropic แล้วใช้ได้เลย และ pstack ที่ port มาก็วางบนฐานนั้นได้
  • แนะนำแบบที่เหมาะกับ ZCode จากการเทียบ: verification skill (ต่อ repo) + feature map (ต่อ feature) + CLI drive ผ่าน Playwright MCP + ขนานด้วย subagents/background agents ตามโควตา — คือส่วนผสมของทั้งสามแนวทางที่ไม่ต้องมี cloud agents
  • ถ้าโปรเจกต์เล็ก: AGENTS.md ไฟล์เดียว + verification skill หนึ่งตัวพอ — อย่าเริ่มจาก feature map 34 ไฟล์ (แบบตัวอย่าง Atlas) เพราะต้นทุน maintain สูงกว่าประโยชน์
  • ข้อเรียนจาก CVE worktree: ถ้าจะขนานด้วย worktrees บนเครื่อง (แนว Anthropic) ต้องตั้ง guard ให้ verification ไม่ถูกข้าม — ฝั่ง ZCode ใช้ hooks (เช่น gate ที่ Stop) ช่วยได้
  • หลัก CLI ที่ converged ทุกแหล่ง (--dry-run, JSON output, idempotent, error ที่บอกทางออก) ควรถูกใช้เป็น checklist ทุกครั้งที่เขียน tool ใหม่ให้ agent ในโปรเจกต์ของเรา
⏱ Effort และสิ่งที่ต้องมี

การนำแนวทางเทียบมาใช้กับ ZCode ไม่ต้องติดตั้งเพิ่ม — ขั้นต่ำคือเขียน verification skill หนึ่งตัว (ครึ่ง-หนึ่งวัน) และปรับ AGENTS.md ให้มีวิธี verify ของโปรเจกต์ (ชั่วโมงเดียว) ส่วนการขนานตามโควตา subagents ไม่มีต้นทุน setup แต่ต้องออกแบบให้ไม่ชนกัน (ไฟล์/branch แยกกัน)

⚠️ ข้อจำกัดและความเสี่ยง
  • ข้อสรุป CLI หลายอันมาจากบล็อกบุคคล/บริษัทเครื่องมือ (Speakeasy ขาย codegen, arcjet ขาย SDK) — มีแรงจูงใจเชิงธุรกิจให้ตอกย้ำว่า tool สำคัญ แม้หลักการจะสมเหตุผล
  • การเทียบ parallelism ยังไม่มี benchmark อิสระที่ดี (worktrees vs cloud vs subagents) — ตัวเลขที่มีส่วนใหญ่เป็น anecdote ของทีมใหญ่
  • CVE-2026-55607 อ้างจากรายงานเดียว (penligent.ai) ยังไม่ได้ตรวจกับฐานข้อมูล CVE ทางการ — ใช้เป็นสัญญาณความเสี่ยง ไม่ใช่ข้อเท็จจริงตายตัว
  • แนวทางของ Anthropic อิง Claude Code โดยตรง — mapping มา ZCode สมมติว่าความเข้ากันได้ระดับ SKILL.md/hooks ซึ่ง agent zcode-adoption ยืนยันแล้วว่าใช่ แต่ฟีเจอร์ย่อยบางตัว (hook event names) ควรเช็คเอกสาร ZCode อีกครั้งตอน implement
  • แกน 'memory-as-docs' มีความเห็นเยอะ ตัวเลขน้อย — ข้อสรุปว่า feature map ดีกว่า wiki ใช้ได้เฉพาะเมื่อ agent เป็นผู้อ่านหลักจริง ๆ
❓ จุดที่ยังไม่ยืนยัน
  • รายละเอียด CVE-2026-55607 (มาจากผล search แหล่งเดียว — ไม่ได้เปิดหน้ารายงานเต็ม)
  • เนื้อหาเต็มของ 'Lessons from Building Claude Code: How We Use Skills' (อ้างผ่าน search snippet บน X — ไม่ได้เปิดต้นฉบับเต็ม)
🔗 แหล่งอ้างอิง (7)
06

การนำไปใช้กับ ZCode

port ได้เกือบทั้งหมด และติดตั้งจริงบนเครื่องนี้แล้ว — สิ่งที่ไม่มีเทียบเท่าตรงคือ Cloud Agents และ Bugbot ของ Cursor

pstack ของ Lauren Tan (poteto) ปัจจุบัน (นับจาก GitHub API วันที่ 2026-10-06) มี 51 skills (24 เป็น principle-*, 27 ตัวหลัก รวม router ชื่อ poteto-mode และคู่ create/maintain-verification-skill), 2 subagents, 1 automation และ playbooks ~23 ไฟล์ใน skills/poteto-mode/playbooks/ — ชุมชนได้ port ไปยัง harness แบบ Claude Code สำเร็จแล้วอย่างน้อย 3 ราย (pstack-claude, potetos-for-everyone, open-pstack) จึงพิสูจน์ว่าแกนหลักเป็น markdown methodology ที่ port ได้สูง สำหรับ ZCode (Z.ai desktop/CLI harness ของ GLM-5.3) พื้นฐานครบ: skills, hooks 7 events, subagents, MCP, plugin format ที่อ่าน .claude-plugin/plugin.json ของ Claude Code ได้โดยตรง และมี built-in browser (Playwright) สำหรับ runtime verification จึงแนะนำ port แบบเลือกบางส่วน ~2-4 วัน หรือแบบ ZCode-idiomatic เต็มรูปแบบ ~1-2 สัปดาห์ สิ่งที่ port ไม่ได้คือ Cursor Cloud Agents (VM บนคราวด์), Bugbot และ multi-model panel (Opus 5.5/Grok) เพราะ ZCode รัน agents บนเครื่อง local และผูกกับตระกูล GLM

แผนที่ port — ใช้ได้จริงระดับไหน
✅ Port ได้ตรง
  • principles 21 ตัว
  • methodology skills (how, why, tdd, unslop…)
  • playbooks ทั้งหมด
  • agents 2 ตัว
🔧 ต้องดัดแปลง
  • poteto-mode router (ตัดของ built-in Cursor)
  • swarm/arena → dynamic workflows/subagents
  • control-ui → Playwright MCP
  • model panel opus/grok → GLM
❌ ใช้ไม่ได้
  • Cursor Cloud Agents (VM คราวด์)
  • Bugbot (GitHub App)
  • scale 2,000 PRs/เดือน (เป็นผลของระบบรวม)
ข้อค้นพบสำคัญ
01
ZCode primitives ยืนยันจากเอกสารทางการ zcode.z.ai/en/docs ครบทั้ง 8 หมวด: Skill (/en/docs/skill), Hooks (/en/docs/hooks), Subagents (/en/docs/subagents), MCP (/en/docs/mcp-services), Plugin (/en/docs/plugin), Command (/en/docs/commands), ZCode Agent (/en/docs/agents), Browser Automation (/en/docs/browser-use) — ตรงกับของจริงในเครื่อง: skills อยู่ ~/.zcode/skills/, plugins ที่ ~/.zcode/cli/plugins/cache/, มี dynamic-workflows ตั้งแต่ v3.14.0 (19 ก.ย. 2026) เรียกด้วย /workflow
02
โครงสร้าง pstack จริงวันนี้ (cursor/plugins repo, path pstack/ ไม่ใช่ plugins/pstack/): skills/ มี 51 ไดเรกทอรี = principle-* 24 ตัว + 27 ตัวหลัก (architect, arena, automate-me, benchmark-checklist, blast-radius, bro, correct, create-verification-skill, figure-it-out, how, interrogate, maintain-verification-skill, make-bot-ui, no-comments, poteto-help, poteto-mode, recall, reflect, setup-pstack, show-me-your-work, swarm, tdd, teach, technical-writing, typescript-best-practices, unslop, why); agents/ มี poteto-agent.md + comment-sicko.md; automations/benny; docs/guide/ 10 บท; manifest เป็น .cursor-plugin/plugin.json
03
media ช่วงแรกรายงาน '24 skills / 21-23 principles' ซึ่งเก่ากว่าของจริงมาก — repo เติบโตเร็ว (potetos-for-everyone เองวิจารณ์ pstack-claude ว่า pin อยู่ที่ v0.14.8 ขณะที่ upstream คือ v0.15.0+) — แผน port ต้องระบุ commit ที่ pin
04
มี port สำเร็จแล้ว 3 ตัว: (1) michael-denyer/pstack-claude — port หลักสำหรับ Claude Code/Codex/Copilot/Pi/OpenCode/Gemini/Prime Agent ติดตั้งแบบ marketplace plugin, ใส่ routing hook, มี agent pstack:effort-<level> แทน reasoning budget; (2) TheOnlyFusionCube/potetos-for-everyone — universal zero-dependency (ไม่มี runtime npm dependency), อ้าง '100% Upstream Parity (v0.15.0+)', มี SessionStart hooks + subagents (poteto-agent, comment-sicko) สำหรับ Claude Code, มี potetos panel --isolate สำหรับ parallel ผ่าน git worktrees; (3) ericlitman/open-pstack
05
เส้นทางลัดที่สำคัญที่สุด: ZCode docs/plugin ระบุ manifest lookup order '.zcode-plugin/plugin.json (recommended) → .claude-plugin/plugin.json (Claude Code compatible)' และ 'preloads the Claude Code marketplace' — หมายความว่า Claude Code plugin อย่าง pstack-claude น่าจะติดตั้งผ่าน UI (Add marketplace → ชี้ GitHub repo) ได้โดยแก้น้อยที่สุด [uncertain] (ยังไม่มีใครทดสอบติดตั้ง pstack-claude บน ZCode จริง)
06
ZCode docs/skill ยืนยัน import path โดยตรง: 'Import from external tools (Claude Code, Codex CLI, etc.) via symlink or copy modes' — skill ที่เป็น markdown ล้วนจาก pstack จึง copy ไป ~/.zcode/skills/<name>/SKILL.md ได้เลย ขอแค่ frontmatter มี name + description และ description ≤ 1024 ตัวอักษร (เกินแล้ว skill ทั้งตัวถูก drop ไม่ใช่ truncate)
07
ข้อจำกัดเชิงเทคนิคของ ZCode ที่แผน port ต้องออกแบบรอบ: (1) ทุก turn จะ inject metadata ของทุก skill ที่เปิดไว้ (name + description ตัดที่ 250 ตัวอักษร) ภายใต้ 'shared fixed budget' — pstack มี 51 skills จึงมีความเสี่ยง budget ล้นแล้ว injection เสื่อมเหลือแค่ชื่อ ทำให้ auto-trigger ตกฮวบ; (2) hooks ระดับ project ใน <workspace>/.zcode/config.json 'ignored as a whole' เพื่อความปลอดภัย — ต้องตั้งที่ user-level (~/.zcode/cli/config.json, hooks.enabled: true) หรือแพ็กใน plugin (hooks/hooks.json); (3) subagent ที่ตั้ง custom tools allowlist อาจเรียก skills ไม่ได้ถ้าไม่ allowlist Skill tool
08
สิ่งที่ port ไม่ได้เพราะเป็นของ Cursor โดยตรง: Cloud Agents (VM แยกบนคราวด์ รัน parallel ได้ไม่จำกัดโดยเครื่อง local ปิดได้), Bugbot (agentic security review บน GitHub PR), และ multi-model panel — pstack README ระบุ default panel คือ 'opus 5.5 / grok' โดยงาน code ส่งให้ grok งานหนัก/prose ให้ opus 5.5, ZCode เป็น GLM-first จึงต้อง remap เป็น GLM-5.3 / GLM-5.3-Flash หรือ providers อื่นที่ config ไว้ (workflow ของ ZCode ระบุ subagent_model ต่อ subagent ได้จาก ListModels)
09
หัวใจ pstack คือ verification loop ซึ่ง ZCode มีของครบจะยิ่งกว่า: built-in browser (Playwright-based, plugin browser-use มี skills control-browser + web-gui-tester), MCP servers, และ hook event Stop ที่ 'returning block continues the loop, at most 3 times in a row' — ใช้เป็น gate บังคับ prove-it-works ก่อนจบ task ได้แบบ native
10
ประเด็นคุณภาพที่ต้องระวัง: รีวิวจาก superbash.ai 'Don't Use Subagents on ZCode (GLM 5.3 Review)' แนะนำให้ main agent ลงมือเองมากกว่า coordinate subagents [uncertain — แหล่งเดียว] ซึ่งกระทบโดยตรงกับ skills ที่พึ่ง fan-out อย่าง swarm/arena/interrogate
11
ต้นทุนรายเดือนจาก zcode.z.ai: GLM Coding Plan Lite ¥94.4 (10,000 weekly credits), Pro ¥430.4 (6x), Max ¥862.4 (14x); ZCode v3.14.4 รองรับ macOS (Apple Silicon/Intel), Windows (x64/ARM64), Linux Beta
กลไก · artifacts · quotes แบบเต็ม
⚙️ กลไกการทำงาน (how_it_works)
  • กลไก ZCode skills: skill = ไดเรกทอรีที่มี SKILL.md (frontmatter ต้องมี name + description; ขาดตัวใดตัวหนึ่ง = ถูกเมิน) อยู่ที่ ~/.zcode/skills/<skill-name>/ หรือมาจาก plugin แบบ skills/<skill-name>/SKILL.md (ต้อง flat, ซ้อนโฟลเดอร์กลุ่มจะไม่ถูกหยิบ); เรียกใช้โดยพิมพ์ $ แล้วเลือกเป็น tag หรือผ่าน / menu; body เกิน 100KB จะถูก truncate ตอน load
  • กลไก ZCode hooks: config ที่ ~/.zcode/cli/config.json ต้องตั้ง hooks.enabled: true, events รองรับ 7 แบบเรียงตามรอบการทำงาน — SessionStart → UserPromptSubmit → PreToolUse → PermissionRequest → PostToolUse / PostToolUseFailure → Stop; executor 2 แบบ: process (argv ตรง แนะนำ) และ command (ส่งเข้า shell — มีไว้เพื่อ 'Claude Marketplace plugin compatibility'); stdin รับ JSON หนึ่งบรรทัดที่รับทั้ง camelCase และ 'Claude-compatible snake_case fields' (session_id, hook_event_name, tool_name, tool_input); exit code 2 = block shortcut; hook Stop คืน {"decision":"block"} ได้ต่อเนื่องสูงสุด 3 ครั้ง
  • กลไก ZCode plugin: โฟลเดอร์ plugin มี .zcode-plugin/plugin.json (หรือ .claude-plugin/plugin.json) + components 5 แบบ auto-detect: commands/*.md ($ARGUMENTS, $1/$2), skills/, agents/*.md (system prompt ของ subagent), hooks/hooks.json, .mcp.json (MCP keys จะถูก namespaced เป็น plugin:<name>:<server>); เพิ่ม marketplace ผ่าน UI 'Create -> Add marketplace' ชี้ GitHub repo/URL/local dir — ไม่มี slash command /plugin marketplace add แบบ Claude Code; เปิดใช้ plugin = เปิดทุก component พร้อมกัน + ได้ code-execution trust
  • กลไก ZCode subagents/workflows: subagent = ไฟล์ .md (name/description + system prompt) ใน agents/ ของ plugin; dynamic workflow (v3.14.0+) = สคริปต์ TypeScript ที่ orchestrate หลาย subagent พร้อม control flow (loop/condition) และ typed intermediate results, ปรับ concurrency ได้แม้รันอยู่ (v3.14.3) — ใช้แทน Cursor swarm/arena ได้เชิงโครงสร้าง
  • กลไกการ port pstack ที่แนะนำ (3 ทางเลือก): ทางเลือก A เร็วสุด — ใช้ potetos-for-everyone ผ่าน npx potetos (เขียนไฟล์ config/markdown ให้ harness แบบ generic) หรือลอง add marketplace ของ pstack-claude ใน ZCode UI เพราะ ZCode อ่าน .claude-plugin ได้; ทางเลือก B (แนะนำ) — copy skills ที่เลือกเข้า ~/.zcode/skills/ ผ่าน import/symlink แล้วแปลง poteto-mode เป็น command; ทางเลือก C — สร้าง ZCode plugin ของตัวเอง (.zcode-plugin/plugin.json) ที่ bundler skills + agents + hooks + commands แล้วแจกทีมเดียวกัน
  • การแปลง router: poteto-mode ของ pstack ใช้ frontmatter แบบ Cursor (mode: true, disable-model-invocation, icon, color, reminder) และอ้างอิงของ built-in ของ Cursor (create-skill, babysit built-in, deslop จาก cursor-team-kit, Bugbot, control-cli/control-ui) — ต้องลบ/แทนที่ทั้งหมดตอน port; บทบาท router แทน 'Cursor custom modes' ได้ด้วย command/*.md หรือ skill หลักที่ auto-trigger ผ่าน description
  • วงจร verification ที่ port แล้วเข้ากับ ZCode: create-verification-skill สอนให้สร้าง CLI เล็ก ๆ ที่ agent-friendly ไว้สั่ง/ตรวจแอป — บน ZCode ให้เขียน skill ที่หุ้ม Playwright MCP (browser-use) สำหรับ UI หรือเขียน MCP server/verify script ประจำโปรเจกต์ แล้วคุมด้วย hook PreToolUse (deny ถ้ายังไม่รัน verification) + Stop (block ต่อ ≤3 รอบจนกว่าจะ prove) — ตรงกับหลัก principle-prove-it-works และ principle-sequence-verifiable-units ของ pstack
  • การ fan-out แบบขนาน: pstack ใช้ skill swarm (coverage matrices, races, gauntlets) บน multi-agent ของ Cursor — บน ZCode ใช้ dynamic-workflows (CreateWorkflow กำหนด max_concurrency + subagent_model ต่อ agent) หรือ background Agent tool; งานที่ต้อง isolation ระดับ repo ให้ทำเหมือน potetos panel --isolate คือใช้ git worktrees
📎 สิ่งที่จับต้องได้ (concrete_artifacts)
  • repo ต้นทาง: github.com/cursor/plugins → ไดเรกทอรี pstack/ (skills 51 ตัวนับจาก GitHub contents API เมื่อ 2026-10-06, agents/ 2 ไฟล์, automations/benny, docs/guide/README.md + 10 บท 01-setup ถึง 10-recipes-and-pitfalls, skills/poteto-mode/playbooks/*.md เช่น babysit.md, perf-issue.md, hillclimb.md, autopilot-stack.md, multi-phase-plan.md, prototype.md, shipping.md, bug-fix.md, feature.md)
  • ตัวอย่าง frontmatter จริงของ poteto-mode: 'name: Poteto Mode', description ยาว, 'disable-model-invocation: true', 'mode: true', 'icon: crown', 'color: yellow', 'reminder: New task? Playbook match or rigor needed -> apply /poteto-mode.' — เป็นชุด field ที่ต้องแปลงตอน port ไป ZCode
  • ports ที่ใช้ได้จริง: github.com/michael-denyer/pstack-claude (install บน Claude Code ด้วย /plugin marketplace add michael-denyer/pstack-claude + /plugin install pstack@pstack-claude; MIT; มี tools/forks.json ประกาศ policy forks), github.com/TheOnlyFusionCube/potetos-for-everyone (npx potetos / npm install --save-dev potetos-for-everyone / curl installer; sync ถึง commit 71ed0d1 ของ cursor/plugins), github.com/ericlitman/open-pstack (มี UPSTREAM.md ตามรอยต้นทาง)
  • ตัวอย่าง verify-skill ที่ใช้แนวคิด pstack อยู่แล้วบน ZCode ในโปรเจกต์ของผู้ใช้เอง: skill eoffice มีเครื่องมือ eo_verify (deterministic asserts วัดจาก DOM จริง, เขียน evidence chain report.json + screenshot ทุกรอบ) และ eo_convention มี VISION gate — พิสูจน์ว่าแพทเทิร์น verification-skill ของ pstack ทำให้เหมือนกันบน ZCode ได้
  • คำสั่ง/พาธสำคัญฝั่ง ZCode: ~/.zcode/skills/<name>/SKILL.md (user-level), plugin = .zcode-plugin/plugin.json, hooks config = ~/.zcode/cli/config.json (hooks.enabled: true, timeoutMs default 60000, stdout cap 32768 bytes), project hooks ที่ <workspace>/.zcode/config.json ถูก ignore (log: config_project_hooks_ignored), env ของ plugin hooks: ZCODE_PLUGIN_ROOT / ZCODE_PLUGIN_DATA, template ${CLAUDE_PLUGIN_ROOT}
  • ตัวเลขเชิงปริมาณ: pstack ~0.15.x upstream / 51 skills / playbooks ~23 ไฟล์ [จำนวนเป๊ะยังไม่ได้ enumerate ทุกไฟล์ — uncertain]; ZCode v3.14.4 (29 ก.ย. 2026) เป็นรุ่นล่าสุดตอนวิจัย; hook events 7 แบบ; skill description cap 1024 chars; SKILL.md body cap 100KB; description excerpt 250 chars; Stop block ได้สูงสุด 3 ครั้งติด; ราคา GLM Coding Plan Lite/Pro/Max = ¥94.4/¥430.4/¥862.4 ต่อเดือน
  • แหล่งอ่านสำรอง: pstack.nishilfaldu.site (unofficial reader แยกหมวด Playbooks/Skills/Agents/Automations), docs/guide บน repo, Flavio Copes 'A deep dive into pstack', grokbot.sh 'The Complete Guide to pstack Pt. 1', atlasnote.ai 'pstack: Engineering Workflows for Cursor'
💬 คำพูดและ claims
  • pstack README (cursor/plugins, 2026-10): 'pstack is my answer. these are the same skills i use everyday to ship high quality code at Cursor. this turns cursor into a real engineering team. the goal is not to maximize loc, in fact it is the opposite.'
  • pstack README (2026-10) เรื่อง multi-model ซึ่งเป็นจุด port ยากที่สุด: 'the mode splits work by model strength: code delegates (feature, refactoring, bug fix, perf, hillclimb) go to grok, while the hardest changes, prose, and judgment go to opus 5.5. the default panel is opus 5.5 / grok.'
  • ZCode docs /en/docs/skill (ดึง 2026-10-06): 'Import from external tools (Claude Code, Codex CLI, etc.) via symlink or copy modes' และ 'description ห้ามเกิน 1024 ตัวอักษร — เกินแล้วจะ drop skill ทั้งตัว' (paraphrase จากภาษาไทยของ docs)
  • ZCode docs /en/docs/plugin (ดึง 2026-10-06): manifest lookup '.zcode-plugin/plugin.json (recommended) → .claude-plugin/plugin.json (Claude Code compatible)' และ hooks executor 'command — Hands the full string to the system shell' มีไว้เพื่อ Claude Marketplace plugin compatibility
  • ZCode docs /en/docs/hooks (ดึง 2026-10-06) เรื่อง Stop event: 'returning block continues the loop, at most 3 times in a row' และ project-level hooks 'ignored as a whole for security and logged as config_project_hooks_ignored'
  • ZCode changelog v3.14.0 (19 ก.ย. 2026): '新增动态工作流,用一段脚本编排多个子代理协作完成复杂任务' (เพิ่ม dynamic workflow — ใช้สคริปต์หนึ่งชิ้น orchestrate หลาย subagent ทำงานซับซ้อนร่วมกัน)
  • Cursor Docs — Cloud Agents (ดึงผ่าน search 2026-10): 'You can run as many agents as you want in parallel, and they do not require your local machine to be connected to the internet' — คือความสามารถหลักที่ ZCode ยังไม่มีเทียบเท่าตรง ๆ
  • ABAB News (26 ส.ค. 2026) รายงาน claim ของ Lauren Tan: 'Cursor engineer Lauren Tan stated that she merged over 1000 PRs last month, with a target of about double this month' — เป็นตัวเลขที่เจ้าตัวรายงานเอง (self-reported) ไม่มี third-party audit
  • potetos-for-everyone README (ดึง 2026-10): '100% Upstream Parity (v0.15.0+)' synced to commit 71ed0d1 และ 'exactly zero runtime npm dependencies' พร้อมตำหนิ pstack-claude ว่า 'pinned to older v0.14.8'
  • superbash.ai (2026) หัวข้อ 'Don't Use Subagents on ZCode (GLM 5.3 Review)': แนะนำให้ GLM 5.3 บน ZCode ลงมือเขียนเองมากกว่าคุม subagents [uncertain — แหล่งเดียว ยังไม่ cross-verify]
  • ไม่พบ quote ตรง ๆ จาก Lauren Tan เรื่องการรองรับหรือความเห็นต่อ ports เหล่านี้ (X/Twitter ติดขวางการอ่าน ไม่มีการยืนยันจากเจ้าของเดิม — ports ทุกตัวระบุชัดว่า unofficial, MIT with attribution)
🎯 นำไปใช้กับ ZCode
port_ได้_ตรง
  • principle-* ทั้ง 24 ตัวและ skill สาย methodology ส่วนใหญ่ (how, why, figure-it-out, correct, no-comments, unslop, tdd, blast-radius, technical-writing, typescript-best-practices, teach, show-me-your-work, poteto-help, recall, reflect, benchmark-checklist, bro) — เป็น markdown ล้วน ไม่ผูก API ของ Cursor; ZCode มี import path ทางการจาก Claude Code (symlink/copy) และ skill format เดียวกัน (SKILL.md + frontmatter name/description) — วิธี: copy เข้า ~/.zcode/skills/<name>/ หรือรวมเป็น plugin ที่มี skills/ flat
  • playbooks ทั้งหมดใน skills/poteto-mode/playbooks/*.md — markdown ล้วน port ตรง (จัดวางใต้ skill ของ router เดิมได้)
  • agents/poteto-agent.md + comment-sicko.md — แปลงเป็น subagent ของ ZCode ได้ตรง (agents/*.md ใน plugin, frontmatter name/description + body เป็น system prompt) หรือใช้เป็น custom agents ผ่าน plugin
  • หลักการ verification-skill แบบ agent-friendly CLI — โครงสร้างบน ZCode พร้อมกว่า: เขียน verify script/CLI ประจำโปรเจกต์แล้ว expose เป็น MCP server หรือ skill ที่หุ้ม Playwright MCP (browser-use plugin: control-browser, web-gui-tester) สำหรับ UI — โปรเจกต์ของผู้ใช้เอง (eoffice) มี eo_verify เป็น living proof ของแพทเทิร์นนี้
ต้อง_ดัดแปลง
  • poteto-mode (router กลาง): แปลง frontmatter Cursor (mode, icon, color, reminder, disable-model-invocation) เป็น ZCode skill/command ที่เรียกด้วย $ หรือ / menu; ตัด references ถึงของ built-in Cursor — create-skill → skill-creator plugin ของ ZCode, deslop (cursor-team-kit) → ตัดหรือเขียนเอง, babysit built-in/Bugbot → subagent review ที่ทำเอง, control-ui/control-cli → Playwright MCP + terminal ของ ZCode
  • multi-model panel (opus 5.5 / grok) และ skills ที่พึ่งหลาย model (interrogate = 'multi-model adversarial', arena = design/code bakeoffs, setup-pstack เลือก models): remap เป็น GLM-5.3 (หนัก) vs GLM-5.3-Flash (เบา/เร็ว) ผ่าน subagent_model ของ dynamic workflows หรือตั้ง providers เพิ่ม; ถ้าติดตั้ง provider อื่นได้ (config มี enabledBuiltinAgentCliProviders) ก็คงคอนเซ็ปต์ panel ได้บางส่วน
  • swarm/arena (parallel fan-out): ใช้ dynamic-workflows ของ ZCode (มี max_concurrency, typed results, control flow) แทนการปล่อย agent ขนานของ Cursor; งานที่ต้อง isolation ใช้ git worktrees ตามแบบ potetos panel --isolate; ต้องทดสอบคุณภาพก่อนเพราะมีรีวิวว่า GLM 5.3 ควบ subagents ได้ไม่ดีนัก [uncertain]
  • setup-pstack (เลือก reasoning budget + models): ไม่มี config ตรงแบบ Cursor — แปลงเป็น skill ตั้งค่าที่เขียนลง CLAUDE.md-equivalent / agent definitions ของโปรเจกต์ หรือกำหนด per-role reasoning ผ่าน workflow subagent_model
  • routing hook ของ pstack-claude: ZCode hooks format รับ 'command' executor ไว้เพื่อ Claude Marketplace compatibility อยู่แล้ว แต่ต้องเปิด hooks.enabled: true ที่ ~/.zcode/cli/config.json (user-level เท่านั้น — project hooks ถูก ignore) และยังไม่มีใครยืนยันว่า hook ของ pstack-claude รันบน ZCode ได้เป๊ะ [uncertain]
  • การ install: ไม่มี /plugin marketplace add เป็น slash command — ใช้ UI 'Create -> Add marketplace' ชี้ไปที่ GitHub repo แทน; ถ้าใช้ potetos-for-everyone แบบ npx potetos ต้องเลือก branch ที่เขียนไฟล์ระดับ harness มากพอ (มันเขียน CLAUDE.md/GEMINI.md ฯลฯ — สำหรับ ZCode ทางเลือก plugin/skills เหมาะกว่า)
ใช้_ไม่ได้
  • Cursor Cloud Agents: VM แยกบนคราวด์ รัน parallel ไม่จำกัด และเครื่อง local ปิดได้ — ZCode รัน agents/workflows บนเครื่องผู้ใช้เท่านั้น; ทดแทนบางส่วนด้วย background Agent/tasks (run_in_background), dynamic workflows, และ bot control ที่สั่งงาน ZCode ระยะไกลผ่าน WeChat/Feishu/Telegram — แต่เครื่องต้องเปิดอยู่
  • Bugbot (agentic security review ติด PR บน GitHub อัตโนมัติ): ไม่มี GitHub App เทียบเท่า — แทนด้วย subagent code-review รอบ PR ตามปฏิทิน/คำสั่งเอง และ triage ด้วยแนว references/bugbot-triage.md ที่ port มาเป็น playbook
  • Cursor IDE integration (Tab completion, inline edit, control-ui สำหรับ Electron/web UI): ZCode เป็น desktop app + CLI มี built-in browser แทน — skills สาย UI-testing ต้องเขียนใหม่บน Playwright MCP ไม่ใช่ Cursor's browser tool
  • scale 2,000 PRs/เดือน: ผลลัพธ์ของระบบรวม (Cloud Agents ขนาน + Bugbot + multi-model panel + วินัย skills) ไม่ใช่ของ skills ล้วน ๆ — ตั้ง expectation ว่า port ได้แค่ 'วิธีทำงาน' ไม่ใช่ 'ผลลัพธ์เชิง scale'
⏱ Effort และสิ่งที่ต้องมี

Effort 3 ระดับ: (1) ทดลองเร็ว ~0.5-1 วัน — add marketplace ของ pstack-claude (หรือ potetos-for-everyone) ใน ZCode UI → เปิดใช้ → เก็บบั๊ก frontmatter/hook; (2) selective port (แนะนำ) ~2-4 วัน — copy ~20 skills ที่เลือก (ตัดของผูก Cursor) เข้า ~/.zcode/skills/, แปลง poteto-mode เป็น command, port agents 2 ตัว, เปิด subset เพื่อไม่ล้น skill injection budget; (3) full ZCode-idiomatic ~1-2 สัปดาห์ — เขียน verification skill ใหม่บน Playwright MCP/verify CLI, ใส่ hooks gates (PreToolUse deny + Stop block ≤3), ทำ workflow แทน swarm/arena และทดสอบกับโปรเจกต์จริง 1-2 งาน สิ่งที่ต้องมีก่อน: ZCode desktop v3.14+ ติดตั้งแล้ว, เปิด skills ใน Settings, ยอมรับ code-execution trust ของ plugin, hooks.enabled: true ที่ user config, โปรเจกต์ปลายทางต้องมีเป้า verification ที่จับต้องได้ (test suite/CLI/UI ที่ Playwright เข้าถึงได้), และบัญชี GLM Coding Plan (Lite ¥94.4/เดือน ขึ้นไป); ถ้าจะใช้ import ผ่าน symlink ต้องเตรียมโฟลเดอร์ pstack ที่ clone ไว้ในเครื่อง (หรือจาก repo pstack-claude/potetos-for-everyone ที่แตกมาแล้ว)

⚠️ ข้อจำกัดและความเสี่ยง
  • skill injection budget เป็นความเสี่ยงจริง: ZCode inject name + description 250 ตัวอักษรของ 'ทุก' skill ที่เปิดไว้ในทุก turn ภายใต้ shared fixed budget — เปิด pstack ครบ 51 ตัวมีโอกาสล้น budget แล้วระบบ 'degrades to names only' ทำให้ auto-trigger แย่ลงมาก — ต้องเปิดแบบ subset (~15-20 ตัวหลัก) และวัดผล
  • คุณภาพ subagent บน GLM 5.3 ยังเป็นข้อถกเถียง (มีรีวิวแนะนำเลี่ยง) [uncertain — แหล่งเดียว] — skills ที่พึ่ง fan-out หนัก (swarm, arena, interrogate) อาจให้ผลแย่กว่าบน Cursor ที่ panel เป็น Opus/Grok; ควร benchmark ก่อนใช้จริงจัง
  • upstream เปลี่ยนเร็ว: skills โตจาก ~24 เป็น 51 ในไม่กี่เดือน — port ที่ pin ไว้จะเก่า (เห็นชัดจาก pstack-claude ถูกตำหนิว่า pin v0.14.8); แผนต้องระบุ commit และกำหนดรอบ sync
  • ความน่าเชื่อถือของแหล่ง: ข้อมูลฝั่ง ZCode มาจากเอกสารทางการ (zcode.z.ai) + การตรวจของจริงในเครื่อง = แข็งแรง; ฝั่ง pstack มาจาก GitHub API (primary) = แข็งแรง; แต่ claim ปริมาณงาน (1,000-2,000 PRs/เดือน) เป็น self-reported โดยเจ้าของ และ claim ด้านคุณภาพ subagent/GLM มาจาก third-party รายเดียว — ยังไม่มี cross-verification อิสระ
  • ความเสี่ยงด้าน hooks: แม้ format เข้ากับ Claude Code แต่ ZCode fail-open (hook crash/ช้า ปล่อย action ผ่าน — จากรีวิว vibecoding.app [uncertain — แหล่งเดียว]) จึงใช้ hooks เป็น 'กันพลาด' ไม่ได้ระดับ security guarantee; และ project-level hooks ถูก ignore ทั้งหมด — ทีมต้องแจกผ่าน plugin ที่เปิดด้วยตนเอง
  • เปิด plugin = มอบ code-execution trust (เอกสาร ZCode เตือนเอง) — ต้อง review hooks/scripts ของ port ก่อนเปิดใช้ โดยเฉพาะ pstack-claude ที่ติดตั้ง routing hook อัตโนมัติ
  • การยืนยันว่า pstack-claude ติดตั้งบน ZCode ได้จริงยังเป็นการอนุมานจาก compatibility ที่เอกสารระบุ (.claude-plugin + Claude-compatible hooks) ไม่ใช่ผลทดสอบ — ควรจัดสรรเวลาไว้แก้ edge case (frontmatter แปลก ๆ อย่าง mode/icon/color/reminder, ความยาว description, relative path ของ playbooks)
  • ZCode เป็นผลิตภัณฑ์วัยเยาว์ที่เปลี่ยนไว (v3.10→v3.14 ภายใน ~1 เดือน, Linux ยัง Beta) — รายละเอียด format อาจเปลี่ยน; ควร pin รุ่นและอ่าน changelog ก่อน upgrade
❓ จุดที่ยังไม่ยืนยัน
  • key_findings
  • quotes_and_claims
  • applicability_to_zcode
  • risks_limitations
🔗 แหล่งอ้างอิง (18)
07

คู่มือใช้งานทั้ง 44 Skills

ทุก card มาจากการอ่าน SKILL.md ต้นฉบับเต็มไฟล์ — กดเพื่อกาง · ค้นหา/กรองตามหมวดได้
5
auto เอง
39
คุณเรียก / poteto-mode เรียกให้
Auto เอง 5 ตัว: how why unslop setup-pstack typescript-best-practices — ที่เหลือถูกตั้งใจให้ "คุณ route, agent ลงมือ" และ auto ที่แท้จริงคือ /poteto-mode เรียกทุกอย่างให้
A · งานปกติ (ใช้บ่อยสุด)
/poteto-mode <โจทย์>
มันจัดการทุกอย่างเอง
B · งานเล็ก เอาวินัยเฉพาะจุด
เรียกตรง เช่น /tdd · /no-comments · /blast-radius
C · สืบปัญหา
/how → /why → คลุมเครือ /figure-it-out → ไม่เชื่อ agent /show-me-your-work
D · ก่อน merge ของเสี่ยง
/blast-radius → /interrogate → PR

🚦 Setup & Routing (3)

👤 เรียกเอง/poteto-modeจุดเข้าหลัก (router) ของ pstack ทั้งชุด คือ style การทำงานของ Lauren Tan ('poteto') สำหรับ coding agent: ตอบก…
📦 ทำอะไร

จุดเข้าหลัก (router) ของ pstack ทั้งชุด คือ style การทำงานของ Lauren Tan ('poteto') สำหรับ coding agent: ตอบกระชับแบบประโยคสั้น, ใช้ subagent อย่างมีสติ, prose ไม่เป็น AI-slop, โค้ดเรียบง่าย และงานทุกชิ้นต้องพิสูจน์ว่าใช้งานได้จริง เมื่อ user เรียก /poteto-mode <task> มันจะอ่าน Principles ให้ครบก่อน จับคู่งานกับ playbook 1 แบบจาก 23 แบบ คัดลอก step ของ playbook ลง todolist แบบคำต่อคำ แล้ว chain skill อื่นของ pstack (how, why, tdd, swarm, arena, interrogate ฯลฯ) ตามที่ playbook กำหนด

💡 เจ๋งยังไง

แก่นคือสถาปัตยกรรม router + playbook ที่ทำให้เจตนาของมนุษย์กลายเป็นขั้นตอนที่ agent ทำตามได้แบบ deterministic ไม่ใช่แค่ 'สไตล์ที่ดี' ลอย ๆ กลไกที่จับต้องได้: (1) step ของ playbook ต้องถูก copy เข้า todolist แบบคำต่อคำ 'ก่อน' จะคิดวิเคราะห์ task — ตัด failure mode ที่ agent อ่าน playbook แล้วเขียนแผนเองจนทำ step สำคัญ (architect, throughput checkpoint) หายไป และห้าม skip เงียบ ๆ ต้องใส่บรรทัด skip: <เหตุผล> (2) principle-* ทั้ง 21 ตัว (Core 9, Architecture 6, Verification 3, Delegation 2, Meta 1) เป็นกฎแบบมี trigger กำกับ และทุก citation ต้อง trace ไปยังการตัดสินใจจริง — 'citation ที่ไม่มี decision อยู่เบื้องหลัง = คุณข้าม leaf skill ไป' (3) กฎตัดสิน fork แบบ empirical ก่อนถามมนุษย์: ถ้าคำตอบหาได้จากการรันจริง (พฤติกรรม, timing, layout, perf) ห้าม AskUserQuestion — sketch ผ่าน Prototype playbook แล้วให้ผลลัพธ์เป็นคนตัดสิน (4) กฎ Autonomy ชัดเจน: reversible ทำเลย หยุดเฉพาะ irreversible writes (force-push, deploy, ลบ data, ส่งข้อความถึงลูกค้า) (5) กฎเขียน reply แบบ anti-slop ที่วัดผลได้ เช่น ห้ามอักษร long-dash, ห้าม colon คั่นกลางประโยค, สร้างประโยคให้สะอาดตั้งแต่แรกเพราะ 'cleanup ทีหลังถูกวัดแล้วว่าพัง' (6) เจ้าของงานคือ agent หลักเสมอ — review diff ของ subagent เอง ไม่ pass-through คำพูดของ subagent

🧭 วิธีใช้

เรียกใช้: /poteto-mode <task description> เช่น /poteto-mode fix this bug หรืออธิบายงานต่อท้ายคำสั่ง สิ่งที่เกิดขึ้นตามลำดับ: (1) todolist แรกของทุก task หลายขั้นคืออ่านส่วน Principles ให้ครบทั้งไฟล์ (2) agent จับคู่ task กับ playbook จาก 23 แบบ: Investigation (คำถาม read-only), Bug fix (repro ก่อนแก้), Perf issue (วัด baseline), Hillclimb (ปรับ metric ต่อเนื่อง), Runtime forensics (diagnose จาก live instrumentation), Trace forensics (diagnose จาก cpuprofile/trace ที่ capture แล้ว), Feature, Refactoring, Prototype (sketch ตัดสินใจถูก), Visual parity (UI เป๊ะพิกเซล), Authoring a skill, Eval, Babysit (ดัน PR/stack ให้ green), Shipping (verify แล้ว land), Autonomous run (รุดจนจบ), Orchestrate (โปรเจกต์หลายวันหลาย PR ใต้ coordinator เดียว), Autopilot-full (คิว PR รันถึง merged), Autopilot-stack (สร้าง stack ให้ operator กด land เอง), Session pickup (รับงานค้างจาก agent ก่อนหน้า), Pause safely (พักงานให้ resume ได้), Multi-phase or multi-PR plan, Worktree and simulator cleanup, Opening a PR — ถ้าไม่มี playbook ไหน fit ใช้ skill figure-it-out ออกแบบ playbook เฉพาะงานให้ (3) copy step ของ playbook ลง todolist คำตอบคำ แล้วทำตาม โดย playbook แต่ละแบบจะสั่ง chain skill เอง เช่น Bug fix จะ fan-out how + why แบบขนาน, ใช้ architect ก่อนแก้ที่ข้าม function boundary, ใช้ tdd เมื่อมี test path ถูก, แล้วจบด้วย Opening a PR (4) redirect กลางทางได้: เรียก principle ด้วยชื่อตรง ๆ (เช่น สั่ง 'apply principle-laziness-protocol') หรือสั่งเปลี่ยน playbook (5) แนวคิด sticky mode: บน Cursor field mode: true ทำให้ style นี้ค้างเป็นโหมดของ session และ field reminder จะยัด nudge 'New task? Playbook match or rigor needed -> apply /poteto-mode. Casual turn or user opts out -> don't.' เพื่อดันให้ agent กลับมาใช้โหมดนี้กับ task ใหม่ทุกครั้ง บน ZCode กลไก nudge นี้อาศัย description ของ skill แทน

💬 ตัวอย่าง prompt

/poteto-mode this pr has a subtle bug where the scroll drifts every 750ms even when idle. repro first, then fix and verify. หรือแบบ autonomous: /poteto-mode i'm going to bed. land the stack even if ci flakes. i want everything merged by morning.

⚡ Auto-invoke

ไม่ auto-invoke — frontmatter มี disable-model-invocation: true ชัดเจน เจตนาออกแบบคือ USER เป็นคนเรียก /poteto-mode เอง แล้วมันจะเป็น entry point ที่ chain skill อื่นของ pstack ทั้งหมดตาม playbook อย่าง deterministic (skill ที่ถูก chain อย่าง how/why/tdd/swarm ฯลฯ จะถูกเรียกโดย poteto-mode ไม่ใช่เรียกเอง) field mode: true, icon, color และ reminder เป็น field เฉพาะของ Cursor ซึ่งบน ZCode น่าจะถูก ignore [uncertain] แต่เจตนายังใช้ได้: บน Cursor reminder ทำหน้าที่ nudge ให้ agent กลับมา apply /poteto-mode กับทุก task ใหม่ ส่วนบน ZCode trigger นั้นอาศัย description ('Use for poteto, /poteto-mode, or requests to work in this style') ที่ user ต้องพิมพ์เรียกเอง

🔗 คู่กันกับ

เป็น hub ที่ chain skill อื่นเกือบทั้งชุด: how, why, tdd, swarm (fan-out ขนาน), arena (design/code bakeoff), interrogate (adversarial review), architect, reflect, unslop, no-comments, technical-writing, show-me-your-work, figure-it-out, create-verification-skill โหลด principle-* 21 ตัวเป็นกฎพื้นฐาน ใช้ subagent_type poteto-agent เป็น default สำหรับ code delegate (ปรับได้ผ่าน ~/.zcode/pstack-roles.md ที่ /setup-pstack เขียนให้) คู่ควรรัน /setup-pstack ก่อนหนึ่งครั้งเมื่อติดตั้ง pstack

⚠️ เคล็ดลับ/ข้อควรระวัง

(1) อย่าเขียนแผน bespoke ทับ step ของ playbook — ต้อง copy คำตอบคำ และ step ที่ข้ามต้องเหลือใน list พร้อมบรรทัด skip: <เหตุผล> (2) อย่า citation principle ลอย ๆ ทุก principle ที่อ้างต้องเปลี่ยนการตัดสินใจจริงอย่างน้อยหนึ่งจุด (3) อย่า override subagent_type ของ skill ที่ถูก route (how, why, interrogate, reflect, swarm ตั้ง type เองเพื่อ diversity) (4) กฎเขียน: ห้าม long-dash ทุกกรณี, ห้าม colon คั่นกลางประโยค (colon นำ list ได้), ประโยคสั้น declarative แต่ห้ามตัดเนื้อหา — ทุก section ที่ playbook สั่งยังต้องอยู่ (5) อย่าถามมนุษย์เรื่องที่รันหาคำตอบได้ — AskUserQuestion เก็บไว้เฉพาะ product/preference call จริง ๆ (6) 'No is an acceptable answer' — agent ต้องตอบด้วยความเห็นจริง ไม่ validate ให้ user ฟรี ๆ (7) เจอ task ใหญ่ข้าม domain หรือ user ไม่อยู่ route ไป figure-it-out หรือ Orchestrate ไม่ใช่ฝืนใส่ playbook แคบ (8) comment ในโค้ดก็ใต้กฎเดียวกับ prose: เขียนสะอาดตั้งแต่แรก เก็บเฉพาะ why ที่โค้ดเล่าเองไม่ได้

❓ ยังไม่ยืนยัน
  • auto_invoke
✨ auto/setup-pstackSkill ตั้งค่าครั้งแรก (first-run configuration) ของ pstack บน ZCode: ตรวจจับ subagent type ที่มีอยู่จริงใน se…
📦 ทำอะไร

Skill ตั้งค่าครั้งแรก (first-run configuration) ของ pstack บน ZCode: ตรวจจับ subagent type ที่มีอยู่จริงใน session, map มันเข้ากับ role ต่าง ๆ ที่ pstack ใช้ (code delegate, critic, investigator, arena runner ฯลฯ) แล้วเขียนลง ~/.zcode/pstack-roles.md เป็น override layer ที่ skill อื่นอ่านเมื่อต้อง spawn subagent และปิดท้ายด้วยการเสนอสร้าง project-local verification skill ถ้าโปรเจกต์ยังไม่มี

💡 เจ๋งยังไง

แก้ปัญหาเฉพาะตัวของ ZCode: ทุก subagent รันบน session model เดียวกัน จึงไม่มี per-role model choice — diversity ที่เหลืออยู่คือ subagent type และ prompt ซึ่ง skill นี้ทำให้ 'ตั้งค่าได้' อย่างเป็นระบบ เพราะ reviewer type อ่าน diff ไม่เหมือน architect type หรือ general-purpose delegate จุดเด่นจากกลไกจริง: (1) เป็น override layer ไม่ใช่ข้อบังคับ — บรรทัดไหนไม่มีในไฟล์ = fallback ไป inline default ของ skill นั้นทันที (2) กัน config พังตั้งแต่ต้นทาง: 'Never write a type you have not confirmed exists' เพราะ role line ที่ชี้ type ที่ Agent tool ไม่รู้จักจะ break ทุก delegation ที่อ่านไฟล์นี้ (3) idempotent — rerun แล้ว overwrite ทั้งไฟล์ ไม่มี state ค้าง (4) panel role (how critics, arena runners, interrogate reviewers) รับเป็น list ทำให้ความยาว list = จำนวน fan-out และใส่ type เดิมซ้ำได้ถ้าอยากได้ volume แทน diversity (5) step 7 เสนอ /create-verification-skill ให้โปรเจกต์ที่ยังไม่มีวิธี drive app จริงเพื่อพิสูจน์งาน — ปิด loop verification ตั้งแต่วันแรก

🧭 วิธีใช้

เรียกใช้: /setup-pstack หรือพิมพ์ 'configure pstack agents' / บอกว่าจะเปลี่ยน subagent choice ของ pstack ขั้นตอนที่ agent ทำ 7 ขั้น: (1) Detect — enumerate subagent_type ที่ session นี้ใช้ได้จริง จาก 3 แหล่งตามลำดับ: built-ins ที่มีเสมอ (general-purpose, Explore, code-architect, code-explorer, code-reviewer), type จาก plugin ที่เปิดใช้ (pstack ให้ poteto-agent และ comment-sicko), และ agent definitions นอก plugin เช่น ~/.zcode/cli/agents/ (2) Load current state — ถ้า ~/.zcode/pstack-roles.md มีอยู่แล้วใช้ค่าเดิมเป็น current, ไม่มีก็เริ่มจาก default mapping (3) Map and confirm — โชว์ตาราง role ทุกตัวกับ type ปัจจุบัน ถาม user ผ่าน AskUserQuestion ว่ารับตามนี้หรือจะเปลี่ยน role ไหน โดย role แบบ panel เป็น list (4) Validate — type ทุกตัวที่จะเขียนต้องอยู่ในชุดที่ detect เจอ ไม่งั้นหยุดถามใหม่ (5) Write — เขียน ~/.zcode/pstack-roles.md ทับทั้งไฟล์ รูปแบบบรรทัดละ role: 'how critics: code-reviewer, poteto-agent, general-purpose' (6) Confirm — อ่านไฟล์ซ้ำแล้ว echo ตารางสุดท้าย พร้อมแจ้ง fallback rule และว่า rerun ปลอดภัย (7) Offer verification skill — เช็คว่าโปรเจกต์มี verify-* skill หรือ harness อยู่ไหม ถ้าไม่มีเสนอสร้างผ่าน /create-verification-skill หนึ่งครั้ง ไม่ต้องดันถ้าปฏิเสธ

💬 ตัวอย่าง prompt

/setup-pstack — ขอให้ how critics ใช้ code-reviewer, poteto-agent, general-purpose และ swarm workers ใช้ poteto-agent

⚡ Auto-invoke

auto-invoke ได้ — ไม่มี disable-model-invocation ใน frontmatter โมเดลจึงหยิบไปใช้เองได้เมื่อ context เข้าที่ เช่น ตอน user เพิ่งติดตั้ง pstack และเริ่มใช้ครั้งแรก, เมื่อ user พูดถึงการ configure subagent, หรือเมื่อผลลัพธ์ delegation บ่งชี้ว่ายังไม่เคยตั้งค่า role อย่างไรก็ตามมันเป็น setup skill ความถี่ต่ำ ไม่ได้ถูก chain เป็น step ของ /poteto-mode — poteto-mode เพียงอ้างว่า per-role lines ใน ~/.zcode/pstack-roles.md ที่ /setup-pstack เขียน จะ override ค่า default และ type choice ใน routed skills (how, why, arena, swarm, architect, interrogate, reflect)

🔗 คู่กันกับ

ควรรันหลังติดตั้ง pstack ก่อน /poteto-mode ครั้งแรก เพราะกำหนด subagent type ให้ role ทั้งหมดที่ poteto-mode และ routed skills (how, why, arena, swarm, architect, interrogate, reflect) ใช้ spawn งาน และใน step สุดท้ายจะเสนอ /create-verification-skill ต่อเองถ้าโปรเจกต์ไม่มี verification skill

⚠️ เคล็ดลับ/ข้อควรระวัง

(1) ห้ามเขียน subagent type ที่ยังไม่ได้ยืนยันว่ามีจริง — บรรทัด role ที่ชี้ type ที่ Agent tool ปฏิเสธ จะพังทุก delegation ที่อ่านไฟล์นั้น (2) ลบบรรทัดไหนออกจาก pstack-roles.md = role นั้นกลับไปใช้ skill default ไม่ต้องแก้ skill (3) rerun /setup-pstack ปลอดภัยเสมอเพราะ overwrite ทั้งไฟล์ (4) อย่าลืมว่าบน ZCode ตั้งได้แค่ type ไม่ใช่ model — อย่าคาดหวัง per-role model choice (5) panel role เป็น list: ซ้ำ type เดิมหลายรอบได้ถ้าต้องการ fan-out เยอะ, arena cross-judge pool เป็น list ที่ Arena เลือกมา 1 ค่า, swarm workers เป็น default type ของ worker ทุกตัวเว้นแต่ race/comparison จะกำหนดเอง (6) จุดที่คนพลาด: คิดว่าต้องตั้งค่าก่อนถึงใช้ pstack ได้ — จริง ๆ ไม่ตั้งก็ใช้ได้ เพราะทุก skill มี inline default ของตัวเอง

❓ ยังไม่ยืนยัน
  • auto_invoke
👤 เรียกเอง/broSkill จิ๋ว (267 bytes) แบบ one-shot ของ pstack: สั่งให้ agent พูดซ้ำข้อความล่าสุดของตัวเองใหม่เป็นภาษามนุษย์ธ…
📦 ทำอะไร

Skill จิ๋ว (267 bytes) แบบ one-shot ของ pstack: สั่งให้ agent พูดซ้ำข้อความล่าสุดของตัวเองใหม่เป็นภาษามนุษย์ธรรมดา ตัด jargon ทิ้ง เล่าให้กระชับและเข้าใจง่ายเหมือนคนสองคนคุยกัน ไม่มี logic หรือขั้นตอนใด ๆ ทั้งไฟล์

💡 เจ๋งยังไง

ความเจ๋งคือความเรียบ: เป็น 'ปุ่ม reset น้ำเสียง' ที่แก้ failure mode ที่พบบ่อยที่สุดของ agent คือตอบยาว เทคนิคัล และอ่านไม่รู้เรื่อง ทั้งที่ตัวไฟล์ไม่มีอะไรเลยนอกจากคำสั่งตรง ๆ หนึ่งประโยค (Restate your last message. Stop using jargon and speak coherently.) — เป็นตัวอย่างว่า skill ไม่จำเป็นต้องซับซ้อนจึงมีประโยชน์ และเข้ากับปรัชญา anti-slop ของ pstack ที่เห็นเต็ม ๆ ใน poteto-mode แต่ถูกหั่นมาให้เรียกใช้แยกได้ทันทีเมื่อคำตอบก่อนหน้าอ่านแล้วปวดหัว

🧭 วิธีใช้

เรียกใช้: พิมพ์ /bro ทันทีหลังจากได้รับคำตอบที่เต็มไปด้วย jargon หรืออธิบายซับซ้อนเกินไป ไม่มีพารามิเตอร์และไม่มีโหมดย่อย — agent จะ restate ข้อความล่าสุดของตัวเองใหม่ทั้งก้อน ใช้เมื่อไหร่: หลังคำตอบเทคนิคแน่น ๆ เช่น ผลจาก investigation ยาว ๆ หรืออธิบายสถาปัตยกรรมที่ใช้ศัพท์เยอะ และต้องการย่อเป็นภาษาพูดให้คนไม่ใช่ engineer ฟังเข้าใจ

💬 ตัวอย่าง prompt

/bro

⚡ Auto-invoke

ไม่ auto-invoke — frontmatter มี disable-model-invocation: true จึงเป็น user-invoked command เท่านั้น โมเดลจะไม่หยิบไปใช้เองและจะไม่ถูก chain โดย poteto-mode; เจตนาคือให้ user เป็นคนกดเมื่อรู้สึกว่าคำตอบล่าสุดยังอ่านยากอยู่

🔗 คู่กันกับ

standalone ไม่ผูกกับ skill ใด ใช้ต่อท้าย reply ใดก็ได้ และมักสมเหตุสมผลหลัง output หนา ๆ ของ pstack เช่น คำตอบจาก Investigation playbook ของ /poteto-mode

⚠️ เคล็ดลับ/ข้อควรระวัง

มัน restate เฉพาะ 'ข้อความล่าสุด' เท่านั้น ดังนั้นต้องพิมพ์ /bro ทันทีหลัง reply ที่อยากให้ย่อ ถ้าคุยต่อไปอีกสองสามเทริดแล้วค่อยเรียก มันจะย่อข้อความล่าสุดแทนที่จะเป็นก้อนที่ตั้งใจ; ไม่มี parameter ไว้ชี้ไปที่ message ก่อนหน้า

🔍 Understand & Explore (7)

✨ auto/howSkill สำรวจ codebase เพื่อตอบคำถาม "how does X work" — ผลิตคำอธิบายสถาปัตยกรรมระดับที่ senior engineer onboar…
📦 ทำอะไร

Skill สำรวจ codebase เพื่อตอบคำถาม "how does X work" — ผลิตคำอธิบายสถาปัตยกรรมระดับที่ senior engineer onboarding เข้า subsystem ใหม่จะได้ mental model ที่ใช้งานได้จริง (ไม่ใช่ annotated source code) มี 2 โหมด: Explain (default) สำรวจแล้วอธิบายอย่างเดียว และ Critique ที่อธิบายให้เข้าใจก่อน แล้วค่อยปล่อย critic หลายตัวแบบอิสระออกมาหาปัญหาเชิงสถาปัตยกรรม

💡 เจ๋งยังไง

กลไกที่เจ๋งคือการประเมินความซับซ้อนของคำถามก่อนเสียทรัพยากร: คำถาม Simple (โมดูลเดียว, คำถามแคบ) ส่ง explainer ตัวเดียวทำ explore+explain ใน pass เดียว, คำถาม Complex (ข้ามไฟล์/services) แตกเป็น 2-4 explorer agents ขนานกัน แต่ละตัวรับ slice ไม่ซ้ำกัน (เช่น data model / request path / config+metrics) แล้วให้ explainer ตัวเดียว synthesize + reconcile สิ่งที่ทับซ้อน โดยมีกติกา "When in doubt, lean simple" กันการ over-engineer นอกจากนี้ Critique mode บังคับ "Explain First" — ห้ามวิจารณ์โค้ดก่อนเข้าใจมัน และ critic แต่ละตัวได้คำอธิบายจากขั้นก่อน + รูปแบบเดียวกันจาก critique-rubric ทำให้วิจารณ์ตั้งอยู่บนความเข้าใจเดียวกัน สุดท้าย lead judgment แบ่งผลเป็น Act on / Consider / Noted / Dismissed แบบ "pragmatic lead, not an aggregator" และกฎ "The explainer's communication is the product" ห้าม lead rewrite คำอธิบายจนเสียรสของผู้เขียน

🧭 วิธีใช้

เรียกใช้ด้วย /how หรือถามตรง ๆ "how does X work" — ใช้เมื่อ: จะ walkthrough โค้ดก่อนแก้อะไร, ถาม placement/ownership/layering ("where should this live", "which package owns this", "is this the right layer"), อยากได้ subsystem architecture / runtime flow / onboarding mental model. ขั้นตอนภายใน: (1) parse คำถาม (subsystem / feature flow / architectural overview / runtime trace) — ถ้ากำกวม state best-guess แล้วลุย ห้ามถามกลับ (2a) คำถาม complex → spawn 2-4 Explore agents (read-only) พร้อมกันใน message เดียว แต่ละตัวได้ base prompt จาก references/explorer-prompt.md + angle เฉพาะ โดย explorer ต้องไล่ call chain จนอธิบาย input-to-output ได้ครบไม่มี hand-waving (2b) คำถาม simple → spawn general-purpose explainer ตัวเดียว (3) complex case: ส่ง findings รวมให้ explainer synthesize ตาม references/explainer-prompt.md (4) present ตามโครง: Overview → Key Concepts → How It Works → Where Things Live → Gotchas. โหมดย่อย Critique: trigger เมื่อผู้ใช้ขอ "architectural issues/problems/improvements" โดยรัน explain ครบก่อน แล้ว spawn critic หนึ่งตัวต่อ entry ใน configured how-critics list (default: code-architect, code-reviewer, poteto-agent, general-purpose) ทุกตัวได้คำอธิบาย + file paths + rubric จาก references/critique-rubric.md จบด้วย lead judgment แบ่ง Act on / Consider / Noted / Dismissed โดยนำเสนอคำอธิบายก่อน แล้ววาง verdict วิจารณ์ไว้ข้างล่าง

💬 ตัวอย่าง prompt

/how rate limiter ของเราทำงานยังไง — เดี๋ยวผมจะเพิ่ม per-tenant quota ในส่วนนี้

⚡ Auto-invoke

ได้ — frontmatter ไม่มี disable-model-invocation จึงเป็น 1 ใน 2 skill ของกลุ่ม understand (คู่กับ why) ที่ออกแบบให้ ZCode agent จุดเองได้เมื่อ context ร้องขอ เช่น ผู้ใช้ถาม how/structure/layer หรือ poteto-mode route เข้ามาตาม trigger "Nontrivial change, architecture decision, or 'are we sure?' → the how skill"

🔗 คู่กันกับ

คู่ตรงข้ามกับ why: how ตอบ what/how ส่วน why ตอบแรงจูงใจ (description ระบุ "Use why for motivation") — ไม่มี requirement ว่าต้องเรียก skill อื่นก่อน; Critique mode ใช้ lead-judgment framework เดียวกับ interrogate และใช้ critic types เดียวกับที่ interrogate ใช้ (code-reviewer, code-architect, poteto-agent, general-purpose)

⚠️ เคล็ดลับ/ข้อควรระวัง
  • ถ้าคำถามกำกวม ห้ามถามกลับ — state best-guess interpretation แล้วเริ่มสำรวจ ให้ user redirect เอาทีหลัง ("Don't ask. Let the user redirect if you're off.")
  • เมื่อไม่แน่ใจให้เอียงไปทาง simple ก่อน — spawn explorers เพิ่มได้ตลอดถ้า explainer เจอ wall
  • explorer ต้องอ่านโค้ดจริง ห้ามเดาจากชื่อไฟล์; ให้ overlap ระหว่าง explorer ได้ — explainer เป็นคน reconcile
  • ตอน present แก้เล็กน้อยเพื่อความชัดได้ แต่ห้าม rewrite ใหญ่ — "The explainer's communication is the product"
  • Critique mode ต้องรัน explain ครบก่อนเสมอ และเรียง output เป็น explanation ก่อน แล้ว critique verdict ทีหลัง เพื่อให้คนที่แค่อยากเข้าใจระบบไม่ต้องไล่อ่านวิจารณ์
  • output format ไม่ต้องใช้ครบทุก section ทุกครั้ง — เลือกให้เหมาะกับคำถาม
✨ auto/whySkill สืบสวนแรงจูงใจและ intent ที่อยู่เบื้องหลังโค้ด — ทำไมถึงสร้างแบบนี้, พิจารณา edge case อะไร, product/bu…
📦 ทำอะไร

Skill สืบสวนแรงจูงใจและ intent ที่อยู่เบื้องหลังโค้ด — ทำไมถึงสร้างแบบนี้, พิจารณา edge case อะไร, product/business/operational constraint ไหนหล่อหลอม design, alternative ไหนถูกปฏิเสธและเพราะอะไร ทำงานโดย enumerate MCP ที่มีใน session แล้ว map เข้า 7 evidence categories (source control, issue tracker, long-form docs, real-time chat, infra observability, error tracking, product analytics warehouse) ไล่ query ขนานกันทุก category แล้ว synthesize เป็น narrative ที่อ้าง citation ทุกข้อความพร้อม confidence calibration ชัดเจน

💡 เจ๋งยังไง

หัวใจคือกลไก "coverage ไม่ใช่ minimalism": ไม่เดาจากคำถามว่าคำตอบซ่อนอยู่ category ไหน แต่ query ทุก category ที่มี MCP ขนานกัน และยก null result เป็น evidence ชั้นหนึ่ง ("a null result from an issue tracker is evidence the decision was not ticketed") — กัน blind spot จากการตัดสินล่วงหน้า ประกอบกับ epistemics framework ที่เข้มงวด: "Evidence before narrative", "Cite everything — ถ้าอ้างไม่ได้คือ inference ไม่ใช่ fact", "Prefer 'appears to' over 'because'" และกฎ "No shortcut by code-reading" (โค้ดบอกว่าทำอะไร แต่แทบไม่เคยบอกว่าทำไม) โครงสร้าง output ก็บังคับความซื่อสัตย์โดย design ผ่านการแยก "What We Found (direct evidence)" / "What We Can Reasonably Infer" / "Competing Hypotheses" / "What We Don't Know" / "Sources Consulted" โดย Sources Consulted ต้องรายงานทุก investigator รวมตัวที่ว่างเปล่าและตัวที่ skip พร้อมเหตุผล ทำให้ผู้อ่านตัดสินความกว้างของการสืบค้นได้ทันที ส่วน failure modes ที่ระบุไว้ตรงจุดมาก: confident storytelling, อ้างโค้ดเป็นหลักฐาน intent ของตัวเอง, recency bias และ sycophantic agreement

🧭 วิธีใช้

เรียกด้วย /why หรือถามตรง ๆ "why does X work this way" — ใช้กับ design rationale, "why we picked Y", regressions, postmortems และ data-backed thresholds (เช่น "ตัวเลข limit นี้มาจากไหน") ขั้นตอนภายใน: (1) parse target + question (design rationale / tradeoff / defensive reasoning / ทำไมโค้ดนี้ยังอยู่ / history sweep) — target กำกวมให้เดาจากบริบทแล้วแจ้ง interpretation (2) สร้าง code anchor inline ก่อน spawn: git blame -L, git log --follow -p, git log --oneline -20 และดึง PR bodies/discussion ด้วย gh pr view เพื่อเก็บ PR numbers + ticket IDs เป็น seed context (3) enumerate MCP ทั้งหมด แล้ว map เข้า 7 evidence categories จากนั้น spawn investigator หนึ่งตัวต่อหนึ่ง category ขนานกันใน message เดียว (general-purpose เพราะต้องใช้ MCP tools; ห้ามรวมหลาย MCP ให้ agent เดียว) — ทุกตัวได้ base prompt จาก references/investigator-prompt.md + category playbook จาก references/sources/*.md; ถ้า target ดู defensive (null check, retry, timeout, rate limit, feature flag) ให้แนบ references/sources/incident-postmortem.md เพิ่ม (4) synthesizer ตัวเดียวรวมผลตาม epistemics framework (references/epistemics.md) (5) present โดยห้ามแตะ confidence language เลย. การ skip investigator ทำได้เฉพาะ 2 เหตุผลที่ต้องเขียนลง Sources Consulted: ไม่มี MCP ของ category นั้น หรือ source นั้น provably irrelevant (bar สูงมาก) — ถ้าคำตอบจะถูกใช้แก้โค้ดต่อ ให้แปลง findings เป็น Preserve / Change / Avoid / Risk constraint set

💬 ตัวอย่าง prompt

/why ทำไม retry logic ของเราถึงใช้ exponential backoff — และค่า timeout 128KB นี่มาจากไหน

⚡ Auto-invoke

ได้ — frontmatter ไม่มี disable-model-invocation จึงเป็น 1 ใน 2 skill ของกลุ่ม understand (คู่กับ how) ที่ออกแบบให้ ZCode agent จุดเองเมื่อคำถามถาม motivation/regression/postmortem และเป็น skill ที่ถูก chain บ่อย: teach เรียก why ขนานกับ how (โดยแนะนำให้ narrow scope เพราะ full sweep ช้า) และ recall ยืม source investigators + per-source playbooks ของ why ไป sweep shared record

🔗 คู่กันกับ

Companion ของ how (เนื้อไฟล์ระบุ "Companion to the how skill" — เป็นคู่เสริม ไม่ใช่ hard requirement): how ตอบ "what the code does and how it works", why ตอบ "what forces led to its shape"; เป็น dependency ของ teach (teach รัน how+why เป็น real skill invocations) และ recall (recall ใช้ investigator + playbooks ของ why ใน step 4)

⚠️ เคล็ดลับ/ข้อควรระวัง
  • ห้ามสรุป intent จากรูปโค้ด — "Handles the null case because it checks for null" คือ mechanics ไม่ใช่ motivation; motivation ต้องมาจากแหล่งภายนอก (PR discussion, ticket, comment) หรือถูก label ว่า inference
  • ห้าม skip investigator ด้วยการคาดเดา ("long-form docs probably don't have this" ไม่ผ่าน) — null result คือ data point, skipped search คือ blind spot
  • ห้ามรวมหลาย evidence category ให้ agent เดียว — แต่ละ MCP มี query vocabulary, result shape และ pitfall ต่างกัน
  • ถ้า user เดาเหตุผลมาเอง ("I assume this is for performance?") ให้ถือเป็น hypothesis ที่ต้องเช็คกับ evidence อิสระ ๆ — ห้ามเห็นด้วยเฉย ๆ (sycophancy)
  • recency bias: commit ล่าสุดไม่ใช่คำตอบ — รูปร่างปัจจุบันมักเป็นผลรวมของการตัดสินใจเก่า ๆ หลายชั้น ต้องไล่ย้อน
  • ตอน present ห้าม rewrite confidence language — การตัด hedge ("appears to", "likely") ทิ้งเพื่อความน่าเชื่อถือ คือ failure mode ที่ skill นี้มีตัวตนเพื่อกัน
  • full sweep ทุก category ช้าและแพง — ถ้าเหตุผลไม่ใช่ประเด็นหลักของงาน (เช่นถูกเรียกจาก teach) ให้ narrow ผ่านตัวคำถามเอง เพื่อให้ why บันทึก skipped categories ตาม contract ของมัน
👤 เรียกเอง/teachSkill อธิบายงานหรือ subsystem หนึ่ง ๆ อย่างเรียบง่ายจนคนอ่านเข้าใจจริง — รัน skill how และ why ให้ทำการสืบค้น…
📦 ทำอะไร

Skill อธิบายงานหรือ subsystem หนึ่ง ๆ อย่างเรียบง่ายจนคนอ่านเข้าใจจริง — รัน skill how และ why ให้ทำการสืบค้นเบื้องหลัง แล้วถักสิ่งที่ได้มาเป็นคำอธิบายเดียวที่ปรับตามจังหวะของผู้ฟัง เป้าหมายคือให้เข้าใจ ไม่ใช่แก้ไขอะไร ใช้กับ "teach me this", "help me really understand X", "explain this change or subsystem to me"

💡 เจ๋งยังไง

สถาปัตยกรรมของมันคือการแยกบทบาทให้ขาด: การสืบค้นเป็นของ how + why ทั้งคู่ ("Those are real skill invocations that do their own digging... Let those skills do the investigation. Don't redo it by hand") teach เป็นเพียงชั้นการสื่อสารที่ blend ผลลัพธ์ ทำให้คุณภาพการสืบค้นได้มาตรฐานเดียวกับ skill จริงและไม่ duplicate งาน ด้านการสอนมีกลไกต่อต้าน lecture-theater ที่เจาะจงมาก: อ่านว่าผู้ฟังถามเพราะอะไร (กำลังจะแก้โค้ด / กำลังรีวิว / debug / มือใหม่) จากบทสนทนาแทนการสอบถาม, "Put the depth where their question is", ห้าม print framing labels ("the key insight", "TL;DR", "Pause") และห้าม quiz ที่สำคัญคือกฎ diagram ที่ซื่อสัตย์ต่อ cognitive load — อะไรที่มี 3+ ส่วนต้องห้ามวาดรวดเดียว ให้วาดซ้ำเป็นชุดทีละส่วนจนผู้อ่านเห็นระบบประกอบขึ้น ("Three small growing diagrams beat one crowded diagram") และคุมคุณภาพภาษาด้วย unslop skill ทั้งฉบับ รวมถึงตัวอย่างความหนาแน่นที่ยกมาเป็นตัวอย่าง (virtualization อธิบายเป็นสองส่วน rendering/loading)

🧭 วิธีใช้

เรียกด้วย /teach (มี disable-model-invocation: true ใน frontmatter) ขั้นตอนภายใน: (1) ตัดสินว่าผู้ฟังควร walk away ด้วยความเข้าใจอะไรบ้าง โดยอ่านเจตนา (จะแก้/รีวิว/debug/มือใหม่) และพื้นฐานเดิมของเขาจากบทสนทนา ข้ามสิ่งที่เขารู้แล้วชัดเจน (2) อ่านโค้ดเองเพื่อ orient ก่อน แล้วรัน how (ทำงานยังไง) กับ why (ทำไมเป็นแบบนี้) ขนานกันแล้ว combine — subsystem ใหญ่รันทั้งคู่ เปลี่ยนเล็กอาจตัวเดียวพอ; ให้ keep why narrow โดย default เพราะ full sweep ของ why ช้า โดยใส่การ narrow ไว้ในตัวคำถามที่เรียก (scoped question, git + source คู่) เพื่อให้ why บันทึก skipped categories ตาม contract ของมันเอง (3) เริ่มด้วย plain definition — ตั้งชื่อสิ่งนั้นและบอกว่ามันคืออะไรแบบที่ senior engineer พูดปากเปล่า แล้วผูกกับกรณีตรงหน้า ต่อด้วย how it works → deeper reasons → edge cases; ให้คำตอบที่สมบูรณ์ที่สุดที่เล็กที่สุดก่อน (1-2 ประโยค) แล้วหยุด เพิ่ม layer เมื่อถูกถาม (4) รักษาบรรยากาศบทสนทนา เสนอเจาะลึกหรือย้ายประเด็น ตามที่ผู้ฟังนำ; ไม่มี quiz ไม่มี pacing theater (5) show don't tell — เปิด diff/โค้ด/debugger เมื่อเร็วกว่า และวาด diagram แบบ build-up ทีละชั้น (flow ที่ผ่าน A→B→C วาดสามรอบ: A→B ก่อน แล้ว redraw เพิ่ม C แล้ว redraw เพิ่ม return edge); ไอเดียเชิง spatial (layout, overlap, scroll position, before/after) ใช้ image generation สไตล์ marker-on-whiteboard พร้อม label สั้น ๆ แทน mermaid

💬 ตัวอย่าง prompt

/teach ขอเข้าใจ virtualization ใน chat list ของเราให้ลึกจริง ๆ — เดี๋ยวผมต้องแก้ scroll bug ในส่วนนี้

⚡ Auto-invoke

ออกแบบให้ไม่ auto-invoke: frontmatter มี disable-model-invocation: true — เจตนาคือผู้ใช้เรียกเองชัดเจน หรือ agent เรียกตามที่ playbook/workflow กำหนด ไม่ใช่ skill ที่ agent จุดเองจากการเดา context (ต่างจาก how/why ที่ไม่มี flag นี้) ข้อสังเกต: ใน ZCode flag นี้อาจถูก ignore เพราะเป็น unknown field แต่ design intent ถือเป็นตัวตั้ง [uncertain]

🔗 คู่กันกับ

ต้องมี how + why — ไฟล์ระบุชัด "Teach sits on top of how and why" และให้รันทั้งคู่เป็น real skill invocations ขนานกัน (โดยยืดหยุ่นได้: เปลี่ยนเล็กอาจเหลือตัวเดียวพอ); คำตอบต้องเขียนผ่าน unslop skill; ต้องคง confidence language ของ why ไว้ครบเวลา reword เพื่อการสอน

⚠️ เคล็ดลับ/ข้อควรระวัง
  • ข้อยกเว้นเดียวของกฎ "reword freely for teaching" คือห้ามแตะ hedge ของ why — "its hedges are findings, not style"
  • ห้าม lecture/performance: ห้าม quiz, ห้าม print "Pause", ห้ามให้ผู้ฟังพูดตาม, ห้าม flag ว่า "ส่วนนี้ยากนะ" / "the key insight" — อยากให้เขาหยุดคือหยุดพูดแล้วปล่อยให้เขาตอบ
  • การไล่ list ฟังก์ชันและค่าคงที่คือ reference ไม่ใช่การสอน — อธิบาย concrete mechanism แทน
  • ห้าม wall of text — ประโยคสั้น 1-2 comma ต่อประโยค, ให้ concept หนึ่งมีชื่อเดียวตลอด, ห้าม em dash, ห้าม mirror sentences และ tidy closers
  • ห้าม echo scaffolding — หัวข้อและกฎใน steps เป็นคำสั่งให้ agent ทำ ไม่ใช่ label ให้ print ออกมา
  • reply ต้องเป็นคำอธิบายเอง ไม่ใช่รายงานว่า "ทำอะไรไปแล้ว"; ตอน one-shot ไม่มีคน live ให้ส่งมอบเนียน ๆ แล้ววางข้อเสนอเจาะลึกไว้ท้าย
  • diagram ชิ้นเดียวรวมทุกส่วน (โดยเฉพาะที่โยนไว้ท้าย) คือ reference ไม่ใช่การสอน — ใช้ชุด diagram ที่โตทีละส่วนแทน
❓ ยังไม่ยืนยัน
  • auto_invoke
👤 เรียกเอง/recallSkill สร้าง working context ล่าสุดของผู้ใช้ขึ้นใหม่ก่อนเริ่มหรือกลับมาทำงาน โดยขุดจากสองแหล่ง: chat history ข…
📦 ทำอะไร

Skill สร้าง working context ล่าสุดของผู้ใช้ขึ้นใหม่ก่อนเริ่มหรือกลับมาทำงาน โดยขุดจากสองแหล่ง: chat history ของตัวเอง (session transcripts ใต้ ~/.zcode/cli/, JSONL หนึ่งไฟล์ต่อหนึ่ง session) และ shared record (symptom ที่ user รายงาน, fix ที่ ship แล้วถูก revert, error ที่ยังยิงใน prod) แล้วส่งมอบ brief กระชับ: capsule สถานะรวม, สถานะราย thread, ปัญหาที่พบซ้ำ และ next move เดียวที่ควรทำต่อ

💡 เจ๋งยังไง

Insight หลักคือการยอมรับว่า transcript ของตัวเองไม่พอ: "A feature with a long bug tail keeps most of its story there, so don't reconstruct it from your transcripts alone" — เรื่องราวส่วนใหญ่ของ feature อยู่ใน shared record ซึ่งเป็นสิ่งที่ why skill ค้นอยู่แล้ว recall จึงยืมกลไกของ why มาใช้แทนการสร้างใหม่ (reuse per-source playbooks, one investigator per source, null results are findings) อีกจุดเด่นคือ context budget ที่ออกแบบดี: heavy reading กระจายไป parallel subagents บน fast cheap model และ "The raw transcripts stay in the subagents. The main thread gets only their findings" รวมถึงการ verify กับ live state เพราะ "A transcript or a stale ticket is history, not current truth" ต้องเช็ค PR/branch/ticket กับ git/gh จริง ส่วน output contract บังคับ status tag ต่อ thread ([merged #N], [open PR #N], [in flight <branch>], [verified, uncommitted], [reverted #N], [planned, not started]) — "A thread with no tag is not done yet, so tag it" และต้องรายงาน fix ที่โดน revert เพื่อให้ attempt ถัดไปเริ่มจากจุดที่ attempt ก่อนล้ม

🧭 วิธีใช้

เรียกด้วย /recall (มี disable-model-invocation: true) — ใช้กับ "recall my work on X", "catch me up", "what have I been working on", "where did I leave off" ก่อนเริ่มหรือ resume งาน ขั้นตอนภายใน: (1) classify ก่อน route: resume จาก chat เดียวจำเพาะคือ playbook session-pickup ไม่ใช่ skill นี้, แปลง habit เป็น skill คือ automate-me, ถ้า user ให้ state capsule มาแล้ว (paths, branch, change) ให้ใช้เลยไม่ต้องขุด (2) lock scope ก่อนค้น: หน้าต่างเวลา (default 7 วัน), topic ถ้าระบุ, workspace (default อันปัจจุบัน — ห้ามอ่าน transcript project อื่นถ้าไม่ได้ร้องขอ) แล้ว state scope กลับให้ user เห็น (3) fan-out ขนานล่า transcript ด้วย subagents บน fast cheap model: จัดลำดับ candidate ด้วย real modification time (ls -t) ไม่ใช่ชื่อ UUID, grep topic ก่อนแล้วอ่านเฉพาะ chat ที่ match เฉพาะช่วงที่เกี่ยว, ข้าม current chat + subagent/eval/test chats; ถ้าเหลือ 1-2 chats ค้นตรงไม่ต้อง fan-out; ทุก subagent รายงาน schema เดียวกัน (topic, goal, decisions, open threads, struggles/corrections, artifacts) พร้อม cite chat UUID (4) sweep shared record ทุกครั้งที่ topic ระบุ feature/file/subsystem/area/bug — เป็น default ไม่ใช่ judgment call: ส่งให้ why skill's source investigators แต่เปลี่ยนคำถามจาก "why was this built" เป็น "สถานะปัจจุบันคืออะไร, เคยลองอะไรที่ไม่อยู่, user ยัง report อะไร" ขนานกับการล่า transcript; ข้ามได้เฉพาะ pure activity recall ที่ไม่มี target ("what did I do this week") (5) verify กับ live state ด้วย git/gh (6) เขียน brief ตาม output contract: Capsule ≤5 bullets, Threads บรรทัดละหนึ่งพร้อม status tag บังคับ, Problems ≤5 อันที่เกิดซ้ำ, Next move หนึ่งอย่างที่ concrete

💬 ตัวอย่าง prompt

/recall payment retries — ผมทำถึงไหนแล้ว เดี๋ยวขอ resume ต่อ

⚡ Auto-invoke

ออกแบบให้ไม่ auto-invoke: frontmatter มี disable-model-invocation: true — เจตนาคือผู้ใช้พิมพ์เรียกเองเมื่อจะเริ่ม/resume งาน; poteto-mode ไม่มี trigger route เข้า skill นี้โดยตรง (งาน resume จาก agent ที่ทำค้างมี playbook แยกคือ session-pickup ซึ่ง recall ขอให้แยกขาดจากกัน) ข้อสังเกต: ใน ZCode flag นี้อาจถูก ignore เพราะเป็น unknown field แต่ design intent ถือเป็นตัวตั้ง [uncertain]

🔗 คู่กันกับ

ยืม source investigators + per-source playbooks จาก why skill ใน step 4 (sweep shared record) แต่เปลี่ยนคำถาม; เขียน brief ผ่าน unslop skill; แยกขาดจาก playbook session-pickup (resume chat เดียวจำเพาะ) และ automate-me (แปลง habit เป็น skill)

⚠️ เคล็ดลับ/ข้อควรระวัง
  • ห้ามขยาย scope เงียบ ๆ — "Never quietly turn 'all' into 'recent N'" ต้อง state scope กลับให้ user ยืนยันก่อน
  • จัดลำดับ transcript candidate ด้วย modification time เท่านั้น ห้ามใช้ชื่อ UUID เพราะชื่อไฟล์ session ไม่สะท้อนเวลา
  • ประวัติไม่ใช่ความจริงปัจจุบัน — PR/branch/ticket ที่ได้มาต้อง verify กับ git/gh ก่อนเขียนลง brief และถ้าคำตอบแกว่งที่สิ่งที่ agent เคยทำจริง ให้อ่าน full transcript ไม่ใช่ copy ที่ตัดแล้ว
  • แยก noise ออก: current chat, subagent chats, eval chats, test chats ไม่นับ
  • ตัด detail ก่อนตัด threads เมื่อ capsule เริ่มเกินจอ; feature/ticket ข้างเคียงไม่ใส่เว้นแต่ block งานหลัก
  • cite แหล่ง: chat findings อ้างด้วย UUID, shared-record findings อ้างด้วย source (PR #, ticket ID, chat permalink, error-tracker issue)
  • sanitize private context ก่อน output ใด ๆ ที่เป็นสาธารณะ
❓ ยังไม่ยืนยัน
  • auto_invoke
👤 เรียกเอง/figure-it-outSkill ออกแบบ playbook แบบ auditable ขึ้นใหม่เมื่อไม่มี playbook ที่แคบกว่าให้พอดี — งานพวก migration ขนาดใหญ่…
📦 ทำอะไร

Skill ออกแบบ playbook แบบ auditable ขึ้นใหม่เมื่อไม่มี playbook ที่แคบกว่าให้พอดี — งานพวก migration ขนาดใหญ่, งาน multi-part ที่ทะเยอทะยาน หรืองานที่มนุษย์จะรีวิวหลังจากหายตัวไป ตัว deliverable ก่อนโค้ดแรกคือ workflow เอง: ลำดับ phase ที่ scale rigor ตามความเสี่ยง รัน scientific method (hypothesis loop) และทิ้ง decision trail ที่มนุษย์ audit ย้อนได้

💡 เจ๋งยังไง

มุมมองที่คมคือมอง "งานที่ไม่มี playbook" เป็นงาน design workflow ไม่ใช่โอกาส improvise — "The deliverable before any code is the workflow itself" พร้อมหลัก "Bias toward more rigor. The cost of building the wrong thing dwarfs the cost of being careful" และ rigor ถูกทำให้เป็นรูปธรรม: "Rigor is gates and artifacts, not 'try harder'" — one-way door กับ high blast radius ได้มากขึ้น, reversible low-stakes ได้น้อยลง จุดเด่นกลางคือ Phase C hypothesis loop ที่จริงจัง: ทุก unit เป็น experiment (state hypothesis → smallest change → measure กับ predicate บน real artifact → keep ถ้า advance, revert ถ้าไม่) โดย verdict มีแค่ VERIFIED / NOT VERIFIED / INCONCLUSIVE และ "Inconclusive is not a pass. Don't hide a negative" กลไกกัน fake-verification ตรงจุดมาก: "Verify by inspecting the artifact, never a self-report. When something passes too easily, suspect the observation method before the system. A blank screenshot passes a lazy gate" และ delegated work ต้องจับคู่ judge + audit artifact ด้วยตัวเองก่อนเชื่อ worker สุดท้าย มันประกอบร่างจาก principle skills ของ poteto-mode อย่างเป็นระบบ (prove-it-works, never-block-on-the-human, foundational-thinking, laziness-protocol, separate-before-serializing-shared-state, sequence-verifiable-units, encode-lessons-in-structure) โดยเปิดท้ายด้วยการ encode การแก้ซ้ำเป็น gate/lint/script เพื่อให้ชัยชนะไม่ถูกกลืนกลับ

🧭 วิธีใช้

เรียกด้วย /figure-it-out หรือพูด "figure it out" (มี disable-model-invocation: true) — ใช้เมื่อ: migration ข้าม call sites จำนวนมาก, งาน multi-part ทะเยอทะยาน, งานที่ user รีวิวหลัง step away หรือเมื่อไม่มี playbook ไหนให้พอดี ขั้นตอนภายใน: Start — เปิด todolist ที่ item แรกคืออ่านส่วน Principles ของ poteto-mode แล้วใส่ทุก phase เป็น todos Phase A Frame — นิยาม definition of done เป็น falsifiable predicate (prove-it-works), quantify scope + ปลุก blockers ก่อนเสียเวลาหลายชั่วโมง, เลือก rigor level เอียงสูง แล้ว present framing/tradeoffs ก่อน commit รันยาว (งาน reversible เดินต่อเองตาม never-block-on-the-human แต่ run หลายชั่วโมงได้ checkpoint เดียว) Phase B Design — แตกงานเป็น atomic units ที่ land แยกได้ เรียง riskiest-unknown-first, สร้าง verification harness + baseline จาก pre-change state ก่อนแก้ (ผลอ่านเป็น "old vs new"), one-way-door design ให้รัน architect skill (ที่รัน arena ด้วย candidates หลากหลาย + read-only judge), ตัดสินว่าอะไร fan out ขนานโดยให้ worker แต่ละตัวมี worktree/branch ของตัวเอง, เขียน phase list ลง todolist — นั่นคือสิ่งที่มนุษย์จะรีวิว Phase C Run — ทุก unit เป็น experiment ใน hypothesis loop, verify ทีละ unit ก่อน unit ถัดไป (ไม่ batch ท้าย), delegated work จับคู่ judge และ audit ของ delegate ด้วยตัวเอง; worker เล่นกล gate → reset และ harden contract, gate เองผิด → fix gate เป็น change แยก Phase D — log ผ่าน show-me-your-work (TSV แถวต่อ decision/unit, เขียน row ไปพร้อมกับงานไม่ใช่กองท้าย) Phase E — ตรวจทั้งงานกับ predicate บน real product ไม่ใช่แค่ harness แล้ว encode recurring correction เป็น gate/lint/script; reply: playbook ที่ออกแบบ, rigor level + เหตุผล, path ของ decision trail, สิ่งที่ verified กับ predicate และสิ่งที่ยังเปิด

💬 ตัวอย่าง prompt

/figure-it-out — ย้ายทั้ง repo จาก REST client เก่าไป client ใหม่ทุก call site เดี๋ยวเย็นผมกลับมารีวิว

⚡ Auto-invoke

ออกแบบให้ไม่ auto-invoke ในงานธรรมดา: frontmatter มี disable-model-invocation: true — ผู้ใช้เรียกเองชัดเจน หรือถูก poteto-mode route เข้ามาตาม trigger "A large or cross-cutting effort ... routes to the figure-it-out skill even when a narrower playbook like Feature fits. Use figure-it-out whenever no bundled playbook fits" ข้อสังเกต: ใน ZCode flag นี้อาจถูก ignore เพราะเป็น unknown field แต่ design intent ถือเป็นตัวตั้ง [uncertain]

🔗 คู่กันกับ

พึ่ง poteto-mode เป็นฐาน (todo แรกบังคับคืออ่าน Principles section ของ poteto-mode); ใช้ show-me-your-work เป็น audit trail ใน Phase D; ใช้ architect skill (ซึ่งรัน arena) สำหรับ one-way-door design decision; อ้าง principle skills ตรง ๆ: prove-it-works, never-block-on-the-human, foundational-thinking, laziness-protocol, separate-before-serializing-shared-state, sequence-verifiable-units, encode-lessons-in-structure; ขอบเขต: ต่างจาก Orchestrate playbook — figure-it-out ออกแบบ run เดียว (one bespoke run) ส่วน Orchestrate รันโปรแกรมระดับโปรเจกต์

⚠️ เคล็ดลับ/ข้อควรระวัง
  • อย่าใช้ทับทาน playbook ที่มีอยู่ — งาน single-unit ที่ match Bug fix / Perf / Feature / Visual parity / Eval / Multi-phase plan ให้ไปที่นั่น; แต่ "เวอร์ชันใหญ่ข้ามส่วน" ของงานเหล่านั้น (migration ข้าม call sites เยอะ) อยู่ที่นี่
  • อย่า over-engineer: "A second arena over a settled design is over-engineering" — ข้าม architect/arena สำหรับงาน mechanical ที่รูปร่างชัดแล้ว
  • เมื่อ pass ง่ายเกินไป ให้สงสัยวิธีวัดก่อนระบบ — gate ที่หลวม (เช่น blank screenshot ผ่าน) คือสัญญาณเตือน ไม่ใช่ชัยชนะ
  • ห้ามซ่อน negative: INCONCLUSIVE ไม่ใช่ pass และห้ามยอมให้เป็น pass
  • baseline ของ harness ต้อง capture จาก pre-change state เพื่อให้การเช็คอ่านเป็น old value vs new value
  • นำ blockers ขึ้นก่อนเริ่ม run ยาว ไม่ใช่หลัง 50 commits ที่ดวลไม่ได้
  • commit decision trail เมื่อ confidence ต้องแสดงให้ reviewer เห็น — ปกติ log เป็น local artifact เท่านั้น (ดู show-me-your-work)
❓ ยังไม่ยืนยัน
  • auto_invoke
👤 เรียกเอง/show-me-your-workSkill เก็บ decision trail ที่รีวิวย้อนได้สำหรับงานรันยาวหรือ unattended — log TSV ไฟล์เดียว canonical แถวต่อห…
📦 ทำอะไร

Skill เก็บ decision trail ที่รีวิวย้อนได้สำหรับงานรันยาวหรือ unattended — log TSV ไฟล์เดียว canonical แถวต่อหนึ่ง decision (ts, phase, decision, why, evidence, result) local เป็น default และ commit เมื่อ reviewer ต้องใช้ trail นี้เพื่อเชื่อผล จบทุก run ด้วย self-audit กับ transcript จริงและการ cross-model review ที่ผลลัพธ์ต้องปรากฏเป็น section "Attention" ในทุก reply ของ run นั้น

💡 เจ๋งยังไง

Format ถูกเลือกด้วยเหตุผลเชิงปฏิบัติครบ: TSV เพราะ GitHub render เป็น sortable table, column -s$'\t' -t อ่านได้ใน terminal, spreadsheet เปิดได้ และ append ได้ด้วยคำสั่งเดียว — evidence ต้องเป็น pointer (commit SHA, PR #, file:line) ไม่ใช่ย่อหน้า จุดที่คมที่สุดคือวินัยเชิงประวัติศาสตร์ของ log: append-only, "A wrong call gets a new row that supersedes it. Never edit or delete history" และตอนจบ run ต้องเปิด transcript จริงมา walk เทียบ — "Every row maps to a real action. Cut invented or aspirational entries" พร้อมคติ "Fix the log, not the story" อีกกลไกที่เกินตัวคือ cross-model review แบบบังคับ: ต้อง spawn subagent คนละ model family มา scan trail + transcript เพื่อหา row ที่ evidence อ่อน, verification ที่อ้างไว้แต่ transcript ไม่มีหลักฐาน และ decision ที่ risky ย้อนหลัง โดยทุก reply จบด้วย Attention section ที่นำด้วย "reviewed by <model>" — "No flags" เป็นค่าที่ valid แต่การไม่ใส่ชื่อ model ไม่ใช่ นอกจากนี้มันถูก design ให้เป็น composable audit trail: skill อื่น (figure-it-out) route มาที่นี่แทนการ invent format ใหม่ ("Reference it by name and let it own the format")

🧭 วิธีใช้

เรียกด้วย /show-me-your-work (มี disable-model-invocation: true) — ใช้กับ autonomous run, multi-phase run หรืองานที่มนุษย์จะรีวิวหลัง step away วิธีใช้: (1) copy header row จาก references/decision-log-template.tsv เพื่อเริ่ม log ใหม่ (2) append แถวด้วย helper: scripts/log.sh <logfile> <phase> <decision> <why> <evidence> <result> — helper ประทับ ts, เขียน header ครั้งแรก, strip tab/newline ที่หลุดมา และ prefix cell ที่ขึ้นด้วย =,+,-,@ ด้วย single quote กัน formula execution เวลา reviewer เปิดใน spreadsheet (printf ตรง ๆ ก็ได้แต่ต้องระวัง byte เหล่านั้นเอง ถ้า cell มาจาก generated/user-supplied text) (3) คอลัมน์: ts (ISO8601), phase, decision (หนึ่งบรรทัด), why (ภาษาพูด plain — ถ้า principle ผลักดันก็บอกเป็นคำพูด เช่น "explored options first, this was a one-way door" ไม่ใช่ jargon tag), evidence (pointer เดียว), result (เช่น tests green, reverted, pixel-diff 0, INCONCLUSIVE, open) (4) log เฉพาะ decision points + checkpoints: fork ที่เลือก, unit จบพร้อมผล verify, pivot/revert พร้อม trigger, blocker ที่ปลุก, gate ที่แก้; loop run = หนึ่ง row ต่อ iteration; ข้ามเรื่อง trivial (5) เก็บที่ decisions.tsv ใน work dir หรือ .audit/<task-slug>.tsv เมื่อหลายงานรันพร้อมกัน — default ไม่ commit; commit เมื่องานใหญ่พงที่ reviewer ต้องใช้ trail เชื่อผล (port ข้ามภาษาใหญ่, migration หลายสัปดาห์) ซึ่ง TSV ใน PR จะ render เป็นตาราง (6) ท้าย run: audit log กับ transcript ของ workspace ปัจจุบัน แล้วรัน cross-model review subagent

💬 ตัวอย่าง prompt

/show-me-your-work — run นี้ยาวทั้งคืน เก็บ decision log ไว้ให้เดี๋ยวเช้าผมกลับมาอ่านเอง

⚡ Auto-invoke

ออกแบบให้ไม่ auto-invoke เองจากงานสั้น ๆ: frontmatter มี disable-model-invocation: true — ผู้ใช้เรียกเอง หรือถูก chain โดย skill อื่น: figure-it-out ใช้เป็น audit trail บังคับใน Phase D และ poteto-mode route เข้ามาเมื่อ "Long, autonomous, or multi-phase work, or any task the user steps away from to review later" ข้อสังเกต: ใน ZCode flag นี้อาจถูก ignore เพราะเป็น unknown field แต่ design intent ถือเป็นตัวตั้ง [uncertain]

🔗 คู่กันกับ

เป็น audit-trail skill ที่ skill อื่น route มาหา: figure-it-out (Phase D) เรียกใช้ และ poteto-mode route งาน long/autonomous เข้า skill นี้; log text ต้องผ่าน unslop; อ้าง principle-encode-lessons-in-structure (evidence จาก committed script ที่ reviewer รันซ้ำได้ ดีกว่า hand-made one-off); ใช้ references/decision-log-template.tsv + scripts/log.sh ในตัวเอง

⚠️ เคล็ดลับ/ข้อควรระวัง
  • หนึ่งแถว = หนึ่ง decision หรือ checkpoint — ถ้าไม่ fit หนึ่งบรรทัดแปลว่า decision ยังไม่ crisp ให้คิดให้ชัดก่อนเขียน
  • Append-only เด็ดขาด — ห้ามแก้หรือลบ history; call ที่ผิดให้เพิ่มแถวใหม่ที่ supersede
  • เขียนเหมือนบอกเพื่อนร่วมงาน: plain words, concrete actions, no AI speak, no abstract jargon — reviewer ต้องเข้าใจแถวโดยไม่ต้อง decode
  • cells ต้อง single-line — ระวัง tab/newline จากข้อความที่ generate หรือมาจาก user; ใช้ helper ช่วยกัน formula execution เวลาเปิดใน spreadsheet
  • ตอน audit: แถวที่ไม่มีใครจะ audit ให้ตัดทิ้ง ("If nobody would audit a row, it doesn't earn its place") แต่ fork/pivot/ทางที่ abandon ที่ shape งานแต่ไม่ได้ log คือ gap ต้องเติม
  • ตอนอ่าน transcript เพื่อ audit ห้าม glob ข้าม workspace อื่น — อ่าน session ของ workspace ปัจจุบันเท่านั้น (transcript โปรเจกต์อื่นคือ private chat)
  • กฎ cross-model review ระบุให้ spawn subagent "คนละ model family" แต่บน ZCode นี้ poteto-mode ระบุว่า subagent ทุกตัวรันบน session model — ความต่างจริงจึงมาจาก subagent type/prompt มากกว่า model family [uncertain]
  • ตัว log เองไม่คือผลงาน — "The trail plus the diff is what lets the human come back and trust the work" (ตาม figure-it-out ที่ใช้ตัวนี้)
❓ ยังไม่ยืนยัน
  • auto_invoke
  • tips
👤 เรียกเอง/interrogateSkill adversarial review ด้วย reviewer subagent หลายตัวที่ spawn อิสระต่อกัน — หนึ่ง reviewer ต่อหนึ่ง subage…
📦 ทำอะไร

Skill adversarial review ด้วย reviewer subagent หลายตัวที่ spawn อิสระต่อกัน — หนึ่ง reviewer ต่อหนึ่ง subagent type ที่ configured ไว้ ทุกตัวได้ prompt + rubric เดียวกัน แล้ว synthesize เป็น verdict เดียวที่แบ่ง finding เป็น Act on / Consider / Noted / Dismissed พร้อม Agreement Map; deliverable คือ verdict เท่านั้น — ห้าม auto-apply การแก้โค้ด

💡 เจ๋งยังไง

Insight เฉพาะของ port นี้ (ต่างจากแนว multi-model แบบ upstream) คือ: ZCode รัน subagent ทุกตัวบน session model ดังนั้น adversarial signal ไม่ได้มาจาก model คนละตระกูล แต่มาจาก "the reviewers' distinct postures and tool scopes (a reviewer type reads a diff differently than an architect type), not assigned personas" — ใช้ type จริงที่มี tool scope ต่างกันแทนการสวม persona ประกอบกับ statistical framing ของ review: "Agreement across independently spawned reviewers is high-confidence signal; single-reviewer findings are worth reading but lower confidence" และ synthesize step ทำมันจริง (consensus = raised by 2+ reviewers, dedup, จด disagreement) มี Step 2 "State the Intent" บังคับเขียน intent หนึ่งย่อหน้าก่อน spawn เพื่อชี้ว่า reviewer ท้าทาย "งานบรรลุ intent ดีไหม" ไม่ใช่ "intent ถูกไหม" และ lead judgment ประกาศชัดว่า "You are the lead reviewer, a pragmatic senior engineer, not a neutral aggregator" — ใช้บริบทเต็ม (goal, constraints, timeline, tradeoffs ที่เคยพิจารณาแล้ว) เพราะ reviewer เห็นเพียง slice ของ codebase นอกจากนี้ section Dismissed ที่บังคับแสดงว่าอะไรถูกกรองทิ้งเพราะอะไร (เพื่อให้ user override ได้) และ robust ต่อ misconfiguration: type ที่ spawn ไม่ได้ให้หยิบตัวใกล้เคียงจาก error แล้วเปิด PR แก้ config แยก — "Do not block the review on the type issue"

🧭 วิธีใช้

เรียกด้วย /interrogate (มี disable-model-invocation: true) — trigger ตาม description: "interrogate", "adversarial review", "multi-subagent review", "challenge this", "stress test this code", "find blind spots", "tear this apart" ขั้นตอนภายใน: (1) Determine scope — ใช้ไฟล์/diff ที่ user ชี้, ถ้าอยู่ feature branch รัน git diff main...HEAD, หรือเก็บไฟล์จาก recent work; package diff + context files ที่ reviewer ต้องใช้ (2) State the intent — เขียนย่อหน้าเดียวจาก user message / commit messages / PR description / ตัวโค้ด; ถ้า intent ไม่ชัด ให้ถาม user ก่อน (3) Spawn reviewer ทุกตัวพร้อมกันใน message เดียว — ใช้ list "interrogate reviewers" จาก ~/.zcode/pstack-roles.md ถ้ามี (หนึ่ง reviewer ต่อ entry, ปรับ label A/B/C/D ตามจำนวน) ไม่งั้นใช้ default: Reviewer A = code-reviewer, B = code-architect, C = poteto-agent, D = general-purpose; ทุกตัวได้ template เดียวกัน (references/reviewer-prompt.md) ใส่ intent + diff + rubric (references/rubric.md) + code-quality lens (references/code-quality-review.md) (4) Synthesize — parse ทุก finding, หา consensus (ยกโดย 2+ reviewers = สัญญาณสูงสุด), deduplicate การพูดต่างคำเรื่องเดียวกัน, จด disagreement ที่ reviewer หนึ่ง flag และอีกตัวบอกตรงข้าม (5) Lead judgment — แบ่งทุก finding เป็น Act on / Consider / Noted / Dismissed พร้อมบอก reviewer ที่ยก + rationale หนึ่งบรรทัด (6) Present ตาม format: Intent / Reviewers / Act On / Consider / Noted / Dismissed / Agreement Map

💬 ตัวอย่าง prompt

/interrogate — challenge การ design cache invalidation ใน branch นี้หน่อย ผมอยากรู้ blind spots ก่อนเปิด PR

⚡ Auto-invoke

ออกแบบให้ไม่ auto-invoke ในงานปกติ: frontmatter มี disable-model-invocation: true — ผู้ใช้เรียกเองด้วย trigger phrases หรือถูก poteto-mode chain เข้ามาเมื่อ "Contested design → the interrogate skill (adversarial multi-subagent review) before shipping" ข้อสังเกต: ใน ZCode flag นี้อาจถูก ignore เพราะเป็น unknown field แต่ design intent ถือเป็นตัวตั้ง [uncertain]

🔗 คู่กันกับ

how skill (Critique mode) ใช้ lead-judgment framework เดียวกันกับ interrogate (ไฟล์ how ระบุ "Same framework as the interrogate skill") และใช้ critic types ชุดเดียวกัน (code-reviewer, code-architect, poteto-agent, general-purpose); configurable ผ่าน ~/.zcode/pstack-roles.md ซึ่งเขียนโดย /setup-pstack; ต่างจาก review ทั่วไปตรง deliverable เป็น synthesized verdict — ไม่แก้โค้ดตาม

⚠️ เคล็ดลับ/ข้อควรระวัง
  • อย่าให้ reviewer ตัดสิน intent — reviewer challenge ว่างานบรรลุ intent ไหม; intent เองไม่ชัดให้ถาม user ก่อน spawn
  • ทุก reviewer ต้องได้ template + rubric + code-quality lens ชุดเดียวกัน เพื่อให้ความต่างของมุมมองมาจาก subagent type จริง ไม่ใช่ prompt คนละแบบ
  • ห้ามแต่ง persona ให้ reviewer — บน ZCode ความ adversarial มาจาก postures/tool scopes ของ type ไม่ใช่การสวมบท
  • ห้าม auto-apply การแก้ — deliverable คือ synthesized verdict เท่านั้น
  • น้ำหนัก finding ขึ้นกับจำนวน reviewer ที่ยกแบบอิสระ: ยกข้าม 2 ตัว = high confidence, ตัวเดียว = อ่านได้แต่ weight ต่ำกว่า
  • การที่ reviewer หนึ่ง flag และอีกตัวบอกตรงข้าม คือบริบทที่มีค่า — ใส่ไว้ใน verdict ไม่ใช่ซ่อน
  • ถ้า configured type spawn ไม่ผ่าน อย่า block review — อ่าน valid types จาก error, ใช้ตัวใกล้เคียง, แล้วเปิด PR แก้ค่า config แยกต่างหาก
❓ ยังไม่ยืนยัน
  • auto_invoke

🏛 Design (4)

👤 เรียกเอง/architectDesign-before-implement skill ของ pstack: sketch types, function signatures, และ module boundary เป็น 'not im…
📦 ทำอะไร

Design-before-implement skill ของ pstack: sketch types, function signatures, และ module boundary เป็น 'not implemented' bodies ก่อนเขียนโค้ดจริง ผ่าน 5 phase — Ground, Sketch, Agree, Implement, Scrap — โดย Phase Sketch จะรัน arena แบบ 'design it twice' ให้ได้ candidate โครงสร้างคนละแบบอย่างน้อย 2 ตัวมา synthesize เป็น design package เดียว แล้วค่อย implement เติมเนื้อใส่ skeleton ทีหลัง

💡 เจ๋งยังไง

กลไกเด็ด ๆ จากไฟล์: (1) บังคับ 'Design it twice' — ต้องมี candidate ที่ structurally distinct อย่างน้อยสองตัวก่อน synthesis เสมอ 'even when the first looks sufficient' คือ principle exhaust-the-design-space ทำเป็นขั้นตอนจริง (2) grounding มีมาตรฐานวัด: 'Naming a file isn't grounding' — ต้องได้ traced model ที่ how skill กำหนด ถ้า design แตะ ownership/layering ต้องรัน why ให้ rationale กลายเป็น constraint ไม่ใช่เดา (3) เกณฑ์เทียบ candidate คือ 'interface depth' — เลือก design ที่ซ่อน complexity ได้มากหลัง public surface ที่เล็กและง่าย (4) มีกลไกจับ architecture ที่ผิดตั้งแต่ Phase E ผ่าน pattern ของ friction: workaround รูปเดียวกันเกิดซ้ำข้ามโค้ด, types ที่ต้องใช้ escape hatch (any, casts, optional fields), 'we need a lock' reflex, callers ต้องรู้กฎภายใน — 'The signal is a pattern, not single instances' แล้ว scrap ด้วยหลัก subtract-before-you-add ให้ sketch ใหม่เล็กกว่าเดิม (5) deviation จาก sketch เป็นสัญญาณต้อง surface ไม่ใช่ดูดซับเงียบ ๆ — ต้องถามว่า sketch ผิด, requirement หลุด, หรือ implementation ฝืนเกิน

🧭 วิธีใช้

เรียกด้วย /architect หรือพูด 'architect this' / 'design this' — ใช้กับงาน non-trivial ที่กระโดดลงโค้ดเลยอาจล็อก shape ผิด ขั้นตอน: (A) Ground — รัน how บน subsystem ที่แตะ (critique mode ถ้าโครงเดิมคือ constraint หรือ design ต้อง push back) ข้ามได้เฉพาะ greenfield แท้ ๆ (B) Sketch — รัน arena พร้อม references/runner-prompt.md ให้ runner (default: code-architect, poteto-agent, code-reviewer, general-purpose ปรับได้ที่ ~/.zcode/pstack-roles.md) ผลิต design package ตาม references/rationale-template.md: เขียน caller's usage ก่อนแล้วค่อย derive type sketch, function signatures, module map — คัดทุก candidate ผ่าน references/design-red-flags.md (shallow module, information leakage, temporal decomposition, pass-through method) (C) Agree — default ลุย implement ต่อเลยไม่มี human checkpoint ถ้าจะให้หยุดขอ sign-off ต้องพิมพ์ '/architect with checkpoint' หรือ 'stop and show me before implementing' (D) Implement — เติม body แทน 'not implemented' โดย synthesized sketch คือ contract deviation ต้อง surface (E) Scrap — friction ซ้ำรูปเดียวกัน 2+ จุด = ทิ้ง sketch, รัน how ใหม่บนสิ่งที่สร้างไปแล้ว, ออกแบบใหม่ให้เหมือน constraint ใหม่เป็น day-one assumption แล้วกลับไป Phase B

💬 ตัวอย่าง prompt

/architect with checkpoint — ออกแบบระบบ per-tenant rate limiting ให้ API gateway เรา อยากเห็น type sketch กับ module map ก่อน แล้วค่อยให้เติมโค้ด

⚡ Auto-invoke

ไม่ auto-invoke — frontmatter ระบุ disable-model-invocation: true ต้อง user เรียก /architect เอง หรือถูก chain: poteto-mode (playbook Bug fix ใช้ architect ก่อนแก้ที่ข้าม function boundary) และ no-comments (step 3 รัน /architect หนึ่งครั้งสำหรับ accepted set ถ้า fix ต้องการ shape) — ใน checkpoint mode มันจะหยุดรอ sign-off ตามที่ผู้เรียกสั่ง นอกนั้นทำต่อเองถึงจบ

🔗 คู่กันกับ

ต้องมี how (Phase A grounding — 'Naming a file isn't grounding') และ arena (Phase B sketch เรียก arena พร้อม runner-prompt) — ทางเลือกเสริม: why (เมื่อ design แตะ ownership/layering), interrogate (กดดัน design แบบ adversarial ก่อน implement), และ principle skills ที่อ้างในไฟล์: exhaust-the-design-space, foundational-thinking, outcome-oriented-execution, redesign-from-first-principles, fix-root-causes, subtract-before-you-add

⚠️ เคล็ดลับ/ข้อควรระวัง
  • 'Naming a file isn't grounding' — ต้องได้ traced model ที่ how กำหนดจริง ห้ามผ่าน Phase A เพราะชี้ชื่อไฟล์ได้
  • ข้าม Phase A ได้เฉพาะงาน greenfield แท้ ๆ ที่ไม่มีระบบรอบ ๆ ให้ integrate
  • เทียบ candidate ด้วย interface depth ไม่ใช่ความคุ้นเคย — design ที่ซ่อน complexity หลัง surface เล็กชนะ design ที่กระจายของข้าม layer
  • สัญญาณ scrap ต้องเป็น pattern ซ้ำ (2+ deviation รูปเดียวกัน, escape hatch ใน types, lock reflex) ไม่ใช่ edge case เดียว — 'complexity in the data is not complexity in the design'
  • sketch ใหม่หลัง scrap ต้องเล็กกว่าเดิมก่อนโต (subtract-before-you-add) และ implementation lessons ต้องเป็น input ของ design ใหม่ ไม่ใช่ vibes
  • human pushback ต่อ shape (ตอน checkpoint หรือหลังจากนั้น) = Phase A evidence ใหม่ — ต้อง re-ground และ re-run arena ก่อนเขียนโค้ดต่อ
👤 เรียกเอง/arenaSkill แข่งขัน prototype: fan-out N parallel candidates ลง task เดียวกัน แล้วอ่านทุก candidate จบ เลือกตัวแข็ง…
📦 ทำอะไร

Skill แข่งขัน prototype: fan-out N parallel candidates ลง task เดียวกัน แล้วอ่านทุก candidate จบ เลือกตัวแข็งสุดเป็น base graft ไอเดียเด็ดของตัวแพ้อีกที และ verify ผลรวม — 6 phase: Frame, Fan out, Cross-judge, Pick, Graft, Verify ใช้ได้กับทุก artifact ที่พยายามครั้งเดียวอาจล็อก shape ผิด ทั้ง design, โค้ด และเอกสาร

💡 เจ๋งยังไง

แก้ single-attempt lock-in ด้วยการแข่งจริงและกลไกกันหลอกตัวเองครบ: (1) 'the prompt is the contract' — Phase A ต้อง derive rubric เป็น 3-6 criteria ที่วัดได้จริง ('Adds a --dry-run flag that skips writes' ไม่ใช่ 'code is correct') และ rubric เห็นเฉพาะ picker ไม่ให้ candidate เห็น กันการทำตาม rubric แบบจำใจ (2) mandatory rationale ต่อ candidate — 'Without it, the parent cannot tell whether a candidate's structure is principled or accidental' ทำให้ graft ไม่พัง (3) cross-judge แยกจาก parent: spawn readonly judge (subagent_type: "Explore") โดย prefer type ที่ไม่ได้ผลิต candidate เอง แล้วเทียบกับการอ่านของ parent — 'Disagreement means one of you is biased or the rubric was ambiguous' (4) กฎ graft ที่คม: ตัวแพ้มักมีของดีแค่ 1-2 อย่าง ต้อง fold ด้วยมือไม่ paste กล และ 'The rejection notes are the highest-signal part of the record' (5) จับสัญญาณ meta ได้: candidate เกิดรวมกันเป็น shape เดียว = consensus ส่งได้เลยไม่ต้อง graft, diverge สุด ๆ = Phase A under-specified ให้ reframe แล้ว rerun ไม่ใช่เฉลี่ยความเห็น (6) Phase F ยังต้องพิสูจน์ตามปกติ — 'The arena does not earn you a pass'

🧭 วิธีใช้

เรียกด้วย /arena หรือ 'arena this' / 'throw it in the arena' — ใช้เมื่อ non-trivial artifact ควรลองหลายแบบก่อน หรือถูกเรียกเป็น sub-step ของ architect (design sketch) และ blast-radius (change ใหญ่ให้ถามหลาย model แล้ว merge) ขั้นตอน: (A) Frame — ประกาศ artifact, rubric 3-6 ข้อ, เลือก runner จาก 'arena runners' ใน ~/.zcode/pstack-roles.md (default อย่างละหนึ่ง: poteto-agent, code-reviewer, code-architect, general-purpose — ซ้ำ type ได้ถ้างาน generation-bound), กำหนด output path แยกต่อ candidate (worktree หรือ /tmp/arena-<slug>/candidate-<n>/) (B) Fan out — spawn N ตัวใน message เดียว run_in_background: true พร้อม task + grounding path + output path + บังคับ rationale ตัวไหนตายใช้ N-1 แล้ว note dropout (C) Cross-judge — หลัง candidate จบครบ spawn judge ตัวเดียวแบบ readonly (subagent_type: "Explore") จาก 'arena cross-judge pool' ให้คะแนนตาม rubric และเสนอ base (D) Pick — parent อ่าน candidate ทุกตัวจบ ให้คะแนน criterion ต่อ criterion เทียบกับ judge แล้วบันทึก synthesis note (E) Graft — เด็ดของดีจากตัวแพ้ fold ด้วยมือ (F) Verify — พิสูจน์ artifact สุดท้าย

💬 ตัวอย่าง prompt

/arena เขียน validation layer ให้ signup flow ของเรา 4 แบบขนานกัน เทียบกันจริงจังแล้ว graft ของดีรวมเป็นตัวเดียว

⚡ Auto-invoke

ไม่ auto-invoke — frontmatter มี disable-model-invocation: true user เรียก /arena เอง หรือถูก chain: architect Phase B (sketch แข่งกัน), blast-radius step 6 ('For a big or wide change, run it as an arena'), และ poteto-mode playbook ฝั่ง design/code bakeoff — บน ZCode port นี้ runner เป็น subagent type (poteto-agent, code-reviewer, code-architect, general-purpose และ cross-judge ใช้ Explore) ปรับได้ผ่าน ~/.zcode/pstack-roles.md ที่ /setup-pstack เขียน

🔗 คู่กันกับ

ถูกเรียกใช้โดย architect (ต้องมี arena เป็น Phase B) และ blast-radius (change ใหญ่ให้รันเป็น arena) — ภายในเรียก principle prove-it-works ตอน Phase F, redesign-from-first-principles ตอน graft, separate-before-serializing-shared-state ผ่าน output path แยกต่อ candidate; arena ของ design ต้องมี grounding artifacts จาก how (ส่งผ่าน architect)

⚠️ เคล็ดลับ/ข้อควรระวัง
  • ห้าม spawn judge ก่อน candidate เขียนจบ — 'Spawning while candidates are still writing means the judge sees partial or empty outputs and reports them as dropouts'
  • ห้ามข้าม rationale — ไม่มี rationale parent แยก 'principled vs accidental' ไม่ได้ graft จะพังทันที
  • อ่าน candidate ทุกตัวจนจบก่อนเลือก — 'Skimming N candidates surfaces only the candidate whose surface looks most familiar'
  • candidate ทุกตัวต้องเขียนคนละ path — N ตัวเขียน path เดียวกันคือ shared mutable state ที่พัง
  • candidate diverge สุด ๆ = โจทย์ under-specified ให้ reframe แล้ว rerun อย่าเฉลี่ยความเห็นกัน
  • จบแล้วต้อง verify ตามปกติ ชนะ arena ไม่ยกโทษให้ — ถ้า verify เจอบั๊กที่ arena ไม่เห็น แปลว่า Phase A ผิด หรือมี candidate เห็นแต่ graft หลุด ให้ย้อนกลับตามกฎ 'Don't paper over'
👤 เรียกเอง/swarmSkill fan-out แบบเบาสุดของ pstack: ปล่อย N background workers ขนานกัน (แบ่ง slice คนละส่วน, แข่ง brief เดียวก…
📦 ทำอะไร

Skill fan-out แบบเบาสุดของ pstack: ปล่อย N background workers ขนานกัน (แบ่ง slice คนละส่วน, แข่ง brief เดียวกัน, หรือผสมทั้งสอง) แล้ว parent รอ รวมผล ส่ง report เดียวกลับ — 4 phase: Frame, Fan out, Aggregate, Report ต่างจาก arena ตรงที่ arena เน้นเลือก-graft artifact เดียวที่ชนะ ส่วน swarm เน้น coverage และรายงานรวม

💡 เจ๋งยังไง

เรียบแต่คมตรงจุดที่ fan-out พังบ่อย: (1) 'N is total workers, not a concurrency limit' — ตัดสับสนระหว่างจำนวนงานกับ concurrency ตั้งแต่ต้น (2) กำหนด selection rule ก่อน spawn สำหรับ race — ต้องประกาศ 'first pass', 'rank all', หรือ 'best-of' ล่วงหน้า ไม่ใช่มาตัดสินทีหลังให้เข้าทางผล (3) 'Every brief stands alone' — แต่ละ worker ต้องได้ goal, scope, slice/arm, วิธี verify, และรูปแบบ report ตายตัว PASS / ISSUES / BLOCKED พร้อม evidence ทำให้ aggregation อ่าน terminal result จบในตัว (4) แยกผล coverage (ทุก slice ต้องมี result) ออกจาก race (ใช้กฎที่ประกาศไว้) และห้าม paste ดิบ — 'Do not paste raw worker dumps' ต้องย่อเป็น compact result table + one-line evidenced issues + gaps/dropouts (5) dropout ไม่ล้มระบบ: ทำต่อด้วย N-1 แล้ว note ไว้ใน report

🧭 วิธีใช้

เรียกด้วย /swarm หรือ 'swarm this' — เหมาะกับ parallel coverage (ตรวจหลายหน้า/หลาย module), race (หลายวิธีทำงานเดียวกัน), gauntlet และ exploration ขั้นตอน: (A) Frame — ประกาศ done predicate + artifact/report ที่ต้องได้, เลือก shape (partition slices / race N workers / ผสม), ถ้า race ประกาศ selection rule ก่อน spawn, ตั้ง N, เลือก worker type จาก 'swarm workers' ใน ~/.zcode/pstack-roles.md (default poteto-agent), race ต้องตั้งชื่อ approach ของแต่ละ arm ล่วงหน้า, แยก writable output ต่อ worker (worktree / branch / /tmp/swarm-<slug>/worker-<n>/) (B) Fan out — spawn N ตัวพร้อมกันใน message เดียว run_in_background: true, brief ทุกใบ standalone, worker ที่ต้องเริ่มจาก non-default pushed branch จะ checkout เอง (C) Aggregate — อ่าน terminal result, เช็คว่าทุก required slice มี result, ใช้ race rule ที่ประกาศไว้ (D) Report — รายงานเดียวใน chat: table + issue one-liner + gaps/dropouts + race rule ที่ใช้

💬 ตัวอย่าง prompt

/swarm รีวิว 6 หน้าของ admin dashboard แบบขนาน หน้าละหนึ่ง worker ตรวจ broken link กับ layout บน viewport เล็ก รายงาน PASS/ISSUES รวมทีเดียว

⚡ Auto-invoke

ไม่ auto-invoke — frontmatter มี disable-model-invocation: true user เรียก /swarm เอง หรือถูก chain โดย poteto-mode (playbook ที่ต้อง fan-out ขนาน เช่น coverage/eval จะ route เข้า swarm) — บน ZCode port นี้ worker default คือ subagent_type poteto-agent (ปรับผ่าน 'swarm workers' ใน ~/.zcode/pstack-roles.md ที่ /setup-pstack เขียน) และ spawn ด้วย run_in_background: true

🔗 คู่กันกับ

เป็น skill ปลายทางที่ parent รวมผลเอง ไฟล์ไม่มี require pairing ตรง ๆ — ใช้คู่กับ arena เมื่ออยากได้ artifact ที่ชนะแล้ว graft ต่อ (swarm = coverage/report, arena = race + synthesize artifact เดียว), ถูก poteto-mode chain เป็นหัวหอก fan-out ขนาน และใช้ worker type จาก ~/.zcode/pstack-roles.md ชุดเดียวกับ skill อื่นของ pstack

⚠️ เคล็ดลับ/ข้อควรระวัง
  • ประกาศ race selection rule ('first pass' / 'rank all' / 'best-of') ก่อน spawn เสมอ — ห้ามเลือกเกณฑ์ทีหลัง
  • ตั้ง N ให้ชัดว่าคือจำนวน worker รวม ไม่ใช่ concurrency limit
  • brief ทุกใบต้อง standalone — worker ไม่เห็นบทสนทนาของ parent ต้องรู้ goal, scope, วิธี verify, รูปแบบ report (PASS/ISSUES/BLOCKED + evidence) ครบในใบเดียว
  • worker ที่เขียนไฟล์ต้องได้ output ของตัวเอง (worktree/branch//tmp) ห้ามใช้ path ร่วมกัน
  • report ห้าม dump ผลดิบ — ต้องย่อเป็น compact table + issue บรรทัดเดียวต่อข้อพร้อม evidence + ระบุ gap/dropout ตรง ๆ
👤 เรียกเอง/blast-radiusSkill วิเคราะห์ก่อน ship ว่า change ไปโดนอะไรข้างนอก diff: 'Find what a change breaks somewhere else, before …
📦 ทำอะไร

Skill วิเคราะห์ก่อน ship ว่า change ไปโดนอะไรข้างนอก diff: 'Find what a change breaks somewhere else, before it ships' — เน้นหา 'the one fact it's safe because of' แล้วพิสูจน์ fact เดียวนั้นด้วยการรันโค้ดจริง แทนการเขียน writeup ยาว เป็น companion ของ how (โค้ดทำอะไร) กับ why (ทำไมโครงแบบนี้) โดย blast-radius ตอบว่ามันทำให้อะไรพังที่อื่น

💡 เจ๋งยังไง

กลไกที่ต่างจาก impact analysis ทั่วไป: (1) 'Listing the callers is not the job. The agent can grep those in a second. The job is the breakage grep won't show you' — สนใจ JSON ที่ API ส่งกลับ, DB column, wire format, ภาษาอื่นที่อ่าน byte เดียวกัน, feature flag, โค้ดสาม hop ถัดไป (2) ต่อต้าน writeup ที่หลอกตัวเอง: 'A blast-radius writeup that sounds right is worthless' — ต้องหา fact ที่ความปลอดภัยทั้งก้อนพิงอยู่แล้ว prove ด้วย script/test ที่เรียกโค้ดจริง 'Words are where you start, not what you ship' (3) มีบันไดความมั่นใจ 5 ขั้น: You said so → pointed at the line (file:line จริง) → showed the bad case can't happen → ran it → reproduced it in the running app โดย 'Step 4 is usually one small script' และ fact ไหนไปไม่ถึงขั้น 4 ต้องบอกตรง ๆ ว่า unproven ห้ามเขียนเป็น settled (4) ทุ่มเวลากับ fact เดียวไม่ใช่ list of maybes — 'Most changes that look scary are safe because of a single fact' (5) change ใหญ่ให้รันเป็น arena เพราะ 'Different models catch different real bugs'

🧭 วิธีใช้

เรียกด้วย 'blast radius of X', 'what could this break' หรือเวลารีวิว diff เล็กที่ยังไม่เชื่อ ขั้นตอน 6 ข้อ: (1) อ่าน change — symbols ที่ add/change/delete และสิ่งที่ diff ไม่พิมพ์บอก (pull PR/commits ด้วย why step 2) (2) หา the one safe fact (3) ลงลึกจุดที่ grep ไม่ถึง: อ่าน source ของ library ที่เรียก เช็ค pinned version กับ local patch, ไล่เวลาการรัน (microtasks, unmount/teardown, Solid versus React) (4) ให้ risk แต่ละตัวมี chance จริงกับ cost จริง แยก 'risks ที่ยืนยันแล้ว' ออกจาก 'cleared ที่เช็คแล้วผ่าน' อ้าง file:line เสมอ ห้ามแต่ง caller/API (5) prove the one fact — เขียน script/test รันโค้ดจริง พร้อม paste ผล พิสูจน์ไม่ได้ถูก ๆ ก็ mark unproven ห้าม round up (6) change ใหญ่/กว้าง → รันเป็น arena ส่งกลับ 5 ส่วน: What it does / The one fact (proven หรือ unproven) / Risks / Cleared / Before you merge (test หรือ repro ที่ถูกที่สุดที่จับบั๊กจริง)

💬 ตัวอย่าง prompt

blast radius ของ PR นี้หน่อย — มันลบ cache eviction ใน batch job ออก ช่วยหา the one fact ที่ทำให้มันปลอดภัยแล้ว prove ด้วย script จริงให้ด้วย

⚡ Auto-invoke

ไม่ auto-invoke — frontmatter ระบุ disable-model-invocation: true ต้อง user พิมพ์เรียก ('blast radius of X', 'what could this break') หรือถูก chain โดย poteto-mode/playbook ตอนรีวิว diff เสี่ยง — เพราะมันถูกออกแบบเป็น companion skill ของ how กับ why ที่ตอบคำถามเฉพาะ ไม่ใช่รันเองกับทุก diff

🔗 คู่กันกับ

บริบทกลุ่ม understand: how (อธิบายว่าโค้ดทำอะไร) + why (pull PR/commits และ rationale ของโครง — ไฟล์ระบุ 'Use why step 2 to pull the PR and commits') — ตัวมันเองถูกกำหนดให้: change ใหญ่รันเป็น arena (step 6), writeup ต้องเขียนผ่าน unslop, และใช้แนวการอ้างหลักฐานเดียวกับ why (file:line จริง, การ search ไม่เจอก็ยังเป็นคำตอบ, ห้ามแต่ง caller)

⚠️ เคล็ดลับ/ข้อควรระวัง
  • อย่าส่ง writeup ที่แค่ฟังดูน่าเชื่อ — ต้องพา fact สำคัญไปถึงขั้น 4 'You ran it' เท่าที่ถูก และบอกตรงว่า fact ไหนหยุดอยู่ขั้นไหน
  • unproven คือคำตอบที่ถูกกฎ — 'Don't round up' พิสูจน์ไม่ได้ให้เขียนว่า unproven ห้ามฟังดูเหมือนผ่าน
  • ใช้เวลากับ 'the one fact' ไม่ใช่ไล่ list ความเสี่ยงยาว ๆ — fact เดียวที่ hold จะทำให้ case น่ากลัวส่วนใหญ่ตายพร้อมกัน
  • ความเสี่ยงที่เช็คแล้วใช้ได้แยกไว้หมวด Cleared ไม่คุยปนกับ real risks
  • การ search แล้วไม่เจอก็นับเป็นคำตอบ และห้ามสมมติ caller หรือ API ที่ไม่มีจริง
  • ก่อน merge ต้องส่ง repro ที่ถูกที่สุดที่จับบั๊กได้ พร้อม script ที่เขียนไว้แนบมาด้วย

🔨 Build & Craft (5)

👤 เรียกเอง/tddSkill แก้บั๊กแบบ test-first: เมื่อบั๊กมี test path ที่ชัดและถูก ให้ทำ broken behavior ให้ executable ก่อนแตะ …
📦 ทำอะไร

Skill แก้บั๊กแบบ test-first: เมื่อบั๊กมี test path ที่ชัดและถูก ให้ทำ broken behavior ให้ executable ก่อนแตะ production code — เขียน regression test เล็กที่สุดที่ fail ก่อน fix แล้ว pass หลัง fix พร้อม guardrail ชัดว่าเมื่อไหร่ไม่ควรฝืนเขียน test (แพง, brittle mocks, integration-heavy) ให้ใช้ verification ที่ใกล้ตัวสุดแทน

💡 เจ๋งยังไง

เป็น TDD ที่รู้จักยอมอย่างมีเหตุผล กันการผลิต test ขยะ: (1) เงื่อนไข fail-for-the-right-reason — 'Run the new test before fixing. Confirm it fails for the intended reason. If it passes or fails for an unrelated reason, correct the test or reproduction before editing the implementation' คือหัวใจ red-green ที่หลายคนข้าม (2) test ต้อง encode intended behavior ไม่ใช่ mirror implementation ปัจจุบัน (3) นิยาม bad test ตรงตัว: test ที่ test mocks เป็นหลัก, encode implementation details, ขึ้นกับ timing หรือ unrelated global state, ต้องใช้ infra แพง หรือจะถูกลบทันทีหลังพิสูจน์ fix — 'Prefer no new test over a bad test' (4) ห้าม skip เงียบ ๆ ถ้า failing test impractical ต้องอธิบายเหตุผลก่อนแล้วเลือก closest executable regression check (script, manual repro, browser automation, snapshot, log assertion) (5) Final response รายงาน evidence ไม่ใช่แค่ผลลัพธ์: ชื่อ test ที่ fail ก่อน, ที่ pass หลัง, และ validation ข้างเคียงที่รันไป

🧭 วิธีใช้

เรียกด้วย /tdd หรือใช้เมื่อ user ขอ TDD, failing test, หรือ regression test โดยตรง หรือเมื่อบั๊กมีจุด test ที่ชัดและถูก — ตัว skill เองกำหนดให้ skip เมื่อ test path ไม่ชัด แพง หรือ integration-heavy ขั้นตอน 7 ข้อ: (1) เข้าใจบั๊ก: intended vs current behavior, affected path, repro เล็กสุดที่ observable (2) เลือก check ที่แคบที่สุดที่ executable ได้ โดยเอา unit/component/integration test ที่ codepath นั้นมีอยู่แล้วถ้าไม่มี practical path อย่าสร้างขึ้นใหม่เพื่อผ่าน workflow (3) เขียน failing test ก่อน (4) รันยืนยันว่า fail ด้วยเหตุผลที่ตั้งใจ (5) แก้ production เปลี่ยนน้อยที่สุดให้ตรง intended behavior โดยคง contract รอบ ๆ (6) รัน regression test ซ้ำให้ผ่าน (7) รัน validation ข้างเคียง (test ใกล้เคียง, type check, lint, scenario check) ถ้า change เสี่ยงกว้าง — ถ้า failing test impractical ห้ามข้ามเงียบ ต้องอธิบายเหตุผลก่อนแล้วใช้ closest executable check แทน

💬 ตัวอย่าง prompt

/tdd บั๊กนี้ครับ: coupon เดียวกันใช้ซ้ำได้ถ้ากด apply สองครั้งติด — เขียน regression test ให้ fail ก่อนแล้วค่อยแก้

⚡ Auto-invoke

ไม่ auto-invoke — frontmatter มี disable-model-invocation: true และ description บอกเจตนาชัด 'Use only when the user explicitly asks for TDD, a failing test, or a regression test, OR when the bug has an obvious cheap local test target' — ในทางปฏิบัติ user เรียกตรง ๆ หรือถูก chain โดย poteto-mode (playbook Bug fix จะใช้ tdd เมื่อมี test path ที่ถูก) แม้ flag จะกันการ trigger อัตโนมัติเมื่อ description match

🔗 คู่กันกับ

จับคู่ธรรมชาติกับ playbook Bug fix ของ poteto-mode (fan-out how + why แบบขนานก่อน แล้ว tdd ตรงขั้นแก้) — ในไฟล์ไม่มีการเรียก pstack skill อื่นต่อ แต่ guardrail ของมันอยู่หมู่เดียวกับ principle prove-it-works (ต้องมี failing-before/passing-after evidence) และเสริมได้ด้วย verification skill ของโปรเจกต์ (verify-*) เวลาต้อง repro บน app จริง

⚠️ เคล็ดลับ/ข้อควรระวัง
  • ห้ามแก้ test ให้เข้ากับ implementation ที่ผิด และห้ามอ่อน assertion เดิมเว้นแต่ expected behavior เปลี่ยนจริงพร้อมเหตุผลชัด
  • test ที่จะโดนลบทันทีหลังพิสูจน์ fix = bad test อย่าเขียน
  • บั๊ก flaky ให้ทำ test ให้ deterministic เท่าที่ทำได้ แล้ว document สัญญาณที่ lock ไว้
  • ถ้าบั๊กเปิดเผย bug class กว้าง ให้ land focused regression path ก่อน แล้วค่อยพิจารณา sibling coverage เพิ่ม
  • ถ้าไม่มี practical test path อย่าสร้าง test ขึ้นเพื่อผ่าน workflow — อธิบายเหตุผลแล้วใช้ closest executable check (script, snapshot, log assertion) แทน
✨ auto/unslopSkill กำจัด AI-slop ออกจากงานเขียนทุกชนิด: scan pattern 31 กลุ่ม (puffery, AI vocabulary อย่าง delve/landscap…
📦 ทำอะไร

Skill กำจัด AI-slop ออกจากงานเขียนทุกชนิด: scan pattern 31 กลุ่ม (puffery, AI vocabulary อย่าง delve/landscape/testament, 'Not just X, but Y', rule of three, em dash overuse, chatbot phrases, filler, สำนวนเมตาฟอร์เทคนิคอย่าง substrate/ratchet) แล้วเขียนใหม่ให้คงความหมาย เพิ่มเสียงมนุษย์ และ self-audit ต่อ ใน frontmatter ระบุว่า 'Must always apply' — เป็น skill ที่ทั้ง pstack ชุดและชุมชนภายนอกหยิบไปใช้มากที่สุด

💡 เจ๋งยังไง

ไม่ใช่แค่ ban list: (1) ครึ่งหลังของงานคือ 'Adding soul' — 'Removing patterns is half the job. Sterile, voiceless writing is just as obvious' ให้มี opinion, ผสมจังหวะประโยค, ยอมรับความซับซ้อน ('Impressive but also kind of unsettling' beats 'impressive'), ปล่อยความไม่เป๊ะบ้างเล็กน้อย (2) กติกาละเอียดถึงระดับ anti-tell ซ้อนกัน: ห้าม em dash ทั้งหมด และห้ามหนีไปใช้ parenthesis แทน ('reaching for parentheses instead just trades one tell for another') (3) กลุ่ม 'Plain speech' เปลี่ยนจากความรู้สึกเป็นกลไก/ตัวเลข: 'Say what it does, not how it feels' — 'SQL you can read' ต้องกลายเป็น '.toSQL() returns the exact string sent to the database' พร้อมเทสความจำเจ: 'if the sentence could appear unchanged in another project's docs, it says nothing about this one. Cut it.' (4) จับ adverb ที่หนุน verb อ่อน ('An adverb propping up a weak verb means the verb is wrong') และ jargon เมตาฟอร์ ('Gold-plating' becomes 'more than the job needs', 'Ratchet' becomes 'a limit that only tightens') (5) เป็น dependency ของ skill หลายตัว — technical-writing บังคับใช้กับทุก doc ที่แตะ, blast-radius ให้เขียนผ่าน unslop, automate-me ใช้เป็น prose discipline ทุกบรรทัด

🧭 วิธีใช้

ไม่ต้องเรียกเองก็ถูกหยิบใช้เวลา agent เขียน/แก้ prose แต่เรียกตรงได้ด้วย /unslop หรือขอให้ไล่ slop ในไฟล์ ขั้นตอน 4 ข้อ: (1) Scan pattern ตามหมวดในไฟล์ — Content (puffery, name-dropping, superficial -ing phrases, promotional language, vague attributions, formulaic challenges), Language (AI vocabulary, fancy ways to say 'is', 'Not just X, but Y', rule of three, synonym cycling, false ranges), Style (em dash, colon คั่นกลางประโยค, boldface, inline-header lists, title case, emoji, curly quotes), Communication artifacts (chatbot phrases, cutoff disclaimers, sycophancy), Filler (In order to → To, excessive hedging), Jargon (abstract metaphor nouns), Plain speech (feelings over mechanisms, dense sentences, passive voice, adverbs, fancy synonyms) (2) Rewrite คงความหมาย ตรง intended tone (3) Add soul ตามหัวข้อ Adding soul (4) Self-audit ด้วยคำถาม 'What makes this obviously AI generated?' แล้วแก้ที่เหลือ

💬 ตัวอย่าง prompt

unslop README.md หน่อย — ไล่ AI-slop ออกทั้งไฟล์ แล้วเขียนใหม่ให้อ่านรู้เรื่องเหมือนคนเขียนจริง อย่าให้เสียความหมายเดิม

⚡ Auto-invoke

auto-invoke ได้และถูกออกแบบให้เป็นแบบนั้น — frontmatter ไม่มี disable-model-invocation และ description ระบุ 'Must always apply' เจตนาคือให้ agent ใช้เองทุกครั้งที่ผลิต/แก้ prose (doc, README, PR description, reply) — นอกจากนี้ถูกเรียกซ้ำอีกชั้นจาก skill ที่ต้องผ่านมันก่อนส่ง เช่น technical-writing ('Apply the unslop skill to every doc this skill touches'), blast-radius ('Write it through unslop'), และ automate-me (prose discipline กับทุกบรรทัดของ -mode skill)

🔗 คู่กันกับ

ตัวเสริมกลางของทั้งชุด: technical-writing (คู่บังคับ — unslop เป็นเจ้าของ catalog ของ slop pattern, technical-writing เป็นเจ้าของโครงสร้าง 4 ชั้น และให้เพิ่ม metaphor ใหม่เข้ากฎ abstract-metaphor ของ unslop ต่อ), blast-radius (writeup ต้องผ่าน unslop), automate-me (ใช้กับทุกบรรทัดของ mode skill), poteto-mode (กฎเขียน anti-slop อ้างแนวคิดเดียวกัน) — ตัวมันเองไม่เรียก skill อื่นต่อ

⚠️ เคล็ดลับ/ข้อควรระวัง
  • อย่าแค่ลบ pattern — งานครึ่งหนึ่งคือเติมเสียง งานเขียนที่สะอาดแต่ไร้ความเห็นก็ยังดูเป็น AI
  • ห้าม em dash โดยเด็ดขาด และอย่าหนีไปใช้ parentheses/en dash แทน — จบประโยคหรือใช้ comma
  • colon ใช้ได้เฉพาะนำ list หรือตัวอย่าง ไม่ใช้เป็นตัวเชื่อมกลางประโยค
  • เทสความจำเจ: ถ้าประโยคนั้นวางใน docs ของโปรเจกต์อื่นได้ไม่เปลี่ยน = มันไม่ได้บอกอะไรเรื่องนี้ ตัดทิ้ง
  • adverb เป็นสัญญาณว่า verb ผิด — 'runs quickly' กลายเป็น 'is fast' หรือใส่ตัวเลขจริง
  • คำเมตาฟอร์เทคนิค (substrate, wedge, vector, ratchet, endgame, north star, flywheel) ต้องแทนด้วยคำ concret จริง ๆ ตามที่ไฟล์ให้ตัวอย่างคู่แทนไว้
👤 เรียกเอง/no-commentsSkill จัดการ comment ในโค้ดแบบ adversarial: spawn subagent 'Comment Sicko' (subagent_type: comment-sicko จาก …
📦 ทำอะไร

Skill จัดการ comment ในโค้ดแบบ adversarial: spawn subagent 'Comment Sicko' (subagent_type: comment-sicko จาก plugin) มาล่า comment ทั้ง scope/diff, parent ตรวจ report ของมันอย่างไม่เชื่อ, ลบ comment ที่ยอมรับ, แก้ที่ root cause จริงแทนการเก็บ workaround, และถ้า comment ขังข้อจำกัดที่เปลี่ยนไม่ได้จริง ให้ offer encoding เป็น type/lint/test/CI ก่อนแล้วค่อยลบ

💡 เจ๋งยังไง

สถาปัตยกรรมที่จับ failure mode ของงาน 'ลบ comment' ได้ครบ: (1) 'Authoring agents defend comments. Defer to Comment Sicko's fresh perspective' — แยกผู้โจมตีออกจากผู้เขียนโค้ดเดิม (2) parent ไม่รับของมาทั้งก้อน: ต้อง reject application-code edits, scope escapes, exception-protected deletions, MUST KILL ที่ให้เหตุผลผิด และ flag ที่เล็ง intentional code และกติกาที่คมที่สุด: 'A keep survives only with proof it is about something we cannot change' (3) มีวงจรลงโทษต่อ report พัง: revert แล้ว rerun report ที่ถูก reject หนึ่งครั้งพร้อมระบุความผิด reject ซ้ำ = report เป็น open และ fail /no-comments ทั้ง run (4) ข้อจำกัดที่ลบไม่ได้ต้อง encode ไม่ใช่เก็บ comment ไว้ — offer lint/type/test/CI ที่ถูกที่สุดรอ interactive approval ผ่านแล้ว encode แล้วลบ comment (5) แก้ที่ root cause ใน scope ถ้า root cause อยู่นอก scope ให้ land fix เล็กสุดใน scope แล้ว report ที่เหลือ open — 'Neither authorizes widening the fence nor fixing instances outside it'

🧭 วิธีใช้

เรียกด้วย /no-comments หรือคำตรง ๆ เรื่อง comment ขั้นตอน 6 ข้อ: (1) Spawn Agent ด้วย subagent_type: "comment-sicko" ส่ง scope (ไฟล์/diff ที่ caller ให้ หรือ diff กับ base branch default main รวม working tree) ห้ามสรุปกฎของ Comment Sicko ซ้ำใน prompt (2) ตรวจ report และ diff — reject สิ่งที่เกินขอบเขต/ผิดเหตุผล, audit lint กับ TypeScript suppressions ที่มันข้าม, รัน /how หรือ /why ก่อนยอมรับ kill/keep แบบบาง ๆ (IMPORTANT, do not remove), rerun หนึ่งครั้งถ้า reject (3) Fix accepted flags ง่าย ๆ ตรง ๆ (ลบ dead path, ตัด parameter, ใช้ API จริง) ถ้า fix ต้องการ shape รัน /architect หนึ่งครั้งสำหรับ accepted set — หยุดที่ sketch ให้ขั้น 4 implement (4) Implement fix ที่ root cause ที่เล็กที่สุดใน scope ถอด workaround ทุกตัวที่ถูกชี้ (5) Constraint comment (do not remove / do not change wording / talk to X before changing) → offer encoding ที่ถูกสุด รอ interactive approval ถ้าผ่าน encode แล้วลบ ไม่ผ่านก็ลบ comment แล้ว report constraint เป็น open (6) Report: จำนวนที่ลบ, ที่ restore, reruns, architect sketch, fixes, encoding ที่ offer/ที่ทำ, unenforced constraints และงาน open อื่น

💬 ตัวอย่าง prompt

/no-comments — ไล่ทุก comment ใน diff ของ branch นี้เทียบกับ main ข้อจำกัดไหนลบไม่ได้จริงช่วย propose lint เข้ามาแทน

⚡ Auto-invoke

ไม่ auto-invoke — frontmatter มี disable-model-invocation: true user เรียก /no-comments เอง หรือถูก chain โดย poteto-mode (playbook ฝั่ง code discipline) — ภายในมันจะ spawn comment-sicko subagent (type มาจาก pstack plugin ซึ่ง /setup-pstack ระบุว่า detect ได้เฉพาะเมื่อเปิด plugin นี้) และขั้นตรวจ report อาจเรียก /how, /why, /architect ต่อเอง

🔗 คู่กันกับ

เรียกต่อเอง: comment-sicko (subagent ผู้ล่า comment), /architect (เมื่อ accepted fixes ต้องการ shape — sketch เท่านั้น), /how หรือ /why (ไขข้อกำกวมของ kill/keep ก่อนตัดสิน) — อ้าง principle-fix-root-causes และ principle-redesign-from-first-principles เป็นเจตนาของขั้น implement โดยทั้งคู่ไม่ให้อำนาจขยาย scope

⚠️ เคล็ดลับ/ข้อควรระวัง
  • อย่า restore comment เพราะเสียดาย — keep จะรอดได้ต่อเมื่อพิสูจน์ได้ว่าเป็นเรื่องที่เราเปลี่ยนไม่ได้จริง
  • report ที่ถูก reject ให้ revert แล้ว rerun หนึ่งครั้งพร้อมชี้ความผิด — reject ซ้ำ = report เป็น open และ fail skill ทั้ง run
  • suppression ด้าน correctness/safety (lint, TypeScript) ยังเป็น MUST KILL ที่ actionable — exception ของ intentional code ไม่ครอบคลุมพวกนี้
  • ขั้น architect ห้ามลงมือ implement — 'Architect shapes. Step 4 implements.' หยุดที่ sketch
  • encoding ต้องรอ interactive approval — รัน unattended หรือใน eval ต้องมี caller pre-approval ก่อน
  • ถ้า root cause อยู่นอก scope ให้ land fix เล็กที่สุดใน scope แล้ว report ที่เหลือ open ห้ามขยายรั้วเอง
👤 เรียกเอง/technical-writingมาตรฐานการเขียนเทคนิค 4 ชั้นที่รวมวิธีเด่นของโลกไว้ในไฟล์เดียว: Diátaxis (เลือกโหมดเอกสารก่อน — tutorial/how-…
📦 ทำอะไร

มาตรฐานการเขียนเทคนิค 4 ชั้นที่รวมวิธีเด่นของโลกไว้ในไฟล์เดียว: Diátaxis (เลือกโหมดเอกสารก่อน — tutorial/how-to/reference/explanation), Google developer style (ประโยคคุยกับ reader), STE (โหลด statement ทีละหนึ่ง), และ Global English (ประโยคตีความได้ทางเดียว) มี 3 กฎครอบหัว: ตัดทุกคำที่ไม่ทำงาน, ใช้คำสั้นทั่วไป, และเมื่อกฎทำให้ประโยคแย่ให้แก้อีกทางหรือปล่อย — ใช้กับ docs, RFC, readme, PR description และ commit message

💡 เจ๋งยังไง

(1) มีเป้าวัดรูปธรรม: 'The goal is writing a tired engineer understands on the first read' และแต่ละชั้นตอบคำถามเดียว (ชนิดเอกสาร / ประโยคพูดกับใคร / ประโยคแบกเท่าไหร่ / อ่านสองทางได้ไหม) (2) กันเอกสารที่ 'ถูกกฎทุกข้อแต่ดูเหมือนเครื่องเขียน' ด้วยส่วน Vary the rhythm — ผสมความยาวประโยคตั้งใจ, มีความเห็นได้เฉพาะใน explanation, เจาะจงเหนือ sterile ('a column rename fails the build' ไม่ใช่ 'schema changes can cause issues') (3) กฎต่อต้าน jargon ปลอม: 'The codebase is the word list' — ใช้ symbol/ไฟล์/flag/command จริง ไม่ใช้คำบรรยายแทน และเจอ metaphor ใหม่ให้เพิ่มเข้ากฎ abstract-metaphor ของ unslop (4) มี worked example ที่ตรวจซอยกลับทีละข้อต่อชั้น ทำให้เป็น checklist ที่สอนได้จริง (5) ขอบเขตยุติธรรมชัด: PR description กับ commit message นับเป็นงานเขียน (ทุกชั้นยกเว้น Diátaxis), UI copy ไม่อยู่ใน scope, ทุก claim จำนวน/tree ต้องจริง ณ commit ที่ land พร้อมคำสั่ง regenerate

🧭 วิธีใช้

เรียกด้วย /technical-writing หรือใช้เวลาเขียน/รีวิว docs, RFC, readme, PR description, commit message กระบวนการ: (1) เลือกโหมด Diátaxis ก่อนด้วยสองคำถาม: บอก action (doing) หรือ understanding (thinking) × ใช้เพื่อ learning หรือ work — หนึ่งเอกสารหนึ่งโหมด ห้ามผสม (ตาราง reference ไม่มีใน tutorial, การโต้แย้งไม่มีใน how-to) ให้ split แล้ว link (2) เขียนประโยคแบบ Google: คุย reader ว่า 'you' present tense, คำสั่งเป็น imperative, เงื่อนไขวางหน้าคำสั่ง ('To delete the document, click Delete.'), common case มาก่อน, ห้าม 'simply/easy/quickly' ใน procedure, heading บอกประเด็น ('Pick the mode first' ไม่ใช่ 'Modes') (3) โหลด statement ทีละอันแบบ STE: หนึ่งคำสั่งต่อประโยค ตัด instruction เกิน ~20 คำ ประโยคอื่นเกิน ~25, เตือนวางหน้า step ที่มันป้องกัน, คง article 'the/a' ที่ตัดแล้วอ่านกำกวม, หนึ่งคำหนึ่งความหมายหนึ่งงาน (4) ปิดช่องสองความหมายแบบ Global English: only/not แนบคำที่มันแก้, แตก noun string ยาว, this/it ชี้สิ่งเดียวที่ชัด, ไม่ใช้ and/or กับ (s), เรียกสิ่งเดียวกันด้วยชื่อเดียวทั้งเอกสาร (5) ผ่าน unslop กับทุก doc ที่แตะ (6) ปิดท้ายด้วย review checklist 8 ข้อในไฟล์

💬 ตัวอย่าง prompt

/technical-writing รีวิว docs/api-gateway ทั้งโฟลเดอร์หน่อย แล้วช่วยเขียน PR description ของงานนี้ใหม่ตามมาตรฐานเดียวกัน

⚡ Auto-invoke

ไม่ auto-invoke — frontmatter มี disable-model-invocation: true user เรียก /technical-writing เอง หรือถูก chain โดย poteto-mode เมื่อ task เป็นงานเขียนเอกสาร/PR description — แต่เมื่อ active แล้วมันจะดึง unslop มาบังคับใช้เองตามหัวข้อ Voice and repo specifics ('Apply the unslop skill to every doc this skill touches')

🔗 คู่กันกับ

คู่บังคับกับ unslop (unslop ถือ catalog ของ slop pattern, ตัวนี้ถือโครงสร้าง 4 ชั้น — ต้องใช้ร่วมกันเสมอ) — อ้างแหล่งตรง: diataxis.fr, developers.google.com/style, asd-ste100.org (STE Issue 9, 2025), Kohl The Global English Style Guide — ใช้ร่วมกับ automate-me เวลาเขียน mode skill และเป็นมาตรฐานของ writeup ที่ skill ฝั่ง review เช่น blast-radius ผลิต

⚠️ เคล็ดลับ/ข้อควรระวัง
  • 'When a rule makes a sentence worse, fix the sentence another way or leave it alone' — กฎรับใช้ reader ไม่ใช่รับใช้ตัวเอง
  • อย่าผสมโหมด Diátaxis — ให้ split เป็นไฟล์แยกแล้ว link ที่จุดที่โหมดเจอกัน
  • ห้าม 'simply', 'easy', 'quickly' ใน procedure และห้าม 'please' — 'If it were simple, the reader would not be here'
  • อย่า pre-announce ('we will soon support') และอย่าเริ่มประโยคติดกันด้วยวลีเดียวกัน
  • เรียกสิ่งเดียวกันด้วยชื่อเดียวทั้งเอกสาร และอย่า reword ประโยคที่ไม่ได้แก้อะไรระหว่าง edit
  • ตัวเลข/จำนวน/tree ที่อ้างใน doc ต้องจริง ณ commit ที่ land พร้อมแนบคำสั่งที่ regenerate มัน
  • code snippet ใช้ tab indent เขียน path/symbol จริง และ UI element อยู่ใน bold ตาม Google style
✨ auto/typescript-best-practicesคู่มือกฎ TypeScript แบบตารางสั้น 16 ข้อที่ให้ apply principle type-system-discipline ก่อนแล้วเกาะ syntax ของ …
📦 ทำอะไร

คู่มือกฎ TypeScript แบบตารางสั้น 16 ข้อที่ให้ apply principle type-system-discipline ก่อนแล้วเกาะ syntax ของ TS: discriminated unions, branded types, constructive modeling (ออกแบบ shape ให้ค่าผิดสร้างไม่ได้), unknown แทน any, ห้าม as cast, narrowing hierarchy, exhaustiveness ด้วย never, satisfies, boundary validation, schema-derived types และอื่น ๆ — examples เต็มอยู่ที่ references/patterns.md

💡 เจ๋งยังไง

ปรัชญาชัดตั้งแต่แถวแรกของตาราง: 'No optional-field bags' — model ให้ impossible states แทนค่าไม่ได้ (1) Constructive modeling คือหัวใจ: '[T, ...T[]] for non-empty, [T, T][] for even length, start plus duration for a range. Not a runtime guard, not a wish for refinement types' — ให้ type ป้องกันแทนการเขียน guard รอบด้าน (2) ปานกลางอย่างมีเหตุผล: 'Keep T[] while every operation on it stays total. Strengthen to NonEmpty<T> only where the loose type forces !, a cast, or a should never happen throw' — ไม่บังคับ hyper-strict ทุกที่ (3) จัดอันดับ narrowing ให้เลือกได้: discriminant switch > in operator > typeof/instanceof > user-defined type guard > as และเตือนว่า type guard ที่โกหกแย่กว่า as เพราะ 'the bug hides behind a name that says it is safe' (4) ผูกกับงานจริง: 'Don't mock what you can run' ในกฎ Real tests, ไม่มี console.log ใน shipped code (ใช้ structured telemetry), object args ยกเว้น hot path — รวมทั้ง type discipline และ testability ไว้ไฟล์เดียว

🧭 วิธีใช้

ไม่ต้องเรียก — skill ออกแบบให้ apply เมื่ออ่านหรือแก้ไฟล์ .ts/.tsx ใด ๆ (auto-invoke) วิธีใช้ตามไฟล์: apply principle type-system-discipline ก่อน แล้วเดินตารางกฎ: (1) model variant ด้วย discriminated union ที่มี kind literal discriminant (2) brand primitive ด้วย '& { readonly __brand: "X" }' เมื่อต้องกันการสลับกัน โดย validate ครั้งเดียวตอนสร้าง (3) เลือก simplest total type ก่อน ค่อย strengthen เมื่อถูกบังคับ (4) external data เป็น unknown, cast เฉพาะหลัง validation, ใช้ satisfies ก่อน as (5) inline 'const _exhaustive: never = x;' ใน default arm เพื่อให้ compiler error เมื่อเพิ่ม variant ใหม่ (6) derive type ด้วย Pick/Omit/Parameters/ReturnType/Awaited/typeof ก่อนประกาศ interface ใหม่ (7) ส่ง object args แทน positional ยกเว้น hot path (per-frame render, tokenizers, parsers) (8) test ด้วยของจริงของ framework พร้อม leak/disposable checks ลด mock เฉพาะของที่รัน local ไม่ได้

💬 ตัวอย่าง prompt

รีวิว src/payments ให้ที อยากให้จัดการ discriminated union กับ as cast ตาม typescript-best-practices โดยเฉพาะจุดที่ยังใช้ optional-field bag อยู่

⚡ Auto-invoke

auto-invoke — frontmatter ไม่มี disable-model-invocation และ description ระบุตรง 'Use when reading or editing any .ts or .tsx file' เจตนาคือให้ ZCode agent ดึงไปใช้เองทุกครั้งที่แตะ TypeScript ไม่ต้องสั่ง — เป็น 1 ใน 2 skill ของกลุ่มนี้ที่ไม่มี flag กัน (อีกตัวคือ unslop)

🔗 คู่กันกับ

apply principle-type-system-discipline ก่อนเสมอ (ระบุในบรรทัดแรกของ body) และอ้าง principle-boundary-discipline ในกฎ boundary validation — ใช้คู่กับ tdd (กฎ Real tests: 'Don't mock what you can run') และ verification skill ของโปรเจกต์เมื่อต้อง 'verify UI in a running build'; ตัวอย่าง pattern อยู่ที่ references/patterns.md ของมันเอง

⚠️ เคล็ดลับ/ข้อควรระวัง
  • 'Every as is a runtime crash waiting' — cast ได้เฉพาะหลัง validation และชอบ satisfies กว่า as เพราะ validates ค่าโดยไม่ widen literal type
  • type guard ต้อง verify claim จริง — guard ที่โกหกแย่กว่า as เพราะซ่อนบั๊กหลังชื่อที่บอกว่าปลอดภัย ตั้งชื่อ isX/hasX เท่านั้น
  • อย่ารีบเสริม NonEmpty<T> ทุกที่ — เก็บ T[] ไว้ตราบใดที่ operation ทั้งหมด total เสริมเมื่อ loose type บังคับ ! / cast / 'should never happen' throw
  • any ทำลาย type checking ทุกจุดที่มันแตะ — external data เป็น unknown เสมอ
  • object args ข้ามได้ใน hot path (per-frame render, tokenizer, parser)
  • ห้าม console.log ใน shipped code — ใช้ structured logger ที่มี context พอ debug ต่อจาก id เดียวได้

♻️ Meta & Improve (2)

👤 เรียกเอง/reflectSkill เก็บเกี่ยวบทเรียนจาก session ปัจจุบัน: spawn reviewer 3 มุมแบบขนาน (Judgment, Tooling, Divergent) ไล่อ่…
📦 ทำอะไร

Skill เก็บเกี่ยวบทเรียนจาก session ปัจจุบัน: spawn reviewer 3 มุมแบบขนาน (Judgment, Tooling, Divergent) ไล่อ่าน transcript ที่ active อยู่, ส่งผลให้ synthesizer แยกเป็น Accepted/Rejected/Backlog, เช็คซ้ำว่า lesson ไหนควร encode เป็น lint/script/flag มากกว่าข้อความใน skill, แล้วเสนอ user approve ก่อน route เป็น edit ของ skill ที่มีอยู่จริงหรือสร้าง skill ใหม่ผ่าน skill-creator

💡 เจ๋งยังไง

ปิดวงจร 'เรียนรู้แล้วเก็บจริง' ของทั้งชุด โดยมีราวกันขยะครบ: (1) เกณฑ์ skip ก่อนแรง — 'Skip when the conversation is trivial, off-topic, or already covered by an existing skill the parent followed correctly. One-offs are not learnings' พร้อม trigger ตรวจได้ เช่น 'A complex task (5+ tool calls) just landed cleanly' หรือ 'The user corrected the agent's approach mid-task' (2) กันการข้าม workspace: หา transcript เฉพาะ session ปัจจุบัน 'Do not glob across other projects' session directories. That crosses workspace boundaries and reads private chats from unrelated projects' (3) แบ่ง 3 lens พร้อม reviewer type กำหนดเอง (default: Judgment = general-purpose, Tooling = code-reviewer, Divergent = general-purpose) และ reviewer ห้ามเขียนไฟล์ — parent เป็นคน apply เอง (4) Structural enforcement check: lesson ที่ lint rule/script/metadata flag/runtime check จับได้ดีกว่า ให้เลื่อนจาก Accepted ไป Backlog ตาม principle encode-lessons-in-structure — prose ใน skill เป็นทางสุดท้าย (5) 'Skill changes affect every future agent in the org; do not auto-apply' — ต้อง user เลือก subset ก่อนเสมอ (6) มี routing แยกความยาก: edit เล็ก parent ทำเอง, edit ใหญ่เกิน ~10 บรรทัด / tune description / สร้าง skill ใหม่ ให้ skill-creator ทำตาม draft/test/iterate loop

🧭 วิธีใช้

เรียกด้วย /reflect หรือพิมพ์ 'reflect' — ใช้เมื่องานซับซ้อนเพิ่งจบสวยและสูตรควรเก็บ, agent เจอ dead ends แล้วหาทางที่ใช้ได้ซึ่ง generalize ได้, user แก้วิธีทำงานกลางทาง, หรือ workflow ที่ non-trivial ยังไม่ถูกจับใน skill ไหน ขั้นตอน: (1) หา transcript ของ session นี้ (เช็คบรรทัดแรกของ JSONL ว่าตรง opening user prompt; ถ้าหาไม่เจอให้เขียน digest ย่อแทน) (2) Spawn 3 reviewer ใน message เดียว แต่ละตัวอ่าน references/judgment-reviewer.md, tooling-reviewer.md, divergent-reviewer.md แบบ verbatim (reviewer ต้องเป็น type ที่มี full tool access เพราะอาจต้องเช็ค ticket/trace ผ่าน MCP) (3) ส่งผลทั้งหมดให้ synthesizer 1 ตัว (general-purpose) อ่าน references/synthesizer.md ได้ Accepted/Rejected/Backlog list (4) Structural enforcement check — ย้ายข้อที่ lint/script จับได้ดีกว่าไป Backlog (5) นำเสนอรอ user approve แล้ว apply ตาม Routing field หรือส่ง Backlog เข้า tracker (6) สรุปสั้น ๆ ไม่มี preamble: edit ที่ apply, skill ใหม่, backlog ที่ไฟล์, ที่ drop พร้อมเหตุผล

💬 ตัวอย่าง prompt

/reflect — session นี้ผมต้องแก้วิธีทำงานของแกหลายจุดเรื่องการ build ทีละ module เก็บ lesson เหล่านี้เข้า skill ที่เกี่ยวข้องหน่อย

⚡ Auto-invoke

ไม่ auto-invoke — frontmatter มี disable-model-invocation: true และ description ระบุ trigger ตรง 'Use when the user says reflect' ต้อง user พิมพ์เอง หรือถูก chain โดย poteto-mode ช่วงปิดงาน — ภายในมันจะ spawn reviewer/synthesizer subagents เอง (general-purpose กับ code-reviewer ตาม default หรือค่า reflect-judgment/tooling ใน ~/.zcode/pstack-roles.md) และมอบ edit ใหญ่ให้ skill-creator skill จาก plugin ต่อ

🔗 คู่กันกับ

ปลายทางหลักคือ skill-creator (plugin skill): edit เล็ก parent ทำตรง, edit ใหญ่ (~10+ บรรทัด) / tune description / new skill → skill-creator draft/test/iterate loop — อ้าง principle-encode-lessons-in-structure ที่ขั้น structural check, ใช้ reviewer types จาก pstack-roles.md เหมือน skill fan-out อื่น, และมี SKILL.md validator ให้รันบน skill ที่แตะถ้า environment มี

⚠️ เคล็ดลับ/ข้อควรระวัง
  • one-off ไม่ใช่ lesson — ถ้า skill ที่มีอยู่ครอบอยู่แล้วและ parent ทำถูกตาม ให้ skip
  • reviewer ห้ามแก้ไฟล์ — เขียน findings กลับมาใน Agent response แล้ว parent เป็นคน apply
  • ห้าม glob transcript ข้ามโปรเจกต์ — ผิดขอบเขต workspace และอ่านแชทส่วนตัวของงานอื่น
  • lesson ที่ lint rule/flag/runtime check จับได้ดีกว่า อย่ายัดเป็น prose ใน skill — ส่ง Backlog ให้กลายเป็น tool
  • ห้าม auto-apply — ทุก Accepted edit ต้องผ่าน user approve ก่อน เพราะกระทบ agent ทุกตัวในองค์กร
  • Backlog ให้ไฟล์เข้า devex tracker เป็น issue อัตโนมัติ ไม่ใช่แก้ skill — มันคือ tracker submission ไม่ใช่ skill edit
👤 เรียกเอง/automate-meFlow แนะนำเป็นขั้นเพื่อเปลี่ยนนิสัยการทำงานของ user ให้เป็น -mode skill ส่วนตัว (เช่น jay-mode, priya-mode): …
📦 ทำอะไร

Flow แนะนำเป็นขั้นเพื่อเปลี่ยนนิสัยการทำงานของ user ให้เป็น -mode skill ส่วนตัว (เช่น jay-mode, priya-mode): เช็ค skill เดิมก่อน, ขุด transcript ของ workspace นี้ด้วย parallel mining subagents, ถามตรงด้วย AskUserQuestion, cluster ผลเป็นหมวด, ร่างผ่าน skill-creator + unslop, แล้ว land ผ่าน worktree/PR — update ซ้ำได้โดยขุดเฉพาะประวัติหลังแก้ครั้งล่าสุด

💡 เจ๋งยังไง

(1) ตัดสิน update vs สร้างใหม่ด้วยหลักฐาน: ดู git log -1 ของ skill เดิมเพื่อขุดเฉพาะ history หลังจากนั้น และถามแค่ 'what's changed or missing' ไม่ใช่เริ่มจากศูนย์ (2) การขุดประวัติมีเกณฑ์ความเชื่อถือ: 'Patterns seen in 2+ slices are high-confidence; lone signals are weak and usually get dropped' และแบ่งประวัติเป็น 3 slices (เช่น สองสี่สัปดาห์ล่าสุด) กัน overfit บทสนทนาเดียว (3) ใช้ AskUserQuestion แบบมีเหตุผล: หนึ่งสองคำถาม, 4-6 options, allow_multiple สำหรับคำถามหมวด, ตามด้วย free-form หนึ่งคำถาม — 'Don't dump 20 questions' (4) มีตัวอย่างรูปร่างให้อ่าน: poteto-mode คือตัวอย่าง shape แต่ห้ามลอกเนื้อหา ('the user's rules are not the same as poteto-mode's') (5) Guardrail ต่อต้าน skill ฟุ่มเฟือยตรงจุด: 'Reference, don't inline', 'Don't force symmetry' (ไม่มี process rules จริงก็ไม่ต้องมีหมวด Process), '"Communicate clearly" is not a section' แต่ 'Short paragraphs. Tables when comparing options.' ใช่ (6) รู้ว่า output ของตัวเองเป็น subjective — ไม่ทำ benchmark loop เอา vibe-check กับ user แทน และ default ตั้ง disable-model-invocation: true ให้ mode skill เพราะ 'they should only apply when the user explicitly invokes them'

🧭 วิธีใช้

เรียกด้วย 'automate me', 'create/update/refresh my -mode skill' หรือ 'turn my working style into a skill' ขั้นตอน 6 ขั้น: (0) เช็ค -mode skill เดิมใน .zcode/skills/**/*-mode/ และ ~/.zcode/skills/ ถ้ามี ยืนยันผ่าน AskUserQuestion ว่า update หรือเริ่มใหม่ (1) Mine history — หา transcript เฉพาะ workspace ปัจจุบัน (ZCode เก็บใต้ ~/.zcode/cli/) spawn parallel mining subagents ต่อ slice ไล่หา response preferences, delegation habits, verification posture, code/prose discipline, process conventions, meta preferences พร้อม evidence pointer (2) ถาม user ตรงด้วย structured questions สองรอบ + open question หนึ่ง (3) Cluster เป็น sections (Response style, Autonomy, Understand first, Subagents, Prose/code discipline, Review and verify, Process, Skills — ใช้เท่าที่ apply) (4) Draft ด้วย skill-creator — วางที่ .zcode/skills/<handle>-mode/ (หรือ personal ~/.zcode/skills/), description trigger ที่ชื่อ + /<handle>-mode + 'work in their style' ไม่ใช่ keyword generic, ตั้ง disable-model-invocation: true เป็น default (5) Iterate ผ่าน unslop + create-skill writing guidelines ให้ user ตรวจหลายรอบ ตัดอย่างโหด (6) Land ใน worktree เปิด PR ห้าม push main ตรง

💬 ตัวอย่าง prompt

automate me — สร้าง zic-mode skill จากวิธีทำงานของผมใน 3 สัปดาห์ที่ผ่านมา ถ้ามีตัวเดิมอยู่แล้วให้อัปเดตทับ

⚡ Auto-invoke

ไม่ auto-invoke — frontmatter มี disable-model-invocation: true และ description กำหนด trigger เป็นคำพูดของ user ('automate me', 'create/update/refresh my -mode skill', 'turn my working style into a skill') เพราะงานนี้แตะไฟล์ skill ส่วนตัวและต้องโต้ตอบถามหลายรอบ — และมันเป็น orchestrator ที่เรียก inline mining pass + skill-creator (จาก plugin) + unslop เองภายใน flow โดย 'It sequences them; it doesn't replace them'

🔗 คู่กันกับ

Orchestrates สามอย่างตามที่ไฟล์ประกาศ: inline mining pass (ขุด transcript), skill-creator จาก skill-creator plugin (authoring + YAML rules + description-optimization loop), unslop (prose discipline ทุกบรรทัด) — ใช้ poteto-mode เป็นตัวอย่างรูปร่าง output (อ่านอย่างเดียว ห้ามลอกเนื้อหา) และ -mode skill ที่ได้ออกมาเป็น skill ประเภท disable-model-invocation ที่ user เรียกด้วยชื่อตัวเอง

⚠️ เคล็ดลับ/ข้อควรระวัง
  • 'Don't overfit to one conversation' — preference ที่พูดครั้งเดียวและถูกขัดครั้งอื่นคือ noise ต้องเห็นซ้ำหลาย instance ก่อนเขียนลง skill
  • 'Don't be clever' — ห้ามแต่ง metaphor หรือ prose กวีใส่ skill สำหรับ agent อ่าน เก็บเป็น operational
  • skill อื่นที่ user พึ่ง ให้อ้างเป็น path ไม่ใช่ก๊อปเนื้อหามาวาง
  • 'Communicate clearly' ไม่ใช่ section — section จะมีได้ต้องเป็นกฎ specific ที่ไม่ใช่ default
  • เรียก user ในเนื้อ skill ด้วย 'the user'/'the human' ไม่ใช่ชื่อจริง เพราะคนอื่นอาจอ่านหรือเอาไปใช้
  • หมวดไหนไม่มีกฎที่มีความหมายให้ข้ามไป — 'Sparse is fine; bloated is not'
  • จุดพลาดบ่อย: update mode ไม่ใช่เขียนใหม่ — preserve section ที่ user ยังไม่ขัด แก้เฉพาะจุดที่มี evidence ใหม่

✅ Verify (2)

👤 เรียกเอง/create-verification-skillGenerator สร้าง project-local verification skill (.zcode/skills/verify-<app>/) ที่ drive app จริงแบบที่ user …
📦 ทำอะไร

Generator สร้าง project-local verification skill (.zcode/skills/verify-<app>/) ที่ drive app จริงแบบที่ user ใช้: 6 sections คือ Launch, Doctor, Drive, Evidence, Cleanup, Helpers พร้อม seed feature map (README + ไฟล์ต่อ feature 3-5 ตัวแรก) และบังคับให้รันคำสั่งของ skill ที่ generate ออกมาจนจบหนึ่งรอบก่อนส่งมอบ — รองรับทุก surface: web UI, CLI/TUI, desktop, API, mobile, library

💡 เจ๋งยังไง

(1) 'Interview the repo, not the user' — ตอบจาก codebase ก่อน ถาม user เฉพาะสิ่งที่มองไม่เห็น ครบ 5 มิติ: Surface (user แตะอะไร), Run (dev command ของ repo + ports/env/seed/auth), Drive (harness ที่มีอยู่ก่อนเสมอ: Playwright specs, expect scripts, PTY helpers แล้วค่อย generic recipe), Observe (evidence อะไรจับได้), Isolate (double-instance ได้ไหม — ถ้าไม่ได้เขียนไว้เลยว่า 'refusing to double-drive a shared instance beats corrupting the user's session') (2) เป้าผู้อ่านชัด: 'You write the generator's output for the next agent, not for a human: it will be read cold, mid-task, by an agent that has never seen the app' — จึงห้าม placeholder ตกค้าง ต้องเป็น selector/command จริงจาก repo นี้ (3) Doctor section คือไอเดียแหลม: one read-only check ตอบ 'is this instance worth driving?' ให้ agent รันก่อนทุกครั้งที่อะไรดูแปลก (4) Evidence มี proof standard ตายตัว: ขับ real user path ไม่ใช่ internal setters/test-only endpoints, จับ side effects (ไฟล์, rows, ข้อความ) คู่กับของที่เห็น และ verify dry-run ด้วยการ observe ไม่ใช่เชื่อชื่อ ('some dry-runs still touch the network or open a browser') (5) กลไกคุณภาพ: 'A generated skill that was never executed is a draft, not a deliverable' และ 'a cleanup that eats the proof fails this step'

🧭 วิธีใช้

เรียกด้วย /create-verification-skill หรือ 'make a control skill for this repo' — ใช้เมื่อโปรเจกต์ยังไม่มีวิธี scripted prove พฤติกรรม UI/CLI/service ขั้นตอน 5 ขั้น: (1) Interview the repo ตาม 5 มิติ — ถ้า checkout build/start ไม่ผ่านให้แก้ก่อน (หรือรายงานให้เป๊ะ) (2) Generate .zcode/skills/verify-<app>/SKILL.md พร้อม frontmatter (name: verify-<app> + description ที่ระบุ app, surface, และเมื่อไหร่ใช้ — ไม่มี frontmatter = skill ไม่ register) ครบ 6 sections: Launch (คำสั่งเริ่ม + สัญญาณ ready + teardown; CLI/TUI สั้นไม่มี server ให้ build ครั้งเดียวแล้ว drive ต่อ PTY/tmux session แยก), Doctor (read-only health check), Drive (harness recipe ด้วย stable handle จริง: ARIA labels, data attributes, prompt strings, route paths — ไม่ใช่ coordinate), Evidence (จับอะไร เก็บที่ไหน, มาตรฐาน proof), Cleanup (teardown สิ่งที่เรา start, ไม่ kill by process name, evidence รอด), Helpers (script ต้อง executable และโชว์ invocation ใน body) (3) Seed feature map: features/README.md + ไฟล์ต่อ feature (top 3-5) ตาม references/feature-map-example มี 4 H2: Sub-features, How to get to it (user POV), Driving it with <harness>, Gotchas (4) Prove: รันคำสั่งของ skill เองจบหนึ่งรอบ — launch, doctor, drive 1 mapped feature, จับ evidence, cleanup แล้วเช็คว่า evidence ยังอยู่; ทำ cleanup หลังทุก iteration ที่ล้มด้วย (5) Offer maintenance loop: ชี้ /maintain-verification-skill

💬 ตัวอย่าง prompt

/create-verification-skill — repo นี้ยังไม่มีทาง prove ว่าหน้า checkout ใช้งานได้จริง สร้าง verify skill ให้ที แล้ว seed feature map 5 ตัวแรก

⚡ Auto-invoke

ไม่ auto-invoke — frontmatter มี disable-model-invocation: true user เรียก /create-verification-skill เอง (หรือถูกเสนอโดย /setup-pstack ใน step สุดท้ายเมื่อโปรเจกต์ยังไม่มี verification skill) — ตัวที่มัน generate ออกมา (verify-<app>) จะเป็น project-local skill ที่ agent คนต่อไปหยิบใช้เวลาต้อง prove พฤติกรรม app และอยู่ในวงจรคู่กับ /maintain-verification-skill

🔗 คู่กันกับ

คู่ lifecycle ตรง ๆ กับ maintain-verification-skill (create ครั้งเดียว, maintain เป็นงานประจำ — ไฟล์ปิดท้ายเองว่า 'Point the user at /maintain-verification-skill for keeping the map honest as the app changes') — ถูก setup-pstack เสนอตอนติดตั้ง pstack, ถูก poteto-mode chain ใน playbook ฝั่ง Shipping/verify, และมี principle-prove-it-works เป็นแนวคิดรากของทั้ง skill

⚠️ เคล็ดลับ/ข้อควรระวัง
  • เขียนให้ agent ตัวต่อไปอ่านเจอ cold กลาง task — ห้าม placeholder เหลือใน generated skill
  • แก้ build/start ที่พังก่อน generate — 'a skill written against a broken base teaches wrong steps'
  • asset ที่หายแต่ไม่เกี่ยว (static dir, sample config) ให้ skill สร้างได้แต่ต้อง mark ชัดว่า verification scaffolding และลบใน cleanup
  • จับ stable handle (ARIA label, data attribute, route) ไม่จับ coordinate หรือลำดับ tab
  • cleanup ห้าม kill by process name — kill สิ่งที่เรา start เอง และต้องยืนยัน evidence ยังอยู่ที่ path ที่ระบุหลัง cleanup
  • drive 1 feature ตอน prove ก็พอ — feature map มีไว้ให้ run ถัดไปคลุมที่เหลือ
  • dry-run ที่ควรปลอดภัยให้ verify ด้วยการ observe จริง (ไฟล์, network, git refs) ไม่ใช่เชื่อชื่อมัน
👤 เรียกเอง/maintain-verification-skillงานประจำที่ทำให้ verification skill และ feature map ที่สร้างโดย /create-verification-skill ไม่โกหก: ส่ง read-…
📦 ทำอะไร

งานประจำที่ทำให้ verification skill และ feature map ที่สร้างโดย /create-verification-skill ไม่โกหก: ส่ง read-only subagent ต่อ feature ไล่ source, จากนั้น coordinator ขับ live ครบทุก feature หนึ่งรอบ, แล้วปิดด้วย 1 PR ของ correction ที่พิสูจน์แล้ว — จบด้วย 1 ใน 3 outcomes ที่ต้องประกาศ: clean / changed / blocked

💡 เจ๋งยังไง

จับความจริงเรื่อง rot ไว้ตรงเป้า: 'A feature map rots the moment the app changes' และกำหนดหน่วยวัดให้เหมาะ — 'The unit of rigor is the feature, not every sentence: cover every feature file from source and exercise every feature live, without terminalising every bullet' (1) ขอบเขตแก้ไขเป็นเกราะ: แก้ได้เฉพาะ directory ของ verification skill 'Never edit product code during a run: a behavior the map describes that the app no longer does is either doc drift (fix the map) or a product regression (report it, don't paper over it in docs)' (2) แยก wave ชัด: source wave เป็น subagent อ่านอย่างเดียว (ต่อ feature ไฟล์ หนึ่งตัว) ที่ไม่ขับ app ไม่แก้ไฟล์ ส่ง summary + drift + recipe, ส่วน live pass อยู่กับ coordinator คนเดียว (3) live pass คุมสาม invariant ตลอดแม้จะพัง: doctor ก่อน drive และ reset/relaunch แทนการหวัง, evidence ที่จับแล้วต้องรอดทุก cleanup (เช็คที่ path ที่ระบุ ไม่ assume), 'nothing a drive started outlives that drive's usefulness' รวมซาก iteration ที่พัง (4) มาตรฐาน unreachable: verified-unreachable ได้เฉพาะเมื่อระบุ prerequisite จริง (auth, entitlement, OS, external state) และ route ที่ลองแล้ว (5) triage แยกสามถังไม่ปน: doc drift → แก้ map, harness gap → แก้ harness แล้ว re-drive, product gap → record ให้ user เท่านั้นห้ามเอาเข้า PR

🧭 วิธีใช้

เรียกด้วย /maintain-verification-skill หรือ 'audit the verify skill' — เป็นงาน periodic หลัง app เปลี่ยน ขั้นตอน: (0) Locate target — หา project-local skill ที่มี launch/drive sections + feature map (ปกติ .zcode/skills/verify-*/) หลายตัวให้ถามเลือก ไม่มีเลยให้หยุดและชี้ /create-verification-skill แทนการแต่งเอง (1) Index hygiene — อ่าน feature map README + glob sibling แก้ entry หาย/เกิน/ซ้ำ/ตาย (2) Source wave — spawn read-only subagent ต่อ feature พร้อมกัน แต่ละตัวตอบ 'how does this user-facing feature work?' จาก source พร้อม cite drift และส่ง live-verification recipe หนึ่งชุด รูปแบบ: summary / source entry points / drift or none / one recipe (3) Reconcile — ทุก feature ต้องมี summary ครบ รวม recipe ที่ซ้อนให้เหลือ app state น้อยที่สุด spot-check drift ที่ถูก cite และสแกน churn ล่าสุดหา user-facing surface ที่หลุดจาก map (ต้องมี source path จริงก่อนเรียกว่า missing) (4) Live pass — บังคับแม้ source ดูสะอาด: ทำตาม launch model ของ skill ตัวเอง (server/UI = instance เดียวขับ serial, CLI สั้น = session สดต่อ drive) ขับทุก feature อย่างน้อยหนึ่งครั้ง, doctor fail ที่มาจาก drift ให้แก้ใต้ edit scope แล้ว retry หนึ่งครั้งก่อนยอม blocked, teardown หลัง drive สุดท้ายของ run รวม re-proof (5) Triage สามถัง (6) Ship or stop — changed: PR เดียวของ proven corrections โดยอ่านทุกไฟล์ที่แก้ซ้ำก่อน, clean/blocked: ไม่มี PR รายงาน outcome กับ coverage อย่างตรงไปตรงมา

💬 ตัวอย่าง prompt

/maintain-verification-skill — เรา refactor routes ไปสองสัปดาห์แล้ว ตรวจ verify-shop ทั้ง feature map ว่ายังตรงจริงไหม ปิดด้วย PR เดียวถ้ามีอะไรต้องแก้

⚡ Auto-invoke

ไม่ auto-invoke — frontmatter มี disable-model-invocation: true user เรียก /maintain-verification-skill เองเมื่ออยาก audit หรือจะตั้ง cadence (ไฟล์บอก 'Suggest a cadence only if they ask') — เป็นครึ่งหลังของวงจรที่ /create-verification-skill เปิด และมันจะชี้กลับไป /create-verification-skill เมื่อหา target ไม่เจอ แทนการแต่ง verification skill ขึ้นเอง

🔗 คู่กันกับ

ต้องมี verification skill ที่มี launch/drive sections + feature map อยู่ก่อน (จาก /create-verification-skill หรือ project-local skill แบบเดียวกัน) — ภายใน spawn read-only subagents ต่อ feature ใน source wave (ประเภทตาม pstack-roles.md แบบเดียวกับ fan-out อื่น) และอาศัย Doctor/Drive/Evidence/Cleanup sections ของ verification skill เป้าหมายเป็นกติกาการขับใน live pass

⚠️ เคล็ดลับ/ข้อควรระวัง
  • 'Required even when source looks clean' — ห้ามข้าม live pass เพราะอ่าน source แล้วรู้สึกโอเค
  • ห้ามแก้ product code ระหว่าง run — พฤติกรรมเพี้ยนแปลว่า doc drift (แก้ map) หรือ product regression (รายงาน) เท่านั้น
  • อย่า drive instance ที่ยังไม่ได้ health-check หลังมันทำอะไรแปลก ๆ — ถ้า doctor มองไม่เห็นปัญหา (UI wedged บน process ที่ healthy) ให้ reset เป็น known state หรือ relaunch ไม่ใช่หวังให้หายเอง
  • evidence ต้องยืนยันว่าอยู่ครบที่ path ที่ระบุหลังทุก cleanup ไม่ใช่ assume
  • verified-unreachable ต้องมี prerequisite ชัด (auth/entitlement/OS/external state) และ route ที่ลองแล้ว — ถ้า map ไม่บอก prerequisite นั่นคือ drift ต้องแก้
  • harness fix จาก triage ต้อง re-drive จริงก่อนขึ้น PR และ teardown ต้องเกิดหลัง drive สุดท้ายของ run รวม re-proof
  • product gap ให้ record แยกให้ user ไม่เอาไปปนใน PR ของ verification corrections — run notes เก็บไว้ที่ scratch อย่า commit

📜 Principles (21) — โหลดให้เองตอน /poteto-mode · เรียก redirect กลางงานด้วยชื่อ

boundary-discipline
📦 สาระ

กติกาเรื่องการจัดวาง validation และ error handling: กระจุก guard ไว้ที่ system boundary (CLI args, config, network, external API) ตัวเดียว แล้วเชื่อ internal types โดยไม่ตรวจซ้ำ ส่วน business logic ให้อยู่ใน pure functions และ shell (framework wiring) ต้องบางและ mechanical

💡 เจ๋งยังไง

Insight หลักคือ validation ที่กระจัดกระจายให้แค่ 'ความรู้สึกปลอดภัยหลอก ๆ' — มัน noise, redundant และ validate ข้อมูลเดิมซ้ำไปมา ส่วนที่แหลมคือ 3 ชั้นของ pattern: (1) ที่ boundary validate แบบ defensive, (2) ข้างในเชื่อ types ไม่ nil-check ซ้ำ, (3) 'Across the boundary' — expose domain concept ออกไป ไม่ปล่อยให้ transport/storage/wire type ของ framework รั่วผ่าน public surface พร้อมเกณฑ์ตัดสินใจตัวเอง: 'Is this data crossing a system boundary right now?' ถ้าไม่ใช่ = validation นั้น redundant และ 'Can this be a pure function that the shell just calls?' ถ้าใช่ = extract ออกมา

🧭 ใช้เมื่อไหร่ + วิธี

ใช้เมื่อ wire validation, error handling หรือ framework adapter ขั้นตอนภายใน: (1) ระบุ boundary จริงของระบบ (CLI, config file, external API, network) (2) เขียน parse/validate ณ จุดนั้นครั้งเดียว แปลง raw data เป็น domain types (3) ลบ nil-check/re-validation ที่ซ้ำใน call chain ลึก (4) ดึง business logic ออกเป็น pure functions ที่ไม่พึ่ง framework เช่น parse function เป็น pure transform จาก raw bytes สู่ typed state, prompt construction รับ structured state ให้ string ออก (5) กันไม่ให้ boundary type รั่วออกนอก module

💬 ตัวอย่าง

แก้ validation ทั้งระบบนี้ — apply boundary-discipline: validate config ที่ parse time ที่เดียว ลบ nil-check ซ้ำใน call chain แล้วดึง business logic ออกเป็น pure function ที่ shell เรียกได้

⚠️ เคล็ดลับ

อย่า re-export transport, storage, framework หรือ wire types ผ่าน public surface — นี่คือจุดที่ boundary รั่วบ่อยสุด และอย่าใส่ redundant nil check ลึกใน call chain ถ้า boundary validate แล้ว ทดสอบตัวเองด้วยคำถาม 'ข้อมูลนี้กำลังข้าม boundary จริงไหม' ก่อนเขียน guard ทุกครั้ง

build-the-lever
📦 สาระ

เมื่องานไม่ trivial ให้สร้าง tool ที่ทำงานนั้น (codemod, script, generator, หรือ delegate skill ที่ subagent อ่าน) แทนการทำมือ — lever คือ artifact เดียวที่ reviewer อ่านและ rerun ได้เพื่อตรวจงาน แทนการ redo งานที่ทำมือ

💡 เจ๋งยังไง

ให้ payoff สองชั้นจาก artifact เดียว: Throughput — script ทำงานเหมือนเดิมทุกครั้งและ rerun ฟรี, Confidence — 'trust me' กลายเป็น 'run this' เพราะ deterministic script แทนการเชื่อคำสรุป จุดที่แหลมคือกติกาปิดท้าย: 'Applying this principle produces a file. If you cited it and there is no codemod, script, generator, or delegate skill in the diff, you didn't apply it.' — บังคับให้ principle นี้ตรวจสอบได้ว่า apply จริงหรือแค่พูด และกติกา 'deterministic lever beats fan-out': ถ้า script ทำได้ในรอบเดียว ห้าม fan out delegate ไปทำมือแทนสิ่งที่ script ทำได้

🧭 ใช้เมื่อไหร่ + วิธี

ใช้กับงานที่ไม่ trivial: edits, migrations, analyses, checks ขั้นตอน: (1) ทำ unit แรกด้วยมือเพื่อเรียนรู้ recipe (2) สร้าง lever (codemod สำหรับ edit, generator สำหรับไฟล์ซ้ำ ๆ, dump-to-sqlite query สำหรับ analysis, rerunnable check สำหรับ verification) (3) พิสูจน์โดย rerun lever บน unit นั้นแล้ว diff เทียบกับเวอร์ชันมือ (4) ทำให้ rerun ได้อย่างปลอดภัยเพราะ reviewer จะรันจริง (5) ถ้า fan out ไป subagent ให้เขียน lever เป็น skill ที่ทุก delegate อ่าน — recipe + verification contract + do-not-touch fences ใน artifact เดียว และวางไว้นอก write scope ของ delegate (6) งานที่ outlive session ให้ commit lever

💬 ตัวอย่าง

ต้องแก้ field นี้ใน 200 ไฟล์ — apply build-the-lever: ทำ 1 ไฟล์ด้วยมือก่อน แล้วเขียน codemod รันทั้งชุด แล้ว diff เทียบไฟล์ที่ทำมือเพื่อพิสูจน์ว่าถูก

⚠️ เคล็ดลับ

Bar คือ triviality ไม่ใช่ repetition — งาน one-off ก็ยังคุ้มถ้า lever คือสิ่งที่ทำให้งาน checkable ได้ จุดที่คนพลาด: สร้าง lever แล้วแต่ไม่ diff เทียบกับงานมือตัวอย่าง, วาง delegate skill ไว้ใน write scope ของ delegate จน contract ถูกแก้เงียบ ๆ และลืม commit lever เมื่องานจะถูกใช้รอบหน้า

encode-lessons-in-structure
📦 สาระ

เมื่อจับตัวเองได้ว่าเขียน instruction เดิมเป็นครั้งที่สอง หรือเจอ correction ซ้ำ ๆ ให้ encode กฎเป็น mechanism (lint rule, metadata flag, runtime check, script) แทนข้อความ — ทุก error, human correction และ unexpected outcome คือ learning signal ที่ต้อง capture, route และปิด loop

💡 เจ๋งยังไง

Insight หลัก: textual instruction พึ่งความสมัครใจของผู้อ่าน (ต้อง notice, จำ, ทำตาม) แต่ structural mechanism enforce กฎโดยไม่ต้องขอความร่วมมือ จุดที่แหลมคือ 'Pick the strongest rung' — เมื่อหลาย mechanism ทำได้ ให้เลือกอันแข็งแรงสุดที่สถานการณ์ยอม (state ที่ compile ไม่ผ่าน > lint/banned API ที่ fail CI > canonical helper > runtime check) เพราะ 'agents copy whatever the surrounding code already does' — guard ที่อ่อนจะกลายเป็น template ให้คนต่อไปลอกและยิ่งอ่อนลง Corollary ที่ดุดัน: 'The instruction IS the symptom' — ถ้า fix เชิงโครงสร้างทำได้ ห้ามทาด้วย instruction

🧭 ใช้เมื่อไหร่ + วิธี

Trigger: จับตัวเองเขียน instruction เดิมครั้งที่สอง หรือเห็น correction ที่เกิดซ้ำ ขั้นตอน: (1) ถามว่ากฎนี้เป็น lint rule / metadata flag / runtime check / script ได้ไหม (2) ได้ → encode แล้ว 'ลบ' instruction เดิมทิ้ง (3) ไม่ได้ (ต้องใช้ judgment จริง) → ทำ instruction ให้เด่นขึ้นและใส่ตัวอย่าง failure mode Feedback loop: capture ทุก correction → แยกแยะ one-off (brain note) vs recurring fix (skill/lint) vs systemic issue (principle) → close the loop ด้วยการ apply ทันทีหรือสร้าง todo ที่ชัดเจน

💬 ตัวอย่าง

เรื่องนี้เตือนไปแล้ว 2 รอบ — apply encode-lessons-in-structure: ทำเป็น lint rule ที่ fail CI แล้วลบ instruction เก่าออกจาก docs

⚠️ เคล็ดลับ

Anti-pattern ที่ไฟล์ระบุชัด: (1) Acknowledging without recording — ตอบ 'I'll keep that in mind' ไม่ persist (2) Recording without routing — จด brain note เรื่อง lint ที่ควรมีแต่ไม่ implement คือขยะ (3) Fixing without generalizing — แก้จุดเดียวทิ้ง pattern ไว้ และเลือก rung ผิด: ถอยไปใช้ runtime check ทั้งที่ type กันได้ คือความผิดพลาดที่ซ้ำได้ทุก codebase

exhaust-the-design-space
📦 สาระ

เมื่อเจอ novel UI interaction หรือ architectural decision ที่ไม่มี precedent ใน codebase ให้สร้าง 2-3 competing prototypes หรือ sketches แล้วเทียบ side by side ก่อน commit — เพราะการสร้างสิ่งผิดแพงกว่าการสำรวจสามทางเลือก

💡 เจ๋งยังไง

แก้ bias ที่ agent (และคน) มักตกหลุม: ตอบคำถาม design ด้วยทางเลือกแรกที่ขึ้นในหัว กติกา 'Design it twice is this rule by another name' ทำให้ principle นี้จับต้องได้ และมีเงื่อนไขกันการโกง: 'A second flavor of the first shape does not count' — prototype ที่สองที่เป็นแค่รสชาติเดียวกับแบบแรกไม่นับว่าได้สำรวจ design space นี่คือ forcing function ให้ thinking แบบ divergent ก่อน convergent โดยมีขอบเขตเมื่อไหร่ไม่ต้องใช้ (งาน mechanical ที่ pattern ชัดอยู่แล้ว) เพื่อไม่ให้กลายเป็นการเสียเวลา

🧭 ใช้เมื่อไหร่ + วิธี

ใช้เมื่อ: novel UI interaction ที่ไม่มี prior art, architectural choice ที่มีหลายทางที่ viable, product design ที่ UX ขึ้นกับ feel ไม่ใช่ logic ไม่ใช้เมื่อ: implementation แบบ mechanical ที่ pattern ตั้งไว้แล้ว, bug fix/refactor ที่มี target state ชัด, งานที่ constraint บังคับทางเดียว ขั้นตอน: (1) ยืนยันว่ายังไม่มี precedent (2) sketch/build 2-3 ทางที่ต่างกันจริง (3) เทียบ side by side (4) ค่อย commit กับทางที่ชนะ

💬 ตัวอย่าง

Flow drag-and-drop นี้ไม่มีใครเคยทำในระบบ — apply exhaust-the-design-space: ทำ prototype 3 แบบที่ต่างกันจริงมาเทียบก่อนเลือก อย่าให้แบบที่ 2 เป็นแค่ flavor ของแบบแรก

⚠️ เคล็ดลับ

อย่าใช้กับงานที่มี precedent หรือ constraint ชัด — ไฟล์มีหัวข้อ 'When it doesn't apply' กันการใช้เกินขอบเขต จุด trap ที่ระบุตรง ๆ: prototype ที่สองที่เป็น 'a second flavor of the first shape' ไม่นับ ต้องต่างกันในระดับ shape ไม่ใช่ระดับรายละเอียด

experience-first
📦 สาระ

The product is the experience — ทุก decision ทางเทคนิคช่วยหรือทำร้ายประสบการณ์ผู้ใช้อย่างใดอย่างหนึ่ง เมื่อความสะดวกของ implementation ปะทะกับความดีใจของผู้ใช้ ให้เลือกผู้ใช้ และเลือกส่ง feature น้อยแต่ polish มากกว่าส่งเยอะแบบหยาบ

💡 เจ๋งยังไง

จุดที่แหลมกว่า principle 'UX first' ทั่วไปคือการขยายนิยาม 'user': 'The user is whoever consumes the work' — ปลายทางของ UI, เพื่อนร่วมงานที่ import library, และ engineer ที่มา maintain ต่อ ล้วนเป็น user ที่ต้องชั่งน้ำหนักเท่ากัน ทำให้ principle ผลิตภัณฑ์ใช้ตัดสิน API design และ code quality ได้ด้วย พร้อมกติกาปฏิบัติ 5 ข้อที่เป็นรูปธรรม: say no to 1,000 things (ทุก feature/control/option ต้อง 'earn its place'), ship less ship better, prototype ก่อน commit, sweat the details (transitions, alignment, spacing, feedback, error states), tighten the core loop และวางความสัมพันธ์กับ foundational-thinking ชัด: foundational กำหนด 'ลำดับ' ของงาน ตัวนี้กำหนด 'เป้าหมาย'

🧭 ใช้เมื่อไหร่ + วิธี

ใช้เมื่อ product/UX หรือ feature-scope tradeoff เข้ามาเกี่ยวข้อง ขั้นตอนที่มันสั่ง: (1) ถามว่าทุก feature/control/option คุ้มที่จะมีอยู่จริงไหม — ไม่คุ้มตัดทิ้ง (2) เลือก ship น้อยแต่ polish (3) ทำ throwaway prototype (เช่น HTML) ก่อนเขียน production (4) ลงลึกรายละเอียด: transitions, alignment, spacing, feedback, error states (5) ทุก feature ต้อง serve workflow กลาง ไม่งั้นให้ 'get out of the way' (6) เวลาอธิบาย impact ให้ยืนในมุมของ user ที่ consume งานนั้น

💬 ตัวอย่าง

ออกแบบ feature นี้ — apply experience-first: เลือกทางที่ user ดีใจที่สุดแม้ implement ยากกว่า ตัด option ที่ไม่ earn its place ทิ้ง และจัด error state ให้สวยด้วย

⚠️ เคล็ดลับ

กับดักที่พบบ่อยคือนึกว่า principle นี้ใช้เฉพาะ UI — ไฟล์ระบุว่า user คือ 'whoever consumes the work' รวมถึงคน maintain ต่อ ดังนั้น library/API design ก็อยู่ในขอบเขตด้วย และ 'say no to 1,000 things' หมายถึงการตัด feature ที่คิดว่าดีอยู่แล้วด้วย ไม่ใช่แค่ตัดของชั่วชัด ๆ

fix-root-causes
📦 สาระ

กติกา debug: ห้ามอุด symptom — trace ทุกปัญหาจนเจอ root cause แล้วแก้ตรงนั้น reproduce ก่อน, ถาม why ไล่จนสุด, ต้านแรงล่อใจที่จะเติม guard (nil check) เพื่อปิด crash

💡 เจ๋งยังไง

ให้เหตุผลทาง economics ที่ชัด: symptom fix ที่เร็ววันนี้คือดอกเบี้ยทบต้น — แต่ละ workaround ทำให้ระบบยิ่งตามใจยากและ bug จริงยังอยู่ จุดที่แหลมที่สุดคือหัวข้อ 'Restart bugs: suspect state before code': 'Code doesn't change between runs. State does.' — อาการพังหลัง restart ส่วนใหญ่คือ persistent state เก่า (config, cache, lock file, serialized state) ไม่ใช่ code ซึ่งเป็น reframing ที่ช่วยลัด debug ได้จริง และกติกา 'If a workaround needs a paragraph-long comment to justify it, the code is wrong (fix the code, not the comment)' ทำหน้าที่เป็น smell detector ที่จับต้องได้

🧭 ใช้เมื่อไหร่ + วิธี

ใช้เมื่อ debug ทุกประเภท ขั้นตอน: (1) Reproduce ก่อน — ถ้า reproduce ไม่ได้จะ verify fix ไม่ได้ (2) ถาม 'why' ไล่ลงไปจนถึง root cause (3) ห้ามเติม guard เพื่อเงียบ crash — nil check ที่ปิด crash คือ symptom fix (4) ถ้าเจอ workaround ที่ต้องเขียน comment ยาวมา justify = code ผิด (5) แก้ 'pattern ไม่ใช่ instance' — grep หา pattern เดียวกันแล้วแก้ทุกจุด (6) ติดแล้วอย่าเดา — instrument: เพิ่ม logging, อ่าน error จริง (7) กรณีพังหลัง restart: สงสัย stale persistent state ก่อน ถ้าลบ state file แล้วกลับมาปกติ ให้ตั้ง state validation เป็น fix หลัก

💬 ตัวอย่าง

Service นี้ crash หลัง restart ทุกครั้ง — apply fix-root-causes: reproduce ให้ได้ก่อน สงสัย stale state file ก่อน code แล้ว trace จนถึง root cause ห้ามเติม nil-check ปิดปาก

⚠️ เคล็ดลับ

ข้อที่คน (และ agent) พลาดบ่อยสุด: ข้ามขั้น reproduce ไปแก้เลย — ถ้า reproduce ไม่ได้ คุณจะ verify fix ไม่ได้เช่นกัน และแก้เฉพาะ instance ที่เจอโดยไม่ grep หา pattern เดียวกันทั้ง codebase กับอาการ 'พังหลัง restart' อย่าเปิด debugger ไล่ code ก่อน — เช็ค state file/cache/lock ก่อนเสมอ

foundational-thinking
📦 สาระ

กติกาเรื่องการตัดสินใจเชิงโครงสร้าง: โครงสร้างที่ถูก (structural decision) ปกป้อง option value, decision ระดับ code ปกป้องความเรียบง่าย — เริ่มจาก data structures ก่อนเขียน logic, ทำ scaffold ก่อน feature, และลำดับงานเพื่อไม่ปิดประตูทางเลือกในอนาคต

💡 เจ๋งยังไง

Insight แกน: 'Over-engineering is often a premature decision that closes doors. The right foundational data structure keeps doors open.' — ทำให้ over-engineering กับ under-engineering เป็นความผิดพลาดเดียวกัน (decision เร็วเกินไป) และให้เหตุผลเชิง economics ของ data shape: 'A data-structure change late is a rewrite. Early, it is often a one-line diff.' รวมถึง Concurrency corollary ที่ใช้ได้ทันที: ก่อน share state ระหว่าง actors ถามว่า 'ถ้าอีก actor แก้พร้อมกันจะเกิดอะไร' — ถ้าคำตอบไม่ใช่ 'nothing' ให้ isolate และปิดท้ายด้วยลำดับที่ชัด: 'Subtraction comes before scaffolding'

🧭 ใช้เมื่อไหร่ + วิธี

ใช้ก่อนเขียน logic: ตอนเลือก core types/data structures, ตอนลำดับ scaffold vs feature, ตอนถามว่า concurrent actors share อะไร ขั้นตอน: (1) กำหนด core types ให้เร็วแล้ว trace access pattern ทุกเส้น เลือก structure ที่ match dominant paths (2) ระดับ code: DRY โครงสร้าง ไม่ใช่ทุกบรรทัด — types/data models ควร converge, สาม statement คล้ายกันยังดีกว่า premature abstraction, prefer explicit over clever (3) ก่อน share state ระหว่าง actor ให้ isolate ถ้าตอบ concurrent ไม่ได้ว่า 'nothing' (4) ถ้าอะไรช่วยทุก phase หลัง ให้ทำเป็น scaffold ก่อน (CI, linting, test infra, shared types) — setup ก่อน feature, tests ก่อน fixes, commit เล็ก single-purpose (5) แต่ละ increment ต้อง land abstraction ที่ coherent หรือ deepens อันเดิม ไม่กระจาย capability ใหม่เป็น special-case ตาม callers (6) ลบ dead weight ก่อนวาง foundation

💬 ตัวอย่าง

เริ่มโปรเจกต์ใหม่ — apply foundational-thinking: กำหนด core types กับ access patterns ก่อนเขียน logic แล้ววาง scaffold (CI, test infra) ก่อน feature แรก

⚠️ เคล็ดลับ

กับดักที่ไฟล์เตือน: DRY ทุกบรรทัดจนเกิด premature abstraction — สาม statement ที่คล้ายกันยังดีกว่า abstraction ที่ยังไม่ถึงเวลา และอย่าลืมว่า late data-structure change = rewrite เต็มรูป ทำให้ช่วงเวลา 'เลือก data shape' คือ decision ที่แพงที่สุดใน project ควรใช้เวลาตรงนั้นมากที่สุด

guard-the-context-window
📦 สาระ

Context window มีจำกัดและไม่ renewable ภายใน session — ทุก token ที่เข้ามาต้องคุ้มที่จะอยู่ กติกาคือ route ของชิ้นใหญ่ (output ยาว, screenshot, เอกสาร) ไปที่ subagent แล้วเก็บแต่ summary ใน main thread ไม่ใช่ raw payload

💡 เจ๋งยังไง

จัดการ context เป็นทรัพยากรเชิง economics: 'Unlike compute or time, context spent inside a session cannot be reclaimed' — เวลากับ compute ชดเชยได้แต่ context ไม่ได้ และอธิบายผลเสียเมื่อล้น: reasoning ถดถอย, เกิด compression artifacts, งานหยุด จุดที่แหลมคือกติกา 'Keep frequently used content inline': template/reference ที่ใช้ทุกครั้งควรอยู่ในตัว skill file เอง ไม่ใช่ไฟล์แยกที่เสีย read ทุกครั้ง — แปลว่าการจัด layout ของข้อมูลเองคือการประหยัด context และกติกา 'Don't read what you won't use' ทำให้เป็นนิสัยระดับการอ่านไฟล์ได้ทันที

🧭 ใช้เมื่อไหร่ + วิธี

ใช้เมื่อ context เริ่มเต็ม: ผลลัพธ์ใหญ่, ไฟล์ยาว, repeated reads, fan-out planning ขั้นตอน: (1) Isolate large payloads — ส่ง verbose output/screenshot/เอกสารใหญ่ให้ subagent ประมวล แล้วรับแต่ summary (2) เลือกอ่านเฉพาะที่ relevant — ไฟล์ที่ไม่เกี่ยวกับ task ปัจจุบันข้ามไป (3) เก็บ template/reference ที่ใช้บ่อย inline ใน skill file (4) Size phases and cap scope — จำกัดจำนวนไฟล์ต่อ phase, ตั้ง turn budget, คิดค่าใช้จ่าย mechanism เข้าไปด้วย

💬 ตัวอย่าง

ต้องตรวจ log สรุป 500 MB — apply guard-the-context-window: ส่งงานไป subagent แล้วเอาแต่ summary และจุดที่น่าสงสัยกลับมา ห้าม dump raw log ใน main context

⚠️ เคล็ดลับ

จุดพลาดที่พบบ่อย: อ่านไฟล์ทั้งไฟล์ทั้งที่ต้องการแค่ส่วนเดียว และวาง template ไว้ไฟล์นอก skill ทำให้เสียค่า read ทุกครั้งที่ใช้ อย่าลืมว่า mechanism ที่เพิ่มเข้ามา (เช่น tool call, fan-out) ก็มีค่า context ของมันเอง — ต้อง account ใน turn budget ด้วย

laziness-protocol
📦 สาระ

เพราะการเขียน code ถูกสำหรับ agent จน over-engineering ง่ายเกินไป ให้ยืม 'ความเหนื่อยล้าของ human maintainer' มาใช้เป็นแรงต้าน — bias ไปทางการลบ และการเปลี่ยนแปลงที่เล็กที่สุดที่แก้ปัญหาได้ หาผลลัพธ์มากที่สุดจาก code กับ complexity น้อยที่สุด

💡 เจ๋งยังไง

แก้ bias เฉพาะของ AI coding โดยตรง: 'Writing code is cheap for you, which makes over-engineering easy' — ความ lazy ที่นี่ไม่ใช่ขี้เกียจแต่คือ discipline ที่เลียนแบบ fatigue ของคน maintain ต่อ จุดที่แหลมคือเกณฑ์วัดที่จับต้องได้: 'If answering a question requires tracing through more than 3 files or layers, flatten it' และ 'Maintain a flat call hierarchy' โดยเจาะแยกว่า 'A rich interface that hides substantial work is not a deep call chain' (ความ lazy ไม่ได้แปลว่าเลี่ยง abstraction ที่คุ้ม) ปิดด้วย Prime directive ที่เด็ดมาก: 'If a human developer would find the code exhausting to maintain, it is a bad solution.' — ความเหนื่อยของคนอ่านคือเกณฑ์ตัดสินสุดท้าย

🧭 ใช้เมื่อไหร่ + วิธี

ใช้เมื่อ refactor, ประเมินขนาด diff หรือรู้สึกอยากเพิ่ม abstraction/layer/signal threading กติกา 6 ข้อ: (1) Prefer deletion — ตอนถูกขอ refactor หาสิ่งที่ลบได้ก่อนเพิ่ม (2) รักษา call hierarchy ให้แบน — trace เกิน 3 ไฟล์/layers = ต้อง flatten (3) Consolidate decisions — choice เดียวห้ามซ้ำหลายที่ ให้อยู่หลัง source of truth เดียวแล้วส่ง flag ง่าย ๆ (4) Minimize the diff — เล็กที่สุดที่แก้ได้ น้อยบรรทัดชนะ boilerplate 'สวย' (5) Question the threading — ถ้า task บังคับให้ thread signal ใหม่ผ่าน types/schemas/pipelines หลายชั้น ให้หยุดหาทางตรงกว่า (6) Sweat the small leaks — กำจัด pass-through เล็ก ๆ, representation leak, duplicated choice ก่อนมันแพร่เพราะ small leaks compound เป็น coordination cost ถาวร

💬 ตัวอย่าง

Refactor module นี้ — apply laziness-protocol: หาอะไรที่ลบทิ้งได้ก่อน ทำ diff เล็กที่สุด และถ้า task นี้ต้อง thread signal ใหม่ผ่าน 4 ชั้น ให้เสนอทางตรงกว่ามาให้ผมดู

⚠️ เคล็ดลับ

อย่าตีความ lazy ผิด: rich interface ที่ซ่อนงานหนักไว้ข้างหลังไม่ใช่ deep call chain ที่ต้อง flatten — สิ่งที่ต้อง flatten คือ indirection ที่ไม่ซ่อนอะไร และระวัง 'small leaks': pass-through จิ๋ว ๆ ที่คิดว่าไม่เป็นไร มันคือจุดที่ coordination cost เริ่มทบ ควรเก็บก่อนแพร่

make-operations-idempotent
📦 สาระ

ออกแบบ operation ทุกชนิดที่ mutate state (command, lifecycle step, processing loop) ให้ converge ไปที่ end state เดิมเสมอ ไม่ว่าจะรันกี่ครั้งหรือเริ่มจาก state ไหน — ทุก operation ต้องตอบได้ว่า 'รันสองรอบจะเป็นอย่างไร, รอบก่อน crash กลางทางแล้วจะเป็นอย่างไร'

💡 เจ๋งยังไง

หยิบ crash/restart/retry ออกจากหมวด 'เหตุการณ์ผิดปกติ' มาเป็นเงื่อนไข design ปกติ: 'If partial state changes the next run's outcome, every restart becomes a debugging session' — idempotency คือสิ่งที่ทำให้ restart กลับมาเป็นเรื่องน่าเบื่อ (ดี) จุดที่แหลมคือ 3 คำถามทดสอบที่ตรวจได้ทันที และเกณฑ์ตัดสินที่ชัดเจน: ถ้าคำตอบอันไหนคือ 'ขึ้นกับ state ที่เหลือไว้' operation นั้นต้องเพิ่ม reconciliation step พร้อม pattern ที่ลอกได้จริง เช่น self-healing locks ที่ใช้ PID-based stale lock detection และ content-based cleanup ที่เทียบด้วย content equivalence ไม่ใช่ creation order

🧭 ใช้เมื่อไหร่ + วิธี

ใช้ตอน design command, lifecycle step หรือ processing loop ที่ต้องทำงานท่ามกลาง crash, restart, retry ขั้นตอน: (1) ถามคำถามตั้งต้น: รันซ้ำจะเป็นอย่างไร, crash กลางทางแล้วรอบหน้าจะเป็นอย่างไร (2) Convergent startup — ตอนเริ่มให้ scan existing state, clean stale artifacts, adopt live sessions (3) Content-based cleanup — เทียบด้วยเนื้อหา ไม่ใช่ลำดับการสร้าง (4) Self-healing locks — ตรวจ stale lock ด้วย PID (5) Idempotent scheduling — งานที่ fail respawn ได้สะอาด, regenerate input ใหม่ทุก cycle (6) รัน 3 คำถามทดสอบ: รันสองรอบติดกัน? crash ได้ทุกจุดแล้วรอบหน้าโอเค? re-execution converge ที่ state เดิม? ถ้าไม่เวิร์ก ให้เพิ่ม reconciliation step

💬 ตัวอย่าง

เขียน startup script ที่อาจโดน restart กลางทาง — apply make-operations-idempotent: ทำให้รันซ้ำได้โดยไม่สะสม state ค้าง มี stale lock detection และ cleanup แบบเทียบ content

⚠️ เคล็ดลับ

คำถามทดสอบ 3 ข้อต้องตอบครบทุก operation ที่ mutate state — โดยเฉพาะข้อ 2 (crash ทุกจุดที่เป็นไปได้) ซึ่งคนมักข้าม และระวังการ cleanup ด้วย creation order/timestamp เพราะมันพังทันทีเมื่อรอบก่อน crash — ต้องใช้ content-based เสมอ

migrate-callers-then-delete-legacy-apis
📦 สาระ

เมื่อตัดสินใจแล้วว่า internal API ใหม่คือ design ที่ถูก ให้ inventory callers, migrate ทั้งหมด และลบ API เก่าใน refactor wave เดียวกัน — ไม่รักษา compatibility layer ไว้เพียงเพราะ caller เดิมยังมีอยู่

💡 เจ๋งยังไง

ตรงประเด็นกับนิสัยเฉพาะของ AI และ monorepo: ค่า default ของ agent มักเป็น 'สร้างใหม่เพิ่มเข้าไปโดยไม่แตะของเก่า' ซึ่งไฟล์นี้ตั้งชื่อผลร้ายได้แม่น: 'Keeping both old and new APIs creates dual-path complexity, slows cleanup, and makes the codebase feel append-only' — codebase กลายเป็นกอง append-only ที่ไม่มีใครกล้าลบ จุดที่แหลมคือการวางเงื่อนไขใช้ที่ชัดเจน (ใช้ได้เมื่อไม่มี external users ต้อง backward compatibility, project รับ coordinated breaking change ได้, อยู่ใน simplification initiative) พร้อมกติกาเรื่อง test: อัปเดต test ให้ assert new contract และ 'ลบ' test ที่ปกป้องแค่ implementation ก่อน refactor

🧭 ใช้เมื่อไหร่ + วิธี

ใช้เมื่อเปิดตัว internal API ใหม่ทั้งที่ caller เก่ายังอยู่ ขั้นตอน: (1) ยืนยันเงื่อนไข: ไม่มี external user พึ่ง compatibility, project รับ breaking change พร้อมกันได้ (2) Inventory callers ทั้งหมด (3) Migrate caller ทุกตัวไป API ใหม่ (4) ลบ API เก่าทันทีใน wave เดียวกัน — ไม่ทิ้ง dual path (5) Adapter ชั่วคราวเป็นข้อยกเว้นที่ต้อง time-box เท่านั้น ไม่ใช่ default architecture (6) อัปเดต test ให้ assert new contract และลบ test ที่เกาะ implementation เดิม

💬 ตัวอย่าง

API ใหม่พร้อมใช้แล้ว — apply migrate-callers-then-delete-legacy-apis: หา caller เก่าทั้งหมด ย้ายให้ครบ แล้วลบ API เก่าใน PR เดียวกัน ไม่ต้องทิ้ง wrapper แบบ deprecated

⚠️ เคล็ดลับ

ขอบเขตสำคัญ: principle นี้ใช้กับ internal API เท่านั้น — ถ้ามี external users ที่พึ่ง backward compatibility อยู่ ไม่ต้องใช้ และระวังคำว่า 'temporary adapter': มันต้อง time-boxed เสมอ มิฉะนั้น adapter ชั่วคราวจะกลายเป็นสถาปัตยกรรมถาวรโดยไม่มีใครตัดสินใจ

minimize-reader-load
📦 สาระ

Maintainability คือ 'งานที่คนอ่านต้องทำเพื่อเข้าใจ code' — วัดด้วยสองแกน: (1) จำนวน layers/indirection ที่ต้อง trace ระหว่างคำถามกับคำตอบ (2) ปริมาณ hidden/mutable state ที่คนอ่านต้องจำในหัว แล้ว collapse layer ที่ไม่คุ้มและ shrink state scope

💡 เจ๋งยังไง

แทนที่ proxy metrics ทั้งหมด (LOC, cyclomatic complexity, 'clean architecture') ด้วยสิ่งเดียวที่สำคัญจริง: reader load — 'Code is read far more than it is written' จุดที่แหลมที่สุดคือการชี้ว่าสองแกนเป็นอิสระต่อกัน: 'A flat file with 50 globals can be as hard to reason about as a 6-layer adapter stack' — แก้แค่แกนเดียวไม่พอ และเกณฑ์ทดสอบที่จับต้องได้: 'Can a new reader answer "where does X come from?" and "what can change X?" in under 30 seconds?' พร้อมกติกากันเกินงามตอนเพิ่มของ: 'does this reduce reader load somewhere else by at least as much?'

🧭 ใช้เมื่อไหร่ + วิธี

ใช้ตอน review หรือ reshape code ที่ trace ยาก ขั้นตอน: (1) นับ layers ระหว่างคำถามกับคำตอบ และนับ hidden state ในหัวผู้อ่าน (2) Collapse layer ที่ไม่ earn their keep: wrapper ที่มี caller เดียว, adapter ที่ไม่มี implementation ที่สอง, indirection สำหรับอนาคตที่ไม่มา — inline ทิ้ง (3) บังคับให้ adjacent layers เปลี่ยน abstraction — layer ที่ repeat method/argument เดิมคือ pass-through ที่ collapse ได้ (4) Demand interface compression — interface กว้างที่ซ่อน complexity น้อยทำให้คนอ่านต้องเรียนทั้ง surface และ implementation (5) Shrink state scope: pure functions > locals > fields > module state > globals, derive แทน sync (6) Name invariant ที่ boundary จุดเดียว (7) ก่อนเพิ่ม layer หรือ state ใด ๆ ถามว่าลด reader load ที่อื่นอย่างน้อยเท่ากันหรือไม่

💬 ตัวอย่าง

โค้ดชุดนี้ผม trace ไม่ออก — apply minimize-reader-load: collapse wrapper ที่มี caller เดียว inline pass-through layer แล้วดึง mutable state ลง scope เล็กสุดให้ผมดู diff ก่อน

⚠️ เคล็ดลับ

กับดัก: โค้ดที่ดู 'แบน' อาจแย่เท่ากับ code ซ้อน 6 ชั้นถ้ามี global 50 ตัว — ต้อง guard สองแกนพร้อมกัน และอย่าเข้าใจว่า 'clean architecture' อัตโนมัติ = ดี: interface กว้างที่ซ่อนอะไรไม่กี่อย่างบังคับให้ผู้อ่านเรียนรู้ทั้ง surface และ implementation ทำให้แย่กว่าไม่มี interface เลย

model-the-domain
📦 สาระ

Encode domain จริงลงใน data structure เดียว (state machine, typed model, registry, reducer ฯลฯ) แทนการกระจายกฎเป็น conditionals ข้ามไฟล์ — structure ที่ match domain ทำให้ invalid states สร้างไม่ได้และลบ branches ทิ้งได้เป็นชุด

💡 เจ๋งยังไง

ให้ economics ที่ชัดของการเลือก structure: 'Choosing it at write time is cheap; recovering it later reads as a refactor and gets deferred' — คือเหตุผลที่ codebase เน่าแบบ scatter boolean เพราะการแก้ทีหลังหน้าตาเป็น refactor ใหญ่ที่ถูกเลี่ยงตลอด จุดที่แหลมคือ taxonomy ของ structure ที่ map กับอาการตรง ๆ (boolean กระจัดกระจาย → state machine, branch ข้ามไฟล์ → registry/discriminated union, module แบบ load-validate-transform-save → module รอบ domain knowledge เพราะ 'Execution order is not ownership') และ tell ที่วินิจฉัยตัวเองได้: feature ใหม่ทำให้ if/else โตขึ้น 1 branch หรือต้องเพิ่ม 'boolean ที่สองที่ต้อง sync กับตัวแรก'

🧭 ใช้เมื่อไหร่ + วิธี

ใช้เมื่อเขียน stateful logic, code branch เยอะ หรือเห็น shape assumption ซ้ำข้ามไฟล์ ขั้นตอน: (1) หา structure จาก list ที่ไฟล์ให้: state machine แทน scattered booleans, typed object แทน loose parameters, map/registry/lookup/discriminated union แทน branch ข้ามไฟล์, reducer/command-event แทน ad hoc mutation, module รอบ domain knowledge แทน sequence phase, module boundary เล็ก ๆ รอบ behavior ซ้ำ, queue/cache/index/graph ตาม access pattern (2) ถ้าไม่ตรงกรณีไหน ให้หาว่า 'code ห้ามยอมให้เกิดอะไร' และข้อมูลถูกอ่านอย่างไร แล้วออกแบบ structure ที่ encode พอดีนั้น (3) อย่า force abstraction — code เรียบ ๆ ที่ shape ปัจจุบันชัดและไม่น่าโต ให้อยู่อย่างนั้น (4) ระวัง abstraction ที่เพิ่ม indirection แต่ไม่ลบ branches/duplicated rules/invalid states/lifecycle risk

💬 ตัวอย่าง

Module นี้มี 3 boolean ที่ต้อง sync กัน — apply model-the-domain: เปลี่ยนเป็น state machine ที่ invalid combination สร้างไม่ได้ แล้วลบ branches ที่ไม่จำเป็นทิ้ง

⚠️ เคล็ดลับ

สอง tell ที่ไฟล์ระบุว่า 'คุณข้าม principle นี้ไปแล้ว': (1) feature ใหม่ทำให้ if/else chain โตขึ้นอีก 1 branch (2) boolean ที่สองที่ต้อง sync กับตัวแรก และระวัง temporal decomposition — module ที่ตั้งชื่อตาม phase (load, validate, transform, save) จะ repeat domain rules เดิมทุก phase เพราะ execution order ไม่ใช่ ownership

never-block-on-the-human
📦 สาระ

Human supervise แบบ asynchronous — agent ต้องไม่ค้างรอ: ตัดสินใจที่ reasonable, ทำต่อไป, present ผลแล้วให้ human course-correct ทีหลัง สงวนคำถามยืนยันไว้เฉพาะ action ที่ย้อนกลับไม่ได้

💡 เจ๋งยังไง

อยู่บน tradeoff ที่ตรงมากสำหรับ AI workflow: 'Code is cheap. Waiting is expensive.' — implementation ที่ผิดแก้ใช้เวลาไม่กี่นาที แต่ agent ที่ block รอเสีย attention ของคนเป็นค่าตอบแทนจริง และเหตุผลเบื้องหลังคือ 'ทุก permission pause ทำให้ human เป็น bottleneck ของ pipeline' จุดที่แหลมคือ pattern 'Proceed, then present': ห้ามถาม 'should I do X?' แต่ให้ 'ทำ X แล้วอธิบายว่าทำไม' พร้อม boundaries ที่วาดเส้นไว้ชัดเจนเพื่อไม่ให้ principle นี้กลายเป็นข้ออ้างไร้สมอง: irreversible action ต้องยืนยัน, product direction เป็นของ human — ที่ห้าม block คือ execution เท่านั้น

🧭 ใช้เมื่อไหร่ + วิธี

ใช้เมื่อใจอยากถาม 'should I do X?' ทั้งที่งานเป็น reversible ขั้นตอน: (1) Proceed, then present — ทำงานเสร็จแล้ว show result พร้อมเหตุผล ไม่ถามก่อนทำ (2) ถามเฉพาะ genuine ambiguity — จริง ๆ ไม่สามารถ infer intent จาก context ได้เท่านั้น (3) ทำระบบให้ self-healing — เห็นปัญหาให้ log ไว้และ fix ในรอบถัดไป (4) ออกแบบ workflow สำหรับ review-after-the-fact: human ดู plans/diffs/changes ตาม schedule ของตัวเอง (5) จำกัน: irreversible (force-push, ลบ production data, ส่ง message ภายนอก) ต้องขอ confirmation — reversible (เขียน code, แก้ note, แบ่ง task) ทำต่อเลย

💬 ตัวอย่าง

อย่าหยุดถามผมทีละขั้น — apply never-block-on-the-human: ทำตาม best judgment ไปเลย สรุป diff ทั้งหมดให้ผม review ตอนจบ เว้นแต่จะแตะ production data หรือ action ที่ undo ไม่ได้

⚠️ เคล็ดลับ

เส้นแบ่งที่ต้องจำ: 'Product direction comes from the human; execution should not block' — principle นี้ห้ามใช้เป็นข้ออ้างตัดสิน product หรือทำ irreversible action ตามใจ และ 'Reserve questions for genuine ambiguity' หมายถ้า infer จาก context ได้ ห้ามถาม แต่ถ้า infer ไม่ได้จริง ๆ ให้ถาม ไม่ใช่เดามั่ว

outcome-oriented-execution
📦 สาระ

Optimize ที่ 'end state ที่ตั้งใจและตรวจสอบได้' แทนการรักษา intermediate state ให้ลื่นทุกขั้น — สำหรับ planned rewrite/migration ที่มี phase boundary ชัด อนุญาตให้จุดกลางพังชั่วคราว (planned, scoped, reversible) ขอแค่ต้อง verify ครบที่ boundary

💡 เจ๋งยังไง

แก้ปัญหาที่คนทำ migration ต้องเจอจริง: 'Keeping every intermediate step fully stable often creates temporary compatibility code that becomes long-lived debt' — ความปลอดภัยปลอมของการให้ทุก commit รันได้ แพงกว่าที่คิดเพราะ compatibility code ชั่วคราวกลายเป็นหนี้ถาวร จุดที่แหลมคือการวาง guardrails รอบ principle ที่ดูเสี่ยง: ใช้เฉพาะ rewrite/migration ที่วางแผนไว้ ต้อง declare จุดที่ยอมให้พักล่วงหน้า, เก็บ high-signal checks ในพื้นที่ที่กำลังแตะไว้ระหว่างทาง และ 'Always run final verification before declaring done' เป็นข้อบังคับที่ไม่มีข้อยกเว้น — คือสมดุลระหว่างความเร็วกับความถูกต้องที่ออกแบบมาไม่ให้ถูกใช้ในทางที่เสีย

🧭 ใช้เมื่อไหร่ + วิธี

ใช้กับ planned rewrite และ migration ที่มี explicit phase boundaries ขั้นตอน: (1) ตั้งเป้าที่ end-state integrity ไม่ใช่ transitional stability (2) Declare ล่วงหน้าว่าจุดไหน temporary breakage ยอมรับได้ (3) ระหว่าง migrate เก็บ high-signal checks เฉพาะบริเวณที่กำลังแตะ (4) ที่ plan completion ต้องรัน static + runtime verification ครบก่อน declare done ไม่ใช้กับงานที่ไม่มี phase boundary ชัดหรือไม่มีแผน reverse

💬 ตัวอย่าง

Migration ORM นี้แบ่ง 3 phase — apply outcome-oriented-execution: declare ให้ชัดว่า phase 2 ยอมให้ build พังชั่วคราวได้ แต่ต้องรัน static + runtime verification ครบก่อนปิด phase สุดท้าย

⚠️ เคล็ดลับ

ข้อจำกัดที่คนพลาด: principle นี้ไม่ใช่ใบอนุญาตให้พังมั่ว — breakage ต้อง planned, scoped, reversible ทั้งสามข้อ และต้องเก็บ high-signal check ในพื้นที่ที่กำลังแตะไว้ระหว่างทาง ไม่ใช่ปิด test ทิ้งหมดแล้วรอตรวจตอนจบ

prove-it-works
📦 สาระ

กติกาปิดงาน: verify ผลลัพธ์ทุก task ด้วยการดู 'ของจริง' โดยตรง — รัน feature, อ่านค่าจริง, inspect diff — ห้าม infer จาก proxy, self-report หรือคำว่า 'it compiles' และถ้าทำได้ให้ script เช็คนั้นเป็น deterministic proof

💡 เจ๋งยังไง

จัดหมวดความเสี่ยงได้เป๊ะ: 'Unverified work has unknown correctness' และ 'Agents report what they intended, not always what happened' — ประโยคหลังคือ insight ที่แทบไม่มี guideline ไหนพูดตรงขนาดนี้เรื่อง delegation: ต้อง trust artifacts ไม่ใช่ self-reports ของ delegate จุดที่แหลมคือรายการ proxy ที่หลอกตัวเอง (file mtimes, output freshness, agent self-reports, cached screenshots) พร้อมเหตุผลว่าทำไมคนเลือกใช้ (รู้สึกถูกกว่า) เทียบกับค่าจริงเมื่อ inference ผิด และลำดับไขปัญหาที่ดี: 'When verification fails, suspect the observation method before suspecting the system' ปิดด้วยการยกระดับ: หนึ่งครั้งที่ตาดูดีที่สุดก็แพ้ script ที่ re-run comparison เดิมได้

🧭 ใช้เมื่อไหร่ + วิธี

ใช้หลังจบ task ทุกชนิด ก่อน declare done ขั้นตอน: (1) ถามตัวเอง 'how do I prove this actually works?' (2) เช็คของจริงไม่ใช่ proxy — ตรวจ process liveness ตรง ๆ, อ่าน actual value ไม่ใช่ cached/derived (3) Code/feature: build (จำเป็นแต่ไม่พอ) → รันและ exercise feature path จริง → เช็ค full chain ว่า data flow จาก input ถึง output → integration ทดสอบ communication path แบบ end-to-end (4) งาน delegate: inspect artifact จริง (git diff, file contents, runtime behavior) ไม่ใช่คำสรุปของ delegate (5) ถ้าทำได้ เขียน script ที่ re-run comparison เดิมได้ แล้วเก็บ output เป็น artifact ให้คน re-run ต่อ (6) เก็บ artifact ให้มองเห็นเสมอ; commit เฉพาะงานใหญ่/ซับซ้อนที่ต้อง audit ย้อนหลัง (เช่น big port/migration — งานของ skill show-me-your-work)

💬 ตัวอย่าง

เสร็จแล้วใช่ไหม — apply prove-it-works: รัน feature จริงให้ผมเห็น data flow ตั้งแต่ input ถึง output และเขียน script เช็คเทียบผลเก่ากับใหม่ไว้ให้ผม rerun ได้ อย่าบอกแค่ว่า compile ผ่าน

⚠️ เคล็ดลับ

สองจุดที่พลาดบ่อย: (1) พอ verify fail ก็สงสัยระบบทันที — ไฟล์สั่งให้สงสัย observation method ของตัวเองก่อน (2) เชื่อ summary ของ subagent — ต้องเปิด git diff/ไฟล์/runtime ดูเองเสมอ และอย่าลืมว่า build ผ่านคือ 'necessary but not sufficient' เท่านั้น

redesign-from-first-principles
📦 สาระ

เมื่อต้อง integrate requirement ใหม่เข้า design เดิม ห้าม bolt-on ทับ — ให้ redesign ใหม่ราวกับ requirement นั้นเป็น assumption พื้นฐานตั้งแต่วันแรก ผลลัพธ์ควรหน้าเหมือนสิ่งที่เราจะสร้างถ้ารู้ตั้งแต่แรก

💡 เจ๋งยังไง

แก้รูปแบบความเสื่อมโทรมของ codebase ที่ชัดที่สุดรูปหนึ่ง: การ bolt-on feature ใหม่ทับ design เดิมจนโครงสร้างบิดเป็นเหนียวเหนอะ จุดที่แหลมคือการกำหนดคุณภาพผลลัพธ์เป็นเงื่อนไขตรวจได้: 'The result should look like what we would have built if we'd known on day one' — นี่คือ bar ที่ทำให้ตัดสินได้ว่างานสำเร็จหรือแค่ 'มันทำงานได้' และกระบวนการที่มีระเบียบครบทั้งอ่านก่อน (holistic understanding), คิดใหม่ (thought experiment วันแรก), propagate ครบทุก reference (types, docs, examples, rationale sections) และหัวใจที่ซ่อนอยู่: 'Think about the redesign holistically, then deliver it incrementally' — คิดทีเดียวใหญ่ ส่งเป็นชิ้นเล็ก

🧭 ใช้เมื่อไหร่ + วิธี

ใช้เมื่อต้องรวม requirement/design ใหม่เข้ากับระบบที่มีอยู่ ขั้นตอนตามไฟล์: (1) อ่านไฟล์ที่กระทบทั้งหมด แล้วเข้าใจ current design แบบ holistic ก่อน (2) ถามคำถามแกน: 'ถ้าเราเขียนจากศูนย์โดยมี requirement นี้ตั้งแต่ต้น เราจะสร้างอะไร?' (3) Propagate change ผ่านทุก reference — types, docs, examples และ rationale sections ด้วย (4) คิด redesign แบบ holistic แต่ส่งมอบเป็น increments และไฟล์ระบุว่านี่คือวิธี 'preserve option value' เวลา integrate change

💬 ตัวอย่าง

ต้องรองรับ multi-tenant — apply redesign-from-first-principles: อ่าน design เดิมทั้งหมดก่อน แล้วเสนอแบบที่ดูเหมือนเราออกแบบ multi-tenant ตั้งแต่วันแรก ไม่ใช่ if-tenant ทับทุกจุด

⚠️ เคล็ดลับ

ความผิดพลาดที่พบคือคิดใหม่แล้วลงมือทำทีเดียวทั้งก้อน — ไฟล์ระบุชัด 'think about the redesign holistically, then deliver it incrementally' อีกจุดคือลืม propagate ไปที่เอกสาร/rationale ซึ่งไฟล์ระบุเองว่าต้องไปถึง types, docs, examples, rationale sections ครบ

separate-before-serializing-shared-state
📦 สาระ

เมื่อ actors หลายตัวอาจเขียนลง state เดียวกัน (ไฟล์, branch, key, state object) ให้กำจัด sharing ออกก่อนเป็น default — serialize เชิงโครงสร้าง (lockfile, sequential phase, single writer) เฉพาะเมื่อ single shared writer เป็น invariant จริง

💡 เจ๋งยังไง

เปลี่ยนลำดับคิด default ของเรื่อง concurrency: คนมักกระโดดไป 'ต้องมี lock' ทันที แต่ไฟล์นี้กำหนดให้ lock เป็น design smell ที่ต้องเช็ค ('Treat "we need a lock" as a design smell to check, not as the default answer') และตัดทางลัดที่ทุกคนหลงใช้ทิ้งเลย: 'Instructions and conventions are not concurrency control' — บอก agent หรือ goroutine ให้ 'take turns' ไม่เวิร์ก เพราะ race มัน intermittent และแก้ยาก จุดที่แหลมที่สุดคือตัวอย่างเฉพาะเจาะจง: สอง worker เขียน field lastX ของตัวเองลง state.json ไฟล์เดียว 'ยังเป็น shared mutation อยู่' แต่ indexer-state.json + metrics-state.json แยกไฟล์ไม่ใช่ — ทำให้เส้นแบ่งจับต้องได้ทันที

🧭 ใช้เมื่อไหร่ + วิธี

ใช้เมื่อ concurrent actors อาจเขียนร่วมกัน ขั้นตอนตามไฟล์: (1) Identify shared mutable state — ไฟล์ที่ทั้งอ่านทั้งเขียน, branch ที่ทั้ง push, API ที่ทั้ง define และ consume (2) Default: eliminate the shared write target — ถามว่า actors ต้องการ object เดียวกันจริงหรือแค่ publish independent facts ให้แต่ละตัวมี file/key/branch/state directory ของตัวเอง แล้ว merge ที่ read/reporting boundary เท่านั้น (3) เฉพาะเมื่อ single shared write target เป็น invariant จริง ค่อย serialize structurally: lockfiles, sequential phases, single-writer actor, atomic compare-and-swap ห้ามใช้ instruction/convention เป็นตัวควบคุม

💬 ตัวอย่าง

Worker 2 ตัวเขียนลง state.json ไฟล์เดียวกันแล้วพังเป็นบางครั้ง — apply separate-before-serializing-shared-state: แยกเป็นไฟล์ของใครของมันแล้ว merge ตอนอ่าน ถ้าต้องใช้ lock จริงค่อยให้เสนอเหตุผล

⚠️ เคล็ดลับ

จุดที่ผิดบ่อยสุดคือนับขั้นตอนข้าม: ดำดิ่งไปแก้ด้วย lock ทันทีที่เจอ race — ต้องถามก่อนว่า actors เนี้ย 'publish independent facts' หรือเปล่า และระวัง false comfort จากการแยก field ภายในไฟล์เดียว: lastX ของสอง worker ใน state.json เดียวยังเป็น shared mutation เต็มรูป

sequence-verifiable-units
📦 สาระ

จัดงานหลายขั้น (sweep, migration, ชุด edit คล้ายกัน) เป็น sequence ของ unit เล็กที่แต่ละ unit จบด้วย state ที่ตรวจได้ — verify ก่อนไปต่อ ห้าม batch แล้วตรวจรวมตอนจบ และจัดลำดับ commit/PR ให้ sequence เล่าเรื่องพิสูจน์ตัวเอง (failing test ก่อน fix)

💡 เจ๋งยังไง

อยู่บน observation เชิง debug economics ที่แม่น: 'A break caught at the unit that caused it is cheap to localize. A break caught after a batch is buried, and you have already built further on a broken base.' — เหตุผลที่การ verify ทีละ unit ถูกกว่าไม่ใช่ความขยัน แต่คือต้นทุนการ localize จุดที่แหลมคือการ map discipline เดียวกันขึ้น 'สองระดับความสูง' (two altitudes): execution (ทำงานจบด้วย check) และ delivery (จัด stack commit/PR ให้ reviewer replay ได้) พร้อม canonical shape ที่ชัด: 'The canonical shape is the failing test first, then the fix on top' — red ก่อน green ทำให้ reviewer เห็นทั้งปัญหาและหลักฐานว่าแก้จริง ปิดด้วยประโยคสวย: 'Each commit lands on its own and the sequence reads as an argument.'

🧭 ใช้เมื่อไหร่ + วิธี

ใช้กับงาน multi-step และการจัด commit/PR ขั้นตอน: (1) เลือก unit เล็กสุดที่จบด้วย check — edit หนึ่งคู่ test ของมัน หรือ commit ที่ยืนเดี่ยวได้ (2) Verify ก่อนไปต่อ — red-to-green ต่อ unit ห้ามเลื่อนไปรวมตอนท้าย (3) Rebase บน clean trunk ก่อนเพื่อให้ทุก check วัดกับ baseline จริง (4) แม้ lever (script) ทำ edit ให้และ check แทบฟรี ก็ต้องรันทุก unit (5) ฝั่ง delivery: เรียง commit ให้ sequence สร้าง confidence ด้วยตัวเอง — failing test ก่อน fix, subtraction ก่อน reshape, baseline capture ก่อน treatment, scaffold ก่อน feature

💬 ตัวอย่าง

Migration 300 ไฟล์ — apply sequence-verifiable-units: แบ่งเป็น unit เล็ก แต่ละอันจบด้วยการรัน check แล้วถึงไปต่อ เรียง commit ให้เป็น failing test ก่อนแล้ว fix ตาม

⚠️ เคล็ดลับ

กับดักที่ไฟล์จับโดยตรง: การ batch edits แล้ว verify ครั้งเดียวท้ายงาน — ถ้าพังจะ localize ไม่ได้และ build ต่อบน base พังไปแล้ว และอย่าลืม rebase บน clean trunk ก่อน มิฉะนั้นทุก check วัดกับ baseline ปลอม ส่วนฝั่ง delivery อย่าเรียง commit ตามลำดับที่ตัวเองแก้สะดวก แต่ตามลำดับที่ reviewer อ่านแล้วเชื่อ

subtract-before-you-add
📦 สาระ

เมื่อระบบกำลังจะโต ให้ลบ complexity ออกก่อนแล้วค่อยสร้างเพิ่ม — deletion ให้ base ที่เรียบง่ายกว่า ทำให้ addition ต่อไปเล็กลงและเปราะน้อยลง พร้อมทำ simplification เป็นการลงทุนต่อเนื่อง ไม่ใช่ event ครั้งเดียว

💡 เจ๋งยังไง

เหตุผลของมันเป็นคณิตศาสตร์ของ complexity: 'Adding to a complex system compounds complexity. Removing first cuts the surface area, reveals the essential structure, and usually makes the next design obvious.' — การลบก่อนไม่ได้แค่ทำให้เล็ก แต่ทำให้ design ที่ถูก 'เห็นตัวเอง' จุดที่แหลมคือกติกา 'Design for observed usage, not speculative edge cases' และการเชื่อมโยงที่ลึก: 'Out-of-spec features drag validators behind them. Persistence, retry-on-startup, and schema migration each need guards to defend their inputs.' — ฟีเจอร์เกิน spec ไม่ได้แค่ฟรี แต่ลาก guard ทั้งชุดตามมาที่ต้อง maintain ตลอดไป และกติกาปิดที่ดุ: reference ที่ไม่มีเนื้อหาใหม่ ให้ 'ลบ' ไม่ใช่เก็บเป็น stub

🧭 ใช้เมื่อไหร่ + วิธี

ใช้ตอนวางลำดับ addition, refactor หรือ rewrite ขั้นตอนตามไฟล์: (1) Sequence removal before construction — ลบก่อนสร้างเสมอ (2) Cut before you polish — ลดให้ถึง minimum ก่อนลงทุนกับคุณภาพ (3) Design ตาม observed usage ไม่ใช่ speculative edge cases (4) ห้าม speculative validators/parsers/guards เกินที่ spec ต้องการ (5) ตัดฟีเจอร์นอก spec — มันลาก guard ตามมา (6) Simplify prompts — ลบ instruction ซ้ำและ template เกินจำเป็น (7) reference ที่ไม่มีเนื้อหาใหม่ → ลบทิ้ง ไม่ทิ้ง stub และยึดจุดหมายถาวร: ทิ้ง design ไว้ 'เรียบง่ายขึ้นและ capable ขึ้นหลัง surface เดิมหรือเล็กกว่า' เทียบกับตอนรับงาน

💬 ตัวอย่าง

ก่อน implement feature ใหม่ใน module นี้ — apply subtract-before-you-add: ลบ dead code, validator ที่ spec ไม่ต้องการ และ stub ค้าง ๆ ออกก่อน แล้วค่อยวางโค้ดใหม่บน base ที่เล็กลง

⚠️ เคล็ดลับ

กับดักคือคิดว่า principle นี้ใช้เฉพาะ code — ไฟล์ใช้กับ prompt ด้วย ('Simplify prompts: remove redundant instructions, excessive templates') และระวังการเก็บ stub ไว้ 'เผื่อดี' — ไฟล์ระบุว่าให้ลบทิ้งเลยถ้าไม่มีเนื้อหาใหม่ ส่วน validator/parser ที่เขียนเผื่อ edge case ที่ไม่มีใน spec คือหนี้ที่ต้องจ่ายด้วย guard ตลอดอายุฟีเจอร์

type-system-discipline
📦 สาระ

ใช้ type checker เป็น proof assistant: กำจัด impossible state, mismatched primitives และ unhandled variant ที่ compile time — ทำ illegal states ให้สร้างไม่ได้, brand semantic primitives, parse ข้อมูลภายนอกที่ boundary, ไม่โกหก compiler, exhaust variants ทุกตัว และ derive type จาก authoritative schema

💡 เจ๋งยังไง

กรอบคิดที่เข้มกว่า 'ใช้ TypeScript ให้ถูก' ปกติ: 'The type checker is a proof assistant. A case the types let you ignore becomes a runtime failure the compiler could have stopped.' จุดที่แหลมคือมี anti-pattern ที่ยกมาเฉพาะเจาะจงจำง่าย: `{ completed: boolean; completedAt?: Date }` ยอมให้ `completed: true; completedAt: undefined` ซึ่งไร้ความหมาย — ต้อง derive boolean จาก `completedAt !== null` หรือแยก variant ชัด รวมถึง idea 'Types are constructions, not restrictions' ที่พลิกวิธีคิด: non-empty list คือ head + rest ไม่ใช่ list ที่เช็ค length; valid time range คือ start + duration ไม่ใช่ timestamp สองตัวที่ต้องเชื่อว่าเรียงกัน และเกณฑ์ปิดที่กันการ over-engineer: 'Am I strengthening this type to keep an operation total, or just to be more precise?' — precision เกินที่ไม่ทำให้ total คือ ceremony ฟรี

🧭 ใช้เมื่อไหร่ + วิธี

ใช้ตอน design types, review function signature หรือเขียน code ในภาษา typed ใด ๆ (grounded ใน syntax เฉพาะด้วย skill typescript-best-practices) patterns 7 ข้อ: (1) Make illegal states unrepresentable — sum types/discriminated unions แทน bag of optional fields (2) Types are constructions — สร้าง type จากค่าที่ต้องการ แทน carve ออกจาก type หลวมด้วย checks (3) Brand semantic primitives — UserId/OrderId ไม่ interchangeable, validate ครั้งเดียวตอนสร้าง (4) External data untyped จนกระทั่ง parsed — มี parse function ที่ทุก boundary (RPC, JSON, CLI args, config, env, DB rows) (5) Don't lie to the type system — cast/unsafe coercion คือ postmortem วันหน้า ให้ validate/narrow/refine model แทน (6) Exhaustive matching เป็นหน้าที่ compiler — never-typed binding (TS), unannotated match (Rust), -Wincomplete-patterns (Haskell) (7) Derive types จาก authoritative schema (protobuf, OpenAPI, GraphQL, migration) แทน hand-roll type คู่ขนาน และ (8) Strengthen a type เฉพาะจุดที่ partiality ปรากฏ — runtime assertion/null check คือสัญญาณว่า type อ่อนเกิน แล้วหยุดตรงนั้น

💬 ตัวอย่าง

ออกแบบ type ให้ order flow — apply type-system-discipline: ทำ illegal states สร้างไม่ได้ด้วย discriminated union, brand UserId/OrderId แยก แล้ว derive type จาก OpenAPI schema แทนเขียน interface คู่ขนาน

⚠️ เคล็ดลับ

การทดสอบตัวเอง 6 ข้อจากไฟล์: ถ้าต้องเขียน comment อธิบายว่า combination ไหน valid = type หลวมเกิน (split เป็น sum type) · สอง argument type primitive เดียวกันแต่ความหมายต่าง = brand · เจอ any/as/assertNotNull = trace กลับไป validate ที่ boundary · เพิ่ม variant ใหม่แล้ว compiler ไม่ตะโกน = match ไม่ exhaustive · type ก๊อป shape ที่ไฟล์อื่นเป็นเจ้าของ = derive แทน · เข้ม type เพราะอยาก precise ทั้งที่ไม่มีอะไร panic = หยุด เก็บ type เดิม และ 'The cast you bury today is the postmortem you write next week'

🗓 เส้นทางสัปดาห์แรก

1
วัน 1
/setup-pstack → ลอง /poteto-mode กับงานจริง 1 งาน → ลอง /bro
2
วัน 2-3
/how + /why กับโค้ดที่ไม่คุ้น · /tdd กับงานเขียนใหม่
3
วัน 4-5
/create-verification-skill กับ project ที่มี dev server (แนะนำ eoffice — มี eo_verify พร้อม)
4
สุดสัปดาห์
/reflect งานทั้งสัปดาห์ → workflow ที่ซ้ำใช้ /automate-me ทำ mode ของคุณเอง
§

วิธีวิจัย

deep-research pipeline · ทุกข้อความ trace กลับไปไฟล์ JSON ที่ผ่าน validation ได้
🧪 Pipeline
  • รอบ 1: /deep-research — ตั้งกรอบ 6 items (ยืนยันกับผู้ใช้: กรอบ/ช่วงเวลา/ภาษา/batch) → agents วิจัย → JSON ต่อ item
  • รอบ 2: /deep-research-deep — agents 4 ตัวอ่าน SKILL.md จากดิสก์โดยตรง 44 ไฟล์ → JSON ต่อ skill
  • ทุกไฟล์ผ่าน validate_json.py (coverage 100%) และ re-validate อิสระก่อนสรุป
📚 แหล่งข้อมูลหลัก
  • cursor/plugins repo (ทางการ) · poteto/verification-skill-example · mirror backnotprop/pstack
  • The Neuron · Flavio Copes · tenten.dev · HN Algolia API · GitHub API · ZCode docs ทางการ + ตรวจของจริงในเครื่อง
⚠️ ข้อจำกัดของข้อมูล
  • X บล็อก scrape ตรง — เนื้อหาบทความมาจาก mirror + รายงานบุคคลที่สาม (attribution ไม่ชัดที่ grokbot.sh)
  • ตัวเลขผลิตภาพทั้งหมดเป็น self-report · plan จำกัด concurrent agents 2 ตัว ทำให้บาง item ผู้วิจัยลงมือเอง