Conversation
A plan is a directory: 00-header.md plus one NN-<task>.md per task. The brief for task N is the header plus that file; sdd-workspace, task-brief, review-package, task-start and task-done take a file or a directory. Before starting the next plan in a set, executors run scripts/plan-boundary, which names every identifier the plan consumes from earlier plans that the code as built does not contain, in the next plan and in every later plan's Plan Set entry, and fix the plan until it prints clean. Measured: the directory form executes the same as the file form (inline 9/9, SDD 6/6, same cost) and plans at the same volume; the gate kept a five-plan set consistent with the code 2/2 against a planted cross-plan naming conflict, where a per-ruling duty managed 0/2.
Three reproducible cases where 1. Task-level 2. Step 2 uses alphabetical order, not Plan Set order ( 3. A later plan with no 4. The test covers none of these: |
Correction to finding 1 above — I could not reproduce it, and I believe it is wrong. Please disregard that item; findings 2, 3 and 4 are unaffected. I built the exact shape the claim describes: a directory plan with The reason is in the awk: The PR's own fixture agrees: Apologies for the noise. The |
Who is submitting this PR? (required)
What problem are you trying to solve?
Two problems with the same root: a plan is one document, so it grows into the thing that gets re-read after every compaction, has to be kept consistent end to end, and is edited in place when a ruling changes something a later task consumes. Both 2026-09-17 field reports ran into it (a 5,953-line single plan for one phase of a small game; a parent session that had to reconstruct plan state nine times across compactions). And when a spec needs several plans, a ruling in plan 1 that renames an interface leaves plans 2-5 saying the old name, and nothing checks: in eval, sessions given a "plans touched" duty on every ruling did not fill it (0/2) and edited only an index line (1/2), and both said in interview they would run a check that had the same force as the intra-plan pre-flight scan, ideally scripted.
What does this PR change?
docs/superpowers/plans/YYYY-MM-DD-<feature>/holds00-header.md(goal, constraints, Plan Set, Review Focus) and oneNN-<task-name>.mdper task in execution order; nothing is repeated between them. writing-plans writes this form. A task file is what an implementer reads; the header is what every task shares.task-briefassembles header plus the one task file;sdd-workspacenames the workspace after the directory;review-package,task-start,task-donepass through. Single-file plans keep working.scripts/plan-boundary NEXT_PLAN(new, executing-plans): collects the identifiers the next plan's Consumes lines and its own Plan Set entry take from earlier plans, drops what the plan itself produces, and reports each one absent from the code as built; then checks every later plan's Plan Set entry for names attributed to plans already complete. Exit 1 until it printsboundary: clean.tests/claude-code/test-plan-directories.sh(brief assembly, workspace naming, the gate flagging a missing name in the next plan and in a later plan's attributed names, accepting a present name, reporting clean after the fix), registered in the runner.Is this change appropriate for the core library?
Yes: it changes the shape of the plan every core skill reads and writes, and adds the one check that makes multi-plan execution safe.
What alternatives did you consider?
Does this PR contain multiple unrelated changes?
No. The directory and the gate are one change: the gate's inputs (Consumes lines, Plan Set entries, per-task files) are the directory's structure, and the directory without the gate leaves the multi-plan problem where it was.
Existing PRs
Environment tested
New harness support (required if this PR adds a new harness)
Not applicable.
Evaluation
test-plan-directories.sh,test-executing-plans-scripts.sh,test-sdd-workspace.shpass. The suite's "Read at beginning" test fails identically on untouched dev.Rigor
superpowers:writing-skillsand completed adversarial pressure testing (paste results below)On the first box: RED baselines were run for both halves (single-file execution; the per-ruling duty failing on the planted conflict), the failing sessions were resumed and interviewed for what would have changed their behavior, and the form followed the answer. The pressure-scenario style test with verbatim rationalizations was not run.
Human review