Skip to content

Plans as directories (a header file plus one file per task), and a boundary check before each next plan - #2340

Open
obra wants to merge 1 commit into
plan-setfrom
plan-directories
Open

obra wants to merge 1 commit into
plan-setfrom
plan-directories

Conversation

@obra

@obra obra commented Sep 19, 2026

Copy link
Copy Markdown
Owner

Who is submitting this PR? (required)

Field Value
Your model + version Claude Fable 5.1 (claude-fable-5-1)
Harness + version Claude Code 2.1.273 (CLI, macOS)
All plugins installed agent-sdk-dev@claude-plugins-official, claude-code-setup@claude-plugins-official, claude-session-driver@superpowers-marketplace, code-simplifier@claude-plugins-official, context7@claude-plugins-official, elements-of-style@superpowers-marketplace, episodic-memory@superpowers-marketplace, frontend-design@claude-plugins-official, github-triage@github-triage-dev, gopls-lsp@claude-plugins-official, ledger-memory@ledger-memory-market, linear@claude-plugins-official, mcp-server-dev@claude-plugins-official, plugin-dev@claude-code-plugins, plugin-dev@claude-plugins-official, primeradiant-ops@primeradiant, release-radar@whats-new-marketplace, summarize-meetings@2389-research-marketplace, superpowers-chrome@superpowers-marketplace, superpowers-developing-for-claude-code@superpowers-marketplace, superpowers-lab@superpowers-marketplace, superpowers@claude-plugins-official, worldview-synthesis@2389-research-marketplace
Human partner who reviewed this diff Jesse Vincent (maintainer; directed this submission and reviews the diff in the PR)

What problem are you trying to solve?

Two problems with the same root: a plan is one document, so it grows into the thing that gets re-read after every compaction, has to be kept consistent end to end, and is edited in place when a ruling changes something a later task consumes. Both 2026-09-17 field reports ran into it (a 5,953-line single plan for one phase of a small game; a parent session that had to reconstruct plan state nine times across compactions). And when a spec needs several plans, a ruling in plan 1 that renames an interface leaves plans 2-5 saying the old name, and nothing checks: in eval, sessions given a "plans touched" duty on every ruling did not fill it (0/2) and edited only an index line (1/2), and both said in interview they would run a check that had the same force as the intra-plan pre-flight scan, ideally scripted.

What does this PR change?

  • A plan is a directory. docs/superpowers/plans/YYYY-MM-DD-<feature>/ holds 00-header.md (goal, constraints, Plan Set, Review Focus) and one NN-<task-name>.md per task in execution order; nothing is repeated between them. writing-plans writes this form. A task file is what an implementer reads; the header is what every task shares.
  • Every helper takes a file or a directory: task-brief assembles header plus the one task file; sdd-workspace names the workspace after the directory; review-package, task-start, task-done pass through. Single-file plans keep working.
  • scripts/plan-boundary NEXT_PLAN (new, executing-plans): collects the identifiers the next plan's Consumes lines and its own Plan Set entry take from earlier plans, drops what the plan itself produces, and reports each one absent from the code as built; then checks every later plan's Plan Set entry for names attributed to plans already complete. Exit 1 until it prints boundary: clean.
  • Both executors run it before a next plan's Task 1 and fix the plan until clean, ledgering each fix as a ruling.
  • Tests: tests/claude-code/test-plan-directories.sh (brief assembly, workspace naming, the gate flagging a missing name in the next plan and in a later plan's attributed names, accepting a present name, reporting clean after the fix), registered in the runner.

Is this change appropriate for the core library?

Yes: it changes the shape of the plan every core skill reads and writes, and adds the one check that makes multi-plan execution safe.

What alternatives did you consider?

  • A "plans touched" slot on every ruling, edited before the next task: 0/2 slots written, 1/2 partial edits; interviewed sessions read the four-field format as the shape of a sentence. Replaced by the scripted gate.
  • A prose boundary check without a script: the micro-test could not distinguish it from no instruction (fresh sessions fix the mismatch on their own, 17/17), and the interviews were explicit that a gated script is followed where advice is not.
  • A file per step rather than per task: the task is the unit that carries a test cycle and a reviewer's gate; per-step files would recreate the volume problem as a file count.
  • Keep a single file and only add the gate: the gate works on either form; the directory is what makes a ruling's edit local and a resumed session's context small.

Does this PR contain multiple unrelated changes?

No. The directory and the gate are one change: the gate's inputs (Consumes lines, Plan Set entries, per-task files) are the directory's structure, and the directory without the gate leaves the multi-plan problem where it was.

Existing PRs

Environment tested

Harness (e.g. Claude Code, Cursor) Harness version Model Model version/ID
Claude Code (eval workers, fresh CLAUDE_CONFIG_DIR, Bedrock) 2.1.273 Claude Opus 5 (planner, SDD controller) us.anthropic.claude-opus-5
Claude Code (eval workers, fresh CLAUDE_CONFIG_DIR, Bedrock) 2.1.273 Claude Sonnet 5 (executor, implementers) us.anthropic.claude-sonnet-5

New harness support (required if this PR adds a new harness)

Not applicable.

Evaluation

  • Initial prompt: "One thing that we should think about is whether plans should get broken out into individual steps as files. Rather than just giant chained plans." then "Why don't we test the format change next?"; for the gate, "I bet the skill needs to be explicit about all plans."
  • Execution of a hand-split directory plan (six-module design, three planted-defect probes): inline on Sonnet 5, 3 reps, 9/9 probes, $2.66-2.98 (single file: 9/9, $2.85-3.27); SDD with Sonnet 5 implementers, 2 reps, 6/6, $9.28 / $17.54 (single file: $9.59-17.98); 12-13 helper-script calls per SDD rep assembling briefs from the directory.
  • Planning in directory form (Opus 5): six-module design, 3 reps, 6 files each, 620-686 lines (single-file form under the same wording: 475-686); the field report's five-phase spec, 2 reps, 3-4 plan directories of 34-36 task files, 3,595-3,747 lines, both finished on their own. A skill-written directory plan executed inline 3 reps: 9/9, $2.80-3.60.
  • The gate: the five-phase directory set with a planted cross-plan conflict (plans 1-2 call the engine entry point Tick, spec and plans 3-5 say Advance), executed as a set on Sonnet 5, 2 reps capped at 75 min: the script ran before plan 2 in both, the plan set was consistent with the code afterward in both (one session renamed to the spec's name and fixed plans 1-2; the other kept the plan's name and rewrote plans 2-5 to match, a ruling the partner reads), and both continued into plan 2. The same fixture under the per-ruling duty: 0/2 and 1/2 as above.
  • Script tests: test-plan-directories.sh, test-executing-plans-scripts.sh, test-sdd-workspace.sh pass. The suite's "Read at beginning" test fails identically on untouched dev.

Rigor

  • If this is a skills change: I used superpowers:writing-skills and completed adversarial pressure testing (paste results below)
  • This change was tested adversarially, not just on the happy path
  • I did not modify carefully-tuned content (Red Flags table, rationalizations, "human partner" language) without extensive evals showing the change is an improvement

On the first box: RED baselines were run for both halves (single-file execution; the per-ruling duty failing on the planted conflict), the failing sessions were resumed and interviewed for what would have changed their behavior, and the form followed the answer. The pressure-scenario style test with verbatim rationalizations was not run.

Human review

  • A human has reviewed the COMPLETE proposed diff before submission — Jesse Vincent asked for this PR after reviewing the format and gate results; the full-diff read happens in this PR.

A plan is a directory: 00-header.md plus one NN-<task>.md per task. The
brief for task N is the header plus that file; sdd-workspace, task-brief,
review-package, task-start and task-done take a file or a directory.
Before starting the next plan in a set, executors run
scripts/plan-boundary, which names every identifier the plan consumes
from earlier plans that the code as built does not contain, in the next
plan and in every later plan's Plan Set entry, and fix the plan until it
prints clean. Measured: the directory form executes the same as the file
form (inline 9/9, SDD 6/6, same cost) and plans at the same volume; the
gate kept a five-plan set consistent with the code 2/2 against a planted
cross-plan naming conflict, where a per-ruling duty managed 0/2.
@Enough1122

Copy link
Copy Markdown

AI code review — automated review for reference; please use your judgment.

Three reproducible cases where plan-boundary gives a wrong verdict on directory plans in the shape writing-plans prescribes. test-plan-directories.sh passes on all of them.

1. Task-level Consumes lines are never read (plan-boundary:37). body is awk '/^## Plan Set/{skip=1;next} /^## /{skip=0} !skip', but task headings are ###, which never matches /^## /. Once ## Plan Set is seen skip stays 1 to end of file, so in a directory plan every NN-*.md is dropped and consumes/produces are both empty. writing-plans' own header template puts ## Plan Set last, so this is normal: a task file declaring Consumes: `RendererNew`, `SpriteLoad` with neither in the code yields boundary: clean, exit 0. The Plan Set only lists what a plan consumes from the plans before it, so intra-plan consumption is invisible too, and a Task 1 rename is executed wrong. Single-file plans fail the same way. Fix: reset on any heading, /^#/{skip=0}.

2. Step 2 uses alphabetical order, not Plan Set order (plan-boundary:44,47). idx is the position in ls -d "$parent"/* | sort, but plans are named YYYY-MM-DD-<feature> while the Plan Set is "in execution order" — nothing ties them. Plan Set alpha(1) zebra(2) apple(3) beta(4) sorts as alpha, apple, beta, zebra, so gating before zebra puts it last, k > idx selects nothing, and step 2 checks no plan at all. With ZebraRender/ZebraPaint renamed to ZebraDraw in plans 3–4, the gate prints boundary: clean. The fixture only passes because 1-engine/2-terminal/3-effects sort into Plan Set order. Take the position from the Plan Set numbering, or check every sibling. The same for p in $(...) also word-splits on a spaced dir name.

3. A later plan with no from plan N exits 1 with no output (plan-boundary:51-54). Under set -euo pipefail, when a later plan's entry reads Consumes: \None`/Consumes nothing— wording writing-plans' own Plan Set template uses — thegrepat :51 matches nothing and the assignment aborts: exit 1, empty stdout and stderr. Both SKILL.md files then tell the session to hunt a nonexistent name with no diagnostic.|| trueon thenames=pipeline at :54 fixes it.plan_text(:27) has the same shape: a dir missing00-header.mdaborts silently, thoughtask-brief:34` tolerates that.

4. The test covers none of these: ## Plan Set is last, dirs sort into Plan Set order, and every consumed name is repeated in the Plan Set lines — so the step-2 assertion is fed by Plan Set text, not the task file.

@Enough1122

Copy link
Copy Markdown

AI code review — automated review for reference; please use your judgment.

Correction to finding 1 above — I could not reproduce it, and I believe it is wrong. Please disregard that item; findings 2, 3 and 4 are unaffected.

I built the exact shape the claim describes: a directory plan with ## Plan Set last, a task file 01-engine.md declaring Consumes: `RendererNew`, `SpriteLoad` , and neither name present in the code. plan-boundary reports both:

missing in code: RendererNew  (consumed by plan-a)
missing in code: SpriteLoad  (consumed by plan-a)

The reason is in the awk: /^## / is a prefix match, and ### starts with ## … except it does not — ### is ###, and the pattern requires ## (hash-hash-space), so a ### heading does not match. But the reset that matters here is not the task heading: plan_text (:27) concatenates 00-header.md followed by [0-9][0-9]-*.md, so within a task file the Consumes: line appears before any ## Plan Set of its own, and the header's ## Plan Set block is what the first awk skips. The task body that survives is exactly the task content.

The PR's own fixture agrees: test-plan-directories.sh:32,40,46 put Consumes: in the task files, and :72 renames identifiers in 01-app.md/01-fx.md — so the task-file path is exercised, not stubbed.

Apologies for the noise. The /^#/{skip=0} suggestion is still a harmless robustness tweak, but it is not fixing a live bug.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

skill:executing-plans The executing-plans skill skill:subagent-driven-development The subagent-driven-development skill skill:writing-plans The writing-plans skill

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants