Skip to content

task-brief: track fence character and length so fenced examples cannot end a task early - #2304

Open
mavericksea-ai wants to merge 2 commits into
obra:devfrom
mavericksea-ai:fix/task-brief-fence-tracking
Open

mavericksea-ai wants to merge 2 commits into
obra:devfrom
mavericksea-ai:fix/task-brief-fence-tracking

Conversation

@mavericksea-ai

@mavericksea-ai mavericksea-ai commented Sep 16, 2026 •

Copy link
Copy Markdown

Who is submitting this PR? (required)

Field Value
Your model + version Fix and tests: Claude Opus 4.6 in Claude Code on macOS. Investigation that found the bug: Codex desktop, 13 Sep 2026, against b36e082. Follow-up commit 80f8ed1 (CRLF fix and two fixtures): Claude Opus 5.5 in the Claude app, 25 Sep 2026
Harness + version Claude Code 2.1.90 on macOS for the fix; the Codex investigation ran the shipped script directly with no model calls; the follow-up was tested in a Linux shell under gawk, mawk and BWK awk
All plugins installed driftproof@driftproofhq v0.10.1 (user scope); Superpowers itself was not installed in the authoring session, the scripts were run from the clone
Human partner who reviewed this diff @mavericksea-ai

What problem are you trying to solve?

skills/subagent-driven-development/scripts/task-brief extracts one task from a plan for the implementer and the reviewer. Its awk toggles an in-fence flag on any line starting with three backticks, so a ~~~ fence is not tracked at all, a four-backtick fence containing a three-backtick block toggles the flag back off inside the example, and an indented fence is not seen. In each case a heading such as ### Task 99: Example only inside a fenced example is read as a real task boundary, extraction stops there, and the script exits 0 with a shorter brief. Anything after the example in that task, including its final requirements, never reaches the implementer or the reviewer.

Reproduced with task-brief PLAN 1 OUT on the shipped script at b36e082 (task-brief is identical on dev at 5940bd8):

Plan Exit Lines Line after the example kept?
``` fenced example 0 6 yes
~~~ fenced example 0 3 no
```` fence with a ``` block inside 0 4 no

The 13 existing sdd-workspace checks pass on the current script; none exercises this boundary.

Two more effects came up after this PR opened. The same boundary also sends content to the wrong task: when the fenced heading carries a real task number, the tail of Task 1 lands in Task 2's brief (oiler's reproduction below, on a real 7-task plan). And the first commit of this PR regressed plans with CRLF line endings: its closing-fence check allowed only spaces and tabs, so a fence ended by a CRLF line never closed and every later task was reported not found. The shipped toggle handles a simple fenced block in a CRLF file, and Git for Windows checks files out with CRLF by default.

What does this PR change?

The fence rule now follows the CommonMark definition: it records the opening fence character and length, accepts up to three spaces of indentation, ignores a backtick line whose info string contains a backtick, and closes only on the same character at the same or greater length with nothing else on the line. Fence lines that belong to the task are still printed. Four fixtures are added to tests/claude-code/test-sdd-workspace.sh: a three-backtick control, a tilde case, a four-backtick fence wrapping a three-backtick block, and an indented three-backtick case, each asserting the line after the example is still in task 1's brief.

A second commit, 80f8ed1, allows a trailing carriage return in the closing-fence check and adds two fixtures: a fenced ### Task 2 heading inside Task 1, asserting Task 1 keeps its end and Task 2's brief starts at its real heading with none of Task 1; and a CRLF plan whose fenced block must close so Task 2 is found.

Is this change appropriate for the core library?

Yes. It is a correctness fix to a shipped core script, no new skill, no project-specific behaviour and no third-party integration.

What alternatives did you consider?

Rejecting tilde fences outright, or changing the task-heading regex to ignore headings inside examples: both leave the nested four-backtick case broken and the second cannot tell an example heading from a real one. A full Markdown parser is out of proportion for a 20-line awk script. Tracking the fence delimiter the way CommonMark defines it is the smallest change that closes every case in the fixtures.

Does this PR contain multiple unrelated changes?

No. One script fix and its tests.

Existing PRs

Environment tested

Harness (e.g. Claude Code, Cursor) Harness version Model Model version/ID
Claude Code (macOS, macOS awk) 2.1.90 Claude Opus 4.6
Linux shell, gawk, mawk and BWK awk (follow-up commit) n/a, script test only Claude Opus 5.5

New harness support (required if this PR adds a new harness)

Not applicable, no harness added.

Evaluation

  • Initial prompt: a request to review obra/superpowers for local correctness defects in the shipped SDD scripts; the fence boundary came out of reading task-brief and testing it with valid Markdown plans.
  • Eval sessions after the change: none. This does not change any skill text or agent-facing prose; it changes what a bash script writes to a file. The measure is the test script: the three new failing cases fail on the current script and pass with the change, the control is unchanged, and all 17 checks pass with the fix. With 80f8ed1 the test has 20 checks: on the shipped script 5 fail (the three original fence cases and the two wrong-task checks), on this PR's first commit 1 fails (CRLF), and all 20 pass with both commits, under gawk, mawk and BWK awk.
  • Outcome difference: a task whose body contains a tilde-fenced, nested-fenced or indented-fenced example now reaches the implementer whole instead of truncated at the example. A task that follows such an example no longer receives the previous task's tail, and plans with CRLF line endings keep working.

Rigor

  • If this is a skills change: not a skills change, no skill content touched
  • This change was tested adversarially, not just on the happy path (tilde, nested, indented and control fixtures, a fenced heading carrying a real task number, and a CRLF plan; the closing rule requires a bare line so an info-string line cannot close a fence; a backtick opening fence with a backtick in its info string is not treated as a fence)
  • I did not modify carefully-tuned content

Human review

  • A human has reviewed the COMPLETE proposed diff before submission

No CI in the repository; bash tests/claude-code/test-sdd-workspace.sh was run locally on macOS, all checks passing. The follow-up commit was also run on Linux under gawk, mawk and BWK awk (which macOS's awk is based on), 20 of 20 passing.

@oiler

oiler commented Sep 25, 2026

Copy link
Copy Markdown

Independent repro from a real plan, on Superpowers 6.4.1 (Claude Code 2.1.282, macOS awk, found by a Claude Opus 5.5 session). task-brief is byte-identical on main today.

Beyond truncation, the same bug also sends content to the wrong task. When a fenced fixture holds a heading like ### Task 2: …, the tail of Task 1 ends up in Task 2's brief:

### Task 1: write a fixture

````markdown
```markdown
### Task 2: heading inside the fixture
```
````

Step 2 of Task 1.

### Task 2: real task two

task-brief plan.md 1 stops after the ```markdown line. task-brief plan.md 2 starts at the fixture heading and includes "Step 2 of Task 1." On our 7-task plan, whose tasks embed sample plans as fixtures, Task 1's brief grew from 63 to 426 lines and picked up other tasks' fixture headings. Tasks 5 and 6 were affected too. A fence rule that tracks character and length, like this PR's, splits all seven correctly.

The fence rule accepted only spaces and tabs after a closing fence, so in a
plan with CRLF line endings the first fenced block never closed and every
later task was reported not found. Allow a trailing carriage return.

Add two cases to the SDD workspace test: a fenced heading that carries a
real task number, where the next task's brief must start at its real
heading (the shape oiler reproduced in obra#2304), and a CRLF plan whose fenced
block must close so the next task is found.
@mavericksea-ai

Copy link
Copy Markdown
Author

Thanks @oiler, this is a really useful reproduction. The wrong-task effect is worse than the truncation I described, and a real 7-task plan makes it concrete.

Testing your case properly also turned up something my first commit got wrong: in a plan with CRLF line endings, the closing fence never closed, so every task after the first code block came back "not found". The shipped script handles that case, so it was a regression in this PR. 80f8ed1 fixes it and adds your shape as a fixture (Task 2's brief must start at its real heading, with none of Task 1), plus a CRLF plan. The test now has 20 checks, all passing under gawk, mawk and BWK awk.

@Enough1122

Copy link
Copy Markdown

AI code review — automated review for reference; please use your judgment.

Strong fix. The closing-fence rule is correct on both counts: it requires the same character and l >= fl, and the backtick-info-string exclusion. Verified against dev on tilde, nested-4-backtick, indented and CRLF plans; the CRLF fixture in 80f8ed1 is a good catch. I ran your six fixtures pre-fix and five fail, so the tests do exercise the fixed path, and all 20 pass on the head.

Answering the question @oiler raised, since it is checkable: this does cover the wrong-task effect, not just the truncation. On oiler's exact shape, dev gives Task 1 = 4 lines and Task 2 = 8 lines starting at ### Task 2: heading inside the fixture; this head gives Task 1 = 10 lines and Task 2 = 2 lines starting at ### Task 2: real task two. The leak is gone. Same on a 7-task plan: the mechanism is per-line, not per-plan.

Two things worth knowing before merge:

Minor — an unterminated fence now makes later tasks unfindable. infence opens on ~~~/``` and only closes on a bare matching line. If a plan has an unclosed fence, the new head reports task 2: not found (exit 3) where dev found it. That is a new failure mode introduced by tracking the fence, and the "example was never closed" case is exactly what a half-written plan looks like. skills/subagent-driven-development/scripts/task-brief:36 — consider treating EOF while infence is still set as "not a fence" so a truncated plan degrades to today's behaviour instead of exit 3.

Minor — this and #2359 conflict. Both rewrite the same awk block in task-brief; merging both as-is conflicts on the script and on the test file. I resolved it locally as a single strict parse driving both the fence state and a same-or-shallower terminator, and that satisfies both PRs' fixture sets — so this is resolvable, but the two need to be sequenced rather than merged independently. Worth agreeing which lands first, since #2304's fc/fl variables already carry the CRLF fix that #2359's parallel cfence parse lacks.

akalongman added a commit to akalongman/superpowers that referenced this pull request Sep 27, 2026
The strict fence parser's close check now allows a trailing carriage
return, the form obra#2304 landed. Two test cases make the same-level
heading rule explicit: a fenced same-level heading stays in its task
(the shape real plans use to quote text for insertion), an unfenced one
ends it (a sibling section such as Self-review or Done when).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

skill:subagent-driven-development The subagent-driven-development skill

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants