Skip to content

writing-plans: the plan's code was built twice during planning, then the handoff offered to implement it a third time #2408

Description

@tamaro-skaljic

Environment (required)

Field Value Provenance / supporting evidence
Superpowers version 6.4.1 (not a checkout) current observation of package.json; historical evidence: the brainstorming and writing-plans bodies injected at transcript lines 34 and 1644 equal the installed SKILL.md files
Harness (Claude Code, Cursor, etc.) Claude Code, VS Code extension (entrypoint claude-vscode) historical evidence; entrypoint field of the records
Harness version 2.1.281, then 2.1.282 historical evidence; version field, transcript lines 3–632 and 638 onward
Your model + version claude-opus-5-5 historical evidence; assistant records, transcript lines 30–4961
All plugins installed superpowers, caveman, anthropic-skills; MCP servers context7 and jean; a command-rewriting shell hook (rtk) current observation of the session's configuration
OS + shell Windows 11 Pro 10.0.26100, Git Bash and PowerShell 7 current observation

Is this a Superpowers issue or a platform issue?

  • I confirmed this issue does not occur without Superpowers installed

The reporter has not tried reproducing without superpowers. Evidence for
involvement is below; it does not establish cause.

What happened?

One long session: brainstorming a design, then /superpowers:writing-plans for plan 1 of it (a Rust crate family), with three manual compactions in between.

In the reporter's words, the brainstorming and writing-plans skills led to three things:

  1. A fully functional prototype was built while the plan was written. So the plan became step-by-step instructions to recreate an implementation that already existed.
  2. The plan was verified by a step-by-step "replay", which built everything a second time.
  3. The execution handoff proposed implementing the plan again, natively or with subagents, a third time.

"A user who doesn't read and think carefully about questions or the steps an agent is executing while following these skills will lead to the plan executed 3 times, the first 2 times thrown away."

What the transcript shows (session diagnosis, confidence high):

  • The prototype question. After writing-plans started (transcript lines 1643–1644), the agent asked: "The plan must contain complete code for every step. Should I first build the core as a throwaway prototype in my scratchpad … so the plan's code is known to compile, pass clippy and pass its tests?" The recommended option gave the cost only as "Writing the plan takes considerably longer" (line 1849). The reporter chose it (line 1854).
  • The first build. The prototype of five crates with 70 tests was built at lines 1969–2864: 91 API responses and 48.6M tokens, 32% of the session's tokens. The agent recorded it as "= final code of plan 1" (line 2939) and generated the plan from its files (line 2947).
  • The replay, started unasked. Without a question, the agent decided to "replay it task by task in a throwaway clone of the repo" (line 2864) and replayed Task 1 (lines 2865–2929). Its own list of remaining steps ended in "Execution handoff: ask the user to review the plan and choose Subagent-driven vs Native". That list survived a compaction (line 3691).
  • The replay question. After the compaction, the agent asked whether to replay Tasks 2–6. The recommended option said "I apply each task step by step to a fresh clone and observe each red and green step … about half an hour" (line 4014). It didn't say that this builds the whole implementation again, or what happens to the result. The reporter chose it (line 4019).
  • The second build. All six tasks were applied, with commits: 41 responses, 11.7M tokens, 13.4 minutes (lines 4261–4590).
  • The handoff. It reported "I replayed all six tasks step by step on a fresh clone" and "the clone's crates ended up identical to the prototype's". It then asked "Please review the plan. Which execution approach would you prefer?", and recommended "Native" partly because the code "has been replayed" (line 4688).
  • The reporter's objection. "What's the point of throwing all that away just to recreate everything manually from scratch?" (line 4691). The agent then said the replay came from no skill (line 4696) and kept the replay's commits as a branch.

Steps to reproduce

  1. Brainstorm a design for a new multi-crate library with superpowers:brainstorming, until the specs are approved.
  2. Run /superpowers:writing-plans for the first plan. The agent proposes building a "throwaway prototype" first, so that the plan's complete code is known to compile and pass its tests. Accept the recommended option.
  3. The prototype is built completely, and the plan's code blocks are generated from it.
  4. The agent verifies the plan by replaying its tasks on a clone of the repository. Accept the recommended "Replay" option.
  5. Observable: at the handoff, the agent asks whether to implement the plan with subagents or natively, although the same message reports that the replay already produced the identical implementation.

Expected behavior

From the reporter:

  • Say what it costs: when proposing the prototype and the replay, the agent should say plainly that each is a complete implementation, and what will happen to it afterwards.
  • The handoff knows the state: the handoff should take into account that an implementation already exists.
  • Plan only, no code; build once and keep it. In the reporter's words: "If the task is to plan a prototype, it should be a plan. If the task is to plan a non-prototype, a prototype may be okay, but it should stay a prototype and not a fully featured implementation. The plan in this case doesn't make sense anymore and it should be more of a 'implementation details / micro architecture' document or a 'This is what I have done and why I have done it' document rather than a step-by-step instruction how to recreate it … If an agent needs to verify something because its unsure about whether it will work, a temporary scratchpad prototype is fine."

Actual behavior

  • The plan's code was built completely twice with identical results, the prototype and the six-task replay, and Task 1 once more in between.
  • A third, full implementation was offered at the handoff.
  • Neither proposal said it was a complete implementation, or what would happen to it.
  • The prototype and the Task 1 replay together took 55M of the session's 152M tokens (about 36%, lines 1969–2937). The six-task replay took another 11.7M.

Debug log or conversation transcript

Session id: e1adc835-b508-49fb-acb3-0c425a5a202b. Delivered local archive: none built. Attached bundle: session-diagnosis-e1adc835.tar.gz.
Superpowers involvement per the diagnosis report: likely, with evidence at transcript lines 1644, 1849, 2939, 3691, 4688 and 4696. This report does not propose a fix.


Filed with the diagnosing-superpowers skill. Model, harness, harness version, and installed plugins are listed above.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    automated-issue-reportFiled by an agent using the diagnosing-superpowers skillbugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions