feat(proxy): add turn-hook extension point for buffered model turns - #1891
Merged
Merged
Conversation
Add headroom/proxy/turn_hooks.py — a neutral extension contract that lets an opt-in proxy extension observe and optionally re-drive a single buffered model turn. A hook's on_request may rewrite the outbound tools/messages; its on_response may call the model again (via call_model) and return a replacement response, transparently to the client. Wire it at four seams in the Anthropic and OpenAI handlers — the pre-send point (alongside the existing tool-search deferral) and the CCR response seam — each gated so an empty registry is a byte-identical no-op. With no extension installed the proxy behaves exactly as before. Covered by tests/test_turn_hooks.py (registry, request mutation, response replacement + re-drive loop, never-raise). Existing CCR/handler suites pass unchanged, proving the no-op property.
Contributor
PR governanceThis PR does not yet satisfy the required template fields:
Please update the PR body, or move the PR back to draft while it is still in progress. |
14 of 21 tasks
chopratejas
added a commit
that referenced
this pull request
Jul 9, 2026
…hema savings (#1896) ## Description **Stacked on #1891** (base = `tejas/turn-hooks-extension`; the diff here is only the wiring + dashboard, and GitHub will retarget to `main` once #1891 merges). Makes the turn-hook seam actually useful on the OpenAI `/v1/chat/completions` **direct path** (the route an `opencode → OpenAI` session uses), and surfaces the resulting tool-schema token savings in the dashboard. - **Shrink** (`on_request`) at the shared pre-send point, gated on the registry and non-streaming — a hook can rewrite the outbound `tools`/`messages`; the net tool-schema token delta is recorded. - **Reload** (`on_response`) after the buffered send, with a `_retry_request`-based `call_model` — a hook can resolve a tool the model asked to load, re-drive the turn, and have the proxy return the final response transparently. No-op when no hook is registered. - **Dashboard**: a new `savings.by_layer.tool_search` aggregate and a per-request `tool_schema_saved_tokens` field sum both the native tool-search deferral and the hook-driven tools rewrite — attributed to Headroom only, so a client that already had tool search on (e.g. Claude Code / Codex) contributes zero. Adds a "Tool-Schema Deferral" card (rendered only when there is a saving) + a per-call badge. Recorded **per turn**: tools are re-sent on every turn, so the saving is logged on each request and summed across the recent window. Closes # ## Type of Change - [ ] Bug fix (non-breaking change that fixes an issue) - [x] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - `handlers/openai.py`: `on_request` shrink at the pre-send point + `on_response` reload loop on the direct buffered path; record the net tool-schema token delta as a saving. Savings are measured from the FINAL tools object regardless of whether the hook replaced or mutated `ctx.tools` in place (identity is used only to decide `body["tools"]` reassignment). - `server.py`: `_tool_schema_saved_from_tags()` helper; aggregate `savings.by_layer.tool_search` in `_build_stats_payload`; expose `tool_schema_saved_tokens` per request. - `dashboard/templates/dashboard.html`: "Tool-Schema Deferral" card (only renders when there is a saving) + per-call badge. - `tests/test_openai_chat_turn_hooks.py`: end-to-end shrink+reload, per-turn tag + `/stats` aggregation, an **in-place-mutating** hook regression (pins the extension contract + dashboard path), and the no-op-when-unregistered invariant. ## Testing - [x] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check .`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added for new functionality - [x] Manual testing performed (real extension end-to-end, incl. live OpenAI + Kimi/Fireworks) ### Test Output ```text $ ruff check headroom/proxy/handlers/openai.py headroom/proxy/server.py tests/test_openai_chat_turn_hooks.py All checks passed! $ ruff format --check <same 3 files> 3 files already formatted $ mypy headroom Success: no issues found in 408 source files $ pytest tests/test_openai_chat_turn_hooks.py -q 4 passed # regression (no-op safety + payload shape) $ pytest tests/test_turn_hooks.py tests/test_ccr_response_handler.py tests/test_proxy_ccr.py \ tests/test_openai_responses_compression_units.py tests/test_handler_outcome_tag_invariant.py -q 81 passed $ pytest tests/test_proxy_savings_history.py tests/test_telemetry_warning.py -q 69 passed ``` ## Real Behavior Proof - Environment: local macOS, project `.venv` (Python 3.12.6); `ruff` pinned to CI's `0.15.17` via `uvx ruff@0.15.17`; `mypy` from the venv. - Exact command / steps: drove `handle_openai_chat` via FastAPI `TestClient` with a registered hook. First with a mocked upstream (`_retry_request`) for deterministic assertions, then against the real out-of-tree tool-router extension over the LIVE OpenAI API (gpt-4o-mini) and the LIVE Fireworks API (Kimi K2.6, OpenAI-compatible), each with a 13-tool belt where the needed tool is deferred. - Observed result: round-1 outbound tools shrank to `[6 core] + search_tools`; the model chose to call `search_tools`; the extension resolved the right tool via BM25; round-2 re-drove with the resolved tools appended and the model called `github_create_issue`; the proxy returned the final response; `x-headroom-transforms` carried `turn_hook:tools:<N>tok` (188 tok on the live Kimi run) and `/stats` `by_layer.tool_search` summed it per turn. The in-place-mutation regression test confirms savings are recorded when a hook mutates `ctx.tools` without replacing it. - Not tested: a full standalone `headroom proxy` process with a real opencode client (used `TestClient` + real upstream); the exercised handler path is identical, only the client differs. Streaming is intentionally out of scope (a streamed turn can't be re-driven). ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [ ] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [ ] I have updated the CHANGELOG.md if applicable
Merged
chopratejas
pushed a commit
that referenced
this pull request
Jul 9, 2026
🤖 I have created a release *beep* *boop* --- <details><summary>0.31.0</summary> ## [0.31.0](v0.30.0...v0.31.0) (2026-07-09) ### Features * **cache:** provider-agnostic cache-mode delta + cc-agnostic prefix comparison ([#1868](#1868)) ([7c2f0ea](7c2f0ea)) * **ccr:** wire retrieve-tool interception into OpenAI Responses handler ([#1898](#1898)) ([62cd307](62cd307)) * **compression:** add audit-safe mode with protected pattern matching ([#1899](#1899)) ([bb112dd](bb112dd)) * **content-router:** accept any real compression (remove min-savings floor) ([#1771](#1771)) ([6c31db9](6c31db9)) * **content-router:** lossless-first dispatch, cross-turn dedup, and A7 lossy-after-fold ([#1818](#1818)) ([60af15f](60af15f)) * **proxy:** add provider-only HTTP proxy ([#1807](#1807)) ([ebe0a3b](ebe0a3b)) * **proxy:** add turn-hook extension point for buffered model turns ([#1891](#1891)) ([ec950f7](ec950f7)) ### Bug Fixes * **build:** enable Intel macOS pip installs via ort-load-dynamic ([#1538](#1538)) ([32ce99e](32ce99e)) * **cache:** avoid fallback session collisions ([#1827](#1827)) ([0f606b6](0f606b6)) * **ccr:** make expired retrieve misses terminal ([#1781](#1781)) ([9cbdba4](9cbdba4)) * **ccr:** preserve Anthropic re-stream shape ([#1854](#1854)) ([f663894](f663894)) * **ccr:** preserve thinking blocks in buffered stream re-synthesis ([#1897](#1897)) ([ede085c](ede085c)) * **cli/proxy:** preserve explicit HEADROOM_MIN_TOKENS=0 / MAX_ITEMS=0 ([#1886](#1886)) ([3a33af1](3a33af1)) * **code-compressor:** CJK-aware relevance-query symbol matching ([#1747](#1747)) ([b38315c](b38315c)) * **codex:** discover updated Codex state stores ([#1889](#1889)) ([9d42eba](9d42eba)) * **codex:** OpenCode Zen telemetry attribution ([#1648](#1648)) ([f18c6bd](f18c6bd)) * **content-detector:** detect and compress space-separated JSON objects ([#1742](#1742)) ([5194bdc](5194bdc)) * **content-router:** token-measure lossless folds at the acceptance gate ([#1772](#1772)) ([c5493ea](c5493ea)) * **copilot:** normalize subscription routing host ([#1836](#1836)) ([afd9cbd](afd9cbd)) * **copilot:** route mixed-model requests per model ([#1785](#1785)) ([5af5e22](5af5e22)) * **dashboard:** deduplicate repeated savings metrics ([#1804](#1804)) ([88f935a](88f935a)) * **dashboard:** distinguish unavailable RTK from zero stats in Docker ([#1900](#1900)) ([87f6e93](87f6e93)) * **dashboard:** distinguish unavailable RTK from zero stats in Docker ([#1901](#1901)) ([361adcd](361adcd)) * **dashboard:** price proxy savings without litellm ([#1728](#1728)) ([188e382](188e382)) * detect and clear stale ANTHROPIC_BASE_URL from crashed wrap sessions ([#1768](#1768)) ([#1837](#1837)) ([84509a4](84509a4)) * **docker:** persist headroom workspace in compose ([#1839](#1839)) ([5e29c06](5e29c06)) * **docker:** report source build version ([#1862](#1862)) ([3807488](3807488)) * **evals:** default unparseable judge scores below pass threshold ([#1892](#1892)) ([42ebbc6](42ebbc6)) * **install:** pass sc.exe create as raw command line so binPath= quoting survives ([#1654](#1654)) ([#1702](#1702)) ([d6e0710](d6e0710)) * **install:** persist --no-http2 override through install apply ([#1676](#1676)) ([6fb5f3b](6fb5f3b)) * **mcp:** isolate ClaudeRegistrar CLI config env ([#1888](#1888)) ([1c947b1](1c947b1)) * **mcp:** surface dead proxy state ([#1786](#1786)) ([931eed8](931eed8)) * **memory:** resolve Trae cwd metadata from user reminders ([#1737](#1737)) ([#1887](#1887)) ([3e85eb1](3e85eb1)) * **opencode:** use local MCP config ([#1383](#1383)) ([4bd3ddf](4bd3ddf)) * **proxy/openai:** thread savings-profile kwargs into chat completions ([#1606](#1606)) ([7ff842d](7ff842d)) * **proxy/openai:** translate max_tokens -> max_completion_tokens on chat path ([#1774](#1774)) ([285808b](285808b)) * **proxy:** bound Codex WS compression fallback latency ([#1802](#1802)) ([d24a3f8](d24a3f8)) * **proxy:** bound HF tokenizer load and offload token counting off event loop ([#1738](#1738)) ([46d5d68](46d5d68)) * **proxy:** cancel retry backoff on shutdown ([#1834](#1834)) ([da2d8dc](da2d8dc)) * **proxy:** compress Anthropic user text blocks when enabled ([#1875](#1875)) ([e36439a](e36439a)) * **proxy:** freeze must forward cached (compressed) prefix byte-identical — stop token-mode cache busting ([#1850](#1850)) ([248ae0f](248ae0f)) * **proxy:** fsync savings dir after atomic rename ([#1764](#1764)) ([7de2c1e](7de2c1e)) * **proxy:** keep cache_control bounded + stable so the freeze overlay stops busting ([#1852](#1852)) ([4820134](4820134)) * **proxy:** persist lifetime cache-read savings across restarts ([#1665](#1665)) ([908997e](908997e)) * **proxy:** preserve streaming passthrough beta headers ([#1783](#1783)) ([0f553a8](0f553a8)) * **proxy:** release _active_streams session lock on setup-phase errors ([#1864](#1864)) ([2ccd831](2ccd831)) * **proxy:** retry HTTP/2 stream resets instead of 502ing ([#1645](#1645)) ([2ce19c2](2ce19c2)) * **proxy:** retry passthrough on transient upstream connection close ([#1513](#1513)) ([5d14080](5d14080)) * **proxy:** route Foundry Anthropic messages ([#1878](#1878)) ([739f654](739f654)) * **proxy:** serve /favicon.ico locally instead of tunneling upstream ([#1787](#1787)) ([#1847](#1847)) ([3076e32](3076e32)) * **proxy:** stop rtk stat failures from corrupting session baseline ([#1693](#1693)) ([681b9a8](681b9a8)) * **proxy:** strip 1m model suffix before upstream forwarding ([#1840](#1840)) ([e22d745](e22d745)) * **proxy:** subtract cache write premiums from net savings ([#1800](#1800)) ([53a465b](53a465b)) * **router:** honor MCP aliases in excluded tools ([#1822](#1822)) ([#1863](#1863)) ([140d6e4](140d6e4)) * **rtk:** link managed rtk onto PATH instead of mutating the hook ([#1698](#1698)) ([140cb05](140cb05)) * **streaming:** preserve server_tool_use sse blocks ([#1826](#1826)) ([4ac5493](4ac5493)) * **toin:** publish skip compression recommendations ([#1782](#1782)) ([be51008](be51008)) * **transforms:** normalize diff compressor context ([#1801](#1801)) ([838c523](838c523)) * **transforms:** pass through ragged tables instead of misaligning columns ([#1713](#1713)) ([c7665ca](c7665ca)) * use rtk native Cursor hook instead of injecting .cursorrules ([#756](#756)) ([#1846](#1846)) ([1573f1f](1573f1f)) * **wrap:** replace stale-proxy detection with Vite-style port fallback ([#1406](#1406)) ([b4205c6](b4205c6)) ### Performance Improvements * **proxy:** cap compression workers to CPU count ([#1803](#1803)) ([0a3851b](0a3851b)) * **savings:** batch tracker persistence off the request hot path ([#1817](#1817)) ([451b9f0](451b9f0)) ### Dependencies * bump the cargo-minor-patch group across 1 directory with 7 updates ([#1909](#1909)) ([45601d9](45601d9)) * bump the npm-minor-patch group across 4 directories with 18 updates ([#1907](#1907)) ([8872bbc](8872bbc)) </details> --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
JerrettDavis
pushed a commit
to peterlodri-sec/headroom
that referenced
this pull request
Jul 14, 2026
…eadroomlabs-ai#1891) ## Description Adds a small, neutral **extension point** to the proxy: a "turn hook" that lets an opt-in extension observe and optionally re-drive a single buffered model turn, without touching the core request/response flow for anyone who has no extension installed. A hook can: - `on_request(ctx)` — inspect or rewrite the outbound tools/messages before they go upstream (the extensible counterpart to the built-in tool-search deferral that already lives at that point). - `on_response(ctx, response, call_model)` — inspect the model's response and, if it wants, call the model again (via `call_model`) and return a **replacement** response — transparently to the client. This is the capability that can't be done from ASGI middleware: it reuses the proxy-internal re-call path (the same `api_call_fn` the CCR handler already drives). The module is **inert unless a hook is registered**: the runners return their input unchanged and are gated on the registry, so with no extension the proxy is byte-identical to today. A failing hook is logged and skipped — it can never take the proxy down. Closes # ## Type of Change - [ ] Bug fix (non-breaking change that fixes an issue) - [x] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - Add `headroom/proxy/turn_hooks.py`: `TurnContext`, the `TurnHook` protocol (`on_request` / `on_response`), a module registry (`register_turn_hook` / `registered_turn_hooks` / `clear_turn_hooks`), and the runners `run_request_hooks` / `run_response_hooks`. Inert when empty; never raises. - Wire it at four seams, each gated so an empty registry is a byte-identical no-op: - Anthropic — pre-send (right after the existing tool-search deferral) + the CCR response seam. - OpenAI — the Responses tool-shaping point (right after the existing tool-search deferral, copy-on-write-safe) + the CCR response seam. - Add `tests/test_turn_hooks.py`. ## Testing - [x] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check .`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added for new functionality - [x] Manual testing performed (no-op regression across CCR/handler suites) ### Test Output ```text $ ruff check headroom/proxy/turn_hooks.py tests/test_turn_hooks.py \ headroom/proxy/handlers/anthropic.py headroom/proxy/handlers/openai.py All checks passed! $ ruff format --check <same 4 files> 4 files already formatted $ mypy headroom Success: no issues found in 408 source files $ pytest tests/test_turn_hooks.py -q 9 passed in 0.17s $ pytest tests/test_turn_hooks.py tests/test_ccr_response_handler.py \ tests/test_ccr_tool_injection.py tests/test_proxy_ccr.py \ tests/test_openai_tool_search_deferral.py \ tests/test_openai_responses_compression_units.py \ tests/test_handler_outcome_tag_invariant.py -q 135 passed (+ 1 pre-existing cross-file flake in test_proxy_ccr::test_health_endpoint, which passes in isolation and in its own file: `pytest tests/test_proxy_ccr.py` -> 19 passed) ``` ## Real Behavior Proof - **Environment:** local macOS, project `.venv` (Python 3.12.6); `ruff` pinned to CI's `0.15.17` via `uvx ruff@0.15.17`; `mypy` from the venv. - **Exact command / steps:** branched off `upstream/main`; added the hook module + wired the four handler seams; ran the ruff/format/mypy/pytest commands above. - **Observed result:** The unit tests exercise the whole contract — registry, `on_request` mutating `ctx.tools`, `on_response` returning a replacement, the `await call_model(...)` re-drive loop, replacement-chaining across hooks, and the never-raise guarantee. The existing CCR + handler suites pass unchanged, which is the point: with no hook registered the added code is a no-op (the runners short-circuit on an empty registry). - **Not tested:** the live interactive re-drive path with a *registered* hook against a real upstream — no hook ships in this repo, so that path is covered here only by the unit test's fake `call_model`. The `on_request` seam fires on the Anthropic pre-send and OpenAI Responses paths (where the existing tool-search deferral runs); other send paths (e.g. chat-completions, streaming) are not wired in this PR. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [ ] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [ ] I have updated the CHANGELOG.md if applicable
JerrettDavis
pushed a commit
to peterlodri-sec/headroom
that referenced
this pull request
Jul 14, 2026
🤖 I have created a release *beep* *boop* --- <details><summary>0.31.0</summary> ## [0.31.0](headroomlabs-ai/headroom@v0.30.0...v0.31.0) (2026-07-09) ### Features * **cache:** provider-agnostic cache-mode delta + cc-agnostic prefix comparison ([headroomlabs-ai#1868](headroomlabs-ai#1868)) ([7c2f0ea](headroomlabs-ai@7c2f0ea)) * **ccr:** wire retrieve-tool interception into OpenAI Responses handler ([headroomlabs-ai#1898](headroomlabs-ai#1898)) ([62cd307](headroomlabs-ai@62cd307)) * **compression:** add audit-safe mode with protected pattern matching ([headroomlabs-ai#1899](headroomlabs-ai#1899)) ([bb112dd](headroomlabs-ai@bb112dd)) * **content-router:** accept any real compression (remove min-savings floor) ([headroomlabs-ai#1771](headroomlabs-ai#1771)) ([6c31db9](headroomlabs-ai@6c31db9)) * **content-router:** lossless-first dispatch, cross-turn dedup, and A7 lossy-after-fold ([headroomlabs-ai#1818](headroomlabs-ai#1818)) ([60af15f](headroomlabs-ai@60af15f)) * **proxy:** add provider-only HTTP proxy ([headroomlabs-ai#1807](headroomlabs-ai#1807)) ([ebe0a3b](headroomlabs-ai@ebe0a3b)) * **proxy:** add turn-hook extension point for buffered model turns ([headroomlabs-ai#1891](headroomlabs-ai#1891)) ([ec950f7](headroomlabs-ai@ec950f7)) ### Bug Fixes * **build:** enable Intel macOS pip installs via ort-load-dynamic ([headroomlabs-ai#1538](headroomlabs-ai#1538)) ([32ce99e](headroomlabs-ai@32ce99e)) * **cache:** avoid fallback session collisions ([headroomlabs-ai#1827](headroomlabs-ai#1827)) ([0f606b6](headroomlabs-ai@0f606b6)) * **ccr:** make expired retrieve misses terminal ([headroomlabs-ai#1781](headroomlabs-ai#1781)) ([9cbdba4](headroomlabs-ai@9cbdba4)) * **ccr:** preserve Anthropic re-stream shape ([headroomlabs-ai#1854](headroomlabs-ai#1854)) ([f663894](headroomlabs-ai@f663894)) * **ccr:** preserve thinking blocks in buffered stream re-synthesis ([headroomlabs-ai#1897](headroomlabs-ai#1897)) ([ede085c](headroomlabs-ai@ede085c)) * **cli/proxy:** preserve explicit HEADROOM_MIN_TOKENS=0 / MAX_ITEMS=0 ([headroomlabs-ai#1886](headroomlabs-ai#1886)) ([3a33af1](headroomlabs-ai@3a33af1)) * **code-compressor:** CJK-aware relevance-query symbol matching ([headroomlabs-ai#1747](headroomlabs-ai#1747)) ([b38315c](headroomlabs-ai@b38315c)) * **codex:** discover updated Codex state stores ([headroomlabs-ai#1889](headroomlabs-ai#1889)) ([9d42eba](headroomlabs-ai@9d42eba)) * **codex:** OpenCode Zen telemetry attribution ([headroomlabs-ai#1648](headroomlabs-ai#1648)) ([f18c6bd](headroomlabs-ai@f18c6bd)) * **content-detector:** detect and compress space-separated JSON objects ([headroomlabs-ai#1742](headroomlabs-ai#1742)) ([5194bdc](headroomlabs-ai@5194bdc)) * **content-router:** token-measure lossless folds at the acceptance gate ([headroomlabs-ai#1772](headroomlabs-ai#1772)) ([c5493ea](headroomlabs-ai@c5493ea)) * **copilot:** normalize subscription routing host ([headroomlabs-ai#1836](headroomlabs-ai#1836)) ([afd9cbd](headroomlabs-ai@afd9cbd)) * **copilot:** route mixed-model requests per model ([headroomlabs-ai#1785](headroomlabs-ai#1785)) ([5af5e22](headroomlabs-ai@5af5e22)) * **dashboard:** deduplicate repeated savings metrics ([headroomlabs-ai#1804](headroomlabs-ai#1804)) ([88f935a](headroomlabs-ai@88f935a)) * **dashboard:** distinguish unavailable RTK from zero stats in Docker ([headroomlabs-ai#1900](headroomlabs-ai#1900)) ([87f6e93](headroomlabs-ai@87f6e93)) * **dashboard:** distinguish unavailable RTK from zero stats in Docker ([headroomlabs-ai#1901](headroomlabs-ai#1901)) ([361adcd](headroomlabs-ai@361adcd)) * **dashboard:** price proxy savings without litellm ([headroomlabs-ai#1728](headroomlabs-ai#1728)) ([188e382](headroomlabs-ai@188e382)) * detect and clear stale ANTHROPIC_BASE_URL from crashed wrap sessions ([headroomlabs-ai#1768](headroomlabs-ai#1768)) ([headroomlabs-ai#1837](headroomlabs-ai#1837)) ([84509a4](headroomlabs-ai@84509a4)) * **docker:** persist headroom workspace in compose ([headroomlabs-ai#1839](headroomlabs-ai#1839)) ([5e29c06](headroomlabs-ai@5e29c06)) * **docker:** report source build version ([headroomlabs-ai#1862](headroomlabs-ai#1862)) ([3807488](headroomlabs-ai@3807488)) * **evals:** default unparseable judge scores below pass threshold ([headroomlabs-ai#1892](headroomlabs-ai#1892)) ([42ebbc6](headroomlabs-ai@42ebbc6)) * **install:** pass sc.exe create as raw command line so binPath= quoting survives ([headroomlabs-ai#1654](headroomlabs-ai#1654)) ([headroomlabs-ai#1702](headroomlabs-ai#1702)) ([d6e0710](headroomlabs-ai@d6e0710)) * **install:** persist --no-http2 override through install apply ([headroomlabs-ai#1676](headroomlabs-ai#1676)) ([6fb5f3b](headroomlabs-ai@6fb5f3b)) * **mcp:** isolate ClaudeRegistrar CLI config env ([headroomlabs-ai#1888](headroomlabs-ai#1888)) ([1c947b1](headroomlabs-ai@1c947b1)) * **mcp:** surface dead proxy state ([headroomlabs-ai#1786](headroomlabs-ai#1786)) ([931eed8](headroomlabs-ai@931eed8)) * **memory:** resolve Trae cwd metadata from user reminders ([headroomlabs-ai#1737](headroomlabs-ai#1737)) ([headroomlabs-ai#1887](headroomlabs-ai#1887)) ([3e85eb1](headroomlabs-ai@3e85eb1)) * **opencode:** use local MCP config ([headroomlabs-ai#1383](headroomlabs-ai#1383)) ([4bd3ddf](headroomlabs-ai@4bd3ddf)) * **proxy/openai:** thread savings-profile kwargs into chat completions ([headroomlabs-ai#1606](headroomlabs-ai#1606)) ([7ff842d](headroomlabs-ai@7ff842d)) * **proxy/openai:** translate max_tokens -> max_completion_tokens on chat path ([headroomlabs-ai#1774](headroomlabs-ai#1774)) ([285808b](headroomlabs-ai@285808b)) * **proxy:** bound Codex WS compression fallback latency ([headroomlabs-ai#1802](headroomlabs-ai#1802)) ([d24a3f8](headroomlabs-ai@d24a3f8)) * **proxy:** bound HF tokenizer load and offload token counting off event loop ([headroomlabs-ai#1738](headroomlabs-ai#1738)) ([46d5d68](headroomlabs-ai@46d5d68)) * **proxy:** cancel retry backoff on shutdown ([headroomlabs-ai#1834](headroomlabs-ai#1834)) ([da2d8dc](headroomlabs-ai@da2d8dc)) * **proxy:** compress Anthropic user text blocks when enabled ([headroomlabs-ai#1875](headroomlabs-ai#1875)) ([e36439a](headroomlabs-ai@e36439a)) * **proxy:** freeze must forward cached (compressed) prefix byte-identical — stop token-mode cache busting ([headroomlabs-ai#1850](headroomlabs-ai#1850)) ([248ae0f](headroomlabs-ai@248ae0f)) * **proxy:** fsync savings dir after atomic rename ([headroomlabs-ai#1764](headroomlabs-ai#1764)) ([7de2c1e](headroomlabs-ai@7de2c1e)) * **proxy:** keep cache_control bounded + stable so the freeze overlay stops busting ([headroomlabs-ai#1852](headroomlabs-ai#1852)) ([4820134](headroomlabs-ai@4820134)) * **proxy:** persist lifetime cache-read savings across restarts ([headroomlabs-ai#1665](headroomlabs-ai#1665)) ([908997e](headroomlabs-ai@908997e)) * **proxy:** preserve streaming passthrough beta headers ([headroomlabs-ai#1783](headroomlabs-ai#1783)) ([0f553a8](headroomlabs-ai@0f553a8)) * **proxy:** release _active_streams session lock on setup-phase errors ([headroomlabs-ai#1864](headroomlabs-ai#1864)) ([2ccd831](headroomlabs-ai@2ccd831)) * **proxy:** retry HTTP/2 stream resets instead of 502ing ([headroomlabs-ai#1645](headroomlabs-ai#1645)) ([2ce19c2](headroomlabs-ai@2ce19c2)) * **proxy:** retry passthrough on transient upstream connection close ([headroomlabs-ai#1513](headroomlabs-ai#1513)) ([5d14080](headroomlabs-ai@5d14080)) * **proxy:** route Foundry Anthropic messages ([headroomlabs-ai#1878](headroomlabs-ai#1878)) ([739f654](headroomlabs-ai@739f654)) * **proxy:** serve /favicon.ico locally instead of tunneling upstream ([headroomlabs-ai#1787](headroomlabs-ai#1787)) ([headroomlabs-ai#1847](headroomlabs-ai#1847)) ([3076e32](headroomlabs-ai@3076e32)) * **proxy:** stop rtk stat failures from corrupting session baseline ([headroomlabs-ai#1693](headroomlabs-ai#1693)) ([681b9a8](headroomlabs-ai@681b9a8)) * **proxy:** strip 1m model suffix before upstream forwarding ([headroomlabs-ai#1840](headroomlabs-ai#1840)) ([e22d745](headroomlabs-ai@e22d745)) * **proxy:** subtract cache write premiums from net savings ([headroomlabs-ai#1800](headroomlabs-ai#1800)) ([53a465b](headroomlabs-ai@53a465b)) * **router:** honor MCP aliases in excluded tools ([headroomlabs-ai#1822](headroomlabs-ai#1822)) ([headroomlabs-ai#1863](headroomlabs-ai#1863)) ([140d6e4](headroomlabs-ai@140d6e4)) * **rtk:** link managed rtk onto PATH instead of mutating the hook ([headroomlabs-ai#1698](headroomlabs-ai#1698)) ([140cb05](headroomlabs-ai@140cb05)) * **streaming:** preserve server_tool_use sse blocks ([headroomlabs-ai#1826](headroomlabs-ai#1826)) ([4ac5493](headroomlabs-ai@4ac5493)) * **toin:** publish skip compression recommendations ([headroomlabs-ai#1782](headroomlabs-ai#1782)) ([be51008](headroomlabs-ai@be51008)) * **transforms:** normalize diff compressor context ([headroomlabs-ai#1801](headroomlabs-ai#1801)) ([838c523](headroomlabs-ai@838c523)) * **transforms:** pass through ragged tables instead of misaligning columns ([headroomlabs-ai#1713](headroomlabs-ai#1713)) ([c7665ca](headroomlabs-ai@c7665ca)) * use rtk native Cursor hook instead of injecting .cursorrules ([headroomlabs-ai#756](headroomlabs-ai#756)) ([headroomlabs-ai#1846](headroomlabs-ai#1846)) ([1573f1f](headroomlabs-ai@1573f1f)) * **wrap:** replace stale-proxy detection with Vite-style port fallback ([headroomlabs-ai#1406](headroomlabs-ai#1406)) ([b4205c6](headroomlabs-ai@b4205c6)) ### Performance Improvements * **proxy:** cap compression workers to CPU count ([headroomlabs-ai#1803](headroomlabs-ai#1803)) ([0a3851b](headroomlabs-ai@0a3851b)) * **savings:** batch tracker persistence off the request hot path ([headroomlabs-ai#1817](headroomlabs-ai#1817)) ([451b9f0](headroomlabs-ai@451b9f0)) ### Dependencies * bump the cargo-minor-patch group across 1 directory with 7 updates ([headroomlabs-ai#1909](headroomlabs-ai#1909)) ([45601d9](headroomlabs-ai@45601d9)) * bump the npm-minor-patch group across 4 directories with 18 updates ([headroomlabs-ai#1907](headroomlabs-ai#1907)) ([8872bbc](headroomlabs-ai@8872bbc)) </details> --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Adds a small, neutral extension point to the proxy: a "turn hook" that lets an opt-in extension observe and optionally re-drive a single buffered model turn, without touching the core request/response flow for anyone who has no extension installed.
A hook can:
on_request(ctx)— inspect or rewrite the outbound tools/messages before they go upstream (the extensible counterpart to the built-in tool-search deferral that already lives at that point).on_response(ctx, response, call_model)— inspect the model's response and, if it wants, call the model again (viacall_model) and return a replacement response — transparently to the client. This is the capability that can't be done from ASGI middleware: it reuses the proxy-internal re-call path (the sameapi_call_fnthe CCR handler already drives).The module is inert unless a hook is registered: the runners return their input unchanged and are gated on the registry, so with no extension the proxy is byte-identical to today. A failing hook is logged and skipped — it can never take the proxy down.
Closes #
Type of Change
Changes Made
headroom/proxy/turn_hooks.py:TurnContext, theTurnHookprotocol (on_request/on_response), a module registry (register_turn_hook/registered_turn_hooks/clear_turn_hooks), and the runnersrun_request_hooks/run_response_hooks. Inert when empty; never raises.tests/test_turn_hooks.py.Testing
pytest)ruff check .)mypy headroom)Test Output
Real Behavior Proof
.venv(Python 3.12.6);ruffpinned to CI's0.15.17viauvx ruff@0.15.17;mypyfrom the venv.upstream/main; added the hook module + wired the four handler seams; ran the ruff/format/mypy/pytest commands above.on_requestmutatingctx.tools,on_responsereturning a replacement, theawait call_model(...)re-drive loop, replacement-chaining across hooks, and the never-raise guarantee. The existing CCR + handler suites pass unchanged, which is the point: with no hook registered the added code is a no-op (the runners short-circuit on an empty registry).call_model. Theon_requestseam fires on the Anthropic pre-send and OpenAI Responses paths (where the existing tool-search deferral runs); other send paths (e.g. chat-completions, streaming) are not wired in this PR.Review Readiness
Checklist