Skip to content

feat(proxy): add turn-hook extension point for buffered model turns - #1891

Merged
chopratejas merged 1 commit into
mainfrom
tejas/turn-hooks-extension
Jul 9, 2026
Merged

chopratejas merged 1 commit into
mainfrom
tejas/turn-hooks-extension

Conversation

@chopratejas

Copy link
Copy Markdown
Collaborator

Description

Adds a small, neutral extension point to the proxy: a "turn hook" that lets an opt-in extension observe and optionally re-drive a single buffered model turn, without touching the core request/response flow for anyone who has no extension installed.

A hook can:

  • on_request(ctx) — inspect or rewrite the outbound tools/messages before they go upstream (the extensible counterpart to the built-in tool-search deferral that already lives at that point).
  • on_response(ctx, response, call_model) — inspect the model's response and, if it wants, call the model again (via call_model) and return a replacement response — transparently to the client. This is the capability that can't be done from ASGI middleware: it reuses the proxy-internal re-call path (the same api_call_fn the CCR handler already drives).

The module is inert unless a hook is registered: the runners return their input unchanged and are gated on the registry, so with no extension the proxy is byte-identical to today. A failing hook is logged and skipped — it can never take the proxy down.

Closes #

Type of Change

  • Bug fix (non-breaking change that fixes an issue)
  • New feature (non-breaking change that adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update
  • Performance improvement
  • Code refactoring (no functional changes)

Changes Made

  • Add headroom/proxy/turn_hooks.py: TurnContext, the TurnHook protocol (on_request / on_response), a module registry (register_turn_hook / registered_turn_hooks / clear_turn_hooks), and the runners run_request_hooks / run_response_hooks. Inert when empty; never raises.
  • Wire it at four seams, each gated so an empty registry is a byte-identical no-op:
    • Anthropic — pre-send (right after the existing tool-search deferral) + the CCR response seam.
    • OpenAI — the Responses tool-shaping point (right after the existing tool-search deferral, copy-on-write-safe) + the CCR response seam.
  • Add tests/test_turn_hooks.py.

Testing

  • Unit tests pass (pytest)
  • Linting passes (ruff check .)
  • Type checking passes (mypy headroom)
  • New tests added for new functionality
  • Manual testing performed (no-op regression across CCR/handler suites)

Test Output

$ ruff check headroom/proxy/turn_hooks.py tests/test_turn_hooks.py \
      headroom/proxy/handlers/anthropic.py headroom/proxy/handlers/openai.py
All checks passed!

$ ruff format --check <same 4 files>
4 files already formatted

$ mypy headroom
Success: no issues found in 408 source files

$ pytest tests/test_turn_hooks.py -q
9 passed in 0.17s

$ pytest tests/test_turn_hooks.py tests/test_ccr_response_handler.py \
      tests/test_ccr_tool_injection.py tests/test_proxy_ccr.py \
      tests/test_openai_tool_search_deferral.py \
      tests/test_openai_responses_compression_units.py \
      tests/test_handler_outcome_tag_invariant.py -q
135 passed  (+ 1 pre-existing cross-file flake in test_proxy_ccr::test_health_endpoint,
             which passes in isolation and in its own file: `pytest tests/test_proxy_ccr.py` -> 19 passed)

Real Behavior Proof

  • Environment: local macOS, project .venv (Python 3.12.6); ruff pinned to CI's 0.15.17 via uvx ruff@0.15.17; mypy from the venv.
  • Exact command / steps: branched off upstream/main; added the hook module + wired the four handler seams; ran the ruff/format/mypy/pytest commands above.
  • Observed result: The unit tests exercise the whole contract — registry, on_request mutating ctx.tools, on_response returning a replacement, the await call_model(...) re-drive loop, replacement-chaining across hooks, and the never-raise guarantee. The existing CCR + handler suites pass unchanged, which is the point: with no hook registered the added code is a no-op (the runners short-circuit on an empty registry).
  • Not tested: the live interactive re-drive path with a registered hook against a real upstream — no hook ships in this repo, so that path is covered here only by the unit test's fake call_model. The on_request seam fires on the Anthropic pre-send and OpenAI Responses paths (where the existing tool-search deferral runs); other send paths (e.g. chat-completions, streaming) are not wired in this PR.

Review Readiness

  • I have performed a self-review
  • This PR is ready for human review

Checklist

  • My code follows the project's style guidelines
  • I have performed a self-review of my code
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes
  • I have updated the CHANGELOG.md if applicable

Add headroom/proxy/turn_hooks.py — a neutral extension contract that lets an
opt-in proxy extension observe and optionally re-drive a single buffered model
turn. A hook's on_request may rewrite the outbound tools/messages; its
on_response may call the model again (via call_model) and return a replacement
response, transparently to the client.

Wire it at four seams in the Anthropic and OpenAI handlers — the pre-send point
(alongside the existing tool-search deferral) and the CCR response seam — each
gated so an empty registry is a byte-identical no-op. With no extension
installed the proxy behaves exactly as before.

Covered by tests/test_turn_hooks.py (registry, request mutation, response
replacement + re-drive loop, never-raise). Existing CCR/handler suites pass
unchanged, proving the no-op property.
@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown
Contributor

PR governance

This PR does not yet satisfy the required template fields:

  • Fill in Real Behavior Proof → Environment.
  • Fill in Real Behavior Proof → Exact command / steps.
  • Fill in Real Behavior Proof → Observed result.
  • Fill in Real Behavior Proof → Not tested.

Please update the PR body, or move the PR back to draft while it is still in progress.

@github-actions github-actions Bot added the status: needs author action Pull request body or readiness checklist still needs author updates label Jul 9, 2026
@chopratejas
chopratejas merged commit ec950f7 into main Jul 9, 2026
27 checks passed
chopratejas added a commit that referenced this pull request Jul 9, 2026
…hema savings (#1896)

## Description

**Stacked on #1891** (base = `tejas/turn-hooks-extension`; the diff here
is only the wiring + dashboard, and GitHub will retarget to `main` once
#1891 merges).

Makes the turn-hook seam actually useful on the OpenAI
`/v1/chat/completions` **direct path** (the route an `opencode → OpenAI`
session uses), and surfaces the resulting tool-schema token savings in
the dashboard.

- **Shrink** (`on_request`) at the shared pre-send point, gated on the
registry and non-streaming — a hook can rewrite the outbound
`tools`/`messages`; the net tool-schema token delta is recorded.
- **Reload** (`on_response`) after the buffered send, with a
`_retry_request`-based `call_model` — a hook can resolve a tool the
model asked to load, re-drive the turn, and have the proxy return the
final response transparently. No-op when no hook is registered.
- **Dashboard**: a new `savings.by_layer.tool_search` aggregate and a
per-request `tool_schema_saved_tokens` field sum both the native
tool-search deferral and the hook-driven tools rewrite — attributed to
Headroom only, so a client that already had tool search on (e.g. Claude
Code / Codex) contributes zero. Adds a "Tool-Schema Deferral" card
(rendered only when there is a saving) + a per-call badge.

Recorded **per turn**: tools are re-sent on every turn, so the saving is
logged on each request and summed across the recent window.

Closes #

## Type of Change

- [ ] Bug fix (non-breaking change that fixes an issue)
- [x] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)

## Changes Made

- `handlers/openai.py`: `on_request` shrink at the pre-send point +
`on_response` reload loop on the direct buffered path; record the net
tool-schema token delta as a saving. Savings are measured from the FINAL
tools object regardless of whether the hook replaced or mutated
`ctx.tools` in place (identity is used only to decide `body["tools"]`
reassignment).
- `server.py`: `_tool_schema_saved_from_tags()` helper; aggregate
`savings.by_layer.tool_search` in `_build_stats_payload`; expose
`tool_schema_saved_tokens` per request.
- `dashboard/templates/dashboard.html`: "Tool-Schema Deferral" card
(only renders when there is a saving) + per-call badge.
- `tests/test_openai_chat_turn_hooks.py`: end-to-end shrink+reload,
per-turn tag + `/stats` aggregation, an **in-place-mutating** hook
regression (pins the extension contract + dashboard path), and the
no-op-when-unregistered invariant.

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed (real extension end-to-end, incl. live
OpenAI + Kimi/Fireworks)

### Test Output

```text
$ ruff check headroom/proxy/handlers/openai.py headroom/proxy/server.py tests/test_openai_chat_turn_hooks.py
All checks passed!

$ ruff format --check <same 3 files>
3 files already formatted

$ mypy headroom
Success: no issues found in 408 source files

$ pytest tests/test_openai_chat_turn_hooks.py -q
4 passed

# regression (no-op safety + payload shape)
$ pytest tests/test_turn_hooks.py tests/test_ccr_response_handler.py tests/test_proxy_ccr.py \
        tests/test_openai_responses_compression_units.py tests/test_handler_outcome_tag_invariant.py -q
81 passed
$ pytest tests/test_proxy_savings_history.py tests/test_telemetry_warning.py -q
69 passed
```

## Real Behavior Proof

- Environment: local macOS, project `.venv` (Python 3.12.6); `ruff`
pinned to CI's `0.15.17` via `uvx ruff@0.15.17`; `mypy` from the venv.
- Exact command / steps: drove `handle_openai_chat` via FastAPI
`TestClient` with a registered hook. First with a mocked upstream
(`_retry_request`) for deterministic assertions, then against the real
out-of-tree tool-router extension over the LIVE OpenAI API (gpt-4o-mini)
and the LIVE Fireworks API (Kimi K2.6, OpenAI-compatible), each with a
13-tool belt where the needed tool is deferred.
- Observed result: round-1 outbound tools shrank to `[6 core] +
search_tools`; the model chose to call `search_tools`; the extension
resolved the right tool via BM25; round-2 re-drove with the resolved
tools appended and the model called `github_create_issue`; the proxy
returned the final response; `x-headroom-transforms` carried
`turn_hook:tools:<N>tok` (188 tok on the live Kimi run) and `/stats`
`by_layer.tool_search` summed it per turn. The in-place-mutation
regression test confirms savings are recorded when a hook mutates
`ctx.tools` without replacing it.
- Not tested: a full standalone `headroom proxy` process with a real
opencode client (used `TestClient` + real upstream); the exercised
handler path is identical, only the client differs. Streaming is
intentionally out of scope (a streamed turn can't be re-driven).

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable
@github-actions github-actions Bot mentioned this pull request Jul 9, 2026
chopratejas pushed a commit that referenced this pull request Jul 9, 2026
🤖 I have created a release *beep* *boop*
---


<details><summary>0.31.0</summary>

##
[0.31.0](v0.30.0...v0.31.0)
(2026-07-09)


### Features

* **cache:** provider-agnostic cache-mode delta + cc-agnostic prefix
comparison
([#1868](#1868))
([7c2f0ea](7c2f0ea))
* **ccr:** wire retrieve-tool interception into OpenAI Responses handler
([#1898](#1898))
([62cd307](62cd307))
* **compression:** add audit-safe mode with protected pattern matching
([#1899](#1899))
([bb112dd](bb112dd))
* **content-router:** accept any real compression (remove min-savings
floor)
([#1771](#1771))
([6c31db9](6c31db9))
* **content-router:** lossless-first dispatch, cross-turn dedup, and A7
lossy-after-fold
([#1818](#1818))
([60af15f](60af15f))
* **proxy:** add provider-only HTTP proxy
([#1807](#1807))
([ebe0a3b](ebe0a3b))
* **proxy:** add turn-hook extension point for buffered model turns
([#1891](#1891))
([ec950f7](ec950f7))


### Bug Fixes

* **build:** enable Intel macOS pip installs via ort-load-dynamic
([#1538](#1538))
([32ce99e](32ce99e))
* **cache:** avoid fallback session collisions
([#1827](#1827))
([0f606b6](0f606b6))
* **ccr:** make expired retrieve misses terminal
([#1781](#1781))
([9cbdba4](9cbdba4))
* **ccr:** preserve Anthropic re-stream shape
([#1854](#1854))
([f663894](f663894))
* **ccr:** preserve thinking blocks in buffered stream re-synthesis
([#1897](#1897))
([ede085c](ede085c))
* **cli/proxy:** preserve explicit HEADROOM_MIN_TOKENS=0 / MAX_ITEMS=0
([#1886](#1886))
([3a33af1](3a33af1))
* **code-compressor:** CJK-aware relevance-query symbol matching
([#1747](#1747))
([b38315c](b38315c))
* **codex:** discover updated Codex state stores
([#1889](#1889))
([9d42eba](9d42eba))
* **codex:** OpenCode Zen telemetry attribution
([#1648](#1648))
([f18c6bd](f18c6bd))
* **content-detector:** detect and compress space-separated JSON objects
([#1742](#1742))
([5194bdc](5194bdc))
* **content-router:** token-measure lossless folds at the acceptance
gate ([#1772](#1772))
([c5493ea](c5493ea))
* **copilot:** normalize subscription routing host
([#1836](#1836))
([afd9cbd](afd9cbd))
* **copilot:** route mixed-model requests per model
([#1785](#1785))
([5af5e22](5af5e22))
* **dashboard:** deduplicate repeated savings metrics
([#1804](#1804))
([88f935a](88f935a))
* **dashboard:** distinguish unavailable RTK from zero stats in Docker
([#1900](#1900))
([87f6e93](87f6e93))
* **dashboard:** distinguish unavailable RTK from zero stats in Docker
([#1901](#1901))
([361adcd](361adcd))
* **dashboard:** price proxy savings without litellm
([#1728](#1728))
([188e382](188e382))
* detect and clear stale ANTHROPIC_BASE_URL from crashed wrap sessions
([#1768](#1768))
([#1837](#1837))
([84509a4](84509a4))
* **docker:** persist headroom workspace in compose
([#1839](#1839))
([5e29c06](5e29c06))
* **docker:** report source build version
([#1862](#1862))
([3807488](3807488))
* **evals:** default unparseable judge scores below pass threshold
([#1892](#1892))
([42ebbc6](42ebbc6))
* **install:** pass sc.exe create as raw command line so binPath=
quoting survives
([#1654](#1654))
([#1702](#1702))
([d6e0710](d6e0710))
* **install:** persist --no-http2 override through install apply
([#1676](#1676))
([6fb5f3b](6fb5f3b))
* **mcp:** isolate ClaudeRegistrar CLI config env
([#1888](#1888))
([1c947b1](1c947b1))
* **mcp:** surface dead proxy state
([#1786](#1786))
([931eed8](931eed8))
* **memory:** resolve Trae cwd metadata from user reminders
([#1737](#1737))
([#1887](#1887))
([3e85eb1](3e85eb1))
* **opencode:** use local MCP config
([#1383](#1383))
([4bd3ddf](4bd3ddf))
* **proxy/openai:** thread savings-profile kwargs into chat completions
([#1606](#1606))
([7ff842d](7ff842d))
* **proxy/openai:** translate max_tokens -&gt; max_completion_tokens on
chat path
([#1774](#1774))
([285808b](285808b))
* **proxy:** bound Codex WS compression fallback latency
([#1802](#1802))
([d24a3f8](d24a3f8))
* **proxy:** bound HF tokenizer load and offload token counting off
event loop
([#1738](#1738))
([46d5d68](46d5d68))
* **proxy:** cancel retry backoff on shutdown
([#1834](#1834))
([da2d8dc](da2d8dc))
* **proxy:** compress Anthropic user text blocks when enabled
([#1875](#1875))
([e36439a](e36439a))
* **proxy:** freeze must forward cached (compressed) prefix
byte-identical — stop token-mode cache busting
([#1850](#1850))
([248ae0f](248ae0f))
* **proxy:** fsync savings dir after atomic rename
([#1764](#1764))
([7de2c1e](7de2c1e))
* **proxy:** keep cache_control bounded + stable so the freeze overlay
stops busting
([#1852](#1852))
([4820134](4820134))
* **proxy:** persist lifetime cache-read savings across restarts
([#1665](#1665))
([908997e](908997e))
* **proxy:** preserve streaming passthrough beta headers
([#1783](#1783))
([0f553a8](0f553a8))
* **proxy:** release _active_streams session lock on setup-phase errors
([#1864](#1864))
([2ccd831](2ccd831))
* **proxy:** retry HTTP/2 stream resets instead of 502ing
([#1645](#1645))
([2ce19c2](2ce19c2))
* **proxy:** retry passthrough on transient upstream connection close
([#1513](#1513))
([5d14080](5d14080))
* **proxy:** route Foundry Anthropic messages
([#1878](#1878))
([739f654](739f654))
* **proxy:** serve /favicon.ico locally instead of tunneling upstream
([#1787](#1787))
([#1847](#1847))
([3076e32](3076e32))
* **proxy:** stop rtk stat failures from corrupting session baseline
([#1693](#1693))
([681b9a8](681b9a8))
* **proxy:** strip 1m model suffix before upstream forwarding
([#1840](#1840))
([e22d745](e22d745))
* **proxy:** subtract cache write premiums from net savings
([#1800](#1800))
([53a465b](53a465b))
* **router:** honor MCP aliases in excluded tools
([#1822](#1822))
([#1863](#1863))
([140d6e4](140d6e4))
* **rtk:** link managed rtk onto PATH instead of mutating the hook
([#1698](#1698))
([140cb05](140cb05))
* **streaming:** preserve server_tool_use sse blocks
([#1826](#1826))
([4ac5493](4ac5493))
* **toin:** publish skip compression recommendations
([#1782](#1782))
([be51008](be51008))
* **transforms:** normalize diff compressor context
([#1801](#1801))
([838c523](838c523))
* **transforms:** pass through ragged tables instead of misaligning
columns
([#1713](#1713))
([c7665ca](c7665ca))
* use rtk native Cursor hook instead of injecting .cursorrules
([#756](#756))
([#1846](#1846))
([1573f1f](1573f1f))
* **wrap:** replace stale-proxy detection with Vite-style port fallback
([#1406](#1406))
([b4205c6](b4205c6))


### Performance Improvements

* **proxy:** cap compression workers to CPU count
([#1803](#1803))
([0a3851b](0a3851b))
* **savings:** batch tracker persistence off the request hot path
([#1817](#1817))
([451b9f0](451b9f0))


### Dependencies

* bump the cargo-minor-patch group across 1 directory with 7 updates
([#1909](#1909))
([45601d9](45601d9))
* bump the npm-minor-patch group across 4 directories with 18 updates
([#1907](#1907))
([8872bbc](8872bbc))
</details>

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
JerrettDavis pushed a commit to peterlodri-sec/headroom that referenced this pull request Jul 14, 2026
…eadroomlabs-ai#1891)

## Description

Adds a small, neutral **extension point** to the proxy: a "turn hook"
that lets an opt-in extension observe and optionally re-drive a single
buffered model turn, without touching the core request/response flow for
anyone who has no extension installed.

A hook can:
- `on_request(ctx)` — inspect or rewrite the outbound tools/messages
before they go upstream (the extensible counterpart to the built-in
tool-search deferral that already lives at that point).
- `on_response(ctx, response, call_model)` — inspect the model's
response and, if it wants, call the model again (via `call_model`) and
return a **replacement** response — transparently to the client. This is
the capability that can't be done from ASGI middleware: it reuses the
proxy-internal re-call path (the same `api_call_fn` the CCR handler
already drives).

The module is **inert unless a hook is registered**: the runners return
their input unchanged and are gated on the registry, so with no
extension the proxy is byte-identical to today. A failing hook is logged
and skipped — it can never take the proxy down.

Closes #

## Type of Change

- [ ] Bug fix (non-breaking change that fixes an issue)
- [x] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)

## Changes Made

- Add `headroom/proxy/turn_hooks.py`: `TurnContext`, the `TurnHook`
protocol (`on_request` / `on_response`), a module registry
(`register_turn_hook` / `registered_turn_hooks` / `clear_turn_hooks`),
and the runners `run_request_hooks` / `run_response_hooks`. Inert when
empty; never raises.
- Wire it at four seams, each gated so an empty registry is a
byte-identical no-op:
- Anthropic — pre-send (right after the existing tool-search deferral) +
the CCR response seam.
- OpenAI — the Responses tool-shaping point (right after the existing
tool-search deferral, copy-on-write-safe) + the CCR response seam.
- Add `tests/test_turn_hooks.py`.

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed (no-op regression across CCR/handler
suites)

### Test Output

```text
$ ruff check headroom/proxy/turn_hooks.py tests/test_turn_hooks.py \
      headroom/proxy/handlers/anthropic.py headroom/proxy/handlers/openai.py
All checks passed!

$ ruff format --check <same 4 files>
4 files already formatted

$ mypy headroom
Success: no issues found in 408 source files

$ pytest tests/test_turn_hooks.py -q
9 passed in 0.17s

$ pytest tests/test_turn_hooks.py tests/test_ccr_response_handler.py \
      tests/test_ccr_tool_injection.py tests/test_proxy_ccr.py \
      tests/test_openai_tool_search_deferral.py \
      tests/test_openai_responses_compression_units.py \
      tests/test_handler_outcome_tag_invariant.py -q
135 passed  (+ 1 pre-existing cross-file flake in test_proxy_ccr::test_health_endpoint,
             which passes in isolation and in its own file: `pytest tests/test_proxy_ccr.py` -> 19 passed)
```

## Real Behavior Proof

- **Environment:** local macOS, project `.venv` (Python 3.12.6); `ruff`
pinned to CI's `0.15.17` via `uvx ruff@0.15.17`; `mypy` from the venv.
- **Exact command / steps:** branched off `upstream/main`; added the
hook module + wired the four handler seams; ran the
ruff/format/mypy/pytest commands above.
- **Observed result:** The unit tests exercise the whole contract —
registry, `on_request` mutating `ctx.tools`, `on_response` returning a
replacement, the `await call_model(...)` re-drive loop,
replacement-chaining across hooks, and the never-raise guarantee. The
existing CCR + handler suites pass unchanged, which is the point: with
no hook registered the added code is a no-op (the runners short-circuit
on an empty registry).
- **Not tested:** the live interactive re-drive path with a *registered*
hook against a real upstream — no hook ships in this repo, so that path
is covered here only by the unit test's fake `call_model`. The
`on_request` seam fires on the Anthropic pre-send and OpenAI Responses
paths (where the existing tool-search deferral runs); other send paths
(e.g. chat-completions, streaming) are not wired in this PR.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable
JerrettDavis pushed a commit to peterlodri-sec/headroom that referenced this pull request Jul 14, 2026
🤖 I have created a release *beep* *boop*
---


<details><summary>0.31.0</summary>

##
[0.31.0](headroomlabs-ai/headroom@v0.30.0...v0.31.0)
(2026-07-09)


### Features

* **cache:** provider-agnostic cache-mode delta + cc-agnostic prefix
comparison
([headroomlabs-ai#1868](headroomlabs-ai#1868))
([7c2f0ea](headroomlabs-ai@7c2f0ea))
* **ccr:** wire retrieve-tool interception into OpenAI Responses handler
([headroomlabs-ai#1898](headroomlabs-ai#1898))
([62cd307](headroomlabs-ai@62cd307))
* **compression:** add audit-safe mode with protected pattern matching
([headroomlabs-ai#1899](headroomlabs-ai#1899))
([bb112dd](headroomlabs-ai@bb112dd))
* **content-router:** accept any real compression (remove min-savings
floor)
([headroomlabs-ai#1771](headroomlabs-ai#1771))
([6c31db9](headroomlabs-ai@6c31db9))
* **content-router:** lossless-first dispatch, cross-turn dedup, and A7
lossy-after-fold
([headroomlabs-ai#1818](headroomlabs-ai#1818))
([60af15f](headroomlabs-ai@60af15f))
* **proxy:** add provider-only HTTP proxy
([headroomlabs-ai#1807](headroomlabs-ai#1807))
([ebe0a3b](headroomlabs-ai@ebe0a3b))
* **proxy:** add turn-hook extension point for buffered model turns
([headroomlabs-ai#1891](headroomlabs-ai#1891))
([ec950f7](headroomlabs-ai@ec950f7))


### Bug Fixes

* **build:** enable Intel macOS pip installs via ort-load-dynamic
([headroomlabs-ai#1538](headroomlabs-ai#1538))
([32ce99e](headroomlabs-ai@32ce99e))
* **cache:** avoid fallback session collisions
([headroomlabs-ai#1827](headroomlabs-ai#1827))
([0f606b6](headroomlabs-ai@0f606b6))
* **ccr:** make expired retrieve misses terminal
([headroomlabs-ai#1781](headroomlabs-ai#1781))
([9cbdba4](headroomlabs-ai@9cbdba4))
* **ccr:** preserve Anthropic re-stream shape
([headroomlabs-ai#1854](headroomlabs-ai#1854))
([f663894](headroomlabs-ai@f663894))
* **ccr:** preserve thinking blocks in buffered stream re-synthesis
([headroomlabs-ai#1897](headroomlabs-ai#1897))
([ede085c](headroomlabs-ai@ede085c))
* **cli/proxy:** preserve explicit HEADROOM_MIN_TOKENS=0 / MAX_ITEMS=0
([headroomlabs-ai#1886](headroomlabs-ai#1886))
([3a33af1](headroomlabs-ai@3a33af1))
* **code-compressor:** CJK-aware relevance-query symbol matching
([headroomlabs-ai#1747](headroomlabs-ai#1747))
([b38315c](headroomlabs-ai@b38315c))
* **codex:** discover updated Codex state stores
([headroomlabs-ai#1889](headroomlabs-ai#1889))
([9d42eba](headroomlabs-ai@9d42eba))
* **codex:** OpenCode Zen telemetry attribution
([headroomlabs-ai#1648](headroomlabs-ai#1648))
([f18c6bd](headroomlabs-ai@f18c6bd))
* **content-detector:** detect and compress space-separated JSON objects
([headroomlabs-ai#1742](headroomlabs-ai#1742))
([5194bdc](headroomlabs-ai@5194bdc))
* **content-router:** token-measure lossless folds at the acceptance
gate ([headroomlabs-ai#1772](headroomlabs-ai#1772))
([c5493ea](headroomlabs-ai@c5493ea))
* **copilot:** normalize subscription routing host
([headroomlabs-ai#1836](headroomlabs-ai#1836))
([afd9cbd](headroomlabs-ai@afd9cbd))
* **copilot:** route mixed-model requests per model
([headroomlabs-ai#1785](headroomlabs-ai#1785))
([5af5e22](headroomlabs-ai@5af5e22))
* **dashboard:** deduplicate repeated savings metrics
([headroomlabs-ai#1804](headroomlabs-ai#1804))
([88f935a](headroomlabs-ai@88f935a))
* **dashboard:** distinguish unavailable RTK from zero stats in Docker
([headroomlabs-ai#1900](headroomlabs-ai#1900))
([87f6e93](headroomlabs-ai@87f6e93))
* **dashboard:** distinguish unavailable RTK from zero stats in Docker
([headroomlabs-ai#1901](headroomlabs-ai#1901))
([361adcd](headroomlabs-ai@361adcd))
* **dashboard:** price proxy savings without litellm
([headroomlabs-ai#1728](headroomlabs-ai#1728))
([188e382](headroomlabs-ai@188e382))
* detect and clear stale ANTHROPIC_BASE_URL from crashed wrap sessions
([headroomlabs-ai#1768](headroomlabs-ai#1768))
([headroomlabs-ai#1837](headroomlabs-ai#1837))
([84509a4](headroomlabs-ai@84509a4))
* **docker:** persist headroom workspace in compose
([headroomlabs-ai#1839](headroomlabs-ai#1839))
([5e29c06](headroomlabs-ai@5e29c06))
* **docker:** report source build version
([headroomlabs-ai#1862](headroomlabs-ai#1862))
([3807488](headroomlabs-ai@3807488))
* **evals:** default unparseable judge scores below pass threshold
([headroomlabs-ai#1892](headroomlabs-ai#1892))
([42ebbc6](headroomlabs-ai@42ebbc6))
* **install:** pass sc.exe create as raw command line so binPath=
quoting survives
([headroomlabs-ai#1654](headroomlabs-ai#1654))
([headroomlabs-ai#1702](headroomlabs-ai#1702))
([d6e0710](headroomlabs-ai@d6e0710))
* **install:** persist --no-http2 override through install apply
([headroomlabs-ai#1676](headroomlabs-ai#1676))
([6fb5f3b](headroomlabs-ai@6fb5f3b))
* **mcp:** isolate ClaudeRegistrar CLI config env
([headroomlabs-ai#1888](headroomlabs-ai#1888))
([1c947b1](headroomlabs-ai@1c947b1))
* **mcp:** surface dead proxy state
([headroomlabs-ai#1786](headroomlabs-ai#1786))
([931eed8](headroomlabs-ai@931eed8))
* **memory:** resolve Trae cwd metadata from user reminders
([headroomlabs-ai#1737](headroomlabs-ai#1737))
([headroomlabs-ai#1887](headroomlabs-ai#1887))
([3e85eb1](headroomlabs-ai@3e85eb1))
* **opencode:** use local MCP config
([headroomlabs-ai#1383](headroomlabs-ai#1383))
([4bd3ddf](headroomlabs-ai@4bd3ddf))
* **proxy/openai:** thread savings-profile kwargs into chat completions
([headroomlabs-ai#1606](headroomlabs-ai#1606))
([7ff842d](headroomlabs-ai@7ff842d))
* **proxy/openai:** translate max_tokens -&gt; max_completion_tokens on
chat path
([headroomlabs-ai#1774](headroomlabs-ai#1774))
([285808b](headroomlabs-ai@285808b))
* **proxy:** bound Codex WS compression fallback latency
([headroomlabs-ai#1802](headroomlabs-ai#1802))
([d24a3f8](headroomlabs-ai@d24a3f8))
* **proxy:** bound HF tokenizer load and offload token counting off
event loop
([headroomlabs-ai#1738](headroomlabs-ai#1738))
([46d5d68](headroomlabs-ai@46d5d68))
* **proxy:** cancel retry backoff on shutdown
([headroomlabs-ai#1834](headroomlabs-ai#1834))
([da2d8dc](headroomlabs-ai@da2d8dc))
* **proxy:** compress Anthropic user text blocks when enabled
([headroomlabs-ai#1875](headroomlabs-ai#1875))
([e36439a](headroomlabs-ai@e36439a))
* **proxy:** freeze must forward cached (compressed) prefix
byte-identical — stop token-mode cache busting
([headroomlabs-ai#1850](headroomlabs-ai#1850))
([248ae0f](headroomlabs-ai@248ae0f))
* **proxy:** fsync savings dir after atomic rename
([headroomlabs-ai#1764](headroomlabs-ai#1764))
([7de2c1e](headroomlabs-ai@7de2c1e))
* **proxy:** keep cache_control bounded + stable so the freeze overlay
stops busting
([headroomlabs-ai#1852](headroomlabs-ai#1852))
([4820134](headroomlabs-ai@4820134))
* **proxy:** persist lifetime cache-read savings across restarts
([headroomlabs-ai#1665](headroomlabs-ai#1665))
([908997e](headroomlabs-ai@908997e))
* **proxy:** preserve streaming passthrough beta headers
([headroomlabs-ai#1783](headroomlabs-ai#1783))
([0f553a8](headroomlabs-ai@0f553a8))
* **proxy:** release _active_streams session lock on setup-phase errors
([headroomlabs-ai#1864](headroomlabs-ai#1864))
([2ccd831](headroomlabs-ai@2ccd831))
* **proxy:** retry HTTP/2 stream resets instead of 502ing
([headroomlabs-ai#1645](headroomlabs-ai#1645))
([2ce19c2](headroomlabs-ai@2ce19c2))
* **proxy:** retry passthrough on transient upstream connection close
([headroomlabs-ai#1513](headroomlabs-ai#1513))
([5d14080](headroomlabs-ai@5d14080))
* **proxy:** route Foundry Anthropic messages
([headroomlabs-ai#1878](headroomlabs-ai#1878))
([739f654](headroomlabs-ai@739f654))
* **proxy:** serve /favicon.ico locally instead of tunneling upstream
([headroomlabs-ai#1787](headroomlabs-ai#1787))
([headroomlabs-ai#1847](headroomlabs-ai#1847))
([3076e32](headroomlabs-ai@3076e32))
* **proxy:** stop rtk stat failures from corrupting session baseline
([headroomlabs-ai#1693](headroomlabs-ai#1693))
([681b9a8](headroomlabs-ai@681b9a8))
* **proxy:** strip 1m model suffix before upstream forwarding
([headroomlabs-ai#1840](headroomlabs-ai#1840))
([e22d745](headroomlabs-ai@e22d745))
* **proxy:** subtract cache write premiums from net savings
([headroomlabs-ai#1800](headroomlabs-ai#1800))
([53a465b](headroomlabs-ai@53a465b))
* **router:** honor MCP aliases in excluded tools
([headroomlabs-ai#1822](headroomlabs-ai#1822))
([headroomlabs-ai#1863](headroomlabs-ai#1863))
([140d6e4](headroomlabs-ai@140d6e4))
* **rtk:** link managed rtk onto PATH instead of mutating the hook
([headroomlabs-ai#1698](headroomlabs-ai#1698))
([140cb05](headroomlabs-ai@140cb05))
* **streaming:** preserve server_tool_use sse blocks
([headroomlabs-ai#1826](headroomlabs-ai#1826))
([4ac5493](headroomlabs-ai@4ac5493))
* **toin:** publish skip compression recommendations
([headroomlabs-ai#1782](headroomlabs-ai#1782))
([be51008](headroomlabs-ai@be51008))
* **transforms:** normalize diff compressor context
([headroomlabs-ai#1801](headroomlabs-ai#1801))
([838c523](headroomlabs-ai@838c523))
* **transforms:** pass through ragged tables instead of misaligning
columns
([headroomlabs-ai#1713](headroomlabs-ai#1713))
([c7665ca](headroomlabs-ai@c7665ca))
* use rtk native Cursor hook instead of injecting .cursorrules
([headroomlabs-ai#756](headroomlabs-ai#756))
([headroomlabs-ai#1846](headroomlabs-ai#1846))
([1573f1f](headroomlabs-ai@1573f1f))
* **wrap:** replace stale-proxy detection with Vite-style port fallback
([headroomlabs-ai#1406](headroomlabs-ai#1406))
([b4205c6](headroomlabs-ai@b4205c6))


### Performance Improvements

* **proxy:** cap compression workers to CPU count
([headroomlabs-ai#1803](headroomlabs-ai#1803))
([0a3851b](headroomlabs-ai@0a3851b))
* **savings:** batch tracker persistence off the request hot path
([headroomlabs-ai#1817](headroomlabs-ai#1817))
([451b9f0](headroomlabs-ai@451b9f0))


### Dependencies

* bump the cargo-minor-patch group across 1 directory with 7 updates
([headroomlabs-ai#1909](headroomlabs-ai#1909))
([45601d9](headroomlabs-ai@45601d9))
* bump the npm-minor-patch group across 4 directories with 18 updates
([headroomlabs-ai#1907](headroomlabs-ai#1907))
([8872bbc](headroomlabs-ai@8872bbc))
</details>

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

status: needs author action Pull request body or readiness checklist still needs author updates

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant