Skip to content

feat(content-router): accept any real compression (remove min-savings floor) - #1771

Merged
chopratejas merged 1 commit into
mainfrom
tejas/accept-any-compression
Jul 3, 2026
Merged

chopratejas merged 1 commit into
mainfrom
tejas/accept-any-compression

Conversation

@chopratejas

Copy link
Copy Markdown
Collaborator

Description

The compression acceptance gate rejected any compression that saved less than ~15% (min_ratio interpolated 0.85 at low context pressure → 0.65 under pressure). That floor was a crude proxy for "big enough to justify busting the prefix cache," but it dropped genuine token savings — notably lossless code/log folds that shrink <15% (the ratio_too_high rejections).

This makes the gate accept any real shrink (ratio < 1.0): any token saved is worth taking. The two guards that actually protect correctness are untouched:

Lowering the two values back to 0.85/0.65 restores the savings floor.

Closes #

Type of Change

  • New feature (non-breaking change that adds functionality)
  • Performance improvement

Changes Made

  • ContentRouterConfig.min_ratio_relaxed: 0.85 → 1.0
  • ContentRouterConfig.min_ratio_aggressive: 0.65 → 1.0
  • Gate now accepts any compression_ratio < 1.0 at every context pressure; reversibility + net-cost guards unchanged.

Testing

  • Unit tests pass (pytest)
  • Linting passes (ruff check)
  • Type checking passes (mypy headroom)
  • New tests added for new functionality (existing gate-mechanism tests cover it; they pass explicit min_ratio values and are unaffected)
  • Manual testing performed

Test Output

# content-router + compression suites (default-config paths):
tests/test_transforms/test_content_router.py ......... 139 passed
# broad compression sweep (-k compress/router/crush/kompress/lossless/ccr/savings/...):
1 failed, 1940 passed, 54 skipped in 181.25s
#   the 1 failure = test_lossless_mode::test_router_lossless_search_no_marker_and_recoverable
#   — a local fastembed-cache state-leak flake; passes in isolation (1 passed in 3.15s),
#   and is in lossless mode which bypasses this gate entirely.
ruff check headroom/transforms/content_router.py  -> All checks passed!
mypy headroom/transforms/content_router.py         -> Success: no issues found

Real Behavior Proof

  • Environment: local worktree, Python 3.12, PYTHONPATH pinned to the branch checkout.
  • Exact command / steps: ran the content-router acceptance-gate suites and a compression-adjacent sweep against the branch; verified the flaky test passes in isolation.
  • Observed result: gate-mechanism tests (explicit min_ratio) unaffected; no default-floor test regressed; blocks that previously produced ratio_too_high at ratios in [0.85, 1.0) are now accepted.
  • Not tested: no live end-to-end proxy run was performed for this specific change; the behavioral effect (more router:* acceptances, fewer ratio_too_high) is inferred from the gate logic + suite.

Review Readiness

  • I have performed a self-review
  • This PR is ready for human review

Checklist

  • My code follows the project's style guidelines
  • I have performed a self-review of my code
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective (existing gate tests cover the mechanism)
  • New and existing unit tests pass locally with my changes
  • I have updated the CHANGELOG.md if applicable (handled at release time)

Additional Notes

Deliberate tradeoff (discussed and chosen): without the net-cost policy enabled, accepting sub-15% wins can be net-negative on prompt-cached sessions, because compressing a block invalidates the cached suffix (a one-time re-write, at 1.25× on Anthropic). If that shows up in practice, enable HEADROOM_NET_COST_POLICY=1 (the precise economics guard) or restore a floor by lowering the two min_ratio_* values.

… floor)

The acceptance gate rejected compressions saving <15% (min_ratio 0.85 at low
context pressure, 0.65 under pressure) — a crude proxy for "big enough to
justify busting the prefix cache." It dropped genuine token savings, including
lossless code/log folds that shrink <15% (the ratio_too_high rejects).

Set min_ratio to 1.0 at every pressure: accept ANY real shrink (ratio < 1.0);
any token saved is worth taking. The two guards that actually matter are
untouched — the reversibility gate keeps lossy-unmarked tool output verbatim
(accuracy), and the opt-in net-cost policy (HEADROOM_NET_COST_POLICY=1)
precisely accounts for cache-bust economics when enabled. Lower the values to
restore a savings floor.
@github-actions

github-actions Bot commented Jul 3, 2026

Copy link
Copy Markdown
Contributor

PR governance

This PR follows the template and is marked ready for human review.

@github-actions github-actions Bot added the status: ready for review Pull request body is complete and the author marked it ready for human review label Jul 3, 2026
@chopratejas
chopratejas merged commit 6c31db9 into main Jul 3, 2026
29 checks passed
chopratejas added a commit that referenced this pull request Jul 3, 2026
Unit mismatch: the apply() acceptance gate computes compression_ratio from
len(text.split()) (word count), but a lossless search/log fold cuts TOKENS by
collapsing a repeated path prefix into one heading — word count stays flat or
rises (the heading adds a word). So the gate saw ratio >= 1.0 and discarded
every free, recoverable win as ratio_too_high. Raising the floor to 1.0 (#1771)
didn't help; the word-ratio was already >= 1.0.

Measure lossless results (strategy_chain has a lossless_* entry) by REAL TOKEN
count via the tokenizer already in scope — not words, not bytes — so a fold is
accepted iff it genuinely reduces tokens. Gate + result cache use this ratio.
Lossy strategies are unchanged (word count tracks their savings) and the
reversibility gate is untouched (LOG/SEARCH/DIFF aren't lossy-unmarked). The
excluded and bash-search paths already bypass this gate; this fixes the main
strategy dispatch.

Regression test drives the full router.apply() path and asserts fewer TOKENS
(compress()/_apply_strategy_to_content bypass the gate, which is why prior unit
tests missed it).
chopratejas added a commit that referenced this pull request Jul 3, 2026
…ate (#1772)

## Description

Unit-mismatch bug in the compression acceptance gate. `router.apply()`
computes `compression_ratio` from `len(text.split())` (word count), but
a **lossless** search/log fold (`compact_lossless`) saves **bytes** by
collapsing a repeated path prefix into a single heading — word count
stays flat or even *rises* (the heading adds a word). So the gate saw
`ratio ≥ 1.0` and discarded every free, byte-recoverable win as
`ratio_too_high`. (Raising the floor to 1.0 in #1771 did **not** fix
this — the word-ratio was already ≥ 1.0.)

Measure lossless results (those whose `strategy_chain` carries a
`lossless_*` entry) by **byte ratio** at the gate and in the result
cache — the real saving. Lossy strategies are unchanged (word count
tracks their token savings), and the reversibility gate is untouched
(`LOG`/`SEARCH`/`DIFF` aren't in `LOSSY_UNMARKED_STRATEGIES`). The
excluded-tool and bash-search paths already bypass this gate via
`continue`; this fixes the **main strategy dispatch** (the lossless-mode
`LOG`/`SEARCH`/`DIFF` path).

Follow-up to #1771.

Closes #

## Type of Change

- [x] Bug fix (non-breaking change that fixes an issue)

## Changes Made

- At the `apply()` acceptance gate: compute `accept_ratio` = byte ratio
for lossless results (`strategy_chain` has `lossless_*`), else the
existing word ratio. Gate + result-cache entry now use `accept_ratio`.
- Added an end-to-end regression test that drives the full
`router.apply()` path.

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [ ] Manual testing performed

### Test Output

```text
tests/test_lossless_mode.py::test_router_apply_accepts_lossless_search_byte_measured PASSED
tests/test_content_router_tool_role_reversibility.py .......... (10 passed)
# broader (pre-move) sweep on the same change:
tests/test_lossless_mode.py / test_transforms/test_content_router.py /
test_lossless_excluded_compaction.py / test_bash_search_lossless_fold.py — 121 passed
ruff check headroom/transforms/content_router.py  -> All checks passed!
mypy headroom/transforms/content_router.py         -> Success: no issues found
```

## Real Behavior Proof

- Environment: local worktree, Python 3.12, `PYTHONPATH` pinned to the
branch.
- Exact command / steps: new regression test constructs a single-file
grep result, runs it through `ContentRouter(lossless=True).apply(...)`,
and asserts the tool output is byte-smaller and recovers exactly
(`search_unheading(out) == original`).
- Observed result: before this fix the fold was rejected (`out ==
original`, counted `ratio_too_high`); after, it's applied (`len(out) <
len(original)`, marker-free, byte-exact recovery). The test also asserts
the fold's word count is ≥ the original's, so the test is meaningless if
"fixed" by word count.
- Not tested: no live end-to-end proxy run; validated via the full
`apply()` path in unit tests.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable (handled at release
time)

## Additional Notes

Why prior tests missed it: `compress()` and `_apply_strategy_to_content`
return the folded result directly and never touch the `apply()`
acceptance gate, so the existing lossless-mode unit tests (which call
those) passed while the real proxy path silently discarded the fold. The
new test exercises `apply()` end-to-end.
@github-actions github-actions Bot mentioned this pull request Jul 9, 2026
chopratejas pushed a commit that referenced this pull request Jul 9, 2026
🤖 I have created a release *beep* *boop*
---


<details><summary>0.31.0</summary>

##
[0.31.0](v0.30.0...v0.31.0)
(2026-07-09)


### Features

* **cache:** provider-agnostic cache-mode delta + cc-agnostic prefix
comparison
([#1868](#1868))
([7c2f0ea](7c2f0ea))
* **ccr:** wire retrieve-tool interception into OpenAI Responses handler
([#1898](#1898))
([62cd307](62cd307))
* **compression:** add audit-safe mode with protected pattern matching
([#1899](#1899))
([bb112dd](bb112dd))
* **content-router:** accept any real compression (remove min-savings
floor)
([#1771](#1771))
([6c31db9](6c31db9))
* **content-router:** lossless-first dispatch, cross-turn dedup, and A7
lossy-after-fold
([#1818](#1818))
([60af15f](60af15f))
* **proxy:** add provider-only HTTP proxy
([#1807](#1807))
([ebe0a3b](ebe0a3b))
* **proxy:** add turn-hook extension point for buffered model turns
([#1891](#1891))
([ec950f7](ec950f7))


### Bug Fixes

* **build:** enable Intel macOS pip installs via ort-load-dynamic
([#1538](#1538))
([32ce99e](32ce99e))
* **cache:** avoid fallback session collisions
([#1827](#1827))
([0f606b6](0f606b6))
* **ccr:** make expired retrieve misses terminal
([#1781](#1781))
([9cbdba4](9cbdba4))
* **ccr:** preserve Anthropic re-stream shape
([#1854](#1854))
([f663894](f663894))
* **ccr:** preserve thinking blocks in buffered stream re-synthesis
([#1897](#1897))
([ede085c](ede085c))
* **cli/proxy:** preserve explicit HEADROOM_MIN_TOKENS=0 / MAX_ITEMS=0
([#1886](#1886))
([3a33af1](3a33af1))
* **code-compressor:** CJK-aware relevance-query symbol matching
([#1747](#1747))
([b38315c](b38315c))
* **codex:** discover updated Codex state stores
([#1889](#1889))
([9d42eba](9d42eba))
* **codex:** OpenCode Zen telemetry attribution
([#1648](#1648))
([f18c6bd](f18c6bd))
* **content-detector:** detect and compress space-separated JSON objects
([#1742](#1742))
([5194bdc](5194bdc))
* **content-router:** token-measure lossless folds at the acceptance
gate ([#1772](#1772))
([c5493ea](c5493ea))
* **copilot:** normalize subscription routing host
([#1836](#1836))
([afd9cbd](afd9cbd))
* **copilot:** route mixed-model requests per model
([#1785](#1785))
([5af5e22](5af5e22))
* **dashboard:** deduplicate repeated savings metrics
([#1804](#1804))
([88f935a](88f935a))
* **dashboard:** distinguish unavailable RTK from zero stats in Docker
([#1900](#1900))
([87f6e93](87f6e93))
* **dashboard:** distinguish unavailable RTK from zero stats in Docker
([#1901](#1901))
([361adcd](361adcd))
* **dashboard:** price proxy savings without litellm
([#1728](#1728))
([188e382](188e382))
* detect and clear stale ANTHROPIC_BASE_URL from crashed wrap sessions
([#1768](#1768))
([#1837](#1837))
([84509a4](84509a4))
* **docker:** persist headroom workspace in compose
([#1839](#1839))
([5e29c06](5e29c06))
* **docker:** report source build version
([#1862](#1862))
([3807488](3807488))
* **evals:** default unparseable judge scores below pass threshold
([#1892](#1892))
([42ebbc6](42ebbc6))
* **install:** pass sc.exe create as raw command line so binPath=
quoting survives
([#1654](#1654))
([#1702](#1702))
([d6e0710](d6e0710))
* **install:** persist --no-http2 override through install apply
([#1676](#1676))
([6fb5f3b](6fb5f3b))
* **mcp:** isolate ClaudeRegistrar CLI config env
([#1888](#1888))
([1c947b1](1c947b1))
* **mcp:** surface dead proxy state
([#1786](#1786))
([931eed8](931eed8))
* **memory:** resolve Trae cwd metadata from user reminders
([#1737](#1737))
([#1887](#1887))
([3e85eb1](3e85eb1))
* **opencode:** use local MCP config
([#1383](#1383))
([4bd3ddf](4bd3ddf))
* **proxy/openai:** thread savings-profile kwargs into chat completions
([#1606](#1606))
([7ff842d](7ff842d))
* **proxy/openai:** translate max_tokens -&gt; max_completion_tokens on
chat path
([#1774](#1774))
([285808b](285808b))
* **proxy:** bound Codex WS compression fallback latency
([#1802](#1802))
([d24a3f8](d24a3f8))
* **proxy:** bound HF tokenizer load and offload token counting off
event loop
([#1738](#1738))
([46d5d68](46d5d68))
* **proxy:** cancel retry backoff on shutdown
([#1834](#1834))
([da2d8dc](da2d8dc))
* **proxy:** compress Anthropic user text blocks when enabled
([#1875](#1875))
([e36439a](e36439a))
* **proxy:** freeze must forward cached (compressed) prefix
byte-identical — stop token-mode cache busting
([#1850](#1850))
([248ae0f](248ae0f))
* **proxy:** fsync savings dir after atomic rename
([#1764](#1764))
([7de2c1e](7de2c1e))
* **proxy:** keep cache_control bounded + stable so the freeze overlay
stops busting
([#1852](#1852))
([4820134](4820134))
* **proxy:** persist lifetime cache-read savings across restarts
([#1665](#1665))
([908997e](908997e))
* **proxy:** preserve streaming passthrough beta headers
([#1783](#1783))
([0f553a8](0f553a8))
* **proxy:** release _active_streams session lock on setup-phase errors
([#1864](#1864))
([2ccd831](2ccd831))
* **proxy:** retry HTTP/2 stream resets instead of 502ing
([#1645](#1645))
([2ce19c2](2ce19c2))
* **proxy:** retry passthrough on transient upstream connection close
([#1513](#1513))
([5d14080](5d14080))
* **proxy:** route Foundry Anthropic messages
([#1878](#1878))
([739f654](739f654))
* **proxy:** serve /favicon.ico locally instead of tunneling upstream
([#1787](#1787))
([#1847](#1847))
([3076e32](3076e32))
* **proxy:** stop rtk stat failures from corrupting session baseline
([#1693](#1693))
([681b9a8](681b9a8))
* **proxy:** strip 1m model suffix before upstream forwarding
([#1840](#1840))
([e22d745](e22d745))
* **proxy:** subtract cache write premiums from net savings
([#1800](#1800))
([53a465b](53a465b))
* **router:** honor MCP aliases in excluded tools
([#1822](#1822))
([#1863](#1863))
([140d6e4](140d6e4))
* **rtk:** link managed rtk onto PATH instead of mutating the hook
([#1698](#1698))
([140cb05](140cb05))
* **streaming:** preserve server_tool_use sse blocks
([#1826](#1826))
([4ac5493](4ac5493))
* **toin:** publish skip compression recommendations
([#1782](#1782))
([be51008](be51008))
* **transforms:** normalize diff compressor context
([#1801](#1801))
([838c523](838c523))
* **transforms:** pass through ragged tables instead of misaligning
columns
([#1713](#1713))
([c7665ca](c7665ca))
* use rtk native Cursor hook instead of injecting .cursorrules
([#756](#756))
([#1846](#1846))
([1573f1f](1573f1f))
* **wrap:** replace stale-proxy detection with Vite-style port fallback
([#1406](#1406))
([b4205c6](b4205c6))


### Performance Improvements

* **proxy:** cap compression workers to CPU count
([#1803](#1803))
([0a3851b](0a3851b))
* **savings:** batch tracker persistence off the request hot path
([#1817](#1817))
([451b9f0](451b9f0))


### Dependencies

* bump the cargo-minor-patch group across 1 directory with 7 updates
([#1909](#1909))
([45601d9](45601d9))
* bump the npm-minor-patch group across 4 directories with 18 updates
([#1907](#1907))
([8872bbc](8872bbc))
</details>

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
JerrettDavis pushed a commit to peterlodri-sec/headroom that referenced this pull request Jul 14, 2026
… floor) (headroomlabs-ai#1771)

## Description

The compression acceptance gate rejected any compression that saved less
than ~15% (`min_ratio` interpolated 0.85 at low context pressure → 0.65
under pressure). That floor was a crude proxy for "big enough to justify
busting the prefix cache," but it dropped genuine token savings —
notably lossless code/log folds that shrink <15% (the `ratio_too_high`
rejections).

This makes the gate accept **any real shrink** (`ratio < 1.0`): any
token saved is worth taking. The two guards that actually protect
correctness are untouched:
- **Reversibility gate** — lossy, unmarked tool output still stays
verbatim (accuracy; headroomlabs-ai#1307).
- **Net-cost policy** (`HEADROOM_NET_COST_POLICY=1`, opt-in) — precisely
accounts for the prefix-cache-bust economics (savings × expected-reads
vs one-time suffix re-write) when a session wants that protection.

Lowering the two values back to `0.85`/`0.65` restores the savings
floor.

Closes #

## Type of Change

- [x] New feature (non-breaking change that adds functionality)
- [x] Performance improvement

## Changes Made

- `ContentRouterConfig.min_ratio_relaxed`: `0.85 → 1.0`
- `ContentRouterConfig.min_ratio_aggressive`: `0.65 → 1.0`
- Gate now accepts any `compression_ratio < 1.0` at every context
pressure; reversibility + net-cost guards unchanged.

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check`)
- [x] Type checking passes (`mypy headroom`)
- [ ] New tests added for new functionality (existing gate-mechanism
tests cover it; they pass explicit `min_ratio` values and are
unaffected)
- [ ] Manual testing performed

### Test Output

```text
# content-router + compression suites (default-config paths):
tests/test_transforms/test_content_router.py ......... 139 passed
# broad compression sweep (-k compress/router/crush/kompress/lossless/ccr/savings/...):
1 failed, 1940 passed, 54 skipped in 181.25s
#   the 1 failure = test_lossless_mode::test_router_lossless_search_no_marker_and_recoverable
#   — a local fastembed-cache state-leak flake; passes in isolation (1 passed in 3.15s),
#   and is in lossless mode which bypasses this gate entirely.
ruff check headroom/transforms/content_router.py  -> All checks passed!
mypy headroom/transforms/content_router.py         -> Success: no issues found
```

## Real Behavior Proof

- Environment: local worktree, Python 3.12, `PYTHONPATH` pinned to the
branch checkout.
- Exact command / steps: ran the content-router acceptance-gate suites
and a compression-adjacent sweep against the branch; verified the flaky
test passes in isolation.
- Observed result: gate-mechanism tests (explicit `min_ratio`)
unaffected; no default-floor test regressed; blocks that previously
produced `ratio_too_high` at ratios in `[0.85, 1.0)` are now accepted.
- Not tested: no live end-to-end proxy run was performed for this
specific change; the behavioral effect (more `router:*` acceptances,
fewer `ratio_too_high`) is inferred from the gate logic + suite.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [ ] I have added tests that prove my fix is effective (existing gate
tests cover the mechanism)
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable (handled at release
time)

## Additional Notes

Deliberate tradeoff (discussed and chosen): without the net-cost policy
enabled, accepting sub-15% wins can be net-negative on prompt-cached
sessions, because compressing a block invalidates the cached suffix (a
one-time re-write, at 1.25× on Anthropic). If that shows up in practice,
enable `HEADROOM_NET_COST_POLICY=1` (the precise economics guard) or
restore a floor by lowering the two `min_ratio_*` values.
JerrettDavis pushed a commit to peterlodri-sec/headroom that referenced this pull request Jul 14, 2026
…ate (headroomlabs-ai#1772)

## Description

Unit-mismatch bug in the compression acceptance gate. `router.apply()`
computes `compression_ratio` from `len(text.split())` (word count), but
a **lossless** search/log fold (`compact_lossless`) saves **bytes** by
collapsing a repeated path prefix into a single heading — word count
stays flat or even *rises* (the heading adds a word). So the gate saw
`ratio ≥ 1.0` and discarded every free, byte-recoverable win as
`ratio_too_high`. (Raising the floor to 1.0 in headroomlabs-ai#1771 did **not** fix
this — the word-ratio was already ≥ 1.0.)

Measure lossless results (those whose `strategy_chain` carries a
`lossless_*` entry) by **byte ratio** at the gate and in the result
cache — the real saving. Lossy strategies are unchanged (word count
tracks their token savings), and the reversibility gate is untouched
(`LOG`/`SEARCH`/`DIFF` aren't in `LOSSY_UNMARKED_STRATEGIES`). The
excluded-tool and bash-search paths already bypass this gate via
`continue`; this fixes the **main strategy dispatch** (the lossless-mode
`LOG`/`SEARCH`/`DIFF` path).

Follow-up to headroomlabs-ai#1771.

Closes #

## Type of Change

- [x] Bug fix (non-breaking change that fixes an issue)

## Changes Made

- At the `apply()` acceptance gate: compute `accept_ratio` = byte ratio
for lossless results (`strategy_chain` has `lossless_*`), else the
existing word ratio. Gate + result-cache entry now use `accept_ratio`.
- Added an end-to-end regression test that drives the full
`router.apply()` path.

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [ ] Manual testing performed

### Test Output

```text
tests/test_lossless_mode.py::test_router_apply_accepts_lossless_search_byte_measured PASSED
tests/test_content_router_tool_role_reversibility.py .......... (10 passed)
# broader (pre-move) sweep on the same change:
tests/test_lossless_mode.py / test_transforms/test_content_router.py /
test_lossless_excluded_compaction.py / test_bash_search_lossless_fold.py — 121 passed
ruff check headroom/transforms/content_router.py  -> All checks passed!
mypy headroom/transforms/content_router.py         -> Success: no issues found
```

## Real Behavior Proof

- Environment: local worktree, Python 3.12, `PYTHONPATH` pinned to the
branch.
- Exact command / steps: new regression test constructs a single-file
grep result, runs it through `ContentRouter(lossless=True).apply(...)`,
and asserts the tool output is byte-smaller and recovers exactly
(`search_unheading(out) == original`).
- Observed result: before this fix the fold was rejected (`out ==
original`, counted `ratio_too_high`); after, it's applied (`len(out) <
len(original)`, marker-free, byte-exact recovery). The test also asserts
the fold's word count is ≥ the original's, so the test is meaningless if
"fixed" by word count.
- Not tested: no live end-to-end proxy run; validated via the full
`apply()` path in unit tests.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable (handled at release
time)

## Additional Notes

Why prior tests missed it: `compress()` and `_apply_strategy_to_content`
return the folded result directly and never touch the `apply()`
acceptance gate, so the existing lossless-mode unit tests (which call
those) passed while the real proxy path silently discarded the fold. The
new test exercises `apply()` end-to-end.
JerrettDavis pushed a commit to peterlodri-sec/headroom that referenced this pull request Jul 14, 2026
🤖 I have created a release *beep* *boop*
---


<details><summary>0.31.0</summary>

##
[0.31.0](headroomlabs-ai/headroom@v0.30.0...v0.31.0)
(2026-07-09)


### Features

* **cache:** provider-agnostic cache-mode delta + cc-agnostic prefix
comparison
([headroomlabs-ai#1868](headroomlabs-ai#1868))
([7c2f0ea](headroomlabs-ai@7c2f0ea))
* **ccr:** wire retrieve-tool interception into OpenAI Responses handler
([headroomlabs-ai#1898](headroomlabs-ai#1898))
([62cd307](headroomlabs-ai@62cd307))
* **compression:** add audit-safe mode with protected pattern matching
([headroomlabs-ai#1899](headroomlabs-ai#1899))
([bb112dd](headroomlabs-ai@bb112dd))
* **content-router:** accept any real compression (remove min-savings
floor)
([headroomlabs-ai#1771](headroomlabs-ai#1771))
([6c31db9](headroomlabs-ai@6c31db9))
* **content-router:** lossless-first dispatch, cross-turn dedup, and A7
lossy-after-fold
([headroomlabs-ai#1818](headroomlabs-ai#1818))
([60af15f](headroomlabs-ai@60af15f))
* **proxy:** add provider-only HTTP proxy
([headroomlabs-ai#1807](headroomlabs-ai#1807))
([ebe0a3b](headroomlabs-ai@ebe0a3b))
* **proxy:** add turn-hook extension point for buffered model turns
([headroomlabs-ai#1891](headroomlabs-ai#1891))
([ec950f7](headroomlabs-ai@ec950f7))


### Bug Fixes

* **build:** enable Intel macOS pip installs via ort-load-dynamic
([headroomlabs-ai#1538](headroomlabs-ai#1538))
([32ce99e](headroomlabs-ai@32ce99e))
* **cache:** avoid fallback session collisions
([headroomlabs-ai#1827](headroomlabs-ai#1827))
([0f606b6](headroomlabs-ai@0f606b6))
* **ccr:** make expired retrieve misses terminal
([headroomlabs-ai#1781](headroomlabs-ai#1781))
([9cbdba4](headroomlabs-ai@9cbdba4))
* **ccr:** preserve Anthropic re-stream shape
([headroomlabs-ai#1854](headroomlabs-ai#1854))
([f663894](headroomlabs-ai@f663894))
* **ccr:** preserve thinking blocks in buffered stream re-synthesis
([headroomlabs-ai#1897](headroomlabs-ai#1897))
([ede085c](headroomlabs-ai@ede085c))
* **cli/proxy:** preserve explicit HEADROOM_MIN_TOKENS=0 / MAX_ITEMS=0
([headroomlabs-ai#1886](headroomlabs-ai#1886))
([3a33af1](headroomlabs-ai@3a33af1))
* **code-compressor:** CJK-aware relevance-query symbol matching
([headroomlabs-ai#1747](headroomlabs-ai#1747))
([b38315c](headroomlabs-ai@b38315c))
* **codex:** discover updated Codex state stores
([headroomlabs-ai#1889](headroomlabs-ai#1889))
([9d42eba](headroomlabs-ai@9d42eba))
* **codex:** OpenCode Zen telemetry attribution
([headroomlabs-ai#1648](headroomlabs-ai#1648))
([f18c6bd](headroomlabs-ai@f18c6bd))
* **content-detector:** detect and compress space-separated JSON objects
([headroomlabs-ai#1742](headroomlabs-ai#1742))
([5194bdc](headroomlabs-ai@5194bdc))
* **content-router:** token-measure lossless folds at the acceptance
gate ([headroomlabs-ai#1772](headroomlabs-ai#1772))
([c5493ea](headroomlabs-ai@c5493ea))
* **copilot:** normalize subscription routing host
([headroomlabs-ai#1836](headroomlabs-ai#1836))
([afd9cbd](headroomlabs-ai@afd9cbd))
* **copilot:** route mixed-model requests per model
([headroomlabs-ai#1785](headroomlabs-ai#1785))
([5af5e22](headroomlabs-ai@5af5e22))
* **dashboard:** deduplicate repeated savings metrics
([headroomlabs-ai#1804](headroomlabs-ai#1804))
([88f935a](headroomlabs-ai@88f935a))
* **dashboard:** distinguish unavailable RTK from zero stats in Docker
([headroomlabs-ai#1900](headroomlabs-ai#1900))
([87f6e93](headroomlabs-ai@87f6e93))
* **dashboard:** distinguish unavailable RTK from zero stats in Docker
([headroomlabs-ai#1901](headroomlabs-ai#1901))
([361adcd](headroomlabs-ai@361adcd))
* **dashboard:** price proxy savings without litellm
([headroomlabs-ai#1728](headroomlabs-ai#1728))
([188e382](headroomlabs-ai@188e382))
* detect and clear stale ANTHROPIC_BASE_URL from crashed wrap sessions
([headroomlabs-ai#1768](headroomlabs-ai#1768))
([headroomlabs-ai#1837](headroomlabs-ai#1837))
([84509a4](headroomlabs-ai@84509a4))
* **docker:** persist headroom workspace in compose
([headroomlabs-ai#1839](headroomlabs-ai#1839))
([5e29c06](headroomlabs-ai@5e29c06))
* **docker:** report source build version
([headroomlabs-ai#1862](headroomlabs-ai#1862))
([3807488](headroomlabs-ai@3807488))
* **evals:** default unparseable judge scores below pass threshold
([headroomlabs-ai#1892](headroomlabs-ai#1892))
([42ebbc6](headroomlabs-ai@42ebbc6))
* **install:** pass sc.exe create as raw command line so binPath=
quoting survives
([headroomlabs-ai#1654](headroomlabs-ai#1654))
([headroomlabs-ai#1702](headroomlabs-ai#1702))
([d6e0710](headroomlabs-ai@d6e0710))
* **install:** persist --no-http2 override through install apply
([headroomlabs-ai#1676](headroomlabs-ai#1676))
([6fb5f3b](headroomlabs-ai@6fb5f3b))
* **mcp:** isolate ClaudeRegistrar CLI config env
([headroomlabs-ai#1888](headroomlabs-ai#1888))
([1c947b1](headroomlabs-ai@1c947b1))
* **mcp:** surface dead proxy state
([headroomlabs-ai#1786](headroomlabs-ai#1786))
([931eed8](headroomlabs-ai@931eed8))
* **memory:** resolve Trae cwd metadata from user reminders
([headroomlabs-ai#1737](headroomlabs-ai#1737))
([headroomlabs-ai#1887](headroomlabs-ai#1887))
([3e85eb1](headroomlabs-ai@3e85eb1))
* **opencode:** use local MCP config
([headroomlabs-ai#1383](headroomlabs-ai#1383))
([4bd3ddf](headroomlabs-ai@4bd3ddf))
* **proxy/openai:** thread savings-profile kwargs into chat completions
([headroomlabs-ai#1606](headroomlabs-ai#1606))
([7ff842d](headroomlabs-ai@7ff842d))
* **proxy/openai:** translate max_tokens -&gt; max_completion_tokens on
chat path
([headroomlabs-ai#1774](headroomlabs-ai#1774))
([285808b](headroomlabs-ai@285808b))
* **proxy:** bound Codex WS compression fallback latency
([headroomlabs-ai#1802](headroomlabs-ai#1802))
([d24a3f8](headroomlabs-ai@d24a3f8))
* **proxy:** bound HF tokenizer load and offload token counting off
event loop
([headroomlabs-ai#1738](headroomlabs-ai#1738))
([46d5d68](headroomlabs-ai@46d5d68))
* **proxy:** cancel retry backoff on shutdown
([headroomlabs-ai#1834](headroomlabs-ai#1834))
([da2d8dc](headroomlabs-ai@da2d8dc))
* **proxy:** compress Anthropic user text blocks when enabled
([headroomlabs-ai#1875](headroomlabs-ai#1875))
([e36439a](headroomlabs-ai@e36439a))
* **proxy:** freeze must forward cached (compressed) prefix
byte-identical — stop token-mode cache busting
([headroomlabs-ai#1850](headroomlabs-ai#1850))
([248ae0f](headroomlabs-ai@248ae0f))
* **proxy:** fsync savings dir after atomic rename
([headroomlabs-ai#1764](headroomlabs-ai#1764))
([7de2c1e](headroomlabs-ai@7de2c1e))
* **proxy:** keep cache_control bounded + stable so the freeze overlay
stops busting
([headroomlabs-ai#1852](headroomlabs-ai#1852))
([4820134](headroomlabs-ai@4820134))
* **proxy:** persist lifetime cache-read savings across restarts
([headroomlabs-ai#1665](headroomlabs-ai#1665))
([908997e](headroomlabs-ai@908997e))
* **proxy:** preserve streaming passthrough beta headers
([headroomlabs-ai#1783](headroomlabs-ai#1783))
([0f553a8](headroomlabs-ai@0f553a8))
* **proxy:** release _active_streams session lock on setup-phase errors
([headroomlabs-ai#1864](headroomlabs-ai#1864))
([2ccd831](headroomlabs-ai@2ccd831))
* **proxy:** retry HTTP/2 stream resets instead of 502ing
([headroomlabs-ai#1645](headroomlabs-ai#1645))
([2ce19c2](headroomlabs-ai@2ce19c2))
* **proxy:** retry passthrough on transient upstream connection close
([headroomlabs-ai#1513](headroomlabs-ai#1513))
([5d14080](headroomlabs-ai@5d14080))
* **proxy:** route Foundry Anthropic messages
([headroomlabs-ai#1878](headroomlabs-ai#1878))
([739f654](headroomlabs-ai@739f654))
* **proxy:** serve /favicon.ico locally instead of tunneling upstream
([headroomlabs-ai#1787](headroomlabs-ai#1787))
([headroomlabs-ai#1847](headroomlabs-ai#1847))
([3076e32](headroomlabs-ai@3076e32))
* **proxy:** stop rtk stat failures from corrupting session baseline
([headroomlabs-ai#1693](headroomlabs-ai#1693))
([681b9a8](headroomlabs-ai@681b9a8))
* **proxy:** strip 1m model suffix before upstream forwarding
([headroomlabs-ai#1840](headroomlabs-ai#1840))
([e22d745](headroomlabs-ai@e22d745))
* **proxy:** subtract cache write premiums from net savings
([headroomlabs-ai#1800](headroomlabs-ai#1800))
([53a465b](headroomlabs-ai@53a465b))
* **router:** honor MCP aliases in excluded tools
([headroomlabs-ai#1822](headroomlabs-ai#1822))
([headroomlabs-ai#1863](headroomlabs-ai#1863))
([140d6e4](headroomlabs-ai@140d6e4))
* **rtk:** link managed rtk onto PATH instead of mutating the hook
([headroomlabs-ai#1698](headroomlabs-ai#1698))
([140cb05](headroomlabs-ai@140cb05))
* **streaming:** preserve server_tool_use sse blocks
([headroomlabs-ai#1826](headroomlabs-ai#1826))
([4ac5493](headroomlabs-ai@4ac5493))
* **toin:** publish skip compression recommendations
([headroomlabs-ai#1782](headroomlabs-ai#1782))
([be51008](headroomlabs-ai@be51008))
* **transforms:** normalize diff compressor context
([headroomlabs-ai#1801](headroomlabs-ai#1801))
([838c523](headroomlabs-ai@838c523))
* **transforms:** pass through ragged tables instead of misaligning
columns
([headroomlabs-ai#1713](headroomlabs-ai#1713))
([c7665ca](headroomlabs-ai@c7665ca))
* use rtk native Cursor hook instead of injecting .cursorrules
([headroomlabs-ai#756](headroomlabs-ai#756))
([headroomlabs-ai#1846](headroomlabs-ai#1846))
([1573f1f](headroomlabs-ai@1573f1f))
* **wrap:** replace stale-proxy detection with Vite-style port fallback
([headroomlabs-ai#1406](headroomlabs-ai#1406))
([b4205c6](headroomlabs-ai@b4205c6))


### Performance Improvements

* **proxy:** cap compression workers to CPU count
([headroomlabs-ai#1803](headroomlabs-ai#1803))
([0a3851b](headroomlabs-ai@0a3851b))
* **savings:** batch tracker persistence off the request hot path
([headroomlabs-ai#1817](headroomlabs-ai#1817))
([451b9f0](headroomlabs-ai@451b9f0))


### Dependencies

* bump the cargo-minor-patch group across 1 directory with 7 updates
([headroomlabs-ai#1909](headroomlabs-ai#1909))
([45601d9](headroomlabs-ai@45601d9))
* bump the npm-minor-patch group across 4 directories with 18 updates
([headroomlabs-ai#1907](headroomlabs-ai#1907))
([8872bbc](headroomlabs-ai@8872bbc))
</details>

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

status: ready for review Pull request body is complete and the author marked it ready for human review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant