feat(content-router): accept any real compression (remove min-savings floor) - #1771
Merged
Merged
Conversation
… floor) The acceptance gate rejected compressions saving <15% (min_ratio 0.85 at low context pressure, 0.65 under pressure) — a crude proxy for "big enough to justify busting the prefix cache." It dropped genuine token savings, including lossless code/log folds that shrink <15% (the ratio_too_high rejects). Set min_ratio to 1.0 at every pressure: accept ANY real shrink (ratio < 1.0); any token saved is worth taking. The two guards that actually matter are untouched — the reversibility gate keeps lossy-unmarked tool output verbatim (accuracy), and the opt-in net-cost policy (HEADROOM_NET_COST_POLICY=1) precisely accounts for cache-bust economics when enabled. Lower the values to restore a savings floor.
Contributor
PR governanceThis PR follows the template and is marked ready for human review. |
13 of 16 tasks
chopratejas
added a commit
that referenced
this pull request
Jul 3, 2026
Unit mismatch: the apply() acceptance gate computes compression_ratio from len(text.split()) (word count), but a lossless search/log fold cuts TOKENS by collapsing a repeated path prefix into one heading — word count stays flat or rises (the heading adds a word). So the gate saw ratio >= 1.0 and discarded every free, recoverable win as ratio_too_high. Raising the floor to 1.0 (#1771) didn't help; the word-ratio was already >= 1.0. Measure lossless results (strategy_chain has a lossless_* entry) by REAL TOKEN count via the tokenizer already in scope — not words, not bytes — so a fold is accepted iff it genuinely reduces tokens. Gate + result cache use this ratio. Lossy strategies are unchanged (word count tracks their savings) and the reversibility gate is untouched (LOG/SEARCH/DIFF aren't lossy-unmarked). The excluded and bash-search paths already bypass this gate; this fixes the main strategy dispatch. Regression test drives the full router.apply() path and asserts fewer TOKENS (compress()/_apply_strategy_to_content bypass the gate, which is why prior unit tests missed it).
chopratejas
added a commit
that referenced
this pull request
Jul 3, 2026
…ate (#1772) ## Description Unit-mismatch bug in the compression acceptance gate. `router.apply()` computes `compression_ratio` from `len(text.split())` (word count), but a **lossless** search/log fold (`compact_lossless`) saves **bytes** by collapsing a repeated path prefix into a single heading — word count stays flat or even *rises* (the heading adds a word). So the gate saw `ratio ≥ 1.0` and discarded every free, byte-recoverable win as `ratio_too_high`. (Raising the floor to 1.0 in #1771 did **not** fix this — the word-ratio was already ≥ 1.0.) Measure lossless results (those whose `strategy_chain` carries a `lossless_*` entry) by **byte ratio** at the gate and in the result cache — the real saving. Lossy strategies are unchanged (word count tracks their token savings), and the reversibility gate is untouched (`LOG`/`SEARCH`/`DIFF` aren't in `LOSSY_UNMARKED_STRATEGIES`). The excluded-tool and bash-search paths already bypass this gate via `continue`; this fixes the **main strategy dispatch** (the lossless-mode `LOG`/`SEARCH`/`DIFF` path). Follow-up to #1771. Closes # ## Type of Change - [x] Bug fix (non-breaking change that fixes an issue) ## Changes Made - At the `apply()` acceptance gate: compute `accept_ratio` = byte ratio for lossless results (`strategy_chain` has `lossless_*`), else the existing word ratio. Gate + result-cache entry now use `accept_ratio`. - Added an end-to-end regression test that drives the full `router.apply()` path. ## Testing - [x] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added for new functionality - [ ] Manual testing performed ### Test Output ```text tests/test_lossless_mode.py::test_router_apply_accepts_lossless_search_byte_measured PASSED tests/test_content_router_tool_role_reversibility.py .......... (10 passed) # broader (pre-move) sweep on the same change: tests/test_lossless_mode.py / test_transforms/test_content_router.py / test_lossless_excluded_compaction.py / test_bash_search_lossless_fold.py — 121 passed ruff check headroom/transforms/content_router.py -> All checks passed! mypy headroom/transforms/content_router.py -> Success: no issues found ``` ## Real Behavior Proof - Environment: local worktree, Python 3.12, `PYTHONPATH` pinned to the branch. - Exact command / steps: new regression test constructs a single-file grep result, runs it through `ContentRouter(lossless=True).apply(...)`, and asserts the tool output is byte-smaller and recovers exactly (`search_unheading(out) == original`). - Observed result: before this fix the fold was rejected (`out == original`, counted `ratio_too_high`); after, it's applied (`len(out) < len(original)`, marker-free, byte-exact recovery). The test also asserts the fold's word count is ≥ the original's, so the test is meaningless if "fixed" by word count. - Not tested: no live end-to-end proxy run; validated via the full `apply()` path in unit tests. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [ ] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective - [x] New and existing unit tests pass locally with my changes - [ ] I have updated the CHANGELOG.md if applicable (handled at release time) ## Additional Notes Why prior tests missed it: `compress()` and `_apply_strategy_to_content` return the folded result directly and never touch the `apply()` acceptance gate, so the existing lossless-mode unit tests (which call those) passed while the real proxy path silently discarded the fold. The new test exercises `apply()` end-to-end.
Merged
chopratejas
pushed a commit
that referenced
this pull request
Jul 9, 2026
🤖 I have created a release *beep* *boop* --- <details><summary>0.31.0</summary> ## [0.31.0](v0.30.0...v0.31.0) (2026-07-09) ### Features * **cache:** provider-agnostic cache-mode delta + cc-agnostic prefix comparison ([#1868](#1868)) ([7c2f0ea](7c2f0ea)) * **ccr:** wire retrieve-tool interception into OpenAI Responses handler ([#1898](#1898)) ([62cd307](62cd307)) * **compression:** add audit-safe mode with protected pattern matching ([#1899](#1899)) ([bb112dd](bb112dd)) * **content-router:** accept any real compression (remove min-savings floor) ([#1771](#1771)) ([6c31db9](6c31db9)) * **content-router:** lossless-first dispatch, cross-turn dedup, and A7 lossy-after-fold ([#1818](#1818)) ([60af15f](60af15f)) * **proxy:** add provider-only HTTP proxy ([#1807](#1807)) ([ebe0a3b](ebe0a3b)) * **proxy:** add turn-hook extension point for buffered model turns ([#1891](#1891)) ([ec950f7](ec950f7)) ### Bug Fixes * **build:** enable Intel macOS pip installs via ort-load-dynamic ([#1538](#1538)) ([32ce99e](32ce99e)) * **cache:** avoid fallback session collisions ([#1827](#1827)) ([0f606b6](0f606b6)) * **ccr:** make expired retrieve misses terminal ([#1781](#1781)) ([9cbdba4](9cbdba4)) * **ccr:** preserve Anthropic re-stream shape ([#1854](#1854)) ([f663894](f663894)) * **ccr:** preserve thinking blocks in buffered stream re-synthesis ([#1897](#1897)) ([ede085c](ede085c)) * **cli/proxy:** preserve explicit HEADROOM_MIN_TOKENS=0 / MAX_ITEMS=0 ([#1886](#1886)) ([3a33af1](3a33af1)) * **code-compressor:** CJK-aware relevance-query symbol matching ([#1747](#1747)) ([b38315c](b38315c)) * **codex:** discover updated Codex state stores ([#1889](#1889)) ([9d42eba](9d42eba)) * **codex:** OpenCode Zen telemetry attribution ([#1648](#1648)) ([f18c6bd](f18c6bd)) * **content-detector:** detect and compress space-separated JSON objects ([#1742](#1742)) ([5194bdc](5194bdc)) * **content-router:** token-measure lossless folds at the acceptance gate ([#1772](#1772)) ([c5493ea](c5493ea)) * **copilot:** normalize subscription routing host ([#1836](#1836)) ([afd9cbd](afd9cbd)) * **copilot:** route mixed-model requests per model ([#1785](#1785)) ([5af5e22](5af5e22)) * **dashboard:** deduplicate repeated savings metrics ([#1804](#1804)) ([88f935a](88f935a)) * **dashboard:** distinguish unavailable RTK from zero stats in Docker ([#1900](#1900)) ([87f6e93](87f6e93)) * **dashboard:** distinguish unavailable RTK from zero stats in Docker ([#1901](#1901)) ([361adcd](361adcd)) * **dashboard:** price proxy savings without litellm ([#1728](#1728)) ([188e382](188e382)) * detect and clear stale ANTHROPIC_BASE_URL from crashed wrap sessions ([#1768](#1768)) ([#1837](#1837)) ([84509a4](84509a4)) * **docker:** persist headroom workspace in compose ([#1839](#1839)) ([5e29c06](5e29c06)) * **docker:** report source build version ([#1862](#1862)) ([3807488](3807488)) * **evals:** default unparseable judge scores below pass threshold ([#1892](#1892)) ([42ebbc6](42ebbc6)) * **install:** pass sc.exe create as raw command line so binPath= quoting survives ([#1654](#1654)) ([#1702](#1702)) ([d6e0710](d6e0710)) * **install:** persist --no-http2 override through install apply ([#1676](#1676)) ([6fb5f3b](6fb5f3b)) * **mcp:** isolate ClaudeRegistrar CLI config env ([#1888](#1888)) ([1c947b1](1c947b1)) * **mcp:** surface dead proxy state ([#1786](#1786)) ([931eed8](931eed8)) * **memory:** resolve Trae cwd metadata from user reminders ([#1737](#1737)) ([#1887](#1887)) ([3e85eb1](3e85eb1)) * **opencode:** use local MCP config ([#1383](#1383)) ([4bd3ddf](4bd3ddf)) * **proxy/openai:** thread savings-profile kwargs into chat completions ([#1606](#1606)) ([7ff842d](7ff842d)) * **proxy/openai:** translate max_tokens -> max_completion_tokens on chat path ([#1774](#1774)) ([285808b](285808b)) * **proxy:** bound Codex WS compression fallback latency ([#1802](#1802)) ([d24a3f8](d24a3f8)) * **proxy:** bound HF tokenizer load and offload token counting off event loop ([#1738](#1738)) ([46d5d68](46d5d68)) * **proxy:** cancel retry backoff on shutdown ([#1834](#1834)) ([da2d8dc](da2d8dc)) * **proxy:** compress Anthropic user text blocks when enabled ([#1875](#1875)) ([e36439a](e36439a)) * **proxy:** freeze must forward cached (compressed) prefix byte-identical — stop token-mode cache busting ([#1850](#1850)) ([248ae0f](248ae0f)) * **proxy:** fsync savings dir after atomic rename ([#1764](#1764)) ([7de2c1e](7de2c1e)) * **proxy:** keep cache_control bounded + stable so the freeze overlay stops busting ([#1852](#1852)) ([4820134](4820134)) * **proxy:** persist lifetime cache-read savings across restarts ([#1665](#1665)) ([908997e](908997e)) * **proxy:** preserve streaming passthrough beta headers ([#1783](#1783)) ([0f553a8](0f553a8)) * **proxy:** release _active_streams session lock on setup-phase errors ([#1864](#1864)) ([2ccd831](2ccd831)) * **proxy:** retry HTTP/2 stream resets instead of 502ing ([#1645](#1645)) ([2ce19c2](2ce19c2)) * **proxy:** retry passthrough on transient upstream connection close ([#1513](#1513)) ([5d14080](5d14080)) * **proxy:** route Foundry Anthropic messages ([#1878](#1878)) ([739f654](739f654)) * **proxy:** serve /favicon.ico locally instead of tunneling upstream ([#1787](#1787)) ([#1847](#1847)) ([3076e32](3076e32)) * **proxy:** stop rtk stat failures from corrupting session baseline ([#1693](#1693)) ([681b9a8](681b9a8)) * **proxy:** strip 1m model suffix before upstream forwarding ([#1840](#1840)) ([e22d745](e22d745)) * **proxy:** subtract cache write premiums from net savings ([#1800](#1800)) ([53a465b](53a465b)) * **router:** honor MCP aliases in excluded tools ([#1822](#1822)) ([#1863](#1863)) ([140d6e4](140d6e4)) * **rtk:** link managed rtk onto PATH instead of mutating the hook ([#1698](#1698)) ([140cb05](140cb05)) * **streaming:** preserve server_tool_use sse blocks ([#1826](#1826)) ([4ac5493](4ac5493)) * **toin:** publish skip compression recommendations ([#1782](#1782)) ([be51008](be51008)) * **transforms:** normalize diff compressor context ([#1801](#1801)) ([838c523](838c523)) * **transforms:** pass through ragged tables instead of misaligning columns ([#1713](#1713)) ([c7665ca](c7665ca)) * use rtk native Cursor hook instead of injecting .cursorrules ([#756](#756)) ([#1846](#1846)) ([1573f1f](1573f1f)) * **wrap:** replace stale-proxy detection with Vite-style port fallback ([#1406](#1406)) ([b4205c6](b4205c6)) ### Performance Improvements * **proxy:** cap compression workers to CPU count ([#1803](#1803)) ([0a3851b](0a3851b)) * **savings:** batch tracker persistence off the request hot path ([#1817](#1817)) ([451b9f0](451b9f0)) ### Dependencies * bump the cargo-minor-patch group across 1 directory with 7 updates ([#1909](#1909)) ([45601d9](45601d9)) * bump the npm-minor-patch group across 4 directories with 18 updates ([#1907](#1907)) ([8872bbc](8872bbc)) </details> --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
JerrettDavis
pushed a commit
to peterlodri-sec/headroom
that referenced
this pull request
Jul 14, 2026
… floor) (headroomlabs-ai#1771) ## Description The compression acceptance gate rejected any compression that saved less than ~15% (`min_ratio` interpolated 0.85 at low context pressure → 0.65 under pressure). That floor was a crude proxy for "big enough to justify busting the prefix cache," but it dropped genuine token savings — notably lossless code/log folds that shrink <15% (the `ratio_too_high` rejections). This makes the gate accept **any real shrink** (`ratio < 1.0`): any token saved is worth taking. The two guards that actually protect correctness are untouched: - **Reversibility gate** — lossy, unmarked tool output still stays verbatim (accuracy; headroomlabs-ai#1307). - **Net-cost policy** (`HEADROOM_NET_COST_POLICY=1`, opt-in) — precisely accounts for the prefix-cache-bust economics (savings × expected-reads vs one-time suffix re-write) when a session wants that protection. Lowering the two values back to `0.85`/`0.65` restores the savings floor. Closes # ## Type of Change - [x] New feature (non-breaking change that adds functionality) - [x] Performance improvement ## Changes Made - `ContentRouterConfig.min_ratio_relaxed`: `0.85 → 1.0` - `ContentRouterConfig.min_ratio_aggressive`: `0.65 → 1.0` - Gate now accepts any `compression_ratio < 1.0` at every context pressure; reversibility + net-cost guards unchanged. ## Testing - [x] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check`) - [x] Type checking passes (`mypy headroom`) - [ ] New tests added for new functionality (existing gate-mechanism tests cover it; they pass explicit `min_ratio` values and are unaffected) - [ ] Manual testing performed ### Test Output ```text # content-router + compression suites (default-config paths): tests/test_transforms/test_content_router.py ......... 139 passed # broad compression sweep (-k compress/router/crush/kompress/lossless/ccr/savings/...): 1 failed, 1940 passed, 54 skipped in 181.25s # the 1 failure = test_lossless_mode::test_router_lossless_search_no_marker_and_recoverable # — a local fastembed-cache state-leak flake; passes in isolation (1 passed in 3.15s), # and is in lossless mode which bypasses this gate entirely. ruff check headroom/transforms/content_router.py -> All checks passed! mypy headroom/transforms/content_router.py -> Success: no issues found ``` ## Real Behavior Proof - Environment: local worktree, Python 3.12, `PYTHONPATH` pinned to the branch checkout. - Exact command / steps: ran the content-router acceptance-gate suites and a compression-adjacent sweep against the branch; verified the flaky test passes in isolation. - Observed result: gate-mechanism tests (explicit `min_ratio`) unaffected; no default-floor test regressed; blocks that previously produced `ratio_too_high` at ratios in `[0.85, 1.0)` are now accepted. - Not tested: no live end-to-end proxy run was performed for this specific change; the behavioral effect (more `router:*` acceptances, fewer `ratio_too_high`) is inferred from the gate logic + suite. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [ ] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [ ] I have added tests that prove my fix is effective (existing gate tests cover the mechanism) - [x] New and existing unit tests pass locally with my changes - [ ] I have updated the CHANGELOG.md if applicable (handled at release time) ## Additional Notes Deliberate tradeoff (discussed and chosen): without the net-cost policy enabled, accepting sub-15% wins can be net-negative on prompt-cached sessions, because compressing a block invalidates the cached suffix (a one-time re-write, at 1.25× on Anthropic). If that shows up in practice, enable `HEADROOM_NET_COST_POLICY=1` (the precise economics guard) or restore a floor by lowering the two `min_ratio_*` values.
JerrettDavis
pushed a commit
to peterlodri-sec/headroom
that referenced
this pull request
Jul 14, 2026
…ate (headroomlabs-ai#1772) ## Description Unit-mismatch bug in the compression acceptance gate. `router.apply()` computes `compression_ratio` from `len(text.split())` (word count), but a **lossless** search/log fold (`compact_lossless`) saves **bytes** by collapsing a repeated path prefix into a single heading — word count stays flat or even *rises* (the heading adds a word). So the gate saw `ratio ≥ 1.0` and discarded every free, byte-recoverable win as `ratio_too_high`. (Raising the floor to 1.0 in headroomlabs-ai#1771 did **not** fix this — the word-ratio was already ≥ 1.0.) Measure lossless results (those whose `strategy_chain` carries a `lossless_*` entry) by **byte ratio** at the gate and in the result cache — the real saving. Lossy strategies are unchanged (word count tracks their token savings), and the reversibility gate is untouched (`LOG`/`SEARCH`/`DIFF` aren't in `LOSSY_UNMARKED_STRATEGIES`). The excluded-tool and bash-search paths already bypass this gate via `continue`; this fixes the **main strategy dispatch** (the lossless-mode `LOG`/`SEARCH`/`DIFF` path). Follow-up to headroomlabs-ai#1771. Closes # ## Type of Change - [x] Bug fix (non-breaking change that fixes an issue) ## Changes Made - At the `apply()` acceptance gate: compute `accept_ratio` = byte ratio for lossless results (`strategy_chain` has `lossless_*`), else the existing word ratio. Gate + result-cache entry now use `accept_ratio`. - Added an end-to-end regression test that drives the full `router.apply()` path. ## Testing - [x] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added for new functionality - [ ] Manual testing performed ### Test Output ```text tests/test_lossless_mode.py::test_router_apply_accepts_lossless_search_byte_measured PASSED tests/test_content_router_tool_role_reversibility.py .......... (10 passed) # broader (pre-move) sweep on the same change: tests/test_lossless_mode.py / test_transforms/test_content_router.py / test_lossless_excluded_compaction.py / test_bash_search_lossless_fold.py — 121 passed ruff check headroom/transforms/content_router.py -> All checks passed! mypy headroom/transforms/content_router.py -> Success: no issues found ``` ## Real Behavior Proof - Environment: local worktree, Python 3.12, `PYTHONPATH` pinned to the branch. - Exact command / steps: new regression test constructs a single-file grep result, runs it through `ContentRouter(lossless=True).apply(...)`, and asserts the tool output is byte-smaller and recovers exactly (`search_unheading(out) == original`). - Observed result: before this fix the fold was rejected (`out == original`, counted `ratio_too_high`); after, it's applied (`len(out) < len(original)`, marker-free, byte-exact recovery). The test also asserts the fold's word count is ≥ the original's, so the test is meaningless if "fixed" by word count. - Not tested: no live end-to-end proxy run; validated via the full `apply()` path in unit tests. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [ ] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective - [x] New and existing unit tests pass locally with my changes - [ ] I have updated the CHANGELOG.md if applicable (handled at release time) ## Additional Notes Why prior tests missed it: `compress()` and `_apply_strategy_to_content` return the folded result directly and never touch the `apply()` acceptance gate, so the existing lossless-mode unit tests (which call those) passed while the real proxy path silently discarded the fold. The new test exercises `apply()` end-to-end.
JerrettDavis
pushed a commit
to peterlodri-sec/headroom
that referenced
this pull request
Jul 14, 2026
🤖 I have created a release *beep* *boop* --- <details><summary>0.31.0</summary> ## [0.31.0](headroomlabs-ai/headroom@v0.30.0...v0.31.0) (2026-07-09) ### Features * **cache:** provider-agnostic cache-mode delta + cc-agnostic prefix comparison ([headroomlabs-ai#1868](headroomlabs-ai#1868)) ([7c2f0ea](headroomlabs-ai@7c2f0ea)) * **ccr:** wire retrieve-tool interception into OpenAI Responses handler ([headroomlabs-ai#1898](headroomlabs-ai#1898)) ([62cd307](headroomlabs-ai@62cd307)) * **compression:** add audit-safe mode with protected pattern matching ([headroomlabs-ai#1899](headroomlabs-ai#1899)) ([bb112dd](headroomlabs-ai@bb112dd)) * **content-router:** accept any real compression (remove min-savings floor) ([headroomlabs-ai#1771](headroomlabs-ai#1771)) ([6c31db9](headroomlabs-ai@6c31db9)) * **content-router:** lossless-first dispatch, cross-turn dedup, and A7 lossy-after-fold ([headroomlabs-ai#1818](headroomlabs-ai#1818)) ([60af15f](headroomlabs-ai@60af15f)) * **proxy:** add provider-only HTTP proxy ([headroomlabs-ai#1807](headroomlabs-ai#1807)) ([ebe0a3b](headroomlabs-ai@ebe0a3b)) * **proxy:** add turn-hook extension point for buffered model turns ([headroomlabs-ai#1891](headroomlabs-ai#1891)) ([ec950f7](headroomlabs-ai@ec950f7)) ### Bug Fixes * **build:** enable Intel macOS pip installs via ort-load-dynamic ([headroomlabs-ai#1538](headroomlabs-ai#1538)) ([32ce99e](headroomlabs-ai@32ce99e)) * **cache:** avoid fallback session collisions ([headroomlabs-ai#1827](headroomlabs-ai#1827)) ([0f606b6](headroomlabs-ai@0f606b6)) * **ccr:** make expired retrieve misses terminal ([headroomlabs-ai#1781](headroomlabs-ai#1781)) ([9cbdba4](headroomlabs-ai@9cbdba4)) * **ccr:** preserve Anthropic re-stream shape ([headroomlabs-ai#1854](headroomlabs-ai#1854)) ([f663894](headroomlabs-ai@f663894)) * **ccr:** preserve thinking blocks in buffered stream re-synthesis ([headroomlabs-ai#1897](headroomlabs-ai#1897)) ([ede085c](headroomlabs-ai@ede085c)) * **cli/proxy:** preserve explicit HEADROOM_MIN_TOKENS=0 / MAX_ITEMS=0 ([headroomlabs-ai#1886](headroomlabs-ai#1886)) ([3a33af1](headroomlabs-ai@3a33af1)) * **code-compressor:** CJK-aware relevance-query symbol matching ([headroomlabs-ai#1747](headroomlabs-ai#1747)) ([b38315c](headroomlabs-ai@b38315c)) * **codex:** discover updated Codex state stores ([headroomlabs-ai#1889](headroomlabs-ai#1889)) ([9d42eba](headroomlabs-ai@9d42eba)) * **codex:** OpenCode Zen telemetry attribution ([headroomlabs-ai#1648](headroomlabs-ai#1648)) ([f18c6bd](headroomlabs-ai@f18c6bd)) * **content-detector:** detect and compress space-separated JSON objects ([headroomlabs-ai#1742](headroomlabs-ai#1742)) ([5194bdc](headroomlabs-ai@5194bdc)) * **content-router:** token-measure lossless folds at the acceptance gate ([headroomlabs-ai#1772](headroomlabs-ai#1772)) ([c5493ea](headroomlabs-ai@c5493ea)) * **copilot:** normalize subscription routing host ([headroomlabs-ai#1836](headroomlabs-ai#1836)) ([afd9cbd](headroomlabs-ai@afd9cbd)) * **copilot:** route mixed-model requests per model ([headroomlabs-ai#1785](headroomlabs-ai#1785)) ([5af5e22](headroomlabs-ai@5af5e22)) * **dashboard:** deduplicate repeated savings metrics ([headroomlabs-ai#1804](headroomlabs-ai#1804)) ([88f935a](headroomlabs-ai@88f935a)) * **dashboard:** distinguish unavailable RTK from zero stats in Docker ([headroomlabs-ai#1900](headroomlabs-ai#1900)) ([87f6e93](headroomlabs-ai@87f6e93)) * **dashboard:** distinguish unavailable RTK from zero stats in Docker ([headroomlabs-ai#1901](headroomlabs-ai#1901)) ([361adcd](headroomlabs-ai@361adcd)) * **dashboard:** price proxy savings without litellm ([headroomlabs-ai#1728](headroomlabs-ai#1728)) ([188e382](headroomlabs-ai@188e382)) * detect and clear stale ANTHROPIC_BASE_URL from crashed wrap sessions ([headroomlabs-ai#1768](headroomlabs-ai#1768)) ([headroomlabs-ai#1837](headroomlabs-ai#1837)) ([84509a4](headroomlabs-ai@84509a4)) * **docker:** persist headroom workspace in compose ([headroomlabs-ai#1839](headroomlabs-ai#1839)) ([5e29c06](headroomlabs-ai@5e29c06)) * **docker:** report source build version ([headroomlabs-ai#1862](headroomlabs-ai#1862)) ([3807488](headroomlabs-ai@3807488)) * **evals:** default unparseable judge scores below pass threshold ([headroomlabs-ai#1892](headroomlabs-ai#1892)) ([42ebbc6](headroomlabs-ai@42ebbc6)) * **install:** pass sc.exe create as raw command line so binPath= quoting survives ([headroomlabs-ai#1654](headroomlabs-ai#1654)) ([headroomlabs-ai#1702](headroomlabs-ai#1702)) ([d6e0710](headroomlabs-ai@d6e0710)) * **install:** persist --no-http2 override through install apply ([headroomlabs-ai#1676](headroomlabs-ai#1676)) ([6fb5f3b](headroomlabs-ai@6fb5f3b)) * **mcp:** isolate ClaudeRegistrar CLI config env ([headroomlabs-ai#1888](headroomlabs-ai#1888)) ([1c947b1](headroomlabs-ai@1c947b1)) * **mcp:** surface dead proxy state ([headroomlabs-ai#1786](headroomlabs-ai#1786)) ([931eed8](headroomlabs-ai@931eed8)) * **memory:** resolve Trae cwd metadata from user reminders ([headroomlabs-ai#1737](headroomlabs-ai#1737)) ([headroomlabs-ai#1887](headroomlabs-ai#1887)) ([3e85eb1](headroomlabs-ai@3e85eb1)) * **opencode:** use local MCP config ([headroomlabs-ai#1383](headroomlabs-ai#1383)) ([4bd3ddf](headroomlabs-ai@4bd3ddf)) * **proxy/openai:** thread savings-profile kwargs into chat completions ([headroomlabs-ai#1606](headroomlabs-ai#1606)) ([7ff842d](headroomlabs-ai@7ff842d)) * **proxy/openai:** translate max_tokens -> max_completion_tokens on chat path ([headroomlabs-ai#1774](headroomlabs-ai#1774)) ([285808b](headroomlabs-ai@285808b)) * **proxy:** bound Codex WS compression fallback latency ([headroomlabs-ai#1802](headroomlabs-ai#1802)) ([d24a3f8](headroomlabs-ai@d24a3f8)) * **proxy:** bound HF tokenizer load and offload token counting off event loop ([headroomlabs-ai#1738](headroomlabs-ai#1738)) ([46d5d68](headroomlabs-ai@46d5d68)) * **proxy:** cancel retry backoff on shutdown ([headroomlabs-ai#1834](headroomlabs-ai#1834)) ([da2d8dc](headroomlabs-ai@da2d8dc)) * **proxy:** compress Anthropic user text blocks when enabled ([headroomlabs-ai#1875](headroomlabs-ai#1875)) ([e36439a](headroomlabs-ai@e36439a)) * **proxy:** freeze must forward cached (compressed) prefix byte-identical — stop token-mode cache busting ([headroomlabs-ai#1850](headroomlabs-ai#1850)) ([248ae0f](headroomlabs-ai@248ae0f)) * **proxy:** fsync savings dir after atomic rename ([headroomlabs-ai#1764](headroomlabs-ai#1764)) ([7de2c1e](headroomlabs-ai@7de2c1e)) * **proxy:** keep cache_control bounded + stable so the freeze overlay stops busting ([headroomlabs-ai#1852](headroomlabs-ai#1852)) ([4820134](headroomlabs-ai@4820134)) * **proxy:** persist lifetime cache-read savings across restarts ([headroomlabs-ai#1665](headroomlabs-ai#1665)) ([908997e](headroomlabs-ai@908997e)) * **proxy:** preserve streaming passthrough beta headers ([headroomlabs-ai#1783](headroomlabs-ai#1783)) ([0f553a8](headroomlabs-ai@0f553a8)) * **proxy:** release _active_streams session lock on setup-phase errors ([headroomlabs-ai#1864](headroomlabs-ai#1864)) ([2ccd831](headroomlabs-ai@2ccd831)) * **proxy:** retry HTTP/2 stream resets instead of 502ing ([headroomlabs-ai#1645](headroomlabs-ai#1645)) ([2ce19c2](headroomlabs-ai@2ce19c2)) * **proxy:** retry passthrough on transient upstream connection close ([headroomlabs-ai#1513](headroomlabs-ai#1513)) ([5d14080](headroomlabs-ai@5d14080)) * **proxy:** route Foundry Anthropic messages ([headroomlabs-ai#1878](headroomlabs-ai#1878)) ([739f654](headroomlabs-ai@739f654)) * **proxy:** serve /favicon.ico locally instead of tunneling upstream ([headroomlabs-ai#1787](headroomlabs-ai#1787)) ([headroomlabs-ai#1847](headroomlabs-ai#1847)) ([3076e32](headroomlabs-ai@3076e32)) * **proxy:** stop rtk stat failures from corrupting session baseline ([headroomlabs-ai#1693](headroomlabs-ai#1693)) ([681b9a8](headroomlabs-ai@681b9a8)) * **proxy:** strip 1m model suffix before upstream forwarding ([headroomlabs-ai#1840](headroomlabs-ai#1840)) ([e22d745](headroomlabs-ai@e22d745)) * **proxy:** subtract cache write premiums from net savings ([headroomlabs-ai#1800](headroomlabs-ai#1800)) ([53a465b](headroomlabs-ai@53a465b)) * **router:** honor MCP aliases in excluded tools ([headroomlabs-ai#1822](headroomlabs-ai#1822)) ([headroomlabs-ai#1863](headroomlabs-ai#1863)) ([140d6e4](headroomlabs-ai@140d6e4)) * **rtk:** link managed rtk onto PATH instead of mutating the hook ([headroomlabs-ai#1698](headroomlabs-ai#1698)) ([140cb05](headroomlabs-ai@140cb05)) * **streaming:** preserve server_tool_use sse blocks ([headroomlabs-ai#1826](headroomlabs-ai#1826)) ([4ac5493](headroomlabs-ai@4ac5493)) * **toin:** publish skip compression recommendations ([headroomlabs-ai#1782](headroomlabs-ai#1782)) ([be51008](headroomlabs-ai@be51008)) * **transforms:** normalize diff compressor context ([headroomlabs-ai#1801](headroomlabs-ai#1801)) ([838c523](headroomlabs-ai@838c523)) * **transforms:** pass through ragged tables instead of misaligning columns ([headroomlabs-ai#1713](headroomlabs-ai#1713)) ([c7665ca](headroomlabs-ai@c7665ca)) * use rtk native Cursor hook instead of injecting .cursorrules ([headroomlabs-ai#756](headroomlabs-ai#756)) ([headroomlabs-ai#1846](headroomlabs-ai#1846)) ([1573f1f](headroomlabs-ai@1573f1f)) * **wrap:** replace stale-proxy detection with Vite-style port fallback ([headroomlabs-ai#1406](headroomlabs-ai#1406)) ([b4205c6](headroomlabs-ai@b4205c6)) ### Performance Improvements * **proxy:** cap compression workers to CPU count ([headroomlabs-ai#1803](headroomlabs-ai#1803)) ([0a3851b](headroomlabs-ai@0a3851b)) * **savings:** batch tracker persistence off the request hot path ([headroomlabs-ai#1817](headroomlabs-ai#1817)) ([451b9f0](headroomlabs-ai@451b9f0)) ### Dependencies * bump the cargo-minor-patch group across 1 directory with 7 updates ([headroomlabs-ai#1909](headroomlabs-ai#1909)) ([45601d9](headroomlabs-ai@45601d9)) * bump the npm-minor-patch group across 4 directories with 18 updates ([headroomlabs-ai#1907](headroomlabs-ai#1907)) ([8872bbc](headroomlabs-ai@8872bbc)) </details> --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
The compression acceptance gate rejected any compression that saved less than ~15% (
min_ratiointerpolated 0.85 at low context pressure → 0.65 under pressure). That floor was a crude proxy for "big enough to justify busting the prefix cache," but it dropped genuine token savings — notably lossless code/log folds that shrink <15% (theratio_too_highrejections).This makes the gate accept any real shrink (
ratio < 1.0): any token saved is worth taking. The two guards that actually protect correctness are untouched:HEADROOM_NET_COST_POLICY=1, opt-in) — precisely accounts for the prefix-cache-bust economics (savings × expected-reads vs one-time suffix re-write) when a session wants that protection.Lowering the two values back to
0.85/0.65restores the savings floor.Closes #
Type of Change
Changes Made
ContentRouterConfig.min_ratio_relaxed:0.85 → 1.0ContentRouterConfig.min_ratio_aggressive:0.65 → 1.0compression_ratio < 1.0at every context pressure; reversibility + net-cost guards unchanged.Testing
pytest)ruff check)mypy headroom)min_ratiovalues and are unaffected)Test Output
Real Behavior Proof
PYTHONPATHpinned to the branch checkout.min_ratio) unaffected; no default-floor test regressed; blocks that previously producedratio_too_highat ratios in[0.85, 1.0)are now accepted.router:*acceptances, fewerratio_too_high) is inferred from the gate logic + suite.Review Readiness
Checklist
Additional Notes
Deliberate tradeoff (discussed and chosen): without the net-cost policy enabled, accepting sub-15% wins can be net-negative on prompt-cached sessions, because compressing a block invalidates the cached suffix (a one-time re-write, at 1.25× on Anthropic). If that shows up in practice, enable
HEADROOM_NET_COST_POLICY=1(the precise economics guard) or restore a floor by lowering the twomin_ratio_*values.