Skip to content

Commit 1cdcf8a

Browse files
committed
Merge remote-tracking branch 'headroomlabs/main' into fix-1513-ci
2 parents 7b4aa7a + 6c48ac8 commit 1cdcf8a

14 files changed

Lines changed: 548 additions & 30 deletions

‎CHANGELOG.md‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -9,6 +9,13 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
99
## Unreleased
1010

1111
### Fixed
12+
- `--backend bedrock` now fails fast with an actionable error when temporary
13+
AWS credentials (`AWS_SESSION_TOKEN`) are used but botocore is not installed
14+
(e.g. the slim default Docker image). litellm's session-token auth path
15+
imports botocore, so the missing dependency previously surfaced only at
16+
request time as a misleading `authentication_error: No module named
17+
'botocore'`. The proxy now tells the user to install the `bedrock` extra up
18+
front ([#1551](https://github.com/headroomlabs-ai/headroom/issues/1551)).
1219
- Content detection no longer crashes the proxy on text containing an
1320
orphaned `+++ ` target line with no preceding `--- ` source line (common in
1421
`set -x` xtrace output and partial diffs). The bundled `unidiff` 0.4.0 parser
@@ -63,6 +70,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
6370
* **proxy:** include the system prompt, tools, and the response-shaping request fields in the SemanticCache key. `_compute_key` hashed only `{model, messages}`, so two non-streaming requests with identical messages but a different top-level `system` prompt, tool set, sampling config, or output-shaping field collided on one key and the second caller was served the first's cached response — generated under different request semantics, in the default config (`cache_enabled` defaults on). The key now folds the request fields that shape generation — `temperature`/`top_p`/`top_k`/`max_tokens`/`stop`, plus OpenAI `tool_choice`/`response_format`/`parallel_tool_calls`/`seed`/`presence_penalty`/`frequency_penalty`/`logit_bias`/`n`/`logprobs`/`top_logprobs`/`reasoning_effort`/`verbosity`/`modalities` and Anthropic `thinking`/`tool_choice`/`output_config` — canonicalizing `system`/`tools` so a moved `cache_control` breakpoint does not fragment it, and the handlers snapshot the fields once at the cache read and reuse them at write so a body mutated by the pipeline cannot diverge the key. Non-streaming path only.
6471
* **learn (verbosity):** `--verbosity --apply --all` now aggregates the savings baseline across every project instead of overwriting it per project (last-project-wins), which previously left the output shaper with a tiny, unrepresentative baseline. The applied verbosity level comes from the project with the most samples ([#1288](https://github.com/headroomlabs-ai/headroom/pull/1288)).
6572
* **proxy/anthropic:** restore token-mode compression on continued Claude Code turns with a frozen prefix and deferred CCR tool injection. Token mode now runs request-side compression even when the client did not pre-register `headroom_retrieve`, relying on the existing marker-triggered injection override to keep emitted CCR markers redeemable ([#1487](https://github.com/headroomlabs-ai/headroom/issues/1487)).
73+
* **proxy:** the dedicated OpenAI handlers (`/v1/chat/completions`, `/v1/responses`) now honor the `x-headroom-base-url` request header, matching the generic passthrough route. Previously only the catch-all passthrough honored it, so OpenAI-compatible gateways (LiteLLM, CPA, self-hosted vLLM, Azure OpenAI) routed correctly for passthrough traffic but the dedicated chat/responses handlers ignored the header and fell back to the default `OPENAI_API_URL`, sending requests (and the user's provider key) to the wrong upstream.
6674
* **subscription:** stop zeroing the 5-hour headroom contribution counters on every poll. The rollover check compared `five_hour.resets_at` with a bare `!=`, but the usage API reports that timestamp with second-level jitter (observed flapping between `01:59:59Z` and `02:00:00Z` on consecutive polls within the same window), so a spurious "5h window rolled over" reset fired every poll interval (~5 min) and the dashboard's per-window savings stuck near 0%. Only a forward jump larger than `_ROLLOVER_MIN_ADVANCE` (1 minute) now counts as a real rollover.
6775
* **wrap:** keep the shared proxy alive when the agent that launched it closes *ungracefully* on Windows. `_start_proxy` spawned the proxy without detaching it, so it stayed in the launcher's console and Job object; closing that terminal window (or `taskkill`/a crash) tree-killed the proxy, bypassing the marker-based reference counting in `_make_cleanup` and breaking every other `headroom wrap` instance routed through the same port. The proxy is now created with `CREATE_NO_WINDOW | CREATE_NEW_PROCESS_GROUP | CREATE_BREAKAWAY_FROM_JOB` (with a graceful fallback when the launcher's Job forbids breakaway); POSIX behavior is unchanged. `CREATE_NO_WINDOW` (rather than `DETACHED_PROCESS`) gives the proxy its own *hidden* console: `DETACHED_PROCESS` leaves a console-subsystem exe (`python.exe`) consoleless, so Windows surfaces a visible console window whose close button kills the proxy.
6876
* **transforms/content_router:** stop replacing `role="tool"` output with a lossy-unrecoverable summary on the live compression path (refs [#1307](https://github.com/chopratejas/headroom/issues/1307)). `ContentRouter.apply()` routed OpenAI-style `role="tool"` string messages — `Bash`/`grep`/`ls`/`cat` output — through the ML/word-drop summarizers; when the result carried no CCR retrieve marker (CCR off, ratio >= 0.8, or the size-gate fallback) the original was unrecoverable and the agent acted on a fabricated summary. Tool-role string content is now kept verbatim unless the compressed form is CCR-recoverable. Assistant/user text is unaffected, and structurally-lossless passes (SmartCrusher/Log/Search) still apply. The Anthropic `tool_result` block path is tracked separately.

‎headroom/backends/litellm.py‎

Lines changed: 15 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -12,8 +12,10 @@
1212

1313
from __future__ import annotations
1414

15+
import importlib.util
1516
import json
1617
import logging
18+
import os
1719
import uuid
1820
from collections.abc import AsyncIterator
1921
from dataclasses import dataclass, field
@@ -423,6 +425,19 @@ def __init__(
423425

424426
# For Bedrock, fetch model map dynamically from AWS API
425427
if provider == "bedrock":
428+
# litellm takes the botocore-backed `_auth_with_aws_session_token`
429+
# path as soon as temporary credentials (AWS_SESSION_TOKEN) are
430+
# present. botocore is an optional dependency (the `bedrock`
431+
# extra); when it is absent — as in the slim default Docker image —
432+
# the failure only surfaces at request time as a misleading
433+
# `authentication_error: No module named 'botocore'` (#1551). Fail
434+
# fast at startup with an actionable message instead.
435+
if os.environ.get("AWS_SESSION_TOKEN") and importlib.util.find_spec("botocore") is None:
436+
raise ImportError(
437+
"Bedrock with temporary credentials (AWS_SESSION_TOKEN) requires "
438+
"botocore, which is not installed. Install the bedrock extra: "
439+
"pip install 'headroom-ai[bedrock]' (or pip install botocore)."
440+
)
426441
self._model_map = _fetch_bedrock_inference_profiles(region)
427442
litellm.set_verbose = False # Reduce noise
428443
else:

‎headroom/cache/backends/__init__.py‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -13,7 +13,7 @@
1313
from headroom.cache.backends import SQLiteBackend, CompressionStoreBackend
1414
from headroom.cache.compression_store import CompressionStore, get_compression_store
1515
16-
# Env-driven default (SQLite at ~/.headroom/ccr_store.db)
16+
# Env-driven default (SQLite at workspace_dir()/ccr_store.db)
1717
store = get_compression_store()
1818
1919
# Direct construction defaults to in-memory; pass a backend for persistence

‎headroom/cache/backends/sqlite.py‎

Lines changed: 5 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -13,7 +13,7 @@
1313
1414
Set ``HEADROOM_CCR_BACKEND=memory`` to opt back into the in-memory
1515
backend, or ``HEADROOM_CCR_SQLITE_PATH`` to relocate the database file
16-
(default ``~/.headroom/ccr_store.db``).
16+
(default ``workspace_dir()/ccr_store.db``).
1717
"""
1818

1919
from __future__ import annotations
@@ -49,11 +49,13 @@
4949

5050

5151
def default_db_path() -> Path:
52-
"""Resolve the database path (env override or ~/.headroom/)."""
52+
"""Resolve the database path (env override, else workspace root)."""
5353
env = os.environ.get("HEADROOM_CCR_SQLITE_PATH", "").strip()
5454
if env:
5555
return Path(env).expanduser()
56-
return Path.home() / ".headroom" / "ccr_store.db"
56+
from ...paths import workspace_dir
57+
58+
return workspace_dir() / "ccr_store.db"
5759

5860

5961
class SQLiteBackend:

‎headroom/cache/compression_store.py‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -946,9 +946,9 @@ def clear_request_compression_store() -> None:
946946
def _create_default_ccr_backend() -> CompressionStoreBackend | None:
947947
"""Create a CCR backend from env (e.g. HEADROOM_CCR_BACKEND=redis).
948948
949-
Default (env unset or "sqlite"): SQLiteBackend at
950-
~/.headroom/ccr_store.db — restart-safe and shared across worker
951-
processes, which the session-scale 30-minute TTL assumes.
949+
Default (env unset or "sqlite"): SQLiteBackend at workspace_dir()/ccr_store.db
950+
— restart-safe and shared across worker processes, which the
951+
session-scale 30-minute TTL assumes.
952952
"memory" opts back into the in-process dict. Other values load
953953
adapters via setuptools entry point 'headroom.ccr_backend'.
954954
Returns None to use InMemoryBackend.

‎headroom/cli/wrap.py‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -447,6 +447,10 @@ def _start_proxy(
447447
# Ensure proxy subprocess uses UTF-8 (Windows defaults to cp1252)
448448
proxy_env = os.environ.copy()
449449
proxy_env["PYTHONIOENCODING"] = "utf-8"
450+
# Vertex AI RST_STREAMs HTTP/2 connections (error_code:2). Force HTTP/1.1
451+
# when wrapping a Vertex-mode client so upstream requests succeed.
452+
if os.environ.get("CLAUDE_CODE_USE_VERTEX") or os.environ.get("ANTHROPIC_VERTEX_PROJECT_ID"):
453+
proxy_env.setdefault("HEADROOM_HTTP2", "false")
450454
# Tell the proxy which agent is being wrapped (for traffic learning output)
451455
if agent_type != "unknown":
452456
proxy_env["HEADROOM_AGENT_TYPE"] = agent_type

‎headroom/providers/proxy_routes.py‎

Lines changed: 25 additions & 19 deletions
Original file line numberDiff line numberDiff line change
@@ -733,22 +733,25 @@ async def vertex_raw_predict(
733733
return await vertex_publisher_passthrough(request, publisher, "rawPredict")
734734

735735
@app.post(
736-
"/projects/{project}/locations/{location}/publishers/anthropic/models/{model}:rawPredict"
736+
"/projects/{project}/locations/{location}/publishers/{publisher}/models/{model}:rawPredict"
737737
)
738738
async def vertex_raw_predict_no_version(
739739
request: Request,
740740
project: str,
741741
location: str,
742+
publisher: str,
742743
model: str,
743744
):
744-
del project
745-
target = _vertex_target_for_location(proxy, location).rstrip("/") + "/v1"
746-
return await proxy.handle_anthropic_messages(
747-
request,
748-
target,
749-
"vertex:anthropic",
750-
model,
751-
)
745+
if publisher == "anthropic":
746+
del project
747+
target = _vertex_target_for_location(proxy, location).rstrip("/") + "/v1"
748+
return await proxy.handle_anthropic_messages(
749+
request,
750+
target,
751+
"vertex:anthropic",
752+
model,
753+
)
754+
return await vertex_publisher_passthrough(request, publisher, "rawPredict")
752755

753756
@app.post(
754757
"/{api_version}/projects/{project}/locations/{location}/publishers/{publisher}/models/{model}:streamRawPredict"
@@ -773,23 +776,26 @@ async def vertex_stream_raw_predict(
773776
return await vertex_publisher_passthrough(request, publisher, "streamRawPredict")
774777

775778
@app.post(
776-
"/projects/{project}/locations/{location}/publishers/anthropic/models/{model}:streamRawPredict"
779+
"/projects/{project}/locations/{location}/publishers/{publisher}/models/{model}:streamRawPredict"
777780
)
778781
async def vertex_stream_raw_predict_no_version(
779782
request: Request,
780783
project: str,
781784
location: str,
785+
publisher: str,
782786
model: str,
783787
):
784-
del project
785-
target = _vertex_target_for_location(proxy, location).rstrip("/") + "/v1"
786-
return await proxy.handle_anthropic_messages(
787-
request,
788-
target,
789-
"vertex:anthropic",
790-
model,
791-
True,
792-
)
788+
if publisher == "anthropic":
789+
del project
790+
target = _vertex_target_for_location(proxy, location).rstrip("/") + "/v1"
791+
return await proxy.handle_anthropic_messages(
792+
request,
793+
target,
794+
"vertex:anthropic",
795+
model,
796+
True,
797+
)
798+
return await vertex_publisher_passthrough(request, publisher, "streamRawPredict")
793799

794800
@app.get("/v1/models")
795801
async def list_models(request: Request):

‎headroom/proxy/handlers/openai.py‎

Lines changed: 12 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -784,6 +784,18 @@ def _headroom_bypass_enabled(headers: Any) -> bool:
784784
"""Return True when inbound headers request full passthrough."""
785785
return _headroom_bypass_enabled(headers)
786786

787+
def _resolve_openai_upstream(self, request: Request) -> str:
788+
"""Return the OpenAI upstream base URL for ``request``.
789+
790+
Honors the ``x-headroom-base-url`` request header so OpenAI-compatible
791+
gateways (LiteLLM, CPA, self-hosted vLLM, Azure OpenAI) route through
792+
the dedicated ``/v1/chat/completions`` and ``/v1/responses`` handlers,
793+
not just the generic passthrough route that already honors it. Falls
794+
back to the configured ``OPENAI_API_URL`` (``OPENAI_TARGET_API_URL``).
795+
"""
796+
custom = request.headers.get("x-headroom-base-url", "").strip()
797+
return custom or self.OPENAI_API_URL
798+
787799
@staticmethod
788800
def _strict_previous_turn_frozen_count(
789801
messages: list[dict[str, Any]],

‎headroom/transforms/content_router.py‎

Lines changed: 82 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -42,6 +42,7 @@
4242
import os
4343
import re
4444
import sys
45+
import threading
4546
import time
4647
from concurrent.futures import ThreadPoolExecutor
4748
from dataclasses import dataclass, field
@@ -66,6 +67,7 @@
6667

6768
_detect_backend_warned = False
6869
_detect_panic_warned = False
70+
_detect_native_unhealthy = False # circuit breaker: native detect hung once (#575)
6971

7072

7173
def _router_debug_dumps(value: Any) -> str:
@@ -129,6 +131,60 @@ def _resolve_detect_backend() -> str:
129131
return "python" if sys.platform == "win32" else "rust"
130132

131133

134+
_DETECT_TIMEOUT_ENV = "HEADROOM_DETECT_TIMEOUT_SECS"
135+
_DEFAULT_DETECT_TIMEOUT_SECS = 5.0
136+
137+
138+
def _detect_timeout_secs() -> float:
139+
"""Watchdog budget (seconds) for one native detect call.
140+
141+
Override with ``HEADROOM_DETECT_TIMEOUT_SECS``; blank, non-numeric, or
142+
non-positive values fall back to the default.
143+
"""
144+
raw = os.environ.get(_DETECT_TIMEOUT_ENV, "").strip()
145+
if not raw:
146+
return _DEFAULT_DETECT_TIMEOUT_SECS
147+
try:
148+
secs = float(raw)
149+
except ValueError:
150+
return _DEFAULT_DETECT_TIMEOUT_SECS
151+
return secs if secs > 0 else _DEFAULT_DETECT_TIMEOUT_SECS
152+
153+
154+
def _rust_detect_watchdogged(rust_detect: Any, content: str, timeout: float) -> Any:
155+
"""Run the native detector under a watchdog thread, bounding the caller's wait.
156+
157+
On Windows the first native ``detect_content_type`` can park forever in an
158+
ort/``Once`` init (``WaitOnAddress``) at 0% CPU, and a wedged native call
159+
cannot be cancelled from Python (#575). The native call releases the GIL
160+
while parked, so a watchdog thread runs it and the caller waits at most
161+
``timeout`` seconds before raising ``TimeoutError`` — letting
162+
``_detect_content`` degrade to the pure-Python detector instead of
163+
deadlocking (and, in the proxy, instead of permanently consuming a
164+
compression-executor worker — see #575's executor-saturation report).
165+
166+
# ponytail: can't kill a GIL-released native call; the watchdog frees the
167+
# caller and the stuck daemon thread is left to die with the process. The
168+
# upgrade path is the Rust-side fix that makes first-call init non-blocking.
169+
"""
170+
box: dict[str, Any] = {}
171+
172+
def _run() -> None:
173+
try:
174+
box["result"] = rust_detect(content)
175+
except BaseException as exc: # noqa: BLE001 — relayed to the caller's degrade path
176+
box["error"] = exc
177+
178+
worker = threading.Thread(target=_run, name="headroom-detect-watchdog", daemon=True)
179+
worker.start()
180+
worker.join(timeout)
181+
if worker.is_alive():
182+
raise TimeoutError(f"native detect_content_type exceeded {timeout:.1f}s watchdog")
183+
if "error" in box:
184+
raise box["error"]
185+
return box["result"]
186+
187+
132188
def _detect_content(content: str) -> DetectionResult:
133189
"""Detect content type via the native chain, with a safe Windows default.
134190
@@ -144,7 +200,7 @@ def _detect_content(content: str) -> DetectionResult:
144200
only consumed `.content_type` from it; the strategy mapping in
145201
`_strategy_from_detection` keys off that field alone.
146202
"""
147-
global _detect_backend_warned
203+
global _detect_backend_warned, _detect_panic_warned, _detect_native_unhealthy
148204

149205
backend = _resolve_detect_backend()
150206
if backend == "python":
@@ -157,11 +213,23 @@ def _detect_content(content: str) -> DetectionResult:
157213
)
158214
return _regex_detect_content_type(content)
159215

216+
if _detect_native_unhealthy:
217+
# Circuit breaker (#575): the native detector hung once under the
218+
# watchdog; every later call would wait the full budget and strand
219+
# another stuck daemon thread, so route straight to pure-Python.
220+
return _regex_detect_content_type(content)
221+
160222
from headroom._core import detect_content_type as _rust_detect
161223

162-
global _detect_panic_warned
163224
try:
164-
rust_result = _rust_detect(content)
225+
if sys.platform == "win32":
226+
# Windows is the only platform where the native detector can deadlock
227+
# on first use (#575); bound it with a watchdog so a hang degrades to
228+
# the pure-Python detector below. Elsewhere it is the trusted default
229+
# hot path — call it directly, with no per-call thread overhead.
230+
rust_result = _rust_detect_watchdogged(_rust_detect, content, _detect_timeout_secs())
231+
else:
232+
rust_result = _rust_detect(content)
165233
# Rust's `content_type` is the lowercase string tag (e.g.
166234
# "json_array"); translate to the Python `ContentType` enum so
167235
# downstream mapping keys match.
@@ -178,7 +246,17 @@ def _detect_content(content: str) -> DetectionResult:
178246
# as asyncio.CancelledError — keep them propagating.
179247
if isinstance(exc, asyncio.CancelledError):
180248
raise
181-
if not _detect_panic_warned:
249+
if isinstance(exc, TimeoutError):
250+
# Watchdog tripped: the native detector hung (#575). Disable it
251+
# process-wide so later calls don't each wait the full budget and
252+
# strand another daemon thread in the wedged native call.
253+
_detect_native_unhealthy = True
254+
logger.warning(
255+
"Native content detector hung (%s); disabling it for this process "
256+
"and using pure-Python detection.",
257+
exc,
258+
)
259+
elif not _detect_panic_warned:
182260
_detect_panic_warned = True
183261
logger.warning(
184262
"Native content detector failed (%s); falling back to pure-Python detection.",

0 commit comments

Comments
 (0)