Summary
Headroom proxy returns HTTP 502 when proxying to an OpenAI-compatible upstream that works correctly when called directly. The failure happens on GET /v1/models, non-streaming POST /v1/responses, and streaming POST /v1/responses.
The Headroom-side traceback shows httpx.RemoteProtocolError: peer closed connection without sending complete message body (incomplete chunked read).
Environment
- Headroom:
0.26.0
- Python runtime used by Headroom: Python 3.13
httpx: 0.28.1
httpcore: 1.0.9
h11: 0.16.0
fastapi: 0.137.1
uvicorn: 0.49.0
starlette: 1.3.1
- OS: macOS
- Upstream: OpenAI-compatible local endpoint exposed on
http://127.0.0.1:8317/v1
Headroom is launched as:
headroom proxy \
--host 127.0.0.1 \
--port 8787 \
--mode token \
--backend anthropic \
--no-telemetry
Relevant env vars:
OPENAI_TARGET_API_URL=http://127.0.0.1:8317/v1
ANTHROPIC_TARGET_API_URL=http://127.0.0.1:8317
HEADROOM_TELEMETRY=off
HEADROOM_DISABLE_KOMPRESS=1
HEADROOM_WS_FAIL_OPEN_ON_COMPRESSION_FAILURE=1
GET /readyz reports the upstream as healthy:
{
"status": "healthy",
"ready": true,
"checks": {
"upstream": {
"enabled": true,
"ready": true,
"status": "healthy",
"url": "http://127.0.0.1:8317",
"error": null
}
}
}
Reproduction
With a valid API key in KEY, direct upstream calls succeed:
curl -sS -o /tmp/models-test.json -w "%{http_code} %{content_type}\n" \
"http://127.0.0.1:8317/v1/models" \
-H "Authorization: Bearer $KEY"
# 200 application/json; charset=utf-8
curl -sS -o /tmp/resp-ns.json -w "%{http_code} %{content_type}\n" \
-X POST "http://127.0.0.1:8317/v1/responses" \
-H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" \
-d '{"model":"gpt-5.5","input":"Please reply OK only","max_output_tokens":20,"stream":false}'
# 200 application/json
curl -sS -o /tmp/resp-s.json -w "%{http_code} %{content_type}\n" \
-X POST "http://127.0.0.1:8317/v1/responses" \
-H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" \
-d '{"model":"gpt-5.5","input":"Please reply OK only","max_output_tokens":20,"stream":true}'
# 200 text/event-stream
The same calls through Headroom fail:
curl -sS -o /tmp/models-test.json -w "%{http_code} %{content_type}\n" \
"http://127.0.0.1:8787/v1/models" \
-H "Authorization: Bearer $KEY"
# 502
curl -sS -o /tmp/resp-ns.json -w "%{http_code} %{content_type}\n" \
-X POST "http://127.0.0.1:8787/v1/responses" \
-H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" \
-d '{"model":"gpt-5.5","input":"Please reply OK only","max_output_tokens":20,"stream":false}'
# 502 application/json
curl -sS -o /tmp/resp-s.json -w "%{http_code} %{content_type}\n" \
-X POST "http://127.0.0.1:8787/v1/responses" \
-H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" \
-d '{"model":"gpt-5.5","input":"Please reply OK only","max_output_tokens":20,"stream":true}'
# 502
Observed Behavior
Headroom returns HTTP 502 for OpenAI-compatible passthrough/proxy requests.
The upstream server logs these same requests as HTTP 200 in the direct path. Other clients using the same upstream directly, including Codex/OpenAI-compatible clients and Anthropic-compatible clients, work as expected.
Headroom Traceback
Relevant excerpt:
httpcore.RemoteProtocolError: peer closed connection without sending complete message body (incomplete chunked read)
The above exception was the direct cause of the following exception:
httpx.RemoteProtocolError: peer closed connection without sending complete message body (incomplete chunked read)
File ".../site-packages/headroom/providers/proxy_routes.py", line 690, in list_models
return await proxy.handle_passthrough(...)
File ".../site-packages/headroom/proxy/handlers/openai.py", line 6074, in handle_passthrough
response = await passthrough_client.request(...)
File ".../site-packages/httpx/_models.py", line 979, in aread
self._content = b"".join([part async for part in self.aiter_bytes()])
Expected Behavior
Headroom should either:
- Successfully proxy OpenAI-compatible responses from this upstream, or
- Gracefully handle upstream chunked-transfer edge cases without converting all
/v1/models and /v1/responses calls into 502s.
Notes / Hypothesis
This looks like a compatibility issue between Headroom/httpx passthrough handling and an OpenAI-compatible upstream's chunked response behavior.
Direct curl calls to the upstream succeed, so this does not appear to be an auth, DNS, or connectivity issue. It may be related to how Headroom fully reads passthrough responses via response.aread() and how httpx/httpcore treats incomplete chunked termination.
Would it be possible for Headroom to tolerate this case, normalize passthrough responses, or provide a config flag to disable full-body buffering for /v1/models and/or OpenAI Responses passthrough?
Summary
Headroom proxy returns HTTP 502 when proxying to an OpenAI-compatible upstream that works correctly when called directly. The failure happens on
GET /v1/models, non-streamingPOST /v1/responses, and streamingPOST /v1/responses.The Headroom-side traceback shows
httpx.RemoteProtocolError: peer closed connection without sending complete message body (incomplete chunked read).Environment
0.26.0httpx:0.28.1httpcore:1.0.9h11:0.16.0fastapi:0.137.1uvicorn:0.49.0starlette:1.3.1http://127.0.0.1:8317/v1Headroom is launched as:
Relevant env vars:
GET /readyzreports the upstream as healthy:{ "status": "healthy", "ready": true, "checks": { "upstream": { "enabled": true, "ready": true, "status": "healthy", "url": "http://127.0.0.1:8317", "error": null } } }Reproduction
With a valid API key in
KEY, direct upstream calls succeed:The same calls through Headroom fail:
Observed Behavior
Headroom returns HTTP 502 for OpenAI-compatible passthrough/proxy requests.
The upstream server logs these same requests as HTTP 200 in the direct path. Other clients using the same upstream directly, including Codex/OpenAI-compatible clients and Anthropic-compatible clients, work as expected.
Headroom Traceback
Relevant excerpt:
Expected Behavior
Headroom should either:
/v1/modelsand/v1/responsescalls into 502s.Notes / Hypothesis
This looks like a compatibility issue between Headroom/httpx passthrough handling and an OpenAI-compatible upstream's chunked response behavior.
Direct
curlcalls to the upstream succeed, so this does not appear to be an auth, DNS, or connectivity issue. It may be related to how Headroom fully reads passthrough responses viaresponse.aread()and howhttpx/httpcoretreats incomplete chunked termination.Would it be possible for Headroom to tolerate this case, normalize passthrough responses, or provide a config flag to disable full-body buffering for
/v1/modelsand/or OpenAI Responses passthrough?