Skip to content

[BUG] Proxy returns 502 with httpx incomplete chunked read against OpenAI-compatible upstream #1112

Description

@983033995

Summary

Headroom proxy returns HTTP 502 when proxying to an OpenAI-compatible upstream that works correctly when called directly. The failure happens on GET /v1/models, non-streaming POST /v1/responses, and streaming POST /v1/responses.

The Headroom-side traceback shows httpx.RemoteProtocolError: peer closed connection without sending complete message body (incomplete chunked read).

Environment

  • Headroom: 0.26.0
  • Python runtime used by Headroom: Python 3.13
  • httpx: 0.28.1
  • httpcore: 1.0.9
  • h11: 0.16.0
  • fastapi: 0.137.1
  • uvicorn: 0.49.0
  • starlette: 1.3.1
  • OS: macOS
  • Upstream: OpenAI-compatible local endpoint exposed on http://127.0.0.1:8317/v1

Headroom is launched as:

headroom proxy \
  --host 127.0.0.1 \
  --port 8787 \
  --mode token \
  --backend anthropic \
  --no-telemetry

Relevant env vars:

OPENAI_TARGET_API_URL=http://127.0.0.1:8317/v1
ANTHROPIC_TARGET_API_URL=http://127.0.0.1:8317
HEADROOM_TELEMETRY=off
HEADROOM_DISABLE_KOMPRESS=1
HEADROOM_WS_FAIL_OPEN_ON_COMPRESSION_FAILURE=1

GET /readyz reports the upstream as healthy:

{
  "status": "healthy",
  "ready": true,
  "checks": {
    "upstream": {
      "enabled": true,
      "ready": true,
      "status": "healthy",
      "url": "http://127.0.0.1:8317",
      "error": null
    }
  }
}

Reproduction

With a valid API key in KEY, direct upstream calls succeed:

curl -sS -o /tmp/models-test.json -w "%{http_code} %{content_type}\n" \
  "http://127.0.0.1:8317/v1/models" \
  -H "Authorization: Bearer $KEY"

# 200 application/json; charset=utf-8
curl -sS -o /tmp/resp-ns.json -w "%{http_code} %{content_type}\n" \
  -X POST "http://127.0.0.1:8317/v1/responses" \
  -H "Authorization: Bearer $KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-5.5","input":"Please reply OK only","max_output_tokens":20,"stream":false}'

# 200 application/json
curl -sS -o /tmp/resp-s.json -w "%{http_code} %{content_type}\n" \
  -X POST "http://127.0.0.1:8317/v1/responses" \
  -H "Authorization: Bearer $KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-5.5","input":"Please reply OK only","max_output_tokens":20,"stream":true}'

# 200 text/event-stream

The same calls through Headroom fail:

curl -sS -o /tmp/models-test.json -w "%{http_code} %{content_type}\n" \
  "http://127.0.0.1:8787/v1/models" \
  -H "Authorization: Bearer $KEY"

# 502
curl -sS -o /tmp/resp-ns.json -w "%{http_code} %{content_type}\n" \
  -X POST "http://127.0.0.1:8787/v1/responses" \
  -H "Authorization: Bearer $KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-5.5","input":"Please reply OK only","max_output_tokens":20,"stream":false}'

# 502 application/json
curl -sS -o /tmp/resp-s.json -w "%{http_code} %{content_type}\n" \
  -X POST "http://127.0.0.1:8787/v1/responses" \
  -H "Authorization: Bearer $KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-5.5","input":"Please reply OK only","max_output_tokens":20,"stream":true}'

# 502

Observed Behavior

Headroom returns HTTP 502 for OpenAI-compatible passthrough/proxy requests.

The upstream server logs these same requests as HTTP 200 in the direct path. Other clients using the same upstream directly, including Codex/OpenAI-compatible clients and Anthropic-compatible clients, work as expected.

Headroom Traceback

Relevant excerpt:

httpcore.RemoteProtocolError: peer closed connection without sending complete message body (incomplete chunked read)

The above exception was the direct cause of the following exception:

httpx.RemoteProtocolError: peer closed connection without sending complete message body (incomplete chunked read)

File ".../site-packages/headroom/providers/proxy_routes.py", line 690, in list_models
    return await proxy.handle_passthrough(...)

File ".../site-packages/headroom/proxy/handlers/openai.py", line 6074, in handle_passthrough
    response = await passthrough_client.request(...)

File ".../site-packages/httpx/_models.py", line 979, in aread
    self._content = b"".join([part async for part in self.aiter_bytes()])

Expected Behavior

Headroom should either:

  1. Successfully proxy OpenAI-compatible responses from this upstream, or
  2. Gracefully handle upstream chunked-transfer edge cases without converting all /v1/models and /v1/responses calls into 502s.

Notes / Hypothesis

This looks like a compatibility issue between Headroom/httpx passthrough handling and an OpenAI-compatible upstream's chunked response behavior.

Direct curl calls to the upstream succeed, so this does not appear to be an auth, DNS, or connectivity issue. It may be related to how Headroom fully reads passthrough responses via response.aread() and how httpx/httpcore treats incomplete chunked termination.

Would it be possible for Headroom to tolerate this case, normalize passthrough responses, or provide a config flag to disable full-body buffering for /v1/models and/or OpenAI Responses passthrough?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions