Skip to content

Commit 285808b

Browse files
authored
fix(proxy/openai): translate max_tokens -> max_completion_tokens on chat path (#1774)
## Description GPT-5 / o-series chat models reject the legacy `max_tokens` — `AI_APICallError: Unsupported parameter: 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead.` — while gpt-4o/4.1 accept `max_completion_tokens` too. openai-compatible clients (opencode via `@ai-sdk/openai-compatible`, older SDKs) still send `max_tokens`, so requests for GPT-5 models fail at the proxy's OpenAI upstream. This is a blocker for any such client pointed at a GPT-5 model through Headroom. The proxy already owns the outbound `/v1/chat/completions` body (it rewrites `messages` to compress them), so translate the token param there: rename `max_tokens` → `max_completion_tokens` when the newer form isn't already set, then drop the rejected legacy key. One-way, safe for current OpenAI models; no-op when the client already sends `max_completion_tokens`. The Responses path (`max_output_tokens`) is unaffected. Closes # ## Type of Change - [x] Bug fix (non-breaking change that fixes an issue) ## Changes Made - New `_normalize_openai_max_tokens(body)` helper + call in `handle_openai_chat` after body finalization, before upstream forward. ## Testing - [x] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added ### Test Output ```text tests/test_openai_max_completion_tokens.py .... 6 passed ruff check ... All checks passed! mypy headroom/proxy/handlers/openai.py ... Success: no issues found ``` ## Real Behavior Proof - Environment: local worktree, Python 3.12. - Exact command / steps: reproduced live — opencode (`@ai-sdk/openai-compatible` → Headroom proxy) targeting `gpt-5.3-chat-latest` failed with `Unsupported parameter: 'max_tokens' ... Use 'max_completion_tokens'` in the DEBUG stream log. The shim renames the param on the outbound body. - Observed result: unit tests confirm the rename/drop/no-op cases. - Not tested: full live opencode completion (its headless `run` stalls for unrelated reasons in this env — separate from this param fix). ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Additional Notes Discovered while debugging why opencode wouldn't run through the proxy: three layered blockers — (1) missing `models` map in the injected provider config [PR #1716], (2) no `apiKey` in the injected config / HTTP path doesn't inject `OPENAI_API_KEY` like the WS path does, (3) this `max_tokens` vs `max_completion_tokens` mismatch. This PR addresses (3).
1 parent 37a12dd commit 285808b

2 files changed

Lines changed: 78 additions & 0 deletions

File tree

‎headroom/proxy/handlers/openai.py‎

Lines changed: 27 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -84,6 +84,24 @@ def _codex_ws_compression_timeout_seconds() -> float:
8484
_OPENCODE_ZEN_HOSTS = {"opencode.ai", "www.opencode.ai"}
8585

8686

87+
def _normalize_openai_max_tokens(body: dict[str, Any]) -> None:
88+
"""Rename the legacy ``max_tokens`` to ``max_completion_tokens`` in-place.
89+
90+
GPT-5 / o-series chat models reject ``max_tokens`` and require
91+
``max_completion_tokens``; gpt-4o/4.1 accept the latter too. So translating
92+
is a safe, one-way shim for current OpenAI models that lets openai-compatible
93+
clients (opencode, older SDKs) which still send ``max_tokens`` work unchanged.
94+
No-op when there is no ``max_tokens``; keeps an already-set
95+
``max_completion_tokens`` and just drops the rejected legacy key.
96+
"""
97+
if not isinstance(body, dict) or "max_tokens" not in body:
98+
return
99+
legacy = body.get("max_tokens")
100+
if legacy is not None and body.get("max_completion_tokens") is None:
101+
body["max_completion_tokens"] = legacy
102+
body.pop("max_tokens", None)
103+
104+
87105
def _header_get(headers: dict[str, str], name: str) -> str | None:
88106
"""Case-insensitive header lookup for plain dicts."""
89107
lowered = name.lower()
@@ -2593,6 +2611,15 @@ async def handle_openai_chat(
25932611
optimized_tokens = tokenizer.count_messages(body["messages"])
25942612
tokens_saved = original_tokens - optimized_tokens
25952613

2614+
# Compatibility shim: GPT-5 / o-series chat models REJECT the legacy
2615+
# `max_tokens` ("Unsupported parameter … Use 'max_completion_tokens'
2616+
# instead"); gpt-4o/4.1 accept `max_completion_tokens` too. openai-
2617+
# compatible clients (opencode, older SDKs) still send `max_tokens`, so
2618+
# translate it here — the proxy already owns the outbound body — and
2619+
# those requests work unchanged. No-op when the caller already set
2620+
# `max_completion_tokens`.
2621+
_normalize_openai_max_tokens(body)
2622+
25962623
# Route through LiteLLM/any-llm backend if configured
25972624
if self.anthropic_backend is not None:
25982625
try:
Lines changed: 51 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,51 @@
1+
"""OpenAI chat-path compatibility shim: max_tokens -> max_completion_tokens.
2+
3+
GPT-5 / o-series chat models reject the legacy ``max_tokens`` and require
4+
``max_completion_tokens`` ("Unsupported parameter: 'max_tokens' is not supported
5+
with this model. Use 'max_completion_tokens' instead."). openai-compatible
6+
clients (opencode, older SDKs) still send ``max_tokens``, so the proxy — which
7+
already owns the outbound request body — translates it.
8+
"""
9+
10+
from __future__ import annotations
11+
12+
from headroom.proxy.handlers.openai import _normalize_openai_max_tokens
13+
14+
15+
def test_renames_legacy_max_tokens():
16+
body = {"model": "gpt-5.3-chat-latest", "max_tokens": 256, "messages": []}
17+
_normalize_openai_max_tokens(body)
18+
assert "max_tokens" not in body
19+
assert body["max_completion_tokens"] == 256
20+
21+
22+
def test_preserves_existing_max_completion_tokens_and_drops_legacy():
23+
body = {"max_tokens": 256, "max_completion_tokens": 100}
24+
_normalize_openai_max_tokens(body)
25+
assert "max_tokens" not in body
26+
assert body["max_completion_tokens"] == 100 # explicit value wins
27+
28+
29+
def test_noop_when_only_max_completion_tokens():
30+
body = {"max_completion_tokens": 128}
31+
_normalize_openai_max_tokens(body)
32+
assert body == {"max_completion_tokens": 128}
33+
34+
35+
def test_noop_when_neither_present():
36+
body = {"model": "gpt-4o", "messages": []}
37+
_normalize_openai_max_tokens(body)
38+
assert "max_completion_tokens" not in body
39+
assert "max_tokens" not in body
40+
41+
42+
def test_null_max_tokens_is_dropped_without_setting_completion():
43+
body = {"max_tokens": None}
44+
_normalize_openai_max_tokens(body)
45+
assert "max_tokens" not in body
46+
assert body.get("max_completion_tokens") is None
47+
48+
49+
def test_non_dict_is_safe():
50+
_normalize_openai_max_tokens(None) # type: ignore[arg-type]
51+
_normalize_openai_max_tokens("nope") # type: ignore[arg-type]

0 commit comments

Comments
 (0)