Summary
On Headroom 0.28.0 (also reproduced on 0.25.0), Hermes agent traffic using the anthropic_messages transport through the Z.ai GLM proxy (:8788) does not resolve tool names for compression routing or TOIN learning. Every pattern in ~/.headroom/toin.json is keyed unknown|unknown|<hash>.
0.28.0 boundary (most useful isolation for maintainers): native Hermes tool exclusion works for built-in tools — proxy logs show router:excluded:tool on read_file and terminal, and the model receives those outputs verbatim. MCP tool results still bypass exclusion (e.g. mcp_CursorTaskRegistry_cursor_list_tasks compressed to ccr: markers). TOIN / name-resolution is unchanged — all 112 patterns remain unknown|unknown|<hash> before and after 0.28.0 tests. So the fix is partial: content-router native-tool path vs MCP anthropic_messages tool path are not aligned.
prefix_cache on :8788 shows zero lifetime cache_read_tokens.
Environment
| Item |
Value |
| Headroom |
0.28.0 (headroom-ai[proxy], macOS arm64, Python 3.13 venv) |
| Proxy |
headroom proxy --port 8788 --anthropic-api-url https://api.z.ai/api/anthropic |
| Client |
Hermes Agent CLI (hermes -z ... --provider headroom-zai -m glm-5.2) |
| Transport |
anthropic_messages (not Claude Code CLI) |
| Env |
HEADROOM_EXCLUDE_TOOLS=headroom_retrieve,mcp_HeadroomZai_headroom_retrieve,mcp_Headroom_headroom_retrieve,terminal,execute_code,read_file,search_files |
Reproduction
- Install
headroom-ai[proxy]==0.28.0, start Z.ai proxy on :8788 with exclusions above.
- Run Hermes one-shot against
headroom-zai / glm-5.2 with file + terminal + MCP tools.
- Inspect
~/.headroom/toin.json pattern keys and proxy PERF lines.
Observed (2026-07-03 WS8 session)
0.28.0 vs 0.25.0 — native tools fixed, MCP + TOIN not:
| Path |
0.25.0 |
0.28.0 |
read_file / terminal (native) |
Compressed despite HEADROOM_EXCLUDE_TOOLS |
Excluded — router:excluded:tool in proxy log; verbatim to model |
MCP tools (e.g. cursor_list_tasks) |
Compressed to ccr: |
Still compressed to ccr: |
| TOIN pattern keys |
unknown|unknown (112/112) |
Unchanged unknown|unknown (112/112) |
read_file + terminal (0.28.0 pass): Proxy log shows exclusions honored:
content_router: 5 msgs — 3 excluded (Read/Glob), ...
transforms=router:excluded:tool*3
Agent received first line of SOUL.md verbatim (# Pepper — Jim Vetter's Personal Operating Agent) and terminal integer 234 with no ccr: markers.
MCP (fail): cursor_list_tasks result contained compressed summary:
"summary":"<<ccr:888241e845ce,string,603B>>"
TOIN (fail): 112/112 patterns keyed unknown|unknown|<hash> before and after test; no new named keys.
prefix_cache (fail): /stats on :8788 after test:
"prefix_cache": {
"totals": {
"cache_read_tokens": 0,
"requests": 0,
"hit_rate": 0
}
}
Proxy PERF lines for GLM requests show cache_read=0 cache_write=0.
Expected
- Tool names from Hermes
anthropic_messages tool_use blocks resolve to real names (e.g. read_file, mcp_CursorTaskRegistry_cursor_list_tasks) for TOIN keys and HEADROOM_EXCLUDE_TOOLS — including MCP-prefixed names, not only native read_file/terminal.
- Excluded tools never emit
ccr: markers to the model — same rule for native and MCP tool results.
- Stable prefixes on Z.ai path accumulate nonzero
cache_read when tool results are not compressed away.
Impact
Custom agents (Hermes, Pepper) on metered GLM cannot use Headroom safely: small outputs may pass, but MCP JSON blobs compress and trigger headroom_retrieve doom loops. Workaround in production: bypass Headroom on GLM lane (zai-glm direct to api.z.ai); keep Headroom only on Sonnet escalation (:8787).
Evidence files (local)
- Registry task
fcbca026 incident log (Fable Window Plan)
- Proxy log excerpts:
~/.headroom/logs/proxy.log (2026-07-03 18:14–18:15 PT)
- Test transcripts:
/tmp/ws8_headroom_glm_test.log, /tmp/ws8_headroom_mcp_test.log
- TOIN snapshot:
~/.headroom/toin.json.bak.20260703_pre028
Notes
HEADROOM_EXCLUDE_TOOLS is a supported env var (_parse_exclude_tools in proxy); the failure is upstream of exclusion matching.
- Sonnet lane on
:8787 at 0.28.0 passes basic smoke (headroom-claude / claude-sonnet-5 → OK).
Summary
On Headroom 0.28.0 (also reproduced on 0.25.0), Hermes agent traffic using the
anthropic_messagestransport through the Z.ai GLM proxy (:8788) does not resolve tool names for compression routing or TOIN learning. Every pattern in~/.headroom/toin.jsonis keyedunknown|unknown|<hash>.0.28.0 boundary (most useful isolation for maintainers): native Hermes tool exclusion works for built-in tools — proxy logs show
router:excluded:toolonread_fileandterminal, and the model receives those outputs verbatim. MCP tool results still bypass exclusion (e.g.mcp_CursorTaskRegistry_cursor_list_taskscompressed toccr:markers). TOIN / name-resolution is unchanged — all 112 patterns remainunknown|unknown|<hash>before and after 0.28.0 tests. So the fix is partial: content-router native-tool path vs MCPanthropic_messagestool path are not aligned.prefix_cacheon:8788shows zero lifetimecache_read_tokens.Environment
headroom-ai[proxy], macOS arm64, Python 3.13 venv)headroom proxy --port 8788 --anthropic-api-url https://api.z.ai/api/anthropichermes -z ... --provider headroom-zai -m glm-5.2)anthropic_messages(not Claude Code CLI)HEADROOM_EXCLUDE_TOOLS=headroom_retrieve,mcp_HeadroomZai_headroom_retrieve,mcp_Headroom_headroom_retrieve,terminal,execute_code,read_file,search_filesReproduction
headroom-ai[proxy]==0.28.0, start Z.ai proxy on:8788with exclusions above.headroom-zai/glm-5.2with file + terminal + MCP tools.~/.headroom/toin.jsonpattern keys and proxyPERFlines.Observed (2026-07-03 WS8 session)
0.28.0 vs 0.25.0 — native tools fixed, MCP + TOIN not:
read_file/terminal(native)HEADROOM_EXCLUDE_TOOLSrouter:excluded:toolin proxy log; verbatim to modelcursor_list_tasks)ccr:ccr:unknown|unknown(112/112)unknown|unknown(112/112)read_file + terminal (0.28.0 pass): Proxy log shows exclusions honored:
Agent received first line of
SOUL.mdverbatim (# Pepper — Jim Vetter's Personal Operating Agent) and terminal integer234with noccr:markers.MCP (fail):
cursor_list_tasksresult contained compressed summary:TOIN (fail): 112/112 patterns keyed
unknown|unknown|<hash>before and after test; no new named keys.prefix_cache (fail):
/statson:8788after test:Proxy PERF lines for GLM requests show
cache_read=0 cache_write=0.Expected
anthropic_messagestool_use blocks resolve to real names (e.g.read_file,mcp_CursorTaskRegistry_cursor_list_tasks) for TOIN keys andHEADROOM_EXCLUDE_TOOLS— including MCP-prefixed names, not only nativeread_file/terminal.ccr:markers to the model — same rule for native and MCP tool results.cache_readwhen tool results are not compressed away.Impact
Custom agents (Hermes, Pepper) on metered GLM cannot use Headroom safely: small outputs may pass, but MCP JSON blobs compress and trigger
headroom_retrievedoom loops. Workaround in production: bypass Headroom on GLM lane (zai-glmdirect toapi.z.ai); keep Headroom only on Sonnet escalation (:8787).Evidence files (local)
fcbca026incident log (Fable Window Plan)~/.headroom/logs/proxy.log(2026-07-03 18:14–18:15 PT)/tmp/ws8_headroom_glm_test.log,/tmp/ws8_headroom_mcp_test.log~/.headroom/toin.json.bak.20260703_pre028Notes
HEADROOM_EXCLUDE_TOOLSis a supported env var (_parse_exclude_toolsin proxy); the failure is upstream of exclusion matching.:8787at 0.28.0 passes basic smoke (headroom-claude/claude-sonnet-5→OK).