Skip to content

Commit 5194bdc

Browse files
fix(content-detector): detect and compress space-separated JSON objects (headroomlabs-ai#1742)
## Description Headroom's `detect_content_type()` only recognizes content starting with `[` as a `JSON array. Many web search tools (SerpAPI, Tavily, custom backends) return space-separated JSON objects instead of a real array like follows ```json {"title": "Result 1", "url": "..."} {"title": "Result 2", "url": "..."} {"title": "Result 3", "url": "..."} ``` That shape is detected as `PLAIN_TEXT` (confidence 0.5), so SmartCrusher never processes it and web-search results compress 0%. Closes headroomlabs-ai#1741 ## Type of Change - [x] Bug fix (non-breaking change that fixes an issue) - [ ] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [x] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - `content_detector.py`: `_try_detect_json` now recognizes a run of ≥2 whitespace-separated (space- or newline-separated) JSON objects and returns `JSON_ARRAY` with `metadata["concatenated"] = True`. The router already falls back to the Python regex detector when the native detector returns `PLAIN_TEXT` (`content_router.py`), so this fixes routing on the default backend too. - `content_detector.py`: added `normalize_concatenated_json()` (and a `_decode_concatenated_json()` helper) that rewrites the space-separated shape into a canonical `[{…}, {…}]` array string. - `smart_crusher.py`: `SmartCrusher.crush()` normalizes concatenated JSON to a real array before handing it to the Rust crusher, so it actually compresses. - The change is deliberately conservative: a single object stays unclaimed (`_try_detect_json('{"id": 1}')` → `None`), and any non-JSON token between objects disqualifies the run. Existing `[`-array detection is unchanged. - Added tests and a CHANGELOG entry. ## Testing - [x] Unit tests pass (`pytest`) — affected suites (full suite has network-dependent ML tests that can't run offline; see note) - [x] Linting passes (`ruff check .`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added for new functionality - [x] Manual testing performed ### Test Output ```text $ ruff check . All checks passed! $ pytest tests/test_transforms_content_detection.py -q ............ [100%] 12 passed $ pytest tests/test_transforms_content_router.py \ tests/test_smart_crusher_toin_attachment.py \ tests/test_transforms_tabular.py -q 96 passed, 2 skipped # + SmartCrusher passthrough tests in test_text_compressors.py: 2 passed ``` ## Real Behavior Proof - Environment: macOS 26.5, Python 3.12.11, editable source build (`uv pip install -e .`) with the Rust `_core` compiled locally; default detection backend (native Rust → Python-regex fallback on PLAIN_TEXT). - Exact command / steps: ran a 100-object space-separated `web_search` payload through `detect_content_type()` and `ContentRouter().compress()`, before and after the patch (repro below). - Observed result: detection flips `PLAIN_TEXT` (conf 0.5) → `JSON_ARRAY` (conf 1.0) and SmartCrusher compression goes from 0.0% to 34.2% (10369 → 6819 bytes) on the identical payload. - Not tested: the native Rust *detector* path in isolation (the fix relies on the existing documented Python-regex fallback for `PLAIN_TEXT`); separators other than whitespace (comma-separated-without-brackets is intentionally not claimed). Before: ``` detected : ContentType.PLAIN_TEXT conf 0.5 strategy : CompressionStrategy.SMART_CRUSHER orig bytes: 10369 comp bytes: 10369 reduction : 0.0% ``` After: ``` detected : ContentType.JSON_ARRAY conf 1.0 strategy : CompressionStrategy.SMART_CRUSHER orig bytes: 10369 comp bytes: 6819 reduction : 34.2% ``` ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation (CHANGELOG) - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [x] I have updated the CHANGELOG.md if applicable ## Additional Notes Co-authored-by: JD Davis <mxjerrett@gmail.com>
1 parent 46d5d68 commit 5194bdc

3 files changed

Lines changed: 112 additions & 8 deletions

File tree

‎headroom/transforms/content_detector.py‎

Lines changed: 61 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -181,16 +181,57 @@ def detect_content_type(content: str) -> DetectionResult:
181181
return DetectionResult(ContentType.PLAIN_TEXT, 0.5, {})
182182

183183

184+
def _decode_concatenated_json(content: str) -> list | None:
185+
"""Decode a run of whitespace-separated top-level JSON values.
186+
187+
Web search tools (SerpAPI, Tavily, custom backends) commonly emit
188+
back-to-back JSON objects separated only by whitespace rather than a real
189+
array: ``{"title": ...} {"title": ...} {"title": ...}``. Returns the list
190+
of decoded values, or None if the text isn't a clean run of JSON values
191+
separated only by whitespace.
192+
"""
193+
decoder = json.JSONDecoder()
194+
idx, length = 0, len(content)
195+
items: list = []
196+
while idx < length:
197+
while idx < length and content[idx].isspace():
198+
idx += 1
199+
if idx >= length:
200+
break
201+
try:
202+
value, idx = decoder.raw_decode(content, idx)
203+
except ValueError:
204+
return None
205+
items.append(value)
206+
return items or None
207+
208+
209+
def normalize_concatenated_json(content: str) -> str | None:
210+
"""Convert whitespace-separated JSON objects into a canonical JSON array.
211+
212+
SmartCrusher only compresses JSON arrays, so this rewrites the
213+
space-separated web_search shape (``{...} {...} {...}``) into
214+
``[{...}, {...}, {...}]``. Returns None unless the content is two or more
215+
whitespace-separated JSON objects.
216+
"""
217+
stripped = content.strip()
218+
if not stripped.startswith("{"):
219+
return None
220+
items = _decode_concatenated_json(stripped)
221+
if items and len(items) >= 2 and all(isinstance(item, dict) for item in items):
222+
return json.dumps(items)
223+
return None
224+
225+
184226
def _try_detect_json(content: str) -> DetectionResult | None:
185227
"""Try to detect JSON array content."""
186228
content = content.strip()
187229

188-
# Quick check: must start with [ for array
189-
if not content.startswith("["):
190-
return None
191-
192-
try:
193-
parsed = json.loads(content)
230+
if content.startswith("["):
231+
try:
232+
parsed = json.loads(content)
233+
except json.JSONDecodeError:
234+
return None
194235
if isinstance(parsed, list):
195236
# Check if it's a list of dicts (SmartCrusher compatible)
196237
if parsed and all(isinstance(item, dict) for item in parsed):
@@ -205,8 +246,20 @@ def _try_detect_json(content: str) -> DetectionResult | None:
205246
0.8,
206247
{"item_count": len(parsed), "is_dict_array": False},
207248
)
208-
except json.JSONDecodeError:
209-
pass
249+
return None
250+
251+
# Space-separated JSON objects (typical web_search output) aren't a valid
252+
# array, so they'd fall through to PLAIN_TEXT and skip SmartCrusher at 0%
253+
# compression. SmartCrusher normalizes this shape to a real array before
254+
# crushing (#1741).
255+
if content.startswith("{"):
256+
items = _decode_concatenated_json(content)
257+
if items and len(items) >= 2 and all(isinstance(item, dict) for item in items):
258+
return DetectionResult(
259+
ContentType.JSON_ARRAY,
260+
1.0,
261+
{"item_count": len(items), "is_dict_array": True, "concatenated": True},
262+
)
210263

211264
return None
212265

‎headroom/transforms/smart_crusher.py‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -54,6 +54,7 @@
5454
from ..tokenizer import Tokenizer
5555
from ..utils import compute_short_hash, create_tool_digest_marker, deep_copy_messages
5656
from .base import Transform
57+
from .content_detector import normalize_concatenated_json
5758

5859
logger = logging.getLogger(__name__)
5960

@@ -446,6 +447,13 @@ def crush(
446447
opaque-blob offload) leaves the content uncompacted instead.
447448
`None` (default) uses the instance's configured value.
448449
"""
450+
# Web search tools often return space-separated JSON objects
451+
# (``{...} {...} {...}``) rather than a real array. The Rust crusher
452+
# only compresses JSON arrays, so normalize that shape first —
453+
# otherwise it passes through at 0% compression (#1741).
454+
normalized = normalize_concatenated_json(content)
455+
if normalized is not None:
456+
content = normalized
449457
rust = (
450458
self._rust
451459
if lossless_only is None or bool(lossless_only) == self._lossless_only

‎tests/test_transforms_content_detection.py‎

Lines changed: 43 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,7 @@
11
from __future__ import annotations
22

3+
import json
4+
35
from headroom.transforms.content_detector import (
46
ContentType,
57
_try_detect_code,
@@ -10,6 +12,7 @@
1012
_try_detect_search,
1113
detect_content_type,
1214
is_json_array_of_dicts,
15+
normalize_concatenated_json,
1316
)
1417
from headroom.transforms.error_detection import (
1518
ERROR_INDICATOR_KEYWORDS,
@@ -59,6 +62,46 @@ def test_json_detection_distinguishes_dict_arrays_and_other_lists() -> None:
5962
assert is_json_array_of_dicts('["value"]') is False
6063

6164

65+
def test_space_separated_json_objects_detected_as_array() -> None:
66+
# Typical web_search output: back-to-back JSON objects, no array brackets.
67+
content = " ".join(
68+
json.dumps({"title": f"Result {i}", "url": f"http://example.com/{i}"}) for i in range(3)
69+
)
70+
result = _try_detect_json(content)
71+
assert result is not None
72+
assert result.content_type is ContentType.JSON_ARRAY
73+
assert result.confidence == 1.0
74+
assert result.metadata == {"item_count": 3, "is_dict_array": True, "concatenated": True}
75+
76+
# Reaches the same verdict through the top-level detector (not PLAIN_TEXT).
77+
assert detect_content_type(content).content_type is ContentType.JSON_ARRAY
78+
assert is_json_array_of_dicts(content) is True
79+
80+
# Newline separation is just as common and must also be recognized.
81+
newline_sep = "\n".join(json.dumps({"id": i, "snippet": "x"}) for i in range(2))
82+
assert _try_detect_json(newline_sep).content_type is ContentType.JSON_ARRAY
83+
84+
85+
def test_space_separated_json_detection_is_conservative() -> None:
86+
# A single object is not an array — must not be claimed.
87+
assert _try_detect_json('{"id": 1}') is None
88+
# Objects interleaved with prose are not clean concatenated JSON.
89+
assert _try_detect_json('{"id": 1} then some prose {"id": 2}') is None
90+
# Scalars/strings between objects disqualify the run of dicts.
91+
assert _try_detect_json('{"id": 1} "loose string"') is None
92+
93+
94+
def test_normalize_concatenated_json_roundtrips_to_array() -> None:
95+
content = '{"a": 1} {"b": 2}'
96+
normalized = normalize_concatenated_json(content)
97+
assert normalized is not None
98+
assert json.loads(normalized) == [{"a": 1}, {"b": 2}]
99+
100+
# Already-valid arrays and single objects are left for the caller as-is.
101+
assert normalize_concatenated_json('[{"a": 1}]') is None
102+
assert normalize_concatenated_json('{"a": 1}') is None
103+
104+
62105
def test_diff_detection_tracks_headers_and_changes() -> None:
63106
diff = "\n".join(
64107
[

0 commit comments

Comments
 (0)