Skip to content

Commit 188e382

Browse files
authored
fix(dashboard): price proxy savings without litellm (#1728)
## Description The dashboard's main `Proxy $ Saved` tile can stay at `$0` on Python 3.14 because the durable proxy savings tracker records `0.0` whenever LiteLLM is unavailable or cannot price a model. The token counters keep moving, but `proxy_savings.json` stores zero-dollar `compression_savings_usd` and `total_input_cost_usd` values for new entries, so `/stats` and the dashboard read a permanent zero for those rows. This fixes the proxy savings pricing authority so positive token deltas use LiteLLM list pricing when available and fall back to the existing Headroom savings fallback when exact pricing is unavailable. Existing historical rows keep their stored write-time values; this changes new savings entries going forward. Closes #1718. ## Type of Change - [x] Bug fix (non-breaking change that fixes an issue) - [ ] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - Added `DEFAULT_FALLBACK_INPUT_COST_PER_TOKEN = 3.0 / 1_000_000` constant to `headroom/proxy/savings_tracker.py`. - Fixed `_estimate_compression_savings_usd()`: removed the early `litellm is None` zero-return; changed missing-pricing path from `return 0.0` to `raise RuntimeError`; fallback `except` now returns `tokens_saved * DEFAULT_FALLBACK_INPUT_COST_PER_TOKEN` instead of `0.0`. - Fixed `_estimate_input_cost_usd()`: moved `use_breakdown` computation before the `litellm is None` guard; introduced `chargeable_tokens` which equals the breakdown sum when a breakdown exists, or `input_tokens` otherwise; both the `litellm is None` path and the `except Exception` path now use `chargeable_tokens` to avoid double-counting when breakdown tokens and `input_tokens` are both provided; exact LiteLLM cache metadata remains authoritative when present. - Added focused regression coverage in `tests/test_proxy_savings_history.py` for the LiteLLM-unavailable path, exact-price preservation, and the historical no-backfill boundary. ## Testing - [x] Unit tests pass (`uv run pytest tests/test_proxy_savings_history.py tests/test_savings_ledger.py -q`) - [x] Linting passes (`uv run ruff check headroom/proxy/savings_tracker.py tests/test_proxy_savings_history.py tests/test_savings_ledger.py`) - [ ] Type checking passes (`uv run mypy headroom`) - [x] New tests added for new functionality when applicable - [ ] Manual testing performed ### Test Output ```text Pytest command: uv run pytest tests/test_proxy_savings_history.py tests/test_savings_ledger.py -q Run through: conhost --headless cmd /v:on /c ============================= test session starts ============================= platform win32 -- Python 3.12.13, pytest-9.0.3, pluggy-1.6.0 rootdir: D:\Repos\headroom-pr-1718-fallback-savings-cost-zero configfile: pyproject.toml plugins: anyio-4.12.1, langsmith-0.9.3, asyncio-1.3.0, cov-7.0.0 asyncio: mode=Mode.AUTO, debug=False, asyncio_default_fixture_loop_scope=None, asyncio_default_test_loop_scope=function collected 37 items tests\test_proxy_savings_history.py ...................... [ 59%] tests\test_savings_ledger.py ............ss. [100%] ============================== warnings summary =============================== tests/test_savings_ledger.py::test_proxy_record_request_appends_ledger_event D:\Repos\headroom-pr-1718-fallback-savings-cost-zero\.venv\Lib\site-packages\fastapi\testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead. from starlette.testclient import TestClient as TestClient # noqa -- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html ================== 35 passed, 2 skipped, 1 warning in 16.47s ================== Ruff command: uv run ruff check headroom/proxy/savings_tracker.py tests/test_proxy_savings_history.py tests/test_savings_ledger.py Run through: conhost --headless cmd /v:on /c All checks passed! ``` ## Real Behavior Proof - Environment: Python proxy savings tracker with LiteLLM forced unavailable (`LITELLM_AVAILABLE=False`, `litellm=None`), using a temporary `proxy_savings.json`. - Exact command / steps: run `uv run pytest tests/test_proxy_savings_history.py tests/test_savings_ledger.py -q`, then inspect `test_fallback_request_pricing_stays_nonzero_with_litellm_unavailable_and_preserves_historic_zeros` and `test_fallback_input_cost_uses_breakdown_sum_not_input_tokens_when_litellm_unavailable`, which load a pre-existing file with zero-dollar historical rows, call `record_request()` with LiteLLM unavailable, and call `_estimate_input_cost_usd()` with both `input_tokens` and a nonzero breakdown. - Observed result: new lifetime, display-session, project, and history entries receive nonzero fallback-priced dollar values while the original zero-dollar history row remains unchanged, and the fallback input-cost path prices only the breakdown sum instead of `input_tokens + breakdown_sum`. - `test_litellm_resolution_and_savings_estimation_fallbacks` verifies that `_estimate_compression_savings_usd` and `_estimate_input_cost_usd` return fallback amounts (not `0.0`) for all three paths: LiteLLM available but metadata missing, LiteLLM available but pricing lookup raises, and `LITELLM_AVAILABLE=False`. - `test_input_cost_counts_cache_reads_when_uncached_input_is_zero` verifies that a fully prefix-cached request (`input_tokens=0, cache_read_tokens=1000`) prices the cache reads at the provider cache rate, not zero. - `test_fallback_input_cost_uses_breakdown_sum_not_input_tokens_when_litellm_unavailable` verifies that when LiteLLM is unavailable and both `input_tokens` and a nonzero cache breakdown are supplied, the fallback prices only the breakdown sum and not `input_tokens + breakdown_sum`, preventing double-counting. - `tests/test_savings_ledger.py` still passes locally, proving the sibling ledger consumer stays compatible with the helper fallback change. - Not tested: live provider traffic and historical backfill. Existing zero-dollar rows remain stored as they were written. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes ## Additional Notes `CHANGELOG.md` is unchanged because changelog generation is release-managed. The subscription contribution panel still has a separate USD wiring mismatch; this PR fixes the dashboard-facing `proxy_savings.json` path named in the latest issue follow-up and keeps historical backfill out of scope.
1 parent 728b330 commit 188e382

2 files changed

Lines changed: 169 additions & 18 deletions

File tree

‎headroom/proxy/savings_tracker.py‎

Lines changed: 24 additions & 13 deletions
Original file line numberDiff line numberDiff line change
@@ -34,6 +34,7 @@
3434
DEFAULT_MAX_HISTORY_AGE_DAYS = 365
3535
DEFAULT_MAX_RESPONSE_HISTORY_POINTS = 500
3636
DEFAULT_DISPLAY_SESSION_INACTIVITY_MINUTES = 60
37+
DEFAULT_FALLBACK_INPUT_COST_PER_TOKEN = 3.0 / 1_000_000
3738

3839
LITELLM_AVAILABLE = importlib.util.find_spec("litellm") is not None
3940
litellm: Any | None = None
@@ -189,18 +190,20 @@ def _resolve_litellm_model(model: str) -> str:
189190
def _estimate_compression_savings_usd(model: str, tokens_saved: int) -> float:
190191
"""Estimate compression savings in USD from saved input tokens."""
191192
litellm = _get_litellm_module()
192-
if tokens_saved <= 0 or litellm is None:
193+
if tokens_saved <= 0:
193194
return 0.0
195+
if litellm is None:
196+
return float(tokens_saved) * float(DEFAULT_FALLBACK_INPUT_COST_PER_TOKEN)
194197

195198
try:
196199
resolved = _resolve_litellm_model(model)
197200
info = litellm.model_cost.get(resolved, {})
198201
input_cost_per_token = info.get("input_cost_per_token")
199202
if not input_cost_per_token:
200-
return 0.0
203+
raise RuntimeError("input cost unavailable")
201204
return float(tokens_saved) * float(input_cost_per_token)
202205
except Exception:
203-
return 0.0
206+
return float(tokens_saved) * float(DEFAULT_FALLBACK_INPUT_COST_PER_TOKEN)
204207

205208

206209
def _estimate_input_cost_usd(
@@ -220,23 +223,31 @@ def _estimate_input_cost_usd(
220223
cache_read = _coerce_int(cache_read_tokens)
221224
cache_write = _coerce_int(cache_write_tokens)
222225
uncached = _coerce_int(uncached_input_tokens)
223-
litellm = _get_litellm_module()
224-
# Gate on tokens actually sent. Providers like Anthropic report cache
225-
# reads/writes separately from `input_tokens` (the uncached portion), so a
226-
# fully prefix-cached request has input_tokens == 0 while cache_read > 0.
227-
# Bailing on `input_tokens <= 0` alone dropped the real cache-read cost,
228-
# leaving days with compression savings but zero recorded spend.
229-
if total_input_tokens + cache_read + cache_write + uncached <= 0 or litellm is None:
226+
227+
# Prefer the breakdown when callers supply segmented token counts.
228+
# Never add `input_tokens` on top of the breakdown to avoid double-counting.
229+
use_breakdown = (cache_read + cache_write + uncached) > 0
230+
chargeable_tokens = (
231+
(cache_read + cache_write + uncached) if use_breakdown else total_input_tokens
232+
)
233+
if chargeable_tokens <= 0:
230234
return 0.0
231235

236+
litellm = _get_litellm_module()
237+
# Keep exact provider pricing authoritative when available.
238+
# `litellm` can be present but lack an entry for the resolved model,
239+
# in which case we fall back to a blended rate instead of zeroing usage.
240+
if litellm is None:
241+
return float(chargeable_tokens) * float(DEFAULT_FALLBACK_INPUT_COST_PER_TOKEN)
242+
232243
try:
233244
resolved = _resolve_litellm_model(model)
234245
info = litellm.model_cost.get(resolved, {})
235246
input_cost_per_token = info.get("input_cost_per_token")
236247
if not input_cost_per_token:
237-
return 0.0
248+
raise RuntimeError("input cost unavailable")
238249

239-
if cache_read + cache_write + uncached > 0:
250+
if use_breakdown:
240251
cache_read_cost = info.get(
241252
"cache_read_input_token_cost",
242253
input_cost_per_token,
@@ -253,7 +264,7 @@ def _estimate_input_cost_usd(
253264

254265
return float(total_input_tokens) * float(input_cost_per_token)
255266
except Exception:
256-
return 0.0
267+
return float(chargeable_tokens) * float(DEFAULT_FALLBACK_INPUT_COST_PER_TOKEN)
257268

258269

259270
def _normalize_history_entry(entry: Any) -> dict[str, Any] | None:

‎tests/test_proxy_savings_history.py‎

Lines changed: 145 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -327,20 +327,128 @@ def fake_cost_per_token(*, model, prompt_tokens, completion_tokens):
327327
) == pytest.approx(0.2)
328328

329329
fake_litellm.model_cost = {}
330-
assert savings_tracker_module._estimate_compression_savings_usd("gpt-4o", 100) == 0.0
331-
assert savings_tracker_module._estimate_input_cost_usd("gpt-4o", 100) == 0.0
330+
assert savings_tracker_module._estimate_compression_savings_usd("gpt-4o", 100) == pytest.approx(
331+
100 * savings_tracker_module.DEFAULT_FALLBACK_INPUT_COST_PER_TOKEN
332+
)
333+
assert savings_tracker_module._estimate_input_cost_usd("gpt-4o", 100) == pytest.approx(
334+
100 * savings_tracker_module.DEFAULT_FALLBACK_INPUT_COST_PER_TOKEN
335+
)
332336

333337
monkeypatch.setattr(
334338
fake_litellm,
335339
"cost_per_token",
336340
lambda **kwargs: (_ for _ in ()).throw(RuntimeError("boom")),
337341
)
338342
assert savings_tracker_module._resolve_litellm_model("mystery-model") == "mystery-model"
339-
assert savings_tracker_module._estimate_compression_savings_usd("mystery-model", 100) == 0.0
343+
assert savings_tracker_module._estimate_compression_savings_usd(
344+
"mystery-model", 100
345+
) == pytest.approx(100 * savings_tracker_module.DEFAULT_FALLBACK_INPUT_COST_PER_TOKEN)
346+
assert savings_tracker_module._estimate_input_cost_usd("mystery-model", 100) == pytest.approx(
347+
100 * savings_tracker_module.DEFAULT_FALLBACK_INPUT_COST_PER_TOKEN
348+
)
349+
# Explicitly force the unavailable path for the whole tracker.
350+
monkeypatch.setattr(savings_tracker_module, "LITELLM_AVAILABLE", False)
351+
assert savings_tracker_module._estimate_compression_savings_usd("gpt-4o", 100) == pytest.approx(
352+
100 * savings_tracker_module.DEFAULT_FALLBACK_INPUT_COST_PER_TOKEN
353+
)
354+
assert savings_tracker_module._estimate_input_cost_usd("gpt-4o", 100) == pytest.approx(
355+
100 * savings_tracker_module.DEFAULT_FALLBACK_INPUT_COST_PER_TOKEN
356+
)
357+
358+
359+
def test_fallback_request_pricing_stays_nonzero_with_litellm_unavailable_and_preserves_historic_zeros(
360+
tmp_path, monkeypatch
361+
):
362+
# Legacy proxy_savings rows can legitimately store zero-dollar values.
363+
savings_path = tmp_path / "proxy_savings.json"
364+
savings_path.write_text(
365+
json.dumps(
366+
{
367+
"schema_version": 3,
368+
"lifetime": {
369+
"requests": 1,
370+
"tokens_saved": 10,
371+
"compression_savings_usd": 0.0,
372+
"total_input_tokens": 120,
373+
"total_input_cost_usd": 0.0,
374+
},
375+
"display_session": {},
376+
"history": [
377+
{
378+
"timestamp": "2026-03-27T09:00:00Z",
379+
"provider": "openai",
380+
"model": "gpt-4o",
381+
"total_tokens_saved": 10,
382+
"compression_savings_usd": 0.0,
383+
"total_input_tokens": 120,
384+
"total_input_cost_usd": 0.0,
385+
}
386+
],
387+
"projects": {
388+
"fallback-demo": {
389+
"requests": 1,
390+
"tokens_saved": 10,
391+
"compression_savings_usd": 0.0,
392+
"total_input_tokens": 120,
393+
"total_input_cost_usd": 0.0,
394+
"last_activity_at": "2026-03-27T09:00:00Z",
395+
}
396+
},
397+
}
398+
),
399+
encoding="utf-8",
400+
)
401+
402+
tracker = SavingsTracker(path=str(savings_path))
403+
initial_snapshot = tracker.snapshot()
404+
assert initial_snapshot["lifetime"]["compression_savings_usd"] == 0.0
405+
assert initial_snapshot["display_session"]["compression_savings_usd"] == 0.0
406+
assert initial_snapshot["projects"]["fallback-demo"]["compression_savings_usd"] == 0.0
407+
assert initial_snapshot["history"][-1]["compression_savings_usd"] == 0.0
340408

341409
monkeypatch.setattr(savings_tracker_module, "LITELLM_AVAILABLE", False)
342-
assert savings_tracker_module._estimate_compression_savings_usd("gpt-4o", 100) == 0.0
343-
assert savings_tracker_module._estimate_input_cost_usd("gpt-4o", 100) == 0.0
410+
monkeypatch.setattr(savings_tracker_module, "litellm", None)
411+
assert tracker.record_request(
412+
model="gpt-4o",
413+
input_tokens=100,
414+
tokens_saved=50,
415+
project="fallback-demo",
416+
timestamp="2026-03-27T09:10:00Z",
417+
)
418+
monkeypatch.setattr(
419+
savings_tracker_module,
420+
"_utc_now",
421+
lambda: datetime(2026, 3, 27, 9, 10, 30, tzinfo=timezone.utc),
422+
)
423+
424+
snapshot = tracker.snapshot()
425+
expected_savings_fallback = 50 * savings_tracker_module.DEFAULT_FALLBACK_INPUT_COST_PER_TOKEN
426+
expected_input_fallback = 100 * savings_tracker_module.DEFAULT_FALLBACK_INPUT_COST_PER_TOKEN
427+
assert snapshot["lifetime"]["compression_savings_usd"] == pytest.approx(
428+
expected_savings_fallback
429+
)
430+
assert snapshot["lifetime"]["total_input_cost_usd"] == pytest.approx(expected_input_fallback)
431+
assert snapshot["display_session"]["compression_savings_usd"] == pytest.approx(
432+
expected_savings_fallback
433+
)
434+
assert snapshot["display_session"]["total_input_cost_usd"] == pytest.approx(
435+
expected_input_fallback
436+
)
437+
assert snapshot["projects"]["fallback-demo"]["compression_savings_usd"] == pytest.approx(
438+
expected_savings_fallback
439+
)
440+
assert snapshot["projects"]["fallback-demo"]["total_input_cost_usd"] == pytest.approx(
441+
expected_input_fallback
442+
)
443+
assert snapshot["history"][-1]["compression_savings_usd"] == pytest.approx(
444+
expected_savings_fallback
445+
)
446+
447+
persisted = json.loads(savings_path.read_text(encoding="utf-8"))
448+
assert persisted["history"][0]["compression_savings_usd"] == 0.0
449+
assert persisted["history"][-1]["compression_savings_usd"] == pytest.approx(
450+
expected_savings_fallback
451+
)
344452

345453

346454
def test_input_cost_counts_cache_reads_when_uncached_input_is_zero(monkeypatch):
@@ -374,6 +482,38 @@ def fake_cost_per_token(*, model, prompt_tokens, completion_tokens):
374482
assert cost == pytest.approx(0.3)
375483

376484

485+
def test_fallback_input_cost_uses_breakdown_sum_not_input_tokens_when_litellm_unavailable(
486+
monkeypatch,
487+
):
488+
# Regression: when both `input_tokens` and a nonzero cache breakdown are
489+
# present and LiteLLM is unavailable, the fallback must price only the
490+
# breakdown sum — never input_tokens + breakdown_sum — to avoid
491+
# double-counting the tokens that the breakdown already covers.
492+
monkeypatch.setattr(savings_tracker_module, "LITELLM_AVAILABLE", False)
493+
monkeypatch.setattr(savings_tracker_module, "litellm", None)
494+
495+
input_tokens = 1000
496+
cache_read = 200
497+
cache_write = 100
498+
uncached = 300
499+
breakdown_sum = cache_read + cache_write + uncached # 600
500+
501+
result = savings_tracker_module._estimate_input_cost_usd(
502+
"gpt-4o",
503+
input_tokens,
504+
cache_read_tokens=cache_read,
505+
cache_write_tokens=cache_write,
506+
uncached_input_tokens=uncached,
507+
)
508+
509+
fallback_rate = savings_tracker_module.DEFAULT_FALLBACK_INPUT_COST_PER_TOKEN
510+
expected = breakdown_sum * fallback_rate
511+
double_counted = (input_tokens + breakdown_sum) * fallback_rate
512+
513+
assert result == pytest.approx(expected)
514+
assert result != pytest.approx(double_counted)
515+
516+
377517
def test_display_session_rolls_after_inactivity_and_counts_zero_savings_requests(
378518
tmp_path, monkeypatch
379519
):

0 commit comments

Comments
 (0)