Add mohit67890/imajev-4b DecisionBench results (1.0 and phase-3 revisions) - #68
Conversation
|
The DecisionBench result comparisonSubmitted results are compared with the pinned Jev ( Overall
Pinned reference baselines
Per primitive
Per family
|
…557982671a7d) and attach row-level artifacts to both records
9844008 to
52e6d82
Compare
|
Filed the evaluation request through the official tracker so this can be run on your side if that is easier than reviewing a fork PR: Hanno-Labs/decision-bench#47. Both records here are unchanged. |
DecisionBench result comparisonSubmitted results are compared with the pinned Jev ( Overall| Model | Revision | Tags | Primary acc | Δ Jev | Δ Luna | Supported acc | Coverage | ECE | NLL | Errors | Unsupported | Pinned reference baselines
Per primitive| Model | Tags | View | Acc | Acc Δ Jev | Acc Δ Luna | ECE | ECE Δ Jev | ECE Δ Luna | NLL | NLL Δ Jev | NLL Δ Luna | Rows | Per family| Model | Tags | View | Acc | Acc Δ Jev | Acc Δ Luna | ECE | ECE Δ Jev | ECE Δ Luna | NLL | NLL Δ Jev | NLL Δ Luna | Rows | |
Hi! This adds DecisionBench 1.0 results for
mohit67890/imajev-4b: the 1.0 release at revision712891d1192c6491441a19bdea254d8a6e1048d0(the original submission, unchanged) and the new phase-3 adapter at revisionc9e5f132465da85d31735ec502d5557982671a7d. Thanks for building the harness; both runs used it without any changes.712891d1)c9e5f132)validate_resultsandpytestpass locally for both records.What changed in the phase-3 adapter
--max-input-tokens 65536) so the 255-candidate rows render fully; nothing is truncated (the harness'sinput_truncation_policyis unchanged).--concurrency 32 --max-candidates 255, 4 option rotations averaged,calibration-rot4.json. Server code:github.com/mohit67890/imajev@9b66a662dd52bbd5bfe0febf3792359f24db606f,scripts/playground/server.py --backend torch.The data-integrity statement from the first submission holds for this adapter too: DecisionBench was never used to train, select or calibrate it; overlap with upstream public datasets (e.g. banking77) is disclosed in the repo's
results/benchmarks/decisionbench/contamination-check.md.The "unknown is always an option" rationale below is kept from the original submission; it still describes why one readout code is reserved.
Why 537 rows are unsupported: "unknown" is always an option
Every unsupported row is a CanonicalEntity (or similar) row with exactly 255 candidates. Our model tops out at 254, and that's deliberate.
imajev picks its answer through a readout with 255 slots. One of them is permanently set aside for "unknown". That leaves 254 for real options, so a row with 255 candidates can't be scored without dropping one. We'd rather the harness record it as unsupported than quietly truncate the list. It counts as a miss, which is fair.
Why spend a slot on "unknown" at all? Because in the business settings we built this for, "I can't tell" is often the most useful answer the model can give:
On DecisionBench every row has a gold answer among the candidates, so "unknown" never wins a point here. Most of the time the model just doesn't use it. The only cost on this benchmark is those 537 rows with 255 candidates, and we're comfortable with that trade.
Model and serving
Qwen/Qwen3.5-4B@851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, plus the trained 255-code readout. The adapter file isadapter_model.safetensors, sha256d8d328f85dcda4459a2331c47aba2a89ec808210e712248afdb69e38b9e0c25d. That's 4,690,992,128 parameters in total, counting the base model's vision tower, which isn't used here.decision-model: it's trained to answernoul,choiceandscorestraight from a readout at the decision position. There's no text generation.calibration.jsonthat ships with the model. That temperature was fitted on 150 of our own items, never on DecisionBench.--max-input-tokens 32768. Nothing is truncated.How it was run
47ea5a479e35fd5ac7fde5c72103143071d7d93f,decision-bench run-system-one-http, unmodified, with--max-candidates 254 --concurrency 3.github.com/mohit67890/imajev@4e623a8a51aa319405a0f32e6b73a8ce73416800. Command:scripts/playground/server.py --backend torch --rotations 4 --calibration <adapter>/calibration.json --max-input-tokens 32768. Three processes behind nginx on a single H100 80GB.pod_decisionbench.sh.raw.jsonl(377 MB, sha256 in the record) is available if you want it.One small thing:
run-system-one-httpwrites XOR'sadapterandprobability_sourcestrings intosummary.jsonwhatever the endpoint. The record declaresimajev-serving-systemone-v1andtrained_readout_4_rotation_mean_temperature_calibrated_v1, which describe this model. We passed our own values for--model-repo,--model-revision,--serving-bundle-sha256(the adapter sha256) and--inference-image(the server command above).Data integrity
To be clear: DecisionBench was not used in any way to build or tune imajev. No rows, task files or outputs from it went into training, checkpoint selection or calibration. The dataset first reached our machines after our last training run had finished, and nothing in our pipeline references it.
We did find overlap, but only because some DecisionBench tasks are built from public datasets that were also in our training mix. It's a coincidence of shared public sources, not a leak. We checked every row against our training manifests anyway, with item-level joins plus 13-gram matching. The full report has the details:
For what it's worth, the overlap doesn't seem to have helped: we score 87.1% on RouteFinancial and 96.7% on RouteGeneralAssistant, so the banking77 task is actually the weaker of the two. The MuSiQue tasks share some Wikipedia passages with public QA sets we trained on, but none of their questions appear in our data. The other 36 tasks show no overlap.
Happy to answer questions or re-run anything. Thanks!
🤖 Generated with Claude Code
Row-level artifacts for both runs (immutable,
artifact.uriin each record): https://huggingface.co/datasets/mohit67890/imajev-decisionbench-runs/tree/d28a139cf197034e185132c9ca97be3f4593be51