Skip to content

Add mohit67890/imajev-4b DecisionBench results (1.0 and phase-3 revisions) - #68

Merged
trashhalo merged 3 commits into
Hanno-Labs:mainfrom
mohit67890:results/imajev-4b-decisionbench
Sep 27, 2026
Merged

trashhalo merged 3 commits into
Hanno-Labs:mainfrom
mohit67890:results/imajev-4b-decisionbench

Conversation

@mohit67890

@mohit67890 mohit67890 commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Hi! This adds DecisionBench 1.0 results for mohit67890/imajev-4b: the 1.0 release at revision 712891d1192c6491441a19bdea254d8a6e1048d0 (the original submission, unchanged) and the new phase-3 adapter at revision c9e5f132465da85d31735ec502d5557982671a7d. Thanks for building the harness; both runs used it without any changes.

1.0 (712891d1) phase 3 (c9e5f132)
Primary accuracy 77.5% 79.7%
Accuracy on scored rows 79.3% 79.7%
Coverage 23,363 / 23,900 (537 rows unsupported, 0 errors) 23,900 / 23,900 (0 unsupported, 0 errors)
ECE (15 bins) 0.024 0.069
Mean NLL 0.663 0.694

validate_results and pytest pass locally for both records.

What changed in the phase-3 adapter

  • The readout grew from 255 to 256 codes, so 255 candidates plus an explicit "unknown" now fit; the 537 rows with 255 candidates that were unsupported for 1.0 are scored (95.7% on those rows). The reasoning family went from 68.0% to 80.6% (FinQA 68 → 92, MuSiQue multihop 83 → 97) after a round of training on decisions the 1.0 model got wrong; ordinal scoring 40.3% → 46.2%. Calibration is worse than 1.0's on this suite (ECE 0.069 vs 0.024): the new temperature softens less. Both effects are reported as measured.
  • The serving limit was raised to 65,536 processed tokens (--max-input-tokens 65536) so the 255-candidate rows render fully; nothing is truncated (the harness's input_truncation_policy is unchanged).
  • Run: 8×H100, 32 server processes behind an nginx least-connections proxy, harness --concurrency 32 --max-candidates 255, 4 option rotations averaged, calibration-rot4.json. Server code: github.com/mohit67890/imajev@9b66a662dd52bbd5bfe0febf3792359f24db606f, scripts/playground/server.py --backend torch.

The data-integrity statement from the first submission holds for this adapter too: DecisionBench was never used to train, select or calibrate it; overlap with upstream public datasets (e.g. banking77) is disclosed in the repo's results/benchmarks/decisionbench/contamination-check.md.

The "unknown is always an option" rationale below is kept from the original submission; it still describes why one readout code is reserved.


Why 537 rows are unsupported: "unknown" is always an option

Every unsupported row is a CanonicalEntity (or similar) row with exactly 255 candidates. Our model tops out at 254, and that's deliberate.

imajev picks its answer through a readout with 255 slots. One of them is permanently set aside for "unknown". That leaves 254 for real options, so a row with 255 candidates can't be scored without dropping one. We'd rather the harness record it as unsupported than quietly truncate the list. It counts as a miss, which is fair.

Why spend a slot on "unknown" at all? Because in the business settings we built this for, "I can't tell" is often the most useful answer the model can give:

  • The answer isn't in the input. A refund request with no order number, an invoice with a blank payment field, a product photo that doesn't show the label: all of these happen constantly in real traffic. A model forced to pick from the listed options will pick something, usually with confidence it hasn't earned.
  • A wrong answer costs more than no answer. If a ticket is routed to the wrong queue, a claim auto-approved on a guess, or a listing tagged with the wrong attribute, a person has to find and undo it later. An "unknown" can go straight to a human, which is what most teams already do with uncertain cases.
  • It keeps the other probabilities honest. When the model can put its doubt somewhere, the confidence on the real options means more. We think that's part of why calibration on this run is good (ECE 0.024). An app can set one threshold and trust it.
  • Nobody has to invent a "none of the above" option for every question. Businesses shouldn't need to design their label sets around the model's inability to abstain.

On DecisionBench every row has a gold answer among the candidates, so "unknown" never wins a point here. Most of the time the model just doesn't use it. The only cost on this benchmark is those 537 rows with 255 candidates, and we're comfortable with that trade.

Model and serving

  • Base model: open-weights LoRA (Apache-2.0) on Qwen/Qwen3.5-4B @ 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, plus the trained 255-code readout. The adapter file is adapter_model.safetensors, sha256 d8d328f85dcda4459a2331c47aba2a89ec808210e712248afdb69e38b9e0c25d. That's 4,690,992,128 parameters in total, counting the base model's vision tower, which isn't used here.
  • decision-model: it's trained to answer noul, choice and score straight from a readout at the decision position. There's no text generation.
  • How probabilities are produced:
    1. One forward pass per option order.
    2. Four cyclic option orders, averaged.
    3. Divided by a single temperature from the calibration.json that ships with the model. That temperature was fitted on 150 of our own items, never on DecisionBench.
  • Long inputs: served with --max-input-tokens 32768. Nothing is truncated.

How it was run

  • Harness: DecisionBench @ 47ea5a479e35fd5ac7fde5c72103143071d7d93f, decision-bench run-system-one-http, unmodified, with --max-candidates 254 --concurrency 3.
  • Endpoint: the model's own server, github.com/mohit67890/imajev @ 4e623a8a51aa319405a0f32e6b73a8ce73416800. Command: scripts/playground/server.py --backend torch --rotations 4 --calibration <adapter>/calibration.json --max-input-tokens 32768. Three processes behind nginx on a single H100 80GB.
  • Reproduction script: everything comes from public pinned sources: pod_decisionbench.sh.
  • Run files: the summary, manifest, pod log and per-task table are here. The full raw.jsonl (377 MB, sha256 in the record) is available if you want it.

One small thing: run-system-one-http writes XOR's adapter and probability_source strings into summary.json whatever the endpoint. The record declares imajev-serving-systemone-v1 and trained_readout_4_rotation_mean_temperature_calibrated_v1, which describe this model. We passed our own values for --model-repo, --model-revision, --serving-bundle-sha256 (the adapter sha256) and --inference-image (the server command above).

Data integrity

To be clear: DecisionBench was not used in any way to build or tune imajev. No rows, task files or outputs from it went into training, checkpoint selection or calibration. The dataset first reached our machines after our last training run had finished, and nothing in our pipeline references it.

We did find overlap, but only because some DecisionBench tasks are built from public datasets that were also in our training mix. It's a coincidence of shared public sources, not a leak. We checked every row against our training manifests anyway, with item-level joins plus 13-gram matching. The full report has the details:

DecisionBench task Public dataset we also trained on Rows whose upstream item was in our training data
RouteFinancial banking77 (train) 1,030 of 1,145
RouteGeneralAssistant CLINC150 (train) 137 of 1,077
RelevanceScore Amazon ESCI 54 of 2,222
ContainsThreat / ToxicitySeverity civil_comments (train) 8 / 11

For what it's worth, the overlap doesn't seem to have helped: we score 87.1% on RouteFinancial and 96.7% on RouteGeneralAssistant, so the banking77 task is actually the weaker of the two. The MuSiQue tasks share some Wikipedia passages with public QA sets we trained on, but none of their questions appear in our data. The other 36 tasks show no overlap.

  • I am not aware of DecisionBench evaluation rows being used for training.

Happy to answer questions or re-run anything. Thanks!

🤖 Generated with Claude Code

Row-level artifacts for both runs (immutable, artifact.uri in each record): https://huggingface.co/datasets/mohit67890/imajev-decisionbench-runs/tree/d28a139cf197034e185132c9ca97be3f4593be51

@mohit67890

mohit67890 commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor Author

The compare check can't run on this PR because it comes from a fork (actions/checkout refuses fork code under pull_request_target). To save you a step, here is the output of your own scripts/compare_results.py, run locally against main @ 87f3825. validate-results and pytest also pass locally.

DecisionBench result comparison

Submitted results are compared with the pinned Jev (typesafe/jev-1.13) and Luna (openai/gpt-5.6-luna) records from main.
Accuracy deltas are percentage points; positive is better. For ECE and NLL deltas, negative is better. Per-view accuracy is calculated on supported rows.

Overall

Model Revision Primary acc Δ Jev Δ Luna Supported acc Coverage ECE NLL Errors Unsupported
mohit67890/imajev-4b 712891d1192c6491441a19bdea254d8a6e1048d0 77.55% +5.52 pp +7.65 pp 79.33% 97.75% 0.024 0.663 0 537
Pinned reference baselines
Reference Revision Primary acc Supported acc Coverage ECE NLL Errors Unsupported
Jev (typesafe/jev-1.13) openrouter-service-snapshot-2026-09-21 72.03% 72.03% 100.00% 0.128 2.433 0 0
Luna (openai/gpt-5.6-luna) openrouter-service-snapshot-2026-09-21-reasoning-aware-v1 69.90% 69.93% 99.96% 0.203 3.154 10 0

Per primitive

Model View Acc Acc Δ Jev Acc Δ Luna ECE ECE Δ Jev ECE Δ Luna NLL NLL Δ Jev NLL Δ Luna Rows
mohit67890/imajev-4b binary_classification 86.58% +27.36 pp +2.86 pp 0.092 -0.065 -0.014 0.377 -0.278 -0.147 6388
mohit67890/imajev-4b candidate_selection 80.64% +0.54 pp +13.71 pp 0.012 -0.088 -0.208 0.645 -1.600 -3.570 15277
mohit67890/imajev-4b ordinal_scoring 40.28% -4.83 pp -5.65 pp 0.321 -0.034 -0.138 1.896 -8.974 -1.271 1698

Per family

Model View Acc Acc Δ Jev Acc Δ Luna ECE ECE Δ Jev ECE Δ Luna NLL NLL Δ Jev NLL Δ Luna Rows
mohit67890/imajev-4b action_selection 63.71% -1.57 pp +3.29 pp 0.072 -0.055 -0.040 1.185 -2.093 -0.159 700
mohit67890/imajev-4b bounded_extraction 93.88% -1.75 pp +0.99 pp 0.008 -0.024 -0.051 0.223 -0.484 -1.505 2223
mohit67890/imajev-4b completion_assessment 90.00% -4.00 pp -3.00 pp 0.078 +0.044 +0.039 0.269 -0.519 -0.031 100
mohit67890/imajev-4b content_classification 80.00% -1.00 pp +0.00 pp 0.131 -0.045 -0.036 0.933 -4.670 -0.163 100
mohit67890/imajev-4b content_moderation 100.00% +55.00 pp +1.00 pp 0.061 -0.359 +0.059 0.067 -0.745 +0.012 100
mohit67890/imajev-4b context_management 88.50% +37.50 pp -2.00 pp 0.068 -0.193 -0.006 0.281 -0.482 -0.050 200
mohit67890/imajev-4b document_classification 39.77% +0.77 pp +1.77 pp 0.257 -0.306 -0.279 3.192 -15.123 -12.873 88
mohit67890/imajev-4b document_record_classification 56.10% -0.54 pp +9.49 pp 0.037 -0.163 -0.263 1.300 -3.711 -0.800 2223
mohit67890/imajev-4b document_workflows 81.91% +9.00 pp +8.55 pp 0.186 +0.072 +0.033 0.550 -0.052 -0.233 2222
mohit67890/imajev-4b entity_alignment 100.00% +0.00 pp +18.56 pp 0.137 +0.136 -0.031 0.161 +0.160 -5.481 1971
mohit67890/imajev-4b function_agent_skill_routing 99.13% +0.12 pp +17.96 pp 0.045 +0.041 -0.064 0.078 +0.019 -4.008 1948
mohit67890/imajev-4b guardrails_moderation 62.38% +23.31 pp +0.54 pp 0.212 -0.235 -0.113 1.278 -6.830 -0.937 2222
mohit67890/imajev-4b navigation 72.00% +0.00 pp -2.00 pp 0.192 -0.049 +0.031 1.434 -5.065 -0.697 100
mohit67890/imajev-4b reasoning 68.00% -6.50 pp -18.17 pp 0.031 +0.004 -0.093 0.700 +0.130 -0.165 1200
mohit67890/imajev-4b record_classification 48.00% -1.00 pp +5.00 pp 0.381 -0.052 -0.159 2.272 -9.713 -12.992 100
mohit67890/imajev-4b retrieval_ranking 62.00% +2.00 pp +2.00 pp 0.174 -0.081 -0.011 1.820 -6.012 -1.242 100
mohit67890/imajev-4b retrieval_verification 87.89% +35.19 pp +0.72 pp 0.031 -0.193 -0.062 0.326 -0.252 -0.089 2222
mohit67890/imajev-4b review_scoring 59.00% -3.00 pp -2.00 pp 0.098 -0.160 -0.137 0.991 -4.878 -0.606 100
mohit67890/imajev-4b risk_scoring 70.00% +4.00 pp +18.00 pp 0.100 -0.141 -0.263 1.038 -5.321 -2.618 100
mohit67890/imajev-4b routing 78.00% +1.00 pp +1.67 pp 0.122 -0.050 -0.013 1.084 -3.517 -0.420 300
mohit67890/imajev-4b routing_triage 91.76% +6.98 pp +39.96 pp 0.083 +0.024 -0.258 0.408 -1.460 -10.525 2222
mohit67890/imajev-4b rubric_scoring_prioritization 62.33% +6.66 pp +17.46 pp 0.077 -0.170 -0.349 0.946 -2.004 -1.387 2222
mohit67890/imajev-4b safety_gating 95.00% +38.00 pp -2.00 pp 0.075 -0.241 +0.049 0.190 -0.571 +0.024 100
mohit67890/imajev-4b semantic_filtering 84.00% +30.00 pp +1.00 pp 0.044 -0.236 -0.121 0.386 -0.331 -0.725 100
mohit67890/imajev-4b target_selection 47.00% -3.00 pp -2.00 pp 0.306 -0.060 -0.109 2.729 -9.813 -5.171 100
mohit67890/imajev-4b triage 85.00% -8.00 pp -4.00 pp 0.117 +0.056 +0.039 0.483 -0.082 +0.096 100
mohit67890/imajev-4b verification 57.50% -3.50 pp -5.50 pp 0.280 -0.025 -0.055 1.207 -6.702 -0.960 200

@mohit67890
mohit67890 force-pushed the results/imajev-4b-decisionbench branch from 9844008 to 52e6d82 Compare September 26, 2026 14:05
@mohit67890 mohit67890 changed the title Add mohit67890/imajev-4b DecisionBench result Add mohit67890/imajev-4b DecisionBench results (1.0 and phase-3 revisions) Sep 26, 2026
@mohit67890

Copy link
Copy Markdown
Contributor Author

Filed the evaluation request through the official tracker so this can be run on your side if that is easier than reviewing a fork PR: Hanno-Labs/decision-bench#47. Both records here are unchanged.

@github-actions

Copy link
Copy Markdown

DecisionBench result comparison

Submitted results are compared with the pinned Jev (typesafe/jev-1.13) and Luna (openai/gpt-5.6-luna) records from main.
Accuracy deltas are percentage points; positive is better. For ECE and NLL deltas, negative is better. Per-view accuracy is calculated on supported rows.

Overall

| Model | Revision | Tags | Primary acc | Δ Jev | Δ Luna | Supported acc | Coverage | ECE | NLL | Errors | Unsupported |
|---|---|---|---:|---:|---:|---:|---:|---:|---:|---:|
| mohit67890/imajev-4b | 712891d1192c6491441a19bdea254d8a6e1048d0 | | 77.55% | +5.52 pp | +7.65 pp | 79.33% | 97.75% | 0.024 | 0.663 | 0 | 537 |
| mohit67890/imajev-4b | c9e5f132465da85d31735ec502d5557982671a7d | | 79.69% | +7.67 pp | +9.79 pp | 79.69% | 100.00% | 0.069 | 0.694 | 0 | 0 |

Pinned reference baselines
Reference Revision Primary acc Supported acc Coverage ECE NLL Errors Unsupported
Jev (typesafe/jev-1.13) openrouter-service-snapshot-2026-09-21 72.03% 72.03% 100.00% 0.128 2.433 0 0
Luna (openai/gpt-5.6-luna) openrouter-service-snapshot-2026-09-21-reasoning-aware-v1 69.90% 69.93% 99.96% 0.203 3.154 10 0

Per primitive

| Model | Tags | View | Acc | Acc Δ Jev | Acc Δ Luna | ECE | ECE Δ Jev | ECE Δ Luna | NLL | NLL Δ Jev | NLL Δ Luna | Rows |
|---|---|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| mohit67890/imajev-4b | | binary_classification | 86.58% | +27.36 pp | +2.86 pp | 0.092 | -0.065 | -0.014 | 0.377 | -0.278 | -0.147 | 6388 |
| mohit67890/imajev-4b | | candidate_selection | 80.64% | +0.54 pp | +13.71 pp | 0.012 | -0.088 | -0.208 | 0.645 | -1.600 | -3.570 | 15277 |
| mohit67890/imajev-4b | | ordinal_scoring | 40.28% | -4.83 pp | -5.65 pp | 0.321 | -0.034 | -0.138 | 1.896 | -8.974 | -1.271 | 1698 |
| mohit67890/imajev-4b | | binary_classification | 86.51% | +27.29 pp | +2.79 pp | 0.080 | -0.077 | -0.026 | 0.373 | -0.282 | -0.152 | 6388 |
| mohit67890/imajev-4b | | candidate_selection | 80.54% | +0.45 pp | +13.61 pp | 0.068 | -0.031 | -0.151 | 0.671 | -1.574 | -3.544 | 15814 |
| mohit67890/imajev-4b | | ordinal_scoring | 46.17% | +1.06 pp | +0.24 pp | 0.380 | +0.026 | -0.079 | 2.111 | -8.760 | -1.057 | 1698 |

Per family

| Model | Tags | View | Acc | Acc Δ Jev | Acc Δ Luna | ECE | ECE Δ Jev | ECE Δ Luna | NLL | NLL Δ Jev | NLL Δ Luna | Rows |
|---|---|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| mohit67890/imajev-4b | | action_selection | 63.71% | -1.57 pp | +3.29 pp | 0.072 | -0.055 | -0.040 | 1.185 | -2.093 | -0.159 | 700 |
| mohit67890/imajev-4b | | bounded_extraction | 93.88% | -1.75 pp | +0.99 pp | 0.008 | -0.024 | -0.051 | 0.223 | -0.484 | -1.505 | 2223 |
| mohit67890/imajev-4b | | completion_assessment | 90.00% | -4.00 pp | -3.00 pp | 0.078 | +0.044 | +0.039 | 0.269 | -0.519 | -0.031 | 100 |
| mohit67890/imajev-4b | | content_classification | 80.00% | -1.00 pp | +0.00 pp | 0.131 | -0.045 | -0.036 | 0.933 | -4.670 | -0.163 | 100 |
| mohit67890/imajev-4b | | content_moderation | 100.00% | +55.00 pp | +1.00 pp | 0.061 | -0.359 | +0.059 | 0.067 | -0.745 | +0.012 | 100 |
| mohit67890/imajev-4b | | context_management | 88.50% | +37.50 pp | -2.00 pp | 0.068 | -0.193 | -0.006 | 0.281 | -0.482 | -0.050 | 200 |
| mohit67890/imajev-4b | | document_classification | 39.77% | +0.77 pp | +1.77 pp | 0.257 | -0.306 | -0.279 | 3.192 | -15.123 | -12.873 | 88 |
| mohit67890/imajev-4b | | document_record_classification | 56.10% | -0.54 pp | +9.49 pp | 0.037 | -0.163 | -0.263 | 1.300 | -3.711 | -0.800 | 2223 |
| mohit67890/imajev-4b | | document_workflows | 81.91% | +9.00 pp | +8.55 pp | 0.186 | +0.072 | +0.033 | 0.550 | -0.052 | -0.233 | 2222 |
| mohit67890/imajev-4b | | entity_alignment | 100.00% | +0.00 pp | +18.56 pp | 0.137 | +0.136 | -0.031 | 0.161 | +0.160 | -5.481 | 1971 |
| mohit67890/imajev-4b | | function_agent_skill_routing | 99.13% | +0.12 pp | +17.96 pp | 0.045 | +0.041 | -0.064 | 0.078 | +0.019 | -4.008 | 1948 |
| mohit67890/imajev-4b | | guardrails_moderation | 62.38% | +23.31 pp | +0.54 pp | 0.212 | -0.235 | -0.113 | 1.278 | -6.830 | -0.937 | 2222 |
| mohit67890/imajev-4b | | navigation | 72.00% | +0.00 pp | -2.00 pp | 0.192 | -0.049 | +0.031 | 1.434 | -5.065 | -0.697 | 100 |
| mohit67890/imajev-4b | | reasoning | 68.00% | -6.50 pp | -18.17 pp | 0.031 | +0.004 | -0.093 | 0.700 | +0.130 | -0.165 | 1200 |
| mohit67890/imajev-4b | | record_classification | 48.00% | -1.00 pp | +5.00 pp | 0.381 | -0.052 | -0.159 | 2.272 | -9.713 | -12.992 | 100 |
| mohit67890/imajev-4b | | retrieval_ranking | 62.00% | +2.00 pp | +2.00 pp | 0.174 | -0.081 | -0.011 | 1.820 | -6.012 | -1.242 | 100 |
| mohit67890/imajev-4b | | retrieval_verification | 87.89% | +35.19 pp | +0.72 pp | 0.031 | -0.193 | -0.062 | 0.326 | -0.252 | -0.089 | 2222 |
| mohit67890/imajev-4b | | review_scoring | 59.00% | -3.00 pp | -2.00 pp | 0.098 | -0.160 | -0.137 | 0.991 | -4.878 | -0.606 | 100 |
| mohit67890/imajev-4b | | risk_scoring | 70.00% | +4.00 pp | +18.00 pp | 0.100 | -0.141 | -0.263 | 1.038 | -5.321 | -2.618 | 100 |
| mohit67890/imajev-4b | | routing | 78.00% | +1.00 pp | +1.67 pp | 0.122 | -0.050 | -0.013 | 1.084 | -3.517 | -0.420 | 300 |
| mohit67890/imajev-4b | | routing_triage | 91.76% | +6.98 pp | +39.96 pp | 0.083 | +0.024 | -0.258 | 0.408 | -1.460 | -10.525 | 2222 |
| mohit67890/imajev-4b | | rubric_scoring_prioritization | 62.33% | +6.66 pp | +17.46 pp | 0.077 | -0.170 | -0.349 | 0.946 | -2.004 | -1.387 | 2222 |
| mohit67890/imajev-4b | | safety_gating | 95.00% | +38.00 pp | -2.00 pp | 0.075 | -0.241 | +0.049 | 0.190 | -0.571 | +0.024 | 100 |
| mohit67890/imajev-4b | | semantic_filtering | 84.00% | +30.00 pp | +1.00 pp | 0.044 | -0.236 | -0.121 | 0.386 | -0.331 | -0.725 | 100 |
| mohit67890/imajev-4b | | target_selection | 47.00% | -3.00 pp | -2.00 pp | 0.306 | -0.060 | -0.109 | 2.729 | -9.813 | -5.171 | 100 |
| mohit67890/imajev-4b | | triage | 85.00% | -8.00 pp | -4.00 pp | 0.117 | +0.056 | +0.039 | 0.483 | -0.082 | +0.096 | 100 |
| mohit67890/imajev-4b | | verification | 57.50% | -3.50 pp | -5.50 pp | 0.280 | -0.025 | -0.055 | 1.207 | -6.702 | -0.960 | 200 |
| mohit67890/imajev-4b | | action_selection | 61.57% | -3.71 pp | +1.14 pp | 0.162 | +0.036 | +0.050 | 1.376 | -1.902 | +0.031 | 700 |
| mohit67890/imajev-4b | | bounded_extraction | 94.06% | -1.57 pp | +1.17 pp | 0.017 | -0.015 | -0.042 | 0.226 | -0.481 | -1.503 | 2223 |
| mohit67890/imajev-4b | | completion_assessment | 90.00% | -4.00 pp | -3.00 pp | 0.069 | +0.035 | +0.030 | 0.351 | -0.437 | +0.051 | 100 |
| mohit67890/imajev-4b | | content_classification | 79.00% | -2.00 pp | -1.00 pp | 0.172 | -0.004 | +0.006 | 1.048 | -4.555 | -0.047 | 100 |
| mohit67890/imajev-4b | | content_moderation | 100.00% | +55.00 pp | +1.00 pp | 0.057 | -0.364 | +0.054 | 0.063 | -0.750 | +0.007 | 100 |
| mohit67890/imajev-4b | | context_management | 87.00% | +36.00 pp | -3.50 pp | 0.059 | -0.202 | -0.015 | 0.294 | -0.470 | -0.038 | 200 |
| mohit67890/imajev-4b | | document_classification | 39.00% | +0.00 pp | +1.00 pp | 0.458 | -0.105 | -0.078 | 4.299 | -14.015 | -11.766 | 100 |
| mohit67890/imajev-4b | | document_record_classification | 54.75% | -1.89 pp | +8.14 pp | 0.151 | -0.049 | -0.149 | 1.409 | -3.603 | -0.692 | 2223 |
| mohit67890/imajev-4b | | document_workflows | 80.47% | +7.56 pp | +7.11 pp | 0.189 | +0.075 | +0.036 | 0.566 | -0.036 | -0.218 | 2222 |
| mohit67890/imajev-4b | | entity_alignment | 99.91% | -0.09 pp | +18.47 pp | 0.067 | +0.066 | -0.101 | 0.085 | +0.085 | -5.557 | 2222 |
| mohit67890/imajev-4b | | function_agent_skill_routing | 98.83% | -0.18 pp | +17.66 pp | 0.008 | +0.004 | -0.101 | 0.054 | -0.006 | -4.032 | 2222 |
| mohit67890/imajev-4b | | guardrails_moderation | 62.29% | +23.22 pp | +0.45 pp | 0.263 | -0.184 | -0.062 | 1.523 | -6.585 | -0.693 | 2222 |
| mohit67890/imajev-4b | | navigation | 71.00% | -1.00 pp | -3.00 pp | 0.225 | -0.016 | +0.064 | 1.780 | -4.718 | -0.350 | 100 |
| mohit67890/imajev-4b | | reasoning | 80.58% | +6.08 pp | -5.58 pp | 0.026 | -0.001 | -0.098 | 0.452 | -0.119 | -0.413 | 1200 |
| mohit67890/imajev-4b | | record_classification | 44.00% | -5.00 pp | +1.00 pp | 0.409 | -0.024 | -0.131 | 2.365 | -9.620 | -12.899 | 100 |
| mohit67890/imajev-4b | | retrieval_ranking | 55.00% | -5.00 pp | -5.00 pp | 0.328 | +0.073 | +0.143 | 2.590 | -5.242 | -0.472 | 100 |
| mohit67890/imajev-4b | | retrieval_verification | 88.30% | +35.60 pp | +1.13 pp | 0.020 | -0.203 | -0.073 | 0.306 | -0.272 | -0.108 | 2222 |
| mohit67890/imajev-4b | | review_scoring | 63.00% | +1.00 pp | +2.00 pp | 0.163 | -0.095 | -0.072 | 1.037 | -4.831 | -0.559 | 100 |
| mohit67890/imajev-4b | | risk_scoring | 70.00% | +4.00 pp | +18.00 pp | 0.137 | -0.104 | -0.226 | 1.250 | -5.109 | -2.406 | 100 |
| mohit67890/imajev-4b | | routing | 76.67% | -0.33 pp | +0.33 pp | 0.170 | -0.001 | +0.036 | 1.320 | -3.281 | -0.184 | 300 |
| mohit67890/imajev-4b | | routing_triage | 90.68% | +5.90 pp | +38.88 pp | 0.035 | -0.023 | -0.305 | 0.380 | -1.489 | -10.554 | 2222 |
| mohit67890/imajev-4b | | rubric_scoring_prioritization | 60.04% | +4.37 pp | +15.17 pp | 0.151 | -0.096 | -0.275 | 1.050 | -1.901 | -1.283 | 2222 |
| mohit67890/imajev-4b | | safety_gating | 95.00% | +38.00 pp | -2.00 pp | 0.048 | -0.267 | +0.022 | 0.211 | -0.550 | +0.045 | 100 |
| mohit67890/imajev-4b | | semantic_filtering | 85.00% | +31.00 pp | +2.00 pp | 0.071 | -0.210 | -0.094 | 0.420 | -0.298 | -0.692 | 100 |
| mohit67890/imajev-4b | | target_selection | 45.00% | -5.00 pp | -4.00 pp | 0.409 | +0.043 | -0.006 | 3.404 | -9.138 | -4.496 | 100 |
| mohit67890/imajev-4b | | triage | 85.00% | -8.00 pp | -4.00 pp | 0.066 | +0.005 | -0.012 | 0.441 | -0.124 | +0.055 | 100 |
| mohit67890/imajev-4b | | verification | 58.50% | -2.50 pp | -4.50 pp | 0.287 | -0.018 | -0.048 | 1.286 | -6.623 | -0.881 | 200 |

@trashhalo
trashhalo merged commit af98205 into Hanno-Labs:main Sep 27, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants