benchgap
Calibration

JevBench 1.5 → JevBench 1.4

JevBench 1.4 is estimated from JevBench 1.5 with a offset logistic curve fitted on 90 models measured on both: y = 0.0886 + (1.2000 − 0.0886) / (1 + exp(−3.18·(x − 0.8121))), R² = 0.63, cross-validated error 10.7 pp. It is used for 18 estimates.

Estimated modelJevBench 1.5JevBench 1.4Source
AutoJev-27B19.5%22.6%estimated ± 10.7 pp, low confidence
AutoJev-27B (RTX PRO 6000)19.5%22.6%estimated ± 10.7 pp, low confidence
Bev / Bonsai 27B15.8%21.2%estimated ± 10.7 pp, low confidence
Bosun v3.1 0.6B2.5%17.3%estimated ± 10.7 pp, low confidence
Decision 4B v1.170.4%55.0%estimated ± 10.7 pp, low confidence
Decision 4B v1.270.8%55.4%estimated ± 10.7 pp, low confidence
Deem 0.8B v12.1%17.2%estimated ± 10.7 pp, low confidence
Eikos-27B18.5%22.2%estimated ± 10.7 pp, low confidence
Imajev-4B (RTX 5090)70.4%55.0%estimated ± 10.7 pp, low confidence
Instinct Dual 4B47.0%36.9%estimated ± 10.7 pp, low confidence
JevK5 v0.3 4B71.9%56.3%estimated ± 10.7 pp, low confidence
Laya multilingual0.0%16.7%estimated ± 10.7 pp, low confidence
Laya typed-decisions0.0%16.7%estimated ± 10.7 pp, low confidence
Manchego v2.168.8%53.6%estimated ± 10.7 pp, low confidence
Needle 3 (2-bit)0.0%16.7%estimated ± 10.7 pp, low confidence
Needle 3 (options as tools)0.0%16.7%estimated ± 10.7 pp, low confidence
Nemotron Diffusion 8B (optimized vLLM)25.7%25.1%estimated ± 10.7 pp, low confidence
Surogate Rune 26B-A4B v3 (RTX PRO 6000)66.5%51.7%estimated ± 10.7 pp, low confidence