benchgap
Other · model

LFM2.5-350M benchmark scores

As of 2026-10-10, LFM2.5-350M (Other) has measured scores on 3 benchmarks and estimated scores on 23 more.

BenchmarkScoreSource
AA-GPQA Diamond30.1%estimated ± 4.7 pp, high confidence
AA-HLE4.9%estimated ± 5.3 pp, medium confidence
Artificial Analysis Intelligence Index8.8%estimated ± 6.3 pp, medium confidence
CritPt0.6%estimated ± 3.8 pp, high confidence
GPQA Diamond30.7%estimated ± 0.4 pp, high confidence
HLE0.0%estimated ± 6.4 pp, low confidence
IFBench37.1%estimated ± 9.6 pp, medium confidence
AA-Omniscience Accuracy3.1%estimated ± 8.5 pp, low confidence
IFEval77.0%measured
MMLU-Pro20.0%measured
GPQA30.6%measured
AA-IFBench22.0%estimated ± 7.4 pp, medium confidence
ARC-AGI-10.0%estimated ± 6.9 pp, low confidence
ARC-AGI-20.0%estimated ± 11.6 pp, low confidence
Vals GPQA Diamond13.6%estimated ± 4.8 pp, medium confidence
Vals MMLU-Pro0.0%estimated ± 1.2 pp, medium confidence
HLE w/o tools6.1%estimated ± 4.9 pp, medium confidence
C-Eval0.1%estimated ± 1.3 pp, low confidence
MMLU-Redux44.6%estimated ± 0.6 pp, medium confidence
SuperGPQA36.7%estimated ± 8.2 pp, low confidence
HealthBench Hard0.0%estimated ± 5.1 pp, low confidence
ARC-AGI-30.0%estimated ± 2.1 pp, medium confidence
HealthBench Professional0.1%estimated ± 4.8 pp, low confidence
MMLU44.5%estimated ± 2.2 pp, high confidence
SimpleQA14.4%estimated ± 10.9 pp, low confidence
BBH42.1%estimated ± 3.2 pp, medium confidence