benchgap
Other · model

LFM2.5-VL-3B benchmark scores

As of 2026-10-07, LFM2.5-VL-3B (Other) has measured scores on 8 benchmarks and estimated scores on 17 more.

BenchmarkScoreSource
GDPval-AA0.0%estimated ± 5.8 pp, low confidence
IFBench25.8%measured
IFEval82.3%measured
τ²-bench results55.1%estimated ± 14.1 pp, low confidence
AA-MMMU-Pro16.2%estimated ± 1.2 pp, medium confidence
BFCL v432.5%measured
AA-IFBench26.5%estimated ± 8.2 pp, low confidence
CharXiv41.4%estimated ± 4.1 pp, low confidence
CharXiv w/o tools30.8%estimated ± 1.8 pp, low confidence
Claw-Eval28.1%estimated ± 2.3 pp, low confidence
MathVision24.3%estimated ± 3.6 pp, medium confidence
MMMU-Pro30.5%measured
ScreenSpot Pro85.1%estimated ± 2.5 pp, low confidence
V*68.2%estimated ± 1.4 pp, low confidence
VideoMMMU84.1%estimated ± 1.0 pp, medium confidence
ERQA39.7%estimated ± 3.2 pp, low confidence
MedXpertQA (MM)0.7%estimated ± 2.4 pp, medium confidence
MMMU48.4%measured
SimpleVQA35.4%measured
ZeroBench74.8%estimated ± 6.4 pp, low confidence
MMMU-Pro w/ Python60.9%estimated ± 0.6 pp, medium confidence
RefCOCO (avg)87.9%measured
RealWorldQA73.1%measured
Video-MME (with subtitle)80.4%estimated ± 1.8 pp, low confidence
OmniDocBench 1.50.0%estimated ± 5.9 pp, low confidence