benchgap
Multilingual

MAXIFE leaderboard

As of 2026-10-10, the highest measured score on MAXIFE is 89.2% by Qwen3.7 Max. 13 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Qwen3.7 Max89.2%measured
2Qwen3.7 Plus88.8%measured
3Qwen3.5 397B88.2%measured
4K-EXAONE 2.084.0%estimated ± 4.7 pp, medium confidence
5Qwen3.6 Plus83.5%estimated ± 2.2 pp, medium confidence
6Claude Opus 4.579.2%measured
7Nemotron 3 Ultra77.3%estimated ± 6.5 pp, low confidence
8Qwen3.5-122B-A10B74.8%estimated ± 6.5 pp, low confidence
9Qwen3.5-27B74.8%estimated ± 6.5 pp, low confidence
10Kimi K2.572.8%measured
11Qwen3 235B 2507 (Reasoning)71.2%estimated ± 6.5 pp, low confidence
12Qwen3.5-35B-A3B71.2%estimated ± 6.5 pp, low confidence
13GLM-569.7%estimated ± 2.2 pp, low confidence
14Qwen3 235B 250766.4%estimated ± 6.5 pp, low confidence
15GPT-4.147.2%estimated ± 6.5 pp, low confidence
16DeepSeek V3 032441.3%estimated ± 6.5 pp, low confidence
17GPT-4o17.2%estimated ± 6.5 pp, low confidence
18Phi-40.0%estimated ± 6.5 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General