benchgap
Other · model

Inkling benchmark scores

As of 2026-10-07, Inkling (Other) has measured scores on 39 benchmarks and estimated scores on 89 more.

BenchmarkScoreSource
AA-GPQA Diamond87.2%measured
AA-HLE31.9%measured
AA-SciCode47.0%measured
AIME2697.1%measured
Artificial Analysis Intelligence Index25.0%measured
CritPt5.4%measured
GDPval-AA28.9%measured
GPQA Diamond87.9%measured
HLE46.0%measured
IFBench79.8%measured
AA-LCR77.3%measured
LiveCodeBench v686.6%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy41.6%measured
SciCode47.7%estimated ± 2.9 pp, medium confidence
Terminal-Bench 2.163.8%measured
BrowseComp77.1%measured
HLE w/ tools49.8%estimated ± 4.6 pp, high confidence
IFEval92.8%estimated ± 3.5 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, medium confidence
VITA-Bench45.8%estimated ± 9.9 pp, low confidence
τ²-bench results91.1%estimated ± 2.4 pp, medium confidence
AA Coding Index52.1%measured
AA-MMMU-Pro73.5%measured
APEX-Agents29.1%estimated ± 0.6 pp, medium confidence
APEX-Agents-AA22.9%estimated ± 4.6 pp, high confidence
BioMysteryBench (human-difficult)41.8%estimated ± 5.9 pp, low confidence
DeepSearchQA86.8%estimated ± 4.5 pp, high confidence
SWE-bench Verified77.6%measured
AutomationBench44.9%estimated ± 7.0 pp, low confidence
BFCL v474.9%estimated ± 2.6 pp, high confidence
CyberGym75.4%estimated ± 4.6 pp, high confidence
JobBench44.5%estimated ± 6.8 pp, medium confidence
SWE-bench Pro54.3%measured
τ³-bench results69.6%estimated ± 1.8 pp, low confidence
WideResearch76.5%estimated ± 4.6 pp, high confidence
AA Agentic Index24.3%measured
MMLU-Pro85.2%estimated ± 2.5 pp, medium confidence
GPQA87.9%measured
AA-IFBench80.3%estimated ± 8.5 pp, medium confidence
Gert Labs58.1%estimated ± 7.4 pp, medium confidence
AA-AnalystAgent23.8%measured
AA EnterpriseOps-Gym38.0%measured
AA Harvey LAB83.4%estimated ± 1.2 pp, medium confidence
ARC-AGI-185.7%estimated ± 2.7 pp, medium confidence
ARC-AGI-248.0%estimated ± 10.1 pp, low confidence
CursorBench 3.136.6%estimated ± 3.4 pp, low confidence
CursorBench 3.249.1%estimated ± 3.6 pp, medium confidence
Terminal-Bench 3.01.5%estimated ± 6.9 pp, low confidence
FrontierCode 1.1 Main17.4%estimated ± 2.7 pp, low confidence
FrontierSWE v24.1%measured
OfficeQA Pro50.3%estimated ± 7.2 pp, low confidence
OSWorld-Verified68.7%estimated ± 5.4 pp, medium confidence
Vals GPQA Diamond87.1%measured
Vals LiveCodeBench85.5%measured
Vals MMLU-Pro86.3%measured
Vals SWE-bench77.6%measured
Terminal-Bench 2.1 (Vals AI)47.6%measured
VulcanBench v383.9%estimated ± 5.6 pp, low confidence
AA AutomationBench5.0%measured
GDP.pdf12.8%measured
AA ITBench42.1%estimated ± 3.6 pp, low confidence
AA Tau3 Banking29.1%measured
AA Terminal-Bench 2.155.1%measured
AA Terminal-Bench 4.01.0%measured
CursorBench 4.017.2%estimated ± 1.6 pp, medium confidence
CWE-bench v137.0%measured
DeepSWE40.8%estimated ± 4.2 pp, medium confidence
HLE w/o tools30.0%measured
OSWorld 2.06.2%estimated ± 13.3 pp, low confidence
PostTrainBench v1.123.4%estimated ± 2.7 pp, high confidence
Terminal-Bench 4.05.0%estimated ± 5.1 pp, low confidence
Terminal-Bench-Science 0.10.0%estimated ± 4.1 pp, low confidence
Toolathlon-Verified71.0%estimated ± 1.9 pp, low confidence
Vibe Code Bench18.6%estimated ± 8.1 pp, medium confidence
CharXiv82.0%measured
CharXiv w/o tools78.1%measured
ExploitGym13.7%estimated ± 5.0 pp, low confidence
C-Eval93.3%estimated ± 0.8 pp, low confidence
Claw-Eval69.3%estimated ± 5.7 pp, medium confidence
DeepPlanning47.2%estimated ± 13.1 pp, low confidence
HMMT Feb 202597.6%estimated ± 1.9 pp, low confidence
HMMT Feb 202688.9%estimated ± 4.0 pp, high confidence
MathVision82.0%estimated ± 2.1 pp, low confidence
MCP Atlas74.1%measured
MCP-Tasks77.0%estimated ± 0.6 pp, low confidence
MMAnswerBench86.5%estimated ± 1.3 pp, high confidence
MMLU-Redux96.8%estimated ± 1.1 pp, low confidence
MMMU-Pro73.5%measured
NL2Repo44.4%estimated ± 4.7 pp, high confidence
QwenClawBench61.8%estimated ± 2.6 pp, high confidence
ScreenSpot Pro75.2%estimated ± 7.1 pp, medium confidence
SuperGPQA82.0%estimated ± 7.5 pp, medium confidence
Terminal-Bench 2.056.7%estimated ± 5.9 pp, medium confidence
Toolathlon50.2%estimated ± 3.5 pp, high confidence
VideoMMMU84.1%estimated ± 1.0 pp, high confidence
ERQA67.1%estimated ± 3.7 pp, high confidence
HealthBench Hard0.0%estimated ± 5.1 pp, low confidence
LiveCodeBench Pro79.1%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)59.6%estimated ± 2.4 pp, high confidence
MedXpertQA (Text)48.1%estimated ± 4.5 pp, low confidence
MMLU-Pro (Arcee)86.2%estimated ± 3.4 pp, medium confidence
React Native Evals77.3%estimated ± 2.4 pp, medium confidence
ResearchClawBench18.5%estimated ± 1.8 pp, medium confidence
SWE-Rebench60.9%estimated ± 3.6 pp, medium confidence
ARC-AGI-30.0%estimated ± 2.1 pp, medium confidence
BioMysteryBench (human-solvable)81.8%estimated ± 0.6 pp, low confidence
DRACO73.4%estimated ± 3.0 pp, medium confidence
FrontierCode 1.1 Extended26.4%estimated ± 2.4 pp, low confidence
HealthBench (raw)41.2%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional56.6%estimated ± 1.3 pp, medium confidence
HealthBench Professional (raw)52.1%estimated ± 5.0 pp, low confidence
HLE-Verified15.3%estimated ± 2.0 pp, low confidence
LABBench281.1%estimated ± 1.8 pp, low confidence
AIME 202597.5%estimated ± 1.5 pp, low confidence
MMMU81.8%estimated ± 0.9 pp, high confidence
LiveCodeBench100.0%estimated ± 8.9 pp, low confidence
Agents' Last Exam28.5%estimated ± 6.3 pp, medium confidence
OpenHarmony Bench51.0%estimated ± 2.4 pp, high confidence
SimpleVQA67.6%estimated ± 7.8 pp, low confidence
MMMLU89.0%estimated ± 1.1 pp, medium confidence
MMLU93.7%estimated ± 1.1 pp, low confidence
MMMU-Pro w/ Python75.1%estimated ± 0.6 pp, high confidence
MATH-50097.9%estimated ± 3.1 pp, low confidence
MathVision w/ Python76.8%estimated ± 0.9 pp, low confidence
RealWorldQA85.6%estimated ± 1.1 pp, medium confidence
Video-MME (with subtitle)87.8%estimated ± 0.9 pp, medium confidence
OmniDocBench 1.560.5%estimated ± 5.9 pp, low confidence