$ipbr-rank · live llm coding-role score
refreshed · 14 sources · updated frequently — models drift and degrade
[ idea ]
1gemini-3.1-pro-preview96.996.9
2claude-opus-4.690.590.5
3claude-opus-4.786.085.6
[ plan ]
1gemini-3.1-pro-preview89.489.4
2gpt-5.583.883.8
3claude-opus-4.779.378.4
[ build ]
1gemini-3.1-pro-preview84.484.4
2gpt-5.583.183.1
3claude-opus-4.680.380.3
[ review ]
1gemini-3.1-pro-preview86.586.5
2kimi-k2.683.783.7
3claude-opus-4.782.282.2
how scoring works

Each model gets four role scores from public benchmarks. Idea measures open-ended creativity. Plan measures structured reasoning, function-calling, and multi-step decomposition. Build measures implementation skill — SWE-bench, LiveCodeBench, terminal tasks. Review measures preference judgment.

raw vs adjusted

The raw score is the benchmark composite, normalized to 0-100. The adjusted score subtracts a reviewer-reservation penalty: when a vendor's models lead the direct LM Arena search/document review proxy, that proxy lead gets discounted from their Idea, Plan, and Build scores so vendors can't game their own preference evaluations.

Penalty coefficients differ by role: Build is penalized hardest (0.32), Plan moderately (0.18), Idea lightly (0.08). Review is never adjusted.

missing data

If a model is missing some metrics within a group, the group score blends from shrink-to-50 to trusting the present metrics across 60-80% group coverage. At 80% coverage and above, the present-weight mean is trusted directly.

Full math, role definitions, and source list →

gemini-3.1-pro-previewgoogle96.996.989.489.484.484.486.5

group breakdown

A_B85.15 / 24A_I92.15 / 24A_P71.95 / 24A_R78.317 / 24BUILD83.54 / 24CRE99.52 / 24GEN99.91 / 24LM_ARENA_REVIEW_PROXY92.34 / 24OPS_long88.47 / 24OPS_precision83.210 / 24OPS_review86.29 / 24PLAN93.92 / 24

metrics

AI_code89.85 / 22AI_complexity92.55 / 22AI_context_awareness7.55 / 24AI_correctness92.517 / 22AI_edge_cases92.55 / 22AI_efficiency89.05 / 22AI_hallucination_resistance10.922 / 24AI_memory_retention92.58 / 24AI_parameter_accuracy78.118 / 24AI_plan_coherence92.58 / 24AI_recovery92.516 / 22AI_refusal92.520 / 22AI_spec92.520 / 22AI_stability92.011 / 22AI_task_completion27.419 / 24AI_tool_selection11.619 / 24ARC_AGI_2100.01 / 17ArtificialAnalysisCoding100.01 / 21ArtificialAnalysisIntelligence100.02 / 21ArtificialAnalysisReasoning100.01 / 21BlendedCost77.313 / 24ContextWindow100.06 / 24CopilotArenaOrLMArenaCode73.47 / 22GDPval24.712 / 16GPQA_HLE_Reasoning100.01 / 21GSO51.38 / 15IFBench97.53 / 21LMArenaCreativeOrOpenEnded99.52 / 23LMArenaSearchDocument92.34 / 19LMArenaText99.52 / 23LongContextRecall100.02 / 21MCPAtlas71.16 / 13OutputSpeed93.94 / 19SWEBenchPro89.15 / 15SWEBenchVerified95.04 / 18SWEComposite94.82 / 24SWERebench99.82 / 21SciCode100.02 / 21SonarBugDensity52.712 / 17SonarComposite54.212 / 24SonarFunctionalSkill78.98 / 17SonarIssueDensity13.213 / 17SonarVulnerabilityDensity58.211 / 17TTFT70.118 / 19Tau2Bench99.34 / 21TerminalBench89.43 / 22
sources arc_agiartificial_analysisgsolmarenamcp_atlasopenrouteroverridessonarswebench_proswerebenchterminal_benchmissing SWEComposite/SWEBenchMultilingual
gpt-5.5openai72.472.483.883.883.183.181.7

group breakdown

A_B66.711 / 24A_I76.711 / 24A_P62.210 / 24A_R88.58 / 24BUILD89.11 / 24CRE60.87 / 24GEN88.94 / 24LM_ARENA_REVIEW_PROXY28.214 / 24OPS_long81.312 / 24OPS_precision77.515 / 24OPS_review79.615 / 24PLAN94.31 / 24

metrics

AI_code23.118 / 22AI_complexity33.219 / 22AI_context_awareness0.020 / 24AI_correctness94.114 / 22AI_edge_cases86.516 / 22AI_efficiency60.110 / 22AI_hallucination_resistance100.014 / 24AI_memory_retention0.024 / 24AI_parameter_accuracy99.13 / 24AI_plan_coherence0.123 / 24AI_recovery98.713 / 22AI_refusal100.016 / 22AI_spec100.016 / 22AI_stability92.78 / 22AI_task_completion100.08 / 24AI_tool_selection88.46 / 24ARC_AGI_296.72 / 17ArtificialAnalysisCoding100.02 / 21ArtificialAnalysisIntelligence98.13 / 21ArtificialAnalysisReasoning100.02 / 21BlendedCost50.623 / 24ContextWindow100.02 / 24CopilotArenaOrLMArenaCode71.98 / 22GDPval95.01 / 16GPQA_HLE_Reasoning100.02 / 21GSO94.02 / 15IFBench80.76 / 21LMArenaCreativeOrOpenEnded60.87 / 23LMArenaSearchDocument28.29 / 19LMArenaText60.87 / 23LongContextRecall98.03 / 21OutputSpeed82.811 / 19SWEBenchPro95.03 / 15SWEBenchVerified95.07 / 18SWEComposite89.94 / 24SWERebench83.58 / 21SciCode94.54 / 21SonarBugDensity94.52 / 17SonarComposite65.55 / 24SonarFunctionalSkill46.513 / 17SonarIssueDensity52.74 / 17SonarVulnerabilityDensity99.22 / 17TTFT78.611 / 19Tau2Bench90.57 / 21TerminalBench100.01 / 22
sources aistupidlevelarc_agiartificial_analysisgsolmarenaopenrouteroverridessonarterminal_benchmissing BUILD/MCPAtlasPLAN/MCPAtlasSWEComposite/SWEBenchMultilingual
claude-opus-4.6anthropic90.590.575.375.380.380.376.1

group breakdown

A_B67.58 / 24A_I76.513 / 24A_P60.613 / 24A_R89.03 / 24BUILD86.22 / 24CRE100.01 / 24GEN89.53 / 24LM_ARENA_REVIEW_PROXY33.113 / 24OPS_long80.514 / 24OPS_precision77.316 / 24OPS_review79.914 / 24PLAN74.26 / 24

metrics

AI_canary_health83.35 / 7AI_code27.410 / 22AI_complexity33.210 / 22AI_context_awareness0.09 / 24AI_correctness94.16 / 22AI_edge_cases86.58 / 22AI_efficiency58.512 / 22AI_hallucination_resistance100.03 / 24AI_memory_retention0.013 / 24AI_parameter_accuracy85.112 / 24AI_plan_coherence0.024 / 24AI_recovery98.75 / 22AI_refusal100.03 / 22AI_spec100.03 / 22AI_stability92.75 / 22AI_task_completion83.310 / 24AI_tool_selection99.92 / 24ARC_AGI_290.94 / 17ArtificialAnalysisCoding76.15 / 21ArtificialAnalysisIntelligence84.05 / 21ArtificialAnalysisReasoning86.35 / 21BlendedCost61.921 / 24ContextWindow99.37 / 24CopilotArenaOrLMArenaCode99.82 / 22GDPval72.87 / 16GPQA_HLE_Reasoning86.35 / 21GSO75.33 / 15IFBench31.416 / 21LMArenaCreativeOrOpenEnded100.01 / 23LMArenaSearchDocument33.18 / 19LMArenaText100.01 / 23LongContextRecall90.24 / 21OutputSpeed81.514 / 19SWEBenchMultilingual90.92 / 6SWEBenchPro100.01 / 15SWEBenchVerified99.72 / 18SWEComposite95.71 / 24SWERebench91.64 / 21SciCode85.85 / 21SonarBugDensity59.58 / 17SonarComposite70.54 / 24SonarFunctionalSkill92.23 / 17SonarIssueDensity46.86 / 17SonarVulnerabilityDensity66.67 / 17TTFT73.216 / 19Tau2Bench91.26 / 21TerminalBench64.27 / 22
sources aistupidlevelarc_agiartificial_analysisgsolmarenamcp_atlasopenrouteroverridessonarswebenchswebench_proswerebenchterminal_benchmissing BUILD/MCPAtlasPLAN/MCPAtlas
claude-opus-4.7anthropic86.085.679.378.478.576.882.2

group breakdown

A_B58.418 / 24A_I76.712 / 24A_P61.211 / 24A_R71.118 / 24BUILD86.23 / 24CRE87.44 / 24GEN95.02 / 24LM_ARENA_REVIEW_PROXY100.01 / 24OPS_long81.213 / 24OPS_precision77.814 / 24OPS_review80.413 / 24PLAN79.94 / 24

metrics

AI_code23.114 / 22AI_complexity33.211 / 22AI_context_awareness0.010 / 24AI_correctness94.17 / 22AI_edge_cases86.59 / 22AI_efficiency61.09 / 22AI_hallucination_resistance0.024 / 24AI_memory_retention0.014 / 24AI_parameter_accuracy88.510 / 24AI_plan_coherence5.414 / 24AI_recovery98.76 / 22AI_refusal100.04 / 22AI_spec100.04 / 22AI_stability88.914 / 22AI_task_completion100.02 / 24AI_tool_selection67.913 / 24ARC_AGI_292.73 / 17ArtificialAnalysisCoding90.33 / 21ArtificialAnalysisIntelligence100.01 / 21ArtificialAnalysisReasoning95.63 / 21BlendedCost61.922 / 24ContextWindow99.38 / 24CopilotArenaOrLMArenaCode100.01 / 22GDPval93.92 / 16GPQA_HLE_Reasoning95.63 / 21GSO100.01 / 15IFBench46.610 / 21LMArenaCreativeOrOpenEnded87.44 / 23LMArenaSearchDocument100.01 / 19LMArenaText87.44 / 23LongContextRecall88.26 / 21OutputSpeed82.512 / 19SWEBenchPro95.02 / 15SWEBenchVerified95.03 / 18SWEComposite90.73 / 24SWERebench85.36 / 21SciCode100.01 / 21SonarBugDensity50.114 / 17SonarComposite51.413 / 24SonarFunctionalSkill93.92 / 17SonarIssueDensity0.017 / 17SonarVulnerabilityDensity25.314 / 17TTFT73.815 / 19Tau2Bench83.19 / 21TerminalBench78.24 / 22
sources aistupidlevelarc_agiartificial_analysisgsolmarenamcp_atlasopenrouteroverridessonarmissing BUILD/MCPAtlasPLAN/MCPAtlasSWEComposite/SWEBenchMultilingual
claude-opus-4.5anthropic58.458.465.865.876.176.169.6

group breakdown

A_B69.57 / 24A_I76.910 / 24A_P62.98 / 24A_R89.82 / 24BUILD79.55 / 24CRE43.614 / 24GEN61.66 / 24LM_ARENA_REVIEW_PROXY11.221 / 24OPS_long79.017 / 24OPS_precision75.717 / 24OPS_review75.517 / 24PLAN68.28 / 24

metrics

AI_canary_health88.52 / 7AI_code35.97 / 22AI_complexity33.29 / 22AI_context_awareness0.08 / 24AI_correctness94.15 / 22AI_edge_cases86.57 / 22AI_efficiency60.011 / 22AI_hallucination_resistance100.02 / 24AI_memory_retention0.012 / 24AI_parameter_accuracy99.82 / 24AI_plan_coherence2.817 / 24AI_recovery98.74 / 22AI_refusal100.02 / 22AI_spec100.02 / 22AI_stability92.74 / 22AI_task_completion100.01 / 24AI_tool_selection92.25 / 24ArtificialAnalysisCoding75.16 / 21ArtificialAnalysisIntelligence71.57 / 21ArtificialAnalysisReasoning63.79 / 21BlendedCost61.920 / 24ContextWindow74.721 / 24CopilotArenaOrLMArenaCode76.86 / 22GDPval71.58 / 16GPQA_HLE_Reasoning63.79 / 21GSO59.35 / 15IFBench44.911 / 21LMArenaCreativeOrOpenEnded43.613 / 23LMArenaSearchDocument11.216 / 19LMArenaText43.613 / 23LongContextRecall100.01 / 21OutputSpeed84.110 / 19SWEBenchPro88.46 / 15SWEBenchVerified92.29 / 18SWEComposite83.87 / 24SWERebench76.59 / 21SciCode72.77 / 21SonarBugDensity73.75 / 17SonarComposite87.11 / 24SonarFunctionalSkill100.01 / 17SonarIssueDensity77.23 / 17SonarVulnerabilityDensity87.23 / 17TTFT76.914 / 19Tau2Bench85.28 / 21TerminalBench54.811 / 22
sources aistupidlevelartificial_analysisgsolmarenamcp_atlasopenroutersonarswebench_proswerebenchterminal_benchmissing BUILD/MCPAtlasGEN/ARC_AGI_2PLAN/MCPAtlasSWEComposite/SWEBenchMultilingual
glm-5.1zai68.568.565.865.872.472.478.5

group breakdown

A_B64.515 / 24A_I73.317 / 24A_P59.015 / 24A_R82.814 / 24BUILD73.39 / 24CRE70.45 / 24GEN53.510 / 24LM_ARENA_REVIEW_PROXY88.05 / 24OPS_long84.59 / 24OPS_precision88.85 / 24OPS_review86.38 / 24PLAN73.37 / 24

metrics

AI_code30.89 / 22AI_complexity35.87 / 22AI_context_awareness7.56 / 24AI_correctness87.518 / 22AI_edge_cases81.017 / 22AI_efficiency56.316 / 22AI_hallucination_resistance92.519 / 24AI_memory_retention7.510 / 24AI_parameter_accuracy74.519 / 24AI_plan_coherence21.011 / 24AI_recovery91.417 / 22AI_refusal92.521 / 22AI_spec92.521 / 22AI_stability83.017 / 22AI_task_completion78.314 / 24AI_tool_selection59.715 / 24ARC_AGI_25.211 / 17ArtificialAnalysisCoding39.513 / 21ArtificialAnalysisIntelligence60.58 / 21ArtificialAnalysisReasoning54.013 / 21BlendedCost93.06 / 24ContextWindow74.919 / 24CopilotArenaOrLMArenaCode95.93 / 22GDPval59.59 / 16GPQA_HLE_Reasoning54.013 / 21IFBench86.85 / 21LMArenaCreativeOrOpenEnded70.45 / 23LMArenaSearchDocument88.05 / 19LMArenaText70.45 / 23LongContextRecall41.217 / 21MCPAtlas100.01 / 13OutputSpeed80.017 / 19SWEBenchMultilingual50.93 / 6SWEBenchVerified91.910 / 18SWEComposite78.68 / 24SWERebench100.01 / 21SciCode40.414 / 21SonarBugDensity100.01 / 17SonarComposite86.02 / 24SonarFunctionalSkill69.89 / 17SonarIssueDensity100.01 / 17SonarVulnerabilityDensity87.24 / 17TTFT100.01 / 19Tau2Bench100.03 / 21TerminalBench55.810 / 22
sources arc_agiartificial_analysislmarenamcp_atlasopenrouteroverridessonarswebenchswerebenchterminal_benchmissing BUILD/GSOSWEComposite/SWEBenchPro
kimi-k2.6moonshot61.961.974.274.272.472.483.7

group breakdown

A_B67.19 / 24A_I77.47 / 24A_P60.514 / 24A_R88.64 / 24BUILD74.27 / 24CRE51.99 / 24GEN67.35 / 24LM_ARENA_REVIEW_PROXY94.82 / 24OPS_long58.220 / 24OPS_precision62.118 / 24OPS_review65.019 / 24PLAN89.13 / 24

metrics

AI_code27.411 / 22AI_complexity33.216 / 22AI_context_awareness0.016 / 24AI_correctness94.111 / 22AI_edge_cases86.513 / 22AI_efficiency57.414 / 22AI_hallucination_resistance100.010 / 24AI_memory_retention0.020 / 24AI_parameter_accuracy78.915 / 24AI_plan_coherence15.913 / 24AI_recovery98.710 / 22AI_refusal100.012 / 22AI_spec100.012 / 22AI_stability88.915 / 22AI_task_completion83.313 / 24AI_tool_selection61.514 / 24ARC_AGI_211.99 / 17ArtificialAnalysisCoding72.87 / 21ArtificialAnalysisIntelligence87.54 / 21ArtificialAnalysisReasoning87.64 / 21BlendedCost89.09 / 24ContextWindow78.814 / 24CopilotArenaOrLMArenaCode94.44 / 22GDPval52.110 / 16GPQA_HLE_Reasoning87.64 / 21IFBench94.54 / 21LMArenaCreativeOrOpenEnded51.99 / 23LMArenaSearchDocument94.82 / 19LMArenaText51.99 / 23LongContextRecall85.37 / 21MCPAtlas92.52 / 13SWEBenchVerified95.05 / 18SWEComposite66.014 / 24SWERebench73.112 / 21SciCode94.53 / 21SonarBugDensity92.53 / 17SonarComposite80.63 / 24SonarFunctionalSkill66.812 / 17SonarIssueDensity92.52 / 17SonarVulnerabilityDensity81.65 / 17Tau2Bench100.01 / 21TerminalBench74.65 / 22
sources aistupidlevelartificial_analysislmarenaopenrouteroverridesmissing BUILD/GSOOPS_long/OutputSpeedOPS_long/TTFTOPS_precision/OutputSpeedOPS_precision/TTFTOPS_review/OutputSpeedOPS_review/TTFTSWEComposite/SWEBenchMultilingualSWEComposite/SWEBenchPro
claude-sonnet-4.6anthropic56.456.459.559.571.371.366.8

group breakdown

A_B66.910 / 24A_I77.09 / 24A_P62.59 / 24A_R88.57 / 24BUILD76.26 / 24CRE42.615 / 24GEN57.88 / 24LM_ARENA_REVIEW_PROXY23.315 / 24OPS_long70.218 / 24OPS_precision55.922 / 24OPS_review65.818 / 24PLAN59.510 / 24

metrics

AI_canary_health88.23 / 7AI_code23.117 / 22AI_complexity33.214 / 22AI_context_awareness0.013 / 24AI_correctness94.19 / 22AI_edge_cases86.511 / 22AI_efficiency63.67 / 22AI_hallucination_resistance100.06 / 24AI_memory_retention0.017 / 24AI_parameter_accuracy90.78 / 24AI_plan_coherence0.121 / 24AI_recovery98.78 / 22AI_refusal100.07 / 22AI_spec100.07 / 22AI_stability92.76 / 22AI_task_completion100.05 / 24AI_tool_selection96.03 / 24ARC_AGI_210.610 / 17ArtificialAnalysisCoding85.14 / 21ArtificialAnalysisIntelligence79.16 / 21ArtificialAnalysisReasoning68.78 / 21BlendedCost74.418 / 24ContextWindow99.311 / 24CopilotArenaOrLMArenaCode93.25 / 22GDPval80.16 / 16GPQA_HLE_Reasoning68.78 / 21GSO30.710 / 15IFBench41.013 / 21LMArenaCreativeOrOpenEnded42.614 / 23LMArenaSearchDocument23.310 / 19LMArenaText42.614 / 23LongContextRecall90.25 / 21MCPAtlas69.87 / 13OutputSpeed87.18 / 19SWEBenchPro76.510 / 15SWEBenchVerified90.311 / 18SWEComposite87.46 / 24SWERebench95.73 / 21SciCode57.98 / 21SonarBugDensity65.86 / 17SonarComposite55.88 / 24SonarFunctionalSkill84.54 / 17SonarIssueDensity22.310 / 17SonarVulnerabilityDensity21.815 / 17TTFT0.019 / 19Tau2Bench53.312 / 21TerminalBench47.414 / 22
sources aistupidlevelarc_agiartificial_analysislmarenamcp_atlasopenroutersonarswerebenchmissing SWEComposite/SWEBenchMultilingual
gemini-3-progoogle80.580.560.460.469.069.058.9

group breakdown

A_B91.22 / 24A_I99.61 / 24A_P75.72 / 24A_R83.313 / 24BUILD64.411 / 24CRE87.43 / 24GEN58.17 / 24LM_ARENA_REVIEW_PROXY19.918 / 24OPS_long45.223 / 24OPS_precision48.023 / 24OPS_review43.024 / 24PLAN55.213 / 24

metrics

AI_code96.82 / 22AI_complexity100.01 / 22AI_context_awareness0.014 / 24AI_correctness100.02 / 22AI_edge_cases100.02 / 22AI_efficiency95.82 / 22AI_hallucination_resistance4.023 / 24AI_memory_retention100.01 / 24AI_parameter_accuracy83.013 / 24AI_plan_coherence100.01 / 24AI_recovery100.02 / 22AI_refusal100.010 / 22AI_spec100.010 / 22AI_stability99.42 / 22AI_task_completion23.420 / 24AI_tool_selection4.820 / 24ARC_AGI_241.96 / 17BlendedCost77.312 / 24ContextWindow0.024 / 24CopilotArenaOrLMArenaCode68.410 / 22GDPval5.016 / 16GSO40.79 / 15LMArenaCreativeOrOpenEnded87.43 / 23LMArenaSearchDocument19.913 / 19LMArenaText87.43 / 23MCPAtlas74.93 / 13SWEBenchMultilingual33.54 / 6SWEBenchPro80.38 / 15SWEBenchVerified82.913 / 18SWEComposite72.111 / 24SWERebench70.614 / 21SonarBugDensity53.29 / 17SonarComposite54.99 / 24SonarFunctionalSkill84.15 / 17SonarIssueDensity6.716 / 17SonarVulnerabilityDensity59.78 / 17TerminalBench61.28 / 22
sources aistupidlevelarc_agiartificial_analysisgsolmarenamcp_atlasopenrouteroverridessonarswebenchswebench_proswerebenchterminal_benchmissing BUILD/ArtificialAnalysisCodingBUILD/LongContextRecallBUILD/SciCodeGEN/ArtificialAnalysisIntelligenceGEN/GPQA_HLE_ReasoningOPS_long/OutputSpeedOPS_long/TTFTOPS_precision/OutputSpeedOPS_precision/TTFTOPS_review/OutputSpeedOPS_review/TTFTPLAN/ArtificialAnalysisReasoningPLAN/IFBenchPLAN/LongContextRecallPLAN/Tau2Bench
gemini-3-flashgoogle74.374.367.067.068.368.363.2

group breakdown

A_B85.14 / 24A_I92.14 / 24A_P71.94 / 24A_R78.316 / 24BUILD59.012 / 24CRE69.46 / 24GEN57.49 / 24LM_ARENA_REVIEW_PROXY20.017 / 24OPS_long95.02 / 24OPS_precision91.61 / 24OPS_review93.52 / 24PLAN65.59 / 24

metrics

AI_code89.84 / 22AI_complexity92.54 / 22AI_context_awareness7.54 / 24AI_correctness92.516 / 22AI_edge_cases92.54 / 22AI_efficiency89.04 / 22AI_hallucination_resistance10.921 / 24AI_memory_retention92.57 / 24AI_parameter_accuracy78.117 / 24AI_plan_coherence92.57 / 24AI_recovery92.515 / 22AI_refusal92.519 / 22AI_spec92.519 / 22AI_stability92.010 / 22AI_task_completion27.418 / 24AI_tool_selection11.618 / 24ARC_AGI_23.114 / 17ArtificialAnalysisCoding58.39 / 21ArtificialAnalysisIntelligence58.911 / 21ArtificialAnalysisReasoning82.76 / 21BlendedCost91.58 / 24ContextWindow100.05 / 24CopilotArenaOrLMArenaCode68.012 / 22GDPval8.014 / 16GPQA_HLE_Reasoning82.76 / 21GSO14.013 / 15IFBench100.02 / 21LMArenaCreativeOrOpenEnded69.46 / 23LMArenaSearchDocument20.012 / 19LMArenaText69.46 / 23LongContextRecall68.69 / 21MCPAtlas22.49 / 13OutputSpeed99.12 / 19SWEBenchMultilingual100.01 / 6SWEBenchPro53.012 / 15SWEBenchVerified100.01 / 18SWEComposite74.19 / 24SWERebench76.310 / 21SciCode78.76 / 21SonarBugDensity52.711 / 17SonarComposite54.211 / 24SonarFunctionalSkill78.97 / 17SonarIssueDensity13.212 / 17SonarVulnerabilityDensity58.210 / 17TTFT81.78 / 19Tau2Bench64.210 / 21TerminalBench48.312 / 22
sources arc_agiartificial_analysisgsolmarenamcp_atlasopenrouteroverridessonarswebenchswebench_proswerebenchterminal_benchmissing none
gpt-5.3-codexopenai57.757.754.754.764.264.271.2

group breakdown

A_B62.617 / 24A_I76.414 / 24A_P57.317 / 24A_R86.812 / 24BUILD66.210 / 24CRE51.411 / 24GEN50.311 / 24LM_ARENA_REVIEW_PROXY92.53 / 24OPS_long58.021 / 24OPS_precision60.620 / 24OPS_review64.121 / 24PLAN54.914 / 24

metrics

AI_code6.020 / 22AI_complexity33.218 / 22AI_context_awareness0.018 / 24AI_correctness94.113 / 22AI_edge_cases86.515 / 22AI_efficiency56.515 / 22AI_hallucination_resistance100.012 / 24AI_memory_retention0.022 / 24AI_parameter_accuracy91.46 / 24AI_plan_coherence0.122 / 24AI_recovery98.712 / 22AI_refusal100.014 / 22AI_spec100.014 / 22AI_stability92.77 / 22AI_task_completion66.715 / 24AI_tool_selection69.112 / 24BlendedCost76.614 / 24ContextWindow85.313 / 24CopilotArenaOrLMArenaCode59.314 / 22GDPval51.511 / 16GSO53.47 / 15LMArenaCreativeOrOpenEnded51.411 / 23LMArenaSearchDocument92.53 / 19LMArenaText51.411 / 23SWEBenchVerified92.58 / 18SWEComposite72.210 / 24SWERebench89.55 / 21SonarComposite50.018 / 24TerminalBench74.36 / 22
sources aistupidlevelartificial_analysislmarenaopenrouteroverridessonarswerebenchterminal_benchmissing BUILD/ArtificialAnalysisCodingBUILD/LongContextRecallBUILD/MCPAtlasBUILD/SciCodeGEN/ARC_AGI_2GEN/ArtificialAnalysisIntelligenceGEN/GPQA_HLE_ReasoningOPS_long/OutputSpeedOPS_long/TTFTOPS_precision/OutputSpeedOPS_precision/TTFTOPS_review/OutputSpeedOPS_review/TTFTPLAN/ArtificialAnalysisReasoningPLAN/IFBenchPLAN/LongContextRecallPLAN/MCPAtlasPLAN/Tau2BenchSWEComposite/SWEBenchMultilingualSWEComposite/SWEBenchProSonarComposite/SonarBugDensitySonarComposite/SonarFunctionalSkillSonarComposite/SonarIssueDensitySonarComposite/SonarVulnerabilityDensity
gpt-5.4openai45.145.144.044.061.561.549.1

group breakdown

A_B25.624 / 24A_I22.924 / 24A_P36.224 / 24A_R31.724 / 24BUILD73.78 / 24CRE51.610 / 24GEN38.616 / 24LM_ARENA_REVIEW_PROXY17.620 / 24OPS_long92.23 / 24OPS_precision88.67 / 24OPS_review90.03 / 24PLAN43.317 / 24

metrics

AI_code6.021 / 22AI_complexity0.221 / 22AI_context_awareness0.019 / 24AI_correctness0.022 / 22AI_edge_cases0.022 / 22AI_efficiency52.217 / 22AI_hallucination_resistance100.013 / 24AI_memory_retention0.023 / 24AI_parameter_accuracy96.65 / 24AI_plan_coherence5.416 / 24AI_recovery0.022 / 22AI_refusal100.015 / 22AI_spec100.015 / 22AI_stability0.921 / 22AI_task_completion100.07 / 24AI_tool_selection87.18 / 24ARC_AGI_275.85 / 17ArtificialAnalysisCoding33.715 / 21ArtificialAnalysisIntelligence27.416 / 21ArtificialAnalysisReasoning15.518 / 21BlendedCost75.015 / 24ContextWindow100.01 / 24CopilotArenaOrLMArenaCode68.011 / 22GDPval81.44 / 16GPQA_HLE_Reasoning15.518 / 21GSO54.06 / 15IFBench62.59 / 21LMArenaCreativeOrOpenEnded51.610 / 23LMArenaSearchDocument17.615 / 19LMArenaText51.610 / 23LongContextRecall24.518 / 21MCPAtlas72.84 / 13OutputSpeed95.83 / 19SWEBenchPro92.54 / 15SWEBenchVerified95.06 / 18SWEComposite88.95 / 24SWERebench83.57 / 21SciCode12.018 / 21SonarBugDensity84.74 / 17SonarComposite60.46 / 24SonarFunctionalSkill66.811 / 17SonarIssueDensity6.815 / 17SonarVulnerabilityDensity100.01 / 17TTFT85.16 / 19Tau2Bench0.021 / 21TerminalBench100.02 / 22
sources aistupidlevelarc_agiartificial_analysisgsolmarenamcp_atlasopenrouteroverridessonarswebench_proterminal_benchmissing SWEComposite/SWEBenchMultilingual
claude-sonnet-4.5anthropic42.842.846.946.957.457.453.7

group breakdown

A_B66.013 / 24A_I75.316 / 24A_P60.912 / 24A_R87.610 / 24BUILD52.513 / 24CRE23.920 / 24GEN32.218 / 24LM_ARENA_REVIEW_PROXY1.822 / 24OPS_long82.410 / 24OPS_precision81.411 / 24OPS_review83.511 / 24PLAN41.618 / 24

metrics

AI_canary_health80.36 / 7AI_code23.116 / 22AI_complexity33.213 / 22AI_context_awareness0.012 / 24AI_correctness94.18 / 22AI_edge_cases86.510 / 22AI_efficiency62.08 / 22AI_hallucination_resistance100.05 / 24AI_memory_retention0.016 / 24AI_parameter_accuracy85.711 / 24AI_plan_coherence0.120 / 24AI_recovery98.77 / 22AI_refusal100.06 / 22AI_spec100.06 / 22AI_stability83.018 / 22AI_task_completion100.04 / 24AI_tool_selection87.17 / 24ARC_AGI_23.712 / 17ArtificialAnalysisCoding45.312 / 21ArtificialAnalysisIntelligence46.012 / 21ArtificialAnalysisReasoning35.315 / 21BlendedCost74.417 / 24ContextWindow99.310 / 24CopilotArenaOrLMArenaCode53.416 / 22GDPval81.93 / 16GPQA_HLE_Reasoning35.315 / 21GSO27.311 / 15IFBench43.012 / 21LMArenaCreativeOrOpenEnded23.919 / 23LMArenaSearchDocument1.817 / 19LMArenaText23.919 / 23LongContextRecall65.711 / 21MCPAtlas6.612 / 13OutputSpeed80.816 / 19SWEBenchMultilingual3.95 / 6SWEBenchPro81.27 / 15SWEBenchVerified85.712 / 18SWEComposite71.612 / 24SWERebench74.911 / 21SciCode46.413 / 21SonarBugDensity2.816 / 17SonarComposite15.623 / 24SonarFunctionalSkill17.215 / 17SonarIssueDensity30.09 / 17SonarVulnerabilityDensity4.616 / 17TTFT78.412 / 19Tau2Bench58.911 / 21TerminalBench37.415 / 22
sources aistupidlevelarc_agiartificial_analysisgsolmarenamcp_atlasopenrouteroverridessonarswebenchswebench_proswerebenchterminal_benchmissing none
gpt-5.2openai47.447.455.455.455.655.658.7

group breakdown

A_B66.012 / 24A_I77.28 / 24A_P65.46 / 24A_R88.65 / 24BUILD50.814 / 24CRE31.917 / 24GEN43.213 / 24LM_ARENA_REVIEW_PROXY21.216 / 24OPS_long58.319 / 24OPS_precision61.319 / 24OPS_review64.820 / 24PLAN56.311 / 24

metrics

AI_code27.412 / 22AI_complexity33.217 / 22AI_context_awareness0.017 / 24AI_correctness94.112 / 22AI_edge_cases86.514 / 22AI_efficiency44.519 / 22AI_hallucination_resistance100.011 / 24AI_memory_retention0.021 / 24AI_parameter_accuracy91.17 / 24AI_plan_coherence23.810 / 24AI_recovery98.711 / 22AI_refusal100.013 / 22AI_spec100.013 / 22AI_stability88.916 / 22AI_task_completion100.06 / 24AI_tool_selection92.24 / 24ARC_AGI_20.017 / 17ArtificialAnalysisCoding63.48 / 21ArtificialAnalysisIntelligence59.79 / 21ArtificialAnalysisReasoning56.411 / 21BlendedCost80.111 / 24ContextWindow85.312 / 24CopilotArenaOrLMArenaCode38.720 / 22GPQA_HLE_Reasoning56.411 / 21GSO64.74 / 15IFBench64.78 / 21LMArenaCreativeOrOpenEnded31.916 / 23LMArenaSearchDocument21.211 / 19LMArenaText31.916 / 23LongContextRecall53.915 / 21SWEBenchMultilingual0.06 / 6SWEBenchPro38.214 / 15SWEBenchVerified81.314 / 18SWEComposite45.620 / 24SciCode54.69 / 21SonarBugDensity64.27 / 17SonarComposite59.77 / 24SonarFunctionalSkill67.210 / 17SonarIssueDensity35.78 / 17SonarVulnerabilityDensity73.46 / 17Tau2Bench50.115 / 21TerminalBench58.29 / 22
sources aistupidlevelarc_agiartificial_analysisgsolmarenamcp_atlasopenroutersonarswebenchswebench_proterminal_benchmissing BUILD/GDPvalBUILD/MCPAtlasOPS_long/OutputSpeedOPS_long/TTFTOPS_precision/OutputSpeedOPS_precision/TTFTOPS_review/OutputSpeedOPS_review/TTFTPLAN/MCPAtlasSWEComposite/SWERebench
grok-4-latestxai57.957.951.251.255.255.254.5

group breakdown

A_B74.86 / 24A_I81.26 / 24A_P64.57 / 24A_R87.99 / 24BUILD45.718 / 24CRE49.113 / 24GEN42.614 / 24LM_ARENA_REVIEW_PROXY19.219 / 24OPS_long80.215 / 24OPS_precision78.913 / 24OPS_review78.916 / 24PLAN43.816 / 24

metrics

AI_code64.46 / 22AI_complexity54.26 / 22AI_context_awareness0.021 / 24AI_correctness100.03 / 22AI_edge_cases58.719 / 22AI_efficiency0.022 / 22AI_hallucination_resistance100.015 / 24AI_memory_retention98.92 / 24AI_parameter_accuracy0.021 / 24AI_plan_coherence100.02 / 24AI_recovery87.618 / 22AI_refusal100.017 / 22AI_spec100.017 / 22AI_stability91.212 / 22AI_task_completion0.021 / 24AI_tool_selection0.021 / 24ARC_AGI_220.78 / 17ArtificialAnalysisCoding51.510 / 21ArtificialAnalysisIntelligence40.314 / 21ArtificialAnalysisReasoning57.010 / 21BlendedCost74.419 / 24ContextWindow78.415 / 24CopilotArenaOrLMArenaCode58.015 / 22GPQA_HLE_Reasoning57.010 / 21IFBench33.115 / 21LMArenaCreativeOrOpenEnded49.112 / 23LMArenaSearchDocument19.214 / 19LMArenaText49.112 / 23LongContextRecall77.08 / 21OutputSpeed82.113 / 19SWEComposite45.619 / 24SWERebench39.117 / 21SciCode51.910 / 21SonarComposite50.019 / 24TTFT79.09 / 19Tau2Bench51.514 / 21TerminalBench11.819 / 22
sources aistupidlevelarc_agiartificial_analysislmarenaopenrouterswerebenchterminal_benchmissing BUILD/GDPvalBUILD/GSOBUILD/MCPAtlasPLAN/MCPAtlasSWEComposite/SWEBenchMultilingualSWEComposite/SWEBenchProSWEComposite/SWEBenchVerifiedSonarComposite/SonarBugDensitySonarComposite/SonarFunctionalSkillSonarComposite/SonarIssueDensitySonarComposite/SonarVulnerabilityDensity
claude-opus-4.1anthropic30.730.746.646.652.152.150.9

group breakdown

A_B65.714 / 24A_I75.715 / 24A_P58.616 / 24A_R88.56 / 24BUILD48.517 / 24CRE0.723 / 24GEN37.717 / 24LM_ARENA_REVIEW_PROXY0.023 / 24OPS_long48.722 / 24OPS_precision43.724 / 24OPS_review46.223 / 24PLAN45.915 / 24

metrics

AI_canary_health68.17 / 7AI_code23.113 / 22AI_complexity33.28 / 22AI_context_awareness0.07 / 24AI_correctness94.14 / 22AI_edge_cases86.56 / 22AI_efficiency48.418 / 22AI_hallucination_resistance100.01 / 24AI_memory_retention0.011 / 24AI_parameter_accuracy71.020 / 24AI_plan_coherence0.119 / 24AI_recovery98.73 / 22AI_refusal100.01 / 22AI_spec100.01 / 22AI_stability92.73 / 22AI_task_completion83.39 / 24AI_tool_selection83.211 / 24BlendedCost0.024 / 24ContextWindow74.720 / 24CopilotArenaOrLMArenaCode53.217 / 22LMArenaCreativeOrOpenEnded0.722 / 23LMArenaSearchDocument0.018 / 19LMArenaText0.722 / 23SWEComposite50.916 / 24SWERebench52.316 / 21SonarComposite50.014 / 24TerminalBench29.416 / 22
sources aistupidlevellmarenaopenrouterswerebenchterminal_benchmissing BUILD/ArtificialAnalysisCodingBUILD/GDPvalBUILD/GSOBUILD/LongContextRecallBUILD/MCPAtlasBUILD/SciCodeGEN/ARC_AGI_2GEN/ArtificialAnalysisIntelligenceGEN/GPQA_HLE_ReasoningOPS_long/OutputSpeedOPS_long/TTFTOPS_precision/OutputSpeedOPS_precision/TTFTOPS_review/OutputSpeedOPS_review/TTFTPLAN/ArtificialAnalysisReasoningPLAN/IFBenchPLAN/LongContextRecallPLAN/MCPAtlasPLAN/Tau2BenchSWEComposite/SWEBenchMultilingualSWEComposite/SWEBenchProSWEComposite/SWEBenchVerifiedSonarComposite/SonarBugDensitySonarComposite/SonarFunctionalSkillSonarComposite/SonarIssueDensitySonarComposite/SonarVulnerabilityDensity
glm-4.7zai41.641.652.352.351.651.655.2

group breakdown

A_B56.020 / 24A_I55.020 / 24A_P46.920 / 24A_R58.520 / 24BUILD44.519 / 24CRE26.819 / 24GEN40.715 / 24LM_ARENA_REVIEW_PROXY50.012 / 24OPS_long88.85 / 24OPS_precision91.14 / 24OPS_review88.84 / 24PLAN55.512 / 24

metrics

AI_context_awareness0.024 / 24AI_hallucination_resistance100.018 / 24AI_memory_retention98.95 / 24AI_parameter_accuracy0.024 / 24AI_plan_coherence100.05 / 24AI_task_completion0.024 / 24AI_tool_selection0.024 / 24ArtificialAnalysisCoding37.914 / 21ArtificialAnalysisIntelligence42.613 / 21ArtificialAnalysisReasoning55.812 / 21BlendedCost96.13 / 24ContextWindow74.918 / 24CopilotArenaOrLMArenaCode68.89 / 22GPQA_HLE_Reasoning55.812 / 21IFBench72.27 / 21LMArenaCreativeOrOpenEnded26.818 / 23LMArenaText26.818 / 23LongContextRecall57.414 / 21MCPAtlas0.013 / 13OutputSpeed88.07 / 19SWEComposite58.415 / 24SWERebench70.913 / 21SciCode48.611 / 21SonarBugDensity51.613 / 17SonarComposite27.321 / 24SonarFunctionalSkill0.017 / 17SonarIssueDensity50.85 / 17SonarVulnerabilityDensity28.713 / 17TTFT97.83 / 19Tau2Bench100.02 / 21TerminalBench27.117 / 22
sources aistupidlevelartificial_analysislmarenamcp_atlasopenroutersonarswerebenchterminal_benchmissing A_B/AI_codeA_B/AI_complexityA_B/AI_correctnessA_B/AI_edge_casesA_B/AI_efficiencyA_B/AI_recoveryA_B/AI_specA_B/AI_stabilityA_I/AI_complexityA_I/AI_correctnessA_I/AI_edge_casesA_I/AI_efficiencyA_I/AI_recoveryA_I/AI_specA_I/AI_stabilityA_P/AI_correctnessA_P/AI_efficiencyA_P/AI_recoveryA_P/AI_specA_P/AI_stabilityA_R/AI_codeA_R/AI_correctnessA_R/AI_edge_casesA_R/AI_recoveryA_R/AI_specA_R/AI_stabilityBUILD/GDPvalBUILD/GSOGEN/ARC_AGI_2LM_ARENA_REVIEW_PROXY/LMArenaSearchDocumentSWEComposite/SWEBenchMultilingualSWEComposite/SWEBenchProSWEComposite/SWEBenchVerified
gemini-2.5-progoogle63.663.643.943.951.251.243.6

group breakdown

A_B85.13 / 24A_I92.13 / 24A_P71.93 / 24A_R78.315 / 24BUILD35.921 / 24CRE60.68 / 24GEN29.719 / 24LM_ARENA_REVIEW_PROXY0.024 / 24OPS_long88.76 / 24OPS_precision84.29 / 24OPS_review87.07 / 24PLAN28.920 / 24

metrics

AI_code89.83 / 22AI_complexity92.53 / 22AI_context_awareness7.53 / 24AI_correctness92.515 / 22AI_edge_cases92.53 / 22AI_efficiency89.03 / 22AI_hallucination_resistance10.920 / 24AI_memory_retention92.56 / 24AI_parameter_accuracy78.116 / 24AI_plan_coherence92.56 / 24AI_recovery92.514 / 22AI_refusal92.518 / 22AI_spec92.518 / 22AI_stability92.09 / 22AI_task_completion27.417 / 24AI_tool_selection11.617 / 24ARC_AGI_23.713 / 17ArtificialAnalysisCoding23.617 / 21ArtificialAnalysisIntelligence14.117 / 21ArtificialAnalysisReasoning44.814 / 21BlendedCost80.110 / 24ContextWindow100.04 / 24CopilotArenaOrLMArenaCode0.921 / 22GDPval7.515 / 16GPQA_HLE_Reasoning44.814 / 21GSO0.015 / 15IFBench19.318 / 21LMArenaCreativeOrOpenEnded60.68 / 23LMArenaSearchDocument0.019 / 19LMArenaText60.68 / 23LongContextRecall67.210 / 21MCPAtlas71.15 / 13OutputSpeed93.25 / 19SWEBenchPro75.711 / 15SWEBenchVerified38.217 / 18SWEComposite36.622 / 24SWERebench1.820 / 21SciCode36.115 / 21SonarBugDensity52.710 / 17SonarComposite54.210 / 24SonarFunctionalSkill78.96 / 17SonarIssueDensity13.211 / 17SonarVulnerabilityDensity58.29 / 17TTFT72.217 / 19Tau2Bench3.519 / 21TerminalBench1.820 / 22
sources arc_agiartificial_analysisgsolmarenaopenrouterswebenchswerebenchterminal_benchmissing SWEComposite/SWEBenchMultilingual
claude-sonnet-4anthropic35.335.335.735.750.250.253.3

group breakdown

A_B47.021 / 24A_I46.821 / 24A_P46.621 / 24A_R57.021 / 24BUILD49.416 / 24CRE27.818 / 24GEN21.020 / 24LM_ARENA_REVIEW_PROXY86.26 / 24OPS_long82.411 / 24OPS_precision81.112 / 24OPS_review83.312 / 24PLAN30.119 / 24

metrics

AI_code23.115 / 22AI_complexity33.212 / 22AI_context_awareness0.011 / 24AI_correctness43.419 / 22AI_edge_cases35.920 / 22AI_efficiency58.413 / 22AI_hallucination_resistance100.04 / 24AI_memory_retention0.015 / 24AI_parameter_accuracy89.79 / 24AI_plan_coherence2.818 / 24AI_recovery48.020 / 22AI_refusal100.05 / 22AI_spec100.05 / 22AI_stability25.620 / 22AI_task_completion100.03 / 24AI_tool_selection85.89 / 24ARC_AGI_20.216 / 17ArtificialAnalysisCoding30.716 / 21ArtificialAnalysisIntelligence29.715 / 21ArtificialAnalysisReasoning8.619 / 21BlendedCost74.416 / 24ContextWindow99.39 / 24CopilotArenaOrLMArenaCode52.918 / 22GDPval80.15 / 16GPQA_HLE_Reasoning8.619 / 21GSO6.014 / 15IFBench35.814 / 21LMArenaCreativeOrOpenEnded27.817 / 23LMArenaSearchDocument86.26 / 19LMArenaText27.817 / 23LiveCodeBench0.02 / 2LongContextRecall60.812 / 21MCPAtlas13.110 / 13OutputSpeed81.215 / 19SWEBenchPro78.49 / 15SWEBenchVerified69.916 / 18SWEComposite66.613 / 24SWERebench55.115 / 21SciCode20.817 / 21SonarBugDensity0.017 / 17SonarComposite19.522 / 24SonarFunctionalSkill26.414 / 17SonarIssueDensity35.87 / 17SonarVulnerabilityDensity0.017 / 17TTFT77.013 / 19Tau2Bench27.718 / 21TerminalBench47.413 / 22
sources aistupidlevelarc_agiartificial_analysisgsolivecodebenchopenroutersonarswebenchswebench_proswerebenchmissing SWEComposite/SWEBenchMultilingual
deepseek-v4-flashdeepseek30.030.062.762.749.449.456.4

group breakdown

A_B27.423 / 24A_I25.023 / 24A_P42.922 / 24A_R32.023 / 24BUILD49.715 / 24CRE13.121 / 24GEN49.412 / 24LM_ARENA_REVIEW_PROXY50.08 / 24OPS_long88.28 / 24OPS_precision91.33 / 24OPS_review88.65 / 24PLAN77.65 / 24

metrics

AI_canary_health83.44 / 7AI_code10.319 / 22AI_complexity0.220 / 22AI_context_awareness47.92 / 24AI_correctness0.021 / 22AI_edge_cases0.021 / 22AI_efficiency64.36 / 22AI_hallucination_resistance100.07 / 24AI_memory_retention0.018 / 24AI_parameter_accuracy97.14 / 24AI_plan_coherence18.512 / 24AI_recovery0.021 / 22AI_refusal100.08 / 22AI_spec100.08 / 22AI_stability0.022 / 22AI_task_completion83.311 / 24AI_tool_selection100.01 / 24ArtificialAnalysisCoding45.611 / 21ArtificialAnalysisIntelligence59.310 / 21ArtificialAnalysisReasoning76.77 / 21BlendedCost100.01 / 24ContextWindow71.622 / 24GPQA_HLE_Reasoning76.77 / 21IFBench100.01 / 21LMArenaCreativeOrOpenEnded13.120 / 23LMArenaText13.120 / 23LongContextRecall52.516 / 21OutputSpeed86.89 / 19SWEComposite50.017 / 24SciCode47.512 / 21SonarComposite50.015 / 24TTFT98.62 / 19Tau2Bench97.95 / 21
sources aistupidlevelartificial_analysislmarenaopenroutermissing BUILD/CopilotArenaOrLMArenaCodeBUILD/GDPvalBUILD/GSOBUILD/MCPAtlasBUILD/TerminalBenchGEN/ARC_AGI_2LM_ARENA_REVIEW_PROXY/LMArenaSearchDocumentPLAN/MCPAtlasPLAN/TerminalBenchSWEComposite/SWEBenchMultilingualSWEComposite/SWEBenchProSWEComposite/SWEBenchVerifiedSWEComposite/SWERebenchSonarComposite/SonarBugDensitySonarComposite/SonarFunctionalSkillSonarComposite/SonarIssueDensitySonarComposite/SonarVulnerabilityDensity
gemini-2.5-flashgoogle31.331.334.434.447.147.153.2

group breakdown

A_B98.01 / 24A_I95.32 / 24A_P79.31 / 24A_R100.01 / 24BUILD24.523 / 24CRE0.024 / 24GEN3.724 / 24LM_ARENA_REVIEW_PROXY78.87 / 24OPS_long95.21 / 24OPS_precision91.52 / 24OPS_review93.61 / 24PLAN17.223 / 24

metrics

AI_code100.01 / 22AI_complexity99.62 / 22AI_context_awareness100.01 / 24AI_correctness100.01 / 22AI_edge_cases100.01 / 22AI_efficiency100.01 / 22AI_hallucination_resistance100.08 / 24AI_memory_retention50.99 / 24AI_parameter_accuracy100.01 / 24AI_plan_coherence53.89 / 24AI_recovery100.01 / 22AI_refusal100.09 / 22AI_spec100.09 / 22AI_stability100.01 / 22AI_task_completion29.716 / 24AI_tool_selection59.616 / 24ARC_AGI_20.815 / 17ArtificialAnalysisCoding0.020 / 21ArtificialAnalysisIntelligence0.819 / 21ArtificialAnalysisReasoning17.916 / 21BlendedCost94.45 / 24ContextWindow100.03 / 24CopilotArenaOrLMArenaCode65.313 / 22GDPval10.313 / 16GPQA_HLE_Reasoning17.916 / 21GSO19.412 / 15IFBench29.217 / 21LMArenaCreativeOrOpenEnded0.023 / 23LMArenaSearchDocument78.87 / 19LMArenaText0.023 / 23LiveCodeBench100.01 / 2LongContextRecall58.813 / 21MCPAtlas26.68 / 13OutputSpeed100.01 / 19SWEBenchPro52.513 / 15SWEBenchVerified0.018 / 18SWEComposite20.424 / 24SWERebench0.021 / 21SciCode23.516 / 21SonarComposite50.016 / 24TTFT78.910 / 19Tau2Bench0.020 / 21TerminalBench0.321 / 22
sources aistupidlevelarc_agiartificial_analysislivecodebenchlmarenaopenrouterswebenchswerebenchterminal_benchmissing SWEComposite/SWEBenchMultilingualSonarComposite/SonarBugDensitySonarComposite/SonarFunctionalSkillSonarComposite/SonarIssueDensitySonarComposite/SonarVulnerabilityDensity
kimi-k2-0905moonshot46.946.933.033.046.646.649.3

group breakdown

A_B62.716 / 24A_I69.618 / 24A_P56.718 / 24A_R87.211 / 24BUILD41.120 / 24CRE50.012 / 24GEN20.021 / 24LM_ARENA_REVIEW_PROXY50.09 / 24OPS_long35.324 / 24OPS_precision58.121 / 24OPS_review54.422 / 24PLAN22.721 / 24

metrics

AI_canary_health88.91 / 7AI_code31.78 / 22AI_complexity33.215 / 22AI_context_awareness0.015 / 24AI_correctness94.110 / 22AI_edge_cases86.512 / 22AI_efficiency8.520 / 22AI_hallucination_resistance100.09 / 24AI_memory_retention0.019 / 24AI_parameter_accuracy82.514 / 24AI_plan_coherence5.415 / 24AI_recovery98.79 / 22AI_refusal100.011 / 22AI_spec100.011 / 22AI_stability71.219 / 22AI_task_completion83.312 / 24AI_tool_selection83.210 / 24ArtificialAnalysisCoding4.219 / 21ArtificialAnalysisIntelligence0.020 / 21ArtificialAnalysisReasoning0.020 / 21BlendedCost92.77 / 24ContextWindow53.423 / 24GPQA_HLE_Reasoning0.020 / 21IFBench0.020 / 21LongContextRecall0.020 / 21OutputSpeed0.019 / 19SWEComposite50.018 / 24SciCode0.020 / 21SonarComposite50.017 / 24TTFT90.25 / 19Tau2Bench48.016 / 21
sources aistupidlevelartificial_analysisopenroutermissing BUILD/CopilotArenaOrLMArenaCodeBUILD/GDPvalBUILD/GSOBUILD/MCPAtlasBUILD/TerminalBenchCRE/LMArenaCreativeOrOpenEndedCRE/LMArenaTextGEN/ARC_AGI_2GEN/LMArenaTextLM_ARENA_REVIEW_PROXY/LMArenaSearchDocumentPLAN/MCPAtlasPLAN/TerminalBenchSWEComposite/SWEBenchMultilingualSWEComposite/SWEBenchProSWEComposite/SWEBenchVerifiedSWEComposite/SWERebenchSonarComposite/SonarBugDensitySonarComposite/SonarFunctionalSkillSonarComposite/SonarIssueDensitySonarComposite/SonarVulnerabilityDensity
glm-4.6zai38.338.329.929.934.934.938.2

group breakdown

A_B56.019 / 24A_I55.019 / 24A_P46.919 / 24A_R58.519 / 24BUILD21.924 / 24CRE32.916 / 24GEN16.222 / 24LM_ARENA_REVIEW_PROXY50.011 / 24OPS_long80.116 / 24OPS_precision86.08 / 24OPS_review83.810 / 24PLAN18.122 / 24

metrics

AI_context_awareness0.023 / 24AI_hallucination_resistance100.017 / 24AI_memory_retention98.94 / 24AI_parameter_accuracy0.023 / 24AI_plan_coherence100.04 / 24AI_task_completion0.023 / 24AI_tool_selection0.023 / 24ArtificialAnalysisCoding15.918 / 21ArtificialAnalysisIntelligence6.118 / 21ArtificialAnalysisReasoning16.517 / 21BlendedCost95.44 / 24ContextWindow75.017 / 24CopilotArenaOrLMArenaCode44.419 / 22GPQA_HLE_Reasoning16.517 / 21IFBench4.719 / 21LMArenaCreativeOrOpenEnded32.915 / 23LMArenaText32.915 / 23LongContextRecall9.819 / 21MCPAtlas7.511 / 13OutputSpeed72.618 / 19SWEBenchPro0.015 / 15SWEBenchVerified79.015 / 18SWEComposite30.223 / 24SWERebench38.418 / 21SciCode12.019 / 21SonarBugDensity7.515 / 17SonarComposite10.724 / 24SonarFunctionalSkill7.516 / 17SonarIssueDensity7.514 / 17SonarVulnerabilityDensity29.012 / 17TTFT96.74 / 19Tau2Bench41.317 / 21TerminalBench13.918 / 22
sources aistupidlevelartificial_analysislmarenaopenrouterswebenchswebench_proswerebenchterminal_benchmissing A_B/AI_codeA_B/AI_complexityA_B/AI_correctnessA_B/AI_edge_casesA_B/AI_efficiencyA_B/AI_recoveryA_B/AI_specA_B/AI_stabilityA_I/AI_complexityA_I/AI_correctnessA_I/AI_edge_casesA_I/AI_efficiencyA_I/AI_recoveryA_I/AI_specA_I/AI_stabilityA_P/AI_correctnessA_P/AI_efficiencyA_P/AI_recoveryA_P/AI_specA_P/AI_stabilityA_R/AI_codeA_R/AI_correctnessA_R/AI_edge_casesA_R/AI_recoveryA_R/AI_specA_R/AI_stabilityBUILD/GDPvalBUILD/GSOGEN/ARC_AGI_2LM_ARENA_REVIEW_PROXY/LMArenaSearchDocumentSWEComposite/SWEBenchMultilingual
grok-code-fast-1xai22.922.923.823.833.933.938.4

group breakdown

A_B35.022 / 24A_I46.122 / 24A_P40.623 / 24A_R55.922 / 24BUILD28.322 / 24CRE7.522 / 24GEN5.623 / 24LM_ARENA_REVIEW_PROXY50.010 / 24OPS_long89.54 / 24OPS_precision88.76 / 24OPS_review88.26 / 24PLAN13.324 / 24

metrics

AI_code0.022 / 22AI_complexity0.022 / 22AI_context_awareness0.022 / 24AI_correctness35.320 / 22AI_edge_cases79.618 / 22AI_efficiency1.921 / 22AI_hallucination_resistance100.016 / 24AI_memory_retention98.93 / 24AI_parameter_accuracy0.022 / 24AI_plan_coherence100.03 / 24AI_recovery73.319 / 22AI_refusal0.022 / 22AI_spec0.022 / 22AI_stability90.013 / 22AI_task_completion0.022 / 24AI_tool_selection0.022 / 24ARC_AGI_225.17 / 17ArtificialAnalysisCoding0.021 / 21ArtificialAnalysisIntelligence0.021 / 21ArtificialAnalysisReasoning0.021 / 21BlendedCost99.32 / 24ContextWindow78.416 / 24CopilotArenaOrLMArenaCode0.022 / 22GPQA_HLE_Reasoning0.021 / 21IFBench0.021 / 21LMArenaCreativeOrOpenEnded7.521 / 23LMArenaText7.521 / 23LongContextRecall0.021 / 21OutputSpeed93.16 / 19SWEComposite41.221 / 24SWERebench27.919 / 21SciCode0.021 / 21SonarComposite50.020 / 24TTFT83.27 / 19Tau2Bench53.313 / 21TerminalBench0.022 / 22
sources aistupidlevelartificial_analysislmarenaopenrouterswerebenchterminal_benchmissing BUILD/GDPvalBUILD/GSOBUILD/MCPAtlasLM_ARENA_REVIEW_PROXY/LMArenaSearchDocumentPLAN/MCPAtlasSWEComposite/SWEBenchMultilingualSWEComposite/SWEBenchProSWEComposite/SWEBenchVerifiedSonarComposite/SonarBugDensitySonarComposite/SonarFunctionalSkillSonarComposite/SonarIssueDensitySonarComposite/SonarVulnerabilityDensity