MODELLEDGERML Verified cut · 21 Jul 2026

Independent AI model intelligence

The scoreboard, without the sales pitch.

Human preference, API prices and subscription plans in one traceable view. Every important number carries a source and date.

Models indexed9
Top model battles14,646
Price sources8
Leaderboard licenseCC BY 4.0

01 / Human preference

Text Arena leaderboard

Style-controlled Bradley–Terry ratings from anonymous head-to-head votes. Confidence intervals show where ranks may overlap.

Verified fallback · live source unavailable
Arena text, style-control snapshot.
RankModelProviderArena score95% intervalBattlesLicense
01Claude Fable 5claude-fable-5Anthropic1507.31500.91513.714,646Proprietary
10Kimi K3kimi-k3Moonshot1485.71475.71495.63,619Proprietary
11GPT-5.6 Sol (xHigh)gpt-5-6-sol-xhighOpenAI1485.11477.21493.06,221Proprietary
17Gemini 3.5 Flash (High)gemini-3-5-flash-highGoogle1476.31469.71482.810,092Proprietary
19Qwen3.7 Max Previewqwen3.7-max-previewAlibaba1475.21465.11485.23,714Proprietary
38Claude Sonnet 5 (High)claude-sonnet-5-highAnthropic1461.01455.01467.013,521Proprietary
46DeepSeek V4 Prodeepseek-v4-proDeepSeek1457.01453.01461.045,278MIT
69Grok 4.3grok-4-3xAI1442.01438.01446.046,389Proprietary
87Mistral Medium 3.5mistral-medium-3-5Mistral1427.01420.01434.011,019Modified MIT
Verified fallback loaded · refreshes on every visit and every 30 minutesInspect source ↗

Coverage watch / unranked

Qwen3.8 Max PreviewAlibaba · released 19 Jul 2026

Available in Qwen Cloud, but not yet present in Arena's current text or code datasets. No rank is invented.

Provider source ↗

02 / API economics · Verified 22 Jul 2026

What your workload costs

Set a simple monthly workload. The estimate uses standard text input and output rates only—tools, caching, tax and batch discounts are excluded.

Monthly scenario10,000 requests

Prices shown in USD per 1M tokens. API usage is separate from consumer subscriptions.

Lowest estimated standard cost

DeepSeek V4 Pro

DeepSeek
$6.09/mo
  1. Grok 4.3$17.50
  2. Mistral Medium 3.5$34.50
  3. Gemini 3.5 Flash$39.00
  4. Claude Sonnet 5$46.00

Cost is not a quality ranking. Choose for the task, then optimize price.

Standard text API list prices, verified 22 July 2026.
ModelProviderInput / 1MOutput / 1MScenario / moContextSource
DeepSeek V4 Prodeepseek-v4-proDeepSeek$0.435$0.870$6.091MOfficial ↗
Grok 4.3Higher rates apply to prompts of 200K+ tokensxAI$1.25$2.50$17.501MOfficial ↗
Mistral Medium 3.5mistral-medium-3-5Mistral$1.50$7.50$34.50256KOfficial ↗
Gemini 3.5 Flashgemini-3.5-flashGoogle$1.50$9.00$39.001.05MOfficial ↗
Claude Sonnet 5Introductory price through 31 Aug 2026Anthropic$2.00$10.00$46.001MOfficial ↗
Kimi K3Live catalog rate; cache pricing excludedMoonshot$3.00$15.00$69.001.05MOfficial ↗
GPT-5.6 Sol>272K input triggers long-context multipliersOpenAI$5.00$30.00$130.001.05MOfficial ↗
Claude Fable 5claude-fable-5Anthropic$10.00$50.00$230.001MOfficial ↗

03 / Consumer plans

Subscriptions are not API credit

These plans buy access to a provider’s application. They usually do not include metered developer API usage.

OpenAI$20 / monthly

ChatGPT Plus

General work, research and coding

Usage limits apply; API billed separatelyVerify ↗
Anthropic$20 / monthly

Claude Pro

Writing, coding and long-form analysis

Google$19.99 / monthly · US

Google AI Pro

Gemini plus Google apps and storage

Alibaba$6 / monthly · launch promotion

Qwen Token Plan Lite

Qwen3.8 Max Preview and coding agents

Credit quotas apply; standard list price is $8Verify ↗
xAI$30 / monthly

SuperGrok

Higher Grok access outside the API

Mistral$14.99 / monthly

Vibe Pro

Chat, research and agentic work

04 / Field notes

AI news with consequences

Product launches, price changes and research—summarized briefly, linked to the original source, and stripped of launch-day hype.

22 Jul 2026ModelsOpenAI

OpenAI launches Presence for governed enterprise agents

Presence brings voice and chat agents into a managed enterprise product with a limited general-availability rollout.

Read source
21 Jul 2026ModelsGoogle

Google announces Gemini 3.6 Flash and two 3.5 variants

A broader Flash family targets general speed, lower-cost inference and specialist cyber workloads.

Read source
19 Jul 2026ModelsQwen Cloud

Qwen3.8 Max Preview enters paid preview

Alibaba's newest Max model is available through Token Plan, but it does not yet have an independent Arena rank. Model Ledger keeps it visible without inventing a score.

Read source
16 Jul 2026ModelsKimi

Moonshot releases Kimi K3

Kimi K3 arrives with native vision and a one-million-token context window—and now appears automatically wherever Arena publishes a score.

Read source
16 Jul 2026ModelsxAI

xAI launches Grok 4.5

The release is positioned around coding, agentic tasks and knowledge work, adding another top-tier option to compare.

Read source
9 Jul 2026BusinessMistral

Mistral adds prompt and skill version control to Studio

Teams can now manage prompts and reusable skills as versioned assets rather than loose production configuration.

Read source
1 Jul 2026PolicyAnthropic

Anthropic redeploys Fable 5 and Mythos 5

The models return after export restrictions were lifted, making policy a direct input into model availability.

Read source

Methodology note 001

Trust the trail, not the badge.

Model Ledger loads Arena's current top 100 text and code rows on every visit and refreshes them every 30 minutes. Exact variants are preserved, missing results never become zero, and unranked new releases stay visibly unranked. Arena data is attributed under CC BY 4.0.

Prices are stored with their source, currency and verification date. A detected change should be reviewed before it replaces the last known valid price. Sponsored placements never affect rank.