Available in Qwen Cloud, but not yet present in Arena's current text or code datasets. No rank is invented.
Provider source ↗Independent AI model intelligence
The scoreboard, without the sales pitch.
Human preference, API prices and subscription plans in one traceable view. Every important number carries a source and date.
01 / Human preference
Text Arena leaderboard
Style-controlled Bradley–Terry ratings from anonymous head-to-head votes. Confidence intervals show where ranks may overlap.
| Rank | Model | Provider | Arena score | 95% interval | Battles | License |
|---|---|---|---|---|---|---|
| 01 | Claude Fable 5claude-fable-5 | Anthropic | 1507.3 | 1500.9–1513.7 | 14,646 | Proprietary |
| 10 | Kimi K3kimi-k3 | Moonshot | 1485.7 | 1475.7–1495.6 | 3,619 | Proprietary |
| 11 | GPT-5.6 Sol (xHigh)gpt-5-6-sol-xhigh | OpenAI | 1485.1 | 1477.2–1493.0 | 6,221 | Proprietary |
| 17 | Gemini 3.5 Flash (High)gemini-3-5-flash-high | 1476.3 | 1469.7–1482.8 | 10,092 | Proprietary | |
| 19 | Qwen3.7 Max Previewqwen3.7-max-preview | Alibaba | 1475.2 | 1465.1–1485.2 | 3,714 | Proprietary |
| 38 | Claude Sonnet 5 (High)claude-sonnet-5-high | Anthropic | 1461.0 | 1455.0–1467.0 | 13,521 | Proprietary |
| 46 | DeepSeek V4 Prodeepseek-v4-pro | DeepSeek | 1457.0 | 1453.0–1461.0 | 45,278 | MIT |
| 69 | Grok 4.3grok-4-3 | xAI | 1442.0 | 1438.0–1446.0 | 46,389 | Proprietary |
| 87 | Mistral Medium 3.5mistral-medium-3-5 | Mistral | 1427.0 | 1420.0–1434.0 | 11,019 | Modified MIT |
Coverage watch / unranked
02 / API economics · Verified 22 Jul 2026
What your workload costs
Set a simple monthly workload. The estimate uses standard text input and output rates only—tools, caching, tax and batch discounts are excluded.
Lowest estimated standard cost
DeepSeek V4 Pro
DeepSeek- Grok 4.3$17.50
- Mistral Medium 3.5$34.50
- Gemini 3.5 Flash$39.00
- Claude Sonnet 5$46.00
Cost is not a quality ranking. Choose for the task, then optimize price.
| Model | Provider | Input / 1M | Output / 1M | Scenario / mo | Context | Source |
|---|---|---|---|---|---|---|
| DeepSeek V4 Prodeepseek-v4-pro | DeepSeek | $0.435 | $0.870 | $6.09 | 1M | Official ↗ |
| Grok 4.3Higher rates apply to prompts of 200K+ tokens | xAI | $1.25 | $2.50 | $17.50 | 1M | Official ↗ |
| Mistral Medium 3.5mistral-medium-3-5 | Mistral | $1.50 | $7.50 | $34.50 | 256K | Official ↗ |
| Gemini 3.5 Flashgemini-3.5-flash | $1.50 | $9.00 | $39.00 | 1.05M | Official ↗ | |
| Claude Sonnet 5Introductory price through 31 Aug 2026 | Anthropic | $2.00 | $10.00 | $46.00 | 1M | Official ↗ |
| Kimi K3Live catalog rate; cache pricing excluded | Moonshot | $3.00 | $15.00 | $69.00 | 1.05M | Official ↗ |
| GPT-5.6 Sol>272K input triggers long-context multipliers | OpenAI | $5.00 | $30.00 | $130.00 | 1.05M | Official ↗ |
| Claude Fable 5claude-fable-5 | Anthropic | $10.00 | $50.00 | $230.00 | 1M | Official ↗ |
03 / Consumer plans
Subscriptions are not API credit
These plans buy access to a provider’s application. They usually do not include metered developer API usage.
ChatGPT Plus
General work, research and coding
Claude Pro
Writing, coding and long-form analysis
Google AI Pro
Gemini plus Google apps and storage
Qwen Token Plan Lite
Qwen3.8 Max Preview and coding agents
SuperGrok
Higher Grok access outside the API
Vibe Pro
Chat, research and agentic work
04 / Field notes
AI news with consequences
Product launches, price changes and research—summarized briefly, linked to the original source, and stripped of launch-day hype.
OpenAI launches Presence for governed enterprise agents
Presence brings voice and chat agents into a managed enterprise product with a limited general-availability rollout.
Google announces Gemini 3.6 Flash and two 3.5 variants
A broader Flash family targets general speed, lower-cost inference and specialist cyber workloads.
Qwen3.8 Max Preview enters paid preview
Alibaba's newest Max model is available through Token Plan, but it does not yet have an independent Arena rank. Model Ledger keeps it visible without inventing a score.
Moonshot releases Kimi K3
Kimi K3 arrives with native vision and a one-million-token context window—and now appears automatically wherever Arena publishes a score.
xAI launches Grok 4.5
The release is positioned around coding, agentic tasks and knowledge work, adding another top-tier option to compare.
Mistral adds prompt and skill version control to Studio
Teams can now manage prompts and reusable skills as versioned assets rather than loose production configuration.
Anthropic redeploys Fable 5 and Mythos 5
The models return after export restrictions were lifted, making policy a direct input into model availability.
Methodology note 001
Trust the trail, not the badge.
Model Ledger loads Arena's current top 100 text and code rows on every visit and refreshes them every 30 minutes. Exact variants are preserved, missing results never become zero, and unranked new releases stay visibly unranked. Arena data is attributed under CC BY 4.0.
Prices are stored with their source, currency and verification date. A detected change should be reviewed before it replaces the last known valid price. Sponsored placements never affect rank.