~/coding-agent-pricing

Coding Agent Pricing

Which coding-agent subscription, model and reasoning effort to use — and how long your quota lasts.

Built for subscriptions (Claude Pro/Max, ChatGPT Plus/Pro for Codex, Google AI Pro/Ultra for Antigravity, SuperGrok, Kimi, GLM Coding Plan), not API keys. API prices appear only because quotas burn in proportion to them.

Data checked on September 28, 2026 · 96 sources

What to use, when

Pick what you are about to do: you get the provider, model, reasoning effort and subscription to use, with the evidence behind the choice.

What do you want to do?

Everyday coding

Features, bug fixes, reviews, small refactors — the work you do all day.

  • Fix a failing test or a reported bug
  • Add an endpoint, a form, a CLI option
  • Review a pull request
  • Small refactor in a few files
Use
Claude Opus 5.5· Anthropic
Effort
Medium
Agent
Claude Code
Subscription
Claude Pro ($20) or Max ($100 / $200)

How to work

  1. No plan needed when you can describe the change in one sentence; otherwise press Shift+Tab to enter plan mode first.
  2. Stay on Medium (the default). For a stubborn bug, raise it with /effort high, or add ultrathink to a single prompt.
  3. Give the agent a way to check itself (tests, build, lint) and run /clear between unrelated tasks.
  4. In Codex: /model → GPT-6 Sol, effort High; structure the prompt as Goal, Context, Constraints, Done when.

Why

  • Best quality per unit of quota: 51.2 on the AA Intelligence Index for $1.34 per task, ahead of GPT-6 Astra at High (50.9 for $1.73).
  • Anthropic's own SWE-bench Pro subset: 92.8% solved at $0.22 per solved task, versus 92.3% at $1.19 for Fable 5.1 at its default.
  • It is Claude Code's default model and effort since 2026-09-22. Step up to High for a hard bug: +4 points on Terminal-Bench 4.0 (52.5% → 56.6%) for about 27% more quota.

Alternatives

  • GPT-6 Sol · HighCodex · ChatGPT Plus ($20) or Pro ($100 / $200)Codex on ChatGPT Plus: DeepSWE 65.3% at $0.64 per task — +8.6 points over Medium for +$0.26; Plus allows about 15–150 Sol messages per 5 hours.
  • GLM-5.3 · MaxClaude Code or ZCode · GLM Coding Plan ($18 / $80 / $168)Budget: GLM Coding Plan Lite ($18/month, 2,000 credits per 5 hours) inside Claude Code; DeepSWE 69.0% at max.

Sources: Artificial Analysis model leaderboard · Optimizing for cost and intelligence · Model configuration · DeepSWE v1.1 leaderboard · Introducing GPT-6 Sol and Luna · Codex pricing · GLM Coding Plan

Not worth it right now

  • Claude Fable 5.1 — Dominated by Opus 5.5 on every independent benchmark, at 2.5–5× the cost: AA index 53.4 at Max for $7.63 vs Opus 5.5 High 53.6 for $1.82. On Max and Team Premium it is also capped at 50% of the weekly limit.
  • Claude Sonnet 5 — Opus 5.5 at Low beats Sonnet 5 at Max on the AA index (42.3 vs 38.2) for a ninth of the cost ($0.55 vs $5.09), and Opus 5.5 is now Claude Code's default.
  • GPT-6 Astra · Max — Max rarely pays on Astra: on Terminal-Bench 4.0 and DeepSWE, Extra high scores as well or better (59.6% vs 59.1%; 74.1% vs 73.2%) for 30–40% less.

How long will your subscriptions last?

Enter the subscriptions you have and how you work: you get the agent hours per 5-hour window and per week, and when the quota runs out.

Quick start
Your subscriptions
How you work
Agent hours per 5-hour window
2.3 h (1.2–18.1) Claude Pro
Agent hours per week (all subscriptions)
17.6 h (7.4–89.9)
Your week
16.6 h / 30 h
First blocked: Mon 11:15

When does the quota run out?

Weekly quota left across your subscriptions over the week, for the pessimistic, central and optimistic calibrations.

Assumptions: budgets are API-dollar equivalents measured by third parties (vendors publish no absolute quota, except Z.ai credits); burn per busy hour comes from the Artificial Analysis Coding Agent Index; effort changes burn per hour only slightly (×0.67 at Low to ×1 at Max) but changes cost per task a lot. Pessimistic = low budget and high burn; optimistic = the opposite. Headless use (claude -p) burns the Claude 5-hour window 1.5–2.3× faster.

Performance against quota burn

Each line is one model across its reasoning-effort levels. Further left = the task burns less of your quota; higher = better result. The thin step line is the efficient frontier: no other model × effort does better for less.

Benchmark

IndependentAgentic terminal tasks run by Artificial Analysis with the same harness for every model, at every effort level.

Source updated continuously · Retrieved Sep 28, 2026

Highlight models (thicker line and label; use the chart legend to hide models)

Score vs cost per task

API-equivalent cost at list prices, log scale — the same ratio at which subscription quotas burn.

Scroll or pinch to zoom (Shift + scroll for the vertical axis), drag the slider, or use the zoom tool at the top right. Click a legend entry to hide or show it.

17 results without a published cost are shown only in the effort chart and the table.

Score by reasoning effort

Where the line flattens, the next effort level costs more quota for little or no gain.

Show the data table
ModelEffortScoreCost per task (USD, API list prices)Efficient frontier
Claude Fable 5.1Low40.4%$7.32
Claude Fable 5.1Medium44.95%$9.12
Claude Fable 5.1High52.02%$11.64
Claude Fable 5.1Extra high55.05%$15.78
Claude Fable 5.1Max52.02%$19.22
Claude Opus 5.5Low31.31%$2.08
Claude Opus 5.5Medium52.53%$4.04✓
Claude Opus 5.5High56.57%$5.12✓
Claude Opus 5.5Extra high59.6%$8.78
Claude Opus 5.5Max59.6%$13.11
Claude Sonnet 5Low2.53%$2.66
Claude Sonnet 5Medium2.02%$5.53
Claude Sonnet 5High5.05%$9.4
Claude Sonnet 5Extra high7.07%$12.8
Claude Sonnet 5Max14.14%$19.69
Claude Haiku 4.5High0%$1.27
GPT-6 AstraLow41.92%$2.25✓
GPT-6 AstraMedium49.49%$4.43
GPT-6 AstraHigh54.04%$4.05✓
GPT-6 AstraExtra high59.6%$5.86✓
GPT-6 AstraMax59.09%$8.5
GPT-6 SolNo reasoning13.13%$1.94
GPT-6 SolLow9.09%$0.49
GPT-6 SolMedium18.69%$1.12
GPT-6 SolHigh26.26%$1.6
GPT-6 SolExtra high30.3%$1.91
GPT-6 SolMax43.94%$4✓
GPT-6 LunaNo reasoning1.52%$0.045✓
GPT-6 LunaLow0%$0.005✓
GPT-6 LunaMedium2.53%$0.072✓
GPT-6 LunaHigh4.55%$0.12✓
GPT-6 LunaExtra high8.08%$0.2✓
GPT-6 LunaMax12.63%$0.24
Gemini 3.8 FlashLow10.1%—
Gemini 3.8 FlashMedium19.7%$4.61
Gemini 3.8 FlashHigh19.7%$5.84
Gemini 3.7 FlashHigh13.64%$4.31
Gemini 3.1 Pro (preview)High4.04%$4.15
Grok 4.7High24.75%$8.95
Grok 4.7Extra high25.76%$14.6
Grok 4.6Low3.03%$2.45
Grok 4.6Medium13.13%$6.7
Grok 4.6High21.21%$7.51
Grok 4.6Extra high17.17%$8.98
Kimi K3Low12.63%—
Kimi K3Max12.63%$5.09
GLM-5.3Low34.85%$6.4
GLM-5.3Max41.92%$8.06
GLM-5.3-FlashMax32.83%$0.79
Qwen3 Coder NextNo reasoning0%$2.83
Qwen3.5 122B A10BHigh0%$1.78
Qwen3.5 397B A17BHigh0%$2.6
Qwen3.6 35B A3BHigh0%$2.76
Qwen3.7 PlusHigh1.01%$1.27
Qwen3.8 2.4T A95BHigh11.11%$11.49
Qwen3.8 27BNo reasoning0%$9.39
Qwen3.8 27BLow2.53%$6.14
Qwen3.8 27BMedium5.05%$5.76
Qwen3.8 27BExtra high5.56%$3.94
Qwen3.8 MaxHigh38.89%$18.68
Qwen3.8-Flash-NextHigh25.25%$0.88
Claude Fable 5Max42.42%$34.91
Claude Opus 4.8Max21.72%$18.15
Claude Opus 5Low26.26%$5.18
Claude Opus 5Medium34.34%$9.18
Claude Opus 5High45.96%$13.14
Claude Opus 5Extra high46.46%$17.35
Claude Opus 5Max48.99%$19.29
Claude Sonnet 4.6Max3.03%$13.28
Apodex 1.1High0%$2.05
Trinity Large ThinkingHigh0.51%$0.62
Celeris-1No reasoning0%$0.17
Command A+High0.51%$0
North Mini CodeHigh0.51%$0
DeepSeek V4 Flash 0731Max12.12%$0.92
DeepSeek V4 Flash VisionMax12.12%$1.28
DeepSeek V4 Pro 0813Max14.14%$3.76
DeepSeek V4.1 FlashNo reasoning5.56%$1.06
DeepSeek V4.1 FlashMax26.77%$1.16
Gemini 3.5 FlashHigh6.57%$6.47
Gemini 3.5 Flash-LiteHigh1.01%$0.71
Gemini 3.6 FlashHigh7.07%$6.67
Gemma 4 31BHigh0%—
Granite 4.2 3BHigh0%$0.018
Granite 4.2 8BHigh0%$0.04
Mercury 2.5High0%$0.82
Ling 3.0 FlashHigh0%—
Ling 3.0 TinyHigh0%—
Ling-3.0-flash-FinHigh0%—
Ling-3.0-flash-VLHigh0%—
Ring-2.6-1TExtra high0.51%$0.92
K2 Horizon 0.9BHigh0%—
K2 Horizon 3.7BHigh0%—
K2 Horizon 375B A23BHigh1.52%—
K2 Horizon 7BHigh1.01%—
K2 Horizon MoVA 36B A4BHigh0%—
LongCat 2.0High0%$0.078
Muse GlimmerHigh0.51%$0.27
Muse Spark 1.1Extra high6.06%$8.57
Muse Spark 1.2Extra high7.07%$4.98
Muse Spark 1.3Extra high16.67%$6.23
Muse Spark 1.3Max33.33%$6.94
MiniMax-M3High2.02%$3.04
Ministral 3 14BNo reasoning0%$0.048
Ministral 3 3BNo reasoning0%$0.01
Ministral 3 8BNo reasoning0%$0.042
Mistral Large 3No reasoning0%$0.13
Mistral Medium 3.5High0%$2.14
Mistral Small 4High0%$0.048
Kimi K2.6High0.51%$4.15
Kimi K2.7 CodeHigh1.01%$3.05
Quasar 438BMax1.01%$12.14
NVIDIA Nemotron 3 Nano 30B A3BHigh0%$0.012
Nemotron 3 Super 120B A12BHigh0%$2.07
Nemotron 3 Ultra 550B A55BHigh0.51%$3.06
Nemotron 3.5 LightningHigh0.51%$0.37
GPT-5.5Medium5.05%$2.99
GPT-5.5High9.09%$5.59
GPT-5.5Extra high14.65%$11.56
GPT-5.5 Instant (June 2026)High12.63%$3.8
GPT-5.6 LunaNo reasoning1.01%$0.016✓
GPT-5.6 LunaLow0%$0.012
GPT-5.6 LunaMedium0.51%$0.025
GPT-5.6 LunaHigh2.53%$0.11
GPT-5.6 LunaExtra high3.54%$0.3
GPT-5.6 LunaMax11.62%$0.85
GPT-5.6 SolLow1.01%$0.52
GPT-5.6 SolMedium14.65%$1.64
GPT-5.6 SolHigh20.71%$2.22
GPT-5.6 SolExtra high24.75%$3.46
GPT-5.6 SolMax39.9%$8.09
GPT-5.6 TerraNo reasoning0.51%$0.23
GPT-5.6 TerraLow1.52%$0.24
GPT-5.6 TerraMedium1.01%$0.34
GPT-5.6 TerraHigh1.52%$0.63
GPT-5.6 TerraExtra high10.1%$1.57
GPT-5.6 TerraMax35.35%$6.39
gpt-oss-120bHigh0%$0.064
gpt-oss-20bHigh0%$0.009
MiniCPM5-2BHigh0%—
Step 5 PreviewHigh33.33%$3.48
Hy3High0.51%$0.41
InklingExtra high1.01%—
Inkling SmallHigh1.01%—
Solar Pro 4High0.51%—
Grok 4.5High10.61%$6.95
MiMo-V2.5High0%—
MiMo-V2.5-ProHigh0%$0.24
MiMo-V2.6-FlashHigh22.73%$0.23✓
MiMo-V2.6-ProHigh34.85%$0.44✓
GLM-5.2Max1.01%$7.81

Coding leaderboards

Independent rankings of coding agents and web development. Switch each leaderboard on or off; scroll or zoom inside a chart, and click a provider in its legend to hide it.

Leaderboards to show

Artificial Analysis Coding Agent Index v1.5

Each coding agent with its vendor's model (Claude Code, Codex, Grok Build, Kimi Code, Antigravity, OpenCode) on DeepSWE, Terminal-Bench 4.0 and SWE-Atlas-QnA — the closest proxy for what a subscription gives you.

Source updated Sep 6, 2026 · Retrieved Sep 28, 2026 · 20 models

Incomplete coverage — this data does not include Claude Sonnet 5. Key current models are missing, so the ranking here is not fully relevant.

Arena Code: WebDev

Blind pairwise human votes on web-development tasks (795,514 votes, cutoff 2026-09-25).

Source updated Sep 25, 2026 · Retrieved Sep 28, 2026 · 134 models

Design Arena: agentic web app, full-stack

Two agents build the same full-stack app (React, Supabase, deployed to Vercel) in a sandbox with 23 tools; people vote for the better one.

Source updated Sep 28, 2026 · Retrieved Sep 28, 2026 · 49 models

Incomplete coverage — this data does not include Claude Opus 5.5, Grok 4.7, GLM-5.3. Key current models are missing, so the ranking here is not fully relevant.

Official launch figures

What each vendor published for its own model at launch, with the effort it used. Vendors pick their own agent, harness and settings, so these numbers usually run higher than independent measurements — use them to see each vendor's claims, and the charts above to compare.

ModelSWE-bench VerifiedSWE-bench ProTerminal-Bench 2.xTerminal-Bench 4.0DeepSWE v1.1FrontierCode 1.1
Claude Fable 5.1Anthropic—81.2%↗Max—55.8%↗Max—50.3%↗Max
Claude Opus 5.5Anthropic—89.9%↗Max—66.4%↗Extra high74.2%↗Max54.4%↗Max
Claude Sonnet 5Anthropic85.2%↗Max63.2%↗Max80.4%↗Extra high · 2.1———
Claude Haiku 4.5Anthropic73.3%↗—————
GPT-6 AstraOpenAI———57.88%↗High74.12%↗Extra high53.3%↗Max
GPT-6 SolOpenAI————68.81%↗Max49.27%↗Max
GPT-6 LunaOpenAI————66.59%↗Max42.42%↗Max
Gemini 3.8 FlashGoogle——89.4%↗Medium · 2.119.1%↗73.7%↗High—
Gemini 3.7 FlashGoogle——85.8%↗Medium · 2.111.2%↗65.3%↗—
Gemini 3.1 Pro (preview)Google80.6%↗High54.2%↗High68.5%↗High · 2.0———
Grok 4.7xAI———38%↗Extra high71%↗High—
Grok 4.6xAI———20.3%↗High65.9%↗High—
Kimi K3Moonshot AI——88.3%↗Max · 2.1—67.5%↗Max—
GLM-5.3Z.ai——88.2%↗Max · 2.1—66.9%↗Max—
GLM-5.3-FlashZ.ai——84.3%↗2.1—63.4%↗—

Remote or local?

Beyond the six providers above: the other subscriptions that are worth it, and when running a model on your own machine makes sense.

Other subscriptions worth knowing

SubscriptionPrice / monthQuotaVerdict
GitHub Copilot↗↗$10 / $39 / $100Dollar credits at API prices: $15 on Pro, $70 on Pro+, $200 on Max; completions free.worth itSecond most used coding agent (21%); GPT-6, Opus 5.5, Fable 5.1, Gemini 3.8 Flash, Grok 4.7 and Kimi K3 in VS Code, JetBrains and the CLI.
Alibaba Qwen Coding Plan↗↗↗$506,000 requests per 5 hours, 45,000 per week, 90,000 per month.worth itPublished quotas; Qwen3.8 Max is #6 on Arena WebDev and 43.3 on the AA Coding Agent Index; works in Claude Code, Codex and OpenCode.
OpenCode Go↗$10Per 5 hours: Kimi K3 110 requests, Qwen3.8 Max 160, GLM-5.3 220, DeepSeek V4.1 Flash 26,000.worth itThe cheapest way to use the best open-weight models in a good agent (OpenCode is the harness AA uses for GLM).
MiniMax Token Plan↗↗↗$22 / $55 / $1325-hour and weekly windows, in tokens.watchClear plan and 47.2% on SWE-rebench, but only 2.0% on AA's Terminal-Bench 4.0 run: watch the next model.
Meta Muse Code↗↗$5 / $15 / $5010–50 prompts per 5 hours on Everyday; 5× and 20× above.niche54.3 on the AA Coding Agent Index (above Kimi K3) for a small price.
Xiaomi MiMo Token Plan↗↗$6 – $100Credits, 20% off during off-peak hours.nicheVery cheap; MiMo-V2.6-Pro reaches 34.8% on AA Terminal-Bench 4.0.
Cursor↗$20 / $60 / $200Two pools (Cursor models, other models), sizes not published.nicheAn IDE surface for many models, including Grok 4.7 and Composer 2.5.
Devin↗↗$20 / $200Not published.nicheDevin Fusion (Fable 5.1 + SWE-2) scores 61.7 on the AA Coding Agent Index.
Factory Droid↗$20 / $100 / $200Per-model multipliers (Opus 5.5 1.6×, GPT-6 Astra 4×).nicheMulti-model agent with explicit multipliers.
Mistral Vibe↗↗$14.99 / $24.99Not published.nicheEuropean option; Mistral Medium 3.5 scores 0% on AA's Terminal-Bench 4.0 run.

Running models locally

Local models are a separate tier: they cost no quota, but the best of them reach 5–33% on Terminal-Bench 4.0 versus about 60% for frontier models. Pair a local model for routine work with a remote frontier model for hard tasks.

HardwareRecommended modelMemoryDecode speedQuality (AA index · Terminal-Bench 4.0)
24 GB GPU (RTX 4090 / 3090)Qwen3.8-27B Q4_K_M16.5 GB38–46 tok/sAA 33.7 · TB4 5.6%↗↗↗
32 GB GPU (RTX 5090)Qwen3.8-27B Q6 / NVFP422–25 GB23–75 tok/s · MTP 98–134AA 33.7 · TB4 5.6%↗↗
128 GB unified memory (Strix Halo, DGX Spark, M5 Max)Qwen3.8-Flash-Next IQ4_XS93.7 GB11–37 tok/s · MTP 47–84AA 39.8 · TB4 25.3%↗↗
256 GB unified memory (M5 Ultra, M3 Ultra)GLM-5.3-Flash MLX 4-bit204 GB26–41 tok/sAA 41.8 · TB4 32.8%↗↗

Worth it for

  • Private or regulated code that must not leave the machine
  • Working offline
  • Hardware you already own
  • High-volume routine work: boilerplate, tests, subagents
  • A 256 GB machine shared by a team

Not worth it for

  • Saving money versus a $20 subscription (hardware and power cost $10–270 a month)
  • Hard agentic tasks: the quality gap to frontier models is large

References

Every figure on this page comes from these pages. For each one: what we took from it, and when.

  1. Model Studio Coding PlanAlibaba CloudOfficial

    Qwen Coding Plan Pro price and request quotas per 5 hours, week and month.

    Published Sep 11, 2026 · Accessed Sep 28, 2026

  2. CursorBench 3.2.0 and Terminal-Bench 4.0 score and cost at each effort level for Fable 5.1, Fable 5 and Mythos 5.1.

    Published Sep 1, 2026 · Accessed Sep 28, 2026

  3. Fable models use limits faster and are capped at 50% of the weekly limit on Max and Team Premium; credits only on Pro and Team Standard.

    Accessed Sep 28, 2026

  4. Claude Opus 5.5AnthropicOfficial

    Score and cost per task at each effort level for Opus 5.5, Fable 5.1 and Opus 5 on Terminal-Bench 4.0, FrontierCode v1.1 and CursorBench 4.0; five-hour limit change of 2026-09-22.

    Published Sep 22, 2026 · Accessed Sep 28, 2026

  5. Headline coding scores (SWE-bench Pro, Terminal-Bench 4.0 at xhigh and max, DeepSWE v1.1, CursorBench 4.0) versus Opus 5, Fable 5.1 and GPT-6 Astra.

    Published Sep 22, 2026 · Accessed Sep 28, 2026

  6. Claude pricingAnthropicOfficial

    Prices of Free, Pro, Max 5x, Max 20x, Team and Enterprise plans and what each includes.

    Accessed Sep 28, 2026

  7. Headline coding scores of Sonnet 5 at max effort (SWE-bench Verified, SWE-bench Pro, Terminal-Bench 2.1).

    Published Jun 30, 2026 · Accessed Sep 28, 2026

  8. EffortAnthropicOfficial

    Effort levels available per model and when to use each (xhigh for agentic coding over 30 minutes, low for subagents).

    Accessed Sep 28, 2026

  9. Five-hour session and weekly limits; no fixed message count; what drives consumption (model, effort, length, tools).

    Accessed Sep 28, 2026

  10. SWE-bench Verified score of Haiku 4.5.

    Published Oct 15, 2025 · Accessed Sep 28, 2026

  11. Release date, positioning and headline coding scores of Sonnet 5.

    Published Jun 30, 2026 · Accessed Sep 28, 2026

  12. Model configurationAnthropicOfficial

    Default model (Opus 5.5 on every plan since v2.1.280) and default effort per model in Claude Code.

    Accessed Sep 28, 2026

  13. Models overviewAnthropicOfficial

    Model IDs, release dates, context windows, max output and status of every current Claude model.

    Accessed Sep 28, 2026

  14. Opus uses meaningfully more quota than Sonnet; higher effort reaches limits faster.

    Accessed Sep 28, 2026

  15. Cost per solved task on a 478-problem SWE-bench Pro subset for Opus 5.5, Fable 5.1 and Sonnet 5 configurations; effort trade-offs for Opus 5.5.

    Accessed Sep 28, 2026

  16. PricingAnthropicOfficial

    Input, cache write, cache read and output prices per million tokens; batch and fast-mode prices.

    Accessed Sep 28, 2026

  17. What is the Max plan?AnthropicOfficial

    Max 5x and Max 20x usage relative to Pro per five-hour session; monthly billing only.

    Accessed Sep 28, 2026

  18. 5-hour session limits increaseAnthropic (ClaudeDevs on X)Official

    Five-hour limits up 20% on 2026-09-22; Opus 5.5 goes 25% further within limits because it is priced lower.

    Published Sep 22, 2026 · Accessed Sep 28, 2026

  19. Arena Code: WebDev leaderboardArena (LMArena)Third party

    Rating, 95% confidence interval, votes and rank of every model on blind pairwise web-development votes (entries[] in the server-rendered payload; vote cutoff 2026-09-25).

    Published Sep 25, 2026 · Accessed Sep 28, 2026

  20. Artificial Analysis Coding Agent IndexArtificial AnalysisThird party

    Coding Agent Index v1.5 per harness × model (Claude Code, Codex, Grok Build, Kimi Code, Antigravity, OpenCode): score, cost, tokens and time per task.

    Accessed Sep 28, 2026

  21. Artificial Analysis evaluation detailsArtificial AnalysisThird party

    Per-model sub-scores embedded in the page payload: Humanity's Last Exam, CritPt, SciCode, GPQA and Terminal-Bench 4.0 for each model × effort.

    Accessed Sep 28, 2026

  22. Artificial Analysis model leaderboardArtificial AnalysisThird party

    Intelligence Index v4.3 and AA's own Terminal-Bench 4.0 run for every model × effort, with cost per task and output tokens per task (JSON embedded in the page).

    Accessed Sep 28, 2026

  23. Grok 4.7 usage on SuperGrok HeavyBigGo Finance (reporting a user post)Third party

    About $40 of Grok 4.7 usage in Grok Build burned about 8% of the Heavy weekly limit.

    Published Sep 22, 2026 · Accessed Sep 28, 2026

  24. Weekly budget of Max 5x and Max 20x in API-equivalent dollars (two accounts, /usage vs ccusage).

    Published Aug 20, 2026 · Accessed Sep 28, 2026

  25. Codex usage researchcodexusage.devThird party

    Codex Plus 5-hour and weekly budgets in credits and dollars, cost per busy hour by effort level.

    Published Sep 22, 2026 · Accessed Sep 28, 2026

  26. Coding plan trackercodingplan.orgThird party

    Current Kimi plan names and CNY prices (Go, Plus, Pro, Max) and the mapping from legacy names.

    Published Sep 24, 2026 · Accessed Sep 28, 2026

  27. Devin pricingCognitionOfficial

    Devin Pro and Max prices.

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  28. FrontierCode leaderboard dataCognitionThird party

    FrontierCode v1.1 score (new_score), cost, tokens and duration for every model × effort, main and extended subsets.

    Published Sep 22, 2026 · Accessed Sep 28, 2026

  29. Cursor pricingCursorOfficial

    Cursor Pro, Pro+ and Ultra prices and their two usage pools.

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  30. CursorBench 4.0CursorThird party

    Score, cost per task, tokens and steps for every model × effort (server-rendered results table).

    Published Sep 10, 2026 · Accessed Sep 28, 2026

  31. DeepSWE v1.1 leaderboardDatacurveThird party

    pass@1 and mean cost per task for every model × effort (mini-swe-agent harness, 4 runs).

    Accessed Sep 28, 2026

  32. Lite plan in Claude Code: a 3-hour session used 70% of a 2,000-credit window; one endpoint with tests took 18 minutes and 210 credits.

    Published Aug 29, 2026 · Accessed Sep 28, 2026

  33. Design Arena: agentic web devDesign ArenaThird party

    Elo, standard error, win rate and battles for the full-stack and frontend agentic web-app arenas (embedded boards and /api/leaderboard).

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  34. Factory pricingFactoryOfficial

    Droid Pro, Plus and Max prices and per-model multipliers.

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  35. Kimi plans in USDgeotoolbox.aiThird party

    International USD prices of Kimi plans and their Kimi Code usage multipliers (1×, 5×, 15×, 30×).

    Published Sep 5, 2026 · Accessed Sep 28, 2026

  36. GitHub Copilot plansGitHubOfficial

    Copilot Pro, Pro+, Max, Business and Enterprise prices and the AI Credits included with each.

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  37. Pro 20x weekly meter measurementGitHub user reportThird party

    Weekly meter moved from 64% to 95% for $609 of API-equivalent usage on Pro 20x.

    Published Sep 10, 2026 · Accessed Sep 28, 2026

  38. Three weekly meters on Pro 5xGitHub user reportThird party

    Three full weekly Pro 5x meters = $959.87 API-equivalent (about $320 per week) and 708 M tokens on GPT-5.6 Sol and Astra.

    Published Sep 7, 2026 · Accessed Sep 28, 2026

  39. API-equivalent value of a full weekly Pro 20x allowance: about $1,925 on GPT-5.6 Sol, $1,191–1,234 on GPT-6 Astra.

    Published Sep 8, 2026 · Accessed Sep 28, 2026

  40. Claude Code usage meter measurementsGitHub user reportsThird party

    Controlled A/B measurements of API-equivalent dollars per 1% of the 5-hour and weekly meters (interactive vs headless), Pro and Max.

    Published Sep 27, 2026 · Accessed Sep 28, 2026

  41. Antigravity modelsGoogleOfficial

    Models available per plan and thinking options (Low/Medium/High) in Antigravity; two usage pools with five-hour and weekly remaining percentages.

    Accessed Sep 28, 2026

  42. Antigravity plansGoogleOfficial

    Quota wording per plan: weekly refresh on Free, five-hour refresh until a weekly limit on Pro and Ultra.

    Accessed Sep 28, 2026

  43. Ultra $100 = 5× and $200 = 20× the Pro token allowance; the Gemini pool is drawn down at API prices; Claude and GPT models use a separate fixed pool.

    Published May 19, 2026 · Accessed Sep 28, 2026

  44. Gemini API pricingGoogleOfficial

    Prices per million tokens for Gemini 3.8/3.7 Flash (introductory until 2026-12-31) and 3.1 Pro Preview.

    Published Sep 24, 2026 · Accessed Sep 28, 2026

  45. Gemini thinkingGoogleOfficial

    Thinking levels and defaults per Gemini model; guidance per level.

    Published Sep 25, 2026 · Accessed Sep 28, 2026

  46. Prices of Google AI Plus, Pro, Ultra 5x and Ultra 20x; Gemini app usage multipliers.

    Accessed Sep 28, 2026

  47. Gemini 3.1 Pro model cardGoogle DeepMindOfficial

    SWE-bench Verified, SWE-bench Pro and Terminal-Bench 2.0 scores of Gemini 3.1 Pro (high thinking).

    Published Feb 19, 2026 · Accessed Sep 28, 2026

  48. Gemini 3.8 Flash evaluationGoogle DeepMindOfficial

    Coding benchmark table of Gemini 3.8 Flash versus Opus 5, Sonnet 5 and GPT-5.6, with the settings used.

    Published Sep 2, 2026 · Accessed Sep 28, 2026

  49. SuperGrok Grok Build usage reportHacker NewsThird party

    Four terminals at full speed for one hour used 9% of the SuperGrok weekly pool during a 2× promotion.

    Published Aug 12, 2026 · Accessed Sep 28, 2026

  50. Qwen3.8-27B hardware testsHardware CornerThird party

    Local decode speeds of Qwen3.8-27B on RTX 4090, 3090, 5090 and M5 Max at several context lengths.

    Published Aug 17, 2026 · Accessed Sep 28, 2026

  51. AI coding agent adoption 2026JetBrains ResearchThird party

    Share of developers using each coding agent (Claude Code 39%, Copilot 21%, Codex 16%, Cursor 12%, OpenCode 7%, Antigravity 6%; 15,000+ respondents).

    Published Aug 15, 2026 · Accessed Sep 28, 2026

  52. KiloBenchKiloThird party

    Terminal-Bench 2.0 completion and cost per full attempt in the Kilo agent (embedded chart points).

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  53. LiveBenchLiveBenchThird party

    Coding and agentic-coding category scores and cost per question (table_2026_06_25.csv and category files).

    Published Sep 10, 2026 · Accessed Sep 28, 2026

  54. Best AI for codingllm-stats (ZeroEval)Third party

    TrueSkill coding rating (μ − 3σ) over 60 benchmark games for every model (initialIndexData in the payload). Needs a cache-busting query string: the plain URL served a stale 2026-09-22 copy.

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  55. Mac Studio M5 Ultra reviewMacStoriesThird party

    Local decode speed of GLM-5.3-Flash on the M5 Ultra.

    Published Sep 21, 2026 · Accessed Sep 28, 2026

  56. Muse CodeMetaOfficial

    Muse Code plans (Everyday, High, Power) and prompts per 5 hours.

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  57. MiniMax Token PlanMiniMaxOfficial

    Token Plan prices and 5-hour / weekly windows.

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  58. Mistral pricingMistral AIOfficial

    Vibe Pro and Team prices (no published quota).

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  59. Kimi API pricingMoonshot AIOfficial

    Cache hit, cache miss and output prices per million tokens for Kimi K3 and K2.7 Code.

    Accessed Sep 28, 2026

  60. Kimi Code membershipMoonshot AIOfficial

    Five-hour rolling window, monthly cap shared across Kimi services, extra usage top-ups.

    Accessed Sep 28, 2026

  61. Kimi Code modelsMoonshot AIOfficial

    Models in Kimi Code (K3, K2.8 Preview, HighSpeed at 3× quota), effort levels and plan access.

    Accessed Sep 28, 2026

  62. Kimi K3Moonshot AIOfficial

    Coding benchmark table of Kimi K3 (max effort) versus Fable 5, GPT-5.6 Sol, Opus 4.8 and GLM-5.2, with harnesses used.

    Published Jul 16, 2026 · Accessed Sep 28, 2026

  63. Kimi membership pricingMoonshot AIOfficial

    CNY monthly prices of the membership tiers (legacy names).

    Accessed Sep 28, 2026

  64. SWE-rebenchNebiusThird party

    Resolved rate, pass@5 and cost per task on the 2026-05-15 → 2026-07-01 window (embedded items).

    Published Jul 24, 2026 · Accessed Sep 28, 2026

  65. Cross-provider quota calibrationoh-my-pi (GitHub)Third party

    Tokens and API-equivalent dollars per 1% of the Anthropic, OpenAI and Antigravity meters, measured on one machine.

    Published Sep 26, 2026 · Accessed Sep 28, 2026

  66. Pro $100 (5× Plus) and Pro $200 (20× Plus); new Pro $200 sign-ups paused since 2026-09-10.

    Accessed Sep 28, 2026

  67. API pricingOpenAIOfficial

    Standard, batch, flex and fast-mode prices per million tokens for GPT-6 Astra, Sol and Luna.

    Accessed Sep 28, 2026

  68. Credits per million tokens for each model (25× the API dollar price); fast mode 2.5×; Ultra billing.

    Accessed Sep 28, 2026

  69. Codex modelsOpenAIOfficial

    Effort controls per surface (Light to Ultra), recommended starting presets (Sol Medium, Luna High, Astra Light), surfaces per model.

    Accessed Sep 28, 2026

  70. Codex pricingOpenAIOfficial

    Plan prices and the estimated local messages per 5 hours for each model on Plus, Pro 5x, Pro 20x and Business.

    Accessed Sep 28, 2026

  71. GPT-6 AstraOpenAIOfficial

    Terminal-Bench 4.0 score, cost, output tokens and latency at each effort for GPT-6 Astra, GPT-5.6 Sol and Fable 5.1; FrontierCode 1.1 Extended; output tokens per DeepSWE task.

    Published Sep 3, 2026 · Accessed Sep 28, 2026

  72. Release dates, context windows, knowledge cutoffs and supported effort values.

    Accessed Sep 28, 2026

  73. DeepSWE v1.1 and FrontierCode 1.1 Main score and cost per task at each effort for GPT-6 Astra, Sol, Luna, GPT-5.6, Opus 5 and Fable 5.

    Published Sep 22, 2026 · Accessed Sep 28, 2026

  74. Business Standard and Premium seat prices, annual and monthly.

    Accessed Sep 28, 2026

  75. Five-hour and weekly windows, shared allowance across Codex and Work, effort advice (Astra Low can beat Sol High).

    Accessed Sep 28, 2026

  76. OpenCode GoOpenCodeOfficial

    OpenCode Go price and requests per 5 hours for each open-weight model.

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  77. FrontierSWE V2Proximal LabsThird party

    Mean, best and worst of 5 trials, cost per trial and duration on 34 long-horizon tasks.

    Published Sep 22, 2026 · Accessed Sep 28, 2026

  78. Quality of Qwen3.8-27B by quantization level on Terminal-Bench 2.1 (Q4_K_M matches BF16, Q2 loses about 10 points).

    Published Aug 26, 2026 · Accessed Sep 28, 2026

  79. Terminal-Bench 4.0 leaderboardTerminal-Bench (tbench.ai)Third party

    Accuracy with confidence interval, total cost and tokens per agent × model × effort over 330 trials.

    Published Sep 21, 2026 · Accessed Sep 28, 2026

  80. Vals AI benchmarksVals AIThird party

    Accuracy, cost per test and latency for Terminal-Bench 4, Vibe Code Bench, Vals Index and other coding benchmarks (astro-island props).

    Published Sep 23, 2026 · Accessed Sep 28, 2026

  81. Vals IndexVals AIThird party

    Vals Index score, cost per test and latency for Opus 5.5, Fable 5.1, GPT-6 Astra and others.

    Accessed Sep 28, 2026

  82. Grok 4.6xAIOfficial

    Coding scores of Grok 4.6 at high effort (DeepSWE v1.1, Terminal-Bench 3.0, CursorBench 3.2).

    Published Aug 12, 2026 · Accessed Sep 28, 2026

  83. Grok 4.7xAIOfficial

    CursorBench 4.0 score, cost per task, output tokens and steps at each effort for Grok 4.7 and for Fable 5.1, Opus 5, Sonnet 5, GPT-5.6 Sol/Terra/Luna and Gemini 3.8 Flash (JS dataset behind the chart).

    Published Sep 21, 2026 · Accessed Sep 28, 2026

  84. Terminal-Bench 4.0, FrontierSWE V2 and SWE-Marathon scores of Grok 4.7 and 4.6; CursorBench 4.0 chart (vector geometry) for Grok 4.6.

    Published Sep 21, 2026 · Accessed Sep 28, 2026

  85. Grok FAQxAIOfficial

    One shared weekly usage pool across Chat, Imagine, Voice and Build since June 2026; shown only as a percentage; extra usage credits.

    Published Aug 27, 2026 · Accessed Sep 28, 2026

  86. Grok plansxAIOfficial

    Monthly and annual prices of SuperGrok Lite, SuperGrok, Plus and Heavy, from the product JSON the page loads (grok.com/rest/products).

    Accessed Sep 28, 2026

  87. ReasoningxAIOfficial

    Effort levels per Grok model (low to xhigh, default high) and guidance per level.

    Accessed Sep 28, 2026

  88. xAI API pricingxAIOfficial

    Input, cached input and output prices per million tokens for Grok 4.7, 4.6 and grok-build-0.1.

    Accessed Sep 28, 2026

  89. MiMo Token PlanXiaomiOfficial

    MiMo Token Plan prices and credits.

    Published Sep 22, 2026 · Accessed Sep 28, 2026

  90. GLM Coding PlanZ.aiOfficial

    Prices of Lite, Pro and Max (monthly, quarterly, yearly) and their five-hour and weekly credit allowances.

    Accessed Sep 28, 2026

  91. Credit formula and per-model token multipliers, peak and off-peak rates, and the official weekly token estimates per tier.

    Accessed Sep 28, 2026

  92. GLM-5.3Z.aiOfficial

    Effort levels (low, high, max; default max) and the recommendation to use max for coding.

    Accessed Sep 28, 2026

  93. GLM-5.3Z.aiOfficial

    Z.ai Code Bench accuracy and output tokens per task at each effort for GLM-5.3, Fable 5 and Opus 4.8.

    Published Aug 14, 2026 · Accessed Sep 28, 2026

  94. GLM-5.3 model cardZ.aiOfficial

    Coding benchmark table of GLM-5.3 at max effort in Claude Code (Terminal-Bench 2.1 and 3.0, DeepSWE v1.1, NL2Repo, FrontierSWE).

    Accessed Sep 28, 2026

  95. GLM-5.3-FlashZ.aiOfficial

    Coding scores of GLM-5.3-Flash (Terminal-Bench 2.1, DeepSWE v1.1) from printed chart labels.

    Published Aug 26, 2026 · Accessed Sep 28, 2026

  96. Z.ai API pricingZ.aiOfficial

    Input, cached input and output prices per million tokens for GLM-5.3 and GLM-5.3-Flash.

    Accessed Sep 28, 2026