Claude Opus 5 Cuts the Price in Half
Meanwhile Kimi K3 Gets Caught Calling Itself Claude

If you are deciding whether to downgrade from Claude Fable 5 to Claude Opus 5, or trying to figure out whether the Kimi K3 distillation controversy has real evidence behind it, this week delivered two stories about the same underlying question: what frontier intelligence actually costs, and where cheap capability really comes from. Anthropic launched Opus 5 on July 24, 2026 and made it the Claude Max default; Moonshot's K3 faces a White House accusation of industrial-scale distillation, while independent researchers found K3 identifying itself as Claude and leaking internal deployment IDs. This article covers Opus 5 pricing and benchmark comparisons, the K3 White House timeline, Ryan Greenblatt's technical findings, a decision matrix, a six-step rollout guide, and FAQ. Last updated: 2026-07-25

01

Two AI Headlines This Week: Six Pain Points for Model Selection

Opus 5 and the K3 distillation story look unrelated on the surface, but both force the same tradeoff: how much frontier intelligence is worth, and whether a low price signals a problem. These six friction points dominate technical discussion right now:

  1. 01

    Flagship vs daily driver: Fable 5 is the strongest tier but costs roughly twice as much; Opus 5 claims CursorBench within 0.5% at half the price — teams must rethink default routing.

  2. 02

    Data retention compliance: Fable 5 and Mythos 5 require accepting 30-day data retention; Opus 5 continues the Opus line's default of no mandatory retention — a real enterprise swing factor.

  3. 03

    K3 capability cannot be independently verified yet: full weights are promised for July 27; at controversy peak, outside researchers could not reproduce the 2.8T architecture or scores (see K3 open weights countdown).

  4. 04

    Distillation accusations lack public proof: White House OSTP director Michael Kratsios posted accusations on July 22–23 without releasing supporting materials; Moonshot has not answered detailed training-process questions.

  5. 05

    Timeline logic is shaky: Fable 5 went publicly available July 1; K3 launched July 16 — only two weeks. Independent experts argue deep distillation is not feasible on that schedule given compute and API costs.

  6. 06

    Supply chain and policy risk: accusations also involve unlicensed Nvidia GB300 chips; if export controls tighten, compliance costs for open-weight models from China could rise.

"Read both stories together: Opus 5 answers the value question with half-price near-flagship performance; the K3 fight is the industry's first public reckoning over whether your cheap model is real engineering or borrowed intelligence."

02

Claude Opus 5 vs Fable 5: Same Price Tier, Flagship-Level Jump

On July 24, 2026 (US Pacific time), Anthropic released Claude Opus 5 (model ID: claude-opus-5), set it as the Claude Max default, and made it the strongest tier available to Claude Pro users. Pricing is unchanged from Opus 4.8: $5 / million input tokens, $25 / million output tokens; 1M-token context window (default and only tier); 128K max output; Thinking enabled by default.

DimensionClaude Opus 5Claude Fable 5
Input / output pricing$5 / $25 per million tokens~$10 / $50 per million tokens (roughly 2× Opus 5)
CursorBench 3.2 (max effort)0.5% below Fable 5 peakPeak benchmark reference
Frontier-Bench v0.1Leads all models; 2×+ Opus 4.8
ARC-AGI 3 the next-best model
OSWorld 2.0Beats Fable 5 best score at under one-third the costComputer-use benchmark leader
Data retentionNo mandatory retention by defaultRequires 30-day data retention policy
Product positioningDaily driver / Claude Max defaultFlagship / extreme agent workloads

Key benchmarks and customer feedback

  • Cursor team: "Opus 5 delivers near-Fable 5 intelligence at Opus speed and cost — CursorBench only slightly below Fable 5."
  • Zapier AutomationBench: Opus 5 tops the leaderboard with 100% pass rate on end-to-end account-health workflows — no prior model cleared it.
  • Box production tests: +11% on data-analysis workflows, +17% on due-diligence scenarios, +8% overall accuracy.
  • Life sciences: internal evals show +10.2 percentage points on inferring molecular structure from spectra vs Opus 4.8; +7.7 points on protein variant function prediction.

Safety: most aligned generation, deliberately not chasing cyber frontier

Anthropic's automated behavior audits call Opus 5 the most aligned model to date — lowest deception rate, hardest to steer into misuse. On high-risk dual-use domains like cyber offense and biosecurity, Anthropic deliberately did not push Opus 5 to the frontier; restricted Mythos 5 still holds that slot. Opus 5's cyber classifier is roughly 85% more permissive than Fable 5's — usable for source-level vulnerability discovery, but still blocks binary exploit scanning, penetration testing, and exploit generation.

info

Sources: Anthropic official launch (anthropic.com/news/claude-opus-5), CNBC, The Verge, The New Stack.

03

Kimi K3 distillation controversy: White House accusation, timeline pushback, new evidence

If Opus 5 is a straightforward product launch, Kimi K3's arc is messier. Moonshot released Kimi K3 on July 16 — 2.8 trillion total parameters, the first open model to enter the "3T class"; sparse MoE (896 experts, 16 active per token, ~50B active parameters); 1M-token context plus native vision. GPQA-Diamond 93.5% and BrowseComp 91.2% were open-model highs at launch (see Kimi K3 deep review).

Controversy timeline

DateEvent
2026-02Anthropic first publicly accused Moonshot / DeepSeek / MiniMax of "industrial-scale distillation attacks," citing 3.4M+ anomalous API interactions
2026-07-01Claude Fable 5 publicly available
2026-07-16Kimi K3 API and product launch
2026-07-22/23White House OSTP director Michael Kratsios accuses "large-scale, covert industrial distillation" plus alleged unlicensed Nvidia GB300 use
2026-07-23TechCrunch reports independent researchers questioning whether two-week distillation is realistic
2026-07-24Ryan Greenblatt (Redwood Research) publishes "K3 identifies as Claude" analysis (GitHub: rgreenblatt/which_claude_is_k3)
2026-07-27 (planned)Kimi K3 full open weights — external verification possible

White House accusation vs expert pushback

Kratsios: "Large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology… is unacceptable." Treasury Secretary Scott Bessent followed with claims of "watermarks from U.S. large models" on many Chinese models — without specifying what "watermark" means.

"Fable only became publicly available on July 1. You cannot distill that much data, finish training, and ship a model in two weeks." — Braden Hancock, Snorkel AI co-founder

Allen Institute for AI researcher Nathan Lambert adds: as Chinese models close the frontier gap, simple supervised fine-tuning distillation yields diminishing returns; the real separator is slower reinforcement learning. "If distillation worked this well, every chaser would have caught up by now."

warning

Industry context: Elon Musk admitted in court that xAI distilled OpenAI models while training Grok and called it common industry practice — distillation itself is not illegal; the fight is over the line between normal technical borrowing and covert industrial extraction.

04

Why does Kimi K3 say it's Claude? The strongest indirect evidence so far

This is the angle worth reading closely — and the one English search traffic actually uses: Why does Kimi K3 say it's Claude.

Around July 24, Ryan Greenblatt published a statistical comparison of how models answer "who are you?" The results:

  • K3 abnormally often identifies as Claude, sometimes outputting internal deployment ID strings like claude-opus-4-5-20250929 and claude-sonnet-4-5-20250929
  • Real Claude models do not report this way — Sonnet 4.5 says "I'm Claude Sonnet 4.5"; Opus 4.5 may error or skip version IDs entirely
  • K3's self-reported versions lock to the Claude 4.5 era (late 2025), not current Fable/Mythos; prior K2 signals pointed at earlier Claude Sonnet 4 — a "chasing the teacher generation" pattern
example
# Internal deployment IDs K3 may output (real Claude models rarely volunteer these)
claude-opus-4-5-20250929
claude-sonnet-4-5-20250929

# Greenblatt's read: the model is surfacing teacher-model deployment metadata
# more accurately than the teacher itself — more consistent with training data
# contaminated by Claude API logs or labeled synthetic samples than with
# a clean independent identity.

Greenblatt stresses this does not directly prove distillation happened — identity confusion could come from data contamination, system-prompt leakage, or public dataset synthesis. Combined with Anthropic's February accusation and the statistical pattern, though, the distillation narrative finally has technical support beyond political statements.

Opus 5 vs Kimi K3: decision matrix for this week

DimensionClaude Opus 5Kimi K3
Pricing tierClosed API, roughly half of Fable 5Open weights + low API; Moonshot claims far below Fable 5 / GPT-5.6 Sol
Compliance / retentionNo mandatory retention; Anthropic official channelOpen weights, but distillation accusations add supply-chain risk
Independent verificationAPI live day one; benchmarks reproducibleFull weights unavailable until July 27
Moderation / refusalsStrict alignment; cyber classifier still blocks high-risk opsCommunity sees lighter content restrictions as a selling point
Local deploymentClosed source — API only2.8T parameters — essentially no one runs full weights locally
Best fitCompliance-sensitive enterprise, daily agent / coding driverExtreme cost sensitivity, acceptable provenance risk, wait for weight verification
05

Six-Step Model Selection Playbook: Scenario First, Stance Second

  1. 01

    Draw compliance boundaries first: if retention, export controls, or supply-chain audits matter, weigh Opus 5's default no-retention policy and Anthropic SLA before benchmark scores.

  2. 02

    Calculate real token cost: Opus 5 matches Opus 4.8 pricing with a capability jump — if Fable 5 price blocked adoption, run CursorBench / AutomationBench A/B on your tasks before changing default routing.

  3. 03

    Wait on K3 until July 27: before full weights drop, architecture and scores cannot be independently verified; use the open weights release checklist to plan verification.

  4. 04

    Watch identity-confusion probes: run simple "who are you?" tests on open models; abnormal deployment ID output may signal training-data contamination — especially for production agent routing.

  5. 05

    Match SEO to native search phrasing: English content should target queries like Why does Kimi K3 say it's Claude and Claude Opus 5 vs Fable 5, not literal translations of Chinese controversy labels.

  6. 06

    Separate execution layer from model layer: whether you route Opus 5 or K3, iOS CI/CD, long CLI agent sessions, and Keychain isolation still need stable macOS compute — cheaper models do not replace infrastructure.

Hard numbers worth citing (EEAT)

  • Opus 5 value: CursorBench 3.2 max effort within 0.5% of Fable 5 peak at roughly 50% of Fable 5 cost
  • K3 scale: 2.8T total parameters, 896-expert MoE, 16 experts active per token (~50B active parameters)
  • Distillation window: Fable 5 public (July 1) to K3 launch (July 16) = 14 days — multiple independent researchers argue insufficient for RL-grade deep distillation
  • Anthropic prior claim: February 2026 cited 3.4M+ anomalous API interactions tied to Moonshot, attributed to deliberate capability extraction

Opus 5 upgrades the "daily driver" API tier; K3 compresses the open-vs-closed gap to a matter of days — but a cheap Linux VPS cannot run xcodebuild, notarytool, or the macOS toolchain Claude Code expects, and a local Mac often hits memory and disk ceilings. Teams juggling multi-model API routing with iOS CI/CD and long CLI agent sessions usually get more reliability from dedicated, SSH-stable remote macOS nodes than from betting on one local machine — cheaper model APIs do not fix Keychain isolation or session stability. For production workloads that need stable SSH sessions and predictable bandwidth, NodeMini Mac Mini cloud rental is usually the better fit. See rental rates and the help center for setup.

Last updated: 2026-07-25. K3 full weights expected July 27 — revisit after independent verification lands.

FAQ

Frequently Asked Questions

Opus 5 stays at $5/$25 per million input/output tokens — roughly half of Fable 5 ($10/$50 tier). On CursorBench 3.2, it trails Fable 5's peak by less than 1%. For stable macOS execution alongside agent workflows, see Mac Mini rental rates.

Yes. From the July 24, 2026 launch, Opus 5 became the Claude Max default and the strongest tier for Claude Pro. Access via Claude API, AWS Bedrock, Google Vertex AI, and Microsoft Foundry; model ID claude-opus-5.

As of publication, the claim remains disputed and unproven. The White House has not released evidence; experts question the two-week timeline; Ryan Greenblatt's finding that K3 identifies as Claude and leaks internal version IDs is the strongest technical indirect signal. More K3 detail in our Kimi K3 review.

Moonshot committed to July 27, 2026 for full weight release. At publication, weights were not yet public and outside researchers could not fully verify architecture or scores. Use our open weights countdown for release-day checks.

Greenblatt's analysis shows K3 abnormally often identifies as Claude when asked, sometimes outputting claude-opus-4-5-20250929 and similar internal deployment IDs — more accurately than real Claude models. That likely points to training data with Claude deployment metadata, but does not by itself prove distillation. For remote dev setup questions, see the help center.