NVDA, AMD, GOOGL, META, MSFT · US
Are Open Models Catching Up? Three eras | SemiAnalysis note
SemiAnalysis: scaling / reasoning / agentic eras; open weights close gap in ~half the time each era (Kimi K2.6 vs Opus 4.5 ~4.8 mo), but benchmarks ≠ product. See AgentX v3, CUDA moat, NVDA.
Cite a section with a deep link, e.g. /en/r/sa-open-models-catching-up-2026#thesis
As-of 2026-09-25 (weekly refresh; equities aligned to §A. Missing series are N/A/null. Not investment advice.)
Are Open Models Catching Up? Three eras of convergence
Source: SemiAnalysis — Are Open Models Catching Up? (2026-08-21). Investor Research digest; not a full republish.
Thesis
SemiAnalysis frames LLM progress as scaling → reasoning → agentic eras and finds open weights close the gap to the closed frontier in roughly half the time each era—while benchmark scores still diverge from daily product quality; the team still prefers Anthropic Claude for routine work.
Agentic inference stack reality is captured in the companion AgentX InferenceX v3 release: once workloads shift to long-context, multiturn, sub-agent traces, hardware perf/$ and software composability explain investability better than leaderboards (see CUDA moat and TileRT InferenceX).
Three eras & convergence
Era 1 — Scaling
- Llama 3.1 405B vs GPT-4o: ~6 months gap (SA framing).
- Open models catch up quickly on dense scaling, but data, compute, and post-training resources remain concentrated in closed labs.
Era 2 — Reasoning
- DeepSeek R1 vs OpenAI o1: ~6 months.
- Competition shifts to RL data, verifiers, and inference-time compute—parallel to CPU industry orchestration narratives and ISA landscape edge inference themes.
Era 3 — Agentic
- Kimi K2.6 vs Claude Opus 4.5: ~4.8 months.
- GLM-5.2 vs GPT-5.2: ~6 months.
- Production traits—trace shape, KV reuse, routing—are quantified in AgentX v3 (e.g. P90 ISL 317k, 95%+ KV hit rate).
Product vs benchmark caveat
- SA still rates Claude higher for daily coding/agent tasks; open models may match public leaderboards yet lag on tool stability and long-session coherence.
- RL hill-climbing risk: open community may overfit public evals—contrasts with closed labs’ private eval + product loops (Meta superintelligence, Gemini/GCP).
Compute & stack implications
Training vs inference
- Scaling era favored training GPU capex (NVDA, AMD, custom ASIC via Google–AMD TPU chatter and AVGO).
- Agentic era shifts margin to inference perf/$, KV memory, disaggregation—see GPU utilization, Vera/Rubin NVL72 TCO, Cerebras CS-4.
Cloud & distribution
- Closed APIs (GOOGL Gemini, MSFT Copilot, AMZN Bedrock) still own distribution and billing; open weights compress model-layer margin and raise infra / neo-cloud value (CRWV, NBIS, IREN).
China open vs Western closed
- Kimi, GLM, DeepSeek, Qwen accelerate convergence—consistent with AgentX MI355X / Kimi K3 / MiniMax M3 matrix; export controls remain friction for Western cloud adoption of certain weights.
Outlook
- Gap may keep narrowing in the agentic era, but whether half-life per era holds depends on closed labs widening private data and product RL advantages.
- Track three signals: (1) open leaderboards, (2) production workloads like AgentX, (3) inference stack PR velocity (CUDA moat, TileRT).
- INTC and CPU players may gain in edge/local agents as small open models improve; datacenter agentic battleground remains GPU + full stack.
Risks
- Benchmark vs product divergence; RL overfitting public evals.
- Closed labs retain top post-training compute and user feedback loops.
- Geopolitics, copyright, weight export limits.
- Open acceleration commoditizes the model layer (infra up, undifferentiated API down).
References
- SemiAnalysis — Are Open Models Catching Up? (2026-08-21)
- Related notes: AgentX InferenceX v3 · CUDA moat · TileRT InferenceX · Gemini/GCP · Meta superintelligence
- Related research: NVDA · AMD · GOOGL · MSFT · GPU utilization · CPU industry
Comments
Sign in to comment