NVDA, AMD · US
TileRT: ultra-high interactivity on NVIDIA GPUs? | SemiAnalysis note
Fast modes show users pay for latency. TileRT statically compiles decode into a persistent kernel; on InferenceX B200 can reach hundreds of tok/s/user—materially faster interactivity iso-cost vs classic engines. See CS-4, Rubin TCO, AgentX v3.
Cite a section with a deep link, e.g. /en/r/sa-tilert-inferencex-2026#thesis
Snapshot
- Date
- 2026-08-10
- 基准场景
- InferenceX / GLM5 等
- B200 互动性量级
- 最高约 ~500 tok/s/user(文中)
- 相对传统引擎
- 约 2–3× 互动性(场景依赖)
As-of 2026-09-25 (weekly refresh; equities aligned to §A. Missing series are N/A/null. Not investment advice.)
Structured research note on a SemiAnalysis piece (summary + investable mapping)—not a reprint. Defer to the original for detail.
Thesis
Fast modes show users pay for latency. TileRT statically compiles decode into a persistent kernel; on InferenceX B200 can reach hundreds of tok/s/user—materially faster interactivity iso-cost vs classic engines. See CS-4, Rubin TCO.
Analysis
Why it matters
Paid fast modes prove latency premium. Labs evaluate Cerebras/Groq etc., so NVIDIA software must answer whether GPU fleets can reach the same interactivity frontier.
What TileRT does
Statically compiles the decode graph into a persistent kernel, maximizing overlap of compute/memory/comms. On InferenceX, a B200 decode server can hit ~500 tok/s/user in cited cases (~2–3× vs classic engines depending on workload). Edge is mainly decode tail; TTFT is fine, not magical.
Ecosystem
Competes with vLLM/SGLang/Dynamo and specialty silicon. AMD MI455X / TPUv7 will land in InferenceX. Read iso-cost/iso-power and agentic traces—not peak tok/s alone—using AgentX InferenceX v3 (393 Claude Code traces, 95%+ KV reuse) as the production-proxy benchmark, alongside frontier open weights in open models catching up.
Implications
| Angle | Implication | Site map |
|---|---|---|
| NVDA 软件护城河 | 抬高 GPU 互动性上限,延缓份额流失到专用硅 | NVDA · Rubin |
| 专用推理硅 | 仍可能在极低 batch 占优,但差距可被软件压缩 | CS-4 |
| AMD / 开源栈 | SGLang 等需持续对标 | AMD |
Outlook
Watch InferenceX verified Rubin/TPU/MI455X numbers, TileRT production adoption, Fast-tier pricing, vs CS-4/Groq on same loads; AgentX v3 extends the benchmark from single-node interactivity to multiturn agentic full-stack composability.
Risks
References
- 原文(SemiAnalysis): Ultra-High Interactivity on NVIDIA GPUs? - TileRT InferenceX
- InferenceX v2: NVIDIA Blackwell Vs AMD vs Hopper
- 对照 CS-4: Cerebras CS-4
- AgentX v3 (agentic full stack): AgentX InferenceX v3 · original
Not investment advice. Copyright remains with SemiAnalysis / authors; this is an index note.
Comments
Sign in to comment