In the news
22,580: GPT-2 to Kimi K3, explained
Baseten · Published · 3 min read
In 30 seconds
- What happened
- Kimi K3 contains 22,580 times more parameters than GPT-2, but architectural innovations like linear attention and DeltaNet fundamentally changed how models process sequences.
- Why it matters
- Engineers building or optimizing large language models need to understand how efficiency techniques evolved from 2019 to 2026 to make informed architecture choices.
- Watch out
- Linear attention trades softmax expressiveness for fixed-size state, reducing memory bandwidth but potentially losing fidelity. DeltaNet addresses information interference in fixed caches but adds complexity.
Listen to this summary
- gpt
- kimi
The patterns behind this
- Infini-Attention Architecture
- Process Reward Models & Verifier-Guided Search
- Progressive Consent & Communication
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.