In the news
Scaling Laws for Mixture Pretraining Under Data Constraints
Apple Machine Learning Research · Published · 3 min read
In 30 seconds
- What happened
- Apple researchers studied how to mix scarce target data with abundant generic data during language model pretraining, finding optimal repetition rates.
- Why it matters
- Engineers training models on low-resource languages or specialized domains with limited data need guidance on mixture ratios and repetition.
- Watch out
- Results span 2000 runs but optimal repetition rates vary by target data size, compute budget, and model scale, requiring case-by-case analysis.
Listen to this summary
- language model
The patterns behind this
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.