Pleias introduced Telco Common Corpus, an open dataset of 10 billion tokens for telecommunications AI development.
What actually shipped in agent engineering, pulled from the labs, arXiv and Hacker News.
See who we follow →Pleias introduced Telco Common Corpus, an open dataset of 10 billion tokens for telecommunications AI development.
Pleias trained a 600M parameter model for RATP subway safety signal detection.
Pleias and NVIDIA release Nemotron-Personas-Belgium, a synthetic persona dataset covering Belgian population demographics.
Europe lost access to Fable series models, exposing reliance on foreign AI infrastructure.
Synth generates specialized training data from custom documents to improve model knowledge.
Three research papers confirm SYNTH synthetic training data matches or beats larger models at lower cost for reasoning tasks.
Pleias and Nvidia release French synthetic personas dataset for training without real sensitive data.
Common Corpus expanded to 2.2 trillion tokens globally, with 53% from non-western countries including China, Japan, India, and Africa.
Pleias and SpineDAO partner to test small structured-reasoning models for spine care against large generic LLMs.
Pleias launched Stratum, an AI-native data layer for parsing, anonymizing, and indexing data on-premise for agentic AI applications.
Pleias released SYNTH, a synthetic dataset from 50,000+ web pages for e-commerce AI, plus two small language models for retrieval.
SYNTH synthetic dataset from Wikipedia trains models with 10-50x less data while achieving top benchmark results.
Recent LLM models can plan, remember, and act independently across long tasks.
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.