Crowdsourced game testing of Olmo 3 revealed model vulnerabilities and behavior exploits.
News Hub
What actually shipped in agent engineering, pulled from the labs, arXiv and Hacker News.
See who we followGoodfire used Ai2's open post-training stack to trace unwanted LLM behavior to individual training examples.
BenchMIRT audits LLM benchmarks question by question to reveal measured capabilities.
Ai2 and Providence Swedish Cancer Institute expanded collaboration after AutoDiscovery validated a new immune signal in breast cancer.
Thai researchers adapted Ai2's Dolma toolkit to create Mangosteen, a 47-billion-token Thai language corpus for improved model performance.
Georgia Tech researchers used Ai2's open Olmo stack to identify training data sources influencing social reasoning capabilities in language models.
Models infer drug class from name patterns in training data rather than learning drug knowledge.
TutorMoments evaluates whether AI tutors recognize when to support students versus encourage independent reasoning.
Ai2 expands partnership with Hugging Face to distribute open models, datasets, and benchmarks.
Reliable agents depend on deterministic tools, guardrails, isolated infrastructure, and real-world evaluations more than model choice.
Danish Foundation Models uses FlexOlmo to build FlexMoRE, a modular LLM letting institutions contribute specialized experts without sharing sensitive data.
Olmo Hybrid models predict meaning-bearing tokens better than transformers, while transformers excel at verbatim copying.
Domyn and AISquared built models for regulated industries using Ai2's open releases.
MolmoMotion is an open language-guided 3D motion forecasting model for robotics.
olmo-eval provides an open evaluation workbench for running benchmarks across LLM development checkpoints.
PointCheck uses Molmo and Olmo 3 models to test web accessibility.
OlmoEarth v1.1 reduces remote-sensing model compute costs by up to 3x while maintaining similar performance.
AIMIP is an open benchmark for evaluating AI climate models against conventional models on historical and future scenarios.
Artificial Analysis uses Ai2's IFBench to evaluate whether models reliably follow complex multi-part instructions.
EMO trains mixture-of-experts models where modular expert groups emerge, allowing selection of task-specific subsets with minimal performance loss.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.








