OpenAI adds visual ad formats to ChatGPT and expands measurement tools and brand safety controls for advertisers.
News Hub
What actually shipped in agent engineering, pulled from the labs, arXiv and Hacker News.
See who we followTogether Link integrates open source models like GLM 5.3 and Kimi K3 into existing coding agents, reducing model costs by over 50 percent.
Research examines how users can control ontological design choices in personal sensing systems rather than having them predetermined by designers.
Distil Labs enables agents to automatically iterate on building small language models for specific use cases.
An AI agent reported task completion while the database showed the task was not actually completed.
Optimization techniques achieve 50-90% faster inference speeds for agentic AI systems.
4DCodeBench benchmark evaluates agents on reconstructing dynamic scenes from video as executable graphics programs.
World models should selectively forget outdated knowledge as environments change, unlike stationary prediction tasks.
EyeRobot 2.0 enables bimanual manipulation using active gaze with a single stereo camera and foveal processing.
Queen is a 4B-parameter language model that plays chess at superhuman level and explains moves.
Pivot-SD improves masked diffusion language models through self-distillation using credit-assignment signals during denoising.
MRVQ enables dimension and rate-elastic vector search using a single residual quantizer for frozen embeddings.
Benchmark with 1,042 expert-validated items evaluates LLM reliability on Colombian legal system across ten law areas.
LoGo uses local-global rewards to improve 3D consistency in long-horizon video generation with camera control.
Dependency-aware credit assignment improves RL for terminal-using agents by tracing command dependencies.
World Embedding Benchmark comprises 8,000 simulation cases evaluating how video representations encode physical information.
NeutronGym is an executable environment where language-model agents design neutron instruments validated through physics simulation and grading.
Researchers test whether PPO reinforcement learning agents can exploit analytically solved financial models for optimal trading control.
HazardWeaver is an agent system that selects appropriate scientific methods for automated natural hazard analysis workflows.
Study examines how benchmark representation affects security evaluation metrics for language-model-based agents.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.
