Hugging Face released LFM2.5-VL-DSpark for accelerating vision-language models.
News Hub
What actually shipped in agent engineering, pulled from the labs, arXiv and Hacker News.
See who we followNVIDIA Nemotron 3 Diarization enables real-time multi-speaker identification for AI applications.
UK AISI and EvalEval work to make AI benchmark results reproducible.
Hugging Face Transformers library now supports running llama.cpp quantized models.
LLM pruning can be formulated as an Ising optimization problem for block removal.
Tokenizers v1 measures encode, decode performance and scaling characteristics.
Article examines whether AI agents consistently reproduce successful task performance.
AUTOMATIC1111 is being rebuilt using Gradio Workflow.
Async GRPO with LoRA training runs across Hugging Face Jobs using a bucket and proxy without NCCL.
IBM released Granite Time Series PatchTST-FM-r2 model with commercial-friendly license.
NeoMME is an efficient multimodal-native and multilingual encoder model.
Coding agents can be equipped with persistent memory systems owned and controlled by users.
A coding model was trained to generate watercolor paintings using TRL and OpenEnv frameworks.
A 350M parameter model was fine-tuned using 100 GRPO steps to improve structured output generation.
Hugging Face released 200+ WebGPU kernels for running AI models locally.
Sentence Transformers enables training and finetuning of multi-vector embedding models.
Granite 4.2 LLMs are large language models with documented construction methods.
A 4-bit quantized model achieves better performance than its original full-precision version through quantization-aware healing techniques.
Hugging Face Inference Endpoints, Jobs, and Buckets provide infrastructure for the Papers with Code search functionality.
Measuring benchmark optimization in speech recognition.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.







