Researchers decomposed speech naturalness into 10 linguistically grounded perceptual dimensions to evaluate automated TTS evaluation methods.
What actually shipped in agent engineering, pulled from the labs, arXiv and Hacker News.
See who we follow →Researchers decomposed speech naturalness into 10 linguistically grounded perceptual dimensions to evaluate automated TTS evaluation methods.
Sparse autoencoders identify interpretable feature directions in multimodal LLMs to isolate changes from multimodal training.
Researchers created a Dutch governmental LLM evaluation framework reflecting public administration values and non-English linguistic requirements.
Dark Souls Learning Environment provides 22 boss encounters as game-playing agent benchmarks with real-time combat and sparse rewards.
Decoding-level taboo tests LLM robustness when system prompts and constraints force models away from nominal generation paths.
Consilience enables verifier-free test-time scaling for LLM reasoning without external verifiers like compilers or value functions.
Thinking Mode Fusion training systematically studies data ratios and schedules between concise and long-form reasoning modes.
BDH-CQ combines in-context learning with recurrent latent reasoning, updating memory iteratively without verbalizing intermediate steps.
SHE evolves LLM agent safety harnesses through trajectories to manage context, memory, tools, and runtime control.
ArchAgent v2 scales automated microarchitecture search to multi-level data prefetching despite vast search spaces and long simulations.
Sci-VBench benchmark contains 1,253 expert-annotated examples for evaluating video generation across 60 scientific subjects in four disciplines.
Researchers demonstrate methods to extract encrypted chain-of-thought reasoning traces from proprietary LLM APIs.
LLM-driven verification layers check feasibility and ethics of robot autonomy planning before execution.
Autonomous research agents operate like greybox fuzzers, generating experiments faster than humans can validate them.
CEAA architecture enables embodied intelligent virtual agents with cognitive capabilities in real-time interactive environments.
On-policy distillation fails when students exploit repetitive loops for token agreement despite globally flawed responses.
RA-FinBERT combines LoRA adaptation with rule-derived features for low-resource financial sentiment classification.
POLIS research program studies institutional design of multi-agent AI systems to understand safety mechanisms.
SKALD framework uses abstract skills as privileged signals for on-policy self-distillation in reinforcement learning.
Macaron-V1 agent model learns from experience in real environments and improves after deployment using mixture-of-LoRA.
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.