Новости
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance
arXiv cs.AI · Опубликовано · 3 мин чтения
За 30 секунд
- Что произошло
- Researchers released BioSecBench-Surveillance, a benchmark with 100 evaluations testing whether AI agents can choose correct analysis pipelines for pathogen genomic sequencing data.
- Почему это важно
- Bioinformaticians and biosecurity engineers evaluating AI for outbreak response need to assess whether models can reliably perform genomic surveillance analysis.
- На что обратить внимание
- Top models achieved only 50 percent accuracy; even when agents selected correct workflows, they failed on critical details like reference selection, thresholds, and normalization choices.
Послушать это резюме
- agent
The Agent Architect
Один паттерн, один компромисс, одна история сбоя в продакшене. Короткий еженедельный брифинг для тех, кто строит агентные системы.
Одно письмо в неделю, отписка в один клик. Адрес используется только для рассылки брифинга.