新闻
D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language Models
arXiv cs.AI · 发布于 · 阅读约3分钟
30秒读懂
- 发生了什么
- Researchers introduced D-Score, a spectral method to detect hallucinations in large language models by analyzing hidden activation patterns from a single forward pass.
- 为何重要
- Matters for engineers building LLM applications who need hallucination detection without external verifiers, retrieval systems, or multiple model generations.
- 注意
- Paper is under review and not yet peer-reviewed. Evaluation limited to two datasets. Requires per-model and per-layer calibration with tolerance parameters.
收听本摘要
文章节选
-->
Computer Science > Computation and Language
arXiv:2607.24586v1 (cs)
[Submitted on 27 Jul 2026]
Title: D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language Models
Authors: Bianca Raimondi , Davide Evangelista , Maurizio Gabbrielli , Elena Loli Piccolomini
View a PDF of the paper titled D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language Models, by Bianca Raimondi and 3 other authors
View PDF HTML (experimental)
Abstract: Large Language Models can produce fluent text that is false, unsupported by the available evidence, or inconsistent with information that appears to be internally represented by the model. We study hallucination detection from the geometry of hidden activations and introduce the D-Score, a simple spectral statistic computed from a single forward pass. For a fixed model, layer, and tolerance parameter, the D-Score counts how many singular directions of the hidden activation matrix have singular values that remain close to the leading one. We use this quantity as a hallucination score, classifying an input text as hallucinated when its D-Score is larger than a pre-defined quantity. The motivation is that, when a model processes a text that conflicts with information available in its own internal state, the hidden representation may encode both the asserted content and some form of count
节选自原文。请前往来源阅读全文。
- language model
The Agent Architect
每周一个模式、一个权衡、一个生产事故案例。为构建智能体系统的人准备的每周简报。
每周一封邮件,一键退订。您的地址仅用于发送简报。