新闻
Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
arXiv cs.AI · 发布于 · 阅读约3分钟
30秒读懂
- 发生了什么
- Researchers found that LLM deployment configurations, not model weights alone, determine how models validate pseudo-scientific claims, with Grok showing inconsistent behavior across interfaces.
- 为何重要
- Engineers deploying LLMs in production should care, especially when models serve as knowledge references or validation systems for contested claims.
- 注意
- The study tested only one pseudo-scientific framework across four months; findings may not generalize to other false claims or deployment contexts.
收听本摘要
文章节选
-->
Computer Science > Computers and Society
arXiv:2607.22513v1 (cs)
[Submitted on 24 Jul 2026]
Title: Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
Authors: Davide Scarso , Hugo Noronha de Almeida , Joaquim Pina
View a PDF of the paper titled Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science, by Davide Scarso and 1 other authors
View PDF
Abstract: Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. We tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate ethnonationalist pseudo-science derived from Frank Salter's biosocial framework across four temporal snapshots (October 2025-February 2026), via both API and web interfaces. Grok's Fast versions (which power the default user experience on X) consistently assigned credibility scores of 70-75, two to five times higher than all other models (which scored 15-40). This pattern was absent from control prompts testing basic evolutionary consensus and refuted Lamarckian claims, where all models performed comparably. Three additional findings emerged: (1) a silent patch reversed Grok's behaviour from chaotic to stably high validation overnight, without any public documentation; (2) the same Grok m
节选自原文。请前往来源阅读全文。
- llm
- language model
- claude
- gpt
- gemini
The Agent Architect
每周一个模式、一个权衡、一个生产事故案例。为构建智能体系统的人准备的每周简报。
每周一封邮件,一键退订。您的地址仅用于发送简报。