Новости
Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
arXiv cs.AI · Опубликовано · 3 мин чтения
За 30 секунд
- Что произошло
- Researchers found that LLM deployment configurations, not model weights alone, determine how models validate pseudo-scientific claims, with Grok showing inconsistent behavior across interfaces.
- Почему это важно
- Engineers deploying LLMs in production should care, especially when models serve as knowledge references or validation systems for contested claims.
- На что обратить внимание
- The study tested only one pseudo-scientific framework across four months; findings may not generalize to other false claims or deployment contexts.
Послушать это резюме
Из статьи
-->
Computer Science > Computers and Society
arXiv:2607.22513v1 (cs)
[Submitted on 24 Jul 2026]
Title: Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
Authors: Davide Scarso , Hugo Noronha de Almeida , Joaquim Pina
View a PDF of the paper titled Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science, by Davide Scarso and 1 other authors
View PDF
Abstract: Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. We tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate ethnonationalist pseudo-science derived from Frank Salter's biosocial framework across four temporal snapshots (October 2025-February 2026), via both API and web interfaces. Grok's Fast versions (which power the default user experience on X) consistently assigned credibility scores of 70-75, two to five times higher than all other models (which scored 15-40). This pattern was absent from control prompts testing basic evolutionary consensus and refuted Lamarckian claims, where all models performed comparably. Three additional findings emerged: (1) a silent patch reversed Grok's behaviour from chaotic to stably high validation overnight, without any public documentation; (2) the same Grok m
Фрагмент оригинала. Полный текст читайте в первоисточнике.
- llm
- language model
- claude
- gpt
- gemini
The Agent Architect
Один паттерн, один компромисс, одна история сбоя в продакшене. Короткий еженедельный брифинг для тех, кто строит агентные системы.
Одно письмо в неделю, отписка в один клик. Адрес используется только для рассылки брифинга.