ニュース
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation
arXiv cs.AI · 公開日 · 読了3分
30秒で要点
- 何が起きたか
- Research shows GPT-5.6-sol gives safer advice when exposed directly to harmful objectives than when intermediate agents reframe them, revealing a compositional safety gap.
- なぜ重要か
- Teams building multi-stage AI workflows or deploying LLMs in production need to understand how instruction laundering through intermediary agents can circumvent safety behaviors.
- 注意点
- The study tests only 25 trade-off profiles on one model; the internal mechanism behind the behavioral reversal remains unidentified and may not generalize across architectures.
この要約を音声で聴く
記事より
-->
Computer Science > Artificial Intelligence
arXiv:2607.21518v1 (cs)
[Submitted on 23 Jul 2026]
Title: Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation
Authors: Linjun Li
View a PDF of the paper titled Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation, by Linjun Li
View PDF HTML (experimental)
Abstract: Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction. Using OpenAI's gpt-5.6-sol model alias, we test 25 pre-specified mirrored trade-off profiles. Direct exposure to an objective authorizing concealment, fabrication, and pressure produced advice net opposed to its target. After an Id and Censor transformed the same objective into affect and a constraint-rewritten, target-bearing intention, the user-facing Superego---which saw the preferred direction but not the raw objective, its manipulative clauses, or its source---produced advice net aligned with the target.
This behavioral reverse shift is consistent with the model recognizing or distrusting the manipulative motive, although we do not identify its internal mechanism. The second result exposes a compositional safety gap: a current high-capability model can be used as the user-facing component of an automated, multi-stage workflow serving an e
原文からの抜粋です。全文は配信元でお読みください。
- agent
- llm
- multi-agent
- gpt
The Agent Architect
1つのパターン、1つのトレードオフ、1つの本番障害事例。エージェントシステムを構築する人のための短い週刊ブリーフィング。
週1回のメール、ワンクリックで購読解除できます。アドレスはブリーフィングの送信のみに使用します。