Loading...
Data Exfiltration via Goal Inference
DEGIManipulating an agent's goal inference mechanisms to extract sensitive data by framing data access as necessary to achieve legitimate-seeming objectives.
Example Scenario
Manipulating an AI assistant to infer that accessing all customer records is necessary to "provide excellent personalized service," causing it to exfiltrate sensitive data while believing it's pursuing a legitimate service improvement goal.
Testing Objectives
- Test goal-based access controls
- Assess objective validation
- Evaluate data access justification
- Validate exfiltration detection
Defensive Strategies
- Explicit data access policies
- Goal-independent access controls
- Data minimization enforcement
- Access justification validation
- Exfiltration pattern detection
Key Features
- Goal-justified data access
- Objective manipulation for exfiltration
- Legitimate-appearing data requests
- Inference system exploitation
Use Cases
- Goal inference security testing
- Data access justification validation
- Objective alignment assessment
- Exfiltration prevention evaluation
Tools & Frameworks
Security Risks
Ethical Guidelines
- •Only test with authorized data access
- •Never exfiltrate production data
- •Report inference vulnerabilities responsibly
- •Focus on improving access controls
- •Consider privacy and compliance impact
Remember: This information is for educational and defensive security purposes only. Always ensure you have proper authorization before testing any techniques.
From the engineer behind this catalog
Get your agent system red-teamed
The attacks documented here work on production agent systems every day. Have yours tested before someone else does: prompt injection, jailbreaks, tool misuse and data exfiltration, with every finding written up next to its fix.
€750 instead of €1,500, one week, written report and walkthrough call, until 30 September