InfoResearchPreprintLLM-specific
Robust Decentralized Fairness Auditing
- Published
- Record updated
Summary
Auditopus is a decentralized method for auditing a large language model's fairness, where multiple auditors each query the LLM and share only cumulative statistics vectors instead of raw queries. The authors show that a single adversarial auditor can fabricate these vectors to make an unfair LLM appear fair, and that the scheme counters this by down-weighting auditors whose vectors are statistically inconsistent with earlier ones. Against an optimizing attacker, it reduces audit error by up to 78% on average relative to no defense and at least 62% relative to robust aggregation baselines.
Related items
- LowAnthropic Cuts Live Internet Access for Internal AI Tests After Claude Exploits Injection FlawsSimilar attack · The Hacker News
- CriticalCVE-2026-108263: Astron Agent code-node execution as root through workflow run endpointsSimilar attack · NVD/CVE Database
- MediumHackers abuse Google Ads, Bing redirects to push Claude ClickFix attacksSimilar attack · BleepingComputer
- CriticalHermes Agent - PKCE Session Takeover via Redirect-URI Parser ConfusionSimilar attack · Tenable Research Advisories
- LowSocial Engineering AI Agents: The New BEC for 2026Similar attack · Dark Reading