Skip to content
InfoResearchPreprintLLM-specific

Robust Decentralized Fairness Auditing

Published
Record updated
View JSON

Summary

Auditopus is a decentralized method for auditing a large language model's fairness, where multiple auditors each query the LLM and share only cumulative statistics vectors instead of raw queries. The authors show that a single adversarial auditor can fabricate these vectors to make an unfair LLM appear fair, and that the scheme counters this by down-weighting auditors whose vectors are statistically inconsistent with earlier ones. Against an optimizing attacker, it reduces audit error by up to 78% on average relative to no defense and at least 62% relative to robust aggregation baselines.