CIRCUS: A Causal Intervention-Based Framework for Enhancing Counterfactual Fairness in Trained Classifiers
inforesearchPeer-Reviewed
researchsafety
Source: IEEE Xplore (Security & AI Journals)March 16, 2026
Summary
This research paper presents CIRCUS, a framework designed to reduce bias in AI classifiers by using causal intervention (a technique that simulates how changing certain input features affects predictions). The framework generates synthetic examples where sensitive attributes like race or gender are modified, then retrains the classifier on these examples to ensure predictions remain fair and consistent regardless of those sensitive attributes. Experimental results show the method significantly reduces bias metrics while maintaining classifier accuracy.
Classification
Attack SophisticationModerate
Impact (CIA+S)
safety
AI Component TargetedModel
Monthly digest — independent AI security research
Original source: http://ieeexplore.ieee.org/document/11435129
First tracked: September 26, 2026 at 02:01 AM
Classified by LLM (prompt v3) · confidence: 85%