InfoResearchPeer-reviewedLLM-specific
LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
- Published
- Record updated
Summary
RobustJudge is an automated, modular framework that systematically tests how robust LLM-as-a-Judge systems are across datasets, prompt templates, judge models, attacks and defenses. The study covers 15 attack methods and 8 defense strategies across 13 models, and finds that LLM-based judges stay susceptible under both pointwise and pairwise protocols. Robustness is highly sensitive to prompt-template and judge-model choice, and optimization-based attacks with long suffixes can substantially inflate scores from both PAI-Judge variants.
Related items
- InfoDoes Target Alignment Mean Target Recovery? An Evidence-Ladder Study of Adversarial Claims on Contrastive EncodersSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoBRANCH: Bypassing Multi-Scanner AI GuardrailsSimilar attack · Arxiv (cs.CR + cs.CL + cs.LG)
- InfoDetecting Adversarial Images through Response Profiles of Vision-Language ModelsSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoGraphRectify: Graph-Based Transfer of Adversarial Example Detectors Across Neural NetworksSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoVCR-Bench: A Modular Open-Source Benchmark for Video Classification RobustnessSimilar attack · Arxiv (cs.RO + cs.CV security)