Skip to content
InfoResearchIndustryLLM-specific

Sidestepping Evaluation Awareness and Anticipating Misalignment with Production Evaluations

Published
Record updated
View JSON

Summary

Marcus Williams, Cameron Raymond and Micah Carroll describe a pipeline that uses de-identified ChatGPT production traffic to build realistic alignment evaluations. The method resamples the final model response from each conversation and labels the new responses with LLM monitors, either to discover unknown misaligned behaviors or to estimate how often known ones occur. The authors note the pipeline depends on a monitor's ability to detect undesirable behaviors and cannot guarantee catching all of them.