{"data":{"id":"d9ec75b8-5203-413b-b5ac-1c93dbd34938","title":"LLMs and Contextual Integrity","summary":"Researchers have identified a significant problem with how large language models (LLMs) handle sensitive information when they remember details from past conversations. When LLMs use persistent memory (stored information from previous interactions), they often leak private details in situations where those details shouldn't be shared, even in frontier models (the most advanced current systems) which showed up to 69% attribute-level violations (inappropriate leaking of specific user details). One research team found that simply asking an LLM to be more careful about privacy doesn't work well because models tend to either share everything or nothing instead of making thoughtful, context-specific decisions.","solution":"One paper explicitly describes a solution: develop a reinforcement learning (RL, a machine learning technique where a system learns by receiving rewards for good behavior) framework that teaches models to reason explicitly about contextual integrity when deciding what information to disclose. The authors demonstrate that this approach \"substantially reduces inappropriate information disclosure while maintaining task performance\" using a synthetic dataset of only 700 examples with diverse contexts. The improvements from this method transfer to established privacy benchmarks with human annotations, showing the approach works across multiple model sizes and families.","labels":["research","privacy"],"sourceUrl":"https://www.schneier.com/blog/archives/2026/08/llms-and-contextual-integrity.html","publishedAt":"2026-08-18T10:40:16.000Z","cveId":null,"cweIds":null,"cvssScore":null,"cvssSeverity":null,"severity":"info","attackType":["pii_leakage"],"issueType":"news","affectedPackages":null,"affectedVendors":["OpenAI"],"affectedVendorsRaw":["GPT-5","OpenAI"],"classifierModel":"claude-haiku-4-5-20251001","classifierPromptVersion":"v3","cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"patchAvailable":null,"disclosureDate":"2026-08-18T10:40:16.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"moderate","impactType":["confidentiality","safety"],"aiComponentTargeted":"model","llmSpecific":true,"classifierConfidence":0.92,"researchCategory":null,"atlasIds":null}}