Skip to content
InfoResearchPeer-reviewedLLM-specific

A novel privacy-preserving large language model integrating trust-weighted and ethical gradient masking

Published
Record updated
View JSON

Summary

Researchers propose a privacy-preserving large language model training approach, the trust-weighted and ethical gradient model, to address the privacy utility tradeoff in LLMs trained on personally identifiable information. The method combines trust-weighted memory, ethical gradient masking under a tracked (ϵ, δ) privacy budget, and an ethical boundary layer that screens risky outputs. Experiments on the Pile dataset report reduced PII leakage with baseline utility maintained, and the authors claim resistance to membership inference, extraction and linkage attacks.