Systematic Evaluation of Dataset Watermarking for Intellectual Protection
Summary
This research evaluates nine different dataset watermarking methods (techniques for embedding hidden markers into training data to prove ownership of AI models) by testing how well they preserve prediction accuracy, verify ownership, and resist attacks. The study finds that while most watermarking methods cause minimal overall accuracy loss, they can significantly harm performance on specific classes, fail completely when training data is mixed together, and have various other robustness vulnerabilities.
Classification
Related Issues
CVE-2024-37052: Deserialization of untrusted data can occur in versions of the MLflow platform running version 1.1.0 or newer, enabling
CVE-2025-45150: Insecure permissions in LangChain-ChatGLM-Webui commit ef829 allows attackers to arbitrarily view and download sensitive
Original source: http://ieeexplore.ieee.org/document/11596568
First tracked: September 3, 2026 at 08:02 PM
Classified by LLM (prompt v3) · confidence: 85%