Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
ZUMA is a training-free framework for multimodal anomaly detection (MAD, identifying unusual patterns using both image and 3D data together) that works without needing labeled training examples, addressing privacy concerns. It uses CLIP (a model trained on images and text) and introduces cross-domain calibration (a technique that bridges differences between how CLIP was trained and how 3D point cloud data works) and dynamic semantic interaction (using natural language descriptions as reference points to spot anomalies) to detect defects in 2D images, 3D objects, or both together without requiring training.
This research introduces DeSA, a protocol for secure aggregation (a privacy technique that protects individual data while combining results) in federated learning (a machine learning approach where multiple devices train a shared model without sending raw data to a central server) across decentralized device-to-device networks. The protocol addresses challenges in zero-trust networks (environments where no participant is automatically trusted) by using zero-knowledge proofs (cryptographic methods that verify information is correct without revealing the information itself) to verify model training, protecting against Byzantine attacks (attacks where malicious nodes send false information to disrupt the system), and employing a one-time masking method to maintain privacy while allowing model aggregation.
Subgraph Federated Learning (FL, a system where pieces of a graph are distributed across multiple devices to protect data privacy) is vulnerable to backdoor attacks (hidden malicious functions that cause a model to behave incorrectly when triggered). Researchers developed BEEF, an attack method that uses adversarial perturbations (carefully crafted small changes to input data that fool the model) as hidden triggers while keeping the model's internal parameters unchanged, making the attack harder to detect than existing methods.
When multiple AI agents work together, they can amplify hallucinations (false or made-up information generated by LLMs), creating a 'snowballing effect' where errors get worse as agents collaborate. This paper proposes a framework that uses semantic reasoning (understanding meaning and logical relationships) to detect and reduce hallucination spread in multiagent systems without changing how the underlying models work.
Fix: The paper proposes a semantic reasoning-based mitigation strategy using bidirectional entailment clustering (a technique that checks logical relationships between statements in both directions). According to the source, this approach 'mitigates hallucination propagation caused by the model itself' and alleviates errors caused by external knowledge deficiencies, and 'effectively reduces the propagation of hallucinations' according to their experiments.
IEEE Xplore (Security & AI Journals)This paper introduces ATRNet-STAR, a new large dataset for training AI models to recognize objects in Synthetic Aperture Radar (SAR) images, which are pictures taken using microwave radar rather than cameras. The dataset contains over 190,000 labeled images of 40 different vehicle types collected under realistic conditions, making it about 10 times larger than the previous standard dataset (MSTAR from the 1990s). The authors evaluated 15 different AI methods on this dataset to show how well current techniques work and identify directions for future research.
Researchers discovered that speech translation (ST) systems, which convert spoken words from one language to another, can be tricked by specially crafted audio manipulations that are imperceptible to human ears. They demonstrated two attack methods: adapting techniques from ASR (automatic speech recognition) attacks and using music-based perturbations to guide the system toward producing harmful outputs. These attacks worked across multiple languages and models, revealing a fundamental weakness in how current speech translation systems process and understand audio.
This paper presents S2WIB, a new method for multiview clustering (organizing data points into groups when information comes from multiple different sources or perspectives). The method improves on existing approaches by using both view quality and self-supervised learning (learning from patterns the model finds in data itself) to determine how much weight to give each data source, and by considering both complementary information and consistency between individual view clustering results and the final result.
This research addresses multi-object navigation (MON), where an AI agent must find multiple targets in unknown environments by balancing immediate actions with long-term planning. Current methods focus too much on local path optimization, causing slow learning and getting stuck in trap states. The researchers propose GDWO (Goal-oriented Dynamic Weight Optimization), an algorithm that dynamically adjusts how much each target task contributes to the overall optimization by using gradient-based updates (mathematical techniques that improve decisions step-by-step) and normalizing weights based on navigation success rates, which improves learning efficiency and path planning.
Adversarial imitation learning (AIL, a technique where an AI learns to mimic expert behavior by competing against a discriminator network) has worked well in practice but lacked solid theoretical foundations except in oversimplified settings. This paper introduces OPT-AIL (optimization-based adversarial imitation learning), a new framework that works with general function approximation (flexible neural network models rather than simple lookup tables), and proves it can learn expert-level policies efficiently while remaining practical to implement.
This survey examines how rewards (scoring systems that guide AI behavior) are designed to align LLMs (large language models, or AI systems trained on massive amounts of text) with what humans want them to do. The review organizes the field by asking how rewards are mathematically defined, how they are built using different data sources and methods, how they work with different training approaches like reinforcement learning from human feedback (RLHF, a technique where humans rate AI outputs to improve performance), and how they are tested for safety and effectiveness.
Mobile super apps (large platforms that host smaller third-party applications, called miniapps, which share the same underlying services) create new security risks because multiple apps can access shared resources and data. Researchers studied how these ecosystems work, identified security vulnerabilities and potential abuses, and developed recommendations to make super app platforms safer while keeping them easy to use.
This research analyzes how discussions about Generative AI spread across different industries (like media, healthcare, and finance) in the six months after ChatGPT's release, using social media data and innovation theory. The study found that different industries had different concerns: media and marketing focused on content generation with positive views, while healthcare and finance were more cautious and focused on analysis. Misinformation was the biggest concern overall, and the research showed that emotional reactions (sentiment) were the main factor driving how quickly information about AI spread between people.
Generative artificial intelligence (GAI, AI systems that create new text, images, or code) is significantly changing how information systems are taught in universities. IS educators are discussing both the benefits and risks of GAI, including concerns about academic integrity (students using AI to cheat), and they are developing recommendations for how to responsibly teach with and about GAI in the classroom.
Industrial process monitoring systems often perform poorly when dealing with noise (unwanted signal disturbances) and outliers (unusual data points) across different working conditions. This research proposes a transfer learning (a technique where a model trained on one task is adapted for a different but related task) framework that combines several advanced neural network approaches, including a variational autoencoder generative adversarial network (a type of AI that learns to generate and discriminate realistic data patterns) and dictionary learning (a method that finds the simplest way to represent complex data), to make monitoring systems more robust and reliable across different industrial scenarios.
The article demonstrates how attackers can use crafted prompts to trick AI assistants into running harmful database queries through prompt-to-SQL injection attacks (where malicious instructions hidden in user input cause an AI to generate dangerous database commands). It identifies vulnerabilities in real systems and describes practical defenses including query filtering, rewriting, data preloading, and using another AI model as a security guard.
Fix: The source mentions four explicit defenses: query filtering, query rewriting, data preloading, and large-language-model-based guards (using another AI model to validate or block dangerous queries).
IEEE Xplore (Security & AI Journals)This article presents FedMC-ADMM, a new algorithm for federated matrix completion (MC, the process of predicting missing values in datasets split across multiple computers). The algorithm combines ADMM (alternating direction method of multipliers, an optimization technique that breaks complex problems into simpler parts) with privacy-preserving federated learning (FL, collaborative AI training where data stays on users' devices rather than being sent to a central server), and the researchers show it converges faster and performs better than existing methods on real-world datasets like Netflix and MovieLens.
Federated learning (a training method where multiple computers learn together while keeping their data private) struggles when some classes have much more training data than others, a problem called head-tail imbalance. Researchers introduced FedGRE, a new approach that improves the shared learning signals (called gradients, which guide how the model updates) using two techniques: accumulation diffusion (mixing stored gradient information with random noise to balance classes) and accumulation refinement (using stored information as a reference point to keep updates stable). Tests on six datasets showed FedGRE outperformed 14 competing methods while protecting privacy.
Fix: The paper proposes FedGRE as a solution, which uses two mechanisms: (1) accumulation diffusion, which "amalgamates accumulated gradients with stochastic gradient perturbations to alleviate class imbalance," and (2) accumulation refinement, which "utilizes the accumulation as an anchor to calibrate global gradient updates, ensuring consistency and mitigating oscillations." The approach also implements "a consistency integration technique to incorporate the refined accumulation into the global model, guaranteeing privacy-preserving and class-balanced global optimization."
IEEE Xplore (Security & AI Journals)This research paper analyzes how companies that invest in digital technologies, including AI, affect their greenhouse gas emissions and natural resource use. The study found that companies investing in these technologies tend to reduce their emissions and consume fewer natural resources, suggesting that digital tools can help address environmental challenges.
This paper addresses white-box attacks (scenarios where attackers can see all the inner workings of an encryption system and control the computer it runs on), which are harder to defend against than black-box attacks (where attackers cannot see the implementation). The authors propose a new method to protect symmetric encryption algorithms that use substitution-permutation networks (a common encryption structure that substitutes and rearranges data) by adding secret components to lookup tables, making the encryption stronger without changing the final encrypted message.
Researchers studied how multimodal generative AI systems (AI that creates both text and images) could be misused to generate fake news by testing AI image generators like Stable Diffusion 3 and evaluating whether GPT-4 (a large language model) could detect these false images. The findings show that current AI has both strengths and weaknesses in creating misinformation, and highlight that stronger safeguards, transparent monitoring, and human oversight are needed to prevent AI from spreading false information.