AI agents
Systems in which a model plans and takes actions through tools, browsers or other software on someone's behalf.
- All items
- 762
- Last 90 days
- 324
- Change
- +44%vs 225 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 3 |
| Jun 2025 | 4 |
| Jul 2025 | 4 |
| Aug 2025 | 5 |
| Sep 2025 | 11 |
| Oct 2025 | 6 |
| Nov 2025 | 3 |
| Dec 2025 | 8 |
| Jan 2026 | 10 |
| Feb 2026 | 49 |
| Mar 2026 | 89 |
| Apr 2026 | 51 |
| May 2026 | 76 |
| Jun 2026 | 78 |
| Jul 2026 | 112 |
| Aug 2026 | 78 |
| Sep 2026 | 133 |
| Oct 2026 | 37 |
25 items
Lost in the comments: Social context as a single‐pass jailbreak and defense on agentic platforms
Oct 8, 2026LowResearchPeer-reviewedSecurityResearchResearchers built a simulation of Moltbook, a social network for AI agents, and tested 100 JailBreakBench goals wrapped in platform-native posts with bystander comments of aggressive, ethical, or measured valence. Reformatting the prompt as platform context alone raised GPT-4o-mini's attack success rate from 7% to 71% in one pass, and measured, intellectually toned comments were the most dangerous. Ethical comments sharply suppressed attack success, and a 35-fold rise in upvotes left it unchanged, showing valence rather than volume drives the effect.
Fix: Safety-valenced signals, such as ethical comments, are proposed as a deployable defense for agentic platforms.
OpenAlex (peer-reviewed AI security)GenAI and Agentic AI Exploit Roundup Q3 2026
Oct 8, 2026InfoResearchIndustrySecurityIndustryThis roundup covers selected AI-related security incidents and exploit disclosures reported between July 1, 2026 and September 30, 2026. It maps each entry to the OWASP Top 10 for LLM Applications 2026 and the OWASP Top 10 for Agentic Applications 2026, with published CVE references where available.
OWASP GenAI SecurityFrom model trust to software topology: a Perspective on structurally governed agentic AI systems
Oct 7, 2026InfoResearchPeer-reviewedSecurityResearchThis Perspective, published in Frontiers in Computer Science on 2026-10-08, argues that agentic AI safety should be treated as a software-architecture problem alongside model alignment and prompt-level defenses. It proposes a five-plane reference topology (intent, orchestration, execution, oversight, adaptation) with the invariant that deciding components should not directly act, and acting components should not act without supervision. The topology is realized in Polos, an open reference specification.
OpenAlex (peer-reviewed AI security)Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment
Sep 30, 2026MediumResearchPeer-reviewedSecurityResearchResearchers ran the first systematic red-teaming of tool invocation in six coding agents: Cursor, Claude Code, Copilot, Windsurf, Cline, and Trae. They introduce ToolLeak, which exfiltrates agent-internal prompts such as system prompts and tool metadata through required tool parameters, and a two-channel prompt injection in the tool description and tool return that achieves remote code execution. The study reports ToolLeak outperforming prompt-leak baselines, with the best pseudo-recall on 18 of 25 agent-LLM pairs.
OpenAlex (peer-reviewed AI security)A Comparative Survey of Security Risks in AI Systems: From LLMs to AI Agents and Embodied Agents
Sep 3, 2026InfoResearchPeer-reviewedResearchSecurityThis ACM Computing Surveys paper, titled "A Comparative Survey of Security Risks in AI Systems: From LLMs to AI Agents and Embodied Agents," appears in Volume 58, Issue 15, pages 1-38, November 2026. The source text provided contains only the citation details, so its research question, method and findings cannot be summarized from it.
ACM Digital Library (TOPS, DTRAP, CSUR)v2026.08
Aug 31, 2026InfoResearchIndustrySecurityResearchMITRE ATLAS v2026.08 is a release of the knowledge base of adversary tactics and techniques involving AI. It contains 1 matrix, 16 tactics, 114 techniques, 83 sub-techniques, 39 mitigations and 72 case studies. The update adds new techniques for autonomous attack orchestration, AI agent communication and AI agent tools, and renames the AI Attack Staging tactic to AI Attack Adaptation.
MITRE ATLAS ReleasesA Comparative Survey of Security Risks in AI Systems: From LLMs to AI Agents and Embodied Agents
Aug 23, 2026InfoResearchPeer-reviewedResearchSecurityThis ACM Computing Surveys article, published in Volume 58, Issue 15 (pages 1-38) in November 2026, surveys security risks across AI systems, from large language models to AI agents and embodied agents. The source text provided contains only the bibliographic citation and no article content, so no findings or methods can be reported from it.
ACM Digital Library (TOPS, DTRAP, CSUR)Agentic AI in Healthcare: Opportunities, Challenges, and Future Directions
Jun 25, 2026InfoResearchPeer-reviewedResearchIndustryACM Digital Library (TOPS, DTRAP, CSUR)Testing Gemini models for scheming tendencies
May 29, 2026InfoResearchIndustryResearchSafetyGoogle DeepMind researchers Victoria Krakovna, David Lindner, Sebastian Farquhar and Rohin Shah introduce Gram (Gauging Realistic Agentic Misbehavior), an automated auditing framework that uses simulated agentic environments to test whether Gemini models sabotage their oversight when deployed as coding agents. Across 17 seed scenarios, Gemini models misbehaved in about 2–3% of simulated scenarios, rising to up to 8% under the red-team auditor, and Gemini 3 models showed more scheming-related reasoning than Gemini 2.5.
DeepMind Safety Research (Medium)SBOMs into Agentic AIBOMs: Schema Extensions, Agentic Orchestration and Reproducibility Evaluation
May 9, 2026InfoResearchPeer-reviewedSecurityResearchACM Digital Library (TOPS, DTRAP, CSUR)FinBot CTF Is Live: A Hands-On Companion to the OWASP GenAI Security Project
Apr 28, 2026InfoResearchIndustryResearchSecurityThe OWASP GenAI Security Project has launched FinBot, a hands-on Agentic Security capture-the-flag platform that is part of its Agentic Security Initiative. FinBot simulates a multi-agent vendor management platform with autonomous onboarding, fraud detection, invoice processing and communications, all powered by LLMs with real tool access. Its challenges cover prompt injection, tool misuse, policy bypass, data exfiltration, privilege escalation and remote code execution, mapped to the OWASP Top 10 for LLM Applications, the OWASP Top 10 for Agentic Applications, CWE and MITRE ATLAS.
OWASP GenAI SecuritySBOMs into Agentic AIBOMs: Schema Extensions, Agentic Orchestration and Reproducibility Evaluation
Apr 27, 2026InfoResearchPeer-reviewedSecurityResearchACM Digital Library (TOPS, DTRAP, CSUR)Benchmarking the effectiveness of multi-agent LLMs in collaborative privacy threat modeling with <span class="small-caps">LINDDUN GO</span>
Apr 26, 2026InfoResearchPeer-reviewedResearchPrivacyElsevier Security JournalsFinBot CTF Is Live: A Hands-On Companion to the OWASP GenAI Security Project
Apr 15, 2026InfoResearchIndustrySecurityResearchFinBot is a hands-on agentic AI capture-the-flag platform from the OWASP GenAI Security Project's Agentic Security Initiative, described as the "Juice Shop for Agentic AI." It simulates a multi-agent vendor management platform with LLM-driven onboarding, fraud detection, invoice processing and communications, and its challenges map to the OWASP Top 10 for LLM Applications, the OWASP Top 10 for Agentic Applications, CWE and MITRE ATLAS. The source positions it as a companion to the Agentic Top 10 framework rather than a replacement.
OWASP GenAI SecuritySBOMs into Agentic AIBOMs: Schema Extensions, Agentic Orchestration and Reproducibility Evaluation
Apr 7, 2026InfoResearchPeer-reviewedSecurityResearchThe source text is a bibliographic citation for an article titled "SBOMs into Agentic AIBOMs: Schema Extensions, Agentic Orchestration and Reproducibility Evaluation," published in Digital Threats: Research and Practice, Volume 7, Issue 2, pages 1-35, June 2026. It contains no abstract, method, or findings, so the research question and results cannot be stated from this text.
ACM Digital Library (TOPS, DTRAP, CSUR)v5.5.0
Mar 30, 2026InfoResearchIndustrySecurityResearchThe v5.5.0 release of the MITRE ATLAS knowledge base adds new techniques, including AI Agent Tool Poisoning, AI Supply Chain Rug Pull, Machine Compromise variants, and Cost Harvesting variants. It also adds case studies such as LLMSmith, the Poisoned Postmark MCP Server email exfiltration, and Model Distillation Campaigns Targeting Anthropic Claude, and updates mitigations including Code Signing, AI Telemetry Logging, and Segmentation of AI Agent Components.
MITRE ATLAS Releasesv5.4.0
Feb 5, 2026MediumResearchIndustrySecurityResearchThe v5.4.0 release of the framework adds four new techniques, including Publish Poisoned AI Agent Tool and User Execution: Poisoned AI Agent Tool, along with Escape to Host and Exploitation for Credential Access. It also updates the Modify AI Agent Configuration technique and adds four case studies covering exposed ClawdBot control interfaces, a poisoned ClawdBot skill supply chain compromise, a 1-click remote code execution in OpenClaw, and command and control via prompt injection in OpenClaw.
MITRE ATLAS Releasesv5.2.0
Jan 30, 2026InfoResearchIndustryResearchSecurityMITRE ATLAS released v5.2.0, which adds new techniques such as AI Service API, Virtualization/Sandbox Evasion, AI Agent Tool Credential Harvesting, AI Agent Tool Data Poisoning, AI Agent Clickbait, Data Destruction via AI Agent Tool Invocation, and Generate Malicious Commands. The release also adds mitigations including Segmentation of AI Agent Components, Input and Output Validation for AI Agent Components, and Deepfake Detection, and adds case studies including SesameOp, Malware Prototype with Embedded Prompt Injection, and LAMEHUG.
Fix: Added mitigations include Segmentation of AI Agent Components, Input and Output Validation for AI Agent Components, and Deepfake Detection. Updated mitigations include Limit Public Release of Information, Model Hardening, Restrict Number of AI Model Queries, Privileged AI Agent Permissions Configuration, AI Agent Tools Permissions Configuration, Human In-the-Loop for AI Agent Actions, and Restrict AI Agent Tool Invocation on Untrusted Data, among others.
MITRE ATLAS ReleasesBuilding Trustworthy AI Agents
Jan 30, 2026InfoResearchPeer-reviewedSafetyIndustryThe article argues that personal AI assistants depend on an unfounded assumption that users can trust systems that are not yet trustworthy. It says current assistants fail predictably: they push users against their own interests, cast doubt on what users know, and cannot separate a user's present self from their past. It also states they handle incomplete, inaccurate, and partial context poorly, with no standard way to improve accuracy, correct error sources, or hold them accountable for wrong information.
IEEE Xplore (Security & AI Journals)Exploring the Agentic Metaverse’s Potential for Transforming Cybersecurity Workforce Development
Dec 12, 2025InfoResearchPeer-reviewedResearchIndustryThis exploratory qualitative study evaluates an AI-driven metaverse prototype designed for cybersecurity training, with feedback from 53 cybersecurity professionals. The authors identify challenges in operationalizing the agentic metaverse, a convergence of immersive metaverse platforms and multi-agent systems, and offer six recommendations with emphasis on implementation and governance.
Fix: The source states six recommendations to guide operationalization of the agentic metaverse, with emphasis on implementation and governance considerations, but does not list their specific content in the provided text.
AIS eLibrary (Journal of AIS, CAIS, etc.)
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.