Methods
How each part of the site is produced, which parts are automatic, and what each one cannot tell you. The linked source of every record is authoritative; this site adds structure, links and context.
Collection
A scheduled job fetches 84 sources every six hours: vulnerability databases (NVD, GitHub Advisory Database, OSV, CISA Known Exploited Vulnerabilities and others), research repositories and journals, standards and government publications, and security and technology news. The vulnerability databases are fetched again every hour. Feeds are filtered to AI and machine-learning terms. Records are deduplicated by URL and by CVE identifier, so one CVE has one record however many sources report it. The full list and each source's state are on Sources and Status.
Classification
Each new record is labelled by a language model (claude-haiku-5-5) with a versioned prompt (v4). The model assigns relevance to AI security, record type, labels, attack types, the AI component targeted, impact, attack sophistication and a self-assessed confidence between 0 and 1. The model and prompt version are stored on every record, and the prompt is published with the dataset.
Scope is part of the prompt. An advisory is kept only when the affected product is AI or machine-learning software (model SDKs, agent frameworks, MCP servers, inference servers, ML libraries, vector databases, AI coding tools) or the flaw sits in an AI feature of a general product. A flaw in general software that merely runs beside AI workloads, such as an operating system, a container platform or a web framework, is left out. Research is kept when it studies the security, safety, privacy or governance of AI systems, and left out when machine learning is only the method for another task. News, incidents and policy about AI companies, products and regulation are all kept. When the prompt or the model changes, the whole archive is labelled again, so every record carries labels from one prompt version on one model.
Summaries and mitigations are written separately (claude-haiku-5-5, prompt v2) for developers and security engineers. They keep product names, versions, identifiers and file paths as the source writes them, state only what the source states, and report a mitigation only when the source gives one. A record whose source text has no article content carries no summary. Summaries and labels can be wrong; they are aids to reading, not findings.
Classifier audit
After every change of prompt or model, a fixed sample of records is labelled again, once by the same model to measure how stable its labels are and once by a second, larger model that acts as an independent rater. The sample takes up to 50 records of each record type and 50 records the classifier judged not relevant. Fields other than relevance are compared only where both runs judged the record relevant, and severity and record type are compared with the stored record only where the model set them rather than CVSS or a source rule. Kappa corrects agreement for chance (1 is perfect agreement and 0 is what chance would give), with quadratic weights for the ordinal fields severity and attack sophistication. Agreement between two models is not accuracy against a human expert, so read these figures as a measure of how settled each label is.
The latest audit (2026-10-09) covers prompt v3 on claude-haiku-4-5-20251001, with claude-sonnet-4-5-20250929 as the second rater. Both runs labelled 286 of the 287 sampled records, drawn with the seed audit-v3-2026-10-09.
| Field | Same model, second run | Second model | ||||
|---|---|---|---|---|---|---|
| Records | Agree | Kappa | Records | Agree | Kappa | |
| Relevant to AI security | 286 | 97.2% | 0.900.83 to 0.96 | 286 | 92.0% | 0.720.60 to 0.81 |
| Record type | 81 | 86.4% | 0.790.68 to 0.89 | 225 | 79.6% | 0.740.67 to 0.80 |
| Severity | 192 | 97.4% | 0.960.90 to 1.00 | 225 | 87.6% | 0.860.78 to 0.92 |
| Attack sophistication | 233 | 87.6% | 0.820.75 to 0.88 | 225 | 76.0% | 0.550.43 to 0.65 |
| AI component targeted | 233 | 90.1% | 0.880.83 to 0.93 | 225 | 78.2% | 0.740.67 to 0.80 |
| LLM-specific | 233 | 95.7% | 0.910.86 to 0.96 | 225 | 89.8% | 0.800.72 to 0.87 |
| Attack types (Jaccard) | 233 | 87.6% | 0.900.86 to 0.93 | 225 | 73.8% | 0.770.71 to 0.82 |
Sample by stratum: Vulnerabilities 50 of 3,021, Incidents 37 of 37, Research 50 of 801, News 50 of 4,373, Policy and regulation 50 of 128, Judged not relevant 50 of 1,900. Agreement is the share of identical labels. For attack types it is the share of identical sets, and the last column gives the mean Jaccard similarity instead of kappa. The line under each statistic is its 95% percentile bootstrap interval. Download the sample with all three sets of labels (CSV).
Severity
When the source gives a CVSS severity, the record uses it. Otherwise the severity is the classifier's estimate, and the record page says so. Dataset exports carry a severity_source field with the value cvss or the model's.
Severity measures technical exploitability, so high and critical are reserved for vulnerabilities and incidents. A news report, research publication or policy document is capped at medium, whatever the model assigned.
Enrichment
- EPSS: the FIRST Exploit Prediction Scoring System's probability that a CVE is exploited in the next 30 days, fetched for every record with a CVE. A CVE that FIRST has not scored shows as not scored.
- Known exploitation: whether the CVE is in the CISA Known Exploited Vulnerabilities catalog, and whether the catalog records ransomware use.
- CVSS vector: parsed into attack vector, complexity, privileges and user interaction (CVSS v3.x).
- CAPEC: attack patterns mapped from the record's CWE identifiers with a fixed table.
- MITRE ATLAS: techniques mapped from the classifier's attack types with a fixed table.
Links between records
News and research items that mention a CVE identifier in their title, summary or text are linked to the vulnerability record for that CVE. The vulnerability page lists them, and its coverage count is the number of such items.
Topics
Topics sit beside the classifier's fixed labels so that the vocabulary can follow the field without changing past classifications. The catalog has 16 topics (version fd2418e0bd1e), each defined by public matching rules. A record belongs to a topic when its title matches one of the topic's rules or its summary matches them at least twice. Whenever the catalog changes, the whole archive is tagged again, so a new topic's trend starts from the first relevant record. The rules are in the topics API.
Exposure Registry
The registry tracks open-source packages on PyPI, npm, crates.io and the Go module proxy. For each package it reads the dependencies the package itself declares in its registry metadata, release by release: PyPI requirements, npm dependency maps, crates.io dependency lists and the direct requirements in each Go module's go.mod.
- A package is LLM-integrated when a release declares a dependency in the published catalog: 17 model SDKs, 25 agent and orchestration frameworks, 2 Model Context Protocol implementations and 8 retrieval stores. Runtime dependencies count, and so do optional extras and Cargo features, which declare an integration a user can switch on; development and build dependencies do not. SDK packages themselves count from their first release.
- The integration date is the first such release. It is found by a search over the release history that assumes a dependency, once added, is kept; a package that added and later removed an SDK may be dated too late or not at all.
- The authority profile lists what the latest release's dependencies let code do: run code, run shell commands, use the file system, drive a browser, make outbound HTTP requests, or expose MCP tools. It is drawn from a catalog of 61 dependency names. It shows reachable capability, not that the package misuses it.
- A package reaches an LLM component through dependencies when a chain of latest-release runtime dependencies, at most five steps long, ends in an integrated package. The registry covers the packages it tracks (seed lists and their direct dependencies, plus packages people look up), so chains through untracked packages are missed.
- Advisories are linked to registry packages by name, because the stored advisory records do not state an ecosystem. Packages that share a name across ecosystems therefore share advisories on their pages. Stack Check uses OSV and is exact.
Stack Check
Stack Check reads the manifests, lockfiles and SBOMs you send, in memory, and does not store them. From a CycloneDX or SPDX SBOM it takes every component with a PyPI, npm, Go or crates.io package URL. For every dependency with an exact version it asks OSV (osv.dev) for known vulnerabilities, which means the package names and versions are sent to that service; you can turn this off. It marks LLM components from the registry catalog, looks up indirect exposure in the registry, derives the authority the dependencies grant, and lists controls whose triggers are met, each with the fact that triggered it. Without a lockfile or SBOM only exact pins are checked. The report downloads as JSON or as a CycloneDX 1.6 AI-BOM, with the fields CycloneDX has no slot for (component category, indirect exposure path, authority, EPSS) as properties named aisecwatch:.
Corrections
Every record page has a form to report a wrong label, severity or relevance. Reports are reviewed, and the record is corrected at the source of the error (data, rule or prompt) rather than by hand where possible. Every change to a record is listed with its reason in the corrections log and on the record page; reporters' comments are not published. Dataset releases are versioned, so a correction never changes a release that has been cited.
So far: 0 reports received, 0 led to a correction, 0 declined and 0 open.