Using AI Sec Watch in research
What a record is, where each value comes from, how reliable the labels are, and how to cite the exact version an analysis used.
What a record is
A record is one item from one public source: a security advisory, a research paper, a policy document, an incident report or a news item. Advisories for AI and machine learning software are the core of the data, and the other records give the research and reporting of the same period. A record is stored once per source address, and one record holds each CVE identifier. Each record has 48 fields in the current export, and its id stays the same from one release to the next.
Where each value comes from
An analysis should know which values the source stated and which a model assigned. The documentation of each release gives the origin of every field, in five groups.
| Origin | Meaning | Examples |
|---|---|---|
| Source | Taken from the source's own feed or interface. | title, source_url, published_at, cve_id, cwe_ids, cvss_score, affected_packages |
| Classifier | Assigned by a language model under a versioned prompt. | labels, issue_type, attack_type, ai_component_targeted, llm_specific |
| Summarizer | Written by a language model from the source text. | summary, solution |
| Enrichment | Joined from a public catalog. | epss_score, epss_checked_at, exploit_maturity, kev_date_added, capec_ids, atlas_ids |
| Derived | Computed when the file is written. | severity_source, issue_type_source, source_category, raw_content_length |
Two fields can come from either side, and the data say which. severity_source tells whether the severity is the source's own rating or the model's, and issue_type_source tells whether the record type was set by a rule for the source or chosen by the model. Each record also names the model and the prompt version behind its labels (classifier_model, classifier_prompt_version) and the date on which its exploitation score was read (epss_checked_at).
A release or the live data
The site changes several times a day. A release is a frozen copy with its own DOI, so an analysis built on a release can be repeated. Use a release for anything that will be published, and the API or the live export for monitoring.
| Release | Published | Records | DOI |
|---|---|---|---|
| 4.0 | 2026-10-10 | 8,260 | 10.5281/zenodo.23282466 |
| 3.0 | 2026-04-11 | 3,020 | 10.5281/zenodo.19519470 |
How reliable the labels are
The classifier fields are checked on a stratified sample of records. Each sampled record is labelled a second time by the same model and once by a second model, and the agreement is reported for each field with an interval. The table, the sample and what the figures do and do not show are on the methods page. Records that were corrected after publication are listed, with the reason, in the corrections log.
Limits to state in a paper
- Labels and summaries are produced automatically and can be wrong. The linked source is the authority.
- Coverage follows the sources that are read. A missing record is not evidence that a problem does not exist.
- EPSS scores and the status in the Known Exploited Vulnerabilities catalog are readings taken on dates that the data record. Both change over time.
- Some advisory fields are filled for part of the advisories only. The documentation of a release gives the counts.
- The text of the sources is not redistributed. A record holds a summary written for the dataset and a link.
Questions the data can support
- How well does a predicted exploitation score agree with exploitation that was later observed?
- Do research papers study the attack types that advisories report?
- How far can a language model be trusted to label security text, and on which fields does it fail?
Loading a release
The CSV file of a release can be read straight from the repository. In the CSV, the values of a list are joined by a semicolon and a space, and a text cell can hold a line break.
import pandas as pd # Release 4.0, https://doi.org/10.5281/zenodo.23282466 url = "https://zenodo.org/records/23282466/files/aisecwatch-dataset-2026-10-10.csv?download=1" df = pd.read_csv(url) advisories = df[df["issue_type"] == "vulnerability"]
# Release 4.0, https://doi.org/10.5281/zenodo.23282466 url <- "https://zenodo.org/records/23282466/files/aisecwatch-dataset-2026-10-10.csv?download=1" df <- read.csv(url, stringsAsFactors = FALSE) advisories <- subset(df, issue_type == "vulnerability")
Citing the dataset
Cite the release that was used, by its DOI in the table above. The DOI 10.5281/zenodo.19519469 always resolves to the latest release.
@misc{luu2026aisecwatch_dataset,
author = {Luu, Truong (Jack)},
title = {{AI Sec Watch Dataset}: {AI/LLM} Security Threat Intelligence},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.19519469},
url = {https://doi.org/10.5281/zenodo.19519469},
note = {Data set. 48 fields per record, in CSV, JSON and JSONL. Accessed: 2026-10-10}
}Papers and theses that use the data are listed on this page when their authors send a reference, through the contact details on the maintainer's site. Exercises for courses are on the teaching page.