Teaching with AI Sec Watch
Four lab exercises on the security of AI software. Each one says what students should be able to do afterwards, what they hand in and how long it takes. All of them run on public pages and files, with no login.
Where the labs fit
The labs suit courses in information security, systems analysis and design, software engineering and applied AI, in the later undergraduate years or at graduate level. The first two need no programming. The third needs basic Python or R, and the fourth needs a spreadsheet. Each lab stands alone, and the four in order make a short unit of about four hours.
Lab 1. Read one advisory
40 minutes. No programming.
Students should be able to
- Tell apart how severe a flaw is, how likely it is to be exploited and whether it has been exploited.
- Find the affected versions of a package and the first fixed version.
- Check a summary against its source.
Steps
- Pick one advisory from the known exploited list and a second one of the same severity from the advisory list that is not on it.
- For each, write down the affected package and versions, the fixed version, the CVSS severity, the EPSS score, and whether CISA lists the flaw as exploited.
- Open the linked source of each advisory and check two of those values against it.
Students hand in
A half-page note that says which of the two a team should patch first and why, with the three signals as evidence.
Questions for discussion
- Why can a critical CVSS score sit beside a low EPSS score?
- A page says a flaw is not listed as exploited. What would have to be true for that statement to be wrong?
Lab 2. Check a software stack
50 minutes. No programming.
Students should be able to
- Read a dependency manifest and say which entries call a language model.
- Explain the difference between a dependency a project chose and one it inherited.
- Turn a list of findings into an order of work.
Steps
- Paste the sample manifest below, or a manifest of your own, into Stack Check. Nothing that is pasted is stored.
- Read the report in its own order: advisories for the pinned versions with their fixed versions, the components that call a language model, the dependencies that reach one indirectly, and the authority those dependencies hold on the host.
- Group the advisories by package and find, for each package, the lowest version that clears all of them.
Sample manifest (requirements.txt)
langchain==0.0.300 llama-index==0.9.0 gradio==3.50.0 transformers==4.30.0 mlflow==2.8.0
The sample pins releases from 2023 on purpose, so the report is long. Grouping it by package is part of the exercise.
Students hand in
A one-page memo to a product owner: what to upgrade first, to which version, and what is left after the upgrades.
Questions for discussion
- Why does a tool that lists only direct dependencies miss part of the exposure?
- Which finding would you accept as a risk for now, and what would you need to know before you did?
Lab 3. Analyse the dataset
75 minutes. Basic Python or R.
Students should be able to
- Load a versioned public dataset and state which release an analysis used.
- Compute descriptive statistics on advisories and read them with care.
- Name one limit of labels that a model assigned.
Steps
- Load the CSV file of a release (the snippet below loads release 4.0) and keep the advisory records, where
issue_typeisvulnerability. - Count the advisories by
attack_type. A record can hold more than one type, joined by a semicolon and a space. - Compare the median
epss_scoreacross the levels ofcvss_severity. - Among advisories with a CVE identifier, find the share that CISA lists as exploited (
kev_date_addedis filled) and compare their EPSS scores with the rest. - Look up, in the label reliability table, how consistent the attack type labels are, and say what that means for the count in the second step.
Starting point (Python)
import pandas as pd
# Release 4.0, https://doi.org/10.5281/zenodo.23282466
url = "https://zenodo.org/records/23282466/files/aisecwatch-dataset-2026-10-10.csv?download=1"
df = pd.read_csv(url)
advisories = df[df["issue_type"] == "vulnerability"]
print(len(df), "records,", len(advisories), "advisories")
# A record can hold more than one attack type, joined by "; "
print(advisories["attack_type"].dropna().str.split("; ").explode().value_counts())
print(advisories.groupby("cvss_severity")["epss_score"].median())The fields are described on the dataset page and in the documentation that ships with each release.
Students hand in
A notebook with the four results, the DOI of the release in its first cell, and one paragraph on the limits of the answers.
Questions for discussion
- The severity of a record can come from the source or from the model, and the field severity_source says which. Why should an analysis keep the two apart?
- The live site changes every day. What would go wrong if two students downloaded the data on different days?
Lab 4. Test the labels
60 minutes. Pairs of students, a spreadsheet.
Students should be able to
- Apply a written rule to real cases and notice where it runs out.
- Compute the agreement between two raters, as a share and as Cohen's kappa.
- Compare human ratings with labels that a model assigned.
Steps
- The instructor draws 20 rows from a release and gives each pair only the
titleandsource_urlcolumns, so nobody sees the labels in the dataset. - Each student reads the sources and answers two questions per record, alone: is the record in scope under the scope rule, and what type of record is it?
- The pair compares answers, computes the share of identical answers and Cohen's kappa for each question, and settles the records on which they differ.
- The pair then compares its settled answers with the labels in the dataset.
Students hand in
The two rating sheets, the agreement figures, and a paragraph on the two records that were hardest to rate.
Questions for discussion
- Where did you differ from each other, and where from the dataset? Are those the same records?
- What would you change in the rule to remove one of your disagreements, and what new problem would the change create?
Notes for instructors
- The site changes every day, so counts on its pages move. For numbers that every student can reproduce, use a dataset release, which is frozen and has its own DOI.
- A language model writes the summaries and assigns the labels. The labs use this: students check a summary against its source in the first lab and measure the labels in the fourth. The methods page says how each field is made.
- A written note or memo can be graded on three things: the evidence it cites, the reasoning from the evidence to the recommendation, and whether the reader could act on it.
- The exercises on this page are free to use and adapt under the CC BY 4.0 licence. A note on how an exercise worked in a course is welcome, through the contact details on the maintainer's site.