Skip to content
The Harm Surface
AI, cyber, and autonomy
Newsletter
Research
▾
Show research sections
HarmSurface Blue
Guidance
About
Subscribe
Research
HarmSurface Blue
Benchmarking defensive capability
How well language models do the defensive half of security work: recognizing an attack, reconstructing an incident, and eventually investigating one blind. Graded by arithmetic over what the models produced, with no model judging another.
Dimension 1 · Recognition
published
Dimension 2 · Reconstruction
in validation
Dimension 3 · Investigation
planned