The Harm SurfaceAI, cyber, and autonomy

Research · a benchmark in development

Nearly every AI-security benchmark measures the attack. This one measures the defense.

Can a model recognize an intrusion, reconstruct what happened from raw telemetry, and investigate it without inventing anything? Blue measures that half of the work, and nothing else.

Self-hostable open-weight models are measured beside frontier ones, because the organizations most exposed are usually the ones choosing between the two.

HarmSurface
Blue Benchmarking defensive capability

The capability spine 1 of 3 published · 1 in validation · 1 planned
  1. Dimension 1 Recognition Does it see that something is wrong, and can it say what? measured · 5 runs · 12 August 2026 Read the results →
  2. Dimension 2 Reconstruction Can it rebuild the incident from the telemetry it is given? In validation Scenarios built and validated. Not yet run across the roster, so there is nothing to publish.
  3. Dimension 3 Investigation Does it know what to look at next? Planned Nothing is claimed here yet. Listed so the shape of the program is visible, not to imply a result.

What D1 found

9 models · 96 probes

Nine models, one frozen set of 96 events, graded by arithmetic over the produced bytes. Four numbers rather than one, because the top five recognition scores overlap inside their own intervals and no single figure separates a model that knows from a model that flags.

A high recognition score can belong to the model you would least want triaging your alerts.

Blue-readiness · top 3

  1. claude-opus-5 0.898
  2. gemini-3.1-pro 0.870
  3. kimi-k3 0.836 1 pass

A composite built to punish a broken axis, not to reward a loud one.

Run data is published under CC BY 4.0: reuse and redistribution are allowed with attribution.

How it scores

Arithmetic over produced bytes

Every score is computed from what a model actually wrote. No model grades another model's work, and no vendor's self-report is used as evidence.

Repeats, and their absence

Every row states its own pass count, and a row measured once carries no interval and says so.

Ties, not false ordinals

Where intervals overlap the page says so instead of printing a rank the data cannot support.

Nobody pays to appear

No vendor sees a result before publication, and no vendor can buy inclusion, exclusion, or a re-run.

Scenarios are driven against real software inside a sealed environment and the telemetry they emit is captured there. The scenarios are ours and are not published, which is what keeps the corpus out of a training set. Blue measures defense: it does not measure, publish, or develop offensive capability.

What this is not, yet

Blue is in development. One dimension is published, one is in validation, and one is planned. The probe set is small, the model list is partial, and two rows in D1 rest on fewer passes than the run. None of that is hidden behind a headline number, and nothing on these pages claims further than the run behind it.

Corrections are annotated on the original, never silently fixed.