The Harm SurfaceAI, cyber, and autonomy

HarmSurface Blue · Capability spine · Dimension D1

Recognition: does it see that something is wrong, and can it say what?

Nine models, one frozen set of 96 events, graded by arithmetic over the produced bytes. Four numbers rather than one, because the top five recognition scores overlap inside their own intervals and no single figure separates a model that knows from a model that flags.

The capability spine 1 of 3 published · 1 in validation · 1 planned
  1. Dimension 1 Recognition you are here Does it see that something is wrong, and can it say what? measured · 5 runs · 12 August 2026 · this page
  2. Dimension 2 Reconstruction Can it rebuild the incident from the telemetry it is given? In validation Scenarios built and validated. Not yet run across the roster, so there is nothing to publish.
  3. Dimension 3 Investigation Does it know what to look at next? Planned Nothing is claimed here yet. Listed so the shape of the program is visible, not to imply a result.

Blue readiness · ordered best first

A single number is what crowns the wrong model, so this one is built to punish what a plain average rewards. A weak axis drags the whole score down, which is the point: a model that spots malice but cannot name it is not ready for blue work, whatever its accuracy. The archetype travels with the number, and the four axes are one scroll down.

  1. 1 claude-opus-5The analysts 0.898 avg 0.898 2 up on recognition
  2. 2 gemini-3.1-proThe analysts 0.870 avg 0.870 2 up on recognition
  3. 3 kimi-k3The mirror 0.836 avg 0.837 3 up on recognition
  4. 4 claude-opus-4.8The near analyst 0.747 avg 0.759 2 down on recognition
  5. 5 glm-5.2The mirror 0.699 avg 0.710 2 up on recognition
  6. 6 gpt-5.6-solThe flood 0.386 avg 0.526 5 down on recognition
  7. 7 llama-primusThe flood 0.320 avg 0.487 2 down on recognition
  8. 8 zysec-7bThe floor 0.158 avg 0.301 1 up on recognition
  9. 9 foundation-sec-8bThe floor 0.120 avg 0.161 1 down on recognition

blue-readiness = (recognition · auc′ · calibration · technique)1/4, where auc′ = max(0, (AUC − 0.5) / 0.5) and calibration = 1 − overconfidence. Each axis is floored at 0.05 before the product, so one failed axis cannot erase every other signal. Equal weighting is a value judgment, disclosed rather than tuned. Arrows show each model's move against a ranking by recognition alone, and "avg" is the plain arithmetic mean of the same four axes, printed so the two composites can be compared directly. A plain average lets one strong column carry three weak ones: it puts the two flooding models mid-table, which is the failure this dimension exists to show.

A high recognition score can belong to the model you would least want triaging your alerts.

  • corpus vm-atr
  • probes 96
  • strata 16 malicious · 40 ambiguous · 40 benign
  • grading wire-truth, no model judge
  • chance 0.500
  • runs 5

Supersedes. The five-model run of 11 August 2026, whose figures this run replaces in full.

5 runs, and 1 on some rows. Our own rule is that a score counts as a claim at five or more, because a probe set this size can move several points on chance alone. Read the ordering, not the third decimal.

Figures are transcribed from the run record and are not recomputed at build time. GLM 5.2 and Kimi K3 were measured once, carry no interval, and their exact values are provisional.

Same data, opposite verdict · recognition alone → blue readiness

RANKED BY RECOGNITION ALONE RANKED BY BLUE READINESS 1 · gpt-5.6-sol · 0.906 2 · claude-opus-4.8 · 0.901 3 · claude-opus-5 · 0.890 4 · gemini-3.1-pro · 0.871 5 · llama-primus · 0.859 6 · kimi-k3 · 0.800 7 · glm-5.2 · 0.738 8 · foundation-sec-8b · 0.469 9 · zysec-7b · 0.062 1 · claude-opus-5 · 0.898 2 · gemini-3.1-pro · 0.870 3 · kimi-k3 · 0.836 4 · claude-opus-4.8 · 0.747 5 · glm-5.2 · 0.699 6 · gpt-5.6-sol · 0.386 7 · llama-primus · 0.320 8 · zysec-7b · 0.158 9 · foundation-sec-8b · 0.120

Wider than this screen. Scroll the chart sideways to read every name.

One line per model. Only the two largest moves are colored. gpt-5.6-sol has the highest recognition score in the run and the sixth profile; kimi-k3 recognizes fewer events and is honest about the rest.

What is measured

Recognition

Did it call this event right?

The model's own verdict at its own threshold. Chance is 0.500, and a flag-everything or flag-nothing strategy scores exactly that, so neither buys anything here.

Ranking AUC

Can it put malicious above benign at all?

Threshold-free knowledge, by Mann-Whitney rank sum over the probabilities. It separates a model with knowledge and a broken threshold from a model with no knowledge for any threshold to recover. A void scores chance, never zero.

Overconfidence

Does it assert certainty it has not earned?

Measured on a separate stratum of genuinely ambiguous events: how often the model returns a confident verdict where none is warranted. Lower is better. This is claimed competence without capability.

Technique identification

Can it name the ATT&CK technique?

Credited only on checkable malicious events. It asks whether the model recognizes malice with the taxonomy behind it, or only has a bad feeling about a line of telemetry.

The roster, positioned

chance 0.50 0.00 0.25 0.50 0.75 1.00 0.40 0.50 0.60 0.70 0.80 0.90 1.00 RECOGNITION → TECHNIQUE NAMED → ANALYST · sees it, names it FLOOD · sees it, cannot name it zysec-7b · 0.062, off scale gpt-5.6-sol claude-opus-4.8 claude-opus-5 gemini-3.1-pro llama-primus kimi-k3 glm-5.2 foundation-sec-8b

Wider than this screen. Scroll the chart sideways to read every name.

  • The analysts
  • The near analyst
  • The mirror
  • The flood
  • The floor
  • single pass, no interval
  • 95% interval on recognition
Right is recognizing the event. Up is naming the ATT&CK technique behind it, and color is the archetype, so a dot lifted out of this picture still means something. Recognition is clipped to 0.40 and above, so the models this chart exists to separate are not squeezed into one corner of it. zysec-7b scores below that and is marked at the edge with its figure rather than dropped. The horizontal bar through a dot is its 95% interval on recognition; where two bars overlap, that column does not separate those two models. kimi-k3 and glm-5.2 ran a single pass, so they are drawn hollow and carry no bar rather than a zero-width one. The dashed vertical is chance on recognition, which a flag-everything and a flag-nothing strategy both score. Bottom right is a working alarm attached to nothing that can say what it heard.

The full profile

Rank bands, not ordinals: rows 1 to 5 are tied inside their own intervals. Every model's repeat count is on its own row, and the run is 5 passes except where a row says otherwise.

ModelRankRecognitionRanking AUC TechniqueOverconfidenceVoid
The analysts
claude-opus-5 frontier 1 to 5 tied 0.890 ±0.019 0.947 0.887 0.080 0
gemini-3.1-pro frontier 1 to 5 tied 0.871 ±0.028 0.940 0.875 0.145 0
The near analyst
claude-opus-4.8 frontier 1 to 5 tied 0.901 ±0.008 0.902 0.775 0.445 0
The flood
gpt-5.6-sol frontier 1 to 5 tied 0.906 ±0.078 0.920 0.125 0.767 0
llama-primus local 8B 1 to 5 tied 0.859 ±0.017 0.889 0.312 1.000 0
The mirror
kimi-k3 frontier1 run 6 0.800 0.956 0.812 0.175 1
glm-5.2 frontier1 run 7 0.738 0.757 0.812 0.225 10
The floor
foundation-sec-8b local 8B 8 0.469 ±0.000 0.487 0.000 0.825 12
zysec-7b local 7B 9 0.062 ±0.000 0.590 0.062 0.100 61

A rank band means the recognition intervals overlap, so that column does not separate those models; a model measured once cannot be shown to tie with anything and keeps its ordinal. The archetype names are ours, coined for the patterns in this run rather than taken from a standard. Read against chance. A flag-everything and a flag-nothing strategy both floor at chance, so neither beats the bar. Overconfidence is the share of ambiguous probes called malicious with no verdict warranted; lower is better.

How to use this

  • Do not choose a model on a recognition score. Five of the nine sit inside one interval on that column, so it cannot separate them. Ask for the four numbers.
  • If your telemetry cannot leave your network, llama-primus is inside the top-five tie on recognition. Read the overconfidence column before you deploy it: it called every ambiguous probe malicious, so it will add alerts rather than reduce them.
  • Ask a vendor what their model does with an ambiguous event, not what it does with an obvious one. Both are easy. Only one is most of a working day.
  • Ask whether the model names the technique or only flags the event. The distance between the two is the distance between an alarm and a place to start.
  • Treat a single run as a single run. Two models here were measured once and their figures are marked provisional for that reason alone.

What each model did

The group names below are ours, coined for the patterns this run produced. They are not a standard taxonomy and no model vendor uses them. A count is not written here because groups are read from the piece and can change with a run.

The analysts

Strong on every axis and calibrated with it. They recognize the event, name the technique, and hold their verdict on probes that do not warrant one. The top raw recognition score is not in this group.

claude-opus-5 frontier

Third on recognition and first on everything the recognition column cannot see. Ranking is 0.947, technique identification is the highest in the run, and overconfidence is the lowest of any model that answers. It abstains where abstaining is the right answer. This is the profile the rung was built to find, and it is not the top raw score.

gemini-3.1-pro frontier

The same shape as Opus 5, one step down on every axis and consistent across all of them. Strong ranking, strong technique identification, and it holds its verdict on the ambiguous stratum.

The near analyst

Recognition and technique identification at analyst level, and calibration that is not. It asserts a verdict on nearly half the ambiguous stratum, which is the single axis separating it from the group above and the reason it is not in it.

claude-opus-4.8 frontier

Second by five thousandths, on the tightest interval in the run. It maps events to ATT&CK well. What separates it from its successor is the ambiguous stratum, where it asserts a verdict on nearly half the probes that do not warrant one.

The flood

They post the recognition scores and cannot say what they flagged. Overconfidence is at or near the ceiling and technique identification is at or near the floor. A one-column leaderboard puts them first.

gpt-5.6-sol frontier

The highest recognition on the board, on an interval wide enough to cover the four models beneath it, so the lead is a point estimate rather than a separation. Inside that tie it flags hard: three ambiguous probes in four are called malicious, and it names the ATT&CK technique on one malicious probe in eight, the lowest of any model that recognizes at all. It sees something and cannot say what.

llama-primus local 8B

An 8B you can host on one workstation, inside the top-five tie on recognition, which is the deployment finding of this run. It also calls every ambiguous probe malicious without exception, the worst calibration measured here, and names the technique on under a third of the malicious ones. Real knowledge, no threshold, and a shallow map.

The mirror

The opposite trade. They recognize fewer events and bluff less about the ones they miss, with technique identification level with the top of the board. Both were measured once.

kimi-k3 frontier

The best ranking in the run on a single pass. It orders malicious above benign better than anything else measured while recognizing fewer of them at its own threshold, which is a threshold to repair rather than knowledge to acquire.

glm-5.2 frontier

Recognition and ranking sit together at about three quarters, and technique identification is level with the top of the board. It catches fewer events and bluffs less on the ones it does not catch. Ten voids on a single pass, so read the ordering rather than the figure.

The floor

One is not distinguishable from chance and names no technique at all. The other returns nothing usable on most probes and is reported as silent rather than scored as wrong.

foundation-sec-8b local 8B

Recognition below the chance line, and ranking 0.013 under it, which on 96 probes is indistinguishable from chance rather than evidence that the confidence runs backwards. It names no technique anywhere in the malicious stratum, and it asserts a confident verdict on four ambiguous probes in five. A security-tuned model that is not performing the security task.

zysec-7b local 7B

Returns nothing usable on 61 of 96 probes rather than guessing. Recognition collapses because the answers are absent, not because they are wrong, and overconfidence stays honest for the same reason. Ranking on what it did answer is barely above chance. Not enough signal to rank, which the rung records rather than converting into a number.

Why four numbers

A single detection score collapses different failures into one figure. A model that flags every line has no threshold. A model that cannot tell a malicious event from an ordinary one has no knowledge. A model that flags the right events but cannot name what it flagged has neither a threshold problem nor a knowledge problem, and its score will not tell you so. All three arrive as one number, and the repair for each is nothing like the repair for the others.

The four columns are defined above the chart. What they are worth is in the distances between them. High ranking with low recognition means the knowledge is present and the threshold is broken. Ranking at chance means there is nothing for a better threshold to recover. High recognition with low technique identification means the alarm works and the analysis behind it does not. High recognition with high overconfidence means the alarm is loud rather than right, and the two are indistinguishable in the first column.

What this run changes

No single number ranks the top of the board. The five highest recognition scores fall within 0.047 of each other, and the leading model's own interval is 0.078 wide. It covers all four beneath it. A leaderboard with one column cannot separate these models, and this run says so on its own leading figure rather than about somebody else's.

A self-hostable 8B is inside that tie. llama-primus recognizes at 0.859 against a frontier top of 0.906, on identical bytes. For an organization that cannot send telemetry off its own network, that is the finding. The rest of its profile is why the finding needs the other three columns.

The highest recognition score belongs to a model you would not want triaging alerts. GPT-5.6 Sol leads the column and posts the lowest technique identification of any model that recognizes at all, at 0.125, while calling three ambiguous probes in four malicious. High recognition here is a loud alarm attached to nothing that can say what it heard.

Calibration and technique identification are where the board separates. Opus 5 and Gemini 3.1 Pro hold their verdict on ambiguous events and name the technique near 0.88. llama-primus and GPT-5.6 Sol do neither. The two groups are indistinguishable on recognition alone.

Nothing was invented. A void scores chance rather than zero. A model measured once is marked and its interval is left blank rather than printed as certainty. A model that mostly stayed silent is reported as silent rather than assigned a number that would look like a measurement.

Method. A seeded, stratified probe set drawn from the frozen vm-atr corpus. Each model is asked, per event, for a verdict, a malicious-probability, and an ATT&CK technique, answered from immediate recognition with no deliberation and no shared context. Recognition grades the on-face verdicts. Ranking AUC is a threshold-free Mann-Whitney over the probabilities, where ties share mean rank and a void scores chance. Overconfidence is measured on a separate ambiguous stratum. Technique identification is credited only on checkable malicious events. Every figure is arithmetic over the produced bytes; no model sits in the grading loop. The frozen bytes are identical for all nine. Local models run on the box. Claude is reached through its own command line, Gemini through agy, and GPT-5.6 Sol through codex, each a clean single shot. GLM and Kimi are reached through a raw inference endpoint.

Runs. 5

License. The run data and the figures on this page are published under CC BY 4.0: reuse and redistribution are allowed with attribution.

What these numbers do not say.

  • The intervals are sampling variance over five repeats on one frozen probe set. An interval over many probe sets is a different and wider estimate, and this run does not have one.
  • GLM 5.2 and Kimi K3 were measured once. Their figures are point estimates and are marked provisional for that reason alone.
  • Models reached through a command line run at that tool's default sampling, which this run did not fix or override.
  • Speed is not compared. These models are served in different ways, so a latency figure would measure the serving and not the model, and it is left out rather than reported with a caveat nobody reads.
  • Blue-readiness orders; it does not diagnose. The composite says which model is readier, and only the profile below says why: a low score never explains on its own that a model floods.
  • Recognition and ranking AUC partly double-count detection. A three-axis variant is defensible, and the formula is published either way so the choice can be argued with rather than taken on trust.

Nobody pays to appear here, to be re-run, or to be removed. Corrections go to the address on About.