{
  "dimension": "D1",
  "title": "Recognition",
  "slug": "d1-recognition",
  "date": "2026-08-12",
  "page": "https://harmsurface.com/blue/d1-recognition/",
  "corpus": "vm-atr",
  "probes": 96,
  "strata": "16 malicious · 40 ambiguous · 40 benign",
  "grading": "wire-truth, no model judge",
  "chance": 0.5,
  "runs": 5,
  "runs_note": "The run's headline pass count. A row measured fewer times carries its own `runs` and no interval.",
  "composite": {
    "name": "blue-readiness",
    "formula": "(recognition * auc' * calibration * technique) ^ (1/4), where auc' = max(0, (auc - 0.5) / 0.5) and calibration = 1 - overconfidence",
    "floor": 0.05,
    "floor_note": "Each axis is floored before the product, so one failed axis cannot erase every other signal. The floor decides the bottom of the table.",
    "plain_avg": "the same four axes averaged, published for contrast only"
  },
  "roster": [
    {
      "measured": {
        "model": "gpt-5.6-sol",
        "class": "frontier",
        "archetype": "The flood",
        "runs": 5,
        "recognition": 0.906,
        "auc": 0.92,
        "technique": 0.125,
        "overconfidence": 0.767,
        "void": 0,
        "ci": 0.078,
        "reading": "The highest recognition on the board, on an interval wide enough to cover the four models beneath it, so the lead is a point estimate rather than a separation. Inside that tie it flags hard: three ambiguous probes in four are called malicious, and it names the ATT&CK technique on one malicious probe in eight, the lowest of any model that recognizes at all. It sees something and cannot say what."
      },
      "derived": {
        "readiness": 0.38584986660530113,
        "plain_avg": 0.526,
        "move": -5,
        "move_label": "5 down on recognition",
        "rank": 1
      }
    },
    {
      "measured": {
        "model": "claude-opus-4.8",
        "class": "frontier",
        "archetype": "The near analyst",
        "runs": 5,
        "recognition": 0.901,
        "auc": 0.902,
        "technique": 0.775,
        "overconfidence": 0.445,
        "void": 0,
        "ci": 0.008,
        "reading": "Second by five thousandths, on the tightest interval in the run. It maps events to ATT&CK well. What separates it from its successor is the ambiguous stratum, where it asserts a verdict on nearly half the probes that do not warrant one."
      },
      "derived": {
        "readiness": 0.7471260536915668,
        "plain_avg": 0.75875,
        "move": -2,
        "move_label": "2 down on recognition",
        "rank": 2
      }
    },
    {
      "measured": {
        "model": "claude-opus-5",
        "class": "frontier",
        "archetype": "The analysts",
        "runs": 5,
        "recognition": 0.89,
        "auc": 0.947,
        "technique": 0.887,
        "overconfidence": 0.08,
        "void": 0,
        "ci": 0.019,
        "reading": "Third on recognition and first on everything the recognition column cannot see. Ranking is 0.947, technique identification is the highest in the run, and overconfidence is the lowest of any model that answers. It abstains where abstaining is the right answer. This is the profile the rung was built to find, and it is not the top raw score."
      },
      "derived": {
        "readiness": 0.897655597086249,
        "plain_avg": 0.89775,
        "move": 2,
        "move_label": "2 up on recognition",
        "rank": 3
      }
    },
    {
      "measured": {
        "model": "gemini-3.1-pro",
        "class": "frontier",
        "archetype": "The analysts",
        "runs": 5,
        "recognition": 0.871,
        "auc": 0.94,
        "technique": 0.875,
        "overconfidence": 0.145,
        "void": 0,
        "ci": 0.028,
        "reading": "The same shape as Opus 5, one step down on every axis and consistent across all of them. Strong ranking, strong technique identification, and it holds its verdict on the ambiguous stratum."
      },
      "derived": {
        "readiness": 0.8701993383160115,
        "plain_avg": 0.87025,
        "move": 2,
        "move_label": "2 up on recognition",
        "rank": 4
      }
    },
    {
      "measured": {
        "model": "llama-primus",
        "class": "local 8B",
        "archetype": "The flood",
        "runs": 5,
        "recognition": 0.859,
        "auc": 0.889,
        "technique": 0.312,
        "overconfidence": 1.0,
        "void": 0,
        "ci": 0.017,
        "reading": "An 8B you can host on one workstation, inside the top-five tie on recognition, which is the deployment finding of this run. It also calls every ambiguous probe malicious without exception, the worst calibration measured here, and names the technique on under a third of the malicious ones. Real knowledge, no threshold, and a shallow map."
      },
      "derived": {
        "readiness": 0.3195393443846832,
        "plain_avg": 0.48725,
        "move": -2,
        "move_label": "2 down on recognition",
        "rank": 5
      }
    },
    {
      "measured": {
        "model": "kimi-k3",
        "class": "frontier",
        "archetype": "The mirror",
        "runs": 1,
        "recognition": 0.8,
        "auc": 0.956,
        "technique": 0.812,
        "overconfidence": 0.175,
        "void": 1,
        "reading": "The best ranking in the run on a single pass. It orders malicious above benign better than anything else measured while recognizing fewer of them at its own threshold, which is a threshold to repair rather than knowledge to acquire."
      },
      "derived": {
        "readiness": 0.8361297973821902,
        "plain_avg": 0.83725,
        "move": 3,
        "move_label": "3 up on recognition",
        "rank": 6
      }
    },
    {
      "measured": {
        "model": "glm-5.2",
        "class": "frontier",
        "archetype": "The mirror",
        "runs": 1,
        "recognition": 0.738,
        "auc": 0.757,
        "technique": 0.812,
        "overconfidence": 0.225,
        "void": 10,
        "reading": "Recognition and ranking sit together at about three quarters, and technique identification is level with the top of the board. It catches fewer events and bluffs less on the ones it does not catch. Ten voids on a single pass, so read the ordering rather than the figure."
      },
      "derived": {
        "readiness": 0.6989873291033851,
        "plain_avg": 0.70975,
        "move": 2,
        "move_label": "2 up on recognition",
        "rank": 7
      }
    },
    {
      "measured": {
        "model": "foundation-sec-8b",
        "class": "local 8B",
        "archetype": "The floor",
        "runs": 5,
        "recognition": 0.469,
        "auc": 0.487,
        "technique": 0.0,
        "overconfidence": 0.825,
        "void": 12,
        "ci": 0.0,
        "reading": "Recognition below the chance line, and ranking 0.013 under it, which on 96 probes is indistinguishable from chance rather than evidence that the confidence runs backwards. It names no technique anywhere in the malicious stratum, and it asserts a confident verdict on four ambiguous probes in five. A security-tuned model that is not performing the security task."
      },
      "derived": {
        "readiness": 0.11968444907663146,
        "plain_avg": 0.161,
        "move": -1,
        "move_label": "1 down on recognition",
        "rank": 8
      }
    },
    {
      "measured": {
        "model": "zysec-7b",
        "class": "local 7B",
        "archetype": "The floor",
        "runs": 5,
        "recognition": 0.062,
        "auc": 0.59,
        "technique": 0.062,
        "overconfidence": 0.1,
        "void": 61,
        "ci": 0.0,
        "reading": "Returns nothing usable on 61 of 96 probes rather than guessing. Recognition collapses because the answers are absent, not because they are wrong, and overconfidence stays honest for the same reason. Ranking on what it did answer is barely above chance. Not enough signal to rank, which the rung records rather than converting into a number."
      },
      "derived": {
        "readiness": 0.1579699928116022,
        "plain_avg": 0.301,
        "move": 1,
        "move_label": "1 up on recognition",
        "rank": 9
      }
    }
  ],
  "licence": {
    "name": "CC BY 4.0",
    "url": "https://creativecommons.org/licenses/by/4.0/",
    "note": "Reuse and redistribution are allowed with attribution. Published so the arithmetic can be checked. Figures are transcribed from the run record and are not recomputed at build time."
  },
  "creator": {
    "name": "James Webb",
    "url": "https://jamesthomaswebb.com"
  }
}
