Signal Scout
Signal Scout · Research · September 2026

Who tests frontier models

Anthropic, OpenAI, Google and Meta each publish a system card with every major release. Some name the outside researchers who tried to break the model first. Some name no one.

Cards read 16Evaluators named 16Scope Last 12 monthsMethod Hand-verified
16
system cards read
16
distinct evaluators named
4
of 16 name no one
01
The grid

Who names who

16 system cards from the last 12 months, read by hand. 16 distinct external evaluators named. 4 of 16 cards name none.

Scope: the most recent model releases from Anthropic, OpenAI, Google and Meta, September 2025 through September 2026. This is a one-time reading, not a scraper -- it does not grow on its own the way the rest of the site does. Each row links to the source card, so the claim is checkable in one click.

Presence by carddot = named in this card
CardLabDateSourceSecureBioUK AISIGray SwanDeloitteIrregularApollo ResearchUS CAISITrajectory LabsSignature ScienceFaculty.aiMETRRedwood ResearchFAR.AIFrontier DesignMicrosoft AI Red TeamPattern Labs
Claude Opus 5AnthropicJul 2026
Claude Fable 5.1 / Mythos 5.1AnthropicSep 2026
Claude Fable 5 / Mythos 5AnthropicJun 2026
Claude Sonnet 4.5AnthropicSep 2025
Claude Haiku 4.5AnthropicOct 2025
GPT-5OpenAIAug 2025
GPT-5.2 updateOpenAIDec 2025
GPT-5.3-CodexOpenAIFeb 2026
GPT-5.5OpenAIApr 2026
GPT-5.1 addendumOpenAINov 2025
Gemini 3 ProGoogleNov 2025
Gemini 3.1 ProGoogle2026
Gemini 3 FlashGoogleDec 2025
Muse SparkMeta2026
Muse Spark 1.1Meta2026
Muse Spark ContemplatingMeta2026

Presence only: an evaluator either appears in a card or does not. A model tested once reads the same as one mentioned twenty times, because mention count reflects how a card was written, not how much work was done.

02
Leaderboard

The leaderboard

Ranked by number of distinct cards each evaluator appears in, not mentions.

EvaluatorCards · labs
SecureBio12Anthropic, Meta, OpenAI
UK AISI9Anthropic, Meta, OpenAI
Gray Swan8Anthropic, Meta, OpenAI
Deloitte7Anthropic, Meta
Irregular7Anthropic, Meta, OpenAI
Apollo Research6Anthropic, Meta, OpenAI
US CAISI6Anthropic, Meta, OpenAI
Trajectory Labs5Anthropic, Meta
Signature Science4Anthropic
Faculty.ai3Anthropic, Meta
METR3Anthropic, OpenAI
Redwood Research2Anthropic
FAR.AI1OpenAI
Frontier Design1Meta
Microsoft AI Red Team1OpenAI
Pattern Labs1OpenAI

SecureBio, UK AISI and Gray Swan are the only evaluators named by all three labs that name anyone. Google's three Gemini cards in this set name no external evaluator.

03
Method

Method and what this is not

Every card was downloaded, converted to text, and searched for evaluator names. Every hit was then read in its surrounding paragraph by hand before being counted -- automated keyword matching alone produced false positives that would have been wrong on a public page:

  • "RAND" matched "randomly selected" and a bibliography citation to a RAND paper, never RAND Corporation testing a model.
  • "Scale AI" matched a citation to SWE-bench Pro, a benchmark Scale AI published, cited as something a model was run on -- not Scale AI acting as an evaluator.
  • "Promptfoo" matched promptfoo.dev listed as a site Claude tried to pull answers from during an evaluation. The opposite of being an evaluator.

Cards showing no evaluators were read directly and confirmed empty, not assumed from a missed keyword. That does not mean no external testing happened -- a lab may test without naming who, or use an evaluator under an NDA. Absence here means not named in the card, nothing more.

This page does not update automatically and is not part of the weekly corpus. Re-read by hand when a new system card ships.