OpenAI's Rulebook For Itself
OpenAI published two governance documents in a single day. The Frontier Governance Framework lays out how OpenAI says it will manage safety as models grow. The "shared playbook for trustworthy third-party evaluations" sets out what an external safety evaluation should disclose - what claim, what system, what tooling, what safeguards.
OpenAI is now writing its own RSP. Two documents. One day. The frontier-lab safety race just turned into a credentialing competition.
The Frontier Governance Framework commits OpenAI to update its own rules as models, evaluations, and regulation change. The Trustworthy Evaluations playbook says external assessors should describe: the claim being tested, the evaluation content, the exact system under test (model, reasoning setting, tool access, harness, safeguards). It's a structure - and a soft attack on whoever runs evals without disclosing harness and tool access (read: most public benchmarks).
For PMs: expect every frontier-lab vendor to publish a similar framework within 90 days. For execs: ask your AI vendor which framework they sign off on and which third-party assessor they use. For governance: this is self-regulation racing the EU AI Act.