In December 2025, the Future of Life Institute published a document that should have caused a global panic. Instead, it was met with a collective shrug from an industry moving too fast to look in the mirror.
The document was the Winter 2025 AI Safety Index — an independent, peer-reviewed assessment of the world's eight leading AI companies across 35 indicators spanning risk assessment, current harms, safety frameworks, existential safety, governance, and information sharing.
The headline finding: no company scored higher than a D in existential safety.
Not Anthropic, the company that markets itself as the "safety-first" AI lab. Not OpenAI, which was literally founded to prevent AI catastrophe. Not Google DeepMind, which employs some of the world's foremost alignment researchers. Every single company building toward artificial general intelligence — the most powerful and potentially dangerous technology in human history — received a failing grade on the question of whether they have a credible plan to control it.
This is not a technical footnote. This is the report card for the entities constructing the most consequential technology of the century, evaluated by the people most qualified to assess the risks.
The Scorecard: Where the Giants Stand
The Winter 2025 AI Safety Index evaluated eight companies using a US GPA-based grading system (A–F). The expert panel — including Stuart Russell (UC Berkeley), David Krueger (University of Montreal), Dylan Hadfield-Menell (MIT), Jessica Newman (AI Policy Hub, UC Berkeley), Sharon Li (University of Wisconsin-Madison), and Tegan Maharaj (Mila) — assessed each company against absolute performance standards based on evidence collected through 8 November 2025.
The results:
| Company | Overall Grade | Score (4.0 Scale) | |---|---|---| | Anthropic | C+ | 2.67 | | OpenAI | C+ | 2.31 | | Google DeepMind | C | 2.08 | | xAI | D | 1.17 | | Zhipu AI (Z.ai) | D | 1.12 | | Meta | D | 1.10 | | DeepSeek | D | 1.02 | | Alibaba Cloud | D- | 0.98 |
To contextualise these grades: in any university, a C+ is a mediocre pass. A D is barely above failure. And the industry's best performer — Anthropic, the company that has invested more in alignment research than any other — achieved the equivalent of a student who "meets minimum expectations but shows limited mastery."
For the rest of the industry, the grades translate to: "shows fundamental gaps in understanding and application."
The Six Domains: Where Companies Succeed and Fail
The AI Safety Index evaluates companies across six domains. The pattern across all eight companies reveals a sector with pockets of competence surrounding a vast void of preparation.
Domain 1: Risk Assessment
This domain measures whether companies systematically identify, evaluate, and categorise the risks their AI systems pose. It covers hazard identification, risk measurement, and external evaluation practices.
Top performers: Anthropic and OpenAI both publish detailed model cards and system evaluation reports. Anthropic's Responsible Scaling Policy (RSP) establishes specific capability thresholds ("AI Safety Levels") that trigger additional safety measures as models become more capable.
Weaknesses: Even the best performers lack quantitative, risk-tied thresholds that would trigger automatic deployment pauses. The thresholds are descriptive, not binding. A former safety team member at one of the top-tier labs described the internal dynamic: "We can identify a risk. We can flag a risk. We cannot stop a deployment over a risk."
Domain 2: Current Harms
No company scored higher than a D in existential safety. Every entity building toward AGI received a failing grade on the question of whether they have a credible plan to control it.
This domain assesses how companies address the harms their current AI systems are already causing — bias, discrimination, misinformation, privacy violations, and labour exploitation.
Industry pattern: This is the domain where companies score highest overall. Most have established reporting mechanisms, conducted bias audits, and published transparency reports. The reason is practical: current harms generate lawsuits, regulatory scrutiny, and negative press. The incentive structure aligns safety with self-interest.
Domain 3: Safety Frameworks
This domain evaluates whether companies have published, actionable safety frameworks with clear escalation procedures, red-teaming protocols, and deployment gates.
Notable development: Between the Summer and Winter 2025 editions of the Index, both xAI and Meta published frontier AI safety frameworks for the first time. However, the expert panel judged these frameworks as "limited in scope and independent oversight" — documents designed to demonstrate compliance effort rather than to constrain behaviour.
Google DeepMind's Frontier Safety Framework and Anthropic's RSP remain the most detailed public safety frameworks in the industry. OpenAI's Safety Advisory Group structure was recognised for its whistleblowing policy transparency — one of the few areas where OpenAI led the field.
Domain 4: Existential Safety — The Sector-Wide Failure
This is where the scorecard turns damning.
Existential safety measures whether companies have credible, actionable plans to prevent catastrophic outcomes from advanced AI systems — including loss of control, misuse for weapons development, or the emergence of AI systems that pursue objectives misaligned with human values.
Result: Every company received a D or lower.
For the second consecutive edition of the Index, no company presented an evidence-based plan for controlling or aligning systems at or beyond the AGI frontier. This is not because the companies don't acknowledge the risk. Most explicitly state that they are building toward AGI. OpenAI's charter describes its mission as building "safe and beneficial artificial general intelligence." Anthropic's founding document identifies "the responsible development and maintenance of advanced AI" as its core purpose.
But acknowledgment is not preparation. The expert panel found:
- No company has published quantitative risk thresholds that would trigger automatic capability restrictions
- No company has demonstrated a credible alignment technique that scales to systems significantly more capable than current models
- No company has established independent oversight mechanisms with the authority to halt development or deployment
- No company has addressed the fundamental problem of controlling systems that may be more intelligent than their creators
Stuart Russell, the UC Berkeley professor whose 2019 book Human Compatible crystallised the alignment problem for a global audience, has noted that the gap between corporate rhetoric and corporate action on existential safety is "the most significant disconnect in the history of technology development."
The analogy that emerges is nuclear. Imagine if every nuclear weapons laboratory acknowledged that an uncontrolled chain reaction could destroy a city, explicitly stated their goal was to build the most powerful weapon ever conceived, and simultaneously received failing grades on their containment protocols. That is the current state of the AI industry's relationship with existential safety.
Domain 5: Governance & Accountability
This domain evaluates internal governance structures, board-level oversight, whistleblowing protections, and the independence of safety teams from commercial decision-making.
Industry divide: A clear gap separates Western companies from Chinese ones on governance disclosure. Anthropic, OpenAI, and Google DeepMind all publish governance structures and, to varying degrees, provide information about safety team authority. Chinese firms — DeepSeek, Zhipu AI, Alibaba Cloud — provide substantially less public information about governance structures, though the panel noted their compliance with domestic regulations mandating content labelling and incident reporting.
Every dollar and every day invested in safety is a dollar and a day not spent on capability advancement. The Nash equilibrium is mutual acceleration toward inadequately tested systems.
The whistleblowing question: OpenAI was recognised for having the most transparent public whistleblowing policy, though it also faced criticism for its treatment of employees who raised safety concerns publicly. The tension between welcoming internal dissent and punishing external disclosure remains unresolved across the industry.
Domain 6: Information Sharing
This domain measures transparency with external researchers, governments, and the public — including the sharing of evaluation results, incident reports, and safety research.
Mixed results: Anthropic and Google DeepMind lead in publishing safety research. OpenAI has reduced its publication of frontier research over time, citing competitive concerns. Meta publishes model architecture details but limits information about safety testing processes. The Chinese companies provide the least external information, though cultural and regulatory factors contribute to this pattern.
The Race Dynamics: Why Safety Keeps Losing
The AI Safety Index reveals a structural problem that no individual company can solve: the competitive dynamics of the AI industry actively punish safety investment.
Every dollar and every day invested in safety is a dollar and a day not spent on capability advancement. In a market where being first to a capability threshold can mean billions in revenue and market dominance, the incentive structure rewards speed and punishes caution. Anthropic may be the "safety-first" lab, but it is also in a desperate race against OpenAI, Google, and an increasingly capable set of Chinese competitors.
The result is what the FLI panel describes as a structural "race to the bottom" — not because companies are evil, but because the market they operate in rewards exactly the behaviour that the safety community warns against. Safety is a collective action problem. The company that unilaterally slows down loses market share. The company that accelerates gains it. The Nash equilibrium — the stable outcome of rational actors pursuing self-interest — is mutual acceleration toward inadequately tested systems.
This is why voluntary safety commitments, however well-intentioned, are structurally inadequate. Voluntarism works when the incentives align with the commitments. In the AI industry, they don't.
What the Safety Index Doesn't Measure
The FLI AI Safety Index is, by its own acknowledgment, limited to evaluating what companies claim to do and what public evidence supports. It cannot assess:
- What happens inside labs that isn't disclosed
- The gap between published safety policies and actual internal practice
- The effectiveness of safety measures against adversarial actors
- Whether safety teams have genuine authority to halt deployments
A former employee of one of the top-scored companies, speaking anonymously, described the dynamic: "We have the best safety policies in the industry. We also have the most creative ways of interpreting them when launch timelines are at stake."
The Index measures the structure of safety. It cannot measure the culture of safety. And culture, in high-stakes technology development, is everything.
The Missing Layer: Architectural Safety
The most striking absence in the AI Safety Index — and in the industry it evaluates — is architectural safety. All eight companies treat safety as a property that is added to AI systems after they are built. Safety teams review models. Red teams probe for vulnerabilities. Evaluation frameworks test for harmful outputs. But the fundamental architecture remains: build the capability first, then try to make it safe.
This is analogous to building a nuclear reactor and then adding the containment structure as an afterthought. The order is wrong.
Society OS's Agent Protocol & Safety Charter represents a fundamentally different approach: a Logic Fortress where safety is not a layer applied on top of capability, but a constraint embedded at the deepest architectural level. The three-swarm model — Foundry Swarm (creation with human-in-the-loop sign-off), Guardian Swarm (real-time audit with automated freeze on detected logic contradictions), and Embassy Swarm (external relations with Swiss Neutrality communication protocols) — builds safety into the operating system itself.
Safety is not a layer applied on top of capability. It is a constraint embedded at the deepest architectural level. Safety is the architecture.
The critical innovation is the USI (Universal Schema Interface) Precedence Rule: no agent may suggest or execute a transaction that bypasses the USI reconciliation layer. Safety is not a review process. It is an invariant — a mathematical constraint that cannot be overridden, not even by the system's creator.
The FLI Safety Index's finding that no company has "published quantitative risk thresholds that trigger automatic capability restrictions" is precisely the problem the Equilibrium Test in Society OS's logic gating solves. Before any state change in the 6×7 matrix, the system simulates the ripple effect. If the change creates an inflation state or a conflict state, the action is automatically rejected. No human override. No "creative interpretation." The mathematics decides.
This is what architectural safety looks like: not policies that humans can bend, but invariants that mathematics enforces.
The Path Forward: From Grades to Governance
The FLI AI Safety Index is a necessary tool. It provides the first standardised, peer-reviewed benchmark for comparing how the world's most powerful AI companies approach the most consequential risks. Its contribution to transparency and accountability is invaluable.
But grading is not governing. The Index can identify the problem. It cannot solve it.
Solving the existential safety problem requires three things that no single company can provide:
1. Independent oversight with enforcement authority. Not advisory boards. Not voluntary commitments. Bodies with the legal power to access labs, evaluate systems, and halt deployments. The EU AI Act's AI Office is a step in this direction. But it was designed for high-risk classification, not for evaluating existential capability thresholds.
2. Shared safety infrastructure. The alignment problem is a collective action problem. No single company's safety research protects the world from a competitor's unsafe system. Shared evaluation protocols, common red-teaming methodologies, and open safety research are public goods that the market will not adequately provide.
3. Governance-as-architecture. The ultimate solution is not better oversight of unsafe systems. It is systems that are safe by design — architectures where the governance constraints are embedded in the infrastructure, not applied as external review.
The AI Safety Index gives every company in the industry a grade.
The question is whether the industry will build systems that don't need grading — because safety is no longer a feature that can be added or removed.
It is the operating system itself.
The 42 Categories of One™ that Society OS has pioneered — innovations with zero prior art — include the only known working implementation of this principle: a self-amending governance architecture embedded at the operating system layer, one that adapts based on real-time ethical dilemmas, user feedback, and pre-defined sovereign values.
No company graded by the FLI has built anything like it.
Because none of them started with the premise that safety is not a feature.
Safety is the architecture.
This article is part of the Sovereign Intelligence Hub's safety series. For how the governance gap enables unsafe deployment, see [The Governance Gap](/hub/the-governance-gap-why-ai-regulation-cant-keep-up). For the OWASP taxonomy of agentic threats, see [OWASP Top 10 for Agentic AI](/hub/owasp-top-10-for-agentic-ai-the-security-threats-nobody-is-talking-about). For the international treaty failure, see [The International AI Treaty That Actually Matters](/hub/the-international-ai-treaty-that-actually-matters).
Sources & Further Reading
- 1.Future of Life Institute — AI Safety Index Winter 2025 (Full Report)
- 2.Future of Life Institute — AI Safety Index Summer 2025
- 3.Russell, S. — Human Compatible: Artificial Intelligence and the Problem of Control (Viking, 2019)
- 4.Anthropic — Responsible Scaling Policy
- 5.Axios — AI Companies Get Failing Safety Grades (December 2025)
- 6.Los Angeles Times — AI Company Safety Scorecard (December 2025)
- 7.Society OS — Agent Protocol & Safety Charter (2026)
- 8.QuantumZeitgeist — AI Safety Future of Life Risk Mitigation Analysis (2025)



