AI

OpenAI Anthropic Google Discuss AI Standards Body

2026-09-15 - ABikram Mondal

OpenAI Anthropic Google Discuss AI Standards Body

Three labs explore self regulation

Anthropic, OpenAI and Google have held regular meetings since July to design a voluntary standards body. The group would set benchmarks for safety testing and auditing of frontier models. Executives below CEO level have driven the talks. Sources close to the discussions told The Information the effort aims to create rules without waiting for government action.

OpenAI chief Sam Altman told staff last week that major labs will likely need to build this framework themselves. A draft executive order for a government backed body stalled inside the White House. The three companies see the industry version as a faster route to common testing requirements.

The meetings continued even after Anthropic CEO Dario Amodei published his call to pace frontier development. Working groups have focused on shared benchmarks rather than new limits on release speed. Participants view consistent audits as a way to reduce duplication and raise the floor for everyone.

Google and OpenAI have competed fiercely on model releases. Anthropic has positioned itself as the safety focused alternative. Their joint work on standards marks a shift from public criticism to private coordination on evaluation methods.

The talks come as stock prices for AI linked companies dropped on Monday after the broader slowdown warnings. Market reaction showed investors weighing the chance that stricter self rules could slow product cycles.

What the body would actually test

Early outlines include common protocols for capability evaluation and misuse resistance. Groups would define red team exercises and reporting formats that all signatories follow. The goal is to make results comparable across labs rather than each company inventing its own metrics.

Participants have discussed how to handle long context reasoning tests and agentic task benchmarks. They also looked at ways to audit training data provenance and post training alignment steps. No final list of required tests exists yet.

The body would remain voluntary. Labs could choose participation levels based on their model tier. Smaller releases might face lighter review than frontier systems. The structure leaves room for public release of summary scores while protecting proprietary details.

One open question is enforcement. The draft plans do not include penalties for missing targets. Instead the group would publish participation status and test outcomes. Public pressure and customer demand would supply the main incentives.

Another point under review is how to include input from academic researchers and civil society groups. Early drafts call for an advisory panel that meets quarterly but holds no veto power over lab decisions.

Why the timing matters now

Calls for coordinated action have grown louder in recent days. Amodei published a long essay urging slower frontier progress. Altman and Elon Musk voiced support in separate statements. The standards body talks predate those public remarks but gained urgency from them.

China pushed back hard against any slowdown framing. Its foreign ministry called the warnings fearmongering. Beijing continues its own push for domestic AI leadership. The contrast makes a unified Western industry position more visible.

Inside the companies the pressure to ship remains intense. Each lab has released new models or variants in the past month. The standards effort is an attempt to manage the next phase without ceding ground to regulators who might impose slower timelines.

Partner companies have already started asking for clearer safety attestations. Cloud providers and enterprise buyers want consistent signals before scaling deployments. The standards body could supply one document set that satisfies multiple customers.

ABikram Mondal builds automation for exactly this kind of problem at https://abikrammondal.com/services/automation.

Benchmarks that could become mandatory

Current candidate tests include variants of existing agentic coding suites and cybersecurity vulnerability scans. Labs have shared internal results on some of these already. The new body would require public reporting of scores on a fixed schedule.

Multimodal reasoning benchmarks are also on the list. The group wants measures that track performance on long horizon tasks rather than single turn prompts. This aligns with the direction most frontier models have taken in 2026.

One proposed addition is a data retention and misuse logging standard. The rule would require labs to keep certain interaction records for audit periods. Customers would receive clearer notices about what gets stored.

Cost and speed metrics would stay outside the core requirements. The focus stays on safety and reliability signals. Labs can still compete on price and latency while agreeing on evaluation baselines.

Early participants expect the first set of published standards by early 2027. That timeline depends on reaching consensus among the three founding labs and any additional signatories that join later.

Who gains and who stays out

Enterprise buyers gain a single set of scores they can compare across providers. Security teams get clearer signals on which models have passed which audits. Smaller developers benefit if the standards reduce the need for custom safety reviews.

Independent researchers could gain access to more structured evaluation data. The advisory panel idea includes slots for outside experts. Whether those voices shape the final rules remains to be seen.

Companies that refuse to join would face questions from customers about why their models lack the common certification. The market penalty could prove stronger than any formal rule.

Regulators in Washington and Brussels have watched the talks. A credible industry body might delay new legislation. It could also give governments a ready made framework to reference if they choose to mandate participation later.

Meta and xAI have not joined the initial discussions. Their absence leaves open whether a truly industry wide standard will emerge or whether multiple competing frameworks appear.

What remains unclear

The scope of the first standards round is still narrow. It covers testing and reporting but stops short of release gates or compute caps. Broader questions about training data sources and model weight sharing sit outside the current scope.

Cost of participation has not been disclosed. Running the audits and maintaining the reporting infrastructure will require staff time and compute. Smaller labs may find the overhead high relative to their resources.

International reach is another gap. The initial group is US centric. European and Asian labs have not been invited to the founding table. Any global standard would need later expansion.

Enforcement remains the weakest point. Without binding contracts or regulatory backing the system relies on reputation and customer pressure. Past voluntary initiatives in tech have shown mixed results on compliance.

The next few months will show whether the three labs can turn the discussions into a working body or whether competitive tensions pull the effort apart.

The short version. Anthropic, OpenAI and Google are quietly building a voluntary AI standards body to set shared safety testing rules before regulators step in.

Sources

Reported from the sources above on 2026-09-15. Figures are as published at the time of writing. If something here has moved on, the linked source is the one to trust.

From the desk of ABikram Mondal

If you got here because you are actually thinking about putting models like this to work inside a real business, wired into the tools a team already uses, that is the work I do. I build for founders and small teams who want the thing to exist and work, not a deck about it.

Automation services  ·  everything I do  ·  talk to me

A new version is available.

Install on iPhone

Add ABikram to your Home Screen — it opens fullscreen like an app and stays updated automatically.

  1. Tap the Share icon in Safari's bottom bar.
  2. Scroll and tap Add to Home Screen.
  3. Tap Add — the trishul icon appears on your phone.

Get the ABikram App

Install the app for blogs, stories, Kundli and every tool — with live updates.