Demis Hassabis, CEO and co-founder of Google DeepMind, has proposed in the United States a technical body to define when a model should be considered “frontier,” test its risks before release, and continuously update the evaluation criteria. The first phase would be voluntary. After the protocol proves effective, approval could become a requirement for making these models available in the American market.
The proposal was published on July 14, 2026 in the article A Framework for Frontier AI and the Dawning of a New Age, shared by Hassabis himself on his X profile. It deserves attention not because a new rule is already in effect — it isn’t — but because it shows where the debate is progressing: from isolated lab commitments to comparable tests, external oversight, and accountability before release.
For companies using AI, the consequence is not to wait for AGI in order to act. It is to adopt a simple discipline now: know which models are in operation, what each agent can do, how its behavior is tested, and who is responsible when the system encounters an exception.
What Demis Hassabis is proposing
Hassabis starts with an ambitious premise. In his view, an artificial general intelligence — a system with the broad set of human cognitive abilities — may be just a few years away. He also envisions an exceptionally rapid economic and social impact.
These are Hassabis's predictions, not proven facts or a consensus among researchers. The most concrete part of the text lies in the institutional design he presents to handle uncertainty.
The proposal brings together eight elements:
- A technical standards body for frontier AI. The structure could be a public-private partnership under federal supervision or a self-regulatory organization, with independent experts and representatives of the open-source ecosystem on the council.
- Real funding and technical capacity. Funding would mainly come from industry, according to the suggested design, to hire top-level researchers and fund the computing required for large-scale testing.
- A dynamic definition of “frontier.” Rather than regulating companies by name or using a fixed parameter count, the body would classify models according to thresholds on benchmarks that are regularly updated.
- High-impact risk assessments. The tests would cover cybersecurity, biological threats, and other domains tied to national security. For agents, they could look for attempts to bypass safeguards, deceive evaluators, or act beyond defined limits.
- Pre-release review. In the voluntary phase, labs would share models with the body up to 30 days before public release.
- Evolution into a formal requirement. If the process proves robust, models classified as frontier would need to be approved before entering the American market.
- Independent and renewable testing. Saturated benchmarks would be replaced, held-out assessments would reduce the risk of training to the test, and third-party auditors could expand analysis capacity.
- Same standard for open and closed models. The criterion would be the system’s capability, not country of origin or distribution method. Models below the frontier threshold would be outside this regime.
The stated goal is to avoid two bad outcomes: letting each lab define on its own what it considers safe or freezing innovation with rules that are outdated before they take effect.
The proposal is not starting from scratch
The United States already has a federal institution with part of this mandate. The Center for AI Standards and Innovation (CAISI), connected to NIST, presents itself as the main point of contact between industry and the American government for testing and collaborative research on commercial AI systems.
CAISI develops voluntary practices, signs agreements with companies, conducts assessments, and coordinates work with defense, energy, homeland security, and intelligence agencies. Its focus includes demonstrable risks in cybersecurity, biosecurity, and chemical weapons.
Additionally, an executive order published by the White House on June 2, 2026 mandated the creation of classified benchmarks to measure advanced cyber capabilities and define when a system should be treated as a “covered frontier model.”
So, what is new in Hassabis’s text?
Our reading is that he proposes connecting elements that are currently separate. CAISI already tests and creates standards; the executive order already calls for a threshold for cyber capabilities; labs already publish their own frameworks. Hassabis adds a hybrid governance arrangement, industry funding, pre-release access, tests applicable to different domains, and an explicit move from voluntary to mandatory.
This is a comparative interpretation by XMACNA. The article does not say CAISI will be replaced nor does it present a finished bill.
Why the reference to FINRA matters
Hassabis cites the Financial Industry Regulatory Authority as inspiration. FINRA is a private, nonprofit, self-regulatory organization. It is funded by financial sector participants, but registered and overseen by the SEC. It writes and enforces rules, examines firms, and monitors market risks.
The analogy demonstrates a kind of architecture: technical knowledge and industry resources within an institution subject to public oversight. It does not mean FINRA would regulate artificial intelligence, nor that financial governance can be copied without adaptation.
For AI, the challenge is even greater. Capabilities change rapidly, tests can be optimized for until they lose value, and the same model may carry different risks depending on the tools and data it can access.
NIST itself recognizes part of this problem. In February 2026, the body published work on the statistical validity of AI assessments, warning that results may depend on implicit assumptions, mix performance definitions, and hide uncertainty. A shiny score, alone, is not a guarantee of safety.
From lab to company: the same logic on a smaller scale
A Brazilian company does not need to create an internal regulator. But it must stop treating the choice of a model as a one-off purchase that is resolved forever.
Hassabis’s reasoning can be translated into seven operational controls:
1. Keep a living inventory
Record which models, agents, and vendors are used, in which processes, and with which data. If nobody can answer where AI makes decisions or acts, governance hasn’t even begun.
2. Classify risk by use
AI that summarizes a meeting does not carry the same risk as an agent that changes records, schedules appointments, recommends treatments, analyzes credit, or sends external communications. The model may be the same; the impact is not.
This logic also appears in the Brazilian discussion about the AI legal framework and Bill 2338: the useful question is what the system does, who it affects, and what evidence is left afterward.
3. Test the function, not just the model
Lab benchmarks measure general capabilities. The company must test the real flow: incomplete data, conflicting instructions, API failures, attempted fraud, out-of-policy requests, and transfers to humans.
That’s the difference between a demo and AI in production. The system needs to function on a bad day, not just during demonstrations.
4. Limit tools and autonomy
An agent should only access the systems needed for its function. Sensitive actions require approval, value limits, or dual verification. Autonomy without scope is just a more sophisticated way to create risk.
5. Maintain audit trails and evidence
Record relevant inputs, tools invoked, decisions, exceptions, and outcomes. Without an audit trail, the company can’t investigate failures, explain decisions, or improve the system.
This principle is central in agentic engineering: a professional agent acts, but also provides accountability.
6. Keep human judgment available
Human escalation is not a failure of automation. It is part of the design. AI must recognize when data is missing, when there is conflict, when risk has increased, or when someone requests a review.
The right mix of AI and human judgment avoids two extremes: blindly depending on the system or turning all automation into a queue of useless approvals.
7. Always reassess when the system changes
A change of model, new prompt, additional integration, access to another data source, or increased autonomy changes the risk. The assessment needs to keep up with the actual version in operation.
The Frontier Safety Framework shows how this can work
Google DeepMind already applies an internal version of this logic. The Frontier Safety Framework 3.1 defines tracked and critical capability levels, early warning assessments, and risk-proportional measures. The latest version broadens focus to harmful manipulation and possible misalignment scenarios.
There is, however, an important difference. An internal framework is still defined and carried out by the laboratory itself. The body imagined by Hassabis would create a common, external layer capable of comparing different organizations.
This move — from individual commitment to collective assessment infrastructure — is the political core of the article.
The future is not set, and that’s the point
Hassabis’s text mixes technological enthusiasm, an AGI forecast, and a concrete regulatory proposal. You don’t have to agree with the timeline suggested to see the issue: more capable, agentic systems connected to tools need better assessments than a fixed set of questions.
For governments, the challenge is developing a technical standard that keeps pace without focusing power in the very labs being evaluated. For companies, the challenge is more immediate: transform AI into an auditable operation before a lack of control becomes an incident, vendor dependency, or regulatory debt.
The responsible path is not paralysis. It is to accelerate with tested safeguards, evidence, and clear accountability.
If your company already uses AI but has not yet mapped out functions, autonomy, data, logs, and human escalation, the XMACNA AI assessment helps identify where to start.
Frequently asked questions
What did Demis Hassabis propose?
He proposed a technical body in the United States to define which models are frontier, assess them in high-risk areas, and continuously update tests. The initial phase would be voluntary and could evolve into mandatory approval before availability in the US market.
Does this body already exist?
Not in the form described by Hassabis. CAISI, connected to NIST, already conducts assessments, develops standards, and coordinates federal agencies. The proposal adds a hybrid structure inspired by self-regulation, pre-release access, and a possible mandatory regulatory gateway.
Will AGI arrive in a few years?
Hassabis believes this is likely, but there is no scientific consensus on timing, operational definition, or technical pathway for AGI. The prediction should be treated as an informed opinion from an industry leader, not a confirmed timetable.
Would the proposal affect open source models?
According to the article, the criterion should be model capability, not whether it is open or closed source or its country of origin. Projects below the frontier threshold, including many from startups and universities, would be excluded.
What should a company do now?
Inventory AI systems, classify uses by risk, limit tools and autonomy, test real workflows, maintain logs, define human escalation, and reassess operations whenever the model, data, or integrations change.