The most notable AI news tonight isn’t about a new model.

It’s this:

Governments are beginning to directly test the most advanced AI models themselves.

Reuters confirmed on September 10:

The EU cybersecurity agency ENISA has gained access to Anthropic’s Mythos 5 model.

And:

Testing has already begun.

An EU Commission spokesperson also confirmed:

ENISA has obtained access to OpenAI’s latest model:

GPT-6 Astra.

At first glance, this may seem like just a “government obtained models” story.

But it actually signifies AI security crossing a new threshold.

Previously, many times:

AI developers:

Tested their own models.

Produced their own System Cards.

Illustrated risks themselves.

Governments made regulations based on this data.

Now, the landscape is changing to:

Governments themselves need direct access to the models to understand their capabilities firsthand.

Why does ENISA particularly want access to these models?

ENISA stands for:

European Union Agency for Cybersecurity.

Think of it as the:

EU Cybersecurity Agency.

It’s not just a consumer testing organization.

What it really cares about is:

If the most powerful AI models develop very high cybersecurity capabilities, can they:

Discover vulnerabilities?

Analyze malware?

Assist with incident response?

Help protect critical infrastructure?

And conversely:

If these capabilities are abused, can they enhance offensive attack capabilities?

So ENISA’s focus isn’t:

“Which model writes emails more beautifully?”

It is:

“What level of capability does this AI reach when deployed in real cybersecurity environments?”

GPT-6 Astra has surpassed OpenAI’s own Critical cybersecurity capability threshold

This is why this development is especially important now.

OpenAI officially confirmed in early September:

GPT-6 Astra has reached the company’s Preparedness Framework level of:

Critical Cybersecurity Capability.

This is the first OpenAI model formally classified at this level.

OpenAI’s description of this tier is very specific.

Given proper:

Tools,

and access permissions,

Astra already has the ability to find previously unknown security vulnerabilities without step-by-step human intervention,

and can even develop new exploit techniques.

During testing, OpenAI stated:

Astra found two previously unknown zero-day vulnerabilities,

and was able to chain them into an actual exploit chain.

So the conversation is no longer:

“Will AI someday become good at cybersecurity?”

It’s now:

At least according to the companies’ own tests, this capability is emerging.

But even if companies call it “Critical,” governments still need to conduct their own testing

This is the real point tonight.

OpenAI has released many tests.

Anthropic also has its own:

Model evaluation,

Safeguards,

Access control,

and security research.

But the government’s challenge is:

Public safety cannot rely solely on vendors evaluating themselves.

This doesn’t mean:

Companies necessarily lied.

Rather, their roles are fundamentally different.

Model developers answer questions like:

Can the product be launched?

Can users access it?

What risks need restriction?

Governments need to consider:

What happens if banks use it?

Electricity companies?

Telecoms?

Government agencies?

What if attackers gain similar capabilities?

What impact would that have on national cyber resilience?

These questions go beyond typical pre-release safety tests.

Just days ago, ENISA had not yet received Mythos access

This progress is notable because it’s very recent.

On September 1:

The European Commission publicly said:

ENISA had received an invitation from Anthropic,

but the parties were still working through:

Security standards,

Safety criteria,

and onboarding procedures.

The message then was clear:

The process was ongoing.

Today:

The Commission has confirmed:

ENISA has not only obtained access to Mythos 5,

but also:

Testing is underway.

In other words, moving from:

“Trying to get access”

to:

“Already conducting tests.”

But this is Mythos 5, not Mythos 5.1

This detail is easy to get wrong.

Anthropic released:

Mythos 5.1

on September 1.

It’s an updated Mythos-class model.

Anthropic positions the Mythos series for:

Cybersecurity,

Biology,

and other high-capability research uses.

Access to Mythos 5.1 remains limited to a small group of vetted organizations.

Reuters confirmed today ENISA has access to:

Mythos 5

not Mythos 5.1.

This version distinction is important to preserve.

Why has model access itself become a policy issue?

Because the more capable a model is,

the less suitable it is for unrestricted public access.

Model companies are increasingly using:

Vetting,

Restricted access,

Safety tiering,

and special cybersecurity controls.

This approach has security reasoning behind it.

But another problem arises.

If only:

Model companies,

or a handful of US institutions,

can see the highest capability models,

how will European entities like:

Cybersecurity agencies,

critical infrastructure operators,

and research organizations

know what they might face in the future?

This is a question the EU has been pressing in recent months.

For Europe,

not having model access could itself become a security risk.

Because you cannot test for:

The capabilities potential attackers may soon acquire,

nor conduct early research on:

Defense strategies.

The EU is preparing to build its own AI Cybersecurity Testing Platform

This isn’t a new idea.

The EU’s published Cybersecurity Strategy already plans for:

ENISA

to partner with the European Commission Joint Research Centre (JRC) to establish a:

Secure Testing Platform for AI in Cybersecurity.

The purpose is to enable various European bodies to test AI models with high cyber capabilities in a:

Safe,

Controlled,

and Standardized Security Rule environment.

Testing activities will include:

Vulnerability scanning,

Remediation,

Incident response,

Threat intelligence,

and detection.

In short:

Finding vulnerabilities,

Fixing them,

Handling attacks,

and analyzing threats.

This is not about checking AI Act compliance

This distinction is important.

EU policy documents clearly state the main purpose of this cybersecurity testing environment is not:

To determine if AI models comply with the AI Act.

Its real objective is to:

Study how AI can be applied in cyber defense,

and whether AI’s capabilities will reshape the cybersecurity landscape.

So don’t reduce today’s news to:

“The EU is reviewing OpenAI and Anthropic for AI Act violations.”

That’s not what’s happening.

Today’s focus is:

Cybersecurity capabilities testing.

What does this mean for regular companies?

You might think:

“My company won’t be using Astra to find zero-days.”

But challenges will cascade down.

As frontier models gain stronger:

Computer usage,

Coding,

Browsing abilities,

Cybersecurity capabilities,

and agent functionalities,

businesses choosing AI can no longer evaluate models only by:

How well they answer,

How fast they are,

Or cost.

They must start considering questions like:

What resources can this model access?

Which tools can it call?

Does it have network access?

Where are credentials stored?

Who authorizes sensitive operations?

Are actions logged?

Can the AI be stopped if needed?

These issues were traditionally the cybersecurity team’s concern.

Now they must become:

Part of the AI workflow itself.

As capabilities grow, security can't rely only on the model 'being well-behaved'

This is a common lesson from recent AI incidents.

If AI only:

Answered with text,

Most security concerns focused on content.

Now AI can:

Write code,

Browse the web,

Operate computers,

Connect accounts,

Call tools,

Perform multi-step tasks autonomously,

and even carry out cybersecurity work.

Simply instructing the model:

“Don’t do bad things.”

Is no longer sufficient.

We also need:

Sandboxes,

Permission controls,

Isolation,

Monitoring,

Approval mechanisms,

Credential boundaries,

and network controls.

In other words:

Do not expect AI to always make the right decision by itself.

The system must also restrict it.

Government testing models also serves another crucial purpose

Imagine an AI company claims:

“Our model is very safe.”

And another says:

“Our model doesn’t have this problem.”

If all test methods, data, thresholds, and results differ,

The government has difficulty making meaningful comparisons.

Once independent testing capabilities exist,

it becomes possible to develop:

Common evaluations,

Shared benchmarks,

and consistent capability thresholds.

For example:

What qualifies as a model able to assist in finding vulnerabilities?

When does it begin to autonomously develop exploits?

Which capabilities warrant restricted access?

Which can safely enter critical infrastructure?

If these standards form,

AI security can move from:

Individual company claims

to:

Cross-company comparisons.

A significant step forward.

But ENISA having access doesn’t mean these models are proven dangerous

It’s important not to misinterpret this.

Government testing:

Does not equal

The model being deemed dangerous.

Nor does it mean:

The EU is planning to ban these models.

The only confirmed facts from Reuters today are:

ENISA has obtained access to:

Mythos 5,

and GPT-6 Astra,

and has begun testing Mythos 5.

What they’ve found,

Whether full results will be published,

And what policy impacts there might be,

remain unknown.

So today’s key takeaway isn’t:

“The EU caught dangerous AI.”

It is:

The EU now has AI models it can inspect itself.

This might be the truly critical infrastructure for the next phase of AI regulation

When people think about AI regulation,

They often picture:

Laws,

Fines,

Reporting,

and rules.

But if regulators do not have:

Models,

Compute resources,

Testing environments,

and expert personnel,

Real evaluation capabilities,

Regulations risk becoming little more than paperwork checks.

Mature AI governance will likely require this foundational infrastructure:

Governments having the ability to test AI models themselves.

Not by building their own ChatGPT,

But by conducting tests in controlled environments,

Accessing models,

Running evaluations,

Reviewing logs,

Benchmarking,

Validating claims,

And understanding the real limits of new capabilities.

So what is the real turning point tonight?

It’s not that ENISA gained two new AI accounts.

It is that:

Frontier model capabilities have grown so strong that:

Model access itself becomes part of public safety capabilities.

Model developers need to test.

Users need to test.

Governments are starting to test too.

If the strongest AI can help defenders:

Find vulnerabilities faster,

Patch them,

and handle attacks,

Authorities cannot wait until models are fully public before learning.

On the other hand,

If such capabilities also increase attack risks,

Access cannot be unrestricted.

The real challenge becomes:

How to ensure those who should test can get early access, while preventing unchecked spread of high-risk capabilities?

This is the core significance of Mythos 5 and GPT-6 Astra entering ENISA today.

AI security is moving from:

“Companies telling governments how safe their models are”

Toward:

“Governments verifying for themselves.”

Today, progress with AI continues.

Learn one AI skill a day.

Save a little time daily.

Improve your abilities step by step.

SasaDaily, growing with you.

Recommended Reading

AI Evening Report | 2026/09/02: OpenAI Officially Confirms Astra Passes "Critical" Cybersecurity Threshold, Restricting Access to Strongest Cyberattack Capabilities

AI Quick Q&A | 2026/09/09: With Muse Including Sentinel and Approval, Does It Guarantee Safety from Prompt Injection Attacks?