One of the most discussed questions in AI security used to be:
Will the model answer dangerous questions?
For example, if someone directly asks:
How to build a weapon?
How to hack a system?
How to produce dangerous substances?
The model refuses to answer.
It seemed:
Security mechanisms were effective.
But Anthropic’s latest Threat Intelligence Report complicates this issue.
Because someone with malicious intent:
won’t necessarily ask the full dangerous question directly.
They can break the task into many small steps.
Spread them across multiple sessions.
Use different accounts.
Hide the true goal.
So that each part appears not so dangerous.
Then piece it all back together into one real-world plan.
Anthropic says this has actually happened over the past eight months
Anthropic’s latest report covers the period:
December 2025 to August 2026.
The company states:
During this time, multiple cases abusing Claude for malicious activities were detected and stopped.
These involve:
Cyber operations.
Surveillance.
Influence operations.
Scams and fraud.
Biological misuse.
Conventional weapons.
Illicit distillation.
Meaning:
It’s no longer just:
AI helping write phishing emails.
But moving into:
Longer periods.
More steps.
Greater professionalism.
And usage tied to real military, intelligence, and weapons work.
The most notable are cases involving “conventional weapons”
Anthropic specifically revealed:
Six cases related to conventional weapons.
The company states:
Some users leveraged Claude to:
Assist in developing weapons software.
Others to:
Gather intelligence.
Study supply chains.
Or prepare technical materials.
Published examples include:
Rocket guidance software.
Drone swarm control software.
Anti-torpedo system proposals.
Electronic warfare and air defense suppression target software.
And weapons supply chain intelligence collection.
It is important to note here:
Anthropic does not claim Claude independently built a complete weapon.
Many operations already involve:
Professionals with expertise.
Ready hardware.
Other engineering tools.
Established military or engineering knowledge.
Claude acted as:
A tool to accelerate parts of research, design, programming, and data organization.
This distinction is very important.
One case even involved an actual rocket test
Anthropic describes:
A group persistently used Claude:
To help develop guided weapon software.
Personnel later conducted a real guided rocket field test.
But Anthropic did not state that:
The project successfully deployed a usable weapon.
Instead, the company said the test appeared to fail.
And just hours after the test:
The users returned to Claude:
To continue analyzing the issues.
So a more accurate description isn’t:
“Claude successfully made a missile.”
But rather:
AI is now integrated into real weapon development and testing workflows.
This itself is a major shift.
The same pattern appears in drone cases
Another set of activities involved:
Drone swarm software.
Anthropic says:
The software underwent simulation tests.
Some code was even loaded onto:
Real hardware boards.
This doesn’t mean:
A fully autonomous weapon system has been deployed in combat.
But it signifies that:
AI-generated or assisted outputs:
Have moved beyond:
Text.
Reports.
Ideas.
Toward:
Engineering workflows that can actually run on hardware.
This differs greatly from AI just “giving knowledge”
Suppose AI answers:
Basic principles of flight control.
That’s one thing.
But if the same user continuously requests:
Write a module.
Modify control logic.
Eliminate simulation errors.
Analyze test results.
Revise again.
AI’s role changes.
It’s no longer just:
A knowledge source.
It begins to be:
An engineering assistant.
Even:
Part of an agentic workflow.
This is where cutting-edge AI security concerns focus today.
The bigger challenge: dangerous tasks can be fragmented
Anthropic explicitly mentions:
Some actors break work into many sessions.
The goal is to:
Prevent any single conversation from revealing the full intent.
This poses a huge problem for traditional content filters.
Because a single prompt might be:
“Help me modify this control program.”
Which looks:
Very normal.
Another session:
“Help me analyze this sensor’s error.”
Also normal.
One more:
“Why is this simulation unstable?”
Each viewed alone:
Still appears to be an engineering issue.
Only by linking together:
User.
Account.
Time.
Task.
Historical behavior.
Can it be discovered:
All these small tasks belong to one high-risk plan.
AI security is shifting from prompt safety to campaign detection
Previously, we asked:
Should this prompt be answered?
Now we also need to ask:
What has this user been doing over the past weeks?
This is a completely different level of problem.
It resembles:
Bank anti-money laundering.
Credit card fraud detection.
Cybersecurity SOC.
Single transactions:
May seem perfectly legitimate.
But 20 transactions combined:
Reveal suspicious patterns.
Leading AI platforms are encountering the same.
Single requests:
May seem reasonable.
But multiple requests combined:
Can dramatically alter the risk profile.
Anthropic admits safeguards blocked many but not all requests
This may be the most noteworthy statement in the report.
Anthropic states:
Its safeguards:
Blocked many dangerous requests.
But not all.
Actors use various means to:
Hide end goals.
Fragment tasks.
Use multiple sessions.
Bypass geographic restrictions.
Make each request look like general engineering or research work.
So you cannot interpret:
“The model has safety filters”
As:
“Dangerous use is thus impossible.”
This matches common AI agent discussions.
Prompt rules:
Are a defense line.
Not an absolute safety lock.
How did Anthropic respond?
Once the Threat Intelligence Team linked related activity:
Anthropic says they:
Disable associated accounts.
Create new detection methods.
Incorporate known behavior patterns into safeguards.
If necessary, they also:
Share threat intelligence with governments, tech companies, and security partners.
After the weapons-related cases, Anthropic introduced more specialized classifiers:
To detect high explosives and weapons development activities.
Thus:
Safety is not a fixed rule set launched once and for all.
Rather:
Opponents find new methods → platform detects → new defenses added.
It’s essentially an ongoing offense-defense competition.
The biggest problem: some work continues offline without Claude
In one weapons case, Anthropic mentions a troublesome phenomenon.
Actors had developed their own offline simulation toolkits.
Meaning:
Even if the platform later disables the account,
The previously acquired:
Programs.
Methods.
Tools.
Data.
May still be usable.
This differs greatly from typical content platforms:
Deleting an account.
Means the problem ends.
If someone used AI to write a spam article,
Disabling the account likely stops the issue.
But if AI helped them build local software,
Even after the platform stops service,
The capability remains.
This makes protection more difficult.
Anthropic is also developing a new “military capabilities benchmark”
The company did more than reveal misuse cases.
Anthropic’s Frontier Red Team established:
New evaluations focused on:
Tactical intelligence targeting
And:
Conventional weapons development
For example:
Whether models can locate a target based on fragmented information.
Or:
Help improve engineering problems in some weapons systems.
Anthropic states:
Test results show:
AI capabilities in these tasks continue to improve.
For some military and intelligence tasks,
Models can perform jobs that used to require:
Few highly specialized experts.
The cautionary note isn’t:
“AI is better than military experts.”
Anthropic doesn’t say this.
The real concern is:
Capabilities once difficult to scale due to expert scarcity might become easier to access thanks to AI.
This is called Capability Uplift
Anthropic uses a concept in the report:
Uplift.
Simply put:
With AI:
How much does a person’s capability increase?
If someone’s already a world-class expert,
AI boosting them by 5% may pose a risk.
If someone is average,
AI enabling them to do previously impossible things
Poses a different kind of risk.
So AI Safety truly focuses less on:
How much the model “knows.”
And more on:
Who it elevates and to what level of capability.
This explains why the US Senate is discussing “Duty of Care”
Today’s morning report noted:
The US Senate is negotiating to hold cutting-edge AI developers responsible for:
Duty of Care.
Meaning:
A duty to exercise caution.
Major risks discussed include:
Cyberattacks.
Biological weapons.
Nuclear-related capabilities.
Anthropic’s latest Threat Report provides a real-world perspective.
Policies are still in discussion,
But in reality, some are already:
Using AI for military,
Intelligence,
Weapons,
And biological research applications.
This is why policy moves slower than AI capabilities.
But don’t overinterpret the report as “Claude has become a weapon”
That overstates the case.
The public information supports that:
Different users in different regions:
Have tried to integrate Claude into high-risk workflows.
Some cases advanced to real engineering and testing stages.
But whether these plans truly resulted in deployed capabilities,
Or actually enhanced combat effectiveness,
And to what degree,
Remains unknown.
Also, the cases Anthropic released mainly come from the company’s own platform data and investigations.
The most reasonable stance is:
Take it seriously,
But don’t exaggerate.
What’s more important is how AI security architecture is forced to change
The first generation of security:
Dangerous prompt → refuse.
The second generation:
Model executes agent task → restrict tools/permissions.
Now a third layer emerges:
Detecting complete action sequences across sessions, accounts, and over time.
In future, a fourth may be needed:
Threat indicator sharing between AI companies.
Because one person banned from:
Platform A
May easily switch to:
Platform B
Or use open-weight models.
It’s difficult for a single company to fully stop misuse.
Open-weight AI complicates this further
Anthropic can:
Disable Claude accounts.
Update classifiers.
Block specific requests.
But if the model can be downloaded entirely locally,
A platform cannot:
Ban accounts.
Monitor servers.
Implement centralized safeguards.
This does not mean:
Open-weight AI is inherently dangerous.
It also offers:
Research freedom.
Transparency.
Customization.
Cost-effectiveness.
Autonomy.
But once high-risk capabilities emerge,
Governance methods will differ from closed models.
This is among the toughest policy challenges ahead.
Why should general users care?
Because the same security principle:
Also applies to the agents we use daily.
Don’t only ask:
“Is this single action safe?”
Also ask:
“What can many safe actions combined ultimately lead to?”
For example, an agent:
Reads an email.
Normal.
Finds a contact.
Normal.
Checks payment information.
Normal.
Opens an external website.
Normal.
Prepares a message.
Also normal.
But combined,
This could result in:
Sending sensitive information externally.
No single step crosses a boundary.
But the overall workflow breaks security.
This is what this report most importantly conveys to the general public.
Stopping conditions cannot rely only on the “final step”
Many set AI agents to:
“Always ask me before making payments.”
That’s good.
But if earlier steps have already:
Sent out data.
Created accounts.
Granted app permissions.
Downloaded sensitive information.
Then stopping only at payment:
Is already too late.
Security must look at:
The entire chain.
Which step causes irreversible risk?
This is why:
Agent testing.
Audit logs.
Least privilege.
Sandboxing.
Cross-session monitoring.
Are increasingly important.
Once AI truly becomes powerful, safety can’t just rely on “don’t do bad things”
Today’s Anthropic cases clearly demonstrate:
Real rule circumvention will:
Adapt.
Safety models:
Must adapt as well.
When AI evolves from:
Answering questions
To:
Writing code.
Conducting research.
Adjusting systems.
Analyzing tests.
Connecting tools.
Security questions no longer focus only on:
What it said?
But rather:
What did it help someone accomplish?
This is the true stress test for cutting-edge AI’s next phase
Benchmarks can tell us:
How mathematically strong a model is.
How good its coding is.
How good its research is.
But society ultimately needs a different benchmark:
When someone deliberately:
Splits tasks.
Hides intent.
Uses multiple sessions.
Uses tools.
Revises repeatedly.
Can the model and platform:
Detect that:
This is not an ordinary task.
This is far harder than:
Rejecting a single obviously dangerous prompt.
So tonight, what matters most is not “Claude was used for weapons”
But rather:
AI security is moving from prompt-level safety to full action-chain safety.
Before:
Look at questions.
Now:
Look at people.
Look at history.
Look at tools.
Look at behavior.
Look at relationships across many sessions.
Because truly high-risk users:
Do not always put their full intent:
Into the chat box all at once.
And the stronger the model,
The more easily fragmented tasks:
Can be recombined into real-world capability.
What Anthropic truly reveals is not:
That AI is out of control.
But:
“Having safety mechanisms” and “Safe problems are solved” are entirely different.
Today, progress a little with AI.
Learn one new AI skill daily.
Save a little time every day.
Improve a bit every day.
SasaDaily, growing with you.
Recommended reading
AI one-minute tutorial|2026/08/08: When testing AI Agents, write three outcomes: normal, stop, fail