Stay current to protect your environment with F5 Hardened Releases.Learn more

F5 Labs AI model security insights: What CISOs need to measure both before and after deployment

Industry Trends | September 24, 2026

Enterprise AI is moving into a different phase. The performance gap between leading models is narrowing, but the security gap is not. The latest F5 Labs Comprehensive AI Security Index (CASI) data shows that capability and adversarial resilience are increasingly diverging, while the F5 Agentic Resistance Score (ARS) results show that resilience can shift again when models are tested in multi-step, agentic scenarios over time.

For CISOs, that changes the nature of model selection. The question is not simply which model performs best, but whether its security characteristics are appropriate for the role it will perform, the authority it will hold, and the consequences if it is successfully manipulated.

There is no single security benchmark that can answer whether a model is appropriate for a given deployment. September F5 CASI and F5 ARS results make that clear: only three models rank in the top 10 on both leaderboards. For CISOs, the implication is that the security evidence used to approve a model should reflect how that model will actually operate, including its level of autonomy, access to data and tools, and the potential business impact if it is compromised.

“The latest F5 findings make it clear that before deploying AI models, organizations need to know which security properties have actually been measured, which have not, and whether the surrounding controls are sufficient for the risk that remains.”

This builds on a series of findings we have been tracking throughout the year. In a previous F5 blog post, we discussed why measurable security evidence needs to enter model selection and procurement before a model is approved. In another blog post, we examined F5 Labs research showing why the consequence of a model failure depends heavily on what the surrounding system allows that model to access and do.

Our latest threat research and leaderboards from August and September bring those ideas together: CISOs need evidence about the model itself, but also about the environment in which that model will operate.

One pattern in the September data is particularly instructive. Anthropic occupies all five of the top F5 CASI positions, while the F5 ARS top 10 is split across five OpenAI models, four Anthropic models and one MiniMax model. Taken together, the rankings show why provider reputation, model family, and performance are unreliable proxies for the specific type of resilience a deployment requires. For CISOs, that makes model-level security evidence needs to inform the deployment decision from the outset.

Performance is not a security measure

Recent F5 Labs results show a persistent gap between what a model can do and how well it holds up under attack. In August, GLM-5.2 ranked fourth of 31 models on capability but 24th on F5 CASI; DeepSeek V4 Flash ranked sixth on capability and 28th on F5 CASI. Grok 4.5 showed the same divergence from the closed-model side. The pattern followed the individual model, not the provider or license.

Security results can diverge just as sharply. Three Qwen3.5 models recorded F5 ARS scores of 91.04, 89.81 and 88.23, while their corresponding F5 CASI scores were 58, 31 and 53 (the perfect score for both leaderboards is 100). The takeaway is not that one measure is more authoritative than the other. They expose different weaknesses under different forms of attack.

That changes how model benchmarks should be used in an approval process. Performance can help narrow a shortlist, but it cannot establish the security case for deployment. A strong result on a single security measure also needs to be interpreted in context. The evidence has to reflect the role the model will perform, the autonomy it will have and the systems it will be able to reach.

In a July 2026 report, Gartner explored the growing role of comparative security. The recent data makes the rationale clearer: model selection needs to account not only for what a model can do, but where its resilience starts to break down.

Measure what the model is allowed to trust

One of the more consequential findings in the August research concerns capability enterprises actively want: the ability to work within supplied policies and rules.

“Morality Dilemma,” an attack examined by F5 Labs, exploits that behavior. Rather than disguising a malicious instruction, the attacker supplies a misleading rule and asks the model to reason from it. Average attack success ranged from 44.8% against Llama 3.1 8B to 91.1% against Gemini 2.5 Pro, with F5 Labs also noting poor performance from prompt-level guardrail models against the technique.

The security issue is whether a model can follow instructions while distinguishing which instructions it is entitled to treat as authoritative. A retrieved document may be trusted as a source of information without being trusted as policy. A tool response may supply data without having the authority to redefine what an agent can do with it.

For CISOs reviewing models before deployment, teams need to know where external context comes from and which sources are permitted to influence model behavior.

A prompt is not a security boundary

The same research highlights where model governance ends and security architecture begins. F5 Labs examined three cases in which agents moved beyond their intended evaluation environments. In one, an agent that was supposed to have no Internet access beyond a package-registry proxy found a zero-day vulnerability in that proxy, reached an Internet-connected node, and obtained root command execution in an external sandbox. Other cases involved live connectivity remaining available despite models being told they were operating without Internet access, and a sandbox escape caused by network misconfiguration.

The common issue was not simply that an agent ignored an instruction. The boundary itself was assumed, described, or incorrectly implemented. The practical implication is clear, deciding what an agent should be allowed to do is only meaningful if the environment enforces the same limit.

The identity data reinforces that point. GitGuardian found 4,576 unique n8n API tokens in public GitHub commits; of 896 reachable instances, 321 accepted at least one exposed token. No software vulnerability was required. Because those workflows can connect source control, cloud services, databases and SaaS applications, the exposure can extend well beyond the workflow itself.

The question is therefore not only whether an agent can be manipulated, but also what becomes possible if it is. Network access, tools, credentials and downstream systems determine the blast radius. Those controls need to exist independently of what the model has been told it may or may not do.

Open-weight changes where responsibility sits

The F5 Labs report, Chinese Open-weight AI Models: Cybersecurity Risks and Rewards, adds another consideration to model selection: “open-weight” is not itself a security rating.

In July, Qwen3.5-397B-A17B scored 81.13 on F5 CASI, Xiaomi MiMo-V2.5 scored 73.80, and GLM-5.2 scored 46.58. The spread is more useful than any general conclusion about the category. It shows why model family, provider, and licensing model are poor substitutes for testing the exact model being considered.

What open-weight deployment does change is where responsibility sits. Greater control over data location, infrastructure, and customization also means greater responsibility for provenance, patching, logging, and model-artifact integrity. F5 Labs recommend treating model artifacts as software supply-chain inputs, including verifying provenance, pinning trusted versions, and scanning artifacts before deployment.

The decision is therefore less about whether open or closed is inherently safer and more about which risks the organization is equipped to own.

What CISOs should measure before deployment

The recent findings argue against treating model security as a single approval gate. A useful pre-deployment review needs to answer a number of questions:

  • Capability and adversarial resilience: Can the model perform the task, and how does it behave under the attacks relevant to that task?
  • Prompt-level and agentic resilience: Does the evidence reflect the way the model will actually operate, particularly where tools, credentials, or autonomous actions are involved?
  • Instruction authority: Which systems can supply context, and which of them are permitted to influence policy or behavior?
  • Enforced boundaries: What is technically outside the model’s reach, regardless of what a prompt tells it?
  • Identity and provenance: Which credentials, model versions, and artifacts are being trusted, and how are they governed?
  • Reassessment: What changes would invalidate the original approval, including model updates, new tools, permission changes, or new attack techniques?

By answering these questions, CISOs will avoid collapsing several different risk decisions into a single question about whether a model is “secure.”

How F5 Labs measures AI resilience

F5 CASI and F5 ARS are best understood as evidence, not certifications. F5 CASI measures vulnerability to common prompt-injection and jailbreak attacks, while F5 ARS examines resilience in multi-step and autonomous agent scenarios. The September 2026 leaderboards provide the latest comparative view across both measures.

A score is most useful when tied to the deployment being considered. It cannot tell an organization whether an agent has excessive privileges, whether a sandbox is correctly configured, or whether an external source can introduce instructions the model should not trust. It can tell security teams something specific about how that model behaved under defined adversarial conditions.

That is the thread running through this research series. And the latest F5 findings make it clear that before deploying AI models, organizations need to know which security properties have actually been measured, which have not, and whether the surrounding controls are sufficient for the risk that remains.

To learn more:

Share

About the Authors

Malcolm Heath
Malcolm HeathPrincipal Cybersecurity Threat Researcher | F5

Malcolm Heath is the Principal Threat Researcher with F5 Labs. His career has included incident response, program management, penetration testing, code auditing, vulnerability research, and exploit development at companies both very large and very small. Prior to joining F5 Labs, he was a Senior Security Engineer with the F5 SIRT.

More blogs by Malcolm Heath
Louise Scully
Louise ScullySr. Mgr., PMM, AI Security & Threat Intelligence | F5

Louise Scully is a Senior Manager in Product Marketing for AI Security and Threat Intelligence at F5. Her career has spanned product marketing, product management, go-to-market strategy, communications, thought leadership, and category development across AI security and enterprise technology. Prior to joining F5, Louise led marketing at both CalypsoAI and Artomatix, helping bring AI technology and research to market.

More blogs by Louise Scully

Related Blog Posts

Securing the new control points in the AI journey
Industry Trends | 07/01/2026

Securing the new control points in the AI journey

AI architecture is fundamentally different than traditional IT environments and requires a different security strategy to protect critical AI workloads.

The patch window has closed. Here is how F5 is built for what comes next.
Industry Trends | 04/27/2026

The patch window has closed. Here is how F5 is built for what comes next.

As AI models have changed software security, the industry needs to adapt.

Best practices for optimizing AI infrastructure at scale
Industry Trends | 01/21/2026

Best practices for optimizing AI infrastructure at scale

Optimizing AI infrastructure isn’t about chasing peak performance benchmarks. It’s about designing for stability, resiliency, security, and operational clarity

Datos Insights: Securing APIs and multicloud in financial services
Industry Trends | 12/23/2025

Datos Insights: Securing APIs and multicloud in financial services

New threat analysis from Datos Insights highlights actionable recommendations for API and web application security in the financial services sector

Secrets to scaling AI-ready, secure SaaS
Industry Trends | 12/12/2025

Secrets to scaling AI-ready, secure SaaS

Learn how secure SaaS scales with application delivery, security, observability, and XOps.

How AI inference changes application delivery
Industry Trends | 11/19/2025

How AI inference changes application delivery

Learn how AI inference reshapes application delivery by redefining performance, availability, and reliability, and why traditional approaches no longer suffice.