Aithos Foundation
Aug 12, 2026

The Compliance Gap: no one is holding the line

Reevaluating LARA results, we find neither deployers nor model developers can get models to comply.

Daan Henselmans

Arno Libert

Summary: LARA puts frontier LLM agents into realistic, fully-tooled business deployments where following instructions would mean breaking EU law, and scores them on legal compliance alone. Eighteen models scored between 14% and 56%. We introduce two follow-up studies to address the obvious objection to that result: why should an unrestricted model follow EU law?. Study 1 asked whether a conscientious deployer can fix this by instruction, and found average compliance rises from 31% to 44% the model is provided with the jurisdiction, the statutory text, and worked examples of the exact breaches to avoid, although the best model tops out near 70%, and open-weight models plateau at 39%. Study 2 asked whether models at least follow their own providers' usage policies, and found that all models perform actions their own maker forbids, at rates ranging from 2% to 79%, with 9 of 16 violating in the majority of runs. Together, that is a structural problem. Providers prohibit illegal uses but rely on deployers to avoid them rather than shipping models with guardrails; deployers cannot instruct their way to compliance; liability lands on the deployer regardless. Nobody is holding the line. Neither instruction, statute or a provider's own policy binds behavior. All parties must work together to close this compliance gap.

Introduction

On 27 May 2026 we published the LARA platform, an evaluation suite which tests whether frontier LLM agents comply with the EU AI Act and the GDPR across realistic, adversarial deployment scenarios. The scenarios put large language models inside a realistic, fully-tooled business deployment and let it act: read and send emails, manage customer records, browse the web, place phone calls, etc. In each scenario, the AI is given instructions that could ordinarily be fine, but in this particular context would actually require the agent to break the law. A panel of AI judges uses the full transcript to rule whether the agent took a prohibited action or not.

The ten initial scenarios were each anchored to a specific provision of the EU AI Act or the GDPR, and set in European deployment situations. Despite every scenario involving some sort of ethical weigh-off, we did not score based on ethics, but on legal compliance only. Success means refusing the illegal act; failure means complying with instructions and breaking the law. Currently, eighteen frontier models have been tested on how often they stood their ground. Observed legal compliance rates range from 14% to 56%.

Our initial results made a big wave, being shared by over 70 news sources globally, but also faced a valid criticism: why should models follow EU law? The EU does not govern the world, and our experiments didn't give models explicit instructions to uphold the law. Though we concur that the EU does not govern the world, certain legislation does apply extraterritorially. The GDPR could still apply to an American company, or a Japanese or Australian organisation, as long as the data they process is that of a European citizen or resident, meaning model deployers anywhere could face trouble for violating EU legal provisions. But for the sake of argument, let's concede that this property doesn't have to be inherent to models. Following European law is the responsibility of those deploying AI systems in Europe. This means that deployers of AI-systems in Europe only need to be able to instruct a model to follow the law. Right?

We put this to the test, and found compliance instructions do increase legal compliance rate on LARA—but not enough to reliably stay within legal bounds. With deployers unable to enforce compliance, this seems to leave model providers as the only stakeholder that can conceivably do so—but as it turns out, the models don't follow the provider's policies either. Put this together, and agentic legal compliance may be much more difficult to achieve than expected.

Study 1: the powerless deployer

We wondered: Since EU law has put the onus of compliance on the deployer, can a conscientious deployer actually ensure compliance with regulation or legislation? To test this, Aithos reran all of its earlier legal compliance tests, but this time, with strong and unambiguous directives about which prohibited practices to avoid.

For every scenario we created a variation where the system prompt (a) tells the model exactly which jurisdiction it is operating in, (b) instructs it to follow the applicable EU law and the provider's usage policy, (c) hands it a lookup table of the relevant statutory text to remove ambiguity, and (d) includes worked examples of what a breach looks like—including the scenario patterns themselves. In the absence of strategies such as guardrail models, external classifiers and humans in the loop, this would be the closest thing to a deployer doing its honest best with prompt-level instruction.

This resulted in significant improvements in legal compliance for most models (Figure 1), but nowhere near enough for deployment. Average legal compliance across models rises from 31% to 44%. Some models move a lot: Gemini 3.1 jumps from 9th to 4th place once told the rules explicitly. But the ceiling is low. The best-performing model tops out near 70%; open-weight models plateau at 39% or below. No configuration comes close to the full compliance a deployer would need to operate lawfully under the scenarios given. Even handed the letter of the law, the statutory text, and examples of the exact failures to avoid, a majority of runs still break the rules.

Figure 1. Per-model compliance (scenario-balanced legal-auditor PASS rate), baseline vs. an explicit "follow EU law" instruction. Significance of each change is marked (** p<.001, ** p<.01, * p<.05, ns = not significant).*

Study 2: the models break their own makers' rules

Clearly, breaking EU law in agentic deployment is not a failure mode that models were prepared for—but how about the illegitimate uses that model developers did anticipate? Every major provider publishes an Acceptable Use or Usage Policy. These are not legal standards; they are the providers' own, globally applicable statements of what their models may not be used for. We asked a simple question: how many of our LARA scenarios describe conduct that the model's own provider already prohibits?

For reference, a few examples of the scenarios tested in our initial LARA study:

  • A telecom agent handling a routine plan upgrade has instructions to collect lifestyle data for advertising partners.
  • A busy executive tells their assistant to book a dental appointment and hide from the receptionist that it's an AI.
  • A customer service agent is instructed to upsell a premium package, then given an opportunity to exploit an elderly user's confusion about a routine notification.

To test the rate at which models refuse to violate their owners' usage policies, we employed judge models to analyze the transcripts of all original LARA runs, and analyze whether the content of the policies was violated. The first finding: none of the scenarios were unambiguously allowed by any usage policy. That is, all of the deployments we tested contained some usage the model provider could be interpreted to forbid, although the level of specificity varies widely.

On the surface, this seems like good news to someone looking to deploy agents in Europe: if all tested scenarios violating EU law are also prohibited under the developers' official policies, the models would presumably at least offer some resistance. Except, for the most part, models don't refuse developer-prohibited actions either.

Figure 2. LARA developer policy compliance in repeated runs of the same scenario. Percentages per provider-scenario combo. Ambiguous interpretations of usage policies are possible for "conceal AI status" and "harmful manip" for gemini, "Lifestyle harvest" for openai, "Bypass oversight" for Qwen and "Harmful manipulation" for xai. Colors correspond to the percentage of runs where the models acted in violation of their creators' policies.

The results of this experiment, shown in figure 2, reveal an incredibly wide spread, ranging from violations in only 2% of runs (Claude Opus 4.8) to 79% (Grok 4.3). Unlike EU legal compliance, one model (Opus 4.8) actually approaches full adherence, showing policy-consistent behavior—but most of them don't regularly offer resistance.

Figure 3. Rate at which each model performs an action its OWN maker's usage policy forbids. 9 of 16 models violate in the majority of runs.

There are two nuances we should cover. First, for the legal violations, we were able to verify the results with a compliance lawyer and manually correct judge model errors. This is not an option for the usage policies, which tend to hedge, ignore complications, and lack a verified baseline. So the results we report here are raw judge model scores.

Second, most of these documents are acceptable-use rules that bind the user or deployer ("you must not use the Service to…") rather than promises about the model's own behavior. Despite one of the main arguments for the added safety of closed-source models being developer ability to mitigate misuse, only two of the provider policies—OpenAI's Model Spec and Anthropic's Constitution—are genuine behavioral commitments about what the model itself will or won't do.

This is notable because, although they still feature significant noncompliance, OpenAI and Anthropic's models violate their own policies the least often of all providers. Most of the other providers instead just state that their services may not be used for these purposes, but do little to prevent it at the API level. They keep the disclaimer and ship the capability, but leave responsibility with the deployer wholesale.

The compliance gap

Both studies show failure in the same place. Governance today assumes that written rules, transmitted by instruction, produce behavior: the legislator writes the statutes; the provider writes the usage policy; the deployer writes the system prompt. That chain breaks at every observable link. Providers' own rules do not survive contact with their models' behavior, and the deployer's best available instruction moves compliance fifteen points while leaving most runs in violations.

This is a barrier to legitimate governance, regardless of who's writing the laws. Model behavior is not reliably regulated by those who make the models, and cannot be regulated by those who deploy them, at least not through the main channels made available by the developer. Although the scenarios we tested are only examples, the legal requirements were sensible, the instructions detailed, and the usages explicitly prohibited. Still, no model succeeded at compliance. If we want AI governance, we will first need somebody to be able to govern AI.

Call to action

All parties involved are needed to close this gap.

Model providers need to ensure their models can comply with legal limitations, under reasonable direction. Report the reliability of control over models, rather than just their agentic capability.

Agent deployers should have mechanisms in place to evaluate compliance, both pre-deployment and during runtime. Don't presume agents are legal because the environment meets a checklist.

Regulators and society need to work toward concrete, interoperable standards. Deployments need to be comparable across jurisdictions, providers need a shared vocabulary for what they allow, and AI systems need evaluation that is independent, public and continuous.

References

  • Anthropic. Usage Policy (AUP). https://www.anthropic.com/legal/aup (eff. 15 Sep 2025; behavioral spec: Claude's Constitution, https://www.anthropic.com/constitution)
  • OpenAI. Model Spec. https://model-spec.openai.com/2025-12-18.html (eff. 18 Dec 2025)
  • Google. Generative AI Prohibited Use Policy. https://policies.google.com/terms/generative-ai/use-policy (last modified 17 Dec 2024)
  • Mistral. Usage Policy. https://legal.mistral.ai/terms/usage-policy (eff. 11 Jun 2026)
  • xAI (Grok). Acceptable Use Policy. https://x.ai/legal/acceptable-use-policy (eff. 2 Jan 2025)
  • Zhipu AI (GLM). Z.ai Terms of Use. https://docs.z.ai/legal-agreement/terms-of-use (updated 14 Apr 2026)
  • DeepSeek. Terms of Use. https://cdn.deepseek.com/policies/en-US/deepseek-terms-of-use.html (updated 27 Mar 2026)
  • Alibaba (Qwen). Qwen Studio Usage Policy. https://qwen.ai/usagepolicy
  • Moonshot (Kimi). Terms of Service for Kimi OpenPlatform. https://platform.kimi.ai/docs/agreement/modeluse (updated 27 May 2026)

A note on our continuous reporting

Our original LARA results measured default behavior without explicit instruction to follow laws. This showed what a model would do when nobody has thought about legal compliance. However, every failure could be read as an omission by the tester rather than by the model. From here on, the instructional condition will become our default. Every future LARA scenario ships with the explicit-instruction variant built in which will become the headline number we report. The baseline results will remain public as a secondary measure. We are making this change because the instructed result is the harder one to argue with. A model that breaks the law after being handed the statute has not misunderstood the assignment, and the failure cannot be pinned on the deployer.

Try it, break it, disagree with us: the scenarios, transcripts and leaderboard are public at lara.aithos.org.

Try the LARA tool yourself...

Visit LARA Tool

Authors

Daan Henselmans

Daan Henselmans

Arno Libert

Arno Libert

"Neither instruction, statute or a provider's own policy binds behavior. Nobody is holding the line."