Compliance / EU AI Act / Article 15

EU AI Act Article 15: security testing for high-risk AI.

Article 15 of the EU AI Act requires high-risk AI systems to reach an appropriate level of accuracy, robustness and cybersecurity, and to resist attempts to alter their use, outputs or performance, from data poisoning to adversarial inputs. For Annex III systems it applies from 2 December 2027. Most customer-service chatbots are not high-risk, but Article 50 already applies to them.

Updated
Sources
06

Article 15, in one table.

FIG. 1 / ARTICLE 15 AT A GLANCE
The act
Regulation (EU) 2024/1689, as amended by the AI Omnibus, Regulation (EU) 2026/1744Note 1
The article
Article 15: accuracy, robustness and cybersecurityNote 2
Binds
High-risk AI systems: Annex III uses, and AI in products regulated under Annex I
Applies from
2 December 2027 for Annex III, 2 August 2028 for Annex INote 3
Testing
Article 9: tested against metrics set in advance, before the system goes on the marketNote 4
Names a pentest
No. It names the attacks a system must resist.
Chatbots
Article 50 disclosure has applied since 2 August 2026Note 5

What does Article 15 require?

Paragraph 1 asks for an appropriate level of accuracy, robustness and cybersecurity, and for consistent performance in those respects throughout the system's lifecycle.Note 6 Paragraph 5 is the security clause:

High-risk AI systems shall be resilient against attempts by unauthorised third parties to alter their use, outputs or performance by exploiting system vulnerabilities.Note 7

It then names the AI-specific attacks the technical solutions must address where appropriate: manipulating the training data (data poisoning), manipulating pre-trained components (model poisoning), inputs designed to make the model err (adversarial examples or model evasion), confidentiality attacks and model flaws.Note 8

Article 15 does not say how to show any of this. Article 9 does say that high-risk systems are tested, against metrics and probabilistic thresholds defined in advance, during development and in any event before they are placed on the market or put into service.Note 9 A security test of the attacks paragraph 5 lists is the natural evidence for both articles.

Which AI systems are high-risk?

Article 15 binds high-risk AI systems only. They come in two groups: systems used for the purposes listed in Annex III, and AI that is a safety component of a product already regulated under the laws in Annex I.Note 10

Annex III reads like a list of decisions about people. Risk assessment and pricing for life and health insurance is on it, for example; a customer-service assistant that answers questions about an order or an open claim is not.Note 11

Classification decides everything else, and it is a legal judgement about your intended purpose. We produce test evidence; whether your system is high-risk is for your counsel.

When does it apply?

The AI Omnibus, Regulation (EU) 2026/1744, moved the high-risk dates when it entered into force on 27 July 2026.Note 12 For Annex III systems the high-risk requirements, Articles 9 and 15 included, now apply from 2 December 2027; for AI in Annex I products from 2 August 2028.Note 13

FIG. 2 / WHEN THE AI ACT ASKS WHAT
  1. The AI Omnibus enters into force and moves the high-risk dates
  2. Article 50: people must be told they are interacting with an AI system
  3. Generative systems already on the market must mark their output (Article 50(2))
  4. Annex III high-risk systems: Articles 9 and 15 apply
  5. High-risk AI in products regulated under Annex I

The dates say when the duties become enforceable, not when testing should start. A system that goes on the market in 2027 needs its Article 9 test results before it does.

What about chatbots that are not high-risk?

Article 50 applies to AI systems that talk to people directly: their providers must make sure people are told they are interacting with an AI system, unless that is obvious from the context.Note 14 It has applied since 2 August 2026, with no grace period for that duty.Note 15

Generative systems that were already on the market before 2 August 2026 have until 2 December 2026 to mark their output as AI-generated in a machine-readable format.Note 16

Article 50 is a transparency duty, not a security test. Yet the attacks Article 15 lists are the ones a chatbot meets every day: an injected instruction is an input designed to make the system misbehave, and a leaked system prompt or customer record is a confidentiality failure. OWASP ranks prompt injection first in its Top 10 for LLM Applications 2026.Note 17 Testing for it is good practice whatever the classification.

What evidence should we keep?

  • A test plan with metrics and thresholds set before testing starts, as Article 9 expects.
  • Results per attack class: which inputs, documents or tool calls were tried, how often each attack reproduced, and on which model version and settings.
  • The fixes, and a retest, so the record shows the system after the change as well as before it.
  • A mapping to the frameworks reviewers already use: the OWASP Top 10 for LLM Applications 2026, the OWASP Top 10 for Agentic Applications and MITRE ATLAS.Note 18

Our AI Pentest reports every model finding with a reproduction rate, such as 7 of 10 attempts, the model and settings it reproduced on, and the framework ids it maps to. Models are probabilistic; thresholds set in advance need numbers, not screenshots.

Test the model, and what it can reach.

Article 15 asks about the AI system, and the system includes the stack behind the model.

01 / Model layer

AI Pentest

Chatbots, RAG and agents: prompt injection, data leakage, poisoning through retrieved content and tool misuse, mapped to OWASP and MITRE ATLAS, with reproduction rates.

02 / Stack layer

Stack Pentest

The APIs, identity, storage and cloud the model can reach. A confidentiality attack usually ends in one of them.

03 / Clearance

Launch Clearance

Shipping AI on a SaaS platform? One scope and one report for the feature and the platform under it, including the paths between them.

AI Act questions we hear before a test.

Do we have to wait until December 2027 to test?

No. The dates say when the duties become enforceable, not when attacks start. Article 9 expects testing before a high-risk system is placed on the market, so a system launching in 2027 needs its evidence before then, and chatbots in production today already face prompt injection and data leakage.

Our chatbot is not high-risk. Should we still test it?

If it reads untrusted content, reaches customer data or calls tools, yes. The law asks less of it, but attackers do not check its classification first. A chatbot and RAG test covers the assistant and its retrieval; the agent package adds tools and actions.

Can you tell us whether our AI system is high-risk?

No. Classification is a legal judgement about your intended purpose and belongs with your counsel. What we can do is describe precisely what the system does, technically, from its inputs to the tools it calls, so your counsel classifies the real system rather than its product description.

What does a model finding look like in your report?

A title, the business impact, the framework id it maps to, a reproduction rate such as 7 of 10 attempts with the model and settings recorded, the fix principle and the retest status. Model behaviour is probabilistic, so a single screenshot is not evidence.

Read next