AI Pentest
Chatbots, RAG and agents: prompt injection, data leakage, poisoning through retrieved content and tool misuse, mapped to OWASP and MITRE ATLAS, with reproduction rates.
Compliance / EU AI Act / Article 15
Article 15 of the EU AI Act requires high-risk AI systems to reach an appropriate level of accuracy, robustness and cybersecurity, and to resist attempts to alter their use, outputs or performance, from data poisoning to adversarial inputs. For Annex III systems it applies from 2 December 2027. Most customer-service chatbots are not high-risk, but Article 50 already applies to them.
Paragraph 1 asks for an appropriate level of accuracy, robustness and cybersecurity, and for consistent performance in those respects throughout the system's lifecycle.Note 6 Paragraph 5 is the security clause:
High-risk AI systems shall be resilient against attempts by unauthorised third parties to alter their use, outputs or performance by exploiting system vulnerabilities.Note 7
It then names the AI-specific attacks the technical solutions must address where appropriate: manipulating the training data (data poisoning), manipulating pre-trained components (model poisoning), inputs designed to make the model err (adversarial examples or model evasion), confidentiality attacks and model flaws.Note 8
Article 15 does not say how to show any of this. Article 9 does say that high-risk systems are tested, against metrics and probabilistic thresholds defined in advance, during development and in any event before they are placed on the market or put into service.Note 9 A security test of the attacks paragraph 5 lists is the natural evidence for both articles.
Article 15 binds high-risk AI systems only. They come in two groups: systems used for the purposes listed in Annex III, and AI that is a safety component of a product already regulated under the laws in Annex I.Note 10
Annex III reads like a list of decisions about people. Risk assessment and pricing for life and health insurance is on it, for example; a customer-service assistant that answers questions about an order or an open claim is not.Note 11
Classification decides everything else, and it is a legal judgement about your intended purpose. We produce test evidence; whether your system is high-risk is for your counsel.
The AI Omnibus, Regulation (EU) 2026/1744, moved the high-risk dates when it entered into force on 27 July 2026.Note 12 For Annex III systems the high-risk requirements, Articles 9 and 15 included, now apply from 2 December 2027; for AI in Annex I products from 2 August 2028.Note 13
The dates say when the duties become enforceable, not when testing should start. A system that goes on the market in 2027 needs its Article 9 test results before it does.
Article 50 applies to AI systems that talk to people directly: their providers must make sure people are told they are interacting with an AI system, unless that is obvious from the context.Note 14 It has applied since 2 August 2026, with no grace period for that duty.Note 15
Generative systems that were already on the market before 2 August 2026 have until 2 December 2026 to mark their output as AI-generated in a machine-readable format.Note 16
Article 50 is a transparency duty, not a security test. Yet the attacks Article 15 lists are the ones a chatbot meets every day: an injected instruction is an input designed to make the system misbehave, and a leaked system prompt or customer record is a confidentiality failure. OWASP ranks prompt injection first in its Top 10 for LLM Applications 2026.Note 17 Testing for it is good practice whatever the classification.
Our AI Pentest reports every model finding with a reproduction rate, such as 7 of 10 attempts, the model and settings it reproduced on, and the framework ids it maps to. Models are probabilistic; thresholds set in advance need numbers, not screenshots.
Article 15 asks about the AI system, and the system includes the stack behind the model.
Chatbots, RAG and agents: prompt injection, data leakage, poisoning through retrieved content and tool misuse, mapped to OWASP and MITRE ATLAS, with reproduction rates.
The APIs, identity, storage and cloud the model can reach. A confidentiality attack usually ends in one of them.
Shipping AI on a SaaS platform? One scope and one report for the feature and the platform under it, including the paths between them.
No. The dates say when the duties become enforceable, not when attacks start. Article 9 expects testing before a high-risk system is placed on the market, so a system launching in 2027 needs its evidence before then, and chatbots in production today already face prompt injection and data leakage.
If it reads untrusted content, reaches customer data or calls tools, yes. The law asks less of it, but attackers do not check its classification first. A chatbot and RAG test covers the assistant and its retrieval; the agent package adds tools and actions.
No. Classification is a legal judgement about your intended purpose and belongs with your counsel. What we can do is describe precisely what the system does, technically, from its inputs to the tools it calls, so your counsel classifies the real system rather than its product description.
A title, the business impact, the framework id it maps to, a reproduction rate such as 7 of 10 attempts with the model and settings recorded, the fix principle and the retest status. Model behaviour is probabilistic, so a single screenshot is not evidence.
The first risk on the OWASP list, and the input Article 15 asks a system to resist.
The ten risk classes our AI findings map to.
What Article 21 asks of the systems around the model.