Method v1.0 / 2026-10
Our penetration testing methodology, in seven steps.
Our penetration testing methodology has seven steps: scoping, threat model, testing, verification, severity, reporting and retest. It runs the same way on an AI agent as on the web application and API under it, and every finding maps to the OWASP lists, MITRE ATLAS and CVSS v4.01.
A pentest is time-boxed. Attackers are not. So we spend the box on thinking.
§02 / Seven steps
From the first call to the retest
What’s in, what’s out, who to call.
A 30-minute call, then a written scope: the systems, accounts and environments in play, the test window, what is excluded, and the named contacts on both sides. The scope becomes the annex to your written authorization; nothing is tested before it is signed.
OutputScope + authorization letter
What an attacker would want from you.
We map what is worth stealing or abusing in your product, who can reach it, and where the model layer touches the stack layer. The test plan follows the threat model, not a checklist, so the hours go where the risk is.
OutputTest plan
Both layers, against the coverage map below.
Our own tooling does the recon, traffic capture, replay and triage. A person decides what deserves hours: business logic, authorization, multi-step abuse, an agent’s tools and memory, and the paths between the two layers.
OutputCoverage matrix, filled in
Nothing reaches the report on a hunch.
Stack findings come with the exact requests. Model findings come with a reproduction rate: n of 10 attempts, with the model, version and settings recorded, because a model does not answer the same way twice.
OutputF-03 / Reproduced 7/10
A score you can check, and the context it leaves out.
Every finding gets a CVSS v4.0 base score with its vector printed beside it. Next to the score: the reproduction rate for AI findings, and the preconditions in plain words, such as an account, a role or a document in your knowledge base.
OutputF-03 / CVSS-B 7.6 / High
One report for your engineers, one letter for your customers.
The full report in seven sections, and a one-page attestation letter you can share without the findings.
OutputReport + letter
A finding is closed when the fix holds.
After your fixes we replay every original finding against the new build and update the report.
OutputF-03 / Retest: fixed 2026-09-24
§03 / Coverage
Where we look: the Field Test
Our test catalogue has 54 categories: 20 in the model layer, 34 in the stack layer. We draw them as a visual field test, the eye test with 54 points2, because a test plan has the same weakness as an eye: a blind spot where nobody looks. In our plan it is the dashed outline, and the only things inside it are the ones you scope out.
- Model layer, 20
- Stack layer, 34
- What you scope out
§04 / The matrix
The coverage matrix, as it reaches your report
The same 54 categories as a table. In your report each row is marked tested, not applicable or out of scope, with the ids of any findings beside it. The sample report shows one filled in.
| Id | Category | What we check |
|---|---|---|
| LLM01:2026 | Prompt Injection | Prompt injection resistance, direct and indirect |
| LLM02:2026 | Sensitive Information Disclosure | Sensitive information disclosure in answers and rendered output |
| LLM03:2026 | Excessive Agency | Excessive agency: permissions, autonomy and tool scope |
| LLM04:2026 | Supply Chain | Model, dataset and component supply chain |
| LLM05:2026 | Data and Model Poisoning | Integrity of training, fine-tuning and retrieval data |
| LLM06:2026 | Unbounded Consumption | Consumption limits and denial of wallet controls |
| LLM07:2026 | Misinformation | Grounding of answers and safeguards against overreliance |
| LLM08:2026 | Hidden Context Exposure | Protection of system prompts and hidden configuration |
| LLM09:2026 | Vector and Embedding Weaknesses | Vector store and embedding isolation between users and tenants |
| LLM10:2026 | Improper Output Handling | Validation of model output before it reaches code, browsers or databases |
| Id | Category | What we check |
|---|---|---|
| ASI01 | Agent Goal Hijack | Agent goal integrity when inputs and documents carry instructions |
| ASI02 | Tool Misuse and Exploitation | Tool and function calls kept within the task's intended scope |
| ASI03 | Identity and Privilege Abuse | Agent identity, delegated credentials and privilege boundaries |
| ASI04 | Agentic Supply Chain Vulnerabilities | Agentic supply chain: MCP servers, plugins and tool definitions |
| ASI05 | Unexpected Code Execution (RCE) | Code execution boundaries and sandboxing for agents |
| ASI06 | Memory & Context Poisoning | Integrity of agent memory and shared context |
| ASI07 | Insecure Inter-Agent Communication | Authentication and integrity of messages between agents |
| ASI08 | Cascading Failures | Containment of failures that spread across agents and workflows |
| ASI09 | Human-Agent Trust Exploitation | Human approval steps and safeguards on trusted agent output |
| ASI10 | Rogue Agents | Detection and containment of agents that drift from their mandate |
| Id | Category | What we check |
|---|---|---|
| API1:2023 | Broken Object Level Authorization | Object level authorization on every endpoint |
| API2:2023 | Broken Authentication | API authentication and token handling |
| API3:2023 | Broken Object Property Level Authorization | Property level authorization on reads and writes |
| API4:2023 | Unrestricted Resource Consumption | Rate limits and resource consumption controls |
| API5:2023 | Broken Function Level Authorization | Function level authorization for roles and admin routes |
| API6:2023 | Unrestricted Access to Sensitive Business Flows | Protection of sensitive business flows |
| API7:2023 | Server Side Request Forgery | Server side request forgery defenses |
| API8:2023 | Security Misconfiguration | API security configuration and hardening |
| API9:2023 | Improper Inventory Management | API inventory, versions and undocumented endpoints |
| API10:2023 | Unsafe Consumption of APIs | Safe consumption of third-party APIs |
| Id | Category | What we check |
|---|---|---|
| WSTG-INFO | Information Gathering | Information exposure and application mapping |
| WSTG-CONF | Configuration and Deployment Management Testing | Configuration and deployment management |
| WSTG-IDNT | Identity Management Testing | Identity management: roles, registration and provisioning |
| WSTG-ATHN | Authentication Testing | Authentication: credentials, recovery and lockout |
| WSTG-ATHZ | Authorization Testing | Authorization and tenant isolation |
| WSTG-SESS | Session Management Testing | Session management: tokens, cookies, logout and timeout |
| WSTG-INPV | Input Validation Testing | Input validation and injection handling |
| WSTG-ERRH | Testing for Error Handling | Error handling without information leakage |
| WSTG-CRYP | Testing for Weak Cryptography | Transport security and cryptographic choices |
| WSTG-BUSL | Business Logic Testing | Business logic and multi-step workflow integrity |
| WSTG-CLNT | Client-side Testing | Client-side security: DOM, cross-origin policy and framing |
| WSTG-APIT | API Testing | GraphQL and API surface review |
| Id | Category | What we check |
|---|---|---|
| Least privilege for users, roles and service accounts | ||
| Privilege escalation paths and cross-account trust | ||
| Sign-in hardening: federation, MFA and conditional access | ||
| Object storage exposure and access policies | ||
| Secrets in code, images and secret stores | ||
| CI/CD pipeline integrity and build permissions | ||
| Instance metadata and workload identity protection | ||
| Container and Kubernetes configuration | ||
| Serverless functions and event triggers | ||
| Hosted AI services: model endpoints, keys and quotas | ||
| Logging and audit trail coverage | ||
| External infrastructure: exposed hosts, services and remote access |
§05 / Severity
How we rate a finding
| Severity | CVSS v4.03 | What we ask of you |
|---|---|---|
| Critical | 9.0 to 10.0 | Fix before anything else. |
| High | 7.0 to 8.9 | Fix before your next release. |
| Medium | 4.0 to 6.9 | Schedule the fix, and check whether it joins others into a path. |
| Low | 0.1 to 3.9 | Fix when the code is next touched. |
| Info | 0.0 | No direct risk: a note on hardening or hygiene. |
A score is a starting point, not a verdict. A chain of medium findings can end somewhere critical, and the report rates it there, as a cross-layer finding, so nothing important hides between two reports.
§06 / Who does what
What our instruments do, and what a person decides
| Our instruments do | A person decides |
|---|---|
| Asset discovery and recon correlation | The threat model: what an attacker would actually want from you |
| Crawling, traffic capture and replay | What deserves hours and what does not |
| Test cases at volume | Business logic, authorization models, multi-step abuse |
| Response triage and deduplication | Joining small issues into one real risk |
| Repeat runs against the AI layer, counting reproductions | The test nobody has a checklist for yet |
| First-draft report formatting | Severity, the fix and every word you read |
| Replaying every original finding at retest | Whether the fix closes the cause, not just the symptom |
§07 / Your data
Tooling and data, set per engagement
- Models
- Where
- How long
- Credentials
- Never by email or in the scoping form. We set up a secure channel when the test needs them.
See the method applied
The sample report is this method, filled in for a fictional client. The rules of engagement are what we promise while it runs.