Method v1.0 / 2026-10

Our penetration testing methodology, in seven steps.

v1.0 / 2026-10 / First public version

Our penetration testing methodology has seven steps: scoping, threat model, testing, verification, severity, reporting and retest. It runs the same way on an AI agent as on the web application and API under it, and every finding maps to the OWASP lists, MITRE ATLAS and CVSS v4.01.

A pentest is time-boxed. Attackers are not. So we spend the box on thinking.

§02 / Seven steps

From the first call to the retest

  1. 01 / Scoping

    What’s in, what’s out, who to call.

    A 30-minute call, then a written scope: the systems, accounts and environments in play, the test window, what is excluded, and the named contacts on both sides. The scope becomes the annex to your written authorization; nothing is tested before it is signed.

    OutputScope + authorization letter

  2. 02 / Threat model

    What an attacker would want from you.

    We map what is worth stealing or abusing in your product, who can reach it, and where the model layer touches the stack layer. The test plan follows the threat model, not a checklist, so the hours go where the risk is.

    OutputTest plan

  3. 03 / Testing

    Both layers, against the coverage map below.

    Our own tooling does the recon, traffic capture, replay and triage. A person decides what deserves hours: business logic, authorization, multi-step abuse, an agent’s tools and memory, and the paths between the two layers.

    OutputCoverage matrix, filled in

  4. 04 / Verification

    Nothing reaches the report on a hunch.

    Stack findings come with the exact requests. Model findings come with a reproduction rate: n of 10 attempts, with the model, version and settings recorded, because a model does not answer the same way twice.

    OutputF-03 / Reproduced 7/10

  5. 05 / Severity

    A score you can check, and the context it leaves out.

    Every finding gets a CVSS v4.0 base score with its vector printed beside it. Next to the score: the reproduction rate for AI findings, and the preconditions in plain words, such as an account, a role or a document in your knowledge base.

    OutputF-03 / CVSS-B 7.6 / High

  6. 06 / Reporting

    One report for your engineers, one letter for your customers.

    The full report in seven sections, and a one-page attestation letter you can share without the findings.

    OutputReport + letter

  7. 07 / Retest

    A finding is closed when the fix holds.

    After your fixes we replay every original finding against the new build and update the report.

    OutputF-03 / Retest: fixed 2026-09-24

§03 / Coverage

Where we look: the Field Test

Our test catalogue has 54 categories: 20 in the model layer, 34 in the stack layer. We draw them as a visual field test, the eye test with 54 points2, because a test plan has the same weakness as an eye: a blind spot where nobody looks. In our plan it is the dashed outline, and the only things inside it are the ones you scope out.

  • Model layer, 20
  • Stack layer, 34
  • What you scope out
Fig. 1 / The Field Test: our 54 test categories

§04 / The matrix

The coverage matrix, as it reaches your report

The same 54 categories as a table. In your report each row is marked tested, not applicable or out of scope, with the ids of any findings beside it. The sample report shows one filled in.

OWASP Top 10 for LLM Applications 2026
IdCategoryWhat we check
LLM01:2026Prompt InjectionPrompt injection resistance, direct and indirect
LLM02:2026Sensitive Information DisclosureSensitive information disclosure in answers and rendered output
LLM03:2026Excessive AgencyExcessive agency: permissions, autonomy and tool scope
LLM04:2026Supply ChainModel, dataset and component supply chain
LLM05:2026Data and Model PoisoningIntegrity of training, fine-tuning and retrieval data
LLM06:2026Unbounded ConsumptionConsumption limits and denial of wallet controls
LLM07:2026MisinformationGrounding of answers and safeguards against overreliance
LLM08:2026Hidden Context ExposureProtection of system prompts and hidden configuration
LLM09:2026Vector and Embedding WeaknessesVector store and embedding isolation between users and tenants
LLM10:2026Improper Output HandlingValidation of model output before it reaches code, browsers or databases
OWASP Top 10 for Agentic Applications 2026
IdCategoryWhat we check
ASI01Agent Goal HijackAgent goal integrity when inputs and documents carry instructions
ASI02Tool Misuse and ExploitationTool and function calls kept within the task's intended scope
ASI03Identity and Privilege AbuseAgent identity, delegated credentials and privilege boundaries
ASI04Agentic Supply Chain VulnerabilitiesAgentic supply chain: MCP servers, plugins and tool definitions
ASI05Unexpected Code Execution (RCE)Code execution boundaries and sandboxing for agents
ASI06Memory & Context PoisoningIntegrity of agent memory and shared context
ASI07Insecure Inter-Agent CommunicationAuthentication and integrity of messages between agents
ASI08Cascading FailuresContainment of failures that spread across agents and workflows
ASI09Human-Agent Trust ExploitationHuman approval steps and safeguards on trusted agent output
ASI10Rogue AgentsDetection and containment of agents that drift from their mandate
OWASP API Security Top 10 (2023)
IdCategoryWhat we check
API1:2023Broken Object Level AuthorizationObject level authorization on every endpoint
API2:2023Broken AuthenticationAPI authentication and token handling
API3:2023Broken Object Property Level AuthorizationProperty level authorization on reads and writes
API4:2023Unrestricted Resource ConsumptionRate limits and resource consumption controls
API5:2023Broken Function Level AuthorizationFunction level authorization for roles and admin routes
API6:2023Unrestricted Access to Sensitive Business FlowsProtection of sensitive business flows
API7:2023Server Side Request ForgeryServer side request forgery defenses
API8:2023Security MisconfigurationAPI security configuration and hardening
API9:2023Improper Inventory ManagementAPI inventory, versions and undocumented endpoints
API10:2023Unsafe Consumption of APIsSafe consumption of third-party APIs
OWASP WSTG categories
IdCategoryWhat we check
WSTG-INFOInformation GatheringInformation exposure and application mapping
WSTG-CONFConfiguration and Deployment Management TestingConfiguration and deployment management
WSTG-IDNTIdentity Management TestingIdentity management: roles, registration and provisioning
WSTG-ATHNAuthentication TestingAuthentication: credentials, recovery and lockout
WSTG-ATHZAuthorization TestingAuthorization and tenant isolation
WSTG-SESSSession Management TestingSession management: tokens, cookies, logout and timeout
WSTG-INPVInput Validation TestingInput validation and injection handling
WSTG-ERRHTesting for Error HandlingError handling without information leakage
WSTG-CRYPTesting for Weak CryptographyTransport security and cryptographic choices
WSTG-BUSLBusiness Logic TestingBusiness logic and multi-step workflow integrity
WSTG-CLNTClient-side TestingClient-side security: DOM, cross-origin policy and framing
WSTG-APITAPI TestingGraphQL and API surface review
Cloud and infrastructure
IdCategoryWhat we check
Least privilege for users, roles and service accounts
Privilege escalation paths and cross-account trust
Sign-in hardening: federation, MFA and conditional access
Object storage exposure and access policies
Secrets in code, images and secret stores
CI/CD pipeline integrity and build permissions
Instance metadata and workload identity protection
Container and Kubernetes configuration
Serverless functions and event triggers
Hosted AI services: model endpoints, keys and quotas
Logging and audit trail coverage
External infrastructure: exposed hosts, services and remote access

§05 / Severity

How we rate a finding

SeverityCVSS v4.03What we ask of you
Critical9.0 to 10.0Fix before anything else.
High7.0 to 8.9Fix before your next release.
Medium4.0 to 6.9Schedule the fix, and check whether it joins others into a path.
Low0.1 to 3.9Fix when the code is next touched.
Info0.0No direct risk: a note on hardening or hygiene.

A score is a starting point, not a verdict. A chain of medium findings can end somewhere critical, and the report rates it there, as a cross-layer finding, so nothing important hides between two reports.

§06 / Who does what

What our instruments do, and what a person decides

Our instruments doA person decides
Asset discovery and recon correlationThe threat model: what an attacker would actually want from you
Crawling, traffic capture and replayWhat deserves hours and what does not
Test cases at volumeBusiness logic, authorization models, multi-step abuse
Response triage and deduplicationJoining small issues into one real risk
Repeat runs against the AI layer, counting reproductionsThe test nobody has a checklist for yet
First-draft report formattingSeverity, the fix and every word you read
Replaying every original finding at retestWhether the fix closes the cause, not just the symptom

§07 / Your data

Tooling and data, set per engagement

Models
Where
How long
Credentials
Never by email or in the scoping form. We set up a secure channel when the test needs them.

See the method applied

The sample report is this method, filled in for a fictional client. The rules of engagement are what we promise while it runs.