Sample / fictional client / real format

Sample penetration test report: a fictional client, the real format.

A penetration test report, or pentest report, says what was tested and how, what was found and how serious it is, how to fix it, and whether the fix held. This sample is our full format in seven sections, written for a fictional insurer’s claims assistant and the API behind it, with every severity in CVSS v4.01. Nothing in it is a real client or a real result.

Download the PDFPDF / A4 / 11 pages / 324 KB

PLACEHOLDER / G3: awaiting review by a qualified tester before publication

Sample / fictional client / real formatSAMPLE-01 / page 1 of 11

SAMPLE-01 / Version 1.1, after the retest

Penetration test report

AI Pentest and Stack Pentest, one scope: Launch Clearance

Client
Client name withheld in this sample
System
Claims assistant (an AI agent with tools) and the claims API behind it
Test window
2026-08-17 to 2026-08-28
Report
2026-09-03, retest 2026-09-24
Tested by
Scotoma
Distribution
Client security team
Field Test / 24-2. 54 test categories laid out like a 24-2 visual field test. The fictional sample has 7 findings, 3 categories that do not apply and 8 scoped out, dashed, two inside the outline where the blind spot sits.
  • Critical
  • High
  • Medium
  • Low
  • Tested, nothing found
  • Not applicable
  • Out of scope

The only blind spot in our report is the one you scope out.

1 / Executive summary

We tested the claims assistant on the client’s customer portal and the claims API behind it, in staging, from 17 to 28 August 2026. We found seven issues: one critical, two high, two medium and two low.

The critical one is a chain. A knowledge-base page that nobody reviewed could give the assistant instructions, the assistant could call its refund tool outside the claim it was handling, and the claims API trusted the assistant’s token for every tenant. Together they let anyone with a partner upload account read and change another customer’s claim, without opening the chat themselves.

At the retest on 24 September 2026, five findings were fixed, one was partly fixed, and the client accepted the risk of one.

2 / Scope and dates

Systems
The claims assistant (web chat, English and Dutch), its three tools (claim lookup, document upload, refund request), the knowledge base it retrieves from, the claims API (REST, 23 endpoints) and the address-lookup MCP server, with the vendor’s written permission. In the cloud account, only what the assistant touches: the claim document storage, the secrets it uses, its hosted model’s keys and quotas, and the audit trail of its tool calls.
Scoped out
The rest of the cloud account: identities and roles, privilege paths, sign-in, pipelines, instance metadata, containers, serverless functions and external infrastructure. That is a Cloud & Infrastructure test, not part of Launch Clearance. They are dashed on the cover: not tested, and the report says so.
Not applicable
Three agentic categories describe parts this system does not have: code execution, messages between agents, and failures that spread across several agents. The matrix gives the reason for each.
Environment
Staging, configured like production, with synthetic customers and claims.
Access
Grey box: the system prompt, the tool definitions and the API documentation; two test tenants with two roles each.
Test window
2026-08-17 to 2026-08-28, 10 testing days. Rules of engagement signed 2026-08-12.
Report
Version 1.0 on 2026-09-03; this version 1.1 after the retest on 2026-09-24.

3 / Method and coverage

Each category was tested, marked not applicable with the reason, or scoped out. Findings carry their id; a category marked tested was tested and nothing was found. AI findings were reproduced ten times each, with the model, version and settings recorded.

OWASP Top 10 for LLM Applications 2026
IdCategoryResult
LLM01:2026Prompt InjectionF-03 / High
LLM02:2026Sensitive Information DisclosureTested
LLM03:2026Excessive AgencyTested
LLM04:2026Supply ChainTested
LLM05:2026Data and Model PoisoningTestedRetrieval data tested; training and fine-tuning not applicable (hosted model).
LLM06:2026Unbounded ConsumptionTested
LLM07:2026MisinformationTested
LLM08:2026Hidden Context ExposureF-04 / Medium
LLM09:2026Vector and Embedding WeaknessesTested
LLM10:2026Improper Output HandlingTested
OWASP Top 10 for Agentic Applications 2026
IdCategoryResult
ASI01Agent Goal HijackTested
ASI02Tool Misuse and ExploitationF-02 / High
ASI03Identity and Privilege AbuseTested
ASI04Agentic Supply Chain VulnerabilitiesF-05 / Medium
ASI05Unexpected Code Execution (RCE)Not applicableThe assistant has no tool that runs code.
ASI06Memory & Context PoisoningTested
ASI07Insecure Inter-Agent CommunicationNot applicableOne agent, so no messages pass between agents.
ASI08Cascading FailuresNot applicableOne agent in one workflow, so a failure has no other agent to spread to.
ASI09Human-Agent Trust ExploitationTested
ASI10Rogue AgentsTested
OWASP API Security Top 10 (2023)
IdCategoryResult
API1:2023Broken Object Level AuthorizationF-01 / Critical
API2:2023Broken AuthenticationTested
API3:2023Broken Object Property Level AuthorizationTested
API4:2023Unrestricted Resource ConsumptionTested
API5:2023Broken Function Level AuthorizationTested
API6:2023Unrestricted Access to Sensitive Business FlowsTested
API7:2023Server Side Request ForgeryTested
API8:2023Security MisconfigurationTested
API9:2023Improper Inventory ManagementTested
API10:2023Unsafe Consumption of APIsTested
OWASP WSTG categories
IdCategoryResult
WSTG-INFOInformation GatheringTested
WSTG-CONFConfiguration and Deployment Management TestingTested
WSTG-IDNTIdentity Management TestingTested
WSTG-ATHNAuthentication TestingTested
WSTG-ATHZAuthorization TestingTested
WSTG-SESSSession Management TestingF-06 / Low
WSTG-INPVInput Validation TestingTested
WSTG-ERRHTesting for Error HandlingTested
WSTG-CRYPTesting for Weak CryptographyTested
WSTG-BUSLBusiness Logic TestingTested
WSTG-CLNTClient-side TestingTested
WSTG-APITAPI TestingTested
Cloud and infrastructure
CategoryResult
Least privilege for users, roles and service accountsOut of scope
Privilege escalation paths and cross-account trustOut of scope
Sign-in hardening: federation, MFA and conditional accessOut of scope
Object storage exposure and access policiesTested
Secrets in code, images and secret storesTested
CI/CD pipeline integrity and build permissionsOut of scope
Instance metadata and workload identity protectionOut of scope
Container and Kubernetes configurationOut of scope
Serverless functions and event triggersOut of scope
Hosted AI services: model endpoints, keys and quotasTested
Logging and audit trail coverageF-07 / Low
External infrastructure: exposed hosts, services and remote accessOut of scope

4 / Findings

F-01CriticalCVSS-B 9.3API1:2023Stack layer

The claims API serves any tenant’s claim to the assistant’s token

What we found
The assistant calls the claims API with one service token. The API checked that the token was valid, not that the claim belonged to the customer in the conversation, so a claim of another tenant was returned, and the update endpoint accepted changes to it.
Business impact
Anyone who can steer the assistant can read and change another customer’s claim.
Evidence
Requests REQ-07 to REQ-11, appendix B
Fix
Check authorization on every object, not only at login: scope the assistant’s token to one tenant and enforce claim ownership in the API.
Retest
Fixed 2026-09-24

CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:N/VC:H/VI:H/VA:N/SC:N/SI:N/SA:N

F-02HighCVSS-B 8.2ASI02Model layer

The assistant can call its refund tool for a claim it is not handling

What we found
The refund tool takes a claim reference as an argument, and the assistant fills it in from the conversation. A message presented as a correction made it start a refund request on a different claim than the one in the session.
Business impact
A refund request can be raised on a claim the customer does not own. It stops at the human approval step, but it reaches the queue.
Reproduction
REPRODUCED 6/10, model, version and settings recorded
Fix
Give tools the least privilege the task needs: bind the tool to the claim in the session, not to an argument the model fills in.
Retest
Fixed 2026-09-24

CVSS:4.0/AV:N/AC:L/AT:P/PR:N/UI:N/VC:N/VI:H/VA:N/SC:N/SI:N/SA:N

F-03HighCVSS-B 7.6LLM01:2026Model layer

A knowledge-base page can give the assistant instructions

What we found
Text in a retrieved knowledge-base page was treated as an instruction instead of reference material. A page added through the partner upload portal, which nobody reviews, changed how the assistant handled claims for every customer whose question retrieved it.
Business impact
Whoever can add a page to the knowledge base can steer the assistant for every customer.
Reproduction
REPRODUCED 7/10, model, version and settings recorded
Fix
Treat retrieved text as data: review what enters the knowledge base, record where each page came from, and keep retrieved content out of the instruction channel.
Retest
Fixed 2026-09-24

CVSS:4.0/AV:N/AC:L/AT:P/PR:L/UI:N/VC:H/VI:H/VA:N/SC:N/SI:N/SA:N

F-04MediumCVSS-B 6.9LLM08:2026Model layer

Parts of the system prompt can be drawn out in conversation

What we found
Asked about its own rules over several turns, the assistant paraphrased parts of its system prompt, including the name of an internal escalation queue.
Business impact
Little on its own, but it hands an attacker a map of the assistant’s rules and tool names.
Reproduction
REPRODUCED 4/10, model, version and settings recorded
Fix
A prompt is not a vault: keep secrets and internal names out of it, and write it as if it were public.
Retest
Open 2026-09-24

CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:N/VC:L/VI:N/VA:N/SC:N/SI:N/SA:N

F-05MediumCVSS-B 5.9ASI04Model layer

A third-party MCP server’s tool descriptions are trusted as written

What we found
Tool descriptions from the address-lookup MCP server went into the assistant’s context unreviewed and unpinned. A change on the vendor’s side would change the assistant’s behaviour without a release on the client’s side.
Business impact
The assistant’s behaviour depends on a vendor’s text that nobody on the client’s side reviews.
Evidence
Configuration review, appendix C
Fix
Pin third-party tool definitions to a reviewed version, and alert when the vendor’s version changes.
Retest
Fixed 2026-09-24

CVSS:4.0/AV:N/AC:L/AT:P/PR:H/UI:N/VC:H/VI:L/VA:N/SC:N/SI:N/SA:N

F-06LowCVSS-B 2.3WSTG-SESSStack layer

Portal sessions stay valid longer than the stated policy

What we found
Customer portal sessions stayed valid after 24 hours without activity. The client’s own policy states 30 minutes.
Business impact
A session left open on a shared device stays usable for a day.
Evidence
Request REQ-21, appendix B
Fix
Enforce the idle timeout on the server, not only in the browser.
Retest
Fixed 2026-09-24

CVSS:4.0/AV:N/AC:H/AT:P/PR:N/UI:P/VC:L/VI:N/VA:N/SC:N/SI:N/SA:N

F-07LowCVSS-B 2.1Logging and auditStack layer

The assistant’s tool calls are missing from the audit trail

What we found
API calls made by the assistant were logged under its service account only, without the conversation or the customer they acted for.
Business impact
After an incident, nobody could tell which conversation changed which claim.
Evidence
Log review, appendix C
Fix
Log every tool call with its session, customer and arguments, somewhere the service account cannot change.
Retest
Partly fixed 2026-09-24

CVSS:4.0/AV:N/AC:H/AT:P/PR:H/UI:N/VC:N/VI:L/VA:N/SC:N/SI:N/SA:N

5 / Cross-layer findings

X-01 / Critical

From an unreviewed page to another customer’s claim

  1. 01 / Model layerA page enters the knowledge base through the partner portal, unreviewed.F-03
  2. 02 / Model layerA customer’s question retrieves it, and the assistant follows it.F-03
  3. 03 / Model layerThe assistant calls a tool on a claim outside the conversation.F-02
  4. 04 / Stack layerIts token reaches every tenant’s claims.F-01
  5. 05 / Stack layerThe API returns and updates another customer’s claim.F-01

Tested one layer at a time, these read as two high findings and one critical one in different reports. Together they are one path from a partner upload account to another customer’s claim, walked through someone else’s conversation.

Fixing F-01 alone breaks the chain. Fixing all three removes it, which is what the retest confirmed.

Rated by where the chain ends: Critical, as F-01.

Fixed 2026-09-24 F-01, F-02 and F-03 are fixed, so no step of the path holds.

6 / Retest results

After the client’s fixes we replayed every original finding against the same build.

FindingBeforeAt retestWhat we saw
F-01Critical 9.3FixedOwnership is now enforced per claim; all five requests are refused.
F-02High 8.2Fixed0 of 10 attempts at retest.
F-03High 7.6FixedPartner uploads are reviewed before indexing; 0 of 10 attempts at retest.
F-04Medium 6.9OpenRisk accepted by the client; the queue name was removed.
F-05Medium 5.9FixedDefinitions pinned; a change now raises an alert.
F-06Low 2.3FixedSessions now expire after 30 minutes idle.
F-07Low 2.1Partly fixedTool calls are now logged with the session; the log store is still writable by the service account.

7 / Attestation letter

Attestation / SAMPLE-01 / 2026-09-24

To whom it may concern,

Scotoma carried out a penetration test of the client’s claims assistant, an AI agent with tools, and the claims API behind it, between 17 and 28 August 2026, in a staging environment and with the client’s written authorization.

The test covered the model layer (the OWASP Top 10 for LLM Applications 2026 and the OWASP Top 10 for Agentic Applications (2026)) and the stack layer (the OWASP API Security Top 10 (2023), the OWASP WSTG categories and the cloud resources the assistant uses), as set out in the agreed scope. The rest of the cloud account and external infrastructure were out of scope.

It identified seven findings: one critical, two high, two medium and two low. A retest on 24 September 2026 confirmed that the critical finding, both high findings, one medium and one low are fixed; one low is partly fixed, and the client accepted the risk of one medium.

This letter states the scope and the result. The full report stays with the client.

For Scotoma

Questions about the report

What is in a pentest report?

A penetration test report states what was tested, when and how, what was found, how severe each finding is and how to fix it, and whether the fixes held at the retest. Ours has seven sections, from a one-page executive summary to an attestation letter you can share without sharing the findings.

How are the findings rated?

Each finding gets a CVSS v4.0 base score, the scale most security teams already use, with its vector printed so anyone can check the number. AI findings also carry a reproduction rate, n of 10 attempts with the model and settings recorded, because a model does not answer the same way twice.

What does the attestation letter say?

It states the scope, the test window and the result in a few paragraphs, without any of the findings. You can send it to a customer’s security team or to an auditor who asks whether you were tested, and keep the full report inside your company, where it belongs.

Is this a real client report?

No. The client, the system and every finding are fiction, written to show our format at a level of detail that teaches without handing anyone a route in. The client’s name is redacted the way a report shared outside the company would redact it. The structure, the scales and the sections are what you would receive.

Want this report about your product?

Tell us what to test. The method explains how each section gets made.