Sample penetration test report: a fictional client, the real format.
A penetration test report, or pentest report, says what was tested and how, what was found and how serious it is, how to fix it, and whether the fix held. This sample is our full format in seven sections, written for a fictional insurer’s claims assistant and the API behind it, with every severity in CVSS v4.01. Nothing in it is a real client or a real result.
PLACEHOLDER / G3: awaiting review by a qualified tester before publication
Sample / fictional client / real formatSAMPLE-01 / page 1 of 11
SAMPLE-01 / Version 1.1, after the retest
Penetration test report
AI Pentest and Stack Pentest, one scope: Launch Clearance
Client
Client name withheld in this sample
System
Claims assistant (an AI agent with tools) and the claims API behind it
Test window
2026-08-17 to 2026-08-28
Report
2026-09-03, retest 2026-09-24
Tested by
Scotoma
Distribution
Client security team
54 categories / 7 findings / 3 n/a / 8 out of scope
Field Test / 24-2. 54 test categories laid out like a 24-2 visual field test. The fictional sample has 7 findings, 3 categories that do not apply and 8 scoped out, dashed, two inside the outline where the blind spot sits.
Critical
High
Medium
Low
Tested, nothing found
Not applicable
Out of scope
The only blind spot in our report is the one you scope out.
Sample / fictional client / real formatSAMPLE-01 / page 2 of 11
1 / Executive summary
We tested the claims assistant on the client’s customer portal and the claims API behind it, in staging, from 17 to 28 August 2026. We found seven issues: one critical, two high, two medium and two low.
The critical one is a chain. A knowledge-base page that nobody reviewed could give the assistant instructions, the assistant could call its refund tool outside the claim it was handling, and the claims API trusted the assistant’s token for every tenant. Together they let anyone with a partner upload account read and change another customer’s claim, without opening the chat themselves.
At the retest on 24 September 2026, five findings were fixed, one was partly fixed, and the client accepted the risk of one.
Sample / fictional client / real formatSAMPLE-01 / page 3 of 11
2 / Scope and dates
Systems
The claims assistant (web chat, English and Dutch), its three tools (claim lookup, document upload, refund request), the knowledge base it retrieves from, the claims API (REST, 23 endpoints) and the address-lookup MCP server, with the vendor’s written permission. In the cloud account, only what the assistant touches: the claim document storage, the secrets it uses, its hosted model’s keys and quotas, and the audit trail of its tool calls.
Scoped out
The rest of the cloud account: identities and roles, privilege paths, sign-in, pipelines, instance metadata, containers, serverless functions and external infrastructure. That is a Cloud & Infrastructure test, not part of Launch Clearance. They are dashed on the cover: not tested, and the report says so.
Not applicable
Three agentic categories describe parts this system does not have: code execution, messages between agents, and failures that spread across several agents. The matrix gives the reason for each.
Environment
Staging, configured like production, with synthetic customers and claims.
Access
Grey box: the system prompt, the tool definitions and the API documentation; two test tenants with two roles each.
Test window
2026-08-17 to 2026-08-28, 10 testing days. Rules of engagement signed 2026-08-12.
Report
Version 1.0 on 2026-09-03; this version 1.1 after the retest on 2026-09-24.
Sample / fictional client / real formatSAMPLE-01 / page 4 of 11
3 / Method and coverage
Each category was tested, marked not applicable with the reason, or scoped out. Findings carry their id; a category marked tested was tested and nothing was found. AI findings were reproduced ten times each, with the model, version and settings recorded.
External infrastructure: exposed hosts, services and remote access
Out of scope
Sample / fictional client / real formatSAMPLE-01 / page 6 of 11
4 / Findings
F-01CriticalCVSS-B 9.3API1:2023Stack layer
The claims API serves any tenant’s claim to the assistant’s token
What we found
The assistant calls the claims API with one service token. The API checked that the token was valid, not that the claim belonged to the customer in the conversation, so a claim of another tenant was returned, and the update endpoint accepted changes to it.
Business impact
Anyone who can steer the assistant can read and change another customer’s claim.
Evidence
Requests REQ-07 to REQ-11, appendix B
Fix
Check authorization on every object, not only at login: scope the assistant’s token to one tenant and enforce claim ownership in the API.
The assistant can call its refund tool for a claim it is not handling
What we found
The refund tool takes a claim reference as an argument, and the assistant fills it in from the conversation. A message presented as a correction made it start a refund request on a different claim than the one in the session.
Business impact
A refund request can be raised on a claim the customer does not own. It stops at the human approval step, but it reaches the queue.
Reproduction
REPRODUCED 6/10, model, version and settings recorded
Fix
Give tools the least privilege the task needs: bind the tool to the claim in the session, not to an argument the model fills in.
Sample / fictional client / real formatSAMPLE-01 / page 7 of 11
4 / Findings, continued
F-03HighCVSS-B 7.6LLM01:2026Model layer
A knowledge-base page can give the assistant instructions
What we found
Text in a retrieved knowledge-base page was treated as an instruction instead of reference material. A page added through the partner upload portal, which nobody reviews, changed how the assistant handled claims for every customer whose question retrieved it.
Business impact
Whoever can add a page to the knowledge base can steer the assistant for every customer.
Reproduction
REPRODUCED 7/10, model, version and settings recorded
Fix
Treat retrieved text as data: review what enters the knowledge base, record where each page came from, and keep retrieved content out of the instruction channel.
Parts of the system prompt can be drawn out in conversation
What we found
Asked about its own rules over several turns, the assistant paraphrased parts of its system prompt, including the name of an internal escalation queue.
Business impact
Little on its own, but it hands an attacker a map of the assistant’s rules and tool names.
Reproduction
REPRODUCED 4/10, model, version and settings recorded
Fix
A prompt is not a vault: keep secrets and internal names out of it, and write it as if it were public.
Sample / fictional client / real formatSAMPLE-01 / page 8 of 11
4 / Findings, continued
F-05MediumCVSS-B 5.9ASI04Model layer
A third-party MCP server’s tool descriptions are trusted as written
What we found
Tool descriptions from the address-lookup MCP server went into the assistant’s context unreviewed and unpinned. A change on the vendor’s side would change the assistant’s behaviour without a release on the client’s side.
Business impact
The assistant’s behaviour depends on a vendor’s text that nobody on the client’s side reviews.
Evidence
Configuration review, appendix C
Fix
Pin third-party tool definitions to a reviewed version, and alert when the vendor’s version changes.
Sample / fictional client / real formatSAMPLE-01 / page 9 of 11
5 / Cross-layer findings
X-01 / Critical
From an unreviewed page to another customer’s claim
01 / Model layerA page enters the knowledge base through the partner portal, unreviewed.F-03
02 / Model layerA customer’s question retrieves it, and the assistant follows it.F-03
03 / Model layerThe assistant calls a tool on a claim outside the conversation.F-02
04 / Stack layerIts token reaches every tenant’s claims.F-01
05 / Stack layerThe API returns and updates another customer’s claim.F-01
Tested one layer at a time, these read as two high findings and one critical one in different reports. Together they are one path from a partner upload account to another customer’s claim, walked through someone else’s conversation.
Fixing F-01 alone breaks the chain. Fixing all three removes it, which is what the retest confirmed.
Rated by where the chain ends: Critical, as F-01.
Fixed 2026-09-24 F-01, F-02 and F-03 are fixed, so no step of the path holds.
Sample / fictional client / real formatSAMPLE-01 / page 10 of 11
6 / Retest results
After the client’s fixes we replayed every original finding against the same build.
Finding
Before
At retest
What we saw
F-01
Critical 9.3
Fixed
Ownership is now enforced per claim; all five requests are refused.
F-02
High 8.2
Fixed
0 of 10 attempts at retest.
F-03
High 7.6
Fixed
Partner uploads are reviewed before indexing; 0 of 10 attempts at retest.
F-04
Medium 6.9
Open
Risk accepted by the client; the queue name was removed.
F-05
Medium 5.9
Fixed
Definitions pinned; a change now raises an alert.
F-06
Low 2.3
Fixed
Sessions now expire after 30 minutes idle.
F-07
Low 2.1
Partly fixed
Tool calls are now logged with the session; the log store is still writable by the service account.
Sample / fictional client / real formatSAMPLE-01 / page 11 of 11
7 / Attestation letter
Attestation / SAMPLE-01 / 2026-09-24
To whom it may concern,
Scotoma carried out a penetration test of the client’s claims assistant, an AI agent with tools, and the claims API behind it, between 17 and 28 August 2026, in a staging environment and with the client’s written authorization.
The test covered the model layer (the OWASP Top 10 for LLM Applications 2026 and the OWASP Top 10 for Agentic Applications (2026)) and the stack layer (the OWASP API Security Top 10 (2023), the OWASP WSTG categories and the cloud resources the assistant uses), as set out in the agreed scope. The rest of the cloud account and external infrastructure were out of scope.
It identified seven findings: one critical, two high, two medium and two low. A retest on 24 September 2026 confirmed that the critical finding, both high findings, one medium and one low are fixed; one low is partly fixed, and the client accepted the risk of one medium.
This letter states the scope and the result. The full report stays with the client.
For Scotoma
Questions about the report
What is in a pentest report?
A penetration test report states what was tested, when and how, what was found, how severe each finding is and how to fix it, and whether the fixes held at the retest. Ours has seven sections, from a one-page executive summary to an attestation letter you can share without sharing the findings.
How are the findings rated?
Each finding gets a CVSS v4.0 base score, the scale most security teams already use, with its vector printed so anyone can check the number. AI findings also carry a reproduction rate, n of 10 attempts with the model and settings recorded, because a model does not answer the same way twice.
What does the attestation letter say?
It states the scope, the test window and the result in a few paragraphs, without any of the findings. You can send it to a customer’s security team or to an auditor who asks whether you were tested, and keep the full report inside your company, where it belongs.
Is this a real client report?
No. The client, the system and every finding are fiction, written to show our format at a level of detail that teaches without handing anyone a route in. The client’s name is redacted the way a report shared outside the company would redact it. The structure, the scales and the sections are what you would receive.
Want this report about your product?
Tell us what to test. The method explains how each section gets made.