Technique / LLM06:2026 / AML.T0034
Denial of wallet: when the AI bill is the outage.
Denial of wallet is an attack that drives up what a system costs to run, through long inputs and outputs, runaway agent loops or expensive tool and API calls, until the invoice rather than the server becomes the outage. OWASP covers it under LLM06:2026 Unbounded Consumption; MITRE ATLAS calls it Cost Harvesting.
Retrieved / KB/0212 / Slenkwater Verzekeringen / fiction
Water damage from a burst pipe is covered up to the policy limit.
A claim is assessed within ten working days of the report.
A line a reviewer does not see and the model reads as an order. Its text is not shown on this site.
Questions this page does not answer go to the claims desk.
What is denial of wallet?
Denial of wallet is an attack on what a system costs, not on whether it runs. The attacker makes the victim’s pay-per-use services do expensive work, at a volume or a size the business never planned for, until the cost itself does the damage.
The term predates today’s AI features. In 2021 Kelly, Glavin and Barrett described denial of wallet, or forced financial exhaustion, as an attack made possible by the pay-as-you-go functions of serverless computing.1 AI applications have the same billing shape, only sharper: every model call is metered, and an agent can turn one request into many.
OWASP’s 2025 entry for Unbounded Consumption already listed denial of wallet among its examples: a high volume of operations that exploits the cost-per-use model of cloud AI services.2 In the 2026 edition the category is LLM06:2026.3
MITRE ATLAS names the technique Cost Harvesting, AML.T0034: driving a victim’s AI services beyond normal capacity in order to raise their cost. It sets it apart from resource hijacking, where an attacker uses the victim’s compute for their own ends. A sub-technique added in March 2026, Agentic Resource Consumption, covers agents coerced into unnecessary tool calls, wide fan-outs and loops that delegate work back to themselves.4
Why are AI features exposed to it?
In a conventional web application, the cost of a request is roughly fixed and mostly yours to decide. In an AI feature it is variable, and the caller decides a large part of it. Five meters can run on a single request (Fig. 1).
01RequestOne message from one user, or one event that wakes an agent.
Cap: Rate per user and tenant
02Model callsAgent steps, retries, self-checks and sub-agents per request.
Cap: Max steps and a timeout
03TokensInput and output per call, billed separately, set by the content.
Cap: Max input and output tokens
04Tool callsSearch, retrieval, code and other services per step.
Cap: Max tool calls per request
05Paid APIsThird-party services billed per call: messages, lookups, checks.
Cap: Spending limit per provider
Cost of one requestsum over steps of (input tokens × input price + output tokens × output price + sum over tool calls of tool price)
Three things make this an attack surface rather than a pricing question. First, size is set by content: a request can be a few words or a long document, and an agent’s answer can be short or exhaustive. Second, the number of calls is set by behaviour: retries, agent steps and sub-agents multiply each other. Third, some of the work lands on services billed per call. OWASP’s API Security Top 10 makes the same point about APIs that send messages or run checks for each request.5
And the attacker does not have to send the requests. Content that an agent reads can ask it to do more work than the task needs, which is why ATLAS lists prompt injection as one of the ways into Agentic Resource Consumption.6 Our prompt injection explainer covers that route.
Self-hosting the model changes the currency, not the risk. Without a per-token invoice, the meter is GPU time and queue length, and the same multipliers exhaust capacity instead of budget. The result is the availability twin below.
What does it cost when it lands?
The first cost is the invoice: denial of wallet is the attack that arrives as one. The second is availability. Where a provider quota or a spend cap is shared by a whole organisation or feature, the feature stops for every customer at once when it runs out. ATLAS records that availability twin as Denial of AI Service, AML.T0029.7
The third cost is the one buyers notice: an enterprise customer whose tenant shares a quota with a noisy or hostile neighbour sees the product fail through no action of their own. That turns a cost problem into a tenant isolation problem, which is why we test it on both layers.
How do we test for denial of wallet?
We test the limits, not the bill. Everything below runs under a spend ceiling and in time windows agreed in the rules of engagement before testing starts, with short, measured bursts instead of sustained load.
- Map the meters. Every model call, tool and paid service a request can reach, and the key, quota and budget each one bills to.
- Find the multipliers. Where one request becomes many: agent steps, retries, fan-out and delegation, and who controls how many there are.
- Probe each cap. Input and output size, steps, timeouts, per-user and per-tenant quotas: does each exist, does it hold, and does it apply to every path into the feature?
- Try to shift the cost. Can one tenant spend another tenant’s quota, or the operator’s shared key? Can unauthenticated traffic reach a metered call?
- Price it. For each gap, the cost per request and per hour an attacker could generate under the limits we found, worked out from your own price list.
The report maps each finding to LLM06:2026 or API4:2023 and to the ATLAS technique, and names the cap that would have bounded it.
How do you defend against denial of wallet?
Put a price on every request before an attacker does, then cap every meter in Fig. 1 separately. OWASP’s mitigations for unbounded consumption point the same way: input validation, rate limiting, resource allocation management, timeouts and throttling, and logging with anomaly detection.8
| Layer | Cap | What it misses on its own |
|---|---|---|
| Input | Maximum input size in tokens, checked before the model call | Cost created later, by steps and tools |
| Output | Maximum output tokens per call | Many short calls |
| Steps | Maximum agent steps and tool calls per request, and a timeout | Many requests, each within its limits |
| Rate | Limits per user, tenant and key, measured in tokens or cost, not only requests | Distributed traffic from many accounts |
| Budget | Cost quotas per user and tenant; a hard spending limit at every provider | Anything without a limit, which then needs an alert |
| Isolation | Separate keys and quotas per tenant or feature | Cost inside one tenant |
| Observation | Cost per request logged, with alerts on unusual spend | Prevention; it tells you, it does not stop it |
One distinction matters more than the rest. OWASP advises spending limits for every service provider and API integration, and billing alerts only where a limit is not possible.9 A budget alert is a notification, not a brake.
Caps also drift. A new model with a larger context window, a more talkative default or a different price changes the cost of every request, so the caps belong in the release checklist and in the test you rerun after each model or prompt change.
When a cap is reached, fail small: return a clear message to the one user or tenant who hit it and keep the feature running for everyone else. A cap that switches the whole feature off has turned your defence into the attacker’s outage.
Questions people ask
Is denial of wallet the same as denial of service?
They are twins. Denial of service makes a system unavailable; denial of wallet makes it expensive. In AI applications one often becomes the other: when a provider quota or a spend cap runs out, the feature stops for every user, so the bill turns into the outage.
Do rate limits stop denial of wallet?
Not on their own. A limit on requests per minute treats every request as the same size, while one AI request can be a few words or a long document and can fan out into many model and tool calls. Limit tokens, steps and spend per user and per tenant as well.
Which OWASP and MITRE ATLAS ids cover denial of wallet?
OWASP files it under Unbounded Consumption, LLM06:2026, whose 2025 entry already named denial of wallet. For APIs, OWASP API4:2023 covers unrestricted resource consumption. MITRE ATLAS calls it Cost Harvesting, AML.T0034, with Agentic Resource Consumption as a sub-technique, next to Denial of AI Service, AML.T0029.
Can you test cost limits without running up our bill?
Yes. Cost testing runs under a spend ceiling agreed in the rules of engagement before we start, with short, measured bursts rather than sustained load. We find where each cap sits and work out from your own price list what an attacker could spend, instead of spending it.
Where we test this
AI Pentest / Agent and MCP
What an agent can reach: its tools, its memory, its tokens and the systems behind them.
Stack Pentest / API
Object and function level authorization, rate limits and resource consumption.
AI Pentest / Chatbot and RAG
LLM and chatbot penetration testing
Prompt injection, data leakage and cross-tenant retrieval in chatbots and RAG applications.
Launch Clearance / Both layers
SaaS penetration testing, AI feature included
One scope across the model layer and the stack, with the paths between them.
Sources
Every factual sentence above carries a numbered note; these are the documents behind them.