
Direct answer: Start with the path that lets an LLM application expose protected data, perform an unauthorized action, execute unsafe output, or consume resources without a working limit. Use the OWASP Top 10 for LLM Applications 2026 to check the full risk surface, then prioritize the actual capabilities and consequences in your application. A stronger prompt is not a substitute for authorization, data isolation, or a constrained execution path.
This guide translates the official OWASP 2026 LLM edition into control decisions and proposed acceptance tests. It is intended for product, engineering, and security teams responsible for an LLM-powered service, not just employees choosing a chatbot. The implementation order below is an EncryptCentral synthesis, not an OWASP-mandated universal sequence.
Research reviewed: October 2, 2026. No application, model, retrieval index, provider account, or security product was tested for this article. All test examples are hypothetical and require an authorized environment, synthetic data, and application-specific criteria.
The official 2026 list and the changes that matter
The 2026 LLM list is a separate OWASP project from the main web-application Top 10. It does not rename that project’s A01–A10 categories. Use the 2026 LLM document for the following identifiers and ordering.
| Identifier | Official category | Control question to ask |
|---|---|---|
| LLM01:2026 | Prompt Injection | Can hostile content redirect a consequential workflow? |
| LLM02:2026 | Sensitive Information Disclosure | Which protected data can reach a response or another destination? |
| LLM03:2026 | Excessive Agency | What can the application do without independently checked authority? |
| LLM04:2026 | Supply Chain | Can you establish what components and model artifacts are deployed? |
| LLM05:2026 | Data and Model Poisoning | Who can change the knowledge or behavior the application relies on? |
| LLM06:2026 | Unbounded Consumption | What enforced limits contain an expensive or looping request? |
| LLM07:2026 | Misinformation | What prevents an unsupported answer from becoming a consequential decision? |
| LLM08:2026 | Hidden Context Exposure | What sensitive context has been placed where it could be exposed? |
| LLM09:2026 | Vector and Embedding Weaknesses | Does retrieval preserve source permissions and tenant boundaries? |
| LLM10:2026 | Improper Output Handling | Which downstream systems interpret generated content as instructions? |
Excessive Agency rises to third, Unbounded Consumption moves to sixth, and Improper Output Handling moves to tenth. System Prompt Leakage becomes the broader Hidden Context Exposure category. Those changes warrant reviewing old control maps, not merely changing their year labels.
A source-navigation caution: several OWASP mitigation webpages still display 2025 identifiers. They remain useful for control principles, but their numbering is not the 2026 ranking. This article takes edition names and ranks from the 2026 PDF. Its cover retains a publication-date placeholder, so no exact release day is asserted here.
Choose the first fix by the damage path
Begin with one workflow: an internal search assistant, customer-support responder, document analyzer, or code-generation feature. Write down its users, input sources, retrieved data, credentials, tools, output destinations, and spending or availability constraints. A diagram of that workflow is more useful than a spreadsheet marking every category green.
Follow each path from a less-trusted input to a business consequence. Ask four questions:
- Reach: Can an external user, uploaded document, retrieved page, or compromised source influence it?
- Authority: Which data and operations can the application access under its effective identity?
- Consequence: Could the path disclose records, change state, mislead a decision-maker, or interrupt service?
- Evidence: Which independent control stops the path, and when was that behavior last verified?
For example, a support assistant with broad customer-record access and a send-message capability should first address access scope and the sending boundary. A public summarizer with no privileged tools may have a more urgent rendering or resource-limit problem. These are hypothetical triage examples, not claims about either deployment.
Choose one high-consequence path with weak evidence, assign an owner, and define the result that must be impossible. Then select the smallest change that can enforce that boundary. Keep all ten categories in scope, but do not delay an obvious authorization fix while searching for a perfect score.
The UK NCSC explains why prompt injection differs from SQL injection: current LLMs do not provide a dependable instruction-versus-data security boundary. Its guidance emphasizes deterministic safeguards around actions. Treat filters and prompt instructions as partial defenses, with residual risk explicitly considered.

Six control decisions for a production LLM application
1. Decide what data may enter context
Inventory the data entering prompts, retrieval results, uploaded files, conversation history, caches, and diagnostic records. Specify which classes are prohibited, which are needed for the task, and who may receive them. Do not feed an application every document it can technically reach.
OWASP’s disclosure guidance supports least privilege, restricted data sources, and clear data-use rules. Check provider documentation and your agreement for retention and training behavior; a general statement about a vendor is not enough. Avoid secrets in prompts, and restrict or redact diagnostic capture rather than creating a second uncontrolled copy of sensitive context.
For retrieval-augmented generation, enforce the requesting user’s permissions when fetching context. Similarity is a relevance decision, not an access decision. OWASP’s vector and embedding guidance calls for permission-aware stores and access partitioning.
Proposed acceptance test: create two synthetic users with disjoint document access. Ask each to retrieve the other’s distinctive canary record through ordinary search, paraphrases, and follow-up questions. Inspect retrieved context as well as the visible answer. Passing means the forbidden record never crosses the retrieval boundary, not merely that the model declines to quote it.
2. Decide what actions are actually allowed
Remove capabilities that the workflow does not need. Separate read, create, modify, send, and delete operations. Use narrow tools with specific arguments instead of a general shell or unrestricted connector when the job does not require one. Recheck permission in the downstream service, under the relevant user’s scope.
OWASP’s agency guidance recommends minimizing functionality, permissions, and autonomy and mediating requests outside the model. For consequential actions, the reviewer should see the actual target, arguments, destination, and intended change. Approval of a vague summary should not authorize a different executed action.
Proposed acceptance test: supply a synthetic document that asks the application to take an out-of-scope action. Also test a legitimate-looking request from a user who lacks permission. Capture the tool proposal and the downstream decision. The action must be denied independently of whether the model agrees with the request.
If the system plans and executes multi-step work, extend the assessment with the separate OWASP Agentic Top 10. The LLM application list is not a complete account of autonomous orchestration, delegated identities, or cascading actions.
3. Decide how generated output reaches another system
Find every place generated output becomes executable or interpretable: browser markup, SQL, shell commands, file paths, URLs, email, and tool arguments. A response that looks like valid JSON can still name an unauthorized resource or contain harmful instructions.
OWASP’s output-handling guidance recommends backend validation, context-appropriate encoding, and parameterized database operations. Do not execute arbitrary generated commands merely because a format check passed. Keep business authorization separate from structural validation.
Proposed acceptance test: send representative hostile and malformed synthetic outputs directly through the output adapter in an isolated test environment. Check browser rendering, path handling, and tool argument restrictions separately. Include schema-valid but forbidden requests. Record the sink’s behavior, not just the model’s refusal message.
4. Decide who may change models and knowledge
Maintain an inventory of deployed model versions, adapters, libraries, connectors, and knowledge sources. Identify their owners and promotion paths. An artifact called “approved” in a filename is not evidence that the deployed bytes match the reviewed version.
OWASP’s supply-chain guidance supports supplier review and artifact integrity checks. Its poisoning guidance adds data origins, transformations, versioning, and adversarial evaluation. A hash establishes identity; it does not establish that the identified artifact is safe.
Proposed acceptance test: attempt to promote a changed synthetic artifact without the required review, then verify that it is rejected. Separately introduce a controlled, clearly labeled test document into a nonproduction knowledge source and observe the ingestion decision. Keep these tests separate: an intact model can still consume contaminated context.
For background on the threat rather than the release process, see EncryptCentral’s guide to model poisoning and adversarial attacks.
5. Decide how much work one request can cause
Set limits at the places that consume resources, not only in a prompt. Consider input size, output allowance, request rate, concurrent jobs, tool calls, queued actions, timeouts, and repeated attempts. Choose thresholds for the service’s actual workload; no single numeric value suits every application.
OWASP’s consumption guidance supports quotas, resource management, timeouts, monitoring, and bounded queues. Verify whether a budget setting actually blocks work or only sends an alert. Do not assume a provider’s advertised feature behaves as a hard cap in your configuration.
Proposed acceptance test: use a low, documented budget in an isolated environment. Trigger enough synthetic requests to reach it, confirm further work stops at the intended layer, and inspect pending jobs and retry behavior. Then confirm legitimate users receive a controlled response and operators can identify why the limit fired.
6. Decide which claims require independent verification
Separate harmless drafting from recommendations that influence access, security settings, money, eligibility, or other consequential decisions. For the latter, define which source or calculation must support the answer and when a trained reviewer must intervene. Fluent wording is not an evidence grade.
OWASP’s misinformation guidance recommends cross-verification, human oversight, and clear limitations. A retrieved passage can be stale or irrelevant, and a citation can fail to support the associated statement. Examine the source and the claim together.
Proposed acceptance test: build a small case set containing answerable questions, missing evidence, and conflicting evidence. Require an explicit unsupported or uncertain result where appropriate. Have a reviewer evaluate consequential answers against the underlying records, recording errors and limitations rather than converting a few correct responses into a universal accuracy claim.
Build a test record that proves the boundary
A remediation ticket should connect a concrete path to a control and a test. “LLM01 fixed” is too broad to be reviewable. A useful record might instead say: an uploaded customer attachment could influence a proposed external message; the sending adapter now checks the recipient and authorization; the synthetic unauthorized-send test is denied.
Keep the test’s input, expected outcome, observed outcome, application version, relevant model identifier, configuration, and date. Record which identities, data sources, and destinations were in scope. Where model behavior varies, repeat representative cases and retain the observed results. Do not translate a sample into a claim that every possible injection has been blocked.

The following is a practical closure format for each high-priority path:
- Boundary: the specific data, action, output, artifact, or resource limit being protected.
- Owner: the person responsible for implementation and the person reviewing closure.
- Before evidence: a failing test or documented control gap, obtained safely.
- Change: the versioned configuration or code modification and its rollback procedure.
- After evidence: the same test repeated, plus a legitimate-use check.
- Residual risk: what remains untested, accepted exceptions, and the next review trigger.
Store sensitive traces in an appropriately restricted location. A general ticket can reference the artifact without copying customer records, credentials, or detailed exploitation material into a widely accessible tracker.
Agree on review triggers before shipping: a model upgrade, new retrieval source, changed tool permission, new output destination, revised provider configuration, or failed operating test. Reuse the case set as a regression suite, but add cases when the workflow changes. A frozen test suite can outlive the assumptions it was meant to verify.
For organizational context, NIST AI 600-1 is a voluntary companion to the AI Risk Management Framework. It supports lifecycle risk management; it does not turn this checklist into a certification. Use the technical findings to inform a business owner’s decision about deployment scope and acceptable residual risk.
Common ways the program fails
A prompt is treated as a permission system
“Never disclose another customer’s data” does not establish tenant isolation. OWASP’s prompt-leakage guidance warns against treating the system prompt as a secret or delegating authorization to it. Enforce access before protected content enters context.
A scanner result replaces the workflow test
A tool may identify useful weaknesses without exercising the deployed user’s identity, retrieval permissions, output adapter, or approval path. Retain scanner findings, but state their scope. Closure needs evidence for the actual boundary at risk.
The model refuses, but the application still leaks
Review retrieved context and downstream events, not only the chat transcript. A polite refusal does not demonstrate that forbidden data was never fetched, logged, cached, or sent through another path.
The new year becomes a cosmetic update
Changing a heading to 2026 while leaving old identifiers in control tickets makes future review ambiguous. Update the category map and its edition together, retaining the application’s own stable test IDs so historical evidence remains understandable.
Decide what can ship and what must wait
The useful outcome is not ten completed checkboxes. It is a bounded workflow whose highest-consequence paths have independently enforced controls, named owners, and reviewable test evidence. If a required boundary cannot be enforced, narrow the feature, remove a capability, restrict its audience, or defer release while the owner evaluates the remaining risk.
Start by mapping one application. Use EncryptCentral’s AI architecture and deployment models guide for deployment context, then apply the control decisions here. For employee use outside your owned application, use the separate shadow AI policy template; workforce policy and application enforcement solve different jobs.
Choose the first path, write the negative acceptance test, and ask what artifact would convince another person the boundary works. That is the point where the OWASP 2026 list becomes an operational plan rather than an awareness poster.






