research / owasp-top-10-llm-practitioners-guide

The OWASP Top 10 for LLM Applications, explained for builders

A practical walk through each risk in the OWASP LLM Top 10 (2025), what it looks like in real systems, and the first control worth adding.

The OWASP Top 10 for LLM Applications is the closest thing the industry has to a shared vocabulary for AI application risk. The 2025 edition reorganised the list around what teams actually ship now: retrieval-augmented generation (RAG), agents with tools, and models wired into production systems.

This guide covers each entry briefly: what it means, how it shows up, and the first control worth putting in place.

LLM01: Prompt Injection

Input changes the model's behaviour in ways the developer did not intend. Direct injection comes from the user. Indirect injection hides in content the model reads, such as a web page, a PDF, an email, or a tool result.

First control: treat everything the model reads as untrusted. Keep instructions and data in separate channels where your API allows it. Never let model output alone authorise a sensitive action.

LLM02: Sensitive Information Disclosure

The model reveals data it shouldn't: other users' records pulled in through retrieval, secrets in the system prompt, or memorised training data.

First control: enforce access control before retrieval, not in the prompt. If a user can't read a document, it should never enter their context window.

LLM03: Supply Chain

Models, adapters, datasets, and plugins come from third parties and can be tampered with or simply be poor quality.

First control: pin and verify model artefacts the same way you pin packages. Prefer safe serialisation formats over ones that can execute code when loaded.

LLM04: Data and Model Poisoning

Attackers influence training, fine-tuning, or embedding data to plant backdoors or bias.

First control: track the provenance of every fine-tuning and RAG source. Watch for unexpected writes to the knowledge base.

LLM05: Improper Output Handling

Model output gets passed to a browser, shell, SQL engine, or template without validation. This is classic injection with an LLM in the middle.

First control: encode and validate model output exactly as you would user input. It is user input, one step removed.

LLM06: Excessive Agency

An agent has more tools, permissions, or autonomy than its task needs. When it gets manipulated, the damage matches the permissions.

First control: least privilege per tool. Read-only by default, and require human confirmation for irreversible actions.

LLM07: System Prompt Leakage

The system prompt is extracted. The real problem is usually that it contained something it shouldn't, such as credentials, internal logic, or access rules.

First control: assume the system prompt is public. Move secrets and authorisation decisions out of it.

LLM08: Vector and Embedding Weaknesses

RAG pipelines bring their own risks: cross-tenant leakage in shared vector stores, poisoned documents, and embedding inversion.

First control: partition vector stores per tenant or apply metadata filters, and sanitise documents at ingestion.

LLM09: Misinformation

The model produces confident, plausible, wrong output, and people or systems act on it.

First control: ground answers in retrieved sources, show citations, and design the interface so users can verify claims.

LLM10: Unbounded Consumption

Uncontrolled inference requests drain budgets, degrade service, or enable model extraction.

First control: rate limits and token quotas per user and per key, plus cost alerts.

Where to start

If you can only do three things this quarter:

  1. Map your data flows. Know every source of text that reaches the model and every place its output goes.
  2. Shrink agency. Strip tools and permissions to the minimum and add confirmation steps for anything destructive.
  3. Test adversarially. Run prompt-injection cases against your own application before launch and keep them as regression tests.

The list will keep changing as the attack surface does. The underlying principle won't: a language model is not a security boundary, so build the boundary around it.