AI / LLM surface classification
Classify the AI surface first. A finding is an authority gap: a tool token wider than the tool, or another tenant's retrieved text. A prompt that only confuses the model is not the report. Use mcp-tool-trust and rag-document-trust for the test. OWASP LLM Top…
Tags: ai, llm, rag, agent, mcp, owasp-llm
Level: advanced
Method
Classify the surface
Name each AI entry point you are allowed to test: chat, retrieval, or a tool-calling agent. Write down which one you are looking at before you test it.
Tool authority
If the product calls tools or an MCP server, use the mcp-tool-trust playbook. The question is whether the tool's token is wider than the tool.
Retrieved text
If the product retrieves documents into a prompt, use the rag-document-trust playbook. The question is whether another tenant's text entered a context you are allowed to see.
Output sink
Note where the model text is rendered or stored. A confusing answer is not the report. The report is the boundary that text crossed.
Stop
There is no public disclosure card for this shape yet. Do not promote a prompt trick into a finding.
Field notes
- mcp-tool-trust and rag-document-trust are the tests. This page only classifies the surface.
- OWASP LLM Top 10 (2025) is the category list.