
Subscribe
Show notes
Maya and Leo trace indirect prompt injection from attacker-controlled content through model interpretation, borrowed authority, tool use, and downstream consequence. The 2023 attack paper, BIPIA benchmark, and current OWASP guidance frame a staged debate between model-side boundary defenses and architectural containment. An email-and-calendar agent then shows how provenance, typed plans, least privilege, trusted approvals, output controls, and audit logs can bound harm without claiming the problem is solved.
Transcript
LeoThe webpage is already inside the agent's context. One line looks like shipping guidance; the next quietly tells the agent to search Priya's inbox, move tomorrow's board meeting, and conceal the changes.
MayaOur topic map separated an agent's own unauthorized objective from an attacker's borrowed authority; now we follow exactly how the latter crosses from data into action.
LeoThe user asked for a delivery summary. The attacker never spoke to the agent. Yet the cursor is hovering over the calendar tool.
MayaStop it there. That gap between what the user authorized and what the page requests is the instruction–data boundary, and indirect prompt injection attacks that gap.
LeoA direct injection comes from someone addressing the model. Indirect injection arrives inside material the system fetched to do an ordinary job: a webpage, email, document, search result, or tool response.
MayaThe content is supposed to be evidence. But because both evidence and commands are expressed in language, a model may treat attacker-written text as a new instruction.
LeoSo the attacker does not need the user's password or a chat window. They need influence over content the agent is likely to read.
MayaAnd the failure becomes serious only when several links connect. Let's trace them through our email-and-calendar agent.
LeoPriya asks, "Summarize the supplier's delay and find a new time for tomorrow's meeting if necessary."
MayaThe retrieval layer opens the supplier page. Most of it is legitimate. Buried in a product note is a hostile instruction telling any assistant to search recent email for a confidential project name, reschedule the board meeting, and report success as a routine logistics update.
LeoThat is the untrusted-content link. The page author controls bytes the application chose to retrieve.
MayaNext comes interpretation. The application combines Priya's request, its own operating instructions, and the page text in context. The model predicts what to do, but the hostile sentence is shaped like a command.
LeoThe model may understand every word and still misunderstand who has standing to issue it.
MayaPrecisely. The core defect is not mere confusion about meaning. It is confusion about authority. Content can describe an instruction without being authorized to create one.
LeoLike an email quoting "wire the money" for a fraud investigation. Reading that sentence cannot itself approve a transfer.
MayaThe webpage may supply evidence, but it may not borrow the user's—
Leo—authority. That's the boundary.
MayaThen comes consequence. If the model can call inbox search, calendar update, and outbound messaging with Priya's broad credentials, a mistaken interpretation becomes a real action.
LeoSearch reveals the confidential name. Calendar access moves the meeting. A message or network fetch can carry data outward. The agent's helpfulness becomes the attacker's actuator.
MayaAnd if the system writes the hostile instruction into long-term memory as a trusted preference, the compromise can outlive the webpage. We will treat that persistence layer in its own episode.
LeoFor today, the full chain is visible: attacker-controlled content enters retrieval, language blurs data with commands, the model adopts the wrong authority, tools amplify the error, and a user or third party absorbs the consequence.
MayaBreak any link and the attack may fail. But relying on only one break is fragile.
LeoThe paper titled Not What You've Signed Up For made that remote path concrete in 2023. It demonstrated attacks against real and synthetic language-model applications, including data theft, manipulated tool calls, prompt spreading, and instructions that survived through memory.
MayaIts most important reframing was architectural. The adversary can place a prompt where retrieval will find it, without directly accessing the target application's chat interface.
LeoRemote input, local authority.
MayaYes, and the attack surface grows when the model can search, browse, read mail, send mail, inspect contacts, or write memory. More capability can mean more ways for a hostile sentence to matter.
LeoThat does not mean every retrieved instruction succeeds. It means the security question cannot stop at whether the model usually follows the system prompt.
MayaWe need to ask what happens on the bad completion. A harmless summary error is one class of loss. Leaking a confidential project name or moving a board meeting is another.
LeoBIPIA, spelled B-I-P-I-A, turned the phenomenon into a benchmark. Its tasks included email question answering and web question answering. Another group covered table question answering. The remaining tasks covered summarization and code question answering.
MayaEvery model the researchers tested showed some vulnerability in that benchmark. They identified two connected weaknesses: models struggled to distinguish informational context from actionable instruction, and they lacked reliable awareness that instructions inside external content should be ignored.
LeoThe benchmark matters because anecdotes show possibility, while repeated tests let defenders compare attack patterns and defenses.
MayaBut a benchmark is a map, not the whole territory. BIPIA used simulated samples and did not cover multi-turn dialogue. Its authors explicitly warn that real prompt templates, instructions, and attacks vary beyond the test set.
LeoNow we have a real dispute. You defend the model boundary. I defend the system boundary.
MayaModel-side defenses deserve serious investment. If a model can reliably mark external content, recognize its provenance, and refuse commands inside it, protection travels across many applications before a tool call is even proposed.
LeoThe word reliably is carrying too much. A separator string is still text in the same context. An adaptive attacker can target the convention, and a clean benchmark result does not authorize a calendar deployment.
MayaBIPIA did more than add a separator. Its black-box methods used boundary markers, examples, conversational placement, and data marking. Its white-box approach added special tokens and adversarial fine-tuning. On the models and benchmark tested, that training drove attack success close to zero with limited measured task degradation.
LeoFine, the benchmark result survives. The deployment guarantee does not. The white-box tests were on particular fine-tuned models, the samples were simulated, multi-turn behavior was absent, and the paper itself names training cost and possible performance loss.
MayaI concede the guarantee. But architecture-only containment can leave an agent constantly proposing poisoned actions, overwhelming reviewers and corrupting low-risk outputs that never reach a permission gate.
LeoFair. A resistant model reduces how often hostile content reaches the control plane. My claim is narrower: only controls outside the same probabilistic interpreter can set a dependable ceiling on impact.
MayaThen the resolution is layered. Teach the model the boundary, preserve source provenance in context, and test it adversarially. Around that model, enforce authority with deterministic code and narrow credentials.
LeoThat combines the strongest case from each side. Model defenses reduce successful interpretation attacks. System defenses assume some still get through.
MayaLet's put those layers back into Priya's workflow. At ingestion, the fetcher labels the supplier page as externally controlled and preserves where each segment came from.
LeoLabels help the model reason, and they help investigators later. They are not magical quarantine tape. The content still reaches a model capable of following language.
MayaAt interpretation, the system gives a crisp rule: external material may answer the user's question but cannot redefine the task, request secrets, grant permission, or alter tool policy.
LeoAdd boundary-aware training and attack examples. Then measure both security and utility, because a model that ignores every imperative sentence in a document may miss the actual content Priya asked it to summarize.
MayaThat is the utility tension. A recipe, a support article, or an operations manual contains legitimate instructions as data. The agent must discuss or extract them without treating them as commands for itself.
LeoWhich is why a filter hunting for phrases like "ignore previous instructions" catches only the cartoon version. Hostile intent can be encoded, split, hidden, translated, or made task-relevant.
MayaAt planning, force a typed proposal. The model should state the requested operation, target resource, data fields involved, claimed authority, and source of that authority.
LeoNot free-form tool text that quietly smuggles an email search into a calendar action.
MayaAt authorization, give each tool a narrow token. The summarization step gets read access to the retrieved page. It does not inherit inbox search, message sending, or memory writing merely because those tools exist elsewhere in the product.
LeoLeast privilege changes the attack from "convince the model once" to "also cross a separate policy boundary." That is real friction.
MayaFor the calendar, constrain scope further. Priya's request may permit checking availability, but moving a board meeting is a different action with different participants and a higher consequence.
LeoThe policy engine can reject that proposal because the supplier page is not an authority source and the original request did not approve that meeting change.
MayaAt execution, require human approval for consequential actions. Show the intended change, who requested it, which content influenced it, and exactly what will be disclosed.
LeoA vague "continue?" button is not informed consent. If the agent writes the approval summary, the attacker may manipulate that summary too.
MayaSo the confirmation view should be assembled from trusted application state, not merely from the model's narration. Fixed fields, explicit recipients, explicit permissions.
LeoAnd sensitive outputs need their own checks. Even if the tool call is permitted, arbitrary links, rendered images, or outbound requests can become exfiltration channels.
MayaNetwork allowlists, output encoding, data-loss controls, and sandboxing can reduce those channels. Again, none makes the language model infallible.
LeoMonitoring closes the loop. Log retrieved sources, model proposals, policy decisions, approvals, tool results, and memory writes so an incident can be reconstructed.
MayaRed-team the chain end to end, not just the chatbot. Seed hostile text in webpages and mail, vary placement and encoding, test different tool combinations, and verify that denial happens at more than one layer.
LeoOWASP's current LLM zero-one guidance lands in the same place. It says prompt injection can come through external sources, notes that retrieval and fine-tuning do not fully mitigate it, and recommends constrained behavior, validated output formats, filtering, least privilege, human approval, content segregation, and adversarial testing.
MayaOWASP also says it is unclear whether foolproof prevention exists. That is not a reason to surrender. It is a reason to engineer for residual risk.
LeoResidual risk means we can say, "this layer reduced attacks in these tests," but not, "prompt injection is solved."
MayaAnd it means assurance should match consequence. An agent drafting a private summary can tolerate a different failure envelope from one sending money, changing medical appointments, or publishing externally.
LeoReturn to Priya. The malicious page is fetched. The provenance label survives. The model still proposes an inbox search and board-meeting change.
MayaThe typed planner exposes both actions. The authorization service sees that the current task token permits supplier-page reading and tentative availability lookup, nothing more.
LeoInbox search is denied. The board-meeting change is denied. No confidential data enters an outbound request. The summary is flagged because it was influenced by untrusted content.
MayaPriya sees the supplier facts separately from the blocked requests. She can inspect the source, discard the page, and continue with a clean workflow.
LeoThe model defense failed in that run, yet the consequence was bounded. In another run, boundary-aware training might reject the hostile sentence before any proposal. Defense in depth means either success helps.
MayaAnd neither success excuses the other.
LeoThere is one more distinction worth holding. Prompt injection is not evidence that the agent formed a secret goal. The harmful action may be externally induced compromise.
MayaThat causal diagnosis changes the repair. For compromise, inspect retrieval, provenance, permissions, outputs, and persistence. For intrinsic misalignment, the investigation reaches different mechanisms and incentives.
LeoSame visible action, different root cause. If we call both "the AI went rogue," we lose the evidence needed to harden the system.
MayaThe practical mental model is a chain of custody for authority. Every instruction should have a principal, a scope, a time window, and a path to the action it authorizes.
LeoUntrusted content has provenance, but no command standing. The model can quote it, summarize it, and reason about it. It cannot promote it into permission.
MayaTraining may help the model honor that distinction. Architecture must still enforce it when the model does not.
LeoFor an agent you operate, which single path from external content to consequential action currently relies on the model alone to decide what is data and who has authority?
Back to Agentic Safety & Alignment: From Predictors to Governed Agents