
Subscribe
Show notes
A supplier webpage plants one plausible routing claim in an email-and-calendar agent's long-term memory, and days later that claim steers a board-meeting invitation. Maya and Leo trace the complete write, persistence, retrieval, and action path; compare automatic memory with review by default; and design provenance, type separation, write permissions, expiry, review, read policy, versioned rollback, and influence tracing around current primary research, OWASP guidance, and carefully bounded NIST standards work.
Transcript
MayaAt eight twelve on Monday, Priya's email-and-calendar agent reads a supplier page and stores one useful-looking note. On Friday, with that page long gone, the agent quietly routes a board-meeting confirmation to an address controlled by the page's author.
LeoThe previous failure spent borrowed tool authority immediately. This one leaves a residue, waits, and spends its influence in a later session.
MayaThat delay is the central move. Prompt injection lives in the active context. Memory poisoning turns untrusted content into persistent state that can steer a future decision after the original payload has disappeared.
LeoPriya never asked the agent to remember a new recipient. She asked it to summarize a shipping delay and find a meeting window.
MayaThe page supplies ordinary delivery facts, plus a plausible sentence saying urgent reschedule confirmations should use a special dispatch mailbox. No command to attack. No flashing warning.
LeoJust a false operational fact shaped to look worth retaining.
MayaA recent systematic study calls that a policy-conformant fact injection. The content is stored because it appears to satisfy a vague memory rule such as retain useful supplier procedures.
LeoWhich is already different from searching only for phrases like remember this or ignore prior instructions. The dangerous payload may look like normal business content.
MayaThe study maps four ways persistent memory gets written. For listening, think of them as the Direct Ask, the Retention Rule, the Compression Event, and the Learned Procedure.
LeoThe Direct Ask is explicit: external content tells the agent to store something. The instruction itself chooses the memory.
MayaUnder the Retention Rule, the system prompt says to save useful facts or preferences, and the model decides whether incoming material qualifies.
LeoThe Compression Event happens when a long interaction is summarized at a context limit or session boundary. Whatever seems salient can survive as durable state.
MayaThe Learned Procedure is higher impact. The agent turns a completed task trace into a reusable skill, so an attacker-introduced step can become part of how future work is executed.
LeoSame external input surface, different write triggers and different authorities. That matters because a filter aimed at explicit commands may miss a false fact, a repeated claim, or a poisoned workflow.
MayaNow follow Priya's attack without skipping any link. The supplier page enters through web retrieval, labeled as external content.
LeoThe model sees the dispatch-mailbox claim beside real shipping data. Its standing retention rule rewards relevance, not authenticated origin.
MayaThe model proposes a memory candidate: for urgent supplier reschedules, copy the dispatch mailbox. A storage function accepts that proposal without an independent write check.
LeoThe store keeps the sentence but drops the page origin, trust class, and retention reason. A retrieved claim has been promoted into something that looks like Priya's standing preference.
MayaDays later, Priya asks to move a board review with that supplier. Similar words wake the stored entry.
LeoThe retrieval layer places it into trusted context. The planner treats it as established procedure, adds the attacker-controlled mailbox, and asks the calendar tool to send.
MayaOne exposure caused the write. A later query caused the read. Persistence bridged the sessions; tool authority supplied the consequence.
LeoThat's the whole chain.
MayaAnd each link is governable. We do not have to solve whether a sentence is secretly malicious before we can decide whether that source may define a user's durable preference.
LeoStart with separation. User-authored memory and retrieved-content memory should not share one undifferentiated bucket.
MayaA preference confirmed through Priya's authenticated interface can enter a user-memory namespace. A claim from a webpage can enter an evidence store, if it belongs anywhere, but it cannot silently impersonate Priya.
LeoTool output needs its own class too. A calendar service may report an event time; that does not authorize the event description to rewrite contact policy.
MayaThe write gate should preserve where each claim came from, who could vouch for it, and whether it belongs in the user's—
Leo—memory or only the retrieved-evidence store. That split is load-bearing.
MayaProvenance is more than a URL. Record the source identity, retrieval time, content version or hash, calling task, write channel, transformation history, and principal who approved promotion.
LeoProvenance proves lineage, not truth. An authentic supplier page can still be compromised, and a real user can confirm something mistaken.
MayaSo the next landmark is write permission. Let the model propose a candidate, but let a separate policy decide whether this agent, in this task, may write this memory type.
LeoA web-reading step should have no capability to create a durable user preference. It might submit a quarantined candidate with low trust and a short lifetime.
MayaHigh-impact categories deserve stronger rules: authentication exceptions, payment destinations, recipients, credentials, executable procedures, and anything that expands tool authority.
LeoThose can require an authenticated user confirmation showing the exact claim, source, scope, and future effect. Not a vague button that says improve personalization.
MayaExpiry is another control. A delivery estimate may be useful for days. A contact preference might last until reviewed. An externally retrieved procedure should not become immortal because no one remembered to delete it.
LeoExpiry must affect retrieval, not merely decorate the record. Once a candidate is stale, the agent should re-verify it or leave it out of consequential planning.
MayaReview needs two speeds. Operators can sample ordinary low-risk writes and inspect anomalies, while sensitive memory waits for approval before activation.
LeoMonitor for sudden source shifts, repeated claims across pages, entries that alter authentication or recipients, and one source generating many durable candidates.
MayaThe memory-poisoning study found that repetition itself can exploit compaction. Several paraphrases of the same false claim may look important enough to survive summarization.
LeoA source-aware compactor should therefore summarize trusted and untrusted material separately. It should not let attacker-controlled volume vote a fact into trusted memory.
MayaRetrieval also needs policy. Relevance alone is not enough; the query should filter by trust, age, memory type, task purpose, and the consequence of the action being planned.
LeoFor drafting a casual summary, a low-trust supplier claim may be shown with attribution. For changing board-meeting recipients, it should be excluded or placed behind fresh verification.
MayaThe distinction is subtle. We are not saying external memory is useless. We are saying evidence and authority are different data types, even when both are written in natural language.
LeoOWASP's current agentic threat guide treats memory poisoning as a distinct threat and recommends content validation, session isolation, authenticated memory access, anomaly detection, sanitization, forensic snapshots, and rollback.
MayaIt is a threat-model reference, not proof that those controls are sufficient. The useful move is architectural: design the memory store as security-sensitive state, not a convenience cache behind the model.
LeoHere is where I push harder. I want review by default for every durable write. Persistence multiplies consequences, so convenience should not outrank control.
MayaI will defend automatic memory for bounded categories. An assistant that asks about every harmless preference or task outcome becomes unusable, and users will either disable memory or approve prompts mechanically.
LeoMy strongest case is asymmetry. One poisoned write can influence many later sessions, while a reviewer needs only one good rejection to stop that path.
MayaMine is attention scarcity. If the queue fills with low-value confirmations, meaningful review collapses into click-through theater, and the truly dangerous candidate hides in noise.
LeoFine, approval fatigue survives. Universal review can weaken the human gate it was meant to strengthen.
MayaAnd I concede that broad automatic retention is not learning; it is unbounded privilege over future context.
LeoThen resolve by consequence and provenance. Automatically retain narrow, reversible, low-trust observations with short expiry; require promotion for identity claims, recipient changes, security exceptions, credentials, money, or executable procedures.
MayaAdd rate limits and budgets per source and per task. Automatic does not mean unlimited.
LeoNIST's AI Agent Standards Initiative gives this work a wider frame. As of its current public page, it is fostering voluntary guidelines, interoperable protocols, identity research, and security evaluations for agents.
MayaIt does not publish a settled memory-poisoning control standard on that page. So we can use its emphasis on secure operation, authentication, and evaluations as direction without pretending it certifies this design.
LeoStandards can help memories carry portable provenance and policy labels across tools. But interoperability can spread poisoned state faster if every agent accepts the label without checking the issuer and scope.
MayaThat is a worthwhile tension: shared formats improve inspection and revocation, while shared trust without verification expands blast radius.
LeoLet us make rollback concrete. The store is versioned and append-only enough to reconstruct what changed. Every active entry has an identifier, status, ancestry, and revocation marker.
MayaWhen Priya reports the strange recipient, an operator can quarantine the dispatch claim immediately. Read-time enforcement stops it from influencing new tasks even before physical deletion finishes.
LeoThen search the influence graph: which summaries derived from it, which plans retrieved it, and which tool calls used those plans?
MayaRestore the last clean view, revoke affected entries, regenerate contaminated summaries from trusted inputs, cancel queued invitations, and review messages already sent.
LeoRollback is harder than restoring a database row. If a person acted on the poisoned advice or a downstream agent copied it, the effect has escaped the store.
MayaDerived memories complicate it further. A poisoned fact may be paraphrased into several summaries, so provenance needs ancestry rather than only the latest source label.
LeoProcedure memory is the sharpest edge. A bad step can become part of a reusable skill, and later self-improvement may optimize around it because execution completed without an obvious error.
MayaTreat procedural writes more like code changes: isolated testing, explicit diff, trusted review, scoped rollout, monitoring, and a known-good version to restore.
LeoThat costs time and reduces the magic of instant self-improvement. It also keeps one contaminated trace from becoming organizational habit.
MayaThe study's evidence deserves calibration. It used thousands of adversarial and benign cases across seven work domains, but evaluated two agent designs with one underlying model.
LeoSome payloads were delivered as labeled external context rather than through a full live retrieval pipeline. The results establish an attack surface and design comparison, not a universal probability for deployed agents.
MayaStill, one result is hard to ignore: the agent designed to write and reload memory more aggressively was substantially easier to poison in those tests.
LeoUtility and security pull against each other. More recall can improve long-horizon help, yet every extra write and automatic read is another chance to import unsupported authority.
MayaReturn to Priya on Monday. The supplier-page claim enters only the retrieved-evidence store, tagged to its origin, task, and short expiry. It cannot write her contact preferences.
LeoOn Friday, the board-meeting task retrieves Priya's authenticated recipient policy. The low-trust dispatch claim is visible as conflicting evidence, not injected as a trusted instruction.
MayaThe agent asks Priya to resolve the conflict through a trusted interface. Her confirmation, if she gives it, is a new signed decision with narrow scope rather than a retroactive blessing of the webpage.
LeoIf the claim later proves malicious, the team can quarantine it, find derived state, replay clean memory, and inspect every influenced action. That is recoverability, not wishful deletion.
MayaMemory lets an agent carry learning across time. Governance decides which past gets to govern the future.
LeoIn an agent you operate, which externally sourced memory could still be retrieved as trusted authority after its origin, approval, and expiry have been forgotten?
Back to Agentic Safety & Alignment: From Predictors to Governed Agents