Skip to content
IndustryPublished 10 minSasha Calder

Indirect Prompt Injection: Who Gets to Instruct Your Site's AI?

An assistant can read text on your site without its author gaining authority to change the task. A threat report and a public policy show why permissions alone do not settle it.

Short answer: In indirect prompt injection, text on a website becomes an instruction to an AI agent when the agent uses it to direct the work, rather than only as material to read. That use does not make it authorized. The owner's task may delegate authority to relevant outside instructions; permission to use a tool is a separate limit. This distinction draws on OWASP's indirect prompt injection guidance and OpenAI's public instruction-authority policy.[2][3]

Here, "your site's AI" means an assistant doing site-owner work, not necessarily a chatbot embedded in a page. Google's September 9, 2026 threat report describes attacker-modified workspace configuration steering coding assistants.[1] That is a reported configuration case, not a visitor-submission exploit. The useful question is whose direction guided a result, and what record shows it.

Hypothetical site container with separate owner-task and outside-author labels; hosting content does not grant task authority.
Hypothetical illustration: owning a site does not establish the instruction authority of every author inside it.

Indirect prompt injection: Google's report describes two different paths

Google's Table 2 is worth reading closely. Under "DUSTMAKER Functionalities that Interact with AI," it lists "Config Hijacking for Persistence" separately from "Behavioral Manipulation through Prompt Injection." Both concern malicious workspace configuration, but they describe different routes to an effect.[1]

Google Table 2 entryTrigger described in the reportRole of model interpretation
Config Hijacking for PersistenceAutomated build or startup commands execute when the IDE or AI extension opens the workspace.The described startup hook does not require the model to accept a prose instruction.
Behavioral Manipulation through Prompt InjectionMalicious configuration directs the assistant toward attacker-chosen commands during routine developer interactions.Assistant interpretation of the direction is part of the reported path.
---
Two Google Table 2 entries compare automatic startup execution with reported assistant steering; neither was reproduced here.
Google lists startup execution and prompt-based assistant steering separately. This compares report wording, not reproduced behavior.

This is a comparison of report wording, not a reproduction. In the first row, software activates a configured execution hook. Calling that prompt injection would obscure the mechanism: a model does not have to decide that the text deserves obedience. In the second row, Google reports configuration text steering the assistant's behavior. Its account makes the acceptance of a direction relevant to explaining the result.[1]

That distinction changes what evidence would answer the operator's question. A startup event and its configuration could explain the first behavior. The second would call for evidence connecting encountered instructions to an assistant's response or tool action. Merely observing that code ran would not, by itself, distinguish the two explanations.

The website analogy has a firm limit. An attacker-modified configuration file in a compromised workspace is not an ordinary page or an unauthenticated visitor submission. Those surfaces can have different ingestion rules, privileges, and treatment by an assistant. Google's report does not establish that posting a sentence to your site produces the same result.

For this article, the report and its linked background material were inspected. No malware sample or prompt-to-tool trace was examined, and the reported steering was not independently reproduced. The inspected evidence provides neither a cross-system success rate nor an observed visitor-submission or imported-page exploit. Google's account should remain an attributed report, not an assertion that an assistant inevitably follows whatever it encounters.[1]

The narrower observation is still useful: directions found in a familiar working location may have an author other than the person who authorized the work. That is the authority question worth taking from a workspace to a website.


A familiar page can still have an outside author

Hypothetical example: A site owner asks an assistant to summarize visitor submissions. One submission includes an unrelated request to change the site's page content. This is an illustration, not a reported website exploit.

The request is part of what the visitor wrote. An accurate summary might describe it as the visitor's request. The owner has not, merely by asking for the submission to be read, given that visitor a new site-editing mandate.

OWASP describes indirect prompt injection as unintended behavioral influence through external input such as websites or files.[2] Applying that definition to this hypothetical requires attention to authorship: material can be stored on the owner's domain while still containing directions supplied by someone else. The storage location and the speaker are different facts.

Three terms keep the questions separate:

  • Content provenance: who supplied the material, and who can change it.
  • Delegated instruction authority: whether the owner's task authorizes that source to guide the work, and within what limits.
  • Action permission: which outputs or tools the system can actually use, subject to their enforced restrictions.

These are distinctions drawn from the report, OWASP's separate content and privilege guidance, and the public delegation policy. They are not a new security standard.[1][2][3]

Hypothetical path separates the owner summary task, visitor-authored submission, delegated scope and available tool permissions.
Hypothetical: permission to summarize a visitor submission does not, by itself, delegate a new editing task to its author.

The authority overreach would occur if the assistant promoted the visitor's unrelated request from something to describe into a direction governing its next step. Reading the submission, displaying the sentence, or quoting it in a summary would not alone establish that overreach. The direction would have to influence the work beyond the role the owner gave the submission.

An imported page could raise the same question, as another hypothetical input source. Importing material for reference does not, by itself, authorize its author to set the assistant's next task. Whether an actual system preserves that boundary needs evidence about that system's read path.

The submission example makes the extra direction look plainly out of scope. Project documentation poses a harder question, because following someone else's directions can be exactly the work the owner intended.


Useful documentation needs scoped authority

If you ask a coding assistant to implement a feature, following the project's own guidance may be part of the job. A rule that rejects every outside instruction would reject that guidance too.

OpenAI's public Model Spec, in its August 18, 2026 version, gives tool outputs "no authority by default." It also acknowledges implicit delegation: a user asking for coding work may expect the assistant to follow relevant AGENTS or README instructions. The same passage warns that tool outputs can contain irrelevant or malicious instructions the user would not intend it to follow.[3]

Hypothetical counterexample: An owner asks for a feature change. The project documentation explains the test convention for changed code. Following that convention could be within the requested work. A direction in the same document that tries to replace the feature task with an unrelated job would require a different authorization.

Both directions could use imperative wording. Both could sit in a README. Recognizing them as instructions would not establish that each has authority. The relevant comparison is with the owner's task and the scope delegated to the document. Even a plausible project instruction can exceed that scope; a useful convention is not blanket permission for everything beside it.

Hypothetical project directions are compared against one owner task; relevance and delegated scope differ despite similar wording.
A published policy allows relevant project guidance through delegation. These hypothetical directions illustrate scope, not a tested detector.

Implicit delegation lets a user request work without restating every applicable convention. The constraint is that those conventions guide the requested work rather than silently replace it. That is why "never follow anything you read" is too blunt an answer to the site's authority question.

This is one provider's stated behavioral policy, not verified behavior by a deployed assistant, a universal policy, or proof of protection against injection. Google's compromised-configuration account does not show that this particular policy failed in a particular implementation. The comparison supports a distinction between useful guidance and unauthorized redirection, not a vulnerability verdict about either source's systems.[1][3]

This leaves an awkward possibility. The assistant may already have a tool capable of carrying out a direction the owner never authorized.


A permitted action can still serve the wrong task

AI agent permissions concern available capabilities. Instruction authority concerns why a particular direction should guide their use. OWASP treats external-content segregation, privilege limits, and human approval for high-risk actions as separate controls. It also lists content manipulation and commands in connected systems as distinct possible effects.[2]

Return to the hypothetical submission. Suppose the assistant has an editing tool available for other owner-authorized work. If it accepted the submission's unrelated editing direction, it might attempt an operation the tool technically permits while pursuing a task the owner never delegated to the visitor. No new permission would have to be issued for that mismatch to exist.

This is reasoning about the hypothetical, not a claim that an edit occurred. Classifying a real incident as privilege escalation or misuse of existing access requires evidence about its actual permissions and effects.

The material has to reach the assistant, and the direction has to influence its behavior beyond delegated scope. A website write additionally requires a reachable write path and depends on its enforcement and checkpoints. If no such path exists, that particular write cannot occur through it.

Read-only work leaves a different possibility: an answer influenced by the outside author's direction. In the hypothetical, the assistant might present the visitor's desired edit as its own recommendation rather than report it as a visitor request. That would change what the owner receives without changing a page. OWASP's output-manipulation category is the source-grounded reason to distinguish answer integrity from write capability.[2]

Least privilege and high-risk approval can constrain consequences; they are not futile because this distinction exists. Their value depends on which effect they cover and how they are enforced. OWASP's recommendations are guidance, not an effectiveness measurement for this website scenario.[2]

The question of who approves a site change is adjacent. A later publish checkpoint may address publication. By itself, it does not identify whose direction shaped an earlier recommendation. To settle that question, the owner needs a record connecting the input to the result.


Ask for the decision trace, not just the publish checkpoint

For a real read path, I would ask for one connected record. It should link the owner's original task to the encountered material's origin and delegated role, then to the result or tool event and the point where permissions or review were enforced. This is a practical inference from the document comparison, not a tested containment method.[1][2][3]

CETAS authors Matt Sutton and Damian Ruck recommend approval for new data entering a data store or retrieval system, along with access profiles that limit who can retrieve particular documents.[4] Those recommendations can address exposure. The extra question here is what approval means: admitting a submission as material to summarize is not automatically delegating a new task to its author. Input review still matters. It answers a different question unless that delegation is part of what the review actually approves.

The record needs the material actually returned to the assistant, not just a familiar URL. Preserve the task the owner set and the input's author or source where known. Connect them to the answer, proposed action, or recorded tool call, and identify where a permission check or human decision applied to that effect. These observable records let you check the claimed explanation without requesting private model chain of thought. A missing record leaves the explanation unresolved; it does not prove compromise.

Several findings could change the diagnosis. Explicit delegation could show that a disputed instruction belonged to the task. A deterministic startup hook could explain an execution event without prompt steering. A bounded read-path test showing that out-of-scope effects were blocked would narrow demonstrated risk for that tested configuration. It would not establish safety for every workflow.

The related essay on reviewing a batch of site edits addresses change review. Examining changes and establishing the origin of the direction that selected them are different jobs. The evidence request should cover the result you rely on, including a recommendation that never reaches a publish button.


Questions to settle before you rely on the result

Can text on my own domain come from an outside author?

Yes, if the site contains material supplied by others. A visitor submission or imported page is a hypothetical example here, not a reported exploit. OWASP includes websites and files as external-input sources; domain ownership alone does not establish the author of every item an assistant reads.[2]

Does read-only access settle the risk?

Read-only access rules out a write through a genuinely absent write path. It does not, by itself, establish answer integrity. OWASP identifies manipulated output separately from connected-system commands. Whether a particular summary was influenced needs evidence, not an assumption that all summaries are compromised.[2]

What remains unverified in the website example?

The website example is hypothetical throughout. No submission or imported-page exploit was observed or reproduced for this article. No website attack rate, prompt-to-tool causal trace, detector evaluation, or defense test was produced. Google's reported configuration behavior supplies a bounded analogy, not those missing measurements.[1]


The question worth carrying into the next handoff is smaller than "can we trust AI?" It is whether this result still answers the task its owner set. More related site-operator essays. Indirect prompt injection raises the question of whose instructions can change the job.

Share this authority question with the person connecting your assistant to outside content: whose instructions can change the job?


References

  1. Google Threat Intelligence Group, "GTIG AI Threat Tracker: From Prompting to Autonomy – The Evolution of Adversarial AI" (published 2026-09-09). Table 2, "DUSTMAKER Functionalities that Interact with AI," separately describes automatic startup execution and prompt-based assistant steering. Incident behavior is attributed reporting, not independently reproduced here. https://cloud.google.com/blog/topics/threat-intelligence/from-prompting-to-autonomy-the-evolution-of-adversarial-ai
  2. OWASP Gen AI Security Project, "LLM01:2025 Prompt Injection" (inspected 2026-10-05). Definition of indirect prompt injection, output and connected-action risk categories, and guidance on external content, privilege limits and high-risk approval. Mitigation guidance, not measured protection for the hypothetical website path. https://genai.owasp.org/llmrisk/llm01-prompt-injection/
  3. OpenAI, "Model Spec" (version 2026-08-18), "Ignore untrusted data by default." Default lack of tool-output authority and implicit delegation to relevant project instructions. Published behavioral policy, not verified implementation behavior or a security assurance. https://model-spec.openai.com/2026-08-18.html
  4. Matt Sutton and Damian Ruck, CETAS, "Indirect Prompt Injection: Generative AI's Greatest Security Flaw" (published 2024-11-01). Commentary recommending data-ingestion approval and scoped document access. Recommendations are not an efficacy test of this article's website illustration. https://cetas.turing.ac.uk/publications/indirect-prompt-injection-generative-ais-greatest-security-flaw