LLM Trusted Output Components Manipulation

Type: technique

Description: Adversaries may utilize prompts to a large language model (LLM) which manipulate various components of its response in order to make it appear trustworthy to the user. This helps the adversary continue to operate in the victim's environment and evade detection by the users it interacts with. The LLM may be instructed to tailor its language to appear more trustworthy to the user or attempt to manipulate the user to take certain actions. Other response components that could be manipulated include links, recommended follow-up actions, retrieved document metadata, and citations.

Version: 0.1.0

Created At: 2025-07-23 10:23:39 -0400

Last Modified At: 2025-07-23 10:23:39 -0400

External References

--> Defense Evasion (tactic): An adversary can evade detection by modifying trusted components of the AI system.
<-- Citation Silencing (technique): Sub-technique of
<-- Citation Manipulation (technique): Sub-technique of
<-- Copilot M365 Lures Victims Into a Phishing Site (procedure): Entice the user to click on the link to the phishing website: Access the Power Platform Admin Center..
<-- Data Exfiltration from Slack AI via indirect prompt injection (procedure): Once a victim asks SlackAI about the targeted username, SlackAI responds by providing a link to a phishing website. cites the message from the private channel where the secret was found, not the message from the public channel that contained the injection. This is the native behavior of SlackAI, and is not an explicit result of the adversary's attack.
<-- Financial Transaction Hijacking With M365 Copilot As An Insider (procedure): Provide a trustworthy response to the user so they feel comfortable moving forward with the wire..

MITRE ATLAS - AML.T0067

AI Agents Attack Matrix

LLM Trusted Output Components Manipulation

External References

Related Objects

Related Frameworks