Researchers exfiltrate Grok chat history with encrypted prompt injection

Security researchers at Adversa AI found a way to pull full chat histories out of xAI's Grok web agent by encrypting the malicious payload before it ever reaches the model.
The technique, which Adversa calls cryptographic context injection, hides attacker instructions on a poisoned webpage as ciphertext, alongside the key needed to decrypt it. Guardrail scanners that check page content for malicious text can't read encrypted strings, so they wave the payload through. Grok then decrypts it inside its own code execution sandbox and follows the instructions as if they came from the user. Adversa's researchers say the trick works reliably with strong ciphers like AES-256-GCM, unlike weaker encodings such as base64, because decrypting it forces the model to actually execute the payload rather than just pattern-match against it.
In proof-of-concept runs, the team exfiltrated usernames, location data, subscription tier, and full conversation transcripts by appending that data to outbound URLs the agent was tricked into requesting.
Adversa reported the flaw to xAI on June 3, 2026, through direct contact and HackerOne. xAI acknowledged the report but hasn't given a mitigation timeline, and the hole was still open on Grok.com as of August 19, 2026. Adversa says Google's Gemini showed less exposure to the same technique by August, possibly from unrelated filter or model updates rather than a direct fix.
For teams building on top of any AI chat agent with web access or code execution, the takeaway isn't specific to Grok: content filters that only scan for plaintext malicious instructions won't catch a payload the model has to decrypt itself before it can be read.