Researchers can now reverse-engineer LLM prompts from output text with near-perfect accuracy

The Decoder The Decoder

https://the-decoder.com/wp-content/uploads/2026/08/prompt_extracting.png" style="height: auto; margin-bottom: 10px;" width="1376" />


Researchers at IIT Bombay and Adobe Research have built an inverse language model that reconstructs the original prompt from an LLM's output with near-perfect accuracy.

Their method, called "Previous-Token Prediction," doesn't need access to model weights and works across different models.

For companies relying on proprietary system prompts, this could be a serious security risk.


The article https://the-decoder.com/researchers-can-now-reverse-engineer-llm-prompts-from-output-text-with-near-perfect-accuracy/">Researchers can now reverse-engineer LLM prompts from output text with near-perfect accuracy appeared first on https://the-decoder.com">The Decoder.

Read full article at The Decoder →