Securing Retrieval-Augmented Generation Against Data Poisoning And Indirect Prompt Injection

Uncategorized

Authors: Ravindra Babu Annam

Abstract: RAG enables large language models to give responses that can be enhanced with knowledge extracted from external documents. However, the retrieval of documents exposes the language models to security threats because the information retrieved may be poisoned and used for manipulation of the model. Current security mechanisms are tailored for particular attacks and hence cannot provide end-to-end protection of the entire retrieval to generation process. This study proposes a Secure Retrieval-Augmented Generation Framework (SRAGF) that enables the protection of RAG-based models against data poisoning and prompt injection attacks. It conducts document relevance assessment, source credibility evaluation, semantic anomalies recognition, suspicious pattern recognition, poisoning detection, injection detection, risk classification, context sanitization, knowledge and instruction separation, and output security evaluation. Security scores are calculated for the retrieved documents, which are classified accordingly and included in the context of generation. Documents that pose high risk are stored separately, the ones presenting medium risk undergo sanitization, and those that do not present any threat are retained for context generation. The baseline frameworks for the comparative analysis are Wang-TAD and Neural Cleanse. In the prepared evaluation results, SRAGF exhibits higher detection rates, precision, recall, and F1-score compared to the two baseline frameworks while lowering attack.

DOI: http://doi.org/10.5281/zenodo.23184152

× How can I help you?