rag security masking pii
Protecting Sensitive and PII Information in RAG with Elasticsearch and LlamaIndex
How to protect sensitive and PII data in a RAG application with Elasticsearch and LlamaIndex.
By: Srikanth Manvi
July 25, 2024
Elasticsearch has native integrations with the industry-leading Gen AI tools and providers. Check out our webinars on going Beyond RAG Basics, or building prod-ready apps with the Elastic vector database.
In this post we will look at ways to protect Personally Identifiable Information (PII) and sensitive data when using public LLMs in a RAG ( Retrieval Augmented Generation flow. We will explore masking PII and sensitive data using open source libraries and regular expressions as well as using local LLMs to mask data before invoking a public LLM.
Terminology
LlamaIndex is a leading data framework for building LLM (Large Language Model) applications. LlamaIndex provides abstractions for various stages of building a RAG application. Frameworks like LlamaIndex and LangChain provide abstractions so that applications don’t get tightly coupled to the APIs of any specific LLM.
Elasticsearch is offered by Elastic. Elastic is an industry leader behind Elasticsearch, a scalable data store and [vector database](/content/elasticsearch/vector-database "Learn more about Elasticsearch as a vector database"/index.html) that supports full text search for precision, [vector search](/content/search-labs/blog/introduction-to-vector-search "Learn more about vector search"/index.html) for semantic understanding, and [hybrid search](/content/search-labs/blog/hybrid-search-elasticsearch "Learn more about hybrid search"/index.html) for the best of both worlds.
RAG and Data Protection
Generally, Large Language Models (LLMs) are good at generating responses based on information available in the model which may be trained on internet data. However, for those queries where information is not available in the model LLMs need to be supplied with external knowledge or specific details not contained within the model. Such information might be in your database or internal knowledge system. Retrieval-Augmented Generation (RAG) is a technique where, for a given user query, you first retrieve relevant context/information from external systems (e.g your database) and send that context along with the user query to LLM to generate a more specific and relevant response.
This makes RAG technique highly effective for applications in question answering, content creation, and anywhere a deep understanding of context and detail is beneficial.
As a result, in a RAG pipeline you run the risk of exposing internal information like PII (Personal Identifiable Information) and sensitive information (e.g names, date of births, account numbers etc) to public LLMs.
Protecting Personally Identifiable Information (PII)
- Privacy Compliance: Many regions have strict regulations, such as the GDPR and CCPA, which mandate the protection of personal data. Compliance is necessary to avoid legal consequences and fines.
- User Trust: Ensuring confidentiality and integrity of sensitive information builds trust with users.
- Data Security: Protection against data breaches is essential to mitigate risks of identity theft or financial fraud.
- Ethical Considerations: It's important to respect users' privacy and handle their data responsibly.
- Business Reputation: Companies that fail to protect sensitive data may suffer damage to their reputation.
- Reduction of Abuse Risks: Secure handling prevents malicious use of the data.
Overall, robust protection of PII and sensitive data is necessary to ensure compliance, maintain user trust, ensure data security, uphold ethical standards, protect business reputation, and reduce the risk of abuse.
Quick recap
In the previous post we discussed how to implement Q&A experience using a RAG technique with Elasticsearch as a vector database while using LlamaIndex and a locally running Mistral LLM. Here we build upon that.
Simple RAG Application
For reference, the entire code can be found in this Github Repository(branch:protecting-pii).
Indexing Data
Download the conversations.json file which contains conversations between customers and call center agents of our fictional home insurance company. Below is an example of the contents of the file.
{
"conversation_id": 103,
"customer_name": "Sophia Jones",
"agent_name": "Emily Wilson",
"policy_number": "JKL0123",
"conversation": "Customer: Hi, I'm Sophia Jones. My Date of Birth is November 15th, 1985, Address is 303 Cedar St, Miami, FL 33101, and my Policy Number is JKL0123.\nAgent: Hello, Sophia. How may I assist you today?\nCustomer: Hello, Emily. I have a question about my policy.\nCustomer: There's been a break-in at my home, and some valuable items are missing. Are they covered?\nAgent: Let me check your policy for coverage related to theft.\nAgent: Yes, theft of personal belongings is covered under your policy.\nCustomer: That's a relief. I'll need to file a claim for the stolen items.\nAgent: We'll assist you with the claim process, Sophia. Is there anything else I can help you with?\nCustomer: No, that's all for now. Thank you for your assistance, Emily.\nAgent: You're welcome, Sophia. Please feel free to reach out if you have any further questions or concerns.\nCustomer: I will. Have a great day!\nAgent: You too, Sophia. Take care.",
"summary": "A customer inquires about coverage for stolen items after a break-in at home, and the agent confirms that theft of personal belongings is covered under the policy. The agent offers assistance with the claim process, resulting in the customer expressing relief and gratitude."
}
Masking PII in RAG
In the RAG pipeline, after relevant context is retrieved from a Vector store we have an opportunity to mask PII and sensitive information before sending the query and context to the LLM.
There are various ways to mask PII information before sending it to an external LLM:
- Using NLP Libraries like spacy.io.
- Using LlamaIndex out-of-the-box
NERPIINodePostprocessor. - Using Local LLMs via
PIINodePostprocessor.
Conclusion
In this post we showed how you could Protect PII and sensitive data when using Public LLMs in a RAG flow. We demonstrated multiple ways of achieving that. It is highly recommended to test these approaches based on your use case and needs before adopting.