rag security masking pii

Protecting Sensitive and PII Information in RAG with Elasticsearch and LlamaIndex

How to protect sensitive and PII data in a RAG application with Elasticsearch and LlamaIndex.

By: Srikanth Manvi
July 25, 2024

Elasticsearch has native integrations with the industry-leading Gen AI tools and providers. Check out our webinars on going Beyond RAG Basics, or building prod-ready apps with the Elastic vector database.

In this post we will look at ways to protect Personally Identifiable Information (PII) and sensitive data when using public LLMs in a RAG ( Retrieval Augmented Generation flow. We will explore masking PII and sensitive data using open source libraries and regular expressions as well as using local LLMs to mask data before invoking a public LLM.

Terminology

LlamaIndex is a leading data framework for building LLM (Large Language Model) applications. LlamaIndex provides abstractions for various stages of building a RAG application. Frameworks like LlamaIndex and LangChain provide abstractions so that applications don’t get tightly coupled to the APIs of any specific LLM.

Elasticsearch is offered by Elastic. Elastic is an industry leader behind Elasticsearch, a scalable data store and [vector database](/content/elasticsearch/vector-database "Learn more about Elasticsearch as a vector database"/index.html) that supports full text search for precision, [vector search](/content/search-labs/blog/introduction-to-vector-search "Learn more about vector search"/index.html) for semantic understanding, and [hybrid search](/content/search-labs/blog/hybrid-search-elasticsearch "Learn more about hybrid search"/index.html) for the best of both worlds.

RAG and Data Protection

Generally, Large Language Models (LLMs) are good at generating responses based on information available in the model which may be trained on internet data. However, for those queries where information is not available in the model LLMs need to be supplied with external knowledge or specific details not contained within the model. Such information might be in your database or internal knowledge system. Retrieval-Augmented Generation (RAG) is a technique where, for a given user query, you first retrieve relevant context/information from external systems (e.g your database) and send that context along with the user query to LLM to generate a more specific and relevant response.

This makes RAG technique highly effective for applications in question answering, content creation, and anywhere a deep understanding of context and detail is beneficial.

As a result, in a RAG pipeline you run the risk of exposing internal information like PII (Personal Identifiable Information) and sensitive information (e.g names, date of births, account numbers etc) to public LLMs.

Protecting Personally Identifiable Information (PII)

Overall, robust protection of PII and sensitive data is necessary to ensure compliance, maintain user trust, ensure data security, uphold ethical standards, protect business reputation, and reduce the risk of abuse.

Quick recap

In the previous post we discussed how to implement Q&A experience using a RAG technique with Elasticsearch as a vector database while using LlamaIndex and a locally running Mistral LLM. Here we build upon that.

Simple RAG Application

For reference, the entire code can be found in this Github Repository(branch:protecting-pii).

Indexing Data

Download the conversations.json file which contains conversations between customers and call center agents of our fictional home insurance company. Below is an example of the contents of the file.

{
"conversation_id": 103,
"customer_name": "Sophia Jones",
"agent_name": "Emily Wilson",
"policy_number": "JKL0123",
"conversation": "Customer: Hi, I'm Sophia Jones. My Date of Birth is November 15th, 1985, Address is 303 Cedar St, Miami, FL 33101, and my Policy Number is JKL0123.\nAgent: Hello, Sophia. How may I assist you today?\nCustomer: Hello, Emily. I have a question about my policy.\nCustomer: There's been a break-in at my home, and some valuable items are missing. Are they covered?\nAgent: Let me check your policy for coverage related to theft.\nAgent: Yes, theft of personal belongings is covered under your policy.\nCustomer: That's a relief. I'll need to file a claim for the stolen items.\nAgent: We'll assist you with the claim process, Sophia. Is there anything else I can help you with?\nCustomer: No, that's all for now. Thank you for your assistance, Emily.\nAgent: You're welcome, Sophia. Please feel free to reach out if you have any further questions or concerns.\nCustomer: I will. Have a great day!\nAgent: You too, Sophia. Take care.",
"summary": "A customer inquires about coverage for stolen items after a break-in at home, and the agent confirms that theft of personal belongings is covered under the policy. The agent offers assistance with the claim process, resulting in the customer expressing relief and gratitude."
}

Masking PII in RAG

In the RAG pipeline, after relevant context is retrieved from a Vector store we have an opportunity to mask PII and sensitive information before sending the query and context to the LLM.

There are various ways to mask PII information before sending it to an external LLM:

  1. Using NLP Libraries like spacy.io.
  2. Using LlamaIndex out-of-the-box NERPIINodePostprocessor.
  3. Using Local LLMs via PIINodePostprocessor.

Conclusion

In this post we showed how you could Protect PII and sensitive data when using Public LLMs in a RAG flow. We demonstrated multiple ways of achieving that. It is highly recommended to test these approaches based on your use case and needs before adopting.