What Is Retrieval-Augmented Generation (RAG)?
Imagine you open your bank's mobile application and ask its AI assistant:
“What are the foreclosure charges on my home loan?”
A large language model such as an LLM may understand what a home loan is and what foreclosure means. It may even know common banking practices. But that does not mean it knows your bank's current foreclosure policy, the conditions applicable to a particular loan product, or whether the policy changed recently.
If the model answers only from the knowledge it learned during training, it could provide outdated information, information from another bank, or an answer that simply sounds plausible. In banking, where customers may make financial decisions based on the response, that is not acceptable.
A better approach is to first search the bank's approved knowledge sources.
The AI application could search the bank's current home-loan documentation and retrieve a passage explaining the applicable foreclosure rules. It can then provide that information to the language model along with the customer's question.
The model is no longer being asked:
“What do you know about home-loan foreclosure charges?”
Instead, it is effectively being asked:
“Using this retrieved section from our bank's current home-loan policy, answer the customer's question.”
That change is the basic idea behind Retrieval-Augmented Generation (RAG).
Retrieval-Augmented Generation is an AI architecture in which an application retrieves relevant information from an external knowledge source and provides that information to a generative model as context before the model produces its response.
The model therefore does not have to depend entirely on information stored in its parameters during training. It can work with information retrieved when the question is asked.
In our banking example, the flow can be thought of simply as:
Customer question → Search bank knowledge → Retrieve relevant policy → Give policy to the LLM → Generate grounded answer
This distinction is important. RAG does not make the language model itself permanently learn the bank's documents. Instead, it gives the model the relevant information at the time it needs it.
That makes RAG particularly useful for applications involving frequently changing, organization-specific, or private knowledge—such as banking policies, product documentation, operating procedures, internal knowledge bases, customer-support information, and regulatory guidance.
Throughout this learning track, we will build this simple idea into a complete engineering system. We will learn how documents are prepared, how relevant information is found, how that information is supplied to an LLM, how answers are evaluated, and eventually how a RAG system can be designed for production.
For now, remember one idea:
RAG allows an LLM application to retrieve relevant external knowledge before generating an answer.
Why Can’t We Just Ask the LLM?
It is reasonable to ask why we need RAG at all. If a large language model can already answer questions about loans, credit cards, banking terminology, and financial concepts, why not simply send the customer's question directly to the model?
The problem is that general knowledge is not the same as authoritative, current, organization-specific knowledge.
Consider a customer who asks:
“What interest rate will I get if I open a fixed deposit today?”
An LLM may understand fixed deposits and may even know typical interest-rate ranges. But the correct answer depends on information that can change frequently: the bank's current rate card, deposit duration, customer category, applicable product, effective date, and other conditions.
A response based only on the model's general knowledge could therefore be outdated even if it sounds completely reasonable.
Now consider another question:
“Why was a fee charged to my account last month?”
This is a different problem. The answer may depend on the customer's account type, transaction history, fee schedule, and the terms applicable to that account. A general-purpose LLM was not trained with access to that customer's private banking records, nor should such information simply become part of a public model's general knowledge.
There is also a third problem: language models can generate plausible answers when they do not have sufficient evidence.
Suppose the bank's actual policy says that a particular charge is waived under specific conditions. If the model does not have that policy available, it may still attempt to answer based on patterns learned during training. The response can be fluent and confident while still being wrong.
In banking, this distinction matters. A grammatically perfect answer is not enough. The application needs to answer from the right information.
RAG changes the problem. Instead of expecting the LLM to already know every policy, product rule, fee schedule, procedure, and document, the application retrieves the relevant information when it is needed and provides it to the model as context.
The goal is therefore not to make the LLM memorize the entire bank.
The goal is to help the LLM find and use the right knowledge for the question being asked.
How Does RAG Work?
Now that we understand why an LLM cannot always be trusted to answer from its own knowledge, we can look at what happens when RAG is introduced.
Consider a customer asking the bank's AI assistant:
“What documents do I need to apply for a home loan?”
The bank may already have the correct answer in its approved home-loan documentation. The challenge is finding the relevant information and giving it to the language model at the right time.
A basic RAG system handles this in three main stages: retrieve, augment, and generate.
1. Retrieve the Relevant Information
When the customer submits the question, the application first searches the bank's knowledge base.
That knowledge base might contain home-loan documentation, eligibility rules, product information, FAQs, policy documents, and other approved material.
The retrieval system tries to find the pieces of information most relevant to the customer's question.
For example, it might retrieve a section stating that applicants need identity proof, address proof, income documents, bank statements, and property-related documents.
The important point is that the system does not send every document the bank owns to the LLM. It tries to retrieve only the information that is useful for answering the current question.
2. Augment the LLM's Context
The retrieved information is then added to the context provided to the language model.
Conceptually, the application might construct an instruction like this:
Customer question:
“What documents do I need to apply for a home loan?”
Retrieved bank information:
“Home-loan applicants must provide identity proof, address proof, income documentation, recent bank statements, and applicable property documents.”
Instruction:
“Answer the customer's question using the provided bank information.”
The LLM now has something it did not have when the customer originally asked the question: relevant, bank-specific context.
This is the “augmented” part of Retrieval-Augmented Generation.
3. Generate the Answer
Finally, the LLM uses the customer's question and the retrieved context to produce a natural-language response.
It might answer:
“For a home-loan application, you will generally need identity and address proof, income documents, recent bank statements, and relevant property documents. The exact requirements may depend on the loan and applicant type.”
The LLM is still responsible for understanding the question and generating a useful response, but the factual foundation for that response came from the bank's knowledge source.
The complete flow can therefore be represented as:
Customer Question → Retrieve Bank Knowledge → Add Relevant Context → LLM → Grounded Answer
This separation is fundamental to understanding RAG.
The retrieval system is responsible for finding useful knowledge, while the language model is responsible for using that knowledge to construct the response.
Later, we will discover that each of these stages contains its own engineering challenges. Documents must be processed correctly, relevant passages must be found accurately, the best context must be selected, and the model must be instructed to use that context appropriately.
But at the beginner level, the most important mental model is simple:
RAG first finds the knowledge, then gives that knowledge to the LLM before asking it to answer.
Where Does the Knowledge Come From?
A RAG system needs a trusted source of information to retrieve from. In a banking application, that knowledge does not come from the language model itself. It comes from information the bank has chosen to make available to the RAG system, such as product documentation, operating procedures, customer FAQs, loan policies, fee schedules, regulatory guidance, and internal knowledge.
Suppose a customer asks:
“Can I make a partial prepayment on my home loan?”
The correct answer may already exist in the bank's home-loan documentation. The challenge is that the relevant information could be buried inside a document that is dozens or even hundreds of pages long. The customer needs only the small portion that answers the question, so the RAG system needs a way to prepare and organize the bank's knowledge for efficient retrieval.
Documents Are the Starting Point
The source information can come from many places, including PDF policy documents, web pages, product manuals, internal knowledge bases, FAQs, databases, and approved operational documents.
Imagine that the bank has a 100-page document called Home Loan Product and Servicing Policy. It contains information about eligibility, required documents, interest rates, repayment rules, partial prepayment, foreclosure, charges, and many other topics.
When a customer asks about partial prepayment, sending the entire 100-page document to the LLM would be inefficient because most of the document has nothing to do with the question. Instead, a RAG system prepares the document so that the relevant part can be found when it is needed.
Large Documents Are Broken Into Smaller Pieces
One common step in preparing documents is dividing them into smaller pieces called chunks. Rather than treating the entire home-loan policy as one enormous piece of information, the system can work with smaller, meaningful sections.
For example:
Home Loan Policy → Eligibility → Required Documents → Interest Rates → Repayment Rules → Partial Prepayment → Foreclosure → Charges and Fees
Now, when the customer asks about paying part of a home loan early, the retrieval system can try to find the Partial Prepayment information instead of sending unrelated sections about eligibility, interest rates, or late-payment charges to the LLM.
This process of deciding how documents should be divided is called chunking. The way chunks are created can significantly affect retrieval quality, so we will explore chunking in much greater depth later in the learning track.
How Can the System Find Similar Meaning?
Breaking a document into smaller pieces solves only part of the problem. The system must still determine which chunk is relevant to the customer's question, and this becomes interesting when the customer and the bank use different words to describe the same idea.
For example, the customer might ask:
“Can I pay some of my housing loan early?”
The bank's official document might instead contain the phrase:
“Partial prepayment of the home-loan outstanding principal...”
The wording is different, but the meaning is closely related. A useful retrieval system should be able to recognize that relationship rather than depending entirely on exact keyword matches.
Modern RAG systems often use embeddings for this purpose. An embedding represents text as a collection of numbers, called a vector, that captures aspects of its meaning. This allows pieces of text with similar meanings to be compared mathematically even when they do not contain exactly the same words.
At this stage, you do not need to understand the mathematics behind embeddings. The important idea is simply that embeddings help computers compare text by meaning rather than only by matching words.
The Prepared Knowledge Becomes Searchable
Once the bank's documents have been divided into useful chunks and represented appropriately, those chunks can be stored in a system that allows them to be searched later. Depending on the architecture, this could involve a vector database, a traditional search engine, a relational database with vector capabilities, or a combination of retrieval technologies.
At this point, the specific database technology is less important than understanding the transformation:
Bank Documents → Useful Chunks → Searchable Representations → Knowledge Available for Retrieval
Now return to our customer:
“Can I pay some of my housing loan early?”
The application can search the prepared knowledge and discover that the customer's question is closely related to Home Loan Policy → Partial Prepayment. That section can then be supplied to the LLM as context, allowing the model to construct an answer based on the bank's actual policy rather than relying only on its general knowledge.
This also gives us an important distinction in a RAG system. Preparing and storing knowledge is part of the indexing or ingestion pipeline, while searching that prepared knowledge when a user asks a question is part of the retrieval or query pipeline.
A complete RAG system needs both.
Revision Notes
- RAG stands for Retrieval-Augmented Generation: it allows an LLM application to retrieve relevant external knowledge before generating an answer.
- An LLM does not automatically know everything: its existing knowledge may be outdated, incomplete, or missing private and organization-specific information.
- RAG grounds the answer in relevant external knowledge: instead of asking the LLM to rely only on what it learned during training, the application gives it useful information related to the current question.
- RAG works in three basic stages: Retrieve → Augment → Generate. First, find relevant information; then add it to the LLM's context; finally, use that context to generate the answer.
- External knowledge can come from many sources: PDFs, policies, product documentation, FAQs, websites, databases, internal knowledge bases, and other approved information.
- Large documents can be divided into smaller chunks: this helps the system retrieve only the information relevant to a particular question instead of sending an entire document to the LLM.
- Embeddings help compare meaning: they represent text numerically so that related information can be found even when the user's wording differs from the wording in the source document.
- Prepared knowledge must be searchable: document chunks and their representations can be stored in systems such as vector databases, search engines, relational databases with vector capabilities, or combinations of these technologies.
- A RAG system has two important sides: the indexing or ingestion pipeline prepares and stores knowledge, while the retrieval or query pipeline searches that knowledge when a question is asked.
- The central idea of RAG is simple: the LLM does not need to memorize all available knowledge. The application retrieves the right information at the right time and gives it to the LLM as context for generating a grounded answer.
Share this article
About the author
Raj Kumar
Continue exploring
Go deeper into practical AI engineering systems.