Ask ChatGPT about your product catalog, your internal policies or a specific customer's history. It will answer with total confidence, fluent prose… and made-up data. LLMs know a lot about everything and nothing about your company: their knowledge froze on training day.
RAG (Retrieval-Augmented Generation) is the technique that solves exactly that: instead of retraining the model, you hand it the relevant documents inside the prompt, on every question. It's the architecture behind most serious corporate chatbots — including the support agent we run at HexagonalGuru — and this guide explains how it works, no smoke and mirrors.
The Problem: The Model Doesn't Know (and Doesn't Know It Doesn't Know)
An LLM is a function of its training: if a fact wasn't there, it doesn't exist for the model. And since it has no "I'm making this up" signal, it fills the gaps with plausible answers. That's what we call hallucinations.
The usual alternatives have serious problems:
- Fine-tuning on your data: expensive, outdated the moment a document changes, and the model still can't cite where each claim came from.
- Stuffing everything into the prompt: your knowledge base doesn't fit in a context window (and if it does, you pay for every token on every call).
- Trusting the model's memory: unacceptable when the answer commits your company in front of a customer.
The Core Idea
RAG splits the problem in two phases: first find the document fragments relevant to the question, then generate the answer using those fragments as context. Retrieve, Augment, Generate:
Phase 1: Indexing (once, and every time a document changes)
- Ingestion: your documents (PDFs, website, knowledge base, tickets) are split into reasonably sized chunks — not a lone paragraph without context, not an entire chapter.
- Embeddings: each chunk goes through a model that turns it into a numeric vector: a "semantic fingerprint" where texts with similar meaning land close together.
- Vector store: the vectors are stored in a specialized database (pgvector, Qdrant, Weaviate…) that can search by similarity in milliseconds.
Phase 2: Querying (on every question)
- The user's question is also turned into an embedding.
- The k most similar chunks are retrieved (top-k), ideally with hybrid search and reranking — we'll cover that in the common mistakes.
- Those fragments are inserted into the prompt with a clear instruction: "answer only using this information and cite the source".
- The LLM generates the answer grounded in your documents, with verifiable citations.
Retrieval Is Where You Win or Lose
The LLM is the flashy part, but a RAG system's quality is decided at retrieval: if the right fragment doesn't reach the prompt, the model will hallucinate as elegantly as before. The architecture that works in production has three floors:
- Hybrid search: vector search understands meaning ("how do I return an order?" finds "refund policy"), but fails on exact codes, SKUs and proper nouns; keyword search (BM25) does exactly the opposite. They run in parallel and the results are merged.
- Reranking: a second model re-orders the candidates by true relevance to the question. It's cheap and improves retrieval noticeably.
- Sensible top-k: better 5 good fragments than 20 mediocre ones competing for the model's attention.
A Minimal Example
The core of the query, simplified, fits on one screen:
// 1. Embed the question
$vector = $embeddings->embed($question);
// 2. Retrieve the most similar chunks (hybrid search + rerank)
$chunks = $vectorStore->findSimilar($vector, limit: 5);
// 3. Build the prompt with the retrieved context
$prompt = "Answer ONLY using this information and cite the source.\n\n"
. implode("\n\n", array_map(
fn ($c) => "[Source: {$c->document}]\n{$c->text}",
$chunks,
))
. "\n\nQuestion: {$question}";
// 4. Generate
$answer = $llm->complete($prompt);
Notice what's not there: no training, no custom model. When you update a document, its chunk gets reindexed and the next answer already uses the new version.
What You Gain
- Answers with your data, always up to date. Change the document, change the answer. No retraining.
- Verifiable citations. Every claim can link to its source: the difference between "a chatbot" and a tool your team actually trusts.
- Access control. Retrieval filters by what each user is allowed to see: a sales rep doesn't retrieve boardroom documents.
- Reasonable cost. You pay for embeddings once per document and a few context tokens per query. No recurring fine-tuning.
- Your data trains nobody. Documents travel inside each call's prompt, not into the provider's training pipeline.
When It's Not the Answer
- You want to change behavior, not knowledge. Tone, output format, response style: that's achieved with well-designed prompts or, in extreme cases, fine-tuning. RAG brings data, not personality.
- Your problem is pure search. If the user wants "the March invoice PDF", a good traditional search engine solves it without an LLM. Not every information problem is a generation problem.
- Ultra-specialized domain with its own jargon. If the base model can't understand your industry's vocabulary even with the documents in front of it, it's time to evaluate specific models or fine-tuning on top of RAG.
- 100% structured data and exact queries. "How many orders came in yesterday?" is answered with SQL, not embeddings. (Though you can combine: tool calling for data, RAG for documents.)
Common Mistakes
- Badly cut chunks. Fragments that split ideas in half or mix three topics. Chunking must respect the document's structure (sections, tables) or it destroys retrieval quality.
- Blindly trusting vector search. Without hybrid search and reranking, queries with exact codes and proper nouns fail silently — and the model hallucinates to compensate.
- Not evaluating. Without a set of real questions with their expected answers (a "golden set"), every improvement is a hunch. Retrieval gets measured: did the right fragment make it into the top-5?
- Stuffing everything retrieved into the prompt. Twenty mediocre chunks dilute five good ones and blow up the cost per query.
- Forgetting permissions. If retrieval doesn't filter by the documents the user can see, your chatbot just bypassed the whole company's access control.
The Bottom Line
RAG is today's standard way to connect an LLM to a company's data: index your documents in a vector store, retrieve what's relevant on each question — with hybrid search and reranking — and let the model write grounded in that context. The result: up-to-date, citable answers within the permissions perimeter, without retraining anything. The hard part isn't the LLM: it's the retrieval engineering.
If you're considering an assistant over your documentation, your catalog or your support history, that's exactly what we build: from the ingestion pipeline to production deployment. Let's talk about your case →
Keep reading: How much does custom software cost? An honest guide to budgets.