RAG: a complete overview
How retrieval-augmented generation works: eight steps, RAG compared with fine-tuning and a long prompt, pros and cons, three steps in code, mistakes and rules of a reliable system.
In short
RAG (retrieval-augmented generation) is a way to make a language model answer from your own materials instead of its general memory. Documents are cut into pieces, each piece gets a vector — a numeric fingerprint of its meaning — and is stored in a database. When a question comes, the closest pieces are found by meaning and handed to the model together with the question, and the model answers only from them, citing the sources. It is how AI assistants answer about prices, terms and products without inventing them. If nothing relevant is found, a good RAG system says so and hands over to a person.
RAG at a glance
The main facts in one table — what RAG is, what it is made of and where it is used.
- What it is
- Search through your materials plus an answer from the model based on what was found
- The term
- Introduced in a 2020 paper by Facebook AI Research
- Parts
- Chunking, an embedding model, a vector index, a language model
- Where vectors live
- PostgreSQL with pgvector, SQLite with an extension or a separate vector database
- Models
- Claude, GPT or local models — chosen for the task and the data
- Updates
- A changed document is re-indexed — no retraining
- Main risk
- Poor search: the model answers confidently from the wrong pieces
How RAG works: 8 steps
The first four steps are done once and repeated when documents change; the last four run on every question.
| Step | What happens | What decides quality |
|---|---|---|
| 1. Collect | documents, catalogue, terms, answers to frequent questions | up-to-date and non-contradictory sources |
| 2. Cut | pieces along headings and paragraphs | one thought per piece, with its heading |
| 3. Vectorise | each piece gets a vector of meaning | an embedding model that knows the language |
| 4. Store | pieces and vectors go into the database | source and date next to each piece |
| 5. Find | the closest pieces by meaning and by words | a similarity threshold, filters by section |
| 6. Rerank | a second model reorders the candidates | optional, but improves precision |
| 7. Answer | the model writes only from the numbered pieces | a strict instruction and citations |
| 8. Log | question, pieces found, answer, rating | reviewing the log and filling gaps |
RAG, fine-tuning or a long prompt
Three ways to give a model your knowledge. They solve different problems and are often combined.
| Criterion | RAG | Fine-tuning | Everything in the prompt |
|---|---|---|---|
| Good for | facts that change | style and format of answers | a small fixed set of texts |
| Updating knowledge | re-index a document | train again | edit the prompt |
| Sources in the answer | yes | no | partly |
| Volume of knowledge | practically unlimited | limited by training data | limited by the context window |
| Cost per question | low: only the pieces found | low | high: the whole text each time |
| Access rights | filters at search time | impossible | a separate prompt per role |
| Start | days | weeks and a dataset | hours |
Pros and cons of RAG
RAG makes the answer only as good as the search. Both the strengths and the weaknesses come from that.
Pros · 5
-
Answers from your data
Prices, terms and specifications come from your documents, not from the model’s memory.
-
Sources you can check
Every statement carries a link to the piece it came from.
-
Fresh without retraining
A changed price list is re-indexed in seconds.
-
Rights at search time
A customer and a manager see answers from different sets of documents.
-
Any model
The model can be replaced without rebuilding the knowledge base.
Cons · 4
-
Search decides everything
If the right piece is not found, even the best model answers badly.
-
Garbage in the base
Outdated and contradictory documents turn into confident wrong answers.
-
Questions across the whole base
“How many orders were there in March” is a query to a database, not a search through texts.
-
Needs care
The log has to be reviewed and gaps in the knowledge filled.
What RAG looks like in code: 3 steps
Cutting, indexing and the answer in TypeScript with PostgreSQL and pgvector. The pipeline was run end to end on a test knowledge base of a furniture shop.
Cutting into pieces
Pieces follow headings and paragraphs, so one piece holds one thought and knows where it came from.
// Split a document into pieces of up to ~800 characters along paragraph borders.
// Every piece keeps its source and heading — the answer will cite them.
export type Chunk = { source: string; heading: string; text: string };
export function chunk(source: string, markdown: string, maxChars = 800): Chunk[] {
const chunks: Chunk[] = [];
let heading = '';
let buffer: string[] = [];
const flush = () => {
const text = buffer.join('\n\n').trim();
if (text) chunks.push({ source, heading, text });
buffer = [];
};
for (const block of markdown.split(/\n{2,}/)) {
if (block.startsWith('#')) {
flush(); // a new heading always starts a new piece
heading = block.replace(/^#+\s*/, '');
continue;
}
if (buffer.join('\n\n').length + block.length > maxChars) flush();
buffer.push(block);
}
flush();
return chunks;
}
Indexing
Vectors from any OpenAI-compatible API; re-indexing a document replaces its pieces in one statement, without duplicates.
// Indexing: every piece gets a vector and goes into PostgreSQL with pgvector.
// CREATE TABLE chunks (id bigserial PRIMARY KEY, source text, heading text,
// text text, embedding vector(384));
import pg from 'pg';
import { chunk } from './chunk.ts';
const db = new pg.Pool({ connectionString: process.env.DATABASE_URL });
// Any OpenAI-compatible embeddings API — a cloud provider or a local model
export async function embed(texts: string[]): Promise<number[][]> {
const res = await fetch(`${process.env.EMBEDDINGS_URL}/v1/embeddings`, {
method: 'POST',
headers: {
'Content-Type': 'application/json',
Authorization: `Bearer ${process.env.EMBEDDINGS_KEY}`,
},
body: JSON.stringify({ model: process.env.EMBEDDINGS_MODEL, input: texts }),
});
if (!res.ok) throw new Error(`embeddings: HTTP ${res.status}`);
const { data } = (await res.json()) as { data: { embedding: number[] }[] };
return data.map((d) => d.embedding);
}
export async function indexDocument(source: string, markdown: string): Promise<number> {
const pieces = chunk(source, markdown);
const vectors = await embed(pieces.map((p) => `${p.heading}\n${p.text}`));
// One statement: the old pieces of the document are replaced atomically
await db.query(
`WITH old AS (DELETE FROM chunks WHERE source = $1)
INSERT INTO chunks (source, heading, text, embedding)
SELECT $1, h, t, e::vector FROM unnest($2::text[], $3::text[], $4::text[]) AS u(h, t, e)`,
[source, pieces.map((p) => p.heading), pieces.map((p) => p.text), vectors.map((v) => JSON.stringify(v))],
);
return pieces.length;
}
The answer with sources
In the test, questions about delivery, returns and warranty found the right piece with a similarity of 0.52–0.60; a question about the weather scored 0.14 and never reached the model.
// The answer: find the closest pieces, give them to the model numbered,
// require citations — and do not call the model at all if nothing was found
import pg from 'pg';
import { embed } from './index.ts';
const db = new pg.Pool({ connectionString: process.env.DATABASE_URL });
const API = process.env.ANTHROPIC_BASE_URL ?? 'https://api.anthropic.com';
export async function ask(question: string) {
const [q] = await embed([question]);
const { rows } = await db.query(
`SELECT source, heading, text, 1 - (embedding <=> $1) AS score
FROM chunks ORDER BY embedding <=> $1 LIMIT 4`,
[JSON.stringify(q)],
);
// The threshold is chosen on your own questions; below it — not an answer
const found = rows.filter((r) => r.score > 0.5);
if (found.length === 0) return { answer: null, sources: [] }; // hand over to a person
const context = found.map((r, i) => `[${i + 1}] ${r.heading}\n${r.text}`).join('\n\n');
const res = await fetch(`${API}/v1/messages`, {
method: 'POST',
headers: {
'content-type': 'application/json',
'x-api-key': process.env.ANTHROPIC_API_KEY ?? '',
'anthropic-version': '2023-06-01',
},
body: JSON.stringify({
model: 'claude-sonnet-5',
max_tokens: 600,
system:
'Answer only from the sources below and cite them as [1], [2]. ' +
'If the sources do not contain the answer, say so.\n\n' + context,
messages: [{ role: 'user', content: question }],
}),
});
if (!res.ok) throw new Error(`model: HTTP ${res.status}`);
const data = (await res.json()) as { content: { type: string; text: string }[] };
return { answer: data.content[0].text, sources: found.map((r) => r.source) };
}
Common mistakes in RAG projects
-
Pieces of a fixed length
Cutting every 500 characters splits a rule in half, and neither half answers the question.
-
No threshold
The nearest piece is always found, even for a question about the weather — and the model builds an answer on it.
-
Only vector search
Article numbers, model names and codes are found better by words — the two searches are combined.
-
An English-only embedding model
Questions in other languages find nothing; the model must know the languages of your customers.
-
Indexing once and forgetting
The price list changed, the base did not — the assistant quotes old prices.
-
No set of test questions
Without 30–50 real questions with expected answers, any change is a guess.
7 rules of a reliable RAG
-
01
Clean the sources first
Remove outdated versions and contradictions before indexing.
-
02
Pieces by meaning
Headings and paragraphs as borders, the heading inside each piece.
-
03
Two searches together
By meaning and by words; the results are merged.
-
04
A threshold and an honest “I don’t know”
Below the threshold the model is not called — the question goes to a person.
-
05
Citations are mandatory
An answer without a source is not shown.
-
06
Rights in the query
Filters by role are applied at search time, not in the instruction to the model.
-
07
Measure on real questions
A fixed set of questions is run after every change to the base or the settings.
Questions about RAG
What is RAG in simple words?
First find the answer in your materials, then let the model formulate it — with a link to the source.
Does RAG stop hallucinations?
It reduces them sharply when there is a threshold, a strict instruction and citations; without them it does not.
Do I need a vector database?
Usually not a separate one: pgvector in PostgreSQL holds millions of pieces next to the rest of the data.
Is my data used to train the model?
Not with business API access to Claude or GPT; for the strictest cases there are local models.
What documents can be used?
Texts, tables, PDFs, pages of the site, the catalogue; scans need text recognition first.
How is quality measured?
On a set of real questions: was the right piece found, and is the answer correct and backed by a source.
How does RAG relate to MCP?
RAG searches through texts; MCP gives the model tools — orders, stock, CRM. Assistants often use both.
Online form
An assistant
on your data
I build AI assistants that answer from your documents, catalogue and rules — with links to sources and a handover to a person when the answer is not there. Tell me about the task — I answer within one working day.