Large language models know a lot, but they do not know your data. They cannot read your internal documentation, query your private database, or reference the PDF you uploaded last Tuesday. Retrieval-Augmented Generation (RAG) solves this by fetching relevant documents at query time and injecting them into the prompt so the model can generate answers grounded in your actual content.
RAG has become the standard architecture for building AI-powered search, customer support bots, documentation assistants, and knowledge management systems. Unlike fine-tuning, which bakes knowledge into model weights and requires retraining when data changes, RAG keeps your knowledge layer separate and updatable. You change a document, re-embed it, and the system immediately reflects the update.
This guide walks through every component of a production RAG pipeline with working Python code. By the end, you will have a clear mental model of how the pieces fit together and the practical knowledge to build one yourself.
Working with API responses? The JSON Formatter helps you inspect and debug the structured payloads that flow between your embedding API, vector database, and LLM. Paste any JSON, get formatted output instantly.
What Is RAG and Why It Matters
Retrieval-Augmented Generation is a two-stage architecture: first retrieve relevant information from an external knowledge base, then generate a response using that information as context. The term was introduced in a 2020 paper by Lewis et al. at Meta AI, but the pattern has since evolved far beyond the original formulation.
The core insight is simple: instead of expecting the LLM to memorize everything during training, you give it the right information at inference time. This approach solves several fundamental problems with standalone LLMs:
- Knowledge cutoff. LLMs are frozen at their training date. RAG lets them answer questions about data that did not exist when they were trained.
- Hallucination. When an LLM does not know an answer, it often fabricates one confidently. RAG grounds responses in actual source documents, reducing hallucination rates significantly.
- Private data. You cannot fine-tune a hosted model on proprietary data without uploading it to a third party. RAG keeps your data in your own infrastructure.
- Verifiability. RAG systems can cite specific source passages, letting users verify claims against the original documents.
- Freshness. Updating a RAG knowledge base takes minutes. Retraining or fine-tuning a model takes hours to days and costs significantly more.
In practice, RAG is not a single technique but a family of patterns. The simplest version is "naive RAG": embed a query, find similar documents, stuff them into a prompt. Production systems layer on re-ranking, hybrid search, query transformation, and multi-step retrieval. We will cover the spectrum from basic to advanced.
RAG Architecture: Retrieval, Augmentation, Generation
Every RAG system has three stages, regardless of the specific tools or models used:
Stage 1: Indexing (Offline)
Before any queries happen, you prepare your knowledge base. This involves loading documents, splitting them into chunks, generating embedding vectors for each chunk, and storing those vectors in a database.
- Document loading — Ingest PDFs, web pages, Markdown files, database records, or any text source.
- Chunking — Split documents into passages of 256-1024 tokens. Chunk boundaries matter: splitting mid-sentence degrades retrieval quality.
- Embedding — Convert each chunk into a high-dimensional vector (typically 768-3072 dimensions) using an embedding model.
- Storage — Write vectors + metadata to a vector database for fast similarity search.
Stage 2: Retrieval (Online)
When a user asks a question, you embed their query with the same model, search the vector database for the most similar chunks, and return the top-k results.
Stage 3: Generation (Online)
You construct a prompt containing the user's question and the retrieved chunks, then send it to an LLM. The model generates an answer grounded in the provided context.
Retrieval quality determines generation quality. If the right documents are not in the context, even the best LLM cannot produce a correct answer. Spend 80% of your optimization effort on retrieval.
Setting Up a Vector Database
A vector database stores embedding vectors and performs approximate nearest-neighbor (ANN) search to find the most similar vectors to a query. Here are three widely-used options, each suited to different scenarios.
ChromaDB: Best for Prototyping
Chroma runs in-process with zero infrastructure. It stores data locally and supports persistent storage. Ideal for development and datasets under 100,000 documents.
Python - ChromaDB Setupimport chromadb
# Persistent storage (survives restarts)
client = chromadb.PersistentClient(path="./chroma_data")
# Create a collection (like a table)
collection = client.get_or_create_collection(
name="documents",
metadata={"hnsw:space": "cosine"} # cosine similarity
)
# Add documents with embeddings
collection.add(
ids=["doc1", "doc2", "doc3"],
documents=[
"RAG systems retrieve relevant context before generation.",
"Vector databases store high-dimensional embeddings.",
"Chunking strategy affects retrieval precision."
],
metadatas=[
{"source": "intro.md", "page": 1},
{"source": "architecture.md", "page": 3},
{"source": "best-practices.md", "page": 7}
]
)
# Query: Chroma embeds the query automatically
results = collection.query(
query_texts=["How does retrieval work in RAG?"],
n_results=3
)
print(results["documents"][0]) # Top 3 matching chunks
Pinecone: Best for Managed Production
Pinecone is a fully managed vector database. You do not run servers, manage indexes, or worry about scaling. It handles billions of vectors with single-digit millisecond latency.
Python - Pinecone Setupfrom pinecone import Pinecone
pc = Pinecone(api_key="your-api-key")
# Create an index (do this once)
pc.create_index(
name="rag-docs",
dimension=1536, # Must match your embedding model
metric="cosine",
spec={"serverless": {"cloud": "aws", "region": "us-east-1"}}
)
index = pc.Index("rag-docs")
# Upsert vectors with metadata
index.upsert(vectors=[
{
"id": "chunk-001",
"values": embedding_vector, # List of 1536 floats
"metadata": {
"text": "RAG retrieves documents before generation.",
"source": "intro.md",
"page": 1
}
}
])
# Query
results = index.query(
vector=query_embedding,
top_k=5,
include_metadata=True
)
for match in results["matches"]:
print(f"Score: {match['score']:.3f} - {match['metadata']['text']}")
Weaviate: Best for Hybrid Search
Weaviate is open-source and combines vector search with traditional keyword (BM25) search in a single query. This hybrid approach catches results that pure semantic search misses.
Python - Weaviate Hybrid Searchimport weaviate
client = weaviate.connect_to_local() # or connect_to_weaviate_cloud()
# Define a collection with vectorizer
documents = client.collections.create(
name="Document",
vectorizer_config=weaviate.classes.config.Configure.Vectorizer.text2vec_openai(),
properties=[
weaviate.classes.config.Property(name="content", data_type=weaviate.classes.config.DataType.TEXT),
weaviate.classes.config.Property(name="source", data_type=weaviate.classes.config.DataType.TEXT),
]
)
# Hybrid search: combines vector + BM25
response = documents.query.hybrid(
query="How does chunking affect RAG quality?",
alpha=0.7, # 0 = pure BM25, 1 = pure vector
limit=5
)
for obj in response.objects:
print(obj.properties["content"])
For managing the configuration files that these databases require, the YAML Editor gives you syntax validation and formatting for Docker Compose files, Kubernetes manifests, and application configs.
Creating Embeddings
Embeddings are the numerical representations that make semantic search possible. An embedding model converts text into a dense vector where similar meanings map to nearby points in vector space. The quality of your embeddings directly determines the quality of your retrieval.
OpenAI Embeddings
OpenAI's text-embedding-3-small (1536 dimensions) and text-embedding-3-large (3072 dimensions) are the most widely used commercial options. They support dimension reduction via the dimensions parameter, letting you trade accuracy for storage and speed.
from openai import OpenAI
client = OpenAI() # Uses OPENAI_API_KEY env var
def get_embeddings(texts: list[str], model="text-embedding-3-small"):
"""Embed a batch of texts. Returns list of vectors."""
response = client.embeddings.create(
input=texts,
model=model
)
return [item.embedding for item in response.data]
# Embed documents
chunks = [
"Vector databases enable fast similarity search.",
"RAG reduces hallucination by grounding in source data.",
"Chunk overlap prevents information loss at boundaries."
]
vectors = get_embeddings(chunks)
print(f"Dimensions: {len(vectors[0])}") # 1536
# Embed a query (same model, same dimensions)
query_vector = get_embeddings(["What prevents hallucination?"])[0]
Cohere Embeddings
Cohere's embed-v4.0 model distinguishes between document and query embeddings, which can improve retrieval relevance. It also natively supports multiple languages.
import cohere
co = cohere.ClientV2(api_key="your-api-key")
# Document embeddings
doc_response = co.embed(
texts=["RAG architecture overview...", "Chunking strategies..."],
model="embed-v4.0",
input_type="search_document",
embedding_types=["float"]
)
# Query embedding (different input_type)
query_response = co.embed(
texts=["How should I chunk my documents?"],
model="embed-v4.0",
input_type="search_query",
embedding_types=["float"]
)
Open-Source: Sentence Transformers
For self-hosted embeddings with zero API costs, Sentence Transformers provides excellent models. The all-MiniLM-L6-v2 model is fast and lightweight (384 dimensions). For higher quality, use BAAI/bge-large-en-v1.5 (1024 dimensions).
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("BAAI/bge-large-en-v1.5")
# Embed documents
chunks = [
"Vector search finds semantically similar passages.",
"BM25 matches exact keywords in documents.",
"Hybrid search combines both approaches."
]
doc_embeddings = model.encode(chunks, normalize_embeddings=True)
# Embed query
query_embedding = model.encode(
["What combines keyword and semantic search?"],
normalize_embeddings=True
)
# Compute similarities directly
from sentence_transformers.util import cos_sim
scores = cos_sim(query_embedding, doc_embeddings)
print(scores) # tensor([[0.31, 0.28, 0.89]])
Always use the same embedding model for documents and queries. If you embed documents with text-embedding-3-small and queries with a different model, the vector spaces will not align and retrieval will fail silently. This is the most common RAG setup mistake.
Building the Retrieval Pipeline
With embeddings stored in a vector database, you can build the complete retrieval pipeline. The following example ties together document ingestion, query embedding, and context retrieval into a reusable class.
Python - Complete Retrieval Pipelineimport chromadb
from openai import OpenAI
class RAGPipeline:
def __init__(self, collection_name="knowledge_base"):
self.openai = OpenAI()
self.chroma = chromadb.PersistentClient(path="./rag_data")
self.collection = self.chroma.get_or_create_collection(
name=collection_name,
metadata={"hnsw:space": "cosine"}
)
def embed(self, texts: list[str]) -> list[list[float]]:
"""Generate embeddings for a list of texts."""
response = self.openai.embeddings.create(
input=texts,
model="text-embedding-3-small"
)
return [item.embedding for item in response.data]
def ingest(self, documents: list[dict]):
"""
Ingest documents into the vector store.
Each doc: {"id": str, "text": str, "metadata": dict}
"""
texts = [doc["text"] for doc in documents]
embeddings = self.embed(texts)
self.collection.add(
ids=[doc["id"] for doc in documents],
embeddings=embeddings,
documents=texts,
metadatas=[doc.get("metadata", {}) for doc in documents]
)
print(f"Ingested {len(documents)} documents.")
def retrieve(self, query: str, top_k: int = 5) -> list[dict]:
"""Retrieve the most relevant chunks for a query."""
query_embedding = self.embed([query])[0]
results = self.collection.query(
query_embeddings=[query_embedding],
n_results=top_k,
include=["documents", "metadatas", "distances"]
)
retrieved = []
for i in range(len(results["ids"][0])):
retrieved.append({
"id": results["ids"][0][i],
"text": results["documents"][0][i],
"metadata": results["metadatas"][0][i],
"distance": results["distances"][0][i]
})
return retrieved
The retrieve method returns chunks sorted by cosine similarity. The distance field tells you how relevant each chunk is—lower values mean higher similarity. You can set a threshold (e.g., reject anything above 0.4) to filter out irrelevant results.
When debugging retrieval results, you will often need to inspect the structured data coming back from your vector store. The API Tester lets you send requests to your RAG endpoint and inspect responses in real time, directly in the browser.
Connecting to LLMs: Claude, GPT-4, Open-Source
The generation step takes your retrieved context and produces a final answer. The key is a well-structured prompt that gives the model clear instructions about how to use the provided context.
Claude (Anthropic)
Python - RAG with Claudeimport anthropic
def generate_with_claude(query: str, context_chunks: list[dict]) -> str:
"""Generate an answer using Claude with retrieved context."""
client = anthropic.Anthropic()
# Format context with source attribution
context_text = ""
for i, chunk in enumerate(context_chunks, 1):
source = chunk["metadata"].get("source", "unknown")
context_text += f"[Source {i}: {source}]\n{chunk['text']}\n\n"
message = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
system="""You are a helpful assistant that answers questions based
on the provided context. Rules:
- Only use information from the context below to answer.
- If the context does not contain enough information, say so explicitly.
- Cite sources using [Source N] notation.
- Be concise and accurate.""",
messages=[{
"role": "user",
"content": f"""Context:
{context_text}
Question: {query}
Answer based on the context above:"""
}]
)
return message.content[0].text
GPT-4 (OpenAI)
Python - RAG with GPT-4from openai import OpenAI
def generate_with_gpt4(query: str, context_chunks: list[dict]) -> str:
"""Generate an answer using GPT-4 with retrieved context."""
client = OpenAI()
context_text = "\n\n".join([
f"[{chunk['metadata'].get('source', 'doc')}]: {chunk['text']}"
for chunk in context_chunks
])
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{
"role": "system",
"content": "Answer questions using only the provided context. "
"Cite sources. If unsure, say you don't know."
},
{
"role": "user",
"content": f"Context:\n{context_text}\n\nQuestion: {query}"
}
],
temperature=0.1 # Low temperature for factual accuracy
)
return response.choices[0].message.content
Open-Source with Ollama
For fully local RAG without sending data to external APIs, use Ollama to run open-source models like Llama 3, Mistral, or Qwen locally.
Python - RAG with Ollama (Local)import requests
def generate_with_ollama(query: str, context_chunks: list[dict]) -> str:
"""Generate using a local Ollama model."""
context_text = "\n\n".join([chunk["text"] for chunk in context_chunks])
response = requests.post("http://localhost:11434/api/generate", json={
"model": "llama3.1:8b",
"prompt": f"""Based on the following context, answer the question.
Only use information from the context. If you cannot answer, say so.
Context:
{context_text}
Question: {query}
Answer:""",
"stream": False,
"options": {"temperature": 0.1}
})
return response.json()["response"]
Putting It All Together
Python - End-to-End RAG Query# Initialize pipeline
rag = RAGPipeline()
# Ingest your documents (do this once)
rag.ingest([
{"id": "c1", "text": "RAG reduces hallucination by providing source documents...",
"metadata": {"source": "rag-intro.md"}},
{"id": "c2", "text": "Cosine similarity measures the angle between two vectors...",
"metadata": {"source": "math-primer.md"}},
{"id": "c3", "text": "Production RAG systems should implement relevance thresholds...",
"metadata": {"source": "best-practices.md"}},
])
# Query
query = "How does RAG prevent hallucination?"
chunks = rag.retrieve(query, top_k=3)
answer = generate_with_claude(query, chunks)
print(answer)
Advanced RAG: Re-Ranking, Hybrid Search, Chunking
Naive RAG—embed, search, generate—works for simple use cases. Production systems need more sophisticated retrieval strategies to handle ambiguous queries, diverse document types, and accuracy requirements.
Re-Ranking
The initial vector search returns approximate results quickly but sometimes ranks less relevant passages higher. A re-ranker is a cross-encoder model that takes the query and each candidate passage as a pair and produces a more accurate relevance score. It is slower than vector search but significantly more precise.
Python - Re-Ranking with Cohereimport cohere
co = cohere.ClientV2(api_key="your-api-key")
def rerank(query: str, documents: list[str], top_n: int = 3):
"""Re-rank documents using a cross-encoder model."""
response = co.rerank(
query=query,
documents=documents,
model="rerank-v3.5",
top_n=top_n
)
return [
{"index": r.index, "score": r.relevance_score, "text": documents[r.index]}
for r in response.results
]
# Typical pattern: retrieve 20, re-rank to top 5
initial_results = rag.retrieve(query, top_k=20)
texts = [r["text"] for r in initial_results]
reranked = rerank(query, texts, top_n=5)
The "retrieve many, re-rank to few" pattern is one of the highest-impact optimizations you can add to a RAG system. Retrieving 20 candidates and re-ranking to the top 5 typically improves answer quality by 15-25% compared to directly retrieving 5.
Hybrid Search
Semantic search excels at understanding meaning but can miss exact keyword matches. BM25 keyword search catches exact terms but misses semantic equivalents. Hybrid search combines both, typically using Reciprocal Rank Fusion (RRF) to merge the two result lists.
Python - Reciprocal Rank Fusiondef reciprocal_rank_fusion(
result_lists: list[list[str]],
k: int = 60
) -> list[tuple[str, float]]:
"""
Merge multiple ranked result lists using RRF.
k=60 is the standard constant from the original paper.
"""
scores = {}
for results in result_lists:
for rank, doc_id in enumerate(results, 1):
scores[doc_id] = scores.get(doc_id, 0) + 1 / (k + rank)
return sorted(scores.items(), key=lambda x: x[1], reverse=True)
# Example: merge vector search and BM25 results
vector_results = ["doc3", "doc1", "doc7", "doc2", "doc5"]
bm25_results = ["doc1", "doc3", "doc4", "doc5", "doc8"]
fused = reciprocal_rank_fusion([vector_results, bm25_results])
# doc3 and doc1 appear in both lists, so they rank highest
Chunking Strategies
How you split documents into chunks has an outsized effect on retrieval quality. There is no universal best approach—the right strategy depends on your content.
- Fixed-size chunks (e.g., 512 tokens with 50-token overlap) — Simple and predictable. Works well for homogeneous text like articles or documentation.
- Semantic chunking — Split at natural boundaries: paragraphs, section headers, topic transitions. Preserves meaning units but produces variable-sized chunks.
- Recursive character splitting — Try splitting by paragraphs first, then by sentences, then by words. LangChain's
RecursiveCharacterTextSplitterimplements this. - Parent-child chunking — Index small chunks for precise retrieval but return the larger parent chunk for context. This gives you both precision in search and completeness in generation.
import re
def semantic_chunk(text: str, max_tokens: int = 512) -> list[str]:
"""Split text at paragraph boundaries, respecting max size."""
paragraphs = re.split(r'\n\n+', text)
chunks = []
current_chunk = ""
for para in paragraphs:
# Rough token estimate: 1 token ~ 4 chars
if len(current_chunk + para) / 4 > max_tokens and current_chunk:
chunks.append(current_chunk.strip())
current_chunk = para
else:
current_chunk += "\n\n" + para if current_chunk else para
if current_chunk.strip():
chunks.append(current_chunk.strip())
return chunks
Query Transformation
User queries are often vague, incomplete, or poorly phrased for retrieval. Query transformation rewrites the query before searching to improve matches.
- HyDE (Hypothetical Document Embeddings) — Ask the LLM to generate a hypothetical answer, then use that answer as the search query. This bridges the gap between question-style queries and document-style chunks.
- Multi-query — Generate 3-5 variations of the original query and retrieve results for each. Merge with RRF. This handles ambiguity by covering multiple interpretations.
- Step-back prompting — Rewrite the query at a higher level of abstraction to retrieve broader context, then use the original specific query with that context.
Production Deployment Tips
Moving a RAG prototype to production introduces challenges around performance, reliability, and observability that do not surface during development.
Manage API Keys Securely
RAG systems typically need API keys for the embedding model, vector database, and LLM. Store these in environment variables, never in code. The Env File Editor helps you manage .env files with syntax highlighting and validation, keeping your secrets organized across environments.
OPENAI_API_KEY=sk-proj-...
ANTHROPIC_API_KEY=sk-ant-...
PINECONE_API_KEY=pcsk_...
COHERE_API_KEY=co-...
Implement Caching
Embedding the same text repeatedly wastes money and adds latency. Cache embeddings and LLM responses where appropriate.
Python - Simple Embedding Cacheimport hashlib
import json
from pathlib import Path
class EmbeddingCache:
def __init__(self, cache_dir="./embedding_cache"):
self.cache_dir = Path(cache_dir)
self.cache_dir.mkdir(exist_ok=True)
def _key(self, text: str, model: str) -> str:
return hashlib.sha256(f"{model}:{text}".encode()).hexdigest()
def get(self, text: str, model: str):
path = self.cache_dir / f"{self._key(text, model)}.json"
if path.exists():
return json.loads(path.read_text())
return None
def set(self, text: str, model: str, embedding: list[float]):
path = self.cache_dir / f"{self._key(text, model)}.json"
path.write_text(json.dumps(embedding))
Monitor Retrieval Quality
Log every query, the retrieved chunks, their relevance scores, and the generated response. Without this data, you cannot debug why the system gave a bad answer. Key metrics to track:
- Retrieval recall@k — What fraction of relevant documents appear in the top-k results?
- Mean relevance score — Are your top results actually relevant, or is the system scraping the bottom of the barrel?
- Answer faithfulness — Does the generated answer stick to the retrieved context, or does the model hallucinate beyond it?
- Latency breakdown — How much time is spent on embedding, retrieval, re-ranking, and generation separately?
Use Docker for Reproducible Infrastructure
If you are self-hosting Weaviate, Qdrant, or ChromaDB in server mode, containerize everything. The Docker Compose Generator can scaffold a multi-service configuration for your vector database, API server, and cache layer.
docker-compose.ymlservices:
weaviate:
image: cr.weaviate.io/semitechnologies/weaviate:1.28.4
ports:
- "8080:8080"
- "50051:50051"
environment:
QUERY_DEFAULTS_LIMIT: 25
AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED: "true"
PERSISTENCE_DATA_PATH: "/var/lib/weaviate"
volumes:
- weaviate_data:/var/lib/weaviate
rag-api:
build: .
ports:
- "8000:8000"
env_file: .env
depends_on:
- weaviate
volumes:
weaviate_data:
Common Pitfalls and How to Avoid Them
1. Embedding Model Mismatch
Using different embedding models (or different versions) for documents and queries is the number one silent failure in RAG systems. The vectors live in incompatible spaces, so similarity scores are meaningless. Always version your embedding model and re-embed all documents if you switch models.
2. Ignoring Chunk Overlap
When chunks split a critical piece of information across two passages, neither chunk contains the complete answer. Use 10-20% overlap between chunks. For a 512-token chunk, overlap 50-100 tokens with the previous and next chunks.
3. Not Setting Relevance Thresholds
Without a minimum relevance threshold, the system returns the "least dissimilar" results even when nothing in the knowledge base is actually relevant. This leads to confidently wrong answers. Set a cosine distance threshold (e.g., 0.4) and return a fallback message when no results meet it.
Python - Relevance Thresholddef retrieve_with_threshold(query, top_k=5, max_distance=0.4):
results = rag.retrieve(query, top_k=top_k)
relevant = [r for r in results if r["distance"] < max_distance]
if not relevant:
return None # Trigger fallback response
return relevant
4. Stuffing Too Much Context
More context is not always better. Including marginally relevant chunks dilutes the signal. Models can get confused when contradictory information appears in the context. Start with 3-5 chunks and increase only if you measure improved answer quality.
5. Not Handling Document Updates
When source documents change, the old embeddings become stale. Build an update pipeline that detects changes (via content hashing or modification timestamps), removes old vectors, and ingests updated ones. Without this, your system gradually drifts from reality.
6. Skipping Evaluation
Build an evaluation dataset before optimizing. Create 50-100 question-answer pairs from your actual documents, then measure retrieval recall and answer quality against this dataset as you change chunking strategies, embedding models, or retrieval parameters. Without evaluation, optimization is guesswork.
When comparing outputs across different configurations, the Text Diff Viewer lets you place two RAG responses side-by-side and spot exactly what changed between pipeline versions.
When your RAG system gives a wrong answer, always check retrieval first. In roughly 80% of cases, the problem is that the right document was not retrieved, not that the LLM misunderstood the context. Print the retrieved chunks before blaming the model.
Related Developer Tools
Free browser-based tools to support your RAG development workflow. No signup, no tracking, runs entirely in your browser.
Frequently Asked Questions
RAG retrieves relevant documents at query time and includes them in the prompt, so the model generates answers grounded in your data without changing its weights. Fine-tuning modifies the model's weights by training on your dataset, which bakes knowledge into the model itself. RAG is better when your data changes frequently, you need source attribution, or you want to avoid the cost and complexity of training. Fine-tuning is better for teaching the model a specific style, format, or domain-specific reasoning pattern. Many production systems combine both: fine-tune for tone and reasoning, then use RAG for factual grounding.
Chunk size depends on your content type and retrieval needs. For most use cases, 256 to 512 tokens works well as a starting point. Smaller chunks (128-256 tokens) give more precise retrieval but may lose context. Larger chunks (512-1024 tokens) preserve more context but may include irrelevant information that dilutes the embedding. Use overlap of 10-20% between chunks to avoid splitting important information at boundaries. Test different sizes against your actual queries: split your documents, run representative questions, and measure whether the correct chunk appears in the top-k results. There is no universal best size because it depends on your document structure and query patterns.
For prototyping and small datasets under 100,000 documents, ChromaDB is the simplest choice because it runs in-process with no infrastructure. For production workloads, Pinecone offers a fully managed service with low operational overhead and scales to billions of vectors. Weaviate is a strong open-source option if you want to self-host and need hybrid search combining vector and keyword retrieval. Qdrant is another performant open-source option with a Rust backend. pgvector works well if you already use PostgreSQL and want to avoid adding another database to your stack. Choose based on scale, team expertise, and whether you prefer managed services or self-hosted infrastructure.
Hallucinations in RAG happen when the model generates information not present in the retrieved context. To reduce them: set a relevance threshold on retrieval scores and return a fallback response when no sufficiently relevant documents are found. Include explicit instructions in your system prompt telling the model to only answer based on the provided context and to say it does not know when the context is insufficient. Use citation-forcing prompts that require the model to quote or reference specific passages. Implement a post-generation verification step that checks whether claims in the output are supported by the retrieved chunks. Finally, retrieve more documents (higher k) and use a re-ranker to ensure the most relevant content is in the context window.
Yes. A minimal RAG system requires three components: an embedding function to convert text to vectors, a vector store to index and search those vectors, and an LLM API call that includes retrieved context in the prompt. You can build this with direct API calls to OpenAI or Anthropic for embeddings and generation, plus a vector database client library like chromadb or pinecone-client. The core logic is roughly 50-100 lines of Python. LangChain and LlamaIndex add abstractions for chaining, document loading, and text splitting that can speed up development, but they also add complexity and dependencies. For simple use cases, direct API calls give you more control and easier debugging. For complex pipelines with multiple retrieval strategies, agents, or tool use, a framework can save significant development time.