Why this matters
- StreamHub's creator support bot answers questions about monetisation, content policies, and technical setup — RAG retrieves the right help article instead of hallucinating.
- Vector search is the retrieval layer: given a user question, find the top-K most semantically similar document chunks in milliseconds.
- Production RAG systems need ingestion pipelines, chunking strategies, re-ranking, and evaluation — not just a prompt with context.
RAG platform components
- Ingestion pipeline — parse PDFs, HTML, markdown; chunk into passages; embed and index.
- Embedding model — converts text to dense vectors (OpenAI, Cohere, open-source BGE).
- Vector database — stores embeddings with metadata; supports similarity search (Pinecone, Weaviate, pgvector).
- Retriever — embeds query, searches vector DB, returns top-K chunks with scores.
- Generator — LLM receives query + retrieved context; produces grounded answer with citations.
- Evaluation — measures retrieval precision, answer faithfulness, and latency.
Pipeline architecture
RAG platform (AWS)
Document ingestion
from langchain.text_splitter import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=512,
chunk_overlap=64,
separators=["\n\n", "\n", ". ", " "]
)
def ingest_document(doc_id: str, content: str, metadata: dict):
chunks = splitter.split_text(content)
embeddings = embedding_model.encode(chunks)
for i, (chunk, vector) in enumerate(zip(chunks, embeddings)):
vector_db.upsert(
id=f"{doc_id}_chunk_{i}",
vector=vector,
metadata={
**metadata,
"doc_id": doc_id,
"chunk_index": i,
"text": chunk
}
)
Query flow
def answer_question(query: str, top_k: int = 5) -> dict:
query_vector = embedding_model.encode(query)
results = vector_db.search(
vector=query_vector,
top_k=top_k,
filter={"category": "creator-help"}
)
context = "\n\n".join([
f"[{r.metadata['doc_id']}]: {r.metadata['text']}"
for r in results
])
prompt = f"""Answer the question using only the provided context.
If the context does not contain the answer, say "I don't have information about that."
Context:
{context}
Question: {query}
Answer with citations to the source documents."""
answer = llm.generate(prompt)
return {
"answer": answer,
"sources": [r.metadata["doc_id"] for r in results],
"retrieval_scores": [r.score for r in results]
}
Vector database schema
-- pgvector example
CREATE EXTENSION vector;
CREATE TABLE document_chunks (
id TEXT PRIMARY KEY,
doc_id TEXT NOT NULL,
chunk_index INT NOT NULL,
content TEXT NOT NULL,
embedding vector(1536), -- OpenAI ada-002 dimensions
metadata JSONB,
created_at TIMESTAMPTZ DEFAULT now()
);
CREATE INDEX idx_chunks_embedding ON document_chunks
USING ivfflat (embedding vector_cosine_ops) WITH (lists = 100);
| Aspect | Vector DB | Best for |
|---|---|---|
| pgvector | Postgres extension, SQL joins | < 10M vectors, existing PG infra |
| Pinecone | Managed, serverless | Fast setup, auto-scaling |
| Weaviate | Hybrid search (vector + keyword) | Complex filtering requirements |
| Qdrant | High performance, filtering | Self-hosted, cost-sensitive |
| Elasticsearch | Dense + sparse retrieval | Existing ES cluster, hybrid search |
pgvector
Vector DBPostgres extension, SQL joinsBest for< 10M vectors, existing PG infraPinecone
Vector DBManaged, serverlessBest forFast setup, auto-scalingWeaviate
Vector DBHybrid search (vector + keyword)Best forComplex filtering requirementsQdrant
Vector DBHigh performance, filteringBest forSelf-hosted, cost-sensitiveElasticsearch
Vector DBDense + sparse retrievalBest forExisting ES cluster, hybrid search
For StreamHub's 500K help articles (~2M chunks), pgvector on Aurora is sufficient. Scale to dedicated vector DB above 50M vectors.
Chunking and retrieval quality
Retrieval optimisation
- Chunk size — 256–512 tokens balances context richness vs retrieval precision.
- Overlap — 10–20% overlap prevents losing context at chunk boundaries.
- Hybrid search — combine vector similarity with BM25 keyword matching for better recall.
- Re-ranking — cross-encoder model re-scores top-20 candidates to top-5 for precision.
- Metadata filtering — pre-filter by category, date, or tenant before vector search.
Evaluation metrics
| Metric | What it measures | Target |
|---|---|---|
| Retrieval precision@5 | Relevant chunks in top 5 | > 0.8 |
| Answer faithfulness | Answer grounded in retrieved context | > 0.9 |
| Latency p95 | End-to-end query time | < 3s |
| Hallucination rate | Claims not supported by context | < 5% |
| User satisfaction | Thumbs up/down on answers | > 85% positive |
Retrieval precision@5
What it measuresRelevant chunks in top 5Target> 0.8Answer faithfulness
What it measuresAnswer grounded in retrieved contextTarget> 0.9Latency p95
What it measuresEnd-to-end query timeTarget< 3sHallucination rate
What it measuresClaims not supported by contextTarget< 5%User satisfaction
What it measuresThumbs up/down on answersTarget> 85% positive
Evaluate on a golden dataset of 200+ question-answer pairs before every model or chunking change.
Quick recall
Everything you need if you only revisit this box.
- RAG = retrieve relevant document chunks → inject into LLM prompt → generate grounded answer.
- Ingestion: chunk documents (512 tokens, 64 overlap) → embed → store in vector DB.
- Vector search finds semantically similar chunks; hybrid search adds keyword matching.
- Re-rank top-20 to top-5 with a cross-encoder for better precision.
- Evaluate retrieval precision, answer faithfulness, and hallucination rate on golden datasets.
Test yourself
Answer these before moving on — recall is what makes it stick.