PrepZone Logo
PrepZone

RAG and Vector Search Platform

Embed documents, retrieve relevant chunks and generate answers with retrieval-augmented generation.

Why this matters

  • StreamHub's creator support bot answers questions about monetisation, content policies, and technical setup — RAG retrieves the right help article instead of hallucinating.
  • Vector search is the retrieval layer: given a user question, find the top-K most semantically similar document chunks in milliseconds.
  • Production RAG systems need ingestion pipelines, chunking strategies, re-ranking, and evaluation — not just a prompt with context.

RAG platform components

  • Ingestion pipeline — parse PDFs, HTML, markdown; chunk into passages; embed and index.
  • Embedding model — converts text to dense vectors (OpenAI, Cohere, open-source BGE).
  • Vector database — stores embeddings with metadata; supports similarity search (Pinecone, Weaviate, pgvector).
  • Retriever — embeds query, searches vector DB, returns top-K chunks with scores.
  • Generator — LLM receives query + retrieved context; produces grounded answer with citations.
  • Evaluation — measures retrieval precision, answer faithfulness, and latency.

Pipeline architecture

RAG platform (AWS)

embedretrieveindexanswerCLIENT
User query
STORAGE
Amazon S3document corpus
COMPUTE
Amazon Bedrockembed + generate
ANALYTICS
OpenSearch k-NNvector index
COMPUTE
RAG APIEKS
S3 docs → Bedrock embed → OpenSearch k-NN → EKS query API.

Document ingestion

Java
from langchain.text_splitter import RecursiveCharacterTextSplitter

splitter = RecursiveCharacterTextSplitter(
    chunk_size=512,
    chunk_overlap=64,
    separators=["\n\n", "\n", ". ", " "]
)

def ingest_document(doc_id: str, content: str, metadata: dict):
    chunks = splitter.split_text(content)
    embeddings = embedding_model.encode(chunks)

    for i, (chunk, vector) in enumerate(zip(chunks, embeddings)):
        vector_db.upsert(
            id=f"{doc_id}_chunk_{i}",
            vector=vector,
            metadata={
                **metadata,
                "doc_id": doc_id,
                "chunk_index": i,
                "text": chunk
            }
        )

Query flow

Java
def answer_question(query: str, top_k: int = 5) -> dict:
    query_vector = embedding_model.encode(query)
    results = vector_db.search(
        vector=query_vector,
        top_k=top_k,
        filter={"category": "creator-help"}
    )

    context = "\n\n".join([
        f"[{r.metadata['doc_id']}]: {r.metadata['text']}"
        for r in results
    ])

    prompt = f"""Answer the question using only the provided context.
If the context does not contain the answer, say "I don't have information about that."

Context:
{context}

Question: {query}

Answer with citations to the source documents."""

    answer = llm.generate(prompt)
    return {
        "answer": answer,
        "sources": [r.metadata["doc_id"] for r in results],
        "retrieval_scores": [r.score for r in results]
    }

Vector database schema

Java
-- pgvector example
CREATE EXTENSION vector;

CREATE TABLE document_chunks (
  id          TEXT PRIMARY KEY,
  doc_id      TEXT NOT NULL,
  chunk_index INT NOT NULL,
  content     TEXT NOT NULL,
  embedding   vector(1536),  -- OpenAI ada-002 dimensions
  metadata    JSONB,
  created_at  TIMESTAMPTZ DEFAULT now()
);

CREATE INDEX idx_chunks_embedding ON document_chunks
  USING ivfflat (embedding vector_cosine_ops) WITH (lists = 100);
AspectVector DBBest for
pgvectorPostgres extension, SQL joins< 10M vectors, existing PG infra
PineconeManaged, serverlessFast setup, auto-scaling
WeaviateHybrid search (vector + keyword)Complex filtering requirements
QdrantHigh performance, filteringSelf-hosted, cost-sensitive
ElasticsearchDense + sparse retrievalExisting ES cluster, hybrid search
  • pgvector

    Vector DBPostgres extension, SQL joins
    Best for< 10M vectors, existing PG infra
  • Pinecone

    Vector DBManaged, serverless
    Best forFast setup, auto-scaling
  • Weaviate

    Vector DBHybrid search (vector + keyword)
    Best forComplex filtering requirements
  • Qdrant

    Vector DBHigh performance, filtering
    Best forSelf-hosted, cost-sensitive
  • Elasticsearch

    Vector DBDense + sparse retrieval
    Best forExisting ES cluster, hybrid search

For StreamHub's 500K help articles (~2M chunks), pgvector on Aurora is sufficient. Scale to dedicated vector DB above 50M vectors.

Chunking and retrieval quality

Retrieval optimisation

  • Chunk size — 256–512 tokens balances context richness vs retrieval precision.
  • Overlap — 10–20% overlap prevents losing context at chunk boundaries.
  • Hybrid search — combine vector similarity with BM25 keyword matching for better recall.
  • Re-ranking — cross-encoder model re-scores top-20 candidates to top-5 for precision.
  • Metadata filtering — pre-filter by category, date, or tenant before vector search.

Evaluation metrics

MetricWhat it measuresTarget
Retrieval precision@5Relevant chunks in top 5> 0.8
Answer faithfulnessAnswer grounded in retrieved context> 0.9
Latency p95End-to-end query time< 3s
Hallucination rateClaims not supported by context< 5%
User satisfactionThumbs up/down on answers> 85% positive
  • Retrieval precision@5

    What it measuresRelevant chunks in top 5
    Target> 0.8
  • Answer faithfulness

    What it measuresAnswer grounded in retrieved context
    Target> 0.9
  • Latency p95

    What it measuresEnd-to-end query time
    Target< 3s
  • Hallucination rate

    What it measuresClaims not supported by context
    Target< 5%
  • User satisfaction

    What it measuresThumbs up/down on answers
    Target> 85% positive

Evaluate on a golden dataset of 200+ question-answer pairs before every model or chunking change.

Quick recall

Everything you need if you only revisit this box.

  • RAG = retrieve relevant document chunks → inject into LLM prompt → generate grounded answer.
  • Ingestion: chunk documents (512 tokens, 64 overlap) → embed → store in vector DB.
  • Vector search finds semantically similar chunks; hybrid search adds keyword matching.
  • Re-rank top-20 to top-5 with a cross-encoder for better precision.
  • Evaluate retrieval precision, answer faithfulness, and hallucination rate on golden datasets.

Test yourself

Answer these before moving on — recall is what makes it stick.