Lensa ML
Lensa ML

Retrieval-Augmented Generation

Step 1 of 6

Why Retrieval?

LLMs have knowledge cutoffs — retrieval brings fresh facts

Knowledge Cutoff Year2022
2019Who won the 2019 World…2020What was the 2020 pand…2021Which vaccine rolled o…2022What chatbot launched …2023What model released Ma…2024Which AI won 2024 Nobe…2025What model family laun…2026What tech trended in 2…Without retrieval, LLMs only know data up to their training cutoffRAG bridges the knowledge gap with external documents
Cutoff
2022
Answered
4/8
Coverage
50%

A 2022 cutoff covers 4 of 8 events. Events after 2022 produce hallucinations because the model has no training data for them.

Step 1 of 6: Why Retrieval?

LLMs have knowledge cutoffs — retrieval brings fresh facts

Knowledge Cutoff Year2022
2019Who won the 2019 World…2020What was the 2020 pand…2021Which vaccine rolled o…2022What chatbot launched …2023What model released Ma…2024Which AI won 2024 Nobe…2025What model family laun…2026What tech trended in 2…Without retrieval, LLMs only know data up to their training cutoffRAG bridges the knowledge gap with external documents
Cutoff
2022
Answered
4/8
Coverage
50%

A 2022 cutoff covers 4 of 8 events. Events after 2022 produce hallucinations because the model has no training data for them.

Step 2 of 6: Document Embedding

Split documents into chunks and embed each as a vector

Chunk Size (tokens)200
Full Document (1000 tokens)C1200tC2200tC3200tC4200tC5200t
Chunk Size
200t
Chunks
5
Precision
medium

200-token chunks produce 5 balanced chunks. Each is embedded into a dense vector — enough context per chunk for meaningful retrieval.

Step 3 of 6: Query & Retrieval

Find the nearest chunk embeddings to the query vector

Top-k Retrieved3
QC1C2C3C4C5C6C7C8Nearest 3 chunks retrieved from vector store
Top-k
3
Retrieved
3 chunks
Skipped
5

Top-3 retrieval balances recall and noise. The 3 nearest chunks in embedding space are returned, giving the model multiple evidence sources.

Step 4 of 6: Similarity Search

Filter chunks by cosine similarity threshold

Similarity Threshold0.50
Transformers O0.53RAG Pipeline0.70Vector Databas0.75Prompt Enginee0.71Fine-tuning Gu0.81Embedding Mode0.93Chunking Strat0.86Reranking Meth0.77
Threshold
0.50
Passing
8
Filtered Out
0
Avg Similarity
0.76

Threshold 0.50 filters to 8 chunks. Cosine similarity measures the angle between query and chunk vectors — higher means more semantically aligned.

Step 5 of 6: Augmented Prompt

Inject retrieved context into the prompt template

Context Window Usage0.50
SYSTEM: Answer using only the provided context.[CONTEXT] — 3 chunks (2008 tokens)Chunk 1: "Transformers Overview…"Chunk 2: "RAG Pipeline…"Chunk 3: "Vector Databases…"[QUERY] "How does RAG work?" (50 tokens)51% used
Chunks Used
3
Context Tokens
2008
Remaining
2008

3 chunks using 2008 tokens. Good balance — enough retrieved evidence to ground the answer, with 2008 tokens left for a thorough response.

Step 6 of 6: Faithfulness & Hallucination

Better retrieval grounds more tokens in real documents

Retrieval Quality0.50
Model Output — tokens colored by groundingRAGretrievesrelevantdocumentsfromavectorstorethenthemodelgeneratesananswergroundedinretrievedcontextreducinghallucinationsGrounded in docsHallucinated
Quality
0.50
Grounded
15/20
Hallucinated
5
Grounding Score
75%

Moderate retrieval quality — 15 grounded tokens, 5 hallucinated. The grounding score of 75% shows room for improvement via better chunking or reranking.