TechNova is a fictional company used as a running example throughout this series.

The RAG in Practice series builds a baseline, single-step retrieval over a document set. That baseline is the right default, and many systems never need more.

This page collects six patterns that go further. Use it when your evaluation shows a gap the baseline can’t close, and pick the pattern that matches the gap. Each entry is a signal for where to look, not a tutorial.

Gap you see Pattern to consider
Chunks lose the meaning of the section they came from Parent-child and contextual chunking
Answers are built on weak or off-topic evidence Self-RAG and Corrective RAG
One question needs facts from several places Agentic RAG
The answer is a chain of relationships between things Graph RAG
The knowledge lives in images and diagrams Multimodal RAG
Section structure matters more than wording Vectorless RAG

Parent-Child and Contextual Chunking

Keep track of the section each chunk came from, so retrieval stays precise and generation gets enough context.

The problem. Flat chunking treats every chunk as independent. For documents with strong nested structure, that’s often wrong. A paragraph inside a chapter on chunking strategies means something different from the same paragraph inside a chapter on embeddings.

What changes. Parent-child chunking stores that structure explicitly.

  • The small child chunk is used for retrieval, because it’s precise and searchable.
  • The larger parent section is assembled for generation, so the model sees the surrounding context, not just the isolated paragraph.

A related variant, contextual chunking, gives each child chunk a short summary of the larger section it came from.

When it’s worth it. Documents with deep structure, such as textbooks, manuals, and policies, where a question matches one precise paragraph but the answer needs the section around it. This is a structural choice you make in the design phase, before launch, and it’s one of the decisions that separates RAG demos from production systems.

Example. “Not covered after 30 days” means one thing in a return-policy section and another in a warranty-exceptions section. A section summary on each chunk helps the system tell those similar-looking chunks apart before the model ever sees them.

Self-RAG and Corrective RAG

Add a check on the evidence, so the pipeline notices weak retrieval before it answers.

The problem. Baseline RAG retrieves once and trusts what comes back. When the retrieved chunks are thin, off-topic, or wrong, a one-pass pipeline has no way to notice.

What changes. The two are often mentioned together, but they put the check in different places.

  • Self-RAG builds self-critique into generation. The model is trained to pause at a few points and ask whether it needs to retrieve at all, whether a retrieved passage is relevant, and whether the sentence it just wrote is supported by that passage.
  • Corrective RAG puts the check outside the generator. A separate retrieval evaluator grades the retrieved documents before the answer is written. If the evidence is weak or ambiguous, the system changes the retrieval path before generating, for example by refining the results, rewriting the query, or using another source such as web search. Web search is optional.

Comparison of Self-RAG and Corrective RAG. Self-RAG checks retrieved evidence before generation and checks the answer after generation, with weak evidence able to trigger another retrieval. Corrective RAG checks retrieval quality before generation and, when evidence is weak, refines or rewrites the search and retrieves again before generating.

When it’s worth it. When answers keep resting on weak evidence and better retrieval alone hasn’t fixed it. You don’t need either one when straightforward retrieval already works well.

Example. A support question retrieves three warranty chunks. Self-RAG checks whether those chunks support the answer it’s producing. Corrective RAG checks the chunks first and retrieves better evidence if they’re weak.

Both patterns make a pipeline more adaptive without turning every request into a full agent workflow. Baseline RAG follows a mostly fixed retrieval path. These patterns add evaluation and correction. Agentic RAG, next, goes further by planning several retrieval steps per question.

Agentic RAG

Let the model plan several retrieval steps when one pass can’t gather everything the answer needs.

The problem. Baseline RAG retrieves once. Some questions need facts from more than one place, combined.

What changes. The model plans retrieval steps one after another, deciding what to look up next from what it has already found.

When it’s worth it. Multi-part questions that cross documents, where a single retrieval keeps returning only half of what the answer needs.

Example. A customer asks, “Is my WH-1000 still under warranty if I bought it 18 months ago and updated to firmware v3.2.1?” Answering means retrieving the warranty terms and the firmware requirements, then reasoning across both.

Graph RAG

Store knowledge as entities and relationships when the question is about how things connect.

The problem. Vector similarity finds passages that read like the question. It may not capture a chain of relationships between entities.

What changes. Graph RAG organizes knowledge as entities and the relationships between them, not just document chunks, so the system can follow the links.

When it’s worth it. When relationships between entities matter more than document similarity.

Example. “Which firmware version fixed the ANC issue on the WH-1000?” means following the links from the product to its firmware versions to the fix.

Multimodal RAG

Retrieve images and other non-text content directly when the knowledge isn’t only in text.

The problem. Baseline RAG retrieves text. What a diagram or an annotated photo shows often doesn’t survive as extracted text.

What changes. The pipeline handles images and other non-text content as retrievable objects, not just the text extracted from them.

When it’s worth it. Knowledge that lives in visuals, such as product manuals with diagrams and troubleshooting guides with annotated images.

Example. A troubleshooting question asks which cable goes in which port, and the answer is in a labeled photo rather than in the surrounding text.

Vectorless RAG

Navigate the document’s structure instead of searching by similarity when section hierarchy matters more than wording.

The problem. A question may require following section references across a changelog, a policy document, and a troubleshooting guide. Flat vector chunking can lose those relationships unless the system preserves them explicitly, which is part of why parent-child and contextual chunking exist.

What changes. Vectorless RAG keeps the document’s structure intact and lets the model navigate sections, more like a person following a table of contents. The open-source PageIndex framework is one example.

  • There are no embeddings, no vector database, and no similarity-based chunking.
  • The document is still split into structural units, such as sections and pages, but it isn’t sliced into chunks by embedding similarity.
  • The model walks the structure to find the relevant part.

When it’s worth it. Structured documents such as contracts, filings, manuals, and long policy documents, where section hierarchy matters more than phrase similarity. It isn’t a universal replacement for vector RAG. For most other workloads, vector retrieval is still the right default.

Example. A customer asks whether the opened-headset exception in the returns policy still applies, and the answer depends on a later section that overrides it. The system follows the policy’s section references instead of matching similar wording.

Reading the benchmark. An eye-catching result is attached to this approach, and it’s worth knowing what it actually measures.

  • On FinanceBench, a question-answering benchmark over financial filings, Mafin 2.5, a commercial system built on PageIndex, reports 98.7% accuracy. The vendor reports that figure itself, so read it as a promising signal rather than an independent head-to-head.
  • The benchmark authors measured a traditional vector RAG baseline at roughly 50%, but only in a setup where the system was already told which filing held the answer. That removes most of the actual retrieval problem.
  • When the same baseline had to search across all the filings, closer to real cross-document retrieval, it scored around 19%.

So “98.7% versus 50%” compares two different setups. The fair reading is that structure-aware retrieval looks strong on this kind of document, while ordinary vector search struggles once it has to find the right filing on its own.


Before adding any of these, make sure the baseline is really the limit. Stale data, missing documents, and a wrong metadata rule all look like retrieval gaps, and none of them needs a new pattern. Run your evaluation set before and after, and keep a pattern only if the gap actually closes.

Read deeper: RAG Part 8: RAG in Production · Pattern 06: Diagnose Retrieval Before Blaming the Model

References