Ingest
Bring approved sources into the knowledge system with ownership, version and update context attached.
D2 Automation Knowledge · RAG reliability
A vector database can retrieve similar text. It cannot, by itself, prove that the right evidence was selected, that the answer stayed grounded or that the underlying knowledge is still current. Reliable RAG is an evidence lifecycle, not one model call.
Direct answer
Reliable RAG is a chain of evidence: approved sources are segmented and described with useful metadata, retrieval finds plausible candidates, ranking or filtering improves evidence selection where needed, generation stays grounded in that evidence, evaluation separates retrieval from answer quality, weak evidence triggers fallback, and the knowledge index is refreshed as sources change.
Reliability pipeline
Each stage has a different failure mode. Combining them into one “RAG quality” score makes diagnosis harder and can hide whether the problem is source data, retrieval, ranking, generation or freshness.
Bring approved sources into the knowledge system with ownership, version and update context attached.
Split content into retrieval units that preserve useful semantic boundaries instead of relying on arbitrary chunk windows by default.
Attach metadata needed for source filtering, access boundaries, recency checks and later evaluation.
Produce plausible evidence candidates for the query through vector search or another retrieval method.
Reorder or constrain candidates when evaluation shows the initial retrieval set is too broad, noisy or poorly ordered.
Generate only from the selected evidence and retain enough source context to verify why the answer was produced.
Separate retrieval quality, faithfulness and answer correctness so the failing stage can be identified instead of hidden in one score.
Use feedback, source changes and repeatable evaluation to update the knowledge index and operating rules over time.
Control model
Know which documents are authoritative, who owns them, which version is current and whether the user is allowed to receive that evidence.
owner · version · freshness · access
Ask whether the system found the evidence needed to answer the question before judging the language model that consumed it.
relevant evidence · coverage · misses
When useful evidence is present but poorly ordered, filtering or reranking can improve which context reaches generation.
candidate order · metadata fit · relevance
The answer should remain inside the selected evidence boundary and make unsupported or conflicting context visible instead of smoothing it over.
evidence support · citation trace · contradiction
A reliable system needs explicit answer states for sufficient evidence, ambiguity, missing evidence and review rather than one universal generate path.
answer · clarify · fallback · review
Sources change after launch. Re-indexing, deletion, stale-content detection and regression evaluation are part of RAG operations, not optional maintenance.
refresh · versioning · regression
Evaluation model
A curated evaluation set should record the expected evidence and expected answer behavior. That lets the team identify which stage actually failed instead of tuning prompts for a retrieval problem.
01
Did the system retrieve the evidence needed to answer?
Inspect: Expected source/chunk present, useful evidence coverage, noisy candidates, repeated misses by query type.
Why it matters: If evidence never enters the context window, generation cannot reliably recover it.
02
Is the generated answer supported by the retrieved evidence?
Inspect: Unsupported claims, contradictions, evidence omissions and whether citations or trace links actually support the statement.
Why it matters: A fluent answer can still be ungrounded even when retrieval returned good context.
03
Is the final answer correct for the user question and expected task?
Inspect: Reference answer or reviewer judgment, task completion, required details and acceptable uncertainty behavior.
Why it matters: A faithful answer can still be incomplete or wrong if the source itself is insufficient, stale or misinterpreted.
Answer states
Relevant, sufficiently strong evidence supports the requested answer and the response stays within that evidence.
The query is ambiguous, underspecified or spans multiple plausible meanings; ask for the missing discriminator before retrieving again.
The approved knowledge base does not contain enough evidence. Say so instead of filling the gap with model prior knowledge unless the product explicitly allows it.
Evidence is conflicting, high-risk or repeatedly produces uncertain results; preserve the query and retrieval trace for human review or system improvement.
Failure diagnosis
The prose sounds correct but the selected evidence is outdated, non-authoritative or outside the intended access boundary.
The source exists but segmentation prevents the relevant context from being retrieved as a useful unit.
Initial retrieval contains the answer but weaker candidates dominate the context window; evaluate filtering or reranking.
The model adds claims that are not present in the retrieved evidence; strengthen grounding and faithfulness evaluation.
A heuristic score is presented as if it were validated probability. Tie confidence labels to explicit operating rules and evaluation evidence.
The workflow runs correctly but the knowledge base no longer reflects the current source; freshness needs its own control loop.
Production checklist
Related evidence & guidance
See the workflow-backed architecture case for retrieval, reranking, grounding, confidence controls and the knowledge lifecycle.
ExploreUse AI inside bounded workflows with deterministic validation, fallback and human-review controls around uncertain output.
ExploreDefine authoritative state and ownership before automation starts copying or interpreting information across systems.
ExploreFAQ
Reliable RAG treats answer generation as the final stage of an evidence pipeline. Source ownership, document segmentation, metadata, retrieval, optional reranking or filtering, grounding, evaluation, fallback behavior, feedback and knowledge refresh all need explicit controls.
No. A vector database helps retrieve semantically similar candidates, but similarity is not the same as authoritative evidence. Reliability also depends on source quality, metadata, retrieval coverage, ranking, grounding, evaluation and stale-content handling.
No. Reranking adds another relevance decision and usually adds latency or cost. Use it when evaluation shows that initial retrieval often contains useful evidence but ranks the wrong candidates too highly.
Use a curated evaluation set that records the question, expected evidence and expected answer behavior. Measure retrieval quality separately from answer faithfulness and answer correctness so failures can be attributed to the right stage.
Weak evidence should produce an explicit operating state such as clarify, fallback or human review. The model should not convert missing or conflicting evidence into confident prose simply because it can generate an answer.
Track source ownership, version or update time, re-indexing rules and deletion behavior. A retrieval pipeline can remain technically healthy while serving outdated evidence if the knowledge-refresh lifecycle is not monitored.
Building a knowledge workflow?
Authorship & accountability
D2 AI & Automation TeamProduction automation, APIs, data pipelines and AI-assisted systems
D2 keeps claims, assumptions and evidence separate. Citations are attached only when a relevant source or evidence asset is available; unresolved material is not automatically presented as a verified fact.
Review D2's evidence methodology →