RAG
Retrieve current, permission-aware evidence and place it in the model context
Retrieval changes the evidence available for this request. Fine-tuning changes repeatable model behavior. Start with that distinction, then test whether the product needs either, both, or a simpler prompt.
Retrieve current, permission-aware evidence and place it in the model context
Adapt response style, format, task behavior, or domain-specific decision patterns
“Should we use RAG or fine-tuning?” is often asked before the failure is defined. A model that lacks current policy documents has a knowledge-access problem. A model that receives the right evidence but repeatedly uses the wrong format has a behavior problem. Those problems call for different controls.
Prompting, retrieval-augmented generation, fine-tuning, and a hybrid system are complementary. Use the least complex option that meets a measured requirement.
Use clear instructions and a few representative examples when the base model can already perform the task.
Fetch relevant external information for each request when facts are private, current, attributable, or permission-bound.
Train on input-output examples when a stable task requires behavior the prompt cannot produce consistently enough.
Retrieve current evidence and use an adapted model to interpret or present that evidence in a specialized, repeatable way.
Policies, prices, inventory, account state, recent research, product documentation, and tenant-specific records should be retrieved or accessed through tools at request time.
Fine-tuning should not be the default way to load a changing knowledge base. Training data changes model parameters; it does not create a dependable database, source link, deletion mechanism, or request-time permission check. A tuned model can still need RAG.
A RAG system prepares a searchable knowledge source, retrieves candidate passages for the user’s request, and supplies selected evidence to the generator. The model answers from that context under application instructions. This can support citations and source-aware review, but only if retrieval and grounding are designed and evaluated.
RAG introduces its own production system. Data must be collected, parsed, segmented, enriched, embedded or indexed, secured, refreshed, retrieved, ranked, and assembled into context. Each step can fail independently.
| RAG layer | Key design question | Useful evidence |
|---|---|---|
| Source | Which content is authoritative, current, permitted, and safe to expose? | Ownership, version, expiry, classification, tenant and access metadata |
| Preparation | How should structure and meaning be preserved during parsing and chunking? | Parse coverage, chunk boundaries, metadata quality, failed-document reports |
| Retrieval | Can the system find the evidence needed for representative questions? | Recall and ranking on labeled query-document pairs, filter correctness |
| Context | Does the model receive enough relevant evidence without noise or conflicts? | Context precision, token use, source diversity, contradiction handling |
| Generation | Does the answer stay within the evidence and represent uncertainty? | Groundedness, citation correctness, completeness, refusal behavior |
| Operation | Can content be refreshed, deleted, traced, and recovered? | Freshness lag, ingestion failures, deletion tests, request-level traces |
A fluent answer can still be grounded in an irrelevant passage. When the system fails, inspect retrieval before changing the prompt or model. Determine whether the answer existed in the source, whether parsing preserved it, whether the query found it, whether ranking selected it, and whether the generator used it correctly.
Access control must happen in the retrieval and application layers, not through an instruction asking the model to ignore restricted content. Carry the authenticated identity and tenant boundary into filtering, and test that prohibited sources cannot enter the context.
Fine-tuning adapts a base model using examples or preference signals. Depending on the supported method, it can improve consistent formatting, classification, extraction, style, tool selection, or a specialized task pattern. It requires a stable task, representative data, and an evaluation that can distinguish improvement from memorization or regression.
Start with an evaluation and the best practical prompt. Save the inputs where the base model fails and label the desired outputs. Review the dataset for duplicates, contradictions, leakage, sensitive data, narrow coverage, and low-quality synthetic examples. Keep a holdout set the training process never sees.
Do not assume fine-tuning makes outputs factual. It can improve how the model performs a task while still generating unsupported claims. If factual correctness depends on external records, supply and verify those records at inference time.
A support copilot may need RAG for current product documentation and account policy, while a fine-tuned model applies the organization’s response structure and escalation categories. A contract workflow may retrieve clauses and precedent, while an adapted model extracts a consistent risk schema. The retriever and tuned generator still need separate evaluation.
| Requirement | Start with | Why |
|---|---|---|
| Answer from changing internal documents with citations | RAG | Knowledge remains external, refreshable, inspectable, and available for source attribution |
| Produce a consistent domain-specific output format from stable inputs | Prompt and schema, then fine-tuning if measured gaps remain | The requirement concerns repeatable behavior rather than changing knowledge |
| Apply a specialized workflow to current tenant data | RAG or tools plus fine-tuning | Runtime data and learned task behavior are separate requirements |
| Answer a question about one provided document | In-context prompting | A retrieval pipeline may add complexity without adding value |
| Perform a deterministic policy calculation | Code or rules | The result should not depend on generative variation |
| Improve a weak RAG answer | Diagnose the failing layer first | The cause may be source quality, parsing, retrieval, ranking, context, or generation |
A hybrid architecture is not automatically more mature. It adds two data lifecycles, more versions, more failure paths, and more cost. Introduce it only when the evaluation shows distinct knowledge and behavior gaps that the combination resolves.
Use the same representative task set to compare prompting, RAG, fine-tuning, and hybrid candidates. Measure the qualities that matter to the user and the operation rather than one generic score.
Evaluate changes by category. A higher overall average can hide a regression for a high-risk workflow or a tenant-specific permission path. Review failed examples and update the correct layer: data, parser, retriever, prompt, model, tool, or application rule.
Botmer’s custom AI solutions teams can help define the evaluation, choose the smallest effective architecture, and build the production controls around it. For role planning, see how to evaluate an AI-native engineer.
This guide uses current cloud and model-provider documentation as technical reference points. Specific platform capabilities can change and should be checked before implementation.
Botmer can help test the smallest architecture that meets the product’s knowledge and behavior requirements.