Skip to content
Botmer International®
Insights/AI architecture

RAG or fine-tuning: choose by what needs to change

Retrieval changes the evidence available for this request. Fine-tuning changes repeatable model behavior. Start with that distinction, then test whether the product needs either, both, or a simpler prompt.

Two different control surfaces
CONTEXT AT REQUEST TIME

RAG

Retrieve current, permission-aware evidence and place it in the model context

PATTERNS LEARNED FROM EXAMPLES

Fine-tuning

Adapt response style, format, task behavior, or domain-specific decision patterns

“Should we use RAG or fine-tuning?” is often asked before the failure is defined. A model that lacks current policy documents has a knowledge-access problem. A model that receives the right evidence but repeatedly uses the wrong format has a behavior problem. Those problems call for different controls.

Begin with four options, not two

Prompting, retrieval-augmented generation, fine-tuning, and a hybrid system are complementary. Use the least complex option that meets a measured requirement.

01 / PROMPT

Instructions and examples

Use clear instructions and a few representative examples when the base model can already perform the task.

02 / RETRIEVE

RAG

Fetch relevant external information for each request when facts are private, current, attributable, or permission-bound.

03 / ADAPT

Fine-tuning

Train on input-output examples when a stable task requires behavior the prompt cannot produce consistently enough.

04 / COMBINE

RAG plus fine-tuning

Retrieve current evidence and use an adapted model to interpret or present that evidence in a specialized, repeatable way.

The first diagnosticIf the correct answer changes when the source changes, the source belongs outside the model

Policies, prices, inventory, account state, recent research, product documentation, and tenant-specific records should be retrieved or accessed through tools at request time.

Fine-tuning should not be the default way to load a changing knowledge base. Training data changes model parameters; it does not create a dependable database, source link, deletion mechanism, or request-time permission check. A tuned model can still need RAG.

Use RAG when the answer depends on current evidence

A RAG system prepares a searchable knowledge source, retrieves candidate passages for the user’s request, and supplies selected evidence to the generator. The model answers from that context under application instructions. This can support citations and source-aware review, but only if retrieval and grounding are designed and evaluated.

RAG is a strong fit when

  • Knowledge changes more frequently than a model training cycle.
  • The content is private, tenant-specific, or filtered by user permissions.
  • Users need references or the ability to inspect the supporting source.
  • Records must be added, corrected, expired, or removed without retraining a model.
  • The task needs a narrow subset of a corpus rather than a summary of everything.

RAG introduces its own production system. Data must be collected, parsed, segmented, enriched, embedded or indexed, secured, refreshed, retrieved, ranked, and assembled into context. Each step can fail independently.

RAG layerKey design questionUseful evidence
SourceWhich content is authoritative, current, permitted, and safe to expose?Ownership, version, expiry, classification, tenant and access metadata
PreparationHow should structure and meaning be preserved during parsing and chunking?Parse coverage, chunk boundaries, metadata quality, failed-document reports
RetrievalCan the system find the evidence needed for representative questions?Recall and ranking on labeled query-document pairs, filter correctness
ContextDoes the model receive enough relevant evidence without noise or conflicts?Context precision, token use, source diversity, contradiction handling
GenerationDoes the answer stay within the evidence and represent uncertainty?Groundedness, citation correctness, completeness, refusal behavior
OperationCan content be refreshed, deleted, traced, and recovered?Freshness lag, ingestion failures, deletion tests, request-level traces

RAG cannot repair missing or wrong evidence

A fluent answer can still be grounded in an irrelevant passage. When the system fails, inspect retrieval before changing the prompt or model. Determine whether the answer existed in the source, whether parsing preserved it, whether the query found it, whether ranking selected it, and whether the generator used it correctly.

Access control must happen in the retrieval and application layers, not through an instruction asking the model to ignore restricted content. Carry the authenticated identity and tenant boundary into filtering, and test that prohibited sources cannot enter the context.

Use fine-tuning when examples define the desired behavior

Fine-tuning adapts a base model using examples or preference signals. Depending on the supported method, it can improve consistent formatting, classification, extraction, style, tool selection, or a specialized task pattern. It requires a stable task, representative data, and an evaluation that can distinguish improvement from memorization or regression.

Fine-tuning is a stronger candidate when

  • The desired behavior is difficult to express reliably through instructions alone but easy to demonstrate with high-quality examples.
  • The task repeats at meaningful volume and prompt examples consume substantial context.
  • A smaller adapted model can be tested against the required quality, latency, and cost target.
  • Inputs and expected outputs form a stable contract, such as a classification, transformation, or constrained domain task.
  • The team can maintain separate training, validation, and holdout data with clear provenance and permissions.

Start with an evaluation and the best practical prompt. Save the inputs where the base model fails and label the desired outputs. Review the dataset for duplicates, contradictions, leakage, sensitive data, narrow coverage, and low-quality synthetic examples. Keep a holdout set the training process never sees.

  1. Define the behavior gap
    Name the failure that prompting, structured output, retrieval, tools, or deterministic code does not solve adequately.
  2. Build a representative baseline
    Run the current model and prompt on real task categories, edge cases, and unacceptable outcomes.
  3. Curate training examples
    Use consistent, authorized, high-quality input-output pairs that demonstrate the target behavior and relevant diversity.
  4. Separate evaluation data
    Keep a holdout set for comparison and include categories likely to expose regressions or unsafe behavior.
  5. Compare complete systems
    Evaluate the tuned model against the same prompt, context, tools, latency, cost, and operational constraints as the baseline.
  6. Version and monitor
    Record the base model, dataset, training configuration, prompt, evaluation result, release, and rollback path.

Do not assume fine-tuning makes outputs factual. It can improve how the model performs a task while still generating unsupported claims. If factual correctness depends on external records, supply and verify those records at inference time.

Combine them when knowledge and behavior both need control

A support copilot may need RAG for current product documentation and account policy, while a fine-tuned model applies the organization’s response structure and escalation categories. A contract workflow may retrieve clauses and precedent, while an adapted model extracts a consistent risk schema. The retriever and tuned generator still need separate evaluation.

RequirementStart withWhy
Answer from changing internal documents with citationsRAGKnowledge remains external, refreshable, inspectable, and available for source attribution
Produce a consistent domain-specific output format from stable inputsPrompt and schema, then fine-tuning if measured gaps remainThe requirement concerns repeatable behavior rather than changing knowledge
Apply a specialized workflow to current tenant dataRAG or tools plus fine-tuningRuntime data and learned task behavior are separate requirements
Answer a question about one provided documentIn-context promptingA retrieval pipeline may add complexity without adding value
Perform a deterministic policy calculationCode or rulesThe result should not depend on generative variation
Improve a weak RAG answerDiagnose the failing layer firstThe cause may be source quality, parsing, retrieval, ranking, context, or generation

A hybrid architecture is not automatically more mature. It adds two data lifecycles, more versions, more failure paths, and more cost. Introduce it only when the evaluation shows distinct knowledge and behavior gaps that the combination resolves.

Evaluate the complete product path

Use the same representative task set to compare prompting, RAG, fine-tuning, and hybrid candidates. Measure the qualities that matter to the user and the operation rather than one generic score.

  • Task result: correctness, completeness, format, action success, or user acceptance.
  • Evidence: retrieval recall, context precision, groundedness, citation support, and permission correctness.
  • Behavior: instruction following, consistency, calibration, refusal, and required escalation.
  • Operation: latency, cost per accepted result, freshness, failure recovery, observability, and maintenance effort.
  • Safety: sensitive-data handling, access boundaries, harmful outputs, prompt injection, and tool authority.

Evaluate changes by category. A higher overall average can hide a regression for a high-risk workflow or a tenant-specific permission path. Review failed examples and update the correct layer: data, parser, retriever, prompt, model, tool, or application rule.

Botmer perspectiveArchitecture should follow the measured failure, not the most fashionable technique

Botmer’s custom AI solutions teams can help define the evaluation, choose the smallest effective architecture, and build the production controls around it. For role planning, see how to evaluate an AI-native engineer.

Reference points

This guide uses current cloud and model-provider documentation as technical reference points. Specific platform capabilities can change and should be checked before implementation.

Choose from evidence

Bring the task, knowledge source, failure examples, and quality bar to measure

Botmer can help test the smallest architecture that meets the product’s knowledge and behavior requirements.

Discuss your product