Skip to content
Botmer International®
Insights/AI unit economics

The operating cost of an AI feature: what to measure before you scale

A token bill describes model consumption. A useful cost model follows the entire workflow, including retrieval, tools, retries, review, and whether the user received an acceptable result.

Cost follows the complete task
01 / INFERENCE

Model

Input, output, reasoning, cached input, embeddings

02 / SYSTEM

Runtime

Retrieval, storage, tools, orchestration, network

03 / EXCEPTIONS

Recovery

Retries, fallbacks, moderation, review, support

04 / VALUE

Outcome

Completed and accepted business task

An AI feature can look inexpensive in a provider dashboard and still have weak unit economics. The model charge may exclude the retrieval pipeline, cloud runtime, third-party tools, repeated attempts, human review, and support generated when the workflow fails. It also says nothing about whether the result created value.

Choose a unit tied to the user outcome

Start with the business task the feature exists to complete. For a support assistant, the unit might be a resolved case with no reopen during an agreed period. For document extraction, it might be an accepted record that passes validation. For a sales research workflow, it might be a reviewed account brief. For an internal coding assistant, it could be an accepted change rather than a generated suggestion.

The denominator mattersCost per request rewards short activity. Cost per accepted outcome rewards useful completion.

Track technical units such as tokens and model calls because engineers can act on them. Pair them with a business unit so product leaders can judge whether the spend creates value.

A single user action can produce several model requests. An agent might plan, retrieve context, select a tool, call an API, inspect the result, retry, and compose an answer. Dividing monthly spend by visible user messages hides that amplification. Instead, assign every step a workflow identifier and roll its cost up to the completed task.

Define success before cost

A result should meet an observable acceptance rule. Depending on the product, that can combine schema validity, grounded evidence, business-rule checks, a user acceptance action, downstream completion, or a sampled human review. If quality is missing from the denominator, cost optimization can make the metric look better while the product becomes less useful.

  • Resource unit: input tokens, output tokens, model calls, tool calls, storage, compute time, or retrieval operations.
  • Workflow unit: conversation, document, case, report, code change, or agent run.
  • Outcome unit: accepted result, resolved case, completed action, qualified lead, or approved record.
  • Quality companion: task success, override rate, escalation rate, defect rate, groundedness, or another task-specific measure.

Map the full operating bill

Model inference is one layer. The rest of the architecture can dominate certain workloads, especially when large document collections, external tools, heavy observability, or human approval are required.

Cost layerWhat belongs in itWhat can drive it unexpectedly
InferenceInput, output, reasoning, cached input, image or audio processing, embeddings, rerankingLong histories, oversized retrieval, verbose outputs, repeated planning, model upgrades, cache misses
Application runtimeAPI compute, queues, functions, containers, databases, vector indexes, object storage, network transferIdle capacity, duplicate work, high-cardinality logs, frequent re-indexing, inefficient polling
Tools and dataSearch, maps, enrichment, browser automation, OCR, transcription, proprietary data, external APIsPer-call fees, rate-limit retries, unused tool definitions, duplicated lookups, premium data tiers
Quality and safetyEvaluations, moderation, guard models, policy checks, human review, red-team exercisesApplying expensive checks to every task instead of matching controls to risk
Failure and recoveryRetries, fallbacks, partial-work cleanup, incident handling, customer support, credits or remediationUnbounded agent loops, non-idempotent tools, weak validation, unavailable dependencies
People and operationPrompt and evaluation maintenance, data curation, model migrations, monitoring, vendor review, on-call workFrequent model changes, undocumented behavior, manual exceptions, unclear ownership

Include the costs relevant to the decision. A product manager comparing two model routes may need variable cost per accepted result. A finance decision about pricing or self-hosting needs a fuller view of engineering, infrastructure, utilization, support, and commitments.

Do not assume self-hosting removes variable cost

Hosted APIs expose consumption clearly, while self-hosted models convert part of that bill into infrastructure and operating responsibility. Hardware or reserved capacity can incur cost while idle. Serving, scaling, security updates, model packaging, observability, and specialist time remain part of the service. Compare realistic utilization and total operating ownership rather than a headline price per token.

Instrument the workflow before changing it

Provider totals are useful for reconciliation but insufficient for product decisions. Emit usage and outcome data at the point where the feature makes a request. The minimum useful record links technical consumption to a tenant, feature, workflow, prompt or configuration version, and outcome.

  1. Assign a workflow identifier
    Carry one identifier across model calls, retrieval, tools, retries, and the final user-visible result.
  2. Record configuration
    Capture provider, model, prompt version, reasoning setting, retrieval version, tool set, and relevant feature flags.
  3. Capture consumption
    Record input, output, reasoning, cached and cache-write tokens where exposed, plus embeddings, tool calls, storage, and runtime usage.
  4. Classify attempts
    Separate first attempts, validation retries, provider retries, user retries, fallbacks, and abandoned runs.
  5. Attach outcome evidence
    Record whether the task completed, passed deterministic checks, was accepted, escalated, corrected, or rejected.
  6. Reconcile and alert
    Compare application records with provider billing and set limits for loops, request volume, per-workflow cost, and unusual changes.

Do not log private content merely to improve cost reporting. Identifiers, usage fields, categories, hashes, and aggregate measures are often enough. Apply the same retention, access, and tenant-boundary rules used for other production telemetry.

Use distributions, not only averages

An average can hide a small set of workflows that repeatedly retrieve too much context or enter tool loops. Track median and upper-percentile cost per outcome, attempts per workflow, and the share of spend by feature, tenant, model, and failure category. Investigate shape changes after model, prompt, retrieval, or product releases.

Optimize in the order that protects quality

The biggest saving is often removing unnecessary work, not compressing every prompt. Establish a representative evaluation set and baseline first. Then change one cost driver at a time and compare quality, latency, and cost on the same tasks.

01 / ELIMINATE

Remove avoidable calls

Use deterministic code for rules, calculations, lookups, validation, and transformations that do not need a model.

02 / BOUND

Stop runaway work

Set step, time, token, retry, and tool budgets. Make repeated actions idempotent and provide a clear fallback.

03 / RETRIEVE

Send relevant evidence

Improve chunking, filters, ranking, and context assembly so the model sees what the task needs rather than the whole corpus.

04 / ROUTE

Match model to task

Use evaluations to determine which tasks can use a smaller or lower-reasoning model and when escalation is justified.

05 / REUSE

Cache stable work

Place reusable instructions and tool definitions in a stable prefix, then measure actual cache hits and retention behavior.

06 / SCHEDULE

Batch non-urgent work

Move eligible offline evaluations, enrichment, summarization, or extraction to provider batch processing where the delay is acceptable.

Constrain output length and request structured fields when the product needs a compact answer. Trim conversation history using a tested memory strategy. Limit the tools exposed to each task. None of these choices should be accepted on token reduction alone; rerun the same evaluations and examine failure categories.

Prompt caching is a measured optimization

Current provider implementations reuse matching prompt prefixes under model-specific eligibility, lifetime, and pricing rules. Keep stable instructions, examples, and tool schemas before dynamic user and retrieval content. Track cached usage and misses rather than assuming a long prompt will be reused. A changed prefix, model, tool definition, or routing pattern can alter cache performance.

Forecast with scenarios before usage scales

A forecast should show how product behavior creates cost. Build low, expected, and high scenarios from active users, tasks per user, attempts per task, calls per attempt, resource consumption per call, tool usage, success rate, and fixed operating cost.

The model should answer practical questions: What happens if conversation history doubles? What if the success rate drops and users retry? What if a larger model is required for one step? Which tenants or plans create negative margin? Where should the product add limits, asynchronous processing, a smaller model, or human escalation?

Botmer perspectiveSet the cost budget beside the quality and latency target during architecture, not after the first large bill

Botmer’s custom AI solutions work can include the evaluation, telemetry, and operating controls needed to move from a promising workflow to sustainable production use.

Reference points

Provider prices and product capabilities change. This guide avoids fixed price claims and uses current documentation for the operating mechanisms described.

Build sustainable AI operations

Bring the workflow, volume assumptions, and quality bar you need to hold

Botmer can help design the architecture, evaluation, and telemetry behind a production AI feature.

Discuss your product