Model
Input, output, reasoning, cached input, embeddings
A token bill describes model consumption. A useful cost model follows the entire workflow, including retrieval, tools, retries, review, and whether the user received an acceptable result.
Input, output, reasoning, cached input, embeddings
Retrieval, storage, tools, orchestration, network
Retries, fallbacks, moderation, review, support
Completed and accepted business task
An AI feature can look inexpensive in a provider dashboard and still have weak unit economics. The model charge may exclude the retrieval pipeline, cloud runtime, third-party tools, repeated attempts, human review, and support generated when the workflow fails. It also says nothing about whether the result created value.
Start with the business task the feature exists to complete. For a support assistant, the unit might be a resolved case with no reopen during an agreed period. For document extraction, it might be an accepted record that passes validation. For a sales research workflow, it might be a reviewed account brief. For an internal coding assistant, it could be an accepted change rather than a generated suggestion.
Track technical units such as tokens and model calls because engineers can act on them. Pair them with a business unit so product leaders can judge whether the spend creates value.
A single user action can produce several model requests. An agent might plan, retrieve context, select a tool, call an API, inspect the result, retry, and compose an answer. Dividing monthly spend by visible user messages hides that amplification. Instead, assign every step a workflow identifier and roll its cost up to the completed task.
A result should meet an observable acceptance rule. Depending on the product, that can combine schema validity, grounded evidence, business-rule checks, a user acceptance action, downstream completion, or a sampled human review. If quality is missing from the denominator, cost optimization can make the metric look better while the product becomes less useful.
Model inference is one layer. The rest of the architecture can dominate certain workloads, especially when large document collections, external tools, heavy observability, or human approval are required.
| Cost layer | What belongs in it | What can drive it unexpectedly |
|---|---|---|
| Inference | Input, output, reasoning, cached input, image or audio processing, embeddings, reranking | Long histories, oversized retrieval, verbose outputs, repeated planning, model upgrades, cache misses |
| Application runtime | API compute, queues, functions, containers, databases, vector indexes, object storage, network transfer | Idle capacity, duplicate work, high-cardinality logs, frequent re-indexing, inefficient polling |
| Tools and data | Search, maps, enrichment, browser automation, OCR, transcription, proprietary data, external APIs | Per-call fees, rate-limit retries, unused tool definitions, duplicated lookups, premium data tiers |
| Quality and safety | Evaluations, moderation, guard models, policy checks, human review, red-team exercises | Applying expensive checks to every task instead of matching controls to risk |
| Failure and recovery | Retries, fallbacks, partial-work cleanup, incident handling, customer support, credits or remediation | Unbounded agent loops, non-idempotent tools, weak validation, unavailable dependencies |
| People and operation | Prompt and evaluation maintenance, data curation, model migrations, monitoring, vendor review, on-call work | Frequent model changes, undocumented behavior, manual exceptions, unclear ownership |
Include the costs relevant to the decision. A product manager comparing two model routes may need variable cost per accepted result. A finance decision about pricing or self-hosting needs a fuller view of engineering, infrastructure, utilization, support, and commitments.
Hosted APIs expose consumption clearly, while self-hosted models convert part of that bill into infrastructure and operating responsibility. Hardware or reserved capacity can incur cost while idle. Serving, scaling, security updates, model packaging, observability, and specialist time remain part of the service. Compare realistic utilization and total operating ownership rather than a headline price per token.
Provider totals are useful for reconciliation but insufficient for product decisions. Emit usage and outcome data at the point where the feature makes a request. The minimum useful record links technical consumption to a tenant, feature, workflow, prompt or configuration version, and outcome.
Do not log private content merely to improve cost reporting. Identifiers, usage fields, categories, hashes, and aggregate measures are often enough. Apply the same retention, access, and tenant-boundary rules used for other production telemetry.
An average can hide a small set of workflows that repeatedly retrieve too much context or enter tool loops. Track median and upper-percentile cost per outcome, attempts per workflow, and the share of spend by feature, tenant, model, and failure category. Investigate shape changes after model, prompt, retrieval, or product releases.
The biggest saving is often removing unnecessary work, not compressing every prompt. Establish a representative evaluation set and baseline first. Then change one cost driver at a time and compare quality, latency, and cost on the same tasks.
Use deterministic code for rules, calculations, lookups, validation, and transformations that do not need a model.
Set step, time, token, retry, and tool budgets. Make repeated actions idempotent and provide a clear fallback.
Improve chunking, filters, ranking, and context assembly so the model sees what the task needs rather than the whole corpus.
Use evaluations to determine which tasks can use a smaller or lower-reasoning model and when escalation is justified.
Place reusable instructions and tool definitions in a stable prefix, then measure actual cache hits and retention behavior.
Move eligible offline evaluations, enrichment, summarization, or extraction to provider batch processing where the delay is acceptable.
Constrain output length and request structured fields when the product needs a compact answer. Trim conversation history using a tested memory strategy. Limit the tools exposed to each task. None of these choices should be accepted on token reduction alone; rerun the same evaluations and examine failure categories.
Current provider implementations reuse matching prompt prefixes under model-specific eligibility, lifetime, and pricing rules. Keep stable instructions, examples, and tool schemas before dynamic user and retrieval content. Track cached usage and misses rather than assuming a long prompt will be reused. A changed prefix, model, tool definition, or routing pattern can alter cache performance.
A forecast should show how product behavior creates cost. Build low, expected, and high scenarios from active users, tasks per user, attempts per task, calls per attempt, resource consumption per call, tool usage, success rate, and fixed operating cost.
The model should answer practical questions: What happens if conversation history doubles? What if the success rate drops and users retry? What if a larger model is required for one step? Which tenants or plans create negative margin? Where should the product add limits, asynchronous processing, a smaller model, or human escalation?
Botmer’s custom AI solutions work can include the evaluation, telemetry, and operating controls needed to move from a promising workflow to sustainable production use.
Provider prices and product capabilities change. This guide avoids fixed price claims and uses current documentation for the operating mechanisms described.
Botmer can help design the architecture, evaluation, and telemetry behind a production AI feature.