Prototype
Test whether the workflow is useful and technically possible.
A prototype earns the next investment by proving that a workflow is useful. Production engineering begins when failure has a real consequence and someone must own what happens next.
Test whether the workflow is useful and technically possible.
Real users and data create a need for controls, measurement, and recovery.
A named team owns reliability, security, releases, cost, and incidents.
A demo can look complete while still depending on manual fixes, broad credentials, friendly test data, one model configuration, and the memory of the person who built it. None of that makes the prototype bad. It means the prototype has done a different job from the production system you may now need.
A prototype should reduce uncertainty quickly. It might answer whether users want an AI-assisted workflow, whether a model can extract the required structure, whether retrieval finds useful context, or whether an integration is technically possible. Fast changes and manual intervention are reasonable while the team is still learning.
A production system must continue working when the easy assumptions stop holding. Inputs become messy. Permissions differ by user. An upstream API slows down. A model version changes. A retry creates a duplicate action. Someone needs to explain why an output was produced and recover from an incident without reconstructing the system from chat history.
The transition should follow the risk and responsibility carried by the workflow, rather than a fixed number of users or prompts.
This is why “we have only 50 users” can still describe a production-critical application. If those 50 people use it to approve payments, handle patient information, route support escalations, or operate a core internal process, the cost of a wrong or unavailable system can be material.
Look for a cluster of signals. One signal may justify a targeted control. Several together usually mean the product needs explicit production ownership.
The feature can create, approve, recommend, or block an action that affects a customer or the business.
Identity, financial, health, employee, customer, or confidential company information crosses the workflow.
Users expect the feature to be available and need a defined fallback when it is not.
More permissions, background jobs, models, data sources, and write actions increase the failure surface.
The answer to “who can diagnose, recover, and communicate?” can no longer be whoever built the demo.
These questions expose the operating requirements. They also keep the team from equating a polished interface with a dependable system.
Moving to production does not automatically require discarding the prototype. Start with an assessment of the code, architecture, data boundaries, provider dependencies, testability, and delivery workflow. Then choose the smallest path that can meet the actual requirements.
| Path | Use it when | What changes |
|---|---|---|
| Harden the current system | The architecture is understandable, the important boundaries are sound, and the code can be tested and operated | Add ownership, version control, review, environments, automated tests, observability, access controls, backup, and release procedures |
| Use a hybrid boundary | The prototype layer remains useful but critical data, workflows, or operations need tighter control | Keep the suitable interface or experimentation surface while moving sensitive logic, durable state, tools, or deployment into managed services |
| Replace selected components | A component cannot satisfy reliability, security, performance, or maintainability requirements without disproportionate work | Define an interface and migrate incrementally, preserving validated product behavior rather than rebuilding every feature at once |
A rewrite is a business decision, not a reaction to how the prototype was made. The team should be able to identify the specific requirement the current design cannot satisfy and explain why replacement is safer or cheaper than remediation.
“Production grade” is not one universal checklist. A private writing assistant and an agent that can issue refunds should not receive the same controls. Begin with the highest-consequence path and design outward from it.
Prompt instructions are product behavior, not an authorization layer. The application must still enforce identity, tenant boundaries, tool permissions, data access, and output handling. OWASP’s guidance treats prompt injection as an application risk that can influence downstream actions; retrieval or fine-tuning alone does not remove it.
For workflows with external content or tools, assume the model can receive hostile input. Keep credentials outside prompts, constrain the functions the model can call, validate arguments, apply least privilege, and add explicit approval where an action cannot be safely undone.
The goal of the first production sprint is evidence and control, not broad feature expansion.
Trace one important user task from input through data access, model calls, tool use, stored state, output, and follow-up action. Mark trust boundaries, sensitive data, manual interventions, and single points of failure. Agree on the success criteria and the unacceptable outcomes.
Create a small but representative evaluation set. Add deterministic checks for schemas, permissions, citations, or business rules where possible. Record the model, prompt, retrieval, and tool configuration with each run so results can be compared rather than remembered.
Establish development and production boundaries, code review, secret management, automated checks, deployment ownership, and rollback. Add limits and approval steps around write actions. Test what happens when dependencies time out or return malformed data.
Expose the workflow to a controlled user group. Watch task success, failure categories, latency, cost, overrides, and support requests. Feed production examples back into evaluation. Expand only when the team can explain both the result and the recovery path.
One well-observed workflow creates better evidence than a wide feature set with unknown failure behavior. If you need help assessing the prototype or delivering that slice, see Botmer’s custom AI solutions.
This guide uses established production and AI risk principles as reference points. The recommendations above are Botmer’s practical synthesis for product teams.
Share the current prototype, users, data boundaries, and hardest failure. Botmer will help define the engineering path around them.