Skip to content
Botmer International®
Insights/Engineering workflow

A responsible Claude Code and Cursor workflow for production teams

AI coding tools can shorten the path from intent to a working diff. Production teams still need a controlled path from that diff to software they can explain, operate, and recover.

Human-controlled delivery loop
01 / INTENT

Define

Task contract, boundaries, evidence, and risk

02 / CHANGE

Generate

Inspect first, then produce a small reviewable diff

03 / EVIDENCE

Verify

Review behavior, tests, security, and edge cases

04 / OPERATION

Release

Approvals, observability, rollback, and ownership

The main production risk with AI-assisted coding is not that every generated line is poor. It is that plausible changes arrive faster than a team can build an accurate mental model of them. A responsible workflow protects review quality while preserving the speed gained during exploration and implementation.

Begin with a task contract

A short task contract gives the engineer and tool the same definition of done. It should describe the problem, the user-visible behavior, the allowed scope, the constraints that must remain true, and the evidence required before the change can be accepted.

Useful prompt structureContext → outcome → boundaries → acceptance → verification

Ask the tool to inspect the relevant implementation and explain its plan before editing. The engineer can correct a wrong understanding before it becomes a wide diff.

A good contract can be brief. For a form change, it might identify the route and component, state the new validation behavior, prohibit changes to the visual system, require the existing submission flow to remain intact, and name the desktop and mobile checks. For a data migration, it should also state reversibility, idempotency, and acceptable downtime.

Include what must remain unchanged

AI tools optimize for the requested outcome, but they do not automatically know which surrounding behavior carries business value. Explicitly preserve public interfaces, analytics events, accessibility semantics, visual behavior, database compatibility, security boundaries, and performance characteristics that matter to the task.

  • Point to the source of truth for the desired behavior instead of describing the entire system from memory.
  • Separate required changes from optional cleanup so the tool cannot use the task as permission for a broad refactor.
  • Name files, directories, commands, or actions that are out of bounds.
  • Define how errors, empty states, retries, and permission failures should behave.
  • State which checks are required and which user flows need manual inspection.

Give the tool bounded, maintained context

Repository guidance can make repeated interactions more consistent. Cursor project rules and a project’s Claude guidance can capture commands, architecture boundaries, conventions, and safety constraints that apply across tasks. Keep those instructions concise, version controlled, and reviewed like other engineering documentation.

Rules are not a substitute for the current task. Stale instructions can produce consistently wrong changes, and a large pile of overlapping rules can hide the constraint that actually matters. Prefer durable facts in repository guidance and put task-specific decisions in the task contract.

REPOSITORY

Stable conventions

Build commands, code organization, supported versions, testing approach, design system, and prohibited patterns.

TASK

Current intent

User outcome, files in scope, acceptance examples, known edge cases, and what must remain unchanged.

SESSION

Observed evidence

Relevant code, command output, failing tests, screenshots, logs, and decisions made while the work progresses.

Protect credentials and sensitive data

Do not paste production secrets, private customer data, access tokens, or unrestricted database exports into a prompt. Use the organization’s approved tool configuration and data policy. Keep credentials in secret management, use test fixtures or sanitized examples, and ensure terminal commands do not print sensitive values into the conversation or logs.

Tool permissions should match the task. Claude Code exposes allowed and disallowed tool settings and permission modes; Cursor exposes review controls for inspecting proposed changes. Avoid bypass modes as a convenience. A tool that can edit files, run arbitrary commands, access a network, or change infrastructure should receive only the authority required for the current work.

Create the smallest useful change

Before editing, ask the tool to locate the relevant paths, trace the current behavior, and identify the tests or call sites that define it. Review that explanation. Then request an implementation with a narrow diff. A small change is easier to understand, test, revert, and compare visually.

  1. Work in an isolated branch or worktree
    Preserve the known-good state and keep unrelated local changes out of the generated diff.
  2. Inspect before editing
    Read local instructions, find the source of behavior, follow imports and call sites, and identify existing patterns before proposing code.
  3. Review the plan
    Check that the proposed files and approach match the task. Correct mistaken assumptions before generation begins.
  4. Implement one coherent unit
    Avoid mixing a feature, dependency upgrades, formatting, and opportunistic refactors in the same change.
  5. Inspect the diff immediately
    Look for deleted behavior, widened scope, duplicated logic, placeholder content, changed public interfaces, and unexplained dependencies.
  6. Commit only understood work
    The engineer should be able to explain every material behavior, test, configuration change, and operational consequence.

Generated tests need the same scrutiny as generated application code. A test that repeats the implementation, mocks away the meaningful behavior, or asserts only the happy path can create confidence without evidence. Start from the task contract and verify externally observable behavior where practical.

Build evidence before accepting the result

Verification should combine human review, automated checks, and direct behavior inspection. None of the three replaces the others. Passing tests do not prove that the task was interpreted correctly. A clean visual check does not prove access control. A reviewer cannot reliably infer every runtime path from a large diff.

LayerQuestions to answerTypical evidence
IntentDoes this solve the stated user or operating problem without changing unrelated behavior?Task contract, acceptance examples, before-and-after behavior
ImplementationIs the code understandable, consistent with the system, and limited to the necessary scope?Line-by-line diff review, architecture review, dependency inspection
CorrectnessWhat happens on normal, empty, invalid, repeated, delayed, and concurrent paths?Focused tests, integration tests, manual scenarios, migration rehearsal
SecurityCan the change cross an authorization, data, secret, input, or supply-chain boundary?Threat-focused review, SAST, dependency and secret scanning, permission tests
ExperienceDoes it preserve layout, accessibility, responsive behavior, content, and interaction states?Same-viewport screenshots, keyboard checks, screen reader semantics, device testing
OperationCan the team observe, release, roll back, and support the change?Logs and metrics, staged release, rollback path, runbook or ownership note

Review the diff as if the author were unavailable

Cursor’s review interface is useful because it presents additions and removals for selective acceptance. The reviewer should still open the surrounding files, trace important call paths, and check generated configuration or migrations in their full context. Claude Code output should be reviewed through the same repository diff and pull request process as any other contribution.

Run the narrowest checks that genuinely exercise the change, followed by the project’s required build, type, lint, and security gates. When appearance matters, compare the same viewport and interaction state against a known-good screenshot. When behavior is asynchronous or stateful, verify retries, duplicate execution, partial failure, and rollback.

Scale controls with the consequence of failure

Not every edit needs the same ceremony. Risk rises with user impact, reversibility, data sensitivity, operational reach, and how difficult the behavior is to verify. The team can move quickly while still making those differences explicit.

Change levelExamplesMinimum control
LowIsolated copy correction, internal documentation, clearly bounded style token useDiff review, focused visual or documentation check, required repository checks
ModerateUser-facing component behavior, API client change, dependency update, background job logicTask contract, focused tests, full diff review, build gates, staging or representative manual check
HighAuthentication, payments, permissions, destructive migration, production infrastructure, sensitive data flowNamed human owner, threat and rollback review, independent approval, security checks, rehearsed release and recovery path

Branch protection and required status checks can enforce parts of the workflow. Code owners can require review from the people responsible for sensitive paths. Those mechanisms are valuable because they apply even when an AI tool produces a convincing explanation or an engineer is moving quickly.

Common failure modes

  • Wide cleanup hidden inside a small task: stop and separate the work before reviewing further.
  • Tests written to match the implementation: return to acceptance examples and observable behavior.
  • Confident claims without command evidence: run the actual check and inspect its complete result.
  • Permission bypass as the default: restore least privilege and approve only the action required.
  • Large generated diff reviewed as a summary: reduce the change or review it file by file with surrounding context.
  • No post-release owner: assign monitoring, rollback, incident, and follow-up responsibility before merge.
Botmer perspectiveAI-assisted delivery is strongest when the team improves both generation speed and the evidence required to release

Senior engineers such as Abdul Hai bring platform judgment to the workflow, while Botmer’s SaaS and web platform engineering teams can own the path from scoped change to production operation.

Reference points

This workflow applies established secure development and pull request controls to AI-assisted coding. Product interfaces and settings can change, so teams should verify current vendor documentation when configuring tools.

Add production judgment

Build a workflow your team can explain, review, and operate

Botmer can share senior engineering profiles or provide a delivery team for AI-native product and platform work.

Request profiles