Your AI agent fixes one bug and breaks another: a founder repair workflow
RAAVHow to turn repeated AI coding fixes into a reproducible investigation, a bounded repair, and evidence you can review without reading every line of code.
Pause the next fix and make the failure reproducible
When an AI coding agent fixes one bug and breaks another, first capture a reproducible failure. Ask it to investigate before editing, explain the evidence for its suspected cause, and propose a narrow repair. Verify the original failure and relevant neighboring behavior, then save the result with the task.
The frustrating part is often the rhythm: you report a broken screen, the agent says it fixed the issue, and your next attempt reveals another problem. A longer instruction to “check everything” rarely gives you a clear way to judge what happened. The repair needs a specific starting condition, expected behavior, and evidence requirement.
You do not have to diagnose the code yourself. Your contribution is to describe the experience precisely and review the result. The agent can inspect the implementation and propose a technical explanation. Ask it to distinguish what it observed from what it suspects, so an appealing explanation does not become an untested fact.
The customer-portal example below is hypothetical. This is a project-management workflow, not a claim that one agent or model is uniquely responsible for recurring bugs. Use it with the coding tool you already have, and adapt the checks to the behavior that matters.
Prompt to give your agent
Pause implementation. Investigate this failure: [starting state], [steps], [expected behavior], [actual behavior], [environment]. Reproduce it if possible. Separate observations from hypotheses. Explain what evidence would distinguish the likely causes before proposing a repair.
Write a bug report around one failed journey
A useful report includes where you began, what account or sample data you used, the actions you took, and what happened. Describe the environment: local preview, test deployment, or live product. Include when the behavior changed if you know. Remove private customer information from any example you share.
For a customer portal, “the profile is broken” leaves room for several interpretations. “I signed in as the test customer, changed the display name, clicked Save, and refreshed; the old name returned” identifies a journey the agent can inspect. It also distinguishes a visual update from a persisted update.
Keep separate failures separate at first. A save failure and an email-delivery failure may share a cause, but that is something to establish. Combining them into one broad complaint can encourage a broad repair that makes it harder to see which behavior improved.
If you cannot reproduce the issue consistently, describe the attempts. Say what differed and what you do not know. Ask for a diagnostic step that is safe in the relevant environment. Inconsistent behavior is still evidence; it just requires a different investigation than a failure that happens every time.
Prompt to give your agent
Turn my report into a reproducible bug description. Include starting conditions, account role, sample data, exact actions, expected and actual results, and environment. Ask only for missing details that affect reproduction. Keep separate symptoms distinct until evidence establishes a shared cause.
Require a cause supported by evidence
Before accepting a repair plan, ask how the suspected cause explains the observed failure. “The database might be slow” is a hypothesis. A recorded failed save request or a missing persistence operation is evidence to investigate. You need the connection between the symptom and the planned change.
In our example, the displayed name changes immediately, but refreshing restores the old value. The agent might suspect the interface updates locally without successfully saving. It should inspect the save path rather than rewriting the entire profile page simply because the page is where you noticed the problem.
Ask what could disprove the explanation. If a save request succeeds and the correct record changes, the investigation needs to look elsewhere. A repair process should allow the agent to abandon a hypothesis when new evidence contradicts it.
If the agent cannot reproduce or inspect a key part, preserve that limitation. You can approve a diagnostic task to obtain the missing evidence. Calling a speculative rewrite a fix makes it difficult to learn from the result, even if the symptom happens to disappear.
Prompt to give your agent
Explain the suspected cause using the evidence gathered. Show how it accounts for the failed journey, what remains uncertain, and what finding would disprove it. If inspection is blocked, propose the smallest diagnostic task. Do not describe a speculative change as a verified fix.
Give the repair a boundary
A repair task should state the failing behavior, the intended correction, the areas likely to change, and the behaviors that must remain intact. Ask the agent to explain any expansion before it becomes work. This prevents an isolated bug from turning into an unreviewed product redesign.
For the profile save, the objective is that the authorized customer can save their own display name and see it after refreshing. Existing sign-in and profile loading should continue to work. The repair should not introduce a new account system unless the agent can demonstrate why the current one prevents a bounded solution.
A narrow change is not always the smallest possible number of edited lines. Some failures require a coordinated correction across the interface and server. The useful boundary is the user outcome and its prerequisites. Ask why each changed area is necessary rather than judging the repair only by its size.
Keep unrelated cleanup as a separate proposal. Improving names or reorganizing files may be sensible later, but it makes a repair harder to review if it obscures the relevant change. A clear boundary also gives another agent a better account of what was intentionally changed.
Check the original failure and its neighbors
A useful verification plan begins with the exact journey that failed. For the profile example, save a new name, reload, and confirm the value remains. Then check relevant nearby behavior: another customer should retain their own name, and the profile should still load after signing in again.
Ask which checks were automated and which were exercised through the interface. Each provides different evidence. An isolated check can validate part of the implementation while leaving the complete user journey untested. A manual demonstration can show the journey while missing another case. The report should make that coverage visible.
Record the environment and result, including anything blocked. If the agent checked locally but could not access a test deployment, do not let the report imply deployed behavior was verified. If sample data or an external service was unavailable, say which part remains open.
When a failure is important and reproducible, ask the agent whether a regression check would help detect it again. A useful check exercises the behavior that failed, rather than merely confirming that the newly written function exists. Keep the effort proportionate to the risk and frequency of the problem.
Prompt to give your agent
Verify the repair against the original failure steps. Check relevant neighboring behavior and explain why those checks matter. Report each result with its environment and evidence. Distinguish automated checks, interface checks, and untested areas. Recommend a regression check when it would meaningfully detect this failure again.
If the repair fails, keep what you learned
A failed repair should improve the next investigation. Record the proposed cause, the change made, the check that failed, and what the result rules out. Otherwise, a new session may repeat the same explanation and rewrite, especially if the only persistent note says “profile still broken.”
If the original save now works but another customer sees the wrong value, treat that result as a new observable failure. Ask the agent to relate it to the repair and inspect the affected behavior. Avoid stacking additional changes before understanding whether the previous change introduced the new problem.
You may need to restore an earlier working version. Ask for the exact change proposed for restoration and the consequences for unfinished work. Do not give an unrestricted instruction to erase recent changes; other tasks may be mixed into the working directory. Have ownership and scope established before a destructive correction.
When a repair repeatedly crosses areas you cannot evaluate, bring in technical review. A clear issue description and evidence trail make that review more efficient. The reviewer can assess the investigation without reconstructing every chat or accepting the last agent’s explanation on trust.
Prompt to give your agent
Record the failed repair attempt: original symptom, suspected cause, changed areas, checks and results, and what we learned. Identify any new failure separately. Propose the next diagnostic action. If restoration is advisable, explain the specific change and its effect on other work before proceeding.
Close the repair with a record the next agent can use
The closing record should contain the failure, the established cause or remaining uncertainty, the repair, and verification evidence. Mark whether the task is implemented, verified, or awaiting founder review. This vocabulary is a review convention you can adopt; it is not a guarantee supplied by every tool.
Persistent instructions can tell the agent to follow your repair workflow. Claude Code, Codex, Cursor, and Gemini CLI each document native context mechanisms, linked below. Keep the changing bug history with the project work rather than continually enlarging an instruction file.
RAAV holds persistent product and project memory, including tasks and verification records. It can give a later session a place to find the previous investigation. You still need accurate evidence and review; a recorded completion claim alone does not establish that a customer journey works.
Your immediate next step is small: pick the failure you can describe most precisely and ask for investigation before another fix. After the repair, run that same journey again. Close the task only with a result you can explain and the remaining limitations visible.
Continue the non-coding founder series
Start with who this approach is for, why the work needs product management, and why now. Then use the practical guide that matches the next problem in your project.
Keep the repair evidence with the work
Explore RAAV founder access for a shared record of tasks, decisions, and verification. Bring a recurring issue and the evidence your next agent needs to avoid starting the investigation again.
Request founder access