You've probably watched an agent do this: it runs a command, hits a KeyError, and instead of fixing anything, it pastes the exact same broken code back in, maybe with a comment apologizing for the mistake. Then it fails the same way again.
That's not a reasoning problem. It's an architecture problem. Nothing in the loop told the agent to actually look at what broke.
What "Self-Correcting" Actually Requires
Running a command is the easy half. Real environments throw transient network timeouts, schema drift, and small syntax slips that a human would catch in a glance and a model will confidently repeat. A production agent needs to notice the failure, figure out what actually caused it, and try something different, not just try again.
Agent Action ββ> Sandbox Execution ββ> Success? ββ> Done
β
βΌ (fails)
Isolate the actual error, strip the noise
β
βΌ
What broke, and why? (not "please fix this")
β
βΌ
Minimal patch ββ> Retry (capped at 3)
The Three Pieces
Filter before you show the model anything. Don't pass a full traceback with forty lines of node_modules internals. An error-sanitizing layer should strip external library noise and leave only the failing line, the local variables at that point, and the exception type: KeyError, SyntaxError, TypeError, whatever it actually is. A model reasoning over noise reasons badly.
Ask a real question, not "please fix this." "Please fix this error" gets you a guess. A structured prompt gets you a diagnosis: - What the spec actually required - What happened instead, the exact assertion or stdout - Why the previous attempt probably diverged - The single smallest change that would close that gap
Put a hard stop on the loop. This is the part people skip until it bites them:
MAX_ATTEMPTS = 3
attempt = 0
while attempt < MAX_ATTEMPTS:
result = execute_sandboxed_command(agent_action)
if result.exit_code == 0:
return result.output
attempt += 1
isolated_error = parse_traceback(result.stderr)
agent_action = request_reflection_patch(agent_action, isolated_error, attempt)
raise AgentCircuitBreakerTriggered("Exceeded maximum reflection budget.")
Without that counter, a subtly wrong fix can trigger a new error, which triggers another patch, which reintroduces the first error, forever. You want the agent to fail loudly and hand off, not burn your API budget arguing with itself overnight.
How Much This Actually Buys You
| Approach | How it feels to run | Roughly how often it recovers | What it's good for |
|---|---|---|---|
| Raw traceback, no filtering | Fast, cheap per attempt | Low, maybe a third of the time | Quick prototypes where a miss is fine |
| Filtered, structured diagnosis | Slower, costs more per attempt | Noticeably better, most attempts land | Day-to-day code generation |
| Two agents, one proposes and one critiques | Slowest, most expensive | Best, worth it for high-stakes changes | Database migrations, anything hard to undo |
Those recovery numbers move around a lot depending on your codebase and the kind of errors you're seeing, don't take them as guarantees, treat them as which direction to lean.
How You'd Actually Test This
Don't wait for a real production failure to find out if your self-correction loop works. Break things on purpose: corrupt a mock API response, delete a file the agent expects to find, feed it a malformed config in a staging container. Then watch whether it actually diagnoses the failure and adjusts, or just resubmits the same broken attempt with new confidence.
Frequently Asked Questions
What happens if the agent's fix introduces a new, different error?
The loop treats it as a fresh diagnostic state and keeps a running history of what's already been tried, so it doesn't quietly oscillate between two conflicting fixes without you noticing.
Three attempts feels arbitrary. Why not five, or ten?
It's a budget decision more than a technical one. Three catches the errors that are actually fixable through reflection, transient timeouts, small syntax mistakes, obvious type errors, without letting a genuinely broken approach burn a huge chunk of your token spend before a human steps in.
Does this work for errors that aren't code, like a bad API response?
Yes, the same pattern applies: isolate what actually went wrong (a 422 with a specific field error, say), reason about why the request was malformed, and patch the request rather than blindly retrying the same payload.
Comments
Comments are reviewed before appearing publicly.
No comments yet β be the first.