Last month Amir and I built an automated agent to migrate customer SQL database schemas between PostgreSQL versions. During our first end-to-end dry run on test data, the agent generated an invalid foreign key constraint, hit an execution error, panicked, and proceeded to attempt three more nonsensical queries before we manually terminated the process.
Uncontrolled autonomous loops are terrifying in production. If an agent encounters an error or hallucinated output, it needs two architectural capabilities: 1. Self-Correction: It must analyze its own execution traceback and retry with an amended approach. 2. Human-in-the-Loop (HITL) Interruption: For high-risk or irreversible state changes (dropping tables, sending customer emails, executing wire transfers), the state machine must pause, serialize state to a database, and wait for human approval before resuming.
We rebuilt our agent using LangGraph, LangChain's state graph library. Here is the exact architectural blueprint we use to build self-healing agent workflows with persistent human verification.
Graph State and Cyclic Topology
Unlike linear chains (like standard LangChain RunnableSequence), LangGraph treats agent execution as a directed cyclic graph. Nodes represent functions, and edges represent conditional routing logic based on current state:
ββββββββββββββββ
β Draft Node β
ββββββββ¬ββββββββ
β
βΌ
ββββββββββββββββ
βββββΊβ Validate Nodeβ
β ββββββββ¬ββββββββ
β β
Failed β βΌ Passed
[Max 3 Tries]β ββββββββββββββββ
βββββ€ Audit Node β
ββββββββ¬ββββββββ
β
βΌ High Risk?
ββββββββββββββββ
β HITL Pause ββββ Human Approves / Rejects
ββββββββ¬ββββββββ
β
βΌ Approved
ββββββββββββββββ
β Execute Node β
ββββββββββββββββImplementing the State Graph in Python
Here is the production implementation of our self-correcting agent state graph using LangGraph:
# schema_agent_graph.py
from typing import TypedDict, Annotated, List, Optional
import operator
from langgraph.graph import StateGraph, END
from langgraph.checkpoint.memory import MemorySaver
class AgentState(TypedDict):
schema_sql: str
target_dialect: str
attempts: int
validation_errors: Annotated[List[str], operator.add]
is_valid: bool
human_approved: Optional[bool]
def draft_migration_node(state: AgentState) -> dict:
"""Generates the initial SQL migration script or refactors based on errors."""
attempts = state.get("attempts", 0) + 1
errors = state.get("validation_errors", [])
if errors:
# Prompt LLM to fix specific validation errors from previous turn
fixed_sql = f"-- Fixed based on error: {errors[-1]}\nALTER TABLE users ADD COLUMN updated_at TIMESTAMP;"
return {"schema_sql": fixed_sql, "attempts": attempts}
initial_sql = "CREATE TABLE users (id SERIAL PRIMARY KEY, username VARCHAR(50));"
return {"schema_sql": initial_sql, "attempts": attempts}
def validate_sql_node(state: AgentState) -> dict:
"""Simulates dry-run execution against an isolated database container."""
sql = state["schema_sql"]
# Simulate validation failure on first attempt to demonstrate self-healing
if state["attempts"] == 1:
return {
"validation_errors": ["Syntax error: missing semicolon or invalid constraint"],
"is_valid": False
}
return {"validation_errors": [], "is_valid": True}
def should_retry_or_proceed(state: AgentState) -> str:
"""Conditional edge router: Decide whether to self-correct, fail, or proceed."""
if state["is_valid"]:
return "human_approval_checkpoint"
if state["attempts"] >= 3:
return "fail_node"
return "draft_migration"
def execute_node(state: AgentState) -> dict:
"""Executes the verified script against production."""
if not state.get("human_approved"):
raise PermissionError("Cannot execute unapproved migration.")
print("Executing migration on production database...")
return {}
# Build the Graph
builder = StateGraph(AgentState)
builder.add_node("draft_migration", draft_migration_node)
builder.add_node("validate_sql", validate_sql_node)
builder.add_node("execute_migration", execute_node)
builder.set_entry_point("draft_migration")
builder.add_edge("draft_migration", "validate_sql")
builder.add_conditional_edges(
"validate_sql",
should_retry_or_proceed,
{
"draft_migration": "draft_migration",
"human_approval_checkpoint": END, # Pause here for human review
"fail_node": END
}
)
# Persistent checkpointer saves thread state across processes
checkpointer = MemorySaver()
app = builder.compile(checkpointer=checkpointer, interrupt_before=["execute_migration"])Handling Human-in-the-Loop Interruption
The interrupt_before=["execute_migration"] directive is the core safety barrier. When the graph reaches this node, LangGraph halts execution, saves the complete thread state into the checkpointer, and yields control back to the caller:
# run_with_hitl.py
from schema_agent_graph import app
thread_config = {"configurable": {"thread_id": "migration_run_402"}}
# Step 1: Run graph until human approval pause
events = app.stream(
{"schema_sql": "", "target_dialect": "postgresql", "attempts": 0, "validation_errors": [], "is_valid": False, "human_approved": None},
config=thread_config
)
for event in events:
print(f"Step: {event}")
# Inspect current state waiting for approval
current_state = app.get_state(thread_config)
print("\n--- WAITING FOR OPERATOR APPROVAL ---")
print("Pending SQL:")
print(current_state.values["schema_sql"])
# Step 2: Human Operator clicks "Approve" in UI
app.update_state(
thread_config,
{"human_approved": True},
as_node="draft_migration"
)
# Step 3: Resume execution from checkpoint
resume_events = app.stream(None, config=thread_config)
for event in resume_events:
print(f"Resumed Step: {event}")Edge Router Failure Traps
When designing cyclic self-correcting graphs, watch out for these traps:
- Infinite Loop Exhaustion: Every cyclic path must decrement or increment an attempt counter. If an error persists across 3 consecutive cycles, break to an escalation node.
- Loss of Historical Errors: Notice the
Annotated[List[str], operator.add]annotation onvalidation_errors. This preserves previous execution errors in the state history so the LLM does not repeat the exact same mistake twice. - Database Checkpointer in Production: In memory checkpointers are lost if your pod restarts. Use
PostgresSaverin production so human approval links can be clicked hours or days later without losing agent memory.
Production Reliability Metrics
| Architectural Pattern | Autonomous Chain | Self-Correcting LangGraph + HITL |
|---|---|---|
| First-Attempt Pass Rate | 62% | 62% |
| Final Success Rate (with retries) | 62% (Crashes on failure) | 94% (Self-corrects in 2-3 passes) |
| Accidental Destructive Actions | High risk | 0% (Gated by human checkpoint) |
| Mean Time to Recovery | Manual intervention required | 8 seconds automated retry |
Autonomous agents become production-grade only when they know how to catch their own mistakes and know when to stop and ask for human verification. LangGraph provides the deterministic rails that make agentic workflows safe.
Comments
Comments are reviewed before appearing publicly.
No comments yet β be the first.