Assigning multiple AI agents to collaborate on a single task looks impressive in concept demos. You spin up a researcher, a writer, and a code reviewer, connect them in a loop, and expect production-ready code.
In reality, multi-agent systems without formal orchestration contracts break quickly.
The researcher passes vague bullet points to the coder. The coder generates an incomplete script and passes it to the reviewer. The reviewer rejects it and sends it back to the researcher. Within five minutes, your agents enter an endless circular ping-pong loop, burning tens of dollars in API credits while making zero actual progress.
This is the orchestration deadlock. To build multi-agent teams that actually complete work, you need strict topology design, typed state contracts, and bounded iteration limits.
Here is an engineering blueprint for designing resilient multi-agent crews with CrewAI and LangGraph.
Sequential vs Graph Topologies: Choosing the Right Engine
Two primary orchestration patterns dominate production agent architectures:
| Architectural Pattern | Best Framework | Strengths | Failure Modes |
|---|---|---|---|
| Linear Assembly Line | CrewAI (Process.sequential) | Simple mental model, clear task handoffs, minimal boilerplate | Cannot handle non-linear retries or conditional branches |
| Stateful Directed Graph | LangGraph (StateGraph) | Cyclic execution, fine-grained state control, human-in-the-loop checkpoints | Higher setup complexity, requires explicit state schemas |
| Hierarchical Crew | CrewAI (Process.hierarchical) | Manager agent delegates subtasks dynamically | Manager model hallucinating task assignments or double-delegating |
For deterministic pipelines (e.g., Code Review $\to$ Static Analysis $\to$ Documentation), sequential execution is safest. For iterative refactoring, LangGraph's state machine provides the necessary control boundaries.
Defining Typed State Contracts with Pydantic
Never rely on unstructured text strings to pass context between agents. Define an explicit shared state model using Pydantic:
from pydantic import BaseModel, Field
from typing import List, Optional, Dict
import time
class AgentTaskResult(BaseModel):
agent_id: str
output_content: str
tokens_consumed: int
completed_at: float = Field(default_factory=time.time)
class OrchestrationState(BaseModel):
task_id: str
original_goal: str
current_iteration: int = 0
max_iterations: int = 5
is_terminal: bool = False
# Typed memory slots
research_summary: Optional[str] = None
source_code: Optional[str] = None
review_feedback: Optional[str] = None
# Audit trail
history: List[AgentTaskResult] = Field(default_factory=list)
errors: List[str] = Field(default_factory=list)By enforcing current_iteration < max_iterations, the system guarantees an exit path even if the reviewer agent repeatedly rejects the output.
Implementing Deadlock-Free LangGraph Workflows
To prevent feedback loops from running indefinitely, construct conditional edges that inspect the iteration counter:
from langgraph.graph import StateGraph, END
def router_node(state: OrchestrationState) -> str:
# Circuit breaker: Force termination on iteration cap
if state.current_iteration >= state.max_iterations:
print("[WARNING] Max iterations reached. Forcing termination.")
return "finalize"
if state.is_terminal:
return "finalize"
if state.review_feedback and "REVISE" in state.review_feedback:
return "refactor_code"
return "finalize"
# Construct the stateful graph
workflow = StateGraph(OrchestrationState)
workflow.add_node("research", research_agent_node)
workflow.add_node("write_code", code_generator_node)
workflow.add_node("review_code", code_reviewer_node)
workflow.add_node("finalize", finalizer_node)
workflow.set_entry_point("research")
workflow.add_edge("research", "write_code")
workflow.add_edge("write_code", "review_code")
# Conditional handoff with deadlock prevention
workflow.add_conditional_edges(
"review_code",
router_node,
{
"refactor_code": "write_code",
"finalize": "finalize"
}
)
workflow.add_edge("finalize", END)
app = workflow.compile()Operational Best Practices for Multi-Agent Fleets
When moving multi-agent crews into production, enforce these three engineering rules:
- Explicit Role Backstories: In CrewAI, keep agent roles narrow and mutually exclusive. An agent responsible for security reviews should not have write access to feature files.
- Context Compression on Handoff: Do not pass the entire conversation history of previous agents. Pass only the structured schema output of the preceding step to avoid blowing up context windows.
- Dead-Letter Queue (DLQ): When an agent fails three times consecutively, shunt the task to an operator alert queue rather than terminating the whole cluster.
Frequently Asked Questions
When should I use CrewAI versus LangGraph?
Use CrewAI when building role-based teams where agents collaborate in clear organizational hierarchies. Use LangGraph when you need complex cyclic graph logic, explicit state machines, and human approval nodes.
How do I stop agents from fighting over task definitions?
Use a centralized orchestrator or router agent that holds the ground truth task specification. Individual worker agents should never negotiate task requirements among themselves.
What is the biggest hidden cost of multi-agent systems?
Token inflation from repeated message context. Every time an agent receives context from three previous teammates, you are paying token fees for redundant information.
Comments
Comments are reviewed before appearing publicly.
No comments yet — be the first.