In LangGraph vs raw Python agent loops, the question is whether to manage the agent loop by hand or reach for an orchestration framework. At the simplest level, every agent is a loop: send messages to a language model, check for tool calls, execute them, and feed the results back. You can write that loop by hand in raw Python. Or you can use a framework like LangGraph to declare the same behavior as a graph.
This tutorial builds both implementations side by side. You will run a raw Python agent loop using an OpenAI-compatible API and an equivalent LangGraph StateGraph. The comparison is concrete: same tools, same system prompt, same model. You will see where each approach wins and when the framework overhead is worth it.
What You Will Build
You will implement a small research agent that answers a single question by searching the web and summarizing results. The agent has two tools:
search(query)returns a list of web results for a query.summarize(text)returns a concise summary of a block of text.
The agent loops until it produces a final answer (a message with no tool calls). Both implementations share the same tool definitions and system prompt, so differences in code and behavior come from the architecture choice, not the task.
| Aspect | Raw Python Loop | LangGraph |
|---|---|---|
| Loop control | Manual while loop | StateGraph with START and conditional edges |
| State management | Python list of messages | MessagesState with checkpointing |
| Tool execution | Manual dispatch | ToolNode |
| Persistence | None | InMemorySaver / PostgresSaver |
| Streaming | Manual SSE handling | Built-in .stream() |
| Human-in-the-loop | Manual interrupt checks | interrupt() node |
| Testing | Mock the model client | Mock at any node |
Set Up The Project
The project uses Python 3.14 and uv for dependency management. It runs entirely inside Docker Compose so the host does not need Python packages installed.
Directory Structure
langgraph-vs-raw-python/
├── src/
│ ├── Dockerfile
│ ├── mock_service.py # Mock search/summarize HTTP service
│ └── agents/
│ ├── __init__.py
│ ├── mock.py # Mock model and client for tests and demos
│ ├── tools.py # Shared tool definitions
│ ├── raw_loop.py # Raw Python agent loop
│ └── langgraph_agent.py # LangGraph StateGraph agent
├── tests/
│ ├── test_raw_loop.py
│ ├── test_langgraph_agent.py
│ └── conftest.py
├── Makefile
├── docker-compose.yml
└── pyproject.tomlInitialize The Project
uv init langgraph-vs-raw-python
cd langgraph-vs-raw-python
uv add langgraph==1.2.11 langchain-core==1.6.3 openai==3.16.2 fastapi uvicorn
uv add --dev pytest pytest-asyncio
Note: The Docker sandbox uses a mock search and summarize service (
mock_service.py). You do not need real API keys to run the code. Both agents accept an injectable model client so tests can substitute a mock.
Define Shared Tools And The Mock Service
Both agents need the same two tools. Create src/agents/tools.py:
"""Shared tool definitions for the research agent."""
from langchain_core.tools import tool
@tool
def search(query: str) -> list[str]:
"""Search the web for a query and return matching result snippets.
Args:
query: The search query string.
Returns:
A list of result snippet strings.
"""
raise NotImplementedError("replace with real search backend")
@tool
def summarize(text: str) -> str:
"""Summarize a block of text into a concise paragraph.
Args:
text: The text to summarize.
Returns:
A concise summary string.
"""
raise NotImplementedError("replace with real summarization backend")The langchain_core.tools.tool decorator turns a plain function into a StructuredTool with an auto-inferred schema. Both LangGraph’s ToolNode and the raw loop can consume these tools.
The mock service at src/mock_service.py provides a local FastAPI endpoint that returns canned results for known queries:
"""Mock HTTP service that simulates web search and summarization."""
from fastapi import FastAPI
from pydantic import BaseModel
class SearchRequest(BaseModel):
query: str
class SummarizeRequest(BaseModel):
text: str
app = FastAPI(title="Mock Research Service")
_FAKE_INDEX = {
"langgraph": [
"LangGraph is a state graph orchestration framework for building agents.",
"It provides checkpointing, persistence, and human-in-the-loop support.",
],
"raw agent loop": [
"A raw agent loop manually manages message history and tool dispatch.",
"It gives full control but requires custom error handling and state.",
],
}
@app.post("/search")
def mock_search(request: SearchRequest) -> dict:
"""Return canned search results for known queries."""
snippets = _FAKE_INDEX.get(request.query.lower(), ["No results found."])
return {"results": snippets}
@app.post("/summarize")
def mock_summarize(request: SummarizeRequest) -> dict:
"""Return a canned summary."""
text = request.text.strip()
if not text:
return {"summary": "No content to summarize."}
return {"summary": f"Summary: {text[:100]}..."}Raw Python Agent Loop
The raw loop manages everything manually. Create src/agents/raw_loop.py. It accepts a model client (any object with a chat.completions.create-compatible interface), a list of tools, and an optional system prompt. The function returns the final assistant message content.
"""Raw Python agent loop using an OpenAI-compatible client."""
import json
from typing import Any, Protocol
class ToolFunc(Protocol):
"""A callable tool compatible with the raw loop."""
def __call__(self, **kwargs: Any) -> str:
...
def _format_tool_for_openai(func: ToolFunc, description: str) -> dict:
"""Build an OpenAI tool schema from a tool function."""
import inspect
sig = inspect.signature(func)
properties = {}
for name, param in sig.parameters.items():
properties[name] = {
"type": "string",
"description": param.description or "",
}
return {
"type": "function",
"function": {
"name": func.__name__,
"description": description,
"parameters": {
"type": "object",
"properties": properties,
"required": list(properties),
},
},
}
def run_raw_agent(
client: Any,
tools: list[tuple[ToolFunc, str]],
question: str,
*,
system_prompt: str = (
"You are a research agent. Use search and summarize tools "
"to answer questions. Stop when you have enough information."
),
model: str = "gpt-4o-mini",
max_turns: int = 10,
) -> str:
"""Run a raw tool-calling loop against an OpenAI-compatible client.
Args:
client: An OpenAI-compatible client with ``chat.completions.create``.
tools: List of (callable, description) pairs.
question: The user's question.
system_prompt: System message for the model.
model: Model name to pass to the client.
max_turns: Maximum number of tool-calling iterations.
Returns:
The final assistant message content.
"""
tool_map = {t.__name__: t for t, _ in tools}
tool_schemas = [_format_tool_for_openai(t, d) for t, d in tools]
messages: list[dict] = [
{"role": "system", "content": system_prompt},
{"role": "user", "content": question},
]
for _ in range(max_turns):
response = client.chat.completions.create(
model=model,
messages=messages,
tools=tool_schemas,
)
msg = response.choices[0].message
if not msg.tool_calls:
return msg.content or ""
messages.append(
{
"role": "assistant",
"content": msg.content or "",
"tool_calls": msg.tool_calls,
}
)
for call in msg.tool_calls:
func = tool_map.get(call.function.name)
if func is None:
result = f"Error: tool {call.function.name!r} is not available."
else:
args = json.loads(call.function.arguments)
try:
result = str(func(**args))
except Exception as exc:
result = f"Error calling {call.function.name}: {exc}"
messages.append(
{
"role": "tool",
"tool_call_id": call.id,
"content": result,
}
)
return msg.content or ""Key observations about the raw loop:
- You manually format the tool schema as an OpenAI-compatible dict. If you switch providers from OpenAI to Anthropic, you must rewrite this conversion.
- You manually manage the message list — appending assistant messages with
toolcalls, then tool messages withtoolcall_id. - You handle errors inline — a tool failure is caught and turned into a string error message.
- No persistence — if the process crashes after a tool call, the work is lost.
- No streaming — the loop calls
createsynchronously and processes the full response.
LangGraph StateGraph Agent
The LangGraph version declares the same behavior as a graph. Create src/agents/langgraph_agent.py:
"""LangGraph agent built with StateGraph and ToolNode."""
from typing import Any
from langchain_core.messages import AIMessage, HumanMessage, SystemMessage
from langgraph.checkpoint.memory import InMemorySaver
from langgraph.graph import MessagesState, StateGraph, START, END
from langgraph.prebuilt import ToolNode
from agents.tools import search, summarize
SYSTEM_PROMPT = (
"You are a research agent. Use search and summarize tools "
"to answer questions. Stop when you have enough information."
)
_TOOLS = [search, summarize]
def _call_model(state: MessagesState, config: dict[str, Any]) -> dict[str, Any]:
"""Invoke the model with the current message list."""
model = config["configurable"]["model"]
response = model.invoke(
[SystemMessage(content=SYSTEM_PROMPT)] + state["messages"]
)
return {"messages": response}
def _should_continue(state: MessagesState) -> str:
"""Route to 'tools' if the last message has tool calls, else end."""
last = state["messages"][-1]
if isinstance(last, AIMessage) and last.tool_calls:
return "tools"
return END
def create_agent_graph(model: Any) -> Any:
"""Build and compile a LangGraph agent.
Args:
model: A LangChain-compatible chat model with ``invoke``.
Returns:
A compiled CompiledStateGraph ready to call.
"""
graph = StateGraph(MessagesState)
graph.add_node("agent", _call_model)
graph.add_node("tools", ToolNode(_TOOLS))
graph.add_edge(START, "agent")
graph.add_conditional_edges(
"agent", _should_continue, {"tools": "tools", END: END}
)
graph.add_edge("tools", "agent")
checkpointer = InMemorySaver()
return graph.compile(checkpointer=checkpointer)
def run_langgraph_agent(model: Any, question: str) -> str:
"""Run the LangGraph agent and return the final answer.
Args:
model: A LangChain-compatible chat model.
question: The user's question.
Returns:
The final assistant message content.
"""
graph = create_agent_graph(model)
result = graph.invoke(
{"messages": [HumanMessage(content=question)]},
{"configurable": {"model": model, "thread_id": "thread-1"}},
)
return result["messages"][-1].contentKey observations about the LangGraph version:
- Declarative structure — the graph topology (agent -> tools -> agent -> conditional) is defined separately from the node logic. Reading the
addedgeandaddconditional_edgescalls, you understand the flow at a glance. - Automatic tool execution —
ToolNodehandles schema conversion, function dispatch, error handling, and message formatting. The node functioncallmodelnever touches tool schemas. - Built-in checkpointing —
InMemorySaverpersists state between turns. WithPostgresSaveryou get durability across restarts and distributed execution. See the LangGraph checkpointing guide for details. - Streaming —
graph.stream()yields each message as it is produced, including intermediate tool results, without extra code. Learn about streaming in LangGraph. - Human-in-the-loop — add `interrupt()` calls in any node to pause execution for human review. No manual signal handling needed.
Mock Model For Testing
Both agents accept an injectable model so tests do not need real API calls. The mock classes live in src/agents/mock.py, and tests/conftest.py re-exports them as pytest fixtures:
"""Mock model and client implementations for testing and demos."""
import json
from typing import Any
from langchain_core.messages import AIMessage
class MockToolCall:
"""Mimics an OpenAI-style tool call response object."""
def __init__(self, name: str, args: dict, call_id: str = "call_1"):
self.id = call_id
self.function = type(
"Func", (), {"name": name, "arguments": json.dumps(args)}
)()
class MockChoice:
def __init__(self, message: Any):
self.message = message
class MockResponse:
def __init__(self, message: Any):
self.choices = [MockChoice(message)]
class MockOpenAIMessage:
"""Mimics an OpenAI-style message with optional tool calls."""
def __init__(self, content: str = "", tool_calls: list | None = None):
self.content = content
self.tool_calls = tool_calls or []
class MockClient:
"""Mock OpenAI-compatible client that simulates a tool-calling agent."""
def __init__(self, responses: list[Any]):
self._responses = responses
self._index = 0
self.chat = type(
"Chat",
(),
{"completions": type("Completions", (), {"create": self._create})()},
)()
def _create(self, **kwargs):
if self._index >= len(self._responses):
resp = MockResponse(MockOpenAIMessage(content="I don't know."))
else:
resp = self._responses[self._index]
self._index += 1
return resp
class MockLangChainModel:
"""Mock LangChain chat model that returns canned messages."""
def __init__(self, responses: list[Any]):
self._responses = responses
self._index = 0
def invoke(self, messages: list, **kwargs):
if self._index >= len(self._responses):
msg = AIMessage(content="I don't know.")
else:
msg = self._responses[self._index]
self._index += 1
return msgThe tests/conftest.py file re-exports these classes and provides pytest fixtures:
"""Shared test fixtures for the LangGraph vs raw loop comparison."""
import pytest
from agents.mock import (
MockClient,
MockLangChainModel,
MockOpenAIMessage,
MockResponse,
MockToolCall,
)
__all__ = [
"MockClient",
"MockLangChainModel",
"MockOpenAIClient",
"MockOpenAIMessage",
"MockResponse",
"MockToolCall",
]
# Re-export for backward compatibility with test imports
MockOpenAIClient = MockClient
@pytest.fixture
def raw_client_factory():
"""Return the MockClient class for constructing test clients."""
return MockClient
@pytest.fixture
def langchain_model_factory():
"""Return the MockLangChainModel class for constructing test models."""
return MockLangChainModelWrite The Tests
Create tests/testrawloop.py:
"""Tests for the raw Python agent loop."""
import json
from agents.raw_loop import run_raw_agent
from tests.conftest import (
MockOpenAIClient,
MockOpenAIMessage,
MockResponse,
MockToolCall,
)
def _make_search(name: str = "search"):
"""Create a fake search function with the given name."""
def fake(**kwargs):
"""Search mock."""
return json.dumps(
[
"LangGraph is a state graph framework.",
"It provides checkpointing and persistence.",
]
)
fake.__name__ = name
return fake
def _make_summarize(name: str = "summarize"):
"""Create a fake summarize function with the given name."""
def fake(**kwargs):
"""Summarize mock."""
return "LangGraph provides checkpointing and persistence."
fake.__name__ = name
return fake
def test_raw_agent_calls_tools_then_answers(raw_client_factory):
"""The raw loop calls tools then produces a final answer."""
tools = [
(_make_search(), "Search the web for information about a query."),
(_make_summarize(), "Summarize a block of text concisely."),
]
client = raw_client_factory(
[
# Turn 1: call search
MockResponse(
MockOpenAIMessage(
tool_calls=[MockToolCall("search", {"query": "langgraph"})],
)
),
# Turn 2: call summarize
MockResponse(
MockOpenAIMessage(
tool_calls=[MockToolCall("summarize", {"text": "some results"})],
)
),
# Turn 3: final answer
MockResponse(
MockOpenAIMessage(
content="LangGraph is an agent framework."
)
),
]
)
result = run_raw_agent(client, tools, "What is LangGraph?")
assert "LangGraph is an agent framework." in result
def test_raw_agent_handles_max_turns(raw_client_factory):
"""The raw loop returns the last message when max_turns is reached."""
tools = [(_make_search(), "Search the web.")]
client = raw_client_factory(
[
MockResponse(
MockOpenAIMessage(
tool_calls=[MockToolCall("search", {"query": "x"})],
)
),
MockResponse(
MockOpenAIMessage(
tool_calls=[MockToolCall("search", {"query": "y"})],
)
),
]
)
result = run_raw_agent(client, tools, "test", max_turns=2)
assert result == ""Create tests/testlanggraphagent.py:
"""Tests for the LangGraph agent."""
from langchain_core.messages import AIMessage, HumanMessage, ToolCall
from langchain_core.tools import tool as lc_tool
from agents.langgraph_agent import create_agent_graph, run_langgraph_agent
from tests.conftest import MockLangChainModel
def _make_tools():
"""Create fresh tool instances for testing."""
@lc_tool
def search(query: str) -> list[str]:
"""Search the web for a query."""
return ["LangGraph is a state graph framework."]
@lc_tool
def summarize(text: str) -> str:
"""Summarize a block of text."""
return "LangGraph provides checkpointing."
return search, summarize
def test_langgraph_agent_runs_to_completion():
"""The LangGraph agent runs through tool calls and produces an answer."""
search, summarize = _make_tools()
import agents.langgraph_agent as lg
original_tools = lg._TOOLS
lg._TOOLS = [search, summarize]
msg1 = AIMessage(
content="",
tool_calls=[ToolCall(name="search", args={"query": "langgraph"}, id="call_1")],
)
msg2 = AIMessage(
content="",
tool_calls=[ToolCall(name="summarize", args={"text": "results"}, id="call_2")],
)
msg3 = AIMessage(content="LangGraph is an agent framework.")
model = MockLangChainModel([msg1, msg2, msg3])
try:
result = run_langgraph_agent(model, "What is LangGraph?")
assert "LangGraph is an agent framework." in result
finally:
lg._TOOLS = original_tools
def test_langgraph_agent_graph_structure():
"""The compiled graph has the expected nodes and edges."""
msg = AIMessage(content="Done.")
model = MockLangChainModel([msg, msg, msg])
search, summarize = _make_tools()
import agents.langgraph_agent as lg
original_tools = lg._TOOLS
lg._TOOLS = [search, summarize]
try:
graph = create_agent_graph(model)
assert "agent" in graph.nodes
assert "tools" in graph.nodes
finally:
lg._TOOLS = original_toolsRun The Code
Build and run inside Docker:
docker compose build
docker compose run --rm app make run
docker compose run --rm app make testThe rawloop entry point constructs a mock OpenAI client, and langgraphagent constructs a mock LangChain model. Both produce the same final answer from the same tools and prompt.
When The Framework Pays Off
| Scenario | Raw Loop | LangGraph | Why |
|---|---|---|---|
| Simple single-turn tool call | Good fit | Overkill | One client.chat.completions.create call |
| Multi-step tool chains | Ad-hoc | Natural | addconditionaledges declares the loop |
| Crash recovery | Lost | Resume from checkpoint | InMemorySaver or PostgresSaver persists state |
| Human review mid-flow | Manual signal | interrupt() | Framework pauses and resumes state |
| Observability | Custom logging | LangSmith traces | Built-in tracing of every node and tool call |
| Streaming to UI | Manual SSE | .stream() | Yields messages as they are produced |
Use the raw loop when you need a single tool call, no persistence, and full control over the request format. Use LangGraph when your agent survives crashes, involves human decisions, runs for many turns, or needs observability. For a deeper comparison, see the side-by-side implementation above.
Conclusion
LangGraph is not a replacement for understanding how raw Python agent loops work. It is a tool that eliminates boilerplate when your agent needs persistence, human oversight, or complex control flow. For a simple ask-and-answer tool loop, the raw Python approach is fewer lines and fewer dependencies. For anything that survives a restart, involves human review, or grows beyond a single loop, LangGraph’s checkpointing and graph primitives pay for their overhead.
Build the project, run the tests, and see for yourself which approach fits your 2026 agent use case.



