Building Resilient LangGraph Workflows with Async Postgres Checkpointing
S L Manikanta
Sep 13, 2026 • 6 min read
bolt Key Takeaways
- Use AsyncPostgresSaver (not the sync MemorySaver) for any LangGraph workflow that runs longer than a single HTTP request.
- Pair it with an asyncpg connection pool (min_size=2, max_size=10) to avoid connection exhaustion under concurrent threads.
- Call await checkpointer.setup() once at startup to create the checkpoint tables before compiling your graph.
- Resume interrupted runs by passing the same thread_id in config — LangGraph replays from the last committed checkpoint automatically.
list On this page expand_more
- Environment
- 1. How LangGraph Checkpointing Works
- 2. Postgres Schema Setup
- 3. Building the Graph with a Durable Checkpointer
- 4. Running Threads and Resuming After Failures
- 5. Concurrent Threads and Connection Pool Sizing
- 6. FastAPI Integration Pattern
- 7. Checkpoint Retention and Cleanup
- 8. Failure Modes and Fixes
- Next Steps
Want to build production-ready AI?
Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.
[!NOTE] Quick Setup (2 minutes): Add durable Postgres checkpointing to any LangGraph graph:
pip install langgraph-checkpoint-postgres asyncpg psycopg2-binaryimport asyncpg from langgraph.checkpoint.postgres.aio import AsyncPostgresSaver pool = await asyncpg.create_pool( "postgresql://user:password@localhost:5432/langgraph", min_size=2, max_size=10, ) checkpointer = AsyncPostgresSaver(pool) await checkpointer.setup() # creates checkpoint tables once graph = your_graph_builder.compile(checkpointer=checkpointer)
In-memory state works fine for toy demos. The moment your LangGraph agent hits a network timeout on the 8th tool call of a 12-step pipeline, you lose everything and restart from zero. That’s the production failure mode AsyncPostgresSaver exists to solve.
This guide covers the exact setup: asyncpg pool configuration, schema migration, thread-safe concurrent runs, and failure recovery.
Environment
| Package | Version |
|---|---|
langgraph | 0.2.14 |
langgraph-checkpoint-postgres | 2.0.1 |
asyncpg | 0.29.0 |
psycopg2-binary | 2.9.9 |
| Python | 3.11+ |
| PostgreSQL | 15+ |
1. How LangGraph Checkpointing Works
Every time a graph node completes, LangGraph serializes the entire state and writes a checkpoint. On the next invocation with the same thread_id, LangGraph loads the latest checkpoint and resumes from where it left off.
sequenceDiagram
participant Client
participant Graph
participant Checkpointer
participant Postgres
Client->>Graph: invoke(input, config={thread_id: "run-42"})
Graph->>Checkpointer: load_checkpoint(thread_id="run-42")
Checkpointer->>Postgres: SELECT * FROM checkpoints WHERE thread_id=...
Postgres-->>Checkpointer: last_checkpoint (or None)
Checkpointer-->>Graph: resume state
loop Each node
Graph->>Graph: execute node
Graph->>Checkpointer: put_checkpoint(state)
Checkpointer->>Postgres: INSERT INTO checkpoints ...
end
Graph-->>Client: final state
The checkpoint store schema uses three tables: checkpoints, checkpoint_blobs, and checkpoint_writes. You don’t manage these directly — await checkpointer.setup() handles creation.
2. Postgres Schema Setup
Run this once at application startup, before compiling the graph:
import asyncio
import asyncpg
from langgraph.checkpoint.postgres.aio import AsyncPostgresSaver
async def create_checkpointer() -> AsyncPostgresSaver:
pool = await asyncpg.create_pool(
dsn="postgresql://langgraph_user:secret@localhost:5432/langgraph_db",
min_size=2,
max_size=10,
command_timeout=60,
max_inactive_connection_lifetime=300.0,
)
checkpointer = AsyncPostgresSaver(pool)
# Creates checkpoint tables if they don't exist.
# Safe to call on every startup — it's idempotent.
await checkpointer.setup()
return checkpointer
The setup() call is idempotent: calling it multiple times on the same database is safe. Run it during application startup, not on every request.
3. Building the Graph with a Durable Checkpointer
from langgraph.graph import StateGraph, MessagesState, END
from langchain_anthropic import ChatAnthropic
from langchain_core.messages import SystemMessage
llm = ChatAnthropic(model="claude-3-5-sonnet-20241022")
def call_model(state: MessagesState):
messages = [SystemMessage(content="You are a helpful assistant.")] + state["messages"]
response = llm.invoke(messages)
return {"messages": [response]}
builder = StateGraph(MessagesState)
builder.add_node("agent", call_model)
builder.set_entry_point("agent")
builder.add_edge("agent", END)
# Inject the checkpointer at compile time
checkpointer = await create_checkpointer()
graph = builder.compile(checkpointer=checkpointer)
4. Running Threads and Resuming After Failures
Each unique thread_id in the run config is an isolated conversation or workflow run. LangGraph persists checkpoints keyed by this ID.
import uuid
# Start a new run
thread_id = str(uuid.uuid4())
config = {"configurable": {"thread_id": thread_id}}
result = await graph.ainvoke(
{"messages": [{"role": "user", "content": "Analyze the Q3 sales report."}]},
config=config,
)
print(f"Run {thread_id} completed: {result['messages'][-1].content[:100]}")
If the process crashes mid-run, resume with the same thread_id:
# On restart — graph loads from last committed checkpoint automatically
resumed_result = await graph.ainvoke(
{"messages": []}, # empty input; state is loaded from checkpoint
config={"configurable": {"thread_id": thread_id}},
)
LangGraph compares the loaded checkpoint against the graph topology, identifies which nodes already completed, and continues execution from the first uncommitted node.
5. Concurrent Threads and Connection Pool Sizing
Under concurrent load, each ainvoke call opens a database connection to read and write checkpoints. If your pool is too small, requests queue behind connection acquisition and spike your P99 latency.
# Production pool sizing formula:
# max_size >= expected_concurrent_graph_runs * checkpoint_writes_per_run / avg_write_duration_sec
#
# For 50 concurrent runs, 5 checkpoint writes each, 10ms per write:
# max_size >= 50 * 5 * 0.010 = 2.5 → use max_size=10 with headroom
pool = await asyncpg.create_pool(
dsn=DATABASE_URL,
min_size=2, # keep alive for low-traffic periods
max_size=10, # scale to this under load
max_queries=50_000, # recycle after 50k queries to prevent bloat
max_inactive_connection_lifetime=300.0, # close idle connections after 5 min
)
Test your pool under realistic concurrency before deploying:
# Simulate 20 concurrent LangGraph threads hitting Postgres
python -c "
import asyncio, asyncpg
from langgraph.checkpoint.postgres.aio import AsyncPostgresSaver
async def stress_test():
pool = await asyncpg.create_pool('postgresql://...', min_size=2, max_size=10)
checkpointer = AsyncPostgresSaver(pool)
await checkpointer.setup()
# ... compile graph and run 20 concurrent ainvoke calls
await pool.close()
asyncio.run(stress_test())
"
6. FastAPI Integration Pattern
The recommended pattern for FastAPI: create the pool at app startup and close it on shutdown.
from contextlib import asynccontextmanager
from fastapi import FastAPI
import asyncpg
from langgraph.checkpoint.postgres.aio import AsyncPostgresSaver
checkpointer: AsyncPostgresSaver = None
@asynccontextmanager
async def lifespan(app: FastAPI):
global checkpointer
pool = await asyncpg.create_pool(
dsn="postgresql://user:pass@localhost/langgraph",
min_size=2,
max_size=10,
)
checkpointer = AsyncPostgresSaver(pool)
await checkpointer.setup()
yield
await pool.close()
app = FastAPI(lifespan=lifespan)
@app.post("/run")
async def run_agent(thread_id: str, message: str):
config = {"configurable": {"thread_id": thread_id}}
result = await graph.ainvoke(
{"messages": [{"role": "user", "content": message}]},
config=config,
)
return {"output": result["messages"][-1].content}
7. Checkpoint Retention and Cleanup
Checkpoints accumulate indefinitely if you don’t prune them. Add a background cleanup job:
import asyncpg
from datetime import datetime, timedelta, timezone
async def prune_old_checkpoints(pool: asyncpg.Pool, retention_days: int = 7):
cutoff = datetime.now(timezone.utc) - timedelta(days=retention_days)
async with pool.acquire() as conn:
deleted = await conn.execute(
"""
DELETE FROM checkpoints
WHERE created_at < $1
""",
cutoff,
)
print(f"Pruned checkpoints older than {retention_days} days: {deleted}")
Run this as a scheduled task (e.g., daily via APScheduler or a cron job hitting a /admin/prune endpoint).
8. Failure Modes and Fixes
| Symptom | Root Cause | Fix |
|---|---|---|
asyncpg.exceptions.TooManyConnectionsError | Pool max_size too low for concurrent runs | Increase max_size or add a read replica |
asyncpg.exceptions.InterfaceError: connection is closed | Idle connection closed by Postgres idle_in_transaction_session_timeout | Set max_inactive_connection_lifetime < Postgres timeout |
| State missing on resume | thread_id mismatch between runs | Always pass the same thread ID from durable storage (e.g., a database row) |
KeyError on state resume | Graph state schema changed after checkpoint was written | Handle missing keys with .get() defaults in node functions |
setup() raises PermissionError | Postgres user lacks CREATE TABLE rights | Grant CREATE on the target schema to the app user |
Next Steps
For workflows that require human approval between nodes, the checkpointer is the foundation for interrupt queues. See Implementing Dynamic Human-in-the-Loop Approval Queues in LangGraph with FastAPI for the full pattern.
To benchmark LangGraph against PydanticAI under concurrent load with Postgres-backed state, see LangGraph vs PydanticAI: State Machine Architecture, Latency, and Memory Footprint.
If you’re building the agent state schema itself, the LangGraph Complete Guide covers TypedDict state design, edge conditions, and subgraph composition.
Want to build production-ready AI?
Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.
Written by S L Manikanta
AI Engineer specializing in agentic workflows, multi-step LLM validation pipelines, and secure cloud environments. Sharing practical lessons from building software.
Related Articles
AI Agent Architecture Patterns: A Guide for Platform Engineers (2026)
A technical comparison of AI agent architectures. Learn when to use Prompt Chaining, Routing, Orchestrator-Workers, and Cyclic State Graphs (LangGraph).
What is an AI Agent Harness? Complete Technical Reference (2026)
A comprehensive technical reference on AI Agent Harnesses. Learn architecture, security, cost optimization, and how to deploy LangGraph agents into production with custom harnesses.
Building StoxFlow: Hybrid Local/Cloud AI Agents Architecture Complete Guide (2026)
How to build a decoupled three-tier AI stock research agent using LangGraph, FastAPI, and hybrid LLM routing (Ollama/Gemini) for optimal cost and performance.