Agentic AI #langgraph#postgres#checkpointing#async#state-machines#production#agents

Building Resilient LangGraph Workflows with Async Postgres Checkpointing

S

S L Manikanta

Sep 13, 2026 • 6 min read

bolt Key Takeaways

  • Use AsyncPostgresSaver (not the sync MemorySaver) for any LangGraph workflow that runs longer than a single HTTP request.
  • Pair it with an asyncpg connection pool (min_size=2, max_size=10) to avoid connection exhaustion under concurrent threads.
  • Call await checkpointer.setup() once at startup to create the checkpoint tables before compiling your graph.
  • Resume interrupted runs by passing the same thread_id in config — LangGraph replays from the last committed checkpoint automatically.
✉ Newsletter

Want to build production-ready AI?

Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.

[!NOTE] Quick Setup (2 minutes): Add durable Postgres checkpointing to any LangGraph graph:

pip install langgraph-checkpoint-postgres asyncpg psycopg2-binary
import asyncpg
from langgraph.checkpoint.postgres.aio import AsyncPostgresSaver

pool = await asyncpg.create_pool(
    "postgresql://user:password@localhost:5432/langgraph",
    min_size=2,
    max_size=10,
)
checkpointer = AsyncPostgresSaver(pool)
await checkpointer.setup()  # creates checkpoint tables once
graph = your_graph_builder.compile(checkpointer=checkpointer)

In-memory state works fine for toy demos. The moment your LangGraph agent hits a network timeout on the 8th tool call of a 12-step pipeline, you lose everything and restart from zero. That’s the production failure mode AsyncPostgresSaver exists to solve.

This guide covers the exact setup: asyncpg pool configuration, schema migration, thread-safe concurrent runs, and failure recovery.


Environment

PackageVersion
langgraph0.2.14
langgraph-checkpoint-postgres2.0.1
asyncpg0.29.0
psycopg2-binary2.9.9
Python3.11+
PostgreSQL15+

1. How LangGraph Checkpointing Works

Every time a graph node completes, LangGraph serializes the entire state and writes a checkpoint. On the next invocation with the same thread_id, LangGraph loads the latest checkpoint and resumes from where it left off.

sequenceDiagram
    participant Client
    participant Graph
    participant Checkpointer
    participant Postgres

    Client->>Graph: invoke(input, config={thread_id: "run-42"})
    Graph->>Checkpointer: load_checkpoint(thread_id="run-42")
    Checkpointer->>Postgres: SELECT * FROM checkpoints WHERE thread_id=...
    Postgres-->>Checkpointer: last_checkpoint (or None)
    Checkpointer-->>Graph: resume state

    loop Each node
        Graph->>Graph: execute node
        Graph->>Checkpointer: put_checkpoint(state)
        Checkpointer->>Postgres: INSERT INTO checkpoints ...
    end

    Graph-->>Client: final state

The checkpoint store schema uses three tables: checkpoints, checkpoint_blobs, and checkpoint_writes. You don’t manage these directly — await checkpointer.setup() handles creation.


2. Postgres Schema Setup

Run this once at application startup, before compiling the graph:

import asyncio
import asyncpg
from langgraph.checkpoint.postgres.aio import AsyncPostgresSaver

async def create_checkpointer() -> AsyncPostgresSaver:
    pool = await asyncpg.create_pool(
        dsn="postgresql://langgraph_user:secret@localhost:5432/langgraph_db",
        min_size=2,
        max_size=10,
        command_timeout=60,
        max_inactive_connection_lifetime=300.0,
    )
    checkpointer = AsyncPostgresSaver(pool)
    # Creates checkpoint tables if they don't exist.
    # Safe to call on every startup — it's idempotent.
    await checkpointer.setup()
    return checkpointer

The setup() call is idempotent: calling it multiple times on the same database is safe. Run it during application startup, not on every request.


3. Building the Graph with a Durable Checkpointer

from langgraph.graph import StateGraph, MessagesState, END
from langchain_anthropic import ChatAnthropic
from langchain_core.messages import SystemMessage

llm = ChatAnthropic(model="claude-3-5-sonnet-20241022")

def call_model(state: MessagesState):
    messages = [SystemMessage(content="You are a helpful assistant.")] + state["messages"]
    response = llm.invoke(messages)
    return {"messages": [response]}

builder = StateGraph(MessagesState)
builder.add_node("agent", call_model)
builder.set_entry_point("agent")
builder.add_edge("agent", END)

# Inject the checkpointer at compile time
checkpointer = await create_checkpointer()
graph = builder.compile(checkpointer=checkpointer)

4. Running Threads and Resuming After Failures

Each unique thread_id in the run config is an isolated conversation or workflow run. LangGraph persists checkpoints keyed by this ID.

import uuid

# Start a new run
thread_id = str(uuid.uuid4())
config = {"configurable": {"thread_id": thread_id}}

result = await graph.ainvoke(
    {"messages": [{"role": "user", "content": "Analyze the Q3 sales report."}]},
    config=config,
)
print(f"Run {thread_id} completed: {result['messages'][-1].content[:100]}")

If the process crashes mid-run, resume with the same thread_id:

# On restart — graph loads from last committed checkpoint automatically
resumed_result = await graph.ainvoke(
    {"messages": []},  # empty input; state is loaded from checkpoint
    config={"configurable": {"thread_id": thread_id}},
)

LangGraph compares the loaded checkpoint against the graph topology, identifies which nodes already completed, and continues execution from the first uncommitted node.


5. Concurrent Threads and Connection Pool Sizing

Under concurrent load, each ainvoke call opens a database connection to read and write checkpoints. If your pool is too small, requests queue behind connection acquisition and spike your P99 latency.

# Production pool sizing formula:
# max_size >= expected_concurrent_graph_runs * checkpoint_writes_per_run / avg_write_duration_sec
#
# For 50 concurrent runs, 5 checkpoint writes each, 10ms per write:
# max_size >= 50 * 5 * 0.010 = 2.5 → use max_size=10 with headroom

pool = await asyncpg.create_pool(
    dsn=DATABASE_URL,
    min_size=2,      # keep alive for low-traffic periods
    max_size=10,     # scale to this under load
    max_queries=50_000,             # recycle after 50k queries to prevent bloat
    max_inactive_connection_lifetime=300.0,  # close idle connections after 5 min
)

Test your pool under realistic concurrency before deploying:

# Simulate 20 concurrent LangGraph threads hitting Postgres
python -c "
import asyncio, asyncpg
from langgraph.checkpoint.postgres.aio import AsyncPostgresSaver

async def stress_test():
    pool = await asyncpg.create_pool('postgresql://...', min_size=2, max_size=10)
    checkpointer = AsyncPostgresSaver(pool)
    await checkpointer.setup()
    # ... compile graph and run 20 concurrent ainvoke calls
    await pool.close()

asyncio.run(stress_test())
"

6. FastAPI Integration Pattern

The recommended pattern for FastAPI: create the pool at app startup and close it on shutdown.

from contextlib import asynccontextmanager
from fastapi import FastAPI
import asyncpg
from langgraph.checkpoint.postgres.aio import AsyncPostgresSaver

checkpointer: AsyncPostgresSaver = None

@asynccontextmanager
async def lifespan(app: FastAPI):
    global checkpointer
    pool = await asyncpg.create_pool(
        dsn="postgresql://user:pass@localhost/langgraph",
        min_size=2,
        max_size=10,
    )
    checkpointer = AsyncPostgresSaver(pool)
    await checkpointer.setup()
    yield
    await pool.close()

app = FastAPI(lifespan=lifespan)

@app.post("/run")
async def run_agent(thread_id: str, message: str):
    config = {"configurable": {"thread_id": thread_id}}
    result = await graph.ainvoke(
        {"messages": [{"role": "user", "content": message}]},
        config=config,
    )
    return {"output": result["messages"][-1].content}

7. Checkpoint Retention and Cleanup

Checkpoints accumulate indefinitely if you don’t prune them. Add a background cleanup job:

import asyncpg
from datetime import datetime, timedelta, timezone

async def prune_old_checkpoints(pool: asyncpg.Pool, retention_days: int = 7):
    cutoff = datetime.now(timezone.utc) - timedelta(days=retention_days)
    async with pool.acquire() as conn:
        deleted = await conn.execute(
            """
            DELETE FROM checkpoints
            WHERE created_at < $1
            """,
            cutoff,
        )
    print(f"Pruned checkpoints older than {retention_days} days: {deleted}")

Run this as a scheduled task (e.g., daily via APScheduler or a cron job hitting a /admin/prune endpoint).


8. Failure Modes and Fixes

SymptomRoot CauseFix
asyncpg.exceptions.TooManyConnectionsErrorPool max_size too low for concurrent runsIncrease max_size or add a read replica
asyncpg.exceptions.InterfaceError: connection is closedIdle connection closed by Postgres idle_in_transaction_session_timeoutSet max_inactive_connection_lifetime < Postgres timeout
State missing on resumethread_id mismatch between runsAlways pass the same thread ID from durable storage (e.g., a database row)
KeyError on state resumeGraph state schema changed after checkpoint was writtenHandle missing keys with .get() defaults in node functions
setup() raises PermissionErrorPostgres user lacks CREATE TABLE rightsGrant CREATE on the target schema to the app user

Next Steps

For workflows that require human approval between nodes, the checkpointer is the foundation for interrupt queues. See Implementing Dynamic Human-in-the-Loop Approval Queues in LangGraph with FastAPI for the full pattern.

To benchmark LangGraph against PydanticAI under concurrent load with Postgres-backed state, see LangGraph vs PydanticAI: State Machine Architecture, Latency, and Memory Footprint.

If you’re building the agent state schema itself, the LangGraph Complete Guide covers TypedDict state design, edge conditions, and subgraph composition.

✉ Newsletter

Want to build production-ready AI?

Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.

S

Written by S L Manikanta

AI Engineer specializing in agentic workflows, multi-step LLM validation pipelines, and secure cloud environments. Sharing practical lessons from building software.

Related Articles

Agentic AI
AI Agent Architecture Patterns: A Guide for Platform Engineers (2026)

A technical comparison of AI agent architectures. Learn when to use Prompt Chaining, Routing, Orchestrator-Workers, and Cyclic State Graphs (LangGraph).

Agentic AI
What is an AI Agent Harness? Complete Technical Reference (2026)

A comprehensive technical reference on AI Agent Harnesses. Learn architecture, security, cost optimization, and how to deploy LangGraph agents into production with custom harnesses.

Agentic AI
Building StoxFlow: Hybrid Local/Cloud AI Agents Architecture Complete Guide (2026)

How to build a decoupled three-tier AI stock research agent using LangGraph, FastAPI, and hybrid LLM routing (Ollama/Gemini) for optimal cost and performance.