Building Production MCP Servers: Python and TypeScript Guide
S L Manikanta
Aug 24, 2026 • 7 min read
list On this page expand_more
- How MCP Works: Transports and Primitives
- Choosing Your Transport: Stdio vs SSE
- Building an MCP Server in Python with FastMCP
- 1. Installation
- 2. Implementation: Database Tool Server
- Building an MCP Server in TypeScript
- 1. Installation
- 2. Implementation: TypeScript Server
- Connecting Your MCP Server to Claude Desktop and Cursor
- For Claude Desktop (claude_desktop_config.json):
- For Cursor IDE:
- Production Best Practices and Security
- 1. Never Trust LLM Generated SQL Directly
- 2. Enforce Strict Output Limits
- 3. Add Request Timeouts
Want to build production-ready AI?
Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.
When building AI tools and agents, writing custom glue code for every model provider quickly turns into a mess. Anthropic introduced the Model Context Protocol (MCP) to fix this by giving AI clients a standard way to discover tools, fetch data, and read system prompts.
Cursor, Claude Desktop, Claude Code, and custom agent frameworks now speak MCP natively.
If you want your internal Postgres database, Redis cache, or REST APIs accessible to any AI coding tool without rewriting custom integrations, building an MCP server is the cleanest approach.
Here is how MCP works under the hood, how to build production servers in both Python and TypeScript, and how to handle security, auth, and deployment.
flowchart LR
subgraph Clients [MCP Clients]
Claude[Claude Code / Desktop]
Cursor[Cursor IDE]
CustomAgent[Custom LangGraph / PydanticAI Agent]
end
subgraph Transport [Transport Layer]
Stdio[Standard I/O: Local Process]
SSE[Server Sent Events: Remote HTTP]
end
subgraph MCPServer [Production MCP Server]
Router[MCP Protocol Router]
Tools[Exposed Tools]
Resources[Read Only Resources]
Prompts[Prompt Templates]
end
subgraph Internal [Internal Systems]
DB[(PostgreSQL Database)]
API[Internal REST / GraphQL APIs]
end
Clients --> Transport
Transport --> Router
Router --> Tools
Router --> Resources
Router --> Prompts
Tools --> DB
Tools --> API
Resources --> DB
How MCP Works: Transports and Primitives
MCP uses JSON-RPC 2.0 messages. It defines three core capabilities that your server can expose:
- Tools: Functions that the LLM can call with arguments to perform actions or mutate data (e.g. running a SQL query or creating a ticket).
- Resources: Read only data streams that the client can attach to context, like file contents, database schemas, or logs.
- Prompts: Pre-built prompt templates and workflows that users can trigger inside their AI interface.
Choosing Your Transport: Stdio vs SSE
- Standard I/O (stdio): The client (e.g. Cursor or Claude Desktop) launches your server as a local child process and talks to it over standard input and standard output. This is best for local developer tools, command line utilities, and private scripts running on your machine.
- Server Sent Events (SSE): The server runs as an HTTP service. The client opens an SSE connection to receive streaming messages from the server and sends HTTP POST requests to send commands. This is what you need for shared team servers, cloud deployments, and Docker containers.
Building an MCP Server in Python with FastMCP
The official mcp Python SDK includes a high level helper called FastMCP (inspired by FastAPI) that handles schema generation from type hints automatically.
1. Installation
pip install "mcp[cli]" asyncpg pydantic
2. Implementation: Database Tool Server
Here is a complete, production ready MCP server that connects to PostgreSQL and exposes a safe query tool and a schema resource:
import os
from typing import Any, Dict, List
import asyncpg
from mcp.server.fastmcp import FastMCP
from pydantic import BaseModel, Field
# Initialize FastMCP server
mcp = FastMCP("postgres-analytics-server")
DATABASE_URL = os.getenv("DATABASE_URL", "postgresql://user:password@localhost:5432/analytics")
pool: asyncpg.Pool | None = None
@mcp.resource("schema://analytics")
async def get_database_schema() -> str:
"""Returns the DDL table schema for the analytics database."""
async with pool.acquire() as conn:
rows = await conn.fetch("""
SELECT table_name, column_name, data_type
FROM information_schema.columns
WHERE table_schema = 'public'
ORDER BY table_name, ordinal_position;
""")
schema_text = "Database Schema:\n"
current_table = ""
for row in rows:
if row["table_name"] != current_table:
current_table = row["table_name"]
schema_text += f"\nTable: {current_table}\n"
schema_text += f" - {row['column_name']}: {row['data_type']}\n"
return schema_text
class QueryInput(BaseModel):
query: str = Field(description="Read-only SELECT SQL query to execute")
max_rows: int = Field(default=50, ge=1, le=500, description="Maximum number of rows to return")
@mcp.tool()
async def run_analytics_query(params: QueryInput) -> List[Dict[str, Any]]:
"""Execute a read-only SQL query against the analytics database."""
# Basic safety check to reject write operations
clean_query = params.query.strip().lower()
if not clean_query.startswith("select") and not clean_query.startswith("with"):
raise ValueError("Only SELECT and WITH queries are permitted on this endpoint.")
async with pool.acquire() as conn:
# Enforce query timeout to prevent runaway table scans
async with conn.transaction(readonly=True):
await conn.execute("SET LOCAL statement_timeout = '5000ms'")
records = await conn.fetch(params.query)
# Slice results to prevent context window explosion
limited_records = records[:params.max_rows]
return [dict(record) for record in limited_records]
if __name__ == "__main__":
import asyncio
async def main():
global pool
pool = await asyncpg.create_pool(DATABASE_URL, min_size=2, max_size=10)
try:
# Runs using standard I/O for local tools like Claude or Cursor
await mcp.run_stdio_async()
finally:
await pool.close()
asyncio.run(main())
Building an MCP Server in TypeScript
If your team runs Node.js or Bun in production, the official TypeScript SDK (@modelcontextprotocol/sdk) provides full type safety.
1. Installation
npm install @modelcontextprotocol/sdk zod
npm install -D typescript @types/node
2. Implementation: TypeScript Server
import { Server } from "@modelcontextprotocol/sdk/server/index.js";
import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js";
import {
CallToolRequestSchema,
ListToolsRequestSchema,
} from "@modelcontextprotocol/sdk/types.js";
import { z } from "zod";
const server = new Server(
{
name: "github-issue-manager",
version: "1.0.0",
},
{
capabilities: {
tools: {},
},
}
);
// Define tool argument schema
const CreateIssueSchema = z.object({
repo: z.string().describe("Repository name in owner/repo format"),
title: z.string().min(3).describe("Issue title"),
body: z.string().describe("Detailed description of the issue"),
labels: z.array(z.string()).optional().describe("Issue labels"),
});
// List available tools
server.setRequestHandler(ListToolsRequestSchema, async () => {
return {
tools: [
{
name: "create_github_issue",
description: "Creates a new issue in a GitHub repository",
inputSchema: {
type: "object",
properties: {
repo: { type: "string", description: "Repository in owner/repo format" },
title: { type: "string", description: "Issue title" },
body: { type: "string", description: "Issue description" },
labels: { type: "array", items: { type: "string" }, description: "Labels" },
},
required: ["repo", "title", "body"],
},
},
],
};
});
// Handle tool execution
server.setRequestHandler(CallToolRequestSchema, async (request) => {
if (request.params.name === "create_github_issue") {
const args = CreateIssueSchema.parse(request.params.arguments);
// Simulate GitHub API call
const issueUrl = `https://github.com/${args.repo}/issues/101`;
return {
content: [
{
type: "text",
text: `Issue created successfully: ${issueUrl}\nTitle: ${args.title}`,
},
],
};
}
throw new Error(`Tool not found: ${request.params.name}`);
});
async function run() {
const transport = new StdioServerTransport();
await server.connect(transport);
}
run().catch((err) => {
console.error("Fatal error running MCP server:", err);
process.exit(1);
});
Connecting Your MCP Server to Claude Desktop and Cursor
To use your local server inside Claude Desktop or Cursor, add it to your configuration file.
For Claude Desktop (claude_desktop_config.json):
On macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
On Windows: %APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"postgres-analytics": {
"command": "python",
"args": ["-m", "src.mcp_server"],
"env": {
"DATABASE_URL": "postgresql://user:secret@localhost:5432/analytics"
}
},
"github-tools": {
"command": "node",
"args": ["/absolute/path/to/dist/index.js"],
"env": {
"GITHUB_TOKEN": "ghp_xxxxxxxxxxxx"
}
}
}
}
For Cursor IDE:
In Cursor, go to Settings > Features > MCP Servers > Add New MCP Server, select command, and enter the execution command.
Production Best Practices and Security
When taking MCP servers from local development to production setups, keep these rules in mind:
1. Never Trust LLM Generated SQL Directly
Always enforce read only transactions (SET LOCAL default_transaction_read_only = on) and use dedicated database users with GRANT SELECT only. Never give an MCP tool raw DROP or DELETE permissions without human confirmation.
2. Enforce Strict Output Limits
LLMs have finite context windows. If an SQL query returns 10,000 rows, serializing that into JSON will exhaust the model’s memory and cost significant token fees. Always hard cap query results at 50 to 100 rows.
3. Add Request Timeouts
Database queries and third party APIs can hang. Set explicit statement timeouts (e.g. 5 seconds) on every network call so your MCP server never blocks the agent indefinitely.
Want to build production-ready AI?
Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.
Written by S L Manikanta
AI Engineer specializing in agentic workflows, multi-step LLM validation pipelines, and secure cloud environments. Sharing practical lessons from building software.
Related Articles
The Shift to Agentic AI Workflows in Production
Why engineering teams are moving away from simple copilots to autonomous agentic workflows, and the technical challenges of long-running state management.
AI Agent Memory: Short-Term vs Long-Term Memory
A complete architectural breakdown of how AI agents manage state, covering short-term conversational context and long-term persistent memory systems.
AI Agent Observability: Logs, Traces, and Metrics in Production
A complete technical reference and implementation guide to observing agentic workflows, tracking LLM token costs, logging reasoning trajectories, tracing nested tool calls, and monitoring system metrics in production.