Skip to main content

Overview

This page documents real-world performance benchmarks for MCP Server with LangGraph and competitor frameworks. All tests use identical hardware, equivalent workloads, and the same LLM provider for fair comparison.
Benchmark Disclaimer: Performance varies based on hardware, configuration, LLM provider, workflow complexity, and network conditions. These benchmarks provide relative comparisons, not absolute guarantees. Your results may differ.

Benchmark Methodology

Test Environment

All benchmarks run on standardized hardware:

Test Configuration

Common Parameters:
  • LLM Provider: Gemini 2.0 Flash (for cost-effective testing)
  • Load Testing Tool: k6 (k6.io)
  • Test Duration: 5 minutes per scenario (after 1-minute ramp-up)
  • Runs: Average of 3 test runs
  • Monitoring: Prometheus + Grafana for metrics collection
Scenarios Tested:
  1. Simple Agent: Single-node workflow, basic tool execution
  2. Multi-Agent: 3-agent sequential coordination
  3. Complex Workflow: 5-node graph with conditional branching
  4. High Concurrency: 100+ concurrent requests

Benchmark Results

Simple Agent Workflow

Scenario: Single-node agent answering factual questions using Gemini Flash.

MCP Server with LangGraph (Self-Hosted on Cloud Run)

LangGraph Cloud (Managed Platform)

Winner: MCP Server (self-hosted) - 5% higher throughput, lower latency, zero platform fees

Multi-Agent Coordination

Scenario: 3-agent workflow (researcher → analyzer → writer) with sequential execution.

MCP Server with LangGraph (Kubernetes on GKE)

CrewAI (Self-Hosted)

Winner: CrewAI (marginally faster) for simple sequential workflows. MCP Server wins for production features (observability, scaling, persistence).
Why CrewAI is Faster: CrewAI’s role-based delegation model has lower orchestration overhead for simple sequential tasks. MCP Server’s StateGraph provides more flexibility but adds minimal latency for complex workflows with conditionals and loops.

Complex Workflow with Conditionals

Scenario: 5-node graph with conditional branching, error handling, and state persistence.

MCP Server with LangGraph

OpenAI AgentKit (Platform)

Winner: MCP Server (only framework suitable for complex programmatic workflows with state persistence)

High Concurrency Load Test

Scenario: 100 concurrent virtual users sending requests continuously for 10 minutes.

MCP Server with LangGraph (Kubernetes with HPA)

Google ADK (Vertex AI Agent Engine)

Winner: MCP Server (12% higher throughput, lower latency, cost-effective at scale)

Cost-Performance Analysis

Cost per 1M Requests (Complex Workflow)

Best Value: MCP Server on Kubernetes (16x cheaper than managed platforms)

Scaling Characteristics

Horizontal Scaling Efficiency

Efficiency: MCP Server scales near-linearly up to 8 pods with minimal coordination overhead.

Vertical Scaling

Recommendation: Horizontal scaling (more pods) is more cost-effective than vertical scaling (bigger pods).

Real-World Performance Expectations

Production Deployment Estimates

Scenario 1: Healthcare Startup (HIPAA-compliant)
  • Load: 50K requests/month
  • Configuration: MCP Server on GKE (2 pods, n2-standard-2)
  • Performance: p95 latency under 1s, 99.95% uptime
  • Cost: ~$150/month (infrastructure + LLM)
Scenario 2: Enterprise Customer Support (High Volume)
  • Load: 10M requests/month
  • Configuration: MCP Server on GKE (12 pods, n2-standard-4, multi-region)
  • Performance: p95 latency under 2s, 99.99% uptime
  • Cost: ~$1,200/month (infrastructure + LLM)
Scenario 3: Financial Services (Compliance-Critical)
  • Load: 500K requests/month
  • Configuration: MCP Server on GKE (4 pods, n2-highmem-4, private cluster)
  • Performance: p95 latency under 1.5s, 99.99% uptime
  • Cost: ~$500/month (infrastructure + LLM)

Benchmark Limitations

What These Benchmarks Don’t Measure

  • Cold Start Performance: All tests measured warm instances (excluding cold starts)
  • Network Latency: Tests run in same region as LLM provider (minimal network overhead)
  • Complex Tool Execution: Simple tools used (web search, database queries not benchmarked)
  • Long-Running Workflows: Tests limited to under 10 second workflows
  • Memory-Intensive Workloads: Benchmarks focus on CPU/network, not memory-bound operations

Factors That Impact Your Performance

  • LLM Provider Response Time: Gemini/Claude/GPT have different latencies
  • Tool Execution Time: Complex tool calls (database queries, API calls) add latency
  • Network Geography: Distance to LLM provider affects response time
  • Workflow Complexity: More nodes, conditionals, and loops increase latency
  • State Persistence: Checkpointing to Redis/PostgreSQL adds overhead

Running Your Own Benchmarks

Want to validate these results in your environment?
1

Clone Repository

2

Install k6

3

Configure Test

Edit scenarios/simple-agent.js with your endpoint:
4

Run Benchmark

See tests/benchmarks/README.md for detailed instructions.

Frequently Asked Questions

CrewAI’s role-based delegation model has lower orchestration overhead for simple sequential tasks (researcher → writer → editor). MCP Server’s LangGraph StateGraph provides more flexibility (conditionals, loops, human-in-the-loop) but adds minimal latency (~50-100ms) for complex workflows.Use CrewAI if: Simple sequential workflows, prototyping, learning Use MCP Server if: Production deployments, complex workflows, enterprise features needed
These benchmarks provide relative comparisons with controlled variables (same hardware, LLM provider, duration). Absolute numbers will vary in your environment based on:
  • Network latency to LLM provider
  • Tool execution time (database queries, API calls)
  • Workflow complexity (our tests use simple workflows)
  • Instance type and configuration
Run your own benchmarks with realistic workflows for accurate predictions.
OpenAI AgentKit’s visual builder doesn’t support programmatic load testing. Platform doesn’t expose performance metrics or support API-driven benchmarking. Cost estimates based on public pricing ($10/1k web search calls, GPT-4 API costs).
Reasons:
  • Cost-effective: 20x cheaper than GPT-4 (0.075vs0.075 vs 10-30 per 1M tokens)
  • Fast: Low latency for responsive testing
  • Fair comparison: Available across all frameworks (via LiteLLM)
  • Representative: Realistic for production use cases balancing cost/performance
Production deployments often use cheaper models (Gemini Flash, Claude Haiku) for high-volume workloads and premium models (GPT-4, Claude Opus) for complex reasoning.
Yes, benchmarks include end-to-end latency:
  • Network round-trip to LLM provider
  • LLM inference time
  • Framework orchestration overhead
  • State persistence (if applicable)
This represents real-world user experience, not just framework overhead.

Contributing Benchmarks

Have benchmark results to share? We welcome contributions:
  1. Fork the repository
  2. Run benchmarks using our methodology (same hardware, config)
  3. Document all parameters (instance type, LLM provider, workflow)
  4. Submit PR with results to tests/benchmarks/RESULTS.md
See CONTRIBUTING.md for guidelines.

Run Benchmarks Yourself

Clone repo and validate results

Compare Frameworks

Full framework comparison guide

Production Deployment

Deploy to production

Multi-LLM Setup

Optimize costs with multiple providers

Benchmark Transparency: All benchmark scripts, configurations, and raw results are open-source in our GitHub repository. We encourage independent validation and welcome corrections.