Browse by type

MCP-Bench is a comprehensive evaluation framework designed to assess Large Language Models' (LLMs) capabilities in tool-use scenarios through the Model Context Protocol (MCP). This benchmark provides an end-to-end pipeline for evaluating how effectively different LLMs can discover, select, and utilize tools to solve real-world tasks.
| Rank | Model | Overall Score |
|---|---|---|
| 1 | gpt-5 | 0.749 |
| 2 | o3 | 0.715 |
| 3 | gpt-oss-120b | 0.692 |
| 4 | gemini-2.5-pro | 0.690 |
| 5 | claude-sonnet-4 | 0.681 |
| 6 | qwen3-235b-a22b-2507 | 0.678 |
| 7 | glm-4.5 | 0.668 |
| 8 | gpt-oss-20b | 0.654 |
| 9 | kimi-k2 | 0.629 |
| 10 | qwen3-30b-a3b-instruct-2507 | 0.627 |
| 11 | gemini-2.5-flash-lite | 0.598 |
| 12 | gpt-4o | 0.595 |
| 13 | gemma-3-27b-it | 0.582 |
| 14 | llama-3-3-70b-instruct | 0.558 |
| 15 | gpt-4o-mini | 0.557 |
| 16 | mistral-small-2503 | 0.530 |
| 17 | llama-3-1-70b-instruct | 0.510 |
| 18 | nova-micro-v1 | 0.508 |
| 19 | llama-3-2-90b-vision-instruct | 0.495 |
| 20 | llama-3-1-8b-instruct | 0.428 |
Overall Score represents the average performance across all evaluation dimensions including rule-based schema understanding, LLM-judged (o4-mini as judge model) task completion, tool usage, and planning effectiveness. Scores are averaged across single-server and multi-server settings.
git clone https://github.com/accenture/mcp-bench.git
cd mcp-bench
conda create -n mcpbench python=3.10
conda activate mcpbench
cd mcp_servers
# Install MCP server dependencies
bash ./install.sh
cd ..
# Create .env file with API keys
# Default setup uses both OpenRouter and Azure OpenAI
# For Azure OpenAI, you also need to set your API version in file benchmark_config.yaml (line205)
# For OpenRouter-only setup, see "Optional: Using only OpenRouter API" section below
cat > .env << EOF
export OPENROUTER_API_KEY="your_openrouterkey_here"
export AZURE_OPENAI_API_KEY="your_azureopenai_apikey_here"
export AZURE_OPENAI_ENDPOINT="your_azureopenai_endpoint_here"
EOF
Some MCP servers require external API keys to function properly. These keys are automatically loaded from ./mcp_servers/api_key. You should set these keys by yourself in file ./mcp_servers/api_key:
# View configured API keys
cat ./mcp_servers/api_key
Required API keys include (These API keys are free and easy to get. You can get all of them within 10 mins):
- NPS_API_KEY: National Park Service API key (for nationalparks server) - Get API key
- NASA_API_KEY: NASA Open Data API key (for nasa-mcp server) - Get API key
- HF_TOKEN: Hugging Face token (for huggingface-mcp-server) - Get token
- GOOGLE_MAPS_API_KEY: Google Maps API key (for mcp-google-map server) - Get API key
- NCI_API_KEY: National Cancer Institute API key (for biomcp server) - Get API key This api key registration website might require US IP to open, see Issue #10 if you have difficulies for getting this api key.
# 1. Verify all MCP servers can be connected
##You should see "28/28 servers connected"
##and "All successfully connected servers returned tools!" after running this
python ./utils/collect_mcp_info.py
# 2. List available models
source .env
python run_benchmark.py --list-models
# 3. Run benchmark (gpt-oss-20b as an example)
##Must use o4-mini as judge model (hard-coded in line 429-436 in ./benchmark/runner.py) to reproduce the results.
## run all tasks
source .env
python run_benchmark.py --models gpt-oss-20b
## single server tasks
source .env
python run_benchmark.py --models gpt-oss-20b \
--tasks-file tasks/mcpbench_tasks_single_runner_format.json
## two server tasks
source .env
python run_benchmark.py --models gpt-oss-20b \
--tasks-file tasks/mcpbench_tasks_multi_2server_runner_format.json
## three server tasks
source .env
python run_benchmark.py --models gpt-oss-20b \
--tasks-file tasks/mcpbench_tasks_multi_3server_runner_format.json
To add new models from OpenRouter:
Copy the model ID (e.g., anthropic/claude-sonnet-4 or meta-llama/llama-3.3-70b-instruct)
Add the model configuration
llm/factory.py and add your model in the OpenRouter section (around line 152)Follow this pattern:
python
configs["your-model-name"] = ModelConfig(
name="your-model-name",
provider_type="openrouter",
api_key=os.getenv("OPENROUTER_API_KEY"),
base_url="https://openrouter.ai/api/v1",
model_name="provider/model-id" # The exact model ID from OpenRouter
)
Verify the model is available
bash
source .env
python run_benchmark.py --list-models
# Your new model should appear in the list
Run benchmark with your model
bash
source .env
python run_benchmark.py --models your-model-name
If you only want to use OpenRouter without Azure:
cat > .env << EOF
OPENROUTER_API_KEY=your_openrouterkey_here
EOF
Edit llm/factory.py and comment out the Azure section (lines 69-101), then add Azure models through OpenRouter instead:
# Comment out or remove the Azure section (lines 69-109)
# if os.getenv("AZURE_OPENAI_API_KEY") and os.getenv("AZURE_OPENAI_ENDPOINT"):
# configs["o4-mini"] = ModelConfig(...)
# ...
# Add Azure models through OpenRouter (in the OpenRouter section around line 106)
if os.getenv("OPENROUTER_API_KEY"):
# Add OpenAI models via OpenRouter
configs["gpt-4o"] = ModelConfig(
name="gpt-4o",
provider_type="openrouter",
api_key=os.getenv("OPENROUTER_API_KEY"),
base_url="https://openrouter.ai/api/v1",
model_name="openai/gpt-4o"
)
configs["gpt-4o-mini"] = ModelConfig(
name="gpt-4o-mini",
provider_type="openrouter",
api_key=os.getenv("OPENROUTER_API_KEY"),
base_url="https://openrouter.ai/api/v1",
model_name="openai/gpt-4o-mini"
)
configs["o3"] = ModelConfig(
name="o3",
provider_type="openrouter",
api_key=os.getenv("OPENROUTER_API_KEY"),
base_url="https://openrouter.ai/api/v1",
model_name="openai/o3"
)
configs["o4-mini"] = ModelConfig(
name="o4-mini",
provider_type="openrouter",
api_key=os.getenv("OPENROUTER_API_KEY"),
base_url="https://openrouter.ai/api/v1",
model_name="openai/o4-mini"
)
configs["gpt-5"] = ModelConfig(
name="gpt-5",
provider_type="openrouter",
api_key=os.getenv("OPENROUTER_API_KEY"),
base_url="https://openrouter.ai/api/v1",
model_name="openai/gpt-5"
)
# Keep existing OpenRouter models...
This way all models will be accessed through OpenRouter's unified API.
MCP-Bench includes 28 diverse MCP servers:
``` mcp-bench/ ├── agent/ # Task execution agents │ ├── init.py │ ├── executor.py # Multi-round task executor with retry logic │ └── execution_context.py # Execution context management ├── benchmark/ # Evaluation framework │ ├── init.py │ ├── evaluator.py # LLM-as-judge evaluation metrics │ ├── runner.py # Benchmark orchestrator │ ├── results_aggregator.py # Results aggregation and statistics │ └── results_formatter.py # Results formatting and display ├── config/ # Configuration management │ ├── init.py │ ├── benchmark_config.yaml # Benchmark configuration │ └── config_loader.py # Configuration loader ├── llm/ # LLM provider abstractions │ ├── init.py │ ├── factory.py # Model factory for multiple providers │ └── provider.py # Unified provider interface ├── mcp_modules/ # MCP server management │ ├── init.py │ ├── connector.py # Server connection handling │ ├── server_manager.py # Multi-server orchestration │ ├── server_manager_persistent.py # Persistent connection manager │ └── tool_cache.py # Tool call caching mechanism ├── synthesis/ # Task generation │ ├── init.py │ ├── task_synthesis.py # Task generation with fuzzy conversion │ ├── generate_benchmark_tasks.py # Batch task generation script │ ├── benchmark_generator.py # Unified benchmark task generator │ ├── README.md # Task synthesis documentation │ └── split_combinations/ # Server combination splits │ ├── mcp_2server_combinations.json │ └── mcp_3server_combinations.json ├── utils/ # Utilities │ ├── init.py │ ├── collect_mcp_info.py # Server discovery and tool collection │ ├── local_server_config.py # Local server configuration │ └── error_handler.py # Error handling utilities ├── tasks/ # Benchmark task files │ ├── mcpbench_tasks_single_runner_format.json │ ├── mcpbench_tasks_multi_2server_runner_format.json │ └── mcpbench_tasks_multi_3server_runner_format.json ├── mcp_servers/ # MCP server implementations (28 servers) │ ├── api_key # API keys configuration file │ ├── commands
browse all types & interfaces →
$ claude mcp add mcp-bench \
-- python -m otcore.mcp_server <graph>