A production-style multi-agent orchestration system built with CrewAI and the Model Context Protocol (MCP). Demonstrates role-based agent collaboration, LLM-as-judge evaluation, and enterprise-grade tool integration patterns.
v2.0 is a complete architectural rebuild focused on production patterns over demos:
- Multi-Agent Architecture: Role-based CrewAI agents (Task Planner, Action Executor) with strict JSON schemas
- LLM-as-Judge Evaluation: Every execution auto-scored on success, plan quality, and reasoning
- Enterprise Tool Integration: GitHub API (issues, PRs, repos), Tavily search (4 tools), custom Weather MCP server
- Production Observability: Persistent metrics tracking, performance visualization, goal-type inference
- Robust Error Handling: Parameter validation, connection retries, graceful fallbacks
Results from ~60+ test runs:
- β4/5 average plan quality across all dimensions
- ~15-20s end-to-end execution for multi-step workflows
- 90%+ tool routing accuracy with explicit failure reporting
- 100% success rate on weather/search queries
User Query β Task Planner β Tool Execution β Action Executor β LLM Judge β Final Answer
β β β
(Pydantic (Async parallel (Score + persist
TaskPlan) execution) metrics)
Three specialized agents:
- Task Planner: Creates validated execution plans with strict tool constraints
- Action Executor: Synthesizes tool results into user-facing answers
- Research Coordinator: Optional deep-dive research (disabled by default for speed)
- GitHub MCP (
@modelcontextprotocol/server-github): 20+ tools for repo management, issues, PRs - Tavily MCP (
mcp-remotebridge): Reliable web search with 4 specialized tools (search, extract, crawl, map) - Weather MCP (custom FastAPI): Built from scratch with SSE transport for real-time data
Centralized, generic tool router with:
- Smart parameter extraction (regex-based parsing for owner/repo, city names, file paths)
- Async parallel execution for multi-step workflows
- Type-safe tool result passing between agents
- Graceful error handling with explicit failure messages
- LLM-as-Judge (
eval/judge.py): Auto-scores every run on 3 dimensions (0-5 scale) - Metrics Tracking (
orchestrai/metrics.py): Persistent JSON logs with goal type inference - Performance Visualization (
view_metrics.py): Aggregates, trends, success rates
- Python 3.11+
- Node.js 18+ (for MCP servers)
- OpenAI API key
- GitHub Personal Access Token (for GitHub MCP)
- Tavily API key (for search)
# Clone and install
git clone https://github.com/deepmehta27/MCP_Navigator.git
cd MCP_Navigator
pip install -r requirements.txt
# Install MCP servers
npm install -g @modelcontextprotocol/server-github
npm install -g mcp-remote
# Configure environment
cp .env.example .env
# Add: OPENAI_API_KEY, GITHUB_TOKEN, TAVILY_API_KEYEdit servers/browser_mcp.json for MCP server connections:
{
"mcpServers": {
"github": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": {
"GITHUB_PERSONAL_ACCESS_TOKEN": "${GITHUB_TOKEN}"
}
},
"tavily": {
"command": "npx",
"args": ["-y", "mcp-remote", "tavily"],
"env": {
"TAVILY_API_KEY": "${TAVILY_API_KEY}"
}
}
}
}python orchestrai/cli.pyYou: Search for the latest multi-agent AI frameworks in 2026
β Plan: 1-step (tavily_search)
β Result: CrewAI, LangGraph, AutoGen, LlamaIndex...
β Judge: Success=4/5, Plan=4/5, Reasoning=4/5
You: What's the weather in San Francisco?
β Plan: 1-step (get_weather)
β Result: 15.9Β°C, wind 9.1 km/h
β Judge: Success=5/5, Plan=5/5, Reasoning=5/5
You: List issues for deepmehta27/mcp-navigator-test
β Plan: 1-step (list_issues)
β Result: [Issue #1..... ]
β Judge: Success=5/5
You: Create an issue titled "Add streaming support"
β Plan: 1-step (create_issue)
β Result: Issue #10 created
β Judge: Success=5/5
You: Search for trending AI repos and create GitHub issue summary
β Plan: 2 steps (tavily_search β create_issue)
β Tool 1: Finds 12 trending repos
β Tool 2: Creates issue with formatted summary
β Judge: Success=5/5, Plan=4/5
python view_metrics.py
# Output:
# Total Runs: 60
# Success Rate: 93.3%
# Avg Success Score: 4.5/5
# Avg Plan Score: 4.0/5
# Performance Trend: β IMPROVING- Role-based separation: Cleaner agent boundaries vs. monolithic ReAct loop
- Built-in schema validation: Pydantic TaskPlan enforcement at agent boundaries
- Easier debugging: Sequential agent execution = linear trace inspection
- Reliability: DuckDuckGo MCP had anti-bot issues causing 40%+ failures
- Quality: Tavily returns structured results with relevance scores
- Speed: Sub-second response times vs. 3-5s for DuckDuckGo
- Existing weather tools lacked proper MCP transport implementation
- Built with FastAPI + SSE for streaming real-time data
- Full control over error handling and retry logic
- Automated quality tracking: No manual evaluation needed across 60+ runs
- Regression detection: Catches plan quality degradation early
- Systematic debugging: Judge notes pinpoint exact failure reasons
MCP_Navigator/
βββ orchestrai/
β βββ cli.py # Main CLI application
β βββ workflow.py # Multi-agent orchestration logic
β βββ agents.py # CrewAI agent definitions
β βββ schemas.py # Pydantic models (TaskPlan, ExecutionResult)
β βββ mcp_tools.py # MCP server connection management
β βββ tool_runner.py # Generic tool execution engine
β βββ metrics.py # Metrics tracking and persistence
βββ eval/
β βββ judge.py # LLM-as-judge evaluation
βββ servers/
β βββ weather.py # Custom Weather MCP server
β βββ browser_mcp.json # MCP server configuration
βββ data/
β βββ metrics.json # Persistent execution metrics
βββ view_metrics.py # Metrics visualization CLI
# Manual test suite (no pytest yet - realistic for v2.0)
python orchestrai/cli.py
> Search for AI frameworks
> Weather in NYC
> Create issue in deepmehta27/mcp-navigator-test titled "Test"
> metrics- Add server to
servers/browser_mcp.json - Update
TOOL SELECTION RULESinworkflow.pyplanner prompt - Add parameter extraction logic to
execute_plan_tools()if needed - Test tool routing with sample queries
- Notes server removed: Inconsistent parameter contracts caused failures (pragmatic cut)
- Research agent disabled by default: Adds 10-15s latency with minimal quality gain
- No streaming UI: CLI shows final results only (batch mode)
- Single-user only: No auth, rate limiting, or multi-tenancy tes deployment guide for production scale
MIT License - see LICENSE for details.
Contributions welcome! Focus areas:
- New MCP integrations: Slack, Calendar, Database tools
- Evaluation improvements: Add precision/recall metrics for search quality
- Performance optimization: Reduce cold-start latency (<10s target)
- Documentation: Production deployment guides
Made with β€οΈ by Deep Mehta