REAL MCP CALLS · SAMPLE DATA
Route an MCP tool.
Give Jev a request. Inspect its tool selection, then execute it through MCP.
Optionally compare the same request with an LLM.
SINGLE EXECUTION
Your choice, one tool call.
Route your request, then execute a valid selection below.
Inspect shared input & provider payloads
The catalog and request are identical; provider-specific wrappers differ. Expected labels are never sent to either model.
Run a comparison to inspect its inputs.
Recent runs
Saved locallyNo comparisons yet.
Compare models · 20-request benchmark
A shared test set. A fair starting point.
20 labeled requests · 4 per tool · 40 provider calls · zero tool executions
This is a small, fixed-label smoke benchmark with enum arguments—not a general MCP capability evaluation. API calls may incur charges.
No benchmark run yet.
| # | User request / expected tool | LLM | Jev | Agreement |
|---|
Available MCP tools
Discovered through MCP · fixed sample resultsArguments: service = auth / billing / search · environment = production / staging. Production is the shared default. Unsupported requests can return no_tool.