inferwatch-mcp
Enables agents to query real-time and historical metrics for locally served Ollama and vLLM instances, including request rates, latency, token counts, and GPU utilization, over stdio.
- Score
- Unranked
- Evidence
- No signals yet
- Last commit
- as last read from GitHub; most reads are from 2 Sep 2026 or later
- Listed
Install
No one-command install. Set it up from its source.
Alternatives · MCPs
- berlinbra-alpha-vantage104 stars · 39 forks92.512
- gpu-mcp-server15 stars · 13 forks81.185
- anirbanbasu-frankfurter6 stars · 10 forks75.326
What it is
Enables agents to query real-time and historical metrics for locally served Ollama and vLLM instances, including request rates, latency, token counts, and GPU utilization, over stdio.
When to use it
Enables agents to query real-time and historical metrics for locally served Ollama and vLLM instances, including request rates, latency, token counts, and GPU utilization, over stdio.
How to install / invoke
See Glama for the install config.
Notes
Listed from the Glama MCP registry.