Armory
Source
Browse
MCPs

inferbench

InferBench's MCP server lets coding agents run, serve and benchmark local LLMs (text + image, llama.cpp + Stable Diffusion) on your own hardware on demand — measuring real tokens/sec and picking the optimal quant for your GPU from a 124-model catalog. Local-first, no cloud required.

Score
32.2231 signal
Evidence
2 stars
Last commit
as last read from GitHub; most reads are from 2 Sep 2026 or later
Listed

Install

No one-command install. Set it up from its source.

Alternatives · MCPs

  1. lalanikarim-comfy46 stars · 16 forks86.781
  2. ichigo3766-stable-diffusion43 stars · 13 forks85.815
  3. chentyjpm-stable-diffusion-ncnn18 stars · 4 forks74.510

What it is

InferBench's MCP server lets coding agents run, serve and benchmark local LLMs (text + image, llama.cpp + Stable Diffusion) on your own hardware on demand — measuring real tokens/sec and picking the optimal quant for your GPU from a 124-model catalog. Local-first, no cloud required.

When to use it

InferBench's MCP server lets coding agents run, serve and benchmark local LLMs (text + image, llama.cpp + Stable Diffusion) on your own hardware on demand — measuring real tokens/sec and picking the optimal quant for your GPU from a 124-model catalog. Local-first, no cloud required.

How to install / invoke

See Glama for the install config.

Notes

Listed from the Glama MCP registry.