visual-understanding
MCP server for multimodal understanding and object grounding (bounding boxes) across images, videos, and documents, with support for multiple AI providers (Zhipu GLM-V, OpenAI GPT-4o, Anthropic Claude, or any OpenAI-compatible endpoint).
- Score
- 22.5571 signal
- Evidence
- 1 star
- Last commit
- Not known
- Listed
Install
No one-command install. Set it up from its source.
Alternatives · MCPs
- disler-just-prompt738 stars · 132 forks97.439
- nesquikm-rubber-duck176 stars · 24 forks93.457
- minimax-multimodal128 stars · 41 forks93.269
What it is
MCP server for multimodal understanding and object grounding (bounding boxes) across images, videos, and documents, with support for multiple AI providers (Zhipu GLM-V, OpenAI GPT-4o, Anthropic Claude, or any OpenAI-compatible endpoint).
When to use it
MCP server for multimodal understanding and object grounding (bounding boxes) across images, videos, and documents, with support for multiple AI providers (Zhipu GLM-V, OpenAI GPT-4o, Anthropic Claude, or any OpenAI-compatible endpoint).
How to install / invoke
See Glama for the install config.
Notes
Listed from the Glama MCP registry.