Atlas / MCP servers / houtini-ai / Houtini LM

Houtini LMSAFE

mcp/houtini-ai/houtini-lm-1

MCP server that saves Claude Code tokens by delegating bounded tasks to local or cloud LLMs. Works with LM Studio, Ollama, vLLM, DeepSeek, Groq, Cerebras.

Verdict
SAFE
Grade
B
Trust score
89 /100
Exposed tools
8 6r · 2w · 0d
Transport
stdio
License
Apache-2.0
Stars
120
01

Overview

From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.

[](https://www.npmjs.com/package/@houtini/lm) [](https://registry.modelcontextprotocol.io) [](https://opensource.org/licenses/Apache-2.0) [](https://snyk.io/test/github/houtini-ai/houtini-lm)

Houtini LM is an MCP server that lets Claude (or any MCP client) hand bounded work to another model - a local LLM on your GPU, OpenAI's latest GPT models, a LiteLLM router, OpenRouter or a cheap cloud API - while you carry on working in the AI platform you already like. It cuts your token bill, and it gives you a second model to review your code whenever you want one.

Quick Navigation What's new | Why use it | Install | How it handles different models | What to hand over | Tools | Reading the footer | Configuration | Endpoints | The manual

I built this because I kept leaving Claude Code running overnight on big refactors and the token bill was painful. A huge chunk of that spend went on bounded tasks any decent model handles fine - generat

Read from source at commit b3fa9959e480OBSERVED · 2026-10-07
02

Connect

Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control. Replace the environment placeholders with a token scoped to the least it needs.

claude-code (npm)
claude mcp add lm --env HOUTINI_LM_API_KEY=${HOUTINI_LM_API_KEY} -- npx -y @houtini/[email protected]
03

Exposed tools (8)

6 read · 2 write · 0 destructive.

ToolRiskDescription
chatwriteSend a task to a local LLM - a sidekick running on the user\
code_taskwriteSend a code-specific task to the local LLM, wrapped with an optimised code-review system prompt. Temperature is locked low (0.2 or the routed model\
code_task_filesreadLike code_task, but the local LLM reads files directly from disk - source never passes through the MCP client\
custom_promptreadStructured analysis via the local LLM with explicit system/context/instruction separation.
discoverreadCheck whether the local LLM is online and what model is loaded. Returns model name, context window size,
embedreadGenerate text embeddings via the local LLM server. Requires an embedding model to be loaded
list_modelsreadList all models on the local LLM server - both loaded (ready) and available (downloaded but not active).
statsreadShow user stats: tokens offloaded, calls made, per-model performance - for the current session AND
04

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codeWARN
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
declared (5 observation(s))
Shell
none-observed
Dependencies
not all pinned
Secrets in source
found

Findings (10)

MEDIUMHard-coded secrets · secret.generic · CWE-798, CWE-321
scripts/verify-inference-lock.mjs:91
at: Date.now(), token: 'held-by-someone-else',
LOWNetwork egress · net.raw_ip · CWE-200, CWE-319
README.md:88
-e HOUTINI_LM_ENDPOINT_URL=http://192.168.1.50:1234 \
LOWNetwork egress · net.raw_ip · CWE-200, CWE-319
docs/GETTING-STARTED.md:40
-e HOUTINI_LM_ENDPOINT_URL=http://192.168.1.50:1234 \
LOWNetwork egress · net.raw_ip · CWE-200, CWE-319
docs/SETUP-LMSTUDIO.md:33
the host's LAN IP (e.g. `http://192.168.1.50:1234`).
LOWNetwork egress · net.raw_ip · CWE-200, CWE-319
docs/SETUP-LMSTUDIO.md:61
- On a different machine: use its LAN URL (`http://192.168.1.50:1234`).
LOWNetwork egress · net.raw_ip · CWE-200, CWE-319
docs/SETUP-OLLAMA.md:87
**Remote hosts need `OLLAMA_HOST`.** Ollama binds to localhost by default. To reach it from another machine, start it with `OLLAMA_HOST=0.0.0.0` and point `HOUTINI_LM_ENDPOINT_URL` at the box's LAN ad
LOWSupply chain · supply.unpinned · CWE-829, CWE-1357
package.json
@modelcontextprotocol/sdk, @types/node, typescript
Why it matters. 3 dependency range(s) float
Fix. pin exact versions or ship a lockfile
LOWPrompt injection · prompt.transfer_instruction · CWE-94, CWE-1427
README.md:219
When something went wrong, a quality line says so: `TRUNCATED` for a partial result (a stalled connection gives you what arrived rather than a timeout error), `hit-max-tokens` when the budget ran out,
Why it matters. an instruction to move sensitive data to an outside destination
Fix. remove; a skill never needs the user's secrets off the machine
LOWPrompt injection · prompt.transfer_instruction · CWE-94, CWE-1427
manual/models.md:25
**`on`** forces thinking on. It's worth it for work where the model's own reasoning improves the answer, like hunting a subtle bug, checking an argument or planning something with several moving parts
Why it matters. an instruction to move sensitive data to an outside destination
Fix. remove; a skill never needs the user's secrets off the machine
LOWPrompt injection · prompt.transfer_instruction · CWE-94, CWE-1427
manual/models.md:68
A plain OpenAI endpoint also doesn't report context windows or output caps in its model list, so houtini-lm starts from its 100,000-token fallback. When a budget overshoots a model's real output cap (
Why it matters. an instruction to move sensitive data to an outside destination
Fix. remove; a skill never needs the user's secrets off the machine

Gates applied: no_behavioural_pass.

Audited 2026-10-07 · audit v0.4.1 · source sha b3fa9959e480full audit observations/trust-audit/mcp-server/houtini-ai__houtini-lm-1.json · Report an issue / request a re-scan
05

Audit history

Every audit this server has had. A grade with a past is a grade somebody is still checking.

DateSourceVerdictGradeScoreChange
2026-10-07b3fa9959e480SAFEB89first audit
06

Questions

What is the Houtini LM MCP server?

MCP server that saves Claude Code tokens by delegating bounded tasks to local or cloud LLMs. Works with LM Studio, Ollama, vLLM, DeepSeek, Groq, Cerebras.

What tools does Houtini LM expose?

8 in total: 6 read-only, 2 that write, and 0 that can delete or overwrite. Every one is listed on this page with its risk.

Is Houtini LM safe to connect to an agent?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean server reads B.

What credentials does Houtini LM need?

It reads HF_TOKEN, HOUTINI_LM_API_KEY, HOUTINI_LM_MIN_TOKENS, LM_PASSWORD, LM_STUDIO_PASSWORD and OPENROUTER_API_KEY from the environment. Give it a token scoped to the least it needs — an agent that can be talked into calling a tool can be talked into calling it with your credentials.

How does Houtini LM run?

It speaks stdio, so it runs as a local process your client starts. It is published on npm as @houtini/lm at 3.3.4.

How current is this page?

The grade is for one exact copy of the source (b3fa9959e480), read on 2026-10-07. The repository is watched and re-audited when it changes.

Advertisement