Houtini LMSAFE
MCP server that saves Claude Code tokens by delegating bounded tasks to local or cloud LLMs. Works with LM Studio, Ollama, vLLM, DeepSeek, Groq, Cerebras.
Overview
From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.
[](https://www.npmjs.com/package/@houtini/lm) [](https://registry.modelcontextprotocol.io) [](https://opensource.org/licenses/Apache-2.0) [](https://snyk.io/test/github/houtini-ai/houtini-lm)
Houtini LM is an MCP server that lets Claude (or any MCP client) hand bounded work to another model - a local LLM on your GPU, OpenAI's latest GPT models, a LiteLLM router, OpenRouter or a cheap cloud API - while you carry on working in the AI platform you already like. It cuts your token bill, and it gives you a second model to review your code whenever you want one.
Quick Navigation What's new | Why use it | Install | How it handles different models | What to hand over | Tools | Reading the footer | Configuration | Endpoints | The manual
I built this because I kept leaving Claude Code running overnight on big refactors and the token bill was painful. A huge chunk of that spend went on bounded tasks any decent model handles fine - generat
b3fa9959e480OBSERVED · 2026-10-07Connect
Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control. Replace the environment placeholders with a token scoped to the least it needs.
claude mcp add lm --env HOUTINI_LM_API_KEY=${HOUTINI_LM_API_KEY} -- npx -y @houtini/[email protected]Exposed tools (8)
6 read · 2 write · 0 destructive.
| Tool | Risk | Description |
|---|---|---|
chat | write | Send a task to a local LLM - a sidekick running on the user\ |
code_task | write | Send a code-specific task to the local LLM, wrapped with an optimised code-review system prompt. Temperature is locked low (0.2 or the routed model\ |
code_task_files | read | Like code_task, but the local LLM reads files directly from disk - source never passes through the MCP client\ |
custom_prompt | read | Structured analysis via the local LLM with explicit system/context/instruction separation. |
discover | read | Check whether the local LLM is online and what model is loaded. Returns model name, context window size, |
embed | read | Generate text embeddings via the local LLM server. Requires an embedding model to be loaded |
list_models | read | List all models on the local LLM server - both loaded (ready) and available (downloaded but not active). |
stats | read | Show user stats: tokens offloaded, calls made, per-model performance - for the current session AND |
Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | WARN |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- declared (5 observation(s))
- Shell
- none-observed
- Dependencies
- not all pinned
- Secrets in source
- found
Findings (10)
at: Date.now(), token: 'held-by-someone-else',
-e HOUTINI_LM_ENDPOINT_URL=http://192.168.1.50:1234 \
-e HOUTINI_LM_ENDPOINT_URL=http://192.168.1.50:1234 \
the host's LAN IP (e.g. `http://192.168.1.50:1234`).
- On a different machine: use its LAN URL (`http://192.168.1.50:1234`).
**Remote hosts need `OLLAMA_HOST`.** Ollama binds to localhost by default. To reach it from another machine, start it with `OLLAMA_HOST=0.0.0.0` and point `HOUTINI_LM_ENDPOINT_URL` at the box's LAN ad
@modelcontextprotocol/sdk, @types/node, typescript
When something went wrong, a quality line says so: `TRUNCATED` for a partial result (a stalled connection gives you what arrived rather than a timeout error), `hit-max-tokens` when the budget ran out,
**`on`** forces thinking on. It's worth it for work where the model's own reasoning improves the answer, like hunting a subtle bug, checking an argument or planning something with several moving parts
A plain OpenAI endpoint also doesn't report context windows or output caps in its model list, so houtini-lm starts from its 100,000-token fallback. When a budget overshoots a model's real output cap (
Gates applied: no_behavioural_pass.
b3fa9959e480full audit observations/trust-audit/mcp-server/houtini-ai__houtini-lm-1.json · Report an issue / request a re-scanAudit history
Every audit this server has had. A grade with a past is a grade somebody is still checking.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-07 | b3fa9959e480 | SAFE | B | 89 | first audit |
Questions
What is the Houtini LM MCP server?
MCP server that saves Claude Code tokens by delegating bounded tasks to local or cloud LLMs. Works with LM Studio, Ollama, vLLM, DeepSeek, Groq, Cerebras.
What tools does Houtini LM expose?
8 in total: 6 read-only, 2 that write, and 0 that can delete or overwrite. Every one is listed on this page with its risk.
Is Houtini LM safe to connect to an agent?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean server reads B.
What credentials does Houtini LM need?
It reads HF_TOKEN, HOUTINI_LM_API_KEY, HOUTINI_LM_MIN_TOKENS, LM_PASSWORD, LM_STUDIO_PASSWORD and OPENROUTER_API_KEY from the environment. Give it a token scoped to the least it needs — an agent that can be talked into calling a tool can be talked into calling it with your credentials.
How does Houtini LM run?
It speaks stdio, so it runs as a local process your client starts. It is published on npm as @houtini/lm at 3.3.4.
How current is this page?
The grade is for one exact copy of the source (b3fa9959e480), read on 2026-10-07. The repository is watched and re-audited when it changes.