vMLXBLOCK
vMLX - Use MLX models easily - JANGQ (GGUF for MLX) - Not dependant on mlx_vlm
Overview
From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.
MLX Inference Server for Apple Silicon
Self-hosted inference server for LLMs, VLMs, and image generation on Apple Silicon. OpenAI + Anthropic + Ollama compatible HTTP API. Self-hosted; no third-party API keys required. Native MTP artifact detection and family-specific cache policy gates keep speculative/cache settings explicit and model-safe.
Looking for a native Swift macOS app or Swift inference engine? See osaurus.ai.
4b985ff8753aOBSERVED · 2026-09-27Connect
Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control. Replace the environment placeholders with a token scoped to the least it needs.
claude mcp add vmlx --env BRAVE_API_KEY=${BRAVE_API_KEY} --env DSV4_MAX_PREFILL_TOKENS=${DSV4_MAX_PREFILL_TOKENS} --env DSV4_PROMPT_SNAPSHOT_MIN_TOKENS=${DSV4_PROMPT_SNAPSHOT_MIN_TOKENS} --env HF_TOKEN=${HF_TOKEN} -- npx -y [email protected]{
"mcpServers": {
"vmlx": {
"command": "npx",
"args": [
"-y",
"[email protected]"
],
"env": {
"BRAVE_API_KEY": "${BRAVE_API_KEY}",
"DSV4_MAX_PREFILL_TOKENS": "${DSV4_MAX_PREFILL_TOKENS}",
"DSV4_PROMPT_SNAPSHOT_MIN_TOKENS": "${DSV4_PROMPT_SNAPSHOT_MIN_TOKENS}",
"HF_TOKEN": "${HF_TOKEN}"
}
}
}
}Exposed tools (41)
25 read · 15 write · 1 destructive. Blast radius: 1 tool can delete or overwrite — an agent that can be talked into calling a tool can be talked into calling this one.
| Tool | Risk | Description |
|---|---|---|
add | write | Add two integers. |
apply_regex | write | Apply a regex find-and-replace across one or more files. Returns the number of replacements made per file. |
ask_user | read | Ask the user a question and wait for their response. Use when you need clarification, confirmation, or user input to proceed with a task. |
batch_edit | write | Apply multiple find-and-replace edits to a file in a single call. More efficient than multiple edit_file calls. Each edit is applied sequentially. |
clipboard_read | read | Read the current contents of the system clipboard. |
clipboard_write | write | Write text to the system clipboard. |
copy_file | read | Copy a file to a new location. |
count_tokens | read | Estimate the token count of a text string using character and word heuristics. |
create_directory | write | Create a directory and any necessary parent directories. |
delete_file | destructive | Delete a file or empty directory. Use with caution — this cannot be undone. |
diff_files | read | Show a unified diff between two files, or between a file and its git HEAD version (if path_b is omitted). |
echo | read | Return the provided text. |
edit_file | write | Edit a file by finding and replacing text. The search_text must match exactly (including indentation). Use read_file first to see the current content. Set replace_all=true to replace ALL occurrences (useful for renaming variables/functions). |
fetch_url | read | Fetch a URL and return its content as text. HTML is automatically converted to readable text. Useful for reading documentation, API responses, or web pages. |
file_info | read | Get metadata about a file or directory: size, type (file/directory/symlink), last modified time, permissions. |
find_files | read | Find files by name pattern. Returns matching file paths. Useful for finding files when you don\ |
get_current_datetime | read | Get the current date, time, and timezone. Use this when you need to know the current date or time. |
get_diagnostics | write | Run diagnostics on a file or project: type checking (TypeScript), linting, or syntax validation. Returns errors and warnings with file locations. |
get_process_output | read | Read stdout/stderr from a previously spawned background process. Returns latest output and whether the process is still running. |
get_tree | read | Get a project directory tree respecting .gitignore rules. Shows the hierarchical file and directory structure. |
git | write | Run git commands in the working directory. Supports status, diff, log, blame, add, commit, branch, checkout, stash, show, and more. Blocks destructive operations (push --force, reset --hard). |
insert_text | write | Insert text at a specific line number in a file. The new text is inserted BEFORE the specified line. Use read_file first to see line numbers. |
list_directory | read | List files and directories at a path. Shows file types and sizes. |
move_file | write | Move or rename a file or directory. |
patch_file | write | Apply a unified diff patch to a file. Useful for making multiple related edits at once. The patch should be in standard unified diff format (--- a/file, +++ b/file, @@ hunks). |
read_file | read | Read the contents of a file with line numbers. Returns up to 2000 lines by default. Use offset/limit for large files. |
read_image | read | Read an image file and return its base64-encoded data with MIME type. Supports png, jpg, gif, webp, svg. Max 10MB. |
read_video | read | Read a local video file and attach it for VL model analysis. Supports mp4, mov, m4v, webm, mkv. Max 50MB. |
record_gemma_label | read | Record a label for the Gemma API stress test. |
record_gemma_stream_label | read | Record a Gemma streaming label. |
record_gemma_stream_response_label | read | Record a Gemma Responses streaming label. |
record_mm3_label | read | Record a label for the MM3 API stress test. |
record_mm3_stream_label | read | Record an MM3 streaming label. |
record_mm3_stream_response_label | read | Record an MM3 Responses streaming label. |
replace_lines | read | Replace a range of lines in a file with new content. Use read_file first to see line numbers. |
run_applescript | write | Execute AppleScript on this Mac with /usr/bin/osascript and return its output. Use this for macOS app automation and system scripting. |
run_command | write | Execute a shell command in the working directory. Has a 60 second timeout. Returns stdout, stderr, and exit code. |
search_files | read | Search file contents for a pattern using ripgrep. Returns matching lines with file paths and line numbers. |
spawn_process | write | Start a long-running background process (e.g., dev server, watcher). Returns a process ID for checking output later. Auto-kills after 5 minutes. |
web_search | read | Search the web using Brave Search. Returns titles, URLs, and descriptions of matching pages. |
write_file | write | Write content to a file. Creates the file and parent directories if they don\ |
Trust audit
BLOCKgrade F · trust 45/100 Do not install this without reading the findings. The audit found something that could harm you or your machine.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | FAIL |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | WARN |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- declared (17 observation(s))
- Network
- declared (10 observation(s))
- Shell
- declared (10 observation(s))
- Dependencies
- not all pinned
- Secrets in source
- found
Findings (25)
opaque_decoded[k] = pickle.loads(base64.b16decode(hexed))
const result = await exec(`"${path}" --version 2>&1`)const result = await exec(`"${pipPath}" --version 2>&1`)basename in {".pypirc", ".npmrc", ".netrc"}__import__(mod)
_mod = importlib.import_module(_mod_name)
print(f" {'(sum sublayers/token)':<28} {'':>18} {sub_total_token_ms:>17.2f} ms")logger.info(" Auth: cluster secret set (%d bytes)", len(secret))print(f"Request {resp.request_id}: token={resp.token}")base_url = f"http://127.0.0.1:{port}"base_url = f"http://127.0.0.1:{port}"base_url = f"http://127.0.0.1:{args.port}"BASE = sys.argv[1] if len(sys.argv) > 1 else "http://127.0.0.1:8000"
token: 'REAL_UI_LIVE_TOOL_ONE',
token: 'REAL_UI_LIVE_TOOL_TWO',
delete_file
path.name: (path.read_text(), yaml.load(path.read_text(), Loader=yaml.BaseLoader))
workflow = yaml.load((ROOT / ".github/workflows/publish-release.yml").read_text(), Loader=yaml.BaseLoader)
workflow = yaml.load((ROOT / ".github/workflows/publish-release.yml").read_text(), Loader=yaml.BaseLoader)
module = importlib.import_module(f"tests.cross_matrix.{name}")module = importlib.import_module(module_path)
module = importlib.import_module(f"vmlx_engine.reasoning.{mod.name}")const handler = new Function(...Object.keys(dependencies), code)(...Object.values(dependencies))
await new Function(...Object.keys(context), `return (async () => {${source.slice(start, end)}})()`)(...Object.values(context))const run = new Function('window', 'sessionId', 'activeSessionId', 'showToast', 't', code)(Gates applied: no_behavioural_pass.
4b985ff8753afull audit observations/trust-audit/mcp-server/jjang-ai__vmlx.json · Report an issue / request a re-scanAudit history
Every audit this server has had. A grade with a past is a grade somebody is still checking.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-09-27 | 4b985ff8753a | BLOCK | F | 45 | first audit |
Questions
What is the vMLX MCP server?
vMLX - Use MLX models easily - JANGQ (GGUF for MLX) - Not dependant on mlx_vlm
What tools does vMLX expose?
41 in total: 25 read-only, 15 that write, and 1 that can delete or overwrite (delete_file). Every one is listed on this page with its risk.
Is vMLX safe to connect to an agent?
No — not without reading the findings first. The audit graded it F (45/100) and found 4 critical or high issues in the source. Each one is listed on this page with the file and line it is on. Separately from the audit: 1 of its tools can destroy data, so scope the token you give it to what you actually need.
What credentials does vMLX need?
It reads BRAVE_API_KEY, DSV4_MAX_PREFILL_TOKENS, DSV4_PROMPT_SNAPSHOT_MIN_TOKENS, HF_TOKEN, TOKENIZERS_PARALLELISM, VLLM_API_KEY, VMLINUX_API_KEY, VMLINUX_CACHE_SELECTION_HOT_ADVANTAGE_TOKENS, VMLINUX_DSV4_PROMPT_SNAPSHOT_MIN_TOKENS, VMLINUX_MIMO_TEXT_PREFILL_REJECT_TOKENS, VMLINUX_MIMO_TEXT_PREFILL_TOTAL_TOKENS and VMLINUX_MIMO_V2_TOKEN_TRACE from the environment. Give it a token scoped to the least it needs — an agent that can be talked into calling a tool can be talked into calling it with your credentials.
How does vMLX run?
It speaks sse, stdio and streamable-http, so it runs as a local process your client starts. It is published on npm as vmlx at 1.6.67.
How current is this page?
The grade is for one exact copy of the source (4b985ff8753a), read on 2026-09-27. The repository is watched and re-audited when it changes.