Clawd CursorBLOCK
clawdcursor compiles whatever's on screen into one UI map — accessibility tree and OCR fused into stable, addressable elements, with a screenshot only when needed — then drives apps through reusable scripts, verifying every action and routing it through a single safety gate.
Overview
From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.
Clawd Cursor
Safe desktop control for any AI agent. Compiles the screen into a UI map and acts on elements by stable id (screenshot/vision only as a last resort), verifies its own actions, and gates everything through one safety checkpoint. Local · cross-OS · any model.
Quickstart · Why it's different · The engine · How it works · Tools · Platforms · Changelog
1048e1073155OBSERVED · 2026-10-01Connect
Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control. Replace the environment placeholders with a token scoped to the least it needs.
claude mcp add clawdcursor -- npx -y [email protected] mcp
Exposed tools (119)
97 read · 20 write · 2 destructive. Blast radius: 2 tools can delete or overwrite — an agent that can be talked into calling a tool can be talked into calling this one.
| Tool | Risk | Description |
|---|---|---|
__nonexistent_tool__ | read | test |
a11y_collapse | read | Collapse a tree node / combo / disclosure by a11y name. |
a11y_expand | read | Expand a tree node / combo / disclosure by a11y name (UIA ExpandCollapsePattern, AX AXExpanded). |
a11y_get_element | read | Fetch a single element by accessibility name — returns name, role, |
a11y_get_value | read | Read the current value of a named field (UIA ValuePattern / AX AXValue). Useful to verify before typing. |
a11y_list_children | read | List a11y elements geometrically contained within a named parent element. |
a11y_select | read | Select a list item / tab / radio by a11y name (UIA SelectionItemPattern, AX AXSelected). |
a11y_toggle | read | Toggle a checkbox / switch / toggle-button by a11y name. Returns new state (On/Off/Indeterminate). |
abort_task | read | Signal the running task to abort. The pipeline checks |
accessibility | read | Interact with the OS accessibility tree — read element names, find by name/role, invoke, toggle, expand/collapse, set value, query state. |
agent_status | read | Return the autonomous agent\ |
batch | write | Run SEVERAL known next actions in ONE call instead of one per turn (saves round-trips). |
browser | read | Chrome DevTools Protocol control — operates on DOM elements by CSS selector rather than screen pixels. Requires Chrome/Edge launched with remote debugging (see |
browser_click | read | Click a page element by visible text or CSS selector (DOM click — no coordinates). Requires browser_connect first. |
browser_connect | read | Open/attach a dedicated browser the agent controls via the DOM (reliable for web pages — no pixels). Call this FIRST for any website task, then use browser_navigate/read/click/type. If it fails, fall back to read_text/smart_click. |
browser_navigate | read | Navigate the agent-owned browser to a URL (waits for load). Requires browser_connect first. |
browser_read | read | Read the current page as structured DOM: interactive elements (links/buttons/inputs with selectors), or text for a CSS selector. Use instead of read_text on web pages. Requires browser_connect first. |
browser_type | read | Type text into a page input by CSS selector or associated label (DOM input — no coordinates). Requires browser_connect first. |
build_uri | read | Build a properly-encoded URI from a scheme + path + query JSON. Returns the URI text; pair with open_uri to dispatch. Examples: scheme= |
cannot_read | read | Escalate from blind mode to vision — the a11y snapshot doesn\ |
cdp_click | read | Click a DOM element by CSS selector or by visible text content. |
cdp_connect | read | Connect to Edge/Chrome browser via Chrome DevTools Protocol (port ${DEFAULT_CDP_PORT}). Must be called before other cdp_* tools. Use navigate_browser to launch Edge with CDP enabled. |
cdp_evaluate | write | Execute JavaScript in the browser page context. Returns the result. |
cdp_list_tabs | read | List all open browser tabs with their URLs and titles. |
cdp_page_context | read | Get a structured list of interactive elements on the current browser page (inputs, buttons, links with selectors and positions). |
cdp_read_text | read | Read text content from a DOM element. Useful for extracting information from web pages. |
cdp_scroll | read | Scroll the browser page via DOM (window.scrollBy). Works regardless of mouse position — use for reliable page scrolling. |
cdp_select_option | read | Select an option in a <select> dropdown by value or visible text. |
cdp_switch_tab | read | Switch CDP connection to a different browser tab by URL or title substring. |
cdp_type | read | Type text into a DOM input field by CSS selector or by associated label text. |
cdp_wait_for_selector | read | Wait for a DOM element matching a CSS selector to appear and become visible. |
click | read | Click at (x,y). The default coordinate space follows context (image-space while a screenshot is in your context, else screen-space) — pass |
close_window | read | Polite close request (WM_CLOSE / AXCloseAction / _NET_CLOSE_WINDOW). App may prompt. |
computer | read | Direct mouse/keyboard/screenshot control (Anthropic Computer-Use style). |
cursor_position | read | Read the current mouse cursor position in image-space coordinates (the same space the mouse_* tools accept — a cursor_position read round-trips with mouse_hover). Computer-use parity: pointer-state read, no side effects. |
delegate_to_agent | read | Hand a whole task to clawdcursor |
desktop_screenshot | read | LAST RESORT — take a screenshot only when the accessibility tree and OCR are both insufficient (custom canvas, icon-only UI, pixel-level verification). Prefer read_screen first, then ocr_read_screen; escalate to screenshot only when those fail. Returns the image resized to 1280px wide. |
desktop_screenshot_region | read | Take a zoomed screenshot of a specific screen region for detailed inspection. Coordinates are in image-space (from desktop_screenshot). |
detect_webview_apps | read | Enumerate running Electron / WebView2 / Chromium-shell apps (e.g. New Outlook, Teams, Discord, Slack, VS Code) whose |
done | read | Declare the task complete. Provide SPECIFIC screen evidence — a window title, a value visible in the document, a status bar message. Do NOT use hedging words ( |
drag | read | Drag the mouse from (startX,startY) to (endX,endY) — select text, draw, resize. To TRACE A CURVE/PATH (gesture, curved track, drawing), pass |
favorites_add | write | Add a task string to the favorites list. No-op if already present. |
favorites_list | read | Return the list of starred ( |
favorites_remove | destructive | Remove a task string from the favorites list. Returns the updated list. |
find_action_button | read | Semantically locate the best clickable element for an intent (e.g. |
find_element | read | Search for UI elements by name, control type, or automation ID within a process. Returns matching elements with bounds. For browser windows with CDP attached, falls back to a DOM query when UIA returns empty so canvas / SPA content is reachable too (results are flagged with a |
find_input_field | read | Semantically locate the best editable field for a purpose (e.g. |
focus_element | read | Keyboard-focus an element by a11y name. Does NOT raise window — use focus_window first if needed. |
focus_window | read | Bring a window to the foreground. Match by processName, pid, or title substring. |
get_active_window | read | Get the currently focused/foreground window. |
get_element_state | read | Get state flags of a named element (focused/enabled/disabled/selected/busy/offscreen/expandable/expanded). |
get_focused_element | read | Get the currently focused UI element (keyboard focus). Returns name, control type, value, bounds, and process ID. |
get_screen_size | read | Get the screen dimensions and scale factor. |
get_system_prompt | read | Return the canonical agent system prompt. Use this to mirror |
get_system_time | read | Return current system time (ISO, epoch, timezone). Zero I/O. |
get_windows | read | List all visible windows with their title, process name, PID, and bounds. |
give_up | read | Abandon the task when it\ |
invoke_element | write | Click/activate a UI element by its accessibility name. MORE RELIABLE than coord clicks — use this when the snapshot shows a named target. |
key | read | Press a key or key combo. Use |
key_down | read | Press a key without releasing. Pair with key_up. Use to hold modifiers (shift, ctrl) during clicks. |
key_press | read | Press a keyboard key or key combination. Use |
key_up | read | Release a key previously pressed with key_down. |
list_displays | read | Enumerate connected displays with logical bounds + DPI ratio. Use before display-specific screenshots. |
list_windows | read | List visible top-level windows with title, process, and bounds. Useful when the active window is wrong or missing. |
logs_recent | read | Return the last 200 captured console log entries from the daemon |
maximize_window | read | Maximize the foreground window (or a matched window). Polite request; WM may interpret. |
minimize_window | read | Minimize the foreground or matched window to the taskbar / Dock. |
minimize_window_to_taskbar | read | Minimize a window (hide to taskbar/Dock). Counterpart to the existing |
mouse_click | read | Click the left mouse button at the given image-space coordinates. |
mouse_double_click | read | Double-click the left mouse button at the given image-space coordinates. |
mouse_down | read | Press a mouse button without releasing. Pair with mouse_up. Enables hold-and-drag + modifier clicks. |
mouse_drag | write | Drag from one image-space coordinate to another (click-hold-move-release). Useful for selecting text, moving objects, or resizing. |
mouse_drag_stepped | read | Drag the mouse along a multi-point path in image-space. |
mouse_hover | write | Move the mouse to the given image-space coordinates without clicking. Useful for revealing tooltips or hover menus. |
mouse_middle_click | read | Middle-click (wheel-click) at image-space (x, y). Opens links in new tab in most browsers; pans in some apps. |
mouse_move_relative | write | Move cursor by a relative offset (dx, dy). Wayland-safe via cursor cache. |
mouse_right_click | read | Right-click at the given image-space coordinates (opens context menu). |
mouse_scroll | read | Scroll the mouse wheel at the given image-space coordinates. |
mouse_scroll_horizontal | read | Scroll horizontally at image-space (x, y). On Windows uses Shift+wheel synthesis; |
mouse_triple_click | read | Triple-click the left mouse button at image-space (x, y) — selects a paragraph in most text editors. In single-line edit fields (dialog filename boxes etc.) a select-all follows automatically when the triple-click left nothing selected, so subsequent typing reliably REPLACES the existing text. |
mouse_up | read | Release a mouse button previously pressed with mouse_down. |
move | write | Move the cursor to (x,y) WITHOUT clicking — hover/dwell over a target (pair with wait(ms) for a required dwell time). The default coordinate space follows context; pass space: |
navigate_browser | read | Open a URL in the browser. Launches with CDP enabled (port ${DEFAULT_CDP_PORT}) for DOM interaction. Call cdp_connect after. Tier 2 (mutation): triggers network egress to an arbitrary destination + spawns/attaches to a browser process. |
ocr_read_screen | read | Step 2 of cheap-first perception: use when the a11y tree (read_screen) is empty or too sparse to identify your target. OS-level OCR returns text elements with pixel coordinates — no image bytes, no vision model. Much cheaper than a screenshot. Coordinates are in real screen pixels. |
open_app | read | Open an application by name (e.g. |
open_file | read | Open a file or folder in the OS default app (explorer / open / xdg-open). |
open_url | read | Open a URL in the default browser. Use instead of navigate_browser when you don\ |
read_clipboard | read | Read the OS clipboard. |
read_text | read | OCR the screen and return visible text + positions. Use when the a11y snapshot is empty/sparse (webview, canvas, PDF, game) to READ on-screen content. Cheaper than a screenshot (no image bytes). May take 1–3s. |
relaunch_with_cdp | read | Relaunch a WebView2 / Electron app with a remote-debugging port enabled so |
resize_window | write | Set the foreground (or matched) window bounds in logical pixels. Omitted fields preserved. |
restore_window | read | Restore a minimized or maximized window to its previous bounds. |
scheduled_task_create | write | Create a recurring task that runs on a cron schedule. The task |
scheduled_task_delete | destructive | Delete one scheduled task by id. No-op if id is unknown. |
scheduled_task_list | read | Return every scheduled task: id, cron expression, task text, enabled |
scheduled_task_toggle | read | Pause/resume a scheduled task without deleting it. When disabled the |
screenshot | write | LAST RESORT — expensive: sends image bytes into LLM context. Escalation order: read_screen (a11y tree, free) → read_text (OCR, cheap) → screenshot (this, expensive). Only call this when both a11y and OCR failed to provide what you need (canvas-only app, icon-only UI, pixel-level verification). |
screenshot_full | read | Capture the full primary display. Returns base64 image bytes plus a |
scroll | read | Scroll at (x,y) in a direction. Omit x,y to scroll at the screen center. If you read x,y off the SCREENSHOT, pass space: |
set_field_value | write | Set an editable field\ |
shortcuts_execute | write | Execute a keyboard shortcut by describing what you want to do (e.g. |
shortcuts_list | write | List available keyboard shortcuts. Filter by category (navigation, browser, editing, social, window, file, view, quick) and/or context (e.g. |
smart_click | read | OCR-locate visible text on screen and click its center. Use when the a11y tree is empty and invoke_element fails (webview/canvas). Pass the exact visible text (e.g. |
smart_read | read | Read text from the screen with automatic fallback. |
smart_type | read | Type text into a UI element. If target is specified, finds and focuses the element first. |
submit_report | write | Submit a redacted task-log report to clawdcursor\ |
submit_task | write | Submit a natural-language task to the autonomous agent. The agent |
switch_tab_os | read | Cycle next/previous browser tab (mod+Tab / mod+Shift+Tab) or jump to tab N (mod+1..9). |
system | write | System integration — clipboard read/write, system time, OCR screen-reading, undo shortcut, named shortcuts registry, delegate to a sub-agent. |
task | read | **Requires the |
task_logs_current | read | Return the structured log entries for the currently-executing task |
task_logs_list | read | List the last 50 task summaries from the daemon\ |
type | read | Type text into the currently focused input. Prefer set_field_value when a field has an a11y name. |
type_text | read | Type text into the currently focused element. Internally uses clipboard paste for reliability (no dropped chars). The user clipboard is saved before and restored after, so calling type_text never clobbers any text the caller had previously placed on the clipboard. |
undo_last | write | Send the OS Undo keystroke (mod+Z). |
wait | write | Pause for N milliseconds (max 5000). Use after actions that trigger animations or page loads. |
wait_for_element | read | Poll the a11y tree until an element matching name/controlType appears. Useful after an action spawns a dialog. |
window | read | Window, app, and display management. Open/focus/maximize/minimize/restore/close/resize windows; enumerate displays; switch browser tabs at the OS level; open apps/files/URLs. |
write_clipboard | write | Write text to the OS clipboard. |
Trust audit
BLOCKgrade F · trust 50/100 Do not install this without reading the findings. The audit found something that could harm you or your machine.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | WARN |
| L1 | Static analysis of the code | WARN |
| L2 | Instruction surface (what it tells the agent) | FAIL |
| L3 | Class-specific surface | WARN |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- declared (5 observation(s))
- Network
- declared (6 observation(s))
- Shell
- declared (5 observation(s))
- Dependencies
- not all pinned
- Secrets in source
- found
Findings (25)
The HTTP transport uses **MCP's streamable-HTTP envelope** (JSON-RPC over POST), not REST. All requests go to a single endpoint, `POST /mcp`, with `Authorization: Bearer <token>` from `~/.clawdcursor/
The HTTP transport uses **MCP's streamable-HTTP envelope** (JSON-RPC over POST), not REST. All requests go to a single endpoint, `POST /mcp`, with `Authorization: Bearer <token>` from `~/.clawdcursor/
el_NN ref clicks (`invoke_element({element_id, snapshot_id})`) reach the safety gate with NO `targetLabel`, tripping the blunt "sensitive app + no target label → confirm" dead-end. Resolve the ref to find-element.jxa
console.log(`${pc.yellow(`${e('🔑', '[KEY]')} Auth token:`)} ${serverToken.slice(0, 8)}...`);console.log(pc.gray(` (full token saved to ${tokenPath})`));console.log(` Stdio transport — no daemon, no port, no token.\n`);
clawdcursor agent # starts on http://127.0.0.1:3847; built-in agent lights up if an LLM is configured
const TEST_IMAGE = '/9j/2wBDAAYEBQYFBAYGBQYHBwYIChAKCgkJChQODwwQFxQYGBcUFhYaHSUfGhsjHBYWICwgIyYnKSopGR8tMC0oMCUoKSj/2wBDAQcHBwoIChMKChMoGhYaKCgoKCgoKCgoKCgoKCgoKCgoKCgoKCgoKCgoKCgoKCgoKCgoKCgoKCgoKCgo
apiKey: 'anthropic-auth-profile-key',
apiKey: 'anthropic-auth-profile-key',
favorites_remove, scheduled_task_delete
.nojekyll
const hash = crypto.createHash('md5').update(sample).digest('hex');const digest = createHash('sha1').update(tr.screenshot.buffer).digest('hex');const val = e.value ? crypto.createHash('sha1').update(e.value).digest('hex').slice(0, 8) : '';return crypto.createHash('sha1').update(payload).digest('hex').slice(0, 16);import type { PlatformAdapter } from '../../platform/types';vi.mock('../../src/platform/ocr-engine', async () => ({import type { ScreenshotResult } from '../../platform/types';import { OcrEngine } from '../../platform/ocr-engine';} from '../../llm/client';
- **`clawdcursor dashboard` removed** from the README CLI block — that command never existed; the dashboard is reachable at `http://127.0.0.1:3847` while `serve` or `start` is running. `status` and `c
| **HTTP MCP** | Headless agents, daemons, orchestration, Agent SDK. POST JSON-RPC to `http://127.0.0.1:3847/mcp`. | Run `clawdcursor agent`. Bearer token at `~/.clawdcursor/token`. |
curl -s -X POST http://127.0.0.1:3847/mcp \
Gates applied: critical_finding, instruction_override, no_behavioural_pass, undeclared_transfer.
1048e1073155full audit observations/trust-audit/mcp-server/amrdab__clawd-cursor.json · Report an issue / request a re-scanAudit history
Every audit this server has had. A grade with a past is a grade somebody is still checking.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-01 | 1048e1073155 | BLOCK | F | 50 | first audit |
Questions
What is the Clawd Cursor MCP server?
clawdcursor compiles whatever's on screen into one UI map — accessibility tree and OCR fused into stable, addressable elements, with a screenshot only when needed — then drives apps through reusable scripts, verifying every action and routing it through a single safety gate.
What tools does Clawd Cursor expose?
119 in total: 97 read-only, 20 that write, and 2 that can delete or overwrite (favorites_remove, scheduled_task_delete). Every one is listed on this page with its risk.
Is Clawd Cursor safe to connect to an agent?
No — not without reading the findings first. The audit graded it F (50/100) and found 3 critical or high issues in the source. Each one is listed on this page with the file and line it is on. Separately from the audit: 2 of its tools can destroy data, so scope the token you give it to what you actually need.
What credentials does Clawd Cursor need?
It reads AI_API_KEY, ANTHROPIC_API_KEY, CLAWD_ALLOW_DISK_TOKEN_DRIFT, GEMINI_API_KEY, GOOGLE_API_KEY, GROQ_API_KEY, KIMI_API_KEY, MOONSHOT_API_KEY, MY_CUSTOM_PROVIDER_API_KEY, OPENAI_API_KEY, OPENCLAW_AGENT_API_KEY and OPENCLAW_AI_API_KEY from the environment. Give it a token scoped to the least it needs — an agent that can be talked into calling a tool can be talked into calling it with your credentials.
How does Clawd Cursor run?
It speaks stdio and streamable-http, so it runs as a local process your client starts. It is published on npm as clawdcursor at 1.5.11.
How current is this page?
The grade is for one exact copy of the source (1048e1073155), read on 2026-10-01. The repository is watched and re-audited when it changes.