Atlas / MCP servers / withrefresh / Web Eval Agent

Web Eval AgentCAUTION

mcp/withrefresh/web-eval-agent

An MCP server that autonomously evaluates web applications.

Verdict
CAUTION
Grade
B
Trust score
88 /100
Exposed tools
2 2r · 0w · 0d
Transport
stdio
License
Apache-2.0
Stars
1,235
01

Overview

From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.

This project has been discontinued. We're building something new at withrefresh.com

🚀 operative.sh web-eval-agent MCP Server

Let the coding agent debug itself, you've got better things to do.

🔥 Supercharge Your Debugging

operative.sh's MCP Server launches a browser-use powered agent to autonomously execute and debug web apps directly in your code editor.

⚡ Features

  • 🌐 Navigate your webapp using BrowserUse (2x faster with operative backend)
  • 📊 Capture network traffic - requests are intelligently filtered and returned into the context window
  • 🚨 Collect console errors - captures logs & errors
  • 🤖 Autonomous debugging - the Cursor agent calls the web QA agent mcp server to test if the code it wrote works as epected end-to-end.

🧰 MCP Tool Reference

Key arguments

  • web_eval_agent
  • url (required) – address of the running app (e.g. http://localhost:3000)
  • task (required) – natural-language description of what to test ("run through the signup flow and note any UX issues")
  • headless_browser (optional, default `false`) – set to true to hide the browser window
  • setup_browser_state
  • url (optional) – page to open first (handy to land directly on a login screen)

You can trigger these tools straight from your IDE chat, for example:

Evaluate my app at http://localhost:3000 – run web_eval_agent with the task "Try the full signup flow and report UX issues".

🏁 Quick Start

Easy Setup with One-Cli

Read from source at commit 5dac76872dc1OBSERVED · 2026-09-25
02

Connect

Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control. Replace the environment placeholders with a token scoped to the least it needs.

claude-code
claude mcp add web-eval-agent --env ANTHROPIC_API_KEY=${ANTHROPIC_API_KEY} --env OPERATIVE_API_KEY=${OPERATIVE_API_KEY} -- uvx web-eval-agent
claude-desktop
{
  "mcpServers": {
    "web-eval-agent": {
      "command": "uvx",
      "args": [
        "web-eval-agent"
      ],
      "env": {
        "ANTHROPIC_API_KEY": "${ANTHROPIC_API_KEY}",
        "OPERATIVE_API_KEY": "${OPERATIVE_API_KEY}"
      }
    }
  }
}
03

Exposed tools (2)

2 read · 0 write · 0 destructive.

ToolRiskDescription
setup_browser_statereadSets up and saves browser state for future use.
web_eval_agentreadEvaluate the user experience / interface of a web application.
04

Trust audit

CAUTIONgrade B · trust 88/100 Install with care. The audit found things worth knowing before you trust its output.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codeWARN
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
declared (5 observation(s))
Shell
declared (1 observation(s))
Dependencies
pinned
Secrets in source
none-found

Findings (9)

MEDIUMNetwork egress · net.raw_ip · CWE-200, CWE-319
webEvalAgent/src/browser_utils.py:824
disable_security=True, headless=headless, cdp_url="http://127.0.0.1:9222"
MEDIUMNetwork egress · net.raw_ip · CWE-200, CWE-319
webEvalAgent/src/env_utils.py:27
base_url = "http://0.0.0.0:8000"
MEDIUMNetwork egress · net.raw_ip · CWE-200, CWE-319
webEvalAgent/src/log_server.py:313
def open_log_dashboard(url='http://127.0.0.1:5009'):
MEDIUMNetwork egress · net.raw_ip · CWE-200, CWE-319
webEvalAgent/src/log_server.py:342
open_log_dashboard(url='http://127.0.0.1:5009')
LOWInventory / provenance · inv.hidden_file · CWE-1104
.pre-commit-config.yaml
.pre-commit-config.yaml
Why it matters. hidden member outside the usual dotfiles
Fix. review its purpose
LOWSupply chain · prompt.pipe_to_shell · CWE-829, CWE-1357
README.md:80
curl -LsSf https://astral.sh/uv/install.sh | sh
LOWSupply chain · prompt.pipe_to_shell · CWE-829, CWE-1357
README.md:127
curl -LsSf https://astral.sh/uv/install.sh | sh)
LOWSupply chain · prompt.pipe_to_shell · CWE-829, CWE-1357
README.md:142
4. Install uv `(curl -LsSf https://astral.sh/uv/install.sh | sh)`
INFOInventory / provenance · inv.oversize · CWE-1104
demo.gif
demo.gif
Why it matters. 32031071 bytes not read

Gates applied: no_behavioural_pass.

Audited 2026-09-25 · audit v0.4.1 · source sha 5dac76872dc1full audit observations/trust-audit/mcp-server/withrefresh__web-eval-agent.json · Report an issue / request a re-scan
05

Audit history

Every audit this server has had. A grade with a past is a grade somebody is still checking.

DateSourceVerdictGradeScoreChange
2026-09-255dac76872dc1CAUTIONB88first audit
06

Questions

What is the Web Eval Agent MCP server?

An MCP server that autonomously evaluates web applications.

What tools does Web Eval Agent expose?

2 in total: 2 read-only, 0 that write, and 0 that can delete or overwrite. Every one is listed on this page with its risk.

Is Web Eval Agent safe to connect to an agent?

With care. The audit graded it B (88/100) and found 9 things worth knowing before you trust this server, listed below with the exact line each was found on.

What credentials does Web Eval Agent need?

It reads ANTHROPIC_API_KEY and OPERATIVE_API_KEY from the environment. Give it a token scoped to the least it needs — an agent that can be talked into calling a tool can be talked into calling it with your credentials.

How does Web Eval Agent run?

It speaks stdio, so it runs as a local process your client starts. It is published on PyPI as web-eval-agent.

How current is this page?

The grade is for one exact copy of the source (5dac76872dc1), read on 2026-09-25. The repository is watched and re-audited when it changes.

Advertisement