SkillAtlasSkill 详情

verify-ui-change

Your AI agent says "done." Reticle checks whether that's true.

审核状态:已审核Quality 72Security 52

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年9月20日
Reticle

Your AI agent says "done." Reticle checks whether that's true.

It drives your real running app, reads what actually happened, and hands back pass · fail · couldn't tell with the file:line to fix.


npm downloads stars license OpenSSF Discord

Install · Use it · How it works · vs Playwright · Benchmarks · Limits · Docs


Watch: an agent drives a real app, reads the network and the store, and returns a verdict with the file:line to fix

▶ Two minutes. An agent finds a bug the screen was hiding, and proves the fix.


Install

Three steps. The first one is the whole install.

1. Run the installer

macOS · Linux

curl -fsSL https://raw.githubusercontent.com/reticlehq/reticle/main/install/install.sh | sh

Windows (PowerShell)

irm https://raw.githubusercontent.com/reticlehq/reticle/main/install/install.ps1 | iex

It puts reticle on your PATH and registers the MCP server with every coding agent it can reach: Claude Code, Cursor, Windsurf, VS Code, Zed, Gemini CLI, Copilot CLI, OpenCode, Warp, Kiro, Amazon Q, Cline, Amp, Continue and Factory Droid. Nothing is asked. Nothing is written outside your agent configs and ~/.reticle.

Codex CLI keeps a TOML config we don't rewrite, so the installer prints the four lines to paste and where they go. It tells you; it doesn't pretend.

2. Open your coding agent

That's it. The tools are already there, because the installer wrote the config before your agent started.

Run the installer first, in a terminal. A coding agent reads its MCP server list once at startup. Install while it's closed and there is nothing to restart. If it was already open, quit and reopen it once.

3. Wire your app

In your project, let your agent run:

RETICLE_INSTALL_SOURCE=readme npx @reticlehq/server init

Like npm init or git init, this is the per-project step. It installs the dev-only SDK, wires your build config, restarts your dev server, and then opens your app and proves a session connected. A config file is not an install, so init doesn't stop until it has seen your app.

Step 1 is once per machine. This step is once per project.

Check it worked

npx @reticlehq/server doctor

Or just ask your agent: "Is Reticle connected to my app?"

Manual install (no pipe to shell)
npm install -g @reticlehq/server   # 1. the CLI
npx @reticlehq/server setup mcp    # 2. register it with your agents

Step 2 registers the same agents as the installer, and writes the /reticle skill where the agent supports one. On Claude Code it also pre-approves the Reticle tools, so there is no Accept prompt on every call.

Any other MCP client: point it at npx @reticlehq/server mcp.

{ "mcpServers": { "reticle": { "command": "npx", "args": ["@reticlehq/server", "mcp"] } } }

Then continue at step 2 above: open your agent, and run npx @reticlehq/server init in your project.

The RETICLE_INSTALL_SOURCE prefix above is optional. It tells us which page somebody installed from, so we know which docs are working.

Requires Node 20.11+.

Claude Code plugin (skill + MCP in one step)
/plugin marketplace add reticlehq/reticle
/plugin install reticle@reticlehq

Registers the MCP server and installs the Reticle skill together. Reopen Claude Code once and the tools are there.

For other agents that support the skills CLI:

npx skills add reticlehq/reticle

Use it

You never write test syntax. You say what should be true, in plain English.

Verify what you just built

"I changed checkout. Verify it with Reticle before you tell me it's done."

Find what the screen is hiding

"The page looks fine but something's off. Use Reticle to check what's happening underneath."

Prove a bug is fixed

"Reproduce the bug with Reticle, fix it, then prove the fix with the same steps."

Lock a flow so it can't break

"Record the login flow with Reticle, then re-verify it after every change."

Sweep before you ship

"Walk the main routes with Reticle. Tell me anything broken."

Reticle answers with evidence: the request that fired, the state that changed, the console line, and the file to open.


Why not Playwright?

Playwright, DevTools and browser agents all stand outside the browser looking in. For a site you don't own, that's right. For the app you're building, the bugs that matter never reach the pixels.

BugLooks fine on screen?Reticle reads
Pay button silently returns 500yesthe network response, tied to the click
Badge shows "12", the store holds 0yesyour app's state
The form fired the request twiceyesrequest count
"Deploy succeeded", the deploy failedyesthe store's real status
A console error slipped inyesthe console since the action
Component re-renders 60×/secyesthe React commit stream

Use both. Playwright for sites you don't own, many browsers, real pixels. Reticle for the app you're building, inside your agent's loop.


The problem

Your agent writes code, assumes it worked, and moves on. It never opens the app.

So the broken modal, the silent 500, the "Deploy succeeded" over a failed deploy: they all ship, and you find them by clicking around afterwards. You've become your agent's QA.

The truth was in the running app the whole time. It just never reached the screen.

A page that looks shipped, hiding mock data, a dead click and a silent 500.

This isn't something your agent forgot. A coding agent is built to produce a change, and it's optimistic by construction. Verification is the opposite motion: going to find out, and being willing to come back with no.


How it works

You: "Verify login works."

Agent, via Reticle: clicks Sign in → POST /api/login → 200 (14 ms) → dashboard rendered → store holds auth: { email: "admin@…" } → PASS, evidence attached.

flowchart LR
    A["Your agent<br/>(Claude Code, Cursor…)"] -->|"look · act · observe · assert"| B(("Reticle"))
    B <-->|"structured events,<br/>not pixels"| C["Your running app<br/>DOM · network · console<br/>store · React fiber"]
    B -->|"verdict + evidence<br/>+ file:line"| A
    style B fill:#8b7bff,stroke:#5b4bd0,color:#fff
    style A fill:#15131f,stroke:#3a3550,color:#fff
    style C fill:#1c2433,stroke:#2f3d57,color:#fff

One call checks many things at once. Say "save that as a flow" and it replays on every later edit with no model in the loop, so today's fix can't quietly break last week's feature.

What one call looks like underneath
// The agent clicked "Pay". Did the right things actually happen?
reticle_assert({
  predicate: { allOf: [
    { kind: "net",     method: "POST", urlContains: "/api/order", status: 200 },
    { kind: "element", query: { role: "dialog", name: "Order confirmed" }, state: "visible" },
    { kind: "signal",  name: "order:saved" },          // the charge actually committed
    { kind: "console", level: "error", absent: true }  // …and nothing errored
  ]}
})
// → { pass: false,
//     failureReason: "POST /api/order returned 500, expected 200",
//     source: { file: "src/checkout/PayButton.tsx", line: 42 } }

Benchmarks

88 real regressions injected into a controlled app, Reticle against a Playwright script. Every number comes from a committed harness. Reproduce it with pnpm bench.

Reticle catches 14x more bugs where the screen looks right: 28 versus 2 across the six categories the two tools disagree on, and 85/86 versus 59/86 overall.

Cumulative tokens to re-verify a four-flow suite over 100 runs: Reticle 128k, Playwright MCP 12.1M.

Re-verification has no model in the loop, so a recorded suite is a fixed, tiny read. Reticle is ahead from the second run even when charged a full LLM drive to author the suite.

Wall-clock time to a verdict: a 2.6 second time-gated transition verified in 176 ms versus a 2,978 ms real wait, and a 16-flow batch in 5.2 seconds versus 31.7 seconds one at a time.

Faster for a structural reason rather than a browser-speed one: a time-gated transition is verified from the event stream instead of waited out, and a batch of flows runs as a batch.


Limits

A verification tool that oversells its reach is worse than none.

Strongsilent failed requests, state that disagrees with the screen, stale caches, double-submits, a write that failed while the UI moved on
Partialraces around a single action. It detects request-never-settled and duplicate-request; it is not a scheduler-level race analyser
Not the toolcross-browser rendering, visual regressions, sites you don't own. That's Playwright
Can't see yetIndexedDB, Web Workers, closed shadow roots, cross-origin iframes

When Reticle can't see something, it says so. A verdict is yes, no, or unknown, where unknown means the evidence couldn't decide. Never a quiet pass.

Cost: zero bytes in production. The SDK sits behind import.meta.env.DEV and is dead-code eliminated; a runtime guard refuses to connect under NODE_ENV=production. No app data leaves your machine.


Supported

WebNext.js (App + Pages), Vite + React, CRA, SvelteKit, Svelte, Astro, Vue 3, Preact, plain HTML
DesktopElectron, Tauri, including the IPC boundary a browser-only tool can't see
Agentsanything that speaks MCP. Config written automatically for Claude Code, Cursor, Windsurf, VS Code, Zed, Gemini CLI, Copilot CLI, OpenCode, Warp, Kiro, Amazon Q, Cline, Amp, Continue, Factory Droid. Codex CLI is a printed four-line paste
BrowsersChrome, Edge, Arc, Brave, Opera, Firefox, Safari, plus Electron and Tauri webviews
Statezustand and Redux need no adapter. Shipped: TanStack Query, Jotai, XState, Valtio, MobX, Recoil, Svelte stores, Pinia
OSmacOS, Linux, Windows

On the roadmap

Routing verification flows with TypeSafe AI's Jev. A verification run makes a lot of small decisions — is this page settled, is this finding worth chasing, does this failure warrant a full capture — and today an LLM answers each one at LLM latency and LLM cost. Jev is a System One model: it returns a typed, probabilistic choice from a fixed set instead of prose, in 70–500ms. That is the exact shape of a routing decision inside Reticle's infra, so when we build that layer, Jev is what decides which flow a run takes. Nothing ships against it yet.

Docs

docs.reticle.sh — a page per tool, a page per command, every example captured from a real run.

Quickstart · Frameworks · Troubleshooting · Architecture · Contributing

Community

Join the Discord → Where the work happens in the open: what's being built, what's up for grabs, and design calls before they land.

Contributors

If Reticle proves useful, a ⭐ helps other developers find it.

License

SDK and adapters are Apache-2.0. The server is FSL (source-available, converts to Apache-2.0 after two years). See LICENSE.

dev-only · localhost-only · your app data stays local

其他

高风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 可能需要外部 token、网络权限或第三方服务。
  • 存在潜在风险命令,请谨慎安装。
  • 扫描发现:2 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/reticlehq/reticle.git
  3. 将 "skills/verify-ui-change" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/reticlehq/reticle.git
  3. 将 "skills/verify-ui-change" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/reticlehq/reticle.git
  3. 将 "skills/verify-ui-change" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/reticlehq/reticle.git
  3. 将 "skills/verify-ui-change" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/reticlehq/reticle.git
  3. 将 "skills/verify-ui-change" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: verify-ui-change
description: Check that a change to a web app actually works in the running app before calling it done. Drives the real page and returns a pass/fail verdict with the request that fired, the state that moved, and the file:line to fix. Use after editing a component, a form, a route, or an API call; when you have said "fixed" but have not opened the app; when the user asks "does it actually work?"; or when a change looks right on screen and you cannot prove it.
license: Apache-2.0
metadata:
  version: 3.1.0
  homepage: https://www.reticle.sh
  repository: https://github.com/reticlehq/reticle

Verify the change in the running app

You edited something a user can see. Nothing is proven until the real app has done it and something other than the DOM agrees.

This uses Reticle, which embeds a dev-only SDK in the user's running app and exposes it as reticle_* MCP tools. No screenshots, no browser download for the verify loop, no dev server of its own.

1. Are the tools here?

reticle_session({ action: "list" })
  • A session comes back → go to step 2.
  • Tools do not exist → Reticle is not installed. Run RETICLE_INSTALL_SOURCE=npx_skill npx @reticlehq/server@latest init, then keep going without the tools: fix every ⚠ it printed, start the project's own dev script in the background if nothing is serving the app, and open it with npx @reticlehq/server open <the url the dev server is serving>. Only once the app is running and connected, ask for the one thing you cannot do yourself: a client restart, so it picks up the MCP server. Stopping at the init command leaves the user with config files and an uninstrumented page. Full setup is in the install-and-verify skill.
  • Tools exist, list is empty → read the why field on the response. It distinguishes "no app running" from "an app is running that never dialled this daemon" from "a tab was here and closed", and each has a different fix. If no app is running, start the project's own dev script from package.json in the background yourself and tell the user in one line that it is running and how to stop it: never a second one, never a guessed command, never kill anything, and the permission prompt is your host's. If one IS running, the app is not the missing piece and the SDK is; do not send the user to start what they already started.

2. Name the consequence BEFORE you act

This is the whole method. An expectation written after you see the result can be talked into agreeing with whatever happened; one written before cannot.

reticle_look({ action: "page", sessionId, mode: "interactive" })   // controls only, with refs

reticle_act_and_wait({ sessionId, ref, action: "click", until: { kind: "allOf", predicates: [
  { kind: "net",     method: "POST", urlContains: "/api/...", status: 200 },
  { kind: "element", query: { testid: "..." } },
  { kind: "console", level: "error", absent: true },
]}})

Multi-step journey? Drive it in one call with reticle_act { steps: [...] }, then assert the outcome once. Do not act → snapshot → act → snapshot: it proves the same thing at several times the cost.

Only reticle_act_and_wait and reticle_assert produce a verdict. reticle_act, snapshot, query, navigate, network and console move or read the app and prove nothing. A drive that ends without one of the first two has no result, however many calls it made.

3. Read the verdict honestly

verifiedmeansdo
yesthe consequence you named happenedreport it, with the evidence
noit did not happen, or a channel contradicted the UIa real finding: report it with because
unknownReticle drove the app and could not tellnot a pass. Say unknown and say why

On unknown / unsettled, re-assert rather than re-driving: reticle_assert({ predicate, since, timeout_ms: 8000 }) using the since from the act result. Re-driving repeats a side effect that already happened.

Never weaken a check to turn a verdict green. An assertion edited until it passes proves nothing.

4. Check what you did not touch

reticle_run({ tool: "reticle_verify", sessionId, args: { action: "coverage" } })   // { total, exercised, untouched }

If untouched still holds controls your change affects, the drive is unfinished. One call, and it is the cheapest guard against reporting a pass over the half you never opened.

5. Report

State what you drove, what the verdict was, and the evidence: the request and status, the state that changed, the app's own signal. If something failed, reticle_look({ action: "element", sessionId, ref }) on the failing element gives the file:line: put it in the report.

Then reticle_session({ action: "yield", mode: "waiting" }) so the human's panel stops reading "live".


More detail, fetchable one page at a time: curl https://docs.reticle.sh/llms.txt for the index, then the single page you need (tools/act-and-wait.md, predicates.md, troubleshooting.md). If Reticle itself misbehaves, file it with reticle_session { action: "feedback" }: one call, then carry on.

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!