SkillAtlasSkill 详情

endpoint-validator

Essays and writing behind this toolkit live at vexjoy.com.

审核状态:已审核Quality 72Security 70

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月31日

VexJoy Agent

VexJoy Agent

Essays and writing behind this toolkit live at vexjoy.com.

AI agents skip steps.

"Looks correct" replaces running tests. "Trivial change" replaces verification. The agent confidently ships broken code because nothing structurally prevented it from skipping the work.

Harnesses have a second problem: given only a skill list, they do not route eagerly enough, or correctly enough. Good skills sit unused. So this toolkit connects the skills, agents, and workflows we want directly into the harness, automatically. You don't have to understand what is here. Say what you want in plain English and you get all the value we have put into it: the right specialist with the right methodology, behind gates that demand exit codes, not assertions.

44 domain agents, 122 workflow skills, 78 hooks, 136 scripts. Agents carry knowledge, skills enforce methodology, hooks block incomplete work, scripts handle determinism.

Works across Claude Code (/do), Codex ($do), Factory (/do), Reasonix (/do).

What It Looks Like

$ claude

> /do debug this Go test

  Routing: go-engineer + systematic-debugging
  Phase 1/4: Reproduce: running test, capturing failure...
  Phase 2/4: Hypothesize: 3 candidates from stack trace...
  Phase 3/4: Verify: isolated root cause in connection pool timeout
  Phase 4/4: Fix: patch applied, test passing, PR opened

  ✓ Delivered: PR #847, fix connection pool timeout in health check

The router reads intent, picks a Go agent paired with a debugging skill, and runs the full lifecycle. You typed one sentence. The system did the rest.

The Pipeline

  ROUTE        PLAN         EXECUTE      VERIFY       DELIVER      RECORD
 ┌──────┐    ┌──────┐    ┌──────┐    ┌──────┐    ┌──────┐    ┌──────┐
 │ /do  │───▶│ Task │───▶│Agent │───▶│Tests │───▶│  PR  │───▶│Route │
 │Router│    │ Plan │    │+Skill│    │Gates │    │Branch│    │Result│
 └──────┘    └──────┘    └──────┘    └──────┘    └──────┘    └──────┘

Anti-Rationalization

This is the single thing that separates it from "agent with a system prompt."

Agent SaysWhat Happens
"Code looks correct, skip tests"Exit gate requires test output. Blocked.
"Trivial change, no verification"Hook blocks completion without evidence.
"Similar to before"Skill demands case-specific proof.
"User is in a hurry"Protocol overrides time pressure.
"I'm confident"Gate demands exit code, not assertion.

Hooks fire automatically. Gates block completion. Skills encode counter-arguments at every skip-worthy step. The agent verifies or it doesn't finish.

For what I do, the difference is enormous. If you're doing simple single-file edits, maybe less so.

Knowledge Work Is First-Class

The same routing serves knowledge work. The content engine researches, drafts in a calibrated voice, validates against 397 AI patterns, and repurposes finished pieces for each platform. /html turns any request into a single self-contained HTML file: report, slide deck, prototype, data viz, diagram. Non-engineers who try the toolkit consistently name the HTML artifacts as the thing they love. No code, no setup beyond the installer.

It Proves Its Own Changes

Changes to the toolkit itself ship with evidence. New skills get blind A/B tests against a no-skill baseline before merge. Routing and writing-standard decisions carry measured verdicts; PHILOSOPHY.md cites the numbers. Experiments that lost go into the negative-results registry, what-didnt-work.md; the registry now covers routing reversals, unvalidated A/B citations, and disabled lint rules alongside the original program refutations.

The automated nightly evolution loop (/evolve, writes to evolution-reports/) ran regularly through mid-May 2026. It is currently dormant; recent evidence has come from manual PRs instead.

Installation

git clone https://github.com/notque/vexjoy-agent.git ~/vexjoy-agent
cd ~/vexjoy-agent
./install.sh

Links into ~/.claude/ and mirrors into ~/.codex/, ~/.factory/, ~/.reasonix/ — each mirror only when that runtime is detected (its command on PATH or its home dir already exists). The installer asks symlink (live updates via git pull) or copy (stable snapshot).

Want only part of the toolkit? Run ./install.sh --configure to pick which skills, agents, and hooks install, or copy .local.example/profile.yaml to .local/profile.yaml and edit. No profile file = full install, unchanged behavior. Credit: @thomasvan. Details: .local.example/README.md.

CLIEntry Point
Claude Code/do
Codex$do
Factory/do
Reasonix/do

Full setup: docs/start-here.md

Codex CLI Parity

Mirrors agents, skills, and supported hooks into ~/.codex/. The original six-hook allowlist was correct for Codex v0.114, when tool hooks only intercepted Bash. Current support requires Codex v0.144.1+ and classifies the 74 Claude hook registrations as 26 native, 35 adapter-backed, and 13 unsupported (61 supported). These are registration counts, not unique hook files. The installer also preserves explicit per-subagent model routing for GPT-5.6 Sol by setting the MultiAgent V2 compatibility keys documented in openai/codex#31814.

Codex now exposes apply_patch to tool hooks. VexJoy's adapter converts each patch operation into the Write/Edit payload expected by existing guards, but it cannot intercept writes performed through unified_exec, unmatched MCP tools, WebSearch, or other unsupported tool paths. PreCompact and Stop adapters also receive less telemetry than Claude Code: Codex does not provide Claude's conversation_history or session_data. This is expanded compatibility, not full Claude parity.

After install or any hook-definition change, run /hooks in Codex and review the new definitions before trusting them. Codex hash-trusts hook commands and skips changed, unreviewed definitions.

Gemini CLI / Antigravity CLI Support (removed)

Gemini CLI support removed (deprecated upstream, transitioned to Antigravity CLI); Antigravity support pending CLI maturity. Per Google's transition announcement, Gemini CLI stops serving requests on 2026-06-18 for Google AI Pro / Ultra and free Gemini Code Assist for individuals. Gemini API integrations (image-gen backends, sprite pipeline, GEMINI_API_KEY) are unaffected and stay in the toolkit.

If a prior install mirrored into ~/.gemini/, remove the stale mirrors with:

rm -rf ~/.gemini/skills ~/.gemini/agents ~/.gemini/hooks ~/.gemini/scripts ~/.gemini/antigravity/plugins/vexjoy-agent
Factory CLI Support

Mirrors agents (as "droids"), skills, and all hooks into ~/.factory/. Hook config merges into ~/.factory/settings.json with paths rewritten.

Reasonix Support

Mirrors skills, scripts, and the allowlisted hooks (scripts/reasonix-hooks-allowlist.txt) into ~/.reasonix/ (no agent or custom-command surface, so neither is installed; the /do router rides in as a skill). Reasonix fires only 4 events (PreToolUse, PostToolUse, UserPromptSubmit, Stop), so only hooks for those events are allowlisted. Hook config is written to the hooks key of ~/.reasonix/settings.json in Reasonix's native flat shape (one entry per hook, match regex over the tool name); the generator builds absolute python3 commands, so no path rewrite is applied. MCP/model/permissions in ~/.reasonix/config.json are user-owned and left untouched.

Token-saving mode

The toolkit supplies its own routing, domain knowledge, methodology, and enforcement. The default system prompt duplicates most of that.

claude --system-prompt "."

Strips built-in tool-use instructions. The toolkit's agents, skills, hooks, and CLAUDE.md provide equivalent coverage.

Four Layers

LayerCountDoes
Agents44Domain knowledge: idiom tables, failure mode catalogs, error-to-fix mappings
Skills122Phased methodology with gates. Can't skip steps. Each phase has exit criteria requiring evidence.
Hooks78Fire on lifecycle events. Block incomplete work. Zero LLM cost.
Scripts136Determinism: test runners, linters, validators. No LLM judgment.

Full skill catalog: docs/skills.md.

┌─────────────────────────────────────────────────┐
│  SKILL.md                                       │
│  ┌─ Frontmatter ─────────────────────────────┐  │
│  │ triggers, pairs_with, success-criteria     │  │
│  └────────────────────────────────────────────┘  │
│  Reference Loading Table (conditional imports)   │
│  Phased Instructions (numbered, with gates)      │
│  Verification (evidence requirements)            │
└─────────────────────────────────────────────────┘

Built with the Toolkit

A game built entirely by Claude Code using these agents, skills, and pipelines:

Choose Your Path

I just want to use it Install, learn /do, done.

I do knowledge work Writing, research, data analysis, moderation, HTML artifacts. No code.

I'm a developer Architecture, extension points, adding agents and skills.

I'm an AI power user Routing tables, pipelines, hooks, telemetry DB.

I'm an AI agent Machine-dense inventory. Tables, paths, schemas.

I'm on LinkedIn 🚀 Thought leadership. Agree? 👇

Philosophy

  • Zero-expertise operation. Say what you want. The system classifies, dispatches, enforces, delivers.
  • LLMs orchestrate, programs execute. Deterministic work belongs to scripts. LLM judgment handles design decisions, diagnosis, review.
  • Density. Every word carries instruction, rule, or decision. Cut everything else.
  • Breadth over depth. Right context ensures correctness. Unfocused context adds cost.
  • Structural enforcement. Exit codes enforce what instructions can't. Quality gates are automated, not advisory.
  • Everything pipelines. Complex work decomposes into phases. Phases have gates. Gates prevent cascading failures.

Full design philosophy: PHILOSOPHY.md

Maintenance

One report-only script surfaces upkeep work; it prints a digest and never edits, deletes, or blocks.

  • python3 scripts/stale-skill-scan.py --top 20 ranks stale skills and agents as pruning candidates. Run it quarterly; see docs/deprecation-template.md.

Scheduled work follows the same boundary as everything else: judgment uses agents; repeatable plumbing uses scripts.

NeedUse
Run a deterministic command on a schedulescripts/agent-scheduler.py with runner: "command"
Run an agent judgment on a schedule, webhook, or file changescripts/agent-scheduler.py with the default runner: "claude"
Install or remove a user crontab entry safelyscripts/crontab-manager.py
Audit shell cron reliabilitycron-automation
Keep one interactive objective moving until criteria verifyobjective-loop

Contributing

See CONTRIBUTING.md.

License

MIT. See LICENSE.

其他

中风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 可能需要外部 token、网络权限或第三方服务。
  • 未检测到高风险命令。
  • 扫描发现:3 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/infrastructure/endpoint-validator" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/infrastructure/endpoint-validator" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/infrastructure/endpoint-validator" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/infrastructure/endpoint-validator" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/infrastructure/endpoint-validator" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: endpoint-validator
promoted_to: service-health-check
description: "Deterministic API endpoint validation with pass/fail reporting."
user-invocable: false
allowed-tools:
  - Bash
  - Read
  - Write
  - Glob
  - Edit
routing:
  triggers:
    - "validate endpoints"
    - "smoke test API"
    - "health check endpoints"
    - "test endpoint"
    - "check API"
    - "smoke test"
  category: infrastructure
  not_for: "process/service uptime or daemon liveness (use service-health-check); only HTTP/API endpoint request validation"
  pairs_with:
    - service-health-check
    - e2e-testing

Endpoint Validator Skill

Deterministic HTTP endpoint validation following a Discover, Validate, Report pattern. Finds endpoints, tests each against expectations, and produces machine-readable results with clear pass/fail verdicts and CI-compatible exit codes.

Reference Loading Table

SignalLoad These FilesWhy
Security header WARNs, HSTS/CSP/X-Frame issuessecurity-headers.mdRoutes to the matching deep reference
Config errors, hardcoded IPs, timeout problemsendpoint-config-preferred-patterns.mdRoutes to the matching deep reference
401/403 failures, Bearer/API-key/cookie authauth-endpoint-patterns.mdRoutes to the matching deep reference

Instructions

Phase 1: DISCOVER

Goal: Locate or receive endpoint definitions before making any requests.

Step 1: Read repository CLAUDE.md

Check for and follow any repository-level CLAUDE.md before running validation. It may contain base URL conventions, environment variable names, or endpoint paths relevant to the project.

Step 2: Search for endpoint configuration

Look for definitions in priority order:

  1. endpoints.json in project root
  2. tests/endpoints.json
  3. Inline specification provided by user or calling agent

Prefer config files checked into version control over ad-hoc endpoint lists. Manually listing endpoints every run leads to drift and missed endpoints.

Step 3: Parse and validate configuration

Configuration must contain base_url and at least one endpoint:

{
  "base_url": "http://localhost:8000",
  "endpoints": [
    {"path": "/health", "expect_status": 200},
    {"path": "/api/v1/users", "expect_key": "data", "timeout": 10},
    {"path": "/api/v1/search?q=test", "max_time": 2.0}
  ]
}

Each endpoint supports these fields:

  • path (required): URL path appended to base_url
  • expect_status (default: 200): Expected HTTP status code
  • expect_key (optional): Top-level JSON key that must exist in response. Only top-level key presence is checked -- full JSON schema validation is out of scope.
  • timeout (default: 5): Request timeout in seconds. The 5-second default prevents hanging on unresponsive endpoints.
  • max_time (optional): Fail if response exceeds this threshold in seconds
  • method (optional): HTTP method. Defaults to GET. POST/PUT/DELETE require explicit configuration with a request body -- send mutating requests only when the user explicitly configures them.
  • headers (optional): Additional headers per endpoint (e.g., Accept, Content-Type, Authorization)

If base_url points to a production host and the config includes POST/PUT/DELETE endpoints, warn the user before proceeding. Mutating production data or triggering rate limits during a smoke test is a serious risk. Use staging environments for write operations; reserve production for GET-only health checks.

Use hostnames or environment variables instead of hardcoded IP addresses in base_url (e.g., http://192.168.1.42:8000). They break on every other machine and CI environment. Use localhost with a configurable port or environment variables instead.

Step 4: Confirm base URL is reachable

Make a single request to base_url before running the full suite. If unreachable, report immediately rather than failing every endpoint individually.

Gate: Configuration parsed, base URL reachable, at least one endpoint defined. Proceed only when gate passes.

Phase 2: VALIDATE

Verification means execution, not reasoning. Run the command. Do not reason about whether the command would pass. Do not summarize the expected output. Execute the check, paste the exit code, paste the relevant output. A verification phase that produces a verdict without an observed tool result is not a verification — it is a guess with a rigor aesthetic.

Goal: Test each endpoint against its expected criteria and collect structured results.

Step 1: Execute requests sequentially

Test endpoints one at a time for predictable, reproducible output. For each endpoint:

  1. Construct full URL from base_url + path
  2. Send request with configured method (GET by default) and timeout
  3. Record status code, response time, and body
  4. Display each result as it completes so the user sees progress

This skill sends one request per endpoint. It is not a load tester or stress tester -- it validates contract compliance, not throughput.

Step 2: Evaluate against expectations

For each response, check in order:

  1. Status code: Does it match expect_status? If not, mark FAIL.
  2. JSON key: If expect_key set, parse JSON and check key exists. If missing or not valid JSON, mark FAIL.
  3. Response time: If max_time set and elapsed exceeds it, mark SLOW. Flag slow endpoints -- they indicate degradation that becomes failure under load.
  4. Security headers: Check response headers for common security headers. Report missing headers as WARN (not FAIL):
    • Strict-Transport-Security -- HSTS enforcement (expected on HTTPS endpoints)
    • Content-Security-Policy -- XSS mitigation
    • X-Content-Type-Options -- should be nosniff
    • X-Frame-Options -- clickjacking prevention (or CSP frame-ancestors)

Skip security header checks for localhost/127.0.0.1 endpoints (development environments typically omit these). Only check on non-localhost base URLs unless explicitly configured.

Step 3: Handle failures gracefully

  • Connection refused: Record as FAIL with "Connection refused" error
  • Timeout exceeded: Record as FAIL with "Timeout after Ns" error
  • Invalid JSON when expect_key set: Record as FAIL with "Invalid JSON response"
  • Unexpected exception: Record as FAIL with exception message

Gate: All endpoints tested. Every result has a clear PASS, FAIL, or SLOW verdict. Proceed only when gate passes.

Phase 3: REPORT

Goal: Produce structured, machine-readable output with summary statistics.

Step 1: Format individual results

ENDPOINT VALIDATION REPORT
==========================
Base URL: http://localhost:8000
Endpoints: 15 tested

RESULTS:
  /api/health                    200 OK      45ms
  /api/users                     200 OK     123ms
  /api/products                  500 FAIL   "Internal Server Error"
  /api/slow                      200 SLOW   3.2s > 2.0s threshold

SECURITY HEADERS (non-localhost only):
  /api/health                    WARN  Missing: Content-Security-Policy, X-Frame-Options
  /api/users                     OK    All security headers present
  /api/products                  SKIP  (endpoint failed)

Step 2: Produce summary

SUMMARY:
  Passed: 13/15 (86.7%)
  Failed: 1 (status error)
  Slow: 1 (exceeded threshold)
  Security header warnings: 3 endpoints missing headers

Step 3: Set exit code

  • Exit 0 if all endpoints passed (SLOW counts as pass unless max_time was set)
  • Exit 1 if any endpoint failed

Gate: Report printed, exit code set. Validation complete.

Examples

Example 1: Pre-Deployment Health Check

User says: "Validate all endpoints before we deploy" Actions:

  1. Find endpoints.json in project root (DISCOVER)
  2. Test each endpoint, collect status codes and times (VALIDATE)
  3. Print report, exit 0 if all pass (REPORT) Result: Structured pass/fail report with CI-compatible exit code

Example 2: Smoke Test After Migration

User says: "Check if the API is still working after the database migration" Actions:

  1. Read endpoint config, confirm base URL reachable (DISCOVER)
  2. Hit each endpoint, check status and expected keys (VALIDATE)
  3. Surface any failures with error details (REPORT) Result: Quick verification that migration did not break API contracts

Error Handling

Error: "Base URL Unreachable"

Cause: Service not running, wrong port, or network issue Solution:

  1. Verify service is running (ps aux, docker ps, or equivalent)
  2. Confirm port matches config (netstat -tlnp or ss -tlnp)
  3. Check for firewall rules or container networking issues

Error: "All Endpoints Timeout"

Cause: Service overwhelmed, wrong host, or proxy misconfiguration Solution:

  1. Test a single endpoint manually with curl -v
  2. Increase timeout values in config if service is legitimately slow
  3. Check if a reverse proxy or load balancer is intercepting requests

Error: "JSON Parse Failure on expect_key Check"

Cause: Endpoint returns HTML, XML, or empty body instead of JSON Solution:

  1. Verify endpoint actually returns JSON (check Content-Type header)
  2. Remove expect_key if endpoint legitimately returns non-JSON
  3. Check if authentication is required (HTML login page returned)

Reference Loading

Task TypeLoad This Reference
Security header WARNs, HSTS/CSP/X-Frame issuesreferences/security-headers.md
Config errors, hardcoded IPs, timeout problemsreferences/endpoint-config-preferred-patterns.md
401/403 failures, Bearer/API-key/cookie authreferences/auth-endpoint-patterns.md

References

CI/CD Integration

# GitHub Actions example
# TODO: scripts/validate_endpoints.py not yet implemented
# Manual alternative: use curl to validate endpoints from endpoints.json
- name: Validate API endpoints
  run: |
    jq -r '.endpoints[].path' endpoints.json | while read path; do
      curl -sf "$BASE_URL$path" > /dev/null && echo "PASS: $path" || echo "FAIL: $path"
    done
# Pre-deployment gate
# TODO: scripts/validate_endpoints.py not yet implemented
# Manual alternative: iterate endpoints.json with curl
jq -r '.endpoints[].path' endpoints.json | while read path; do
  curl -sf "http://localhost:8000$path" > /dev/null || { echo "FAIL: $path"; exit 1; }
done

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!