SkillAtlasSkill 详情

agent-comparison

Essays and writing behind this toolkit live at vexjoy.com.

审核状态:已审核Quality 72Security 70

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月31日

VexJoy Agent

VexJoy Agent

Essays and writing behind this toolkit live at vexjoy.com.

AI agents skip steps.

"Looks correct" replaces running tests. "Trivial change" replaces verification. The agent confidently ships broken code because nothing structurally prevented it from skipping the work.

Harnesses have a second problem: given only a skill list, they do not route eagerly enough, or correctly enough. Good skills sit unused. So this toolkit connects the skills, agents, and workflows we want directly into the harness, automatically. You don't have to understand what is here. Say what you want in plain English and you get all the value we have put into it: the right specialist with the right methodology, behind gates that demand exit codes, not assertions.

44 domain agents, 122 workflow skills, 78 hooks, 136 scripts. Agents carry knowledge, skills enforce methodology, hooks block incomplete work, scripts handle determinism.

Works across Claude Code (/do), Codex ($do), Factory (/do), Reasonix (/do).

What It Looks Like

$ claude

> /do debug this Go test

  Routing: go-engineer + systematic-debugging
  Phase 1/4: Reproduce: running test, capturing failure...
  Phase 2/4: Hypothesize: 3 candidates from stack trace...
  Phase 3/4: Verify: isolated root cause in connection pool timeout
  Phase 4/4: Fix: patch applied, test passing, PR opened

  ✓ Delivered: PR #847, fix connection pool timeout in health check

The router reads intent, picks a Go agent paired with a debugging skill, and runs the full lifecycle. You typed one sentence. The system did the rest.

The Pipeline

  ROUTE        PLAN         EXECUTE      VERIFY       DELIVER      RECORD
 ┌──────┐    ┌──────┐    ┌──────┐    ┌──────┐    ┌──────┐    ┌──────┐
 │ /do  │───▶│ Task │───▶│Agent │───▶│Tests │───▶│  PR  │───▶│Route │
 │Router│    │ Plan │    │+Skill│    │Gates │    │Branch│    │Result│
 └──────┘    └──────┘    └──────┘    └──────┘    └──────┘    └──────┘

Anti-Rationalization

This is the single thing that separates it from "agent with a system prompt."

Agent SaysWhat Happens
"Code looks correct, skip tests"Exit gate requires test output. Blocked.
"Trivial change, no verification"Hook blocks completion without evidence.
"Similar to before"Skill demands case-specific proof.
"User is in a hurry"Protocol overrides time pressure.
"I'm confident"Gate demands exit code, not assertion.

Hooks fire automatically. Gates block completion. Skills encode counter-arguments at every skip-worthy step. The agent verifies or it doesn't finish.

For what I do, the difference is enormous. If you're doing simple single-file edits, maybe less so.

Knowledge Work Is First-Class

The same routing serves knowledge work. The content engine researches, drafts in a calibrated voice, validates against 397 AI patterns, and repurposes finished pieces for each platform. /html turns any request into a single self-contained HTML file: report, slide deck, prototype, data viz, diagram. Non-engineers who try the toolkit consistently name the HTML artifacts as the thing they love. No code, no setup beyond the installer.

It Proves Its Own Changes

Changes to the toolkit itself ship with evidence. New skills get blind A/B tests against a no-skill baseline before merge. Routing and writing-standard decisions carry measured verdicts; PHILOSOPHY.md cites the numbers. Experiments that lost go into the negative-results registry, what-didnt-work.md; the registry now covers routing reversals, unvalidated A/B citations, and disabled lint rules alongside the original program refutations.

The automated nightly evolution loop (/evolve, writes to evolution-reports/) ran regularly through mid-May 2026. It is currently dormant; recent evidence has come from manual PRs instead.

Installation

git clone https://github.com/notque/vexjoy-agent.git ~/vexjoy-agent
cd ~/vexjoy-agent
./install.sh

Links into ~/.claude/ and mirrors into ~/.codex/, ~/.factory/, ~/.reasonix/ — each mirror only when that runtime is detected (its command on PATH or its home dir already exists). The installer asks symlink (live updates via git pull) or copy (stable snapshot).

Want only part of the toolkit? Run ./install.sh --configure to pick which skills, agents, and hooks install, or copy .local.example/profile.yaml to .local/profile.yaml and edit. No profile file = full install, unchanged behavior. Credit: @thomasvan. Details: .local.example/README.md.

CLIEntry Point
Claude Code/do
Codex$do
Factory/do
Reasonix/do

Full setup: docs/start-here.md

Codex CLI Parity

Mirrors agents, skills, and supported hooks into ~/.codex/. The original six-hook allowlist was correct for Codex v0.114, when tool hooks only intercepted Bash. Current support requires Codex v0.144.1+ and classifies the 74 Claude hook registrations as 26 native, 35 adapter-backed, and 13 unsupported (61 supported). These are registration counts, not unique hook files. The installer also preserves explicit per-subagent model routing for GPT-5.6 Sol by setting the MultiAgent V2 compatibility keys documented in openai/codex#31814.

Codex now exposes apply_patch to tool hooks. VexJoy's adapter converts each patch operation into the Write/Edit payload expected by existing guards, but it cannot intercept writes performed through unified_exec, unmatched MCP tools, WebSearch, or other unsupported tool paths. PreCompact and Stop adapters also receive less telemetry than Claude Code: Codex does not provide Claude's conversation_history or session_data. This is expanded compatibility, not full Claude parity.

After install or any hook-definition change, run /hooks in Codex and review the new definitions before trusting them. Codex hash-trusts hook commands and skips changed, unreviewed definitions.

Gemini CLI / Antigravity CLI Support (removed)

Gemini CLI support removed (deprecated upstream, transitioned to Antigravity CLI); Antigravity support pending CLI maturity. Per Google's transition announcement, Gemini CLI stops serving requests on 2026-06-18 for Google AI Pro / Ultra and free Gemini Code Assist for individuals. Gemini API integrations (image-gen backends, sprite pipeline, GEMINI_API_KEY) are unaffected and stay in the toolkit.

If a prior install mirrored into ~/.gemini/, remove the stale mirrors with:

rm -rf ~/.gemini/skills ~/.gemini/agents ~/.gemini/hooks ~/.gemini/scripts ~/.gemini/antigravity/plugins/vexjoy-agent
Factory CLI Support

Mirrors agents (as "droids"), skills, and all hooks into ~/.factory/. Hook config merges into ~/.factory/settings.json with paths rewritten.

Reasonix Support

Mirrors skills, scripts, and the allowlisted hooks (scripts/reasonix-hooks-allowlist.txt) into ~/.reasonix/ (no agent or custom-command surface, so neither is installed; the /do router rides in as a skill). Reasonix fires only 4 events (PreToolUse, PostToolUse, UserPromptSubmit, Stop), so only hooks for those events are allowlisted. Hook config is written to the hooks key of ~/.reasonix/settings.json in Reasonix's native flat shape (one entry per hook, match regex over the tool name); the generator builds absolute python3 commands, so no path rewrite is applied. MCP/model/permissions in ~/.reasonix/config.json are user-owned and left untouched.

Token-saving mode

The toolkit supplies its own routing, domain knowledge, methodology, and enforcement. The default system prompt duplicates most of that.

claude --system-prompt "."

Strips built-in tool-use instructions. The toolkit's agents, skills, hooks, and CLAUDE.md provide equivalent coverage.

Four Layers

LayerCountDoes
Agents44Domain knowledge: idiom tables, failure mode catalogs, error-to-fix mappings
Skills122Phased methodology with gates. Can't skip steps. Each phase has exit criteria requiring evidence.
Hooks78Fire on lifecycle events. Block incomplete work. Zero LLM cost.
Scripts136Determinism: test runners, linters, validators. No LLM judgment.

Full skill catalog: docs/skills.md.

┌─────────────────────────────────────────────────┐
│  SKILL.md                                       │
│  ┌─ Frontmatter ─────────────────────────────┐  │
│  │ triggers, pairs_with, success-criteria     │  │
│  └────────────────────────────────────────────┘  │
│  Reference Loading Table (conditional imports)   │
│  Phased Instructions (numbered, with gates)      │
│  Verification (evidence requirements)            │
└─────────────────────────────────────────────────┘

Built with the Toolkit

A game built entirely by Claude Code using these agents, skills, and pipelines:

Choose Your Path

I just want to use it Install, learn /do, done.

I do knowledge work Writing, research, data analysis, moderation, HTML artifacts. No code.

I'm a developer Architecture, extension points, adding agents and skills.

I'm an AI power user Routing tables, pipelines, hooks, telemetry DB.

I'm an AI agent Machine-dense inventory. Tables, paths, schemas.

I'm on LinkedIn 🚀 Thought leadership. Agree? 👇

Philosophy

  • Zero-expertise operation. Say what you want. The system classifies, dispatches, enforces, delivers.
  • LLMs orchestrate, programs execute. Deterministic work belongs to scripts. LLM judgment handles design decisions, diagnosis, review.
  • Density. Every word carries instruction, rule, or decision. Cut everything else.
  • Breadth over depth. Right context ensures correctness. Unfocused context adds cost.
  • Structural enforcement. Exit codes enforce what instructions can't. Quality gates are automated, not advisory.
  • Everything pipelines. Complex work decomposes into phases. Phases have gates. Gates prevent cascading failures.

Full design philosophy: PHILOSOPHY.md

Maintenance

One report-only script surfaces upkeep work; it prints a digest and never edits, deletes, or blocks.

  • python3 scripts/stale-skill-scan.py --top 20 ranks stale skills and agents as pruning candidates. Run it quarterly; see docs/deprecation-template.md.

Scheduled work follows the same boundary as everything else: judgment uses agents; repeatable plumbing uses scripts.

NeedUse
Run a deterministic command on a schedulescripts/agent-scheduler.py with runner: "command"
Run an agent judgment on a schedule, webhook, or file changescripts/agent-scheduler.py with the default runner: "claude"
Install or remove a user crontab entry safelyscripts/crontab-manager.py
Audit shell cron reliabilitycron-automation
Keep one interactive objective moving until criteria verifyobjective-loop

Contributing

See CONTRIBUTING.md.

License

MIT. See LICENSE.

Agent / MCP / Skill 创作测试与质量

中风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 可能需要外部 token、网络权限或第三方服务。
  • 未检测到高风险命令。
  • 扫描发现:2 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/meta/agent-comparison" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/meta/agent-comparison" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/meta/agent-comparison" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/meta/agent-comparison" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/meta/agent-comparison" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: agent-comparison
description: "A/B test agent variants for quality and token cost."
user-invocable: false
allowed-tools:
  - Read
  - Write
  - Edit
  - Bash
  - Glob
  - Grep
  - Task
routing:
  triggers:
    - "compare agents"
    - "A/B test agents"
    - "benchmark agents"
    - "optimize skill"
    - "optimize description"
    - "run autoresearch"
  not_for: "creating new skills from scratch (use skill-creator skill) — this skill compares and optimizes existing agent descriptions and routing"
  category: meta-tooling
  pairs_with:
    - agent-evaluation
    - skill-eval

Agent Comparison Skill

Compare agent variants through controlled A/B benchmarks. Runs identical tasks on both agents, grades output quality with domain-specific checklists, and reports total session token cost to a working solution. This skill is exclusively for agent variant comparison — use agent-evaluation for single-agent assessment, and skill-eval for skill testing.

Reference Loading Table

SignalLoad These FilesWhy
selecting benchmark tasks and directory layout (Phase 1)benchmark-tasks.mdLoads detailed guidance from benchmark-tasks.md.
example-driven tasks, errorsexamples-and-errors.mdLoads detailed guidance from examples-and-errors.md.
scoring solutions: 5-criteria rubric and effective cost calculationgrading-rubric.mdLoads detailed guidance from grading-rubric.md.
deciding when to run comparisons; December 2024 baseline datamethodology.mdLoads detailed guidance from methodology.md.
configuring autoresearch: targets, task formats, eval isolation modesoptimization-guide.mdLoads detailed guidance from optimization-guide.md.
executing Phase 5 OPTIMIZE step by stepoptimize-phase.mdLoads detailed guidance from optimize-phase.md.
writing the Phase 4 comparison reportreport-template.mdLoads detailed guidance from report-template.md.

Instructions

See references/examples-and-errors.md for error handling. See references/optimize-phase.md for Phase 5 OPTIMIZE full procedure. See references/methodology.md for December 2024 benchmark data.

Phase 1: PREPARE

Goal: Create benchmark environment and validate both agent variants exist.

Read and follow the repository CLAUDE.md before starting any execution.

Step 1: Analyze original agent

wc -l agents/{original-agent}.md
grep "^## " agents/{original-agent}.md
grep -c '```' agents/{original-agent}.md

Step 2: Create or validate compact variant

If creating a compact variant, preserve:

  • YAML frontmatter (name, description, routing)
  • Core patterns and principles
  • Error handling philosophy

Remove or condense:

  • Lengthy code examples (keep 1-2 representative per pattern)
  • Verbose explanations (condense to bullet points)
  • Redundant instructions and changelogs

Target 10-15% of original size while keeping essential knowledge. Remove redundancy, not capability — stripping error handling patterns or concurrency guidance creates an unfair comparison because the compact agent is missing essential knowledge rather than expressing it concisely.

Step 3: Validate compact variant structure

head -20 agents/{compact-agent}.md | grep -E "^(name|description):"
echo "Original: $(wc -l < agents/{original-agent}.md) lines"
echo "Compact:  $(wc -l < agents/{compact-agent}.md) lines"

Step 4: Create benchmark directory and prepare prompts

mkdir -p benchmark/{task-name}/{full,compact}

Write the task prompt ONCE, then copy it for both agents. Both agents must receive the exact same task description, character-for-character, because different requirements produce different solutions and invalidate all measurements.

Keep benchmark scripts simple — no speculative features or configurable frameworks that were not requested.

Gate: Both agent variants exist with valid YAML frontmatter. Benchmark directories created. Identical task prompts written. Proceed only when gate passes.

Phase 2: BENCHMARK

Goal: Run identical tasks on both agents, capturing all metrics.

Step 1: Run simple task benchmark (2-3 tasks)

Use algorithmic problems with clear specifications (e.g., Advent of Code Day 1-6). Simple tasks establish a baseline — if an agent fails here, it has fundamental issues. Running multiple simple tasks is necessary because a single data point is sensitive to task selection bias and cannot distinguish luck from systematic quality.

Spawn both agents in parallel using Task tool:

Task(
  prompt="[exact task prompt]\nSave to: benchmark/{task}/full/",
  subagent_type="{full-agent}"
)

Task(
  prompt="[exact task prompt]\nSave to: benchmark/{task}/compact/",
  subagent_type="{compact-agent}"
)

Run in parallel to avoid caching effects or system load variance skewing results.

Step 2: Run complex task benchmark (1-2 tasks)

Use production-style problems that require concurrency, error handling, edge case anticipation — these are where quality differences emerge because simple tasks mask differences in edge case handling. See references/benchmark-tasks.md for standard tasks.

Recommended complex tasks:

  • Worker Pool: Rate limiting, graceful shutdown, panic recovery
  • LRU Cache with TTL: Generics, background goroutines, zero-value semantics
  • HTTP Service: Middleware chains, structured errors, health checks

Step 3: Capture metrics for each run

Record immediately after each agent completes — delayed recording loses precision. Track input/output token counts per turn where visible, since total session cost (not just prompt size) is what matters.

MetricFull AgentCompact Agent
Tests passX/XX/X
Race conditionsXX
Code lines (main)XX
Test linesXX
Session tokensXX
Wall-clock timeXm XsXm Xs
Retry cyclesXX

Step 4: Run tests with race detector

cd benchmark/{task-name}/full && go test -race -v -count=1
cd benchmark/{task-name}/compact && go test -race -v -count=1

Use -count=1 to disable test caching. All generated code must pass the same test suite with the -race flag because race conditions are automatic quality failures.

Gate: Both agents completed all tasks. Metrics captured for every run. Test output saved. Proceed only when gate passes.

Phase 3: GRADE

Goal: Score code quality beyond pass/fail using domain-specific checklists.

Step 1: Create quality checklist BEFORE reviewing code

Define criteria before seeing results to prevent bias — inventing criteria after seeing one agent's output skews the comparison. See references/grading-rubric.md for standard rubrics.

Criterion5/53/51/5
CorrectnessAll tests pass, no race conditionsSome failuresBroken
Error HandlingComprehensive, production-readyAdequateNone
IdiomsExemplary for the languageAcceptableFailure modes
DocumentationThoroughAdequateNone
TestingComprehensive coverageBasicMinimal

Step 2: Score each solution independently

Grade each agent's code on all five criteria. Score one agent completely before starting the other. Report facts and show command output rather than describing it — every claim must be backed by measurable data (tokens, test counts, quality scores).

## {Agent} Solution - {Task}

| Criterion | Score | Notes |
|-----------|-------|-------|
| Correctness | X/5 | |
| Error Handling | X/5 | |
| Idioms | X/5 | |
| Documentation | X/5 | |
| Testing | X/5 | |
| **Total** | **X/25** | |

Step 3: Document specific bugs with production impact

For each bug found, record:

### Bug: {description}
- Agent: {which agent}
- What happened: {behavior}
- Correct behavior: {expected}
- Production impact: {consequence}
- Test coverage: {did tests catch it? why not?}

"Tests pass" is necessary but not sufficient — production bugs often pass tests. Apply the domain-specific quality checklist rather than relying only on test pass rates, because tests can miss goroutine leaks, wrong semantics, and other production issues.

Step 4: Calculate effective cost

effective_cost = total_tokens * (1 + bug_count * 0.25)

An agent using 194k tokens with 0 bugs has better economics than one using 119k tokens with 5 bugs requiring fixes. The metric that matters is total cost to working, production-quality solution — not prompt size, because prompt is a one-time cost while reasoning tokens dominate sessions. Check quality scores before claiming token savings, since savings that come from cutting corners are not real savings.

Gate: Both solutions graded with evidence. Specific bugs documented with production impact. Effective cost calculated. Proceed only when gate passes.

Phase 4: REPORT

Goal: Generate comparison report with evidence-backed verdict.

Step 1: Generate comparison report

Use the report template from references/report-template.md. Include:

  • Executive summary with clear winner per metric
  • Per-task results with metrics tables
  • Token economics analysis (one-time prompt cost vs session cost)
  • Specific bugs found and their production impact
  • Verdict based on total evidence

Step 2: Run comparison analysis

python3 ${CLAUDE_SKILL_DIR}/scripts/compare.py benchmark/{task-name}/

Step 3: Analyze token economics

The key economic insight: agent prompts are a one-time cost per session. Everything after — reasoning, code generation, debugging, retries — costs tokens on every turn. When a micro agent produces correct code, it uses approximately the same total tokens. The savings appear only when it cuts corners.

PatternDescription
Large agent, low churnHigh initial cost, fewer retries, less debugging
Small agent, high churnLow initial cost, more retries, more debugging

Our data showed a 57-line agent used 69.5k tokens vs 69.6k for a 3,529-line agent on the same correct solution — prompt size alone does not determine cost.

Step 4: State verdict with evidence

The verdict must be backed by data. Include:

  • Which agent won on simple tasks (expected: equivalent)
  • Which agent won on complex tasks (expected: full agent)
  • Total session cost comparison
  • Effective cost comparison (with bug penalty)
  • Clear recommendation for when to use each variant

See references/methodology.md for the complete testing methodology with December 2024 data.

Step 5: Clean up

Remove temporary benchmark files and debug outputs. Keep only the comparison report and generated code.

Gate: Report generated with all metrics. Verdict stated with evidence. Report saved to benchmark directory.

Phase 5: OPTIMIZE (optional — invoked explicitly)

Goal: Run an automated optimization loop that improves a markdown target's frontmatter description using trigger-rate eval tasks, then selects the best measured variants through beam search or single-path search.

Invoke when the user says "optimize this skill", "optimize the description", or "run autoresearch". The existing manual A/B comparison (Phases 1-4) remains the path for full agent benchmarking.

See references/optimize-phase.md for the full 9-step procedure, all CLI flags, recommended modes, live eval defaults, current reality check, and optional extensions.

Gate: Optimization complete. Results reviewed. Cherry-picked improvements applied and verified against full task set. Results recorded.


References

  • ${CLAUDE_SKILL_DIR}/references/methodology.md: Complete testing methodology with December 2024 data
  • ${CLAUDE_SKILL_DIR}/references/grading-rubric.md: Detailed grading criteria and quality checklists
  • ${CLAUDE_SKILL_DIR}/references/benchmark-tasks.md: Standard benchmark task descriptions and prompts
  • ${CLAUDE_SKILL_DIR}/references/report-template.md: Comparison report template with all required sections
  • ${CLAUDE_SKILL_DIR}/references/optimize-phase.md: Full Phase 5 OPTIMIZE procedure (autoresearch loop, CLI flags, beam search, reality check)
  • ${CLAUDE_SKILL_DIR}/references/examples-and-errors.md: Error handling for common benchmark failures

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!