复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
Essays and writing behind this toolkit live at vexjoy.com.
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
Essays and writing behind this toolkit live at vexjoy.com.
AI agents skip steps.
"Looks correct" replaces running tests. "Trivial change" replaces verification. The agent confidently ships broken code because nothing structurally prevented it from skipping the work.
Harnesses have a second problem: given only a skill list, they do not route eagerly enough, or correctly enough. Good skills sit unused. So this toolkit connects the skills, agents, and workflows we want directly into the harness, automatically. You don't have to understand what is here. Say what you want in plain English and you get all the value we have put into it: the right specialist with the right methodology, behind gates that demand exit codes, not assertions.
44 domain agents, 122 workflow skills, 78 hooks, 136 scripts. Agents carry knowledge, skills enforce methodology, hooks block incomplete work, scripts handle determinism.
Works across Claude Code (/do), Codex ($do), Factory (/do), Reasonix (/do).
$ claude
> /do debug this Go test
Routing: go-engineer + systematic-debugging
Phase 1/4: Reproduce: running test, capturing failure...
Phase 2/4: Hypothesize: 3 candidates from stack trace...
Phase 3/4: Verify: isolated root cause in connection pool timeout
Phase 4/4: Fix: patch applied, test passing, PR opened
✓ Delivered: PR #847, fix connection pool timeout in health check
The router reads intent, picks a Go agent paired with a debugging skill, and runs the full lifecycle. You typed one sentence. The system did the rest.
ROUTE PLAN EXECUTE VERIFY DELIVER RECORD
┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐
│ /do │───▶│ Task │───▶│Agent │───▶│Tests │───▶│ PR │───▶│Route │
│Router│ │ Plan │ │+Skill│ │Gates │ │Branch│ │Result│
└──────┘ └──────┘ └──────┘ └──────┘ └──────┘ └──────┘
This is the single thing that separates it from "agent with a system prompt."
| Agent Says | What Happens |
|---|---|
| "Code looks correct, skip tests" | Exit gate requires test output. Blocked. |
| "Trivial change, no verification" | Hook blocks completion without evidence. |
| "Similar to before" | Skill demands case-specific proof. |
| "User is in a hurry" | Protocol overrides time pressure. |
| "I'm confident" | Gate demands exit code, not assertion. |
Hooks fire automatically. Gates block completion. Skills encode counter-arguments at every skip-worthy step. The agent verifies or it doesn't finish.
For what I do, the difference is enormous. If you're doing simple single-file edits, maybe less so.
The same routing serves knowledge work. The content engine researches, drafts in a calibrated voice, validates against 397 AI patterns, and repurposes finished pieces for each platform. /html turns any request into a single self-contained HTML file: report, slide deck, prototype, data viz, diagram. Non-engineers who try the toolkit consistently name the HTML artifacts as the thing they love. No code, no setup beyond the installer.
Changes to the toolkit itself ship with evidence. New skills get blind A/B tests against a no-skill baseline before merge. Routing and writing-standard decisions carry measured verdicts; PHILOSOPHY.md cites the numbers. Experiments that lost go into the negative-results registry, what-didnt-work.md; the registry now covers routing reversals, unvalidated A/B citations, and disabled lint rules alongside the original program refutations.
The automated nightly evolution loop (/evolve, writes to evolution-reports/) ran regularly through mid-May 2026. It is currently dormant; recent evidence has come from manual PRs instead.
git clone https://github.com/notque/vexjoy-agent.git ~/vexjoy-agent
cd ~/vexjoy-agent
./install.sh
Links into ~/.claude/ and mirrors into ~/.codex/, ~/.factory/, ~/.reasonix/ — each mirror only when that runtime is detected (its command on PATH or its home dir already exists). The installer asks symlink (live updates via git pull) or copy (stable snapshot).
Want only part of the toolkit? Run ./install.sh --configure to pick which skills, agents, and hooks install, or copy .local.example/profile.yaml to .local/profile.yaml and edit. No profile file = full install, unchanged behavior. Credit: @thomasvan. Details: .local.example/README.md.
| CLI | Entry Point |
|---|---|
| Claude Code | /do |
| Codex | $do |
| Factory | /do |
| Reasonix | /do |
Full setup: docs/start-here.md
Mirrors agents, skills, and supported hooks into ~/.codex/. The original six-hook allowlist was correct for Codex v0.114, when tool hooks only intercepted Bash. Current support requires Codex v0.144.1+ and classifies the 74 Claude hook registrations as 26 native, 35 adapter-backed, and 13 unsupported (61 supported). These are registration counts, not unique hook files. The installer also preserves explicit per-subagent model routing for GPT-5.6 Sol by setting the MultiAgent V2 compatibility keys documented in openai/codex#31814.
Codex now exposes apply_patch to tool hooks. VexJoy's adapter converts each patch operation into the Write/Edit payload expected by existing guards, but it cannot intercept writes performed through unified_exec, unmatched MCP tools, WebSearch, or other unsupported tool paths. PreCompact and Stop adapters also receive less telemetry than Claude Code: Codex does not provide Claude's conversation_history or session_data. This is expanded compatibility, not full Claude parity.
After install or any hook-definition change, run /hooks in Codex and review the new definitions before trusting them. Codex hash-trusts hook commands and skips changed, unreviewed definitions.
Gemini CLI support removed (deprecated upstream, transitioned to Antigravity CLI); Antigravity support pending CLI maturity. Per Google's transition announcement, Gemini CLI stops serving requests on 2026-06-18 for Google AI Pro / Ultra and free Gemini Code Assist for individuals. Gemini API integrations (image-gen backends, sprite pipeline, GEMINI_API_KEY) are unaffected and stay in the toolkit.
If a prior install mirrored into ~/.gemini/, remove the stale mirrors with:
rm -rf ~/.gemini/skills ~/.gemini/agents ~/.gemini/hooks ~/.gemini/scripts ~/.gemini/antigravity/plugins/vexjoy-agent
Mirrors agents (as "droids"), skills, and all hooks into ~/.factory/. Hook config merges into ~/.factory/settings.json with paths rewritten.
Mirrors skills, scripts, and the allowlisted hooks (scripts/reasonix-hooks-allowlist.txt) into ~/.reasonix/ (no agent or custom-command surface, so neither is installed; the /do router rides in as a skill). Reasonix fires only 4 events (PreToolUse, PostToolUse, UserPromptSubmit, Stop), so only hooks for those events are allowlisted. Hook config is written to the hooks key of ~/.reasonix/settings.json in Reasonix's native flat shape (one entry per hook, match regex over the tool name); the generator builds absolute python3 commands, so no path rewrite is applied. MCP/model/permissions in ~/.reasonix/config.json are user-owned and left untouched.
The toolkit supplies its own routing, domain knowledge, methodology, and enforcement. The default system prompt duplicates most of that.
claude --system-prompt "."
Strips built-in tool-use instructions. The toolkit's agents, skills, hooks, and CLAUDE.md provide equivalent coverage.
| Layer | Count | Does |
|---|---|---|
| Agents | 44 | Domain knowledge: idiom tables, failure mode catalogs, error-to-fix mappings |
| Skills | 122 | Phased methodology with gates. Can't skip steps. Each phase has exit criteria requiring evidence. |
| Hooks | 78 | Fire on lifecycle events. Block incomplete work. Zero LLM cost. |
| Scripts | 136 | Determinism: test runners, linters, validators. No LLM judgment. |
Full skill catalog: docs/skills.md.
┌─────────────────────────────────────────────────┐
│ SKILL.md │
│ ┌─ Frontmatter ─────────────────────────────┐ │
│ │ triggers, pairs_with, success-criteria │ │
│ └────────────────────────────────────────────┘ │
│ Reference Loading Table (conditional imports) │
│ Phased Instructions (numbered, with gates) │
│ Verification (evidence requirements) │
└─────────────────────────────────────────────────┘
A game built entirely by Claude Code using these agents, skills, and pipelines:
I just want to use it Install, learn /do, done.
I do knowledge work Writing, research, data analysis, moderation, HTML artifacts. No code.
I'm a developer Architecture, extension points, adding agents and skills.
I'm an AI power user Routing tables, pipelines, hooks, telemetry DB.
I'm an AI agent Machine-dense inventory. Tables, paths, schemas.
I'm on LinkedIn 🚀 Thought leadership. Agree? 👇
Full design philosophy: PHILOSOPHY.md
One report-only script surfaces upkeep work; it prints a digest and never edits, deletes, or blocks.
python3 scripts/stale-skill-scan.py --top 20 ranks stale skills and agents as pruning candidates. Run it quarterly; see docs/deprecation-template.md.Scheduled work follows the same boundary as everything else: judgment uses agents; repeatable plumbing uses scripts.
| Need | Use |
|---|---|
| Run a deterministic command on a schedule | scripts/agent-scheduler.py with runner: "command" |
| Run an agent judgment on a schedule, webhook, or file change | scripts/agent-scheduler.py with the default runner: "claude" |
| Install or remove a user crontab entry safely | scripts/crontab-manager.py |
| Audit shell cron reliability | cron-automation |
| Keep one interactive objective moving until criteria verify | objective-loop |
See CONTRIBUTING.md.
MIT. See LICENSE.
name: explanation-traces
description: "Query and display structured decision traces from routing, agent selection, and skill execution."
user-invocable: true
argument-hint: "<optional: specific decision to explain>"
allowed-tools:
- Read
- Bash
- Glob
- Grep
routing:
triggers:
- "why did you"
- "explain routing"
- "show trace"
- "decision log"
- "why that agent"
- "explain decision"
- "show decisions"
- "trace log"
force_route: true
not_for: "general 'why did the test fail' debugging, explaining concepts to a user, code documentation, stack traces — only for querying recorded routing/agent decisions"
pairs_with: []
complexity: Simple
category: analysisThis skill reads the per-dispatch route event log and presents routing decisions and their outcomes as a human-readable timeline. It answers "why did I get routed here?" from what was recorded at decision time — never from post-hoc reconstruction or rationalization.
The log: <CLAUDE_LEARNING_DIR>/route-events.jsonl, default ~/.claude/learning/route-events.jsonl. Append-only JSONL — one JSON object per line. Written via hooks/lib/route_events.py by two producers:
| Producer | Fires on | Appends |
|---|---|---|
hooks/routing-decision-recorder.py | PostToolUse (Agent dispatch) | One DECISION event per /do-routed dispatch |
hooks/routing-outcome-finalizer.py | UserPromptSubmit | One OUTCOME event when it finalizes a pending dispatch |
The log is auxiliary instrumentation: writes are failure-safe (worst case one lost line), and the aggregate routing rows in learning.db stay authoritative for the confidence loop.
Key constraints baked into the workflow:
ts (epoch seconds) and recorded fields are authoritative; keep their precisionrequest_snippet is private session data: show it to this session's own user, and keep it out of anything that leaves the session (PR bodies, issues, exports) — report counts there insteadGoal: Find the route event log.
Step 1: Resolve the path and check it
LOG="${CLAUDE_LEARNING_DIR:-$HOME/.claude/learning}/route-events.jsonl"
wc -l "$LOG"
CLAUDE_LEARNING_DIR redirects the log (tests and redirected DBs use it); unset means the default ~/.claude/learning/.
Step 2: Handle missing log
If the file is absent or empty, stop and inform the user:
No route event log found at ~/.claude/learning/route-events.jsonl
(or $CLAUDE_LEARNING_DIR/route-events.jsonl when that variable is set).
The log is created on the first /do-routed dispatch by the
routing-decision-recorder hook (hooks/routing-decision-recorder.py).
An empty or missing log means no /do-routed dispatch has been recorded
yet — or merged hook changes were never synced to ~/.claude; run
hooks/sync-to-user-claude.py or restart the session.
Recorded events are the only source this skill reads. Reconstructing decisions from memory or conversation history defeats its purpose — with no log, there is nothing to read, and the honest answer is exactly that.
GATE: Log found and non-empty. Proceed only when gate passes.
Goal: Extract events and filter to the user's query.
Step 1: Read the events
Parse each line as one JSON object. Two event types (full semantics: references/trace-schema.md; source of truth: hooks/lib/route_events.py).
DECISION — one per /do-routed dispatch:
| Field | Meaning |
|---|---|
ts | Epoch seconds (float) when the dispatch was recorded |
session | Session id ("" when unknown) |
request_snippet | First 200 chars of the routed request |
agent, skill, complexity | The chosen route |
health_at_decision | Picked pair's confidence at decision time; null = no weight row or never evaluated (disambiguate with gate_inputs_present) |
n, failure | The other demote-floor inputs, snapshotted with health |
action | Step-1.5 health-gate outcome: keep, demote, or tiebreak |
alternates | Keys offered as alternatives; null when none recorded |
gate_inputs_present | true = the marker carried a health= token; false/absent = legacy marker, health never read |
OUTCOME — one per finalized dispatch:
| Field | Meaning |
|---|---|
ts | Epoch seconds when the outcome was finalized |
session | Session id |
key | Routing key {agent}:{skill} (agent-only {agent}: when skill unknown) |
outcome | success, failure, or neutral |
reason | Short cause (e.g. tool-errors, rejection, acceptance, neutral-new-topic); absent in older events |
routing_relevant | true = a signal the confidence loop acts on; absent = relevance not asserted |
Additive-field history: older lines may lack n, failure, action, alternates, gate_inputs_present, reason, routing_relevant. An absent field means "not recorded then", never corruption.
Step 2: Filter to the user's query
| User signal | Filter strategy |
|---|---|
| Names an agent or skill | DECISION events where agent or skill matches, or the name appears in alternates; OUTCOME events whose key contains it |
| "Why did I get routed here" / latest dispatch | Most recent DECISION events (tail of the log), current session first |
| Asks about outcome ("did it work", "why failure") | OUTCOME events, joined back to their decisions |
| Names a session | Filter both types on session |
| No specific target | Chronological timeline of the most recent session |
Step 3: Join outcomes to decisions
Match an OUTCOME to its DECISION on the same session AND key == "{agent}:{skill}". A decision with no matched outcome is pending (the finalizer runs on a later user prompt) or was never finalized — report that state as-is.
GATE: At least one decision event parsed and filtered. Proceed only when gate passes.
Goal: Format events as a human-readable decision timeline.
Step 1: Build the timeline
Sort by ts (numeric — concurrent appends can interleave lines out of order). Convert ts to local ISO time for display; show raw ts on request. For each decision:
[TIME] {agent} + {skill} ({complexity})
Request: "{request_snippet}"
Health at decision: {health line — see below}
Alternates: {alternates, or "none recorded"}
Outcome: {outcome} ({reason}) — or "pending: not yet finalized"
Health line — three recorded states, rendered distinctly:
| Recorded | Render as |
|---|---|
Numeric health_at_decision | 0.62 (n=7, failure=1) → action=keep |
null + gate_inputs_present: true | no weight row at decision time (new pair) |
null + gate_inputs_present false/absent | health gate not instrumented for this dispatch (legacy marker) |
Group entries by session when the timeline spans more than one, to prevent wall-of-text.
Step 2: Lead with the answer to the user's question
If the user asked "why did I get routed here?", lead with the matching decision, then offer surrounding context:
You asked: "Why did I get the governance agent?"
Decision at [TIME]:
Route: toolkit-governance-engineer + pr-workflow (Complex)
Request: "ship the explanation-traces repoint as a green-CI PR..."
Health at decision: no weight row at decision time (new pair)
Alternates: none recorded
Outcome: pending — not yet finalized
--- Session timeline (3 dispatches) ---
[... remaining entries ...]
Step 3: Flag gaps honestly
When entries lack additive fields or matched outcomes, say so explicitly:
Note: [N] decision(s) predate the health-gate instrumentation — they show
WHAT was routed but carry no health data. [M] decision(s) have no matched
outcome: pending or never finalized.
Incomplete data presented honestly beats complete-looking data that includes fabrication. Leave gaps as gaps.
GATE: Timeline presented. User's question answered from recorded events. Done.
User says: "Show me the decision log"
skill: explanation-traces
Actions:
User says: "Why did I get routed to that agent?"
skill: explanation-traces "why that agent?"
Actions:
User says: "Why was that dispatch marked a failure?"
skill: explanation-traces "failure outcome"
Actions:
outcome: failure events; join each to its decision by session + key (Phase 2)reason (e.g. tool-errors, rejection) with the originating decision (Phase 3)
Result: The recorded failure cause, with the route and request that produced itWrong: Reconstructing "why" from memory when the log is missing. Right: If no log exists, say so, name the real path and producing hook, and stop. Never fabricate an explanation.
Wrong: Rendering every null health as "no data".
Right: null + gate_inputs_present: true = pick had no weight row (new pair). null + false/absent = legacy marker, health never read. Different facts; render them differently.
Wrong: Pairing an outcome with "the decision right above it" in the file.
Right: Match on session + key == "{agent}:{skill}". Interleaved sessions make file adjacency meaningless.
Wrong: Flagging pre-instrumentation lines as malformed because gate_inputs_present is missing.
Right: Fields were added over time; treat absence as "not recorded then" and say so.
Wrong: Always dumping the full timeline regardless of what the user asked. Right: Lead with the specific answer, then offer full context as supplementary detail.
Cause: No /do-routed dispatch recorded yet, hooks never synced to ~/.claude, or CLAUDE_LEARNING_DIR points elsewhere.
Solution: Report the resolved path and the producing hook (hooks/routing-decision-recorder.py). Suggest hooks/sync-to-user-claude.py when hook changes were merged mid-session. Skip any reconstruction from conversation history.
Cause: Truncated append (rare — per-line appends are atomic) or manual edit. Solution: Skip the bad line, keep parsing the rest, and report the count and line numbers of skipped lines. JSONL fails per line, never whole-file.
Cause: File exists but every line is an OUTCOME, or the recorder's marker parsing is failing.
Solution: Report counts by type. Point to references/error-handling.md for the recorder diagnosis steps.
Cause: The recorder only records /do-routed top-level dispatches — nested fan-out and manual Agent calls are deliberately excluded.
Solution: Show what IS recorded and explain the exclusion. Full mapping: references/error-handling.md.
| Task type | Signals | Reference file |
|---|---|---|
| Reading or explaining event fields | "health_at_decision", "gate_inputs_present", "alternates", "key", "schema" | references/trace-schema.md |
| Diagnosing wrong or thin trace data | "health null", "no alternates", "legacy marker", "not instrumented" | references/preferred-patterns.md |
| Handling parse or read errors | "malformed", "missing field", "no decisions", "not found", "pending" | references/error-handling.md |
| Presenting filtered timeline | "why did you", "show trace", "decision log", "explain routing" | references/trace-schema.md |
references/trace-schema.md: Real event schema for route-events.jsonl — DECISION and OUTCOME fields, health states, join rules, examplesreferences/preferred-patterns.md: Failure mode catalog for reading the log — join mistakes, health-state conflation, privacy — with detection commandsreferences/error-handling.md: Error-fix mappings — missing log, malformed lines, no decisions, unmatched outcomes, unrecorded dispatches
评论 (0)
暂无评论,成为第一个评论者吧!