复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
Essays and writing behind this toolkit live at vexjoy.com.
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
Essays and writing behind this toolkit live at vexjoy.com.
AI agents skip steps.
"Looks correct" replaces running tests. "Trivial change" replaces verification. The agent confidently ships broken code because nothing structurally prevented it from skipping the work.
Harnesses have a second problem: given only a skill list, they do not route eagerly enough, or correctly enough. Good skills sit unused. So this toolkit connects the skills, agents, and workflows we want directly into the harness, automatically. You don't have to understand what is here. Say what you want in plain English and you get all the value we have put into it: the right specialist with the right methodology, behind gates that demand exit codes, not assertions.
44 domain agents, 122 workflow skills, 78 hooks, 136 scripts. Agents carry knowledge, skills enforce methodology, hooks block incomplete work, scripts handle determinism.
Works across Claude Code (/do), Codex ($do), Factory (/do), Reasonix (/do).
$ claude
> /do debug this Go test
Routing: go-engineer + systematic-debugging
Phase 1/4: Reproduce: running test, capturing failure...
Phase 2/4: Hypothesize: 3 candidates from stack trace...
Phase 3/4: Verify: isolated root cause in connection pool timeout
Phase 4/4: Fix: patch applied, test passing, PR opened
✓ Delivered: PR #847, fix connection pool timeout in health check
The router reads intent, picks a Go agent paired with a debugging skill, and runs the full lifecycle. You typed one sentence. The system did the rest.
ROUTE PLAN EXECUTE VERIFY DELIVER RECORD
┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐
│ /do │───▶│ Task │───▶│Agent │───▶│Tests │───▶│ PR │───▶│Route │
│Router│ │ Plan │ │+Skill│ │Gates │ │Branch│ │Result│
└──────┘ └──────┘ └──────┘ └──────┘ └──────┘ └──────┘
This is the single thing that separates it from "agent with a system prompt."
| Agent Says | What Happens |
|---|---|
| "Code looks correct, skip tests" | Exit gate requires test output. Blocked. |
| "Trivial change, no verification" | Hook blocks completion without evidence. |
| "Similar to before" | Skill demands case-specific proof. |
| "User is in a hurry" | Protocol overrides time pressure. |
| "I'm confident" | Gate demands exit code, not assertion. |
Hooks fire automatically. Gates block completion. Skills encode counter-arguments at every skip-worthy step. The agent verifies or it doesn't finish.
For what I do, the difference is enormous. If you're doing simple single-file edits, maybe less so.
The same routing serves knowledge work. The content engine researches, drafts in a calibrated voice, validates against 397 AI patterns, and repurposes finished pieces for each platform. /html turns any request into a single self-contained HTML file: report, slide deck, prototype, data viz, diagram. Non-engineers who try the toolkit consistently name the HTML artifacts as the thing they love. No code, no setup beyond the installer.
Changes to the toolkit itself ship with evidence. New skills get blind A/B tests against a no-skill baseline before merge. Routing and writing-standard decisions carry measured verdicts; PHILOSOPHY.md cites the numbers. Experiments that lost go into the negative-results registry, what-didnt-work.md; the registry now covers routing reversals, unvalidated A/B citations, and disabled lint rules alongside the original program refutations.
The automated nightly evolution loop (/evolve, writes to evolution-reports/) ran regularly through mid-May 2026. It is currently dormant; recent evidence has come from manual PRs instead.
git clone https://github.com/notque/vexjoy-agent.git ~/vexjoy-agent
cd ~/vexjoy-agent
./install.sh
Links into ~/.claude/ and mirrors into ~/.codex/, ~/.factory/, ~/.reasonix/ — each mirror only when that runtime is detected (its command on PATH or its home dir already exists). The installer asks symlink (live updates via git pull) or copy (stable snapshot).
Want only part of the toolkit? Run ./install.sh --configure to pick which skills, agents, and hooks install, or copy .local.example/profile.yaml to .local/profile.yaml and edit. No profile file = full install, unchanged behavior. Credit: @thomasvan. Details: .local.example/README.md.
| CLI | Entry Point |
|---|---|
| Claude Code | /do |
| Codex | $do |
| Factory | /do |
| Reasonix | /do |
Full setup: docs/start-here.md
Mirrors agents, skills, and supported hooks into ~/.codex/. The original six-hook allowlist was correct for Codex v0.114, when tool hooks only intercepted Bash. Current support requires Codex v0.144.1+ and classifies the 74 Claude hook registrations as 26 native, 35 adapter-backed, and 13 unsupported (61 supported). These are registration counts, not unique hook files. The installer also preserves explicit per-subagent model routing for GPT-5.6 Sol by setting the MultiAgent V2 compatibility keys documented in openai/codex#31814.
Codex now exposes apply_patch to tool hooks. VexJoy's adapter converts each patch operation into the Write/Edit payload expected by existing guards, but it cannot intercept writes performed through unified_exec, unmatched MCP tools, WebSearch, or other unsupported tool paths. PreCompact and Stop adapters also receive less telemetry than Claude Code: Codex does not provide Claude's conversation_history or session_data. This is expanded compatibility, not full Claude parity.
After install or any hook-definition change, run /hooks in Codex and review the new definitions before trusting them. Codex hash-trusts hook commands and skips changed, unreviewed definitions.
Gemini CLI support removed (deprecated upstream, transitioned to Antigravity CLI); Antigravity support pending CLI maturity. Per Google's transition announcement, Gemini CLI stops serving requests on 2026-06-18 for Google AI Pro / Ultra and free Gemini Code Assist for individuals. Gemini API integrations (image-gen backends, sprite pipeline, GEMINI_API_KEY) are unaffected and stay in the toolkit.
If a prior install mirrored into ~/.gemini/, remove the stale mirrors with:
rm -rf ~/.gemini/skills ~/.gemini/agents ~/.gemini/hooks ~/.gemini/scripts ~/.gemini/antigravity/plugins/vexjoy-agent
Mirrors agents (as "droids"), skills, and all hooks into ~/.factory/. Hook config merges into ~/.factory/settings.json with paths rewritten.
Mirrors skills, scripts, and the allowlisted hooks (scripts/reasonix-hooks-allowlist.txt) into ~/.reasonix/ (no agent or custom-command surface, so neither is installed; the /do router rides in as a skill). Reasonix fires only 4 events (PreToolUse, PostToolUse, UserPromptSubmit, Stop), so only hooks for those events are allowlisted. Hook config is written to the hooks key of ~/.reasonix/settings.json in Reasonix's native flat shape (one entry per hook, match regex over the tool name); the generator builds absolute python3 commands, so no path rewrite is applied. MCP/model/permissions in ~/.reasonix/config.json are user-owned and left untouched.
The toolkit supplies its own routing, domain knowledge, methodology, and enforcement. The default system prompt duplicates most of that.
claude --system-prompt "."
Strips built-in tool-use instructions. The toolkit's agents, skills, hooks, and CLAUDE.md provide equivalent coverage.
| Layer | Count | Does |
|---|---|---|
| Agents | 44 | Domain knowledge: idiom tables, failure mode catalogs, error-to-fix mappings |
| Skills | 122 | Phased methodology with gates. Can't skip steps. Each phase has exit criteria requiring evidence. |
| Hooks | 78 | Fire on lifecycle events. Block incomplete work. Zero LLM cost. |
| Scripts | 136 | Determinism: test runners, linters, validators. No LLM judgment. |
Full skill catalog: docs/skills.md.
┌─────────────────────────────────────────────────┐
│ SKILL.md │
│ ┌─ Frontmatter ─────────────────────────────┐ │
│ │ triggers, pairs_with, success-criteria │ │
│ └────────────────────────────────────────────┘ │
│ Reference Loading Table (conditional imports) │
│ Phased Instructions (numbered, with gates) │
│ Verification (evidence requirements) │
└─────────────────────────────────────────────────┘
A game built entirely by Claude Code using these agents, skills, and pipelines:
I just want to use it Install, learn /do, done.
I do knowledge work Writing, research, data analysis, moderation, HTML artifacts. No code.
I'm a developer Architecture, extension points, adding agents and skills.
I'm an AI power user Routing tables, pipelines, hooks, telemetry DB.
I'm an AI agent Machine-dense inventory. Tables, paths, schemas.
I'm on LinkedIn 🚀 Thought leadership. Agree? 👇
Full design philosophy: PHILOSOPHY.md
One report-only script surfaces upkeep work; it prints a digest and never edits, deletes, or blocks.
python3 scripts/stale-skill-scan.py --top 20 ranks stale skills and agents as pruning candidates. Run it quarterly; see docs/deprecation-template.md.Scheduled work follows the same boundary as everything else: judgment uses agents; repeatable plumbing uses scripts.
| Need | Use |
|---|---|
| Run a deterministic command on a schedule | scripts/agent-scheduler.py with runner: "command" |
| Run an agent judgment on a schedule, webhook, or file change | scripts/agent-scheduler.py with the default runner: "claude" |
| Install or remove a user crontab entry safely | scripts/crontab-manager.py |
| Audit shell cron reliability | cron-automation |
| Keep one interactive objective moving until criteria verify | objective-loop |
See CONTRIBUTING.md.
MIT. See LICENSE.
name: quick
description: "Tracked lightweight execution with composable rigor flags: --trivial, --discuss, --research, --full. Covers zero-ceremony inline fixes (typo, spelling fix, small mistake in a single file, ≤3 edits) through contained multi-file changes."
user-invocable: true
argument-hint: "[--trivial] [--discuss] [--research] [--full] <task>"
allowed-tools:
- Read
- Write
- Edit
- Bash
- Grep
- Glob
- Skill
- Task
routing:
force_route: true
triggers:
- quick task
- small change
- ad hoc task
- add a flag
- small refactor
- targeted fix
- quick fix
- typo fix
- fix typo
- fix the typo
- one-line change
- trivial fix
- rename variable
- rename this variable
- update value
- fix import
- small mistake
- small mistake in
- mistake in spelling
- spelling mistake
- spelling fix
- fix the spelling
- typo in
- small fix in
- small fix
- tiny fix
not_for: "'quick' as speed preference, general bug diagnosis requiring investigation"
complexity: Simple
category: processQuick covers the full lightweight tier from zero-ceremony inline fixes (≤3 edits, --trivial mode) through contained multi-file changes. Full-ceremony Simple+ tasks (task_plan.md, agent routing, quality gates) belong in /do. The key design principle is composable rigor: the base mode is minimal (plan + execute), and users add process incrementally via flags.
Flags (all OFF by default):
| Flag | Effect |
|---|---|
--trivial | Zero-ceremony inline mode for ≤3 file edits: no plan display, no branch, direct commit. Scope-gates strictly — escalates automatically if the task needs more than 3 edits. Use for typo fixes, one-line constant changes, renaming a variable, fixing an import. |
--discuss | Add a pre-planning discussion phase to resolve ambiguities (breadth-first — surfaces all gray areas at once) |
--interview | Add a depth-first decision-tree interview before the edit phase. One question at a time with a recommendation per question. Sibling to --discuss — pick --interview when decisions are interdependent and answer A constrains valid options for B. |
--research | Add a research phase before planning to build context on unfamiliar code |
--full | Add plan verification + full quality gates (tests, lint, diff review) |
--no-branch | Skip feature branch creation, work on current branch |
--no-commit | Skip the commit step (for batching multiple quick tasks) |
| Signal | Load These Files | Why |
|---|---|---|
| usage examples, task ID format, error handling | examples.md | Loads detailed guidance from examples.md. |
| emitting banners, commit format, or STATE.md entries | templates.md | Loads detailed guidance from templates.md. |
When --trivial is passed (or the router recognises a clearly one-line mechanical change), execute inline without plan display, without a feature branch, and without spawning subagents — because the overhead of those steps dwarfs the actual work.
Step 1: Read CLAUDE.md
Read repository CLAUDE.md before any edit, because repo-specific constraints affect how even trivial changes should be made.
Step 2: Scope check
| Question | If Yes |
|---|---|
| Does this need reading docs, investigating behavior, or understanding unfamiliar code? | Drop --trivial, redirect to /quick --research |
| Does this touch more than 3 files? | Drop --trivial, redirect to standard /quick |
| Does this add imports from new packages or modify dependency files? | Drop --trivial, redirect to standard /quick |
| Is the request ambiguous or underspecified? | Ask one clarifying question; if still ambiguous after one round, drop --trivial and use /quick --discuss |
If redirecting, say: This task exceeds --trivial scope ([reason]). Continuing as /quick. Then proceed to Phase 0 with the original request.
Step 3: Locate target files and execute
Read the target file(s). Make edits using the Edit tool. Track edit count. After each edit, check: have we hit 3 edits? If more are needed, stop -- do not rationalize "just one more edit." Say: "Scope exceeded during --trivial execution (3+ edits needed). Preserving work done. Continuing as /quick." Hand off to the standard quick phases with context about what was already done.
Step 4: Check branch
If on main/master, create a short-lived branch: git checkout -b quick/<brief-description>.
Step 5: Stage and commit
Stage specific files with git add <specific-files>. Commit using the format from references/templates.md (conventional commit, type usually fix:, chore:, or refactor:).
Step 6: Display summary
Use the --trivial summary banner format from references/templates.md.
GATE: Edit count is 1-3, commit succeeded. STOP — do not proceed to Phase 0.
Step 1: Read CLAUDE.md
Read and follow the repository's CLAUDE.md before doing anything else, because repo-specific conventions override defaults and skipping this causes style/tooling mismatches.
Step 2: Parse flags
Extract --discuss, --interview, --research, --full, --no-branch, and --no-commit from the invocation. Everything remaining after flag extraction is the task description.
Step 3: Scope check
If the task involves multiple components, architectural changes, or needs parallel execution, redirect to /do instead because quick tasks are single-threaded by design -- parallelism means the task has outgrown this tier.
This phase activates when the user passes --discuss or the request contains signals of uncertainty ("not sure", "maybe", "could be", "what do you think").
Step 1: Identify ambiguities
Read the request and list specific questions:
Step 2: Present questions
Use the DISCUSS banner format from references/templates.md. Wait for user response. Do not proceed until ambiguities are resolved.
GATE: All ambiguities resolved. Proceed to Phase 2 or Phase 3.
This phase activates when the user passes --interview. It is the depth-first sibling to --discuss — pick this mode when decisions are interdependent and answering A would change the valid options for B, so batch-asking forces premature commitments.
Step 1: Load the depth-first reference
Load planning/references/depth-first-interview.md and follow its phases (PRIME → ENUMERATE BRANCHES → TRAVERSE → COMPILE OUTPUT). The reference enforces a hard cap of 5 total questions and 3-level recursion per branch — these limits are infrastructure, not advisory.
Step 2: Treat --interview as an explicit trigger
Per the reference's Phase 0 trigger classification, /quick --interview is an explicit invocation. Skip the Phase 0 opt-out question — the user already opted in by passing the flag. Go directly to ENUMERATE BRANCHES.
Step 3: Compile output to inline context
The reference's Phase 3 emits a structured block (Resolved Decisions / Carried Forward / Scope Boundary / Mode Used). Keep this block inline as the discussion artifact and use it to inform Phase 3 PLAN. Do not write a separate task_plan.md for /quick tier.
GATE: Interview output emitted. Resolved decisions inform the inline plan. Proceed to Phase 2 or Phase 3.
This phase activates when the user passes --research or the task touches code that needs investigation. Use --research when touching unfamiliar code because confidence about code behavior is not the same as correctness — --trivial exists for when you truly know.
Step 1: Identify scope
Determine which files and patterns need reading to understand the change.
Step 2: Read and analyze
Read relevant source files, tests, and configuration. Build a mental model of:
Step 3: Summarize findings
Present a brief (3-5 line) summary of what you learned and how it affects the plan.
GATE: Sufficient understanding to plan the change. Proceed to Phase 3.
Step 1: Generate task ID
Assign the task ID now, not later, because untracked tasks become invisible and "later" never comes.
Format: YYMMDD-xxx where xxx is Base36 sequential (0-9, a-z).
# Check STATE.md for today's tasks to determine next sequence
date_prefix=$(date +%y%m%d)
If STATE.md exists in the repo root, find the highest sequence number for today's date prefix and increment. If no tasks today, start at 001. Use Base36 for the sequence: 001, 002, ... 009, 00a, 00b, ... 00z, 010, ...
If STATE.md is corrupted, scan git log for Quick task YYMMDD- patterns to find the true next ID. If a branch name collision occurs, increment the sequence number and try again.
Step 2: Create inline plan
Always display the inline plan, even for obvious tasks, because the plan catches misunderstandings before they become wrong edits and confirms alignment in 10 seconds that saves minutes. Use the inline plan banner instead of writing a task_plan.md file; that keeps the workflow in the Simple+ tier and preserves the minimum viable ceremony.
Use the inline plan banner format from references/templates.md.
If estimated edits exceed 15, prompt the user to consider /do -- edit count is a scope signal regardless of difficulty.
If the task involves security, payments, or data migration, recommend --full because a one-line auth change can be catastrophic and risk is about impact, not size.
Step 3: Create feature branch (unless --no-branch)
Create a feature branch because small changes on main break the same as big ones:
git checkout -b quick/<task-id>-<brief-kebab-description>
If already on a non-main feature branch and --no-branch is set, stay on the current branch.
GATE: Task ID assigned, plan displayed, branch created. Proceed to Phase 4.
Step 1: Make edits
Execute the changes described in the plan. Track edit count throughout.
Step 2: Scope monitoring
Step 3: Verify changes (base mode)
Run a language-appropriate syntax check on edited files (e.g., python3 -m py_compile, go build ./..., tsc --noEmit). If --full flag is set, run the full quality gate instead (see Phase 5).
GATE: All planned edits complete. Sanity check passes.
Step 1: Run tests for affected packages/modules only -- do not run the full suite unless explicitly requested.
Step 2: Lint check using the repo's configured linter on changed files.
Step 3: Review changes with git diff. Check for unintended changes, missing error handling, and broken imports.
GATE: Tests pass, lint clean, diff reviewed. Proceed to Phase 6.
Stage specific files with git add <specific-files> -- not git add ., to avoid accidental inclusions. Commit using the format from references/templates.md. Include the task ID in the commit body for traceability.
GATE: Commit succeeded. Verify with git log -1 --oneline.
Step 1: Update STATE.md
Log the task to STATE.md because this is how tasks stay visible and cross-referenceable. Use the STATE.md schema from references/templates.md to create the file if it does not exist, and append one row per task. If escalated from --trivial, use tier trivial->quick.
Step 2: Display summary
Use the completion banner format from references/templates.md.
See
references/examples.mdfor worked examples per flag mode, task ID format, and error handling.
评论 (0)
暂无评论,成为第一个评论者吧!