复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
Essays and writing behind this toolkit live at vexjoy.com.
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
Essays and writing behind this toolkit live at vexjoy.com.
AI agents skip steps.
"Looks correct" replaces running tests. "Trivial change" replaces verification. The agent confidently ships broken code because nothing structurally prevented it from skipping the work.
Harnesses have a second problem: given only a skill list, they do not route eagerly enough, or correctly enough. Good skills sit unused. So this toolkit connects the skills, agents, and workflows we want directly into the harness, automatically. You don't have to understand what is here. Say what you want in plain English and you get all the value we have put into it: the right specialist with the right methodology, behind gates that demand exit codes, not assertions.
44 domain agents, 122 workflow skills, 78 hooks, 136 scripts. Agents carry knowledge, skills enforce methodology, hooks block incomplete work, scripts handle determinism.
Works across Claude Code (/do), Codex ($do), Factory (/do), Reasonix (/do).
$ claude
> /do debug this Go test
Routing: go-engineer + systematic-debugging
Phase 1/4: Reproduce: running test, capturing failure...
Phase 2/4: Hypothesize: 3 candidates from stack trace...
Phase 3/4: Verify: isolated root cause in connection pool timeout
Phase 4/4: Fix: patch applied, test passing, PR opened
✓ Delivered: PR #847, fix connection pool timeout in health check
The router reads intent, picks a Go agent paired with a debugging skill, and runs the full lifecycle. You typed one sentence. The system did the rest.
ROUTE PLAN EXECUTE VERIFY DELIVER RECORD
┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐
│ /do │───▶│ Task │───▶│Agent │───▶│Tests │───▶│ PR │───▶│Route │
│Router│ │ Plan │ │+Skill│ │Gates │ │Branch│ │Result│
└──────┘ └──────┘ └──────┘ └──────┘ └──────┘ └──────┘
This is the single thing that separates it from "agent with a system prompt."
| Agent Says | What Happens |
|---|---|
| "Code looks correct, skip tests" | Exit gate requires test output. Blocked. |
| "Trivial change, no verification" | Hook blocks completion without evidence. |
| "Similar to before" | Skill demands case-specific proof. |
| "User is in a hurry" | Protocol overrides time pressure. |
| "I'm confident" | Gate demands exit code, not assertion. |
Hooks fire automatically. Gates block completion. Skills encode counter-arguments at every skip-worthy step. The agent verifies or it doesn't finish.
For what I do, the difference is enormous. If you're doing simple single-file edits, maybe less so.
The same routing serves knowledge work. The content engine researches, drafts in a calibrated voice, validates against 397 AI patterns, and repurposes finished pieces for each platform. /html turns any request into a single self-contained HTML file: report, slide deck, prototype, data viz, diagram. Non-engineers who try the toolkit consistently name the HTML artifacts as the thing they love. No code, no setup beyond the installer.
Changes to the toolkit itself ship with evidence. New skills get blind A/B tests against a no-skill baseline before merge. Routing and writing-standard decisions carry measured verdicts; PHILOSOPHY.md cites the numbers. Experiments that lost go into the negative-results registry, what-didnt-work.md; the registry now covers routing reversals, unvalidated A/B citations, and disabled lint rules alongside the original program refutations.
The automated nightly evolution loop (/evolve, writes to evolution-reports/) ran regularly through mid-May 2026. It is currently dormant; recent evidence has come from manual PRs instead.
git clone https://github.com/notque/vexjoy-agent.git ~/vexjoy-agent
cd ~/vexjoy-agent
./install.sh
Links into ~/.claude/ and mirrors into ~/.codex/, ~/.factory/, ~/.reasonix/ — each mirror only when that runtime is detected (its command on PATH or its home dir already exists). The installer asks symlink (live updates via git pull) or copy (stable snapshot).
Want only part of the toolkit? Run ./install.sh --configure to pick which skills, agents, and hooks install, or copy .local.example/profile.yaml to .local/profile.yaml and edit. No profile file = full install, unchanged behavior. Credit: @thomasvan. Details: .local.example/README.md.
| CLI | Entry Point |
|---|---|
| Claude Code | /do |
| Codex | $do |
| Factory | /do |
| Reasonix | /do |
Full setup: docs/start-here.md
Mirrors agents, skills, and supported hooks into ~/.codex/. The original six-hook allowlist was correct for Codex v0.114, when tool hooks only intercepted Bash. Current support requires Codex v0.144.1+ and classifies the 74 Claude hook registrations as 26 native, 35 adapter-backed, and 13 unsupported (61 supported). These are registration counts, not unique hook files. The installer also preserves explicit per-subagent model routing for GPT-5.6 Sol by setting the MultiAgent V2 compatibility keys documented in openai/codex#31814.
Codex now exposes apply_patch to tool hooks. VexJoy's adapter converts each patch operation into the Write/Edit payload expected by existing guards, but it cannot intercept writes performed through unified_exec, unmatched MCP tools, WebSearch, or other unsupported tool paths. PreCompact and Stop adapters also receive less telemetry than Claude Code: Codex does not provide Claude's conversation_history or session_data. This is expanded compatibility, not full Claude parity.
After install or any hook-definition change, run /hooks in Codex and review the new definitions before trusting them. Codex hash-trusts hook commands and skips changed, unreviewed definitions.
Gemini CLI support removed (deprecated upstream, transitioned to Antigravity CLI); Antigravity support pending CLI maturity. Per Google's transition announcement, Gemini CLI stops serving requests on 2026-06-18 for Google AI Pro / Ultra and free Gemini Code Assist for individuals. Gemini API integrations (image-gen backends, sprite pipeline, GEMINI_API_KEY) are unaffected and stay in the toolkit.
If a prior install mirrored into ~/.gemini/, remove the stale mirrors with:
rm -rf ~/.gemini/skills ~/.gemini/agents ~/.gemini/hooks ~/.gemini/scripts ~/.gemini/antigravity/plugins/vexjoy-agent
Mirrors agents (as "droids"), skills, and all hooks into ~/.factory/. Hook config merges into ~/.factory/settings.json with paths rewritten.
Mirrors skills, scripts, and the allowlisted hooks (scripts/reasonix-hooks-allowlist.txt) into ~/.reasonix/ (no agent or custom-command surface, so neither is installed; the /do router rides in as a skill). Reasonix fires only 4 events (PreToolUse, PostToolUse, UserPromptSubmit, Stop), so only hooks for those events are allowlisted. Hook config is written to the hooks key of ~/.reasonix/settings.json in Reasonix's native flat shape (one entry per hook, match regex over the tool name); the generator builds absolute python3 commands, so no path rewrite is applied. MCP/model/permissions in ~/.reasonix/config.json are user-owned and left untouched.
The toolkit supplies its own routing, domain knowledge, methodology, and enforcement. The default system prompt duplicates most of that.
claude --system-prompt "."
Strips built-in tool-use instructions. The toolkit's agents, skills, hooks, and CLAUDE.md provide equivalent coverage.
| Layer | Count | Does |
|---|---|---|
| Agents | 44 | Domain knowledge: idiom tables, failure mode catalogs, error-to-fix mappings |
| Skills | 122 | Phased methodology with gates. Can't skip steps. Each phase has exit criteria requiring evidence. |
| Hooks | 78 | Fire on lifecycle events. Block incomplete work. Zero LLM cost. |
| Scripts | 136 | Determinism: test runners, linters, validators. No LLM judgment. |
Full skill catalog: docs/skills.md.
┌─────────────────────────────────────────────────┐
│ SKILL.md │
│ ┌─ Frontmatter ─────────────────────────────┐ │
│ │ triggers, pairs_with, success-criteria │ │
│ └────────────────────────────────────────────┘ │
│ Reference Loading Table (conditional imports) │
│ Phased Instructions (numbered, with gates) │
│ Verification (evidence requirements) │
└─────────────────────────────────────────────────┘
A game built entirely by Claude Code using these agents, skills, and pipelines:
I just want to use it Install, learn /do, done.
I do knowledge work Writing, research, data analysis, moderation, HTML artifacts. No code.
I'm a developer Architecture, extension points, adding agents and skills.
I'm an AI power user Routing tables, pipelines, hooks, telemetry DB.
I'm an AI agent Machine-dense inventory. Tables, paths, schemas.
I'm on LinkedIn 🚀 Thought leadership. Agree? 👇
Full design philosophy: PHILOSOPHY.md
One report-only script surfaces upkeep work; it prints a digest and never edits, deletes, or blocks.
python3 scripts/stale-skill-scan.py --top 20 ranks stale skills and agents as pruning candidates. Run it quarterly; see docs/deprecation-template.md.Scheduled work follows the same boundary as everything else: judgment uses agents; repeatable plumbing uses scripts.
| Need | Use |
|---|---|
| Run a deterministic command on a schedule | scripts/agent-scheduler.py with runner: "command" |
| Run an agent judgment on a schedule, webhook, or file change | scripts/agent-scheduler.py with the default runner: "claude" |
| Install or remove a user crontab entry safely | scripts/crontab-manager.py |
| Audit shell cron reliability | cron-automation |
| Keep one interactive objective moving until criteria verify | objective-loop |
See CONTRIBUTING.md.
MIT. See LICENSE.
name: pr-workflow
description: |
Pull request lifecycle: commit, codex review, sync, review, fix, status,
cleanup, and PR mining. Use when user wants to commit changes, get a
second-opinion code review from Codex, push changes, create a PR, check PR
status, fix review comments, clean up branches after merge, or mine tribal
knowledge from PR reviews. Use for "commit my changes", "codex review",
"push my changes", "create a PR", "pr status", "fix PR comments",
"clean up branches", "mine PRs", or "address feedback".
user-invocable: true
allowed-tools:
- Bash
- Read
- Write
- Edit
- Grep
- Glob
- Task
- Skill
- AskUserQuestion
routing:
force_route: true
not_for: "general disagreement ('push back on a design'), committing to an idea ('commit to this approach'), pushing out the door, push notifications, social media reviews, metaphorical commit/merge ('commit to a decision', 'merge ideas in your head', 'merge the branches in your head', 'move forward and commit'), 'commit' meaning resolve/decide rather than git-commit — only for git push/commit/PR operations"
triggers:
- "push changes"
- "push my changes"
- "push to GitHub"
- "push to remote"
- "create PR"
- "sync to GitHub"
- "PR status"
- "branch status"
- "merge readiness"
- "fix PR comments"
- "resolve PR feedback"
- "pr-fix"
- "cleanup branches"
- "clean up branches"
- "merged branches"
- "delete merged branch"
- "prune branches"
- "mine PRs"
- "extract review comments"
- "tribal knowledge"
- "process PR feedback"
- "address review comments"
- "submit PR"
- "create pull request"
- "send for review"
- "open PR"
- "generate branch name"
- "validate branch name"
- "name branch"
- "branch convention"
- "git branch name"
- "check CI"
- "CI status"
- "actions status"
- "did CI pass"
- "build status"
- "CI passed"
- "stage and commit"
- "stage files commit"
- "stage modified commit"
- "commit staged"
- "commit changes"
- "commit these"
- "commit my changes"
- "commit my files"
- "codex review"
- "second opinion"
- "code review codex"
- "gpt review"
- "cross-model review"
- "git push"
- "push to origin"
- "push my branch"
- "push the branch"
- "ship it"
- "ship this"
- "ship this work"
- "merge these fixes"
- "merge this work"
- "merge this in"
- "make a pull request"
- "draft a PR"
- "draft pr"
- "publish my changes"
- "publish this"
- "publish my work"
- "let's get this reviewed"
- "send this to GitHub"
- "send to github"
- "wrap up and merge"
- "wrap this up and merge"
- "land PR"
- "land the PR"
- "land this PR"
- "merge contributor PR"
- "rebase and merge PR"
- "update changelog"
- "release notes"
- "curate changelog"
- "decision brief"
- "owner decision brief"
- "authorization tier"
category: git-workflow
pairs_with:
- verification-before-completion
- code-linting
- systematic-code-reviewUmbrella skill for the entire pull request lifecycle. Routes to the correct reference based on the PR task requested.
Detect the user's intent and load the appropriate reference file:
| Intent | Trigger phrases | Reference |
|---|---|---|
| Sync (default) | "push", "create PR", "sync", "ship this" | ${CLAUDE_SKILL_DIR}/references/sync.md |
| Pipeline | "submit PR", "full PR", "end-to-end PR", "open PR" | ${CLAUDE_SKILL_DIR}/references/pipeline.md |
| Fix | "fix PR comments", "address review", "pr-fix", "resolve feedback" | ${CLAUDE_SKILL_DIR}/references/fix.md |
| Status | "pr status", "branch status", "is my PR ready", "check CI" | ${CLAUDE_SKILL_DIR}/references/status.md |
| Cleanup | "clean up branches", "delete merged branch", "prune" | ${CLAUDE_SKILL_DIR}/references/cleanup.md |
| Feedback | "process PR feedback", "address reviews", "what did reviewers say" | ${CLAUDE_SKILL_DIR}/references/feedback.md |
| Miner | "mine PRs", "extract review comments", "tribal knowledge", "reviewer patterns" | ${CLAUDE_SKILL_DIR}/references/miner.md |
| Branch name | "generate branch name", "validate branch name", "name branch", "branch convention", "git branch name" | ${CLAUDE_SKILL_DIR}/references/branch-name.md |
| CI check | "check CI", "CI status", "actions status", "did CI pass", "build status", "CI passed" | ${CLAUDE_SKILL_DIR}/references/ci-check.md |
| Commit | "commit changes", "stage and commit", "commit my changes", "commit my files", "commit these" | ${CLAUDE_SKILL_DIR}/references/commit.md |
| Codex review | "codex review", "second opinion", "code review codex", "gpt review", "cross-model review" | ${CLAUDE_SKILL_DIR}/references/codex-review.md |
| Land | "land PR", "land the PR", "merge contributor PR", "rebase and merge PR" | ${CLAUDE_SKILL_DIR}/references/land-pr.md |
| Body safety | any gh call writing or reading a PR/issue body | ${CLAUDE_SKILL_DIR}/references/gh-body-safety.md |
| Changelog | "update changelog", "release notes", "curate changelog" | ${CLAUDE_SKILL_DIR}/references/changelog-curation.md |
| Decision brief | "decision brief", "authorization tier", "ask the owner", "is it decision-ready" | ${CLAUDE_SKILL_DIR}/references/owner-decision-briefs.md |
| Risk classify | automatic pre-review step; also "classify PR risk", "pr risk", "risk check" | ${CLAUDE_SKILL_DIR}/references/pr-risk-policy.md |
Default action: When invoked with no arguments or ambiguous intent, load sync.md (the most common PR use case).
| Signal | Load These Files | Why |
|---|---|---|
| "push", "create PR", "sync", "ship this" | sync.md | Sync (default) |
| "submit PR", "full PR", "end-to-end PR", "open PR" | pipeline.md | Pipeline |
| "fix PR comments", "address review", "pr-fix", "resolve feedback" | fix.md | Fix |
| "pr status", "branch status", "is my PR ready", "check CI" | status.md | Status |
| "clean up branches", "delete merged branch", "prune" | cleanup.md | Cleanup |
| "process PR feedback", "address reviews", "what did reviewers say" | feedback.md | Feedback |
| "mine PRs", "extract review comments", "tribal knowledge", "reviewer patterns" | miner.md | Miner |
| "generate branch name", "validate branch name", "name branch", "branch convention", "git branch name" | branch-name.md | Branch name |
| "check CI", "CI status", "actions status", "did CI pass", "build status", "CI passed" | ci-check.md | CI check |
| "commit changes", "stage and commit", "commit my changes", "commit my files", "commit these" | commit.md | Commit |
| "codex review", "second opinion", "code review codex", "gpt review", "cross-model review" | codex-review.md | Codex review |
| "land PR", "land the PR", "merge contributor PR", "rebase and merge PR" | land-pr.md | Land |
any gh call writing or reading a PR/issue body | gh-body-safety.md | Body safety |
| "update changelog", "release notes", "curate changelog" | changelog-curation.md | Changelog |
| "decision brief", "authorization tier", "ask the owner", "is it decision-ready" | owner-decision-briefs.md | Decision brief |
| "INDEX.json conflict", "INDEX conflict on rebase", "two PRs regenerated INDEX", "regenerate INDEX after rebase" | index-conflict-resolution.md | INDEX conflict |
| "classify PR risk", "pr risk", "risk check", or automatic pre-review step | pr-risk-policy.md | Risk classify |
Before dispatching reviewers for any PR, run risk classification:
python3 scripts/pr-risk-classify.py --base "$MAIN_BRANCH" --head HEAD
Route to the appropriate review lane based on the risk field:
| Risk | Lane | Action |
|---|---|---|
low | Quick single review | parallel-code-review (3 agents) — lightweight, fast |
medium | Full right-size-review roster | Run right-size-review.py for tier-appropriate wave composition |
high | Full roster + operator sign-off | Right-size-review roster + add **Operator sign-off required** note to PR body Notes section |
When recommend_split is true, surface the recommendation to the user before dispatching review: "This PR has N lines changed (above the 800-line ceiling). Consider splitting into smaller PRs for faster, higher-quality review." Proceed with review if the user chooses to continue.
gh pr create --body "..." is how every agent in this toolkit opens PRs, and it bypasses .github/pull_request_template.md entirely — GitHub only applies that file to the web UI and to a bare gh pr create with no --body. So the structure must be reproduced in the --body string by hand.
Every agent-authored PR body uses the same three sections as .github/pull_request_template.md, in this order: Summary → Changes → Notes. This keeps PR bodies consistent across models (Opus, Sonnet, and every other harness produce the same shape).
Tests run as GitHub Actions, so the Checks tab is the test record. Let CI carry the proof: keep command output (ruff exit, pytest counts, gate traces, dogfood runs) out of the body. Pasting $ pytest → N passed duplicates the Checks tab and bloats the PR — leave it to CI.
Aim for high meaning per word: each line states one fact about the change, declaratively, so a reviewer understands it fast. Density is the target, not minimal length — a large change keeps the three sections and carries the detail it needs; it earns that length by packing each line with signal. Four rules carry the vibe:
Not covered by CI — terraform plan is manual). State each caveat as one fact and let CI carry command output. Skip what is always true (a PR can be reverted; CI runs the tests).Worked example — the shape to emulate (#608-good vs #710-bad):
| Section | Dense (emulate, like #608) | Bloated (rewrite toward dense, the #710 shape) |
|---|---|---|
| Summary | "Registers 3 hooks in settings.json; integrates pre-route.py into /do Phase 2 as a deterministic pre-filter." | One 5-sentence block stuffed with jargon and four metrics. |
| Changes | "SKILL.md — add 21 PR-creation trigger phrases." | One bullet inlining all 21 phrases verbatim; another a 3-sentence rationale paragraph. |
| Notes | "Not covered by CI — the Terraform plan is applied manually." or "Apply the column migration before deploying." (a required risk/verification caveat) — or omitted entirely when nothing qualifies | "Roll back by reverting the branch commits. Tests run in CI — see Checks." (always-true filler) plus a pasted pytest -v dump and a Scope & Risk wall. |
The dense column reads in seconds because each cell carries facts; the bloated column buries the same facts in volume (and re-states what the Checks tab already shows). Aim every body at the dense column.
Copy this canonical skeleton into --body:
## Summary
<!-- State the goal plainly: 1-3 sentences or a few crisp bullets, one fact per line. Name the ADR/issue if any. Keep metrics to the single number that matters. -->
## Changes
<!-- One line per change: verb + what + where. State shape and count for many sub-items; let the diff enumerate them. -->
- `path/or/area` — what changed
## Notes
<!-- Omit for routine PRs (drop the section when nothing qualifies). Note it, one terse declarative line per point, when a reviewer cannot infer the signal from the diff: a non-obvious decision, a deliberate omission, a follow-up, a gotcha, a "supersedes #N". Note it too when a RISK/VERIFICATION trigger holds — manual verification was performed; part of the change sits outside CI coverage; migration/rollout ordering matters; a security-sensitive surface changed (e.g. `Not covered by CI — terraform plan is manual`). Let CI carry command output. Skip what is always true (a PR can be reverted; CI runs the tests). -->
The sync (sync.md Step 5) and pipeline (pipeline.md Phase 5) references carry this same skeleton at their gh pr create call sites. When either path writes a --body, it uses this structure and the density rules above.
--body from the canonical three-section skeleton above (Summary / Changes / Notes), because --body bypasses the GitHub template file. Write the body via temp file + quoted heredoc + --body-file per ${CLAUDE_SKILL_DIR}/references/gh-body-safety.md. Tests run in CI — the Checks tab is the test record, so keep command output out of the body
评论 (0)
暂无评论,成为第一个评论者吧!