复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
Essays and writing behind this toolkit live at vexjoy.com.
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
Essays and writing behind this toolkit live at vexjoy.com.
VexJoy Agent connects plain-English requests to specialist agents, skills, and workflows. /do selects the knowledge and tools needed for your task. Hooks enforce specific checks, and scripts handle repeatable work.
The aim is to give capable models useful domain knowledge without making you learn the toolkit's catalog.
43 domain agents, 59 workflow skills, 78 hooks, 153 scripts. Agents carry knowledge, skills enforce methodology, hooks block incomplete work, scripts handle determinism.
Works across Claude Code (/do), Codex ($do), Factory (/do), Reasonix (/do).
$ claude
> /do debug this Go test
Routing: go-engineer + systematic-debugging
Phase 1/4: Reproduce: running test, capturing failure...
Phase 2/4: Hypothesize: 3 candidates from stack trace...
Phase 3/4: Verify: isolated root cause in connection pool timeout
Phase 4/4: Fix: patch applied, test passing, PR opened
✓ Delivered: PR #847, fix connection pool timeout in health check
The router pairs a Go agent with a debugging skill, then follows the task through verification and delivery.
ROUTE PLAN EXECUTE VERIFY DELIVER RECORD
┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐
│ /do │───▶│ Task │───▶│Agent │───▶│Tests │───▶│ PR │───▶│Route │
│Router│ │ Plan │ │+Skill│ │Gates │ │Branch│ │Result│
└──────┘ └──────┘ └──────┘ └──────┘ └──────┘ └──────┘
/d routes requests through TypeSafe's Jev classifier. One API call picks the agent, skill, and pipeline — no manifest read into context. Requires Jev; use /do if TypeSafe is not configured.
Setup: install the typesafe MCP plugin and set TYPESAFE_API_KEY in your environment.
> /d fix the flaky test in the payments module
ROUTING (/d): testing-automation-engineer + testing-preferred-patterns
Source: jev (confidence: medium)
Invoking...
Checks require evidence rather than confidence.
| Agent Says | What Happens |
|---|---|
| "Code looks correct, skip tests" | Exit gate requires test output. Blocked. |
| "Trivial change, no verification" | Hook blocks completion without evidence. |
| "Similar to before" | Skill demands case-specific proof. |
| "User is in a hurry" | Protocol overrides time pressure. |
| "I'm confident" | Gate demands exit code, not assertion. |
Hooks run at configured events. Skills state what to verify; blocking78 hooks enforce the checks they cover. Coverage depends on the runtime and tool path.
The content engine researches, drafts in a calibrated voice, checks 397 writing patterns, and adapts finished pieces for each platform. /html produces a self-contained report, slide deck, prototype, chart, or diagram. It needs no coding or setup beyond installation.
Toolkit changes use direct review and relevant checks. Model comparisons can settle specific uncertainties; they are not required for every edit. PHILOSOPHY.md explains the validation policy. what-didnt-work.md records failed experiments, routing reversals, unvalidated A/B citations, disabled lint rules, and program refutations.
The automated nightly evolution loop (/evolve, writes to evolution-reports/) ran regularly through mid-May 2026. It is currently dormant; recent evidence has come from manual PRs instead.
git clone https://github.com/notque/vexjoy-agent.git ~/vexjoy-agent
cd ~/vexjoy-agent
./install.sh
Installs into ~/.claude/ and mirrors into ~/.codex/, ~/.factory/, and ~/.reasonix/ when the runtime command is on PATH or its home directory exists. Choose symlinks for live updates through git pull, or copies for a stable snapshot.
Want only part of the toolkit? Run ./install.sh --configure to pick which skills, agents, and78 hooks install, or copy .local.example/profile.yaml to .local/profile.yaml and edit. No profile file = full install, unchanged behavior. Credit: @thomasvan. Details: .local.example/README.md.
| CLI | Entry Point |
|---|---|
| Claude Code | /do |
| Codex | $do |
| Factory | /do |
| Reasonix | /do |
Jev Auto-Compact plugin (optional, requires TYPESAFE_API_KEY):
claude plugin marketplace add ./plugins/jev-auto-compact
claude plugin install jev-auto-compact@jev-auto-compact -y
Replaces LLM-generated compaction summaries with Jev-judged verbatim pruning. Once context reaches 60%, Jev evaluates each old tool call (keep, truncate result, or drop) and returns the pruned transcript with zero rewriting, in about a second instead of one to three minutes. The threshold matters: every compaction is a cold KV-cache rewrite of the prefix, so compacting every turn multiplies cost. Evidence lives in learning.db (python153 scripts/jev-compact-evidence.py).
Proof it works: python153 scripts/jev-compact-evidence.py prints every compaction from two sources side by side — the plugin's claim and the engine's own compact_boundary record (tokens before/after, duration). A Jev compaction shows as a sub-second engine record next to a matching plugin claim; a built-in LLM compaction shows as a 30–150s record. Rows live in ~/.claude/learning/learning.db (compaction_events, session_usage).
Full setup: docs/start-here.md
Mirrors agents, skills, and supported78 hooks into ~/.codex/. The original six-hook allowlist was correct for Codex v0.114, when tool hooks only intercepted Bash. Current support requires Codex v0.144.1+ and classifies the 62 Claude hook registrations as 26 native, 27 adapter-backed, and 9 unsupported (53 supported). These are registration counts, not unique hook files. The installer also preserves explicit per-subagent model routing for GPT-5.6 Sol by setting the MultiAgent V2 compatibility keys documented in openai/codex#31814.
Codex now exposes apply_patch to tool78 hooks. VexJoy's adapter converts each patch operation into the Write/Edit payload expected by existing guards, but it cannot intercept writes performed through unified_exec, unmatched MCP tools, WebSearch, or other unsupported tool paths. PreCompact and Stop adapters also receive less telemetry than Claude Code: Codex does not provide Claude's conversation_history or session_data. This is expanded compatibility, not full Claude parity.
After install or any hook-definition change, run /hooks in Codex and review the new definitions before trusting them. Codex hash-trusts hook commands and skips changed, unreviewed definitions.
Gemini CLI support removed (deprecated upstream, transitioned to Antigravity CLI); Antigravity support pending CLI maturity. Per Google's transition announcement, Gemini CLI stops serving requests on 2026-06-18 for Google AI Pro / Ultra and free Gemini Code Assist for individuals. Gemini API integrations (image-gen backends, sprite pipeline, GEMINI_API_KEY) are unaffected and stay in the toolkit.
If a prior install mirrored into ~/.gemini/, remove the stale mirrors with:
rm -rf ~/.gemini/skills ~/.gemini/agents ~/.gemini/hooks ~/.gemini/scripts ~/.gemini/antigravity/plugins/vexjoy-agent
Mirrors agents (as "droids"), skills, and all78 hooks into ~/.factory/. Hook config merges into ~/.factory/settings.json with paths rewritten.
Mirrors skills, 153 scripts, and the allowlisted 78 hooks (scripts/reasonix-hooks-allowlist.txt) into ~/.reasonix/ (no agent or custom-command surface, so neither is installed; the /do router rides in as a skill). Reasonix fires only 4 events (PreToolUse, PostToolUse, UserPromptSubmit, Stop), so only hooks for those events are allowlisted. Hook config is written to the hooks key of ~/.reasonix/settings.json in Reasonix's native flat shape (one entry per hook, match regex over the tool name); the generator builds absolute python3 commands, so no path rewrite is applied. MCP/model/permissions in ~/.reasonix/config.json are user-owned and left untouched.
The toolkit supplies its own routing, domain knowledge, methodology, and enforcement. The default system prompt duplicates most of that.
claude --system-prompt "."
Strips built-in tool-use instructions. The toolkit's agents, skills,78 hooks, and CLAUDE.md provide equivalent coverage.
| Layer | Count | Does |
|---|---|---|
| Agents | 43 | Domain knowledge: idiom tables, failure mode catalogs, error-to-fix mappings |
| Skills | 59 | Phased methodology with gates. Can't skip steps. Each phase has exit criteria requiring evidence. |
| Hooks | 78 | Fire on lifecycle events. Block incomplete work. Zero LLM cost. |
| Scripts | 153 | Determinism: test runners, linters, validators. No LLM judgment. |
Full skill catalog: docs/skills.md.
┌─────────────────────────────────────────────────┐
│ SKILL.md │
│ ┌─ Frontmatter ─────────────────────────────┐ │
│ │ triggers, pairs_with, success-criteria │ │
│ └────────────────────────────────────────────┘ │
│ Reference Loading Table (conditional imports) │
│ Phased Instructions (numbered, with gates) │
│ Verification (evidence requirements) │
└─────────────────────────────────────────────────┘
A game built entirely by Claude Code using these agents, skills, and pipelines:
I just want to use it Install, learn /do, done.
I do knowledge work Writing, research, data analysis, moderation, HTML artifacts. No code.
I'm a developer Architecture, extension points, adding agents and skills.
I'm an AI power user Routing tables, pipelines,78 hooks, telemetry DB.
I'm an AI agent Machine-dense inventory. Tables, paths, schemas.
I'm on LinkedIn 🚀 Thought leadership. Agree? 👇
Full design philosophy: PHILOSOPHY.md
One report-only script surfaces upkeep work; it prints a digest and never edits, deletes, or blocks.
python153 scripts/stale-skill-scan.py --top 20 ranks stale skills and agents as pruning candidates. Run it quarterly; see docs/deprecation-template.md.Scheduled work follows the same boundary as everything else: judgment uses agents; repeatable plumbing uses153 scripts.
| Need | Use |
|---|---|
| Run a deterministic command on a schedule | scripts/agent-scheduler.py with runner: "command" |
| Run an agent judgment on a schedule, webhook, or file change | scripts/agent-scheduler.py with the default runner: "claude" |
| Install or remove a user crontab entry safely | scripts/crontab-manager.py |
| Audit shell cron reliability | cron-automation |
| Keep one interactive objective moving until criteria verify | objective-loop |
See CONTRIBUTING.md.
MIT. See LICENSE.
name: workflow
description: "Structured work: multi-phase tasks, feature builds, planning, objective loops, hill climbing."
user-invocable: true
context: fork
allowed-tools:
- Read
- Write
- Edit
- Bash
- Glob
- Grep
- Skill
- Agent
- Task
routing:
force_route: true
not_for: "code review (use review), testing (use testing), security (use security)"
triggers:
- "workflow"
- "multi-phase task"
- "feature design"
- "feature plan"
- "feature implement"
- "build feature end to end"
- "full feature lifecycle"
- "write spec"
- "define requirements"
- "create plan"
- "create tasks"
- "keep working until"
- "iterate until done"
- "drive this to done"
- "make this faster"
- "speed this up"
- "reduce latency"
- "profile and optimize"
- "hill climb on this metric"
- "tidy up"
- "clean up"
- "untangle"
- "reorganize"
- "structured pipeline"
- "phased execution"
category: process
pairs_with:
- review
- testing
- security
- pr-workflowFive modes for structured multi-phase work. Match the request, follow that mode's instructions.
| Request pattern | Mode |
|---|---|
| Feature design/plan/implement/validate/release, end-to-end | Feature Lifecycle |
| Write spec, define requirements, create plan, interview, pause/resume | Planning |
| Keep working until, iterate until done, drive to done | Objective Loop |
| Make faster, speed up, reduce latency, profile, hill climb | Hill Climb |
| All other structured workflows: review, debug, refactor, research, create, explore, upgrade | Ad-Hoc Workflow |
Phase-gated workflow: DESIGN > PLAN > IMPLEMENT > VALIDATE > RELEASE. Each phase must pass its gate before the next begins.
If .feature/ exists, check state: python3 ~/.claude/scripts/feature-state.py status.
Route to the indicated phase.
If no feature state exists, determine entry from intent:
Load the phase reference, then follow it exactly:
| Phase | Reference | Produces |
|---|---|---|
| DESIGN | references/fl-design.md | design.md |
| PLAN | references/fl-plan.md | Wave-ordered task list |
| IMPLEMENT | references/fl-implement.md | Code changes |
| VALIDATE | references/fl-validate.md | Quality gate report |
| RELEASE | references/fl-release.md | Merged PR |
| End-to-end | references/fl-pipeline.md | Full lifecycle |
| State conventions | references/fl-shared.md | -- |
| Error recovery | references/fl-error-handling.md | -- |
State operations use python3 ~/.claude/scripts/feature-state.py only.
Never manipulate state files directly.
Spec writing, plan creation, interviews, ambiguity triage, and session pause/resume. Planning owns specs and saved plans; execution goes through subagent-driven-development or workflow dispatch.
| Signal | Reference |
|---|---|
| Write spec, user stories, define requirements, scope, acceptance criteria | references/pl-spec.md |
| Discuss ambiguities, resolve gray areas, pre-planning discussion | references/pl-pre-plan.md |
| Interview me, depth-first review, "not sure", "where do I start", "poke holes" | references/pl-depth-first-interview.md |
| Implicit ambiguity or unclear implementation choices | references/pl-ambiguity-triage.md |
| Another person holds needed facts or approval | references/pl-human-source-elicitation.md |
| Observation can settle a disputed choice | references/pl-empirical-prototype.md |
| Phase or session transition near | references/pl-context-boundary.md |
| Create plan, task plan, file-backed planning | references/pl-plan-files.md |
| Check plan, validate plan, pre-execution check | references/pl-check.md |
| List plans, show plan, complete plan, manage plans | references/pl-manage.md |
| Pause, save progress, handoff, stopping for now | references/pl-pause.md |
| Resume, continue, pick up where I left off | references/pl-resume.md |
For interviews, batch independent questions into frontier rounds. Ask dependent questions sequentially. Include a recommendation per question.
Iterate-until-verified-done loop. A user states an objective with verifiable done-criteria; each iteration routes one /do cycle, verifies by executing the criteria, and reschedules until verified-done or budget-stop.
Gather from the request: objective statement, DONE-CRITERIA (verifiable checks), iteration budget (default 5), NOT-DONE-YET guardrails (what may never be done to satisfy a criterion).
DONE-CRITERIA types: command (preferred -- deterministic command with expected
exit code/output) or rubric (only when no mechanical check exists -- frozen
at SPEC time, graded by a fresh-context agent).
Write .objective/<slug>/state.md from references/ol-state-file.md. Wakeups
resume from the state file, never conversation memory.
Plan the smallest next step. Route through /do: classify -> route -> dispatch agents -> evaluate. The loop dispatches exclusively through /do -- never edit inline.
Run every done-criterion check. A worker's "passes" claim never substitutes for re-running.
command: run it, paste exit code and output into iteration log.rubric: dispatch fresh-context sub-agent (did NOT produce the work) with
artifact + rubric only. Returns PASS/FAIL with cited evidence.All pass -> final report, STOP. Any unmet -> Phase 5.
Criteria-gaming guard: a criterion may never be satisfied by weakening a hook, gate, test, or safety control. Stop and report the conflict if that is the only visible path.
All pass: stop. Unmet + iterations remain: update state file, call
ScheduleWakeup (delay 270s for active polling, 1200s+ for idle work).
Budget exhausted: honest NOT-DONE report with per-criterion status.
Metric-driven optimization loop. One number moves; everything else stays fixed. Each iteration: hypothesis -> one change -> correctness floor -> re-measure -> accept or revert.
| Field | Required | Default |
|---|---|---|
| METRIC (one number, units, direction) | yes | -- |
| MEASURE (deterministic command) | yes | -- |
| TARGET (value that ends the loop) | yes | -- |
| FLOOR (correctness gate commands, must exit 0) | yes | -- |
| FIXTURE (pinned dataset/workload) | yes | -- |
| Variance tolerance | no | 2x baseline spread |
| Iteration budget | no | 8 |
| Plateau threshold K | no | 3 |
One METRIC per loop. Two numbers with a trade-off: promote one to the FLOOR.
Load references/hc-domain-playbooks.md for pre-filled SPEC blocks per domain
(frame rate, API latency, CI time, bundle size, memory, token cost).
Run MEASURE N times (N >= 5, N >= 10 for wall-clock). Record median and spread. If spread >= target improvement: STOP -- harness too noisy. Report noise sources and offer to stabilize first.
Locate the cost before changing anything. Load references/hc-profiling-tools.md
for per-domain tooling. Guessing at hot spots is the dominant failure mode.
State one hypothesis targeting the profiled hot spot. Make one change. Run FLOOR commands -- revert immediately if any fail.
Run MEASURE N times. Compare median to baseline. Accept only if delta > variance
tolerance. Update ledger (references/hc-ledger.md). If accepted, new baseline.
Target reached: final report. K consecutive non-improving iterations: plateau stop. Budget exhausted: report what worked and what remains.
For structured multi-phase work that does not fit the four modes above. Identify the workflow from the table, load its reference, follow its phases exactly.
Ask first: does this need a multi-agent workflow? Skip the workflow when:
single-file mechanical edit (use quick), one agent satisfies the request
(direct dispatch), lookup/status/count (direct agent). Escalate only when the
request has independent subtasks, needs orthogonal verification, or names
"comprehensive / thorough / adversarial / tournament."
| Pattern | What it does |
|---|---|
| Classify-and-act | Route by type up front; or classify-at-end |
| Fan-out-and-synthesize | Independent agents in parallel, barrier, one synthesizer |
| Adversarial verification | Executor builds, fresh skeptic refutes |
| Generate-and-filter | Over-generate candidates, gate keeps survivors |
| Tournament | N agents attempt same task; pairwise judges pick winner per round |
| Loop-until-done | Repeat until hard completion test passes |
| Quarantine | Read-only triage agent for untrusted content; separate privileged acting agent |
Load the reference for the matched workflow. references/... paths resolve
under ${CLAUDE_SKILL_DIR}.
| Category | Workflow | Reference |
|---|---|---|
| Code Review | Comprehensive multi-wave | references/comprehensive-review.md |
| Debugging | Evidence-based diagnosis | references/systematic-debugging.md |
| Refactoring | Safe refactoring with test gates | references/systematic-refactoring.md |
| Research | Formal research with source gates | references/research-pipeline.md |
| Research | Research to article | references/research-to-article.md |
| Content | Article evaluation | references/article-evaluation-pipeline.md |
| Content | De-AI content | references/de-ai-pipeline.md |
| Content | Documentation | references/doc-pipeline.md |
| Exploration | Codebase exploration | references/explore-pipeline.md |
| Exploration | Multi-perspective analysis | references/do-perspectives.md |
| Creation | Skill creation | references/skill-creation-pipeline.md |
| Creation | Hook development | references/hook-development-pipeline.md |
| Creation | MCP server | references/mcp-pipeline-builder.md |
| Creation | Pipeline scaffolding | references/pipeline-scaffolder.md |
| Creation | Domain research | references/domain-research.md |
| Creation | Chain composition | references/chain-composer.md |
| Creation | Auto-pipeline generation | references/auto-pipeline.md |
| Upgrade | Agent/skill upgrade | references/agent-upgrade.md |
| Upgrade | System upgrade | references/system-upgrade.md |
| Upgrade | Toolkit improvement | references/toolkit-improvement.md |
| Testing | Pipeline test runner | references/pipeline-test-runner.md |
| Testing | Pipeline retro | references/pipeline-retro.md |
| GitHub | Profile rules extraction | references/github-profile-rules.md |
| Orchestration | Task orchestration | references/workflow-orchestrator.md |
| Orchestration | DAG composition | references/dag-composition-patterns.md |
| Orchestration | Compatibility matrix | references/dag-compatibility-matrix.md |
| Orchestration | Common DAG patterns | references/dag-skill-patterns.md |
| Orchestration | DAG examples | references/dag-orchestration-examples.md |
| Orchestration | DAG advanced | references/dag-orchestration-advanced.md |
| Orchestration | Feedback loop | references/feedback-loop-construction.md |
"Workflow" is the canonical term. "Pipeline" is the retained legacy alias --
kept for back-compat in routing keys, meta.name exports, and
pipeline-index.json. Use "workflow" in new prose; do not rename code
identifiers.
| Error | Response |
|---|---|
| Mode ambiguous | Ask the user to clarify intent |
| Phase mismatch | Report current state, suggest correct next phase |
| Missing artifact | Route back to previous phase |
| Noisy harness (hill-climb) | Stop; report spread vs target improvement |
| Budget exhausted | Honest NOT-DONE report with per-criterion status |
| State file missing on wakeup | Report and stop; ask user to restate objective |
评论 (0)
暂无评论,成为第一个评论者吧!