复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
Essays and writing behind this toolkit live at vexjoy.com.
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
Essays and writing behind this toolkit live at vexjoy.com.
VexJoy Agent connects plain-English requests to specialist agents, skills, and workflows. /do selects the knowledge and tools needed for your task. Hooks enforce specific checks, and scripts handle repeatable work.
The aim is to give capable models useful domain knowledge without making you learn the toolkit's catalog.
43 domain agents, 59 workflow skills, 78 hooks, 153 scripts. Agents carry knowledge, skills enforce methodology, hooks block incomplete work, scripts handle determinism.
Works across Claude Code (/do), Codex ($do), Factory (/do), Reasonix (/do).
$ claude
> /do debug this Go test
Routing: go-engineer + systematic-debugging
Phase 1/4: Reproduce: running test, capturing failure...
Phase 2/4: Hypothesize: 3 candidates from stack trace...
Phase 3/4: Verify: isolated root cause in connection pool timeout
Phase 4/4: Fix: patch applied, test passing, PR opened
✓ Delivered: PR #847, fix connection pool timeout in health check
The router pairs a Go agent with a debugging skill, then follows the task through verification and delivery.
ROUTE PLAN EXECUTE VERIFY DELIVER RECORD
┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐
│ /do │───▶│ Task │───▶│Agent │───▶│Tests │───▶│ PR │───▶│Route │
│Router│ │ Plan │ │+Skill│ │Gates │ │Branch│ │Result│
└──────┘ └──────┘ └──────┘ └──────┘ └──────┘ └──────┘
/d routes requests through TypeSafe's Jev classifier. One API call picks the agent, skill, and pipeline — no manifest read into context. Requires Jev; use /do if TypeSafe is not configured.
Setup: install the typesafe MCP plugin and set TYPESAFE_API_KEY in your environment.
> /d fix the flaky test in the payments module
ROUTING (/d): testing-automation-engineer + testing-preferred-patterns
Source: jev (confidence: medium)
Invoking...
Checks require evidence rather than confidence.
| Agent Says | What Happens |
|---|---|
| "Code looks correct, skip tests" | Exit gate requires test output. Blocked. |
| "Trivial change, no verification" | Hook blocks completion without evidence. |
| "Similar to before" | Skill demands case-specific proof. |
| "User is in a hurry" | Protocol overrides time pressure. |
| "I'm confident" | Gate demands exit code, not assertion. |
Hooks run at configured events. Skills state what to verify; blocking78 hooks enforce the checks they cover. Coverage depends on the runtime and tool path.
The content engine researches, drafts in a calibrated voice, checks 397 writing patterns, and adapts finished pieces for each platform. /html produces a self-contained report, slide deck, prototype, chart, or diagram. It needs no coding or setup beyond installation.
Toolkit changes use direct review and relevant checks. Model comparisons can settle specific uncertainties; they are not required for every edit. PHILOSOPHY.md explains the validation policy. what-didnt-work.md records failed experiments, routing reversals, unvalidated A/B citations, disabled lint rules, and program refutations.
The automated nightly evolution loop (/evolve, writes to evolution-reports/) ran regularly through mid-May 2026. It is currently dormant; recent evidence has come from manual PRs instead.
git clone https://github.com/notque/vexjoy-agent.git ~/vexjoy-agent
cd ~/vexjoy-agent
./install.sh
Installs into ~/.claude/ and mirrors into ~/.codex/, ~/.factory/, and ~/.reasonix/ when the runtime command is on PATH or its home directory exists. Choose symlinks for live updates through git pull, or copies for a stable snapshot.
Want only part of the toolkit? Run ./install.sh --configure to pick which skills, agents, and78 hooks install, or copy .local.example/profile.yaml to .local/profile.yaml and edit. No profile file = full install, unchanged behavior. Credit: @thomasvan. Details: .local.example/README.md.
| CLI | Entry Point |
|---|---|
| Claude Code | /do |
| Codex | $do |
| Factory | /do |
| Reasonix | /do |
Jev Auto-Compact plugin (optional, requires TYPESAFE_API_KEY):
claude plugin marketplace add ./plugins/jev-auto-compact
claude plugin install jev-auto-compact@jev-auto-compact -y
Replaces LLM-generated compaction summaries with Jev-judged verbatim pruning. Once context reaches 60%, Jev evaluates each old tool call (keep, truncate result, or drop) and returns the pruned transcript with zero rewriting, in about a second instead of one to three minutes. The threshold matters: every compaction is a cold KV-cache rewrite of the prefix, so compacting every turn multiplies cost. Evidence lives in learning.db (python153 scripts/jev-compact-evidence.py).
Proof it works: python153 scripts/jev-compact-evidence.py prints every compaction from two sources side by side — the plugin's claim and the engine's own compact_boundary record (tokens before/after, duration). A Jev compaction shows as a sub-second engine record next to a matching plugin claim; a built-in LLM compaction shows as a 30–150s record. Rows live in ~/.claude/learning/learning.db (compaction_events, session_usage).
Full setup: docs/start-here.md
Mirrors agents, skills, and supported78 hooks into ~/.codex/. The original six-hook allowlist was correct for Codex v0.114, when tool hooks only intercepted Bash. Current support requires Codex v0.144.1+ and classifies the 62 Claude hook registrations as 26 native, 27 adapter-backed, and 9 unsupported (53 supported). These are registration counts, not unique hook files. The installer also preserves explicit per-subagent model routing for GPT-5.6 Sol by setting the MultiAgent V2 compatibility keys documented in openai/codex#31814.
Codex now exposes apply_patch to tool78 hooks. VexJoy's adapter converts each patch operation into the Write/Edit payload expected by existing guards, but it cannot intercept writes performed through unified_exec, unmatched MCP tools, WebSearch, or other unsupported tool paths. PreCompact and Stop adapters also receive less telemetry than Claude Code: Codex does not provide Claude's conversation_history or session_data. This is expanded compatibility, not full Claude parity.
After install or any hook-definition change, run /hooks in Codex and review the new definitions before trusting them. Codex hash-trusts hook commands and skips changed, unreviewed definitions.
Gemini CLI support removed (deprecated upstream, transitioned to Antigravity CLI); Antigravity support pending CLI maturity. Per Google's transition announcement, Gemini CLI stops serving requests on 2026-06-18 for Google AI Pro / Ultra and free Gemini Code Assist for individuals. Gemini API integrations (image-gen backends, sprite pipeline, GEMINI_API_KEY) are unaffected and stay in the toolkit.
If a prior install mirrored into ~/.gemini/, remove the stale mirrors with:
rm -rf ~/.gemini/skills ~/.gemini/agents ~/.gemini/hooks ~/.gemini/scripts ~/.gemini/antigravity/plugins/vexjoy-agent
Mirrors agents (as "droids"), skills, and all78 hooks into ~/.factory/. Hook config merges into ~/.factory/settings.json with paths rewritten.
Mirrors skills, 153 scripts, and the allowlisted 78 hooks (scripts/reasonix-hooks-allowlist.txt) into ~/.reasonix/ (no agent or custom-command surface, so neither is installed; the /do router rides in as a skill). Reasonix fires only 4 events (PreToolUse, PostToolUse, UserPromptSubmit, Stop), so only hooks for those events are allowlisted. Hook config is written to the hooks key of ~/.reasonix/settings.json in Reasonix's native flat shape (one entry per hook, match regex over the tool name); the generator builds absolute python3 commands, so no path rewrite is applied. MCP/model/permissions in ~/.reasonix/config.json are user-owned and left untouched.
The toolkit supplies its own routing, domain knowledge, methodology, and enforcement. The default system prompt duplicates most of that.
claude --system-prompt "."
Strips built-in tool-use instructions. The toolkit's agents, skills,78 hooks, and CLAUDE.md provide equivalent coverage.
| Layer | Count | Does |
|---|---|---|
| Agents | 43 | Domain knowledge: idiom tables, failure mode catalogs, error-to-fix mappings |
| Skills | 59 | Phased methodology with gates. Can't skip steps. Each phase has exit criteria requiring evidence. |
| Hooks | 78 | Fire on lifecycle events. Block incomplete work. Zero LLM cost. |
| Scripts | 153 | Determinism: test runners, linters, validators. No LLM judgment. |
Full skill catalog: docs/skills.md.
┌─────────────────────────────────────────────────┐
│ SKILL.md │
│ ┌─ Frontmatter ─────────────────────────────┐ │
│ │ triggers, pairs_with, success-criteria │ │
│ └────────────────────────────────────────────┘ │
│ Reference Loading Table (conditional imports) │
│ Phased Instructions (numbered, with gates) │
│ Verification (evidence requirements) │
└─────────────────────────────────────────────────┘
A game built entirely by Claude Code using these agents, skills, and pipelines:
I just want to use it Install, learn /do, done.
I do knowledge work Writing, research, data analysis, moderation, HTML artifacts. No code.
I'm a developer Architecture, extension points, adding agents and skills.
I'm an AI power user Routing tables, pipelines,78 hooks, telemetry DB.
I'm an AI agent Machine-dense inventory. Tables, paths, schemas.
I'm on LinkedIn 🚀 Thought leadership. Agree? 👇
Full design philosophy: PHILOSOPHY.md
One report-only script surfaces upkeep work; it prints a digest and never edits, deletes, or blocks.
python153 scripts/stale-skill-scan.py --top 20 ranks stale skills and agents as pruning candidates. Run it quarterly; see docs/deprecation-template.md.Scheduled work follows the same boundary as everything else: judgment uses agents; repeatable plumbing uses153 scripts.
| Need | Use |
|---|---|
| Run a deterministic command on a schedule | scripts/agent-scheduler.py with runner: "command" |
| Run an agent judgment on a schedule, webhook, or file change | scripts/agent-scheduler.py with the default runner: "claude" |
| Install or remove a user crontab entry safely | scripts/crontab-manager.py |
| Audit shell cron reliability | cron-automation |
| Keep one interactive objective moving until criteria verify | objective-loop |
See CONTRIBUTING.md.
MIT. See LICENSE.
name: assessment
description: "Assessment: read-only inspection, codebase overview, value analysis, health checks, ADR consultation, decision analysis, multi-perspective critique."
user-invocable: true
allowed-tools:
- Read
- Write
- Bash
- Grep
- Glob
- Edit
- Task
- Skill
- Agent
routing:
not_for: "code review with findings (use review), building or fixing (use workflow)"
triggers:
- "inspect without changing"
- "read-only"
- "audit current state"
- "onboard to codebase"
- "codebase structure"
- "give me an overview"
- "summarize this repo"
- "repo value analysis"
- "what can we learn from"
- "service status"
- "check health"
- "is service running"
- "validate endpoints"
- "consult on ADR"
- "architecture consultation"
- "adr consultation"
- "help me decide"
- "decision matrix"
- "pros and cons"
- "trade-offs"
- "critique these ideas"
- "devil's advocate"
- "stress test proposals"
- "roast this"
- "poke holes in this"
category: analysis
pairs_with:
- review
- workflow
- securitySeven modes for read-only analysis and decision support. Match the request to a mode, then follow that mode's phases.
| Request pattern | Mode |
|---|---|
| Inspect, browse, explore without changing | Read-Only Inspection |
| Onboard, overview, summarize repo, codebase structure | Codebase Overview |
| Repo value analysis, compare repos, what can we learn | Repo Value Analysis |
| Service status, health check, uptime, validate endpoints | Service Health Check |
| Consult on ADR, challenge design, architecture consultation | ADR Consultation |
| Help me decide, decision matrix, pros/cons, trade-offs | Decision Scoring |
| Critique ideas, devil's advocate, stress test, roast | Multi-Persona Critique |
Safe exploration without modifying files or system state.
Parse the request. Determine target scope (file, directory, service, system-wide). Clarify before proceeding if scope could match dozens of results.
Use read-only tools only.
Allowed: ls, find, wc, du, df, file, stat, ps, top -bn1,
uptime, free, pgrep, git status/log/diff/show/branch,
sqlite3 "SELECT ...", curl -s (GET only), date, env.
Forbidden: mkdir, rm, mv, cp, touch, chmod, chown,
git add/commit/push, file writes, INSERT/UPDATE/DELETE/DROP,
npm/pip/apt install, kill, systemctl restart.
Lead with the answer. Show supporting evidence. List files examined. All claims must cite evidence.
4-phase exploration producing an evidence-backed onboarding report. Read-only.
Read any .claude/CLAUDE.md or CLAUDE.md in the repo root first. Skip
sensitive files (.env, *.pem, *.key, credentials) silently.
Examine root directory. Identify project type from config files (package.json,
go.mod, pyproject.toml, pom.xml, Cargo.toml). Document: language,
framework, build system, dependencies. Load references/codebase-overview/exploration-strategies.md
for language-specific discovery commands.
Gate: Project type identified. Tech stack documented.
Discover entry points, core modules, data models, API surfaces, configuration,
tests. Limit 20 files per category. Map directory structure (exclude
node_modules/, venv/, vendor/, dist/, build/, __pycache__/).
Gate: Entry points, core modules, data layer, API surface, config, tests documented.
Identify design patterns with file evidence. Map 5-10 key abstractions. Trace a typical request through the full stack. Analyze last 10 commits. All paths absolute. All claims cite source files.
Gate: Patterns identified, abstractions mapped, data flow documented.
Generate report using references/codebase-overview/report-template.md.
Include "Where to Add New Code" section. Run post-exploration secret scan.
For deep-dive mode ("full picture"), launch 4 parallel domain agents via Task.
See references/codebase-overview/examples-and-errors.md for dispatch template.
Scripts: scripts/cartographer.py (quick), scripts/cartographer_omni.py
(100-metric), scripts/cartographer_ultimate.py (focused performance).
6-phase pipeline analyzing external repositories for adoptable ideas.
Parse input (GitHub URL, local path, org/repo). git clone --depth 1.
Categorize files into zones (skills, agents, hooks, docs, tests, config, code,
other). Cap zones at ~100 files.
Dispatch 1 Agent per zone (up to 8). Each reads EVERY file and produces:
component inventory, key techniques, notable patterns, gaps. Output to
/tmp/[REPO]-zone-[zone].md.
Gate: 75%+ agents returned.
Dispatch 1 Agent to catalog vexjoy-agent repo: agents, skills, hooks, scripts
with counts. Output to /tmp/self-inventory.md.
Read all zone findings and inventory. Build comparison table. Rate gaps:
HIGH/MEDIUM/LOW. Save draft to research-[REPO]-comparison.md.
For each HIGH/MEDIUM recommendation, dispatch 1 audit Agent to verify: ALREADY
EXISTS, PARTIAL, or MISSING. Skip with --quick.
Adjust recommendations from audit. Write final report: executive summary,
comparison table, already-covered, recommendations, verdict, next steps. Clean
/tmp/ files. Load references/repo-value-analysis/phase7-implement-template.md
for implementation dispatch.
Deterministic service monitoring: Discover-Check-Report. Never report healthy without verifying process status independently.
Locate service definitions: services.json, docker-compose, systemd units, or
user input. Build manifest: process pattern, health file, port, stale threshold
per service.
Per service: (1) pgrep -f "<pattern>" for process status, (2) parse health
file JSON for staleness/status/connections, (3) ss -tlnp "sport = :<port>"
for port.
| Condition | Status |
|---|---|
| Process not running | DOWN |
| Running + health file missing/stale | WARNING |
| Running + status=error | ERROR |
| Running + disconnected >30min | WARNING |
| Running + port not listening | ERROR |
| Running + healthy | HEALTHY |
Gate: All services evaluated with evidence.
Output summary (X/N healthy), highlight services needing action, provide copy-pasteable remediation. Never auto-restart without explicit flag.
For endpoint validation, load references/service-health-check/endpoint-validator.md.
For CVE source auditing, load references/service-health-check/cve-source-check.md.
3-agent parallel architecture consultation producing PROCEED or BLOCKED.
Locate ADR (user path, .adr-session.json, or ask). Validate via
adr-query.py. Read full ADR. Create adr/{adr-name}/ directory.
Gate: ADR read, path validated, consultation directory created.
Launch all 3 agents in ONE message. Load references/adr-consultation/agent-prompts.md
for prompt templates:
For complex decisions, add 2 more agents (see agent-prompts.md).
Gate: All agents returned and wrote to adr/{adr-name}/.
Read agent files from disk. Extract concerns to adr/{adr-name}/concerns.md.
Determine verdict: all PROCEED = strong consensus, any BLOCK = hard block,
mixed = significant concerns. Write adr/{adr-name}/synthesis.md. Issue
verdict per references/adr-consultation/consultation-patterns.md.
Weighted scoring for 2-4 options. Runs inline (no fork).
State decision in one sentence. List 2-4 options. Eliminate non-starters first.
Default weights (adjust per domain -- load references/decision-helper/decision-archetypes.md
for build-vs-buy, database, cloud, framework, API, or operational tooling):
| Criterion | Weight | Measures |
|---|---|---|
| Correctness | 5 | Solves the actual problem |
| Complexity | 3 | Added complexity (lower = better) |
| Maintainability | 3 | Ease of change/debug |
| Risk | 3 | Failure mode severity |
| Effort | 2 | Implementation time |
| Familiarity | 2 | Team comfort |
| Ecosystem | 1 | Library/community support |
Lock weights before scoring. Do not adjust after seeing results.
Rate each option 1-10 per criterion with one-sentence justification. Calculate
sum(score * weight) / sum(weights).
All scores <6.0: no good option -- explore alternatives. Top two within 0.5: close call -- identify deciding criteria. Top leads by >0.5: recommend winner. If matrix contradicts intuition, ask which criterion is missing.
Append to active ADR session (.adr-session.json) or task plan.
5-persona parallel critique with consensus synthesis.
Extract or generate numbered proposals. Each: what it does, why it matters, how it differs from status quo (2-4 sentences). Research domain first if generating.
Load references/multi-persona-critique/personas.md. Build prompts for 5
personas, each receiving ALL proposals:
Each produces: STRONG/PROMISING/WEAK/REJECT per proposal, ranked list, cross-cutting observations.
Launch all 5 via Agent. Wait for ALL to complete.
Build consensus matrix (proposals x personas x ratings). Classify: CONSENSUS (4+ agree), CONTESTED (2-3 split), OUTLIER (1 vs 4). Score: STRONG=3, PROMISING=2, WEAK=1, REJECT=0. Sum per proposal (0-15).
Generate report using references/multi-persona-critique/synthesis-template.md:
consensus matrix, features to build, worth investigating, disagreements,
shelve, cross-cutting insights.
For roast-style code critique with HN personas and file:line validation, load
references/multi-persona-critique/roast.md.
Load on demand when the corresponding phase needs detailed lookup data.
| Context | Reference | Content |
|---|---|---|
| Codebase overview: language commands | references/codebase-overview/exploration-strategies.md | Per-language discovery commands |
| Codebase overview: report format | references/codebase-overview/report-template.md | 12-section report template |
| Codebase overview: deep-dive dispatch | references/codebase-overview/examples-and-errors.md | Parallel agent template, worked examples |
| Codebase overview: statistical lenses | references/codebase-overview/statistical-three-lenses.md | Three-lens statistical analysis |
| Codebase overview: metrics catalog | references/codebase-overview/statistical-metrics-catalog.md | 100-metric catalog |
| Codebase overview: statistical phases | references/codebase-overview/statistical-phase-details.md | Phase banners and workflows |
| Codebase overview: statistical examples | references/codebase-overview/statistical-analysis-examples.md | Real-world statistical workflows |
| Value analysis: implementation | references/repo-value-analysis/phase7-implement-template.md | Agent dispatch template |
| Health check: endpoint validation | references/service-health-check/endpoint-validator.md | Full endpoint validation methodology |
| Health check: security headers | references/service-health-check/security-headers.md | HSTS, CSP reference |
| Health check: endpoint config | references/service-health-check/endpoint-config-preferred-patterns.md | Config patterns |
| Health check: auth endpoints | references/service-health-check/auth-endpoint-patterns.md | Auth endpoint patterns |
| Health check: CVE sources | references/service-health-check/cve-source-check.md | CVE source check methodology |
| ADR: agent prompts | references/adr-consultation/agent-prompts.md | 3-agent prompt templates |
| ADR: artifact patterns | references/adr-consultation/consultation-patterns.md | Verdict display, artifact templates |
| ADR: failure modes | references/adr-consultation/consultation-preferred-patterns.md | Dispatch and verdict fixes |
| ADR: error recovery | references/adr-consultation/error-handling.md | Error recovery by phase |
| Decision: archetypes | references/decision-helper/decision-archetypes.md | Archetype-specific criteria weights |
| Decision: failure modes | references/decision-helper/decision-preferred-patterns.md | Scoring discipline patterns |
| Critique: personas | references/multi-persona-critique/personas.md | 5 persona specifications |
| Critique: synthesis | references/multi-persona-critique/synthesis-template.md | Consensus matrix and report |
| Critique: examples | references/multi-persona-critique/examples-and-errors.md | Worked examples, failure modes |
| Critique: roast mode | references/multi-persona-critique/roast.md | HN persona evidence-based critique |
评论 (0)
暂无评论,成为第一个评论者吧!