SkillAtlasSkill 详情

forensics

Essays and writing behind this toolkit live at vexjoy.com.

审核状态:已审核Quality 72Security 78

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月31日

VexJoy Agent

VexJoy Agent

Essays and writing behind this toolkit live at vexjoy.com.

AI agents skip steps.

"Looks correct" replaces running tests. "Trivial change" replaces verification. The agent confidently ships broken code because nothing structurally prevented it from skipping the work.

Harnesses have a second problem: given only a skill list, they do not route eagerly enough, or correctly enough. Good skills sit unused. So this toolkit connects the skills, agents, and workflows we want directly into the harness, automatically. You don't have to understand what is here. Say what you want in plain English and you get all the value we have put into it: the right specialist with the right methodology, behind gates that demand exit codes, not assertions.

44 domain agents, 122 workflow skills, 78 hooks, 136 scripts. Agents carry knowledge, skills enforce methodology, hooks block incomplete work, scripts handle determinism.

Works across Claude Code (/do), Codex ($do), Factory (/do), Reasonix (/do).

What It Looks Like

$ claude

> /do debug this Go test

  Routing: go-engineer + systematic-debugging
  Phase 1/4: Reproduce: running test, capturing failure...
  Phase 2/4: Hypothesize: 3 candidates from stack trace...
  Phase 3/4: Verify: isolated root cause in connection pool timeout
  Phase 4/4: Fix: patch applied, test passing, PR opened

  ✓ Delivered: PR #847, fix connection pool timeout in health check

The router reads intent, picks a Go agent paired with a debugging skill, and runs the full lifecycle. You typed one sentence. The system did the rest.

The Pipeline

  ROUTE        PLAN         EXECUTE      VERIFY       DELIVER      RECORD
 ┌──────┐    ┌──────┐    ┌──────┐    ┌──────┐    ┌──────┐    ┌──────┐
 │ /do  │───▶│ Task │───▶│Agent │───▶│Tests │───▶│  PR  │───▶│Route │
 │Router│    │ Plan │    │+Skill│    │Gates │    │Branch│    │Result│
 └──────┘    └──────┘    └──────┘    └──────┘    └──────┘    └──────┘

Anti-Rationalization

This is the single thing that separates it from "agent with a system prompt."

Agent SaysWhat Happens
"Code looks correct, skip tests"Exit gate requires test output. Blocked.
"Trivial change, no verification"Hook blocks completion without evidence.
"Similar to before"Skill demands case-specific proof.
"User is in a hurry"Protocol overrides time pressure.
"I'm confident"Gate demands exit code, not assertion.

Hooks fire automatically. Gates block completion. Skills encode counter-arguments at every skip-worthy step. The agent verifies or it doesn't finish.

For what I do, the difference is enormous. If you're doing simple single-file edits, maybe less so.

Knowledge Work Is First-Class

The same routing serves knowledge work. The content engine researches, drafts in a calibrated voice, validates against 397 AI patterns, and repurposes finished pieces for each platform. /html turns any request into a single self-contained HTML file: report, slide deck, prototype, data viz, diagram. Non-engineers who try the toolkit consistently name the HTML artifacts as the thing they love. No code, no setup beyond the installer.

It Proves Its Own Changes

Changes to the toolkit itself ship with evidence. New skills get blind A/B tests against a no-skill baseline before merge. Routing and writing-standard decisions carry measured verdicts; PHILOSOPHY.md cites the numbers. Experiments that lost go into the negative-results registry, what-didnt-work.md; the registry now covers routing reversals, unvalidated A/B citations, and disabled lint rules alongside the original program refutations.

The automated nightly evolution loop (/evolve, writes to evolution-reports/) ran regularly through mid-May 2026. It is currently dormant; recent evidence has come from manual PRs instead.

Installation

git clone https://github.com/notque/vexjoy-agent.git ~/vexjoy-agent
cd ~/vexjoy-agent
./install.sh

Links into ~/.claude/ and mirrors into ~/.codex/, ~/.factory/, ~/.reasonix/ — each mirror only when that runtime is detected (its command on PATH or its home dir already exists). The installer asks symlink (live updates via git pull) or copy (stable snapshot).

Want only part of the toolkit? Run ./install.sh --configure to pick which skills, agents, and hooks install, or copy .local.example/profile.yaml to .local/profile.yaml and edit. No profile file = full install, unchanged behavior. Credit: @thomasvan. Details: .local.example/README.md.

CLIEntry Point
Claude Code/do
Codex$do
Factory/do
Reasonix/do

Full setup: docs/start-here.md

Codex CLI Parity

Mirrors agents, skills, and supported hooks into ~/.codex/. The original six-hook allowlist was correct for Codex v0.114, when tool hooks only intercepted Bash. Current support requires Codex v0.144.1+ and classifies the 74 Claude hook registrations as 26 native, 35 adapter-backed, and 13 unsupported (61 supported). These are registration counts, not unique hook files. The installer also preserves explicit per-subagent model routing for GPT-5.6 Sol by setting the MultiAgent V2 compatibility keys documented in openai/codex#31814.

Codex now exposes apply_patch to tool hooks. VexJoy's adapter converts each patch operation into the Write/Edit payload expected by existing guards, but it cannot intercept writes performed through unified_exec, unmatched MCP tools, WebSearch, or other unsupported tool paths. PreCompact and Stop adapters also receive less telemetry than Claude Code: Codex does not provide Claude's conversation_history or session_data. This is expanded compatibility, not full Claude parity.

After install or any hook-definition change, run /hooks in Codex and review the new definitions before trusting them. Codex hash-trusts hook commands and skips changed, unreviewed definitions.

Gemini CLI / Antigravity CLI Support (removed)

Gemini CLI support removed (deprecated upstream, transitioned to Antigravity CLI); Antigravity support pending CLI maturity. Per Google's transition announcement, Gemini CLI stops serving requests on 2026-06-18 for Google AI Pro / Ultra and free Gemini Code Assist for individuals. Gemini API integrations (image-gen backends, sprite pipeline, GEMINI_API_KEY) are unaffected and stay in the toolkit.

If a prior install mirrored into ~/.gemini/, remove the stale mirrors with:

rm -rf ~/.gemini/skills ~/.gemini/agents ~/.gemini/hooks ~/.gemini/scripts ~/.gemini/antigravity/plugins/vexjoy-agent
Factory CLI Support

Mirrors agents (as "droids"), skills, and all hooks into ~/.factory/. Hook config merges into ~/.factory/settings.json with paths rewritten.

Reasonix Support

Mirrors skills, scripts, and the allowlisted hooks (scripts/reasonix-hooks-allowlist.txt) into ~/.reasonix/ (no agent or custom-command surface, so neither is installed; the /do router rides in as a skill). Reasonix fires only 4 events (PreToolUse, PostToolUse, UserPromptSubmit, Stop), so only hooks for those events are allowlisted. Hook config is written to the hooks key of ~/.reasonix/settings.json in Reasonix's native flat shape (one entry per hook, match regex over the tool name); the generator builds absolute python3 commands, so no path rewrite is applied. MCP/model/permissions in ~/.reasonix/config.json are user-owned and left untouched.

Token-saving mode

The toolkit supplies its own routing, domain knowledge, methodology, and enforcement. The default system prompt duplicates most of that.

claude --system-prompt "."

Strips built-in tool-use instructions. The toolkit's agents, skills, hooks, and CLAUDE.md provide equivalent coverage.

Four Layers

LayerCountDoes
Agents44Domain knowledge: idiom tables, failure mode catalogs, error-to-fix mappings
Skills122Phased methodology with gates. Can't skip steps. Each phase has exit criteria requiring evidence.
Hooks78Fire on lifecycle events. Block incomplete work. Zero LLM cost.
Scripts136Determinism: test runners, linters, validators. No LLM judgment.

Full skill catalog: docs/skills.md.

┌─────────────────────────────────────────────────┐
│  SKILL.md                                       │
│  ┌─ Frontmatter ─────────────────────────────┐  │
│  │ triggers, pairs_with, success-criteria     │  │
│  └────────────────────────────────────────────┘  │
│  Reference Loading Table (conditional imports)   │
│  Phased Instructions (numbered, with gates)      │
│  Verification (evidence requirements)            │
└─────────────────────────────────────────────────┘

Built with the Toolkit

A game built entirely by Claude Code using these agents, skills, and pipelines:

Choose Your Path

I just want to use it Install, learn /do, done.

I do knowledge work Writing, research, data analysis, moderation, HTML artifacts. No code.

I'm a developer Architecture, extension points, adding agents and skills.

I'm an AI power user Routing tables, pipelines, hooks, telemetry DB.

I'm an AI agent Machine-dense inventory. Tables, paths, schemas.

I'm on LinkedIn 🚀 Thought leadership. Agree? 👇

Philosophy

  • Zero-expertise operation. Say what you want. The system classifies, dispatches, enforces, delivers.
  • LLMs orchestrate, programs execute. Deterministic work belongs to scripts. LLM judgment handles design decisions, diagnosis, review.
  • Density. Every word carries instruction, rule, or decision. Cut everything else.
  • Breadth over depth. Right context ensures correctness. Unfocused context adds cost.
  • Structural enforcement. Exit codes enforce what instructions can't. Quality gates are automated, not advisory.
  • Everything pipelines. Complex work decomposes into phases. Phases have gates. Gates prevent cascading failures.

Full design philosophy: PHILOSOPHY.md

Maintenance

One report-only script surfaces upkeep work; it prints a digest and never edits, deletes, or blocks.

  • python3 scripts/stale-skill-scan.py --top 20 ranks stale skills and agents as pruning candidates. Run it quarterly; see docs/deprecation-template.md.

Scheduled work follows the same boundary as everything else: judgment uses agents; repeatable plumbing uses scripts.

NeedUse
Run a deterministic command on a schedulescripts/agent-scheduler.py with runner: "command"
Run an agent judgment on a schedule, webhook, or file changescripts/agent-scheduler.py with the default runner: "claude"
Install or remove a user crontab entry safelyscripts/crontab-manager.py
Audit shell cron reliabilitycron-automation
Keep one interactive objective moving until criteria verifyobjective-loop

Contributing

See CONTRIBUTING.md.

License

MIT. See LICENSE.

数据与 AI

中风险

  • 来源需自行核对维护者身份。
  • 未检测到明显脚本安装指令。
  • 可能需要外部 token、网络权限或第三方服务。
  • 未检测到高风险命令。
  • 扫描发现:1 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/process/forensics" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/process/forensics" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/process/forensics" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/process/forensics" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/process/forensics" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: forensics
description: "Post-mortem diagnostic analysis of failed workflows."
user-invocable: false
command: /forensics
allowed-tools:
  - Read
  - Grep
  - Glob
routing:
  triggers:
    - forensics
    - what went wrong
    - why did this fail
    - stuck loop
    - diagnose workflow
    - post-mortem
    - workflow failure
    - session crashed
    - why is this stuck
    - investigate failure
    - "why did this break"
    - "incident review"
  pairs_with:
    - workflow
    - planning
  complexity: Medium
  category: process

Forensics Skill

Investigate failed or stuck workflows through post-mortem analysis of git history, plan files, and session artifacts. Forensics answers "what went wrong and why" -- it detects workflow-level failures that individual tool errors don't reveal.

Key distinction: A tool error is "ruff found 3 lint errors." A workflow failure is "the agent entered a fix/retry loop editing the same file 5 times and never progressed." The harness surfaces tool-level errors. Forensics handles workflow-level patterns.

Reference Loading

TaskLoad
Collecting git evidence, running git log commands, scrubbing credentialsreferences/evidence-collection.md
Identifying failure type from symptoms, causal chain analysisreferences/failure-signatures.md
Running any of the 5 anomaly detectors, scoring confidencereferences/detectors.md

Reference Loading Table

SignalLoad These FilesWhy
Phase 2 DETECT: running the 5 anomaly detectorsdetectors.mdLoads detailed guidance from detectors.md.
Phase 1 GATHER: git extraction, loop queries, credential scrubbingevidence-collection.mdLoads detailed guidance from evidence-collection.md.
matching observed symptoms to the 5 failure typesfailure-signatures.mdLoads detailed guidance from failure-signatures.md.

Instructions

This is a read-only diagnostic. The tool restriction to Read/Grep/Glob enforces this at the platform level. A diagnostic tool that modifies state destroys the evidence it needs to analyze -- forensics examines, it does not fix. Even when the user asks you to fix what you find, complete the report and recommend remediation instead. The wrong fix applied automatically can destroy work.

Phase 1: GATHER

Goal: Collect the raw evidence needed for anomaly detection. Determine what branch, plan, and time range to analyze.

Step 1: Identify the investigation target

Accept the target from one of these sources (in priority order):

  1. Explicit branch: User specifies a branch name to investigate
  2. Current branch: Use the current git branch if no branch specified
  3. Explicit plan: User points to a specific task_plan.md

Before analysis, read the repository's CLAUDE.md if present. Repository conventions inform what "normal" looks like (e.g., expected branch patterns, required artifacts).

Step 2: Locate the plan file

Search for the plan that governed the workflow:

  • Check task_plan.md in the repository root
  • Check .feature/state/plan/ for feature plans
  • Check plan/active/ for workflow-orchestrator plans

Record whether a plan exists. If no plan is found, note this -- it limits scope drift and abandoned work detection but does not block the investigation. Three of the five detectors (stuck loop, crash/interruption, and degraded abandoned work) still function without a plan, so never skip analysis because no plan file was found.

Step 3: Collect git history

Read the git log for the target branch. Extract:

  • Commit hashes, messages, timestamps, and files changed
  • The branch's divergence point from main/master

Use Grep to search git log output for patterns. Focus on:

  • Commits on this branch since divergence from the base branch
  • File change frequency across commits
  • Commit message patterns (similarity, repetition)

If the branch has hundreds of commits, focus on the most recent 50 and note the truncation in the final report.

Step 4: Check working tree state

Examine the current state:

  • Are there uncommitted changes? (look for modified/untracked indicators)
  • Are there orphaned .claude/worktrees/ directories?
  • Is there an active task_plan.md with incomplete phases?

See references/evidence-collection.md for concrete git commands for each evidence type: log extraction, loop detection queries, timestamp analysis, and credential scrubbing patterns.

GATE: Evidence collected. At minimum: git history available, branch identified. Proceed to DETECT only when evidence gathering is complete.


Phase 2: DETECT

Goal: Run all 5 anomaly detectors against the collected evidence. Always run every detector -- anomalies are often correlated (a stuck loop causes missing artifacts causes abandoned work), so partial analysis misses the causal chain. Each detector produces zero or more findings, and every finding must include a confidence level (High/Medium/Low) because false positives erode trust.

See references/detectors.md for full detector specifications: confidence scoring tables, false positive guidance, and per-detector skip conditions when no plan file exists. See references/failure-signatures.md for observable patterns per failure type, detection commands, and causal chain analysis when multiple detectors fire.

Run detectors 1-5 in order: Stuck Loop, Missing Artifacts, Abandoned Work, Scope Drift, Crash/Interruption.

GATE: All 5 detectors have run. Each produced zero or more findings with confidence levels. Proceed to REPORT.


Phase 3: REPORT

Goal: Compile findings into a structured diagnostic report with root cause hypothesis and remediation recommendations. Every claim in the report must trace to specific evidence -- a forensics report without evidence is an opinion piece, not a diagnostic.

Step 1: Scrub sensitive content

Before assembling the report, scan all evidence strings for:

  • API keys, tokens, passwords (patterns: sk-, ghp_, token=, password=, secret=, key=, bearer tokens, base64-encoded credentials)
  • Absolute home directory paths

Replace sensitive values with [REDACTED] and home paths with ~/. Treat all credential-shaped strings as real -- you cannot determine whether a credential is live from its format alone. Reports may be shared or logged, so a leaked credential in a forensics report is worse than the original workflow failure. Redact paths in every report regardless of audience; it costs nothing and prevents future exposure.

Step 2: Compile anomaly table

Order findings by confidence (High first, then by detector number) so the reader gets the strongest signals first:

## Forensics Report: [branch name or session identifier]

### Anomalies Detected
| # | Type | Confidence | Description |
|---|------|------------|-------------|
| 1 | [type] | [High/Medium/Low] | [description with evidence] |
| 2 | [type] | [High/Medium/Low] | [description with evidence] |

If no anomalies detected:

### Anomalies Detected
No anomalies detected. The workflow appears to have executed normally.

Step 3: Synthesize root cause hypothesis

Connect the anomalies into a coherent narrative. Look for causal chains:

  • Stuck loop + scope drift = agent tried to fix a problem, drifted into unrelated files looking for the root cause
  • Missing artifacts + abandoned work = session crashed before producing outputs
  • Crash/interruption + stuck loop = agent exhausted retries and was terminated

The hypothesis must be specific, testable, and grounded in evidence from the anomaly findings -- never speculate beyond what the data supports:

  • BAD: "Something went wrong during execution"
  • GOOD: "Agent entered a lint fix loop on server.go (4 consecutive commits with 'fix lint' messages), which consumed the session's context budget before Phase 3 VERIFY could execute, leaving test artifacts missing"

Step 4: Recommend remediation

Provide specific, actionable recommendations. Each recommendation should reference the anomaly it addresses. Remediation is advisory text only -- never execute fixes, even if the user asks. Remediation requires understanding intent, not just detecting anomalies.

Anomaly TypeTypical Remediation
Stuck loopIdentify the root cause of the loop (often a lint/type error the agent can't resolve). Fix manually, then resume from the last successful phase.
Missing artifactsRe-run the phase that failed to produce artifacts. Check if the phase definition is clear enough for the executor.
Abandoned workResume from the last completed phase. Check .debug-session.md or plan status for where to pick up.
Scope driftReview out-of-scope changes for necessity. Revert unrelated changes. Re-scope the plan if the drift was needed.
Crash/interruptionCheck for uncommitted changes worth preserving. Clean up orphaned worktrees. Resume from last committed state.

Step 5: Format final report

Include relevant git log excerpts, file snippets, and timestamps as evidence for every anomaly. Show git hashes, timestamps, and file paths rather than making unsupported assertions.

================================================================
 FORENSICS REPORT: [branch/session identifier]
================================================================

 Scan completed: [timestamp]
 Branch: [branch name]
 Commits analyzed: [count]
 Plan file: [path or "not found"]

================================================================
 ANOMALIES
================================================================

 | # | Type | Confidence | Description |
 |---|------|------------|-------------|
 | ... | ... | ... | ... |

================================================================
 ROOT CAUSE HYPOTHESIS
================================================================

 [Narrative connecting anomalies into causal explanation]

================================================================
 RECOMMENDED REMEDIATION
================================================================

 1. [Specific action referencing anomaly #N]
 2. [Specific action referencing anomaly #N]

================================================================
 EVIDENCE
================================================================

 [Relevant git log excerpts, file snippets, timestamps]
 [All paths redacted, credentials scrubbed]

================================================================

GATE: Report is complete, scrubbed, and formatted. Deliver to user.


Error Handling

ErrorCauseSolution
No git history on branchBranch has zero commits or just forkedReport "insufficient evidence" -- forensics needs commit history to analyze
No plan file foundWorkflow ran without a planNote limitation in report. Detectors 2 (missing artifacts), 3 (abandoned work), and 4 (scope drift) operate in degraded mode or skip. Detectors 1 (stuck loop) and 5 (crash) still function.
Worktree access failsOrphaned worktree with broken symlinksReport the orphaned worktree as crash/interruption evidence. Do not attempt cleanup.
Git log too largeLong-lived branch with hundreds of commitsFocus analysis on the most recent 50 commits. Note truncation in report.
Ambiguous branch targetUser request doesn't clearly identify which branchAsk: "Which branch should I investigate? Current branch is [X]."

References

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!