SkillAtlasSkill 详情

fact-check

Essays and writing behind this toolkit live at vexjoy.com.

审核状态:已审核Quality 80Security 92

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月31日

VexJoy Agent

VexJoy Agent

Essays and writing behind this toolkit live at vexjoy.com.

AI agents skip steps.

"Looks correct" replaces running tests. "Trivial change" replaces verification. The agent confidently ships broken code because nothing structurally prevented it from skipping the work.

Harnesses have a second problem: given only a skill list, they do not route eagerly enough, or correctly enough. Good skills sit unused. So this toolkit connects the skills, agents, and workflows we want directly into the harness, automatically. You don't have to understand what is here. Say what you want in plain English and you get all the value we have put into it: the right specialist with the right methodology, behind gates that demand exit codes, not assertions.

44 domain agents, 122 workflow skills, 78 hooks, 136 scripts. Agents carry knowledge, skills enforce methodology, hooks block incomplete work, scripts handle determinism.

Works across Claude Code (/do), Codex ($do), Factory (/do), Reasonix (/do).

What It Looks Like

$ claude

> /do debug this Go test

  Routing: go-engineer + systematic-debugging
  Phase 1/4: Reproduce: running test, capturing failure...
  Phase 2/4: Hypothesize: 3 candidates from stack trace...
  Phase 3/4: Verify: isolated root cause in connection pool timeout
  Phase 4/4: Fix: patch applied, test passing, PR opened

  ✓ Delivered: PR #847, fix connection pool timeout in health check

The router reads intent, picks a Go agent paired with a debugging skill, and runs the full lifecycle. You typed one sentence. The system did the rest.

The Pipeline

  ROUTE        PLAN         EXECUTE      VERIFY       DELIVER      RECORD
 ┌──────┐    ┌──────┐    ┌──────┐    ┌──────┐    ┌──────┐    ┌──────┐
 │ /do  │───▶│ Task │───▶│Agent │───▶│Tests │───▶│  PR  │───▶│Route │
 │Router│    │ Plan │    │+Skill│    │Gates │    │Branch│    │Result│
 └──────┘    └──────┘    └──────┘    └──────┘    └──────┘    └──────┘

Anti-Rationalization

This is the single thing that separates it from "agent with a system prompt."

Agent SaysWhat Happens
"Code looks correct, skip tests"Exit gate requires test output. Blocked.
"Trivial change, no verification"Hook blocks completion without evidence.
"Similar to before"Skill demands case-specific proof.
"User is in a hurry"Protocol overrides time pressure.
"I'm confident"Gate demands exit code, not assertion.

Hooks fire automatically. Gates block completion. Skills encode counter-arguments at every skip-worthy step. The agent verifies or it doesn't finish.

For what I do, the difference is enormous. If you're doing simple single-file edits, maybe less so.

Knowledge Work Is First-Class

The same routing serves knowledge work. The content engine researches, drafts in a calibrated voice, validates against 397 AI patterns, and repurposes finished pieces for each platform. /html turns any request into a single self-contained HTML file: report, slide deck, prototype, data viz, diagram. Non-engineers who try the toolkit consistently name the HTML artifacts as the thing they love. No code, no setup beyond the installer.

It Proves Its Own Changes

Changes to the toolkit itself ship with evidence. New skills get blind A/B tests against a no-skill baseline before merge. Routing and writing-standard decisions carry measured verdicts; PHILOSOPHY.md cites the numbers. Experiments that lost go into the negative-results registry, what-didnt-work.md; the registry now covers routing reversals, unvalidated A/B citations, and disabled lint rules alongside the original program refutations.

The automated nightly evolution loop (/evolve, writes to evolution-reports/) ran regularly through mid-May 2026. It is currently dormant; recent evidence has come from manual PRs instead.

Installation

git clone https://github.com/notque/vexjoy-agent.git ~/vexjoy-agent
cd ~/vexjoy-agent
./install.sh

Links into ~/.claude/ and mirrors into ~/.codex/, ~/.factory/, ~/.reasonix/ — each mirror only when that runtime is detected (its command on PATH or its home dir already exists). The installer asks symlink (live updates via git pull) or copy (stable snapshot).

Want only part of the toolkit? Run ./install.sh --configure to pick which skills, agents, and hooks install, or copy .local.example/profile.yaml to .local/profile.yaml and edit. No profile file = full install, unchanged behavior. Credit: @thomasvan. Details: .local.example/README.md.

CLIEntry Point
Claude Code/do
Codex$do
Factory/do
Reasonix/do

Full setup: docs/start-here.md

Codex CLI Parity

Mirrors agents, skills, and supported hooks into ~/.codex/. The original six-hook allowlist was correct for Codex v0.114, when tool hooks only intercepted Bash. Current support requires Codex v0.144.1+ and classifies the 74 Claude hook registrations as 26 native, 35 adapter-backed, and 13 unsupported (61 supported). These are registration counts, not unique hook files. The installer also preserves explicit per-subagent model routing for GPT-5.6 Sol by setting the MultiAgent V2 compatibility keys documented in openai/codex#31814.

Codex now exposes apply_patch to tool hooks. VexJoy's adapter converts each patch operation into the Write/Edit payload expected by existing guards, but it cannot intercept writes performed through unified_exec, unmatched MCP tools, WebSearch, or other unsupported tool paths. PreCompact and Stop adapters also receive less telemetry than Claude Code: Codex does not provide Claude's conversation_history or session_data. This is expanded compatibility, not full Claude parity.

After install or any hook-definition change, run /hooks in Codex and review the new definitions before trusting them. Codex hash-trusts hook commands and skips changed, unreviewed definitions.

Gemini CLI / Antigravity CLI Support (removed)

Gemini CLI support removed (deprecated upstream, transitioned to Antigravity CLI); Antigravity support pending CLI maturity. Per Google's transition announcement, Gemini CLI stops serving requests on 2026-06-18 for Google AI Pro / Ultra and free Gemini Code Assist for individuals. Gemini API integrations (image-gen backends, sprite pipeline, GEMINI_API_KEY) are unaffected and stay in the toolkit.

If a prior install mirrored into ~/.gemini/, remove the stale mirrors with:

rm -rf ~/.gemini/skills ~/.gemini/agents ~/.gemini/hooks ~/.gemini/scripts ~/.gemini/antigravity/plugins/vexjoy-agent
Factory CLI Support

Mirrors agents (as "droids"), skills, and all hooks into ~/.factory/. Hook config merges into ~/.factory/settings.json with paths rewritten.

Reasonix Support

Mirrors skills, scripts, and the allowlisted hooks (scripts/reasonix-hooks-allowlist.txt) into ~/.reasonix/ (no agent or custom-command surface, so neither is installed; the /do router rides in as a skill). Reasonix fires only 4 events (PreToolUse, PostToolUse, UserPromptSubmit, Stop), so only hooks for those events are allowlisted. Hook config is written to the hooks key of ~/.reasonix/settings.json in Reasonix's native flat shape (one entry per hook, match regex over the tool name); the generator builds absolute python3 commands, so no path rewrite is applied. MCP/model/permissions in ~/.reasonix/config.json are user-owned and left untouched.

Token-saving mode

The toolkit supplies its own routing, domain knowledge, methodology, and enforcement. The default system prompt duplicates most of that.

claude --system-prompt "."

Strips built-in tool-use instructions. The toolkit's agents, skills, hooks, and CLAUDE.md provide equivalent coverage.

Four Layers

LayerCountDoes
Agents44Domain knowledge: idiom tables, failure mode catalogs, error-to-fix mappings
Skills122Phased methodology with gates. Can't skip steps. Each phase has exit criteria requiring evidence.
Hooks78Fire on lifecycle events. Block incomplete work. Zero LLM cost.
Scripts136Determinism: test runners, linters, validators. No LLM judgment.

Full skill catalog: docs/skills.md.

┌─────────────────────────────────────────────────┐
│  SKILL.md                                       │
│  ┌─ Frontmatter ─────────────────────────────┐  │
│  │ triggers, pairs_with, success-criteria     │  │
│  └────────────────────────────────────────────┘  │
│  Reference Loading Table (conditional imports)   │
│  Phased Instructions (numbered, with gates)      │
│  Verification (evidence requirements)            │
└─────────────────────────────────────────────────┘

Built with the Toolkit

A game built entirely by Claude Code using these agents, skills, and pipelines:

Choose Your Path

I just want to use it Install, learn /do, done.

I do knowledge work Writing, research, data analysis, moderation, HTML artifacts. No code.

I'm a developer Architecture, extension points, adding agents and skills.

I'm an AI power user Routing tables, pipelines, hooks, telemetry DB.

I'm an AI agent Machine-dense inventory. Tables, paths, schemas.

I'm on LinkedIn 🚀 Thought leadership. Agree? 👇

Philosophy

  • Zero-expertise operation. Say what you want. The system classifies, dispatches, enforces, delivers.
  • LLMs orchestrate, programs execute. Deterministic work belongs to scripts. LLM judgment handles design decisions, diagnosis, review.
  • Density. Every word carries instruction, rule, or decision. Cut everything else.
  • Breadth over depth. Right context ensures correctness. Unfocused context adds cost.
  • Structural enforcement. Exit codes enforce what instructions can't. Quality gates are automated, not advisory.
  • Everything pipelines. Complex work decomposes into phases. Phases have gates. Gates prevent cascading failures.

Full design philosophy: PHILOSOPHY.md

Maintenance

One report-only script surfaces upkeep work; it prints a digest and never edits, deletes, or blocks.

  • python3 scripts/stale-skill-scan.py --top 20 ranks stale skills and agents as pruning candidates. Run it quarterly; see docs/deprecation-template.md.

Scheduled work follows the same boundary as everything else: judgment uses agents; repeatable plumbing uses scripts.

NeedUse
Run a deterministic command on a schedulescripts/agent-scheduler.py with runner: "command"
Run an agent judgment on a schedule, webhook, or file changescripts/agent-scheduler.py with the default runner: "claude"
Install or remove a user crontab entry safelyscripts/crontab-manager.py
Audit shell cron reliabilitycron-automation
Keep one interactive objective moving until criteria verifyobjective-loop

Contributing

See CONTRIBUTING.md.

License

MIT. See LICENSE.

其他

低风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 未检测到明显外部权限要求。
  • 未检测到高风险命令。
  • 扫描发现:0 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/research/fact-check" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/research/fact-check" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/research/fact-check" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/research/fact-check" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/notque/vexjoy-agent.git
  3. 将 "skills/research/fact-check" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: fact-check
description: "Verify factual claims against sources before publish."
user-invocable: false
allowed-tools:
  - Read
  - Grep
  - Glob
  - Bash
  - Write
  - WebFetch
  - WebSearch
routing:
  triggers:
    - "fact check"
    - "fact-check"
    - "verify claims"
    - "check facts"
    - "verify this quote"
    - "is this accurate"
    - "is this true"
    - "check this claim"
    - "verify this"
    - "are these numbers right"
  not_for: "building a research report from scratch (research-pipeline) or open-web adversarial research (deep-research). Pick fact-check when a draft or source set already exists and the question is whether its claims hold."
  category: research
  pairs_with:
    - research-pipeline
    - voice-writer

Fact Check

Overview

Verifies every factual claim in a draft or source set before publish. Workflow: extract claims, verify each against evidence, adjudicate with one of four labels, emit a per-claim report with a Warnings section. Burden of proof sits on the claim, not the checker: a claim without supporting evidence stays unproven.

Works standalone on any document, or as a pre-publish gate for voice-writer and publish flows. The gate is non-blocking: the report warns and lists findings; the caller decides whether to publish. When the caller provides source documents, verify against those first; reach for the web only when no provided source covers a claim and web access is in scope.


Reference Loading Table

SignalLoad These FilesWhy
verifying a claim: lateral reading, source-tier climbing, triangulation, quote checksreferences/verification-methods.mdMethod detail for Phase 2
judging whether a stat, price, title, or event date is currentreferences/staleness-windows.mdPer-type freshness windows and re-check actions

Instructions

Phase 1: EXTRACT

Goal: List every checkable claim in the document.

A claim is checkable when it asserts something a source could confirm or refute: statistics, prices, dates, quotes, attributions (who said or did what), titles and roles, event facts, rankings, and causal assertions presented as established fact. Opinions and clearly-labeled speculation stay out of the claim list.

For each claim, record:

| ID | Claim (verbatim or tight paraphrase) | Type | Location |
|----|--------------------------------------|------|----------|
| C1 | [exact assertion]                    | stat / quote / attribution / event / title / price / causal | [paragraph or line] |

Extract quotes verbatim — quote verification in Phase 2 compares exact words, so a paraphrased extraction would hide a fabrication.

Gate: Every checkable assertion in the document has a claim ID. A claim skipped here is a claim never verified, so sweep the document twice — once for numbers and dates, once for quotes and attributions.


Phase 2: VERIFY

Goal: Gather evidence for and against each claim.

Work claim by claim. Methods in depth: references/verification-methods.md.

Step 1: Check provided sources first. When the caller supplied source documents, search them for each claim before anything else. Record the exact source file and passage that supports or contradicts the claim. A claim no provided source covers is a candidate for Missing-source in Phase 3.

Step 2: Read laterally. Judge a source by what other sources say about it and its claim, rather than by the source's own presentation. A single source repeating itself across pages still counts as one source.

Step 3: Climb the source tier. Follow citations upward toward the origin: a news article citing a study is weaker evidence than the study itself. Verify against the highest-tier source reachable — primary documents, original data, the actual transcript.

Step 4: Triangulate contested claims. A contested or surprising claim needs two independent sources — independent means separate origins, so two outlets quoting the same wire story count as one. One source suffices only for routine, uncontested facts drawn from a primary document.

Step 5: Verify quotes on three axes. A quote passes only when all three hold:

  1. Exact words — the quoted text matches the source verbatim. Trimmed or altered wording is a finding, even when the meaning seems preserved.
  2. Attributed speaker — the source confirms this person said it. Right words from the wrong mouth is misattribution.
  3. Original context — the surrounding source text supports the meaning the document gives the quote. A real quote deployed to mean something its context contradicts is a finding.

Step 6: Check staleness. Time-sensitive claims (stats, prices, titles, roles, event status) expire. Apply the windows in references/staleness-windows.md: when a claim's evidence is older than its window, look for a newer figure. A claim contradicted by a newer source from the same origin is stale, and stale counts against the claim — the burden of proof includes freshness.

For each claim, record the evidence found, the source(s), and the source tier. Evidence written down now becomes the report in Phase 4; thin notes here produce an unauditable report.

Gate: Every claim has an evidence record — supporting passages, contradicting passages, or a note that the search came up empty. An empty record is itself evidence and feeds Phase 3.


Phase 3: ADJUDICATE

Goal: Assign each claim exactly one label. The four labels are disjoint: Unverifiable means sources engage the claim but settle nothing; Missing-source means no source addresses it at all.

LabelAssign when
VerifiedEvidence supports the claim — exact match for quotes and figures, current within its staleness window, from a sufficient source tier, triangulated if contested
DisputedEvidence contradicts the claim: the source says something different, a newer source supersedes the figure, the quote's words or speaker differ from the source, or credible sources conflict
UnverifiableSources engage the claim but settle nothing — results pending, source explicitly declines to confirm, or the available evidence cannot reach the claim's specificity
Missing-sourceNo available source addresses the claim at all

Adjudication rules, applied in this order:

  1. Burden of proof sits on the claim. The default state is unproven; evidence moves a claim to Verified, and absence of evidence moves it to Missing-source — sympathy for the author moves nothing.
  2. Contradiction beats support: when one source supports and a credible source contradicts, the label is Disputed, and the report shows both.
  3. Partial verification gets the weakest applicable label. A quote with exact words but the wrong speaker is Disputed, because one failed axis fails the quote.
  4. A stale figure superseded by a newer one from the same origin is Disputed; a figure merely older than its window with no newer figure found is Unverifiable, flagged for staleness.

Gate: Every claim from Phase 1 carries exactly one label and a one-line justification citing its evidence record.


Phase 4: REPORT

Goal: Emit the verification report, verdict first.

# Fact-Check Report: [document name]

## Summary
Claims checked: [N] | Verified: [n] | Disputed: [n] | Unverifiable: [n] | Missing-source: [n]
Labels: Verified = evidence supports; Disputed = evidence contradicts; Unverifiable = sources engage but settle nothing; Missing-source = no source addresses it.
Unchecked: [n claims unchecked because X — e.g., paywalled source, dead link]

## Per-Claim Findings

### C1 — [label]
Claim: [text]
Evidence: [source + passage, or "no source addresses this claim"]
Reasoning: [one or two lines: why this label]

[... every claim, in ID order ...]

## Warnings
- [Each Disputed claim with its correction, each fabricated or altered quote, each stale figure with the current one]
- [Patterns worth the author's attention: e.g., every stat traces to a single source]

## Publish Recommendation
[Hold / fix-then-publish / clear — with the blocking items listed. Advisory: the caller decides.]

The label legend rides in the report because readers act on the report alone, without this SKILL.md. The Unchecked line keeps source failures visible instead of letting them inflate labels silently. The Warnings section carries every Disputed and Missing-source finding in actionable form: what the document says, what the evidence says, and the fix. A report that buries a fabricated quote in a table row has failed its purpose — surface it.

Gate: Report covers every claim ID from Phase 1; Warnings section lists every Disputed and Missing-source claim.


Error Handling

Error: "Caller provided a draft with no sources and web access is out of scope"

Cause: Nothing to verify against. Solution: Run EXTRACT, then label every claim Missing-source and report that verification needs sources. The claim list itself is useful output — it tells the author what needs sourcing.

Error: "Sources conflict with each other"

Cause: Two credible sources give different figures or accounts. Solution: Label the claim Disputed, present both sources with dates and tiers in the evidence record, and note in Warnings which source is newer or higher-tier.

Error: "Quote is a translation or cleaned-up transcription"

Cause: Exact-words check fails on legitimately edited speech. Solution: Verify against the original-language or raw source when reachable. When the edit is disclosed and meaning-preserving, label Verified and note the edit; an undisclosed edit stays a finding.

Error: "Claim count is very large"

Cause: Long document with dense factual content. Solution: Keep every claim. Batch verification by source — verify all claims touching one source together — rather than trimming the claim list. Report length scales with claim count by design.


References

  • references/verification-methods.md — lateral reading, source-tier climbing, triangulation, quote verification in depth
  • references/staleness-windows.md — freshness windows per claim type and re-check actions
  • evals/ — blind A/B fixture corpus with ground-truth.json and judging rubric.md (maintenance context: load only when evaluating or modifying this skill)

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!