复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
Retriever is an open source document intelligence plugin for Claude Code. It
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
Retriever is an open source document intelligence plugin for Claude Code. It
turns local folders, productions, mailbox exports, and other supported
document collections into a workspace you can search, browse, review, enrich,
analyze, and export with Claude. Users work with Retriever through Claude
Code: install it once, open Claude Code in the target workspace, and use
plain-English requests or /retriever:* commands to ingest, inspect, review,
process, and export. It fits legal teams, investigations, compliance,
diligence, and other document-heavy workflows that need local-first search,
review, analysis, and handoff.
Typical uses include first-pass review and hot-doc triage, contract review, mailbox and production analysis, Bates lookup, OCR or extraction over a selected set, and CSV or archive export for handoff.
Project docs:
IMPORTANT: IF YOU HANDLE SENSITIVE, CLIENT, OR INTERNAL DOCUMENTS, TURN OFF
Help improve Claude IN Settings -> Privacy BEFORE USING RETRIEVER.
Anthropic says that when this setting is off, new Claude chats and coding
sessions are not used for future model training, although flagged
conversations may still be used for trust and safety purposes.
Review the Anthropic privacy policy and
confirm it satisfies your practice, client, regulatory, and organizational
requirements before using Retriever with sensitive material.
Zero Data Retention (ZDR) is available for Claude Code on Claude for Enterprise. Anthropic says ZDR is enabled per organization and covers Claude Code inference on Claude for Enterprise. See Zero data retention for the current scope, limitations, and enablement details.
Open Claude Code and paste this. Claude does the rest.
Install Retriever: run
if [ -d ~/.claude/skills/retriever/.git ]; then git -C ~/.claude/skills/retriever pull --ff-only origin main; else git clone --single-branch --depth 1 https://github.com/sdemyanov/retriever.git ~/.claude/skills/retriever; fi. Then runcd ~/.claude/skills/retriever && ./setup.
To test Retriever with public sample data, open Claude Code and paste this:
Set up Retriever sample data: run
if [ -d ~/retriever-data-public/.git ]; then git -C ~/retriever-data-public pull --ff-only origin main; else git clone --single-branch --depth 1 https://github.com/sdemyanov/retriever-data-public.git ~/retriever-data-public; fi.
Then open Claude Code in ~/retriever-data-public and run
Ingest this folder.
/retriever:* commands to
search, review, analyze, enrich, and export.Example prompts:
Ingest this folderFind documents mentioning indemnificationReview the workspace for hot docsShow emails from Alice in 2023Extract counterparties from the current contractsExport the current resultsIf you are trying Retriever for the first time, after opening Claude Code in the target folder, a good starting point is:
Check the workspace status for this folderIngest this folderShow me the first results for <your first keyword>Only show emails if that helpsAdd dataset name to the visible columns if you want more contextSave this scope as first-pass if you want to reuse itExport the current results if you want a handoff artifactThat path exercises the setup, browse, narrowing, display, persistence, and export surfaces that most users rely on first.
DAT + OPT with TEXT/,
IMAGES/, and optional NATIVES/Ingest-path behaviors worth knowing:
.ics/.ifb/.vcal/.vcs) that arrive as email
attachments are promoted into the parent email — the invite's organizer,
attendees, when, location, join URL, UID, and sequence are rolled into the
email's indexed text and rendered as a structured invite header in the
previewactivate-text-revision).zip, .rar, .7z are not unpacked or indexed automatically.retriever/ in the
workspace root, and leaves original source files in place.control_number values
for referencing, analysis, and export. Production documents use produced
Bates values as the control number.python3 -m retriever entrypoints that drive the resumable backend to a
terminal state.Natural-language instructions in Claude Code are the primary interface. The slash commands below are optional shortcuts, and some natural-language instructions may map to them internally when that is the clearest fit.
Use this when you are starting with a new folder of files.
Primary interface in Claude Code:
If you want to target a processed production root explicitly, ask Retriever to ingest that production root rather than the whole folder.
This is the main interactive workflow, and plain English is usually the best place to start.
Natural-language examples:
Show emails about the NDAOnly show messages from 2023Sort newest first and show 25 at a timeSome requests like these may map to slash commands such as:
/retriever:search nda
/retriever:filter content_type = 'Email'
/retriever:filter date_created BETWEEN '2023-01-01' AND '2023-12-31'
/retriever:sort date_created desc
/retriever:page-size 25
/retriever:next
Retriever treats Bates-like input as a first-class lookup mode.
Natural-language examples:
Show Bates ABC000123Show Bates range ABC000123-ABC000150Claude may use commands like:
/retriever:bates ABC000123
/retriever:bates ABC000123-ABC000150
You can also set Bates scope through /retriever:search because it
auto-detects Bates-shaped input:
/retriever:search ABC000123-ABC000150
Once the scope is right, start with plain-English requests in Claude Code.
Natural-language examples:
Translate the current results into SpanishExtract counterparties from the current contractsOCR the documents with empty textDescribe the images in the current resultsIf a processing command is interrupted, ask Claude Code to resume the run.
If you need plain full-text search for something that looks like a Bates value, force FTS:
/retriever:search --fts ABC000123
A scope is the conjunction of:
from-run selectorNatural-language examples:
Save this view as merger-email-hotdocsLoad the scope merger-email-hotdocsShow the current scopeSome scope-management requests may map to commands like:
/retriever:search merger
/retriever:filter content_type = 'Email'
/retriever:dataset "Hot Docs"
/retriever:scope save merger-email-hotdocs
Later:
/retriever:scope load merger-email-hotdocs
Useful related commands:
/retriever:scope
/retriever:scope list
/retriever:scope clear
Datasets are named document collections. They are useful for saved result sets, source-backed groupings, and repeatable exports.
Natural-language examples:
Create a dataset called Priority SetPut these documents in Priority SetList datasetsClear the active datasetSome dataset requests may map to commands like:
/retriever:dataset
/retriever:dataset list
/retriever:dataset "Priority Set"
/retriever:dataset "Hot Docs", "Witness Files"
/retriever:dataset clear
/retriever:dataset list renders as a compact stats table so you can see each
dataset's document count, top custodians, and activity range at a glance
without drilling in.
Once your scope is right, plain-English export requests are the normal path.
Natural-language examples:
Export the current results to CSVBuild a preview bundle for these documentsCreate a portable archive of the current resultsIf you need specific filenames or export options, say that directly in the request.
Use cases:
Retriever supports user-managed custom fields plus manual corrections to editable built-in fields.
Natural-language examples:
Add a field called privilege_statusDescribe privilege_status as Privilege designationMark DOC001.00000042 as privilegedClear privilege_status on DOC001.00000042Some field and metadata requests may map to commands like:
/retriever:field add privilege_status text
/retriever:field describe privilege_status "Privilege designation"
/retriever:fill privilege_status privileged on DOC001.00000042
/retriever:fill privilege_status clear on DOC001.00000042
/retriever:fill can also populate a value across the active scope. Those bulk
forms require --confirm:
/retriever:search privileged
/retriever:filter content_type = 'Email' AND custodian = 'Garcia'
/retriever:fill privilege_status privileged --confirm
Important details:
/retriever:fill refuses to target derived or system-managed fields
(custodian, dataset_name, production_name, hashes, ids, ingest
timestamps); correct those through the appropriate ingest or conversation
command instead/retriever:field delete is permanent; the slash surface previews the
removal and requires --confirm before actually dropping the fieldRetriever can freeze a selector into a run and process it later. For most users, it is better to describe the outcome you want in plain English and let Claude guide the setup.
High-level flow:
create-run./retriever:from-run <run-id>.Natural-language example:
Create an Issue Tags extraction job, add a primary_issue output, and run it on the current scopeThe detailed slash-command reference, /retriever:search and
/retriever:filter syntax, field/column discovery notes, display and paging
tips, and advanced CLI quick reference now live in
docs/browse-reference.md.
Retriever treats the selected folder as the workspace root. All persistent
state lives under .retriever/:
.retriever/
├── retriever.db
├── previews/
├── text-revisions/
├── jobs/
├── locks/
├── logs/
└── runtime.json
Important consequences:
<repo-root>/.retriever-plugin-runtime/...) when needed, not under
.retriever/; the main python3 -m retriever entrypoint still runs in the
active Python interpreter. See Runtime and dependencies for details.Retriever indexes logical documents, not just files.
That means:
Retriever has a persistent browse session per workspace.
That session keeps three kinds of state:
from-run selectorsScope changes reset paging. Display settings and browse preferences persist until you change them or reset them.
Document listings use a standard table:
Scope, Sort, and Pagetitle cell is the clickable preview linkDocuments 1-10 of 85. Ask for the next page to see more.Default behavior:
10100content_type, title, author, date_created, control_numberrelevance ascbates ascdate_created descRetriever uses two Python layers:
python3 -m retriever entrypoint<repo-root>/.retriever-plugin-runtime/<system>-<machine>-pyX.Y/venv/
Heavy parser dependencies (pdfplumber, python-docx, openpyxl, xlrd,
extract-msg, libpff-python, striprtf, Pillow,
charset-normalizer) are lazy-installed into that shared venv the first
time a command actually needs them. Retriever first uses whatever is already
importable in the active interpreter, then falls back to the shared runtime for
optional parser packages. Non-parsing commands do not pay that cost.
Consequences:
.retriever/ folder stays lightweight — it holds data,
state, and logs, not Python packagesworkspace status will report
the runtime state and warn if something needed is missingThe on-disk directory name still uses .retriever-plugin-runtime/ for
compatibility with older workspaces and generated tooling, but users interact
with Retriever through Claude Code.
pypff backend being available. Use
workspace status if PST ingest is not ready; parser dependencies are
lazy-installed into the shared plugin runtime (see Runtime and
dependencies below), so the status check will also tell you if the runtime
needs to be (re)populated.ingest-production when you want to target a production root explicitly./retriever:scope, /retriever:dataset, /retriever:from-run,
/retriever:sort, and /retriever:page-size before assuming the underlying
data is gone.Open Claude Code and paste this. Claude does the rest.
Uninstall Retriever: remove
~/.claude/commands/retriever,~/.claude/retriever-manifest.json,~/.claude/skills/retriever, and theretriever.pthfile from your Python user site-packages. Remove the Retriever section from~/.claude/CLAUDE.md; if that leaves the file empty, delete the file. Skip anything that does not exist. Then give a short summary of what was removed.
Retriever is licensed under the Elastic License 2.0 (ELv2). The SPDX
identifier is Elastic-2.0. See the Elastic License
2.0 for the license terms.
name: fill
description: >
Use this skill when the user wants to populate, set, tag, mark, label,
classify, annotate, flag, or clear a custom or editable built-in field value
on one document or on a filtered/scoped result set — or when the user types
"/fill".
metadata:
version: "1.1.17"Operates under
retriever:routing. If the user's intent actually fits a different tier — anotherretriever:*skill, a Tier 2 slash, a Tier 3tools.pysubcommand, or (last resort) direct DB access — stop and re-route against the ladder before continuing.
Use this skill for /fill <field> <value>, /fill <field> clear, /fill ... on <doc-ref>, and /fill ... on <doc-ref,doc-ref,...>.
python3 skills/tool-template/tools.py slash . /fill ....on <doc-ref[,doc-ref,...]> form.on ..., rely on the active browse state. If no active selection exists yet, narrow it first with retriever:search, retriever:dataset, retriever:filter, retriever:bates, retriever:from-run, or by asking the user for the target documents.--confirm; single explicit-document fills do not.custodian, dataset_name, production_name, hashes, ids, or ingest timestamps.
评论 (0)
暂无评论,成为第一个评论者吧!