复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
Agent skill + stdlib Python service that strips multi-vendor AI provenance marks from text and f...
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
_ _ _ ____ ___ ____ ____ _ _ ____ ____ _ _ ____ ____ ____ _ _ ____ _ _ ____ ____
| | | |__| | |___ |__/ |\/| |__| |__/ |_/ [__ __ |__/ |___ |\/| | | | | |___ |__/
|_|_| | | | |___ | \ | | | | | \ | \_ ___] | \ |___ | | |__| \/ |___ | \
Agent skill + stdlib Python service that strips multi-vendor AI provenance marks from text and files. For privacy and hygiene on content you own.
The skill is a thin HTTP client — the agent host needs no Python. All work runs in the service.
Author: ShadowAqueduct
Latest release: v0.5.0
| Layer | Target | Method |
|---|---|---|
| A | Invisible Unicode, exotic spaces, bidi, tag chars | Deterministic Python |
| B | Statistical (token-sampling) text watermarks | Agent rewrite + optional rewrite_text.py |
| Files | C2PA / EXIF / XMP / doc props | PNG, JPEG, WebP, AVIF, HEIC, BMP, GIF, TIFF, SVG, PDF, DOCX, XLSX, PPTX, EPUB, ODT, HTML, Markdown, MP4/MOV/M4A/M4V, WAV, MP3, FLAC |
Covers class-level marks from Claude, Gemini/SynthID-Text, OpenAI provenance surfaces, and open-LLM schemes (Kirchenbauer green-list, keyed-Gumbel / Aaronson EXP).
Skill ships no code — it calls the service over HTTP. Install the skill, start the service, set WATERMARKS_SERVICE_URL if it is not http://127.0.0.1:8765.
One installer (Python 3.10+, stdlib only):
python3 install_skill.py --skill remove-ai-marks --target claude-code
| Host | Target | Lands in |
|---|---|---|
| Claude Code (personal) | --target claude-code | ~/.claude/skills/<skill> |
| Claude Code (project) | --target claude-project --project-dir PATH | PATH/.claude/skills/<skill> |
| Cowork / claude.ai / cloud | --target cowork | dist/<skill>.zip (upload under Customize → Skills) |
| Cursor | --target cursor | ~/.cursor/skills/<skill> |
Shipped skills: remove-ai-marks (full, service-backed) and clean-user-facing-text (text only, self-contained). Use --list to see them. Existing installs are kept as backups unless you pass --force. --link symlinks the checkout for live edits.
/plugin marketplace add guillaumemeyer/watermarks-remover
/plugin install watermarks-remover@watermarks-remover
Skills load as /watermarks-remover:remove-ai-marks and /watermarks-remover:clean-user-facing-text. Update with /plugin marketplace update watermarks-remover.
mkdir -p ~/.grok/skills
ln -sfn "$(pwd)/skills/remove-ai-marks" ~/.grok/skills/remove-ai-marks
make serve # http://127.0.0.1:8765
# or:
python3 service/scripts/server.py --host 127.0.0.1 --port 8765
Optional system tools (used when present): c2patool, exiftool, qpdf. Core scripts need only Python 3.10+ stdlib.
A skill is an instruction — the model decides whether to run it. A hook runs on every matching tool call and does not need model cooperation.
The plugin registers a PostToolUse hook on Write|Edit|MultiEdit|NotebookEdit that runs service/scripts/hook_written_file.py:
| Mode | Behaviour |
|---|---|
check (default) | Reports marks, leaves the file alone |
clean | Strips marks in place, notifies the model |
Set mode via plugin settings (Hook mode) or WATERMARKS_HOOK_MODE=clean. Without the plugin, add the hook yourself in ~/.claude/settings.json.
Hooks cover files the agent writes and the pre-commit gate. Chat transcript text still depends on the skill (best-effort).
SCRIPTS=service/scripts
python3 "$SCRIPTS/inspect_file.py" draft.md
python3 "$SCRIPTS/clean_file.py" draft.md -o draft.cleaned.md
python3 "$SCRIPTS/clean_file.py" photo.png -o photo.cleaned.png
python3 "$SCRIPTS/clean_file.py" notes.docx -o notes.cleaned.docx
# Text Layer A
python3 "$SCRIPTS/inspect_text.py" draft.md
python3 "$SCRIPTS/clean_text.py" draft.md -o draft.cleaned.md --stats
# Layer B rewrite (default: print prompt only)
python3 "$SCRIPTS/rewrite_text.py" draft.md --backend print-prompt --strength paraphrase
Text tools refuse binary input (DOCX, PDF, images) and point you at inspect_file.py / clean_file.py. Unrecognized formats are never auto-cleaned.
Same machinery as a stdlib HTTP server (service/scripts/server.py):
| Method | Path | Returns |
|---|---|---|
| GET | /health | {"ok": true, "version": ...} |
| GET | /capabilities | optional tools / backends |
| GET | /openapi.json | OpenAPI 3.0.3 spec |
| POST | /inspect | kind, suspicious, report |
| POST | /detect | detections |
| POST | /clean | cleaned base64 + report |
| POST | /inspect/batch, /clean/batch | per-file results (max 50) |
WM="http://127.0.0.1:8765"
curl -s "$WM/health"
curl -s -X POST "$WM/clean" -H 'Content-Type: application/json' \
-d "{\"file\": \"$(base64 < notes.md | tr -d '\n')\", \"name\": \"notes.md\"}"
Set WATERMARKS_SERVER_API_KEY to require bearer auth. Binds loopback by default.
Detection is separate from cleaning. Text detectors (MarkLLM, keyed-Gumbel, Claude seam) and image SynthID scoring are opt-in and fail-soft.
make docker-core-build
docker run --rm -p 127.0.0.1:8765:8765 --read-only --tmpfs /tmp watermarks-remover
docker compose up -d # core only
docker compose --profile harness up -d # + markllm / markdiffusion
docker compose --profile heavy up -d # + ctrlregen / synthid (local builds)
Published images on GHCR: core, markllm, markdiffusion. CtrlRegen and SynthID scorer stay local-only (upstream licensing). Copy .env.example → .env for optional config; nothing is required for basic text cleaning.
| Format | Clean |
|---|---|
| PNG / JPEG / WebP | Drop C2PA / XMP / EXIF segments |
| AVIF / HEIC | Drop ISOBMFF boxes |
| BMP / GIF / TIFF | Truncate trailing meta / drop extensions & tags |
| SVG | Strip <metadata>, XMP |
| exiftool → qpdf (structural) → optional Ghostscript deep image pass | |
| DOCX / XLSX / PPTX / ODT / EPUB | Scrub props, customXml, OPF, embedded media |
| HTML / Markdown | Strip meta / JSON-LD / AI frontmatter keys + Layer A |
| MP4 / MOV / M4A / M4V / WAV / MP3 / FLAC | Drop C2PA / ID3 / LIST chunks |
PDF needs qpdf for a real strip (exiftool alone is incremental and leaves recoverable bytes). Ghostscript handles metadata inside embedded images. Soft-bound C2PA and pure pixel/audio/video watermarks remain out of scope for the core path.
| Backend | Role | Notes |
|---|---|---|
| reverse-SynthID | Image SynthID score | External checkout; detection only |
| CtrlRegen | Pixel-domain removal | External; heavy; conservative strength default |
| MarkLLM | Text watermark verify (KGW / SynthID) | Same-config only, not a vendor oracle |
| MarkDiffusion | Image watermark harness + DiffusionPurification | Same-config only |
keyed-Gumbel (detect_gumbel.py) | Model-free same-key replay | Stdlib; needs the generation key |
Bootstrap scripts live under service/scripts/ (setup_synthid.sh, setup_ctrlregen.sh, setup_markllm.sh, setup_markdiffusion.sh). Layer B rewrite is iterative and can be driven by these detectors when configured.
No tool can certify that a vendor detector will fail. Prefer a non-origin model for Layer B so you do not re-stamp the text.
Skip Layer B when quality matters more than hygiene: use Layer A + file cleaners and keep the original prose.
# .pre-commit-config.yaml
repos:
- repo: https://github.com/guillaumemeyer/watermarks-remover
rev: v0.5.0
hooks:
- id: watermarks-remover-check # fail on marks
# - id: watermarks-remover-clean # opt-in: clean in place
For privacy and research on content you own or are authorized to process. Not for academic fraud or false “human-written” claims. Users must follow local law. The authors disclaim liability for misuse.
See skills/remove-ai-marks/references/ethics.md.
MIT — see LICENSE.
name: clean-user-facing-text
description: >
Audit and finalize authorized natural-language text meant for readers:
strip suspicious invisible Unicode, then rewrite prose while keeping facts,
meaning, and the writer's voice. Use when the user asks to clean, humanize,
polish, or finalize articles, manuscripts, reports, documentation, emails,
product copy, UI text, Markdown, or HTML prose, or when a project rule
explicitly requires this workflow. Don't use for code-only tasks or
undisclosed authorship evasion; leave code, commands, identifiers, paths,
APIs, formulas, citations, required disclosures, and verbatim quotations
unchanged.Final hygiene pass on prose the user owns or is authorized to process. Unicode cleanup is deterministic. Statistical-watermark reduction is best-effort. Never claim a rewrite proves human authorship or is undetectable. Keep required academic, legal, platform, and regulatory disclosures.
For practical guidance on preserving a writer's voice and removing formulaic prose, read references/writing-in-your-voice.md whenever the user asks to retain or adjust voice.
Resolve SCRIPTS to this skill's scripts/ directory.
Use the available Python 3 launcher for the platform. Replace PYTHON below
with python3 on most macOS/Linux systems, py on Windows, or another verified
Python 3 command.
Inspect first when editing an existing file:
PYTHON "$SCRIPTS/inspect_text.py" --json INPUT
PYTHON "$SCRIPTS/clean_text.py" INPUT -o OUTPUT --stats --no-normalize-spaces
PYTHON "$SCRIPTS/inspect_text.py" --json OUTPUT
Use - for stdin. Prefer a new *.cleaned.* output unless the user explicitly requests in-place editing.
Use --no-normalize-spaces by default so NBSP, narrow no-break spaces, figure spaces, and CJK ideographic spaces retain their layout semantics. Normalize spaces only when the user requests it.
Do not use --aggressive-homoglyphs, --nfkc, or --strip-emoji-glue unless the user requests aggressive normalization and accepts possible changes to multilingual text, emoji, directionality, or typography.
The scripts support plain text, source text, Markdown, and HTML source as text. For mixed Markdown or HTML, inspect hit positions first. If a hit falls inside protected code, attributes, or another non-prose span, do not run whole-file cleanup; clean only the prose segments or leave that hit unchanged. Do not pass binary containers such as PDF, DOCX, images, or archives.
For a chat-only response that is not written to a file, perform the rewrite workflow directly. Do not claim that the chat response received a deterministic post-send Unicode filter.
When prose and code are mixed, rewrite prose only. Never rename variables, alter string literals, reformat code, or change executable output as part of this skill. If a Markdown or HTML file contains executable snippets, preserve those spans byte-for-byte whenever practical.
When the user asks for an audit, distinguish:
For technical background, read references/watermark-notes.md. For misuse or disclosure questions, read references/responsible-use.md.
评论 (0)
暂无评论,成为第一个评论者吧!