SkillAtlasSkill 详情

clean-user-facing-text

Agent skill + stdlib Python service that strips multi-vendor AI provenance marks from text and f...

审核状态:已审核Quality 72Security 70

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月24日
_ _ _ ____ ___ ____ ____ _  _ ____ ____ _  _ ____    ____ ____ _  _ ____ _  _ ____ ____
| | | |__|  |  |___ |__/ |\/| |__| |__/ |_/  [__  __ |__/ |___ |\/| |  | |  | |___ |__/
|_|_| |  |  |  |___ |  \ |  | |  | |  \ | \_ ___]    |  \ |___ |  | |__|  \/  |___ |  \

watermarks-remover

CI Release Stars

Agent skill + stdlib Python service that strips multi-vendor AI provenance marks from text and files. For privacy and hygiene on content you own.

The skill is a thin HTTP client — the agent host needs no Python. All work runs in the service.

Author: ShadowAqueduct

Latest release: v0.5.0

LayerTargetMethod
AInvisible Unicode, exotic spaces, bidi, tag charsDeterministic Python
BStatistical (token-sampling) text watermarksAgent rewrite + optional rewrite_text.py
FilesC2PA / EXIF / XMP / doc propsPNG, JPEG, WebP, AVIF, HEIC, BMP, GIF, TIFF, SVG, PDF, DOCX, XLSX, PPTX, EPUB, ODT, HTML, Markdown, MP4/MOV/M4A/M4V, WAV, MP3, FLAC

Covers class-level marks from Claude, Gemini/SynthID-Text, OpenAI provenance surfaces, and open-LLM schemes (Kirchenbauer green-list, keyed-Gumbel / Aaronson EXP).


Install

Skill ships no code — it calls the service over HTTP. Install the skill, start the service, set WATERMARKS_SERVICE_URL if it is not http://127.0.0.1:8765.

One installer (Python 3.10+, stdlib only):

python3 install_skill.py --skill remove-ai-marks --target claude-code
HostTargetLands in
Claude Code (personal)--target claude-code~/.claude/skills/<skill>
Claude Code (project)--target claude-project --project-dir PATHPATH/.claude/skills/<skill>
Cowork / claude.ai / cloud--target coworkdist/<skill>.zip (upload under Customize → Skills)
Cursor--target cursor~/.cursor/skills/<skill>

Shipped skills: remove-ai-marks (full, service-backed) and clean-user-facing-text (text only, self-contained). Use --list to see them. Existing installs are kept as backups unless you pass --force. --link symlinks the checkout for live edits.

Claude Code plugin (marketplace)

/plugin marketplace add guillaumemeyer/watermarks-remover
/plugin install watermarks-remover@watermarks-remover

Skills load as /watermarks-remover:remove-ai-marks and /watermarks-remover:clean-user-facing-text. Update with /plugin marketplace update watermarks-remover.

Grok

mkdir -p ~/.grok/skills
ln -sfn "$(pwd)/skills/remove-ai-marks" ~/.grok/skills/remove-ai-marks

Start the service

make serve                 # http://127.0.0.1:8765
# or:
python3 service/scripts/server.py --host 127.0.0.1 --port 8765

Optional system tools (used when present): c2patool, exiftool, qpdf. Core scripts need only Python 3.10+ stdlib.


Automatic cleaning via hook

A skill is an instruction — the model decides whether to run it. A hook runs on every matching tool call and does not need model cooperation.

The plugin registers a PostToolUse hook on Write|Edit|MultiEdit|NotebookEdit that runs service/scripts/hook_written_file.py:

ModeBehaviour
check (default)Reports marks, leaves the file alone
cleanStrips marks in place, notifies the model

Set mode via plugin settings (Hook mode) or WATERMARKS_HOOK_MODE=clean. Without the plugin, add the hook yourself in ~/.claude/settings.json.

Hooks cover files the agent writes and the pre-commit gate. Chat transcript text still depends on the skill (best-effort).


Quick use

SCRIPTS=service/scripts

python3 "$SCRIPTS/inspect_file.py" draft.md
python3 "$SCRIPTS/clean_file.py" draft.md -o draft.cleaned.md
python3 "$SCRIPTS/clean_file.py" photo.png -o photo.cleaned.png
python3 "$SCRIPTS/clean_file.py" notes.docx -o notes.cleaned.docx

# Text Layer A
python3 "$SCRIPTS/inspect_text.py" draft.md
python3 "$SCRIPTS/clean_text.py" draft.md -o draft.cleaned.md --stats

# Layer B rewrite (default: print prompt only)
python3 "$SCRIPTS/rewrite_text.py" draft.md --backend print-prompt --strength paraphrase

Text tools refuse binary input (DOCX, PDF, images) and point you at inspect_file.py / clean_file.py. Unrecognized formats are never auto-cleaned.


HTTP service

Same machinery as a stdlib HTTP server (service/scripts/server.py):

MethodPathReturns
GET/health{"ok": true, "version": ...}
GET/capabilitiesoptional tools / backends
GET/openapi.jsonOpenAPI 3.0.3 spec
POST/inspectkind, suspicious, report
POST/detectdetections
POST/cleancleaned base64 + report
POST/inspect/batch, /clean/batchper-file results (max 50)
WM="http://127.0.0.1:8765"
curl -s "$WM/health"
curl -s -X POST "$WM/clean" -H 'Content-Type: application/json' \
  -d "{\"file\": \"$(base64 < notes.md | tr -d '\n')\", \"name\": \"notes.md\"}"

Set WATERMARKS_SERVER_API_KEY to require bearer auth. Binds loopback by default.

Detection is separate from cleaning. Text detectors (MarkLLM, keyed-Gumbel, Claude seam) and image SynthID scoring are opt-in and fail-soft.


Docker

make docker-core-build
docker run --rm -p 127.0.0.1:8765:8765 --read-only --tmpfs /tmp watermarks-remover

docker compose up -d                         # core only
docker compose --profile harness up -d       # + markllm / markdiffusion
docker compose --profile heavy up -d         # + ctrlregen / synthid (local builds)

Published images on GHCR: core, markllm, markdiffusion. CtrlRegen and SynthID scorer stay local-only (upstream licensing). Copy .env.example → .env for optional config; nothing is required for basic text cleaning.


File formats

FormatClean
PNG / JPEG / WebPDrop C2PA / XMP / EXIF segments
AVIF / HEICDrop ISOBMFF boxes
BMP / GIF / TIFFTruncate trailing meta / drop extensions & tags
SVGStrip <metadata>, XMP
PDFexiftool → qpdf (structural) → optional Ghostscript deep image pass
DOCX / XLSX / PPTX / ODT / EPUBScrub props, customXml, OPF, embedded media
HTML / MarkdownStrip meta / JSON-LD / AI frontmatter keys + Layer A
MP4 / MOV / M4A / M4V / WAV / MP3 / FLACDrop C2PA / ID3 / LIST chunks

PDF needs qpdf for a real strip (exiftool alone is incremental and leaves recoverable bytes). Ghostscript handles metadata inside embedded images. Soft-bound C2PA and pure pixel/audio/video watermarks remain out of scope for the core path.


Optional backends

BackendRoleNotes
reverse-SynthIDImage SynthID scoreExternal checkout; detection only
CtrlRegenPixel-domain removalExternal; heavy; conservative strength default
MarkLLMText watermark verify (KGW / SynthID)Same-config only, not a vendor oracle
MarkDiffusionImage watermark harness + DiffusionPurificationSame-config only
keyed-Gumbel (detect_gumbel.py)Model-free same-key replayStdlib; needs the generation key

Bootstrap scripts live under service/scripts/ (setup_synthid.sh, setup_ctrlregen.sh, setup_markllm.sh, setup_markdiffusion.sh). Layer B rewrite is iterative and can be driven by these detectors when configured.


How text marking works

  • Layer A removes edit-based Unicode carriers (testable, lossless).
  • Layer B attacks sampling watermarks via heavy rewrite (best-effort; costs style and voice).
  • File cleaners strip C2PA / XMP / props from supported containers.

No tool can certify that a vendor detector will fail. Prefer a non-origin model for Layer B so you do not re-stamp the text.

Skip Layer B when quality matters more than hygiene: use Layer A + file cleaners and keep the original prose.


Pre-commit

# .pre-commit-config.yaml
repos:
  - repo: https://github.com/guillaumemeyer/watermarks-remover
    rev: v0.5.0
    hooks:
      - id: watermarks-remover-check   # fail on marks
      # - id: watermarks-remover-clean # opt-in: clean in place

Ethics

For privacy and research on content you own or are authorized to process. Not for academic fraud or false “human-written” claims. Users must follow local law. The authors disclaim liability for misuse.

See skills/remove-ai-marks/references/ethics.md.


License

MIT — see LICENSE.

Bibliography (selected)

  • How Claude marks AI-generated content
  • Dathathri et al., Scalable watermarking for identifying large language model outputs (SynthID-Text, Nature 2024)
  • Kirchenbauer et al., A Watermark for Large Language Models (arXiv:2301.10226)
  • C2PA / c2patool
  • Liu et al., Image Watermarks are Removable Using Controllable Regeneration from Clean Noise (CtrlRegen, ICLR 2025)
  • THU-BPM/MarkLLM, MarkDiffusion
其他

中风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 可能需要外部 token、网络权限或第三方服务。
  • 未检测到高风险命令。
  • 扫描发现:2 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/ShadowAqueduct/watermark-remover.git
  3. 将 "skills/clean-user-facing-text" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/ShadowAqueduct/watermark-remover.git
  3. 将 "skills/clean-user-facing-text" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/ShadowAqueduct/watermark-remover.git
  3. 将 "skills/clean-user-facing-text" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/ShadowAqueduct/watermark-remover.git
  3. 将 "skills/clean-user-facing-text" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/ShadowAqueduct/watermark-remover.git
  3. 将 "skills/clean-user-facing-text" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: clean-user-facing-text
description: >
  Audit and finalize authorized natural-language text meant for readers:
  strip suspicious invisible Unicode, then rewrite prose while keeping facts,
  meaning, and the writer's voice. Use when the user asks to clean, humanize,
  polish, or finalize articles, manuscripts, reports, documentation, emails,
  product copy, UI text, Markdown, or HTML prose, or when a project rule
  explicitly requires this workflow. Don't use for code-only tasks or
  undisclosed authorship evasion; leave code, commands, identifiers, paths,
  APIs, formulas, citations, required disclosures, and verbatim quotations
  unchanged.

Clean user-facing text

Final hygiene pass on prose the user owns or is authorized to process. Unicode cleanup is deterministic. Statistical-watermark reduction is best-effort. Never claim a rewrite proves human authorship or is undetectable. Keep required academic, legal, platform, and regulatory disclosures.

Workflow

  1. Identify the prose that readers will see.
  2. Protect non-prose spans:
    • fenced and inline code
    • commands, paths, URLs, identifiers, API names, and exact values
    • formulas, citations, and text the user asks to quote verbatim
  3. Preserve every claim, fact, number, name, citation, and requirement.
  4. Establish the writing brief before changing prose:
    • use a voice sample only when the user owns it or is authorised to use it; don't imitate another named person
    • when there is no sample, make the prose clear and natural without pretending to imitate a particular person
    • keep required disclosures, uncertainty, and the writer's actual point of view
  5. Rewrite the remaining prose once:
    • vary clause order, sentence boundaries, rhythm, connectors, and function words
    • replace formulaic transitions and filler with direct, natural wording
    • keep the concrete details and judgement that make the text recognisable as the writer's
    • treat unusual grammar, repetition, directness, and phrasing as possible voice or accessibility choices; change them only when the user asks or when they create a clear reading problem
    • preserve the requested language, tone, structure, and formatting; never translate unless asked
    • for non-English text, use fluent constructions native to that language rather than English sentence patterns
    • do not add or remove claims merely to increase variation
  6. For text artifacts or supplied text files, run the deterministic Unicode pass after rewriting.
  7. Return only the polished result unless the user asks for an audit or explanation.

For practical guidance on preserving a writer's voice and removing formulaic prose, read references/writing-in-your-voice.md whenever the user asks to retain or adjust voice.

Deterministic Unicode pass

Resolve SCRIPTS to this skill's scripts/ directory. Use the available Python 3 launcher for the platform. Replace PYTHON below with python3 on most macOS/Linux systems, py on Windows, or another verified Python 3 command.

Inspect first when editing an existing file:

PYTHON "$SCRIPTS/inspect_text.py" --json INPUT
PYTHON "$SCRIPTS/clean_text.py" INPUT -o OUTPUT --stats --no-normalize-spaces
PYTHON "$SCRIPTS/inspect_text.py" --json OUTPUT

Use - for stdin. Prefer a new *.cleaned.* output unless the user explicitly requests in-place editing.

Use --no-normalize-spaces by default so NBSP, narrow no-break spaces, figure spaces, and CJK ideographic spaces retain their layout semantics. Normalize spaces only when the user requests it.

Do not use --aggressive-homoglyphs, --nfkc, or --strip-emoji-glue unless the user requests aggressive normalization and accepts possible changes to multilingual text, emoji, directionality, or typography.

The scripts support plain text, source text, Markdown, and HTML source as text. For mixed Markdown or HTML, inspect hit positions first. If a hit falls inside protected code, attributes, or another non-prose span, do not run whole-file cleanup; clean only the prose segments or leave that hit unchanged. Do not pass binary containers such as PDF, DOCX, images, or archives.

For a chat-only response that is not written to a file, perform the rewrite workflow directly. Do not claim that the chat response received a deterministic post-send Unicode filter.

Code boundary

When prose and code are mixed, rewrite prose only. Never rename variables, alter string literals, reformat code, or change executable output as part of this skill. If a Markdown or HTML file contains executable snippets, preserve those spans byte-for-byte whenever practical.

Reporting

When the user asks for an audit, distinguish:

  • Verifiable: Unicode characters removed or replaced, with script counts.
  • Best-effort: prose rewritten to alter token and syntax patterns.
  • Not established: official detector evasion, human authorship, or removal of a vendor's secret-key watermark.

For technical background, read references/watermark-notes.md. For misuse or disclosure questions, read references/responsible-use.md.

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!