SkillAtlasSkill 详情

hipocampus-core

Drop-in proactive memory harness for AI agents. Zero infrastructure — just files.

审核状态:已审核Quality 72Security 78

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年7月29日

hipocampus

Drop-in proactive memory harness for AI agents. Zero infrastructure — just files.

One command to set up. Works immediately with Claude Code, OpenCode, and OpenClaw.

Benchmark

Evaluated on MemAware — 900 implicit context questions across 3 months of conversation history. The agent must proactively surface relevant past context that the user never explicitly asks about.

MethodEasy (n=300)Medium (n=300)Hard (n=300)Overall
No Memory1.0%0.7%0.7%0.8%
BM25 Search4.7%1.7%2.0%2.8%
BM25 + Vector Search6.0%3.7%0.7%3.4%
Hipocampus (tree only)14.7%5.7%7.3%9.2%
Hipocampus + BM2518.7%10.0%5.7%11.4%
Hipocampus + Vector26.0%18.0%8.0%17.3%
Hipocampus + Vector (10K ROOT)34.0%21.0%8.0%21.0%

Hipocampus + Vector is 21.6x better than no memory and 5.1x better than search alone. On hard questions (cross-domain, zero keyword overlap), Hipocampus scores 8.0% vs 0.7% for vector search — 11.4x better. Search structurally cannot find these connections; the compaction tree can.

Increasing the ROOT.md budget from 3K to 10K tokens (120 topics vs 39) improves Easy from 26% to 34% and overall from 17.3% to 21.0% — more topic coverage means more connections found. Hard tier remains at 8.0%, indicating cross-domain reasoning is bottlenecked by the answer model, not the index size.

Install

Claude Code Plugin

/plugin marketplace add kevin-hs-sohn/hipocampus
/plugin install hipocampus@kevin-hs-sohn/hipocampus

Then run npx hipocampus init for full setup.

Standalone (npm)

npx hipocampus init

Options

npx hipocampus init --no-vector    # BM25 only (saves ~2GB disk)
npx hipocampus init --no-search    # Compaction tree only, no qmd
npx hipocampus init --platform claude-code  # Override platform detection

The Problem: You Can't Search for What You Don't Know You Know

AI agents forget everything between sessions. The obvious solutions — RAG, long context windows, memory files — each solve part of the problem. But they all miss the hardest part: knowing that relevant context exists when nobody asked about it.

A concrete example

You ask your agent: "Refactor this API endpoint for the new payment flow."

Three weeks ago, you and the agent had a long discussion about API rate limiting and decided on a token bucket strategy. That decision is recorded in the session logs. But the agent doesn't know it exists — so it refactors the endpoint without considering rate limits. The payment flow starts dropping requests under load a week later.

This isn't a retrieval failure. The agent never searched for "rate limiting" because the user asked about "payment flow." There is no search query that connects these. The connection only exists if the agent has a holistic view of its own knowledge.

Why existing approaches fail

Large context windows (200K–1M tokens): You could dump all history into context. But attention degrades with length — important details from three weeks ago get drowned by noise. And every API call pays for the full context. At 500K tokens per call, costs become prohibitive.

RAG (vector search, BM25): Powerful when you know what to search for. But search requires a query, and a query requires suspecting that relevant context exists. Our MemAware benchmark confirms: BM25 search scores just 2.8% on implicit context — barely better than no memory (0.8%), while consuming 5x the tokens. Search is a precision tool for known unknowns. It cannot help with unknown unknowns.

Memory files (MEMORY.md, auto memory): Good for the first week. After a month, hundreds of decisions and insights can't fit in a system prompt. You're forced to choose what to keep, and the agent doesn't know what it has forgotten.

What hipocampus does differently

Hipocampus maintains a ~3K token topic index (ROOT.md) that compresses your entire conversation history into a scannable overview — like a table of contents for everything the agent has ever discussed. This is auto-loaded into every session.

When a request comes in, the agent already sees all past topics at zero search cost. It notices connections that search would miss — "this refactoring task relates to the rate limiting decision from three weeks ago" — and retrieves specific details on demand via search or tree traversal.

The effect is similar to injecting your full history into every API call, at a fraction of the token cost.

How It Works

3-Tier Memory

Like a CPU cache hierarchy:

Layer 1 — Hot (always loaded, ~3K tokens)

FilePurpose
memory/ROOT.mdCompressed index of ALL past history — the key innovation
SCRATCHPAD.mdActive work state
WORKING.mdTasks in progress
TASK-QUEUE.mdTask backlog

ROOT.md has four sections:

## Active Context (recent ~7 days)
- hipocampus open-source: finalizing spec, ROOT.md format refactor

## Recent Patterns
- compaction design: functional sections outperform chronological

## Historical Summary
- 2026-01~02: initial 3-tier design, clawy.pro K8s launch
- 2026-03: hipocampus open-source, qmd integration

## Topics Index
- hipocampus [project, 2d]: compaction tree, ROOT.md, skills → spec/
- legal [reference, 14d]: Civil Act §750, tort liability → knowledge/legal-750.md
- clawy.pro [project, 30d]: K8s infra, provisioning, 80-bot deployment

Each topic carries a type (project, feedback, user, reference) and age — so the agent knows not just what it knows, but what kind of information it is and how fresh it is. O(1) lookup — no file reads needed.

Layer 2 — Warm (read on demand)

PathPurpose
memory/YYYY-MM-DD.mdRaw daily logs — structured session records
knowledge/*.mdCurated knowledge base
plans/*.mdTask plans

Layer 3 — Cold (search + compaction tree)

Two retrieval mechanisms:

  • RAG (qmd) — semantic search when you know what you're looking for
  • Compaction tree — hierarchical drill-down (ROOT → monthly → weekly → daily → raw) for browsing and discovery
Compaction chain: Raw → Daily → Weekly → Monthly → Root

memory/
├── ROOT.md                     # Auto-loaded topic index
├── 2026-03-15.md               # Raw daily log (permanent)
├── daily/2026-03-15.md         # Daily compaction node
├── weekly/2026-W11.md          # Weekly index node
└── monthly/2026-03.md          # Monthly index node

Smart Compaction

Below threshold, source files are copied verbatim — no information loss. Above threshold, LLM generates keyword-dense summaries.

LevelThresholdBelowAbove
Raw → Daily~200 linesCopy verbatimLLM summary
Daily → Weekly~300 linesConcatLLM summary
Weekly → Monthly~500 linesConcatLLM summary
Monthly → RootAlwaysRecursive recompaction—

Memory Types

Every memory entry is classified into one of four types, controlling how it's preserved over time:

TypePurposeCompaction behavior
projectWork, decisions, technical findingsCompressed when completed
feedbackUser corrections on approachAlways preserved verbatim
userUser identity, expertise, preferencesAlways preserved
referenceExternal pointers (URLs, tools)Preserved with staleness markers

user and feedback memories never get compressed away — they survive indefinitely. project memories compress into Historical Summary after completion. reference entries get a [?] marker after 30 days without verification.

Selective Recall

When a question might relate to past memory, hipocampus uses a 3-step fallback:

  1. ROOT.md triage (O(1)) — Topics Index lookup. Resolves most queries instantly.
  2. Manifest-based LLM selection — For cross-domain queries where keywords don't match. Reads compaction node frontmatter only (<500 tokens), LLM selects top 5 relevant files.
  3. qmd search — BM25/vector hybrid for specific keyword retrieval.

Step 2 solves the keyword mismatch problem: "배포" ↔ "deployment", "CI/CD" ↔ "github-actions" — the LLM understands semantic connections that keyword search misses.

Automatic Operation

Everything runs automatically after npx hipocampus init:

MechanismWhenCost
Session StartFirst message — load hot files, check compactionRead only
End-of-Task CheckpointAfter every task — typed entry to daily logLLM (subagent)
Proactive FlushEvery ~20 messages — prevent context lossLLM (subagent)
Pre-Compaction HookBefore context compression — mechanical compactZero LLM
Secret ScanningDuring compaction — redact API keys, tokensZero LLM
ROOT.md Auto-LoadEvery session start~3K tokens

Memory writes are dispatched to subagents to keep the main session clean.

Adaptive compaction triggers: Compaction runs when any condition is met — cooldown expired (default 3h), raw log exceeds 300 lines, or 5+ checkpoints accumulated. Active sessions compact more frequently; quiet days skip unnecessary work.

Comparison

Ad-hoc MEMORY.mdOpenVikingHipocampus
SetupManualPython server + embedding modelnpx hipocampus init
InfrastructureNoneServer + DBNone — just files
SearchNoneVector + directory recursiveBM25 + vector hybrid (qmd)
Knows what it knowsOnly what fits (~50 lines)No (search required)ROOT.md (~3K tokens)
Scales over monthsNo — overflowsYesYes — self-compressing tree

File Layout

project/
├── SCRATCHPAD.md
├── WORKING.md
├── TASK-QUEUE.md
├── memory/
│   ├── ROOT.md                  # Topic index (auto-loaded)
│   ├── (YYYY-MM-DD.md)         # Raw daily logs
│   ├── daily/                   # Daily compaction nodes
│   ├── weekly/                  # Weekly index nodes
│   └── monthly/                 # Monthly index nodes
├── knowledge/
├── plans/
├── hipocampus.config.json
└── .claude/skills/hipocampus-*  # Agent skills (5 skills)

Configuration

{
  "platform": "claude-code",
  "search": { "vector": true, "embedModel": "auto" },
  "compaction": { "rootMaxTokens": 3000, "cooldownHours": 3 }
}
FieldDefaultDescription
platformauto-detected"claude-code", "opencode", or "openclaw"
search.vectortrueEnable vector embeddings (~2GB disk)
search.embedModel"auto""auto" for embeddinggemma-300M, "qwen3" for CJK
compaction.rootMaxTokens3000Max token budget for ROOT.md
compaction.cooldownHours3Min hours between compaction runs (0 = disable)

Skills

Hipocampus installs five agent skills:

  • hipocampus-core — Session start protocol + typed checkpoints + exclusion rules
  • hipocampus-compaction — 5-level compaction tree with type-aware rules + secret scanning
  • hipocampus-recall — 3-step selective recall (ROOT.md → manifest LLM → qmd search)
  • hipocampus-search — Search guide: ROOT.md lookup, qmd, tree traversal
  • hipocampus-flush — Manual memory flush via subagent

Spec

Formal specification in spec/:

License

MIT

开发与工程Agent / MCP / Skill 创作

中风险

  • 来源需自行核对维护者身份。
  • 未检测到明显脚本安装指令。
  • 可能需要外部 token、网络权限或第三方服务。
  • 未检测到高风险命令。
  • 扫描发现:1 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/kevin-hs-sohn/hipocampus.git
  3. 将 "skills/core" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/kevin-hs-sohn/hipocampus.git
  3. 将 "skills/core" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/kevin-hs-sohn/hipocampus.git
  3. 将 "skills/core" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/kevin-hs-sohn/hipocampus.git
  3. 将 "skills/core" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/kevin-hs-sohn/hipocampus.git
  3. 将 "skills/core" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: hipocampus-core
description: "3-tier agent memory system with 5-level compaction tree. Claude Code version. Defines session start protocol, end-of-task checkpoints, and memory file management. MUST be followed every session."

Hipocampus — Agent Memory Protocol (Claude Code)

Memory Architecture

Layer 1 (System Prompt — auto-loaded via @import):
  SCRATCHPAD.md    ~150 lines  active working state
  WORKING.md       ~100 lines  current tasks
  TASK-QUEUE.md    ~50 lines   task backlog
  memory/ROOT.md   ~100 lines  topic index of all memory (~3K tokens)

  Long-term memory and user profile are managed by Claude Code's platform auto memory.

Layer 2 (On-Demand — read when needed):
  memory/YYYY-MM-DD.md         raw daily logs (permanent, never deleted)
  knowledge/*.md               detailed knowledge (searchable via qmd)
  plans/*.md                   task plans

Layer 3 (Search — via qmd + compaction tree):
  memory/daily/YYYY-MM-DD.md   daily compaction nodes
  memory/weekly/YYYY-WNN.md    weekly compaction nodes
  memory/monthly/YYYY-MM.md    monthly compaction nodes
  Tree traversal: ROOT → monthly → weekly → daily → raw

Session Start (MANDATORY — run on first user message)

FIRST RESPONSE RULE: On the very first user message of every session, before doing ANYTHING else: Run the Session Start protocol below FIRST. This takes priority over ANY user request — even if the user asks you to do something specific. Complete the step below, ONLY THEN respond to the user.

SCRATCHPAD.md, WORKING.md, TASK-QUEUE.md, memory/ROOT.md are auto-loaded via @import in CLAUDE.md. No manual read needed.

This procedure must be completed before responding to the user NO MATTER WHAT

  1. DO NOT SKIP DO NOT COMPROMISE Compaction maintenance (cooldown-gated): Read memory/.compaction-state.json and hipocampus.config.json (compaction.cooldownHours, default 3).

    Compaction triggers (any ONE is sufficient):

    • Cooldown expired: cooldownHours since lastCompactionRun
    • Raw volume: rawLinesSinceLastCompaction > 300
    • Checkpoint count: checkpointsSinceLastCompaction > 5
    • State file missing or cooldownHours is 0

    If no trigger is met: skip compaction subagent. If any trigger is met: write memory/.compaction-state.json with { "lastCompactionRun": "<current ISO timestamp>", "rawLinesSinceLastCompaction": 0, "checkpointsSinceLastCompaction": 0 }, then dispatch compaction subagent.

    State file is written immediately on dispatch (fire-and-forget), not after subagent completion. The cooldown tracks "a compaction was initiated," not "a compaction succeeded."

    This step is MANDATORY every session. You MUST read the state file and make the judgment. The only thing that may be skipped is the subagent dispatch when no trigger is met. This procedure must be completed before responding to the user NO MATTER WHAT

Memory Recall

When the user's question may relate to past memory, use the hipocampus-recall skill for structured retrieval. See hipocampus/skills/recall/SKILL.md.

End-of-Task Checkpoint (MANDATORY)

After completing any task, dispatch a subagent to append a structured log to memory/YYYY-MM-DD.md.

Compose the subagent task:

Append the following to memory/YYYY-MM-DD.md:

[Topic Name] [type]

  • request: [what the user asked]
  • analysis: [what you researched/analyzed]
  • decisions: [choices made with rationale]
  • outcome: [what was done, files changed]
  • references: [knowledge/ files, external sources]

Where type is: project | feedback | user | reference

For feedback entries, use:

[Feedback Topic] [feedback]

  • rule: [the behavioral rule]
  • why: [reason given]
  • how-to-apply: [when/where this applies]

The subagent only needs to do one thing: append to the daily log. This is the source of truth — everything else (SCRATCHPAD, WORKING, TASK-QUEUE) is updated lazily at next session start or by the agent naturally during work.

After appending to the daily log, the subagent should also increment the checkpoint counter in memory/.compaction-state.json: read the file, increment checkpointsSinceLastCompaction by 1, write back. If the file or field is missing, start from 0.

The subagent needs the task summary you provide — it doesn't have access to the conversation.

Priority if timeout imminent (no time for subagent — write directly to memory/YYYY-MM-DD.md)

Proactive Session Dump

Do not wait for task completion to write to the daily log. Proactively dispatch a subagent to append to memory/YYYY-MM-DD.md when:

  • The conversation has been going for ~20+ messages without a checkpoint
  • You sense the context is getting large
  • A significant decision or analysis was just completed, even if the overall task isn't done
  • You're switching between topics within the same task

Compose the subagent task with a summary of what to dump, same as the checkpoint format. The subagent writes the file; the main session stays clean.

This protects against context compression — if the platform compresses your conversation history, undumped details are lost forever. Write early, write often. The daily log is append-only, so multiple dumps in the same session are fine.

What NOT to Save

When composing checkpoint content for the subagent, exclude:

  • Secrets — API keys, tokens, passwords, credentials. If encountered, write [REDACTED].
  • Code snippets >5 lines — use file path + line range instead
  • git diff/log output — use commit hash instead
  • Debugging intermediate attempts — record final solution only
  • File tree / directory listings — derivable from project
  • Stack traces — compress to 1-line error message
  • Content already in SCRATCHPAD/WORKING/TASK-QUEUE — no duplication
  • Ephemeral task state — only useful within current session

File Size Targets

FileTargetWhen Exceeded
ROOT.md~100 lines (~3K tokens)Automatic recursive self-compression
SCRATCHPAD~150 linesRemove completed items
WORKING~100 linesRemove completed tasks
TASK-QUEUE~50 linesArchive completed items

Rules

  • Long-term facts are managed by platform auto memory. No separate MEMORY.md file.
  • Raw daily logs (memory/YYYY-MM-DD.md): permanent. Never delete or edit after session.
  • ROOT.md: managed by compaction process. Do not manually edit.
  • All memory writes via subagent — never pollute main session with memory operations.
  • If this session ends NOW, the next session must be able to continue immediately.
  • Don't skip checkpoints — lost context means you forget.

Edge Cases

  • Midnight-spanning session: Use the session start date for the raw log file name. Do not split across dates.
  • Returning after long absence: "Most recent daily" means the latest file that exists, whether it's from yesterday or last week.

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!