SkillAtlasSkill 详情

cmf

Common Metadata Framework (CMF) is a metadata tracking and versioning system for ML pipelines.

审核状态:已审核Quality 80Security 88

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月13日

Common Metadata Framework (CMF)

Deploy Docs PyPI version Docs License

Common Metadata Framework (CMF) is a metadata tracking and versioning system for ML pipelines. It tracks code, data, and pipeline metrics—offering Git-like metadata management across distributed environments.


🚀 Features

  • ✅ Track artifacts (datasets, models, metrics) using content-based hashes
  • ✅ Automatically logs code versions (Git) and data versions (DVC)
  • ✅ Push/pull metadata via CLI across distributed sites
  • ✅ REST API for direct server interaction
  • ✅ Implicit & explicit tracking of pipeline execution
  • ✅ Fine-grained or coarse-grained metric logging

🏛 Quick Start

Get started with CMF in minutes using our example ML pipeline:

📖 Try the Getting Started Example

This example demonstrates:

  • Initializing a CMF project
  • Tracking an ML pipeline with multiple stages (parse → featurize → train → test)
  • Versioning datasets and models
  • Pushing artifacts and metadata
  • Querying tracked metadata

📦 Installation

Requirements

  • Linux/Ubuntu/Debian
  • Python: Version 3.9 to 3.11 (3.10 recommended)
  • Git (latest)

Virtual Environment

Conda
conda create -n cmf python=3.10
conda activate cmf
Virtualenv
virtualenv --python=3.10 .cmf
source .cmf/bin/activate

Install CMF

Latest from GitHub
pip install git+https://github.com/HewlettPackard/cmf
Stable from PyPI
pip install cmflib

Server Setup

📖 Follow the CMF Server Installation Guide


📘 Documentation


🧠 How It Works

CMF tracks pipeline stages, inputs/outputs, metrics, and code. It supports decentralized execution across datacenters, edge, and cloud.

  • Artifacts are versioned using DVC (.dvc files).
  • Code is tracked with Git.
  • Metadata is logged to relational DB (e.g., SQLite, PostgreSQL)
  • Sync metadata with cmf metadata push and cmf metadata pull.

🏛 Architecture

CMF is composed of:

  • cmflib - Metadata library provides API to log/query metadata
  • CMF Client – CLI to sync metadata with server, push/pull artifacts to the user-specified repo, push/pull code from Git
  • CMF Server – REST API for metadata merge
  • Central Repositories – Git (code), DVC (artifacts), CMF (metadata)


🔧 Sample Usage

from cmflib.cmf import Cmf
from ml_metadata.proto import metadata_store_pb2 as mlpb

metawriter = Cmf(filepath="mlmd", pipeline_name="test_pipeline")

context: mlpb.Context = metawriter.create_context(
    pipeline_stage="prepare",
    custom_properties={"user-metadata1": "metadata_value"}
)

execution: mlpb.Execution = metawriter.create_execution(
    execution_type="Prepare",
    custom_properties={"split": split, "seed": seed}
)

artifact: mlpb.Artifact = metawriter.log_dataset(
    "artifacts/data.xml.gz", "input",
    custom_properties={"user-metadata1": "metadata_value"}
)
cmf                          # CLI to manage metadata and artifacts
cmf init                     # Initialize artifact repository
cmf init show                # Show current CMF config
cmf metadata push            # Push metadata to server
cmf metadata pull            # Pull metadata from server

➡️ For the complete list of commands, please refer to the Command Reference


✅ Benefits

  • Full ML pipeline observability
  • Unified metadata, artifact, and code tracking
  • Scalable metadata syncing
  • Team collaboration on metadata

🎤 Talks & Publications


🌐 Related Projects


🤝 Community


📄 License

Licensed under the Apache 2.0 License


© Hewlett Packard Enterprise. Built for reproducibility in ML.

其他

中风险

  • 来源需自行核对维护者身份。
  • 未检测到明显脚本安装指令。
  • 未检测到明显外部权限要求。
  • 未检测到高风险命令。
  • 扫描发现:1 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/HewlettPackard/cmf.git
  3. 将 "skills/cmf" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/HewlettPackard/cmf.git
  3. 将 "skills/cmf" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/HewlettPackard/cmf.git
  3. 将 "skills/cmf" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/HewlettPackard/cmf.git
  3. 将 "skills/cmf" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/HewlettPackard/cmf.git
  3. 将 "skills/cmf" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: cmf
description: >
  Use when working with the Common Metadata Framework (CMF) — initializing CMF in a project,
  instrumenting ML pipeline code, syncing metadata and artifacts with a CMF Server, querying
  lineage and artifacts, or setting up the CMF MCP server for AI assistant integration.
  Routes to the appropriate specialized skill based on the task.
version: 1.0.0
user_invocable: true
argument_hint: "<task>"

Route to the appropriate sub-skill based on the user's task. If the task is ambiguous, ask one clarifying question before routing.

Quickstart (default path)

If the user wants to add CMF to their project for the first time, or the request is general/unclear, route here:

cmf-init — Install cmflib, configure a storage backend, and create the mlmd metadata store.

Routing Table

TaskSub-skill
Install CMF, configure storage backend (local, S3, MinIO, SSH, OSDF), run cmf initcmf-init
Add Cmf() calls to existing ML pipeline code — contexts, executions, datasets, models, metricscmf-instrument
Push or pull metadata and artifacts using the CMF CLIcmf-sync
Query pipeline history, artifact lineage, and execution metadata using CmfQuerycmf-query
Deploy CMF Server with Docker Compose, or connect to an existing shared servercmf-server
Set up the CMF MCP server so an AI assistant can query CMF metadatacmf-mcp

Key Concepts

  • Pipeline — top-level grouping of stages, identified by pipeline_name passed to Cmf()
  • Context — a stage type (e.g. "train", "evaluate"), created with create_context()
  • Execution — one run of a stage, created with create_execution(); hyperparameters go here
  • Artifact — dataset, model, or metrics file logged as input or output of an execution
  • mlmd — the local SQLite file that stores all metadata; pushed to CMF Server for collaboration

Docs: Getting Started · cmflib API · CLI Reference · MCP Server

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!