复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
🔔 Claude Scientific Skills is now Scientific Agent Skills.
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
🔔 Claude Scientific Skills is now Scientific Agent Skills. Same skills, broader compatibility — now works with any AI agent that supports the open Agent Skills standard, not just Claude.
New: K-Dense BYOK — A free, open-source AI co-scientist that runs on your desktop, powered by Scientific Agent Skills. Bring your own API keys, pick from 40+ models, and get a full research workspace with web search, file handling, 100+ scientific databases, and access to all 165 skills in this repo. Your data stays on your computer, and you can optionally scale to cloud compute via Modal for heavy workloads. Get started here.
🎥 Webinar recording — Getting Started with K-Dense BYOK A hands-on walkthrough of K-Dense BYOK, our free, open-source AI co-scientist that runs locally on your own machine and is powered by Scientific Agent Skills. We cover how to set it up, bring your own API keys, and run real research workflows with these skills. No prior technical experience needed. Watch the recording →
Stay up to date: Follow K-Dense on X, LinkedIn, YouTube, and Reddit for new skills, release announcements, walkthroughs, research workflow demos, and examples you can use with your own AI agent.
📄 Paper: Scientific Agent Skills is described in Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents (arXiv:2609.00065). If you use these skills in your research, please cite the paper.
A comprehensive collection of 165 ready-to-use scientific and research skills (covering cancer genomics, individual-level 1000 Genomes queries, hosted regulatory-sequence prediction, live pathogen-variant surveillance, analytical method validation, PK/PD modelling and dose selection, full-text biomedical and regulatory literature retrieval, drug-target binding, bounded biomedical knowledge graph search, molecular dynamics, RNA velocity, microbiome foundation models, geospatial science, time series forecasting, scientific ML resource discovery via Hugging Science, 78+ scientific databases, and more) for any AI agent that supports the open Agent Skills standard, created by K-Dense. The repository is also a portable Agent Plugins package (plugin.json + skills/), so plugin-capable clients can load the whole collection as one plugin. Works with Cursor, Claude Code, Codex, Google Antigravity, and more. Transform your AI agent into a research assistant capable of executing complex multi-step scientific workflows across biology, chemistry, medicine, and beyond.
⭐ Help make AI for science easier to discover: If Scientific Agent Skills saves you time, teaches your agent a workflow, or helps your lab move faster, please star this repository. A star is a public signal that these open, reusable research skills are worth maintaining: it helps scientists, engineers, and open-source contributors find the project, shows which agent-skill standards are gaining real adoption, and gives us a clear reason to keep expanding the collection for the community.
These skills enable your AI agent to seamlessly work with specialized scientific libraries, databases, and tools across multiple scientific domains. While the agent can use any Python package or API on its own, these explicitly defined skills provide curated documentation and examples that make it significantly stronger and more reliable for the workflows below:
Transform your AI coding agent into an 'AI Scientist' on your desktop!
🎬 New to Scientific Agent Skills? Watch our Getting Started with Scientific Agent Skills video for a quick walkthrough.
Recorded walkthroughs of these skills on real research tasks, from the K-Dense YouTube channel:
| Video | What it covers |
|---|---|
| Skills 101: Build Your Own Scientific Agent Skill | Writing, testing, and packaging a new skill from scratch |
| Literature Review and Hypothesis Generation | Searching the literature and generating grounded hypotheses |
| Draft and Budget an Experimental Protocol | Turning a planned experiment into a costed, written protocol |
| Draft Responses to Reviewer Comments | Building a point-by-point rebuttal from reviewer feedback |
| Can AI Reproduce a Nature Medicine Paper? | An end-to-end reproduction attempt on a published analysis |
This repository provides 165 scientific and research skills organized into the following categories:
Each skill includes:
SKILL.md)scripts/ — CI blocks a pull request that adds bundled tooling without onescripts/ has a suite under tests/, plus a repo-wide structural contract (frontmatter, link resolution, script parsing, --help behavior) that runs on every pull requestInstall Scientific Agent Skills with a single command:
npx skills add K-Dense-AI/scientific-agent-skills
This is a common standards-based installer for supported Agent Skills hosts, including current versions of Claude Code, Claude Cowork, Codex, Gemini CLI, Google Antigravity, and Cursor. Confirm installation paths and optional metadata behavior in your host's current documentation.
gh skill)If you use the GitHub CLI (v2.90.0+), you can install skills with gh skill:
# Browse and install interactively
gh skill install K-Dense-AI/scientific-agent-skills
# Install a specific skill directly
gh skill install K-Dense-AI/scientific-agent-skills scanpy
# Target a specific agent host
gh skill install K-Dense-AI/scientific-agent-skills --agent cursor
gh skill install K-Dense-AI/scientific-agent-skills --agent claude-code
gh skill install K-Dense-AI/scientific-agent-skills --agent codex
gh skill install K-Dense-AI/scientific-agent-skills --agent gemini
gh skill automatically installs to the correct directory for your agent host and records provenance metadata for supply chain integrity.
Pin to a specific release tag or commit SHA for reproducible installs:
# Pin to a release tag
gh skill install K-Dense-AI/scientific-agent-skills --pin v2.66.0
# Pin to a commit SHA
gh skill install K-Dense-AI/scientific-agent-skills --pin abc123def
# Check for updates interactively
gh skill update
# Update all installed skills
gh skill update --all
This repository is a valid Agent Plugins 1.0.0 package: root plugin.json plus Agent Skills under skills/. Clients that support the standard discover every immediate child of skills/ that contains a SKILL.md.
Cursor — symlink or copy the repo into the local plugins directory, then reload:
mkdir -p ~/.cursor/plugins/local
ln -s "$(pwd)" ~/.cursor/plugins/local/scientific-agent-skills
Restart Cursor or run Developer: Reload Window, then confirm the plugin and its skills appear under Customize. See Cursor plugins.
Codex — install from a local checkout (confirm the current CLI flag names in Codex docs):
codex plugins install .
Compatible clients (Cursor, Codex, GitHub Copilot, VS Code, Kiro, and others listed at agent-plugins.org) share the same package layout; installation UX stays client-specific.
Agent hosts differ in install paths, discovery settings, and support for optional frontmatter fields. npx skills add (Option 1) commonly installs into the ~/.agents/skills/ convention, with project-scoped installs under .agents/skills/; confirm both paths against your host's current documentation. To install manually on a host configured to scan one of those locations:
git clone https://github.com/K-Dense-AI/scientific-agent-skills.git ~/.agents/skills/scientific-agent-skills # user-level
git clone https://github.com/K-Dense-AI/scientific-agent-skills.git .agents/skills/scientific-agent-skills # project-level
For Hermes versions that support skill taps, add the repository as a tap:
hermes skills tap add K-Dense-AI/scientific-agent-skills
Every SKILL.md has YAML frontmatter, but legacy and community skills vary in metadata formatting (block or flow style) and optional extension fields. Repository updates must keep metadata.version as a quoted numeric string and pass canonical skills-ref validate ./skills/<skill-name> checks. Hosts may interpret optional metadata and credential prompts differently, so verify behavior on the target host. Because 165 skills add up to a lot of standing context, consider installing a topical subset rather than the whole collection.
NemoClaw note: NemoClaw runs agents inside NVIDIA OpenShell with default-deny outbound networking. Skills are discovered and loaded normally, but any skill that needs the network — package installs via
uv, or API calls (Exa, Parallel, Benchling, NCBI, Materials Project, …) — only works once the operator pre-approves the relevant domains in the OpenShell TUI.
That's it! A compatible host can discover the skills from its configured paths and use them when relevant. You can also invoke any skill manually by mentioning the skill name in your prompt.
Skills can execute code and influence your coding agent's behavior. Review what you install.
Agent Skills are powerful — they can instruct your AI agent to run arbitrary code, install packages, make network requests, and modify files on your system. A malicious or poorly written skill has the potential to steer your coding agent into harmful behavior.
We take security seriously. All contributions go through a review process, and we run LLM-based security scans (via Cisco AI Defense Skill Scanner) on every skill in this repository. However, as a small team with a growing number of community contributions, we cannot guarantee that every skill has been exhaustively reviewed for all possible risks.
It is ultimately your responsibility to review the skills you install and decide which ones to trust.
We recommend the following:
SKILL.md before installing. Each skill's documentation describes what it does, what packages it uses, and what external services it connects to. If something looks suspicious, don't install it.K-Dense-AI) have been through our internal review process. Community-contributed skills have been reviewed to the best of our ability, but with limited resources.uv pip install cisco-ai-skill-scanner
skill-scanner scan /path/to/skill --use-behavioral
Skills are scanned weekly — incrementally, so unchanged skills carry their previous findings forward, with a full rescan of everything at least every 30 days and whenever the scanner or model changes — and the results are published to docs/security-report.md. See SECURITY.md for our security policy, what is in scope, how to report a vulnerability privately, and how to contest a scan finding. We try to address security gaps as they arise.
Scientific Agent Skills is powered by 50+ incredible open source projects maintained by dedicated developers and research communities worldwide. Projects like Biopython, Scanpy, RDKit, scikit-learn, PyTorch Lightning, and many others form the foundation of these skills.
If you find value in this repository, please consider supporting the projects that make it possible:
👉 View the full list of projects to support
The docx, pdf, pptx, and xlsx document skills are created and maintained by Anthropic and vendored here from anthropics/skills. They are used under Anthropic's terms — see each skill's LICENSE.txt — and we track upstream so you get their latest improvements. All credit for those four skills goes to Anthropic.
SKILL.md files for specific requirements)The skills use uv as the package manager for installing Python dependencies. Install it using the instructions for your operating system:
macOS and Linux:
curl -LsSf https://astral.sh/uv/install.sh | sh
Windows:
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
Alternative (via pip):
pip install uv
After installation, verify it works by running:
uv --version
For more installation options and details, visit the official uv documentation.
Once you've installed the skills, you can ask your AI agent to execute complex multi-step scientific workflows. Here are some example prompts:
Goal: Prioritize EGFR inhibitor candidates for preclinical lung-cancer research
Prompt:
Use available skills you have access to whenever possible. Query ChEMBL for EGFR inhibitors (IC50 < 50nM), analyze structure-activity relationships
with RDKit, generate improved analogs with datamol, perform virtual screening with DiffDock
against AlphaFold EGFR structure, search PubMed for resistance mechanisms, check COSMIC for
mutations, and create visualizations and a comprehensive report.
Skills Used: database-lookup, rdkit, datamol, diffdock, paper-lookup, scientific-visualization
Goal: Comprehensive analysis of 10X Genomics data with public data integration
Prompt:
Use available skills you have access to whenever possible. Load 10X dataset with Scanpy, perform QC and doublet removal, integrate with Cellxgene
Census data, identify cell types using NCBI Gene markers, run differential expression with
PyDESeq2, infer gene regulatory networks with Arboreto, enrich pathways via Reactome/KEGG,
and identify therapeutic targets with Open Targets.
Skills Used: scanpy, cellxgene-census, database-lookup, pydeseq2, arboreto
Goal: Integrate RNA-seq, proteomics, and metabolomics to predict patient outcomes
Prompt:
Use available skills you have access to whenever possible. Analyze RNA-seq with PyDESeq2, process mass spec with pyOpenMS, integrate metabolites from
HMDB/Metabolomics Workbench, map proteins to pathways (UniProt/KEGG), find interactions via
STRING, correlate omics layers with statsmodels, build predictive model with scikit-learn,
and search ClinicalTrials.gov for relevant trials.
Skills Used: pydeseq2, pyopenms, database-lookup, statsmodels, scikit-learn
Goal: Discover allosteric modulators for protein-protein interactions
Prompt:
Use available skills you have access to whenever possible. Retrieve AlphaFold structures, identify interaction interface with BioPython, search ZINC
for allosteric candidates (MW 300-500, logP 2-4), filter with RDKit, dock with DiffDock,
rank with DeepChem, check PubChem suppliers, search USPTO patents, and optimize leads with
MedChem/molfeat.
Skills Used: database-lookup, biopython, rdkit, diffdock, deepchem, medchem, molfeat
Goal: Annotate a synthetic or properly de-identified VCF for hereditary-cancer research and qualified review
Prompt:
Use available skills you have access to whenever possible. Work only with authorized synthetic
or de-identified data. Parse the VCF with pysam, annotate variants with Ensembl VEP, retrieve
ClinVar/COSMIC/NCBI Gene/UniProt evidence, and verify literature sources. Build an evidence-
traceable research summary with scientific-writing. If clinical-reports is used, create only a
visibly marked draft structure from a verified source-fact manifest for qualified review; do not
diagnose, assess individual risk, recommend treatment, or determine trial eligibility.
Skills Used: pysam, database-lookup, paper-lookup, scientific-writing, clinical-reports
Goal: Analyze gene regulatory networks from RNA-seq data
Prompt:
Use available skills you have access to whenever possible. Query NCBI Gene for annotations, retrieve sequences from UniProt, identify interactions via
STRING, map to Reactome/KEGG pathways, analyze topology with Torch Geometric, reconstruct
GRNs with Arboreto, assess druggability with Open Targets, model with PyMC, visualize
networks, and search GEO for similar patterns.
Skills Used: database-lookup, torch-geometric, arboreto, pymc, networkx, scientific-visualization
📖 Want more examples? Check out docs/examples.md for comprehensive workflow examples and detailed use cases across all scientific domains.
This repository contains 165 scientific and research skills organized across multiple domains. Each skill provides comprehensive documentation, code examples, and best practices for working with scientific libraries, databases, and tools.
Note: The Python package and integration skills listed below are explicitly defined skills — curated with documentation, examples, and best practices for stronger, more reliable performance. They are not a ceiling: the agent can install and use any Python package or call any API, even without a dedicated skill. The skills listed simply make common workflows faster and more dependable.
A unified database-lookup skill provides deterministic REST API access to 78 public databases across all domains, with retrieval contracts, pagination/count reconciliation, and endpoint provenance. Dedicated skills cover specialized data platforms. Multi-database packages like BioServices (~40 bioinformatics services), BioPython (39 NCBI sub-databases via Entrez), and gget (20+ genomics databases) add further coverage.
datasets, transformers, and gradio_client)--standard profile<1220>/<1225>/<1226>, the CLSI EP series, and ISO/IEC 17025 cited by designation and scope only; stdlib-only statistics, no network access📖 For complete details on all skills, see docs/skills.md
💡 Looking for practical examples? Check out docs/examples.md for comprehensive workflow examples across all scientific domains.
Deep dives, benchmarks, and guides from the K-Dense blog that are directly relevant to using the skills in this repository.
skills/<name>/.SKILL.md and scripts/, scan before installing, and pin versions instead of tracking a branch.AGENTS.md profiles supplying the "how to think" layer alongside the "what to do" procedures in these skills.SKILL.md / AGENTS.md expert profiles by distilling how a given practitioner reasons.We welcome contributions to expand and improve this scientific skills repository!
For detailed instructions on adding or updating a skill, see CONTRIBUTING.md. The guide covers repository structure, required SKILL.md frontmatter, Agent Skills specification requirements, versioning, validation, security scanning, and pull request expectations.
✨ Add New Skills
📚 Improve Existing Skills
🐛 Report Issues
git checkout -b feature/amazing-skill)SKILL.md files with required frontmatter and metadata.versiontests/<skill-name>/ if your skill ships scripts/git commit -m 'Add amazing skill')git push origin feature/amazing-skill)✅ Adhere to the Agent Skills Specification — Every skill must follow the official spec (valid SKILL.md frontmatter, naming conventions, directory structure)
✅ Include a quoted metadata.version value in every SKILL.md
✅ Increment metadata.version when updating an existing skill
✅ Maintain consistency with existing skill documentation format
✅ Ensure all code examples are tested and functional
✅ Follow scientific best practices in examples and workflows
✅ Update relevant documentation when adding new capabilities
✅ Provide clear comments and docstrings in code
✅ Include references to official documentation
Every skill that ships scripts/ must have a test suite under tests/<skill-name>/ and an entry in tests/skill-requirements.toml. This is enforced — tests/_meta fails a pull request that adds bundled tooling without one, and it also runs a repo-wide structural contract over all skills (frontmatter conformance, SKILL.md length, local links resolving, scripts parsing, no shipped bytecode, no hardcoded local paths, --help behavior).
# Structural contract and coverage guard — seconds, no scientific packages needed
uv run python -m pytest tests/_meta -q
# One skill's suite
uv run --with pytest python -m pytest tests/<skill-name> -q
# Every suite, each in its own throwaway environment
uv run python tests/run_all.py --isolated
The Skill Tests workflow runs the contract plus the standard-library-only suites on every pull request; the full --isolated sweep builds ~100 environments and is run locally or on a schedule.
All skills in this repository are security-scanned using Cisco AI Defense Skill Scanner, an open-source tool that detects prompt injection, data exfiltration, and malicious code patterns in Agent Skills.
If you are contributing a new skill, we recommend running the scanner locally before submitting a pull request:
uv pip install cisco-ai-skill-scanner
skill-scanner scan /path/to/your/skill --use-behavioral
Note: A clean scan result reduces noise in review, but does not guarantee a skill is free of all risk. Contributed skills are also reviewed manually before merging.
Contributors are recognized in our community and may be featured in:
Your contributions help make scientific computing more accessible and enable researchers to leverage AI tools more effectively!
This project builds on 50+ amazing open source projects. If you find value in these skills, please consider supporting the projects we depend on.
Problem: Skills not loading
SKILL.md fileProblem: Missing Python dependencies
SKILL.md file for required packagesuv pip install package-nameProblem: API rate limits
Problem: Authentication errors
SKILL.md for authentication setupProblem: Outdated examples
Problem: gh skill install or docs link to scientific-skills/ fails (v2.43.0+)
skills/ (not scientific-skills/) to match the Agent Skills layout expected by GitHub CLIscientific-skills/<name> to skills/<name>gh skill install K-Dense-AI/scientific-agent-skills after pulling the latest releaseQ: Is this free to use?
A: Yes! This repository is MIT licensed. However, each individual skill has its own license specified in the license metadata field within its SKILL.md file—be sure to review and comply with those terms.
Q: Why are all skills grouped together instead of separate packages?
A: We believe good science in the age of AI is inherently interdisciplinary. Bundling all skills together makes it trivial for you (and your agent) to bridge across fields—e.g., combining genomics, cheminformatics, clinical data, and machine learning in one workflow—without worrying about which individual skills to install or wire together.
Q: Can I use this for commercial projects?
A: The repository itself is MIT licensed, which allows commercial use. However, individual skills may have different licenses—check the license field in each skill's SKILL.md file to ensure compliance with your intended use.
Q: Do all skills have the same license?
A: No. Each skill has its own license specified in the license metadata field within its SKILL.md file. These licenses may differ from the repository's MIT License. Users are responsible for reviewing and adhering to the license terms of each individual skill they use.
Q: How often is this updated?
A: We regularly update skills to reflect the latest versions of packages and APIs. Major updates are announced in release notes.
Q: Can I use this with other AI models?
A: The core SKILL.md format follows the open Agent Skills standard. Installation paths, discovery, and optional metadata support vary by host and version, so confirm your target host's current documentation.
Q: Do I need all the Python packages installed?
A: No! Only install the packages you need. Each skill specifies its requirements in its SKILL.md file.
Q: What if a skill doesn't work?
A: First check the Troubleshooting section. If the issue persists, file an issue on GitHub with detailed reproduction steps.
Q: Do the skills work offline?
A: Database skills require internet access to query APIs. Package skills work offline once Python dependencies are installed.
Q: Can I contribute my own skills?
A: Absolutely! We welcome contributions. See the Contributing section for guidelines and best practices.
Q: How do I report bugs or suggest features?
A: Open an issue on GitHub with a clear description. For bugs, include reproduction steps and expected vs actual behavior.
Need help? Here's how to get support:
SKILL.md and references/ foldersIf you use Scientific Agent Skills in your research or project, please cite our paper:
Timothy Kassis, Vinayak Agarwal, Yuhuan He, Darshil Patel, and Aubrey M. Brueckner. Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents. arXiv:2609.00065, 2026. https://arxiv.org/abs/2609.00065
When relevant, also cite the individual skill or skills that materially supported your work.
GitHub's Cite this repository button, backed by CITATION.cff, produces the same paper citation in APA or BibTeX.
The paper citation helps others find the repository, understand the broader skill ecosystem used in your workflow, and credit the maintenance effort behind Scientific Agent Skills. Individual skill citations give more precise credit for the specific package, database, or workflow guidance your agent used.
Recommended practice:
@misc{kassis2026scientificagentskills,
title = {Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents},
author = {Kassis, Timothy and Agarwal, Vinayak and He, Yuhuan and Patel, Darshil and Brueckner, Aubrey M.},
year = {2026},
eprint = {2609.00065},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2609.00065}
}
Kassis, T., Agarwal, V., He, Y., Patel, D., & Brueckner, A. M. (2026). Scientific Agent Skills: A library of procedural knowledge for research agents. arXiv. https://arxiv.org/abs/2609.00065
Kassis, Timothy, et al. "Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents." arXiv, 2026, arxiv.org/abs/2609.00065.
Kassis, T., Agarwal, V., He, Y., Patel, D., & Brueckner, A. M. (2026). Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents. arXiv:2609.00065. https://arxiv.org/abs/2609.00065
If you also need to cite a specific version of the repository itself (for example, to pin the exact skill set an analysis ran against), add a software citation alongside the paper and record the release tag or commit you used:
@software{scientific_agent_skills_2026,
author = {{K-Dense Inc.}},
title = {Scientific Agent Skills: A Comprehensive Collection of Scientific Tools for AI Agents},
year = {2026},
url = {https://github.com/K-Dense-AI/scientific-agent-skills},
note = {165 skills covering databases, packages, integrations, and analysis tools}
}
When citing a specific skill, include the skill name, version from metadata.version in that skill's SKILL.md, and the direct skill URL. For example:
@software{scientific_agent_skills_astropy_2026,
author = {{K-Dense Inc.}},
title = {Astropy Skill for Scientific Agent Skills},
year = {2026},
url = {https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/skills/astropy},
note = {Version 1.0, part of Scientific Agent Skills}
}
Plain text format:
Astropy skill for Scientific Agent Skills, version 1.0.
K-Dense Inc. (2026).
https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/skills/astropy
We appreciate acknowledgment in publications, presentations, or projects that benefit from these skills.
This project is licensed under the MIT License.
Copyright © 2026 K-Dense Inc. (k-dense.ai)
See LICENSE.md for full terms.
⚠️ Important: Each skill has its own license specified in the
licensemetadata field within itsSKILL.mdfile. These licenses may differ from the repository's MIT License and may include additional terms or restrictions. Users are responsible for reviewing and adhering to the license terms of each individual skill they use.
name: datalad
description: Retrieve, version, and publish scientific datasets with DataLad and git-annex, and capture computational provenance with datalad run, rerun, and containers-run. Use when cloning or fetching data from OpenNeuro, DANDI, datasets.datalad.org, or any DataLad dataset; when a file in a dataset reads as a broken symlink or a small pointer instead of real data; when an analysis needs a machine-readable record of how each output was produced so it can be re-executed; or when publishing a dataset to siblings such as a GitHub repository plus a storage remote. Also use to decide between DataLad and plain Git for a data-carrying repository.
compatibility: Needs datalad 1.6.x on Python 3.10+, plus git and git-annex 10.x. git-annex is not written in Python but installs as a prebuilt wheel from PyPI (`uv pip install git-annex`), from a system package manager, or from conda-forge. Container-based provenance also needs datalad-container (1.2.x) and Singularity/Apptainer or Docker. clone, get, and push need network access; credentialed remotes read secrets from the system keyring or from DATALAD_CREDENTIAL_<NAME>_<COMPONENT> environment variables.
license: MIT
allowed-tools: Read Write Edit Bash
metadata:
version: "1.0"
skill-author: Dylan PulverDataLad is a data management layer over Git and git-annex. Git tracks the dataset structure, small text files, and the history. git-annex tracks the content of large files, storing each file as a key and keeping the bytes somewhere that is not necessarily the local repository.
That split is the single most important thing to internalise, because it means a freshly
cloned dataset contains the full history and the full file listing while containing almost
none of the data. A 100 TB dataset clones in seconds and occupies a few megabytes. The
bytes arrive only when asked for, per file, with datalad get.
The second thing DataLad adds is provenance. datalad run executes a command and commits
the result together with a machine-readable record of the command, its inputs, and its
outputs. datalad rerun reads that record back and re-executes it. This turns "how was
this figure produced" from an archaeology problem into a command.
Use DataLad when any of the following holds:
datasets.datalad.org,
which are distributed as DataLad datasets.Use plain Git when the repository is code and text only, everything fits comfortably in Git, and nobody needs partial checkouts. DataLad on top of a small pure-code repository adds indirection without buying anything.
# git-annex is NOT written in Python but is available from PyPI if you already
# have git itself installed:
uv pip install git-annex
# You can also install it first from the system
# (Debian/Ubuntu: apt install git-annex; macOS: brew install git-annex;
# conda-forge: conda install -c conda-forge git-annex)
uv pip install datalad
uv pip install datalad-container # only for containers-run
datalad wtf --section dependencies # confirm git-annex version is visible
The PyPI git-annex package ships the prebuilt binary as a wheel for Linux, macOS, and
Windows rather than building the Haskell sources, so it installs like any other Python
dependency and can be pinned in the same environment as DataLad. It does not bring git
along with it.
datalad wtf prints the resolved environment and is the first thing to run when behaviour
looks impossible. An old or missing git-annex is behind a large share of confusing errors.
DataLad itself is MIT licensed. git-annex is a separate tool under the AGPL, which matters only if you redistribute a modified git-annex rather than call it.
After datalad clone, annexed files exist as symlinks into .git/annex/objects/ (or as
small pointer files where symlinks are unavailable, such as on Windows or a crippled
filesystem). Nothing has downloaded the content yet.
datalad clone https://github.com/OpenNeuroDatasets/ds000001.git
cd ds000001
ls sub-01/anat/ # the file is listed
python -c "import nibabel; nibabel.load('sub-01/anat/sub-01_T1w.nii.gz')" # fails
datalad get sub-01/anat/sub-01_T1w.nii.gz # now it works
The failure mode to recognise: a tool reports the file as empty, truncated, corrupt, "not
a gzip file", or a broken symlink, and the file size on disk is a few hundred bytes. That
is a pointer, not a corrupted download. Run datalad get before reading data, and treat
"file exists" as insufficient evidence that its content is present.
Before an analysis touches a directory, fetch it explicitly:
datalad get sub-01/ # everything under a path
datalad get -r . # everything, including subdatasets
datalad get -n -r . # subdataset structure only, no file content
datalad status --annex reports how much content is present locally, and
git annex whereis <path> reports which repositories hold a given file. whereis reads
recorded state and does not contact the remotes, so it tells you what git-annex last
learned rather than what is true right now.
See data-access.md for finding datasets, subdataset behaviour, dropping content safely, and repairing a dataset.
datalad run is the reason to reach for DataLad in a methods context. It saves the
command alongside its effect, in the same commit:
datalad run -m "extract brain mask" \
--input "sub-01/anat/sub-01_T1w.nii.gz" \
--output "derivatives/sub-01_brain.nii.gz" \
"bet {inputs} {outputs} -m"
What each part does, and why skipping it hurts:
--input retrieves the content before running, so the command does not fail on a
pointer. It also records the dependency, which is what lets rerun fetch the same
inputs on a different machine.--output unlocks or removes the target first, so git-annex does not refuse to write
over content it is protecting. Without it, a second run of the same command commonly
fails with a permission error on an annexed file that looks read-only.{inputs} and {outputs} expand to those values. {pwd}, {dspath}, and {tmpdir}
are also available, and {inputs[0]} indexes individual entries.=== Do not change lines below ===
and ^^^ Do not change lines above ^^^. Do not hand-edit that block; rerun parses it.datalad run refuses to start when the dataset has unsaved modifications, because an
unclean starting state makes the record unreliable. Save or discard first, or pass
--explicit to declare that the listed inputs and outputs are the complete story. Check a
command before committing to it with --dry-run basic or --dry-run command.
A run that changes nothing produces no commit, exactly as datalad save does.
datalad rerun # redo the run recorded at HEAD
datalad rerun --report # show what would be done, change nothing
datalad rerun --script recompute.sh # extract the commands instead of running them
datalad rerun --since <commit> -b check <revision> # replay a range onto a new branch
Rerunning onto a branch (-b) is the safe way to test reproducibility: the replay lands
somewhere else, and a diff against the original branch answers whether the outputs came
back identical.
With the datalad-container extension, register an image once and every subsequent run
records which image produced the outputs:
datalad containers-add fsl --url docker://brainlife/fsl:6.0.4
datalad containers-run -n fsl -m "brain mask in container" \
--input "sub-01/anat/sub-01_T1w.nii.gz" \
--output "derivatives/sub-01_brain.nii.gz" \
"bet {inputs} {outputs} -m"
The image itself is tracked in the dataset, so the software environment travels with the
data and the provenance record rather than living in someone's shell history. When only
one container is configured, -n may be omitted.
See provenance.md for the STAMPED principles and the YODA
project layout, the run record format, --explicit and --assume-ready semantics, and
exporting provenance toward W3C PROV.
datalad status # what changed, including subdataset state
datalad save -m "add QC report" path/to/file
datalad save -m "checkpoint" -r # recurse into subdatasets
datalad save -m "small text file" --to-git notes.md
datalad save decides per file whether content goes to Git or to git-annex, following the
dataset's .gitattributes. Force a file into Git with --to-git, which is the right call
for code and small text files that should stay directly readable. The yoda procedure
(datalad create -c yoda) sets this up for code/, README.md, and CHANGELOG.md
automatically.
datalad create my_dataset # plain dataset
datalad create -c yoda my_analysis # analysis layout (code/ tracked in Git,
# README.md and CHANGELOG.md preconfigured)
datalad create -d . inputs/raw # register a new subdataset under an existing one
-c yoda applies the analysis project layout described in
provenance.md. -d . is what registers a new dataset as a
subdataset of the parent rather than leaving an unrelated repository inside it.
A DataLad dataset is usually published to two places at once: a Git hosting service for the history, and a storage remote for the annexed content.
datalad create-sibling-github myaccount/mydataset
git annex initremote store type=S3 bucket=my-bucket encryption=none autoenable=true
datalad siblings configure -s github --publish-depends store
datalad push --to github
The Git sibling and the storage sibling are created by different tools on purpose. A Git
sibling is a Git remote, and datalad create-sibling-* handles the hosting-service ones.
An S3 bucket (or WebDAV, or an SSH directory) is a git-annex special remote, not a Git
remote, so it is created with git annex initremote. datalad siblings picks the special
remote up afterwards and treats it like any other. Using datalad siblings add --url s3://... here is the mistake this section exists to prevent: --url is a Git remote URL,
S3 is not, and the push --to github below then fails on the --publish-depends hop.
--publish-depends is what stops the common broken publication: a Git repository whose
history references content that was never uploaded, so collaborators clone successfully
and then find every datalad get failing. Declaring the dependency makes the storage
sibling publish first, every time.
datalad push sends both the Git history and, by default (--data auto-if-wanted), the
annexed content the target is configured to want. Pass --data anything to push all
content regardless of the target's preferences.
See publishing.md for RIA stores, special remotes, credential handling, and configuring which sibling holds what.
git annex whereis sub-01/ # confirm another copy exists first
datalad drop sub-01/ # remove local content, keep the pointer
datalad drop --what all --reckless kill <path> # last resort, destroys data
datalad drop refuses by default when it cannot verify another copy of the content
exists, which is a safety check rather than an obstacle. --nocheck and --if-dirty are
deprecated; the current spelling is --reckless availability, and it means what it says.
--what selects between filecontent (the default), allkeys, datasets, and all.
| Symptom | Cause | Fix |
|---|---|---|
| File reads as empty, truncated, or a broken symlink | Content not retrieved; only the pointer is present | datalad get <path> |
| "Permission denied" writing an existing output | git-annex write-protects annexed content | Declare it with --output, or datalad unlock <path> |
datalad run refuses to start | Dataset has unsaved changes | datalad save first, or pass --explicit |
datalad drop refuses | No verified second copy of the content | Push to a sibling first, or accept --reckless availability |
Collaborator clones but every get fails | History published without the content | Publish the storage sibling, and set --publish-depends |
| Clone succeeds, subdataset directories are empty | Subdatasets are not installed by default | datalad get -n -r ., then get the paths you need |
| Commands behave impossibly | git-annex missing or too old | datalad wtf --section dependencies |
registry.datalad.org, OpenNeuro, DANDI, datasets.datalad.org and the ///
shortcut), clone and get options, subdataset handling, annex content states, dropping
and removing, and fsck repair.run and rerun options in full, containers-run, and the
current state of exporting DataLad provenance toward W3C PROV.create-sibling-* variants, RIA stores, special remotes, push semantics, and
credential handling.The bids skill covers the Brain Imaging Data Structure that most of the neuroimaging
datasets distributed through DataLad are organised in. A typical workflow clones a BIDS
dataset with DataLad, validates it with the BIDS tooling, then runs a BIDS-App under
datalad containers-run so the derivatives carry provenance.
datalad run chapter: https://handbook.datalad.org/en/latest/basics/101-108-run.htmlTopic scope for this skill was informed in part by @bcmcpher's MIT-licensed datalad-cli plugin (nineteen per-command slash-command skills). The text here is written independently and grounded in the upstream DataLad documentation; overlap is unavoidable because both cover DataLad, but the structure, style, and specific technical claims are different.
评论 (0)
暂无评论,成为第一个评论者吧!