SkillAtlasSkill 详情

marine-carbonate-chemistry

🔔 Claude Scientific Skills is now Scientific Agent Skills.

审核状态:已审核Quality 72Security 70

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年10月1日

Scientific Agent Skills

License: MIT arXiv Version Skills Databases Agent Skills Agent Plugins Security Scan Skill Tests Works with X LinkedIn YouTube Reddit

🔔 Claude Scientific Skills is now Scientific Agent Skills. Same skills, broader compatibility — now works with any AI agent that supports the open Agent Skills standard, not just Claude.

New: K-Dense BYOK — A free, open-source AI co-scientist that runs on your desktop, powered by Scientific Agent Skills. Bring your own API keys, pick from 40+ models, and get a full research workspace with web search, file handling, 100+ scientific databases, and access to all 181 skills in this repo. Your data stays on your computer, and you can optionally scale to cloud compute via Modal for heavy workloads. Get started here.

🎥 Webinar recording — Getting Started with K-Dense BYOK A hands-on walkthrough of K-Dense BYOK, our free, open-source AI co-scientist that runs locally on your own machine and is powered by Scientific Agent Skills. We cover how to set it up, bring your own API keys, and run real research workflows with these skills. No prior technical experience needed. Watch the recording →

Stay up to date: Follow K-Dense on X, LinkedIn, YouTube, and Reddit for new skills, release announcements, walkthroughs, research workflow demos, and examples you can use with your own AI agent.

📄 Paper: Scientific Agent Skills is described in Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents (arXiv:2609.00065). If you use these skills in your research, please cite the paper.

A collection of 181 scientific and research skills for AI agents, created by K-Dense. The skills cover biology, chemistry, medicine, physics, engineering, Earth science, data analysis, and scientific communication. Each provides guidance for a specific package, data source, or workflow, including the scientific conventions and validation checks needed to use it.

The collection follows the open Agent Skills standard and works with Cursor, Claude Code, Codex, Google Antigravity, and other compatible hosts. It is also a portable Agent Plugins package (plugin.json + skills/), so clients that support that standard can load the collection as one plugin. Browse the illustrated skill guides, skill categories, or complete catalog to choose the skills relevant to your work.

⭐ Help make AI for science easier to discover: If Scientific Agent Skills saves you time, teaches your agent a workflow, or helps your lab move faster, please star this repository. A star is a public signal that these open, reusable research skills are worth maintaining: it helps scientists, engineers, and open-source contributors find the project, shows which agent-skill standards are gaining real adoption, and gives us a clear reason to keep expanding the collection for the community.


These skills enable your AI agent to seamlessly work with specialized scientific libraries, databases, and tools across multiple scientific domains. While the agent can use any Python package or API on its own, these explicitly defined skills provide curated documentation and examples that make it significantly stronger and more reliable for the workflows below:

  • 🧬 Bioinformatics & Genomics - Sequence analysis, single-cell RNA-seq, pooled CRISPR screens, primer design, amplicon microbiomes, variant annotation, phylogenetics
  • 🧪 Cheminformatics & Drug Discovery - Molecular property prediction, virtual screening, ADMET analysis, molecular docking, lead optimization, calibrated 1D NMR processing
  • 🔬 Proteomics & Mass Spectrometry - LC-MS/MS processing, peptide identification, spectral matching, protein quantification
  • 🏥 Clinical Research & Evidence Workflows - Clinical trials, pharmacogenomics, variant evidence review, pharmacokinetic/pharmacodynamic modelling and dose-regimen evaluation, aggregate decision-support evaluation, source-bound draft report structures, and formatting of clinician-authored treatment decisions
  • 🧠 Healthcare AI & Biosignal Research - EHR and model research, physiological signal analysis, and retrospective validation—not patient-specific diagnosis, treatment, alarms, or deployment decisions
  • 🐭 Preclinical Research & Animal Welfare - Multivariate severity scoring and humane-endpoint forecasting for laboratory animal studies, for 3Rs/refinement analysis and EU Directive 2010/63/EU reporting—an aid to severity assessment, never a decision rule
  • 🖼️ Microscopy, Medical Imaging & Digital Pathology - Quantitative fluorescence microscopy, privacy-aware DICOM processing, research-only whole-slide image analysis, computational pathology, and radiology data workflows
  • 🧠 Neuroscience & Electrophysiology - BIDS datasets, NWB conversion and clock alignment, extracellular recordings, and physiological signals
  • 🤖 Machine Learning & AI - Deep learning, reinforcement learning, time series analysis, model interpretability, Bayesian methods
  • 🔮 Materials Science & Chemistry - Crystal structure analysis, CALPHAD phase equilibria, metabolic modeling, computational chemistry
  • 🌌 Physics & Astronomy - Astronomical data analysis, coordinate transformations, cosmological calculations, symbolic mathematics, physics computations
  • ⚙️ Engineering & Simulation - Battery cycling models, chemical kinetics and ignition delay, fluid dynamics, lab hardware CAD and custom-part fabrication, discrete-event simulation, process optimization
  • 📊 Data Analysis & Visualization - Statistical analysis, network analysis, time series, publication-quality figures, large-scale data processing, EDA
  • 🌊 Chemical Oceanography - Seawater carbonate chemistry, ocean acidification, mineral saturation, and measurement uncertainty
  • 🌍 Geospatial Science & Remote Sensing - Satellite imagery processing, GIS analysis, spatial statistics, terrain analysis, machine learning for Earth observation
  • 🧪 Laboratory Automation - Liquid handling protocols, lab equipment control, workflow automation, LIMS integration
  • 📚 Scientific Communication - Evidence-traceable writing, confidential authorized peer review, literature synthesis, document processing, macro-free PPTX posters, slides, schematics, and citation management
  • 🔬 Multi-omics & Systems Biology - Multi-modal data integration, pathway analysis, carbon-13 metabolic flux inference, biochemical kinetic models
  • 🧬 Protein Engineering & Design - Protein language models, structure prediction, cryo-EM refinement, sequence design, function annotation
  • 🧰 Agent Platforms & Infrastructure - Build on Pi with SDK, RPC, extensions, custom providers/models, packages, TUI components, and session tooling
  • 🎓 Research Methodology - Evidence-bounded candidate hypotheses, scientific brainstorming, critical thinking, grant writing, and qualitative low-stakes evaluation of scholarly works
  • ⚖️ Regulatory & Standards - Draft evidence-preparation artifacts for ISO management-system and laboratory standards, plus analytical method validation, verification, and transfer under ICH/USP/CLSI frameworks—prepared for qualified review, never a certification, accreditation, or method-release decision

Transform your AI coding agent into an 'AI Scientist' on your desktop!

🎬 New to Scientific Agent Skills? Watch our Getting Started with Scientific Agent Skills video for a quick walkthrough.

🎥 More tutorials

Recorded walkthroughs of these skills on real research tasks, from the K-Dense YouTube channel:

VideoWhat it covers
Skills 101: Build Your Own Scientific Agent SkillWriting, testing, and packaging a new skill from scratch
Literature Review and Hypothesis GenerationSearching the literature and generating grounded hypotheses
Draft and Budget an Experimental ProtocolTurning a planned experiment into a costed, written protocol
Draft Responses to Reviewer CommentsBuilding a point-by-point rebuttal from reviewer feedback
Can AI Reproduce a Nature Medicine Paper?An end-to-end reproduction attempt on a published analysis

📦 What's Included

This repository provides 181 scientific and research skills organized into the following categories:

  • 100+ Scientific & Financial Databases - The Database Lookup skill documents 80 databases with public, registered, or licensed access (PubChem, ChEMBL, UniProt, COSMIC, ClinicalTrials.gov, FRED, USPTO, and more), including endpoint selection, pagination, and provenance. Dedicated skills cover DepMap, Imaging Data Commons, PrimeKG, NCATS ARAX, U.S. Treasury Fiscal Data, Hugging Science, OneKGPd, Genomic Intelligence, and AlphaGenome. Multi-database packages such as BioServices, Biopython, and gget add further coverage
  • 70+ Optimized Python Package Skills - Explicitly defined, version-aware workflows for RDKit, Scanpy, PyTorch Lightning, scikit-learn, PyTDC, PathML, pydicom, NeuroKit2, PufferLib, QuTiP, GeoPandas, pymatgen, BioPython, Qiskit, Molecular Dynamics (OpenMM/MDAnalysis), and others. The agent can still use any Python package; these skills provide stronger, safer guidance for the packages listed
  • 9 Scientific Integration Skills - Explicitly defined skills for Benchling, DNAnexus, LatchBio, OMERO, Protocols.io, Open Notebook, Ginkgo Cloud Lab, LabArchives, and Opentrons. Again, the agent is not limited to these — any API or platform reachable from Python is fair game; these skills are the optimized, pre-documented paths
  • 30+ Analysis & Communication Tools - Literature review, evidence-traceable scientific writing, confidential peer review, document processing, Paperclip (full-text papers, FDA/PMDA/EMA filings, and trial registries with line-pinned citations), Paperzilla, Exa Search, macro-free PPTX posters, slides, schematics, infographics, Mermaid diagrams, and more
  • 10+ Research & Clinical Tools - Evidence-bounded hypothesis generation, grant writing, aggregate clinical decision-support research, clinician-authored treatment-plan formatting, PK/PD modelling and simulation (NCA, population PK, exposure-response, bioequivalence, first-in-human dose), BIDS, ISO standards-readiness evidence preparation (ISO 13485, ISO 14971, ISO/IEC 17025, ISO 15189), analytical method validation and transfer (ICH Q2(R2)/Q14, ICH M10, USP, CLSI EP), scenario analysis, and workflow-derived skill drafting with Autoskill

Every skill has a SKILL.md with its purpose, workflow, and version metadata. Depending on the workflow, it also includes code examples, reference documentation, executable helpers, or templates. Skills with bundled scripts/ have a corresponding test suite under tests/<skill-name>/ and a dependency entry in tests/skill-requirements.toml.

What's new in 2.71.0

This update refreshes all 181 skills, including package and API guidance, dependency requirements, reference documentation, and validation workflows. Highlights include:

Each skill records its own compatibility requirements and validation scope. Check those details before reusing an older workflow; local tests, live-service checks, and illustrative examples have different coverage.

Recently added workflows

Research taskSkillsWorkflow focus
Design and analyze assaysPrimer Design, FlowKit, MAGeCK, CellProfilerPCR specificity screening, cytometry compensation and gating, pooled CRISPR contrasts, and microscopy segmentation QC
Process spectra and model metabolismnmrglue, 13C Metabolic Flux, TelluriumCalibrated 1D NMR, isotope-tracing inference with identifiability checks, and reproducible biochemical simulations
Reconstruct structures and preserve neural dataRELION, NWB ConversionCryo-EM half-map validation and two-photon imaging/behavior conversion with clock alignment
Analyze microbiomes and seawater chemistryQIIME 2 Amplicon, Marine Carbonate ChemistryPaired-end 16S processing with read-retention checks, carbonate speciation, and measurement uncertainty
Simulate energy and materials systemsPyBaMM, Cantera, pycalphadBattery cycling, ignition delay, and alloy phase equilibria with numerical checks and input provenance

📋 Table of Contents


🚀 Why Use This?

⚡ Accelerate Your Research

  • Save Days of Work - Skip API documentation research and integration setup
  • Reviewed Starting Points - Tested examples with explicit validation, provenance, and safety boundaries; verify them in the target environment
  • Multi-Step Workflows - Execute complex pipelines with a single prompt

🎯 Comprehensive Coverage

  • 181 Skills - Extensive coverage across all major scientific domains
  • 100+ Databases - 80 databases documented by Database Lookup, plus dedicated data access skills and multi-database packages such as BioServices, Biopython, and gget
  • 70+ Optimized Python Package Skills - Current, version-scoped guidance for packages including RDKit, Scanpy, PyTorch Lightning, scikit-learn, PyTDC, pydicom, PufferLib, QuTiP, GeoPandas, pymatgen, Qiskit, Molecular Dynamics (OpenMM/MDAnalysis), scVelo, and TimesFM (the agent can use any Python package; these are the pre-documented paths)

🔧 Easy Integration

  • Simple Setup - Copy skills to your skills directory and start working
  • Configured Discovery - Compatible hosts can find and use relevant skills from their configured skill paths
  • Well Documented - Each skill includes examples, use cases, and best practices

🌟 Maintained & Supported

  • Regular Updates - Continuously maintained and expanded by K-Dense team
  • Validation in CI - Relevant pull requests run the repository-wide structural checks and standard-library-only skill suites. Scientific dependencies are tested in separate environments; see Testing for the full validation workflow
  • Community Driven - Open source with active community contributions
  • Enterprise Ready - Commercial support available for advanced needs

🎯 Getting Started

Option 1: npx (supported hosts)

Install Scientific Agent Skills with a single command:

npx skills add K-Dense-AI/scientific-agent-skills

This is a common standards-based installer for supported Agent Skills hosts, including current versions of Claude Code, Claude Cowork, Codex, Gemini CLI, Google Antigravity, and Cursor. Confirm installation paths and optional metadata behavior in your host's current documentation.

Option 2: GitHub CLI (gh skill)

If you use the GitHub CLI (v2.90.0+), you can install skills with gh skill:

# Browse and install interactively
gh skill install K-Dense-AI/scientific-agent-skills

# Install a specific skill directly
gh skill install K-Dense-AI/scientific-agent-skills scanpy

# Target a specific agent host
gh skill install K-Dense-AI/scientific-agent-skills --agent cursor
gh skill install K-Dense-AI/scientific-agent-skills --agent claude-code
gh skill install K-Dense-AI/scientific-agent-skills --agent codex
gh skill install K-Dense-AI/scientific-agent-skills --agent gemini

gh skill automatically installs to the correct directory for your agent host and records provenance metadata for supply chain integrity.

Version pinning

Pin to a specific release tag or commit SHA for reproducible installs:

# Pin to a release tag
gh skill install K-Dense-AI/scientific-agent-skills --pin v2.66.0

# Pin to a commit SHA
gh skill install K-Dense-AI/scientific-agent-skills --pin abc123def

Keeping skills up to date

# Check for updates interactively
gh skill update

# Update all installed skills
gh skill update --all

Option 3: Agent Plugins (Cursor, Codex, and other plugin clients)

This repository is a valid Agent Plugins 1.0.0 package: root plugin.json plus Agent Skills under skills/. Clients that support the standard discover every immediate child of skills/ that contains a SKILL.md.

Cursor — symlink or copy the repo into the local plugins directory, then reload:

mkdir -p ~/.cursor/plugins/local
ln -s "$(pwd)" ~/.cursor/plugins/local/scientific-agent-skills

Restart Cursor or run Developer: Reload Window, then confirm the plugin and its skills appear under Customize. See Cursor plugins.

Codex — install from a local checkout (confirm the current CLI flag names in Codex docs):

codex plugins install .

Compatible clients (Cursor, Codex, GitHub Copilot, VS Code, Kiro, and others listed at agent-plugins.org) share the same package layout; installation UX stays client-specific.

Other Agent Skills hosts (OpenClaw, NemoClaw, Pi, Hermes, …)

Agent hosts differ in install paths, discovery settings, and support for optional frontmatter fields. npx skills add (Option 1) commonly installs into the ~/.agents/skills/ convention, with project-scoped installs under .agents/skills/; confirm both paths against your host's current documentation. To install manually on a host configured to scan one of those locations:

git clone https://github.com/K-Dense-AI/scientific-agent-skills.git ~/.agents/skills/scientific-agent-skills   # user-level
git clone https://github.com/K-Dense-AI/scientific-agent-skills.git .agents/skills/scientific-agent-skills      # project-level

For Hermes versions that support skill taps, add the repository as a tap:

hermes skills tap add K-Dense-AI/scientific-agent-skills

Every SKILL.md uses YAML frontmatter with a quoted metadata.version. Repository contributions must use block-style YAML; JSON-style flow mappings fail the reference validator. Optional host-specific configuration belongs under metadata, with host manifest blocks kept as nested mappings. See AGENTS.md for the complete rules. Hosts may interpret optional metadata and credential prompts differently, so verify behavior on the target host. Installing a topical subset keeps the available skill catalog focused on your work.

NemoClaw note: NemoClaw runs agents inside NVIDIA OpenShell with default-deny outbound networking. Skills are discovered and loaded normally, but any skill that needs the network — package installs via uv, or API calls (Exa, Parallel, Benchling, NCBI, Materials Project, …) — only works once the operator pre-approves the relevant domains in the OpenShell TUI.

That's it! A compatible host can discover the skills from its configured paths and use them when relevant. You can also invoke any skill manually by mentioning the skill name in your prompt.


⚠️ Security Disclaimer

Skills can execute code and influence your coding agent's behavior. Review what you install.

Agent Skills are powerful — they can instruct your AI agent to run arbitrary code, install packages, make network requests, and modify files on your system. A malicious or poorly written skill has the potential to steer your coding agent into harmful behavior.

We take security seriously. All contributions go through a review process, and we run LLM-based security scans (via Cisco AI Defense Skill Scanner) on every skill in this repository. However, as a small team with a growing number of community contributions, we cannot guarantee that every skill has been exhaustively reviewed for all possible risks.

It is ultimately your responsibility to review the skills you install and decide which ones to trust.

We recommend the following:

  • Do not install everything at once. Only install the skills you actually need for your work. While installing the full collection was reasonable when K-Dense created and maintained every skill, the repository now includes many community contributions that we may not have reviewed as thoroughly.
  • Read the SKILL.md before installing. Each skill's documentation describes what it does, what packages it uses, and what external services it connects to. If something looks suspicious, don't install it.
  • Check the contribution history. Skills authored by K-Dense (K-Dense-AI) have been through our internal review process. Community-contributed skills have been reviewed to the best of our ability, but with limited resources.
  • Run the security scanner yourself. Before installing third-party skills, scan them locally:
    uv pip install cisco-ai-skill-scanner
    skill-scanner scan /path/to/skill --use-behavioral
    
  • Report anything suspicious. If you find a skill that looks malicious or behaves unexpectedly, please open an issue immediately so we can investigate.

Skills are scanned weekly — incrementally, so unchanged skills carry their previous findings forward, with a full rescan of everything at least every 30 days and whenever the scanner or model changes — and the results are published to docs/security-report.md. See SECURITY.md for our security policy, what is in scope, how to report a vulnerability privately, and how to contest a scan finding. We try to address security gaps as they arise.


❤️ Support the Open Source Community

Scientific Agent Skills is powered by 50+ incredible open source projects maintained by dedicated developers and research communities worldwide. Projects like Biopython, Scanpy, RDKit, scikit-learn, PyTorch Lightning, and many others form the foundation of these skills.

If you find value in this repository, please consider supporting the projects that make it possible:

  • ⭐ Star their repositories on GitHub
  • 💰 Sponsor maintainers via GitHub Sponsors or NumFOCUS
  • 📝 Cite projects in your publications
  • 💻 Contribute code, docs, or bug reports

👉 View the full list of projects to support


🙏 Skill Credits

The docx, pdf, pptx, and xlsx document skills are created and maintained by Anthropic and vendored here from anthropics/skills. They are used under Anthropic's terms — see each skill's LICENSE.txt — and we track upstream so you get their latest improvements. All credit for those four skills goes to Anthropic.


⚙️ Prerequisites

  • Python: 3.13+ for repository tooling; individual skill dependencies may support broader Python ranges
  • uv: Python package manager (required for installing skill dependencies)
  • Client: Any agent that supports the Agent Skills standard (Cursor, Claude Code, Gemini CLI, Codex, Google Antigravity, etc.)
  • System: macOS, Linux, or Windows with WSL2
  • Dependencies: Follow each skill's compatibility field and setup instructions for packages, system tools, credentials, and network access. Installing the skill files does not install their dependencies

Use a separate environment for each scientific workflow: some skills require different Python versions or incompatible package pins. uv sync installs the repository's development and validation tools, not every scientific package in the collection.

Installing uv

The skills use uv as the package manager for installing Python dependencies. Install it using the instructions for your operating system:

macOS and Linux:

curl -LsSf https://astral.sh/uv/install.sh | sh

Windows:

powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

Alternative (via pip):

pip install uv

After installation, verify it works by running:

uv --version

For more installation options and details, visit the official uv documentation.


💡 Quick Examples

Once you've installed the skills, you can ask your AI agent to execute complex multi-step scientific workflows. Here are some example prompts:

🧪 Drug Discovery Pipeline

Goal: Prioritize EGFR inhibitor candidates for preclinical lung-cancer research

Prompt:

Use available skills you have access to whenever possible. Query ChEMBL for EGFR inhibitors (IC50 < 50nM), analyze structure-activity relationships 
with RDKit, generate improved analogs with datamol, perform virtual screening with DiffDock 
against AlphaFold EGFR structure, search PubMed for resistance mechanisms, check COSMIC for 
mutations, and create visualizations and a comprehensive report.

Skills Used: database-lookup, rdkit, datamol, diffdock, paper-lookup, scientific-visualization


🔬 Single-Cell RNA-seq Analysis

Goal: Comprehensive analysis of 10X Genomics data with public data integration

Prompt:

Use available skills you have access to whenever possible. Load 10X dataset with Scanpy, perform QC and doublet removal, integrate with Cellxgene 
Census data, identify cell types using NCBI Gene markers, run differential expression with 
PyDESeq2, infer gene regulatory networks with Arboreto, enrich pathways via Reactome/KEGG, 
and identify therapeutic targets with Open Targets.

Skills Used: scanpy, cellxgene-census, database-lookup, pydeseq2, arboreto


🧬 Multi-Omics Biomarker Discovery

Goal: Integrate RNA-seq, proteomics, and metabolomics to predict patient outcomes

Prompt:

Use available skills you have access to whenever possible. Analyze RNA-seq with PyDESeq2, process mass spec with pyOpenMS, integrate metabolites from 
HMDB/Metabolomics Workbench, map proteins to pathways (UniProt/KEGG), find interactions via 
STRING, correlate omics layers with statsmodels, build predictive model with scikit-learn, 
and search ClinicalTrials.gov for relevant trials.

Skills Used: pydeseq2, pyopenms, database-lookup, statsmodels, scikit-learn


🎯 Virtual Screening Campaign

Goal: Discover allosteric modulators for protein-protein interactions

Prompt:

Use available skills you have access to whenever possible. Retrieve AlphaFold structures, identify interaction interface with BioPython, search ZINC 
for allosteric candidates (MW 300-500, logP 2-4), filter with RDKit, dock with DiffDock, 
rank with DeepChem, check PubChem suppliers, search USPTO patents, and optimize leads with 
MedChem/molfeat.

Skills Used: database-lookup, biopython, rdkit, diffdock, deepchem, medchem, molfeat


🏥 Research Variant Evidence Review

Goal: Annotate a synthetic or properly de-identified VCF for hereditary-cancer research and qualified review

Prompt:

Use available skills you have access to whenever possible. Work only with authorized synthetic
or de-identified data. Parse the VCF with pysam, annotate variants with Ensembl VEP, retrieve
ClinVar/COSMIC/NCBI Gene/UniProt evidence, and verify literature sources. Build an evidence-
traceable research summary with scientific-writing. If clinical-reports is used, create only a
visibly marked draft structure from a verified source-fact manifest for qualified review; do not
diagnose, assess individual risk, recommend treatment, or determine trial eligibility.

Skills Used: pysam, database-lookup, paper-lookup, scientific-writing, clinical-reports


🌐 Systems Biology Network Analysis

Goal: Analyze gene regulatory networks from RNA-seq data

Prompt:

Use available skills you have access to whenever possible. Query NCBI Gene for annotations, retrieve sequences from UniProt, identify interactions via 
STRING, map to Reactome/KEGG pathways, analyze topology with Torch Geometric, reconstruct 
GRNs with Arboreto, assess druggability with Open Targets, model with PyMC, visualize 
networks, and search GEO for similar patterns.

Skills Used: database-lookup, torch-geometric, arboreto, pymc, networkx, scientific-visualization

📖 Want more examples? Check out docs/examples.md for comprehensive workflow examples and detailed use cases across all scientific domains.


🔬 Use Cases

🧪 Drug Discovery & Medicinal Chemistry

  • Virtual Screening: Screen millions of compounds from PubChem/ZINC against protein targets
  • Lead Optimization: Analyze structure-activity relationships with RDKit, generate analogs with datamol
  • ADMET Prediction: Predict absorption, distribution, metabolism, excretion, and toxicity with DeepChem
  • Molecular Docking: Predict binding poses with DiffDock and rescore poses with affinity-oriented tools
  • Bioactivity Mining: Query ChEMBL for known inhibitors and analyze SAR patterns

🧬 Bioinformatics & Genomics

  • Sequence Analysis: Process DNA/RNA/protein sequences with BioPython and pysam
  • Single-Cell Analysis: Analyze 10X Genomics data with Scanpy, identify cell types, infer GRNs with Arboreto
  • Variant Annotation: Annotate research VCF files with Ensembl VEP and retrieve ClinVar evidence for qualified interpretation
  • Variant Database Management: Build scalable VCF databases with TileDB-VCF for incremental sample addition, efficient population-scale queries, and compressed storage of genomic variant data
  • Population Genomics: Query variants, cohort sample IDs, and relatedness in the 3,202-person GRCh38 1000 Genomes cohort with OneKGPd
  • Regulatory Sequence Models: Run hosted Genomic Intelligence promoter, splice, enhancer, chromatin, expression, and gene-annotation predictions for research—not clinical or diagnostic decisions
  • Variant Effect Prediction: Look up AlphaGenome Atlas AVI scores, Phred ranks, feature attributions, and per-tissue track effects for any GRCh38 SNV, or score variants and run in silico mutagenesis with the AlphaGenome model — for research prioritisation, not diagnosis
  • Pathogen Surveillance: Track which viral lineages are circulating now and how fast they are growing (SARS-CoV-2, influenza including H5N1, RSV, mpox, measles, dengue) through the GenSpectrum LAPIS API, with reporting lag measured rather than assumed
  • Gene Discovery: Query NCBI Gene, UniProt, and Ensembl for comprehensive gene information
  • Network Analysis: Identify protein-protein interactions via STRING, map to pathways (KEGG, Reactome)

🏥 Clinical Research & Evidence Workflows

  • Clinical Trials: Analyze aggregate trial landscapes and protocol criteria without deciding individual eligibility
  • Variant Evidence Review: Annotate authorized research data with ClinVar, COSMIC, and ClinPGx; qualified professionals retain interpretation responsibility
  • Drug Safety Research: Query FDA databases for aggregate adverse-event, interaction, and recall evidence
  • Clinical Pharmacology: Derive exposure metrics from concentration-time data, fit compartmental and population PK models, relate exposure to effect, and evaluate dosing regimens, bioequivalence, and first-in-human dose
  • Full-Text Evidence Retrieval: Search and read papers, regulatory filings, and trial records end to end with Paperclip, returning citations pinned to line numbers rather than to abstracts
  • Decision-Support Evaluation: Prepare synthetic or aggregate evaluation, evidence-profile, privacy, and governance artifacts—not live clinical decisions
  • Clinician-Authored Documentation: Structure verified source-bound report drafts and format treatment decisions already made by authorized licensed professionals

🔬 Multi-Omics & Systems Biology

  • Multi-Omics Integration: Combine RNA-seq, proteomics, and metabolomics data
  • Pathway Analysis: Enrich differentially expressed genes in KEGG/Reactome pathways
  • Network Biology: Reconstruct gene regulatory networks, identify hub genes
  • Biomarker Discovery: Integrate multi-omics layers to predict patient outcomes

📊 Data Analysis & Visualization

  • Statistical Analysis: Perform hypothesis testing, power analysis, and experimental design
  • Publication Figures: Create publication-quality visualizations with matplotlib and seaborn
  • Network Visualization: Visualize biological networks with NetworkX
  • Report Generation: Produce evidence-traceable research reports with Scientific Writing and document tools; Clinical Reports outputs remain visibly marked drafts built only from verified synthetic, de-identified, or aggregate source facts

🧪 Laboratory Automation

  • Protocol Design: Author and simulate Opentrons or PyLabRobot protocols before trained-operator review
  • LIMS/ELN Integration: Prepare scoped Benchling and LabArchives operations with explicit authorization for remote writes
  • Workflow Automation: Validate and simulate multi-step laboratory workflows offline; physical execution stays behind equipment-specific operator safety gates
  • Custom Lab Hardware Fabrication: Pre-flight STEP files, then quote CNC-machined or 3D-printed parts on Fictiv with DFM feedback and lead-time options relayed back; nothing is ordered without explicit approval of the exact total

📚 Available Skills

This repository contains 181 scientific and research skills organized across multiple domains. Each skill provides comprehensive documentation, code examples, and best practices for working with scientific libraries, databases, and tools.

Skill Categories

Note: The Python package and integration skills listed below are explicitly defined skills — curated with documentation, examples, and best practices for stronger, more reliable performance. They are not a ceiling: the agent can install and use any Python package or call any API, even without a dedicated skill. The skills listed simply make common workflows faster and more dependable.

Categories overlap: a skill may appear in more than one domain, so the category counts do not sum to the 181 unique skills.

Package versions below identify the baselines documented by the skills. Follow each linked SKILL.md for its complete dependency stack, runtime requirements, and validation scope.

🧬 Bioinformatics & Genomics (32 skills)

  • RNA-seq preparation: Bulk RNA-seq (FASTQ and quantification workflows, sample/reference QC, and validated gene-level counts for a PyDESeq2 handoff)
  • Sequence analysis: BioPython, pysam, scikit-bio, BioServices
  • PCR assay design: Primer Design (Primer3 candidates, thermodynamics, variant masking, tails, and bounded local or BLAST-assisted specificity screening)
  • Single-cell analysis: Scanpy, AnnData, scvi-tools, scVelo (RNA velocity), Arboreto, Cellxgene Census
  • Genomic tools: gget, current geniml/Gtars interval workflows, deepTools, Polars-Bio, Zarr, TileDB-VCF
  • Flow cytometry: FlowIO (FCS input/output), FlowKit (compensation, transforms, hierarchical gating, and FlowJo workspace analysis)
  • Coordinate hygiene: Genomic Coordinates (convert intervals across BED/GFF/GTF/VCF/SAM/WIG conventions, normalise variant representations, and catch 0-based vs 1-based and assembly/contig-naming mismatches before they corrupt an analysis)
  • Population and sequence intelligence: OneKGPd (individual-level 1000 Genomes cohort queries) and Genomic Intelligence (hosted regulatory/gene-expression predictions; research only)
  • Variant effect prediction: AlphaGenome (DeepMind's AlphaGenome Atlas of precomputed effects for every GRCh38 SNV — AVI score with Phred rank and 18 SHAP feature attributions, plus raw and quantile scores across 9,440 RNA-seq, DNase, ATAC, ChIP, CAGE, splicing and contact-map tracks — and on-demand model scoring, in silico mutagenesis and REF-vs-ALT track prediction for human and mouse; research only, not clinical)
  • Differential expression: PyDESeq2
  • Pooled CRISPR screens: MAGeCK (guide counting, replicate QC, explicit contrasts, and gene hit ranking)
  • Amplicon microbiomes: QIIME 2 Amplicon (paired-end 16S import, primer trimming, denoising, taxonomy, and read-retention checks)
  • Functional enrichment: Pathway Enrichment (ORA, GSEA/preranked, ssGSEA via gseapy + g:Profiler; GO, KEGG, Reactome, WikiPathways, MSigDB)
  • Phylogenetics: ETE Toolkit, Phylogenetics (MAFFT, IQ-TREE 2, FastTree)
  • Microbiome foundation models: Waypoint (Outpost Bio's open Waypoint-6m/45m/170m checkpoints, the Atlas 539k-sample MGnify pretraining corpus, and the eight-task Compass benchmark — embedding, fine-tuning, benchmarking, and pretraining on taxonomic abundance profiles, with MetaPhlAn/Kraken2/QIIME 2 conversion)

🧪 Cheminformatics & Drug Discovery (11 skills)

  • Molecular manipulation: RDKit, Datamol, Molfeat
  • NMR processing: nmrglue (calibrated 1D FIDs, phase/baseline correction, peak candidates, and signed integration)
  • Deep learning: DeepChem, TorchDrug
  • Docking & screening: DiffDock
  • Molecular dynamics: OpenMM + MDAnalysis (MD simulation & trajectory analysis)
  • Cloud quantum chemistry: Rowan (pKa, docking, cofolding)
  • Drug-likeness: MedChem
  • Benchmarks: PyTDC 1.1.15 on its verified CPython 3.11 compatibility stack

🔬 Proteomics & Mass Spectrometry (2 skills)

  • Spectral processing: matchms, pyOpenMS

🏥 Clinical Research & Evidence Workflows (8 skills)

  • Clinical databases: via Database Lookup (ClinicalTrials.gov, ClinVar, ClinPGx, COSMIC, FDA, cBioPortal, Monarch, and more)
  • Clinical pharmacology: PK/PD Modeling (non-compartmental analysis, compartmental and population PK, exposure-response and Emax, TMDD, PBPK orientation, bioequivalence including RSABE/ABEL, allometric scaling and first-in-human dose, DDI prediction under ICH M12, concentration-QTc, and Bayesian therapeutic drug monitoring — stdlib + numpy/scipy, no proprietary estimation software invoked)
  • Cancer genomics: DepMap (cancer dependency scores, drug sensitivity)
  • Cancer imaging: Imaging Data Commons (NCI radiology & pathology datasets via idc-index)
  • Healthcare AI research: PyHealth
  • Decision-support research: local, aggregate or synthetic Clinical Decision Support evaluation and governance artifacts only
  • Clinical documentation: source-bound Clinical Reports drafts and formatting of verified clinician-authored decisions with Treatment Plans; neither skill diagnoses or recommends care

🐭 Preclinical Research & Animal Welfare (1 skill)

  • Severity assessment: RELSA Severity Assessment (multivariate RELSA scores from body weight, temperature, clinical/nesting scores, biomarkers, activity, heart rate, burrowing and wheel running; ARIMA humane-endpoint forecasting with 95% prediction intervals; KDE-derived attention and danger zones for 3Rs/refinement and EU Directive 2010/63/EU severity reporting) — an aid to severity assessment, never a decision rule

🖼️ Microscopy, Medical Imaging & Digital Pathology (5 skills)

  • DICOM processing: pydicom 3.0.2 with privacy-first local preflight and no diagnostic or de-identification-compliance claims
  • Whole slide imaging: histolab and research-only PathML 3.0.8
  • Quantitative microscopy: CellProfiler (reusable nuclei-measurement pipeline, channel manifests, segmentation overlays, and measurement checks)
  • Virtual spatial transcriptomics: noncommercial DeepSpot-M for transcriptome-wide spatial gene expression from 224x224 H&E tiles

🧠 Neuroscience & Electrophysiology (4 skills)

  • Data standards: BIDS (Brain Imaging Data Structure for neuroscience and biomedical datasets; pairs with DataLad for retrieval — OpenNeuro at github.com/OpenNeuroDatasets and DANDI at github.com/dandisets publish their holdings as DataLad datasets — and for BIDS-App runs under recorded provenance, with BEP028 the extension proposal for provenance records in derivatives)
  • Neural recordings: Neuropixels-Analysis (extracellular spikes, silicon probes, spike sorting)
  • Data conversion: NWB Conversion (two-photon TIFF and behavioral positions, clock alignment, round-trip preservation, and NWB validation)
  • Physiological signals: NeuroKit2 0.2.13 for reproducible research workflows—not diagnosis, monitoring decisions, or medical-device validation

🤖 Machine Learning & AI (14 core skills)

  • Deep learning: PyTorch Lightning, Transformers, Stable Baselines3, and PufferLib workflows for native 5.0, published 3.0.0, and pinned historical 4.0
  • Classical ML: scikit-learn, scikit-survival 0.28, and SHAP
  • Time series: aeon, TimesFM (Google's zero-shot foundation model for univariate forecasting)
  • Bayesian methods: PyMC
  • Optimization: PyMOO
  • Graph ML: Torch Geometric
  • Dimensionality reduction: UMAP-learn
  • Statistical modeling: statsmodels

🔮 Materials Science, Chemistry & Physics (8 skills)

  • Materials: current split pymatgen wrapper/core plus explicitly bounded Materials Project queries
  • Alloy thermodynamics: pycalphad (TDB-driven phase equilibria, phase fractions, composition checks, and database provenance)
  • Metabolic modeling: COBRApy
  • Astronomy: Astropy
  • Quantum computing: Cirq, PennyLane, Qiskit, QuTiP 5.3.1

⚙️ Engineering & Simulation (9 skills)

  • Lab hardware CAD: parametric build123d 0.13.0 workflows for microfluidic chips and molds, optomechanical mounts, microplate and cuvette adapters, and behavior rigs, with interface checks and multi-view renders
  • Custom-part fabrication: Fictiv (browser-driven CNC, 3D printing, sheet metal, urethane casting and molding quotes, DFM review, lead-time and region tiers, and checkout, with CAD pre-flight checks and an explicit approval gate before any order, quote request or share)
  • Numerical computing: MATLAB and GNU Octave planning/review workflows, with R2026b documentation, an explicit R2026a Python-integration baseline, and GNU Octave 11.3.0
  • Computational fluid dynamics: bounded FluidSim 0.9 simulations with numerical-validity and HPC checks
  • Experimental flow measurement: OpenPIV (velocity fields from PIV image pairs, interrogation-window cross-correlation, spurious-vector validation, vorticity/strain-rate/turbulence statistics)
  • Discrete-event simulation: SimPy 4.1.2 with replication, warm-up, and output-analysis guidance
  • Chemical kinetics: Cantera (homogeneous ignition delay, mechanism snapshots, numerical refinement, and conservation checks)
  • Battery experiments: PyBaMM (charge/discharge protocols, parameter provenance, numerical checks, and measured-curve comparisons)
  • Symbolic math: SymPy

🌊 Chemical Oceanography (1 skill)

  • Marine Carbonate Chemistry: paired seawater measurements with PyCO2SYS, carbonate speciation, pH-scale handling, lab-to-in-situ corrections, aragonite/calcite saturation, and uncertainty propagation

📊 Data Analysis & Visualization (22 skills)

  • Visualization: Matplotlib, Seaborn, Scientific Visualization
  • Geospatial analysis: GeoPandas 1.2.0 and GeoMaster (remote sensing, GIS, satellite imagery, and spatial ML)
  • Data processing: Dask, Polars, Vaex
  • Network analysis: NetworkX
  • Document processing: LiteParse (local PDF/document parsing with bounding boxes and OCR), MarkItDown, PDF, DOCX, PPTX, and XLSX
  • Infographics: Infographics (AI-powered professional infographic creation)
  • Diagrams: Markdown & Mermaid Writing (text-based diagrams as default documentation standard)
  • Exploratory data analysis: bounded local EDA for explicitly supported formats, with unknown formats failing closed
  • Statistical analysis: Statistical Analysis workflows
  • Units and measurement uncertainty: Uncertainty & Units (pint dimensional checking, GUM uncertainty budgets, Type A/B evaluation, coverage factors and expanded uncertainty, Monte Carlo propagation, CODATA constants)
  • Experimental design: Experimental Design (randomization, blocking, factorial/fractional-factorial DOE, crossover, cluster, sequential designs; pyDOE3)
  • Statistical power: Statistical Power (sample-size & power for t-tests, ANOVA, proportions, correlation, regression — closed-form plus simulation-based for GLMs, mixed models, and cluster designs)

🧪 Laboratory Automation (6 skills)

  • Liquid handling: offline-first PyLabRobot planning/simulation and Opentrons authoring, with physical execution behind explicit operator safety gates
  • Cloud lab: Ginkgo Cloud Lab (protein expression & purification across cell-free/E. coli/Pichia, IVT RNA synthesis, thermal shift and Echo-MS assays, SPR onboarding, fluorescent pixel art via autonomous RAC infrastructure)
  • Protocol management: bounded protocols.io reads across documented v3/v4 endpoints and non-executing write plans
  • LIMS/ELN integration: Benchling and the separate LabArchives legacy ELN and Inventory v1 APIs

🔬 Multi-omics & Systems Biology (5 skills)

  • Pathway analysis: via Database Lookup (KEGG, Reactome, STRING) and PrimeKG
  • Data management: LaminDB
  • Biochemical dynamics: Tellurium (kinetic models, time-course simulations, perturbations, and reproducible model exchange)
  • Carbon-13 metabolic flux inference: 13C Metabolic Flux (validated atom maps, steady-state isotope simulation, constrained fitting, and flux-identifiability profiles)

🧬 Protein Engineering & Design (5 skills)

  • Protein language models: ESM
  • Experimental structure reconstruction: RELION (extracted-particle refinement, optics/STAR validation, and half-map/FSC checks)
  • Glycoengineering: Glycoengineering (N/O-glycosylation prediction, therapeutic antibody optimization)
  • Cloud laboratory platform: Adaptyv (automated protein testing and validation)
  • Cloud structure & design platform: Tamarind (managed-GPU access to AlphaFold, Boltz, Chai, ESMFold, RFdiffusion, ProteinMPNN, BoltzGen, antibody/nanobody design, DiffDock/Vina docking, binding affinity, and MSA generation via REST API or MCP)

📚 Scientific Communication (27 skills)

  • Literature: Paper Lookup (PubMed, PMC, bioRxiv, medRxiv, arXiv, OpenAlex, Crossref, Semantic Scholar, CORE, Unpaywall), Literature Review, Paperzilla
  • Full-text corpus access: Paperclip (read-only virtual filesystem over ~11M full-text papers, 217K+ FDA/PMDA/EMA regulatory documents, clinical trial registries, and UniProt/PDB/ChEMBL entries — source-scoped semantic search, corpus-wide grep, SQL metadata queries, map/reduce reading across many papers, figure vision analysis, and line-pinned citations)
  • Advanced paper search: BGPT Paper Search (25+ structured fields per paper — methods, results, sample sizes, quality scores — from full text, not just abstracts)
  • Web intelligence: Parallel Web (web search, URL/PDF extraction, deep research, structured enrichment, entity discovery, and recurring monitoring), Exa Search, and Research Lookup
  • Research notebooks: Open Notebook (self-hosted NotebookLM alternative — PDFs, videos, audio, web pages; 16+ AI providers; multi-speaker podcast generation)
  • Writing: evidence-traceable Scientific Writing and local, confidential, authorized Peer Review
  • Document processing: LiteParse, PDF, DOCX, PPTX, XLSX, and MarkItDown
  • Publishing and paper workflows: Venue Templates
  • Presentations: Scientific Slides, LaTeX Posters, and macro-free PPTX Posters generated from author-approved local manifests
  • Diagrams: Scientific Schematics, Markdown & Mermaid Writing
  • Infographics: Infographics (10 types, 8 styles, colorblind-safe palettes)
  • Citations: Citation Management, pyzotero
  • Illustration: Generate Image (generation, editing, and compositing through the OpenRouter Image API, with model discovery, capability checks, and dry runs)

🔬 Scientific Databases & Data Access (13 skills → 100+ databases total)

Database Lookup documents 80 databases with public, registered, or licensed access across scientific and financial domains, with retrieval contracts, pagination/count reconciliation, and endpoint provenance. Dedicated skills cover specialized data platforms. Multi-database packages such as BioServices, Biopython, and gget add further coverage.

  • Unified access: Database Lookup (80 databases spanning chemistry, genomics, clinical, pathways, patents, economics, and more — PubChem, ChEMBL, UniProt, PDB, AlphaFold, KEGG, Reactome, STRING, ClinVar, COSMIC, ClinicalTrials.gov, FDA, FRED, USPTO, SEC EDGAR, and dozens more — with auditable filters and provenance)
  • Cancer genomics: DepMap (cancer cell line dependencies, drug sensitivity, gene effect profiles)
  • Public germline variant evidence: Folklore Variant Evidence (ClinGen gene-disease validity assertions and source-linked evidence for one supported GRCh38 variant, with explicit ambiguity handling and related literature for qualified professional review)
  • Cancer imaging: Imaging Data Commons (NCI radiology & pathology datasets via idc-index)
  • Knowledge graph: PrimeKG (precision medicine knowledge graph — genes, drugs, diseases, phenotypes)
  • Biomedical knowledge graph search: NCATS ARAX (bounded, Biolink-constrained one-hop and endpoint-pinned two-hop queries over knowledge graphs with up to five explicitly selected NCATS Translator providers, with provenance preservation)
  • Fiscal data: U.S. Treasury Fiscal Data (national debt, Treasury statements, auctions, exchange rates)
  • Scientific ML resource catalog: Hugging Science (curated index of datasets, models, blog posts, and interactive Spaces across 17 scientific domains — astronomy, biology, chemistry, climate, genomics, materials science, medicine, physics, scientific reasoning, and more — with usage patterns for datasets, transformers, and gradio_client)
  • Individual-level population genomics: OneKGPd (3,202-person high-coverage 1000 Genomes cohort queries)
  • Hosted regulatory genomics: Genomic Intelligence (promoter, splice, enhancer, chromatin, expression, and gene-annotation predictions for research use)
  • Precomputed variant effects: AlphaGenome (Atlas lookups for every GRCh38 SNV with AVI, Phred, SHAP feature attributions, and per-track scores; on-demand AlphaGenome model scoring for indels, mouse, and custom scorers; Atlas website deep links)
  • Ontology identifiers: Ontology Term Resolution (resolve free-text tissue, cell-type, disease, phenotype, assay, chemical, organism, and developmental-stage labels to term IDs and validate CURIEs against EBI OLS4, for GEO/ENA/BioSamples/CELLxGENE/HCA/ISA-Tab metadata)
  • Live pathogen surveillance: Pathogen Variant Surveillance (which viral lineages are circulating now, how fast they are growing, and what mutations they carry — SARS-CoV-2, influenza including H5N1, RSV, mpox, measles, dengue and more through the GenSpectrum LAPIS API, with lineage names resolved against the live pango-designation nomenclature and reporting lag measured rather than assumed)

🔧 Infrastructure & Platforms (12 skills)

  • Cloud compute: Modal
  • Data distribution and provenance: DataLad (clone and fetch OpenNeuro, DANDI and registry.datalad.org datasets over git-annex, capture re-executable provenance with datalad run/rerun and containers-run, publish to siblings)
  • GPU acceleration: Optimize for GPU (CuPy, Numba CUDA, Warp, cuDF, cuML, cuGraph, KvikIO, cuCIM, cuxfilter, cuVS, cuSpatial, RAFT)
  • Genomics platforms: DNAnexus, LatchBio
  • Workflow engines: Nextflow (build/run/debug Nextflow & nf-core pipelines — DSL2 modules, executors/containers, HPC/cloud scaling) and pacsomatic (operator toolkit for the nf-core/pacsomatic tumor-normal somatic variant-calling workflow)
  • Microscopy: OMERO
  • Automation: Opentrons
  • Resource detection: Get Available Resources on request or before a clearly resource-sensitive local workload; redacted and without stress tests
  • Workflow mining: Autoskill (local screenpipe-based repeated workflow detection and skill drafting)
  • Agent platform development: Pi Agent (Pi 0.99.2 terminal harness, SDK, RPC/JSONL, native MCP, extensions, providers/models, packages, TUI components, and session tooling)

🎓 Research Methodology & Planning (13 skills)

  • Ideation: evidence-aware Scientific Brainstorming and non-scoring Hypothesis Generation that keeps hypotheses labeled as candidates
  • Text-dataset hypothesis software: HypoGeniC/HypoRefine produces candidate textual patterns and task-prediction statistics, not validated scientific hypotheses
  • Autonomous optimization: Arbor (Hypothesis Tree Refinement — iteratively improve a code/model/agent-harness/data artifact against a dev evaluator while a held-out test gate guards against overfitting)
  • Critical analysis: Scientific Critical Thinking and qualitative, low-stakes Scholar Evaluation of works—never ranking people or supporting consequential decisions
  • Scenario analysis: What-If Oracle (4–6 branch possibility exploration, contingency planning, decision stress-testing)
  • Multi-perspective deliberation: Consciousness Council (diverse expert viewpoints, devil's advocate analysis)
  • Cognitive profiling: DHDNA Profiler (extract thinking patterns and cognitive signatures from any text)
  • Funding: Research Grants
  • Discovery: Research Lookup, Paper Lookup (10 academic databases)
  • Market analysis: evidence-traceable Market Research Reports with assumption-led sizing and forecast sensitivity

⚖️ Regulatory & Standards (2 skills)

  • Standards readiness: draft evidence-preparation artifacts for ISO 13485 (medical device QMS), ISO 14971 (device risk management), ISO/IEC 17025 (testing and calibration laboratories), and ISO 15189 (medical laboratories), with per-standard process domains selected by a --standard profile
  • Analytical method validation: plan, evaluate, and document validation, verification, and transfer of analytical procedures (HPLC, LC-MS/MS, GC, CE, ICP-MS, dissolution, qNMR, qPCR, NIR, ligand-binding and cell-based assays) under whichever framework governs — ICH Q2(R2)/Q14 and ICH M10 encoded from their openly licensed text, with USP <1220>/<1225>/<1226>, the CLSI EP series, and ISO/IEC 17025 cited by designation and scope only; stdlib-only statistics, no network access
  • Assurance-lane separation: keeps ISO certification, laboratory accreditation, FDA QMSR inspection, CLIA certification, MDSAP, and EU MDR/IVDR evidence boundaries distinct—laboratories are accredited rather than certified, and ISO 15189 accreditation does not satisfy CLIA
  • Never a compliance, audit, assessment, certification, accreditation, or method-release decision; qualified RA/QA, legal, laboratory-director, assessor, and certification-body review is required

📖 For complete details on all skills, see docs/skills.md

💡 Looking for practical examples? Check out docs/examples.md for comprehensive workflow examples across all scientific domains.


📝 From the Blog

Deep dives, benchmarks, and guides from the K-Dense blog that are directly relevant to using the skills in this repository.

Start here

Skill benchmarks and deep dives

Why the workflow layer matters

Security and safe deployment

Complementary open-source projects


🤝 Contributing

We welcome contributions to expand and improve this scientific skills repository!

Read AGENTS.md before creating or changing a skill, and see CONTRIBUTING.md for the detailed workflow and pull request process. Contributions should focus on a scientific package, database, platform, or research workflow; broad coding, infrastructure, and routing skills are outside the repository's scope.

Ways to Contribute

✨ Add New Skills

  • Create skills for additional scientific packages or databases
  • Add integrations for scientific platforms and tools

📚 Improve Existing Skills

  • Enhance documentation with more examples and use cases
  • Add new workflows and reference materials
  • Improve code examples and scripts
  • Fix bugs or update outdated information

🐛 Report Issues

  • Submit bug reports with detailed reproduction steps
  • Suggest improvements or new features

How to Contribute

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-skill)
  3. Follow CONTRIBUTING.md and the existing directory structure
  4. Ensure all new skills include valid SKILL.md files with required frontmatter and metadata.version
  5. Test your examples and workflows thoroughly, and add a suite under tests/<skill-name>/ if your skill ships scripts/
  6. Commit your changes (git commit -m 'Add amazing skill')
  7. Push to your branch (git push origin feature/amazing-skill)
  8. Submit a pull request with a clear description of your changes

Contribution Guidelines

✅ Adhere to the Agent Skills Specification — Every skill must follow the official spec (valid SKILL.md frontmatter, naming conventions, directory structure)
✅ Include a quoted metadata.version value in every SKILL.md
✅ Increment metadata.version when updating an existing skill
✅ Maintain consistency with existing skill documentation format
✅ Ensure all code examples are tested and functional
✅ Follow scientific best practices in examples and workflows
✅ Update relevant documentation when adding new capabilities
✅ Provide clear comments and docstrings in code
✅ Include references to official documentation

Testing

Every skill that ships scripts/ must have a test suite under tests/<skill-name>/ and an entry in tests/skill-requirements.toml. The tests/_meta guard checks this coverage and the shared structural contract: frontmatter, document length, local links, script syntax, bytecode, and hardcoded local paths. It also validates plugin.json against the bundled Agent Plugins schema and checks that its version matches pyproject.toml. CLI --help and worked-example execution checks belong to the individual skill suites.

# Install repository development and validation tools
uv sync

# Structural contract and coverage guard — seconds, no scientific packages needed
uv run python -m pytest tests/_meta -q

# Validate one skill's Agent Skills frontmatter
uv run skills-ref validate skills/<skill-name>

# One skill's suite in the current environment
uv run --with pytest python -m pytest tests/<skill-name> -q

# One skill with its declared packages and Python version
uv run python tests/run_all.py --isolated <skill-name>

# Every suite, each in its own throwaway environment
uv run python tests/run_all.py --isolated

Run one skill per pytest process: different skills reuse module names such as _common, so collecting several suites together can import the wrong module. tests/run_all.py handles process isolation; --isolated also provides each skill's declared dependency environment. A suite run in the current environment may skip checks when required packages are absent.

The Skill Tests workflow runs the contract and standard-library-only suites for relevant pull requests and pushes to main, and can be triggered manually. The full scientific dependency sweep is not run in CI; run it before a release or when changing the shared contract. Some skills need external runtimes or system tools; installation gaps are documented in tests/skill-requirements.toml.

Security Scanning

All skills in this repository are security-scanned using Cisco AI Defense Skill Scanner, an open-source tool that detects prompt injection, data exfiltration, and malicious code patterns in Agent Skills.

If you are contributing a new skill, we recommend running the scanner locally before submitting a pull request:

uv pip install cisco-ai-skill-scanner
skill-scanner scan /path/to/your/skill --use-behavioral

Note: A clean scan result reduces noise in review, but does not guarantee a skill is free of all risk. Contributed skills are also reviewed manually before merging.

Recognition

Contributors are recognized in our community and may be featured in:

  • Repository contributors list
  • Special mentions in release notes
  • K-Dense community highlights

Your contributions help make scientific computing more accessible and enable researchers to leverage AI tools more effectively!

Support Open Source

This project builds on 50+ amazing open source projects. If you find value in these skills, please consider supporting the projects we depend on.


🔧 Troubleshooting

Common Issues

Problem: Skills not loading

  • Verify skill folders are in the correct directory (see Getting Started)
  • Each skill folder must contain a SKILL.md file
  • Restart your agent/IDE after copying skills
  • In Cursor, check Settings → Rules to confirm skills are discovered

Problem: Missing Python dependencies

  • Solution: Check the specific SKILL.md file for required packages
  • Install dependencies: uv pip install package-name

Problem: API rate limits

  • Solution: Many databases have rate limits. Review the specific database documentation
  • Consider implementing caching or batch requests

Problem: Authentication errors

  • Solution: Some services require API keys. Check the SKILL.md for authentication setup
  • Verify your credentials and permissions

Problem: Outdated examples

  • Solution: Report the issue via GitHub Issues
  • Check the official package documentation for updated syntax

Problem: gh skill install or docs link to scientific-skills/ fails (v2.43.0+)

  • As of v2.43.0, skills live under skills/ (not scientific-skills/) to match the Agent Skills layout expected by GitHub CLI
  • Update manual copy paths, bookmarks, and citations from scientific-skills/<name> to skills/<name>
  • Re-run gh skill install K-Dense-AI/scientific-agent-skills after pulling the latest release

❓ FAQ

General Questions

Q: Is this free to use?
A: Yes! This repository is MIT licensed. However, each individual skill has its own license specified in the license metadata field within its SKILL.md file—be sure to review and comply with those terms.

Q: Why are all skills grouped together instead of separate packages?
A: We believe good science in the age of AI is inherently interdisciplinary. Bundling all skills together makes it trivial for you (and your agent) to bridge across fields—e.g., combining genomics, cheminformatics, clinical data, and machine learning in one workflow—without worrying about which individual skills to install or wire together.

Q: Can I use this for commercial projects?
A: The repository itself is MIT licensed, which allows commercial use. However, individual skills may have different licenses—check the license field in each skill's SKILL.md file to ensure compliance with your intended use.

Q: Do all skills have the same license?
A: No. Each skill has its own license specified in the license metadata field within its SKILL.md file. These licenses may differ from the repository's MIT License. Users are responsible for reviewing and adhering to the license terms of each individual skill they use.

Q: How often is this updated?
A: We regularly update skills to reflect the latest versions of packages and APIs. Major updates are announced in release notes.

Q: Can I use this with other AI models?
A: The core SKILL.md format follows the open Agent Skills standard. Installation paths, discovery, and optional metadata support vary by host and version, so confirm your target host's current documentation.

Installation & Setup

Q: Do I need all the Python packages installed?
A: No. Install only the dependencies for the workflows you use, following each SKILL.md. Keep workflows with incompatible package or Python requirements in separate environments.

Q: What if a skill doesn't work?
A: First check the Troubleshooting section. If the issue persists, file an issue on GitHub with detailed reproduction steps.

Q: Do the skills work offline?
A: Local analysis can work offline once its dependencies, reference data, and model files are available. Database queries, hosted models, and cloud platforms need network access; check the individual skill's requirements.

Contributing

Q: Can I contribute my own skills?
A: Absolutely! We welcome contributions. See the Contributing section for guidelines and best practices.

Q: How do I report bugs or suggest features?
A: Open an issue on GitHub with a clear description. For bugs, include reproduction steps and expected vs actual behavior.


💬 Support

Need help? Here's how to get support:

  • 📖 Documentation: Check the relevant SKILL.md and references/ folders
  • 🐛 Bug Reports: Open an issue
  • 💡 Feature Requests: Submit a feature request
  • 📣 Updates and demos: Follow X, LinkedIn, YouTube, and Reddit to keep up with new skills, tutorials, and Scientific Agent Skills releases
  • 💼 Enterprise Support: Contact K-Dense for commercial support

📖 Citation

If you use Scientific Agent Skills in your research or project, please cite our paper:

Timothy Kassis, Vinayak Agarwal, Yuhuan He, Darshil Patel, and Aubrey M. Brueckner. Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents. arXiv:2609.00065, 2026. https://arxiv.org/abs/2609.00065

When relevant, also cite the individual skill or skills that materially supported your work.

GitHub's Cite this repository button, backed by CITATION.cff, produces the same paper citation in APA or BibTeX.

The paper citation helps others find the repository, understand the broader skill ecosystem used in your workflow, and credit the maintenance effort behind Scientific Agent Skills. Individual skill citations give more precise credit for the specific package, database, or workflow guidance your agent used.

Recommended practice:

  • Always cite the Scientific Agent Skills paper using one of the formats below.
  • Also cite each individual skill that directly contributed to your analysis, code, figures, reports, or research workflow.
  • If a skill wraps or documents an external package, database, or platform, cite that upstream project too when your field's norms require it.

Paper Citation

BibTeX

@misc{kassis2026scientificagentskills,
  title         = {Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents},
  author        = {Kassis, Timothy and Agarwal, Vinayak and He, Yuhuan and Patel, Darshil and Brueckner, Aubrey M.},
  year          = {2026},
  eprint        = {2609.00065},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2609.00065}
}

APA

Kassis, T., Agarwal, V., He, Y., Patel, D., & Brueckner, A. M. (2026). Scientific Agent Skills: A library of procedural knowledge for research agents. arXiv. https://arxiv.org/abs/2609.00065

MLA

Kassis, Timothy, et al. "Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents." arXiv, 2026, arxiv.org/abs/2609.00065.

Plain Text

Kassis, T., Agarwal, V., He, Y., Patel, D., & Brueckner, A. M. (2026). Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents. arXiv:2609.00065. https://arxiv.org/abs/2609.00065

Software Citation

If you also need to cite a specific version of the repository itself (for example, to pin the exact skill set an analysis ran against), add a software citation alongside the paper and record the release tag or commit you used:

@software{scientific_agent_skills_2026,
  author = {{K-Dense Inc.}},
  title = {Scientific Agent Skills: A Comprehensive Collection of Scientific Tools for AI Agents},
  year = {2026},
  url = {https://github.com/K-Dense-AI/scientific-agent-skills},
  note = {181 skills covering databases, packages, integrations, and analysis tools}
}

Individual Skill Citation

When citing a specific skill, include the skill name, version from metadata.version in that skill's SKILL.md, and the direct skill URL. For example:

@software{scientific_agent_skills_astropy_2026,
  author = {{K-Dense Inc.}},
  title = {Astropy Skill for Scientific Agent Skills},
  year = {2026},
  url = {https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/skills/astropy},
  note = {Version 1.5, part of Scientific Agent Skills}
}

Plain text format:

Astropy skill for Scientific Agent Skills, version 1.5.
K-Dense Inc. (2026).
https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/skills/astropy

We appreciate acknowledgment in publications, presentations, or projects that benefit from these skills.


📄 License

This project is licensed under the MIT License.

Copyright © 2026 K-Dense Inc. (k-dense.ai)

Key Points:

  • ✅ Free for any use (commercial and noncommercial)
  • ✅ Open source - modify, distribute, and use freely
  • ✅ Permissive - minimal restrictions on reuse
  • ⚠️ No warranty - provided "as is" without warranty of any kind

See LICENSE.md for full terms.

Individual Skill Licenses

⚠️ Important: Each skill has its own license specified in the license metadata field within its SKILL.md file. These licenses may differ from the repository's MIT License and may include additional terms or restrictions. Users are responsible for reviewing and adhering to the license terms of each individual skill they use.

研究与检索

中风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 可能需要外部 token、网络权限或第三方服务。
  • 未检测到高风险命令。
  • 扫描发现:4 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/K-Dense-AI/scientific-agent-skills.git
  3. 将 "skills/marine-carbonate-chemistry" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/K-Dense-AI/scientific-agent-skills.git
  3. 将 "skills/marine-carbonate-chemistry" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/K-Dense-AI/scientific-agent-skills.git
  3. 将 "skills/marine-carbonate-chemistry" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/K-Dense-AI/scientific-agent-skills.git
  3. 将 "skills/marine-carbonate-chemistry" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/K-Dense-AI/scientific-agent-skills.git
  3. 将 "skills/marine-carbonate-chemistry" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: marine-carbonate-chemistry
description: Solves seawater carbonate chemistry with PyCO2SYS for chemical oceanography, ocean acidification, and marine carbon-cycle research. Use for paired total alkalinity, dissolved inorganic carbon, pH, or seawater pCO2/fCO2 measurements; carbonate speciation; aragonite and calcite saturation; Revelle factors; lab-to-in-situ temperature and pressure corrections; and measurement uncertainty propagation. Applies to carbonate-system calculations, not general aqueous speciation or air-sea gas-flux estimation.
license: MIT
compatibility: Requires Python 3.13 with PyCO2SYS 1.8.3.4 and NumPy. Network access is needed only to install packages or obtain external data; bundled calculations run locally without credentials.
metadata:
  version: "1.1"
  skill-author: K-Dense Inc.
  upstream-version: "PyCO2SYS 1.8.3.4"
  last-reviewed: "2026-10-01"

Marine Carbonate Chemistry

Turn two independent seawater carbonate measurements into a reproducible speciation table, mineral saturation estimates, and a record of the calculation assumptions. Targets PyCO2SYS 1.8.3.4, tested with Python 3.13 and NumPy 2.5.3. As reviewed on 2026-10-01, this remains the stable PyPI release. The v2 documentation is for a beta with breaking changes; use the v1 documentation for this pin.

When to use

  • Analyze bottle samples, shipboard carbonate measurements, or acidification experiments.
  • Calculate total-scale pH, seawater pCO2/fCO2, carbonate ion, aragonite/calcite saturation state, or the Revelle factor from a valid measured pair.
  • Convert a system determined at laboratory conditions to specified ocean conditions.
  • Quantify how stated measurement uncertainties affect the calculated results.

This workflow concerns seawater carbonate equilibria. Freshwater, porewaters with substantial uncharacterized alkalinity, brines outside the selected calibration range, and reaction/transport models require additional chemistry and validation. Do not infer an air-sea flux or atmospheric carbon removal from a carbonate equilibrium alone.

Establish the measurement contract

Before running a solver, identify the two measured variables, their units, quality flags, and their temperature/pressure basis. Retain a separate source table containing station, depth, timestamps, methods, reference materials, and original QC codes, joined by sample ID. Do not turn missing values or rejected measurements into zero.

QuantityRequired convention
Total alkalinity (TA), DIC, nutrientsmicromol per kg seawater, not per litre or kg water
SalinityPractical Salinity, not Absolute Salinity in g/kg
TemperatureIn-situ/measurement temperature in degrees Celsius, not potential or Conservative Temperature
PressureSea pressure in dbar; surface sample is 0, not 1 atmosphere
pHDeclared total, seawater, free, or NBS scale, at the declared measurement conditions
pCO2 / fCO2Seawater partial pressure / fugacity in microatm; these are distinct quantities

TA and DIC remain constant during the solver's temperature/pressure conversion for a closed sample. pH and gas parameters change. Two inputs measured at different conditions cannot simply share one temperature value. Establish a consistent measurement basis first. Temperature correction does not repair sample changes caused by gas exchange, biology, evaporation, or mineral dissolution/precipitation.

Use two independent carbonate parameters. pCO2 plus fCO2 is not an independent pair. Three or more measurements enable an overdetermination check: solve independent pairs and compare predicted versus measured third parameters, including their uncertainty. Do not average inconsistent solutions to hide a calibration or scale mismatch.

Install

Create a dedicated environment in the user's working directory:

uv venv --python 3.13 .venv
uv pip install --python .venv/bin/python "PyCO2SYS==1.8.3.4" "numpy==2.5.3"

On Windows the environment's interpreter is .venv/Scripts/python.exe. The commands below use the POSIX interpreter path. Set the shell variable SKILL_DIR to this installed skill's directory. Keep inputs and generated outputs in the working directory.

Workflow

  1. Prepare paired measurements. Use the schema in references/input-and-results.md. Resolve units and quality flags before creating the input file. Supply phosphate and silicate explicitly; zero is an assumption to justify, not a missing-data code.
  2. Choose equilibrium constants. Read references/chemistry-decisions.md for pH scales, carbonic-acid constants, borate, saturation interpretation, and uncertainty limits. Match the study's validated convention and report it. The helper supports carbonic-acid options 10 and 15; other systems require a separately verified direct PyCO2SYS call.
  3. Solve with scripts/solve_carbonate.py. It validates the full input table, solves the pair, checks finite outputs and DIC species balance, then writes carbonate.csv and provenance.json into a new output directory.
  4. Review flags and consistency. Inspect calibration-range and gas-pressure flags, carbonate balance, measured-third-parameter residuals when available, and controls. A successful solve does not validate the sample, constants, or measurement method.
  5. Report at the intended conditions. Results ending _out describe the supplied output temperature/pressure. Unsuffixed results describe input conditions. Gas results retain the helper's uncorrected hydrostatic gas convention (see below). Include parameter pair, pH scale, units, constants, nutrient assumptions, uncertainty scope, software versions, and excluded/flagged samples with the result table.

Worked example: closed-sample condition correction

The following values are synthetic, not field observations. Save this as samples.csv in a working directory. The two samples differ only in DIC; the second represents a fixed-alkalinity CO2-addition comparison. Their measurements are at 25 C and 0 dbar; results are also requested at 10 C and 1000 dbar.

sample_id,par1,par2,salinity,temperature,pressure,total_phosphate,total_silicate,temperature_out,pressure_out,u_par1,u_par2
baseline,2300,2000,35,25,0,0,0,10,1000,2,2
added_co2,2300,2100,35,25,0,0,0,10,1000,2,2

Run from that working directory:

.venv/bin/python "$SKILL_DIR/scripts/solve_carbonate.py" samples.csv \
  --par1-type alkalinity --par2-type dic --k-carbonic 10 \
  --output-dir carbonate-results

For the baseline, the tested version gives input-condition total pH 8.045886, pCO2 396.958 microatm, and aragonite saturation 3.386201. At the specified output conditions, total pH is 8.241241 and aragonite saturation 2.605691. These rounded values are regression checks for this exact setup, not universal seawater benchmarks. With independent 2 micromol/kg uncertainties in TA and DIC only, u_pH_total is about 0.004580. This excludes equilibrium-constant and other input uncertainty. Both rows carry gas_pressure_correction_disabled_output: the output pH and mineral saturation include pressure effects, but the reported pCO2/fCO2 do not include the hydrostatic corrections to CO2 solubility and fugacity. Do not compare those gas values directly with a pressure-corrected subsurface sensor measurement.

For TA + measured pH, use --par2-type ph --ph-scale total only if the source explicitly identifies total-scale pH; replace par2 and u_par2 with the measured pH and its absolute standard uncertainty. A column named merely pH is insufficient to establish its scale.

Uncertainty and interpretation

Optional u_ input columns contain absolute one-standard-deviation uncertainties. They propagate to total pH, pCO2, and aragonite saturation at each requested condition. The helper assumes independent errors and treats unlisted inputs/constants as exact. For covariance, constants uncertainty, or strongly nonlinear uncertainty, follow the decision guide and validate a tailored propagation instead of calling these outputs a complete uncertainty budget.

Omega < 1 indicates thermodynamic undersaturation with respect to the named mineral. It does not establish a dissolution rate or an organism's response. A lower pH across unmatched samples is not by itself evidence of an anthropogenic acidification trend.

Sources and validation boundary

Repository tests exercise the pinned solver, independent-pair round trips, carbon balance, pH-scale equivalence, condition correction, Revelle-factor derivatives, gas-pressure conventions, uncertainty quadrature, CSV errors, and the worked example. The helper calls the local pyco2.sys Python API; it has no HTTP endpoints or authentication. Tests establish software behavior, not independent field-data validation; upstream's validation page also contains historical examples, including a removed pyco2.test interface.

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!