SkillAtlasSkill 详情

portaljs-check-data-quality

The AI-native framework for building data portals.

审核状态:已审核Quality 80Security 80

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月15日

PortalJS

PortalJS

The AI-native framework for building data portals.
Describe the portal you want — your agent helps you choose an architecture, scaffolds it, and loads your data.

Docs · Discussions · Report a bug

npm version GitHub stars Join our Discord MIT License


Quickstart

Create a portal — one command, nothing to install beyond Node 22+:

npm create portaljs@latest my-portal
cd my-portal
npm run dev      # → http://localhost:3000

You get the three surfaces — Home, a Catalog (/search), and a dataset Showcase (/@<namespace>/<slug>) — over sample data. Plain, editable Next.js, no lock-in. Add your own CSV/JSON to datasets.json and it renders automatically.

Build it with your AI assistant — PortalJS ships Claude Code skills that do the assembly. Install them once (into ~/.claude/commands):

curl -fsSL https://raw.githubusercontent.com/datopian/portaljs/main/scripts/install-portaljs-skills.sh | bash

Then, in a Claude Code session from any directory:

/portaljs-architect    not sure what stack you need? start here
/portaljs-new-portal   "Auckland Council open data portal"
/portaljs-add-dataset  ./data/air-quality.csv

/portaljs-new-portal scaffolds the three surfaces; /portaljs-add-dataset (or /portaljs-add-resource) loads data; /portaljs-connect-ckan points it at a CKAN backend; /portaljs-deploy ships it. (All skills + install →)

Prefer the bare template — plain Next.js, no AI, no lock-in:

npx tiged datopian/portaljs/examples/portaljs-catalog my-portal
cd my-portal && npm install && npm run dev      # → http://localhost:3000

You get Home, a Catalog (/search), and a dataset Showcase (/@<namespace>/<slug>) over sample data. Add your own CSV/JSON to datasets.json and it renders automatically.

⭐ If it's useful, a star helps others find it.

Why PortalJS

Building a data portal has always meant more than a website. You have to decide where the data lives, how it's versioned, how people search it, how it's served, and how it's governed — and then wire a frontend on top. Teams either over-build on a heavy data warehouse they don't need, or under-build on a pile of scripts that doesn't scale.

PortalJS is an open-source, agentic skills framework that helps data teams build, develop, and ship data portals — and the data infrastructure underneath them. It isn't only a frontend. The skills do two jobs:

  • Advise — given what you're building, what your data is, and what it's for, they recommend an architecture: storage, compute, catalog, access, hosting, metadata.
  • Build — they scaffold that stack as plain, editable Next.js code with no lock-in.

It is opinionated but open: the recommended modern path is git + object storage (Cloudflare R2) + Parquet, queried with DuckDB — an open lakehouse instead of a classic warehouse. For living, incremental tables you can layer on DuckLake, and a traditional datastore (CKAN, a warehouse) stays a first-class option when you need it. You always own plain code.

Built and maintained in the open by Datopian and the PortalJS community.

Architecture at a glance

        🧑  you describe what you want to build
        │
        ▼
╭─ 🤖  AGENTIC SKILLS ──────────────────────────────────  decide + build
│   /portaljs-architect · /portaljs-new-portal · /portaljs-add-dataset · /portaljs-add-chart · /portaljs-add-map …
╰─  generates plain, editable Next.js code — no lock-in
        │
        ▼
╭─ 🖥️  SURFACES ────────────────────────────────────────  what users see
│   🏠 Home /      🔎 Catalog /search      📊 Showcase /@ns/slug
╰─  read data through one DataProvider contract
        │
        ▼
╭─ 🔌  PROVIDERS ───────────────────────────────────────  pluggable backends
│   📁 static·git     🐘 CKAN     🔭 OpenMetadata     🗂️ git-LFS + R2
╰─  swap the source without touching a page
        │
        ▼
📦  STORAGE + COMPUTE  —  choose your point on the spectrum:

      flat files  ─▶  Git-LFS + R2  ─▶  Parquet on R2 + 🦆 DuckDB  ─▶  warehouse / CKAN
      simplest                       ⭐ open lakehouse (default)        heaviest
                                     (+ DuckLake for living tables)

☁️  Substrate  —  Cloudflare R2 (storage) · Workers (runtime) · D1 (catalog) · Pages (static)
     object storage stays S3-compatible — R2 is the default, never a lock-in

Three surfaces. Every data portal is built from three: a Home page that explains it and offers search, a Catalog (/search) to discover datasets, and a Showcase (/@<namespace>/<slug>) to explore one dataset — metadata, preview, download/API, and charts/maps. (Core concepts →)

One seam. The surfaces read data only through a DataProvider, so the source — static files today, a CKAN or lakehouse backend tomorrow — can change without touching a page.

See ROADMAP.md for the full model and the architecture decision framework for how /portaljs-architect turns your needs into a stack.

Build a portal with your AI assistant

PortalJS ships Claude Code skills that turn a brief into a working portal.

Setup

Install the skills once into your personal scope so they're available from any directory:

curl -fsSL https://raw.githubusercontent.com/datopian/portaljs/main/scripts/install-portaljs-skills.sh | bash

Restart Claude Code (or open a new session) and type / to see them. See .claude/INSTALL.md for other install options (versioned plugin, or running straight from a clone of this repo).

Use

If you're not sure how to set up your portal, start with the advisor, then build:

/portaljs-architect    we have ~200 public CSVs, updated quarterly, and must publish DCAT-AP
/portaljs-new-portal   "Auckland Council open data portal"
/portaljs-add-dataset  ./data/air-quality.csv
/portaljs-add-dataset  https://example.com/parks.geojson

The skills are interactive — if your brief is thin, they interview you in short rounds rather than erroring. /portaljs-architect recommends a stack and hands off; /portaljs-new-portal scaffolds the three surfaces; /portaljs-add-dataset appends to the datasets.json manifest and the showcase renders automatically at /@<namespace>/<slug>. Run npm run dev and you have a portal.

Prefer to build by hand? The skills are a convenience, not a requirement — scaffold the template directly with the CLI:

npm create portaljs@latest my-portal

(Or grab the bare template with no prompts: npx tiged datopian/portaljs/examples/portaljs-catalog my-portal.)

Available skills

SkillWhat it does
/portaljs-architectAdvisory — turns your needs (data, scale, governance) into a recommended architecture before you build. Start here if you're unsure of the stack.
/portaljs-new-portalScaffold a new portal (Home + Catalog + Showcase) from a brief — copies the template, substitutes your project name and description, installs deps, verifies the build.
/portaljs-add-datasetAdd a CSV, TSV, JSON, or GeoJSON dataset — registers it in the catalog and renders its showcase automatically; large local files are pushed to Cloudflare R2 via Git LFS for you.
/portaljs-add-resourceAttach another file (data dictionary, methodology, extra data) to an existing dataset — it becomes multi-resource and the showcase renders a section per file.
/portaljs-add-chartAdd a line, bar, area, pie, or scatter chart to a dataset's showcase.
/portaljs-add-mapRender a GeoJSON dataset on an interactive map and register it on the home page.
/portaljs-add-geoAuto-ingest a geospatial file (GeoJSON, Shapefile, GeoPackage, KML/KMZ, FlatGeobuf, CSV-with-geometry) on your own machine — no server: normalizes CRS to EPSG:4326, derives a PMTiles render tier and a GeoParquet query tier, pushes both plus the original to R2, and registers one dual-tier dataset the showcase maps and queries in place.
/portaljs-define-schemaInfer a Frictionless Table Schema from a dataset's data, add license/source/keyword metadata, and surface a typed field table on its showcase.
/portaljs-add-dcatMake the portal harvestable — emit standards-compliant DCAT feeds (DCAT 2/3, DCAT-AP, DCAT-US, national profiles) in JSON-LD, Turtle, and RDF/XML so national/EU/US open-data portals can harvest its datasets.
/portaljs-connect-ckanWire the portal to a CKAN backend over its API instead of static files.
/portaljs-check-data-qualityValidate a dataset against its schema and flag quality issues (type mismatches, missing values, constraint violations).
/portaljs-migrateHarvest or migrate a whole catalog into the portal from CKAN, Socrata, OpenDataSoft, ArcGIS, or DCAT-US, over a canonical Frictionless/DCAT model.
/arcgis-to-portaljsMigrate a whole ArcGIS Hub site into the portal end-to-end — harvest its /data.json, export every FeatureService layer via the ArcGIS REST API, convert to the serverless dual tier (PMTiles + GeoParquet), push to R2, and write a source-vs-derived parity report.
/portaljs-deployBuild a static export and publish it to PortalJS Arc — Datopian-managed hosting on Cloudflare — with a live <slug>.arc.portaljs.com URL.

Large-data scaling — big files pushed to Cloudflare R2 via Git LFS — already ships in /portaljs-add-dataset. More skill families — metadata schemas (Frictionless/DCAT), more backends (OpenMetadata), a browser DuckDB query layer, and access control — are on the roadmap. Write your own — see .claude/AUTHORING.md.

What's in this repo

.claude/commands/    the agentic skills (slash commands)
examples/            reference portals — portaljs-catalog is the canonical template
packages/
  core/              layout/UI components            (@portaljs/core)
  ckan/              CKAN catalog UI + React          (@portaljs/ckan)
  ckan-api-client-js/ pure CKAN API client            (@portaljs/ckan-api-client-js)
site/                portaljs.com — the marketing site + docs
ROADMAP.md           direction, the four contracts, sequencing

The canonical template, examples/portaljs-catalog, is where the three surfaces and the DataProvider seam live — read it before building.

What makes it different

  • 🌱 Open source, MIT, no lock-in — every skill emits plain Next.js you can fork and own.
  • 🧭 Advisory, not just generative — /portaljs-architect helps you decide the infrastructure, not only scaffold a UI.
  • 🦆 Open lakehouse by default — git + R2 + Parquet queried with DuckDB, over a heavy warehouse; add DuckLake for living/incremental tables. A datastore/warehouse stays a supported choice.
  • ☁️ Cloudflare-first, portable — R2 / Workers / D1 / Pages as the default substrate, but object storage stays S3-compatible.
  • 🧩 Decoupled, any backend — one DataProvider contract in front of CKAN, DKAN, OpenMetadata, DataHub, GitHub, Frictionless, plain files — or your own.
  • 🎨 Bring your own stack — adopt the template or lift the skills and the three-surface model into an app you already have.

Examples

Reference implementations live in examples/:

ExampleBackend
portaljs-catalogCanonical template — Home + Catalog + Showcase over a static manifest
portaljs-templateMinimal single-page starter
ckan · ckan-ssgCKAN
github-backed-catalogGitHub
dataset-frictionlessFrictionless Data Package
fivethirtyeight · openspending · turingReal-world portals

Community & support

Contributing

PortalJS is built in the open and we welcome contributions of all sizes — new skills, examples, docs, and fixes. See CONTRIBUTING.md to get started, and read ROADMAP.md and VISION.md for where the project is headed.

License

MIT © Datopian

数据与 AI

中风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 未检测到明显外部权限要求。
  • 未检测到高风险命令。
  • 扫描发现:2 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/datopian/portaljs.git
  3. 将 "skills/portaljs-check-data-quality" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/datopian/portaljs.git
  3. 将 "skills/portaljs-check-data-quality" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/datopian/portaljs.git
  3. 将 "skills/portaljs-check-data-quality" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/datopian/portaljs.git
  3. 将 "skills/portaljs-check-data-quality" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/datopian/portaljs.git
  3. 将 "skills/portaljs-check-data-quality" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: portaljs-check-data-quality
description: Audit a local or remote tabular file (CSV/TSV) for common data quality issues — schema, nulls, types, duplicates. Read-only. Use when a dataset needs a quality check before publishing, or a showcase renders wrong (blank cells, garbled numbers, an unsortable date column) and the cause needs isolating.
allowed-tools: Bash(curl:*), Bash(awk:*), Bash(sort:*), Bash(head:*), Bash(wc:*)
version: 1.0.0
author: Datopian <hello@datopian.com>
license: MIT
compatibility: Claude Code with PortalJS portals (Next.js 14, React 18, Node 18+). Runs from any project via the plugin, a personal ~/.claude/commands install, or a portaljs clone.
tags:
  - portaljs
  - data-portal
  - data-quality
  - audit
  - csv
  - validation

PortalJS — Check Data Quality

Overview

Run a read-only quality audit of one CSV or TSV file, local or remote, and return a structured JSON report. The audit profiles every column — null/blank counts, inferred value types, numeric ranges, likely year/date fields — and flags duplicate rows, duplicate values in identifier-like columns, ambiguous overlapping year columns (e.g. calendar year vs fiscal year), and mixed-type columns. It never edits the source file, datasets.json, or any other project file; it only reads the target file (a remote URL is downloaded to a temp file that is deleted before the run ends) and prints a report. Use it before publishing a dataset with portaljs-add-dataset, or to diagnose why a showcase renders wrong.

Prerequisites

  • python3 on PATH — the audit logic runs as an embedded Python script; nothing is installed.
  • One CSV or TSV file, given as a local path or an http/https URL. Only one file per run.

Instructions

The canonical, full step-by-step workflow is .claude/commands/portaljs-check-data-quality.md — the single source of truth. Read and follow it when executing. Summary:

  1. Gather input — the file path or URL to audit. If missing, ask for it; never dead-end.
  2. Resolve the source: if it's an http/https URL, download it to a temp file first; otherwise use the local path as given.
  3. Validate the extension is .csv or .tsv. If not, or the file is missing, or the header row is empty, stop and surface the error JSON as-is — do not guess a fix.
  4. Profile every column: null/blank counts, distinct values, sample values, inferred per-value type (boolean/integer/float/date/string), numeric min/max, and year range for columns whose name looks year-like.
  5. Derive findings from the profiles — duplicate rows, missing-value ratios, invalid year values, mixed types, suspect negative values, duplicate identifier values, and ambiguous overlapping year columns — each tagged critical, warning, or info.
  6. Assemble the JSON report (status, file metadata, findings, recommendations, column_profiles), print it, and clean up the temp file if one was created.
  7. Relay the report to the user as-is; do not modify the source file, datasets.json, or any other project file based on the findings — that's a separate, explicit step.

Output

A single JSON object printed to stdout:

  • status — ok, warning, or critical.
  • file, file_name, source_type (local or url), row_count, column_count.
  • findings — structured issues, most severe first.
  • recommendations — de-duplicated suggested next steps.
  • column_profiles — per-column summary (nulls, blanks, distinct count, sample values, inferred types, numeric/year ranges).

No files are created or modified. A remote URL's temp download is removed on exit, success or failure alike.

Error Handling

SymptomCauseFix
"File ... is not available."Local path is wrong, or the URL download failedVerify the path or URL is reachable and retry.
"Only CSV and TSV files are supported right now."File extension isn't .csv/.tsvConvert the file, or point to its tabular source instead.
"... does not contain tabular headers."File is empty or the header row is malformedOpen the file and confirm it has a valid, non-empty header line.
Command hangs on a URLRemote host is slow or blocks non-browser requestsDownload the file manually and audit the local copy instead.
python3: command not foundPython 3 isn't installed or not on PATHInstall Python 3, or run the audit where it's available.
Report looks truncated in the terminalLarge report wrapped/paginated by the shellRedirect to a file (> report.json) and open it separately.

Examples

Example 1 — Audit a local CSV before publishing

/portaljs-check-data-quality ./public/data/trash.csv

Example 2 — Audit a remote CSV over HTTPS

/portaljs-check-data-quality https://example.com/trash.csv

Example 3 — Audit a TSV and save the report for review

bash scripts/check-data-quality.sh ./data/emissions.tsv > /tmp/emissions-quality.json

Example 4 — Read a critical status report

{
  "status": "critical",
  "findings": [
    { "severity": "critical", "check": "duplicate_rows", "message": "42 duplicate rows found." }
  ],
  "recommendations": ["Review and deduplicate repeated rows if they are not intentional."]
}

Fix the flagged rows/columns, then re-run the audit before publishing.

Resources

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!