SkillAtlasSkill 详情

databricks-cost-leak-hunter

A model-agnostic agent-skills platform.

审核状态:已审核Quality 72Security 52

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月29日

Tons of Skills

A model-agnostic agent-skills platform. The canonical layer is harness-free by construction; Claude Code is currently the verified-native harness. Other harnesses remain engineering candidates until their native-path integration is verified; source research alone is never presented as public support.

Release CLI Plugins Skills GitHub Stars skills.sh Sponsor: Kobiton Buy me a monster

ko-fi

Version semantics: the release badge is this marketplace's display version. npm packages, including the ccpi CLI and publishable plugins, retain their own package versions; they are intentionally not expected to equal the display version. The version-surface checker governs the display surfaces without rewriting package semver.

Install

Inside Claude Code, one command installs the whole marketplace:

/plugin marketplace add jeremylongshore/claude-code-plugins

Or use the CLI:

pnpm add -g @intentsolutionsio/ccpi
ccpi install devops-automation-pack

Browse the marketplace · Explore plugins · Download bundles

Killer Skill of the Week — no-ai-slop by Peter Yang

Strip AI slop from any draft — named-pattern edits that keep the writer's real voice

no-ai-slop does two jobs and refuses to fake a third. In Edit mode it makes the minimum effective edit — cutting throat-clearing, weak verbs, and abstract nouns while deliberately preserving the writer's cadence, bluntness, humor, and honest admissions, so a rough draft still sounds like the same person afterward. In Detect mode it names each AI-slop pattern it finds, quotes the offending line, and gives the fix in a few words — and pointedly does NOT score the draft or guess whether an AI wrote it. That restraint is the whole point: AI detectors guess; named patterns are evidence the reader can check. MIT-licensed, single focused skill, actively maintained by Peter Yang.

"AI detectors guess. Named patterns are evidence the user can check." — Peter Yang

Grade: A | Week of July 22, 2026 (W30) | View on GitHub

Previous picks: tonone, mnemos, databricks-pack, kobiton-automate, skyvern, code-cleanup, web-analytics, token-optimizer, executive-assistant-skills, skill-creator, cursor-pack, crypto-portfolio-tracker. See all at tonsofskills.com.

Scale, labeled

Every number below names the cohort it counts and the command that reproduces it — an unlabeled count is how a corpus ends up with five contradictory answers to "how many skills."

CountCohortReproduce with
442catalog plugins (catalog-entry cohort)node scripts/generate-readme-toc.mjs over marketplace.extended.json
3,067marketplace-visible skills (distinct)node -e "import('./scripts/corpus-resolver.mjs').then(m=>console.log(m.resolveCorpus('marketplace-visible').length))"
347agent definitions in pluginsgit ls-files 'plugins/**' | grep '/agents/.*\.md'
19plugin categoriesls -d plugins/*/

📦 Live npm Downloads

Across 396 published packages in the claude-code-plugins namespace. Updated daily by GitHub Actions.

WindowAll packagesEstablished (>30d)
Last 24 hours962962
Last 7 days2,9202,916
Last 30 days12,86812,779

"Established" excludes packages first published within the last 30 days, so a bulk-publish event doesn't dominate the headline.

Top 10 by last 30 days:

#PackageLast 30d
1@intentsolutionsio/openrouter-pack556
2@intentsolutionsio/groq-pack496
3@intentsolutionsio/databricks-pack274
4@intentsolutionsio/clickhouse-pack273
5@intentsolutionsio/wallet-security-auditor263
6@intentsolutionsio/notion-pack258
7@intentsolutionsio/elevenlabs-pack244
8@intentsolutionsio/freshie-inventory-manager214
9@intentsolutionsio/supabase-pack210
10@intentsolutionsio/agency-os204

Last refreshed 2026-08-19T03:03:05.709Z.

Ways in

Five real questions, five doors — each resolves to a live, generated surface, never a hand-maintained list:

Browse by category

The 19 categories below link into the live marketplace. Plugin counts are the catalog-entry cohort — regenerated from marketplace.extended.json by this generator; the catalog itself lives on tonsofskills.com, never in this file (§ 6A of the platform blueprint).

CategoryPlugins
🤖AI & Machine Learning36
🎭AI Agents & Agency10
🔌API Development26
💼Business Tools6
👥Community21
₿Crypto & Web327
💾Database26
🎨Design2
🔧DevOps & Infrastructure36
📚Examples & Templates5
🧩MCP Servers16
📦Packages5
⚡Performance25
✅Productivity30
🎁SaaS Skill Packs106
🔐Security27
✨Skill Enhancers9
🧪Testing28
📁Analytics1

What the classes mean

Four artifact classes live in this repository, distinguished on sight and never blurred — provenance is a truth requirement here, not a UX nicety:

ClassWhat it isHow the reader can tell
Canonical skillFirst-party, harness-free, the source of truthNo .source.json in its plugin directory
Generated adapterA thin, machine-produced harness projectionLives under a generated path with a "generated — do not edit" header
First-party packageAn Intent Solutions distribution (npm, cowork zip)@intentsolutionsio scope, IS-authored license
Upstream mirrorSomebody else's work, hosted mirror-by-default.source.json present — upstream author, license, and pinned commit recorded

Certification

Not yet certified. The certification program (tiers T0–T4 with retained, hash-matched evidence) is a later epic of the platform blueprint; until its report exists, no artifact on this surface claims a tier. This line is rendered from the absence of certification-report.json — honestly, not cosmetically.

Contribute

Start with the contribution guide, then the intake and review standards every submission passes through:

Governance

Provenance

External plugins are hosted mirror-by-default: the contributor's repository stays the source of truth, every mirrored source is pinned in a content lockfile, and upstream credit — author, license, resolved commit — is recorded in the mirror itself. Improvements flow by upstreaming to the author's repository, never by silently editing the mirror. The full decision record is the external-sync model.

License

MIT for the repository scaffolding and first-party tooling; each plugin carries its own license in its manifest, and mirrored plugins keep their upstream license verbatim.

其他

高风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 可能需要外部 token、网络权限或第三方服务。
  • 存在潜在风险命令,请谨慎安装。
  • 扫描发现:4 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/jeremylongshore/tons-of-skills-marketplace.git
  3. 将 "skills/.curated/databricks-cost-leak-hunter" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/jeremylongshore/tons-of-skills-marketplace.git
  3. 将 "skills/.curated/databricks-cost-leak-hunter" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/jeremylongshore/tons-of-skills-marketplace.git
  3. 将 "skills/.curated/databricks-cost-leak-hunter" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/jeremylongshore/tons-of-skills-marketplace.git
  3. 将 "skills/.curated/databricks-cost-leak-hunter" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/jeremylongshore/tons-of-skills-marketplace.git
  3. 将 "skills/.curated/databricks-cost-leak-hunter" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: databricks-cost-leak-hunter
description: |
  Hunt down Databricks cost leaks — wasted DBUs, idle clusters, oversized SQL
  warehouses, and untagged runaway spend — and produce a FinOps cost report.
  Use when a user asks why their Databricks bill is high, wants to find cost
  leaks / wasted DBUs / idle clusters, or needs a FinOps cost report.
  Trigger with "databricks cost", "why is my databricks bill",
  "find wasted spend", "cost leak".
allowed-tools: Read, Write, Edit, Bash(databricks:*), Bash(jq:*), Glob, mcp__databricks-workspace-mcp__clusters_get, mcp__databricks-workspace-mcp__clusters_events, mcp__databricks-workspace-mcp__clusters_list, mcp__databricks-workspace-mcp__instance_pools_list, mcp__databricks-workspace-mcp__pipelines_get
version: 2.27.0
author: Jeremy Longshore <jeremy@intentsolutions.io>
license: MIT
compatibility: Designed for Claude Code
tags: [saas, databricks, finops, cost]

Databricks Cost Leak Hunter

Audits a Databricks workspace for real-dollar cost leaks — idle compute, jobs on the wrong SKU, overprovisioned clusters, and the Photon premium paid without the speedup — then emits a CFO-grokkable, dollar-ranked FinOps remediation report.

Overview

This skill finds where Databricks money is leaking and how much, in dollars, per month. Confirmed-spend figures come from the customer's own system.billing.usage joined to system.billing.list_prices — never an estimate. Two of the four categories (overprovisioning, Photon premium) are explicitly modeled/at-risk amounts, labeled as such so a CFO never confuses them with recoverable spend. The skill surfaces four named leak categories, ranks them by monthly dollar impact, and explains each root cause in FinOps language a CFO can act on without an engineer to translate.

It is architecturally distinct from the v1 databricks-cost-tuning skill: that one AUTHORS policy (creates cluster policies, spot configs). This one DETECTS leaks and reports them, dollarized and ranked. The math is deterministic — bundled scripts in scripts/ do the arithmetic so the agent never eyeballs numbers — and deep domain knowledge lives in references/ loaded only when a leak needs it.

The skill uses two data planes. Dollar figures come from the Databricks CLI Statement Execution API (databricks api post /api/2.0/sql/statements) reading system.* — authenticated by the CLI's own DATABRICKS_HOST+DATABRICKS_TOKEN / databricks auth login. The live config/event evidence that explains why a leak exists (auto-termination setting, node type, autoscale floor, pool min_idle) comes from the custom databricks-workspace-mcp control-plane tools — the one MCP dependency. The SQL produces the number; the workspace MCP turns it into a verified, single-config-change fix.

Prerequisites

Read access to the billing system tables is the hard dependency and the most common failure. system.billing.usage requires a metastore-admin grant chain — the skill detects a missing grant upfront and reports it, rather than failing mid-flow.

  • Databricks Premium or Enterprise workspace with Unity Catalog (system tables are UC-governed; not available on Standard).
  • Metastore-admin grant chain on the billing schema, granted by a metastore admin to the running principal:
    • GRANT USE CATALOG ON CATALOG system TO <principal>
    • GRANT USE SCHEMA ON SCHEMA system.billing TO <principal>
    • GRANT SELECT ON TABLE system.billing.usage TO <principal>
    • GRANT SELECT ON TABLE system.billing.list_prices TO <principal>
  • Compute system schema for config/utilization corroboration (same chain on system.compute): system.compute.clusters, system.compute.node_timeline.
  • Databricks CLI authenticated (databricks auth login, or DATABRICKS_HOST + DATABRICKS_TOKEN) and jq for parsing JSON tool output. The dollar queries run through the CLI Statement Execution API — UC enforces the grant chain above.
  • DATABRICKS_WAREHOUSE_ID env var set to a running SQL warehouse — every statement-execution call (Step 1's probe included) requires it.
  • databricks-workspace-mcp registered (its own PAT/U2M/M2M auth; PAT unsupported in Databricks-App deployment mode). It reads the live REST API, not system.*, so it needs no system-table grants. If it is absent the skill still produces dollar figures but cannot corroborate live config — it then accepts pasted config input.

Authentication. The CLI Statement Execution API uses DATABRICKS_HOST + DATABRICKS_TOKEN or databricks auth login; UC enforces the metastore grant chain on every system.* read. The custom databricks-workspace-mcp authenticates separately via its own PAT / U2M / M2M token. No secrets are hardcoded — all auth comes from the environment or the registered MCP server.

Run the upfront grant check before any analysis — see Step 1.

Instructions

The pipeline is detect → compute → rank → report. SQL detection runs through the CLI Statement Execution API; config corroboration runs on databricks-workspace-mcp; the dollar arithmetic runs in scripts/; deep knowledge loads from references/ on demand.

Step 1: Verify the Grant Chain (fail fast, not mid-flow)

Probe the billing tables before anything else. If the probe errors with a permission message, STOP and report the exact missing grant — do not continue into the leak scans. Requires DATABRICKS_WAREHOUSE_ID (a running SQL warehouse).

databricks api post /api/2.0/sql/statements --json '{
  "warehouse_id": "'"$DATABRICKS_WAREHOUSE_ID"'",
  "statement": "SELECT 1 FROM system.billing.usage LIMIT 1",
  "wait_timeout": "30s"
}' | jq -r '.status.state, .status.error.message // "ok"'

If state is not SUCCEEDED, load ${CLAUDE_SKILL_DIR}/references/system-tables-setup.md and report the missing grant chain to the user verbatim. Stop here.

Step 2: Pull the Spend Baseline

Establish the trailing-30-day total spend so every leak can be expressed as a share of a real number, and capture the window's MAX(usage_date) to stamp into the report. The price-window join (usage × list_prices.pricing.default, matched on sku_name AND usage_unit within the price-effective window, currency_code='USD') is the dollar primitive reused by every category query.

# The CLI does NOT expand ${VARS} inside a --json @file, so inject the warehouse
# id with jq at call time (the static template carries only wait_timeout + statement).
databricks api post /api/2.0/sql/statements --json "$(
  jq --arg wh "$DATABRICKS_WAREHOUSE_ID" '. + {warehouse_id: $wh}' \
    "${CLAUDE_SKILL_DIR}/scripts/sql/spend-baseline.sql.json"
)"

The canonical CTE and full per-category SQL live in ${CLAUDE_SKILL_DIR}/references/cost-leak-categories.md. Load it now — the four detection queries below all reference its priced CTE.

Step 3: Detect Leak 1 — Clusters That Never Auto-Terminate

Join priced All-Purpose usage to system.compute.clusters; flag clusters whose latest-change auto_termination_minutes = 0. Rank by 30-day idle spend. This is confirmed spend — money actually billed for idle compute.

SELECT p.usage_metadata.cluster_id AS cluster_id,
       COALESCE(c.cluster_name, 'unknown') AS cluster_name,
       c.auto_termination_minutes,
       ROUND(SUM(p.usd), 2) AS spend_30d_usd
FROM priced p
JOIN cluster_cfg c ON p.usage_metadata.cluster_id = c.cluster_id
WHERE p.billing_origin_product = 'ALL_PURPOSE'
  AND c.auto_termination_minutes = 0
GROUP BY p.usage_metadata.cluster_id, c.cluster_name, c.auto_termination_minutes
HAVING SUM(p.usd) > 0
ORDER BY spend_30d_usd DESC;

Corroborate each flagged cluster's live config with databricks-workspace-mcp clusters_get (confirm autotermination_minutes = 0 right now) and clusters_events (measure the idle gap between RUNNING and TERMINATING).

Step 4: Detect Leak 2 — Scheduled Jobs on All-Purpose Compute

The signature leak: a usage row with a job_id in usage_metadata AND billing_origin_product = 'ALL_PURPOSE' ($0.55/DBU) instead of JOBS_COMPUTE ($0.15/DBU). Re-price the same DBUs at the current Jobs rate to compute savings. This is confirmed savings — a deterministic re-pricing delta. The jobs_rate CTE is deduped to one USD rate per usage_unit so the join cannot fan out (see cost-leak-categories.md).

SELECT p.usage_metadata.job_id AS job_id,
       ROUND(SUM(p.usd), 2) AS spend_on_all_purpose_30d_usd,
       ROUND(SUM(p.usd) - SUM(p.usage_quantity * jr.jobs_unit_price), 2)
         AS potential_savings_30d_usd
FROM priced p
JOIN jobs_rate jr ON p.usage_unit = jr.usage_unit
WHERE p.billing_origin_product = 'ALL_PURPOSE'
  AND p.usage_metadata.job_id IS NOT NULL
GROUP BY p.usage_metadata.job_id
HAVING SUM(p.usd) - SUM(p.usage_quantity * jr.jobs_unit_price) > 0
ORDER BY potential_savings_30d_usd DESC;

Confirm the live compute type is All-Purpose (not Jobs) with clusters_list / clusters_get (cluster_source) before recommending the move. The job_id + ALL_PURPOSE billing signal is itself dollar-accurate; the REST check is belt-and-suspenders.

Step 5: Detect Leak 3 — Overprovisioned Clusters Idling Below Floor

Aggregate mean CPU from system.compute.node_timeline, join to 30-day spend, flag clusters burning real dollars at chronically low utilization (< 25%). This figure is an estimate (est_overprovision = spend × (1 − CPU%)), not billed waste — it is the one modeled number in the pipeline and is labeled est_* everywhere.

SELECT s.cluster_id,
       ROUND(u.avg_cpu_pct, 1) AS avg_cpu_pct,
       ROUND(s.spend_30d_usd, 2) AS spend_30d_usd,
       ROUND(s.spend_30d_usd * (1 - LEAST(u.avg_cpu_pct,100)/100.0), 2)
         AS est_overprovision_30d_usd
FROM spend s
JOIN util u ON s.cluster_id = u.cluster_id
WHERE u.avg_cpu_pct < 25 AND s.spend_30d_usd > 0
ORDER BY est_overprovision_30d_usd DESC;

Corroborate the configured floor with clusters_get (REST nested autoscale.min_workers / autoscale.max_workers); for the idle-pool variant use instance_pools_list (min_idle_instances + stats.idle_count) — pool waste is NOT a billing row.

Step 6: Detect Leak 4 — Photon Premium Without the Speedup

Photon is not a column on system.compute.clusters; it is billing-visible via the SKU. Isolate usage whose sku_name ILIKE '%PHOTON%' and surface the ~2× premium portion as the at-risk amount — money for review against actual runtime gain, not confirmed waste.

SELECT p.usage_metadata.cluster_id AS cluster_id,
       ROUND(SUM(p.usd), 2)        AS photon_spend_30d_usd,
       ROUND(SUM(p.usd) / 2.0, 2)  AS photon_premium_at_risk_30d_usd
FROM priced p
WHERE p.sku_name ILIKE '%PHOTON%'
  AND p.billing_origin_product IN ('ALL_PURPOSE','JOBS_COMPUTE')
  AND p.usage_metadata.cluster_id IS NOT NULL
GROUP BY p.usage_metadata.cluster_id
HAVING SUM(p.usd) > 0
ORDER BY photon_premium_at_risk_30d_usd DESC;

Confirm Photon is live and worth keeping with databricks-workspace-mcp clusters_get (REST runtime_engine — a config-plane field, not a system column); for DLT pipelines use pipelines_get (spec.photon / serverless / edition). See ${CLAUDE_SKILL_DIR}/references/dlt-tier-cost-tradeoffs.md when the leak touches DLT/serverless tiers.

Step 7: Compute, Rank, and Write the Report

Pass each category's query result to the deterministic ranker — the LLM does NOT do the arithmetic. Each leak object carries a kind field (confirmed / estimated / at-risk) so the renderer can split the headline into confirmed-recoverable vs estimated/at-risk-pending-review and stamp a Confidence column. The script sums per-category figures by kind, ranks descending by monthly dollar impact, annualizes the headline and #1 line, stamps the trailing-30-day window end date, and renders the CFO-grokkable report.

# Per-category results and the rendered report are RUNTIME outputs — they go to
# a working dir ($OUT), never the skill package. Steps 3–6 wrote leak-*.json here.
OUT="${OUT:-$(pwd)/cost-leak-out}" && mkdir -p "$OUT"
jq -s '.' "$OUT"/leak-*.json | \
  python3 "${CLAUDE_SKILL_DIR}/scripts/rank-and-report.py" \
    --monthly-spend 100000 \
    --window-end "$WINDOW_END_DATE" \
    --out "$OUT/cost-leak-report.md"

Use Glob to collect the per-category leak-*.json results, Write the rendered report, and Edit it if the user wants the headline spend rescaled. Render the output using the verbatim template in ${CLAUDE_SKILL_DIR}/references/cfo-output-format.md.

Output

  • A CFO-grokkable report file ($OUT/cost-leak-report.md in the working dir) leading with a split headline that never sums confirmed and unconfirmed dollars under one verb — ### A $<spend>/month workspace is burning **~$<confirmed>/month** (confirmed), plus up to **~$<at-risk>/month** pending review — each with its ~$<annualized>/year companion.
  • A trailing-30-day window stamp under the headline (Trailing 30 days ending <window-end>) so every figure has an explicit calendar window, not just a /month cadence label.
  • The ranked leak table with a Confidence column (# | Where it's leaking | $/month | Confidence | The fix), one row per category, ranked highest dollar impact first, $/month right-aligned, each fix a single config change. Root-cause cells use plain-business language — no raw DBU unit in the CFO-visible text (DBU detail stays in the per-leak detail artifacts).
  • The #1-line callout — the top leak annualized, named, with its confidence, and stated as fixed in one setting.
  • The assumed-vs-cited disclosure — only the workspace-spend input is assumed; on a live run confirmed figures are computed from system.billing.usage, while overprovision (estimated) and Photon premium (at-risk) are labeled as modeled.
  • Per-leak detail artifacts: idle-cluster list, all-purpose-job migration list with per-job savings, overprovisioned-cluster rightsizing list, Photon-premium at-risk list — each with the corroborating live config from databricks-workspace-mcp and the underlying $/DBU rates for engineers.

Error Handling

ErrorCauseSolution
PERMISSION_DENIED on system.billing.usageMetastore-admin grant chain missingRun Step 1; report the exact GRANT USE CATALOG / USE SCHEMA / SELECT chain from system-tables-setup.md. Stop, do not continue.
CLI not authenticated / token expiredNo valid DATABRICKS_HOST + tokenRe-run databricks auth login; verify with databricks current-user me.
Empty / unset DATABRICKS_WAREHOUSE_IDRequired warehouse for statement execution not setSet DATABRICKS_WAREHOUSE_ID to a running SQL warehouse before Step 1.
list_prices join returns NULL usdCustom/negotiated pricing not in list_prices, or usage_unit mismatchJoin on sku_name AND usage_unit within the price window with currency_code='USD'; if still NULL, use the customer's contracted rate card from references/cost-leak-categories.md.
Workspace MCP missingServer not registeredDegrade gracefully: report it absent, run the dollar half, accept pasted config for corroboration. Never fail silently mid-flow.
node_timeline empty for a clusterServerless/short-lived compute, or monitoring lagSkip Leak 3 for that cluster; note "utilization unavailable" rather than reporting $0 overprovision.
Untagged spend / no cluster_nameClusters lack CostCenter/Team tagsAttribute by cluster_id; flag attribution as incomplete in the report footer.

Examples

Example 1: "Why is my Databricks bill high?"

Runs the full pipeline. The grant check passes, the four scans return rows, and the ranker emits the CFO report with a split, confidence-stamped headline:

### A $100K/month Databricks workspace is burning **~$19,000/month** (confirmed), plus up to **~$8,000/month** pending review

Trailing 30 days ending 2026-06-22. Confirmed ~$228K/year; up to ~$96K/year more pending review. Every line below is one config change.

| # | Where it's leaking | $/month | Confidence | The fix |
|---|---|--:|---|---|
| 1 | Clusters that never shut themselves off — paying around the clock for compute nobody is using | **$12,000** | Confirmed | Set auto-shutoff (e.g. 30 min) |
| 2 | Scheduled batch jobs running on the premium notebook tier — ~3.6× the batch rate for identical work | **$7,000** | Confirmed | Move job clusters to the batch tier |
| 3 | Clusters sized for peak, idling most of the time — typically 30–50% oversized | **$5,000** | Estimated | Turn on autoscaling, drop the floor |
| 4 | Paying a ~2× speed-engine premium on jobs that don't run faster | **$3,000** | At-risk | Turn off the speed engine where it adds no gain |

**The #1 line alone — idle clusters (confirmed) — is ~$144K/year, fixed in one setting.**

Example 2: Idle-Cluster Sweep

User asks "find idle clusters wasting money." The skill runs Step 3 only, joins the spend to clusters_get, and reports each auto_termination_minutes = 0 cluster with its 30-day idle spend and the live idle gap from clusters_events.

Example 3: All-Purpose-Job Rightsizing

User asks "are any jobs on the wrong compute?" Step 4 returns each job_id running on All-Purpose with potential_savings_30d_usd, corroborated by clusters_get confirming cluster_source is not JOB — the single fix is "move to Jobs Compute."

Resources

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!