Agent skill · Data & Analytics

dask

Distributed computing for larger-than-RAM pandas/NumPy workflows. Use when you need to scale existing pandas/NumPy code beyond memory or across clusters. Best for parallel file processing, distributed ML, integration with existing pandas code. For out-of-core analytics on single machine use vaex; for in-memory speed use polars.

K-Dense-AIgithub.com/K-Dense-AIGitHub ↗
claude-codecan modify filesMIT
Install
npx skills add K-Dense-AI/scientific-agent-skills --skill dask --agent claude-code

Same command for any agent — swap --agent for codex, cursor, copilot.

Facts
Files in the skill folder: 7
SKILL.md size: 15 KB
Bundled scripts: none
Version: 1.1
Allowed tools: ReadWriteEditBash
Requires: Requires Python 3.10+ and dask 2025.1+. DataFrame workflows need pandas 2+ and PyArrow 16+. Cloud paths (s3://, gcs://)…
Path: skills/dask/SKILL.md
Open the folder on GitHub →
Where it comes from
Stars: 32,619
Language: Python
Read our review of the source →

Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.

From the SKILL.md

# Dask ## Overview Dask is a Python library for parallel and distributed computing that enables three critical capabilities: - **Larger-than-memory execution** on single machines for data exceeding available RAM - **Parallel processing** for improved computational speed across multiple cores - **Distributed computation** supporting terabyte-scale datasets across multiple machines Dask scales from laptops (processing ~100 GiB) to clusters (processing ~100 TiB) while maintaining familiar Python APIs. **Current upstream:** dask **2026.3.0** (PyPI, March 2026). Docs: [docs.dask.org](https://docs.dask.org/en/stable/). Since **2025.1.0**, the expression-based DataFrame API with query planning is the only implementation — do not install `dask-expr` separately or set `dataframe.query-planning: False`. ## Quick Start ### Installation ```bash uv pip install "dask>=2025.1" ``` For a typical pandas/NumPy workflow with the distributed scheduler and dashboard: ```bash uv pip install "dask[complete]" ``` Remote object storage (S3, GCS, Azure): ```bash uv pip install s3fs # s3:// paths uv pip install gcsfs # gs:// paths ``` Requires **Python 3.10+** (3.9 support dropped in 2024.12). DataFrame I/O

What's inside
Steps it walks through
  1. Overview
  2. Quick Start
  3. Installation
  4. When to Use This Skill
  5. Core Capabilities
  6. 1. DataFrames - Parallel Pandas Operations
  7. 2. Arrays - Parallel NumPy Operations
  8. 3. Bags - Parallel Processing of Unstructured Data
  9. 4. Futures - Task-Based Parallelization
  10. 5. Schedulers - Execution Backends
  11. Best Practices
  12. Start with Simpler Solutions
  13. Critical Performance Rules
  14. Common Workflow Patterns
Ships with 6 files
  • references/arrays.md
  • references/bags.md
  • references/best-practices.md
  • references/dataframes.md
  • references/futures.md
  • references/schedulers.md
Commands it runs
uv pip install "dask>=2025.1"
uv pip install "dask[complete]"
uv pip install s3fs    # s3:// paths
uv pip install gcsfs   # gs:// paths
More from scientific-agent-skills
All skills →
About this skill
What does the dask skill do?

Distributed computing for larger-than-RAM pandas/NumPy workflows. Use when you need to scale existing pandas/NumPy code beyond memory or across clusters. Best for parallel file processing, distributed ML, integration with existing pandas code. For out-of-core analytics on single machine use vaex; for in-memory speed use polars.

How do I install it?

Run `npx skills add K-Dense-AI/scientific-agent-skills --skill dask --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.

Where does this skill come from?

From K-Dense-AI/scientific-agent-skills, a repository with 32,619 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.

Is a popular skill a good skill?

Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.

Keep going