Agent skill · Data & Analytics

data-engineer

Build scalable data pipelines, modern data warehouses, and real-time streaming architectures. Implements Apache Spark, dbt, Airflow, and cloud-native data platforms.

Nick44,086★ · +407/wk · 1 repos on radarProfile →
claude-codecodexcursorMIT
Install
npx skills add sickn33/agentic-awesome-skills --skill data-engineer --agent claude-code

Same command for any agent — swap --agent for codex, cursor, copilot.

Facts
Files in the skill folder: 1
SKILL.md size: 11 KB
Bundled scripts: none
Path: skills/data-engineer/SKILL.md
Open the folder on GitHub →
Where it comes from
Stars: 44,414 · +328 this week
Language: Python
Read our review of the source →

Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.

From the SKILL.md

You are a data engineer specializing in scalable data pipelines, modern data architecture, and analytics infrastructure. ## Use this skill when - Designing batch or streaming data pipelines - Building data warehouses or lakehouse architectures - Implementing data quality, lineage, or governance ## Do not use this skill when - You only need exploratory data analysis - You are doing ML model development without pipelines - You cannot access data sources or storage systems ## Instructions 1. Define sources, SLAs, and data contracts. 2. Choose architecture, storage, and orchestration tools. 3. Implement ingestion, transformation, and validation. 4. Monitor quality, costs, and operational reliability. ## Safety - Protect PII and enforce least-privilege access. - Validate data before writing to production sinks. ## Purpose Expert data engineer specializing in building robust, scalable data pipelines and modern data platforms. Masters the complete modern data stack including batch and streaming processing, data warehousing, lakehouse architectures, and cloud-native data services. Focuses on reliable, performant, and cost-effective data solutions. ## Capabilities ### Modern Data Stack & Ar

What's inside
Steps it walks through
  1. Use this skill when
  2. Do not use this skill when
  3. Instructions
  4. Safety
  5. Purpose
  6. Capabilities
  7. Modern Data Stack & Architecture
  8. Batch Processing & ETL/ELT
  9. Real-Time Streaming & Event Processing
  10. Workflow Orchestration & Pipeline Management
  11. Data Modeling & Warehousing
  12. Cloud Data Platforms & Services
  13. Data Quality & Governance
  14. Performance Optimization & Scaling
More from agentic-awesome-skills
All skills →
About this skill
What does the data-engineer skill do?

Build scalable data pipelines, modern data warehouses, and real-time streaming architectures. Implements Apache Spark, dbt, Airflow, and cloud-native data platforms.

How do I install it?

Run `npx skills add sickn33/agentic-awesome-skills --skill data-engineer --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.

Where does this skill come from?

From sickn33/agentic-awesome-skills, a repository with 44,414 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.

Is a popular skill a good skill?

Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.

Keep going