Agent skill · Testing & QA

splitting-datasets

Split datasets into training, validation, and test partitions with the right stratification and temporal rules. Use as a narrow preprocessing helper once the broader ML workflow is already chosen, not as the main route owner for an end-to-end ML task.

majiayu000github.com/majiayu000GitHub ↗
claude-codecan modify filesMIT
Install
npx skills add majiayu000/claude-skill-registry --skill splitting-datasets-foryourhealth111-pix-vibe-skills --agent claude-code

Same command for any agent — swap --agent for codex, cursor, copilot.

Facts
Files in the skill folder: 2
SKILL.md size: 1 KB
Bundled scripts: none
Version: 1.0.0
Declared author: Jeremy Longshore <jeremy@intentsolutions.io>
Allowed tools: ReadWriteEditGrepGlobBash(cmd:*)
Path: skills/ai-ml/splitting-datasets-foryourhealth111-pix-vibe-skills/SKILL.md
Open the folder on GitHub →
Where it comes from
Stars: 534
Language: HTML

Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.

From the SKILL.md

# Dataset Splitter ## Positioning Treat this skill as a narrow helper for partition strategy. ## When to Use Use this skill when: - Prepare a dataset for machine learning model training. - Create training, validation, and testing sets. - Partition data to evaluate model performance. ## Not For / Boundaries - Full preprocessing-pipeline ownership: use `preprocessing-data-with-automated-pipelines` - Leakage audits and prediction-time checks: use `ml-data-leakage-guard` - Model training and tuning after the split: use `training-machine-learning-models` ## Typical Outputs - Partition strategy with ratios, random seeds, and stratification rules - Notes on temporal or grouped split constraints - Handoff guidance for leakage review and downstream training ## Related Skills - `preprocessing-data-with-automated-pipelines` for the broader preprocessing sequence - `ml-data-leakage-guard` to verify the split does not leak future or test information

What's inside
Steps it walks through
  1. Positioning
  2. When to Use
  3. Not For / Boundaries
  4. Typical Outputs
  5. Related Skills
Ships with 1 file
  • metadata.json
More from claude-skill-registry
All skills →
About this skill
What does the splitting-datasets skill do?

Split datasets into training, validation, and test partitions with the right stratification and temporal rules. Use as a narrow preprocessing helper once the broader ML workflow is already chosen, not as the main route owner for an end-to-end ML task.

How do I install it?

Run `npx skills add majiayu000/claude-skill-registry --skill splitting-datasets-foryourhealth111-pix-vibe-skills --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.

Where does this skill come from?

From majiayu000/claude-skill-registry, a repository with 534 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.

Is a popular skill a good skill?

Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.

Keep going