A curated, post-training focused data hub with thousands of datasets across general, math, science, code, multilingual, and agent/function-calling domains. Includes README content listing dataset categories and examples; repo has 4720 stars and 9 open issues.
Collecting history — the radar snapshots this repo daily. The trend line appears after 3 days of data (1 so far).
What it is
A curated list of datasets and tools for post-training of LLMs. The repository aggregates datasets across various domains (General, Math, Science, Code, Multilingual, Instruction following, and Agent & Function calling) with descriptions and links to HuggingFace datasets. It also includes notes about licensing practices generally, as per the README. The project description states: "Curated list of datasets and tools for post-training." The repository statistics show 4720 stars and 393 forks, with 9 open issues, created 2024-04-27 and last push 2026-04-29.
How it works
The README enumerates dataset collections and individual datasets by category, providing dataset names, sizes (#), thinking traces (Yes/No), and notes. It groups datasets into sections like General, Math, Science, Code, Instruction following, Multilingual, and Agent & Function calling, each with a table of datasets and brief descriptions. It does not include executable code snippets or a single unified data-loading workflow; instead it serves as a catalog with links to the source datasets.
Getting started
The README does not include installation or usage commands. There are no explicit setup instructions or a runnable demo in the provided content beyond links to external dataset pages. No code blocks or install commands are present in the truncated README section provided.
Recent releases
RELEASES (latest 0):
- none
Traction
stars_7d or stars_1d are not provided in the FACTS block. The only numeric metrics available are:
- stars: 4720
- forks: 393
- open_issues: 9
- created: 2024-04-27
- last_push: 2026-04-29 These numbers appear in the repository facts and are cited here as raw counts.
Behind the repo
No company or startup link is provided in the FACTS block.
Caveats
License: none listed Created date: 2024-04-27 Last push: 2026-04-29 Open issues: 9 There is no explicit license mentioned in the README excerpt; the datasets linked often mention permissive licenses in their notes, but the repository license field shows "none listed".






