Agent skill · Data & Analytics

lang-data-and-transparency

Use when preparing the data, annotation, and reproducibility materials for a Language (LSA) manuscript — shared datasets and code, glossed corpora, sound files, and the ethics of working with language consultants and communities. Language values transparent, documented data; over-stating a mandated deposit is as wrong as hiding materials. Documents and shares; it does not run the analysis.

brycew6m4,252★ · +31/wk · 3 repos on radarProfile →
claude-codeMIT
Install
npx skills add brycewang-stanford/Awesome-Journal-Skills --skill lang-data-and-transparency --agent claude-code

Same command for any agent — swap --agent for codex, cursor, copilot.

Facts
Files in the skill folder: 1
SKILL.md size: 6 KB
Bundled scripts: none
Path: Language-Linguistic-Society-Skills/skills/lang-data-and-transparency/SKILL.md
Open the folder on GitHub →
Where it comes from
Stars: 909 · +31 this week
Language: Stata
Read our review of the source →

Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.

From the SKILL.md

# Data & Transparency (lang-data-and-transparency) *Language* increasingly treats **documented, checkable data and reproducible analysis** as a mark of serious work: a reader should be able to see the pattern and, where quantitative, re-run the model. But the requirements differ by subfield and evolve, and linguistic data carry **ethical obligations** to consultants and communities that generic "open data" rhetoric ignores. This skill helps you document, share, and protect your materials appropriately — without over-stating a deposit mandate the journal may not impose. ## When to trigger - Assembling the data/code/annotation to accompany a submission - Deciding what can and cannot be shared (consultant confidentiality, community agreements, licensed corpora) - A reader asked for the dataset, the glossed corpus, the sound files, or the analysis script - Writing a data-availability statement ## What "transparent" means at Language (by data type) ### Quantitative (experiment / corpus) - Share the **analysis-ready data and the script** that reproduces the models, tables, and figures; pin package versions and set seeds. A repository (e.g., OSF) with a readme is the norm. - If the raw co

What's inside
Steps it walks through
  1. When to trigger
  2. What "transparent" means at Language (by data type)
  3. Quantitative (experiment / corpus)
  4. Elicited / fieldwork
  5. Phonetic
  6. Ethics of linguistic data (do not skip)
  7. Calibration (do not over- or under-state, hedged)
  8. Referee/editor conformance check
  9. Anti-patterns
  10. Transparency pass for Language
  11. Output format
  12. Supplementary resources
More from Awesome-Journal-Skills
All skills →
About this skill
What does the lang-data-and-transparency skill do?

Use when preparing the data, annotation, and reproducibility materials for a Language (LSA) manuscript — shared datasets and code, glossed corpora, sound files, and the ethics of working with language consultants and communities. Language values transparent, documented data; over-stating a mandated deposit is as wrong as hiding materials. Documents and shares; it does not run the analysis.

How do I install it?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill lang-data-and-transparency --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.

Where does this skill come from?

From brycewang-stanford/Awesome-Journal-Skills, a repository with 909 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.

Is a popular skill a good skill?

Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.

Keep going