bio-sequence-statistics
Calculate sequence statistics (N50, length distribution, GC content, summary reports) using Biopython. Use when analyzing sequence datasets, generating QC reports, or comparing assemblies.
npx skills add majiayu000/claude-skill-registry --skill sequence-statistics-gptomics-bioskills-450d661b --agent claude-code
Same command for any agent — swap --agent for codex, cursor, copilot.
Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.
# Sequence Statistics Calculate comprehensive statistics for sequence datasets using Biopython. ## Required Imports ```python from Bio import SeqIO from Bio.SeqUtils import gc_fraction import statistics ``` ## Basic Statistics ### Sequence Count and Total Length ```python records = list(SeqIO.parse('sequences.fasta', 'fasta')) total_seqs = len(records) total_bp = sum(len(r.seq) for r in records) print(f'Sequences: {total_seqs}') print(f'Total bp: {total_bp:,}') ``` ### Length Statistics ```python lengths = [len(r.seq) for r in SeqIO.parse('sequences.fasta', 'fasta')] print(f'Count: {len(lengths)}') print(f'Total: {sum(lengths):,} bp') print(f'Min: {min(lengths):,} bp') print(f'Max: {max(lengths):,} bp') print(f'Mean: {statistics.mean(lengths):,.1f} bp') print(f'Median: {statistics.median(lengths):,.1f} bp') print(f'Std Dev: {statistics.stdev(lengths):,.1f} bp') ``` ## N50 and Nx Statistics ### Calculate N50 ```python def calculate_n50(lengths): sorted_lengths = sorted(lengths, reverse=True) total = sum(sorted_lengths) cumsum = 0 for length in sorted_lengths: cumsum += length if cumsum >= total * 0.5: return length return 0 lengths = [len(r.seq) for r in SeqIO.parse('assembly.fasta'
- Required Imports
- Basic Statistics
- Sequence Count and Total Length
- Length Statistics
- N50 and Nx Statistics
- Calculate N50
- Calculate Any Nx Value
- L50 (Number of Sequences in N50)
- GC Content Statistics
- Overall GC Content
- Per-Sequence GC Distribution
- GC Content Histogram Data
- Length Distribution
- Length Histogram Data
What does the bio-sequence-statistics skill do?
Calculate sequence statistics (N50, length distribution, GC content, summary reports) using Biopython. Use when analyzing sequence datasets, generating QC reports, or comparing assemblies.
How do I install it?
Run `npx skills add majiayu000/claude-skill-registry --skill sequence-statistics-gptomics-bioskills-450d661b --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.
Where does this skill come from?
From majiayu000/claude-skill-registry, a repository with 534 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.
Is a popular skill a good skill?
Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.
