hypogenic
Automated LLM-driven hypothesis generation and testing on tabular datasets. Use when you want to systematically explore hypotheses about patterns in empirical data (e.g., deception detection, content analysis). Combines literature insights with data-driven hypothesis testing. For manual hypothesis formulation use hypothesis-generation; for creative ideation use scientific-brainstorming.
npx skills add BioTender-max/awesome-bio-agent-skills --skill hypogenic --agent claude-code
Same command for any agent — swap --agent for codex, cursor, copilot.
Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.
What it does
Automated hypothesis generation and testing using large language models on tabular datasets, combining data-driven hypotheses with literature-backed insights. It supports three approaches: HypoGeniC (data-driven generation), HypoRefine (literature and data integration), and Union methods (combining literature and data-driven hypotheses).
How it works
The skill provides a CLI and Python API to generate hypotheses and run inference:
- Generate hypotheses via commands like
hypogenic_generation --config <config.yaml> --method <hypogenic|hyporefine|union> --num_hypotheses <n>and then infer withhypogenic_inference --config <config.yaml> --hypotheses <path>. - Or use Python API: instantiate a BaseTask with a config and optional extract_label function, call
generate_hypotheses(method, num_hypotheses, output_path), theninference(hypothesis_bank, test_data). - Supports loading datasets in HuggingFace style, requiring keys like
text_features_1andlabel, and templates for observations, generation, and inference withinconfig.yaml. - Optional components: Redis caching, parallel processing, and adaptive refinement to improve hypothesis quality across iterations.
When to use it
Use when you need to generate and test multiple hypotheses from observational data and literature, especially for domains like deception detection, AI-generated content identification, mental health indicators, and predictive modeling where hypothesis-driven analysis is beneficial.
What it can touch
Commands and configuration touch: local files for datasets, config.yaml, output hypothesis banks, and optional literature artifacts. The skill references tools like OpenAI and Anthropic via API and local LLMs, and relies on dependencies such as Redis for caching and GROBID for literature processing when using HypoRefine workflows.
Caveats
License in the skill metadata is MIT license; the description notes optional literature processing steps (PDFs via GROBID) and requires setup steps for literature processing workflows. It emphasizes custom extract_label() functions to parse LLM outputs and ensure labels align with dataset formats.
# Hypogenic ## Overview Hypogenic provides automated hypothesis generation and testing using large language models to accelerate scientific discovery. The framework supports three approaches: HypoGeniC (data-driven hypothesis generation), HypoRefine (synergistic literature and data integration), and Union methods (mechanistic combination of literature and data-driven hypotheses). ## Quick Start Get started with Hypogenic in minutes: ```bash # Install the package uv pip install hypogenic # Clone example datasets git clone https://github.com/ChicagoHAI/HypoGeniC-datasets.git ./data # Run basic hypothesis generation hypogenic_generation --config ./data/your_task/config.yaml --method hypogenic --num_hypotheses 20 # Run inference on generated hypotheses hypogenic_inference --config ./data/your_task/config.yaml --hypotheses output/hypotheses.json ``` **Or use Python API:** ```python from hypogenic import BaseTask # Create task with your configuration task = BaseTask(config_path="./data/your_task/config.yaml") # Generate hypotheses task.generate_hypotheses(method="hypogenic", num_hypotheses=20) # Run inference results = task.inference(hypothesis_bank="./output/hypotheses.json") ``` ## Whe
- Overview
- Quick Start
- When to Use This Skill
- Key Features
- Core Capabilities
- 1. HypoGeniC: Data-Driven Hypothesis Generation
- 2. HypoRefine: Literature and Data Integration
- 3. Union Methods
- Installation
- Dataset Format
- Configuration
- Literature Processing (HypoRefine/Union Methods)
- CLI Usage
- Hypothesis Generation
Install the package uv pip install hypogenic Clone example datasets git clone https://github.com/ChicagoHAI/HypoGeniC-datasets.git ./data Run basic hypothesis generation hypogenic_generation --config ./data/your_task/config.yaml --method hypogenic --num_hypotheses 20 Run inference on generated hypotheses hypogenic_inference --config ./data/your_task/config.yaml --hypotheses output/hypotheses.json For HypoGeniC examples For HypoRefine/Union examples
What does the hypogenic skill do?
Automated LLM-driven hypothesis generation and testing on tabular datasets. Use when you want to systematically explore hypotheses about patterns in empirical data (e.g., deception detection, content analysis). Combines literature insights with data-driven hypothesis testing. For manual hypothesis formulation use hypothesis-generation; for creative ideation use scientific-brainstorming.
How do I install it?
Run `npx skills add BioTender-max/awesome-bio-agent-skills --skill hypogenic --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.
Where does this skill come from?
From BioTender-max/awesome-bio-agent-skills, a repository with 135 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.
Is a popular skill a good skill?
Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.
