hypogenic
Automated LLM-driven hypothesis generation and testing on tabular datasets. Use when you want to systematically explore hypotheses about patterns in empirical data (e.g., deception detection, content analysis). Combines literature insights with data-driven hypothesis testing. For manual hypothesis formulation use hypothesis-generation; for creative ideation use scientific-brainstorming.
npx skills add LeonChaoX/qinyan-academic-skills --skill hypogenic --agent claude-code
Same command for any agent — swap --agent for codex, cursor, copilot.
Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.
What it does
Hypogenic provides automated hypothesis generation and testing using large language models to accelerate scientific discovery. It supports three approaches: HypoGeniC (data-driven hypothesis generation), HypoRefine (synergistic literature and data integration), and Union methods (mechanistic combination of literature and data-driven hypotheses).
How it works
- Quick Start shows commands to install, clone datasets, and run basic hypothesis generation and inference, including a Python API usage example with BaseTask and config.yaml.
- Core capabilities describe three modes:
- HypoGeniC: generate hypotheses solely from observational data through iterative refinement (initialize with a data subset, refine based on performance, replace poorly-performing ones).
- HypoRefine: extract insights from research papers, generate theory-grounded hypotheses, and refine with data-driven hypotheses.
- Union Methods: combine literature-only hypotheses with framework outputs, with variants Literature ∪ HypoGeniC and Literature ∪ HypoRefine.
- Installation and optional dependencies list Redis caching, s2orc-doc2json, and GROBID, plus instructions to clone datasets for different modes.
- CLI usage covers hypogenic_generation and hypogenic_inference with key parameters like config path, method, number of hypotheses, and output locations.
- Python API usage demonstrates loading a task, setting a custom extract_label function, generating hypotheses, and running inference.
- Workflow examples illustrate end-to-end data-driven, literature-informed, and union approaches across domains like deception detection and AI-generated content identification.
- Configuration section describes required elements in config.yaml (dataset paths, prompt templates, and placeholders) and template capabilities (dynamic variables, role-based prompts).
- Literature processing section explains steps to set up GROBID, place PDFs, and process them for HypoRefine workflows.
- Troubleshooting provides common issues and remedies, including prompt refinement, data adequacy, label extraction, and PDF processing.
When to use it
Use when you want to generate scientific hypotheses from observational data, test multiple competing hypotheses, integrate literature with data-driven insights, and accelerate research discovery across domains such as deception detection, AI-generated content identification, mental health indicators, and predictive modeling.
What it can touch
- CLI: hypogenic_generation, hypogenic_inference
- Python API: BaseTask, extract_label customization
- Files and datasets: configuration via config.yaml, dataset JSON files (train/val/test), literature PDFs for HypoRefine workflows
- Optional tools: Redis server, s2orc-doc2json, GROBID (setup and processing scripts)
Caveats
- License: MIT
- Requires configuration of task-specific prompts and an extract_label function aligned to dataset labels
- Literature processing depends on external tools (GROBID) and PDF preprocessing steps
# Hypogenic ## Overview Hypogenic provides automated hypothesis generation and testing using large language models to accelerate scientific discovery. The framework supports three approaches: HypoGeniC (data-driven hypothesis generation), HypoRefine (synergistic literature and data integration), and Union methods (mechanistic combination of literature and data-driven hypotheses). ## Quick Start Get started with Hypogenic in minutes: ```bash # Install the package uv pip install hypogenic # Clone example datasets git clone https://github.com/ChicagoHAI/HypoGeniC-datasets.git ./data # Run basic hypothesis generation hypogenic_generation --config ./data/your_task/config.yaml --method hypogenic --num_hypotheses 20 # Run inference on generated hypotheses hypogenic_inference --config ./data/your_task/config.yaml --hypotheses output/hypotheses.json ``` **Or use Python API:** ```python from hypogenic import BaseTask # Create task with your configuration task = BaseTask(config_path="./data/your_task/config.yaml") # Generate hypotheses task.generate_hypotheses(method="hypogenic", num_hypotheses=20) # Run inference results = task.inference(hypothesis_bank="./output/hypotheses.json") ``` ## Whe
- Overview
- Quick Start
- When to Use This Skill
- Key Features
- Core Capabilities
- 1. HypoGeniC: Data-Driven Hypothesis Generation
- 2. HypoRefine: Literature and Data Integration
- 3. Union Methods
- Installation
- Dataset Format
- Configuration
- Literature Processing (HypoRefine/Union Methods)
- CLI Usage
- Hypothesis Generation
Install the package uv pip install hypogenic Clone example datasets git clone https://github.com/ChicagoHAI/HypoGeniC-datasets.git ./data Run basic hypothesis generation hypogenic_generation --config ./data/your_task/config.yaml --method hypogenic --num_hypotheses 20 Run inference on generated hypotheses hypogenic_inference --config ./data/your_task/config.yaml --hypotheses output/hypotheses.json For HypoGeniC examples For HypoRefine/Union examples
What does the hypogenic skill do?
Automated LLM-driven hypothesis generation and testing on tabular datasets. Use when you want to systematically explore hypotheses about patterns in empirical data (e.g., deception detection, content analysis). Combines literature insights with data-driven hypothesis testing. For manual hypothesis formulation use hypothesis-generation; for creative ideation use scientific-brainstorming.
How do I install it?
Run `npx skills add LeonChaoX/qinyan-academic-skills --skill hypogenic --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.
Where does this skill come from?
From LeonChaoX/qinyan-academic-skills, a repository with 759 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.
Is a popular skill a good skill?
Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.
