Agent skill · Data & Analytics

hypogenic

Automated LLM-driven hypothesis generation and testing on tabular datasets. Use when you want to systematically explore hypotheses about patterns in empirical data (e.g., deception detection, content analysis). Combines literature insights with data-driven hypothesis testing. For manual hypothesis formulation use hypothesis-generation; for creative ideation use scientific-brainstorming.

LeonChaoXgithub.com/LeonChaoXGitHub ↗
claude-codeMIT
Install
npx skills add LeonChaoX/qinyan-academic-skills --skill hypogenic --agent claude-code

Same command for any agent — swap --agent for codex, cursor, copilot.

Facts
Files in the skill folder: 2
SKILL.md size: 21 KB
Bundled scripts: none
Path: skills/09-机器学习与人工智能/hypogenic/SKILL.md
Open the folder on GitHub →
Where it comes from
Stars: 759
Language: Python

Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.

Review
written from the skill's own SKILL.md · Aug 5, 2026

What it does

Hypogenic provides automated hypothesis generation and testing using large language models to accelerate scientific discovery. It supports three approaches: HypoGeniC (data-driven hypothesis generation), HypoRefine (synergistic literature and data integration), and Union methods (mechanistic combination of literature and data-driven hypotheses).

How it works

  • Quick Start shows commands to install, clone datasets, and run basic hypothesis generation and inference, including a Python API usage example with BaseTask and config.yaml.
  • Core capabilities describe three modes:
    • HypoGeniC: generate hypotheses solely from observational data through iterative refinement (initialize with a data subset, refine based on performance, replace poorly-performing ones).
    • HypoRefine: extract insights from research papers, generate theory-grounded hypotheses, and refine with data-driven hypotheses.
    • Union Methods: combine literature-only hypotheses with framework outputs, with variants Literature ∪ HypoGeniC and Literature ∪ HypoRefine.
  • Installation and optional dependencies list Redis caching, s2orc-doc2json, and GROBID, plus instructions to clone datasets for different modes.
  • CLI usage covers hypogenic_generation and hypogenic_inference with key parameters like config path, method, number of hypotheses, and output locations.
  • Python API usage demonstrates loading a task, setting a custom extract_label function, generating hypotheses, and running inference.
  • Workflow examples illustrate end-to-end data-driven, literature-informed, and union approaches across domains like deception detection and AI-generated content identification.
  • Configuration section describes required elements in config.yaml (dataset paths, prompt templates, and placeholders) and template capabilities (dynamic variables, role-based prompts).
  • Literature processing section explains steps to set up GROBID, place PDFs, and process them for HypoRefine workflows.
  • Troubleshooting provides common issues and remedies, including prompt refinement, data adequacy, label extraction, and PDF processing.

When to use it

Use when you want to generate scientific hypotheses from observational data, test multiple competing hypotheses, integrate literature with data-driven insights, and accelerate research discovery across domains such as deception detection, AI-generated content identification, mental health indicators, and predictive modeling.

What it can touch

  • CLI: hypogenic_generation, hypogenic_inference
  • Python API: BaseTask, extract_label customization
  • Files and datasets: configuration via config.yaml, dataset JSON files (train/val/test), literature PDFs for HypoRefine workflows
  • Optional tools: Redis server, s2orc-doc2json, GROBID (setup and processing scripts)

Caveats

  • License: MIT
  • Requires configuration of task-specific prompts and an extract_label function aligned to dataset labels
  • Literature processing depends on external tools (GROBID) and PDF preprocessing steps
From the SKILL.md

# Hypogenic ## Overview Hypogenic provides automated hypothesis generation and testing using large language models to accelerate scientific discovery. The framework supports three approaches: HypoGeniC (data-driven hypothesis generation), HypoRefine (synergistic literature and data integration), and Union methods (mechanistic combination of literature and data-driven hypotheses). ## Quick Start Get started with Hypogenic in minutes: ```bash # Install the package uv pip install hypogenic # Clone example datasets git clone https://github.com/ChicagoHAI/HypoGeniC-datasets.git ./data # Run basic hypothesis generation hypogenic_generation --config ./data/your_task/config.yaml --method hypogenic --num_hypotheses 20 # Run inference on generated hypotheses hypogenic_inference --config ./data/your_task/config.yaml --hypotheses output/hypotheses.json ``` **Or use Python API:** ```python from hypogenic import BaseTask # Create task with your configuration task = BaseTask(config_path="./data/your_task/config.yaml") # Generate hypotheses task.generate_hypotheses(method="hypogenic", num_hypotheses=20) # Run inference results = task.inference(hypothesis_bank="./output/hypotheses.json") ``` ## Whe

What's inside
Steps it walks through
  1. Overview
  2. Quick Start
  3. When to Use This Skill
  4. Key Features
  5. Core Capabilities
  6. 1. HypoGeniC: Data-Driven Hypothesis Generation
  7. 2. HypoRefine: Literature and Data Integration
  8. 3. Union Methods
  9. Installation
  10. Dataset Format
  11. Configuration
  12. Literature Processing (HypoRefine/Union Methods)
  13. CLI Usage
  14. Hypothesis Generation
Ships with 1 file
  • references/config_template.yaml
Commands it runs
Install the package
uv pip install hypogenic
Clone example datasets
git clone https://github.com/ChicagoHAI/HypoGeniC-datasets.git ./data
Run basic hypothesis generation
hypogenic_generation --config ./data/your_task/config.yaml --method hypogenic --num_hypotheses 20
Run inference on generated hypotheses
hypogenic_inference --config ./data/your_task/config.yaml --hypotheses output/hypotheses.json
For HypoGeniC examples
For HypoRefine/Union examples
More from qinyan-academic-skills
All skills →
About this skill
What does the hypogenic skill do?

Automated LLM-driven hypothesis generation and testing on tabular datasets. Use when you want to systematically explore hypotheses about patterns in empirical data (e.g., deception detection, content analysis). Combines literature insights with data-driven hypothesis testing. For manual hypothesis formulation use hypothesis-generation; for creative ideation use scientific-brainstorming.

How do I install it?

Run `npx skills add LeonChaoX/qinyan-academic-skills --skill hypogenic --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.

Where does this skill come from?

From LeonChaoX/qinyan-academic-skills, a repository with 759 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.

Is a popular skill a good skill?

Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.

Keep going