corpus-discovery-dialogue
Guide Socratic discovery of research questions and analytical approaches for text corpus analysis
npx skills add majiayu000/claude-skill-registry --skill corpus-discovery-dialogue --agent claude-code
Same command for any agent — swap --agent for codex, cursor, copilot.
Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.
What it does
Guides users through a multi-phase, Socratic dialogue to turn vague interest in a text corpus into concrete research questions and an analytical roadmap. It structures inquiry into Phase 0 (context and sample review), Phase 1 (corpus understanding with targeted questions), Phase 2 (interest exploration to surface tacit goals and hypotheses), and Phase 3 (formulation of candidate research questions). It prompts for sample texts, corpus characteristics (type, size, time span, authorship, format, metadata), and data readiness, then syntheses outputs like a Corpus Profile and an Interest Map before producing Candidate Research Questions. It relies on structured prompts and scripted blocks to elicit specifics, wait for user responses at each step, and require user confirmation before advancing. The process emphasizes phases, checkpoints, and explicit grounding in actual texts, with announcements and phase descriptions included for clarity. The skill is designed to be invoked when starting corpus analysis with unclear research questions or when structuring exploratory research, and it asks for optional inputs such as corpus_file_path and domain_context.
How it works
- Initiates with a start-announcement describing five phases and expected duration (20-40 minutes).
- Phase 0: Context Establishment & Sample Review:
- Requests 3-5 sample texts if not provided.
- Reviews samples to extract initial observations across length, style, content type, terminology, temporal markers, and structure.
- Presents initial observations in a structured template and waits for user validation before Phase 1.
- Phase 1: Corpus Understanding (5-10 minutes):
- Asks a sequence of targeted questions about basic characteristics, corpus size, temporal span, authorship, format, and data metadata.
- Each sub-question has defined answer options (A–F, etc.) and prompts for confirmation before proceeding.
- Outputs a structured "Corpus Profile" synthesis after Phase 1 concludes with user confirmation.
- Phase 2: Interest Exploration (5-10 minutes):
- Probes motivation, puzzles, surprises, disappointment, and stakes through open-ended prompts.
- Generates an "Interest Map" capturing primary draws, puzzles, surprises, disappointments, stakes, and implicit hypotheses.
- Phase 3: Question Formulation (5-10 minutes):
- Prompts to present 8-12 candidate research questions organized by type, based on the Corpus Profile and Interest Map.
- Output: Each phase yields structured sections (e.g., Corpus Profile, Interest Map) and culminates in Candidate Research Questions tailored to corpus size, temporal span, authorship, and interests.
When to use it
- Invocation phrases include: "Help me explore this corpus", "What questions should I ask of these [texts/reviews/documents]?", "Guide me through corpus analysis", "I have text data but don't know where to start", "Design a research plan for this corpus".
- Suitable for starting corpus analysis with unclear questions or needing a structured exploratory framework.
- Skip Phase 1–3 when research questions are already clearly defined or when using other specialized dialogues (e.g., hypothesis-testing-dialogue, analysis-interpretation-dialogue).
What it can touch
- Requires: Optional inputs corpus_file_path, domain_context, research_constraints.
- Data touches include: corpus samples (3-5 examples), metadata fields (dates, categories, author info), and format (JSON, CSV, TXT, PDF, DOCX, HTML).
- It prompts for and relies on user-provided corpus descriptions, sample texts, and metadata availability to tailor profile and questions.
Caveats
- Based on user-provided responses; outputs depend on accuracy and completeness of inputs.
- The process emphasizes discovery before execution and requires sequential confirmation to proceed between phases.
- No outcomes are promised beyond formulation of questions and a roadmap; actual analysis results depend on later steps and data readiness.
# Corpus Discovery Dialogue ## Purpose Transform vague interest in a text corpus into concrete, answerable research questions with clear analytical roadmaps. Prevents aimless exploration and "fishing expedition" research. **Core innovation:** Uses Socratic questioning to help researchers articulate what they don't yet consciously know they want to discover. **Core principle:** Discovery before execution. Research questions emerge through dialogue, not pre-formed declarations. The researcher makes all interpretive decisions; this skill guides the process. ## When to Use **Invocation Triggers:** - "Help me explore this corpus" - "What questions should I ask of these [texts/reviews/documents]?" - "Guide me through corpus analysis" - "I have text data but don't know where to start" - "Design a research plan for this corpus" **Use This When:** - Starting corpus analysis with unclear research questions - Have data but uncertain what questions it can answer - Need structure for exploratory research - Want to formulate hypothesis before diving into analysis **Skip This When:** - Research questions already clearly defined - Conducting hypothesis testing (use hypothesis-testing-dialogue inst
- Purpose
- When to Use
- Input Requirements
- Announce at Start
- The Process
- Phase 0: Context Establishment & Sample Review
- Phase 1: Corpus Understanding (5-10 minutes)
- 1A. Basic Characteristics
- 1B. Corpus Size
- 1C. Temporal Span
- 1D. Authorship
- 1E. Structural Format
- 1F. Data Access Confirmation
- Phase 2: Interest Exploration (5-10 minutes)
What does the corpus-discovery-dialogue skill do?
Guide Socratic discovery of research questions and analytical approaches for text corpus analysis
How do I install it?
Run `npx skills add majiayu000/claude-skill-registry --skill corpus-discovery-dialogue --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.
Where does this skill come from?
From majiayu000/claude-skill-registry, a repository with 534 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.
Is a popular skill a good skill?
Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.
