Agent skill · Code Review & Quality

embedding-strategies

Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains.

majiayu000github.com/majiayu000GitHub ↗
claude-codeMIT
Install
npx skills add majiayu000/claude-skill-registry --skill embedding-strategies-duanbiao2000-obsidiandoc26 --agent claude-code

Same command for any agent — swap --agent for codex, cursor, copilot.

Facts
Files in the skill folder: 2
SKILL.md size: 19 KB
Bundled scripts: none
Path: skills/ai-ml/embedding-strategies-duanbiao2000-obsidiandoc26/SKILL.md
Open the folder on GitHub →
Where it comes from
Stars: 534
Language: HTML

Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.

Review
written from the skill's own SKILL.md · Aug 5, 2026

What it does

Guide to selecting and optimizing embedding models for vector search applications.

How it works

The skill provides structured guidance for when to use embedding strategies (e.g., choosing embedding models for RAG, optimizing chunking, domain-specific embeddings, comparing performance, reducing dimensions, multilingual handling). It includes core concepts such as Embedding Model Comparison (a table of models with their dimensions, max tokens, and best use), and an Embedding Pipeline diagram showing the flow from Document to Vector with optional steps like overlap, size, and normalization. It also offers multiple Templates (1-4) that demonstrate concrete code patterns for using Voyage AI embeddings, OpenAI embeddings (with optional dimensionality reduction), local embeddings with Sentence Transformers, and chunking strategies (token-based, sentence-based, semantic sections, and recursive splitting). Additionally, it defines a Domain-Specific Embedding Pipeline (DomainEmbeddingPipeline) and a CodeEmbeddingPipeline, illustrating how to preprocess, chunk, and embed documents or code, and how to handle code-specific chunking via tree-sitter. Finally, there is Embedding Quality Evaluation logic to measure retrieval metrics.

When to use it

Use when selecting embedding models for RAG, optimizing chunking strategies, fine-tuning embeddings for specific domains, comparing embedding model performance, reducing embedding dimensions, or handling multilingual content.

What it can touch

Templates show concrete code interactions:

  • Voyage AI Embeddings via VoyageAIEmbeddings, including model variations (voyage-3-large, voyage-code-3, voyage-finance-2, voyage-law-2).
  • OpenAI Embeddings via OpenAI client, with optional dimensions parameter for Matryoshka reduction.
  • Local Embeddings via SentenceTransformer (e.g., BAAI/bge-large-en-v1.5, multilingual-e5-large).
  • Chunking utilities: chunk_by_tokens, chunk_by_sentences, chunk_by_semantic_sections, recursive_character_splitter.
  • DomainEmbeddingPipeline and CodeEmbeddingPipeline demonstrating document processing, chunking, and code-specific embeddings.

Caveats

License is MIT. The skill assumes availability of specific libraries (VoyageAIEmbeddings, OpenAI client, sentence_transformers, tree-sitter for code parsing). No guarantees on performance or outcomes are stated beyond the provided templates and guidance. No explicit risk disclosures beyond typical dependencies and environment requirements are listed in the text.

From the SKILL.md

# Embedding Strategies Guide to selecting and optimizing embedding models for vector search applications. ## When to Use This Skill - Choosing embedding models for RAG - Optimizing chunking strategies - Fine-tuning embeddings for domains - Comparing embedding model performance - Reducing embedding dimensions - Handling multilingual content ## Core Concepts ### 1. Embedding Model Comparison (2026) | Model | Dimensions | Max Tokens | Best For | | -------------------------- | ---------- | ---------- | ----------------------------------- | | **voyage-3-large** | 1024 | 32000 | Claude apps (Anthropic recommended) | | **voyage-3** | 1024 | 32000 | Claude apps, cost-effective | | **voyage-code-3** | 1024 | 32000 | Code search | | **voyage-finance-2** | 1024 | 32000 | Financial documents | | **voyage-law-2** | 1024 | 32000 | Legal documents | | **text-embedding-3-large** | 3072 | 8191 | OpenAI apps, high accuracy | | **text-embedding-3-small** | 1536 | 8191 | OpenAI apps, cost-effective | | **bge-large-en-v1.5** | 1024 | 512 | Open source, local deployment | | **all-MiniLM-L6-v2** | 384 | 256 | Fast, lightweight | | **multilingual-e5-large** | 1024 | 512 | Multi-language | ### 2. Embedding

What's inside
Steps it walks through
  1. When to Use This Skill
  2. Core Concepts
  3. 1. Embedding Model Comparison (2026)
  4. 2. Embedding Pipeline
  5. Templates
  6. Template 1: Voyage AI Embeddings (Recommended for Claude)
  7. Template 2: OpenAI Embeddings
  8. Template 3: Local Embeddings with Sentence Transformers
  9. Template 4: Chunking Strategies
  10. Template 5: Domain-Specific Embedding Pipeline
  11. Template 6: Embedding Quality Evaluation
  12. Best Practices
  13. Do's
  14. Don'ts
Ships with 1 file
  • metadata.json
More from claude-skill-registry
All skills →
About this skill
What does the embedding-strategies skill do?

Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains.

How do I install it?

Run `npx skills add majiayu000/claude-skill-registry --skill embedding-strategies-duanbiao2000-obsidiandoc26 --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.

Where does this skill come from?

From majiayu000/claude-skill-registry, a repository with 534 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.

Is a popular skill a good skill?

Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.

Keep going