Agent skill · Code Review & Quality

embedding-strategies

Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains.

majiayu000github.com/majiayu000GitHub ↗
claude-codeMIT
Install
npx skills add majiayu000/claude-skill-registry --skill embedding-strategies-ericgrill-agents-skills-plugin --agent claude-code

Same command for any agent — swap --agent for codex, cursor, copilot.

Facts
Files in the skill folder: 2
SKILL.md size: 19 KB
Bundled scripts: none
Path: skills/ai-ml/embedding-strategies-ericgrill-agents-skills-plugin/SKILL.md
Open the folder on GitHub →
Where it comes from
Stars: 534
Language: HTML

Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.

Review
written from the skill's own SKILL.md · Aug 5, 2026

What it does

Guides selecting and optimizing embedding models for vector search and RAG applications, covering model comparison, chunking strategies, and domain-specific embeddings. Provides concrete templates and pipelines to implement embeddings across Claude-compatible code, OpenAI, and local models.

How it works

  • Describes embedding model comparison across several named models with dimensions, max tokens, and best-use notes.
  • Specifies an embedding pipeline flow: Document → Chunking → Preprocessing → Embedding Model → Vector, with options for overlap and normalization.
  • Provides multiple Template blocks detailing code snippets for using Voyage AI embeddings, OpenAI embeddings (including optional dimension reduction and batch handling), and local embeddings with Sentence Transformers.
  • Includes chunking strategies: token-based, sentence-based, semantic-sectioning, and recursive character splitting, each with concrete functions and parameters.
  • Defines a DomainEmbeddingPipeline that preprocesses text, chunks, creates embeddings, and stores metadata for documents; also includes a CodeEmbeddingPipeline for code-specific chunking via tree-sitter and code-context embedding.
  • Lists a Template 4 for chunking and a Template 5 for domain-specific embedding workflows, culminating in an EmbeddedDocument data structure used to store embeddings with metadata.

When to use it

Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains. Also recommended for comparing model performance, reducing embedding dimensions, and handling multilingual content.

What it can touch

  • Tools: claude-code
  • Code samples reference: VoyageAIEmbeddings, OpenAI Embeddings, SentenceTransformer, tiktoken, nltk, tree_sitter (via tree_sitter_languages), and specific model names such as "voyage-3-large", "voyage-code-3", "text-embedding-3-small", and "BAAI/bge-large-en-v1.5".

Caveats

License: MIT. The skill provides code templates and model suggestions but does not guarantee retrieval quality or domain-specific success; users should validate embeddings in their own data and pipelines.

From the SKILL.md

# Embedding Strategies Guide to selecting and optimizing embedding models for vector search applications. ## When to Use This Skill - Choosing embedding models for RAG - Optimizing chunking strategies - Fine-tuning embeddings for domains - Comparing embedding model performance - Reducing embedding dimensions - Handling multilingual content ## Core Concepts ### 1. Embedding Model Comparison (2026) | Model | Dimensions | Max Tokens | Best For | | -------------------------- | ---------- | ---------- | ----------------------------------- | | **voyage-3-large** | 1024 | 32000 | Claude apps (Anthropic recommended) | | **voyage-3** | 1024 | 32000 | Claude apps, cost-effective | | **voyage-code-3** | 1024 | 32000 | Code search | | **voyage-finance-2** | 1024 | 32000 | Financial documents | | **voyage-law-2** | 1024 | 32000 | Legal documents | | **text-embedding-3-large** | 3072 | 8191 | OpenAI apps, high accuracy | | **text-embedding-3-small** | 1536 | 8191 | OpenAI apps, cost-effective | | **bge-large-en-v1.5** | 1024 | 512 | Open source, local deployment | | **all-MiniLM-L6-v2** | 384 | 256 | Fast, lightweight | | **multilingual-e5-large** | 1024 | 512 | Multi-language | ### 2. Embedding

What's inside
Steps it walks through
  1. When to Use This Skill
  2. Core Concepts
  3. 1. Embedding Model Comparison (2026)
  4. 2. Embedding Pipeline
  5. Templates
  6. Template 1: Voyage AI Embeddings (Recommended for Claude)
  7. Template 2: OpenAI Embeddings
  8. Template 3: Local Embeddings with Sentence Transformers
  9. Template 4: Chunking Strategies
  10. Template 5: Domain-Specific Embedding Pipeline
  11. Template 6: Embedding Quality Evaluation
  12. Best Practices
  13. Do's
  14. Don'ts
Ships with 1 file
  • metadata.json
More from claude-skill-registry
All skills →
About this skill
What does the embedding-strategies skill do?

Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains.

How do I install it?

Run `npx skills add majiayu000/claude-skill-registry --skill embedding-strategies-ericgrill-agents-skills-plugin --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.

Where does this skill come from?

From majiayu000/claude-skill-registry, a repository with 534 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.

Is a popular skill a good skill?

Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.

Keep going