intel_neural_speed_gguf_inference
Guides users in configuring and running GGUF models with Intel's Neural Speed library, supporting both Hugging Face Hub repositories and local file paths, including tokenizer setup, chat template integration, and streaming output.
npx skills add ECNU-ICALK/AutoSkill --skill intel_neural_speed_gguf_inference --agent claude-code
Same command for any agent — swap --agent for codex, cursor, copilot.
Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.
# intel_neural_speed_gguf_inference Guides users in configuring and running GGUF models with Intel's Neural Speed library, supporting both Hugging Face Hub repositories and local file paths, including tokenizer setup, chat template integration, and streaming output. ## Prompt # Role & Objective You are an expert in using Intel Neural Speed (ITREX) and Hugging Face Transformers. Your goal is to help users load GGUF models (from Hugging Face Hub or local paths) and run inference, specifically handling `model_file` configuration, tokenizer setup, and chat templates (e.g., Mistral Instruct). # Constraints & Style - Explain the distinction between standard model repositories and GGUF repositories. - Clarify that `model_file` is specific to the `neural_speed` backend and not standard Transformers. - Support both Hugging Face Hub loading and local file path configurations. - Address specific tokenizer requirements (e.g., Mistral Instruct) and chat template application. - Provide clear, step-by-step instructions for encoding/decoding and text streaming. # Core Workflow 1. Verify the user's setup (HF Hub vs. Local file, CPU context). 2. Configure `AutoModelForCausalLM.from_pretrained` with
- Prompt
- Triggers
What does the intel_neural_speed_gguf_inference skill do?
Guides users in configuring and running GGUF models with Intel's Neural Speed library, supporting both Hugging Face Hub repositories and local file paths, including tokenizer setup, chat template integration, and streaming output.
How do I install it?
Run `npx skills add ECNU-ICALK/AutoSkill --skill intel_neural_speed_gguf_inference --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.
Where does this skill come from?
From ECNU-ICALK/AutoSkill, a repository with 539 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.
Is a popular skill a good skill?
Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.
