skills/ollama-optimizer/SKILL.md
Optimize Ollama configuration for the current machine's hardware. Use when asked to speed up Ollama, tune local LLM performance, or pick models that fit available GPU/RAM. Don't use for LM Studio, llama.cpp, vLLM, or hosted-API LLM providers.
npx skillsauth add luongnv89/skills ollama-optimizerInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Optimize Ollama configuration based on system hardware analysis.
Use this skill when the user asks to optimize Ollama, configure Ollama, speed up Ollama, fix Ollama running slow, set up a local LLM, tune inference speed, reduce memory usage, or select models that fit their GPU/RAM. The skill analyzes hardware (GPU, VRAM, RAM, CPU) and produces tailored recommendations.
Do not use for LM Studio, llama.cpp, vLLM, or hosted-API LLM providers (OpenAI, Anthropic) — those use different runtimes and tuning surfaces.
Fast path (opt-in only): only skip full hardware analysis if the user explicitly asks to. Otherwise always run Phases 1-4 and follow the tier-based recommendation — do not apply shortcuts by default, and do not let them override a tier decision already made. For the per-platform shortcut commands and env vars, see Platform-Specific Setup and Environment Variables.
Run the detection script to gather hardware information:
python3 scripts/detect_system.py
Parse the JSON output to identify:
hardware_tier — the script's computed category, max_model_size, and recommended_quantUse hardware_tier from Phase 1 as the tier decision. Do not re-derive it; the table below explains what each tier means and which optimizations it implies. Override the script only with an explicit reason (e.g. VRAM shared with a display), and state that reason in the report.
Hardware Tier Classification:
| Tier (category) | Script band | Max Model | Key Optimizations |
|------|----------|-----------|-------------------|
| cpu_only | No GPU detected | 3B | num_thread tuning, Q4_K_M quant |
| low_vram | <6GB VRAM | 3B | Flash attention, KV cache q4_0 |
| entry | 6-10GB VRAM | 8B | Flash attention, KV cache q8_0 |
| prosumer | 10-16GB VRAM | 14B | Flash attention, full offload |
| workstation | 16-48GB VRAM | 32B | Standard config, Q5_K_M option |
| high_end | 48GB+ VRAM | 70B+ | Multiple models, Q5/Q6 quants |
Apple Silicon Special Case:
entryprosumerworkstation; 64GB+ Mac → high_endCreate a structured optimization guide with these sections:
Present detected hardware specs and highlight constraints (e.g., "8GB unified memory limits to 8B models").
List what's needed based on the platform:
Essential environment variables:
# Always recommended
export OLLAMA_FLASH_ATTENTION=1
# Memory-constrained systems (<12GB)
export OLLAMA_KV_CACHE_TYPE=q8_0 # or q4_0 for severe constraints
Model selection guidance:
ollama list outputModelfile tuning (when needed):
PARAMETER num_gpu <layers> # Partial offload for limited VRAM
PARAMETER num_thread <cores> # CPU threads (physical cores, not hyperthreads)
PARAMETER num_ctx <size> # Reduce context for memory savings
Provide copy-paste commands in order:
$SHELL decides: ~/.zshrc, ~/.bashrc, or ~/.bash_profile) and append the env vars:
RC=~/.zshrc # or ~/.bashrc / ~/.bash_profile, matching $SHELL
cp "$RC" "$RC.ollama-bak"
printf '\n# ollama-optimizer start\nexport OLLAMA_FLASH_ATTENTION=1\n<KV cache + other export lines from section 3, per tier>\n# ollama-optimizer end\n' >> "$RC"
ollama run <model> --verbosecp ~/.zshrc.ollama-bak ~/.zshrc — then restart Ollama.# Benchmark current performance
python3 scripts/benchmark_ollama.py --model <model>
# Expected output: tokens/s and generation latency — record as the post-tuning baseline.
# Check GPU memory usage (NVIDIA)
nvidia-smi
# Verify config is applied
ollama run <model> "test" --verbose 2>&1 | head -20
A run passes when all of the following are true:
OLLAMA_FLASH_ATTENTION, KV-cache quantisation) are written to a shell init file the user actually uses, with a backup of the prior file.ollama run <model> with --verbose and captures the actual offload/cache numbers.After completing each major step, output a status report in this format:
◆ [Step Name] ([step N of M] — [context])
··································································
[Check 1]: √ pass
[Check 2]: √ pass (note if relevant)
[Check 3]: × fail — [reason]
[Check 4]: √ pass
[Criteria]: √ N/M met
____________________________
Result: PASS | FAIL | PARTIAL
Adapt the check names to match what the step actually validates. Use √ for pass, × for fail, and — to add brief context. The "Criteria" line summarizes how many acceptance criteria were met. The "Result" line gives the overall verdict.
◆ Detection (step 1 of 4 — hardware profiling)
··································································
Hardware detected: √ pass — macOS 14, Apple M2
GPU identified: √ pass — Apple Metal (unified memory)
RAM measured: √ pass — 16GB unified memory
[Criteria]: √ 3/3 met
____________________________
Result: PASS
◆ Analysis (step 2 of 4 — profile selection)
··································································
Tier classified: √ pass — Prosumer (16GB unified)
Profile selected: √ pass — Flash attention, full offload
Bottlenecks identified: √ pass — memory bandwidth primary constraint
[Criteria]: √ 3/3 met
____________________________
Result: PASS
◆ Plan (step 3 of 4 — optimization guide)
··································································
Guide generated: √ pass — ollama-optimization-guide.md written
Parameters tuned: √ pass — OLLAMA_FLASH_ATTENTION=1, KV_CACHE_TYPE=q8_0
Model recommendations ready: √ pass — llama3.1:14b-instruct-q4_K_M suggested
[Criteria]: √ 3/3 met
____________________________
Result: PASS
◆ Verification (step 4 of 4 — config validation)
··································································
Benchmark commands listed: √ pass — python3 scripts/benchmark_ollama.py
Config verified: √ pass — ollama run --verbose output checked
[Criteria]: √ 2/2 met
____________________________
Result: PASS
Generate an ollama-optimization-guide.md file. Ask the user where to save it (suggest ~/.config/ollama/optimization-guide.md or current directory). Contents:
# Ollama Optimization Guide
**Generated:** <timestamp>
**System:** <OS> | <CPU> | <RAM>GB RAM | <GPU>
## System Overview
<hardware summary and constraints>
## Current Configuration
<existing Ollama setup and env vars>
## Recommendations
### Environment Variables
<shell commands to set vars>
### Model Selection
<recommended models with rationale>
### Performance Tuning
<Modelfile adjustments if needed>
## Execution Checklist
- [ ] <step 1>
- [ ] <step 2>
...
## Verification
<benchmark commands and expected results>
## Rollback
<commands to revert changes if needed>
development
Scan a live site with isitagentready.com, then approve each step: triage the 0-5 agent-readiness score, write agent-ready-plan.md, file issues via /plan-to-issues. Don't use for applying llms.txt/SEO fixes (seo-ai-optimizer) or app-store ASO.
development
Review a product codebase and landing page against 32 viral principles and produce a Virality Score plus ranked fixes. Use to audit virality or prioritize growth. Don't use for SEO, ASO, copywriting, or code review.
development
Generate a Technical Architecture Document (TAD) from a PRD. Use when asked to design system architecture or define how a product is built. Updates tad.md and reports GitHub links. Don't use for PRD authoring, sprint tasks, or code implementation.
development
Check product and brand names for conflicts across trademarks, domains, social handles, and package registries. Returns a risk level and Proceed/Modify/Abandon recommendation. Skip for name brainstorming, logo design, or trademark filings.