Adoption

Agent Skills are supported by leading AI development tools.

VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory

alvarobartt/hf-mem

Name: hf-mem
Author: alvarobartt

skills/hf-mem/SKILL.md

npx skillsauth add alvarobartt/hf-mem hf-mem

Clean

TrivyContainer and dependency vulnerability scanner

Clean

SemgrepStatic code analysis for vulnerabilities

Clean

mcp-scan (Snyk)Model Context Protocol security validation

Skipped

Snyk (dep)Open source security scanning

Skipped

Socket.devSupply chain security analysis

Skipped

VirusTotalMulti-engine malware detection

Skipped

CrowdStrikeAdvanced threat intelligence

Skipped

OSV-ScannerOpen Source Vulnerability database check

Skipped

OWASP Dep-Check

hf-mem

Estimates inference memory (model weights + optional KV cache) for models on the Hugging Face Hub using HTTP Range requests; no weights are downloaded.

When to use

User asks how much VRAM or memory a model needs to run
User wants to know if a model fits on their GPU or a given instance
User references a Hugging Face model ID or URL and asks about inference requirements

Requirements

uv installed (for uvx)
HF_TOKEN env var or --hf-token flag (gated/private models only)

Safetensors models

Auto-detected when the repo contains model.safetensors, model.safetensors.index.json, or model_index.json. Covers Transformers, Diffusers, and Sentence Transformers; no extra flags needed.

uvx hf-mem --model-id <org/model>

GGUF models

Auto-detected when the repo contains only .gguf files. When both Safetensors and GGUF files coexist, pass --gguf-file to target a specific file. Any shard path works for sharded models.

uvx hf-mem --model-id <org/model> --gguf-file <path-in-repo>

KV cache estimation (`--experimental`)

Adds KV cache memory on top of weights. Applies to LLMs (...ForCausalLM), VLMs (...ForConditionalGeneration), and GGUF models. Reads max_model_len from config.json by default; override with --max-model-len. KV cache dtype defaults to auto (reads torch_dtype/dtype from config.json, or the FP8 quantization format if applicable; for GGUF auto = F16).

uvx hf-mem --model-id <org/model> [--gguf-file <path>] \
  --experimental [--max-model-len N] [--batch-size N] \
  [--kv-cache-dtype auto|bfloat16|fp8|fp8_e4m3|fp8_e5m2]

Examples

# Transformers
uvx hf-mem --model-id MiniMaxAI/MiniMax-M2

# Diffusers
uvx hf-mem --model-id Qwen/Qwen-Image

# Sentence Transformers
uvx hf-mem --model-id google/embeddinggemma-300m

# LLM with KV cache
uvx hf-mem --model-id mistralai/Mistral-7B-v0.1 --experimental

# GGUF with KV cache (sharded)
uvx hf-mem --model-id unsloth/Qwen3.5-397B-A17B-GGUF \
  --gguf-file Q4_K_M/Qwen3.5-397B-A17B-Q4_K_M-00001-of-00006.gguf \
  --experimental

Errors

HTTP 401: the model is gated or private; provide HF_TOKEN or --hf-token.
HTTP 404: the model ID not found on the Hub.
RuntimeError: no supported weight format found, or --gguf-file path doesn't match any file in the repository.

alvarobartt/hf-mem

skills/hf-mem/SKILL.md

CLI to estimate the required memory to load either Safetensors or GGUF model weights for inference from the Hugging Face Hub

887 stars

tools

Updated Apr 3, 2026

$ install --global

skillsauth

npx skillsauth add alvarobartt/hf-mem hf-mem

Install this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.

Security Scan Results

3 of 9 scanners reported clean

Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.

Scanners Passed

Scanners in report

Clean

TrivyContainer and dependency vulnerability scanner

95%

Clean

SemgrepStatic code analysis for vulnerabilities

95%

Clean

mcp-scan (Snyk)Model Context Protocol security validation

95%

Skipped

Snyk (dep)Open source security scanning

50%

Skipped

Socket.devSupply chain security analysis

50%

Skipped

VirusTotalMulti-engine malware detection

50%

Skipped

CrowdStrikeAdvanced threat intelligence

50%

Skipped

OSV-ScannerOpen Source Vulnerability database check

50%

Skipped

OWASP Dep-Check

50%

Last scanned: Apr 3, 2026, 4:00 PM4.3s1 file scanned

SKILL.md

name:: hf-mem
description:: CLI to estimate the required memory to load either Safetensors or GGUF model weights for inference from the Hugging Face Hub
license:: mit

hf-mem

Estimates inference memory (model weights + optional KV cache) for models on the Hugging Face Hub using HTTP Range requests; no weights are downloaded.

When to use

User asks how much VRAM or memory a model needs to run
User wants to know if a model fits on their GPU or a given instance
User references a Hugging Face model ID or URL and asks about inference requirements

Requirements

uv installed (for uvx)
HF_TOKEN env var or --hf-token flag (gated/private models only)

Safetensors models

Auto-detected when the repo contains model.safetensors, model.safetensors.index.json, or model_index.json. Covers Transformers, Diffusers, and Sentence Transformers; no extra flags needed.

uvx hf-mem --model-id <org/model>

GGUF models

Auto-detected when the repo contains only .gguf files. When both Safetensors and GGUF files coexist, pass --gguf-file to target a specific file. Any shard path works for sharded models.

uvx hf-mem --model-id <org/model> --gguf-file <path-in-repo>

KV cache estimation (`--experimental`)

uvx hf-mem --model-id <org/model> [--gguf-file <path>] \
  --experimental [--max-model-len N] [--batch-size N] \
  [--kv-cache-dtype auto|bfloat16|fp8|fp8_e4m3|fp8_e5m2]

Examples

# Transformers
uvx hf-mem --model-id MiniMaxAI/MiniMax-M2

# Diffusers
uvx hf-mem --model-id Qwen/Qwen-Image

# Sentence Transformers
uvx hf-mem --model-id google/embeddinggemma-300m

# LLM with KV cache
uvx hf-mem --model-id mistralai/Mistral-7B-v0.1 --experimental

# GGUF with KV cache (sharded)
uvx hf-mem --model-id unsloth/Qwen3.5-397B-A17B-GGUF \
  --gguf-file Q4_K_M/Qwen3.5-397B-A17B-Q4_K_M-00001-of-00006.gguf \
  --experimental

Errors

HTTP 401: the model is gated or private; provide HF_TOKEN or --hf-token.
HTTP 404: the model ID not found on the Hub.
RuntimeError: no supported weight format found, or --gguf-file path doesn't match any file in the repository.

Related Skills

openclaw/taskflow

tools

VerifiedTrustedCommunity

Use when work should span one or more detached tasks but still behave like one job with a single owner context. TaskFlow is the durable flow substrate under authoring layers like Lobster, ACPX, plugins, or plain code. Keep conditional logic in the caller; use TaskFlow for flow identity, child-task linkage, waiting state, revision-checked mutations, and user-facing emergence.

357,764SKILL.mdUpdated Apr 10, 2026

openclaw/extensions/lobster

tools

VerifiedTrustedCommunity

# Lobster Lobster executes multi-step workflows with approval checkpoints. Use it when: - User wants a repeatable automation (triage, monitor, sync) - Actions need human approval before executing (send, post, delete) - Multiple tool calls should run as one deterministic operation ## When to use Lobster | User intent | Use Lobster? | | ------------------------------------------------------ | --------------------------

357,764SKILL.mdUpdated Apr 10, 2026

openclaw/extensions/lobster

steipete/extensions/lobster

tools

VerifiedTrustedCommunity

357,588SKILL.mdUpdated Apr 13, 2026

steipete/extensions/lobster

steipete/xurl

tools

VerifiedTrustedCommunity

A CLI tool for making authenticated requests to the X (Twitter) API. Use this skill when you need to post tweets, reply, quote, search, read posts, manage followers, send DMs, upload media, or interact with any X API v2 endpoint.

356,423SKILL.mdUpdated Apr 13, 2026

Download

For Claude Desktop. Download once, then upload the file in the app — no terminal needed.

Need help? View full Cowork setup guide →

Install manually

Choose your platform

# Clone the repo
git clone https://github.com/alvarobartt/hf-mem.git

# Copy into Claude Code skills folder (global)
cp -r hf-mem/skills/hf-mem ~/.claude/skills/

Claude Code Skills — official skills path docs.

Repository

alvarobartt/hf-mem

887 stars

Compatible with

Claude Code

OpenAI Codex CLI

ChatGPT

Adoption

alvarobartt/hf-mem

$ install --global

Security Scan Results

SKILL.md

hf-mem

When to use

Requirements

Safetensors models

GGUF models

KV cache estimation (--experimental)

Examples

Errors

Related Skills

openclaw/taskflow

openclaw/extensions/lobster

steipete/extensions/lobster

steipete/xurl

alvarobartt/hf-mem

$ install --global

Security Scan Results

SKILL.md

hf-mem

When to use

Requirements

Safetensors models

GGUF models

KV cache estimation (--experimental)

Examples

Errors

Related Skills

openclaw/taskflow

openclaw/extensions/lobster

steipete/extensions/lobster

steipete/xurl

KV cache estimation (`--experimental`)

KV cache estimation (`--experimental`)