Adoption

Agent Skills are supported by leading AI development tools.

VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory

ADu2021/adaptive-token-reduction-image-representation

Name: adaptive-token-reduction-image-representation
Author: ADu2021

skills/skillxiv-v0.0.2-claude-opus-4.6/adaptive-token-reduction-image-representation/SKILL.md

npx skillsauth add ADu2021/skillXiv adaptive-token-reduction-image-representation

Clean

TrivyContainer and dependency vulnerability scanner

Clean

SemgrepStatic code analysis for vulnerabilities

Clean

mcp-scan (Snyk)Model Context Protocol security validation

Skipped

Snyk (dep)Open source security scanning

Skipped

Socket.devSupply chain security analysis

Skipped

VirusTotalMulti-engine malware detection

Skipped

CrowdStrikeAdvanced threat intelligence

Skipped

OSV-ScannerOpen Source Vulnerability database check

Skipped

OWASP Dep-Check

Core Concept

This skill implements adaptive token reduction for vision encoders. The key insight is that many visual tokens are redundant—their information can be reconstructed from a smaller subset of more informative tokens. Rather than random pruning, this approach learns which tokens are essential by training a selector network to identify valuable tokens and a reconstructor to verify that discarded tokens can be faithfully recovered.

Architecture Overview

The system has three main components:

Feature Selector (S): Three Transformer layers with a Gumbel-Softmax head that generates binary masks, choosing which tokens to keep or discard
Feature Reconstructor (R): Three Transformer layers that reconstruct discarded tokens from retained ones plus a shared learnable masked embedding
Optimization Objective: Balances reconstruction fidelity against pruning efficiency using modified regularization

The training uses an autoencoder-like framework where the selector learns to identify redundant tokens, and the reconstructor validates that removed tokens are recoverable.

Implementation

The feature selection mechanism uses Gumbel-Softmax for differentiable discrete choices. The following code shows the core selector and reconstructor modules:

import torch
import torch.nn as nn
from torch.nn.functional import gumbel_softmax

class TokenSelector(nn.Module):
    """Selects which tokens to retain using Gumbel-Softmax."""
    def __init__(self, hidden_dim, num_layers=3, num_heads=8):
        super().__init__()
        self.transformer_layers = nn.ModuleList([
            nn.TransformerEncoderLayer(
                d_model=hidden_dim,
                nhead=num_heads,
                dim_feedforward=hidden_dim * 4,
                batch_first=True
            ) for _ in range(num_layers)
        ])
        self.selection_head = nn.Linear(hidden_dim, 1)

    def forward(self, features, temperature=1.0, training=True):
        """
        Args:
            features: (B, N, D) token embeddings
            temperature: Gumbel-Softmax temperature
            training: if True, use stochastic selection
        Returns:
            masks: (B, N) binary masks
            logits: (B, N) selection logits
        """
        x = features
        for layer in self.transformer_layers:
            x = layer(x)

        logits = self.selection_head(x).squeeze(-1)

        if training:
            # Gumbel-Softmax for differentiable discrete selection
            masks = gumbel_softmax(logits.unsqueeze(-1), tau=temperature, hard=True)
            masks = masks.squeeze(-1)
        else:
            masks = (logits > 0).float()

        return masks, logits

The reconstructor mirrors this structure but focuses on recovering discarded tokens:

class TokenReconstructor(nn.Module):
    """Reconstructs removed tokens from retained ones."""
    def __init__(self, hidden_dim, num_layers=3, num_heads=8):
        super().__init__()
        self.transformer_layers = nn.ModuleList([
            nn.TransformerEncoderLayer(
                d_model=hidden_dim,
                nhead=num_heads,
                dim_feedforward=hidden_dim * 4,
                batch_first=True
            ) for _ in range(num_layers)
        ])
        self.masked_embedding = nn.Parameter(torch.randn(1, 1, hidden_dim))
        self.reconstruction_head = nn.Linear(hidden_dim, hidden_dim)

    def forward(self, features, masks):
        """
        Args:
            features: (B, N, D) original token embeddings
            masks: (B, N) binary masks from selector
        Returns:
            reconstructed: (B, N, D) reconstructed features
        """
        # Create input with masked tokens replaced
        masked_input = features * masks.unsqueeze(-1)
        masked_input = masked_input + self.masked_embedding * (1 - masks.unsqueeze(-1))

        x = masked_input
        for layer in self.transformer_layers:
            x = layer(x)

        reconstructed = self.reconstruction_head(x)
        return reconstructed

Training uses L2 reconstruction loss with a pruning efficiency regularizer:

def compute_loss(original_features, reconstructed_features, masks, pruning_weight=0.1):
    """
    Args:
        original_features: (B, N, D) original tokens
        reconstructed_features: (B, N, D) reconstructed tokens
        masks: (B, N) selection masks
        pruning_weight: weight for pruning efficiency term
    """
    batch_size = masks.shape[0]

    # Reconstruction loss for discarded tokens only
    discard_mask = 1 - masks
    reconstruction_loss = torch.sum(
        ((original_features - reconstructed_features) ** 2) * discard_mask.unsqueeze(-1)
    ) / (torch.sum(discard_mask) + 1e-6)

    # Pruning efficiency: encourage removal of tokens
    pruning_ratio = torch.mean(1 - masks)

    # Modified regularization: max(L_pr, p) prevents trivial solutions
    pruning_loss = torch.max(
        pruning_weight * reconstruction_loss,
        torch.tensor(pruning_ratio)
    )

    total_loss = reconstruction_loss + pruning_loss
    return total_loss, reconstruction_loss, pruning_ratio

Practical Guidance

When to Use:

Reducing inference latency in vision-language models (LLaVA, etc.)
Processing high-resolution images where token count becomes a bottleneck
OCR and image understanding tasks that benefit from aggressive compression
Deployment scenarios where model size or memory is constrained

When NOT to Use:

Complex reasoning tasks that require fine-grained spatial details
Tasks where every pixel matters (e.g., small object detection)
Low-resolution inputs where 50% pruning would be too aggressive
When inference speed is not a critical constraint

Key Hyperparameters:

num_layers (3-4): Depth of selector/reconstructor networks; deeper = more capacity
temperature (0.5-2.0): Gumbel-Softmax temperature; lower = sharper, higher = softer
pruning_weight (0.01-0.5): Balance between reconstruction and efficiency; higher = more aggressive pruning
Training dataset size: Paper uses 100,000 COCO images

Common Pitfalls:

Setting pruning_weight too high leads to trivial solutions that remove nearly all tokens
Using token selection during training but hard masking during inference causes distribution mismatch
Not evaluating on task-specific benchmarks; pruning effectiveness varies by task
Forgetting to freeze the vision encoder while training selector/reconstructor

Performance Notes

OCR tasks: Up to 50% token reduction with negligible degradation
Reasoning-heavy tasks: Minimal benefit from aggressive pruning (suggest 20-30%)
Training time: Approximately 24 hours on 100K COCO images with standard GPUs
Inference overhead: Selection adds ~5-10% latency but saves much more downstream

References

Gumbel-Softmax paper for differentiable discrete sampling
Vision Transformer (ViT) and DeiT architectures
LLaVA and LLaVA-NeXT multimodal models
COCO dataset for training selector networks

ADu2021/adaptive-token-reduction-image-representation

skills/skillxiv-v0.0.2-claude-opus-4.6/adaptive-token-reduction-image-representation/SKILL.md

Adaptively prune visual tokens from vision encoders by reconstructing discarded features from retained ones, reducing computational cost by 50% while maintaining task performance on OCR and image understanding tasks.

2 stars

development

Updated Apr 16, 2026

$ install --global

skillsauth

npx skillsauth add ADu2021/skillXiv adaptive-token-reduction-image-representation

Install this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.

Security Scan Results

3 of 9 scanners reported clean

Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.

Scanners Passed

Scanners in report

Clean

TrivyContainer and dependency vulnerability scanner

95%

Clean

SemgrepStatic code analysis for vulnerabilities

95%

Clean

mcp-scan (Snyk)Model Context Protocol security validation

95%

Skipped

Snyk (dep)Open source security scanning

50%

Skipped

Socket.devSupply chain security analysis

50%

Skipped

VirusTotalMulti-engine malware detection

50%

Skipped

CrowdStrikeAdvanced threat intelligence

50%

Skipped

OSV-ScannerOpen Source Vulnerability database check

50%

Skipped

OWASP Dep-Check

50%

Last scanned: Apr 16, 2026, 3:02 PM4.8s1 file scanned

SKILL.md

name:: adaptive-token-reduction-image-representation
title:: When Less is Enough: Adaptive Token Reduction for Efficient Image Representation
version:: 0.0.2
engine:: skillxiv-v0.0.2-claude-opus-4.6
license:: MIT
url:: https://arxiv.org/abs/2503.16660
keywords:: [Token Reduction, Vision Transformers, Feature Selection, Image Compression, Efficiency]
description:: Adaptively prune visual tokens from vision encoders by reconstructing discarded features from retained ones, reducing computational cost by 50% while maintaining task performance on OCR and image understanding tasks.

Core Concept

Architecture Overview

The system has three main components:

Feature Selector (S): Three Transformer layers with a Gumbel-Softmax head that generates binary masks, choosing which tokens to keep or discard
Feature Reconstructor (R): Three Transformer layers that reconstruct discarded tokens from retained ones plus a shared learnable masked embedding
Optimization Objective: Balances reconstruction fidelity against pruning efficiency using modified regularization

The training uses an autoencoder-like framework where the selector learns to identify redundant tokens, and the reconstructor validates that removed tokens are recoverable.

Implementation

The feature selection mechanism uses Gumbel-Softmax for differentiable discrete choices. The following code shows the core selector and reconstructor modules:

import torch
import torch.nn as nn
from torch.nn.functional import gumbel_softmax

class TokenSelector(nn.Module):
    """Selects which tokens to retain using Gumbel-Softmax."""
    def __init__(self, hidden_dim, num_layers=3, num_heads=8):
        super().__init__()
        self.transformer_layers = nn.ModuleList([
            nn.TransformerEncoderLayer(
                d_model=hidden_dim,
                nhead=num_heads,
                dim_feedforward=hidden_dim * 4,
                batch_first=True
            ) for _ in range(num_layers)
        ])
        self.selection_head = nn.Linear(hidden_dim, 1)

    def forward(self, features, temperature=1.0, training=True):
        """
        Args:
            features: (B, N, D) token embeddings
            temperature: Gumbel-Softmax temperature
            training: if True, use stochastic selection
        Returns:
            masks: (B, N) binary masks
            logits: (B, N) selection logits
        """
        x = features
        for layer in self.transformer_layers:
            x = layer(x)

        logits = self.selection_head(x).squeeze(-1)

        if training:
            # Gumbel-Softmax for differentiable discrete selection
            masks = gumbel_softmax(logits.unsqueeze(-1), tau=temperature, hard=True)
            masks = masks.squeeze(-1)
        else:
            masks = (logits > 0).float()

        return masks, logits

The reconstructor mirrors this structure but focuses on recovering discarded tokens:

class TokenReconstructor(nn.Module):
    """Reconstructs removed tokens from retained ones."""
    def __init__(self, hidden_dim, num_layers=3, num_heads=8):
        super().__init__()
        self.transformer_layers = nn.ModuleList([
            nn.TransformerEncoderLayer(
                d_model=hidden_dim,
                nhead=num_heads,
                dim_feedforward=hidden_dim * 4,
                batch_first=True
            ) for _ in range(num_layers)
        ])
        self.masked_embedding = nn.Parameter(torch.randn(1, 1, hidden_dim))
        self.reconstruction_head = nn.Linear(hidden_dim, hidden_dim)

    def forward(self, features, masks):
        """
        Args:
            features: (B, N, D) original token embeddings
            masks: (B, N) binary masks from selector
        Returns:
            reconstructed: (B, N, D) reconstructed features
        """
        # Create input with masked tokens replaced
        masked_input = features * masks.unsqueeze(-1)
        masked_input = masked_input + self.masked_embedding * (1 - masks.unsqueeze(-1))

        x = masked_input
        for layer in self.transformer_layers:
            x = layer(x)

        reconstructed = self.reconstruction_head(x)
        return reconstructed

Training uses L2 reconstruction loss with a pruning efficiency regularizer:

def compute_loss(original_features, reconstructed_features, masks, pruning_weight=0.1):
    """
    Args:
        original_features: (B, N, D) original tokens
        reconstructed_features: (B, N, D) reconstructed tokens
        masks: (B, N) selection masks
        pruning_weight: weight for pruning efficiency term
    """
    batch_size = masks.shape[0]

    # Reconstruction loss for discarded tokens only
    discard_mask = 1 - masks
    reconstruction_loss = torch.sum(
        ((original_features - reconstructed_features) ** 2) * discard_mask.unsqueeze(-1)
    ) / (torch.sum(discard_mask) + 1e-6)

    # Pruning efficiency: encourage removal of tokens
    pruning_ratio = torch.mean(1 - masks)

    # Modified regularization: max(L_pr, p) prevents trivial solutions
    pruning_loss = torch.max(
        pruning_weight * reconstruction_loss,
        torch.tensor(pruning_ratio)
    )

    total_loss = reconstruction_loss + pruning_loss
    return total_loss, reconstruction_loss, pruning_ratio

Practical Guidance

When to Use:

Reducing inference latency in vision-language models (LLaVA, etc.)
Processing high-resolution images where token count becomes a bottleneck
OCR and image understanding tasks that benefit from aggressive compression
Deployment scenarios where model size or memory is constrained

When NOT to Use:

Complex reasoning tasks that require fine-grained spatial details
Tasks where every pixel matters (e.g., small object detection)
Low-resolution inputs where 50% pruning would be too aggressive
When inference speed is not a critical constraint

Key Hyperparameters:

num_layers (3-4): Depth of selector/reconstructor networks; deeper = more capacity
temperature (0.5-2.0): Gumbel-Softmax temperature; lower = sharper, higher = softer
pruning_weight (0.01-0.5): Balance between reconstruction and efficiency; higher = more aggressive pruning
Training dataset size: Paper uses 100,000 COCO images

Common Pitfalls:

Setting pruning_weight too high leads to trivial solutions that remove nearly all tokens
Using token selection during training but hard masking during inference causes distribution mismatch
Not evaluating on task-specific benchmarks; pruning effectiveness varies by task
Forgetting to freeze the vision encoder while training selector/reconstructor

Performance Notes

OCR tasks: Up to 50% token reduction with negligible degradation
Reasoning-heavy tasks: Minimal benefit from aggressive pruning (suggest 20-30%)
Training time: Approximately 24 hours on 100K COCO images with standard GPUs
Inference overhead: Selection adds ~5-10% latency but saves much more downstream

References

Gumbel-Softmax paper for differentiable discrete sampling
Vision Transformer (ViT) and DeiT architectures
LLaVA and LLaVA-NeXT multimodal models
COCO dataset for training selector networks

Related Skills

ADu2021/flow-map-trajectory-tilting

testing

VerifiedTrustedCommunity

Uses flow maps as look-ahead operators to enable principled reward-guided diffusion by predicting trajectory endpoints at any denoising step. Deploy when applying rewards or preferences to diffusion trajectories with meaningful gradients throughout generation.

2SKILL.mdUpdated Apr 17, 2026

ADu2021/flow-map-trajectory-tilting

ADu2021/flexible-data-mixture-of-experts

testing

VerifiedTrustedCommunity

Train language models where each expert learns independently on closed datasets, enabling flexible inference with selective data inclusion or exclusion. 41% performance improvement while allowing users to opt out of specific data sources without retraining.

2SKILL.mdUpdated Apr 17, 2026

ADu2021/flexible-data-mixture-of-experts

ADu2021/flexibility-trap-diffusion-reasoning

data-ai

VerifiedTrustedCommunity

Understand how token generation flexibility in diffusion LMs paradoxically constrains reasoning, as models exploit ordering flexibility to avoid uncertain tokens, and apply simplified approaches that preserve parallel decoding benefits. Use when optimizing diffusion-based language models for reasoning tasks.

2SKILL.mdUpdated Apr 17, 2026

ADu2021/flexibility-trap-diffusion-reasoning

ADu2021/flex-continuous-agent-evolution

devops

VerifiedTrustedCommunity

Enable LLM agents to improve continuously during deployment by constructing structured experience libraries through self-reflection on successes and failures—achieving 23% improvement on reasoning without gradient-based parameter updates or external training.

2SKILL.mdUpdated Apr 17, 2026

ADu2021/flex-continuous-agent-evolution

Download

For Claude Desktop. Download once, then upload the file in the app — no terminal needed.

Need help? View full Cowork setup guide →

Install manually

Choose your platform

# Clone the repo
git clone https://github.com/ADu2021/skillXiv.git

# Copy into Claude Code skills folder (global)
cp -r skillXiv/skills/skillxiv-v0.0.2-claude-opus-4.6/adaptive-token-reduction-image-representation ~/.claude/skills/

Claude Code Skills — official skills path docs.

Repository

ADu2021/skillXiv

2 stars

Compatible with

Claude Code

OpenAI Codex CLI

ChatGPT