Adoption

Agent Skills are supported by leading AI development tools.

VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory

wentorai/computer-vision-guide

Name: computer-vision-guide
Author: wentorai

skills/domains/ai-ml/computer-vision-guide/SKILL.md

npx skillsauth add wentorai/research-plugins computer-vision-guide

Clean

TrivyContainer and dependency vulnerability scanner

Clean

SemgrepStatic code analysis for vulnerabilities

Clean

mcp-scan (Snyk)Model Context Protocol security validation

Skipped

Snyk (dep)Open source security scanning

Skipped

Socket.devSupply chain security analysis

Skipped

VirusTotalMulti-engine malware detection

Skipped

CrowdStrikeAdvanced threat intelligence

Skipped

OSV-ScannerOpen Source Vulnerability database check

Skipped

OWASP Dep-Check

Computer Vision Guide

A skill for conducting computer vision research, covering model architectures, dataset preparation, training pipelines, evaluation metrics, and common experimental protocols for image classification, object detection, and segmentation tasks.

Core Tasks and Architectures

Computer Vision Task Taxonomy

Image Classification:
  Input: Single image
  Output: Class label(s)
  Models: ResNet, EfficientNet, ViT, ConvNeXt

Object Detection:
  Input: Single image
  Output: Bounding boxes + class labels
  Models: YOLO (v5-v9), Faster R-CNN, DETR, RT-DETR

Semantic Segmentation:
  Input: Single image
  Output: Per-pixel class label
  Models: U-Net, DeepLab, SegFormer, Mask2Former

Instance Segmentation:
  Input: Single image
  Output: Per-pixel labels distinguishing individual objects
  Models: Mask R-CNN, Mask2Former, SAM

Image Generation:
  Input: Text prompt or noise
  Output: Generated image
  Models: Stable Diffusion, DALL-E, Imagen

Model Architecture Evolution

CNNs (Convolutional Neural Networks):
  LeNet (1998) -> AlexNet (2012) -> VGG (2014) -> ResNet (2015)
  -> EfficientNet (2019) -> ConvNeXt (2022)

Vision Transformers:
  ViT (2020) -> DeiT (2021) -> Swin Transformer (2021)
  -> BEiT (2021) -> DINOv2 (2023)

Trend: Transformers are competitive with CNNs at scale.
Hybrid architectures combining convolutions and attention are common.

Dataset Preparation

Building a Research Dataset

import os
from pathlib import Path


def organize_image_dataset(source_dir: str,
                            split_ratios: dict = None) -> dict:
    """
    Organize images into train/val/test splits.

    Args:
        source_dir: Directory containing class subdirectories
        split_ratios: Dict with 'train', 'val', 'test' ratios
    """
    if split_ratios is None:
        split_ratios = {"train": 0.7, "val": 0.15, "test": 0.15}

    import random
    random.seed(42)

    stats = {}
    for class_dir in sorted(Path(source_dir).iterdir()):
        if not class_dir.is_dir():
            continue

        images = list(class_dir.glob("*.jpg")) + list(class_dir.glob("*.png"))
        random.shuffle(images)

        n = len(images)
        n_train = int(n * split_ratios["train"])
        n_val = int(n * split_ratios["val"])

        stats[class_dir.name] = {
            "total": n,
            "train": n_train,
            "val": n_val,
            "test": n - n_train - n_val
        }

    return stats

Data Augmentation

from torchvision import transforms


def get_training_transforms(img_size: int = 224) -> transforms.Compose:
    """
    Standard data augmentation pipeline for training.

    Args:
        img_size: Target image size
    """
    return transforms.Compose([
        transforms.RandomResizedCrop(img_size, scale=(0.8, 1.0)),
        transforms.RandomHorizontalFlip(p=0.5),
        transforms.ColorJitter(brightness=0.2, contrast=0.2,
                               saturation=0.2, hue=0.1),
        transforms.RandomRotation(15),
        transforms.ToTensor(),
        transforms.Normalize(
            mean=[0.485, 0.456, 0.406],
            std=[0.229, 0.224, 0.225]
        )
    ])

Training Pipeline

Transfer Learning Workflow

import torch
import torch.nn as nn
from torchvision import models


def create_classifier(num_classes: int,
                      backbone: str = "resnet50",
                      pretrained: bool = True) -> nn.Module:
    """
    Create an image classifier using transfer learning.

    Args:
        num_classes: Number of target classes
        backbone: Model architecture name
        pretrained: Whether to use ImageNet-pretrained weights
    """
    if backbone == "resnet50":
        weights = models.ResNet50_Weights.DEFAULT if pretrained else None
        model = models.resnet50(weights=weights)
        model.fc = nn.Linear(model.fc.in_features, num_classes)
    elif backbone == "vit_b_16":
        weights = models.ViT_B_16_Weights.DEFAULT if pretrained else None
        model = models.vit_b_16(weights=weights)
        model.heads.head = nn.Linear(
            model.heads.head.in_features, num_classes
        )
    else:
        raise ValueError(f"Unknown backbone: {backbone}")

    return model

Evaluation Metrics

Metrics by Task

Classification:
  - Top-1 Accuracy: Fraction of correct predictions
  - Top-5 Accuracy: Correct class in top 5 predictions
  - Precision, Recall, F1: Per-class and macro-averaged
  - Confusion Matrix: Visualize class-level errors

Object Detection:
  - mAP (mean Average Precision): Standard COCO metric
  - [email protected]: AP at IoU threshold 0.5
  - [email protected]:0.95: AP averaged over IoU thresholds 0.5 to 0.95
  - AP per class: Identifies weak categories

Segmentation:
  - mIoU (mean Intersection over Union): Standard metric
  - Pixel Accuracy: Fraction of correctly classified pixels
  - Dice Coefficient: F1 score at the pixel level

Reproducibility Checklist

What to Report in Papers

1. Architecture: Exact model name, number of parameters
2. Pretraining: Dataset and weights used for initialization
3. Training: Optimizer, learning rate schedule, batch size, epochs
4. Augmentation: Full list of augmentations with parameters
5. Hardware: GPU type, number, training time
6. Evaluation: Exact metrics, test set version, evaluation protocol
7. Code: Link to repository with training and evaluation scripts
8. Random seeds: Report seeds used; ideally report mean over 3+ seeds

Ethical Considerations

When collecting or using image datasets, consider consent (especially for images of people), geographic and demographic representation, potential for bias amplification, and dual-use concerns. Document the dataset's composition and limitations. Follow the Datasheets for Datasets framework. For generative models, implement safeguards against generating harmful content.

wentorai/computer-vision-guide

skills/domains/ai-ml/computer-vision-guide/SKILL.md

Apply computer vision research methods, models, and evaluation tools

198 stars

tools

Updated Apr 16, 2026

$ install --global

skillsauth

npx skillsauth add wentorai/research-plugins computer-vision-guide

Install this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.

Security Scan Results

3 of 9 scanners reported clean

Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.

Scanners Passed

Scanners in report

Clean

TrivyContainer and dependency vulnerability scanner

95%

Clean

SemgrepStatic code analysis for vulnerabilities

95%

Clean

mcp-scan (Snyk)Model Context Protocol security validation

95%

Skipped

Snyk (dep)Open source security scanning

50%

Skipped

Socket.devSupply chain security analysis

50%

Skipped

VirusTotalMulti-engine malware detection

50%

Skipped

CrowdStrikeAdvanced threat intelligence

50%

Skipped

OSV-ScannerOpen Source Vulnerability database check

50%

Skipped

OWASP Dep-Check

50%

Last scanned: Apr 16, 2026, 3:27 PM4.4s1 file scanned

SKILL.md

name:: computer-vision-guide
description:: Apply computer vision research methods, models, and evaluation tools
emoji:: 👁️
category:: domains
subcategory:: ai-ml
keywords:: ["computer vision", "image classification", "object detection", "CNN", "vision transformer", "deep learning"]
source:: wentor-research-plugins

Computer Vision Guide

Core Tasks and Architectures

Computer Vision Task Taxonomy

Image Classification:
  Input: Single image
  Output: Class label(s)
  Models: ResNet, EfficientNet, ViT, ConvNeXt

Object Detection:
  Input: Single image
  Output: Bounding boxes + class labels
  Models: YOLO (v5-v9), Faster R-CNN, DETR, RT-DETR

Semantic Segmentation:
  Input: Single image
  Output: Per-pixel class label
  Models: U-Net, DeepLab, SegFormer, Mask2Former

Instance Segmentation:
  Input: Single image
  Output: Per-pixel labels distinguishing individual objects
  Models: Mask R-CNN, Mask2Former, SAM

Image Generation:
  Input: Text prompt or noise
  Output: Generated image
  Models: Stable Diffusion, DALL-E, Imagen

Model Architecture Evolution

CNNs (Convolutional Neural Networks):
  LeNet (1998) -> AlexNet (2012) -> VGG (2014) -> ResNet (2015)
  -> EfficientNet (2019) -> ConvNeXt (2022)

Vision Transformers:
  ViT (2020) -> DeiT (2021) -> Swin Transformer (2021)
  -> BEiT (2021) -> DINOv2 (2023)

Trend: Transformers are competitive with CNNs at scale.
Hybrid architectures combining convolutions and attention are common.

Dataset Preparation

Building a Research Dataset

import os
from pathlib import Path


def organize_image_dataset(source_dir: str,
                            split_ratios: dict = None) -> dict:
    """
    Organize images into train/val/test splits.

    Args:
        source_dir: Directory containing class subdirectories
        split_ratios: Dict with 'train', 'val', 'test' ratios
    """
    if split_ratios is None:
        split_ratios = {"train": 0.7, "val": 0.15, "test": 0.15}

    import random
    random.seed(42)

    stats = {}
    for class_dir in sorted(Path(source_dir).iterdir()):
        if not class_dir.is_dir():
            continue

        images = list(class_dir.glob("*.jpg")) + list(class_dir.glob("*.png"))
        random.shuffle(images)

        n = len(images)
        n_train = int(n * split_ratios["train"])
        n_val = int(n * split_ratios["val"])

        stats[class_dir.name] = {
            "total": n,
            "train": n_train,
            "val": n_val,
            "test": n - n_train - n_val
        }

    return stats

Data Augmentation

from torchvision import transforms


def get_training_transforms(img_size: int = 224) -> transforms.Compose:
    """
    Standard data augmentation pipeline for training.

    Args:
        img_size: Target image size
    """
    return transforms.Compose([
        transforms.RandomResizedCrop(img_size, scale=(0.8, 1.0)),
        transforms.RandomHorizontalFlip(p=0.5),
        transforms.ColorJitter(brightness=0.2, contrast=0.2,
                               saturation=0.2, hue=0.1),
        transforms.RandomRotation(15),
        transforms.ToTensor(),
        transforms.Normalize(
            mean=[0.485, 0.456, 0.406],
            std=[0.229, 0.224, 0.225]
        )
    ])

Training Pipeline

Transfer Learning Workflow

import torch
import torch.nn as nn
from torchvision import models


def create_classifier(num_classes: int,
                      backbone: str = "resnet50",
                      pretrained: bool = True) -> nn.Module:
    """
    Create an image classifier using transfer learning.

    Args:
        num_classes: Number of target classes
        backbone: Model architecture name
        pretrained: Whether to use ImageNet-pretrained weights
    """
    if backbone == "resnet50":
        weights = models.ResNet50_Weights.DEFAULT if pretrained else None
        model = models.resnet50(weights=weights)
        model.fc = nn.Linear(model.fc.in_features, num_classes)
    elif backbone == "vit_b_16":
        weights = models.ViT_B_16_Weights.DEFAULT if pretrained else None
        model = models.vit_b_16(weights=weights)
        model.heads.head = nn.Linear(
            model.heads.head.in_features, num_classes
        )
    else:
        raise ValueError(f"Unknown backbone: {backbone}")

    return model

Evaluation Metrics

Metrics by Task

Classification:
  - Top-1 Accuracy: Fraction of correct predictions
  - Top-5 Accuracy: Correct class in top 5 predictions
  - Precision, Recall, F1: Per-class and macro-averaged
  - Confusion Matrix: Visualize class-level errors

Object Detection:
  - mAP (mean Average Precision): Standard COCO metric
  - [email protected]: AP at IoU threshold 0.5
  - [email protected]:0.95: AP averaged over IoU thresholds 0.5 to 0.95
  - AP per class: Identifies weak categories

Segmentation:
  - mIoU (mean Intersection over Union): Standard metric
  - Pixel Accuracy: Fraction of correctly classified pixels
  - Dice Coefficient: F1 score at the pixel level

Reproducibility Checklist

What to Report in Papers

1. Architecture: Exact model name, number of parameters
2. Pretraining: Dataset and weights used for initialization
3. Training: Optimizer, learning rate schedule, batch size, epochs
4. Augmentation: Full list of augmentations with parameters
5. Hardware: GPU type, number, training time
6. Evaluation: Exact metrics, test set version, evaluation protocol
7. Code: Link to repository with training and evaluation scripts
8. Random seeds: Report seeds used; ideally report mean over 3+ seeds

Ethical Considerations

Related Skills

wentorai/thuthesis-guide

documentation

VerifiedTrustedCommunity

Write Tsinghua University theses using the ThuThesis LaTeX template

230SKILL.mdUpdated Jun 10, 2026

wentorai/thuthesis-guide

wentorai/thesis-writing-guide

development

VerifiedTrustedCommunity

Templates, formatting rules, and strategies for thesis and dissertation writing

230SKILL.mdUpdated Jun 10, 2026

wentorai/thesis-writing-guide

wentorai/thesis-template-guide

documentation

VerifiedTrustedCommunity

Set up LaTeX templates for PhD and Master's thesis documents

230SKILL.mdUpdated Jun 10, 2026

wentorai/thesis-template-guide

wentorai/sjtuthesis-guide

documentation

VerifiedTrustedCommunity

Write SJTU theses using the SJTUThesis LaTeX template with full compliance

230SKILL.mdUpdated Jun 10, 2026

wentorai/sjtuthesis-guide

Download

For Claude Desktop. Download once, then upload the file in the app — no terminal needed.

Need help? View full Cowork setup guide →

Install manually

Choose your platform

# Clone the repo
git clone https://github.com/wentorai/research-plugins.git

# Copy into Claude Code skills folder (global)
cp -r research-plugins/skills/domains/ai-ml/computer-vision-guide ~/.claude/skills/

Claude Code Skills — official skills path docs.

Repository

wentorai/research-plugins

198 stars

Compatible with

Claude Code

OpenAI Codex CLI

ChatGPT