Adoption

Agent Skills are supported by leading AI development tools.

VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory

github/datanalysis-credit-risk

Name: datanalysis-credit-risk
Author: github

skills/datanalysis-credit-risk/SKILL.md

npx skillsauth add github/awesome-copilot datanalysis-credit-risk

Clean

TrivyContainer and dependency vulnerability scanner

Clean

SemgrepStatic code analysis for vulnerabilities

Clean

mcp-scan (Snyk)Model Context Protocol security validation

Skipped

Snyk (dep)Open source security scanning

Skipped

Socket.devSupply chain security analysis

Skipped

VirusTotalMulti-engine malware detection

Skipped

CrowdStrikeAdvanced threat intelligence

Skipped

OSV-ScannerOpen Source Vulnerability database check

Skipped

OWASP Dep-Check

Data Cleaning and Variable Screening

Quick Start

# Run the complete data cleaning pipeline
python ".github/skills/datanalysis-credit-risk/scripts/example.py"

Complete Process Description

The data cleaning pipeline consists of the following 11 steps, each executed independently without deleting the original data:

Get Data - Load and format raw data
Organization Sample Analysis - Statistics of sample count and bad sample rate for each organization
Separate OOS Data - Separate out-of-sample (OOS) samples from modeling samples
Filter Abnormal Months - Remove months with insufficient bad sample count or total sample count
Calculate Missing Rate - Calculate overall and organization-level missing rates for each feature
Drop High Missing Rate Features - Remove features with overall missing rate exceeding threshold
Drop Low IV Features - Remove features with overall IV too low or IV too low in too many organizations
Drop High PSI Features - Remove features with unstable PSI
Null Importance Denoising - Remove noise features using label permutation method
Drop High Correlation Features - Remove high correlation features based on original gain
Export Report - Generate Excel report containing details and statistics of all steps

Core Functions

| Function | Purpose | Module | |------|------|----------| | get_dataset() | Load and format data | references.func | | org_analysis() | Organization sample analysis | references.func | | missing_check() | Calculate missing rate | references.func | | drop_abnormal_ym() | Filter abnormal months | references.analysis | | drop_highmiss_features() | Drop high missing rate features | references.analysis | | drop_lowiv_features() | Drop low IV features | references.analysis | | drop_highpsi_features() | Drop high PSI features | references.analysis | | drop_highnoise_features() | Null Importance denoising | references.analysis | | drop_highcorr_features() | Drop high correlation features | references.analysis | | iv_distribution_by_org() | IV distribution statistics | references.analysis | | psi_distribution_by_org() | PSI distribution statistics | references.analysis | | value_ratio_distribution_by_org() | Value ratio distribution statistics | references.analysis | | export_cleaning_report() | Export cleaning report | references.analysis |

Parameter Description

Data Loading Parameters

DATA_PATH: Data file path (best are parquet format)
DATE_COL: Date column name
Y_COL: Label column name
ORG_COL: Organization column name
KEY_COLS: Primary key column name list

OOS Organization Configuration

OOS_ORGS: Out-of-sample organization list

Abnormal Month Filtering Parameters

min_ym_bad_sample: Minimum bad sample count per month (default 10)
min_ym_sample: Minimum total sample count per month (default 500)

Missing Rate Parameters

missing_ratio: Overall missing rate threshold (default 0.6)

IV Parameters

overall_iv_threshold: Overall IV threshold (default 0.1)
org_iv_threshold: Single organization IV threshold (default 0.1)
max_org_threshold: Maximum tolerated low IV organization count (default 2)

PSI Parameters

psi_threshold: PSI threshold (default 0.1)
max_months_ratio: Maximum unstable month ratio (default 1/3)
max_orgs: Maximum unstable organization count (default 6)

Null Importance Parameters

n_estimators: Number of trees (default 100)
max_depth: Maximum tree depth (default 5)
gain_threshold: Gain difference threshold (default 50)

High Correlation Parameters

max_corr: Correlation threshold (default 0.9)
top_n_keep: Keep top N features by original gain ranking (default 20)

Output Report

The generated Excel report contains the following sheets:

汇总 - Summary information of all steps, including operation results and conditions
机构样本统计 - Sample count and bad sample rate for each organization
分离OOS数据 - OOS sample and modeling sample counts
Step4-异常月份处理 - Abnormal months that were removed
缺失率明细 - Overall and organization-level missing rates for each feature
Step5-有值率分布统计 - Distribution of features in different value ratio ranges
Step6-高缺失率处理 - High missing rate features that were removed
Step7-IV明细 - IV values of each feature in each organization and overall
Step7-IV处理 - Features that do not meet IV conditions and low IV organizations
Step7-IV分布统计 - Distribution of features in different IV ranges
Step8-PSI明细 - PSI values of each feature in each organization each month
Step8-PSI处理 - Features that do not meet PSI conditions and unstable organizations
Step8-PSI分布统计 - Distribution of features in different PSI ranges
Step9-null importance处理 - Noise features that were removed
Step10-高相关性剔除 - High correlation features that were removed

Features

Interactive Input: Parameters can be input before each step execution, with default values supported
Independent Execution: Each step is executed independently without deleting original data, facilitating comparative analysis
Complete Report: Generate complete Excel report containing details, statistics, and distributions
Multi-process Support: IV and PSI calculations support multi-process acceleration
Organization-level Analysis: Support organization-level statistics and modeling/OOS distinction

github/datanalysis-credit-risk

skills/datanalysis-credit-risk/SKILL.md

Credit risk data cleaning and variable screening pipeline for pre-loan modeling. Use when working with raw credit data that needs quality assessment, missing value analysis, or variable selection before modeling. it covers data loading and formatting, abnormal period filtering, missing rate calculation, high-missing variable removal,low-IV variable filtering, high-PSI variable removal, Null Importance denoising, high-correlation variable removal, and cleaning report generation. Applicable scenarios arecredit risk data cleaning, variable screening, pre-loan modeling preprocessing.

25,862 stars

development

Updated Mar 18, 2026

$ install --global

skillsauth

npx skillsauth add github/awesome-copilot datanalysis-credit-risk

Install this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.

Security Scan Results

3 of 9 scanners reported clean

Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.

Scanners Passed

Scanners in report

Clean

TrivyContainer and dependency vulnerability scanner

95%

Clean

SemgrepStatic code analysis for vulnerabilities

95%

Clean

mcp-scan (Snyk)Model Context Protocol security validation

95%

Skipped

Snyk (dep)Open source security scanning

50%

Skipped

Socket.devSupply chain security analysis

50%

Skipped

VirusTotalMulti-engine malware detection

70%

Skipped

CrowdStrikeAdvanced threat intelligence

50%

Skipped

OSV-ScannerOpen Source Vulnerability database check

50%

Skipped

OWASP Dep-Check

50%

Last scanned: Apr 20, 2026, 11:26 AM212.7s1 file scanned

SKILL.md

name:: datanalysis-credit-risk
description:: Credit risk data cleaning and variable screening pipeline for pre-loan modeling. Use when working with raw credit data that needs quality assessment, missing value analysis, or variable selection before modeling. it covers data loading and formatting, abnormal period filtering, missing rate calculation, high-missing variable removal,low-IV variable filtering, high-PSI variable removal, Null Importance denoising, high-correlation variable removal, and cleaning report generation. Applicable scenarios arecredit risk data cleaning, variable screening, pre-loan modeling preprocessing.

Data Cleaning and Variable Screening

Quick Start

# Run the complete data cleaning pipeline
python ".github/skills/datanalysis-credit-risk/scripts/example.py"

Complete Process Description

The data cleaning pipeline consists of the following 11 steps, each executed independently without deleting the original data:

Get Data - Load and format raw data
Organization Sample Analysis - Statistics of sample count and bad sample rate for each organization
Separate OOS Data - Separate out-of-sample (OOS) samples from modeling samples
Filter Abnormal Months - Remove months with insufficient bad sample count or total sample count
Calculate Missing Rate - Calculate overall and organization-level missing rates for each feature
Drop High Missing Rate Features - Remove features with overall missing rate exceeding threshold
Drop Low IV Features - Remove features with overall IV too low or IV too low in too many organizations
Drop High PSI Features - Remove features with unstable PSI
Null Importance Denoising - Remove noise features using label permutation method
Drop High Correlation Features - Remove high correlation features based on original gain
Export Report - Generate Excel report containing details and statistics of all steps

Core Functions

Parameter Description

Data Loading Parameters

DATA_PATH: Data file path (best are parquet format)
DATE_COL: Date column name
Y_COL: Label column name
ORG_COL: Organization column name
KEY_COLS: Primary key column name list

OOS Organization Configuration

OOS_ORGS: Out-of-sample organization list

Abnormal Month Filtering Parameters

min_ym_bad_sample: Minimum bad sample count per month (default 10)
min_ym_sample: Minimum total sample count per month (default 500)

Missing Rate Parameters

missing_ratio: Overall missing rate threshold (default 0.6)

IV Parameters

overall_iv_threshold: Overall IV threshold (default 0.1)
org_iv_threshold: Single organization IV threshold (default 0.1)
max_org_threshold: Maximum tolerated low IV organization count (default 2)

PSI Parameters

psi_threshold: PSI threshold (default 0.1)
max_months_ratio: Maximum unstable month ratio (default 1/3)
max_orgs: Maximum unstable organization count (default 6)

Null Importance Parameters

n_estimators: Number of trees (default 100)
max_depth: Maximum tree depth (default 5)
gain_threshold: Gain difference threshold (default 50)

High Correlation Parameters

max_corr: Correlation threshold (default 0.9)
top_n_keep: Keep top N features by original gain ranking (default 20)

Output Report

The generated Excel report contains the following sheets:

汇总 - Summary information of all steps, including operation results and conditions
机构样本统计 - Sample count and bad sample rate for each organization
分离OOS数据 - OOS sample and modeling sample counts
Step4-异常月份处理 - Abnormal months that were removed
缺失率明细 - Overall and organization-level missing rates for each feature
Step5-有值率分布统计 - Distribution of features in different value ratio ranges
Step6-高缺失率处理 - High missing rate features that were removed
Step7-IV明细 - IV values of each feature in each organization and overall
Step7-IV处理 - Features that do not meet IV conditions and low IV organizations
Step7-IV分布统计 - Distribution of features in different IV ranges
Step8-PSI明细 - PSI values of each feature in each organization each month
Step8-PSI处理 - Features that do not meet PSI conditions and unstable organizations
Step8-PSI分布统计 - Distribution of features in different PSI ranges
Step9-null importance处理 - Noise features that were removed
Step10-高相关性剔除 - High correlation features that were removed

Features

Interactive Input: Parameters can be input before each step execution, with default values supported
Independent Execution: Each step is executed independently without deleting original data, facilitating comparative analysis
Complete Report: Generate complete Excel report containing details, statistics, and distributions
Multi-process Support: IV and PSI calculations support multi-process acceleration
Organization-level Analysis: Support organization-level statistics and modeling/OOS distinction

Related Skills

github/python-pypi-package-builder

tools

VerifiedTrustedOfficial

End-to-end skill for building, testing, linting, versioning, and publishing a production-grade Python library to PyPI. Covers all four build backends (setuptools+setuptools_scm, hatchling, flit, poetry), PEP 440 versioning, semantic versioning, dynamic git-tag versioning, OOP/SOLID design, type hints (PEP 484/526/544/561), Trusted Publishing (OIDC), and the full PyPA packaging flow. Use for: creating Python packages, pip-installable SDKs, CLI tools, framework plugins, pyproject.toml setup, py.typed, setuptools_scm, semver, mypy, pre-commit, GitHub Actions CI/CD, or PyPI publishing.

29,271SKILL.mdUpdated Apr 11, 2026

github/python-pypi-package-builder

github/mcp-security-audit

tools

VerifiedTrustedOfficial

Audit MCP (Model Context Protocol) server configurations for security issues. Use this skill when: - Reviewing .mcp.json files for security risks - Checking MCP server args for hardcoded secrets or shell injection patterns - Validating that MCP servers use pinned versions (not @latest) - Detecting unpinned dependencies in MCP server configurations - Auditing which MCP servers a project registers and whether they're on an approved list - Checking for environment variable usage vs. hardcoded credentials in MCP configs - Any request like "is my MCP config secure?", "audit my MCP servers", or "check .mcp.json" keywords: [mcp, security, audit, secrets, shell-injection, supply-chain, governance]

29,271SKILL.mdUpdated Apr 11, 2026

github/mcp-security-audit

github/lsp-setup

tools

VerifiedTrustedOfficial

Enable code intelligence (go-to-definition, find-references, hover, type info) for any programming language by installing and configuring an LSP server for Copilot CLI. Detects the OS, installs the right server, and generates the JSON configuration (user-level or repo-level). Use when you need deeper code understanding and no LSP server is configured, or when the user asks to set up, install, or configure an LSP server.

29,271SKILL.mdUpdated Apr 11, 2026

github/gsap-framer-scroll-animation

development

VerifiedTrustedOfficial

Use this skill whenever the user wants to build scroll animations, scroll effects, parallax, scroll-triggered reveals, pinned sections, horizontal scroll, text animations, or any motion tied to scroll position — in vanilla JS, React, or Next.js. Covers GSAP ScrollTrigger (pinning, scrubbing, snapping, timelines, horizontal scroll, ScrollSmoother, matchMedia) and Framer Motion / Motion v12 (useScroll, useTransform, useSpring, whileInView, variants). Use this skill even if the user just says "animate on scroll", "fade in as I scroll", "make it scroll like Apple", "parallax effect", "sticky section", "scroll progress bar", or "entrance animation". Also triggers for Copilot prompt patterns for GSAP or Framer Motion code generation. Pairs with the premium-frontend-ui skill for creative philosophy and design-level polish.

29,271SKILL.mdUpdated Apr 11, 2026

github/gsap-framer-scroll-animation

Download

For Claude Desktop. Download once, then upload the file in the app — no terminal needed.

Need help? View full Cowork setup guide →

Install manually

Choose your platform

# Clone the repo
git clone https://github.com/github/awesome-copilot.git

# Copy into Claude Code skills folder (global)
cp -r awesome-copilot/skills/datanalysis-credit-risk ~/.claude/skills/

Claude Code Skills — official skills path docs.

Repository

github/awesome-copilot

25,862 stars

Compatible with

Claude Code

OpenAI Codex CLI

ChatGPT