Adoption

Agent Skills are supported by leading AI development tools.

VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory

eric861129/pandas-pro

Name: pandas-pro
Author: eric861129

public/SKILLS/Data & Analysis/pandas-pro/SKILL.md

npx skillsauth add eric861129/skills_all-in-one pandas-pro

Clean

TrivyContainer and dependency vulnerability scanner

Clean

SemgrepStatic code analysis for vulnerabilities

Clean

mcp-scan (Snyk)Model Context Protocol security validation

Skipped

Snyk (dep)Open source security scanning

Skipped

Socket.devSupply chain security analysis

Skipped

VirusTotalMulti-engine malware detection

Skipped

CrowdStrikeAdvanced threat intelligence

Skipped

OSV-ScannerOpen Source Vulnerability database check

Skipped

OWASP Dep-Check

Pandas Pro

Expert pandas developer specializing in efficient data manipulation, analysis, and transformation workflows with production-grade performance patterns.

Core Workflow

Assess data structure — Examine dtypes, memory usage, missing values, data quality:

print(df.dtypes)
print(df.memory_usage(deep=True).sum() / 1e6, "MB")
print(df.isna().sum())
print(df.describe(include="all"))

Design transformation — Plan vectorized operations, avoid loops, identify indexing strategy
Implement efficiently — Use vectorized methods, method chaining, proper indexing

Validate results — Check dtypes, shapes, null counts, and row counts:

assert result.shape[0] == expected_rows, f"Row count mismatch: {result.shape[0]}"
assert result.isna().sum().sum() == 0, "Unexpected nulls after transform"
assert set(result.columns) == expected_cols

Optimize — Profile memory, apply categorical types, use chunking if needed

Reference Guide

Load detailed guidance based on context:

| Topic | Reference | Load When | |-------|-----------|-----------| | DataFrame Operations | references/dataframe-operations.md | Indexing, selection, filtering, sorting | | Data Cleaning | references/data-cleaning.md | Missing values, duplicates, type conversion | | Aggregation & GroupBy | references/aggregation-groupby.md | GroupBy, pivot, crosstab, aggregation | | Merging & Joining | references/merging-joining.md | Merge, join, concat, combine strategies | | Performance Optimization | references/performance-optimization.md | Memory usage, vectorization, chunking |

Code Patterns

Vectorized Operations (before/after)

# ❌ AVOID: row-by-row iteration
for i, row in df.iterrows():
    df.at[i, 'tax'] = row['price'] * 0.2

# ✅ USE: vectorized assignment
df['tax'] = df['price'] * 0.2

Safe Subsetting with `.copy()`

# ❌ AVOID: chained indexing triggers SettingWithCopyWarning
df['A']['B'] = 1

# ✅ USE: .loc[] with explicit copy when mutating a subset
subset = df.loc[df['status'] == 'active', :].copy()
subset['score'] = subset['score'].fillna(0)

GroupBy Aggregation

summary = (
    df.groupby(['region', 'category'], observed=True)
    .agg(
        total_sales=('revenue', 'sum'),
        avg_price=('price', 'mean'),
        order_count=('order_id', 'nunique'),
    )
    .reset_index()
)

Merge with Validation

merged = pd.merge(
    left_df, right_df,
    on=['customer_id', 'date'],
    how='left',
    validate='m:1',          # asserts right key is unique
    indicator=True,
)
unmatched = merged[merged['_merge'] != 'both']
print(f"Unmatched rows: {len(unmatched)}")
merged.drop(columns=['_merge'], inplace=True)

Missing Value Handling

# Forward-fill then interpolate numeric gaps
df['price'] = df['price'].ffill().interpolate(method='linear')

# Fill categoricals with mode, numerics with median
for col in df.select_dtypes(include='object'):
    df[col] = df[col].fillna(df[col].mode()[0])
for col in df.select_dtypes(include='number'):
    df[col] = df[col].fillna(df[col].median())

Time Series Resampling

daily = (
    df.set_index('timestamp')
    .resample('D')
    .agg({'revenue': 'sum', 'sessions': 'count'})
    .fillna(0)
)

Pivot Table

pivot = df.pivot_table(
    values='revenue',
    index='region',
    columns='product_line',
    aggfunc='sum',
    fill_value=0,
    margins=True,
)

Memory Optimization

# Downcast numerics and convert low-cardinality strings to categorical
df['category'] = df['category'].astype('category')
df['count'] = pd.to_numeric(df['count'], downcast='integer')
df['score'] = pd.to_numeric(df['score'], downcast='float')
print(df.memory_usage(deep=True).sum() / 1e6, "MB after optimization")

Constraints

MUST DO

Use vectorized operations instead of loops
Set appropriate dtypes (categorical for low-cardinality strings)
Check memory usage with .memory_usage(deep=True)
Handle missing values explicitly (don't silently drop)
Use method chaining for readability
Preserve index integrity through operations
Validate data quality before and after transformations
Use .copy() when modifying subsets to avoid SettingWithCopyWarning

MUST NOT DO

Iterate over DataFrame rows with .iterrows() unless absolutely necessary
Use chained indexing (df['A']['B']) — use .loc[] or .iloc[]
Ignore SettingWithCopyWarning messages
Load entire large datasets without chunking
Use deprecated methods (.ix, .append() — use pd.concat())
Convert to Python lists for operations possible in pandas
Assume data is clean without validation

Output Templates

When implementing pandas solutions, provide:

Code with vectorized operations and proper indexing
Comments explaining complex transformations
Memory/performance considerations if dataset is large
Data validation checks (dtypes, nulls, shapes)

eric861129/pandas-pro

public/SKILLS/Data & Analysis/pandas-pro/SKILL.md

Performs pandas DataFrame operations for data analysis, manipulation, and transformation. Use when working with pandas DataFrames, data cleaning, aggregation, merging, or time series analysis. Invoke for data manipulation tasks such as joining DataFrames on multiple keys, pivoting tables, resampling time series, handling NaN values with interpolation or forward-fill, groupby aggregations, type conversion, or performance optimization of large datasets.

38 stars

development

Updated Apr 4, 2026

$ install --global

skillsauth

npx skillsauth add eric861129/skills_all-in-one pandas-pro

Install this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.

Security Scan Results

3 of 9 scanners reported clean

Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.

Scanners Passed

Scanners in report

Clean

TrivyContainer and dependency vulnerability scanner

95%

Clean

SemgrepStatic code analysis for vulnerabilities

95%

Clean

mcp-scan (Snyk)Model Context Protocol security validation

95%

Skipped

Snyk (dep)Open source security scanning

50%

Skipped

Socket.devSupply chain security analysis

50%

Skipped

VirusTotalMulti-engine malware detection

50%

Skipped

CrowdStrikeAdvanced threat intelligence

50%

Skipped

OSV-ScannerOpen Source Vulnerability database check

50%

Skipped

OWASP Dep-Check

50%

Last scanned: Apr 7, 2026, 3:23 AM23.9s1 file scanned

SKILL.md

name:: pandas-pro
description:: Performs pandas DataFrame operations for data analysis, manipulation, and transformation. Use when working with pandas DataFrames, data cleaning, aggregation, merging, or time series analysis. Invoke for data manipulation tasks such as joining DataFrames on multiple keys, pivoting tables, resampling time series, handling NaN values with interpolation or forward-fill, groupby aggregations, type conversion, or performance optimization of large datasets.
license:: MIT
author:: https://github.com/Jeffallan
version:: 1.1.0
domain:: data-ml
triggers:: pandas, DataFrame, data manipulation, data cleaning, aggregation, groupby, merge, join, time series, data wrangling, pivot table, data transformation
role:: expert
scope:: implementation
output-format:: code
related-skills:: python-pro

Pandas Pro

Expert pandas developer specializing in efficient data manipulation, analysis, and transformation workflows with production-grade performance patterns.

Core Workflow

Assess data structure — Examine dtypes, memory usage, missing values, data quality:

print(df.dtypes)
print(df.memory_usage(deep=True).sum() / 1e6, "MB")
print(df.isna().sum())
print(df.describe(include="all"))

Design transformation — Plan vectorized operations, avoid loops, identify indexing strategy
Implement efficiently — Use vectorized methods, method chaining, proper indexing

Validate results — Check dtypes, shapes, null counts, and row counts:

assert result.shape[0] == expected_rows, f"Row count mismatch: {result.shape[0]}"
assert result.isna().sum().sum() == 0, "Unexpected nulls after transform"
assert set(result.columns) == expected_cols

Optimize — Profile memory, apply categorical types, use chunking if needed

Reference Guide

Load detailed guidance based on context:

Code Patterns

Vectorized Operations (before/after)

# ❌ AVOID: row-by-row iteration
for i, row in df.iterrows():
    df.at[i, 'tax'] = row['price'] * 0.2

# ✅ USE: vectorized assignment
df['tax'] = df['price'] * 0.2

Safe Subsetting with `.copy()`

# ❌ AVOID: chained indexing triggers SettingWithCopyWarning
df['A']['B'] = 1

# ✅ USE: .loc[] with explicit copy when mutating a subset
subset = df.loc[df['status'] == 'active', :].copy()
subset['score'] = subset['score'].fillna(0)

GroupBy Aggregation

summary = (
    df.groupby(['region', 'category'], observed=True)
    .agg(
        total_sales=('revenue', 'sum'),
        avg_price=('price', 'mean'),
        order_count=('order_id', 'nunique'),
    )
    .reset_index()
)

Merge with Validation

merged = pd.merge(
    left_df, right_df,
    on=['customer_id', 'date'],
    how='left',
    validate='m:1',          # asserts right key is unique
    indicator=True,
)
unmatched = merged[merged['_merge'] != 'both']
print(f"Unmatched rows: {len(unmatched)}")
merged.drop(columns=['_merge'], inplace=True)

Missing Value Handling

# Forward-fill then interpolate numeric gaps
df['price'] = df['price'].ffill().interpolate(method='linear')

# Fill categoricals with mode, numerics with median
for col in df.select_dtypes(include='object'):
    df[col] = df[col].fillna(df[col].mode()[0])
for col in df.select_dtypes(include='number'):
    df[col] = df[col].fillna(df[col].median())

Time Series Resampling

daily = (
    df.set_index('timestamp')
    .resample('D')
    .agg({'revenue': 'sum', 'sessions': 'count'})
    .fillna(0)
)

Pivot Table

pivot = df.pivot_table(
    values='revenue',
    index='region',
    columns='product_line',
    aggfunc='sum',
    fill_value=0,
    margins=True,
)

Memory Optimization

# Downcast numerics and convert low-cardinality strings to categorical
df['category'] = df['category'].astype('category')
df['count'] = pd.to_numeric(df['count'], downcast='integer')
df['score'] = pd.to_numeric(df['score'], downcast='float')
print(df.memory_usage(deep=True).sum() / 1e6, "MB after optimization")

Constraints

MUST DO

Use vectorized operations instead of loops
Set appropriate dtypes (categorical for low-cardinality strings)
Check memory usage with .memory_usage(deep=True)
Handle missing values explicitly (don't silently drop)
Use method chaining for readability
Preserve index integrity through operations
Validate data quality before and after transformations
Use .copy() when modifying subsets to avoid SettingWithCopyWarning

MUST NOT DO

Iterate over DataFrame rows with .iterrows() unless absolutely necessary
Use chained indexing (df['A']['B']) — use .loc[] or .iloc[]
Ignore SettingWithCopyWarning messages
Load entire large datasets without chunking
Use deprecated methods (.ix, .append() — use pd.concat())
Convert to Python lists for operations possible in pandas
Assume data is clean without validation

Output Templates

When implementing pandas solutions, provide:

Code with vectorized operations and proper indexing
Comments explaining complex transformations
Memory/performance considerations if dataset is large
Data validation checks (dtypes, nulls, shapes)

Related Skills

eric861129/what-if-oracle

development

VerifiedTrustedCommunity

Run structured What-If scenario analysis with multi-branch possibility exploration. Use this skill when the user asks speculative questions like "what if...", "what would happen if...", "what are the possibilities", "explore scenarios", "scenario analysis", "possibility space", "what could go wrong", "best case / worst case", "risk analysis", "contingency planning", "strategic options", or any question about uncertain futures. Also trigger when the user faces a fork-in-the-road decision, wants to stress-test an idea, or needs to think through consequences before committing.

38SKILL.mdUpdated Apr 4, 2026

eric861129/what-if-oracle

eric861129/venue-templates

development

VerifiedTrustedCommunity

Access comprehensive LaTeX templates, formatting requirements, and submission guidelines for major scientific publication venues (Nature, Science, PLOS, IEEE, ACM), academic conferences (NeurIPS, ICML, CVPR, CHI), research posters, and grant proposals (NSF, NIH, DOE, DARPA). This skill should be used when preparing manuscripts for journal submission, conference papers, research posters, or grant proposals and need venue-specific formatting requirements and templates.

38SKILL.mdUpdated Apr 4, 2026

eric861129/venue-templates

eric861129/the-fool

development

VerifiedTrustedCommunity

Use when challenging ideas, plans, decisions, or proposals using structured critical reasoning. Invoke to play devil's advocate, run a pre-mortem, red team, or audit evidence and assumptions.

38SKILL.mdUpdated Apr 4, 2026

eric861129/scientific-writing

tools

VerifiedTrustedCommunity

Core skill for the deep research and writing tool. Write scientific manuscripts in full paragraphs (never bullet points). Use two-stage process with (1) section outlines with key points using research-lookup then (2) convert to flowing prose. IMRAD structure, citations (APA/AMA/Vancouver), figures/tables, reporting guidelines (CONSORT/STROBE/PRISMA), for research papers and journal submissions.

38SKILL.mdUpdated Apr 4, 2026

eric861129/scientific-writing

Download

For Claude Desktop. Download once, then upload the file in the app — no terminal needed.

Need help? View full Cowork setup guide →

Install manually

Choose your platform

# Clone the repo
git clone https://github.com/eric861129/skills_all-in-one.git

# Copy into Claude Code skills folder (global)
cp -r skills_all-in-one/public/SKILLS/Data & Analysis/pandas-pro ~/.claude/skills/

Claude Code Skills — official skills path docs.

Repository

eric861129/skills_all-in-one

38 stars

Compatible with

Claude Code

OpenAI Codex CLI

ChatGPT

Adoption

eric861129/pandas-pro

$ install --global

Security Scan Results

SKILL.md

Pandas Pro

Core Workflow

Reference Guide

Code Patterns

Vectorized Operations (before/after)

Safe Subsetting with .copy()

GroupBy Aggregation

Merge with Validation

Missing Value Handling

Time Series Resampling

Pivot Table

Memory Optimization

Constraints

MUST DO

MUST NOT DO

Output Templates

Related Skills

eric861129/what-if-oracle

eric861129/venue-templates

eric861129/the-fool

eric861129/scientific-writing

eric861129/pandas-pro

$ install --global

Security Scan Results

SKILL.md

Pandas Pro

Core Workflow

Reference Guide

Code Patterns

Vectorized Operations (before/after)

Safe Subsetting with .copy()

GroupBy Aggregation

Merge with Validation

Missing Value Handling

Time Series Resampling

Pivot Table

Memory Optimization

Constraints

MUST DO

MUST NOT DO

Output Templates

Related Skills

eric861129/what-if-oracle

eric861129/venue-templates

eric861129/the-fool

eric861129/scientific-writing

Safe Subsetting with `.copy()`

Safe Subsetting with `.copy()`