plugins/static-analysis/skills/sarif-parsing/SKILL.md
Parses and processes SARIF files from static analysis tools like CodeQL, Semgrep, or other scanners. Triggers on "parse sarif", "read scan results", "aggregate findings", "deduplicate alerts", or "process sarif output". Handles filtering, deduplication, format conversion, and CI/CD integration of SARIF data. Does NOT run scans — use the Semgrep or CodeQL skills for that.
npx skillsauth add trailofbits/skills sarif-parsingInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
You are a SARIF parsing expert. Your role is to help users effectively read, analyze, and process SARIF files from static analysis tools.
Use this skill when:
Do NOT use this skill for:
SARIF 2.1.0 is the current OASIS standard. Every SARIF file has this hierarchical structure:
sarifLog
├── version: "2.1.0"
├── $schema: (optional, enables IDE validation)
└── runs[] (array of analysis runs)
├── tool
│ ├── driver
│ │ ├── name (required)
│ │ ├── version
│ │ └── rules[] (rule definitions)
│ └── extensions[] (plugins)
├── results[] (findings)
│ ├── ruleId
│ ├── ruleIndex (index into tool.driver.rules[])
│ ├── level (OPTIONAL, inherited from the rule when absent)
│ ├── message.text
│ ├── locations[]
│ │ └── physicalLocation
│ │ ├── artifactLocation.uri
│ │ └── region (startLine, startColumn, etc.)
│ ├── fingerprints{}
│ └── partialFingerprints{}
└── artifacts[] (scanned files metadata)
result.level is optional. CodeQL omits it on every result and records severity on the
rule as defaultConfiguration.level, which the result inherits. Read result.level
directly and a CodeQL run scores as clean however many errors it found, which is how a
severity gate ends up exiting 0 on a failing repo.
Resolve severity in this order (SARIF 2.1.0 section 3.27.10):
kind other than "fail" (a pass/notApplicable record), so "none"result.level, when presentdefaultConfiguration.level, joining ruleIndex into
runs[].tool.driver.rules[], or matching ruleId against rules[].id when the tool
omits ruleIndex"warning", the SARIF defaultEvery severity query in this skill starts from that resolution. In jq it is the
LEVEL_FN definition in {baseDir}/resources/jq-queries.md;
in Python it is resolve_level(result, run) in
{baseDir}/resources/sarif_helpers.py.
Without stable fingerprints, you can't track findings across runs:
Tools report different paths (/path/to/project/ vs /github/workspace/), so path-based matching fails. Fingerprints hash the content (code snippet, rule ID, relative location) to create stable identifiers regardless of environment.
| Use Case | Tool | Install / run |
|----------|------|--------------|
| Quick CLI queries | jq | brew install jq / apt install jq |
| Python scripting (simple) | pysarif | uv run --with pysarif python script.py |
| Python scripting (advanced) | sarif-tools | uv run --with sarif-tools python script.py |
| .NET applications | SARIF SDK | NuGet package |
| JavaScript/Node.js | sarif-js | npm package |
| Go applications | garif | go get github.com/chavacava/garif |
| Validation | SARIF Validator | sarifweb.azurewebsites.net |
For rapid exploration and one-off queries:
# Pretty print the file
jq '.' results.sarif
# Count total findings
jq '[.runs[].results[]] | length' results.sarif
# List all rule IDs triggered
jq '[.runs[].results[].ruleId] | unique' results.sarif
# Severity resolution, needed by every query below that filters on level.
# See resources/jq-queries.md for the annotated version.
LEVEL_FN='
def rule($run):
. as $r
| ($run.tool.driver.rules // []) as $rules
| (if ($r.ruleIndex | type) == "number" and $r.ruleIndex >= 0
then $rules[$r.ruleIndex] else null end)
// first($rules[] | select(.id == $r.ruleId))
// null;
def level($run):
. as $r
| if ($r.kind // "fail") != "fail" then "none"
else ($r.level // rule($run).defaultConfiguration.level // "warning") end;
'
# Extract errors only
jq "$LEVEL_FN"'.runs[] as $run | $run.results[] | select(level($run) == "error")' results.sarif
# Get findings with file locations
jq '.runs[].results[] | {
rule: .ruleId,
message: .message.text,
file: .locations[0].physicalLocation.artifactLocation.uri,
line: .locations[0].physicalLocation.region.startLine
}' results.sarif
# Filter by severity and get count per rule
jq "$LEVEL_FN"'[.runs[] as $run | $run.results[] | select(level($run) == "error")] | group_by(.ruleId) | map({rule: .[0].ruleId, count: length})' results.sarif
# Extract findings for a specific file
jq --arg file "src/auth.py" '.runs[].results[] | select(.locations[].physicalLocation.artifactLocation.uri | contains($file))' results.sarif
For programmatic access with full object model:
from pysarif import load_from_file, save_to_file
# Load SARIF file
sarif = load_from_file("results.sarif")
# Iterate through runs and results
for run in sarif.runs:
tool_name = run.tool.driver.name
print(f"Tool: {tool_name}")
for result in run.results:
# pysarif fills a missing result.level with "warning", so .level here is NOT the
# rule-inherited severity: a CodeQL error (no level on the result, severity on the
# rule) reads as "warning". Gate severity with Strategy 1's level() or with
# resolve_level() in resources/sarif_helpers.py, which resolve it from the rule.
print(f" {result.rule_id}: {result.message.text}")
if result.locations:
loc = result.locations[0].physical_location
if loc and loc.artifact_location:
print(f" File: {loc.artifact_location.uri}")
if loc.region:
print(f" Line: {loc.region.start_line}")
# Save modified SARIF
save_to_file(sarif, "modified.sarif")
For aggregation, reporting, and CI/CD integration:
from sarif import loader
# Load single file
sarif_data = loader.load_sarif_file("results.sarif")
# Or load multiple files
sarif_set = loader.load_sarif_files(["tool1.sarif", "tool2.sarif"])
# Get summary report
report = sarif_data.get_report()
# Get histogram by severity
errors = report.get_issue_type_histogram_for_severity("error")
warnings = report.get_issue_type_histogram_for_severity("warning")
# Filter by severity. sarif-tools hands back raw result dicts, and a result's level may
# live on its rule, so resolve it against the run instead of reading r["level"].
from sarif_helpers import extract_findings, filter_by_level, load_sarif
high_severity = filter_by_level(extract_findings(load_sarif("results.sarif")), "error")
sarif-tools CLI commands:
# Summary of findings
sarif summary results.sarif
# List all results with details
sarif ls results.sarif
# Get results by severity
sarif ls --level error results.sarif
# Diff two SARIF files (find new/fixed issues)
sarif diff baseline.sarif current.sarif
# Convert to other formats
sarif csv results.sarif > results.csv
sarif html results.sarif > report.html
When combining results from multiple tools:
import json
from sarif_helpers import deduplicate, extract_findings
def aggregate_sarif_files(sarif_paths: list[str]) -> dict:
"""Combine multiple SARIF files into one."""
aggregated = {
"version": "2.1.0",
"$schema": "https://json.schemastore.org/sarif-2.1.0.json",
"runs": []
}
for path in sarif_paths:
with open(path) as f:
sarif = json.load(f)
aggregated["runs"].extend(sarif.get("runs", []))
return aggregated
unique = deduplicate(extract_findings(aggregate_sarif_files(["tool1.sarif", "tool2.sarif"])))
deduplicate() prefers whatever fingerprints or partialFingerprints the tool supplied
and falls back to hashing rule ID, the whole normalized path, line, and message. Keep the
directory in that key: the same rule at the same line in auth/login.py and
admin/login.py is two findings, and a basename-only key throws one of them away.
resources/sarif_helpers.py covers this with the standard library alone.
extract_findings() returns Finding objects whose severity is already resolved, and
filter_by_level(), sort_by_severity(), deduplicate() and diff_findings() consume
those:
from sarif_helpers import extract_findings, filter_by_level, load_sarif, sort_by_severity
findings = sort_by_severity(extract_findings(load_sarif("results.sarif")))
for f in filter_by_level(findings, "error"):
print(f"{f.file_path}:{f.start_line} [{f.level}] {f.rule_id}: {f.message}")
Writing your own extractor, severity is the part that goes wrong silently:
def resolve_level(result: dict, run: dict) -> str:
"""Severity of a result: its own level, else its rule's default, else "warning"."""
if result.get("kind", "fail") != "fail":
return "none"
if result.get("level"):
return result["level"]
rules = run.get("tool", {}).get("driver", {}).get("rules", [])
index = result.get("ruleIndex")
rule = rules[index] if isinstance(index, int) and 0 <= index < len(rules) else next(
(r for r in rules if r.get("id") == result.get("ruleId")), {}
)
return rule.get("defaultConfiguration", {}).get("level") or "warning"
Results carry ruleIndex on some tools and only ruleId on others, so a resolver that
joins one way alone silently returns the default for every result the other kind of tool
produces.
Different tools report paths differently (absolute, relative, URI-encoded), so
file:///src/a%20b.py and src/a b.py can be the same file. Strip the file:// scheme,
percent-decode, resolve against a base path, and normalize separators before comparing or
hashing anything: normalize_path() in resources/sarif_helpers.py does all four.
Fingerprints may not match if:
Solution: Use multiple fingerprint strategies:
def compute_stable_fingerprint(result: dict, file_content: str = None) -> str:
"""Compute environment-independent fingerprint."""
import hashlib
components = [
result.get("ruleId", ""),
result.get("message", {}).get("text", "")[:100], # First 100 chars
]
# Add code snippet if available
if file_content and result.get("locations"):
region = result["locations"][0].get("physicalLocation", {}).get("region", {})
if region.get("startLine"):
lines = file_content.split("\n")
line_idx = region["startLine"] - 1
if 0 <= line_idx < len(lines):
# Normalize whitespace
components.append(lines[line_idx].strip())
return hashlib.sha256("".join(components).encode()).hexdigest()[:16]
SARIF allows many optional fields. Always use defensive access:
def safe_get_location(result: dict) -> tuple[str, int]:
"""Safely extract file and line from result."""
try:
loc = result.get("locations", [{}])[0]
phys = loc.get("physicalLocation", {})
file_path = phys.get("artifactLocation", {}).get("uri", "unknown")
line = phys.get("region", {}).get("startLine", 0)
return file_path, line
except (IndexError, KeyError, TypeError):
return "unknown", 0
For very large SARIF files (100MB+):
import ijson # run via: uv run --with ijson
def stream_results(sarif_path: str):
"""Stream results without loading entire file."""
with open(sarif_path, "rb") as f:
# Stream through results arrays
for result in ijson.items(f, "runs.item.results.item"):
yield result
Validate before processing to catch malformed files:
# Using ajv-cli
npm install -g ajv-cli
ajv validate -s sarif-schema-2.1.0.json -d results.sarif
# Using Python jsonschema
uv run --with jsonschema python your_script.py # e.g. the function below
from jsonschema import validate, ValidationError
import json
def validate_sarif(sarif_path: str, schema_path: str) -> bool:
"""Validate SARIF file against schema."""
with open(sarif_path) as f:
sarif = json.load(f)
with open(schema_path) as f:
schema = json.load(f)
try:
validate(sarif, schema)
return True
except ValidationError as e:
print(f"Validation error: {e.message}")
return False
- name: Upload SARIF
uses: github/codeql-action/upload-sarif@v3
with:
sarif_file: results.sarif
- name: Check for high severity
run: |
# select(.level == "error") counts zero on CodeQL output, which records severity on
# the rule instead. Resolve the level or the gate passes on a repo full of errors.
HIGH_COUNT=$(jq '
def rule($run):
. as $r
| ($run.tool.driver.rules // []) as $rules
| (if ($r.ruleIndex | type) == "number" and $r.ruleIndex >= 0
then $rules[$r.ruleIndex] else null end)
// first($rules[] | select(.id == $r.ruleId))
// null;
def level($run):
. as $r
| if ($r.kind // "fail") != "fail" then "none"
else ($r.level // rule($run).defaultConfiguration.level // "warning") end;
[.runs[] as $run | $run.results[] | select(level($run) == "error")] | length
' results.sarif)
if [ "$HIGH_COUNT" -gt 0 ]; then
echo "Found $HIGH_COUNT high severity issues"
exit 1
fi
from sarif import loader
def check_for_regressions(baseline: str, current: str) -> int:
"""Return count of new issues not in baseline."""
baseline_data = loader.load_sarif_file(baseline)
current_data = loader.load_sarif_file(current)
baseline_fps = {get_fingerprint(r) for r in baseline_data.get_results()}
new_issues = [r for r in current_data.get_results()
if get_fingerprint(r) not in baseline_fps]
return len(new_issues)
result.level: it is optional, and CodeQL always omits itFor ready-to-use query templates, see {baseDir}/resources/jq-queries.md:
LEVEL_FN - the severity resolution every filtering query starts fromFor Python utilities, see {baseDir}/resources/sarif_helpers.py:
resolve_level() - Severity from the result or the rule it inherits fromnormalize_path() - Handle tool-specific path formatscompute_fingerprint() - Rule, normalized path, line, and messagededuplicate() - Remove duplicates across runsTwo SARIF fixtures live in {baseDir}/resources/fixtures, one with severity on the rules only and one with severity on the results. Each contains exactly one error, so a gate can be checked against a known answer before it is trusted.
development
Reviews a code target by launching a panel of specialist auditor agents and merging their reports. Use when asked to run a panel review.
development
Reviews the current branch's changes against its base branch as a pull request: correctness of new and modified code, test coverage for it, and documentation accuracy. Use when asked to review a branch, a diff, or a pull request.
tools
Runs an autonomous review-and-fix improvement loop over a Claude Code skill until a review comes back clean, with a cross-round findings ledger, escalation when fixes stop converging, and a mechanical scope guard. Reviews are performed by the plugin-dev skill-reviewer agent. Use to fix skill quality issues, iteratively refine a skill, or resume a loop after an escalation ('fix my skill', 'improve this skill until it passes review', 'skill improvement loop'). NOT for a one-time review — use the plugin-dev skill-reviewer agent directly.
tools
Runs an autonomous review-and-fix improvement loop over the current branch's changes until a PR review comes back clean, scoped mechanically to the directories the branch touched. Reviews are performed by an installed PR-review skill (default: pr-review-toolkit's review-pr). Use to fix review findings on a branch before opening or updating a pull request ('clean up this branch', 'fix this PR until review passes', 'run review-and-fix on my changes'). NOT for a one-time review — run the PR-review skill directly.