skills/exploring-replay-vision-observations/SKILL.md
Guides agents through pulling a Replay Vision scanner's observations, reading the findings, and acting on them — summarizing patterns across sessions, drilling into individual recordings, and turning real, corroborated issues into PostHog tasks, insights, or an investigating-replay hand-off. TRIGGER when: user wants to pull/read/triage Replay Vision observations, asks "what has my scanner found", wants to act on or summarize scanner findings, turn observations into tasks/work, or points at a /replay-vision/<scanner-id> URL. DO NOT TRIGGER when: creating or sizing a scanner (use creating-replay-vision-scanners), running a one-off scan you don't then analyse, or authoring a signals scout.
npx skillsauth add posthog/ai-plugin exploring-replay-vision-observationsInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
A scanner is a standing LLM probe over session recordings; each time it runs against a session it records one observation. This skill is about the other half of the loop — reading what the scanners have found and doing something useful with it. For creating or sizing scanners, use [[creating-replay-vision-scanners]].
(scanner, session).scanner_result. Its shape depends on the scanner's scanner_type, but it always
carries a confidence:
monitor → a verdict (yes / no / inconclusive) plus an open-ended observation.classifier → one or more tags from the scanner's label set.scorer → a numeric score on the scanner's scale.summarizer → a free-text summary (optionally with facet embeddings).succeeded observations carry a finding. Triage the rest by status/error_reason (see below).If a scanner has emits_signals: true, its observations also feed the Signals pipeline and may surface as
Inbox signal reports (clusters of related findings). When the user's intent is "work the reports", that's
the inbox path — see Acting on findings below.
If the user gave a /project/<id>/replay-vision/<scanner-id> URL, that path segment is the scanner ID.
Otherwise list them with vision-scanners-list and pick the relevant one.
Then call vision-scanners-get to read its configuration before reading results — the scanner_type and
scanner_config.prompt tell you how to interpret scanner_result (a verdict field only makes sense once you
know it's a monitor; a score only means something against the scorer's scale).
Pick the axis that matches the question:
vision-scanners-observations-list (the workhorse). Filter to
status=succeeded to get only sessions with a finding, then narrow by verdict (monitors) or tags
(classifiers). Scorers aren't filtered by score — rank them with order_by=-result_score instead. Use
order_by (e.g. -result_score, -completed_at) to surface the strongest hits first.vision-observations-list (the session_id query
parameter is REQUIRED). Use this while investigating a single recording.vision-scanners-observations-get or vision-observations-retrieve —
returns the frozen scanner_snapshot (config at run time) and the complete scanner_result, including any
event citations that link the finding back to specific events in the recording.Triage status so you don't mistake a non-result for "nothing wrong":
| status | meaning | typical error_reason |
| --------------------- | ------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- |
| succeeded | has a scanner_result | — |
| ineligible | session couldn't be analysed — a normal outcome, not an error | too_short, no_recording, too_inactive, too_long, no_events |
| failed | the scan errored | provider_rejected, validation_failed, rasterization_failed, provider_transient, internal_error |
| pending / running | still in flight | — |
A scanner that looks like it "found nothing" is often producing mostly ineligible observations — check the
mix before concluding.
verdict: yes; treat inconclusive as a weak signal. The observation text is the
substance.tags to see the distribution of what's happening across sessions.Weight by confidence, and don't over-index on a single observation. To understand a specific hit, take its
session_id and either cross-reference other scanners (vision-observations-list) or drill into the actual
recording with the [[investigating-replay]] skill and the session-recording MCP tools.
To test a scanner's lens against a specific session that doesn't have an observation yet, trigger one on demand
with vision-scanners-scan-session — it's async (minutes; rasterising the recording + the LLM call are slow)
and, like all observations, runs at most once per (scanner, session).
Match the action to the user's intent, and corroborate before you create work:
session_ids
(e.g. "12 of 40 succeeded observations flagged checkout confusion; sessions A, B, C"). Cite, don't assert.insight or notebook to track its
frequency, bundle the supporting recordings into a session-recording playlist so a human can watch the
evidence, and add an annotation if it marks a regression. There is no MCP tool to open a PostHog
task directly — to route a finding into tracked work, use the Inbox path below (for signal-emitting
scanners) or hand the summary to a human or coding agent to act on. Group by distinct issue, not per
observation.inbox-reports-list + inbox-report-artefacts-list (the report's work log is the
evidence). See the [[inbox-exploration]] skill; that path also records your work against the report.The discipline that matters: a single observation is one model's judgment on one recording. Confirm a finding reproduces across observations (or against the raw recording) before turning it into a task, an alert, or a claim — the same rigor the signals pipeline applies before it promotes observations to a report.
succeeded observations have a scanner_result — everything else is triage metadata.ineligible ≠ failed. Ineligible is a normal terminal outcome (e.g. the recording was too short), not
a bug to chase.(scanner, session) — re-scanning a session that already has any observation
(even ineligible/failed) is a no-op.scanner_snapshot it ran under, so older
observations may reflect a previous prompt/config (scanner_version).vision-quota-retrieve
before triggering a batch of them.data-ai
Signals scout for PostHog Tasks, the agent work items a project runs. Two lenses: delivery health (runs failing, clustered by repository and error class, and retry storms) every run, and on a slower rotation demand (recurring asks across human-authored tasks that point at a product gap). Skips the scout fleet's own run rows.
devops
Signals scout for the PostHog Conversations (support inbox) product. Watches the `$conversation_*` ticket-lifecycle events for support-delivery regressions — SLA breach-rate steps, first-response latency blowouts, backlog inflow-vs-resolution imbalance, and channel / assignment concentration — and files each dated regression as a report. Complements the per-ticket product-feedback signals the emission pipeline already fires; does not re-surface individual ticket content.
development
Populates and maintains a project's data catalog (semantic layer): canonical metrics, trust marks (certifications) on warehouse tables/views, and reviewed table relationships. Use when asked to set up / seed / bootstrap the data catalog or semantic layer, to catalog a project's metrics, to certify or deprecate data sources, to propose or review table joins, or to work through the proposal review queue. To *use* an existing catalog to answer a business-number question, see querying-posthog-data instead. Trigger terms: data catalog, semantic layer, canonical metric, certify table, deprecate source, relationship proposal, metric drift, review queue.
tools
Investigate logs in a PostHog project: verify a service or deployment is healthy, explain an error spike, triage an incident, or understand what a log stream is saying. Use when the user asks to "check the logs", asks whether a service, deploy, release, or change is working or broke anything, asks why errors are up or what changed, or wants the root cause of failures visible in logs. Routes the logs MCP tools (services overview, pattern mining, before/after pattern diffing, bucketed counts, facets, raw rows) so investigations start from summaries instead of raw rows or hand-written SQL over the logs table.