Guidewire Observability And Incident ResponseSAFE
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
Overview
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
4f83675ca38aOBSERVED · 2026-10-08Host compatibility
What the documentation claims. We have not run a compatibility test.
| Host | Status | Notes |
|---|---|---|
| claude-code | mentioned |
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
--- name: guidewire-observability-and-incident-response description: Operate a Guidewire Cloud API integration in production — define SLIs/SLOs for token availability, bind success rate, FNOL p99 latency; route alerts so the on-call gets paged for real outages and never for transient noise; triage 401 spikes, 409 storms, 429 saturation, scope drift, and Gosu OOM cascades from signal to recovery in 15 minutes or less. Use when designing a dashboard for a new integration, writing the on-call runbook, or running a post-incident review. Trigger with "guidewire observability", "guidewire slo", "guidewire on-call", "guidewire 401 spike", "guidewire 409 storm", "guidewire incident". allowed-tools: Read, Write, Edit, Bash(curl:*), Bash(jq:*), Grep, Glob version: 1.26.0 license: MIT author: Jeremy Longshore <[email protected]> compatibility: Designed for Claude Code tags: - guidewire - observability - sli-slo - incident-response - on-call - runbook --- # Guidewire Observability and Incident Response ## Overview Run a production Guidewire Cloud API integration with the dashboards, alerts, and runbooks an on-call engineer can act on at 3am. This skill consolidates the operational layer: what to measure, what to alert on, how to triage the top five incident classes, and how to close the loop with a post-incident review that prevents recurrence rather than performing root-cause theater. Five operational failures this skill prevents: 1. **Vanity dashboards** — graphs of total request count, no per-endpoint p99, no SLI burn rate; on-call sees green during a real outage because the right thing was never measured. 2. **Alert fatigue** — every transient `5xx` pages someone; in three weeks the team mutes the channel; a real incident two weeks later goes unnoticed for an hour. 3. **No triage tree** — on-call wakes up to "401 spike", does not know whether to rotate a secret, restart the integration, or call GCC; loses 20 minutes Googling. 4. **Skipped post-incident review** — the same root cause produces three incidents in a quarter because no one wrote down the action items from the first one. 5. **Common-errors table living in nine Slack threads** — operators cannot find the recovery for an error they have seen before, ask the same question in #ops, the answer takes 45 minutes. ## Prerequisites - A working integration emitting structured logs and metrics to a backend (Datadog, Grafana+Prometheus, New Relic, Splunk, or equivalent) - Access to a paging system (PagerDuty, Opsgenie, VictorOps) with on-call rotations defined - The `integration_audit` table from `guidewire-security-and-rbac` populated — incident triage depends on knowing what the integration tried to do - `correlation_id` propagated end-to-end through every Cloud API call ## Instructions Build the operational layer in this order. Each step targets one of the five operational failures listed in Overview. ### 1. Define SLIs that measure user-visible behavior Track what users care about, not what is easy to measure. For a Guidewire integration, the SLIs that matter: | SLI | Measurement | Target | |---|---|---| | Token-endpoint availability | success rate of `/oauth/token` over a rolling 5min window | ≥99.9% | | Cloud API write success | `2xx` rate on POST/PATCH calls, excluding 4xx caller errors | ≥99.5% | | Bind success rate | `bound` / `quoted` ratio over rolling 1h, excluding referrals | ≥98% | | FNOL intake p99 latency | end-to-end time from inbound event to `claim.created` log | ≤2s | | Quote-to-bind median latency | median time from quote-call to bind-success | ≤30s | `5xx` from upstream Cloud API counts against availability; `4xx` from caller bugs (validation failures, scope mismatches) does not — those are caller errors, not integration outages. Distinguishing the two is the single biggest cause of either alert fatigue or missed incidents, depending on which way the bias goes. ### 2. Burn-rate-based alerting, not threshold alerting Alerting on "error rate > 1%" pages the on-call every time a transient `502` happens. Alert on **SLO burn rate** instead — the rate at which the error budget is being consumed. ``` Fast burn: 2% error budget consumed in 1 hour → page immediately Slow burn: 5% error budget consumed in 6 hours → page during business hours Trickle: 10% error budget consumed in 3 days → ticket, not page ``` A 5-minute outage that consumes 1% of the monthly budget should not page; a 20-minute outage that consumes 4% should. Burn-rate alerts encode this naturally. ### 3. Top-five triage trees Each tree is the ~5-step decision sequence on-call follows from signal to recovery. Memorize the entry signal; the body is in the runbook. **T1: 401 spike on Cloud API calls** ``` 401s > 1% for 5min ├─ Check token age in cache: is the integration refreshing? → if no, restart token-cache process ├─ Decode a recent token, verify exp > now + 60s → if no, clock skew or aggressive proxy caching ├─ Check GCC: is the Service Application enabled? → if no, talk to tenant admin └─ Check secret-rotation history: was a secret rotated in the last 24h? → if yes, run dual-secret swap or restart ``` **T2: 409 storm on PATCH calls** ``` 409s > 5% for 5min ├─ Are concurrent writers expected? → if yes, scale checksum round-trip retry budget ├─ Is one resource-id producing all 409s? → if yes, it's a hot key; coordinate writes via queue └─ Is the client retrying the bare PATCH instead of the GET-PATCH cycle? → fix the retry layer ``` **T3: 429 saturation on Cloud API or Hub** ``` 429s > 1% for 10min ├─ Is the Hub /oauth/token endpoint 429ing? → token cache missing single-flight gate; deploy fix immediately ├─ Is the data-plane API 429ing? → check tenant quota in GCC, request increase if legitimate growth └─ Is one customer driving the saturation? → tenant-side rate limit on the integration's intake ``` **T4: Scope drift detected** ``` Scope-drift alert from auth refresh ├─ What scop
Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | PASS |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
4f83675ca38afull audit observations/trust-audit/skill/jeremylongshore__guidewire-observability-and-incident-response.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | 4f83675ca38a | SAFE | B | 89 | first audit |
Questions
What does the Guidewire Observability And Incident Response skill do?
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
Is Guidewire Observability And Incident Response safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Guidewire Observability And Incident Response access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
Which assistants does Guidewire Observability And Incident Response work with?
Its documentation mentions claude-code. That is what the text claims, not a compatibility test we ran.
How current is this page?
The grade is for one exact copy of the source (4f83675ca38a), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.