Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@
},
"metadata": {
"description": "Tsuga toolkit for AI coding agents: one plugin with the tsuga CLI driver, live-platform investigation, dashboards, incident workflows, OpenTelemetry instrumentation, Collector, signal-choice, telemetry debug, and audit skills.",
"version": "0.9.1"
"version": "0.10.0"
},
"plugins": [
{
Expand Down
2 changes: 1 addition & 1 deletion plugins/tsuga/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"name": "tsuga",
"description": "Tsuga observability plugin: the `tsuga` CLI driver (commands, TQL syntax, aggregation bodies, counter math, deep links, cloud/k8s translators); live-platform investigation for service health, errors, latency, and monitor coverage; dashboard building; incident orchestration; OpenTelemetry SDK, Collector, OTTL, signal-choice, telemetry debug, and audit skills; and meta-skills for building and validating skill bundles.",
"version": "0.9.1",
"version": "0.10.0",
"author": {
"name": "Tsuga Engineering",
"email": "engineering@tsuga.com"
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -21,8 +21,8 @@ Phase 1 — build:
Use the `$build-incident-history` skill. Specifically:
1. Read `${CLAUDE_PLUGIN_ROOT}/skills/build-incident-history/SKILL.md` and every file under `${CLAUDE_PLUGIN_ROOT}/skills/build-incident-history/references/`.
2. Execute the phases in `PROCEDURE.md` in order. Do NOT skip Phase 0 (sanity check) or Phase 5 (verification).
3. Fan out per-incident SUMMARY.md writing to parallel subagents — batches of 10–20. Each subagent gets one INC-id, the template, and the lessons doc. Prompt template is in `SUBAGENT_PROMPT.md`; copy verbatim, substitute `{inc_id}` and `{company}`.
4. Before subagent fan-out, optionally run Phase 2 (per-incident helper extraction) to pre-digest raw inputs into `/tmp/incident-extracts/<inc_id>/`. This is faster than having each subagent parse the raw JSON.
3. Optionally run Phase 2 (per-incident helper extraction) first, pre-digesting raw inputs into `/tmp/incident-extracts/<inc_id>/`. This is faster than having each subagent parse the raw JSON.
4. Then fan out per-incident SUMMARY.md writing to parallel subagents — batches of 10–20. Each subagent gets one INC-id, the template, and the lessons doc. Prompt template is in `SUBAGENT_PROMPT.md`; copy verbatim, substitute `{inc_id}` and `{company}`.
5. After fan-out, run Phase 4 to emit `_inventory.csv`.

Phase 2 — health check:
Expand Down
4 changes: 2 additions & 2 deletions plugins/tsuga/skills/build-incident-history/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,9 +13,9 @@ Procedure for bootstrapping the `incident-history` skill from raw incident mater

```
skills/incident-history/references/incidents/
├── _inventory.csv ← index: incident-id, title, declared_at, last_iso, service, team, severity
├── _inventory.csv ← index: incident_id, title, declared_at, last_iso, severity, affected_team, affected_services
├── INC-0001/
│ ├── SUMMARY.md ← the canonical ~300-line dossier
│ ├── SUMMARY.md ← the canonical dossier (see SUMMARY_TEMPLATE.md for the length target)
│ └── metadata.json ← {incident_id, declared_at, last_iso, title, severity}
├── INC-0002/
│ └── …
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -26,7 +26,7 @@ inputs/incidents/
└── …
```

## `metadata.json` — required fields
## `metadata.json` — fields

```json
{
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@ Translation table the subagent MUST follow:
| `get-metric name=X` | `tsuga metrics get X` |
| `list-monitors` / `get-monitor id=X` | `tsuga monitors list` / `tsuga monitors get X` (note the plural "monitors"!) |
| `list-dashboards` / `get-dashboard id=X` | `tsuga dashboards list` / `tsuga dashboards get X` |
| `list-routes`, `list-teams`, `list-services`, `list-notification-rules` | all singular→plural: `tsuga routes list`, `tsuga teams list`, etc. |
| `list-routes`, `list-teams`, `list-services`, `list-notification-rules` | all singular→plural: `tsuga log-routes list`, `tsuga teams list`, etc. |
| `aggregate-scalar dataSource=logs aggregate=count filter="X"` | heredoc into `/tmp/q.json` + `tsuga aggregation scalar -f /tmp/q.json` |
| `aggregate-timeseries dataSource=metrics aggregationWindow=5m …` | heredoc + `tsuga aggregation timeseries -f /tmp/q.json`, body has `aggregationWindow: "5m"` |

Expand All @@ -46,7 +46,7 @@ The data source is "spans" in the TQL sense but the CLI command is `traces searc

### 4. Singular vs plural resource names

`tsuga monitor get X` is wrong. The CLI follows the pattern `tsuga <resources-plural> <verb>`: `tsuga monitors get`, `tsuga dashboards list`, `tsuga routes get`, `tsuga teams list`, `tsuga services get`, etc. Always plural.
`tsuga monitor get X` is wrong. The CLI follows the pattern `tsuga <resources-plural> <verb>`: `tsuga monitors get`, `tsuga dashboards list`, `tsuga log-routes get`, `tsuga teams list`, `tsuga services get`, etc. Always plural.

### 5. `rtk` prefix is noise in docs

Expand All @@ -56,7 +56,7 @@ The RTK hook rewrites commands transparently at execution time. Writing `rtk tsu

- `timeRange` in the JSON body requires **Unix seconds integers**, not relative strings like `"-1h"`. Use the helper:
```bash
FROM=$(date -u -v-1H +%s); TO=$(date -u +%s) # macOS
TO=$(date -u +%s); FROM=$((TO - 3600))
# Linux: FROM=$(date -u -d '1 hour ago' +%s); TO=$(date -u +%s)
```
- `groupBy` goes at body level, not inside query items: `"groupBy": [{"fields": ["context.cluster_id"], "limit": 10}]`.
Expand All @@ -77,7 +77,7 @@ If the responder's `tsuga/commands.txt` is missing or empty for an incident, the

- Leave the Diagnostic path section empty except for a one-line note:
> _No command log captured for this incident. Reconstruction would be invention — flagged in Confidence._
- Add a Confidence note at the bottom of the SUMMARY.md: "low — Diagnostic path not recoverable from inputs."
- Add a `## Confidence` section at the bottom of the SUMMARY.md: "low — Diagnostic path not recoverable from inputs."

A SUMMARY.md with an honest empty section is far more useful than one with hallucinated probes, because the retrieval layer can filter out low-confidence entries from analogue search.

Expand Down Expand Up @@ -110,7 +110,7 @@ When writing a Diagnostic path probe, use the OR-match idiom:
tsuga logs search --query "(context.service.name:app-order-ingest OR context.service.name:ingest) level:ERROR" --from -1h
```

If the responder's original probe used only one form and that caused them to miss a subset, note this in the Findings — it's the most common source of "we couldn't see half the problem" confusion.
If the responder's original probe used only one form and that caused them to miss a subset, note this in that probe's `Finding:` line — it's the most common source of "we couldn't see half the problem" confusion.

### 13. Engine roles are not first-party services

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -51,13 +51,15 @@ Example helper script skeleton (keep outside this skill; it's your ops-side glue
```bash
for inc in "$INPUTS"/INC-*/; do
inc_id=$(basename "$inc")
rm -rf "/tmp/incident-extracts/$inc_id"
mkdir -p "/tmp/incident-extracts/$inc_id"

# Flatten Slack thread to one line per message, most important first
jq -r '.messages | sort_by(.ts) | .[] | "\(.ts) [\(.user_name)] \(.text)"' \
"$inc/slack/thread-*.json" > "/tmp/incident-extracts/$inc_id/slack-flat.txt" 2>/dev/null
jq -r '.messages | sort_by(.ts) | .[] | "\(.ts) [\(.user_profile.real_name // .username // .user)] \(.text)"' \
"$inc"/slack/thread-*.json > "/tmp/incident-extracts/$inc_id/slack-flat.txt" 2>/dev/null

# Flatten PRs to title/author/merge-date/url
# Flatten PRs to title/author/merge-date/url. `mergedAt` is only present if the capture
# requested it (`gh pr list --json number,state,title,mergedAt,url`).
jq -r '.[] | "#\(.number) [\(.state)] \(.title) (merged=\(.mergedAt // "n/a")) \(.url)"' \
"$inc/github/prs.json" > "/tmp/incident-extracts/$inc_id/prs-flat.txt" 2>/dev/null

Expand All @@ -70,7 +72,7 @@ Result: per-incident helper files the subagent reads instead of raw JSON. Saves

## Phase 3 — fan out subagents

**One subagent per incident. Run in parallel — aim for batches of 10–20 at a time.** The per-service fan-out in `knowledge-company` used 32 in parallel; incident-history can match that or go wider since each subagent has less to do.
**One subagent per incident, run in parallel.** Batch 10–20 at a time; each subagent has less to do than a service dossier, so wider waves are fine if the host tolerates them. Keep the batch size consistent with `SUBAGENT_PROMPT.md`.

Each subagent gets:

Expand All @@ -82,7 +84,7 @@ Each subagent gets:
- The lessons doc: `${CLAUDE_PLUGIN_ROOT}/skills/build-incident-history/references/LESSONS.md`
- The verification doc: `${CLAUDE_PLUGIN_ROOT}/skills/build-incident-history/references/VERIFICATION.md`

The subagent's contract is in `SUBAGENT_PROMPT.md` — do not retype it; copy verbatim and substitute only the `{inc_id}` placeholder.
The subagent's contract is in `SUBAGENT_PROMPT.md` — do not retype it; copy verbatim and substitute the `{inc_id}` and `{company}` placeholders.

## Phase 4 — write `metadata.json` + `_inventory.csv`

Expand All @@ -95,7 +97,7 @@ OUTPUT=./skills/incident-history/references/incidents
{
echo "incident_id,title,declared_at,last_iso,severity,affected_team,affected_services"
for f in "$OUTPUT"/INC-*/metadata.json; do
jq -r '[.incident_id, .title, .declared_at, .last_iso, .severity, .affected_team, (.affected_services | join(";"))] | @csv' "$f"
jq -r '[.incident_id, .title, .declared_at, .last_iso, .severity, .affected_team, ((.affected_services // []) | join(";"))] | @csv' "$f"
done
} > "$OUTPUT/_inventory.csv"
```
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ Copy verbatim. Substitute `{inc_id}` and `{company}`. The orchestrator fans out

## Prompt template

```
````
Write a SUMMARY.md for {company} incident `{inc_id}`. This is one of many per-incident dossiers being generated in a single build of the `incident-history` skill.

**Output file:** `skills/incident-history/references/incidents/{inc_id}/SUMMARY.md` (create parent dir with `mkdir -p`).
Expand Down Expand Up @@ -42,7 +42,7 @@ If the responder's original commands are not recoverable, **do NOT invent them**

> _No command log captured for this incident. Reconstruction would be invention — flagged in Confidence._

And add a Confidence note at the bottom: "low — Diagnostic path not recoverable from inputs."
And add a `## Confidence` section at the bottom: "low — Diagnostic path not recoverable from inputs."

**Required structure (canonical section list):**

Expand Down Expand Up @@ -72,7 +72,7 @@ And add a Confidence note at the bottom: "low — Diagnostic path not recoverabl
F="skills/incident-history/references/incidents/{inc_id}/SUMMARY.md"

# Forbidden MCP-tool shapes
grep -nE "^(search-logs|search-spans|list-metrics|get-metric|list-monitors|get-monitor|list-dashboards|get-dashboard|aggregate-scalar|aggregate-timeseries|list-log-patterns|list-new-error-patterns|list-error-pattern-increases)\b" "$F"
grep -nE "^(search-logs|search-spans|list-metrics|get-metric|list-monitors|get-monitor|list-dashboards|get-dashboard|list-routes|list-teams|list-services|list-notification-rules|aggregate-scalar|aggregate-timeseries|list-log-patterns|list-new-error-patterns|list-error-pattern-increases)\b" "$F"

# Pseudo-syntax arg shape
grep -nE "\bquery=|\bfrom=-|\b to=now\b|\blimit=|\bfilter=|\baggregationWindow=|\bdataSource=" "$F" \
Expand All @@ -93,17 +93,17 @@ done
[ -f "skills/incident-history/references/incidents/{inc_id}/metadata.json" ] || echo "MISSING metadata.json"
```

All four must return zero / clean output.
All five must return zero / clean output.

**Return** a 2–3 sentence summary:
- Line count + whether Diagnostic path was recoverable (N probes) or not.
- Number of timeline events, monitors cited.
- Confidence level you'd assign this SUMMARY (high/medium/low) and the reason.
```
````

## Notes for the orchestrator

- **Batch size:** 10–20 in parallel. Incidents are smaller tasks than service dossiers; wider batches fit.
- **`{company}`:** substitute with the company name (e.g., "Tsuga"). Used in the Incident-at-a-glance framing.
- **Progress tracking:** for batches in the 100+ range, use TodoWrite entries per wave of 20. Mark each wave complete only after the 2-random-file execution gate of `VERIFICATION.md` passes for that wave — not the moment the subagent returns "done".
- **Progress tracking:** for batches in the 100+ range, use TodoWrite entries per wave of 20. Mark each wave complete only after the 5-random-file gate in `VERIFICATION.md` passes for that wave — not the moment the subagent returns "done".
- **Failures:** a subagent claiming success on a low-quality input (empty `tsuga/commands.txt`, no slack thread) must have produced a SUMMARY.md with explicit low-confidence notes — not fabricated content. Spot-check for this.
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ Keep the total length under ~400 lines per incident. If you hit 400, trim — mo

## Template body — copy verbatim, fill in placeholders

```markdown
````markdown
# {incident_id} — {title}

| Field | Value |
Expand Down Expand Up @@ -65,7 +65,7 @@ Finding: {one sentence about what the output revealed}.

```bash
# For aggregations, use the real CLI shape:
FROM=$(date -u -v-1H +%s); TO=$(date -u +%s) # macOS
TO=$(date -u +%s); FROM=$((TO - 3600))
cat > /tmp/q.json <<JSON
{
"timeRange": {"from": $FROM, "to": $TO},
Expand All @@ -84,7 +84,7 @@ tsuga aggregation timeseries -f /tmp/q.json
Finding: {…}.

Rules:
- Every command must parse as real `tsuga` CLI. See `LESSONS.md §"Commands must be tested"` for the full translation table and the forbidden-token grep.
- Every command must parse as real `tsuga` CLI. See `LESSONS.md §"Command-shape mistakes"` for the full translation table and the forbidden-token grep.
- No `rtk` prefix.
- If the responder ran the same probe three times with different time ranges, consolidate to one probe with a note about iteration.
- If a probe returned nothing useful, **keep it** — negative probes are the most valuable signal for analogue search ("tried X, didn't help").
Expand All @@ -108,18 +108,22 @@ Bullets. Each is one sentence about something the team learned or committed to c
- `order-ingest` dropped events for N minutes with no alert because the only P1 was on poll latency, not on batch throughput — add a throughput monitor.
- A customer had a hand-edited ingestion API key that the reconcile code path deleted on a schema migration — add a pre-reconcile diff/confirm step.

## Confidence (optional)

One line: `high` / `medium` / `low` plus the reason. Required when an input was missing or empty — the retrieval layer filters low-confidence entries out of analogue search.

## Commentary (optional)

Italicized one-paragraph running commentary from the responder's scratch notes, if any. Useful context but not load-bearing.
```
````

---

## Exemplar — abbreviated

Here is what a healthy SUMMARY.md looks like in miniature:

```markdown
````markdown
# INC-0001 — Metrics on acme-trading are slow

| Field | Value |
Expand Down Expand Up @@ -157,7 +161,7 @@ Finding: no errors. Logs + spans were healthy — metrics-only regression.

### Probe 2 — is the query engine reporting cold-path saturation?
```bash
FROM=$(date -u -v-1H +%s); TO=$(date -u +%s)
TO=$(date -u +%s); FROM=$((TO - 3600))
cat > /tmp/q.json <<JSON
{"timeRange":{"from":$FROM,"to":$TO},"dataSource":"metrics","queries":[{"aggregate":{"type":"percentile","percentile":95,"field":"query_below_day_duration_milliseconds"},"filter":"context.env:prod context.cluster_id:acme-trading"}],"formula":"q1","aggregationWindow":"5m"}
JSON
Expand All @@ -181,6 +185,6 @@ Finding: compaction runs were completing but `merged_doc_count` per run was 6x n
## Lessons / follow-ups
- Add a per-cluster compaction-lag monitor — the cluster-wide one did not fire because aggregate was fine.
- Consider a "cold-path fraction" metric on query-engine so the next one pages faster.
```
````

Ship at roughly this density.
Original file line number Diff line number Diff line change
Expand Up @@ -61,8 +61,18 @@ required=(
"## Lessons / follow-ups"
)
for f in "$OUTPUT"/INC-*/SUMMARY.md; do
grep -qE '^# INC-[0-9]+ — .+' "$f" || echo "MISSING or malformed H1 '# {incident_id} — {title}' in $f"
# Presence and order: a dossier with the right headings in the wrong order is not canonical.
prev=0
for h in "${required[@]}"; do
grep -qF "$h" "$f" || echo "MISSING '$h' in $f"
n=$(grep -nxF "$h" "$f" | head -1 | cut -d: -f1)
if [ -z "$n" ]; then
echo "MISSING '$h' in $f"
elif [ "$n" -lt "$prev" ]; then
echo "OUT OF ORDER '$h' in $f"
else
prev=$n
fi
done
done
```
Expand All @@ -78,11 +88,13 @@ The Diagnostic path sections must not contain MCP-tool pseudo-syntax or `rtk` pr
grep -rnE "^(search-logs|search-spans|list-metrics|get-metric|list-monitors|get-monitor|list-dashboards|get-dashboard|list-routes|list-teams|list-services|list-notification-rules|aggregate-scalar|aggregate-timeseries|list-log-patterns|list-new-error-patterns|list-error-pattern-increases)\b" "$OUTPUT"

# Pseudo-syntax argument shape
grep -rnE "\bquery=|\bfrom=-|\b to=now\b|\blimit=|\bfilter=|\baggregationWindow=|\bdataSource=" "$OUTPUT" \
| grep -v "/explorer?query=" # URL query params are fine
| grep -v '"aggregationWindow":' # inside JSON bodies is fine
| grep -v '"dataSource":'
| grep -v '"filter":'
# Strip URL params and JSON keys from each line first: dropping whole lines would mask a real
# violation that happens to share a line with a legitimate URL or key.
grep -rl . "$OUTPUT" | while IFS= read -r f; do
sed -E 's#/explorer\?query=[^ )"`]*##g; s#"(aggregationWindow|dataSource|filter)":##g' "$f" \
| grep -nE '\bquery=|\bfrom=-|\bto=now\b|\blimit=|\bfilter=|\baggregationWindow=|\bdataSource=' \
| sed "s#^#$f:#"
done

# rtk prefix
grep -rnE "^rtk |[[:space:]]rtk " "$OUTPUT" | head -20
Expand Down Expand Up @@ -135,11 +147,17 @@ awk -F, 'NR>1 && ($1=="" || $2=="" || $3=="" || $4=="") {print NR": "$0}' "$OUTP
`entrypoint.sh` filters the archive by `SNAPSHOT_AT` at container start. Simulate that filter to confirm `metadata.json` dates are actually usable:

```bash
# Pick an arbitrary incident with a known declared_at
inc=INC-0001
jq -r '.last_iso' "$OUTPUT/$inc/metadata.json" | xargs -I{} date -u -d {} +%s \
&& echo "OK: $inc last_iso parses as Unix seconds" \
|| echo "FAIL: $inc last_iso does not parse"
# Every incident, not just one. Parse with python so the check works on BSD and GNU alike, and
# so an empty last_iso fails instead of being skipped.
bad=0
for f in "$OUTPUT"/INC-*/metadata.json; do
iso=$(jq -r '.last_iso // empty' "$f")
if [ -z "$iso" ] || ! python3 -c 'import sys,datetime; datetime.datetime.fromisoformat(sys.argv[1].replace("Z","+00:00"))' "$iso" 2>/dev/null; then
echo "FAIL: $(dirname "$f") last_iso does not parse: '${iso:-<missing>}'"
bad=$((bad+1))
fi
done
[ "$bad" -eq 0 ] && echo "OK: every last_iso parses as an ISO-8601 timestamp"
```

**Pass:** every `metadata.json`'s `last_iso` parses to Unix seconds. If any don't, `entrypoint.sh` will silently drop those incidents on filter.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -36,7 +36,7 @@ Use the `$build-knowledge-company` skill. Specifically:
Phase 2 — health check:
Use the `$check-skill-health` skill. Specifically:
1. Run `${CLAUDE_PLUGIN_ROOT}/skills/check-skill-health/scripts/lint-all.sh skills/knowledge-company/` (structural checks, offline).
2. Then run with live execution: `${CLAUDE_PLUGIN_ROOT}/skills/check-skill-health/scripts/lint-all.sh --execute skills/knowledge-company/` (samples 5 random SERVICE_KNOWLEDGE.md files and runs the first `tsuga` command from each against prod telemetry).
2. Then run the sampling audit: `${CLAUDE_PLUGIN_ROOT}/skills/check-skill-health/scripts/lint-all.sh --execute skills/knowledge-company/` (samples random SERVICE_KNOWLEDGE.md files and audits whether the first `tsuga` command in each is read-only and well-shaped — it never executes them).
3. If any FAIL: do NOT hand-edit the affected file. Fix the root cause in the template / subagent prompt / lessons doc, regenerate the affected services via subagent, re-run both lint passes. Iterate until `lint-all.sh --execute` returns exit code 0.
4. WARNs are informational — read them, decide whether to fix or annotate.

Expand Down
Loading
Loading