Chuyển bộ slide markdown thành bài thuyết trình HTML một file, điều hướng bằng phím, có chế độ trình chiếu kèm ghi chú người trình bày.
---
name: md-slides
description: Converts a markdown deck (slides separated by `---` HR boundaries or by `# ` H1 headings, with optional `<!-- notes: ... -->` presenter notes blocks) into a single-file HTML presentation with arrow-key / space / PgDn / PgUp / Home / End / P / Esc keyboard navigation, presenter mode (split view with current slide + speaker notes + clock + next-slide preview), URL-hash deep linking, and `@media print` page-per-slide for PDF export. Triggers when the markdown-html-orchestrator classifies an input as SLIDES, or when invoked directly via /cs:md-slides. Reuses md-document's markdown parser for slide-body rendering and reads design-system tokens via config_loader.py. Refuses if input has no clear slide boundaries, produces a 1-slide deck, or `--strict-notes` is on with < 50% notes coverage. Use after orchestrator routing.
version: 2.10.3
author: Alireza Rezvani
license: MIT
tags: [markdown, html, slides, deck, presenter-mode, keyboard-nav, print-to-pdf, single-file, design-system]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# md-slides — Markdown deck → single-file HTML presentation
The slide-deck converter. Reads a markdown deck (HR or H1 boundaries, optional presenter notes), emits a single-file HTML presentation that runs in any browser with keyboard navigation, presenter mode, and print-to-PDF.
Three stdlib tools pipeline together:
```
slide_splitter.py → presenter_notes_parser.py → deck_html_renderer.py
(md → ordered (extract <!-- notes: (slides + design-system
slides with --> blocks, attach tokens → single-file
titles) per slide) HTML with keyboard nav)
```
## When to invoke
| Symptom | Action |
|---|---|
| `markdown-html-orchestrator` routes input as SLIDES | Invoke this skill |
| User runs `/cs:md-slides <path>.md` directly | Invoke this skill |
| Input has 3+ `---` HR lines OR 5+ H1 headings with short bodies | Invoke this skill |
| Input is a long-form spec | Route to `md-document` instead |
| Input is a code review | Route to `md-review` instead |
| Input has no clear slide boundaries | Refuse, route to `md-document` |
| Input would produce 1 slide | Refuse (it's a poster) |
## Pipeline
```bash
# 1. Split slides on --- or H1 (auto-detect by default)
python3 markdown-html/skills/md-slides/scripts/slide_splitter.py \
--input <path>.md --output /tmp/slides.json
# 2. Extract <!-- notes: ... --> blocks from each slide
python3 markdown-html/skills/md-slides/scripts/presenter_notes_parser.py \
--slides /tmp/slides.json --output /tmp/deck.json
# 3. Render single-file HTML deck
python3 markdown-html/skills/md-slides/scripts/deck_html_renderer.py \
--slides /tmp/deck.json --title "My Talk" --output deck.html
```
## What ships in the HTML
- **All slides as `<section class="slide">`** — one visible at a time, controlled by JS
- **Keyboard nav** — `→` / `Space` / `PgDn` advance; `←` / `PgUp` previous; `Home`/`End` jump; `P` presenter mode; `Esc` exits presenter
- **URL-hash deep linking** — `#3` jumps to slide 3; browser back/forward walks slides; share `deck.html#5` to send someone directly there
- **Progress bar** — 3px at top showing position through the deck
- **Slide counter** — bottom-right ("3 / 12")
- **Presenter mode** (P key) — splits the window: current slide on left (60% width), panel on right with clock + speaker notes + next-slide preview
- **Print stylesheet** — `Cmd+P` produces a PDF with one slide per page
- **`@media (prefers-reduced-motion: reduce)`** honored
- **12 brand CSS custom properties** from design-system; design_style affects layout density
- **Reuses md-document's markdown parser** — slide bodies render with consistent paragraph/list/code/table/callout handling
## Hard rules
1. **Refuses input with no clear slide boundaries.** Auto mode needs ≥ 3 HR lines or ≥ 5 H1 headings. Otherwise exit 6 — route to md-document.
2. **Refuses 1-slide decks.** That's a poster, not a deck. Exit 5.
3. **Refuses input < 100 lines.** Same Shihipar threshold as all converters.
4. **Refuses without onboarding.** Same gate as every converter.
5. **`--strict-notes` refuses < 50% notes coverage.** A deck where most slides have no notes isn't set up for presenter mode. Exit 7.
6. **Soft-warns slides > 40 source lines.** Signal-to-noise; renders anyway but surfaces the count.
7. **Single-file output.** All CSS + JS inline. Only external is Google Fonts CSS. Prism.js is opt-in via `--syntax`.
8. **No JS framework runtime.** Vanilla JS + keyboard event handlers, no React/Vue/Svelte.
## Forcing-question library (Matt Pocock grill discipline)
1. **Is this actually a deck, or a long document?** Recommended: if you can't draw clear slide boundaries, it's not a deck. Canon: Tufte *Cognitive Style of PowerPoint*.
2. **HR (`---`) or H1 boundaries?** Recommended: HR for typical decks; H1 for outline-driven decks. Canon: Marp / reveal.js / pandoc convergence.
3. **Will it be presented live or distributed for self-paced reading?** Recommended: live → need presenter notes; self-paced → notes optional. Canon: Weinschenk *100 Things Every Presenter Needs to Know*.
4. **Is there any slide over 40 source lines?** Recommended: split it. Canon: NN/g — audience attention drops past ~6 bullets / 200 words.
5. **Is `--syntax` needed?** Recommended: only for decks with substantial code blocks. Default off. Canon: single-file shareability discipline.
## Distinct from
- **`md-document`** — that's one continuous document. This is N discrete slides.
- **`md-review`** — that renders diff hunks + annotations. This renders prose slides.
- **`marketing/landing/`** — that's a landing page, not a deck.
- **Keynote / PowerPoint** — those are graphic-design tools. This is for markdown-authored decks projected from a browser.
## Output artifact
`{default_output_dir}/deck-{slug}.html` (path resolved by orchestrator's `output_path_resolver.py`; collision suffix `-2`, `-3`, … by default).
## References
- Shihipar — *Claude Code HTML output* (Medium, 2026), Tier 3 use case "Slide Decks"
- Reynolds — *Presentation Zen* (less is more discipline)
- Atkinson — *Beyond Bullet Points* (the bullet-heavy failure mode)
- Tufte — *The Cognitive Style of PowerPoint* (the polemic)
- reveal.js / Big / Marp — convergent markdown-to-deck conventions
- See `references/` for full citations (presentation_ux, keyboard_nav_patterns, single_file_deck_conventions)
FILE:assets/md_slides_template.html
<!DOCTYPE html>
<!--
md_slides_template.html — Reference shape for deck_html_renderer.py output.
Documents the canonical single-file deck layout. The renderer produces this
same shape dynamically from a parsed deck (slide_splitter + presenter_notes
_parser) plus the design-system config.
See: deck_html_renderer.py for the live implementation.
-->
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>{{DECK_TITLE}}</title>
<link rel="stylesheet" href="https://fonts.googleapis.com/css2?family={{HEADING_FONT}}&family={{BODY_FONT}}&display=swap">
<!-- Prism is OPT-IN for decks (off by default) — pass --syntax to enable -->
<!-- <link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/prismjs@1.29.0/themes/prism-tomorrow.min.css"> -->
<style>
:root {
/* 12 brand tokens from design-system.derived_palette */
--md-bg: {{BG}}; --md-surface: {{SURFACE}}; --md-border: {{BORDER}};
--md-text: {{TEXT}}; --md-text-muted: {{TEXT_MUTED}};
--md-accent: {{ACCENT}}; --md-accent-soft: {{ACCENT_SOFT}};
--md-code-bg: {{CODE_BG}}; --md-link: {{LINK}}; --md-link-hover: {{LINK_HOVER}};
--md-success: {{SUCCESS}}; --md-warn: {{WARN}};
--md-font-heading: 'Inter', system-ui, sans-serif;
--md-font-body: 'Inter', system-ui, sans-serif;
}
/* One slide visible at a time; print stylesheet shows them all */
.slide { display: none; position: absolute; inset: 0; padding: 4vh 8vw; }
.slide.active { display: flex; flex-direction: column; justify-content: center; }
.progress { position: fixed; top: 0; height: 3px; background: var(--md-accent); }
.presenter-panel { display: none; position: fixed; right: 0; width: 40vw; }
body.presenter .deck { width: 60vw; }
body.presenter .presenter-panel { display: flex; }
@media print {
.slide { display: flex !important; position: relative; height: 100vh; page-break-after: always; }
.chrome, .progress, .presenter-panel { display: none !important; }
}
/* ... BASE_CSS ... */
</style>
</head>
<body class="style-{{DESIGN_STYLE}}">
<!-- Progress bar at top -->
<div id="progress" class="progress"></div>
<!-- All slides — visibility controlled by JS toggling .active -->
<div class="deck">
<section class="slide active" id="slide-1" data-notes="">
<h1>{{SLIDE_1_TITLE}}</h1>
<!-- body markdown rendered as HTML (paragraphs / lists / code / tables / callouts) -->
</section>
<section class="slide" id="slide-2" data-notes="Speaker notes for slide 2 go here. Multi-line.">
<h1>{{SLIDE_2_TITLE}}</h1>
<ul>
<li>Point one</li>
<li>Point two</li>
</ul>
</section>
<!-- ...more slides... -->
</div>
<!-- Presenter view: appears when P is pressed -->
<aside class="presenter-panel" aria-label="Presenter view (toggle with P)">
<h3>Clock</h3>
<div class="clock">12:34:56</div>
<h3>Speaker notes</h3>
<div class="notes"></div>
<div class="next-preview">
<h4>Up next (#3)</h4>
<h1>{{NEXT_SLIDE_TITLE}}</h1>
</div>
</aside>
<!-- Bottom-right chrome: slide counter + presenter toggle -->
<div class="chrome">
<a href="#" onclick="event.preventDefault();
document.dispatchEvent(new KeyboardEvent('keydown', {key:'P'}))">P · presenter</a>
<span id="counter">1 / 5</span>
</div>
<!-- Inline vanilla-JS payload (~3 KB):
- keyboard nav: ← / → / Space / PgDn / PgUp / Home / End / P / Esc
- URL hash sync (#3 = slide 3; works for deep links + back button)
- presenter panel with clock + notes + next-slide preview
- prefers-reduced-motion honored throughout
-->
<script>
(function () {
"use strict";
var slides = document.querySelectorAll(".slide");
var current = 0;
function show(idx) {
idx = Math.max(0, Math.min(slides.length - 1, idx));
slides.forEach(function (s, i) { s.classList.toggle("active", i === idx); });
current = idx;
history.replaceState(null, "", "#" + (current + 1));
}
document.addEventListener("keydown", function (e) {
if (e.metaKey || e.ctrlKey || e.altKey) return;
switch (e.key) {
case "ArrowRight":
case "PageDown":
case " ": e.preventDefault(); show(current + 1); break;
case "ArrowLeft":
case "PageUp": e.preventDefault(); show(current - 1); break;
case "Home": e.preventDefault(); show(0); break;
case "End": e.preventDefault(); show(slides.length - 1); break;
case "p":
case "P": e.preventDefault();
document.body.classList.toggle("presenter"); break;
}
});
// Honor URL hash on load
if (location.hash) {
var n = parseInt(location.hash.slice(1), 10);
if (!isNaN(n)) show(n - 1);
} else { show(0); }
})();
</script>
</body>
</html>
FILE:references/keyboard_nav_patterns.md
# Keyboard Navigation Patterns
**Why this exists:** Every presenter expects ← / → to advance slides, Space to advance, and Esc to exit presenter mode. These conventions are 20 years old. This document records what `md-slides` honors and why each binding is wired the way it is.
## The keymap
| Key | Action | Source |
|---|---|---|
| **→ / Space / PgDn** | Next slide | reveal.js, Big, Spectacle, Keynote, PowerPoint — universal |
| **← / PgUp** | Previous slide | universal |
| **Home** | First slide | reveal.js / Big convention |
| **End** | Last slide | reveal.js / Big convention |
| **P** | Toggle presenter mode | reveal.js convention |
| **Esc** | Exit presenter mode (if active) | universal accessibility expectation |
| **Ctrl/Cmd + P** | Print to PDF (browser-native) | browser default; we don't intercept |
We deliberately do NOT bind:
- **Number keys (1-9)** for slide jump — too easy to hit accidentally during typing
- **F11** for fullscreen — browser default; we don't override
- **Touch gestures** — out of scope; click-to-advance is sufficient for the click-to-advance case
- **Vim keys (h/j/k/l)** — niche; the arrow keys are the universal expectation
## Implementation discipline
```js
document.addEventListener("keydown", function (e) {
if (e.metaKey || e.ctrlKey || e.altKey) return; // Don't fight browser shortcuts
switch (e.key) {
case "ArrowRight":
case "PageDown":
case " ":
e.preventDefault(); next(); break;
// ...
}
});
```
- **Modifier-aware**: we check `metaKey/ctrlKey/altKey` and bail. This means `Cmd+R` reloads (browser default), `Ctrl+P` prints (browser default), and we don't fight them.
- **`preventDefault()` on every match**: Space normally scrolls; we replace that with advance-slide.
- **`replaceState` for URL hash**: each slide change updates `#N` in the URL so deep links work + back button moves through slides naturally.
## URL hash deep linking
Every slide has `id="slide-N"`. The initial render reads `location.hash` and jumps to that slide on load. Sharing `deck.html#3` puts the reader directly on slide 3. Browser back/forward buttons walk slide history.
This works because we use `history.replaceState` (not `pushState`) for arrow-key navigation — otherwise every arrow click would push a new entry and back-button behavior would feel wrong.
## Accessibility
- **`@media (prefers-reduced-motion: reduce)`** — transitions and animations are suppressed when the user prefers reduced motion. The deck still works; it just doesn't animate.
- **Presenter panel uses `<aside aria-label="Presenter view (toggle with P)">`** — screen readers announce its purpose.
- **Slide counter is plain text** — `<span id="counter">1 / 5</span>` is announced by screen readers as the slide changes (we update the text content).
- **No focus traps** — keyboard users can Tab out of any interactive element naturally.
- **No JS-required content** — the slides are static `<section>` elements; JS just controls visibility. With JS off, the first slide renders correctly and the user can scroll through all of them.
## Sources
### 1. reveal.js — *Keyboard Bindings* (revealjs.com)
The convention-setter for web-based decks. Established: ← / → / Space / PgUp / PgDn / Home / End / Esc / F. We honor the subset our scope requires.
### 2. Tom MacWright — *Big* (github.com/tmcw/big)
Single-file deck tool. Validates the minimum viable keybinding set (arrow / space) and the URL-hash navigation pattern.
### 3. Spectacle (github.com/FormidableLabs/spectacle)
React-based deck framework. Same keymap conventions. Reinforces the arrow + space pairing.
### 4. WCAG 2.2 §2.1.1 *Keyboard* (w3.org/WAI/WCAG22)
All functionality must be operable through a keyboard interface. Our nav, presenter toggle, and Esc-to-exit all satisfy this.
### 5. WCAG 2.2 §2.4.3 *Focus Order* (w3.org/WAI/WCAG22)
Focus order must preserve meaning. We don't trap focus or jump focus across slide changes; default tab order applies.
### 6. NN/g — *Keyboard Accessibility* (Jakob Nielsen, 2024 update)
Best practices for keyboard-only navigation: don't fight modifier shortcuts, always provide a visible escape from modal states (P toggles back, Esc also works).
### 7. MDN — *KeyboardEvent.key* (developer.mozilla.org)
The standardized key value strings we match against ("ArrowRight" not "Right"; " " for space; "PageDown" not "PgDn"). Cross-browser-stable since 2018.
## Applied to `md-slides`
The JS payload in `deck_html_renderer.py` wires all of the above, in ~80 lines of vanilla JS, no framework. The presenter panel and progress bar update reactively as the user presses keys.
FILE:references/presentation_ux.md
# Presentation UX
**Why this exists:** The `md-slides` converter doesn't try to be Keynote or PowerPoint — those tools optimize for elaborate visual production. It optimizes for the case Shihipar's essay sketches: a deck written in markdown, exported to a single .html file, projected from a browser. This document records the UX choices behind that scope.
## What this skill is for
- A talk you wrote in markdown (intent, not graphic design)
- A board meeting deck assembled from a Notion doc you exported
- A workshop walkthrough where the speaker drives the pace
- A training session that needs presenter notes
- A meeting recap distributed as a printable PDF
## What it isn't for
- Marketing decks with elaborate motion graphics — use Figma / Keynote
- Pitch decks meant to dazzle — same
- Slides you'll edit visually after generation — they're generated artifacts, regenerate from the markdown
- Live-collaborative editing — single-author, single-snapshot
## Signal-to-noise discipline
The single biggest failure mode of agent-generated decks is **too much per slide**. A slide with 40+ source lines of markdown becomes 40+ visible lines of text — the audience reads instead of listening; the speaker becomes redundant.
`slide_splitter.py` warns on any slide whose source markdown exceeds 40 lines (configurable via `MAX_SLIDE_LINES`). It doesn't refuse — sometimes a long quote or code block legitimately needs the space — but it surfaces the count so the author can decide.
Default suggested decomposition: each idea = one slide. If a slide has 5+ bullet points or 200+ words of body text, it should probably split.
## What the renderer enforces visually
- **22px base font, 1.45 line-height** — projector-readable from the back row
- **3rem H1, 2.5rem H2** — slide titles are the visual anchor
- **8vw side padding, 4vh top/bottom** — generous margins so text doesn't crash the edges of the projection
- **One slide visible at a time** — no infinite scroll; the slide is the unit of attention
- **No transitions / animations by default** — `prefers-reduced-motion: reduce` honored; no fly-ins or fades to distract
- **Progress bar at top** — 3px high; tells the audience how far through they are
## The presenter-notes contract
`<!-- notes: ... -->` blocks attached to slides serve three audiences:
1. **The speaker** during the talk (visible in presenter view)
2. **The audience** if the deck is distributed afterward (notes stay accessible via the data attribute, can be extracted programmatically)
3. **The author** when revisiting the deck months later (notes preserve intent that the slide alone doesn't)
`--strict-notes` enforces ≥ 50% coverage when presenter mode is in use — a deck where most slides have no notes isn't really set up for presenter mode.
## Sources
### 1. Cliff Atkinson — *Beyond Bullet Points* (Microsoft Press, 2011, 4th ed.)
The case that bullet-heavy slides suppress audience attention. We don't enforce a bullet-cap, but the > 40-line warning targets the same failure pattern.
### 2. Garr Reynolds — *Presentation Zen* (New Riders, 2019, 2nd ed.)
The "less is more" discipline: high signal-to-noise per slide. Reynolds argues for one idea per slide; our default 40-line warning approximates this for markdown-authored decks (a single idea written in markdown rarely exceeds 40 lines).
### 3. Edward Tufte — *The Cognitive Style of PowerPoint* (Graphics Press, 2003)
The polemic against slide-as-document. Tufte's argument is that slides flatten hierarchical information; our split into clear slide units + the optional handout-via-print mode (each slide one page) honors his core complaint by keeping the slide and the handout cleanly separable.
### 4. Jakob Nielsen / NN/g — *PowerPoint Usability* (2011, updated 2024)
Empirical: audience attention drops sharply when a slide exceeds about 6 bullet points or about 200 words. The 40-source-line warning is a markdown-aware proxy for these thresholds.
### 5. Susan Weinschenk — *100 Things Every Presenter Needs to Know About People* (New Riders, 2012)
Specifically on the cognitive cost of reading-while-listening — the audience can't do both well. Our presenter-notes pattern lets the speaker put the depth in notes and keep the slide visually minimal.
### 6. Marp / reveal.js / pandoc-Beamer — markdown-to-slides conventions
The `---` HR boundary and `<!-- notes: ... -->` syntax we accept come from the convergent convention these tools have shipped for a decade. We're compatible by reading the same markup, not innovating new syntax.
### 7. Tom MacWright — *Big* (github.com/tmcw/big, MIT)
A single-HTML-file presentation tool that emphasizes ridiculously-large text (the slide title fills the viewport). We're less aggressive — we render markdown bodies — but we share the discipline of "one slide = one viewport, no scroll."
## Applied to `md-slides`
The renderer ships projector-readable defaults, one-slide-per-viewport layout, presenter-notes contract, soft-warn on > 40 lines per slide, and the print stylesheet for PDF export. No graphics tooling, no motion, no live collaboration — those are out of scope.
FILE:references/single_file_deck_conventions.md
# Single-File Deck Conventions
**Why this exists:** The orchestrator's single-file discipline document establishes the rule for the whole domain. This document records the md-slides-specific decisions — what's different about decks vs documents, and what the single-file constraint means specifically for presentation artifacts.
## The contract
One `.html` file, projected from any browser. Externals limited to:
- **`fonts.googleapis.com`** — Google Fonts CSS for typography
- **`cdn.jsdelivr.net`** — Prism.js, ONLY when `--syntax` is explicitly passed (opt-in; off by default)
Why Prism is opt-in for decks (vs always-on for md-document):
- Most decks have few/short code blocks; the Prism payload is overhead for a typography-heavy artifact
- A speaker projecting a deck doesn't need pixel-perfect syntax fidelity; the audience reads from a distance
- Print-to-PDF benefits from absence of CDN dependencies (the print artifact is the artifact, not a server)
## Why single-file specifically matters for decks
1. **Projector-friendly**: the speaker opens one file from their laptop. No server, no `npm start`, no build artifact directory.
2. **Conference WiFi-tolerant**: even if WiFi fails mid-talk, the deck keeps working. Google Fonts is cached after first paint; Prism (when enabled) gracefully degrades to plain `<pre>`.
3. **PDF-exportable**: the same `.html` file becomes a PDF via the browser's print dialog (`Cmd+P` / `Ctrl+P`). The `@media print` stylesheet makes each slide one page.
4. **Email-attachable**: a 15-30 KB single-file deck attaches to email and renders inline in modern clients.
5. **Reproducible**: the deck-as-artifact is deterministic. Re-running the renderer on the same markdown produces the same HTML (no timestamps, no random IDs).
## What we deliberately don't externalize
- **CSS** — fully inline. Base CSS + design-system tokens + style overrides = ~6-7 KB.
- **JavaScript** — fully inline. Vanilla JS + IntersectionObserver-free navigation = ~3 KB.
- **Slide content** — every `<section>` is in the HTML. No lazy-loading, no fetch.
- **Speaker notes** — stored in the `data-notes` attribute of each `<section>`. Travels with the slide.
- **Images** — currently passed through as URLs; users wanting full single-file portability for image-heavy decks can pre-process to base64 (out of scope for v2.10.3).
## Why no transitions / animations
Three reasons:
1. **Speaker pacing** — transitions add latency between key press and visible response. For a brisk talk, that latency adds up.
2. **Recording-friendly** — many talks get screen-recorded. A snap transition is easier to edit than a fade.
3. **`prefers-reduced-motion`** — about 35% of users have motion-sensitivity preferences enabled (per WCAG WG estimates). A motion-free deck works for everyone by default.
If a deck genuinely needs motion, the user can add a `<style>` override block to the rendered HTML (it's a `.html` file — fully editable).
## Print-to-PDF discipline
```css
@media print {
html, body { height: auto; overflow: visible; }
.chrome, .progress, .presenter-panel { display: none !important; }
.slide {
display: flex !important; /* override "only active is visible" */
position: relative; /* break out of absolute positioning */
height: 100vh; /* one slide = one page */
page-break-after: always;
break-after: page;
}
.slide:last-child { page-break-after: auto; }
}
```
The result: `Cmd+P` from the deck produces a PDF where each slide is one A4/Letter page. No PDF generation pipeline needed; the browser does it.
## Sources
### 1. Tom MacWright — *Big* (github.com/tmcw/big, MIT)
The single-file deck tool that proved the model works. Used at speaking engagements by the JavaScript community for over a decade.
### 2. reveal.js — Single-file export (revealjs.com/installation/#full-setup)
reveal.js can produce a single-file output but typically ships multi-file. We adopt the single-file end of the spectrum exclusively.
### 3. Slides.com — HTML export
Slides.com's "export to HTML" produces a single-file artifact. Validates the user demand for the format independent of any one tool.
### 4. Marp — *Marp CLI HTML output* (marp.app)
Markdown-to-deck tool. Outputs single-file HTML by default. Our `---` boundary + `<!-- notes: -->` syntax is read-compatible with Marp's input.
### 5. Pandoc — `--standalone` flag (pandoc.org)
The pattern of "compile to single self-contained HTML." We honor the same constraint.
### 6. MDN — `@media print` (developer.mozilla.org/en-US/docs/Web/CSS/@media/print)
The browser-native primitive that makes "deck as printable PDF" work without any PDF library.
### 7. WCAG 2.2 §2.3.3 *Animation from Interactions* and `prefers-reduced-motion`
The accessibility-driven case for motion-free defaults. Our deck respects this preference automatically.
## Applied to `md-slides`
`deck_html_renderer.py` emits the single-file shape: inline CSS + inline JS + all `<section>` elements with `data-notes` attributes. The only external is Google Fonts CSS; Prism is opt-in via `--syntax`. Print-to-PDF works out of the box.
FILE:scripts/deck_html_renderer.py
#!/usr/bin/env python3
"""deck_html_renderer.py - Render parsed slides into a single-file HTML deck.
Stdlib-only. Reads slide JSON (from slide_splitter + presenter_notes_parser)
plus the design-system config, emits one .html file with:
- All slides as <section class="slide" id="slide-N"> elements
- One slide visible at a time (driven by URL hash + JS)
- Keyboard nav: ← / → / Space / PgDn / PgUp / Home / End
- Presenter mode toggle (P key): split view with current + notes + clock + next-slide preview
- @media print { section { display: block; page-break-after: always; } } → PDF export
- 12 design-system tokens applied; design_style affects layout density
- Reuses md-document's markdown_parser to render slide-body content (paragraphs,
lists, code, tables, callouts) consistently with md-document
Vanilla JS only — no frameworks. Total payload ~3-4 KB inline.
Single-file output: all CSS + JS inline. Only external is Google Fonts CSS.
No Prism (slide code blocks are short; we color them with the design-system
code background but don't fetch the Prism CDN by default — pass --syntax to
enable Prism for code-heavy decks).
NO LLM CALLS. Pure templating.
Usage:
python deck_html_renderer.py --slides slides.json --output deck.html
python deck_html_renderer.py --sample --output /tmp/sample.html
python deck_html_renderer.py --slides slides.json --syntax --output deck.html
"""
from __future__ import annotations
import argparse
import base64
import html
import json
import os
import sys
from pathlib import Path
from typing import Any
# Bridge to design-system config
_DESIGN_SYSTEM_SCRIPTS = (
Path(__file__).resolve().parent.parent.parent / "design-system" / "scripts"
)
sys.path.insert(0, str(_DESIGN_SYSTEM_SCRIPTS))
try:
import config_loader as _cfg
except ImportError:
_cfg = None
# Reuse md-document's markdown parser for slide-body rendering
_MD_DOCUMENT_SCRIPTS = (
Path(__file__).resolve().parent.parent.parent / "md-document" / "scripts"
)
sys.path.insert(0, str(_MD_DOCUMENT_SCRIPTS))
try:
import markdown_parser as _mp
except ImportError:
_mp = None
CALLOUT_ICONS: dict[str, str] = {
"NOTE": "i", "TIP": "*", "IMPORTANT": "!", "WARNING": "!", "CAUTION": "!",
}
def _palette_to_css(palette: dict[str, str]) -> str:
if not palette:
palette = {
"--md-bg": "#0E1E38", "--md-surface": "#142B50", "--md-border": "#1A3868",
"--md-text": "#F7F7F2", "--md-text-muted": "rgba(247, 247, 242, 0.68)",
"--md-accent": "#00D4AA", "--md-accent-soft": "rgba(0, 212, 170, 0.14)",
"--md-code-bg": "#122648",
"--md-link": "#00D4AA", "--md-link-hover": "#08FECE",
"--md-success": "#10A85C", "--md-warn": "#C87C10",
}
return "\n".join(f" {k}: {v};" for k, v in palette.items())
def _font_url(heading: str, body: str) -> str:
families = sorted({heading, body})
parts = "&".join(f"family={f.replace(' ', '+')}:wght@400;600;700" for f in families)
return f"https://fonts.googleapis.com/css2?{parts}&display=swap"
def _font_stack(name: str) -> str:
fallback = ("Georgia, serif" if name in
("Playfair Display", "Merriweather", "Lora", "Source Serif 4")
else "system-ui, -apple-system, sans-serif")
return f"'{name}', {fallback}"
def _render_block(block: dict[str, Any]) -> str:
"""Render one parsed-markdown block to slide HTML."""
t = block["type"]
if t == "heading":
# Inside a slide, demote: H1 already handled as slide title; H2 stays H2; etc.
level = max(2, block["level"])
text = _mp.render_inline_html(block["text"]) if _mp else html.escape(block["text"])
return f'<h{level}>{text}</h{level}>'
if t == "paragraph":
text = _mp.render_inline_html(block["text"]) if _mp else html.escape(block["text"])
return f'<p>{text}</p>'
if t == "hr":
return '<hr>'
if t == "code":
lang = block.get("language") or "text"
body = html.escape(block["body"])
return f'<pre><code class="language-{html.escape(lang)}">{body}</code></pre>'
if t == "list":
tag = "ol" if block.get("ordered") else "ul"
items = "".join(
f"<li>{_mp.render_inline_html(item) if _mp else html.escape(item)}</li>"
for item in block["items"]
)
return f"<{tag}>{items}</{tag}>"
if t == "table":
headers = block["headers"]
aligns = block.get("aligns") or ["left"] * len(headers)
rows = block["rows"]
thead = "<thead><tr>" + "".join(
f'<th class="align-{a}">{_mp.render_inline_html(h) if _mp else html.escape(h)}</th>'
for h, a in zip(headers, aligns)
) + "</tr></thead>"
tbody = "<tbody>" + "".join(
"<tr>" + "".join(
f'<td class="align-{aligns[i] if i < len(aligns) else "left"}">'
f'{_mp.render_inline_html(cell) if _mp else html.escape(cell)}</td>'
for i, cell in enumerate(row)
) + "</tr>"
for row in rows
) + "</tbody>"
return f"<table>{thead}{tbody}</table>"
if t == "callout":
kind = (block.get("kind") or "NOTE").upper()
icon = CALLOUT_ICONS.get(kind, "i")
body = "<br>".join(
_mp.render_inline_html(ln) if _mp else html.escape(ln)
for ln in (block.get("body_lines") or []) if ln
)
klass = kind.lower()
return (
f'<aside class="callout callout-{klass}" role="note">'
f'<div class="callout-label"><span class="callout-icon" aria-hidden="true">{icon}</span>'
f'{html.escape(kind)}</div>'
f'<div class="callout-body">{body}</div>'
f'</aside>'
)
if t == "blockquote":
body = "<br>".join(
_mp.render_inline_html(ln) if _mp else html.escape(ln)
for ln in block.get("body_lines") or [] if ln
)
return f"<blockquote>{body}</blockquote>"
return ""
def _render_slide_body(body_markdown: str) -> str:
"""Parse a slide's markdown body and render it as HTML blocks."""
if not body_markdown.strip():
return ""
if _mp is None:
return f'<pre>{html.escape(body_markdown)}</pre>'
parsed = _mp.parse_markdown(body_markdown)
return "\n".join(_render_block(b) for b in parsed["blocks"])
# ----- CSS template -----------------------------------------------------------
BASE_CSS = """
:root {
__PALETTE__
--md-font-heading: __HEADING_FONT__;
--md-font-body: __BODY_FONT__;
--md-font-mono: 'JetBrains Mono', ui-monospace, SFMono-Regular, Menlo, monospace;
}
* { box-sizing: border-box; }
html, body { margin: 0; padding: 0; height: 100%; overflow: hidden; }
body {
background: var(--md-bg);
color: var(--md-text);
font-family: var(--md-font-body);
font-size: 22px;
line-height: 1.45;
}
/* Slide layout — each slide is a viewport-sized panel */
.deck { width: 100vw; height: 100vh; position: relative; }
.slide {
position: absolute;
inset: 0;
display: none;
flex-direction: column;
justify-content: center;
padding: 4vh 8vw;
overflow: auto;
}
.slide.active { display: flex; }
.slide h1, .slide h2, .slide h3 {
font-family: var(--md-font-heading);
margin: 0 0 0.6em;
line-height: 1.15;
font-weight: 700;
}
.slide h1 { font-size: 3rem; }
.slide h2 { font-size: 2.5rem; }
.slide h3 { font-size: 1.875rem; }
.slide p { margin: 0.5em 0 1em; font-size: 1.25rem; }
.slide ul, .slide ol { padding-left: 1.5em; font-size: 1.25rem; }
.slide li { margin: 0.4em 0; }
.slide a { color: var(--md-link); }
.slide code {
background: var(--md-code-bg);
padding: 0.1em 0.3em;
border-radius: 4px;
font-family: var(--md-font-mono);
font-size: 0.9em;
}
.slide pre {
background: var(--md-code-bg);
border: 1px solid var(--md-border);
border-radius: 8px;
padding: 1rem 1.25rem;
font-family: var(--md-font-mono);
font-size: 1rem;
overflow: auto;
margin: 0.75em 0;
}
.slide pre code { background: transparent; padding: 0; }
.slide table {
width: 100%;
border-collapse: collapse;
margin: 1em 0;
font-size: 1.125rem;
}
.slide th, .slide td { border: 1px solid var(--md-border); padding: 0.5em 0.75em; text-align: left; }
.slide th { background: var(--md-surface); font-weight: 600; }
.slide td.align-center, .slide th.align-center { text-align: center; }
.slide td.align-right, .slide th.align-right { text-align: right; }
.slide blockquote {
border-left: 4px solid var(--md-accent);
margin: 1em 0;
padding: 0.4em 1em;
color: var(--md-text-muted);
font-style: italic;
font-size: 1.25rem;
}
.slide hr { border: 0; border-top: 1px solid var(--md-border); margin: 1.5em 0; }
.slide .callout {
border-left: 4px solid var(--md-accent);
background: var(--md-accent-soft);
padding: 0.75em 1em;
margin: 1em 0;
border-radius: 0 8px 8px 0;
}
.slide .callout-label {
font-family: var(--md-font-heading);
font-weight: 700;
font-size: 0.875rem;
letter-spacing: 0.05em;
text-transform: uppercase;
color: var(--md-accent);
margin-bottom: 0.25em;
display: flex;
gap: 0.5em;
align-items: center;
}
.slide .callout-note { border-left-color: var(--md-link); }
.slide .callout-tip { border-left-color: var(--md-success); }
.slide .callout-warning { border-left-color: var(--md-warn); }
.slide .callout-caution { border-left-color: var(--md-warn); }
/* Chrome (slide counter + hint) */
.chrome {
position: fixed;
bottom: 1rem;
right: 1.25rem;
font-family: var(--md-font-mono);
font-size: 0.875rem;
color: var(--md-text-muted);
user-select: none;
}
.chrome a { color: var(--md-text-muted); margin-right: 0.5rem; }
.chrome a:hover { color: var(--md-accent); }
/* Progress bar */
.progress {
position: fixed;
top: 0; left: 0;
height: 3px;
background: var(--md-accent);
transition: width 0.2s ease;
z-index: 100;
}
/* Presenter mode (P key toggles body class) */
body.presenter .deck { width: 60vw; }
body.presenter .presenter-panel { display: flex; }
.presenter-panel {
display: none;
position: fixed;
top: 0; right: 0;
width: 40vw;
height: 100vh;
background: var(--md-surface);
border-left: 1px solid var(--md-border);
flex-direction: column;
padding: 1.5rem 1.5rem;
font-size: 1rem;
overflow: auto;
}
.presenter-panel h3 {
font-family: var(--md-font-heading);
font-size: 0.875rem;
text-transform: uppercase;
letter-spacing: 0.05em;
color: var(--md-text-muted);
margin: 0 0 0.5rem;
border-bottom: 1px solid var(--md-border);
padding-bottom: 0.4rem;
}
.presenter-panel .clock { font-family: var(--md-font-mono); font-size: 1.5rem; margin-bottom: 1rem; }
.presenter-panel .notes { flex: 1; line-height: 1.5; color: var(--md-text); margin-bottom: 1rem; white-space: pre-wrap; }
.presenter-panel .next-preview {
background: var(--md-bg);
border: 1px solid var(--md-border);
border-radius: 8px;
padding: 0.75rem 1rem;
font-size: 0.875rem;
max-height: 25vh;
overflow: hidden;
opacity: 0.7;
}
.presenter-panel .next-preview h4 {
font-family: var(--md-font-heading);
margin: 0 0 0.3em;
font-size: 1rem;
color: var(--md-accent);
}
/* Print: all slides visible, one per page */
@media print {
html, body { height: auto; overflow: visible; }
.chrome, .progress, .presenter-panel { display: none !important; }
.deck, body.presenter .deck { width: 100%; height: auto; }
.slide {
display: flex !important;
position: relative;
height: 100vh;
page-break-after: always;
break-after: page;
}
.slide:last-child { page-break-after: auto; }
}
@media (prefers-reduced-motion: reduce) {
* { animation: none !important; transition: none !important; }
}
"""
# Inline vanilla-JS payload
JS_PAYLOAD = r"""
(function () {
"use strict";
var slides = document.querySelectorAll(".slide");
var total = slides.length;
var current = 0;
var presenter = false;
function show(idx) {
idx = Math.max(0, Math.min(total - 1, idx));
slides.forEach(function (s, i) {
s.classList.toggle("active", i === idx);
});
current = idx;
document.getElementById("counter").textContent = (current + 1) + " / " + total;
var prog = document.getElementById("progress");
if (prog) prog.style.width = (100 * (current + 1) / total) + "%";
if (location.hash !== "#" + (current + 1)) {
history.replaceState(null, "", "#" + (current + 1));
}
updatePresenterPanel();
}
function next() { show(current + 1); }
function prev() { show(current - 1); }
function first() { show(0); }
function last() { show(total - 1); }
function togglePresenter() {
presenter = !presenter;
document.body.classList.toggle("presenter", presenter);
updatePresenterPanel();
}
function updatePresenterPanel() {
var panel = document.querySelector(".presenter-panel");
if (!panel) return;
var notes = slides[current].getAttribute("data-notes") || "(no notes for this slide)";
panel.querySelector(".notes").textContent = notes;
var preview = panel.querySelector(".next-preview");
if (current + 1 < total) {
var nextSlide = slides[current + 1];
var title = nextSlide.querySelector("h1, h2") ;
preview.innerHTML = "<h4>Up next (#" + (current + 2) + ")</h4>" +
(title ? title.outerHTML : "<em>(no title)</em>");
preview.style.display = "block";
} else {
preview.innerHTML = "<h4>End of deck</h4>";
}
}
function tickClock() {
var now = new Date();
var hh = String(now.getHours()).padStart(2, "0");
var mm = String(now.getMinutes()).padStart(2, "0");
var ss = String(now.getSeconds()).padStart(2, "0");
var el = document.querySelector(".presenter-panel .clock");
if (el) el.textContent = hh + ":" + mm + ":" + ss;
}
document.addEventListener("keydown", function (e) {
if (e.metaKey || e.ctrlKey || e.altKey) return;
switch (e.key) {
case "ArrowRight":
case "PageDown":
case " ":
e.preventDefault(); next(); break;
case "ArrowLeft":
case "PageUp":
e.preventDefault(); prev(); break;
case "Home": e.preventDefault(); first(); break;
case "End": e.preventDefault(); last(); break;
case "p":
case "P":
e.preventDefault(); togglePresenter(); break;
case "Escape":
if (presenter) { e.preventDefault(); togglePresenter(); }
break;
}
});
// Mount the initial slide (respect URL hash, otherwise slide 1)
var initial = 0;
if (location.hash) {
var n = parseInt(location.hash.slice(1), 10);
if (!isNaN(n) && n >= 1 && n <= total) initial = n - 1;
}
show(initial);
setInterval(tickClock, 1000);
tickClock();
})();
"""
def render(deck_payload: dict[str, Any], config: dict[str, Any],
title: str = "Deck", enable_syntax: bool = False) -> str:
palette = config.get("derived_palette") or {}
typo = config.get("typography") or {}
heading_font = typo.get("heading_font", "Inter")
body_font = typo.get("body_font", "Inter")
style = config.get("design_style", "technical")
company_name = config.get("company_name", "")
css = (BASE_CSS
.replace("__PALETTE__", _palette_to_css(palette))
.replace("__HEADING_FONT__", _font_stack(heading_font))
.replace("__BODY_FONT__", _font_stack(body_font)))
slides = deck_payload.get("slides", [])
slide_html_parts: list[str] = []
for s in slides:
title_html = (f'<h1>{html.escape(s["title"])}</h1>' if s.get("title") else "")
body_html = _render_slide_body(s.get("body_markdown", ""))
notes_attr = html.escape(s.get("notes") or "", quote=True)
slide_html_parts.append(
f'<section class="slide" id="slide-{s["slide_number"]}" '
f'data-notes="{notes_attr}">'
f'{title_html}{body_html}</section>'
)
slides_html = "\n".join(slide_html_parts)
prism_links = ""
if enable_syntax:
prism_links = (
'<link rel="stylesheet" '
'href="https://cdn.jsdelivr.net/npm/prismjs@1.29.0/themes/prism-tomorrow.min.css">\n'
'<script defer src="https://cdn.jsdelivr.net/npm/prismjs@1.29.0/components/prism-core.min.js"></script>\n'
'<script defer src="https://cdn.jsdelivr.net/npm/prismjs@1.29.0/plugins/autoloader/prism-autoloader.min.js"></script>'
)
return f"""<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>{html.escape(title)}</title>
<link rel="preconnect" href="https://fonts.googleapis.com" crossorigin>
<link rel="stylesheet" href="{_font_url(heading_font, body_font)}">
{prism_links}
<style>{css}</style>
</head>
<body class="style-{style}">
<div id="progress" class="progress"></div>
<div class="deck">
{slides_html}
</div>
<aside class="presenter-panel" aria-label="Presenter view (toggle with P)">
<h3>Clock</h3>
<div class="clock">--:--:--</div>
<h3>Speaker notes</h3>
<div class="notes"></div>
<div class="next-preview"></div>
</aside>
<div class="chrome">
<a href="#" onclick="event.preventDefault(); document.dispatchEvent(new KeyboardEvent('keydown', {{key:'P'}}))">P · presenter</a>
<span id="counter">1 / {len(slides)}</span>
</div>
<script>{JS_PAYLOAD}</script>
</body>
</html>"""
def main(argv: list[str]) -> int:
p = argparse.ArgumentParser(description=__doc__.split("\n")[0])
p.add_argument("--slides", help="Path to presenter_notes_parser JSON, or '-' for stdin")
p.add_argument("--output", help="Path to write HTML (else stdout)")
p.add_argument("--title", default="Deck", help="Browser tab title")
p.add_argument("--syntax", action="store_true",
help="Enable Prism.js CDN for code syntax highlighting (off by default)")
p.add_argument("--no-config", action="store_true",
help="Bypass design-system config (use DEFAULTS)")
p.add_argument("--sample", action="store_true",
help="Render the built-in 5-slide sample deck")
p.add_argument("--strict-notes", action="store_true",
help="Refuse to render if < 50%% of slides have presenter notes "
"AND user wants presenter mode")
args = p.parse_args(argv)
if args.sample:
sys.path.insert(0, str(Path(__file__).resolve().parent))
import presenter_notes_parser
import slide_splitter
slides_payload = slide_splitter.split_slides(slide_splitter.SAMPLE_MARKDOWN)
deck_payload = presenter_notes_parser.attach_notes(slides_payload)
title = "Sample Deck — The Case for Single-File HTML"
else:
if not args.slides:
p.print_help()
return 0
raw = sys.stdin.read() if args.slides == "-" else Path(args.slides).read_text(encoding="utf-8")
deck_payload = json.loads(raw)
title = args.title
# Hard rule: if --strict-notes, refuse < 50% coverage
if args.strict_notes:
coverage = deck_payload.get("summary", {}).get("notes_coverage_pct", 0)
if coverage < 50:
print(f"refusing (--strict-notes): only {coverage}% of slides have presenter "
f"notes (need ≥ 50% for a presenter deck). Add more "
f"<!-- notes: ... --> blocks or drop --strict-notes.",
file=sys.stderr)
return 7
if args.no_config or os.environ.get("MARKDOWN_HTML_NO_CONFIG") == "1":
config = _cfg.DEFAULTS if _cfg else {}
else:
config = _cfg.load_config() if _cfg else {}
output = render(deck_payload, config, title=title, enable_syntax=args.syntax)
if args.output and args.output != "-":
Path(args.output).write_text(output, encoding="utf-8")
notes_count = sum(1 for s in deck_payload.get("slides", []) if s.get("has_notes"))
print(f"wrote {args.output}: {len(output):,} bytes, "
f"{deck_payload['summary']['total_slides']} slides "
f"({notes_count} with notes, syntax={args.syntax})")
else:
print(output)
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/presenter_notes_parser.py
#!/usr/bin/env python3
"""presenter_notes_parser.py - Extract <!-- notes: ... --> blocks from slide bodies.
Stdlib-only. Operates on the slide JSON produced by slide_splitter.py. For each
slide, finds any HTML-comment notes block (the convention used by reveal.js,
Marp, Big, Pandoc-Beamer) and:
- Records the notes text separately under `notes`
- Removes the notes block from `body_markdown` so the slide renders cleanly
Accepted notes syntax (case-insensitive on the keyword):
<!-- notes: This is a single-line presenter note. -->
<!-- notes:
Multi-line notes block. Can contain markdown.
- Bullet points
- Multiple paragraphs
-->
<!-- speaker-notes: alias -->
<!-- presenter: alias -->
If a slide has multiple notes blocks, they're concatenated with blank lines
between them.
NO LLM CALLS. Pure regex + slide-by-slide transformation.
Usage:
python presenter_notes_parser.py --slides slides.json --output slides-notes.json
python presenter_notes_parser.py --sample
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Any
# Matches the entire <!-- notes: ... --> block, including multi-line
NOTES_BLOCK_RE = re.compile(
r"<!--\s*(?:notes|speaker-notes|presenter)\s*:\s*(.*?)\s*-->",
re.IGNORECASE | re.DOTALL,
)
def extract_notes_from_slide(slide: dict[str, Any]) -> dict[str, Any]:
"""Return a new slide dict with `notes` populated and `body_markdown`
stripped of any notes blocks."""
body = slide.get("body_markdown", "")
found: list[str] = []
def _capture(m: re.Match) -> str:
found.append(m.group(1).strip())
return "" # remove the block from body
cleaned = NOTES_BLOCK_RE.sub(_capture, body)
# Tidy up: collapse runs of >2 blank lines, trim trailing whitespace
cleaned = re.sub(r"\n{3,}", "\n\n", cleaned).strip("\n")
out = dict(slide)
out["body_markdown"] = cleaned
out["notes"] = "\n\n".join(found) if found else ""
out["has_notes"] = bool(found)
return out
def attach_notes(slides_payload: dict[str, Any]) -> dict[str, Any]:
slides = slides_payload.get("slides", [])
new_slides = [extract_notes_from_slide(s) for s in slides]
notes_count = sum(1 for s in new_slides if s["has_notes"])
out = dict(slides_payload)
out["slides"] = new_slides
out["summary"] = dict(out.get("summary", {}))
out["summary"]["slides_with_notes"] = notes_count
out["summary"]["notes_coverage_pct"] = (
round(100 * notes_count / len(new_slides), 1) if new_slides else 0.0
)
return out
def main(argv: list[str]) -> int:
p = argparse.ArgumentParser(description=__doc__.split("\n")[0])
p.add_argument("--slides", help="Path to slide_splitter JSON output, or '-' for stdin")
p.add_argument("--output", help="Path to write JSON output (else stdout)")
p.add_argument("--sample", action="store_true",
help="Run on the slide_splitter built-in sample")
args = p.parse_args(argv)
if args.sample:
sys.path.insert(0, str(Path(__file__).resolve().parent))
import slide_splitter
slides_payload = slide_splitter.split_slides(slide_splitter.SAMPLE_MARKDOWN)
elif args.slides:
raw = sys.stdin.read() if args.slides == "-" else Path(args.slides).read_text(encoding="utf-8")
slides_payload = json.loads(raw)
else:
p.print_help()
return 0
result = attach_notes(slides_payload)
payload = json.dumps(result, indent=2)
if args.output:
Path(args.output).write_text(payload, encoding="utf-8")
print(f"wrote {args.output}: {result['summary']['slides_with_notes']}/"
f"{result['summary']['total_slides']} slides have presenter notes "
f"({result['summary']['notes_coverage_pct']}% coverage)")
else:
print(payload)
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/slide_splitter.py
#!/usr/bin/env python3
"""slide_splitter.py - Split a markdown deck into ordered slides.
Stdlib-only. Accepts three boundary modes:
--boundary hr Split on `---` HR lines (the most common convention; what
reveal.js / pandoc / Marp / Big all read by default).
--boundary h1 Split on top-level `# ` headings; each H1 starts a new slide.
--boundary auto (default) Pick the better signal automatically:
- HR count ≥ 3 → use HR
- else H1 count ≥ 5 → use H1
- else FAIL — input has no clear slide boundaries
The first slide gets everything from the start of file up to the first boundary.
H1-mode treats the H1 line as part of the slide (it becomes the slide title);
HR-mode does NOT include the `---` line in either slide.
NO LLM CALLS. Pure regex + state machine.
Hard rules (refusals):
1. 1-slide deck → exit 5 (it's a poster — route to md-document)
2. Any slide body > 40 source lines → warning printed to stderr (signal-to-
noise; presenters fail with too much per slide). Soft-fail; renders anyway.
3. No boundaries detectable in auto mode → exit 6 (route to md-document)
Usage:
python slide_splitter.py --input deck.md --output slides.json
python slide_splitter.py --input - --boundary h1
python slide_splitter.py --sample
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Any
HR_RE = re.compile(r"^---\s*$")
H1_RE = re.compile(r"^#\s+(.+?)\s*$")
MAX_SLIDE_LINES = 40 # soft warn above this (signal-to-noise)
def _extract_title(slide_body: list[str]) -> tuple[str, list[str]]:
"""If the first non-blank line is a heading, treat it as the title and
return (title, body_without_title_line). Otherwise return ('', body)."""
for i, line in enumerate(slide_body):
if not line.strip():
continue
m = H1_RE.match(line)
if m:
return (m.group(1).strip(), slide_body[:i] + slide_body[i + 1:])
# Also accept H2 as a slide title if the slide has no H1 (common in
# decks where H1 = whole-deck title and each slide leads with H2)
m2 = re.match(r"^##\s+(.+?)\s*$", line)
if m2:
return (m2.group(1).strip(), slide_body[:i] + slide_body[i + 1:])
break
return ("", slide_body)
def _pick_boundary_auto(lines: list[str]) -> str:
"""Pick HR vs H1 based on signal counts. HR wins if ≥3; else H1 if ≥5."""
hr_count = sum(1 for ln in lines if HR_RE.match(ln))
h1_count = sum(1 for ln in lines if H1_RE.match(ln))
if hr_count >= 3:
return "hr"
if h1_count >= 5:
return "h1"
return "" # caller treats empty as "no boundary detected"
def split_slides(text: str, boundary: str = "auto") -> dict[str, Any]:
lines = text.splitlines()
if boundary == "auto":
chosen = _pick_boundary_auto(lines)
if not chosen:
return {
"slides": [],
"boundary": "auto",
"boundary_used": None,
"summary": {
"total_slides": 0,
"max_slide_lines": 0,
"over_threshold": [],
"error": "no clear slide boundaries (need ≥3 HR or ≥5 H1)",
},
}
boundary = chosen
raw_groups: list[list[str]] = []
current: list[str] = []
boundary_lines: list[int] = [] # source line indices where each slide starts
if boundary == "hr":
current_start = 0
for i, ln in enumerate(lines):
if HR_RE.match(ln):
raw_groups.append(current)
boundary_lines.append(current_start)
current = []
current_start = i + 1
continue
current.append(ln)
# Last slide
raw_groups.append(current)
boundary_lines.append(current_start)
elif boundary == "h1":
current_start = 0
first_h1_seen = False
for i, ln in enumerate(lines):
if H1_RE.match(ln):
if first_h1_seen:
raw_groups.append(current)
boundary_lines.append(current_start)
current = []
current_start = i
else:
# Everything before the first H1 (if non-empty) is the
# opening slide; the first H1 starts slide 2.
if any(s.strip() for s in current):
raw_groups.append(current)
boundary_lines.append(0)
current = []
current_start = i
first_h1_seen = True
current.append(ln)
if current:
raw_groups.append(current)
boundary_lines.append(current_start)
else:
return {
"slides": [],
"boundary": boundary,
"boundary_used": boundary,
"summary": {
"total_slides": 0,
"max_slide_lines": 0,
"over_threshold": [],
"error": f"unknown --boundary mode: {boundary}",
},
}
# Trim leading/trailing blank lines per slide; drop empty slides
cleaned: list[dict[str, Any]] = []
over_threshold: list[int] = []
max_lines = 0
for idx, body_lines in enumerate(raw_groups):
while body_lines and not body_lines[0].strip():
body_lines = body_lines[1:]
while body_lines and not body_lines[-1].strip():
body_lines = body_lines[:-1]
if not body_lines:
continue
title, body_no_title = _extract_title(body_lines)
line_count = len(body_lines)
max_lines = max(max_lines, line_count)
if line_count > MAX_SLIDE_LINES:
over_threshold.append(idx + 1)
cleaned.append({
"slide_number": len(cleaned) + 1,
"title": title,
"body_markdown": "\n".join(body_no_title),
"raw_body_markdown": "\n".join(body_lines),
"source_line": boundary_lines[idx] if idx < len(boundary_lines) else 0,
"line_count": line_count,
})
return {
"slides": cleaned,
"boundary": boundary,
"boundary_used": boundary,
"summary": {
"total_slides": len(cleaned),
"max_slide_lines": max_lines,
"over_threshold": over_threshold,
"max_threshold": MAX_SLIDE_LINES,
},
}
SAMPLE_MARKDOWN = """# The Case for Single-File HTML
Why agent-generated artifacts should ship as one .html file.
---
# Three forces converged
- Outputs got longer
- The editing relationship changed (LLM edits, not human)
- The information became spatial
Markdown can't carry any of those three at length.
<!-- notes: This is the framing slide. Start by asking the audience how
many of them have stopped reading a markdown spec past line 100. -->
---
# What HTML restores
| Dimension | Markdown | HTML |
|-----------|----------|------|
| Hierarchy | Indented `#` | Typography scale |
| Navigation | Linear | Sticky TOC + scrollspy |
| Comparison | Lists | Tables, grids |
| Interaction | None | Search, copy, hover |
<!-- notes: Spend 30 seconds on each row. The hierarchy row is the one
that lands hardest for engineers. -->
---
# Single-file discipline
- All CSS inline
- All JS inline
- Only externals: Google Fonts + Prism CDN
- Falls back gracefully
> One `.html` file uploads to S3, opens in any browser, attaches to email.
<!-- notes: The shareability point is the easiest sell. Skip the technical
details unless someone asks. -->
---
# Try it now
Append "as an HTML file" to your next Claude Code prompt.
That's the whole switch.
"""
def main(argv: list[str]) -> int:
p = argparse.ArgumentParser(description=__doc__.split("\n")[0])
p.add_argument("--input", help="Path to markdown file, or '-' for stdin")
p.add_argument("--output", help="Path to write JSON output (else stdout)")
p.add_argument("--boundary", choices=["auto", "hr", "h1"], default="auto",
help="Slide boundary mode (default: auto-detect)")
p.add_argument("--sample", action="store_true",
help="Run on a built-in 5-slide deck")
args = p.parse_args(argv)
if args.sample:
text = SAMPLE_MARKDOWN
elif args.input:
text = sys.stdin.read() if args.input == "-" else Path(args.input).read_text(encoding="utf-8")
else:
p.print_help()
return 0
result = split_slides(text, args.boundary)
# Hard rule: auto mode with no clear boundaries
if "error" in result["summary"]:
print(f"refusing: {result['summary']['error']}. "
f"Add `---` between slides, or use H1 boundaries, or route to md-document.",
file=sys.stderr)
return 6
# Hard rule: single-slide deck is a poster, not a deck
if result["summary"]["total_slides"] == 1:
print("refusing: input produces a 1-slide deck — that's a poster. "
"Route to md-document or add more --- boundaries.",
file=sys.stderr)
return 5
# Soft warn: slides over 40 source lines
if result["summary"]["over_threshold"]:
offenders = ", ".join(f"#{n}" for n in result["summary"]["over_threshold"])
print(f"warning: slides {offenders} exceed {MAX_SLIDE_LINES} source lines "
f"(signal-to-noise — consider splitting). Rendering anyway.",
file=sys.stderr)
payload = json.dumps(result, indent=2)
if args.output:
Path(args.output).write_text(payload, encoding="utf-8")
print(f"wrote {args.output}: {result['summary']['total_slides']} slides "
f"(boundary={result['boundary_used']}, max_lines={result['summary']['max_slide_lines']})")
else:
print(payload)
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
Tự động hóa trình duyệt để điều khiển Google NotebookLM: đọc, truy vấn notebook, thêm nguồn và tạo các đầu ra Studio như audio, infographic, slide, mind map.
---
name: notebooklm
description: "Browser automation skill for controlling Google's NotebookLM. Handles reading and querying notebooks, adding sources (URLs, text, files, YouTube links, synthesized content), generating Studio outputs (Audio Overview, infographics, slide decks, study guides, briefing docs, mind maps, timelines, FAQs), and creating new notebooks. Triggers on any phrase involving NotebookLM — 'open NotebookLM', 'check my [name] notebook', 'pull info from NotebookLM', 'ask my notebook about X', 'add [source] to NotebookLM', 'create an infographic in NotebookLM', 'use NotebookLM Studio', 'generate a slide deck from my notebook', or any variation where the goal involves NotebookLM. Requires browser automation environment — fails gracefully when unavailable."
license: MIT
metadata:
source_spec: "megaprompts/03-notebooklm-megaprompt.md"
build_pattern: "Path B (direct conversion)"
shape: "browser-automation (distinct from research-pack convention)"
version: 1.0.0
---
# NotebookLM — Browser Automation
> **Requires:** A browser automation environment (Claude Code CLI with computer-use, Claude Chrome Extension, or equivalent). **Skill will gracefully fail in non-automation contexts with a clear "not supported" message.**
> **Critical:** This skill is the only browser-automation skill in the v2 collection. It does NOT follow the research-pack Agent Integrity Rules convention. Different constraints apply (UI dynamics, async generation, login walls).
## Step 0: Browser Context Setup (Mandatory)
Before any other action, verify browser automation is available:
1. Check whether browser-control tools are loaded in the harness (screenshot, click, find-element, navigate)
2. If unavailable → **halt with clear message:** "This skill requires browser automation. Currently in {context}. Cannot proceed. Use Claude Code CLI with computer-use, Claude Chrome Extension, or equivalent."
3. If available → take initial screenshot, navigate to https://notebooklm.google.com
4. **Detect login wall via screenshot.** If login screen detected: halt with "Please log in to NotebookLM in the browser, then re-invoke this skill." **Never attempt to handle login automatically.**
## Phase 0: Grill-Me Intake (Action-Routing)
Up to 4 forcing questions, one at a time, dependency-ordered. Most invocations stop at Q3.
### Q1 (root) — Action
> **What do you want me to do? Pick one:**
>
> 1. **Read / extract** — ask a question of an existing notebook
> 2. **Add a source** — push content (URL, text, file, Google Doc, or synthesized content) into a notebook
> 3. **Generate a Studio output** — Audio Overview, Study Guide, Briefing Doc, Timeline, FAQ, Infographic, Slides, or Mind Map
> 4. **Create a new notebook** — initialize with title + initial sources
>
> *Why I'm asking:* Each action takes a different path through the UI and requires different parameters. Naming the action upfront prevents wasted screenshots and lets me ask only the follow-up questions that apply.
**Forcing choice.** If the user says "open NotebookLM" without specifying an action, **refuse to start** and re-ask Q1.
### Q2 (depends on Q1) — Notebook identity
> **Which notebook?** *(asked for actions 1, 2, 3 — not for "create new")*
>
> *Why I'm asking:* If you give me a name, I'll search the homepage; if you give me a URL, I'll navigate directly. Names that are ambiguous will get a disambiguation prompt with screenshots.
For action 4 (create new): replace with "What's the title for the new notebook?"
### Q3 (depends on Q1) — Action-specific parameter
**Action 1 (read/extract):**
> "What's the question to ask the notebook? Use natural phrasing — the notebook's chat handles it best."
**Action 2 (add source):**
> "What source type? Pick one:
> 1. URL / website / YouTube link
> 2. Copied text (paste here or point at content)
> 3. File upload (provide absolute path)
> 4. Google Doc (link)
> 5. Synthesized content (I'll pre-process and add as 'Copied text')
>
> *Why I'm asking:* Each source type goes through a different sub-flow in the Add Source dialog. Picking upfront saves a step."
**Action 3 (Studio output):**
> "Which Studio output? Audio Overview / Study Guide / Briefing Doc / Timeline / FAQ / Table of Contents / Infographic / Slides / Mind Map. And: any custom-prompt direction? **Default prompts produce mediocre output — I always open the customization menu and write a detailed prompt.** Tell me the angle or audience.
>
> *Why I'm asking:* The output type sets the UI button to find. The custom prompt is mandatory for quality."
**Action 4 (create new):**
> "Initial sources? Provide URLs, file paths, or 'I'll add later'."
### Q4 (depends on Q1 = action 3) — Studio custom prompt detail
> **Tell me the angle, audience, and length for the Studio output. Examples:**
>
> - **Audio Overview:** "Two-host conversation for a non-technical executive, 8–10 min, focus on business implications not technical depth"
> - **Infographic:** "Decision-tree style, action-oriented, 6 panels max, monochrome navy"
> - **Study Guide:** "Undergrad-level, definitions + 3 practice questions per concept"
>
> *Why I'm asking:* This becomes the custom prompt. **Default Studio prompts produce mediocre output — specific direction produces sharp output.**
**Asked only for Studio output generation (Q1=3). Skip otherwise.**
**Stop condition:** After Q4 (or earlier with dependency skips), commit and start the action sequence.
See [`references/studio_output_custom_prompts.md`](references/studio_output_custom_prompts.md) for the canon.
## Notebook Discovery
For actions 1-3 (require existing notebook):
1. Navigate to homepage → screenshot
2. If user provided **URL** → navigate directly
3. If user provided **name**:
- Use semantic find() to locate notebook card by visible title text
- If multiple matches → screenshot homepage, list options, ask user to specify
- If no match → ask user to provide URL or confirm spelling
For action 4 (create new):
1. Locate "New notebook" button on homepage
2. Click → set title from Q2
3. Add initial sources per Q3
## Action 1: Read / Extract
1. Open the notebook (notebook discovery above)
2. Locate chat input (semantic find or screenshot coordinates)
3. Type the question (use the user's natural phrasing from Q3)
4. Submit (Enter or send button)
5. **Wait 3–5 seconds**
6. Screenshot the response area
7. Extract and present in **clean format** (not raw chat dump)
## Action 2: Add Sources
Sub-flows per source type:
| Type | UI flow |
|---|---|
| URL / Website / YouTube | Add Source → Link → paste URL |
| Copied Text | Add Source → Copied text → paste content |
| File Upload | Use file-upload tool with absolute path + input ref (never click native file picker) |
| Google Doc | Add Source → Google Docs → Drive picker |
| Synthesized content | Pre-process content elsewhere, then add as Copied text |
**After every add:** wait for ingestion spinner, screenshot to confirm success.
**Synthesized content pattern (powerful):** instead of asking NotebookLM to ingest a raw URL with potentially noisy content, pre-process the content (extract main article, strip nav/ads/comments), then add as "Copied text". Produces dramatically better summarization.
## Action 3: Studio Outputs
**All 9 output types supported:** Audio Overview, Study Guide, Briefing Doc, Timeline, FAQ, Table of Contents, Infographic, Slides, Mind Map.
**Mandatory workflow:**
1. Locate Studio panel (right side; may need toggle)
2. Find the specific output button for the requested type
3. **Open customization menu** (chevron/arrow next to button) — **NOT the main button**
4. **Write detailed custom prompt** (from Q4)
5. Confirm and submit
6. **Do NOT wait for completion** — confirm generation started, notify user, return
### Custom prompt examples (4 output types)
**Audio Overview:**
> "Two-host conversation between a researcher and an experienced practitioner. Audience: non-technical executive making a budget decision. Length: 8-10 minutes. Focus on business implications, not technical depth. Include one concrete example per major point. Acknowledge counter-arguments briefly."
**Infographic:**
> "Decision-tree style. Action-oriented (each panel ends with a decision or action). 6 panels max. Monochrome navy + amber highlight. Each panel has: title (4-6 words), 1-2 sentence body, decision/action line. No filler panels."
**Study Guide:**
> "Undergraduate-level (define every technical term). Structure: 6 concepts × 4 elements each (definition / why it matters / one worked example / 3 practice questions). Practice questions Bloom-higher-order (apply/analyze), not recall."
**Slides (slide deck):**
> "12 slides max. 1-2 sentences per slide body. Presenter notes per slide with: one concrete example + one likely audience objection + how to address it. No bullet points in slide bodies — prose only. End with one-slide call-to-action."
See [`references/studio_output_custom_prompts.md`](references/studio_output_custom_prompts.md) for more.
## Action 4: Create New Notebook
1. Navigate to homepage
2. Click "New notebook"
3. Set title from Q2
4. Add initial sources from Q3 (use Action 2 sub-flows per source type)
5. **Wait for auto-summary generation** (this one IS synchronous — usually completes in <30 sec)
6. Screenshot final state
## Critical Async Behavior
> **Async output rule:** For Studio generations (especially **Audio Overview** — 5-10 min), DO NOT wait for completion. The user's session will time out.
>
> Workflow: Click Generate → confirm generation has started via screenshot → tell the user "Generation in progress — NotebookLM will notify you when ready" → **end the task.**
This is the **fire-and-notify** pattern. Different from add-source and auto-summary (which are fast enough to wait).
Use `scripts/async_action_classifier.py` to determine wait-or-notify per action:
| Action | Wait? |
|---|---|
| Add Source (URL/text/file) | Yes — wait for ingestion spinner (~5-30s) |
| Read/Extract (chat) | Yes — wait 3-5s for response |
| Studio: Audio Overview | **No** — fire and notify (5-10 min) |
| Studio: Infographic / Slides / Mind Map | **No** — fire and notify (2-5 min) |
| Studio: Study Guide / Briefing Doc / FAQ | Yes — wait ~30-60s |
| Create New Notebook | Yes — wait for auto-summary (<30s) |
See [`references/async_action_discipline.md`](references/async_action_discipline.md) for the canon.
## Screenshot-First Discipline
NotebookLM is a **dynamic SPA** where UI varies by:
- Account tier (free vs Plus vs Enterprise)
- Feature rollout (some Studio types not yet available to all users)
- Recent UI changes (Google iterates the product frequently)
**Every UI action must be preceded by a screenshot.** Reasons:
1. Verify the UI matches expectations before acting
2. Catch login walls early
3. Detect unexpected layout changes
4. Audit trail for debugging
Use `screenshot()` (or equivalent in your browser-automation tool) before every meaningful UI interaction.
See [`references/browser_automation_canon.md`](references/browser_automation_canon.md) for the discipline.
## find()-Before-Click
Use **semantic element finders** before pixel coordinates wherever possible:
- ✅ `find(text="Audio Overview")` → returns element regardless of position
- ❌ `click(x=420, y=380)` → breaks when UI rearranges
Semantic finders survive minor UI changes. Pixel coordinates do not.
Only fall back to coordinates when:
- Semantic find() returns nothing
- Element has no stable text/aria-label/data-attribute
- Visual position is the only reliable signal
## Saving Outputs to Workspace
For Read/Extract actions producing useful information:
1. Extract chat response cleanly (strip UI chrome)
2. Format readably (paragraphs, lists, code blocks as appropriate)
3. If user requested → save to file (`WORKSPACE/notebooklm/<notebook-slug>-<action>-<date>.md`)
4. Otherwise → return in chat as final summary
For Studio outputs:
1. NotebookLM hosts the output (Audio Overview is in-app, Infographic downloadable, etc.)
2. Report the location (URL or in-app navigation path) to user
3. Don't try to download/save Studio outputs to local workspace — that's NotebookLM's job
## Reporting Back Format
After completing any action:
1. Take final screenshot if visually relevant
2. Give **clean summary** (not raw chat dump):
- Notebook used (name)
- Action taken (specific)
- Result (1-2 sentences)
- For generated outputs: what was created + where it is + when ready
3. For fire-and-notify actions: explicit "NotebookLM will notify you when ready"
## Error Handling
| Failure | Behavior |
|---|---|
| Browser automation unavailable | Fail fast with "this skill requires browser automation" message (Step 0 halt) |
| Login wall detected | Stop. Tell user to log in. Don't attempt auto-login. |
| Multiple notebooks match name | Screenshot homepage, list options, ask user to specify |
| Source ingestion spinner stuck > 60s | Note timeout, ask user if they want to retry |
| Studio button not found in panel | Scroll down or look for "Discover more"; if still missing, note feature may not be enabled for this account |
| Chat response doesn't appear in 10s | Screenshot, check for error state, retry once |
| Page layout changed unexpectedly | Screenshot, describe what's visible, ask user for guidance |
## Tooling
| Script | Role |
|---|---|
| `scripts/action_router.py` | Q1-Q4 answers → action plan + UI flow + required parameters |
| `scripts/custom_prompt_template_generator.py` | Studio output type + audience + length → starter custom prompt |
| `scripts/async_action_classifier.py` | Action name → wait-or-notify pattern (fire-and-notify for slow generations) |
## References
- [`references/browser_automation_canon.md`](references/browser_automation_canon.md) — screenshot-first + find-before-click + tool-agnostic patterns (7+ sources)
- [`references/studio_output_custom_prompts.md`](references/studio_output_custom_prompts.md) — why defaults are mediocre + per-output-type templates (7+ sources)
- [`references/async_action_discipline.md`](references/async_action_discipline.md) — fire-and-notify pattern for slow UI ops (7+ sources)
## Anti-Patterns To Reject
- Tool-specific tool names without abstraction (e.g., hardcoding "Claude Chrome Extension")
- Synchronous waiting on Studio generations (especially Audio Overview)
- Skipping screenshots between actions
- Using pixel coordinates when semantic find() is available
- Attempting to handle login flows automatically
- Generating Studio outputs without opening customization menu
- Using default Studio prompts (always write custom)
---
**Version:** 1.0.0
**Source spec:** [`megaprompts/03-notebooklm-megaprompt.md`](../../../../megaprompts/03-notebooklm-megaprompt.md)
**Build pattern:** Path B (direct conversion). Browser-automation shape — distinct from research-pack convention.
FILE:references/async_action_discipline.md
# Async Action Discipline — Fire-and-Notify for Slow UI Operations
This reference answers exactly one decision: **for each NotebookLM action, does the skill wait synchronously OR fire and notify the user?**
## The Core Trade-Off
Browser-automation sessions have practical time limits:
- **Claude Code CLI sessions:** typically time out after ~10 minutes of inactivity
- **Chrome Extension sessions:** depend on the user keeping the tab open
- **API contexts:** strict timeout (60s-5min depending on tier)
NotebookLM's Studio operations vary in completion time:
- **Audio Overview:** 5-10 minutes (model generation + audio synthesis)
- **Infographic / Slides / Mind Map:** 2-5 minutes (complex visual generation)
- **Study Guide / Briefing Doc / FAQ:** 30-60 seconds (text generation)
- **Add Source ingestion:** 5-30 seconds (parsing + indexing)
- **Auto-summary on new notebook:** 10-30 seconds
The mis-match between session timeout and operation time means **synchronous waiting on slow ops is a failure mode**. The skill must fire-and-notify for slow operations instead.
## The Fire-and-Notify Pattern
When triggering a slow operation:
1. Locate the trigger button (via find())
2. Click it
3. **Verify generation started** via screenshot (look for spinner / "Generating" indicator)
4. **Tell the user**: "Generation in progress — NotebookLM will notify you when ready. NotebookLM sends in-app and email notifications when complete."
5. **End the task** — return control to the user
Key: **don't loop waiting for completion**. The browser-automation session will time out before NotebookLM finishes.
The user already knows how NotebookLM works — they'll see the notification when the Audio Overview is ready. The skill's job is to confirm the generation **started**, not to babysit it to completion.
## Per-Action Timing Catalog
| Action | Type | Timing | Wait or Notify? |
|---|---|---|---|
| Chat send (Read/Extract) | Q&A | 3-10s | **Wait** (short) |
| Add Source: URL | Ingestion | 5-15s | **Wait** |
| Add Source: Text | Ingestion | 5-15s | **Wait** |
| Add Source: File upload | Ingestion | 10-30s | **Wait** (with 60s timeout) |
| Add Source: Google Doc | Ingestion | 10-30s | **Wait** (with 60s timeout) |
| Create New Notebook | Init + summary | 15-30s | **Wait** |
| Studio: Study Guide | Generation | 30-60s | **Wait** (with 90s timeout) |
| Studio: Briefing Doc | Generation | 30-60s | **Wait** (with 90s timeout) |
| Studio: FAQ | Generation | 30-60s | **Wait** (with 90s timeout) |
| Studio: Table of Contents | Generation | 20-40s | **Wait** |
| Studio: Timeline | Generation | 30-60s | **Wait** (with 90s timeout) |
| **Studio: Audio Overview** | Audio gen | **5-10 min** | **NOTIFY** (fire-and-notify) |
| **Studio: Infographic** | Visual gen | **2-5 min** | **NOTIFY** |
| **Studio: Slides** | Visual gen | **2-5 min** | **NOTIFY** |
| **Studio: Mind Map** | Visual gen | **2-5 min** | **NOTIFY** |
The dividing line: **>2 minutes → fire-and-notify**. Anything shorter, wait synchronously with appropriate timeout.
## Wait-Discipline Details
For "wait" actions, the discipline:
1. After clicking the trigger, **start a wait loop** with timeout
2. Poll for completion signal:
- Spinner disappearing
- "Done" / "Ready" indicator
- Result content appearing
3. Take screenshot at intervals (every 10-15s) for audit
4. If timeout exceeded → screenshot, note timeout, ask user how to proceed
Example for chat (Read/Extract):
```
1. Submit question
2. Take screenshot at T+3s
3. If response present → extract, return
4. If still generating → take screenshot at T+8s
5. If response present → extract, return
6. If still not present → take screenshot at T+15s, note delay
7. If T+30s and still no response → note error, retry once
8. If second attempt fails → report failure with screenshot
```
## Notify-Discipline Details
For "notify" (fire-and-notify) actions:
1. Click Generate (in the customization menu, after custom prompt is set)
2. Take screenshot **within 5s** to verify generation started
3. Confirm via visible signal:
- Spinner appearing
- Status text changing to "Generating..."
- Generate button disabled or replaced with "Cancel"
4. **Tell the user**:
> "Generation triggered for {output_type}. NotebookLM takes ~{N} minutes for this. NOT waiting in this session — NotebookLM will notify you in-app and via email when ready. Returning control to you."
5. **End the task**. Don't loop.
## What Goes Wrong With Synchronous Waiting
### Failure 1: Session timeout
Skill clicks "Generate Audio Overview" at T+0. Browser-automation session times out at T+10min. Skill never confirms generation completed. User gets confused error message.
### Failure 2: Wasted compute
Skill loops `screenshot()` every 5s waiting for completion. 10 minutes × 12 screenshots/min = 120 wasted screenshots. Compute cost adds up.
### Failure 3: Hides errors
While waiting, if NotebookLM shows a transient error and recovers, the skill's wait loop might not catch it. Better: confirm started, hand off to NotebookLM's own error handling.
### Failure 4: Blocks user
User wanted to do something else while Audio Overview generated. Synchronous wait blocks the session.
## What Goes Wrong With Premature Notify
The opposite failure: applying fire-and-notify to a fast action.
### Anti-example: Read/Extract chat send
If the skill clicks send + immediately notifies "Response is being generated", the user has to come back later to retrieve it. But chat responses complete in 3-10s. Waiting is correct.
### Anti-example: Add Source URL
URL ingestion is 5-15s. The skill should wait for the ingestion spinner to clear, then confirm success. Premature notify means user doesn't know if the source was added successfully.
The 2-minute dividing line is empirically the right boundary. Below it, wait. Above it, notify.
## Tooling
`scripts/async_action_classifier.py` returns the wait-or-notify verdict per action:
```bash
python async_action_classifier.py --action audio_overview
# Returns: FIRE_AND_NOTIFY (estimated 5-10 min)
python async_action_classifier.py --action add_source_url
# Returns: WAIT (estimated 5-15s, timeout 60s)
```
Use this before triggering any Studio or Add-Source action so the skill applies the right pattern.
## Edge Cases
### Generation visibly fails immediately
If the Generate click results in an error toast within 5s → catch it, report, don't notify "in progress."
### Generation queued behind another generation
NotebookLM serializes Studio generations per notebook. If user requests Audio Overview while previous one is generating → either:
- Wait for previous to clear (only if previous is ~30s from completion based on visible progress)
- Notify: "Previous generation in progress; new one will queue"
### User cancels mid-generation
If user types "stop" while skill is in wait loop → stop polling, screenshot final state, report what was triggered.
## Anti-Patterns
### Synchronous wait on Audio Overview
The classic failure. 5-10 min wait exceeds session timeout.
### Fire-and-notify on chat send
Premature notify on fast action. User can't tell if it worked.
### No completion signal verification
Just clicking Generate and notifying without confirming generation started. If the click missed (semantic find on wrong element), the user thinks generation started when it didn't.
### Loop screenshot every 1s during wait
Wastes compute. 10-15s intervals are fine for waits.
### Ignore visible error toasts
If error appears, react. Don't proceed with "in progress" message if generation clearly didn't start.
## Operational Checklist (Per Action)
- [ ] Classify action via `async_action_classifier.py` (wait vs notify)
- [ ] If wait: set appropriate timeout (longer for file upload than chat)
- [ ] If notify: verify generation started via screenshot within 5s
- [ ] Tell user explicitly when fire-and-notify ("NotebookLM will notify you when ready")
- [ ] End task after notify — don't loop
- [ ] On wait timeout: screenshot + ask user how to proceed
- [ ] On visible error: report immediately
## Citations (7 sources)
1. **Anthropic API documentation — session timeout behavior.** Source for the 60s-5min API timeout ranges that motivate fire-and-notify for slow ops. https://docs.anthropic.com/
2. **Google NotebookLM Studio documentation + community findings (2024-2026).** Source for the per-output-type timing estimates (Audio Overview 5-10 min, Infographic 2-5 min, etc.). Empirical from user reports.
3. **Erlang / OTP "let it crash" philosophy (Joe Armstrong).** Source for the broader async pattern: don't wait for slow operations synchronously. Hand off + let the supervising system handle completion notifications.
4. **AWS Step Functions documentation.** Source for the fire-and-notify pattern in distributed systems. Step Functions explicitly distinguishes "Wait" tasks from "Callback" (fire-and-notify) tasks based on duration.
5. **Twelve-Factor App — Factor IX (Disposability).** Source for the "fast startup, graceful shutdown" discipline. Skills should be disposable — return control quickly rather than holding sessions open.
6. **Node.js event loop documentation.** Source for the async-event-loop discipline. Synchronous blocking on slow ops harms throughput; async handoff preserves it.
7. **Playwright + Selenium wait-strategy documentation.** Source for the polling-with-timeout pattern for "wait" actions. Both frameworks document the discipline of bounded waits with explicit timeouts.
FILE:references/browser_automation_canon.md
# Browser Automation Canon — Screenshot-First + find()-Before-Click
This reference answers exactly one decision: **what discipline does the notebooklm skill follow when controlling NotebookLM's UI, and why?**
## The Core Frame
NotebookLM is a **dynamic single-page application** where the UI varies significantly by:
- Account tier (free / Plus / Enterprise)
- Feature rollout (Studio types not yet GA for all users)
- Time (Google iterates the product frequently — buttons move, labels change, layouts shift)
- A/B experiments (different users see different UIs)
Hard-coded coordinates break instantly when these change. Semantic discipline survives.
## The Two Disciplines
### 1. Screenshot-First
**Every meaningful UI action must be preceded by a screenshot.**
Reasons:
1. **Verify UI matches expectations** before acting (account-tier differences, A/B variants)
2. **Catch login walls early** before attempting actions on a not-logged-in page
3. **Detect unexpected layout changes** (Google ships UI changes weekly)
4. **Audit trail** for debugging when actions fail
The performance cost is negligible (screenshots are fast). The safety benefit is large.
### 2. find()-Before-Click
**Use semantic element finders before pixel coordinates.**
Pattern preference:
| Pattern | Survives | Use when |
|---|---|---|
| `find(text="Audio Overview")` | UI rearrangements | Element has stable visible text |
| `find(aria_label="Generate")` | A11y-labeled elements | Element has stable aria-label |
| `find(data_test="studio-btn")` | Internal test IDs | Element has stable data-* attribute (rare in NotebookLM) |
| `find(role="button", name="Send")` | Accessibility tree | Element role + accessible name stable |
| `click(x=420, y=380)` | Nothing | **Last resort only** — UI must visually align exactly |
### Why coordinates break
Google rolled out a UI redesign in Q2 2025 that moved the Studio panel from right-side to a collapsible drawer. Skills using pixel coordinates broke overnight. Skills using `find(text="Studio")` adapted automatically.
## The Tool-Agnostic Vocabulary
NotebookLM skill uses generic terms — NOT hardcoded to a specific tool:
| Skill says | Tool-specific implementation examples |
|---|---|
| "browser automation tool" | Claude computer-use, Claude Chrome Extension, Playwright, Puppeteer |
| "screenshot tool" | `screenshot()`, `page.screenshot()`, `browser.captureScreenshot()` |
| "find tool" | `find_element()`, `page.locator()`, `$$(selector)` |
| "click tool" | `click()`, `page.click()`, `element.click()` |
| "navigate tool" | `navigate()`, `page.goto()`, `browser.open()` |
| "file-upload tool" | Tool-specific file-upload mechanism (not native file picker) |
Why: the skill should work across browser-automation environments. Tool-specific code locks the skill to one harness.
## Critical Discipline: Never Click Native File Pickers
When uploading a file to NotebookLM:
❌ **Do NOT click the "Choose File" button** — this opens a native OS file picker that browser automation cannot control.
✅ **Use the file-upload tool** — your browser-automation environment provides a way to attach files programmatically. Example patterns:
- Playwright: `page.set_input_files(input_ref, file_path)`
- Computer-use: `upload_file(file_input_locator, absolute_path)`
- Puppeteer: `page.$eval(input_selector, (el, path) => el.files = ...)`
The file-upload tool bypasses the native picker entirely.
## Login Wall Discipline
NotebookLM requires Google authentication. Login walls can appear:
- Initial visit (not logged in)
- Session timeout (logged in but stale)
- Account switch (multiple Google accounts)
- 2FA challenge (security verification)
**Hard rule: never attempt to handle login automatically.**
Why:
- Credentials are sensitive
- 2FA requires human interaction
- Account state varies (logged in to wrong account, etc.)
- Browser-automation handling login creates security audit nightmares
Detection pattern:
```
1. Take screenshot
2. Check for login signals: "Sign in to Google", login URL, password field visible
3. If detected → halt with clear message: "Please log in to NotebookLM in your browser, then re-invoke this skill."
4. Do not type credentials, do not click login buttons
```
## Async Wait Discipline (See Also: async_action_discipline.md)
Different actions have different timing:
- **Fast (< 5s)**: chat send, basic clicks — wait synchronously
- **Medium (5-60s)**: source ingestion, auto-summary, Study Guide / Briefing Doc — wait with timeout
- **Slow (1-10 min)**: Audio Overview, Infographic, Slides — **fire-and-notify, don't wait**
Mis-applying the timing causes:
- Session timeout (user waits 10 min for Audio Overview to complete)
- False failure reports (skill thinks it failed because it timed out waiting)
- Wasted compute
## Screenshot Audit Pattern
For debugging + transparency, the skill should produce a screenshot trail per session:
```
~/notebooklm_sessions/<date>-<action>/
001-step0-environment-check.png
002-homepage-loaded.png
003-notebook-found.png
004-chat-input-located.png
005-question-submitted.png
006-response-received.png
```
Per-screenshot rationale:
- Step 0 baseline (proves environment was checked)
- Each UI interaction documented
- Final state captured
This is the audit-log equivalent for browser-automation skills.
## Anti-Patterns
### "Just click where it usually is"
Pixel coordinates. Breaks on any UI change. Most common failure mode.
### "Skip screenshots to save time"
The cost of skipping is unrecoverable: when something goes wrong, no audit trail exists. Always screenshot.
### "Try to handle login programmatically"
Security risk + brittle. Always halt and ask user to log in manually.
### "Wait for Audio Overview synchronously"
5-10 minute generations exceed session timeout. Fire-and-notify pattern.
### "Use the main Studio button"
Default Studio prompts produce mediocre output. Always open the customization menu (chevron next to main button) and write a custom prompt.
### "Hardcode 'Claude Chrome Extension'"
Tool-specific. Use "browser automation tool" or equivalent generic term.
### "Click 'Choose File' for uploads"
Opens native file picker, breaks automation. Use the file-upload tool.
## Operational Checklist (Per UI Action)
- [ ] Screenshot taken before action
- [ ] Semantic find() attempted before pixel coordinates
- [ ] Login wall check after navigation
- [ ] Tool-agnostic vocabulary in skill body
- [ ] File uploads use file-upload tool, not file picker
- [ ] Async timing classified correctly (wait vs fire-and-notify)
- [ ] Screenshot trail saved for audit
## Citations (7 sources)
1. **Anthropic Computer Use documentation (Claude 3.5 Sonnet + later).** Source for the screenshot-first + semantic-find discipline in Claude's official browser automation tool. The patterns in this reference are the production patterns Anthropic recommends.
2. **Playwright documentation — playwright.dev.** Source for the semantic locator patterns (`page.locator()`, role-based selectors, accessibility tree). Playwright's locator model is the modern standard.
3. **Selenium documentation + best practices guides.** Source for the historical context: Selenium pioneered semantic finders 15+ years before Playwright; the patterns are well-established.
4. **W3C Web Driver protocol.** Source for the standardized element-locator strategies (id, name, class, link text, partial link text, tag name, css, xpath). The protocol formalized what to find by.
5. **Microsoft Power Automate Desktop UI automation guidance.** Source for the "image-based selectors are last resort" guidance. Microsoft's enterprise RPA tooling formalized the same discipline.
6. **Google A11y / ARIA patterns — w3.org/WAI/ARIA/.** Source for the role + accessible-name selector strategy. ARIA labels are the most stable selector when present.
7. **Anthropic's "Tool Use" + browser automation cookbook (docs.anthropic.com).** Source for the tool-agnostic vocabulary discipline. Anthropic explicitly recommends generic "screenshot tool" / "click tool" language to keep skills portable across harnesses.
FILE:references/studio_output_custom_prompts.md
# Studio Output Custom Prompts — Why Defaults Are Mediocre
This reference answers exactly one decision: **why does the notebooklm skill always open the Studio customization menu and write a detailed custom prompt, and what does a good custom prompt look like per output type?**
## The Core Claim
NotebookLM's Studio generates 9 output types from your notebook's sources:
- Audio Overview (podcast-style)
- Study Guide
- Briefing Doc
- Timeline
- FAQ
- Table of Contents
- Infographic
- Slides (slide deck)
- Mind Map
**The default prompts produce mediocre output.** They are written to work across all possible source materials → they are generic by design. Generic prompts produce generic output.
The mediocre-output failure mode:
- **Audio Overview default:** generic two-host summary, 12-15 min, undifferentiated voice
- **Infographic default:** title-and-bullet panels, no decision logic, generic palette
- **Study Guide default:** definition list + bland questions, no audience calibration
**Sharp custom prompts produce dramatically better output.** The customization menu (chevron next to the main Studio button) opens a text field where you describe: angle / audience / length / style / specific structure.
## When to Use Custom Prompts
**Always.** This is non-negotiable in the notebooklm skill.
The mandatory workflow for Action 3 (Studio output):
1. Locate the specific output button (find by text — e.g., `find(text="Audio Overview")`)
2. **Open the customization menu** — the chevron/dropdown arrow next to the main button (NOT the main button itself)
3. **Write a detailed custom prompt** in the text field
4. Submit Generate
Skipping step 2-3 (clicking the main button directly) uses the default prompt. The skill refuses to do this.
## Custom Prompt Anatomy
A good custom prompt specifies:
| Element | Why |
|---|---|
| **Audience** | Drives jargon level + assumed background |
| **Angle** | What perspective / lens to apply (e.g., business vs technical) |
| **Length** | Prevents bloat or thinness |
| **Structure** | Specific layout (panels, slides, sections) |
| **Style** | Tone, voice, formality |
| **Examples / counter-examples** | Concrete anchors |
A weak prompt has 1-2 of these. A strong prompt has 4-5.
## Per-Output-Type Custom Prompt Templates
### Audio Overview
**Default fails because:** generic two-host conversational tone, undefined audience, often 12-15 min (too long), surface-level coverage.
**Good custom prompt:**
> "Two-host conversation between a researcher and an experienced practitioner. Audience: [non-technical executive | technical lead | undergraduate student | general public]. Length: 8-10 minutes. Focus on [business implications | technical mechanism | historical evolution | practical applications]. Include one concrete example per major point. Acknowledge counter-arguments briefly. End with one specific takeaway, not a generic summary."
**Pattern:**
- Two-host setup (specify roles, not just "two hosts")
- Audience explicit
- Length tight
- Focus angle picked
- Concrete-example requirement
- Counter-argument requirement
- Specific closing instruction
### Infographic
**Default fails because:** generic title-and-bullets, no decision logic, oversize panel count, neutral palette.
**Good custom prompt:**
> "Decision-tree style. Action-oriented (each panel ends with a decision or action the viewer takes). 6 panels max. Monochrome navy with amber highlight on the action line. Each panel has: title (4-6 words), 1-2 sentence body explaining the situation, decision/action line in amber. No filler panels. Last panel: 'next step' with specific URL/contact/resource."
**Pattern:**
- Style (decision-tree vs storytelling vs comparison vs process)
- Action-orientation
- Panel count cap (6 is the sweet spot; 8+ becomes unreadable)
- Color palette
- Per-panel structure
- Closing call-to-action
### Study Guide
**Default fails because:** definition list + bland recall questions, no audience calibration, no Bloom higher-order discipline.
**Good custom prompt:**
> "[Undergraduate | Graduate | Professional] level — define every technical term if undergrad, assume fluency if grad. Structure: [N] concepts, each with 4 elements: (1) one-paragraph definition, (2) why this matters in practice (concrete example), (3) one worked problem, (4) 3 practice questions Bloom-higher-order (apply / analyze / evaluate; NO recall questions). Discussion questions tied to a specific learning outcome listed at the start of each section."
**Pattern:**
- Audience explicit (drives jargon + question complexity)
- Concept count specific
- Per-concept structure rigid
- Bloom-level discipline
- Learning-outcome anchoring
### Slides (Slide Deck)
**Default fails because:** dense slides, bullet-heavy, no presenter notes, generic structure.
**Good custom prompt:**
> "12 slides max. 1-2 sentences per slide body — NO bullet points in slide bodies (prose only). Per slide: include presenter notes with (a) one concrete example, (b) one likely audience objection, (c) how to address it. Title slide + 10 content slides + closing call-to-action slide. Closing slide: specific next step (not 'thank you')."
**Pattern:**
- Slide count cap (12 max for executive audience)
- Body density rule (1-2 sentences, no bullets)
- Presenter notes mandatory + structured
- Title + content + close structure
- Closing call-to-action (not generic 'thank you')
### Briefing Doc
**Default fails because:** generic memo structure, no key-decisions section, undifferentiated length.
**Good custom prompt:**
> "Audience: [executive / board / investor / partner]. Length: [1 page / 2 pages / 5 pages]. Structure: (1) BLUF (bottom line up front, 2 sentences max), (2) Key findings (3-5 numbered), (3) Decisions needed from this audience (numbered, with options + recommendation), (4) Open questions (3 max), (5) Suggested next step (1 specific action). Tone: [neutral analytical / persuasive / cautionary]."
### Timeline
**Default fails because:** chronological dump with no significance annotation.
**Good custom prompt:**
> "Milestone-focused (not event-dump). [N] milestones max. Per milestone: date, milestone (one phrase), significance (one sentence — why this changed the field). Order: reverse-chronological [or chronological if historical narrative]. Group into 3-4 eras with era-level summary at each boundary. Exclude minor events that don't shift the trajectory."
### FAQ
**Default fails because:** generic Q&A pairs, no audience-calibrated answer depth.
**Good custom prompt:**
> "Audience: [internal team / customer / external stakeholder]. 8-12 questions. Each answer: 2-3 sentences max. Question phrasing: how the audience would actually ask it (not how the topic owner would write it). Group into 3 categories. Include 2-3 'difficult question' entries (objections / concerns) — handle them directly, not evasively."
### Mind Map
**Default fails because:** generic radial spread with no priority hierarchy.
**Good custom prompt:**
> "Central concept: [name]. 3-5 primary branches (the major dimensions). Each branch: 2-4 sub-branches. Max depth: 3 levels (central → branch → sub-branch). Use noun phrases for branches (not full sentences). Mark 2-3 sub-branches as 'critical' (the highest-leverage points). Skip details that don't connect back to a critical sub-branch."
### Table of Contents
**Default fails because:** literal section dump, no annotation.
**Good custom prompt:**
> "Structure: section number + section title + 1-sentence summary of what the section covers. Length: 8-15 sections. Group into 2-3 parts with part-level summary at each boundary. Mark 2-3 sections as 'start here' for newcomers."
## Anti-Patterns
### Click the main Studio button (default prompt)
The most common skill failure. Always open customization menu.
### Generic custom prompt
> "Make a good infographic about my notebook."
This is a default prompt with extra words. Specify audience / angle / length / structure / style.
### Over-detailed custom prompt
> "Use exactly the color #1A3A5C for headers, then #E8F0F8 for table headers, with Arial 12pt body, then..."
NotebookLM's Studio engine doesn't honor pixel-precise design specs. Stay at the level of "monochrome navy with amber highlight," not exact hex codes.
### Custom prompt for the wrong output type
> Asked for Audio Overview, wrote a custom prompt about visual design.
Match prompt content to output type. Audio prompt → audio direction; visual prompt → visual direction.
### Custom prompt that contradicts source material
> "Generate an infographic showing my notebook's positive findings about X" — when the notebook has mixed findings.
Studio outputs source-grounded content. Custom prompts shape **how** content is presented, not what content exists. Force-skew via custom prompt produces inaccurate output.
## Operational Checklist (Per Studio Output Generation)
- [ ] Q4 (custom prompt direction) collected from user
- [ ] Studio panel located via find() (not pixel coordinates)
- [ ] Specific output button located (not the panel header)
- [ ] **Customization menu opened** (chevron, NOT main button)
- [ ] Custom prompt written: audience + angle + length + structure + style
- [ ] Custom prompt run through `scripts/custom_prompt_template_generator.py` for starter if needed
- [ ] Generate button clicked
- [ ] Confirmation screenshot taken
- [ ] User notified: "Generation in progress — NotebookLM will notify you when ready" (for slow ops)
## Citations (7 sources)
1. **Google NotebookLM documentation + product blog posts (2024-2026).** Source for the Studio output catalog and the customization menu existence. Google has progressively added customization since Audio Overview launched.
2. **Anthropic's prompt engineering guide — docs.anthropic.com.** Source for the audience-angle-length-structure-style anatomy of effective prompts. The pattern transfers from Claude prompting to NotebookLM Studio prompting.
3. **Refactoring UI — Adam Wathan & Steve Schoger (2018).** Source for visual-output design discipline (panel-count caps, palette restraint, action-orientation). The Infographic and Slides templates apply this discipline.
4. **Carmine Gallo, *Talk Like TED* (2014).** Source for the slide-deck structure (12-max, 1-2 sentences per slide, presenter notes with examples + objections). Gallo's analysis of top TED talks formalizes the discipline.
5. **Patrick Lencioni, *The Five Dysfunctions of a Team* (2002).** Source for the BLUF (bottom line up front) discipline used in Briefing Doc template. Executive briefings front-load the decision.
6. **Bloom's revised taxonomy — Anderson & Krathwohl (2001).** Source for the Study Guide's Bloom-higher-order discipline (apply / analyze / evaluate / create). Connects to research-pack's discussion question canon.
7. **Edward Tufte, *Visual Display of Quantitative Information* (1983, 2001 2nd ed.).** Source for the data-ink ratio discipline applied to Infographic + Timeline templates. Tufte's "small multiples" and "milestone discipline" inform the per-panel and per-milestone structures.
FILE:scripts/action_router.py
#!/usr/bin/env python3
"""action_router.py — Q1-Q4 answers → action plan + UI flow + required parameters.
Stdlib-only. Routes notebooklm intake answers to one of 4 action flows:
1. read_extract — chat-based extraction
2. add_source — push content via Add Source dialog (5 sub-types)
3. studio — generate Studio output (9 types) with mandatory custom prompt
4. create_new — new notebook with title + initial sources
Returns the action plan: required parameters, UI flow steps, mandatory checks.
NO LLM CALLS. Pure rule-based routing.
Usage:
python action_router.py --action read_extract --notebook "Q3 prep" --question "what are recent trends?"
python action_router.py --action add_source --notebook "Q3 prep" --source-type url --source-value "https://..."
python action_router.py --action studio --notebook "Q3 prep" --studio-type audio_overview --custom-prompt "..."
python action_router.py --action create_new --title "New project" --initial-sources "url1,url2"
python action_router.py --sample
"""
import argparse
import json
import sys
from typing import Any, Dict, List, Optional
VALID_ACTIONS = ["read_extract", "add_source", "studio", "create_new"]
VALID_SOURCE_TYPES = ["url", "text", "file", "google_doc", "synthesized"]
VALID_STUDIO_TYPES = [
"audio_overview", "study_guide", "briefing_doc", "timeline", "faq",
"table_of_contents", "infographic", "slides", "mind_map",
]
ACTION_FLOWS = {
"read_extract": {
"required_params": ["notebook", "question"],
"ui_flow": [
"Step 0: Browser environment check",
"Navigate to homepage → screenshot",
"Login wall check (halt if detected)",
"Notebook discovery (find by name or navigate URL)",
"Open notebook → screenshot",
"Locate chat input via find()",
"Type question (user's natural phrasing)",
"Submit (Enter or send button)",
"Wait 3-5s (synchronous; chat is fast)",
"Screenshot response area",
"Extract response in clean format (not raw chat dump)",
"Report to user",
],
"timing": "WAIT (3-10s)",
"screenshots_required": 4,
},
"add_source": {
"required_params": ["notebook", "source_type", "source_value"],
"ui_flow_by_source_type": {
"url": [
"Open notebook → screenshot",
"Click 'Add Source' → screenshot",
"Click 'Link' option",
"Paste URL",
"Submit → wait for ingestion spinner",
"Screenshot to confirm success",
],
"text": [
"Open notebook → screenshot",
"Click 'Add Source' → screenshot",
"Click 'Copied text' option",
"Paste content",
"Submit → wait for ingestion spinner",
"Screenshot to confirm success",
],
"file": [
"Open notebook → screenshot",
"Click 'Add Source' → screenshot",
"Click 'Upload file' option",
"Use file-upload tool with absolute path (NOT native file picker)",
"Wait for upload + ingestion",
"Screenshot to confirm success",
],
"google_doc": [
"Open notebook → screenshot",
"Click 'Add Source' → screenshot",
"Click 'Google Docs' option",
"Use Drive picker",
"Confirm doc selection",
"Wait for ingestion",
"Screenshot to confirm success",
],
"synthesized": [
"Pre-process content externally (extract main content, strip nav/ads)",
"Open notebook → screenshot",
"Click 'Add Source' → screenshot",
"Click 'Copied text' option",
"Paste synthesized content",
"Submit → wait for ingestion",
"Screenshot to confirm success",
],
},
"timing": "WAIT (5-30s with 60s timeout)",
"screenshots_required": 3,
},
"studio": {
"required_params": ["notebook", "studio_type", "custom_prompt"],
"ui_flow": [
"Step 0: Browser environment check",
"Navigate to notebook → screenshot",
"Locate Studio panel (right side; may need toggle) → screenshot",
"Find specific output button via find(text=studio_type)",
"**Open customization menu** (chevron NEXT to button; NOT main button)",
"**Write detailed custom prompt** in customization field",
"Submit Generate",
"**Verify generation started** via screenshot within 5s",
"**Fire-and-notify if slow** (Audio Overview, Infographic, Slides, Mind Map = 2-10 min)",
"Tell user: 'Generation in progress — NotebookLM will notify you when ready'",
"End task (don't wait for completion on slow ops)",
],
"timing_by_studio_type": {
"audio_overview": "FIRE_AND_NOTIFY (5-10 min)",
"study_guide": "WAIT (30-60s, timeout 90s)",
"briefing_doc": "WAIT (30-60s, timeout 90s)",
"timeline": "WAIT (30-60s, timeout 90s)",
"faq": "WAIT (30-60s, timeout 90s)",
"table_of_contents": "WAIT (20-40s, timeout 60s)",
"infographic": "FIRE_AND_NOTIFY (2-5 min)",
"slides": "FIRE_AND_NOTIFY (2-5 min)",
"mind_map": "FIRE_AND_NOTIFY (2-5 min)",
},
"screenshots_required": 5,
"critical_rule": "ALWAYS open customization menu (chevron) — NEVER click main Studio button (uses default mediocre prompt)",
},
"create_new": {
"required_params": ["title"],
"optional_params": ["initial_sources"],
"ui_flow": [
"Step 0: Browser environment check",
"Navigate to homepage → screenshot",
"Click 'New notebook' button (find via text)",
"Set title from --title argument",
"If initial_sources provided: add each via Action 2 sub-flow (per source type)",
"Wait for auto-summary generation (typically <30s)",
"Screenshot final state",
"Report new notebook URL to user",
],
"timing": "WAIT (15-30s for init + auto-summary)",
"screenshots_required": 3,
},
}
def route(action: str, **params) -> Dict[str, Any]:
if action not in VALID_ACTIONS:
raise ValueError(f"Invalid action '{action}'. Pick from: {VALID_ACTIONS}")
flow_template = ACTION_FLOWS[action].copy()
flow_template["action"] = action
flow_template["parameters"] = params
# Validate required params
required = flow_template.get("required_params", [])
missing = [p for p in required if not params.get(p)]
if missing:
flow_template["validation_errors"] = [f"Missing required parameter: {p}" for p in missing]
# Per-action customization
if action == "add_source":
source_type = params.get("source_type")
if source_type and source_type not in VALID_SOURCE_TYPES:
flow_template["validation_errors"] = flow_template.get("validation_errors", []) + [
f"Invalid source_type '{source_type}'. Pick from: {VALID_SOURCE_TYPES}"
]
elif source_type:
flow_template["ui_flow"] = flow_template["ui_flow_by_source_type"][source_type]
del flow_template["ui_flow_by_source_type"]
if action == "studio":
studio_type = params.get("studio_type")
if studio_type and studio_type not in VALID_STUDIO_TYPES:
flow_template["validation_errors"] = flow_template.get("validation_errors", []) + [
f"Invalid studio_type '{studio_type}'. Pick from: {VALID_STUDIO_TYPES}"
]
elif studio_type:
flow_template["timing"] = flow_template["timing_by_studio_type"][studio_type]
# Custom prompt is mandatory for studio
if not params.get("custom_prompt") or len(params.get("custom_prompt", "")) < 30:
flow_template["validation_errors"] = flow_template.get("validation_errors", []) + [
"Studio output requires DETAILED custom_prompt (min 30 chars). Default prompts produce mediocre output."
]
return flow_template
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"Action: {result['action']}")
out.append("")
out.append("Parameters:")
for k, v in result.get("parameters", {}).items():
if v:
display = v if len(str(v)) < 80 else str(v)[:77] + "..."
out.append(f" {k}: {display}")
out.append("")
if result.get("validation_errors"):
out.append("⚠️ Validation errors:")
for err in result["validation_errors"]:
out.append(f" - {err}")
out.append("")
out.append(f"Timing: {result.get('timing', 'N/A')}")
out.append(f"Screenshots required: {result.get('screenshots_required', 'N/A')}")
if result.get("critical_rule"):
out.append(f"⚠️ Critical rule: {result['critical_rule']}")
out.append("")
out.append("UI flow:")
flow = result.get("ui_flow", [])
for i, step in enumerate(flow, 1):
out.append(f" {i}. {step}")
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--action", choices=VALID_ACTIONS)
parser.add_argument("--notebook")
parser.add_argument("--question")
parser.add_argument("--source-type", choices=VALID_SOURCE_TYPES)
parser.add_argument("--source-value")
parser.add_argument("--studio-type", choices=VALID_STUDIO_TYPES)
parser.add_argument("--custom-prompt")
parser.add_argument("--title")
parser.add_argument("--initial-sources")
parser.add_argument("--sample", action="store_true")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
result = route(
"studio",
notebook="Q3 launch prep",
studio_type="audio_overview",
custom_prompt="Two-host conversation for non-technical executive, 8-10 min, focus on business implications not technical depth",
)
elif args.action:
params: Dict[str, Any] = {}
if args.notebook: params["notebook"] = args.notebook
if args.question: params["question"] = args.question
if args.source_type: params["source_type"] = args.source_type
if args.source_value: params["source_value"] = args.source_value
if args.studio_type: params["studio_type"] = args.studio_type
if args.custom_prompt: params["custom_prompt"] = args.custom_prompt
if args.title: params["title"] = args.title
if args.initial_sources: params["initial_sources"] = args.initial_sources
try:
result = route(args.action, **params)
except ValueError as e:
print(f"error: {e}", file=sys.stderr); return 2
else:
parser.print_help(); return 0
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0 if not result.get("validation_errors") else 1
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/async_action_classifier.py
#!/usr/bin/env python3
"""async_action_classifier.py — Classify NotebookLM action as wait-or-notify.
Stdlib-only. Given an action name, returns whether the skill should wait
synchronously OR fire-and-notify the user. The 2-minute dividing line:
- Under 2 minutes → WAIT (with appropriate timeout)
- Over 2 minutes → FIRE_AND_NOTIFY (browser session would time out)
Returns: action timing classification + estimated duration + recommended
timeout + fire-and-notify message template (if applicable).
NO LLM CALLS. Pure lookup table.
Usage:
python async_action_classifier.py --action audio_overview
python async_action_classifier.py --action add_source_url
python async_action_classifier.py --output json
python async_action_classifier.py --sample
"""
import argparse
import json
import sys
from typing import Any, Dict, List
ACTION_TIMING = {
# Read/Extract
"chat_send": {
"category": "read_extract",
"verdict": "WAIT",
"estimated_duration_seconds": (3, 10),
"timeout_seconds": 30,
"polling_interval_seconds": 3,
},
# Add Source sub-types
"add_source_url": {
"category": "add_source",
"verdict": "WAIT",
"estimated_duration_seconds": (5, 15),
"timeout_seconds": 60,
"polling_interval_seconds": 5,
},
"add_source_text": {
"category": "add_source",
"verdict": "WAIT",
"estimated_duration_seconds": (5, 15),
"timeout_seconds": 60,
"polling_interval_seconds": 5,
},
"add_source_file": {
"category": "add_source",
"verdict": "WAIT",
"estimated_duration_seconds": (10, 30),
"timeout_seconds": 90,
"polling_interval_seconds": 5,
},
"add_source_google_doc": {
"category": "add_source",
"verdict": "WAIT",
"estimated_duration_seconds": (10, 30),
"timeout_seconds": 90,
"polling_interval_seconds": 5,
},
"add_source_synthesized": {
"category": "add_source",
"verdict": "WAIT",
"estimated_duration_seconds": (5, 15),
"timeout_seconds": 60,
"polling_interval_seconds": 5,
},
# Create New
"create_new_notebook": {
"category": "create_new",
"verdict": "WAIT",
"estimated_duration_seconds": (15, 30),
"timeout_seconds": 60,
"polling_interval_seconds": 5,
},
# Studio outputs — fast
"study_guide": {
"category": "studio",
"verdict": "WAIT",
"estimated_duration_seconds": (30, 60),
"timeout_seconds": 120,
"polling_interval_seconds": 10,
},
"briefing_doc": {
"category": "studio",
"verdict": "WAIT",
"estimated_duration_seconds": (30, 60),
"timeout_seconds": 120,
"polling_interval_seconds": 10,
},
"timeline": {
"category": "studio",
"verdict": "WAIT",
"estimated_duration_seconds": (30, 60),
"timeout_seconds": 120,
"polling_interval_seconds": 10,
},
"faq": {
"category": "studio",
"verdict": "WAIT",
"estimated_duration_seconds": (30, 60),
"timeout_seconds": 120,
"polling_interval_seconds": 10,
},
"table_of_contents": {
"category": "studio",
"verdict": "WAIT",
"estimated_duration_seconds": (20, 40),
"timeout_seconds": 90,
"polling_interval_seconds": 5,
},
# Studio outputs — slow (fire-and-notify)
"audio_overview": {
"category": "studio",
"verdict": "FIRE_AND_NOTIFY",
"estimated_duration_seconds": (300, 600),
"estimated_duration_human": "5-10 minutes",
"notify_message": (
"Audio Overview generation triggered. Estimated 5-10 minutes. "
"NotebookLM will notify you in-app and via email when ready. "
"NOT waiting in this session — returning control to you now."
),
},
"infographic": {
"category": "studio",
"verdict": "FIRE_AND_NOTIFY",
"estimated_duration_seconds": (120, 300),
"estimated_duration_human": "2-5 minutes",
"notify_message": (
"Infographic generation triggered. Estimated 2-5 minutes. "
"NotebookLM will notify you when ready. "
"NOT waiting in this session — returning control."
),
},
"slides": {
"category": "studio",
"verdict": "FIRE_AND_NOTIFY",
"estimated_duration_seconds": (120, 300),
"estimated_duration_human": "2-5 minutes",
"notify_message": (
"Slides generation triggered. Estimated 2-5 minutes. "
"NotebookLM will notify you when ready. "
"NOT waiting in this session — returning control."
),
},
"mind_map": {
"category": "studio",
"verdict": "FIRE_AND_NOTIFY",
"estimated_duration_seconds": (120, 300),
"estimated_duration_human": "2-5 minutes",
"notify_message": (
"Mind Map generation triggered. Estimated 2-5 minutes. "
"NotebookLM will notify you when ready. "
"NOT waiting in this session — returning control."
),
},
}
def classify(action: str) -> Dict[str, Any]:
if action not in ACTION_TIMING:
# Try fuzzy match for synonyms
action_lower = action.lower().replace(" ", "_").replace("-", "_")
for known_action in ACTION_TIMING:
if action_lower in known_action or known_action in action_lower:
return {"action": known_action, "matched_from": action, **ACTION_TIMING[known_action]}
raise ValueError(f"Unknown action '{action}'. Known: {sorted(ACTION_TIMING.keys())}")
return {"action": action, **ACTION_TIMING[action]}
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"Action: {result['action']}")
if result.get('matched_from'):
out.append(f" (matched from: {result['matched_from']})")
out.append(f"Category: {result['category']}")
out.append(f"Verdict: {result['verdict']}")
out.append("")
dur = result['estimated_duration_seconds']
if isinstance(dur, (list, tuple)) and len(dur) == 2:
out.append(f"Estimated duration: {dur[0]}-{dur[1]} seconds")
out.append(f"Human duration: {result.get('estimated_duration_human', f'{dur}s')}")
out.append("")
if result['verdict'] == "WAIT":
out.append(f"Timeout: {result['timeout_seconds']}s")
out.append(f"Polling interval: {result['polling_interval_seconds']}s")
out.append("")
out.append("Wait discipline:")
out.append(" 1. Click trigger (after screenshot)")
out.append(" 2. Start wait loop with timeout")
out.append(f" 3. Poll every {result['polling_interval_seconds']}s for completion signal")
out.append(" 4. Screenshot at each poll for audit")
out.append(" 5. On timeout: screenshot, note delay, ask user")
else: # FIRE_AND_NOTIFY
out.append("Fire-and-notify message (paste into skill output):")
out.append("")
out.append(f" {result['notify_message']}")
out.append("")
out.append("Discipline:")
out.append(" 1. Click trigger (in customization menu, not main button)")
out.append(" 2. Verify generation started via screenshot WITHIN 5 seconds")
out.append(" 3. Tell user the notify message above")
out.append(" 4. END TASK — do not loop waiting for completion")
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--action", help="Action name (e.g., audio_overview, chat_send, add_source_url)")
parser.add_argument("--sample", action="store_true")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
result = classify("audio_overview")
elif args.action:
try:
result = classify(args.action)
except ValueError as e:
print(f"error: {e}", file=sys.stderr); return 2
else:
parser.print_help(); return 0
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/custom_prompt_template_generator.py
#!/usr/bin/env python3
"""custom_prompt_template_generator.py — Studio output type + audience → custom prompt starter.
Stdlib-only. Generates a starter custom prompt for NotebookLM Studio outputs
based on output type + audience + length + angle. Default prompts produce
mediocre output; this generator produces a sharper starter the user can refine.
Output types: audio_overview, study_guide, briefing_doc, timeline, faq,
table_of_contents, infographic, slides, mind_map.
NO LLM CALLS. Template-based prompt construction.
Usage:
python custom_prompt_template_generator.py --output-type audio_overview --audience executive --length 8min
python custom_prompt_template_generator.py --output-type infographic --audience consumer --angle decision-tree
python custom_prompt_template_generator.py --sample
"""
import argparse
import json
import sys
from typing import Any, Dict, List, Optional
VALID_OUTPUT_TYPES = [
"audio_overview", "study_guide", "briefing_doc", "timeline", "faq",
"table_of_contents", "infographic", "slides", "mind_map",
]
VALID_AUDIENCES = [
"executive", "technical_lead", "undergraduate", "graduate",
"consumer", "internal_team", "investor", "general_public",
]
VALID_ANGLES = {
"audio_overview": ["business_implications", "technical_mechanism", "historical_evolution", "practical_applications"],
"infographic": ["decision_tree", "process_flow", "comparison", "storytelling"],
"study_guide": ["definitions_first", "problem_solving", "case_study", "review_focused"],
"briefing_doc": ["neutral_analytical", "persuasive", "cautionary"],
"slides": ["narrative", "data_driven", "framework", "case_study"],
"timeline": ["chronological", "reverse_chronological", "era_grouped"],
"faq": ["onboarding", "objection_handling", "deep_dive", "troubleshooting"],
"mind_map": ["hierarchical", "interconnected", "priority_marked"],
"table_of_contents": ["sequential", "thematic", "audience_guided"],
}
def generate_audio_overview(audience: str, length: str, angle: str) -> str:
audience_descriptions = {
"executive": "non-technical executive making a budget or strategic decision",
"technical_lead": "senior technical lead evaluating an implementation approach",
"undergraduate": "undergraduate student new to the subject",
"graduate": "graduate student doing literature review",
"consumer": "interested general public, no specialized background",
"internal_team": "internal team member needing context to do their work",
"investor": "investor evaluating a thesis or opportunity",
"general_public": "intelligent general public reader",
}
angle_focuses = {
"business_implications": "business implications, ROI, market dynamics — NOT technical depth",
"technical_mechanism": "how the mechanism works under the hood, with concrete technical detail",
"historical_evolution": "how the field got to its current state — inflection points, paradigm shifts",
"practical_applications": "how this gets used in practice, with concrete examples per major point",
}
aud = audience_descriptions.get(audience, audience)
foc = angle_focuses.get(angle, angle)
return (
f"Two-host conversation between a researcher and an experienced practitioner. "
f"Audience: {aud}. Length: {length}. Focus: {foc}. "
"Include one concrete example per major point. Acknowledge counter-arguments briefly. "
"End with one specific takeaway, not a generic summary."
)
def generate_infographic(audience: str, length: str, angle: str) -> str:
angle_structures = {
"decision_tree": "Decision-tree style. Action-oriented (each panel ends with a decision/action). Branches lead to next panel.",
"process_flow": "Linear process flow. Panels are sequential steps. Each panel: step name + one-sentence what-happens + visual cue.",
"comparison": "Side-by-side comparison. 3-4 alternatives compared on 4-6 dimensions. Color-coded per alternative.",
"storytelling": "Narrative arc. Panels tell a story: situation → complication → resolution → takeaway.",
}
structure = angle_structures.get(angle, angle_structures["decision_tree"])
panel_count = "4-6 panels max" if length == "compact" else "6-8 panels"
return (
f"{structure} {panel_count}. "
f"Audience: {audience}. Monochrome navy with one accent color (amber highlight on key info). "
"Each panel: title (4-6 words), 1-2 sentence body, action/decision/next-step line. "
"No filler panels. Last panel: specific call-to-action with concrete next step (URL/contact/resource)."
)
def generate_study_guide(audience: str, length: str, angle: str) -> str:
jargon_rule = (
"Define every technical term. Assume zero specialized background."
if audience in ("undergraduate", "general_public", "consumer")
else "Assume technical fluency in the field. Brief context for novel concepts only."
if audience in ("graduate", "technical_lead")
else "Calibrate jargon to audience expertise."
)
concept_count = "4-6 core concepts" if length == "compact" else "6-10 core concepts"
return (
f"Audience: {audience}. {jargon_rule} "
f"Structure: {concept_count}, each with 4 elements: "
"(1) one-paragraph definition, "
"(2) why this matters in practice (concrete example), "
"(3) one worked problem or applied scenario, "
"(4) 3 practice questions Bloom-higher-order (apply / analyze / evaluate; NO recall questions). "
"Discussion questions tied to a specific learning outcome listed at the start of each section."
)
def generate_slides(audience: str, length: str, angle: str) -> str:
slide_count = "8-10 slides" if length == "compact" else "12-15 slides"
return (
f"{slide_count} max. Audience: {audience}. 1-2 sentences per slide body — "
"NO bullet points in slide bodies (prose only). "
"Per slide: include presenter notes with "
"(a) one concrete example, "
"(b) one likely audience objection, "
"(c) how to address it. "
"Title slide + content slides + closing call-to-action slide. "
"Closing slide: specific next step (not generic 'thank you')."
)
def generate_briefing_doc(audience: str, length: str, angle: str) -> str:
tone_descriptions = {
"neutral_analytical": "Neutral analytical tone. Present evidence, let it speak.",
"persuasive": "Persuasive tone with explicit recommendation backed by evidence.",
"cautionary": "Cautionary tone, surface risks and trade-offs prominently.",
}
tone = tone_descriptions.get(angle, tone_descriptions["neutral_analytical"])
length_pages = "1 page" if length == "compact" else "2-3 pages" if length == "standard" else "5 pages"
return (
f"Audience: {audience}. Length: {length_pages}. {tone} "
"Structure: "
"(1) BLUF (bottom line up front, 2 sentences max), "
"(2) Key findings (3-5 numbered), "
"(3) Decisions needed from this audience (numbered, with options + recommendation), "
"(4) Open questions (3 max), "
"(5) Suggested next step (1 specific action)."
)
def generate_timeline(audience: str, length: str, angle: str) -> str:
milestone_count = "5-8 milestones" if length == "compact" else "8-12 milestones"
direction = "Reverse-chronological (most recent first)" if angle == "reverse_chronological" else "Chronological (earliest first)"
return (
f"Milestone-focused (not event-dump). {milestone_count} max. "
"Per milestone: date, milestone (one phrase), significance (one sentence — why this changed the field). "
f"Order: {direction}. "
"Group into 3-4 eras with era-level summary at each boundary. "
"Exclude minor events that don't shift the trajectory."
)
def generate_faq(audience: str, length: str, angle: str) -> str:
question_count = "6-10 questions" if length == "compact" else "10-15 questions"
return (
f"Audience: {audience}. {question_count}. "
"Each answer: 2-3 sentences max. "
"Question phrasing: how the audience would actually ask it (not how the topic owner would write it). "
"Group into 3 categories. "
"Include 2-3 'difficult question' entries (objections / concerns) — handle them directly, not evasively."
)
def generate_mind_map(audience: str, length: str, angle: str) -> str:
return (
f"Audience: {audience}. Central concept clearly named. "
"3-5 primary branches (the major dimensions). Each branch: 2-4 sub-branches. "
"Max depth: 3 levels (central → branch → sub-branch). "
"Use noun phrases for branches (not full sentences). "
"Mark 2-3 sub-branches as 'critical' (the highest-leverage points). "
"Skip details that don't connect back to a critical sub-branch."
)
def generate_table_of_contents(audience: str, length: str, angle: str) -> str:
section_count = "6-10 sections" if length == "compact" else "10-15 sections"
return (
f"{section_count}. Each entry: section number + section title + 1-sentence summary of what the section covers. "
f"Audience: {audience}. "
"Group into 2-3 parts with part-level summary at each boundary. "
"Mark 2-3 sections as 'start here' for newcomers."
)
GENERATORS = {
"audio_overview": generate_audio_overview,
"infographic": generate_infographic,
"study_guide": generate_study_guide,
"slides": generate_slides,
"briefing_doc": generate_briefing_doc,
"timeline": generate_timeline,
"faq": generate_faq,
"mind_map": generate_mind_map,
"table_of_contents": generate_table_of_contents,
}
def generate(output_type: str, audience: str, length: str, angle: Optional[str] = None) -> Dict[str, Any]:
if output_type not in VALID_OUTPUT_TYPES:
raise ValueError(f"Invalid output_type '{output_type}'. Pick from: {VALID_OUTPUT_TYPES}")
if audience not in VALID_AUDIENCES:
raise ValueError(f"Invalid audience '{audience}'. Pick from: {VALID_AUDIENCES}")
valid_angles = VALID_ANGLES.get(output_type, [])
if angle and angle not in valid_angles:
raise ValueError(f"Invalid angle '{angle}' for output_type '{output_type}'. Pick from: {valid_angles}")
if not angle and valid_angles:
angle = valid_angles[0] # default to first
generator_fn = GENERATORS[output_type]
prompt = generator_fn(audience, length, angle or "")
return {
"output_type": output_type,
"audience": audience,
"length": length,
"angle": angle,
"custom_prompt": prompt,
"instructions": "Use this as STARTER text in the NotebookLM customization menu (chevron next to Studio output button, NOT main button). Refine the prompt further based on the specific notebook contents and user's intent.",
}
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"Output type: {result['output_type']}")
out.append(f"Audience: {result['audience']}")
out.append(f"Length: {result['length']}")
out.append(f"Angle: {result['angle'] or '(default)'}")
out.append("")
out.append("Custom prompt (paste into NotebookLM customization menu):")
out.append("")
out.append(result['custom_prompt'])
out.append("")
out.append(f"📝 {result['instructions']}")
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--output-type", choices=VALID_OUTPUT_TYPES)
parser.add_argument("--audience", choices=VALID_AUDIENCES)
parser.add_argument("--length", default="standard", choices=["compact", "standard", "deep"])
parser.add_argument("--angle")
parser.add_argument("--sample", action="store_true")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
result = generate("audio_overview", "executive", "compact", "business_implications")
elif args.output_type and args.audience:
try:
result = generate(args.output_type, args.audience, args.length, args.angle)
except ValueError as e:
print(f"error: {e}", file=sys.stderr); return 2
else:
parser.print_help(); return 0
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
Thiết kế chiến lược observability kết hợp metrics, logs, traces, gồm SLI/SLO, golden signals và tối ưu cảnh báo.
---
name: "observability-designer"
description: "Design production-ready observability strategies combining metrics, logs, and traces. Includes SLI/SLO design, golden-signals monitoring, alert optimization. Use when adding observability to a new service, refactoring alerting that is too noisy, or designing an SLO program before scaling production load."
---
# Observability Designer (POWERFUL)
**Category:** Engineering
**Tier:** POWERFUL
**Description:** Design comprehensive observability strategies for production systems including SLI/SLO frameworks, alerting optimization, and dashboard generation.
## Overview
Observability Designer enables you to create production-ready observability strategies that provide deep insights into system behavior, performance, and reliability. This skill combines the three pillars of observability (metrics, logs, traces) with proven frameworks like SLI/SLO design, golden signals monitoring, and alert optimization to create comprehensive observability solutions.
## Core Competencies
### SLI/SLO/SLA Framework Design
- **Service Level Indicators (SLI):** Define measurable signals that indicate service health
- **Service Level Objectives (SLO):** Set reliability targets based on user experience
- **Service Level Agreements (SLA):** Establish customer-facing commitments with consequences
- **Error Budget Management:** Calculate and track error budget consumption
- **Burn Rate Alerting:** Multi-window burn rate alerts for proactive SLO protection
### Three Pillars of Observability
#### Metrics
- **Golden Signals:** Latency, traffic, errors, and saturation monitoring
- **RED Method:** Rate, Errors, and Duration for request-driven services
- **USE Method:** Utilization, Saturation, and Errors for resource monitoring
- **Business Metrics:** Revenue, user engagement, and feature adoption tracking
- **Infrastructure Metrics:** CPU, memory, disk, network, and custom resource metrics
#### Logs
- **Structured Logging:** JSON-based log formats with consistent fields
- **Log Aggregation:** Centralized log collection and indexing strategies
- **Log Levels:** Appropriate use of DEBUG, INFO, WARN, ERROR, FATAL levels
- **Correlation IDs:** Request tracing through distributed systems
- **Log Sampling:** Volume management for high-throughput systems
#### Traces
- **Distributed Tracing:** End-to-end request flow visualization
- **Span Design:** Meaningful span boundaries and metadata
- **Trace Sampling:** Intelligent sampling strategies for performance and cost
- **Service Maps:** Automatic dependency discovery through traces
- **Root Cause Analysis:** Trace-driven debugging workflows
### Dashboard Design Principles
#### Information Architecture
- **Hierarchy:** Overview → Service → Component → Instance drill-down paths
- **Golden Ratio:** 80% operational metrics, 20% exploratory metrics
- **Cognitive Load:** Maximum 7±2 panels per dashboard screen
- **User Journey:** Role-based dashboard personas (SRE, Developer, Executive)
#### Visualization Best Practices
- **Chart Selection:** Time series for trends, heatmaps for distributions, gauges for status
- **Color Theory:** Red for critical, amber for warning, green for healthy states
- **Reference Lines:** SLO targets, capacity thresholds, and historical baselines
- **Time Ranges:** Default to meaningful windows (4h for incidents, 7d for trends)
#### Panel Design
- **Metric Queries:** Efficient Prometheus/InfluxDB queries with proper aggregation
- **Alerting Integration:** Visual alert state indicators on relevant panels
- **Interactive Elements:** Template variables, drill-down links, and annotation overlays
- **Performance:** Sub-second render times through query optimization
### Alert Design and Optimization
#### Alert Classification
- **Severity Levels:**
- **Critical:** Service down, SLO burn rate high
- **Warning:** Approaching thresholds, non-user-facing issues
- **Info:** Deployment notifications, capacity planning alerts
- **Actionability:** Every alert must have a clear response action
- **Alert Routing:** Escalation policies based on severity and team ownership
#### Alert Fatigue Prevention
- **Signal vs Noise:** High precision (few false positives) over high recall
- **Hysteresis:** Different thresholds for firing and resolving alerts
- **Suppression:** Dependent alert suppression during known outages
- **Grouping:** Related alerts grouped into single notifications
#### Alert Rule Design
- **Threshold Selection:** Statistical methods for threshold determination
- **Window Functions:** Appropriate averaging windows and percentile calculations
- **Alert Lifecycle:** Clear firing conditions and automatic resolution criteria
- **Testing:** Alert rule validation against historical data
### Runbook Generation and Incident Response
#### Runbook Structure
- **Alert Context:** What the alert means and why it fired
- **Impact Assessment:** User-facing vs internal impact evaluation
- **Investigation Steps:** Ordered troubleshooting procedures with time estimates
- **Resolution Actions:** Common fixes and escalation procedures
- **Post-Incident:** Follow-up tasks and prevention measures
#### Incident Detection Patterns
- **Anomaly Detection:** Statistical methods for detecting unusual patterns
- **Composite Alerts:** Multi-signal alerts for complex failure modes
- **Predictive Alerts:** Capacity and trend-based forward-looking alerts
- **Canary Monitoring:** Early detection through progressive deployment monitoring
### Golden Signals Framework
#### Latency Monitoring
- **Request Latency:** P50, P95, P99 response time tracking
- **Queue Latency:** Time spent waiting in processing queues
- **Network Latency:** Inter-service communication delays
- **Database Latency:** Query execution and connection pool metrics
#### Traffic Monitoring
- **Request Rate:** Requests per second with burst detection
- **Bandwidth Usage:** Network throughput and capacity utilization
- **User Sessions:** Active user tracking and session duration
- **Feature Usage:** API endpoint and feature adoption metrics
#### Error Monitoring
- **Error Rate:** 4xx and 5xx HTTP response code tracking
- **Error Budget:** SLO-based error rate targets and consumption
- **Error Distribution:** Error type classification and trending
- **Silent Failures:** Detection of processing failures without HTTP errors
#### Saturation Monitoring
- **Resource Utilization:** CPU, memory, disk, and network usage
- **Queue Depth:** Processing queue length and wait times
- **Connection Pools:** Database and service connection saturation
- **Rate Limiting:** API throttling and quota exhaustion tracking
### Distributed Tracing Strategies
#### Trace Architecture
- **Sampling Strategy:** Head-based, tail-based, and adaptive sampling
- **Trace Propagation:** Context propagation across service boundaries
- **Span Correlation:** Parent-child relationship modeling
- **Trace Storage:** Retention policies and storage optimization
#### Service Instrumentation
- **Auto-Instrumentation:** Framework-based automatic trace generation
- **Manual Instrumentation:** Custom span creation for business logic
- **Baggage Handling:** Cross-cutting concern propagation
- **Performance Impact:** Instrumentation overhead measurement and optimization
### Log Aggregation Patterns
#### Collection Architecture
- **Agent Deployment:** Log shipping agent strategies (push vs pull)
- **Log Routing:** Topic-based routing and filtering
- **Parsing Strategies:** Structured vs unstructured log handling
- **Schema Evolution:** Log format versioning and migration
#### Storage and Indexing
- **Index Design:** Optimized field indexing for common query patterns
- **Retention Policies:** Time and volume-based log retention
- **Compression:** Log data compression and archival strategies
- **Search Performance:** Query optimization and result caching
### Cost Optimization for Observability
#### Data Management
- **Metric Retention:** Tiered retention based on metric importance
- **Log Sampling:** Intelligent sampling to reduce ingestion costs
- **Trace Sampling:** Cost-effective trace collection strategies
- **Data Archival:** Cold storage for historical observability data
#### Resource Optimization
- **Query Efficiency:** Optimized metric and log queries
- **Storage Costs:** Appropriate storage tiers for different data types
- **Ingestion Rate Limiting:** Controlled data ingestion to manage costs
- **Cardinality Management:** High-cardinality metric detection and mitigation
## Scripts Overview
This skill includes three powerful Python scripts for comprehensive observability design:
### 1. SLO Designer (`slo_designer.py`)
Generates complete SLI/SLO frameworks based on service characteristics:
- **Input:** Service description JSON (type, criticality, dependencies)
- **Output:** SLI definitions, SLO targets, error budgets, burn rate alerts, SLA recommendations
- **Features:** Multi-window burn rate calculations, error budget policies, alert rule generation
### 2. Alert Optimizer (`alert_optimizer.py`)
Analyzes and optimizes existing alert configurations:
- **Input:** Alert configuration JSON with rules, thresholds, and routing
- **Output:** Optimization report and improved alert configuration
- **Features:** Noise detection, coverage gaps, duplicate identification, threshold optimization
### 3. Dashboard Generator (`dashboard_generator.py`)
Creates comprehensive dashboard specifications:
- **Input:** Service/system description JSON
- **Output:** Grafana-compatible dashboard JSON and documentation
- **Features:** Golden signals coverage, RED/USE methods, drill-down paths, role-based views
## Integration Patterns
### Monitoring Stack Integration
- **Prometheus:** Metric collection and alerting rule generation
- **Grafana:** Dashboard creation and visualization configuration
- **Elasticsearch/Kibana:** Log analysis and dashboard integration
- **Jaeger/Zipkin:** Distributed tracing configuration and analysis
### CI/CD Integration
- **Pipeline Monitoring:** Build, test, and deployment observability
- **Deployment Correlation:** Release impact tracking and rollback triggers
- **Feature Flag Monitoring:** A/B test and feature rollout observability
- **Performance Regression:** Automated performance monitoring in pipelines
### Incident Management Integration
- **PagerDuty/VictorOps:** Alert routing and escalation policies
- **Slack/Teams:** Notification and collaboration integration
- **JIRA/ServiceNow:** Incident tracking and resolution workflows
- **Post-Mortem:** Automated incident analysis and improvement tracking
## Advanced Patterns
### Multi-Cloud Observability
- **Cross-Cloud Metrics:** Unified metrics across AWS, GCP, Azure
- **Network Observability:** Inter-cloud connectivity monitoring
- **Cost Attribution:** Cloud resource cost tracking and optimization
- **Compliance Monitoring:** Security and compliance posture tracking
### Microservices Observability
- **Service Mesh Integration:** Istio/Linkerd observability configuration
- **API Gateway Monitoring:** Request routing and rate limiting observability
- **Container Orchestration:** Kubernetes cluster and workload monitoring
- **Service Discovery:** Dynamic service monitoring and health checks
### Machine Learning Observability
- **Model Performance:** Accuracy, drift, and bias monitoring
- **Feature Store Monitoring:** Feature quality and freshness tracking
- **Pipeline Observability:** ML pipeline execution and performance monitoring
- **A/B Test Analysis:** Statistical significance and business impact measurement
## Best Practices
### Organizational Alignment
- **SLO Setting:** Collaborative target setting between product and engineering
- **Alert Ownership:** Clear escalation paths and team responsibilities
- **Dashboard Governance:** Centralized dashboard management and standards
- **Training Programs:** Team education on observability tools and practices
### Technical Excellence
- **Infrastructure as Code:** Observability configuration version control
- **Testing Strategy:** Alert rule testing and dashboard validation
- **Performance Monitoring:** Observability system performance tracking
- **Security Considerations:** Access control and data privacy in observability
### Continuous Improvement
- **Metrics Review:** Regular SLI/SLO effectiveness assessment
- **Alert Tuning:** Ongoing alert threshold and routing optimization
- **Dashboard Evolution:** User feedback-driven dashboard improvements
- **Tool Evaluation:** Regular assessment of observability tool effectiveness
## Success Metrics
### Operational Metrics
- **Mean Time to Detection (MTTD):** How quickly issues are identified
- **Mean Time to Resolution (MTTR):** Time from detection to resolution
- **Alert Precision:** Percentage of actionable alerts
- **SLO Achievement:** Percentage of SLO targets met consistently
### Business Metrics
- **System Reliability:** Overall uptime and user experience quality
- **Engineering Velocity:** Development team productivity and deployment frequency
- **Cost Efficiency:** Observability cost as percentage of infrastructure spend
- **Customer Satisfaction:** User-reported reliability and performance satisfaction
This comprehensive observability design skill enables organizations to build robust, scalable monitoring and alerting systems that provide actionable insights while maintaining cost efficiency and operational excellence.
FILE:assets/sample_alerts.json
{
"alerts": [
{
"alert": "HighLatency",
"expr": "histogram_quantile(0.95, rate(http_request_duration_seconds_bucket{service=\"payment-service\"}[5m])) > 0.5",
"for": "5m",
"labels": {
"severity": "warning",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "High request latency detected",
"description": "95th percentile latency is {{ $value }}s for payment-service",
"runbook_url": "https://runbooks.company.com/high-latency"
},
"historical_data": {
"fires_per_day": 2.5,
"false_positive_rate": 0.15,
"average_duration_minutes": 12
}
},
{
"alert": "ServiceDown",
"expr": "up{service=\"payment-service\"} == 0",
"labels": {
"severity": "critical",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "Payment service is down",
"description": "Payment service has been down for more than 1 minute",
"runbook_url": "https://runbooks.company.com/service-down"
},
"historical_data": {
"fires_per_day": 0.1,
"false_positive_rate": 0.05,
"average_duration_minutes": 3
}
},
{
"alert": "HighErrorRate",
"expr": "sum(rate(http_requests_total{service=\"payment-service\",code=~\"5..\"}[5m])) / sum(rate(http_requests_total{service=\"payment-service\"}[5m])) > 0.01",
"for": "2m",
"labels": {
"severity": "warning",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "High error rate detected",
"description": "Error rate is {{ $value | humanizePercentage }} for payment-service",
"runbook_url": "https://runbooks.company.com/high-error-rate"
},
"historical_data": {
"fires_per_day": 1.8,
"false_positive_rate": 0.25,
"average_duration_minutes": 8
}
},
{
"alert": "HighCPUUsage",
"expr": "rate(process_cpu_seconds_total{service=\"payment-service\"}[5m]) * 100 > 80",
"labels": {
"severity": "warning",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "High CPU usage",
"description": "CPU usage is {{ $value }}% for payment-service"
},
"historical_data": {
"fires_per_day": 15.2,
"false_positive_rate": 0.8,
"average_duration_minutes": 45
}
},
{
"alert": "HighMemoryUsage",
"expr": "process_resident_memory_bytes{service=\"payment-service\"} / process_virtual_memory_max_bytes{service=\"payment-service\"} * 100 > 85",
"labels": {
"severity": "info",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "High memory usage",
"description": "Memory usage is {{ $value }}% for payment-service"
},
"historical_data": {
"fires_per_day": 8.5,
"false_positive_rate": 0.6,
"average_duration_minutes": 30
}
},
{
"alert": "DatabaseConnectionPoolExhaustion",
"expr": "db_connections_active{service=\"payment-service\"} / db_connections_max{service=\"payment-service\"} > 0.9",
"for": "1m",
"labels": {
"severity": "critical",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "Database connection pool near exhaustion",
"description": "Connection pool utilization is {{ $value | humanizePercentage }}",
"runbook_url": "https://runbooks.company.com/db-connections"
},
"historical_data": {
"fires_per_day": 0.3,
"false_positive_rate": 0.1,
"average_duration_minutes": 5
}
},
{
"alert": "LowTraffic",
"expr": "sum(rate(http_requests_total{service=\"payment-service\"}[5m])) < 10",
"for": "10m",
"labels": {
"severity": "warning",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "Unusually low traffic",
"description": "Request rate is {{ $value }} RPS, which is unusually low"
},
"historical_data": {
"fires_per_day": 12.0,
"false_positive_rate": 0.9,
"average_duration_minutes": 120
}
},
{
"alert": "HighLatencyDuplicate",
"expr": "histogram_quantile(0.95, rate(http_request_duration_seconds_bucket{service=\"payment-service\"}[5m])) > 0.5",
"for": "5m",
"labels": {
"severity": "warning",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "High request latency detected (duplicate)",
"description": "95th percentile latency is {{ $value }}s for payment-service"
},
"historical_data": {
"fires_per_day": 2.5,
"false_positive_rate": 0.15,
"average_duration_minutes": 12
}
},
{
"alert": "VeryLowErrorRate",
"expr": "sum(rate(http_requests_total{service=\"payment-service\",code=~\"5..\"}[5m])) / sum(rate(http_requests_total{service=\"payment-service\"}[5m])) > 0.001",
"labels": {
"severity": "info",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "Error rate above 0.1%",
"description": "Error rate is {{ $value | humanizePercentage }}"
},
"historical_data": {
"fires_per_day": 25.0,
"false_positive_rate": 0.95,
"average_duration_minutes": 5
}
},
{
"alert": "DiskUsageHigh",
"expr": "disk_usage_percent{service=\"payment-service\"} > 85",
"labels": {
"severity": "warning",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "Disk usage high",
"description": "Disk usage is {{ $value }}%"
},
"historical_data": {
"fires_per_day": 3.2,
"false_positive_rate": 0.4,
"average_duration_minutes": 240
}
}
],
"services": [
{
"name": "payment-service",
"type": "api",
"criticality": "critical",
"team": "payments"
},
{
"name": "user-service",
"type": "api",
"criticality": "high",
"team": "identity"
},
{
"name": "notification-service",
"type": "api",
"criticality": "medium",
"team": "communications"
}
],
"alert_routing": {
"routes": [
{
"match": {
"severity": "critical"
},
"receiver": "pager-critical",
"group_wait": "10s",
"group_interval": "1m",
"repeat_interval": "5m"
},
{
"match": {
"severity": "warning"
},
"receiver": "slack-warnings",
"group_wait": "30s",
"group_interval": "5m",
"repeat_interval": "1h"
},
{
"match": {
"severity": "info"
},
"receiver": "email-info",
"group_wait": "2m",
"group_interval": "10m",
"repeat_interval": "24h"
}
]
},
"receivers": [
{
"name": "pager-critical",
"pagerduty_configs": [
{
"routing_key": "pager-key-critical",
"description": "Critical alert: {{ range .Alerts }}{{ .Annotations.summary }}{{ end }}"
}
]
},
{
"name": "slack-warnings",
"slack_configs": [
{
"api_url": "https://hooks.slack.com/services/warnings",
"channel": "#alerts-warnings",
"title": "Warning Alert",
"text": "{{ range .Alerts }}{{ .Annotations.description }}{{ end }}"
}
]
},
{
"name": "email-info",
"email_configs": [
{
"to": "team-notifications@company.com",
"subject": "Info Alert: {{ .GroupLabels.alertname }}",
"body": "{{ range .Alerts }}{{ .Annotations.description }}{{ end }}"
}
]
}
]
}
FILE:assets/sample_service_api.json
{
"name": "payment-service",
"type": "api",
"criticality": "critical",
"user_facing": true,
"description": "Handles payment processing and transaction management",
"team": "payments",
"environment": "production",
"dependencies": [
{
"name": "user-service",
"type": "api",
"criticality": "high"
},
{
"name": "payment-gateway",
"type": "external",
"criticality": "critical"
},
{
"name": "fraud-detection",
"type": "ml",
"criticality": "high"
}
],
"endpoints": [
{
"path": "/api/v1/payments",
"method": "POST",
"sla_latency_ms": 500,
"expected_tps": 100
},
{
"path": "/api/v1/payments/{id}",
"method": "GET",
"sla_latency_ms": 200,
"expected_tps": 500
},
{
"path": "/api/v1/payments/{id}/refund",
"method": "POST",
"sla_latency_ms": 1000,
"expected_tps": 10
}
],
"business_metrics": {
"revenue_per_hour": {
"metric": "sum(payment_amount * rate(payments_successful_total[1h]))",
"target": 50000,
"unit": "USD"
},
"conversion_rate": {
"metric": "sum(rate(payments_successful_total[5m])) / sum(rate(payment_attempts_total[5m]))",
"target": 0.95,
"unit": "percentage"
}
},
"infrastructure": {
"container_orchestrator": "kubernetes",
"replicas": 6,
"cpu_limit": "2000m",
"memory_limit": "4Gi",
"database": {
"type": "postgresql",
"connection_pool_size": 20
},
"cache": {
"type": "redis",
"cluster_size": 3
}
},
"compliance_requirements": [
"PCI-DSS",
"SOX",
"GDPR"
],
"tags": [
"payment",
"transaction",
"critical-path",
"revenue-generating"
]
}
FILE:assets/sample_service_web.json
{
"name": "customer-portal",
"type": "web",
"criticality": "high",
"user_facing": true,
"description": "Customer-facing web application for account management and billing",
"team": "frontend",
"environment": "production",
"dependencies": [
{
"name": "user-service",
"type": "api",
"criticality": "high"
},
{
"name": "billing-service",
"type": "api",
"criticality": "high"
},
{
"name": "notification-service",
"type": "api",
"criticality": "medium"
},
{
"name": "cdn",
"type": "external",
"criticality": "medium"
}
],
"pages": [
{
"path": "/dashboard",
"sla_load_time_ms": 2000,
"expected_concurrent_users": 1000
},
{
"path": "/billing",
"sla_load_time_ms": 3000,
"expected_concurrent_users": 200
},
{
"path": "/settings",
"sla_load_time_ms": 1500,
"expected_concurrent_users": 100
}
],
"business_metrics": {
"daily_active_users": {
"metric": "count(user_sessions_started_total[1d])",
"target": 10000,
"unit": "users"
},
"session_duration": {
"metric": "avg(user_session_duration_seconds)",
"target": 300,
"unit": "seconds"
},
"bounce_rate": {
"metric": "sum(rate(page_views_bounced_total[1h])) / sum(rate(page_views_total[1h]))",
"target": 0.3,
"unit": "percentage"
}
},
"infrastructure": {
"container_orchestrator": "kubernetes",
"replicas": 4,
"cpu_limit": "1000m",
"memory_limit": "2Gi",
"storage": {
"type": "nfs",
"size": "50Gi"
},
"ingress": {
"type": "nginx",
"ssl_termination": true,
"rate_limiting": {
"requests_per_second": 100,
"burst": 200
}
}
},
"monitoring": {
"synthetic_checks": [
{
"name": "login_flow",
"url": "/auth/login",
"frequency": "1m",
"locations": ["us-east", "eu-west", "ap-south"]
},
{
"name": "checkout_flow",
"url": "/billing/checkout",
"frequency": "5m",
"locations": ["us-east", "eu-west"]
}
],
"rum": {
"enabled": true,
"sampling_rate": 0.1
}
},
"compliance_requirements": [
"GDPR",
"CCPA"
],
"tags": [
"frontend",
"customer-facing",
"billing",
"high-traffic"
]
}
FILE:expected_outputs/sample_dashboard.json
{
"metadata": {
"title": "customer-portal - SRE Dashboard",
"service": {
"name": "customer-portal",
"type": "web",
"criticality": "high",
"user_facing": true,
"description": "Customer-facing web application for account management and billing",
"team": "frontend",
"environment": "production",
"dependencies": [
{
"name": "user-service",
"type": "api",
"criticality": "high"
},
{
"name": "billing-service",
"type": "api",
"criticality": "high"
},
{
"name": "notification-service",
"type": "api",
"criticality": "medium"
},
{
"name": "cdn",
"type": "external",
"criticality": "medium"
}
],
"pages": [
{
"path": "/dashboard",
"sla_load_time_ms": 2000,
"expected_concurrent_users": 1000
},
{
"path": "/billing",
"sla_load_time_ms": 3000,
"expected_concurrent_users": 200
},
{
"path": "/settings",
"sla_load_time_ms": 1500,
"expected_concurrent_users": 100
}
],
"business_metrics": {
"daily_active_users": {
"metric": "count(user_sessions_started_total[1d])",
"target": 10000,
"unit": "users"
},
"session_duration": {
"metric": "avg(user_session_duration_seconds)",
"target": 300,
"unit": "seconds"
},
"bounce_rate": {
"metric": "sum(rate(page_views_bounced_total[1h])) / sum(rate(page_views_total[1h]))",
"target": 0.3,
"unit": "percentage"
}
},
"infrastructure": {
"container_orchestrator": "kubernetes",
"replicas": 4,
"cpu_limit": "1000m",
"memory_limit": "2Gi",
"storage": {
"type": "nfs",
"size": "50Gi"
},
"ingress": {
"type": "nginx",
"ssl_termination": true,
"rate_limiting": {
"requests_per_second": 100,
"burst": 200
}
}
},
"monitoring": {
"synthetic_checks": [
{
"name": "login_flow",
"url": "/auth/login",
"frequency": "1m",
"locations": [
"us-east",
"eu-west",
"ap-south"
]
},
{
"name": "checkout_flow",
"url": "/billing/checkout",
"frequency": "5m",
"locations": [
"us-east",
"eu-west"
]
}
],
"rum": {
"enabled": true,
"sampling_rate": 0.1
}
},
"compliance_requirements": [
"GDPR",
"CCPA"
],
"tags": [
"frontend",
"customer-facing",
"billing",
"high-traffic"
]
},
"target_role": "sre",
"generated_at": "2026-02-16T14:02:03.421248Z",
"version": "1.0"
},
"configuration": {
"time_ranges": [
"1h",
"6h",
"1d",
"7d"
],
"default_time_range": "6h",
"refresh_interval": "30s",
"timezone": "UTC",
"theme": "dark"
},
"layout": {
"grid_settings": {
"width": 24,
"height_unit": "px",
"cell_height": 30
},
"sections": [
{
"title": "Service Overview",
"collapsed": false,
"y_position": 0,
"panels": [
"service_status",
"slo_summary",
"error_budget"
]
},
{
"title": "Golden Signals",
"collapsed": false,
"y_position": 8,
"panels": [
"latency",
"traffic",
"errors",
"saturation"
]
},
{
"title": "Resource Utilization",
"collapsed": false,
"y_position": 16,
"panels": [
"cpu_usage",
"memory_usage",
"network_io",
"disk_io"
]
},
{
"title": "Dependencies & Downstream",
"collapsed": true,
"y_position": 24,
"panels": [
"dependency_status",
"downstream_latency",
"circuit_breakers"
]
}
]
},
"panels": [
{
"id": "service_status",
"title": "Service Status",
"type": "stat",
"grid_pos": {
"x": 0,
"y": 0,
"w": 6,
"h": 4
},
"targets": [
{
"expr": "up{service=\"customer-portal\"}",
"legendFormat": "Status"
}
],
"field_config": {
"overrides": [
{
"matcher": {
"id": "byName",
"options": "Status"
},
"properties": [
{
"id": "color",
"value": {
"mode": "thresholds"
}
},
{
"id": "thresholds",
"value": {
"steps": [
{
"color": "red",
"value": 0
},
{
"color": "green",
"value": 1
}
]
}
},
{
"id": "mappings",
"value": [
{
"options": {
"0": {
"text": "DOWN"
}
},
"type": "value"
},
{
"options": {
"1": {
"text": "UP"
}
},
"type": "value"
}
]
}
]
}
]
},
"options": {
"orientation": "horizontal",
"textMode": "value_and_name"
}
},
{
"id": "slo_summary",
"title": "SLO Achievement (30d)",
"type": "stat",
"grid_pos": {
"x": 6,
"y": 0,
"w": 9,
"h": 4
},
"targets": [
{
"expr": "(1 - (increase(http_requests_total{service=\"customer-portal\",code=~\"5..\"}[30d]) / increase(http_requests_total{service=\"customer-portal\"}[30d]))) * 100",
"legendFormat": "Availability"
},
{
"expr": "histogram_quantile(0.95, increase(http_request_duration_seconds_bucket{service=\"customer-portal\"}[30d])) * 1000",
"legendFormat": "P95 Latency (ms)"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "thresholds"
},
"thresholds": {
"steps": [
{
"color": "red",
"value": 0
},
{
"color": "yellow",
"value": 99.0
},
{
"color": "green",
"value": 99.9
}
]
}
}
},
"options": {
"orientation": "horizontal",
"textMode": "value_and_name"
}
},
{
"id": "error_budget",
"title": "Error Budget Remaining",
"type": "gauge",
"grid_pos": {
"x": 15,
"y": 0,
"w": 9,
"h": 4
},
"targets": [
{
"expr": "(1 - (increase(http_requests_total{service=\"customer-portal\",code=~\"5..\"}[30d]) / increase(http_requests_total{service=\"customer-portal\"}[30d])) - 0.999) / 0.001 * 100",
"legendFormat": "Error Budget %"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "thresholds"
},
"min": 0,
"max": 100,
"thresholds": {
"steps": [
{
"color": "red",
"value": 0
},
{
"color": "yellow",
"value": 25
},
{
"color": "green",
"value": 50
}
]
},
"unit": "percent"
}
},
"options": {
"showThresholdLabels": true,
"showThresholdMarkers": true
}
},
{
"id": "latency",
"title": "Request Latency",
"type": "timeseries",
"grid_pos": {
"x": 0,
"y": 8,
"w": 12,
"h": 6
},
"targets": [
{
"expr": "histogram_quantile(0.50, rate(http_request_duration_seconds_bucket{service=\"customer-portal\"}[5m])) * 1000",
"legendFormat": "P50 Latency"
},
{
"expr": "histogram_quantile(0.95, rate(http_request_duration_seconds_bucket{service=\"customer-portal\"}[5m])) * 1000",
"legendFormat": "P95 Latency"
},
{
"expr": "histogram_quantile(0.99, rate(http_request_duration_seconds_bucket{service=\"customer-portal\"}[5m])) * 1000",
"legendFormat": "P99 Latency"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"unit": "ms",
"custom": {
"drawStyle": "line",
"lineInterpolation": "linear",
"lineWidth": 1,
"fillOpacity": 10
}
}
},
"options": {
"tooltip": {
"mode": "multi",
"sort": "desc"
},
"legend": {
"displayMode": "table",
"placement": "bottom"
}
}
},
{
"id": "traffic",
"title": "Request Rate",
"type": "timeseries",
"grid_pos": {
"x": 12,
"y": 8,
"w": 12,
"h": 6
},
"targets": [
{
"expr": "sum(rate(http_requests_total{service=\"customer-portal\"}[5m]))",
"legendFormat": "Total RPS"
},
{
"expr": "sum(rate(http_requests_total{service=\"customer-portal\",code=~\"2..\"}[5m]))",
"legendFormat": "2xx RPS"
},
{
"expr": "sum(rate(http_requests_total{service=\"customer-portal\",code=~\"4..\"}[5m]))",
"legendFormat": "4xx RPS"
},
{
"expr": "sum(rate(http_requests_total{service=\"customer-portal\",code=~\"5..\"}[5m]))",
"legendFormat": "5xx RPS"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"unit": "reqps",
"custom": {
"drawStyle": "line",
"lineInterpolation": "linear",
"lineWidth": 1,
"fillOpacity": 0
}
}
},
"options": {
"tooltip": {
"mode": "multi",
"sort": "desc"
},
"legend": {
"displayMode": "table",
"placement": "bottom"
}
}
},
{
"id": "errors",
"title": "Error Rate",
"type": "timeseries",
"grid_pos": {
"x": 0,
"y": 14,
"w": 12,
"h": 6
},
"targets": [
{
"expr": "sum(rate(http_requests_total{service=\"customer-portal\",code=~\"5..\"}[5m])) / sum(rate(http_requests_total{service=\"customer-portal\"}[5m])) * 100",
"legendFormat": "5xx Error Rate"
},
{
"expr": "sum(rate(http_requests_total{service=\"customer-portal\",code=~\"4..\"}[5m])) / sum(rate(http_requests_total{service=\"customer-portal\"}[5m])) * 100",
"legendFormat": "4xx Error Rate"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"unit": "percent",
"custom": {
"drawStyle": "line",
"lineInterpolation": "linear",
"lineWidth": 2,
"fillOpacity": 20
}
},
"overrides": [
{
"matcher": {
"id": "byName",
"options": "5xx Error Rate"
},
"properties": [
{
"id": "color",
"value": {
"fixedColor": "red"
}
}
]
}
]
},
"options": {
"tooltip": {
"mode": "multi",
"sort": "desc"
},
"legend": {
"displayMode": "table",
"placement": "bottom"
}
}
},
{
"id": "saturation",
"title": "Saturation Metrics",
"type": "timeseries",
"grid_pos": {
"x": 12,
"y": 14,
"w": 12,
"h": 6
},
"targets": [
{
"expr": "rate(process_cpu_seconds_total{service=\"customer-portal\"}[5m]) * 100",
"legendFormat": "CPU Usage %"
},
{
"expr": "process_resident_memory_bytes{service=\"customer-portal\"} / process_virtual_memory_max_bytes{service=\"customer-portal\"} * 100",
"legendFormat": "Memory Usage %"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"unit": "percent",
"max": 100,
"custom": {
"drawStyle": "line",
"lineInterpolation": "linear",
"lineWidth": 1,
"fillOpacity": 10
}
}
},
"options": {
"tooltip": {
"mode": "multi",
"sort": "desc"
},
"legend": {
"displayMode": "table",
"placement": "bottom"
}
}
},
{
"id": "cpu_usage",
"title": "CPU Usage",
"type": "gauge",
"grid_pos": {
"x": 0,
"y": 20,
"w": 6,
"h": 4
},
"targets": [
{
"expr": "rate(process_cpu_seconds_total{service=\"customer-portal\"}[5m]) * 100",
"legendFormat": "CPU %"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "thresholds"
},
"unit": "percent",
"min": 0,
"max": 100,
"thresholds": {
"steps": [
{
"color": "green",
"value": 0
},
{
"color": "yellow",
"value": 70
},
{
"color": "red",
"value": 90
}
]
}
}
},
"options": {
"showThresholdLabels": true,
"showThresholdMarkers": true
}
},
{
"id": "memory_usage",
"title": "Memory Usage",
"type": "gauge",
"grid_pos": {
"x": 6,
"y": 20,
"w": 6,
"h": 4
},
"targets": [
{
"expr": "process_resident_memory_bytes{service=\"customer-portal\"} / 1024 / 1024",
"legendFormat": "Memory MB"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "thresholds"
},
"unit": "decbytes",
"thresholds": {
"steps": [
{
"color": "green",
"value": 0
},
{
"color": "yellow",
"value": 512000000
},
{
"color": "red",
"value": 1024000000
}
]
}
}
}
},
{
"id": "network_io",
"title": "Network I/O",
"type": "timeseries",
"grid_pos": {
"x": 12,
"y": 20,
"w": 6,
"h": 4
},
"targets": [
{
"expr": "rate(process_network_receive_bytes_total{service=\"customer-portal\"}[5m])",
"legendFormat": "RX Bytes/s"
},
{
"expr": "rate(process_network_transmit_bytes_total{service=\"customer-portal\"}[5m])",
"legendFormat": "TX Bytes/s"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"unit": "binBps"
}
}
},
{
"id": "disk_io",
"title": "Disk I/O",
"type": "timeseries",
"grid_pos": {
"x": 18,
"y": 20,
"w": 6,
"h": 4
},
"targets": [
{
"expr": "rate(process_disk_read_bytes_total{service=\"customer-portal\"}[5m])",
"legendFormat": "Read Bytes/s"
},
{
"expr": "rate(process_disk_write_bytes_total{service=\"customer-portal\"}[5m])",
"legendFormat": "Write Bytes/s"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"unit": "binBps"
}
}
}
],
"variables": [
{
"name": "environment",
"type": "query",
"query": "label_values(environment)",
"current": {
"text": "production",
"value": "production"
},
"includeAll": false,
"multi": false,
"refresh": "on_dashboard_load"
},
{
"name": "instance",
"type": "query",
"query": "label_values(up{service=\"customer-portal\"}, instance)",
"current": {
"text": "All",
"value": "$__all"
},
"includeAll": true,
"multi": true,
"refresh": "on_time_range_change"
},
{
"name": "handler",
"type": "query",
"query": "label_values(http_requests_total{service=\"customer-portal\"}, handler)",
"current": {
"text": "All",
"value": "$__all"
},
"includeAll": true,
"multi": true,
"refresh": "on_time_range_change"
}
],
"alerts_integration": {
"alert_annotations": true,
"alert_rules_query": "ALERTS{service=\"customer-portal\"}",
"alert_panels": [
{
"title": "Active Alerts",
"type": "table",
"query": "ALERTS{service=\"customer-portal\",alertstate=\"firing\"}",
"columns": [
"alertname",
"severity",
"instance",
"description"
]
}
]
},
"drill_down_paths": {
"service_overview": {
"from": "service_status",
"to": "detailed_health_dashboard",
"url": "/d/service-health/customer-portal-health",
"params": [
"var-service",
"var-environment"
]
},
"error_investigation": {
"from": "errors",
"to": "error_details_dashboard",
"url": "/d/errors/customer-portal-errors",
"params": [
"var-service",
"var-time_range"
]
},
"latency_analysis": {
"from": "latency",
"to": "trace_analysis_dashboard",
"url": "/d/traces/customer-portal-traces",
"params": [
"var-service",
"var-handler"
]
},
"capacity_planning": {
"from": "saturation",
"to": "capacity_dashboard",
"url": "/d/capacity/customer-portal-capacity",
"params": [
"var-service",
"var-time_range"
]
}
}
}
FILE:expected_outputs/sample_slo_framework.json
{
"metadata": {
"service": {
"name": "payment-service",
"type": "api",
"criticality": "critical",
"user_facing": true,
"description": "Handles payment processing and transaction management",
"team": "payments",
"environment": "production",
"dependencies": [
{
"name": "user-service",
"type": "api",
"criticality": "high"
},
{
"name": "payment-gateway",
"type": "external",
"criticality": "critical"
},
{
"name": "fraud-detection",
"type": "ml",
"criticality": "high"
}
],
"endpoints": [
{
"path": "/api/v1/payments",
"method": "POST",
"sla_latency_ms": 500,
"expected_tps": 100
},
{
"path": "/api/v1/payments/{id}",
"method": "GET",
"sla_latency_ms": 200,
"expected_tps": 500
},
{
"path": "/api/v1/payments/{id}/refund",
"method": "POST",
"sla_latency_ms": 1000,
"expected_tps": 10
}
],
"business_metrics": {
"revenue_per_hour": {
"metric": "sum(payment_amount * rate(payments_successful_total[1h]))",
"target": 50000,
"unit": "USD"
},
"conversion_rate": {
"metric": "sum(rate(payments_successful_total[5m])) / sum(rate(payment_attempts_total[5m]))",
"target": 0.95,
"unit": "percentage"
}
},
"infrastructure": {
"container_orchestrator": "kubernetes",
"replicas": 6,
"cpu_limit": "2000m",
"memory_limit": "4Gi",
"database": {
"type": "postgresql",
"connection_pool_size": 20
},
"cache": {
"type": "redis",
"cluster_size": 3
}
},
"compliance_requirements": [
"PCI-DSS",
"SOX",
"GDPR"
],
"tags": [
"payment",
"transaction",
"critical-path",
"revenue-generating"
]
},
"generated_at": "2026-02-16T14:01:57.572080Z",
"framework_version": "1.0"
},
"slis": [
{
"name": "Availability",
"description": "Percentage of successful requests",
"type": "ratio",
"good_events": "sum(rate(http_requests_total{service=\"payment-service\",code!~\"5..\"}))",
"total_events": "sum(rate(http_requests_total{service=\"payment-service\"}))",
"unit": "percentage"
},
{
"name": "Request Latency P95",
"description": "95th percentile of request latency",
"type": "threshold",
"query": "histogram_quantile(0.95, rate(http_request_duration_seconds_bucket{service=\"payment-service\"}[5m]))",
"unit": "seconds"
},
{
"name": "Error Rate",
"description": "Rate of 5xx errors",
"type": "ratio",
"good_events": "sum(rate(http_requests_total{service=\"payment-service\",code!~\"5..\"}))",
"total_events": "sum(rate(http_requests_total{service=\"payment-service\"}))",
"unit": "percentage"
},
{
"name": "Request Throughput",
"description": "Requests per second",
"type": "gauge",
"query": "sum(rate(http_requests_total{service=\"payment-service\"}[5m]))",
"unit": "requests/sec"
},
{
"name": "User Journey Success Rate",
"description": "Percentage of successful complete user journeys",
"type": "ratio",
"good_events": "sum(rate(user_journey_total{service=\"payment-service\",status=\"success\"}[5m]))",
"total_events": "sum(rate(user_journey_total{service=\"payment-service\"}[5m]))",
"unit": "percentage"
},
{
"name": "Feature Availability",
"description": "Percentage of time key features are available",
"type": "ratio",
"good_events": "sum(rate(feature_checks_total{service=\"payment-service\",status=\"available\"}[5m]))",
"total_events": "sum(rate(feature_checks_total{service=\"payment-service\"}[5m]))",
"unit": "percentage"
}
],
"slos": [
{
"name": "Availability SLO",
"description": "Service level objective for percentage of successful requests",
"sli_name": "Availability",
"target_value": 0.9999,
"target_display": "99.99%",
"operator": ">=",
"time_windows": [
"1h",
"1d",
"7d",
"30d"
],
"measurement_window": "30d",
"service": "payment-service",
"criticality": "critical"
},
{
"name": "Request Latency P95 SLO",
"description": "Service level objective for 95th percentile of request latency",
"sli_name": "Request Latency P95",
"target_value": 100,
"target_display": "0.1s",
"operator": "<=",
"time_windows": [
"1h",
"1d",
"7d",
"30d"
],
"measurement_window": "30d",
"service": "payment-service",
"criticality": "critical"
},
{
"name": "Error Rate SLO",
"description": "Service level objective for rate of 5xx errors",
"sli_name": "Error Rate",
"target_value": 0.001,
"target_display": "0.1%",
"operator": "<=",
"time_windows": [
"1h",
"1d",
"7d",
"30d"
],
"measurement_window": "30d",
"service": "payment-service",
"criticality": "critical"
},
{
"name": "User Journey Success Rate SLO",
"description": "Service level objective for percentage of successful complete user journeys",
"sli_name": "User Journey Success Rate",
"target_value": 0.9999,
"target_display": "99.99%",
"operator": ">=",
"time_windows": [
"1h",
"1d",
"7d",
"30d"
],
"measurement_window": "30d",
"service": "payment-service",
"criticality": "critical"
},
{
"name": "Feature Availability SLO",
"description": "Service level objective for percentage of time key features are available",
"sli_name": "Feature Availability",
"target_value": 0.9999,
"target_display": "99.99%",
"operator": ">=",
"time_windows": [
"1h",
"1d",
"7d",
"30d"
],
"measurement_window": "30d",
"service": "payment-service",
"criticality": "critical"
}
],
"error_budgets": [
{
"slo_name": "Availability SLO",
"error_budget_rate": 9.999999999998899e-05,
"error_budget_percentage": "0.010%",
"budgets_by_window": {
"1h": "0.4 seconds",
"1d": "8.6 seconds",
"7d": "1.0 minutes",
"30d": "4.3 minutes"
},
"burn_rate_alerts": [
{
"name": "Availability Burn Rate 2% Alert",
"description": "Alert when Availability is consuming error budget at 14.4x rate",
"severity": "critical",
"short_window": "5m",
"long_window": "1h",
"burn_rate_threshold": 14.4,
"budget_consumed": "2%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 14.4) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 14.4)",
"annotations": {
"summary": "High burn rate detected for Availability",
"description": "Error budget consumption rate is 14.4x normal, will exhaust 2% of monthly budget"
}
},
{
"name": "Availability Burn Rate 5% Alert",
"description": "Alert when Availability is consuming error budget at 6x rate",
"severity": "warning",
"short_window": "30m",
"long_window": "6h",
"burn_rate_threshold": 6,
"budget_consumed": "5%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 6) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 6)",
"annotations": {
"summary": "High burn rate detected for Availability",
"description": "Error budget consumption rate is 6x normal, will exhaust 5% of monthly budget"
}
},
{
"name": "Availability Burn Rate 10% Alert",
"description": "Alert when Availability is consuming error budget at 3x rate",
"severity": "info",
"short_window": "2h",
"long_window": "1d",
"burn_rate_threshold": 3,
"budget_consumed": "10%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 3) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 3)",
"annotations": {
"summary": "High burn rate detected for Availability",
"description": "Error budget consumption rate is 3x normal, will exhaust 10% of monthly budget"
}
},
{
"name": "Availability Burn Rate 10% Alert",
"description": "Alert when Availability is consuming error budget at 1x rate",
"severity": "info",
"short_window": "6h",
"long_window": "3d",
"burn_rate_threshold": 1,
"budget_consumed": "10%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 1) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 1)",
"annotations": {
"summary": "High burn rate detected for Availability",
"description": "Error budget consumption rate is 1x normal, will exhaust 10% of monthly budget"
}
}
]
},
{
"slo_name": "User Journey Success Rate SLO",
"error_budget_rate": 9.999999999998899e-05,
"error_budget_percentage": "0.010%",
"budgets_by_window": {
"1h": "0.4 seconds",
"1d": "8.6 seconds",
"7d": "1.0 minutes",
"30d": "4.3 minutes"
},
"burn_rate_alerts": [
{
"name": "User Journey Success Rate Burn Rate 2% Alert",
"description": "Alert when User Journey Success Rate is consuming error budget at 14.4x rate",
"severity": "critical",
"short_window": "5m",
"long_window": "1h",
"burn_rate_threshold": 14.4,
"budget_consumed": "2%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 14.4) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 14.4)",
"annotations": {
"summary": "High burn rate detected for User Journey Success Rate",
"description": "Error budget consumption rate is 14.4x normal, will exhaust 2% of monthly budget"
}
},
{
"name": "User Journey Success Rate Burn Rate 5% Alert",
"description": "Alert when User Journey Success Rate is consuming error budget at 6x rate",
"severity": "warning",
"short_window": "30m",
"long_window": "6h",
"burn_rate_threshold": 6,
"budget_consumed": "5%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 6) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 6)",
"annotations": {
"summary": "High burn rate detected for User Journey Success Rate",
"description": "Error budget consumption rate is 6x normal, will exhaust 5% of monthly budget"
}
},
{
"name": "User Journey Success Rate Burn Rate 10% Alert",
"description": "Alert when User Journey Success Rate is consuming error budget at 3x rate",
"severity": "info",
"short_window": "2h",
"long_window": "1d",
"burn_rate_threshold": 3,
"budget_consumed": "10%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 3) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 3)",
"annotations": {
"summary": "High burn rate detected for User Journey Success Rate",
"description": "Error budget consumption rate is 3x normal, will exhaust 10% of monthly budget"
}
},
{
"name": "User Journey Success Rate Burn Rate 10% Alert",
"description": "Alert when User Journey Success Rate is consuming error budget at 1x rate",
"severity": "info",
"short_window": "6h",
"long_window": "3d",
"burn_rate_threshold": 1,
"budget_consumed": "10%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 1) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 1)",
"annotations": {
"summary": "High burn rate detected for User Journey Success Rate",
"description": "Error budget consumption rate is 1x normal, will exhaust 10% of monthly budget"
}
}
]
},
{
"slo_name": "Feature Availability SLO",
"error_budget_rate": 9.999999999998899e-05,
"error_budget_percentage": "0.010%",
"budgets_by_window": {
"1h": "0.4 seconds",
"1d": "8.6 seconds",
"7d": "1.0 minutes",
"30d": "4.3 minutes"
},
"burn_rate_alerts": [
{
"name": "Feature Availability Burn Rate 2% Alert",
"description": "Alert when Feature Availability is consuming error budget at 14.4x rate",
"severity": "critical",
"short_window": "5m",
"long_window": "1h",
"burn_rate_threshold": 14.4,
"budget_consumed": "2%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 14.4) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 14.4)",
"annotations": {
"summary": "High burn rate detected for Feature Availability",
"description": "Error budget consumption rate is 14.4x normal, will exhaust 2% of monthly budget"
}
},
{
"name": "Feature Availability Burn Rate 5% Alert",
"description": "Alert when Feature Availability is consuming error budget at 6x rate",
"severity": "warning",
"short_window": "30m",
"long_window": "6h",
"burn_rate_threshold": 6,
"budget_consumed": "5%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 6) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 6)",
"annotations": {
"summary": "High burn rate detected for Feature Availability",
"description": "Error budget consumption rate is 6x normal, will exhaust 5% of monthly budget"
}
},
{
"name": "Feature Availability Burn Rate 10% Alert",
"description": "Alert when Feature Availability is consuming error budget at 3x rate",
"severity": "info",
"short_window": "2h",
"long_window": "1d",
"burn_rate_threshold": 3,
"budget_consumed": "10%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 3) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 3)",
"annotations": {
"summary": "High burn rate detected for Feature Availability",
"description": "Error budget consumption rate is 3x normal, will exhaust 10% of monthly budget"
}
},
{
"name": "Feature Availability Burn Rate 10% Alert",
"description": "Alert when Feature Availability is consuming error budget at 1x rate",
"severity": "info",
"short_window": "6h",
"long_window": "3d",
"burn_rate_threshold": 1,
"budget_consumed": "10%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 1) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 1)",
"annotations": {
"summary": "High burn rate detected for Feature Availability",
"description": "Error budget consumption rate is 1x normal, will exhaust 10% of monthly budget"
}
}
]
}
],
"sla_recommendations": {
"applicable": true,
"service": "payment-service",
"commitments": [
{
"metric": "Availability",
"target": 0.9989,
"target_display": "99.89%",
"measurement_window": "monthly",
"measurement_method": "Uptime monitoring with 1-minute granularity"
},
{
"metric": "Feature Availability",
"target": 0.9989,
"target_display": "99.89%",
"measurement_window": "monthly",
"measurement_method": "Uptime monitoring with 1-minute granularity"
}
],
"penalties": [
{
"breach_threshold": "< 99.99%",
"credit_percentage": 10
},
{
"breach_threshold": "< 99.9%",
"credit_percentage": 25
},
{
"breach_threshold": "< 99%",
"credit_percentage": 50
}
],
"measurement_methodology": "External synthetic monitoring from multiple geographic locations",
"exclusions": [
"Planned maintenance windows (with 72h advance notice)",
"Customer-side network or infrastructure issues",
"Force majeure events",
"Third-party service dependencies beyond our control"
]
},
"monitoring_recommendations": {
"metrics": {
"collection": "Prometheus with service discovery",
"retention": "90 days for raw metrics, 1 year for aggregated",
"alerting": "Prometheus Alertmanager with multi-window burn rate alerts"
},
"logging": {
"format": "Structured JSON logs with correlation IDs",
"aggregation": "ELK stack or equivalent with proper indexing",
"retention": "30 days for debug logs, 90 days for error logs"
},
"tracing": {
"sampling": "Adaptive sampling with 1% base rate",
"storage": "Jaeger or Zipkin with 7-day retention",
"integration": "OpenTelemetry instrumentation"
}
},
"implementation_guide": {
"prerequisites": [
"Service instrumented with metrics collection (Prometheus format)",
"Structured logging with correlation IDs",
"Monitoring infrastructure (Prometheus, Grafana, Alertmanager)",
"Incident response processes and escalation policies"
],
"implementation_steps": [
{
"step": 1,
"title": "Instrument Service",
"description": "Add metrics collection for all defined SLIs",
"estimated_effort": "1-2 days"
},
{
"step": 2,
"title": "Configure Recording Rules",
"description": "Set up Prometheus recording rules for SLI calculations",
"estimated_effort": "4-8 hours"
},
{
"step": 3,
"title": "Implement Burn Rate Alerts",
"description": "Configure multi-window burn rate alerting rules",
"estimated_effort": "1 day"
},
{
"step": 4,
"title": "Create SLO Dashboard",
"description": "Build Grafana dashboard for SLO tracking and error budget monitoring",
"estimated_effort": "4-6 hours"
},
{
"step": 5,
"title": "Test and Validate",
"description": "Test alerting and validate SLI measurements against expectations",
"estimated_effort": "1-2 days"
},
{
"step": 6,
"title": "Documentation and Training",
"description": "Document runbooks and train team on SLO monitoring",
"estimated_effort": "1 day"
}
],
"validation_checklist": [
"All SLIs produce expected metric values",
"Burn rate alerts fire correctly during simulated outages",
"Error budget calculations match manual verification",
"Dashboard displays accurate SLO achievement rates",
"Alert routing reaches correct escalation paths",
"Runbooks are complete and tested"
]
}
}
FILE:README.md
# Observability Designer
A comprehensive toolkit for designing production-ready observability strategies including SLI/SLO frameworks, alert optimization, and dashboard generation.
## Overview
The Observability Designer skill provides three powerful Python scripts that help you create, optimize, and maintain observability systems:
- **SLO Designer**: Generate complete SLI/SLO frameworks with error budgets and burn rate alerts
- **Alert Optimizer**: Analyze and optimize existing alert configurations to reduce noise and improve effectiveness
- **Dashboard Generator**: Create comprehensive dashboard specifications with role-based layouts and drill-down paths
## Quick Start
### Prerequisites
- Python 3.7+
- No external dependencies required (uses Python standard library only)
### Basic Usage
```bash
# Generate SLO framework for a service
python3 scripts/slo_designer.py --service-type api --criticality critical --user-facing true --service-name payment-service
# Optimize existing alerts
python3 scripts/alert_optimizer.py --input assets/sample_alerts.json --analyze-only
# Generate a dashboard specification
python3 scripts/dashboard_generator.py --service-type web --name "Customer Portal" --role sre
```
## Scripts Documentation
### SLO Designer (`slo_designer.py`)
Generates comprehensive SLO frameworks based on service characteristics.
#### Features
- **Automatic SLI Selection**: Recommends appropriate SLIs based on service type
- **Target Setting**: Suggests SLO targets based on service criticality
- **Error Budget Calculation**: Computes error budgets and burn rate thresholds
- **Multi-Window Burn Rate Alerts**: Generates 4-window burn rate alerting rules
- **SLA Recommendations**: Provides customer-facing SLA guidance
#### Usage Examples
```bash
# From service definition file
python3 scripts/slo_designer.py --input assets/sample_service_api.json --output slo_framework.json
# From command line parameters
python3 scripts/slo_designer.py \
--service-type api \
--criticality critical \
--user-facing true \
--service-name payment-service \
--output payment_slos.json
# Generate and display summary only
python3 scripts/slo_designer.py --input assets/sample_service_web.json --summary-only
```
#### Service Definition Format
```json
{
"name": "payment-service",
"type": "api",
"criticality": "critical",
"user_facing": true,
"description": "Handles payment processing",
"team": "payments",
"environment": "production",
"dependencies": [
{
"name": "user-service",
"type": "api",
"criticality": "high"
}
]
}
```
#### Supported Service Types
- **api**: REST APIs, GraphQL services
- **web**: Web applications, SPAs
- **database**: Database services, data stores
- **queue**: Message queues, event streams
- **batch**: Batch processing jobs
- **ml**: Machine learning services
#### Criticality Levels
- **critical**: 99.99% availability, <100ms P95 latency, <0.1% error rate
- **high**: 99.9% availability, <200ms P95 latency, <0.5% error rate
- **medium**: 99.5% availability, <500ms P95 latency, <1% error rate
- **low**: 99% availability, <1s P95 latency, <2% error rate
### Alert Optimizer (`alert_optimizer.py`)
Analyzes existing alert configurations and provides optimization recommendations.
#### Features
- **Noise Detection**: Identifies alerts with high false positive rates
- **Coverage Analysis**: Finds gaps in monitoring coverage
- **Duplicate Detection**: Locates redundant or overlapping alerts
- **Threshold Analysis**: Reviews alert thresholds for appropriateness
- **Fatigue Assessment**: Evaluates alert volume and routing
#### Usage Examples
```bash
# Analyze existing alerts
python3 scripts/alert_optimizer.py --input assets/sample_alerts.json --analyze-only
# Generate optimized configuration
python3 scripts/alert_optimizer.py \
--input assets/sample_alerts.json \
--output optimized_alerts.json
# Generate HTML report
python3 scripts/alert_optimizer.py \
--input assets/sample_alerts.json \
--report alert_analysis.html \
--format html
```
#### Alert Configuration Format
```json
{
"alerts": [
{
"alert": "HighLatency",
"expr": "histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m])) > 0.5",
"for": "5m",
"labels": {
"severity": "warning",
"service": "payment-service"
},
"annotations": {
"summary": "High request latency detected",
"runbook_url": "https://runbooks.company.com/high-latency"
},
"historical_data": {
"fires_per_day": 2.5,
"false_positive_rate": 0.15
}
}
],
"services": [
{
"name": "payment-service",
"criticality": "critical"
}
]
}
```
#### Analysis Categories
- **Golden Signals**: Latency, traffic, errors, saturation
- **Resource Utilization**: CPU, memory, disk, network
- **Business Metrics**: Revenue, conversion, user engagement
- **Security**: Auth failures, suspicious activity
- **Availability**: Uptime, health checks
### Dashboard Generator (`dashboard_generator.py`)
Creates comprehensive dashboard specifications with role-based optimization.
#### Features
- **Role-Based Layouts**: Optimized for SRE, Developer, Executive, and Ops personas
- **Golden Signals Coverage**: Automatic inclusion of key monitoring metrics
- **Service-Type Specific Panels**: Tailored panels based on service characteristics
- **Interactive Elements**: Template variables, drill-down paths, time range controls
- **Grafana Compatibility**: Generates Grafana-compatible JSON
#### Usage Examples
```bash
# From service definition
python3 scripts/dashboard_generator.py \
--input assets/sample_service_web.json \
--output dashboard.json
# With specific role optimization
python3 scripts/dashboard_generator.py \
--service-type api \
--name "Payment Service" \
--role developer \
--output payment_dev_dashboard.json
# Generate Grafana-compatible JSON
python3 scripts/dashboard_generator.py \
--input assets/sample_service_api.json \
--output dashboard.json \
--format grafana
# With documentation
python3 scripts/dashboard_generator.py \
--service-type web \
--name "Customer Portal" \
--output portal_dashboard.json \
--doc-output portal_docs.md
```
#### Target Roles
- **sre**: Focus on availability, latency, errors, resource utilization
- **developer**: Emphasize latency, errors, throughput, business metrics
- **executive**: Highlight availability, business metrics, user experience
- **ops**: Priority on resource utilization, capacity, alerts, deployments
#### Panel Types
- **Stat**: Single value displays with thresholds
- **Gauge**: Resource utilization and capacity metrics
- **Timeseries**: Trend analysis and historical data
- **Table**: Top N lists and detailed breakdowns
- **Heatmap**: Distribution and correlation analysis
## Sample Data
The `assets/` directory contains sample configurations for testing:
- `sample_service_api.json`: Critical API service definition
- `sample_service_web.json`: High-priority web application definition
- `sample_alerts.json`: Alert configuration with optimization opportunities
The `expected_outputs/` directory shows example outputs from each script:
- `sample_slo_framework.json`: Complete SLO framework for API service
- `optimized_alerts.json`: Optimized alert configuration
- `sample_dashboard.json`: SRE dashboard specification
## Best Practices
### SLO Design
- Start with 1-2 SLOs per service and iterate
- Choose SLIs that directly impact user experience
- Set targets based on user needs, not technical capabilities
- Use error budgets to balance reliability and velocity
### Alert Optimization
- Every alert must be actionable
- Alert on symptoms, not causes
- Use multi-window burn rate alerts for SLO protection
- Implement proper escalation and routing policies
### Dashboard Design
- Follow the F-pattern for visual hierarchy
- Use consistent color semantics across dashboards
- Include drill-down paths for effective troubleshooting
- Optimize for the target role's specific needs
## Integration Patterns
### CI/CD Integration
```bash
# Generate SLOs during service onboarding
python3 scripts/slo_designer.py --input service-config.json --output slos.json
# Validate alert configurations in pipeline
python3 scripts/alert_optimizer.py --input alerts.json --analyze-only --report validation.html
# Auto-generate dashboards for new services
python3 scripts/dashboard_generator.py --input service-config.json --format grafana --output dashboard.json
```
### Monitoring Stack Integration
- **Prometheus**: Generated alert rules and recording rules
- **Grafana**: Dashboard JSON for direct import
- **Alertmanager**: Routing and escalation policies
- **PagerDuty**: Escalation configuration
### GitOps Workflow
1. Store service definitions in version control
2. Generate observability configurations in CI/CD
3. Deploy configurations via GitOps
4. Monitor effectiveness and iterate
## Advanced Usage
### Custom SLO Targets
Override default targets by including them in service definitions:
```json
{
"name": "special-service",
"type": "api",
"criticality": "high",
"custom_slos": {
"availability_target": 0.9995,
"latency_p95_target_ms": 150,
"error_rate_target": 0.002
}
}
```
### Alert Rule Templates
Use template variables for reusable alert rules:
```yaml
# Generated Prometheus alert rule
- alert: {{ service_name }}_HighLatency
expr: histogram_quantile(0.95, rate(http_request_duration_seconds_bucket{service="{{ service_name }}"}[5m])) > {{ latency_threshold }}
for: 5m
labels:
severity: warning
service: "{{ service_name }}"
```
### Dashboard Variants
Generate multiple dashboard variants for different use cases:
```bash
# SRE operational dashboard
python3 scripts/dashboard_generator.py --input service.json --role sre --output sre-dashboard.json
# Developer debugging dashboard
python3 scripts/dashboard_generator.py --input service.json --role developer --output dev-dashboard.json
# Executive business dashboard
python3 scripts/dashboard_generator.py --input service.json --role executive --output exec-dashboard.json
```
## Troubleshooting
### Common Issues
#### Script Execution Errors
- Ensure Python 3.7+ is installed
- Check file paths and permissions
- Validate JSON syntax in input files
#### Invalid Service Definitions
- Required fields: `name`, `type`, `criticality`
- Valid service types: `api`, `web`, `database`, `queue`, `batch`, `ml`
- Valid criticality levels: `critical`, `high`, `medium`, `low`
#### Missing Historical Data
- Alert historical data is optional but improves analysis
- Include `fires_per_day` and `false_positive_rate` when available
- Use monitoring system APIs to populate historical metrics
### Debug Mode
Enable verbose logging by setting environment variable:
```bash
export DEBUG=1
python3 scripts/slo_designer.py --input service.json
```
## Contributing
### Development Setup
```bash
# Clone the repository
git clone <repository-url>
cd engineering/observability-designer
# Run tests
python3 -m pytest tests/
# Lint code
python3 -m flake8 scripts/
```
### Adding New Features
1. Follow existing code patterns and error handling
2. Include comprehensive docstrings and type hints
3. Add test cases for new functionality
4. Update documentation and examples
## Support
For questions, issues, or feature requests:
- Check existing documentation and examples
- Review the reference materials in `references/`
- Open an issue with detailed reproduction steps
- Include sample configurations when reporting bugs
---
*This skill is part of the Claude Skills marketplace. For more information about observability best practices, see the reference documentation in the `references/` directory.*
FILE:references/alert_design_patterns.md
# Alert Design Patterns: A Guide to Effective Alerting
## Introduction
Well-designed alerts are the difference between a reliable system and 3 AM pages about non-issues. This guide provides patterns and anti-patterns for creating alerts that provide value without causing fatigue.
## Fundamental Principles
### The Golden Rules of Alerting
1. **Every alert should be actionable** - If you can't do something about it, don't alert
2. **Every alert should require human intelligence** - If a script can handle it, automate the response
3. **Every alert should be novel** - Don't alert on known, ongoing issues
4. **Every alert should represent a user-visible impact** - Internal metrics matter only if users are affected
### Alert Classification
#### Critical Alerts
- Service is completely down
- Data loss is occurring
- Security breach detected
- SLO burn rate indicates imminent SLO violation
#### Warning Alerts
- Service degradation affecting some users
- Approaching resource limits
- Dependent service issues
- Elevated error rates within SLO
#### Info Alerts
- Deployment notifications
- Capacity planning triggers
- Configuration changes
- Maintenance windows
## Alert Design Patterns
### Pattern 1: Symptoms, Not Causes
**Good**: Alert on user-visible symptoms
```yaml
- alert: HighLatency
expr: histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m])) > 0.5
for: 5m
annotations:
summary: "API latency is high"
description: "95th percentile latency is {{ $value }}s, above 500ms threshold"
```
**Bad**: Alert on internal metrics that may not affect users
```yaml
- alert: HighCPU
expr: cpu_usage > 80
# This might not affect users at all!
```
### Pattern 2: Multi-Window Alerting
Reduce false positives by requiring sustained problems:
```yaml
- alert: ServiceDown
expr: (
avg_over_time(up[2m]) == 0 # Short window: immediate detection
and
avg_over_time(up[10m]) < 0.8 # Long window: avoid flapping
)
for: 1m
```
### Pattern 3: Burn Rate Alerting
Alert based on error budget consumption rate:
```yaml
# Fast burn: 2% of monthly budget in 1 hour
- alert: ErrorBudgetFastBurn
expr: (
error_rate_5m > (14.4 * error_budget_slo)
and
error_rate_1h > (14.4 * error_budget_slo)
)
for: 2m
labels:
severity: critical
# Slow burn: 10% of monthly budget in 3 days
- alert: ErrorBudgetSlowBurn
expr: (
error_rate_6h > (1.0 * error_budget_slo)
and
error_rate_3d > (1.0 * error_budget_slo)
)
for: 15m
labels:
severity: warning
```
### Pattern 4: Hysteresis
Use different thresholds for firing and resolving to prevent flapping:
```yaml
- alert: HighErrorRate
expr: error_rate > 0.05 # Fire at 5%
for: 5m
# Resolution happens automatically when error_rate < 0.03 (3%)
# This prevents flapping around the 5% threshold
```
### Pattern 5: Composite Alerts
Alert when multiple conditions indicate a problem:
```yaml
- alert: ServiceDegraded
expr: (
(latency_p95 > latency_threshold)
or
(error_rate > error_threshold)
or
(availability < availability_threshold)
) and (
request_rate > min_request_rate # Only alert if we have traffic
)
```
### Pattern 6: Contextual Alerting
Include relevant context in alerts:
```yaml
- alert: DatabaseConnections
expr: db_connections_active / db_connections_max > 0.8
for: 5m
annotations:
summary: "Database connection pool nearly exhausted"
description: "{{ $labels.database }} has {{ $value | humanizePercentage }} connection utilization"
runbook_url: "https://runbooks.company.com/database-connections"
impact: "New requests may be rejected, causing 500 errors"
suggested_action: "Check for connection leaks or increase pool size"
```
## Alert Routing and Escalation
### Routing by Impact and Urgency
#### Critical Path Services
```yaml
route:
group_by: ['service']
routes:
- match:
service: 'payment-api'
severity: 'critical'
receiver: 'payment-team-pager'
continue: true
- match:
service: 'payment-api'
severity: 'warning'
receiver: 'payment-team-slack'
```
#### Time-Based Routing
```yaml
route:
routes:
- match:
severity: 'critical'
receiver: 'oncall-pager'
- match:
severity: 'warning'
time: 'business_hours' # 9 AM - 5 PM
receiver: 'team-slack'
- match:
severity: 'warning'
time: 'after_hours'
receiver: 'team-email' # Lower urgency outside business hours
```
### Escalation Patterns
#### Linear Escalation
```yaml
receivers:
- name: 'primary-oncall'
pagerduty_configs:
- escalation_policy: 'P1-Escalation'
# 0 min: Primary on-call
# 5 min: Secondary on-call
# 15 min: Engineering manager
# 30 min: Director of engineering
```
#### Severity-Based Escalation
```yaml
# Critical: Immediate escalation
- match:
severity: 'critical'
receiver: 'critical-escalation'
# Warning: Team-first escalation
- match:
severity: 'warning'
receiver: 'team-escalation'
```
## Alert Fatigue Prevention
### Grouping and Suppression
#### Time-Based Grouping
```yaml
route:
group_wait: 30s # Wait 30s to group similar alerts
group_interval: 2m # Send grouped alerts every 2 minutes
repeat_interval: 1h # Re-send unresolved alerts every hour
```
#### Dependent Service Suppression
```yaml
- alert: ServiceDown
expr: up == 0
- alert: HighLatency
expr: latency_p95 > 1
# This alert is suppressed when ServiceDown is firing
inhibit_rules:
- source_match:
alertname: 'ServiceDown'
target_match:
alertname: 'HighLatency'
equal: ['service']
```
### Alert Throttling
```yaml
# Limit to 1 alert per 10 minutes for noisy conditions
- alert: HighMemoryUsage
expr: memory_usage_percent > 85
for: 10m # Longer 'for' duration reduces noise
annotations:
summary: "Memory usage has been high for 10+ minutes"
```
### Smart Defaults
```yaml
# Use business logic to set intelligent thresholds
- alert: LowTraffic
expr: request_rate < (
avg_over_time(request_rate[7d]) * 0.1 # 10% of weekly average
)
# Only alert during business hours when low traffic is unusual
for: 30m
```
## Runbook Integration
### Runbook Structure Template
```markdown
# Alert: {{ $labels.alertname }}
## Immediate Actions
1. Check service status dashboard
2. Verify if users are affected
3. Look at recent deployments/changes
## Investigation Steps
1. Check logs for errors in the last 30 minutes
2. Verify dependent services are healthy
3. Check resource utilization (CPU, memory, disk)
4. Review recent alerts for patterns
## Resolution Actions
- If deployment-related: Consider rollback
- If resource-related: Scale up or optimize queries
- If dependency-related: Engage appropriate team
## Escalation
- Primary: @team-oncall
- Secondary: @engineering-manager
- Emergency: @site-reliability-team
```
### Runbook Integration in Alerts
```yaml
annotations:
runbook_url: "https://runbooks.company.com/alerts/{{ $labels.alertname }}"
quick_debug: |
1. curl -s https://{{ $labels.instance }}/health
2. kubectl logs {{ $labels.pod }} --tail=50
3. Check dashboard: https://grafana.company.com/d/service-{{ $labels.service }}
```
## Testing and Validation
### Alert Testing Strategies
#### Chaos Engineering Integration
```python
# Test that alerts fire during controlled failures
def test_alert_during_cpu_spike():
with chaos.cpu_spike(target='payment-api', duration='2m'):
assert wait_for_alert('HighCPU', timeout=180)
def test_alert_during_network_partition():
with chaos.network_partition(target='database'):
assert wait_for_alert('DatabaseUnreachable', timeout=60)
```
#### Historical Alert Analysis
```prometheus
# Query to find alerts that fired without incidents
count by (alertname) (
ALERTS{alertstate="firing"}[30d]
) unless on (alertname) (
count by (alertname) (
incident_created{source="alert"}[30d]
)
)
```
### Alert Quality Metrics
#### Alert Precision
```
Precision = True Positives / (True Positives + False Positives)
```
Track alerts that resulted in actual incidents vs false alarms.
#### Time to Resolution
```prometheus
# Average time from alert firing to resolution
avg_over_time(
(alert_resolved_timestamp - alert_fired_timestamp)[30d]
) by (alertname)
```
#### Alert Fatigue Indicators
```prometheus
# Alerts per day by team
sum by (team) (
increase(alerts_fired_total[1d])
)
# Percentage of alerts acknowledged within 15 minutes
sum(alerts_acked_within_15m) / sum(alerts_fired) * 100
```
## Advanced Patterns
### Machine Learning-Enhanced Alerting
#### Anomaly Detection
```yaml
- alert: AnomalousTraffic
expr: |
abs(request_rate - predict_linear(request_rate[1h], 300)) /
stddev_over_time(request_rate[1h]) > 3
for: 10m
annotations:
summary: "Traffic pattern is anomalous"
description: "Current traffic deviates from predicted pattern by >3 standard deviations"
```
#### Dynamic Thresholds
```yaml
- alert: DynamicHighLatency
expr: |
latency_p95 > (
quantile_over_time(0.95, latency_p95[7d]) + # Historical 95th percentile
2 * stddev_over_time(latency_p95[7d]) # Plus 2 standard deviations
)
```
### Business Hours Awareness
```yaml
# Different thresholds for business vs off hours
- alert: HighLatencyBusinessHours
expr: latency_p95 > 0.2 # Stricter during business hours
for: 2m
# Active 9 AM - 5 PM weekdays
- alert: HighLatencyOffHours
expr: latency_p95 > 0.5 # More lenient after hours
for: 5m
# Active nights and weekends
```
### Progressive Alerting
```yaml
# Escalating alert severity based on duration
- alert: ServiceLatencyElevated
expr: latency_p95 > 0.5
for: 5m
labels:
severity: info
- alert: ServiceLatencyHigh
expr: latency_p95 > 0.5
for: 15m # Same condition, longer duration
labels:
severity: warning
- alert: ServiceLatencyCritical
expr: latency_p95 > 0.5
for: 30m # Same condition, even longer duration
labels:
severity: critical
```
## Anti-Patterns to Avoid
### Anti-Pattern 1: Alerting on Everything
**Problem**: Too many alerts create noise and fatigue
**Solution**: Be selective; only alert on user-impacting issues
### Anti-Pattern 2: Vague Alert Messages
**Problem**: "Service X is down" - which instance? what's the impact?
**Solution**: Include specific details and context
### Anti-Pattern 3: Alerts Without Runbooks
**Problem**: Alerts that don't explain what to do
**Solution**: Every alert must have an associated runbook
### Anti-Pattern 4: Static Thresholds
**Problem**: 80% CPU might be normal during peak hours
**Solution**: Use contextual, adaptive thresholds
### Anti-Pattern 5: Ignoring Alert Quality
**Problem**: Accepting high false positive rates
**Solution**: Regularly review and tune alert precision
## Implementation Checklist
### Pre-Implementation
- [ ] Define alert severity levels and escalation policies
- [ ] Create runbook templates
- [ ] Set up alert routing configuration
- [ ] Define SLOs that alerts will protect
### Alert Development
- [ ] Each alert has clear success criteria
- [ ] Alert conditions tested against historical data
- [ ] Runbook created and accessible
- [ ] Severity and routing configured
- [ ] Context and suggested actions included
### Post-Implementation
- [ ] Monitor alert precision and recall
- [ ] Regular review of alert fatigue metrics
- [ ] Quarterly alert effectiveness review
- [ ] Team training on alert response procedures
### Quality Assurance
- [ ] Test alerts fire during controlled failures
- [ ] Verify alerts resolve when conditions improve
- [ ] Confirm runbooks are accurate and helpful
- [ ] Validate escalation paths work correctly
Remember: Great alerts are invisible when things work and invaluable when things break. Focus on quality over quantity, and always optimize for the human who will respond to the alert at 3 AM.
FILE:references/dashboard_best_practices.md
# Dashboard Best Practices: Design for Insight and Action
## Introduction
A well-designed dashboard is like a good story - it guides you through the data with purpose and clarity. This guide provides practical patterns for creating dashboards that inform decisions and enable quick troubleshooting.
## Design Principles
### The Hierarchy of Information
#### Primary Information (Top Third)
- Service health status
- SLO achievement
- Critical alerts
- Business KPIs
#### Secondary Information (Middle Third)
- Golden signals (latency, traffic, errors, saturation)
- Resource utilization
- Throughput and performance metrics
#### Tertiary Information (Bottom Third)
- Detailed breakdowns
- Historical trends
- Dependency status
- Debug information
### Visual Design Principles
#### Rule of 7±2
- Maximum 7±2 panels per screen
- Group related information together
- Use sections to organize complexity
#### Color Psychology
- **Red**: Critical issues, danger, immediate attention needed
- **Yellow/Orange**: Warnings, caution, degraded state
- **Green**: Healthy, normal operation, success
- **Blue**: Information, neutral metrics, capacity
- **Gray**: Disabled, unknown, or baseline states
#### Chart Selection Guide
- **Line charts**: Time series, trends, comparisons over time
- **Bar charts**: Categorical comparisons, top N lists
- **Gauges**: Single value with defined good/bad ranges
- **Stat panels**: Key metrics, percentages, counts
- **Heatmaps**: Distribution data, correlation analysis
- **Tables**: Detailed breakdowns, multi-dimensional data
## Dashboard Archetypes
### The Overview Dashboard
**Purpose**: High-level health check and business metrics
**Audience**: Executives, managers, cross-team stakeholders
**Update Frequency**: 5-15 minutes
```yaml
sections:
- title: "Business Health"
panels:
- service_availability_summary
- revenue_per_hour
- active_users
- conversion_rate
- title: "System Health"
panels:
- critical_alerts_count
- slo_achievement_summary
- error_budget_remaining
- deployment_status
```
### The SRE Operational Dashboard
**Purpose**: Real-time monitoring and incident response
**Audience**: SRE, on-call engineers
**Update Frequency**: 15-30 seconds
```yaml
sections:
- title: "Service Status"
panels:
- service_up_status
- active_incidents
- recent_deployments
- title: "Golden Signals"
panels:
- latency_percentiles
- request_rate
- error_rate
- resource_saturation
- title: "Infrastructure"
panels:
- cpu_memory_utilization
- network_io
- disk_space
```
### The Developer Debug Dashboard
**Purpose**: Deep-dive troubleshooting and performance analysis
**Audience**: Development teams
**Update Frequency**: 30 seconds - 2 minutes
```yaml
sections:
- title: "Application Performance"
panels:
- endpoint_latency_breakdown
- database_query_performance
- cache_hit_rates
- queue_depths
- title: "Errors and Logs"
panels:
- error_rate_by_endpoint
- log_volume_by_level
- exception_types
- slow_queries
```
## Layout Patterns
### The F-Pattern Layout
Based on eye-tracking studies, users scan in an F-pattern:
```
[Critical Status] [SLO Summary ] [Error Budget ]
[Latency ] [Traffic ] [Errors ]
[Saturation ] [Resource Use ] [Detailed View]
[Historical ] [Dependencies ] [Debug Info ]
```
### The Z-Pattern Layout
For executive dashboards, follow the Z-pattern:
```
[Business KPIs ] → [System Status]
↓ ↓
[Trend Analysis ] ← [Key Metrics ]
```
### Responsive Design
#### Desktop (1920x1080)
- 24-column grid
- Panels can be 6, 8, 12, or 24 units wide
- 4-6 rows visible without scrolling
#### Laptop (1366x768)
- Stack wider panels vertically
- Reduce panel heights
- Prioritize most critical information
#### Mobile (768px width)
- Single column layout
- Simplified panels
- Touch-friendly controls
## Effective Panel Design
### Stat Panels
```yaml
# Good: Clear value with context
- title: "API Availability"
type: stat
targets:
- expr: avg(up{service="api"}) * 100
field_config:
unit: percent
thresholds:
steps:
- color: red
value: 0
- color: yellow
value: 99
- color: green
value: 99.9
options:
color_mode: background
text_mode: value_and_name
```
### Time Series Panels
```yaml
# Good: Multiple related metrics with clear legend
- title: "Request Latency"
type: timeseries
targets:
- expr: histogram_quantile(0.50, rate(http_duration_bucket[5m]))
legend: "P50"
- expr: histogram_quantile(0.95, rate(http_duration_bucket[5m]))
legend: "P95"
- expr: histogram_quantile(0.99, rate(http_duration_bucket[5m]))
legend: "P99"
field_config:
unit: ms
custom:
draw_style: line
fill_opacity: 10
options:
legend:
display_mode: table
placement: bottom
values: [min, max, mean, last]
```
### Table Panels
```yaml
# Good: Top N with relevant columns
- title: "Slowest Endpoints"
type: table
targets:
- expr: topk(10, histogram_quantile(0.95, sum by (handler)(rate(http_duration_bucket[5m]))))
format: table
instant: true
transformations:
- id: organize
options:
exclude_by_name:
Time: true
rename_by_name:
Value: "P95 Latency (ms)"
handler: "Endpoint"
```
## Color and Visualization Best Practices
### Threshold Configuration
```yaml
# Traffic light system with meaningful boundaries
thresholds:
steps:
- color: green # Good performance
value: null # Default
- color: yellow # Degraded performance
value: 95 # 95th percentile of historical normal
- color: orange # Poor performance
value: 99 # 99th percentile of historical normal
- color: red # Critical performance
value: 99.9 # Worst case scenario
```
### Color Blind Friendly Palettes
```yaml
# Use patterns and shapes in addition to color
field_config:
overrides:
- matcher:
id: byName
options: "Critical"
properties:
- id: color
value:
mode: fixed
fixed_color: "#d73027" # Red-orange for protanopia
- id: custom.draw_style
value: "points" # Different shape
```
### Consistent Color Semantics
- **Success/Health**: Green (#28a745)
- **Warning/Degraded**: Yellow (#ffc107)
- **Error/Critical**: Red (#dc3545)
- **Information**: Blue (#007bff)
- **Neutral**: Gray (#6c757d)
## Time Range Strategy
### Default Time Ranges by Dashboard Type
#### Real-time Operational
- **Default**: Last 15 minutes
- **Quick options**: 5m, 15m, 1h, 4h
- **Auto-refresh**: 15-30 seconds
#### Troubleshooting
- **Default**: Last 1 hour
- **Quick options**: 15m, 1h, 4h, 12h, 1d
- **Auto-refresh**: 1 minute
#### Business Review
- **Default**: Last 24 hours
- **Quick options**: 1d, 7d, 30d, 90d
- **Auto-refresh**: 5 minutes
#### Capacity Planning
- **Default**: Last 7 days
- **Quick options**: 7d, 30d, 90d, 1y
- **Auto-refresh**: 15 minutes
### Time Range Annotations
```yaml
# Add context for time-based events
annotations:
- name: "Deployments"
datasource: "Prometheus"
expr: "deployment_timestamp"
title_format: "Deploy {{ version }}"
text_format: "Deployed version {{ version }} to {{ environment }}"
- name: "Incidents"
datasource: "Incident API"
query: "incidents.json?service={{ service }}"
color: "red"
```
## Interactive Features
### Template Variables
```yaml
# Service selector
- name: service
type: query
query: label_values(up, service)
current:
text: All
value: $__all
include_all: true
multi: true
# Environment selector
- name: environment
type: query
query: label_values(up{service="$service"}, environment)
current:
text: production
value: production
```
### Drill-Down Links
```yaml
# Panel-level drill-downs
- title: "Error Rate"
type: timeseries
# ... other config ...
options:
data_links:
- title: "View Error Logs"
url: "/d/logs-dashboard?var-service=__field.labels.service&from=__from&to=__to"
- title: "Error Traces"
url: "/d/traces-dashboard?var-service=__field.labels.service"
```
### Dynamic Panel Titles
```yaml
- title: "service - Request Rate" # Uses template variable
type: timeseries
# Title updates automatically when service variable changes
```
## Performance Optimization
### Query Optimization
#### Use Recording Rules
```yaml
# Instead of complex queries in dashboards
groups:
- name: http_requests
rules:
- record: http_request_rate_5m
expr: sum(rate(http_requests_total[5m])) by (service, method, handler)
- record: http_request_latency_p95_5m
expr: histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (service, le))
```
#### Limit Data Points
```yaml
# Good: Reasonable resolution for dashboard
- expr: http_request_rate_5m[1h]
interval: 15s # One point every 15 seconds
# Bad: Too many points for visualization
- expr: http_request_rate_1s[1h] # 3600 points!
```
### Dashboard Performance
#### Panel Limits
- **Maximum panels per dashboard**: 20-30
- **Maximum queries per panel**: 10
- **Maximum time series per panel**: 50
#### Caching Strategy
```yaml
# Use appropriate cache headers
cache_timeout: 30 # Cache for 30 seconds on fast-changing panels
cache_timeout: 300 # Cache for 5 minutes on slow-changing panels
```
## Accessibility
### Screen Reader Support
```yaml
# Provide text alternatives for visual elements
- title: "Service Health Status"
type: stat
options:
text_mode: value_and_name # Includes both value and description
field_config:
mappings:
- options:
"1":
text: "Healthy"
color: "green"
"0":
text: "Unhealthy"
color: "red"
```
### Keyboard Navigation
- Ensure all interactive elements are keyboard accessible
- Provide logical tab order
- Include skip links for complex dashboards
### High Contrast Mode
```yaml
# Test dashboards work in high contrast mode
theme: high_contrast
colors:
- "#000000" # Pure black
- "#ffffff" # Pure white
- "#ffff00" # Pure yellow
- "#ff0000" # Pure red
```
## Testing and Validation
### Dashboard Testing Checklist
#### Functional Testing
- [ ] All panels load without errors
- [ ] Template variables filter correctly
- [ ] Time range changes update all panels
- [ ] Drill-down links work as expected
- [ ] Auto-refresh functions properly
#### Visual Testing
- [ ] Dashboard renders correctly on different screen sizes
- [ ] Colors are distinguishable and meaningful
- [ ] Text is readable at normal zoom levels
- [ ] Legends and labels are clear
#### Performance Testing
- [ ] Dashboard loads in < 5 seconds
- [ ] No queries timeout under normal load
- [ ] Auto-refresh doesn't cause browser lag
- [ ] Memory usage remains reasonable
#### Usability Testing
- [ ] New team members can understand the dashboard
- [ ] Action items are clear during incidents
- [ ] Key information is quickly discoverable
- [ ] Dashboard supports common troubleshooting workflows
## Maintenance and Governance
### Dashboard Lifecycle
#### Creation
1. Define dashboard purpose and audience
2. Identify key metrics and success criteria
3. Design layout following established patterns
4. Implement with consistent styling
5. Test with real data and user scenarios
#### Maintenance
- **Weekly**: Check for broken panels or queries
- **Monthly**: Review dashboard usage analytics
- **Quarterly**: Gather user feedback and iterate
- **Annually**: Major review and potential redesign
#### Retirement
- Archive dashboards that are no longer used
- Migrate users to replacement dashboards
- Document lessons learned
### Dashboard Standards
```yaml
# Organization dashboard standards
standards:
naming_convention: "[Team] [Service] - [Purpose]"
tags: [team, service_type, environment, purpose]
refresh_intervals: [15s, 30s, 1m, 5m, 15m]
time_ranges: [5m, 15m, 1h, 4h, 1d, 7d, 30d]
color_scheme: "company_standard"
max_panels_per_dashboard: 25
```
## Advanced Patterns
### Composite Dashboards
```yaml
# Dashboard that includes panels from other dashboards
- title: "Service Overview"
type: dashlist
targets:
- "service-health"
- "service-performance"
- "service-business-metrics"
options:
show_headings: true
max_items: 10
```
### Dynamic Dashboard Generation
```python
# Generate dashboards from service definitions
def generate_service_dashboard(service_config):
panels = []
# Always include golden signals
panels.extend(generate_golden_signals_panels(service_config))
# Add service-specific panels
if service_config.type == 'database':
panels.extend(generate_database_panels(service_config))
elif service_config.type == 'queue':
panels.extend(generate_queue_panels(service_config))
return {
'title': f"{service_config.name} - Operational Dashboard",
'panels': panels,
'variables': generate_variables(service_config)
}
```
### A/B Testing for Dashboards
```yaml
# Test different dashboard designs with different teams
experiment:
name: "dashboard_layout_test"
variants:
- name: "traditional_layout"
weight: 50
config: "dashboard_v1.json"
- name: "f_pattern_layout"
weight: 50
config: "dashboard_v2.json"
success_metrics:
- "time_to_insight"
- "user_satisfaction"
- "troubleshooting_efficiency"
```
Remember: A dashboard should tell a story about your system's health and guide users toward the right actions. Focus on clarity over complexity, and always optimize for the person who will use it during a stressful incident.
FILE:references/slo_cookbook.md
# SLO Cookbook: A Practical Guide to Service Level Objectives
## Introduction
Service Level Objectives (SLOs) are a key tool for managing service reliability. This cookbook provides practical guidance for implementing SLOs that actually improve system reliability rather than just creating meaningless metrics.
## Fundamentals
### The SLI/SLO/SLA Hierarchy
- **SLI (Service Level Indicator)**: A quantifiable measure of service quality
- **SLO (Service Level Objective)**: A target range of values for an SLI
- **SLA (Service Level Agreement)**: A business agreement with consequences for missing SLO targets
### Golden Rule of SLOs
**Start simple, iterate based on learning.** Your first SLOs won't be perfect, and that's okay.
## Choosing Good SLIs
### The Four Golden Signals
1. **Latency**: How long requests take to complete
2. **Traffic**: How many requests are coming in
3. **Errors**: How many requests are failing
4. **Saturation**: How "full" your service is
### SLI Selection Criteria
A good SLI should be:
- **Measurable**: You can collect data for it
- **Meaningful**: It reflects user experience
- **Controllable**: You can take action to improve it
- **Proportional**: Changes in the SLI reflect changes in user happiness
### Service Type Specific SLIs
#### HTTP APIs
- **Request latency**: P95 or P99 response time
- **Availability**: Proportion of successful requests (non-5xx)
- **Throughput**: Requests per second capacity
```prometheus
# Availability SLI
sum(rate(http_requests_total{code!~"5.."}[5m])) / sum(rate(http_requests_total[5m]))
# Latency SLI
histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m]))
```
#### Batch Jobs
- **Freshness**: Age of the last successful run
- **Correctness**: Proportion of jobs completing successfully
- **Throughput**: Items processed per unit time
#### Data Pipelines
- **Data freshness**: Time since last successful update
- **Data quality**: Proportion of records passing validation
- **Processing latency**: Time from ingestion to availability
### Anti-Patterns in SLI Selection
❌ **Don't use**: CPU usage, memory usage, disk space as primary SLIs
- These are symptoms, not user-facing impacts
❌ **Don't use**: Counts instead of rates or proportions
- "Number of errors" vs "Error rate"
❌ **Don't use**: Internal metrics that users don't care about
- Queue depth, cache hit rate (unless they directly impact user experience)
## Setting SLO Targets
### The Art of Target Setting
Setting SLO targets is balancing act between:
- **User happiness**: Targets should reflect acceptable user experience
- **Business value**: Tighter SLOs cost more to maintain
- **Current performance**: Targets should be achievable but aspirational
### Target Setting Strategies
#### Historical Performance Method
1. Collect 4-6 weeks of historical data
2. Calculate the worst user-visible performance in that period
3. Set your SLO slightly better than the worst acceptable performance
#### User Journey Mapping
1. Map critical user journeys
2. Identify acceptable performance for each step
3. Work backwards to component SLOs
#### Error Budget Approach
1. Decide how much unreliability you can afford
2. Set SLO targets based on acceptable error budget consumption
3. Example: 99.9% availability = 43.8 minutes downtime per month
### SLO Target Examples by Service Criticality
#### Critical Services (Revenue Impact)
- **Availability**: 99.95% - 99.99%
- **Latency (P95)**: 100-200ms
- **Error Rate**: < 0.1%
#### High Priority Services
- **Availability**: 99.9% - 99.95%
- **Latency (P95)**: 200-500ms
- **Error Rate**: < 0.5%
#### Standard Services
- **Availability**: 99.5% - 99.9%
- **Latency (P95)**: 500ms - 1s
- **Error Rate**: < 1%
## Error Budget Management
### What is an Error Budget?
Your error budget is the maximum amount of unreliability you can accumulate while still meeting your SLO. It's calculated as:
```
Error Budget = (1 - SLO) × Time Window
```
For a 99.9% availability SLO over 30 days:
```
Error Budget = (1 - 0.999) × 30 days = 0.001 × 30 days = 43.8 minutes
```
### Error Budget Policies
Define what happens when you consume your error budget:
#### Conservative Policy (High-Risk Services)
- **> 50% consumed**: Freeze non-critical feature releases
- **> 75% consumed**: Focus entirely on reliability improvements
- **> 90% consumed**: Consider emergency measures (traffic shaping, etc.)
#### Balanced Policy (Standard Services)
- **> 75% consumed**: Increase focus on reliability work
- **> 90% consumed**: Pause feature work, focus on reliability
#### Aggressive Policy (Early Stage Services)
- **> 90% consumed**: Review but continue normal operations
- **100% consumed**: Evaluate SLO appropriateness
### Burn Rate Alerting
Multi-window burn rate alerts help you catch SLO violations before they become critical:
```yaml
# Fast burn: 2% budget consumed in 1 hour
- alert: FastBurnSLOViolation
expr: (
(1 - (sum(rate(http_requests_total{code!~"5.."}[5m])) / sum(rate(http_requests_total[5m])))) > (14.4 * 0.001)
and
(1 - (sum(rate(http_requests_total{code!~"5.."}[1h])) / sum(rate(http_requests_total[1h])))) > (14.4 * 0.001)
)
for: 2m
# Slow burn: 10% budget consumed in 3 days
- alert: SlowBurnSLOViolation
expr: (
(1 - (sum(rate(http_requests_total{code!~"5.."}[6h])) / sum(rate(http_requests_total[6h])))) > (1.0 * 0.001)
and
(1 - (sum(rate(http_requests_total{code!~"5.."}[3d])) / sum(rate(http_requests_total[3d])))) > (1.0 * 0.001)
)
for: 15m
```
## Implementation Patterns
### The SLO Implementation Ladder
#### Level 1: Basic SLOs
- Choose 1-2 SLIs that matter most to users
- Set aspirational but achievable targets
- Implement basic alerting when SLOs are missed
#### Level 2: Operational SLOs
- Add burn rate alerting
- Create error budget dashboards
- Establish error budget policies
- Regular SLO review meetings
#### Level 3: Advanced SLOs
- Multi-window burn rate alerts
- Automated error budget policy enforcement
- SLO-driven incident prioritization
- Integration with CI/CD for deployment decisions
### SLO Measurement Architecture
#### Push vs Pull Metrics
- **Pull** (Prometheus): Good for infrastructure metrics, real-time alerting
- **Push** (StatsD): Good for application metrics, business events
#### Measurement Points
- **Server-side**: More reliable, easier to implement
- **Client-side**: Better reflects user experience
- **Synthetic**: Consistent, predictable, may not reflect real user experience
### SLO Dashboard Design
Essential elements for SLO dashboards:
1. **Current SLO Achievement**: Large, prominent display
2. **Error Budget Remaining**: Visual indicator (gauge, progress bar)
3. **Burn Rate**: Time series showing error budget consumption rate
4. **Historical Trends**: 4-week view of SLO achievement
5. **Alerts**: Current and recent SLO-related alerts
## Advanced Topics
### Dependency SLOs
For services with dependencies:
```
SLO_service ≤ min(SLO_inherent, ∏SLO_dependencies)
```
If your service depends on 3 other services each with 99.9% SLO:
```
Maximum_SLO = 0.999³ = 0.997 = 99.7%
```
### User Journey SLOs
Track end-to-end user experiences:
```prometheus
# Registration success rate
sum(rate(user_registration_success_total[5m])) / sum(rate(user_registration_attempts_total[5m]))
# Purchase completion latency
histogram_quantile(0.95, rate(purchase_completion_duration_seconds_bucket[5m]))
```
### SLOs for Batch Systems
Special considerations for non-request/response systems:
#### Freshness SLO
```prometheus
# Data should be no more than 4 hours old
(time() - last_successful_update_timestamp) < (4 * 3600)
```
#### Throughput SLO
```prometheus
# Should process at least 1000 items per hour
rate(items_processed_total[1h]) >= 1000
```
#### Quality SLO
```prometheus
# At least 99.5% of records should pass validation
sum(rate(records_valid_total[5m])) / sum(rate(records_processed_total[5m])) >= 0.995
```
## Common Mistakes and How to Avoid Them
### Mistake 1: Too Many SLOs
**Problem**: Drowning in metrics, losing focus
**Solution**: Start with 1-2 SLOs per service, add more only when needed
### Mistake 2: Internal Metrics as SLIs
**Problem**: Optimizing for metrics that don't impact users
**Solution**: Always ask "If this metric changes, do users notice?"
### Mistake 3: Perfectionist SLOs
**Problem**: 99.99% SLO when 99.9% would be fine
**Solution**: Higher SLOs cost exponentially more; pick the minimum acceptable level
### Mistake 4: Ignoring Error Budgets
**Problem**: Treating any SLO miss as an emergency
**Solution**: Error budgets exist to be spent; use them to balance feature velocity and reliability
### Mistake 5: Static SLOs
**Problem**: Setting SLOs once and never updating them
**Solution**: Review SLOs quarterly; adjust based on user feedback and business changes
## SLO Review Process
### Monthly SLO Review Agenda
1. **SLO Achievement Review**: Did we meet our SLOs?
2. **Error Budget Analysis**: How did we spend our error budget?
3. **Incident Correlation**: Which incidents impacted our SLOs?
4. **SLI Quality Assessment**: Are our SLIs still meaningful?
5. **Target Adjustment**: Should we change any targets?
### Quarterly SLO Health Check
1. **User Impact Validation**: Survey users about acceptable performance
2. **Business Alignment**: Do SLOs still reflect business priorities?
3. **Measurement Quality**: Are we measuring the right things?
4. **Cost/Benefit Analysis**: Are tighter SLOs worth the investment?
## Tooling and Automation
### Essential Tools
1. **Metrics Collection**: Prometheus, InfluxDB, CloudWatch
2. **Alerting**: Alertmanager, PagerDuty, OpsGenie
3. **Dashboards**: Grafana, DataDog, New Relic
4. **SLO Platforms**: Sloth, Pyrra, Service Level Blue
### Automation Opportunities
- **Burn rate alert generation** from SLO definitions
- **Dashboard creation** from SLO specifications
- **Error budget calculation** and tracking
- **Release blocking** based on error budget consumption
## Getting Started Checklist
- [ ] Identify your service's critical user journeys
- [ ] Choose 1-2 SLIs that best reflect user experience
- [ ] Collect 4-6 weeks of baseline data
- [ ] Set initial SLO targets based on historical performance
- [ ] Implement basic SLO monitoring and alerting
- [ ] Create an SLO dashboard
- [ ] Define error budget policies
- [ ] Schedule monthly SLO reviews
- [ ] Plan for quarterly SLO health checks
Remember: SLOs are a journey, not a destination. Start simple, learn from experience, and iterate toward better reliability management.
FILE:scripts/alert_optimizer.py
#!/usr/bin/env python3
"""
Alert Optimizer - Analyze and optimize alert configurations
This script analyzes existing alert configurations and identifies optimization opportunities:
- Noisy alerts with high false positive rates
- Missing coverage gaps in monitoring
- Duplicate or redundant alerts
- Poor threshold settings and alert fatigue risks
- Missing runbooks and documentation
- Routing and escalation policy improvements
Usage:
python alert_optimizer.py --input alert_config.json --output optimized_config.json
python alert_optimizer.py --input alerts.json --analyze-only --report report.html
"""
import json
import argparse
import sys
import re
import math
from typing import Dict, List, Any, Tuple, Set
from datetime import datetime, timedelta
from collections import defaultdict, Counter
class AlertOptimizer:
"""Analyze and optimize alert configurations."""
# Alert severity priority mapping
SEVERITY_PRIORITY = {
'critical': 1,
'high': 2,
'warning': 3,
'info': 4
}
# Common noisy alert patterns
NOISY_PATTERNS = [
r'disk.*usage.*>.*[89]\d%', # Disk usage > 80% often noisy
r'memory.*>.*[89]\d%', # Memory > 80% often noisy
r'cpu.*>.*[789]\d%', # CPU > 70% can be noisy
r'response.*time.*>.*\d+ms', # Low latency thresholds
r'error.*rate.*>.*0\.[01]%' # Very low error rate thresholds
]
# Essential monitoring categories
COVERAGE_CATEGORIES = [
'availability',
'latency',
'error_rate',
'resource_utilization',
'security',
'business_metrics'
]
# Golden signals that should always be monitored
GOLDEN_SIGNALS = [
'latency',
'traffic',
'errors',
'saturation'
]
def __init__(self):
"""Initialize the Alert Optimizer."""
self.alert_config = {}
self.optimization_results = {}
self.alert_analysis = {}
def load_alert_config(self, file_path: str) -> Dict[str, Any]:
"""Load alert configuration from JSON file."""
try:
with open(file_path, 'r') as f:
return json.load(f)
except FileNotFoundError:
raise ValueError(f"Alert configuration file not found: {file_path}")
except json.JSONDecodeError as e:
raise ValueError(f"Invalid JSON in alert configuration: {e}")
def analyze_alert_noise(self, alerts: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Identify potentially noisy alerts."""
noisy_alerts = []
for alert in alerts:
noise_score = 0
noise_reasons = []
alert_rule = alert.get('expr', alert.get('condition', ''))
alert_name = alert.get('alert', alert.get('name', 'Unknown'))
# Check for common noisy patterns
for pattern in self.NOISY_PATTERNS:
if re.search(pattern, alert_rule, re.IGNORECASE):
noise_score += 3
noise_reasons.append(f"Matches noisy pattern: {pattern}")
# Check for very frequent evaluation intervals
evaluation_interval = alert.get('for', '0s')
if self._parse_duration(evaluation_interval) < 60: # Less than 1 minute
noise_score += 2
noise_reasons.append("Very short evaluation interval")
# Check for lack of 'for' clause
if not alert.get('for') or alert.get('for') == '0s':
noise_score += 2
noise_reasons.append("No 'for' clause - may cause alert flapping")
# Check for overly sensitive thresholds
if self._has_sensitive_threshold(alert_rule):
noise_score += 2
noise_reasons.append("Potentially sensitive threshold")
# Check historical firing rate if available
historical_data = alert.get('historical_data', {})
if historical_data:
firing_rate = historical_data.get('fires_per_day', 0)
if firing_rate > 10: # More than 10 fires per day
noise_score += 3
noise_reasons.append(f"High firing rate: {firing_rate} times/day")
false_positive_rate = historical_data.get('false_positive_rate', 0)
if false_positive_rate > 0.3: # > 30% false positives
noise_score += 4
noise_reasons.append(f"High false positive rate: {false_positive_rate*100:.1f}%")
if noise_score >= 3: # Threshold for considering an alert noisy
noisy_alert = {
'alert_name': alert_name,
'noise_score': noise_score,
'reasons': noise_reasons,
'current_rule': alert_rule,
'recommendations': self._generate_noise_reduction_recommendations(alert, noise_reasons)
}
noisy_alerts.append(noisy_alert)
return sorted(noisy_alerts, key=lambda x: x['noise_score'], reverse=True)
def _parse_duration(self, duration_str: str) -> int:
"""Parse duration string to seconds."""
if not duration_str or duration_str == '0s':
return 0
duration_map = {'s': 1, 'm': 60, 'h': 3600, 'd': 86400}
match = re.match(r'(\d+)([smhd])', duration_str)
if match:
value, unit = match.groups()
return int(value) * duration_map.get(unit, 1)
return 0
def _has_sensitive_threshold(self, rule: str) -> bool:
"""Check if alert rule has potentially sensitive thresholds."""
# Look for very low error rates or very tight latency thresholds
sensitive_patterns = [
r'error.*rate.*>.*0\.0[01]', # Error rate > 0.01% or 0.001%
r'latency.*>.*[12]\d\d?ms', # Latency > 100-299ms
r'response.*time.*>.*0\.[12]', # Response time > 0.1-0.2s
r'cpu.*>.*[456]\d%' # CPU > 40-69% (too sensitive for most cases)
]
for pattern in sensitive_patterns:
if re.search(pattern, rule, re.IGNORECASE):
return True
return False
def _generate_noise_reduction_recommendations(self, alert: Dict[str, Any],
reasons: List[str]) -> List[str]:
"""Generate recommendations to reduce alert noise."""
recommendations = []
if "No 'for' clause" in str(reasons):
recommendations.append("Add 'for: 5m' clause to prevent flapping")
if "Very short evaluation interval" in str(reasons):
recommendations.append("Increase evaluation interval to at least 1 minute")
if "sensitive threshold" in str(reasons):
recommendations.append("Review and increase threshold based on historical data")
if "High firing rate" in str(reasons):
recommendations.append("Analyze historical firing patterns and adjust thresholds")
if "High false positive rate" in str(reasons):
recommendations.append("Implement more specific conditions to reduce false positives")
if "noisy pattern" in str(reasons):
recommendations.append("Consider using percentile-based thresholds instead of absolute values")
return recommendations
def identify_coverage_gaps(self, alerts: List[Dict[str, Any]],
services: List[Dict[str, Any]] = None) -> Dict[str, Any]:
"""Identify gaps in monitoring coverage."""
coverage_analysis = {
'missing_categories': [],
'missing_golden_signals': [],
'service_coverage_gaps': [],
'critical_gaps': [],
'recommendations': []
}
# Analyze coverage by category
covered_categories = set()
alert_categories = []
for alert in alerts:
alert_rule = alert.get('expr', alert.get('condition', ''))
alert_name = alert.get('alert', alert.get('name', ''))
category = self._classify_alert_category(alert_rule, alert_name)
if category:
covered_categories.add(category)
alert_categories.append(category)
# Check for missing essential categories
missing_categories = set(self.COVERAGE_CATEGORIES) - covered_categories
coverage_analysis['missing_categories'] = list(missing_categories)
# Check for missing golden signals
covered_signals = set()
for alert in alerts:
alert_rule = alert.get('expr', alert.get('condition', ''))
signal = self._identify_golden_signal(alert_rule)
if signal:
covered_signals.add(signal)
missing_signals = set(self.GOLDEN_SIGNALS) - covered_signals
coverage_analysis['missing_golden_signals'] = list(missing_signals)
# Analyze service-specific coverage if service list provided
if services:
service_coverage = self._analyze_service_coverage(alerts, services)
coverage_analysis['service_coverage_gaps'] = service_coverage
# Identify critical gaps
critical_gaps = []
if 'availability' in missing_categories:
critical_gaps.append("Missing availability monitoring")
if 'error_rate' in missing_categories:
critical_gaps.append("Missing error rate monitoring")
if 'errors' in missing_signals:
critical_gaps.append("Missing error signal monitoring")
coverage_analysis['critical_gaps'] = critical_gaps
# Generate recommendations
recommendations = self._generate_coverage_recommendations(coverage_analysis)
coverage_analysis['recommendations'] = recommendations
return coverage_analysis
def _classify_alert_category(self, rule: str, alert_name: str) -> str:
"""Classify alert into monitoring category."""
rule_lower = rule.lower()
name_lower = alert_name.lower()
if any(keyword in rule_lower or keyword in name_lower
for keyword in ['up', 'down', 'available', 'reachable']):
return 'availability'
if any(keyword in rule_lower or keyword in name_lower
for keyword in ['latency', 'response_time', 'duration']):
return 'latency'
if any(keyword in rule_lower or keyword in name_lower
for keyword in ['error', 'fail', '5xx', '4xx']):
return 'error_rate'
if any(keyword in rule_lower or keyword in name_lower
for keyword in ['cpu', 'memory', 'disk', 'network', 'utilization']):
return 'resource_utilization'
if any(keyword in rule_lower or keyword in name_lower
for keyword in ['security', 'auth', 'login', 'breach']):
return 'security'
if any(keyword in rule_lower or keyword in name_lower
for keyword in ['revenue', 'conversion', 'user', 'business']):
return 'business_metrics'
return 'other'
def _identify_golden_signal(self, rule: str) -> str:
"""Identify which golden signal an alert covers."""
rule_lower = rule.lower()
if any(keyword in rule_lower for keyword in ['latency', 'response_time', 'duration']):
return 'latency'
if any(keyword in rule_lower for keyword in ['rate', 'rps', 'qps', 'throughput']):
return 'traffic'
if any(keyword in rule_lower for keyword in ['error', 'fail', '5xx']):
return 'errors'
if any(keyword in rule_lower for keyword in ['cpu', 'memory', 'disk', 'utilization']):
return 'saturation'
return None
def _analyze_service_coverage(self, alerts: List[Dict[str, Any]],
services: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Analyze monitoring coverage per service."""
service_coverage = []
for service in services:
service_name = service.get('name', '')
service_alerts = [alert for alert in alerts
if service_name in alert.get('expr', '') or
service_name in alert.get('labels', {}).get('service', '')]
covered_signals = set()
for alert in service_alerts:
signal = self._identify_golden_signal(alert.get('expr', ''))
if signal:
covered_signals.add(signal)
missing_signals = set(self.GOLDEN_SIGNALS) - covered_signals
if missing_signals or len(service_alerts) < 3: # Less than 3 alerts per service
coverage_gap = {
'service': service_name,
'alert_count': len(service_alerts),
'covered_signals': list(covered_signals),
'missing_signals': list(missing_signals),
'criticality': service.get('criticality', 'medium'),
'recommendations': []
}
if len(service_alerts) == 0:
coverage_gap['recommendations'].append("Add basic availability monitoring")
if 'errors' in missing_signals:
coverage_gap['recommendations'].append("Add error rate monitoring")
if 'latency' in missing_signals:
coverage_gap['recommendations'].append("Add latency monitoring")
service_coverage.append(coverage_gap)
return service_coverage
def _generate_coverage_recommendations(self, coverage_analysis: Dict[str, Any]) -> List[str]:
"""Generate recommendations to improve monitoring coverage."""
recommendations = []
for missing_category in coverage_analysis['missing_categories']:
if missing_category == 'availability':
recommendations.append("Add service availability/uptime monitoring")
elif missing_category == 'latency':
recommendations.append("Add response time and latency monitoring")
elif missing_category == 'error_rate':
recommendations.append("Add error rate and HTTP status code monitoring")
elif missing_category == 'resource_utilization':
recommendations.append("Add CPU, memory, and disk utilization monitoring")
elif missing_category == 'security':
recommendations.append("Add security monitoring (auth failures, suspicious activity)")
elif missing_category == 'business_metrics':
recommendations.append("Add business KPI monitoring")
for missing_signal in coverage_analysis['missing_golden_signals']:
recommendations.append(f"Implement {missing_signal} monitoring (Golden Signal)")
if coverage_analysis['critical_gaps']:
recommendations.append("Address critical monitoring gaps as highest priority")
return recommendations
def find_duplicate_alerts(self, alerts: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Identify duplicate or redundant alerts."""
duplicates = []
alert_signatures = defaultdict(list)
# Group alerts by signature
for i, alert in enumerate(alerts):
signature = self._generate_alert_signature(alert)
alert_signatures[signature].append((i, alert))
# Find exact duplicates
for signature, alert_group in alert_signatures.items():
if len(alert_group) > 1:
duplicate_group = {
'type': 'exact_duplicate',
'signature': signature,
'alerts': [{'index': i, 'name': alert.get('alert', alert.get('name', f'Alert_{i}'))}
for i, alert in alert_group],
'recommendation': 'Remove duplicate alerts, keep the most comprehensive one'
}
duplicates.append(duplicate_group)
# Find semantic duplicates (similar but not identical)
semantic_duplicates = self._find_semantic_duplicates(alerts)
duplicates.extend(semantic_duplicates)
return duplicates
def _generate_alert_signature(self, alert: Dict[str, Any]) -> str:
"""Generate a signature for alert comparison."""
expr = alert.get('expr', alert.get('condition', ''))
labels = alert.get('labels', {})
# Normalize the expression by removing whitespace and standardizing
normalized_expr = re.sub(r'\s+', ' ', expr).strip()
# Create signature from expression and key labels
key_labels = {k: v for k, v in labels.items()
if k in ['service', 'severity', 'team']}
return f"{normalized_expr}::{json.dumps(key_labels, sort_keys=True)}"
def _find_semantic_duplicates(self, alerts: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Find semantically similar alerts."""
semantic_duplicates = []
# Group alerts by service and metric type
service_groups = defaultdict(list)
for i, alert in enumerate(alerts):
service = self._extract_service_from_alert(alert)
metric_type = self._extract_metric_type_from_alert(alert)
key = f"{service}::{metric_type}"
service_groups[key].append((i, alert))
# Look for similar alerts within each group
for key, alert_group in service_groups.items():
if len(alert_group) > 1:
similar_alerts = self._identify_similar_alerts(alert_group)
if similar_alerts:
semantic_duplicates.extend(similar_alerts)
return semantic_duplicates
def _extract_service_from_alert(self, alert: Dict[str, Any]) -> str:
"""Extract service name from alert."""
labels = alert.get('labels', {})
if 'service' in labels:
return labels['service']
expr = alert.get('expr', alert.get('condition', ''))
# Try to extract service from metric labels
service_match = re.search(r'service="([^"]+)"', expr)
if service_match:
return service_match.group(1)
return 'unknown'
def _extract_metric_type_from_alert(self, alert: Dict[str, Any]) -> str:
"""Extract metric type from alert."""
expr = alert.get('expr', alert.get('condition', ''))
# Common metric patterns
if 'up' in expr.lower():
return 'availability'
elif any(keyword in expr.lower() for keyword in ['latency', 'duration', 'response_time']):
return 'latency'
elif any(keyword in expr.lower() for keyword in ['error', 'fail', '5xx']):
return 'error_rate'
elif any(keyword in expr.lower() for keyword in ['cpu', 'memory', 'disk']):
return 'resource'
return 'other'
def _identify_similar_alerts(self, alert_group: List[Tuple[int, Dict[str, Any]]]) -> List[Dict[str, Any]]:
"""Identify similar alerts within a group."""
similar_groups = []
# Simple similarity check based on threshold values and conditions
threshold_groups = defaultdict(list)
for index, alert in alert_group:
expr = alert.get('expr', alert.get('condition', ''))
threshold = self._extract_threshold_from_expression(expr)
severity = alert.get('labels', {}).get('severity', 'unknown')
similarity_key = f"{threshold}::{severity}"
threshold_groups[similarity_key].append((index, alert))
# If multiple alerts have very similar thresholds, they might be redundant
for similarity_key, similar_alerts in threshold_groups.items():
if len(similar_alerts) > 1:
similar_group = {
'type': 'semantic_duplicate',
'similarity_key': similarity_key,
'alerts': [{'index': i, 'name': alert.get('alert', alert.get('name', f'Alert_{i}'))}
for i, alert in similar_alerts],
'recommendation': 'Review for potential consolidation - similar thresholds and conditions'
}
similar_groups.append(similar_group)
return similar_groups
def _extract_threshold_from_expression(self, expr: str) -> str:
"""Extract threshold value from alert expression."""
# Look for common threshold patterns
threshold_patterns = [
r'>[\s]*([0-9.]+)',
r'<[\s]*([0-9.]+)',
r'>=[\s]*([0-9.]+)',
r'<=[\s]*([0-9.]+)',
r'==[\s]*([0-9.]+)'
]
for pattern in threshold_patterns:
match = re.search(pattern, expr)
if match:
return match.group(1)
return 'unknown'
def analyze_thresholds(self, alerts: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Analyze alert thresholds for optimization opportunities."""
threshold_analysis = []
for alert in alerts:
alert_name = alert.get('alert', alert.get('name', 'Unknown'))
expr = alert.get('expr', alert.get('condition', ''))
analysis = {
'alert_name': alert_name,
'current_expression': expr,
'threshold_issues': [],
'recommendations': []
}
# Check for hard-coded thresholds
if re.search(r'[><=]\s*[0-9.]+', expr):
analysis['threshold_issues'].append('Hard-coded threshold value')
analysis['recommendations'].append('Consider parameterizing thresholds')
# Check for percentage-based thresholds that might be too strict
percentage_match = re.search(r'([><=])\s*0?\.\d+', expr)
if percentage_match:
operator = percentage_match.group(1)
if operator in ['>', '>='] and 'error' in expr.lower():
analysis['threshold_issues'].append('Very low error rate threshold')
analysis['recommendations'].append('Consider increasing error rate threshold based on SLO')
# Check for missing hysteresis
if '>' in expr and 'for:' not in str(alert):
analysis['threshold_issues'].append('No hysteresis (for clause)')
analysis['recommendations'].append('Add "for" clause to prevent alert flapping')
# Check for resource utilization thresholds
if any(resource in expr.lower() for resource in ['cpu', 'memory', 'disk']):
threshold_value = self._extract_threshold_from_expression(expr)
if threshold_value and threshold_value.replace('.', '').isdigit():
threshold_num = float(threshold_value)
if threshold_num < 0.7: # Less than 70%
analysis['threshold_issues'].append('Low resource utilization threshold')
analysis['recommendations'].append('Consider increasing threshold to reduce noise')
# Add historical data analysis if available
historical_data = alert.get('historical_data', {})
if historical_data:
false_positive_rate = historical_data.get('false_positive_rate', 0)
if false_positive_rate > 0.2:
analysis['threshold_issues'].append(f'High false positive rate: {false_positive_rate*100:.1f}%')
analysis['recommendations'].append('Analyze historical data and adjust threshold')
if analysis['threshold_issues']:
threshold_analysis.append(analysis)
return threshold_analysis
def assess_alert_fatigue_risk(self, alerts: List[Dict[str, Any]]) -> Dict[str, Any]:
"""Assess risk of alert fatigue."""
fatigue_assessment = {
'total_alerts': len(alerts),
'risk_level': 'low',
'risk_factors': [],
'metrics': {},
'recommendations': []
}
# Count alerts by severity
severity_counts = Counter()
for alert in alerts:
severity = alert.get('labels', {}).get('severity', 'unknown')
severity_counts[severity] += 1
fatigue_assessment['metrics']['severity_distribution'] = dict(severity_counts)
# Calculate risk factors
critical_count = severity_counts.get('critical', 0)
warning_count = severity_counts.get('warning', 0) + severity_counts.get('high', 0)
total_high_priority = critical_count + warning_count
# Too many high-priority alerts
if total_high_priority > 50:
fatigue_assessment['risk_factors'].append('High number of critical/warning alerts')
fatigue_assessment['recommendations'].append('Review and reduce number of high-priority alerts')
# Poor critical to warning ratio
if critical_count > 0 and warning_count > 0:
critical_ratio = critical_count / (critical_count + warning_count)
if critical_ratio > 0.3: # More than 30% critical
fatigue_assessment['risk_factors'].append('High ratio of critical alerts')
fatigue_assessment['recommendations'].append('Review critical alert criteria - not everything should be critical')
# Estimate daily alert volume
daily_estimate = self._estimate_daily_alert_volume(alerts)
fatigue_assessment['metrics']['estimated_daily_alerts'] = daily_estimate
if daily_estimate > 100:
fatigue_assessment['risk_factors'].append('High estimated daily alert volume')
fatigue_assessment['recommendations'].append('Implement alert grouping and suppression rules')
# Check for missing runbooks
alerts_without_runbooks = [alert for alert in alerts
if not alert.get('annotations', {}).get('runbook_url')]
runbook_ratio = len(alerts_without_runbooks) / len(alerts) if alerts else 0
if runbook_ratio > 0.5:
fatigue_assessment['risk_factors'].append('Many alerts lack runbooks')
fatigue_assessment['recommendations'].append('Create runbooks for alerts to improve response efficiency')
# Determine overall risk level
risk_score = len(fatigue_assessment['risk_factors'])
if risk_score >= 3:
fatigue_assessment['risk_level'] = 'high'
elif risk_score >= 1:
fatigue_assessment['risk_level'] = 'medium'
return fatigue_assessment
def _estimate_daily_alert_volume(self, alerts: List[Dict[str, Any]]) -> int:
"""Estimate daily alert volume."""
total_estimated = 0
for alert in alerts:
# Use historical data if available
historical_data = alert.get('historical_data', {})
if historical_data and 'fires_per_day' in historical_data:
total_estimated += historical_data['fires_per_day']
continue
# Otherwise estimate based on alert characteristics
expr = alert.get('expr', alert.get('condition', ''))
severity = alert.get('labels', {}).get('severity', 'warning')
# Base estimate by severity
base_estimates = {
'critical': 0.1, # Critical should rarely fire
'high': 0.5,
'warning': 2,
'info': 5
}
estimate = base_estimates.get(severity, 1)
# Adjust based on alert type
if 'error_rate' in expr.lower():
estimate *= 1.5 # Error rate alerts tend to be more frequent
elif 'availability' in expr.lower() or 'up' in expr.lower():
estimate *= 0.5 # Availability alerts should be rare
total_estimated += estimate
return int(total_estimated)
def generate_optimized_config(self, alerts: List[Dict[str, Any]],
analysis_results: Dict[str, Any]) -> Dict[str, Any]:
"""Generate optimized alert configuration."""
optimized_alerts = []
for i, alert in enumerate(alerts):
optimized_alert = alert.copy()
alert_name = alert.get('alert', alert.get('name', f'Alert_{i}'))
# Apply noise reduction optimizations
noisy_alerts = analysis_results.get('noisy_alerts', [])
for noisy_alert in noisy_alerts:
if noisy_alert['alert_name'] == alert_name:
optimized_alert = self._apply_noise_reduction(optimized_alert, noisy_alert)
break
# Apply threshold optimizations
threshold_issues = analysis_results.get('threshold_analysis', [])
for threshold_issue in threshold_issues:
if threshold_issue['alert_name'] == alert_name:
optimized_alert = self._apply_threshold_optimization(optimized_alert, threshold_issue)
break
# Ensure proper alert metadata
optimized_alert = self._ensure_alert_metadata(optimized_alert)
optimized_alerts.append(optimized_alert)
# Remove duplicates based on analysis
if 'duplicate_alerts' in analysis_results:
optimized_alerts = self._remove_duplicate_alerts(optimized_alerts,
analysis_results['duplicate_alerts'])
# Add missing alerts for coverage gaps
if 'coverage_gaps' in analysis_results:
new_alerts = self._generate_missing_alerts(analysis_results['coverage_gaps'])
optimized_alerts.extend(new_alerts)
optimized_config = {
'alerts': optimized_alerts,
'optimization_metadata': {
'optimized_at': datetime.utcnow().isoformat() + 'Z',
'original_count': len(alerts),
'optimized_count': len(optimized_alerts),
'changes_applied': analysis_results.get('optimizations_applied', [])
}
}
return optimized_config
def _apply_noise_reduction(self, alert: Dict[str, Any],
noise_analysis: Dict[str, Any]) -> Dict[str, Any]:
"""Apply noise reduction optimizations to an alert."""
optimized_alert = alert.copy()
for recommendation in noise_analysis['recommendations']:
if 'for:' in recommendation and not alert.get('for'):
optimized_alert['for'] = '5m'
elif 'threshold' in recommendation.lower():
# This would require more sophisticated threshold adjustment
# For now, add annotation for manual review
if 'annotations' not in optimized_alert:
optimized_alert['annotations'] = {}
optimized_alert['annotations']['optimization_note'] = 'Review threshold - potentially too sensitive'
return optimized_alert
def _apply_threshold_optimization(self, alert: Dict[str, Any],
threshold_analysis: Dict[str, Any]) -> Dict[str, Any]:
"""Apply threshold optimizations to an alert."""
optimized_alert = alert.copy()
# Add 'for' clause if missing
if 'No hysteresis' in str(threshold_analysis['threshold_issues']):
if not alert.get('for'):
optimized_alert['for'] = '5m'
# Add optimization annotations
if threshold_analysis['recommendations']:
if 'annotations' not in optimized_alert:
optimized_alert['annotations'] = {}
optimized_alert['annotations']['threshold_recommendations'] = '; '.join(threshold_analysis['recommendations'])
return optimized_alert
def _ensure_alert_metadata(self, alert: Dict[str, Any]) -> Dict[str, Any]:
"""Ensure alert has proper metadata."""
optimized_alert = alert.copy()
# Ensure annotations exist
if 'annotations' not in optimized_alert:
optimized_alert['annotations'] = {}
# Add summary if missing
if 'summary' not in optimized_alert['annotations']:
alert_name = alert.get('alert', alert.get('name', 'Alert'))
optimized_alert['annotations']['summary'] = f"Alert: {alert_name}"
# Add description if missing
if 'description' not in optimized_alert['annotations']:
optimized_alert['annotations']['description'] = 'This alert requires a description. Please update with specific details about the condition and impact.'
# Ensure proper labels
if 'labels' not in optimized_alert:
optimized_alert['labels'] = {}
if 'severity' not in optimized_alert['labels']:
optimized_alert['labels']['severity'] = 'warning'
return optimized_alert
def _remove_duplicate_alerts(self, alerts: List[Dict[str, Any]],
duplicates: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Remove duplicate alerts from the list."""
indices_to_remove = set()
for duplicate_group in duplicates:
if duplicate_group['type'] == 'exact_duplicate':
# Keep the first alert, remove the rest
alert_indices = [alert_info['index'] for alert_info in duplicate_group['alerts']]
indices_to_remove.update(alert_indices[1:]) # Remove all but first
return [alert for i, alert in enumerate(alerts) if i not in indices_to_remove]
def _generate_missing_alerts(self, coverage_gaps: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Generate alerts for missing coverage."""
new_alerts = []
for missing_signal in coverage_gaps.get('missing_golden_signals', []):
if missing_signal == 'latency':
new_alert = {
'alert': 'HighLatency',
'expr': 'histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m])) > 0.5',
'for': '5m',
'labels': {
'severity': 'warning'
},
'annotations': {
'summary': 'High request latency detected',
'description': 'The 95th percentile latency is above 500ms for 5 minutes.',
'generated': 'true'
}
}
new_alerts.append(new_alert)
elif missing_signal == 'errors':
new_alert = {
'alert': 'HighErrorRate',
'expr': 'sum(rate(http_requests_total{code=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) > 0.01',
'for': '5m',
'labels': {
'severity': 'warning'
},
'annotations': {
'summary': 'High error rate detected',
'description': 'Error rate is above 1% for 5 minutes.',
'generated': 'true'
}
}
new_alerts.append(new_alert)
return new_alerts
def analyze_configuration(self, alert_config: Dict[str, Any]) -> Dict[str, Any]:
"""Perform comprehensive analysis of alert configuration."""
alerts = alert_config.get('alerts', alert_config.get('rules', []))
services = alert_config.get('services', [])
analysis_results = {
'summary': {
'total_alerts': len(alerts),
'analysis_timestamp': datetime.utcnow().isoformat() + 'Z'
},
'noisy_alerts': self.analyze_alert_noise(alerts),
'coverage_gaps': self.identify_coverage_gaps(alerts, services),
'duplicate_alerts': self.find_duplicate_alerts(alerts),
'threshold_analysis': self.analyze_thresholds(alerts),
'alert_fatigue_assessment': self.assess_alert_fatigue_risk(alerts)
}
# Generate overall recommendations
analysis_results['overall_recommendations'] = self._generate_overall_recommendations(analysis_results)
return analysis_results
def _generate_overall_recommendations(self, analysis_results: Dict[str, Any]) -> List[str]:
"""Generate overall recommendations based on complete analysis."""
recommendations = []
# High-priority recommendations
if analysis_results['alert_fatigue_assessment']['risk_level'] == 'high':
recommendations.append("HIGH PRIORITY: Address alert fatigue risk by reducing alert volume")
if len(analysis_results['coverage_gaps']['critical_gaps']) > 0:
recommendations.append("HIGH PRIORITY: Address critical monitoring gaps")
# Medium-priority recommendations
if len(analysis_results['noisy_alerts']) > 0:
recommendations.append(f"Optimize {len(analysis_results['noisy_alerts'])} noisy alerts to reduce false positives")
if len(analysis_results['duplicate_alerts']) > 0:
recommendations.append(f"Remove or consolidate {len(analysis_results['duplicate_alerts'])} duplicate alert groups")
# General recommendations
recommendations.append("Implement proper alert routing and escalation policies")
recommendations.append("Create runbooks for all production alerts")
recommendations.append("Set up alert effectiveness monitoring and regular reviews")
return recommendations
def export_analysis(self, analysis_results: Dict[str, Any], output_file: str,
format_type: str = 'json'):
"""Export analysis results."""
if format_type.lower() == 'json':
with open(output_file, 'w') as f:
json.dump(analysis_results, f, indent=2)
elif format_type.lower() == 'html':
self._export_html_report(analysis_results, output_file)
else:
raise ValueError(f"Unsupported format: {format_type}")
def _export_html_report(self, analysis_results: Dict[str, Any], output_file: str):
"""Export analysis as HTML report."""
html_content = self._generate_html_report(analysis_results)
with open(output_file, 'w') as f:
f.write(html_content)
def _generate_html_report(self, analysis_results: Dict[str, Any]) -> str:
"""Generate HTML report of analysis results."""
html = f"""
<!DOCTYPE html>
<html>
<head>
<title>Alert Configuration Analysis Report</title>
<style>
body {{ font-family: Arial, sans-serif; margin: 20px; }}
.header {{ background: #f4f4f4; padding: 20px; border-radius: 5px; }}
.section {{ margin: 20px 0; padding: 15px; border: 1px solid #ddd; border-radius: 5px; }}
.critical {{ border-left: 5px solid #ff0000; }}
.warning {{ border-left: 5px solid #ff9900; }}
.info {{ border-left: 5px solid #0066cc; }}
.success {{ border-left: 5px solid #00aa00; }}
ul {{ margin: 10px 0; }}
li {{ margin: 5px 0; }}
</style>
</head>
<body>
<div class="header">
<h1>Alert Configuration Analysis Report</h1>
<p>Generated: {analysis_results['summary']['analysis_timestamp']}</p>
<p>Total Alerts Analyzed: {analysis_results['summary']['total_alerts']}</p>
</div>
<div class="section critical">
<h2>Overall Recommendations</h2>
<ul>
{''.join(f'<li>{rec}</li>' for rec in analysis_results['overall_recommendations'])}
</ul>
</div>
<div class="section warning">
<h2>Alert Fatigue Assessment</h2>
<p><strong>Risk Level:</strong> {analysis_results['alert_fatigue_assessment']['risk_level'].upper()}</p>
<p><strong>Risk Factors:</strong></p>
<ul>
{''.join(f'<li>{factor}</li>' for factor in analysis_results['alert_fatigue_assessment']['risk_factors'])}
</ul>
</div>
<div class="section info">
<h2>Noisy Alerts ({len(analysis_results['noisy_alerts'])})</h2>
{''.join(f'<div><strong>{alert["alert_name"]}</strong> (Score: {alert["noise_score"]})<ul>{"".join(f"<li>{reason}</li>" for reason in alert["reasons"])}</ul></div>'
for alert in analysis_results['noisy_alerts'][:5])}
</div>
<div class="section info">
<h2>Coverage Gaps</h2>
<p><strong>Missing Categories:</strong> {', '.join(analysis_results['coverage_gaps']['missing_categories']) or 'None'}</p>
<p><strong>Missing Golden Signals:</strong> {', '.join(analysis_results['coverage_gaps']['missing_golden_signals']) or 'None'}</p>
<p><strong>Critical Gaps:</strong> {len(analysis_results['coverage_gaps']['critical_gaps'])}</p>
</div>
</body>
</html>
"""
return html
def print_summary(self, analysis_results: Dict[str, Any]):
"""Print human-readable summary of analysis."""
print(f"\n{'='*60}")
print(f"ALERT CONFIGURATION ANALYSIS SUMMARY")
print(f"{'='*60}")
summary = analysis_results['summary']
print(f"\nOverall Statistics:")
print(f" Total Alerts: {summary['total_alerts']}")
print(f" Analysis Date: {summary['analysis_timestamp']}")
# Alert fatigue assessment
fatigue = analysis_results['alert_fatigue_assessment']
print(f"\nAlert Fatigue Risk: {fatigue['risk_level'].upper()}")
if fatigue['risk_factors']:
print(f" Risk Factors:")
for factor in fatigue['risk_factors']:
print(f" • {factor}")
# Noisy alerts
noisy = analysis_results['noisy_alerts']
print(f"\nNoisy Alerts: {len(noisy)}")
if noisy:
print(f" Top 3 Noisiest:")
for alert in noisy[:3]:
print(f" • {alert['alert_name']} (Score: {alert['noise_score']})")
# Coverage gaps
gaps = analysis_results['coverage_gaps']
print(f"\nMonitoring Coverage:")
print(f" Missing Categories: {len(gaps['missing_categories'])}")
print(f" Missing Golden Signals: {len(gaps['missing_golden_signals'])}")
print(f" Critical Gaps: {len(gaps['critical_gaps'])}")
# Duplicates
duplicates = analysis_results['duplicate_alerts']
print(f"\nDuplicate Alerts: {len(duplicates)} groups")
# Overall recommendations
recommendations = analysis_results['overall_recommendations']
print(f"\nTop Recommendations:")
for i, rec in enumerate(recommendations[:5], 1):
print(f" {i}. {rec}")
print(f"\n{'='*60}\n")
def main():
"""Main function for CLI usage."""
parser = argparse.ArgumentParser(
description='Analyze and optimize alert configurations',
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
# Analyze alert configuration
python alert_optimizer.py --input alerts.json --analyze-only
# Generate optimized configuration
python alert_optimizer.py --input alerts.json --output optimized_alerts.json
# Generate HTML report
python alert_optimizer.py --input alerts.json --report report.html --format html
"""
)
parser.add_argument('--input', '-i', required=True,
help='Input alert configuration JSON file')
parser.add_argument('--output', '-o',
help='Output optimized configuration JSON file')
parser.add_argument('--report', '-r',
help='Generate analysis report file')
parser.add_argument('--format', choices=['json', 'html'], default='json',
help='Report format (json or html)')
parser.add_argument('--analyze-only', action='store_true',
help='Only perform analysis, do not generate optimized config')
args = parser.parse_args()
optimizer = AlertOptimizer()
try:
# Load alert configuration
alert_config = optimizer.load_alert_config(args.input)
# Perform analysis
analysis_results = optimizer.analyze_configuration(alert_config)
# Generate optimized configuration if requested
if not args.analyze_only:
optimized_config = optimizer.generate_optimized_config(
alert_config.get('alerts', alert_config.get('rules', [])),
analysis_results
)
output_file = args.output or 'optimized_alerts.json'
optimizer.export_analysis(optimized_config, output_file, 'json')
print(f"Optimized configuration saved to: {output_file}")
# Generate report if requested
if args.report:
optimizer.export_analysis(analysis_results, args.report, args.format)
print(f"Analysis report saved to: {args.report}")
# Always show summary
optimizer.print_summary(analysis_results)
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
sys.exit(1)
if __name__ == '__main__':
main()
FILE:scripts/dashboard_generator.py
#!/usr/bin/env python3
"""
Dashboard Generator - Generate comprehensive dashboard specifications
This script generates dashboard specifications based on service/system descriptions:
- Panel layout optimized for different screen sizes and roles
- Metric queries (Prometheus-style) for comprehensive monitoring
- Visualization types appropriate for different metric types
- Drill-down paths for effective troubleshooting workflows
- Golden signals coverage (latency, traffic, errors, saturation)
- RED/USE method implementation
- Business metrics integration
Usage:
python dashboard_generator.py --input service_definition.json --output dashboard_spec.json
python dashboard_generator.py --service-type api --name "Payment Service" --output payment_dashboard.json
"""
import json
import argparse
import sys
import math
from typing import Dict, List, Any, Tuple
from datetime import datetime, timedelta
class DashboardGenerator:
"""Generate comprehensive dashboard specifications."""
# Dashboard layout templates by role
ROLE_LAYOUTS = {
'sre': {
'primary_focus': ['availability', 'latency', 'errors', 'resource_utilization'],
'secondary_focus': ['throughput', 'capacity', 'dependencies'],
'time_ranges': ['1h', '6h', '1d', '7d'],
'default_refresh': '30s'
},
'developer': {
'primary_focus': ['latency', 'errors', 'throughput', 'business_metrics'],
'secondary_focus': ['resource_utilization', 'dependencies'],
'time_ranges': ['15m', '1h', '6h', '1d'],
'default_refresh': '1m'
},
'executive': {
'primary_focus': ['availability', 'business_metrics', 'user_experience'],
'secondary_focus': ['cost', 'capacity_trends'],
'time_ranges': ['1d', '7d', '30d'],
'default_refresh': '5m'
},
'ops': {
'primary_focus': ['resource_utilization', 'capacity', 'alerts', 'deployments'],
'secondary_focus': ['throughput', 'latency'],
'time_ranges': ['5m', '30m', '2h', '1d'],
'default_refresh': '15s'
}
}
# Service type specific metric configurations
SERVICE_METRICS = {
'api': {
'golden_signals': ['latency', 'traffic', 'errors', 'saturation'],
'key_metrics': [
'http_requests_total',
'http_request_duration_seconds',
'http_request_size_bytes',
'http_response_size_bytes'
],
'resource_metrics': ['cpu_usage', 'memory_usage', 'goroutines']
},
'web': {
'golden_signals': ['latency', 'traffic', 'errors', 'saturation'],
'key_metrics': [
'http_requests_total',
'http_request_duration_seconds',
'page_load_time',
'user_sessions'
],
'resource_metrics': ['cpu_usage', 'memory_usage', 'connections']
},
'database': {
'golden_signals': ['latency', 'traffic', 'errors', 'saturation'],
'key_metrics': [
'db_connections_active',
'db_query_duration_seconds',
'db_queries_total',
'db_slow_queries_total'
],
'resource_metrics': ['cpu_usage', 'memory_usage', 'disk_io', 'connections']
},
'queue': {
'golden_signals': ['latency', 'traffic', 'errors', 'saturation'],
'key_metrics': [
'queue_depth',
'message_processing_duration',
'messages_published_total',
'messages_consumed_total'
],
'resource_metrics': ['cpu_usage', 'memory_usage', 'disk_usage']
}
}
# Visualization type recommendations
VISUALIZATION_TYPES = {
'latency': 'line_chart',
'throughput': 'line_chart',
'error_rate': 'line_chart',
'success_rate': 'stat',
'resource_utilization': 'gauge',
'queue_depth': 'bar_chart',
'status': 'stat',
'distribution': 'heatmap',
'alerts': 'table',
'logs': 'logs_panel'
}
def __init__(self):
"""Initialize the Dashboard Generator."""
self.service_config = {}
self.dashboard_spec = {}
def load_service_definition(self, file_path: str) -> Dict[str, Any]:
"""Load service definition from JSON file."""
try:
with open(file_path, 'r') as f:
return json.load(f)
except FileNotFoundError:
raise ValueError(f"Service definition file not found: {file_path}")
except json.JSONDecodeError as e:
raise ValueError(f"Invalid JSON in service definition: {e}")
def create_service_definition(self, service_type: str, name: str,
criticality: str = 'medium') -> Dict[str, Any]:
"""Create a service definition from parameters."""
return {
'name': name,
'type': service_type,
'criticality': criticality,
'description': f'{name} - A {criticality} criticality {service_type} service',
'team': 'platform',
'environment': 'production',
'dependencies': [],
'tags': []
}
def generate_dashboard_specification(self, service_def: Dict[str, Any],
target_role: str = 'sre') -> Dict[str, Any]:
"""Generate comprehensive dashboard specification."""
service_name = service_def.get('name', 'Service')
service_type = service_def.get('type', 'api')
# Get role-specific configuration
role_config = self.ROLE_LAYOUTS.get(target_role, self.ROLE_LAYOUTS['sre'])
dashboard_spec = {
'metadata': {
'title': f"{service_name} - {target_role.upper()} Dashboard",
'service': service_def,
'target_role': target_role,
'generated_at': datetime.utcnow().isoformat() + 'Z',
'version': '1.0'
},
'configuration': {
'time_ranges': role_config['time_ranges'],
'default_time_range': role_config['time_ranges'][1], # Second option as default
'refresh_interval': role_config['default_refresh'],
'timezone': 'UTC',
'theme': 'dark'
},
'layout': self._generate_dashboard_layout(service_def, role_config),
'panels': self._generate_panels(service_def, role_config),
'variables': self._generate_template_variables(service_def),
'alerts_integration': self._generate_alerts_integration(service_def),
'drill_down_paths': self._generate_drill_down_paths(service_def)
}
return dashboard_spec
def _generate_dashboard_layout(self, service_def: Dict[str, Any],
role_config: Dict[str, Any]) -> Dict[str, Any]:
"""Generate dashboard layout configuration."""
return {
'grid_settings': {
'width': 24, # Grafana-style 24-column grid
'height_unit': 'px',
'cell_height': 30
},
'sections': [
{
'title': 'Service Overview',
'collapsed': False,
'y_position': 0,
'panels': ['service_status', 'slo_summary', 'error_budget']
},
{
'title': 'Golden Signals',
'collapsed': False,
'y_position': 8,
'panels': ['latency', 'traffic', 'errors', 'saturation']
},
{
'title': 'Resource Utilization',
'collapsed': False,
'y_position': 16,
'panels': ['cpu_usage', 'memory_usage', 'network_io', 'disk_io']
},
{
'title': 'Dependencies & Downstream',
'collapsed': True,
'y_position': 24,
'panels': ['dependency_status', 'downstream_latency', 'circuit_breakers']
}
]
}
def _generate_panels(self, service_def: Dict[str, Any],
role_config: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Generate dashboard panels based on service and role."""
service_name = service_def.get('name', 'service')
service_type = service_def.get('type', 'api')
panels = []
# Service Overview Panels
panels.extend(self._create_overview_panels(service_def))
# Golden Signals Panels
panels.extend(self._create_golden_signals_panels(service_def))
# Resource Utilization Panels
panels.extend(self._create_resource_panels(service_def))
# Service-specific panels
if service_type == 'api':
panels.extend(self._create_api_specific_panels(service_def))
elif service_type == 'database':
panels.extend(self._create_database_specific_panels(service_def))
elif service_type == 'queue':
panels.extend(self._create_queue_specific_panels(service_def))
# Role-specific additional panels
if 'business_metrics' in role_config['primary_focus']:
panels.extend(self._create_business_metrics_panels(service_def))
if 'capacity' in role_config['primary_focus']:
panels.extend(self._create_capacity_panels(service_def))
return panels
def _create_overview_panels(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Create service overview panels."""
service_name = service_def.get('name', 'service')
return [
{
'id': 'service_status',
'title': 'Service Status',
'type': 'stat',
'grid_pos': {'x': 0, 'y': 0, 'w': 6, 'h': 4},
'targets': [
{
'expr': f'up{{service="{service_name}"}}',
'legendFormat': 'Status'
}
],
'field_config': {
'overrides': [
{
'matcher': {'id': 'byName', 'options': 'Status'},
'properties': [
{'id': 'color', 'value': {'mode': 'thresholds'}},
{'id': 'thresholds', 'value': {
'steps': [
{'color': 'red', 'value': 0},
{'color': 'green', 'value': 1}
]
}},
{'id': 'mappings', 'value': [
{'options': {'0': {'text': 'DOWN'}}, 'type': 'value'},
{'options': {'1': {'text': 'UP'}}, 'type': 'value'}
]}
]
}
]
},
'options': {
'orientation': 'horizontal',
'textMode': 'value_and_name'
}
},
{
'id': 'slo_summary',
'title': 'SLO Achievement (30d)',
'type': 'stat',
'grid_pos': {'x': 6, 'y': 0, 'w': 9, 'h': 4},
'targets': [
{
'expr': f'(1 - (increase(http_requests_total{{service="{service_name}",code=~"5.."}}[30d]) / increase(http_requests_total{{service="{service_name}"}}[30d]))) * 100',
'legendFormat': 'Availability'
},
{
'expr': f'histogram_quantile(0.95, increase(http_request_duration_seconds_bucket{{service="{service_name}"}}[30d])) * 1000',
'legendFormat': 'P95 Latency (ms)'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'thresholds'},
'thresholds': {
'steps': [
{'color': 'red', 'value': 0},
{'color': 'yellow', 'value': 99.0},
{'color': 'green', 'value': 99.9}
]
}
}
},
'options': {
'orientation': 'horizontal',
'textMode': 'value_and_name'
}
},
{
'id': 'error_budget',
'title': 'Error Budget Remaining',
'type': 'gauge',
'grid_pos': {'x': 15, 'y': 0, 'w': 9, 'h': 4},
'targets': [
{
'expr': f'(1 - (increase(http_requests_total{{service="{service_name}",code=~"5.."}}[30d]) / increase(http_requests_total{{service="{service_name}"}}[30d])) - 0.999) / 0.001 * 100',
'legendFormat': 'Error Budget %'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'thresholds'},
'min': 0,
'max': 100,
'thresholds': {
'steps': [
{'color': 'red', 'value': 0},
{'color': 'yellow', 'value': 25},
{'color': 'green', 'value': 50}
]
},
'unit': 'percent'
}
},
'options': {
'showThresholdLabels': True,
'showThresholdMarkers': True
}
}
]
def _create_golden_signals_panels(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Create golden signals monitoring panels."""
service_name = service_def.get('name', 'service')
return [
{
'id': 'latency',
'title': 'Request Latency',
'type': 'timeseries',
'grid_pos': {'x': 0, 'y': 8, 'w': 12, 'h': 6},
'targets': [
{
'expr': f'histogram_quantile(0.50, rate(http_request_duration_seconds_bucket{{service="{service_name}"}}[5m])) * 1000',
'legendFormat': 'P50 Latency'
},
{
'expr': f'histogram_quantile(0.95, rate(http_request_duration_seconds_bucket{{service="{service_name}"}}[5m])) * 1000',
'legendFormat': 'P95 Latency'
},
{
'expr': f'histogram_quantile(0.99, rate(http_request_duration_seconds_bucket{{service="{service_name}"}}[5m])) * 1000',
'legendFormat': 'P99 Latency'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'palette-classic'},
'unit': 'ms',
'custom': {
'drawStyle': 'line',
'lineInterpolation': 'linear',
'lineWidth': 1,
'fillOpacity': 10
}
}
},
'options': {
'tooltip': {'mode': 'multi', 'sort': 'desc'},
'legend': {'displayMode': 'table', 'placement': 'bottom'}
}
},
{
'id': 'traffic',
'title': 'Request Rate',
'type': 'timeseries',
'grid_pos': {'x': 12, 'y': 8, 'w': 12, 'h': 6},
'targets': [
{
'expr': f'sum(rate(http_requests_total{{service="{service_name}"}}[5m]))',
'legendFormat': 'Total RPS'
},
{
'expr': f'sum(rate(http_requests_total{{service="{service_name}",code=~"2.."}}[5m]))',
'legendFormat': '2xx RPS'
},
{
'expr': f'sum(rate(http_requests_total{{service="{service_name}",code=~"4.."}}[5m]))',
'legendFormat': '4xx RPS'
},
{
'expr': f'sum(rate(http_requests_total{{service="{service_name}",code=~"5.."}}[5m]))',
'legendFormat': '5xx RPS'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'palette-classic'},
'unit': 'reqps',
'custom': {
'drawStyle': 'line',
'lineInterpolation': 'linear',
'lineWidth': 1,
'fillOpacity': 0
}
}
},
'options': {
'tooltip': {'mode': 'multi', 'sort': 'desc'},
'legend': {'displayMode': 'table', 'placement': 'bottom'}
}
},
{
'id': 'errors',
'title': 'Error Rate',
'type': 'timeseries',
'grid_pos': {'x': 0, 'y': 14, 'w': 12, 'h': 6},
'targets': [
{
'expr': f'sum(rate(http_requests_total{{service="{service_name}",code=~"5.."}}[5m])) / sum(rate(http_requests_total{{service="{service_name}"}}[5m])) * 100',
'legendFormat': '5xx Error Rate'
},
{
'expr': f'sum(rate(http_requests_total{{service="{service_name}",code=~"4.."}}[5m])) / sum(rate(http_requests_total{{service="{service_name}"}}[5m])) * 100',
'legendFormat': '4xx Error Rate'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'palette-classic'},
'unit': 'percent',
'custom': {
'drawStyle': 'line',
'lineInterpolation': 'linear',
'lineWidth': 2,
'fillOpacity': 20
}
},
'overrides': [
{
'matcher': {'id': 'byName', 'options': '5xx Error Rate'},
'properties': [{'id': 'color', 'value': {'fixedColor': 'red'}}]
}
]
},
'options': {
'tooltip': {'mode': 'multi', 'sort': 'desc'},
'legend': {'displayMode': 'table', 'placement': 'bottom'}
}
},
{
'id': 'saturation',
'title': 'Saturation Metrics',
'type': 'timeseries',
'grid_pos': {'x': 12, 'y': 14, 'w': 12, 'h': 6},
'targets': [
{
'expr': f'rate(process_cpu_seconds_total{{service="{service_name}"}}[5m]) * 100',
'legendFormat': 'CPU Usage %'
},
{
'expr': f'process_resident_memory_bytes{{service="{service_name}"}} / process_virtual_memory_max_bytes{{service="{service_name}"}} * 100',
'legendFormat': 'Memory Usage %'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'palette-classic'},
'unit': 'percent',
'max': 100,
'custom': {
'drawStyle': 'line',
'lineInterpolation': 'linear',
'lineWidth': 1,
'fillOpacity': 10
}
}
},
'options': {
'tooltip': {'mode': 'multi', 'sort': 'desc'},
'legend': {'displayMode': 'table', 'placement': 'bottom'}
}
}
]
def _create_resource_panels(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Create resource utilization panels."""
service_name = service_def.get('name', 'service')
return [
{
'id': 'cpu_usage',
'title': 'CPU Usage',
'type': 'gauge',
'grid_pos': {'x': 0, 'y': 20, 'w': 6, 'h': 4},
'targets': [
{
'expr': f'rate(process_cpu_seconds_total{{service="{service_name}"}}[5m]) * 100',
'legendFormat': 'CPU %'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'thresholds'},
'unit': 'percent',
'min': 0,
'max': 100,
'thresholds': {
'steps': [
{'color': 'green', 'value': 0},
{'color': 'yellow', 'value': 70},
{'color': 'red', 'value': 90}
]
}
}
},
'options': {
'showThresholdLabels': True,
'showThresholdMarkers': True
}
},
{
'id': 'memory_usage',
'title': 'Memory Usage',
'type': 'gauge',
'grid_pos': {'x': 6, 'y': 20, 'w': 6, 'h': 4},
'targets': [
{
'expr': f'process_resident_memory_bytes{{service="{service_name}"}} / 1024 / 1024',
'legendFormat': 'Memory MB'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'thresholds'},
'unit': 'decbytes',
'thresholds': {
'steps': [
{'color': 'green', 'value': 0},
{'color': 'yellow', 'value': 512000000}, # 512MB
{'color': 'red', 'value': 1024000000} # 1GB
]
}
}
}
},
{
'id': 'network_io',
'title': 'Network I/O',
'type': 'timeseries',
'grid_pos': {'x': 12, 'y': 20, 'w': 6, 'h': 4},
'targets': [
{
'expr': f'rate(process_network_receive_bytes_total{{service="{service_name}"}}[5m])',
'legendFormat': 'RX Bytes/s'
},
{
'expr': f'rate(process_network_transmit_bytes_total{{service="{service_name}"}}[5m])',
'legendFormat': 'TX Bytes/s'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'palette-classic'},
'unit': 'binBps'
}
}
},
{
'id': 'disk_io',
'title': 'Disk I/O',
'type': 'timeseries',
'grid_pos': {'x': 18, 'y': 20, 'w': 6, 'h': 4},
'targets': [
{
'expr': f'rate(process_disk_read_bytes_total{{service="{service_name}"}}[5m])',
'legendFormat': 'Read Bytes/s'
},
{
'expr': f'rate(process_disk_write_bytes_total{{service="{service_name}"}}[5m])',
'legendFormat': 'Write Bytes/s'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'palette-classic'},
'unit': 'binBps'
}
}
}
]
def _create_api_specific_panels(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Create API-specific panels."""
service_name = service_def.get('name', 'service')
return [
{
'id': 'endpoint_latency',
'title': 'Top Slowest Endpoints',
'type': 'table',
'grid_pos': {'x': 0, 'y': 24, 'w': 12, 'h': 6},
'targets': [
{
'expr': f'topk(10, histogram_quantile(0.95, sum by (handler) (rate(http_request_duration_seconds_bucket{{service="{service_name}"}}[5m])))) * 1000',
'legendFormat': '{{handler}}',
'format': 'table',
'instant': True
}
],
'transformations': [
{
'id': 'organize',
'options': {
'excludeByName': {'Time': True},
'renameByName': {'Value': 'P95 Latency (ms)'}
}
}
],
'field_config': {
'overrides': [
{
'matcher': {'id': 'byName', 'options': 'P95 Latency (ms)'},
'properties': [
{'id': 'color', 'value': {'mode': 'thresholds'}},
{'id': 'thresholds', 'value': {
'steps': [
{'color': 'green', 'value': 0},
{'color': 'yellow', 'value': 100},
{'color': 'red', 'value': 500}
]
}}
]
}
]
}
},
{
'id': 'request_size_distribution',
'title': 'Request Size Distribution',
'type': 'heatmap',
'grid_pos': {'x': 12, 'y': 24, 'w': 12, 'h': 6},
'targets': [
{
'expr': f'sum by (le) (rate(http_request_size_bytes_bucket{{service="{service_name}"}}[5m]))',
'legendFormat': '{{le}}'
}
],
'options': {
'calculate': True,
'yAxis': {'unit': 'bytes'},
'color': {'scheme': 'Spectral'}
}
}
]
def _create_database_specific_panels(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Create database-specific panels."""
service_name = service_def.get('name', 'service')
return [
{
'id': 'db_connections',
'title': 'Database Connections',
'type': 'timeseries',
'grid_pos': {'x': 0, 'y': 24, 'w': 8, 'h': 6},
'targets': [
{
'expr': f'db_connections_active{{service="{service_name}"}}',
'legendFormat': 'Active Connections'
},
{
'expr': f'db_connections_idle{{service="{service_name}"}}',
'legendFormat': 'Idle Connections'
},
{
'expr': f'db_connections_max{{service="{service_name}"}}',
'legendFormat': 'Max Connections'
}
]
},
{
'id': 'query_performance',
'title': 'Query Performance',
'type': 'timeseries',
'grid_pos': {'x': 8, 'y': 24, 'w': 8, 'h': 6},
'targets': [
{
'expr': f'rate(db_queries_total{{service="{service_name}"}}[5m])',
'legendFormat': 'Queries/sec'
},
{
'expr': f'rate(db_slow_queries_total{{service="{service_name}"}}[5m])',
'legendFormat': 'Slow Queries/sec'
}
]
},
{
'id': 'db_locks',
'title': 'Database Locks',
'type': 'stat',
'grid_pos': {'x': 16, 'y': 24, 'w': 8, 'h': 6},
'targets': [
{
'expr': f'db_locks_waiting{{service="{service_name}"}}',
'legendFormat': 'Waiting Locks'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'thresholds'},
'thresholds': {
'steps': [
{'color': 'green', 'value': 0},
{'color': 'yellow', 'value': 1},
{'color': 'red', 'value': 5}
]
}
}
}
}
]
def _create_queue_specific_panels(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Create queue-specific panels."""
service_name = service_def.get('name', 'service')
return [
{
'id': 'queue_depth',
'title': 'Queue Depth',
'type': 'timeseries',
'grid_pos': {'x': 0, 'y': 24, 'w': 12, 'h': 6},
'targets': [
{
'expr': f'queue_depth{{service="{service_name}"}}',
'legendFormat': 'Messages in Queue'
}
]
},
{
'id': 'message_throughput',
'title': 'Message Throughput',
'type': 'timeseries',
'grid_pos': {'x': 12, 'y': 24, 'w': 12, 'h': 6},
'targets': [
{
'expr': f'rate(messages_published_total{{service="{service_name}"}}[5m])',
'legendFormat': 'Published/sec'
},
{
'expr': f'rate(messages_consumed_total{{service="{service_name}"}}[5m])',
'legendFormat': 'Consumed/sec'
}
]
}
]
def _create_business_metrics_panels(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Create business metrics panels."""
service_name = service_def.get('name', 'service')
return [
{
'id': 'business_kpis',
'title': 'Business KPIs',
'type': 'stat',
'grid_pos': {'x': 0, 'y': 30, 'w': 24, 'h': 4},
'targets': [
{
'expr': f'rate(business_transactions_total{{service="{service_name}"}}[1h])',
'legendFormat': 'Transactions/hour'
},
{
'expr': f'avg(business_transaction_value{{service="{service_name}"}}) * rate(business_transactions_total{{service="{service_name}"}}[1h])',
'legendFormat': 'Revenue/hour'
},
{
'expr': f'rate(user_registrations_total{{service="{service_name}"}}[1h])',
'legendFormat': 'New Users/hour'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'palette-classic'},
'custom': {
'displayMode': 'basic'
}
}
},
'options': {
'orientation': 'horizontal',
'textMode': 'value_and_name'
}
}
]
def _create_capacity_panels(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Create capacity planning panels."""
service_name = service_def.get('name', 'service')
return [
{
'id': 'capacity_trends',
'title': 'Capacity Trends (7d)',
'type': 'timeseries',
'grid_pos': {'x': 0, 'y': 34, 'w': 24, 'h': 6},
'targets': [
{
'expr': f'predict_linear(avg_over_time(rate(http_requests_total{{service="{service_name}"}}[5m])[7d:1h]), 7*24*3600)',
'legendFormat': 'Predicted Traffic (7d)'
},
{
'expr': f'predict_linear(avg_over_time(process_resident_memory_bytes{{service="{service_name}"}}[7d:1h]), 7*24*3600)',
'legendFormat': 'Predicted Memory Usage (7d)'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'palette-classic'},
'custom': {
'drawStyle': 'line',
'lineStyle': {'dash': [10, 10]}
}
}
}
}
]
def _generate_template_variables(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Generate template variables for dynamic dashboard filtering."""
service_name = service_def.get('name', 'service')
return [
{
'name': 'environment',
'type': 'query',
'query': 'label_values(environment)',
'current': {'text': 'production', 'value': 'production'},
'includeAll': False,
'multi': False,
'refresh': 'on_dashboard_load'
},
{
'name': 'instance',
'type': 'query',
'query': f'label_values(up{{service="{service_name}"}}, instance)',
'current': {'text': 'All', 'value': '$__all'},
'includeAll': True,
'multi': True,
'refresh': 'on_time_range_change'
},
{
'name': 'handler',
'type': 'query',
'query': f'label_values(http_requests_total{{service="{service_name}"}}, handler)',
'current': {'text': 'All', 'value': '$__all'},
'includeAll': True,
'multi': True,
'refresh': 'on_time_range_change'
}
]
def _generate_alerts_integration(self, service_def: Dict[str, Any]) -> Dict[str, Any]:
"""Generate alerts integration configuration."""
service_name = service_def.get('name', 'service')
return {
'alert_annotations': True,
'alert_rules_query': f'ALERTS{{service="{service_name}"}}',
'alert_panels': [
{
'title': 'Active Alerts',
'type': 'table',
'query': f'ALERTS{{service="{service_name}",alertstate="firing"}}',
'columns': ['alertname', 'severity', 'instance', 'description']
}
]
}
def _generate_drill_down_paths(self, service_def: Dict[str, Any]) -> Dict[str, Any]:
"""Generate drill-down navigation paths."""
service_name = service_def.get('name', 'service')
return {
'service_overview': {
'from': 'service_status',
'to': 'detailed_health_dashboard',
'url': f'/d/service-health/{service_name}-health',
'params': ['var-service', 'var-environment']
},
'error_investigation': {
'from': 'errors',
'to': 'error_details_dashboard',
'url': f'/d/errors/{service_name}-errors',
'params': ['var-service', 'var-time_range']
},
'latency_analysis': {
'from': 'latency',
'to': 'trace_analysis_dashboard',
'url': f'/d/traces/{service_name}-traces',
'params': ['var-service', 'var-handler']
},
'capacity_planning': {
'from': 'saturation',
'to': 'capacity_dashboard',
'url': f'/d/capacity/{service_name}-capacity',
'params': ['var-service', 'var-time_range']
}
}
def generate_grafana_json(self, dashboard_spec: Dict[str, Any]) -> Dict[str, Any]:
"""Convert dashboard specification to Grafana JSON format."""
metadata = dashboard_spec['metadata']
config = dashboard_spec['configuration']
grafana_json = {
'dashboard': {
'id': None,
'title': metadata['title'],
'tags': [metadata['service']['type'], metadata['target_role'], 'generated'],
'timezone': config['timezone'],
'refresh': config['refresh_interval'],
'time': {
'from': 'now-1h',
'to': 'now'
},
'templating': {
'list': dashboard_spec['variables']
},
'panels': self._convert_panels_to_grafana_format(dashboard_spec['panels']),
'version': 1,
'schemaVersion': 30
},
'overwrite': True
}
return grafana_json
def _convert_panels_to_grafana_format(self, panels: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Convert panel specifications to Grafana format."""
grafana_panels = []
for panel in panels:
grafana_panel = {
'id': hash(panel['id']) % 1000, # Generate numeric ID
'title': panel['title'],
'type': panel['type'],
'gridPos': panel['grid_pos'],
'targets': panel['targets'],
'fieldConfig': panel.get('field_config', {}),
'options': panel.get('options', {}),
'transformations': panel.get('transformations', [])
}
grafana_panels.append(grafana_panel)
return grafana_panels
def generate_documentation(self, dashboard_spec: Dict[str, Any]) -> str:
"""Generate documentation for the dashboard."""
metadata = dashboard_spec['metadata']
service = metadata['service']
doc_content = f"""# {metadata['title']} Documentation
## Overview
This dashboard provides comprehensive monitoring for {service['name']}, a {service['type']} service with {service['criticality']} criticality.
**Target Audience:** {metadata['target_role'].upper()} teams
**Generated:** {metadata['generated_at']}
## Dashboard Sections
### Service Overview
- **Service Status**: Real-time availability status
- **SLO Achievement**: 30-day SLO compliance metrics
- **Error Budget**: Remaining error budget visualization
### Golden Signals Monitoring
- **Latency**: P50, P95, P99 response times
- **Traffic**: Request rate by status code
- **Errors**: Error rates for 4xx and 5xx responses
- **Saturation**: CPU and memory utilization
### Resource Utilization
- **CPU Usage**: Process CPU consumption
- **Memory Usage**: Memory utilization tracking
- **Network I/O**: Network throughput metrics
- **Disk I/O**: Disk read/write operations
## Key Metrics
### SLIs Tracked
"""
# Add service-type specific metrics
service_type = service.get('type', 'api')
if service_type in self.SERVICE_METRICS:
metrics = self.SERVICE_METRICS[service_type]['key_metrics']
for metric in metrics:
doc_content += f"- `{metric}`: Core service metric\n"
doc_content += f"""
## Alert Integration
- Active alerts are displayed in context with relevant panels
- Alert annotations show on time series charts
- Click-through to alert management system available
## Drill-Down Paths
"""
drill_downs = dashboard_spec.get('drill_down_paths', {})
for path_name, path_config in drill_downs.items():
doc_content += f"- **{path_name}**: From {path_config['from']} → {path_config['to']}\n"
doc_content += f"""
## Usage Guidelines
### Time Ranges
Use appropriate time ranges for different investigation types:
- **Real-time monitoring**: 15m - 1h
- **Recent incident investigation**: 1h - 6h
- **Trend analysis**: 1d - 7d
- **Capacity planning**: 7d - 30d
### Variables
- **environment**: Filter by deployment environment
- **instance**: Focus on specific service instances
- **handler**: Filter by API endpoint or handler
### Performance Optimization
- Use longer time ranges for capacity planning
- Refresh intervals are optimized per role:
- SRE: 30s for operational awareness
- Developer: 1m for troubleshooting
- Executive: 5m for high-level monitoring
## Maintenance
- Dashboard panels automatically adapt to service changes
- Template variables refresh based on actual metric labels
- Review and update business metrics quarterly
"""
return doc_content
def export_specification(self, dashboard_spec: Dict[str, Any], output_file: str,
format_type: str = 'json'):
"""Export dashboard specification."""
if format_type.lower() == 'json':
with open(output_file, 'w') as f:
json.dump(dashboard_spec, f, indent=2)
elif format_type.lower() == 'grafana':
grafana_json = self.generate_grafana_json(dashboard_spec)
with open(output_file, 'w') as f:
json.dump(grafana_json, f, indent=2)
else:
raise ValueError(f"Unsupported format: {format_type}")
def print_summary(self, dashboard_spec: Dict[str, Any]):
"""Print human-readable summary of dashboard specification."""
metadata = dashboard_spec['metadata']
service = metadata['service']
config = dashboard_spec['configuration']
panels = dashboard_spec['panels']
print(f"\n{'='*60}")
print(f"DASHBOARD SPECIFICATION SUMMARY")
print(f"{'='*60}")
print(f"\nDashboard Details:")
print(f" Title: {metadata['title']}")
print(f" Target Role: {metadata['target_role'].upper()}")
print(f" Service: {service['name']} ({service['type']})")
print(f" Criticality: {service['criticality']}")
print(f" Generated: {metadata['generated_at']}")
print(f"\nConfiguration:")
print(f" Default Time Range: {config['default_time_range']}")
print(f" Refresh Interval: {config['refresh_interval']}")
print(f" Available Time Ranges: {', '.join(config['time_ranges'])}")
print(f"\nPanels ({len(panels)}):")
panel_types = {}
for panel in panels:
panel_type = panel['type']
panel_types[panel_type] = panel_types.get(panel_type, 0) + 1
for panel_type, count in panel_types.items():
print(f" {panel_type}: {count}")
variables = dashboard_spec.get('variables', [])
print(f"\nTemplate Variables ({len(variables)}):")
for var in variables:
print(f" {var['name']} ({var['type']})")
drill_downs = dashboard_spec.get('drill_down_paths', {})
print(f"\nDrill-down Paths: {len(drill_downs)}")
print(f"\nKey Features:")
print(f" • Golden Signals monitoring")
print(f" • Resource utilization tracking")
print(f" • Alert integration")
print(f" • Role-optimized layout")
print(f" • Service-type specific panels")
print(f"\n{'='*60}\n")
def main():
"""Main function for CLI usage."""
parser = argparse.ArgumentParser(
description='Generate comprehensive dashboard specifications',
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
# Generate from service definition file
python dashboard_generator.py --input service.json --output dashboard.json
# Generate from command line parameters
python dashboard_generator.py --service-type api --name "Payment Service" --output payment_dashboard.json
# Generate Grafana-compatible JSON
python dashboard_generator.py --input service.json --output dashboard.json --format grafana
# Generate with specific role focus
python dashboard_generator.py --service-type web --name "Frontend" --role developer --output frontend_dev.json
"""
)
parser.add_argument('--input', '-i',
help='Input service definition JSON file')
parser.add_argument('--output', '-o',
help='Output dashboard specification file')
parser.add_argument('--service-type',
choices=['api', 'web', 'database', 'queue', 'batch', 'ml'],
help='Service type')
parser.add_argument('--name',
help='Service name')
parser.add_argument('--criticality',
choices=['critical', 'high', 'medium', 'low'],
default='medium',
help='Service criticality level')
parser.add_argument('--role',
choices=['sre', 'developer', 'executive', 'ops'],
default='sre',
help='Target role for dashboard optimization')
parser.add_argument('--format',
choices=['json', 'grafana'],
default='json',
help='Output format (json specification or grafana compatible)')
parser.add_argument('--doc-output',
help='Generate documentation file')
parser.add_argument('--summary-only', action='store_true',
help='Only display summary, do not save files')
args = parser.parse_args()
if not args.input and not (args.service_type and args.name):
parser.error("Must provide either --input file or --service-type and --name")
generator = DashboardGenerator()
try:
# Load or create service definition
if args.input:
service_def = generator.load_service_definition(args.input)
else:
service_def = generator.create_service_definition(
args.service_type, args.name, args.criticality
)
# Generate dashboard specification
dashboard_spec = generator.generate_dashboard_specification(service_def, args.role)
# Output results
if not args.summary_only:
output_file = args.output or f"{service_def['name'].replace(' ', '_').lower()}_dashboard.json"
generator.export_specification(dashboard_spec, output_file, args.format)
print(f"Dashboard specification saved to: {output_file}")
# Generate documentation if requested
if args.doc_output:
documentation = generator.generate_documentation(dashboard_spec)
with open(args.doc_output, 'w') as f:
f.write(documentation)
print(f"Documentation saved to: {args.doc_output}")
# Always show summary
generator.print_summary(dashboard_spec)
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
sys.exit(1)
if __name__ == '__main__':
main()
FILE:scripts/slo_designer.py
#!/usr/bin/env python3
"""
SLO Designer - Generate comprehensive SLI/SLO frameworks for services
This script analyzes service descriptions and generates complete SLO frameworks including:
- SLI definitions based on service characteristics
- SLO targets based on criticality and user impact
- Error budget calculations and policies
- Multi-window burn rate alerts
- SLA recommendations for customer-facing services
Usage:
python slo_designer.py --input service_definition.json --output slo_framework.json
python slo_designer.py --service-type api --criticality high --user-facing true
"""
import json
import argparse
import sys
import math
from typing import Dict, List, Any, Tuple
from datetime import datetime, timedelta
class SLODesigner:
"""Design and generate SLO frameworks for services."""
# SLO target recommendations based on service criticality
SLO_TARGETS = {
'critical': {
'availability': 0.9999, # 99.99% - 4.38 minutes downtime/month
'latency_p95': 100, # 95th percentile latency in ms
'latency_p99': 500, # 99th percentile latency in ms
'error_rate': 0.001 # 0.1% error rate
},
'high': {
'availability': 0.999, # 99.9% - 43.8 minutes downtime/month
'latency_p95': 200, # 95th percentile latency in ms
'latency_p99': 1000, # 99th percentile latency in ms
'error_rate': 0.005 # 0.5% error rate
},
'medium': {
'availability': 0.995, # 99.5% - 3.65 hours downtime/month
'latency_p95': 500, # 95th percentile latency in ms
'latency_p99': 2000, # 99th percentile latency in ms
'error_rate': 0.01 # 1% error rate
},
'low': {
'availability': 0.99, # 99% - 7.3 hours downtime/month
'latency_p95': 1000, # 95th percentile latency in ms
'latency_p99': 5000, # 99th percentile latency in ms
'error_rate': 0.02 # 2% error rate
}
}
# Burn rate windows for multi-window alerting
BURN_RATE_WINDOWS = [
{'short': '5m', 'long': '1h', 'burn_rate': 14.4, 'budget_consumed': '2%'},
{'short': '30m', 'long': '6h', 'burn_rate': 6, 'budget_consumed': '5%'},
{'short': '2h', 'long': '1d', 'burn_rate': 3, 'budget_consumed': '10%'},
{'short': '6h', 'long': '3d', 'burn_rate': 1, 'budget_consumed': '10%'}
]
# Service type specific SLI recommendations
SERVICE_TYPE_SLIS = {
'api': ['availability', 'latency', 'error_rate', 'throughput'],
'web': ['availability', 'latency', 'error_rate', 'page_load_time'],
'database': ['availability', 'query_latency', 'connection_success_rate', 'replication_lag'],
'queue': ['availability', 'message_processing_time', 'queue_depth', 'message_loss_rate'],
'batch': ['job_success_rate', 'job_duration', 'data_freshness', 'resource_utilization'],
'ml': ['model_accuracy', 'prediction_latency', 'training_success_rate', 'feature_freshness']
}
def __init__(self):
"""Initialize the SLO Designer."""
self.service_config = {}
self.slo_framework = {}
def load_service_definition(self, file_path: str) -> Dict[str, Any]:
"""Load service definition from JSON file."""
try:
with open(file_path, 'r') as f:
return json.load(f)
except FileNotFoundError:
raise ValueError(f"Service definition file not found: {file_path}")
except json.JSONDecodeError as e:
raise ValueError(f"Invalid JSON in service definition: {e}")
def create_service_definition(self, service_type: str, criticality: str,
user_facing: bool, name: str = None) -> Dict[str, Any]:
"""Create a service definition from parameters."""
return {
'name': name or f'{service_type}_service',
'type': service_type,
'criticality': criticality,
'user_facing': user_facing,
'description': f'A {criticality} criticality {service_type} service',
'dependencies': [],
'team': 'platform',
'environment': 'production'
}
def generate_slis(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Generate Service Level Indicators based on service characteristics."""
service_type = service_def.get('type', 'api')
base_slis = self.SERVICE_TYPE_SLIS.get(service_type, ['availability', 'latency', 'error_rate'])
slis = []
for sli_name in base_slis:
sli = self._create_sli_definition(sli_name, service_def)
if sli:
slis.append(sli)
# Add user-facing specific SLIs
if service_def.get('user_facing', False):
user_slis = self._generate_user_facing_slis(service_def)
slis.extend(user_slis)
return slis
def _create_sli_definition(self, sli_name: str, service_def: Dict[str, Any]) -> Dict[str, Any]:
"""Create detailed SLI definition."""
service_name = service_def.get('name', 'service')
sli_definitions = {
'availability': {
'name': 'Availability',
'description': 'Percentage of successful requests',
'type': 'ratio',
'good_events': f'sum(rate(http_requests_total{{service="{service_name}",code!~"5.."}}))',
'total_events': f'sum(rate(http_requests_total{{service="{service_name}"}}))',
'unit': 'percentage'
},
'latency': {
'name': 'Request Latency P95',
'description': '95th percentile of request latency',
'type': 'threshold',
'query': f'histogram_quantile(0.95, rate(http_request_duration_seconds_bucket{{service="{service_name}"}}[5m]))',
'unit': 'seconds'
},
'error_rate': {
'name': 'Error Rate',
'description': 'Rate of 5xx errors',
'type': 'ratio',
'good_events': f'sum(rate(http_requests_total{{service="{service_name}",code!~"5.."}}))',
'total_events': f'sum(rate(http_requests_total{{service="{service_name}"}}))',
'unit': 'percentage'
},
'throughput': {
'name': 'Request Throughput',
'description': 'Requests per second',
'type': 'gauge',
'query': f'sum(rate(http_requests_total{{service="{service_name}"}}[5m]))',
'unit': 'requests/sec'
},
'page_load_time': {
'name': 'Page Load Time P95',
'description': '95th percentile of page load time',
'type': 'threshold',
'query': f'histogram_quantile(0.95, rate(page_load_duration_seconds_bucket{{service="{service_name}"}}[5m]))',
'unit': 'seconds'
},
'query_latency': {
'name': 'Database Query Latency P95',
'description': '95th percentile of database query latency',
'type': 'threshold',
'query': f'histogram_quantile(0.95, rate(db_query_duration_seconds_bucket{{service="{service_name}"}}[5m]))',
'unit': 'seconds'
},
'connection_success_rate': {
'name': 'Database Connection Success Rate',
'description': 'Percentage of successful database connections',
'type': 'ratio',
'good_events': f'sum(rate(db_connections_total{{service="{service_name}",status="success"}}[5m]))',
'total_events': f'sum(rate(db_connections_total{{service="{service_name}"}}[5m]))',
'unit': 'percentage'
}
}
return sli_definitions.get(sli_name)
def _generate_user_facing_slis(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Generate additional SLIs for user-facing services."""
service_name = service_def.get('name', 'service')
return [
{
'name': 'User Journey Success Rate',
'description': 'Percentage of successful complete user journeys',
'type': 'ratio',
'good_events': f'sum(rate(user_journey_total{{service="{service_name}",status="success"}}[5m]))',
'total_events': f'sum(rate(user_journey_total{{service="{service_name}"}}[5m]))',
'unit': 'percentage'
},
{
'name': 'Feature Availability',
'description': 'Percentage of time key features are available',
'type': 'ratio',
'good_events': f'sum(rate(feature_checks_total{{service="{service_name}",status="available"}}[5m]))',
'total_events': f'sum(rate(feature_checks_total{{service="{service_name}"}}[5m]))',
'unit': 'percentage'
}
]
def generate_slos(self, service_def: Dict[str, Any], slis: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Generate Service Level Objectives based on service criticality."""
criticality = service_def.get('criticality', 'medium')
targets = self.SLO_TARGETS.get(criticality, self.SLO_TARGETS['medium'])
slos = []
for sli in slis:
slo = self._create_slo_from_sli(sli, targets, service_def)
if slo:
slos.append(slo)
return slos
def _create_slo_from_sli(self, sli: Dict[str, Any], targets: Dict[str, float],
service_def: Dict[str, Any]) -> Dict[str, Any]:
"""Create SLO definition from SLI."""
sli_name = sli['name'].lower().replace(' ', '_')
# Map SLI names to target keys
target_mapping = {
'availability': 'availability',
'request_latency_p95': 'latency_p95',
'error_rate': 'error_rate',
'user_journey_success_rate': 'availability',
'feature_availability': 'availability',
'page_load_time_p95': 'latency_p95',
'database_query_latency_p95': 'latency_p95',
'database_connection_success_rate': 'availability'
}
target_key = target_mapping.get(sli_name)
if not target_key:
return None
target_value = targets.get(target_key)
if target_value is None:
return None
# Determine comparison operator and format target
if 'latency' in sli_name or 'duration' in sli_name:
operator = '<='
target_display = f"{target_value}ms" if target_value < 10 else f"{target_value/1000}s"
elif 'rate' in sli_name and 'error' in sli_name:
operator = '<='
target_display = f"{target_value * 100}%"
target_value = target_value # Keep as decimal
else:
operator = '>='
target_display = f"{target_value * 100}%"
# Calculate time windows
time_windows = ['1h', '1d', '7d', '30d']
slo = {
'name': f"{sli['name']} SLO",
'description': f"Service level objective for {sli['description'].lower()}",
'sli_name': sli['name'],
'target_value': target_value,
'target_display': target_display,
'operator': operator,
'time_windows': time_windows,
'measurement_window': '30d',
'service': service_def.get('name', 'service'),
'criticality': service_def.get('criticality', 'medium')
}
return slo
def calculate_error_budgets(self, slos: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Calculate error budgets for SLOs."""
error_budgets = []
for slo in slos:
if slo['operator'] == '>=': # Availability-type SLOs
target = slo['target_value']
error_budget_rate = 1 - target
# Calculate budget for different time windows
time_windows = {
'1h': 3600,
'1d': 86400,
'7d': 604800,
'30d': 2592000
}
budgets = {}
for window, seconds in time_windows.items():
budget_seconds = seconds * error_budget_rate
if budget_seconds < 60:
budgets[window] = f"{budget_seconds:.1f} seconds"
elif budget_seconds < 3600:
budgets[window] = f"{budget_seconds/60:.1f} minutes"
else:
budgets[window] = f"{budget_seconds/3600:.1f} hours"
error_budget = {
'slo_name': slo['name'],
'error_budget_rate': error_budget_rate,
'error_budget_percentage': f"{error_budget_rate * 100:.3f}%",
'budgets_by_window': budgets,
'burn_rate_alerts': self._generate_burn_rate_alerts(slo, error_budget_rate)
}
error_budgets.append(error_budget)
return error_budgets
def _generate_burn_rate_alerts(self, slo: Dict[str, Any], error_budget_rate: float) -> List[Dict[str, Any]]:
"""Generate multi-window burn rate alerts."""
alerts = []
service_name = slo['service']
sli_query = self._get_sli_query_for_burn_rate(slo)
for window_config in self.BURN_RATE_WINDOWS:
alert = {
'name': f"{slo['sli_name']} Burn Rate {window_config['budget_consumed']} Alert",
'description': f"Alert when {slo['sli_name']} is consuming error budget at {window_config['burn_rate']}x rate",
'severity': self._determine_alert_severity(float(window_config['budget_consumed'].rstrip('%'))),
'short_window': window_config['short'],
'long_window': window_config['long'],
'burn_rate_threshold': window_config['burn_rate'],
'budget_consumed': window_config['budget_consumed'],
'condition': f"({sli_query}_short > {window_config['burn_rate']}) and ({sli_query}_long > {window_config['burn_rate']})",
'annotations': {
'summary': f"High burn rate detected for {slo['sli_name']}",
'description': f"Error budget consumption rate is {window_config['burn_rate']}x normal, will exhaust {window_config['budget_consumed']} of monthly budget"
}
}
alerts.append(alert)
return alerts
def _get_sli_query_for_burn_rate(self, slo: Dict[str, Any]) -> str:
"""Generate SLI query fragment for burn rate calculation."""
service_name = slo['service']
sli_name = slo['sli_name'].lower().replace(' ', '_')
if 'availability' in sli_name or 'success' in sli_name:
return f"(1 - (sum(rate(http_requests_total{{service='{service_name}',code!~'5..'}})) / sum(rate(http_requests_total{{service='{service_name}'}}))))"
elif 'error' in sli_name:
return f"(sum(rate(http_requests_total{{service='{service_name}',code=~'5..'}})) / sum(rate(http_requests_total{{service='{service_name}'}})))"
else:
return f"sli_burn_rate_{sli_name}"
def _determine_alert_severity(self, budget_consumed_percent: float) -> str:
"""Determine alert severity based on budget consumption rate."""
if budget_consumed_percent <= 2:
return 'critical'
elif budget_consumed_percent <= 5:
return 'warning'
else:
return 'info'
def generate_sla_recommendations(self, service_def: Dict[str, Any],
slos: List[Dict[str, Any]]) -> Dict[str, Any]:
"""Generate SLA recommendations for customer-facing services."""
if not service_def.get('user_facing', False):
return {
'applicable': False,
'reason': 'SLA not recommended for non-user-facing services'
}
criticality = service_def.get('criticality', 'medium')
# SLA targets should be more conservative than SLO targets
sla_buffer = 0.001 # 0.1% buffer below SLO
sla_recommendations = {
'applicable': True,
'service': service_def.get('name'),
'commitments': [],
'penalties': self._generate_penalty_structure(criticality),
'measurement_methodology': 'External synthetic monitoring from multiple geographic locations',
'exclusions': [
'Planned maintenance windows (with 72h advance notice)',
'Customer-side network or infrastructure issues',
'Force majeure events',
'Third-party service dependencies beyond our control'
]
}
for slo in slos:
if slo['operator'] == '>=' and 'availability' in slo['sli_name'].lower():
sla_target = max(0.9, slo['target_value'] - sla_buffer)
commitment = {
'metric': slo['sli_name'],
'target': sla_target,
'target_display': f"{sla_target * 100:.2f}%",
'measurement_window': 'monthly',
'measurement_method': 'Uptime monitoring with 1-minute granularity'
}
sla_recommendations['commitments'].append(commitment)
return sla_recommendations
def _generate_penalty_structure(self, criticality: str) -> List[Dict[str, Any]]:
"""Generate penalty structure based on service criticality."""
penalty_structures = {
'critical': [
{'breach_threshold': '< 99.99%', 'credit_percentage': 10},
{'breach_threshold': '< 99.9%', 'credit_percentage': 25},
{'breach_threshold': '< 99%', 'credit_percentage': 50}
],
'high': [
{'breach_threshold': '< 99.9%', 'credit_percentage': 10},
{'breach_threshold': '< 99.5%', 'credit_percentage': 25}
],
'medium': [
{'breach_threshold': '< 99.5%', 'credit_percentage': 10}
],
'low': []
}
return penalty_structures.get(criticality, [])
def generate_framework(self, service_def: Dict[str, Any]) -> Dict[str, Any]:
"""Generate complete SLO framework."""
# Generate SLIs
slis = self.generate_slis(service_def)
# Generate SLOs
slos = self.generate_slos(service_def, slis)
# Calculate error budgets
error_budgets = self.calculate_error_budgets(slos)
# Generate SLA recommendations
sla_recommendations = self.generate_sla_recommendations(service_def, slos)
# Create comprehensive framework
framework = {
'metadata': {
'service': service_def,
'generated_at': datetime.utcnow().isoformat() + 'Z',
'framework_version': '1.0'
},
'slis': slis,
'slos': slos,
'error_budgets': error_budgets,
'sla_recommendations': sla_recommendations,
'monitoring_recommendations': self._generate_monitoring_recommendations(service_def),
'implementation_guide': self._generate_implementation_guide(service_def, slis, slos)
}
return framework
def _generate_monitoring_recommendations(self, service_def: Dict[str, Any]) -> Dict[str, Any]:
"""Generate monitoring tool recommendations."""
service_type = service_def.get('type', 'api')
recommendations = {
'metrics': {
'collection': 'Prometheus with service discovery',
'retention': '90 days for raw metrics, 1 year for aggregated',
'alerting': 'Prometheus Alertmanager with multi-window burn rate alerts'
},
'logging': {
'format': 'Structured JSON logs with correlation IDs',
'aggregation': 'ELK stack or equivalent with proper indexing',
'retention': '30 days for debug logs, 90 days for error logs'
},
'tracing': {
'sampling': 'Adaptive sampling with 1% base rate',
'storage': 'Jaeger or Zipkin with 7-day retention',
'integration': 'OpenTelemetry instrumentation'
}
}
if service_type == 'web':
recommendations['synthetic_monitoring'] = {
'frequency': 'Every 1 minute from 3+ geographic locations',
'checks': 'Full user journey simulation',
'tools': 'Pingdom, DataDog Synthetics, or equivalent'
}
return recommendations
def _generate_implementation_guide(self, service_def: Dict[str, Any],
slis: List[Dict[str, Any]],
slos: List[Dict[str, Any]]) -> Dict[str, Any]:
"""Generate implementation guide for the SLO framework."""
return {
'prerequisites': [
'Service instrumented with metrics collection (Prometheus format)',
'Structured logging with correlation IDs',
'Monitoring infrastructure (Prometheus, Grafana, Alertmanager)',
'Incident response processes and escalation policies'
],
'implementation_steps': [
{
'step': 1,
'title': 'Instrument Service',
'description': 'Add metrics collection for all defined SLIs',
'estimated_effort': '1-2 days'
},
{
'step': 2,
'title': 'Configure Recording Rules',
'description': 'Set up Prometheus recording rules for SLI calculations',
'estimated_effort': '4-8 hours'
},
{
'step': 3,
'title': 'Implement Burn Rate Alerts',
'description': 'Configure multi-window burn rate alerting rules',
'estimated_effort': '1 day'
},
{
'step': 4,
'title': 'Create SLO Dashboard',
'description': 'Build Grafana dashboard for SLO tracking and error budget monitoring',
'estimated_effort': '4-6 hours'
},
{
'step': 5,
'title': 'Test and Validate',
'description': 'Test alerting and validate SLI measurements against expectations',
'estimated_effort': '1-2 days'
},
{
'step': 6,
'title': 'Documentation and Training',
'description': 'Document runbooks and train team on SLO monitoring',
'estimated_effort': '1 day'
}
],
'validation_checklist': [
'All SLIs produce expected metric values',
'Burn rate alerts fire correctly during simulated outages',
'Error budget calculations match manual verification',
'Dashboard displays accurate SLO achievement rates',
'Alert routing reaches correct escalation paths',
'Runbooks are complete and tested'
]
}
def export_json(self, framework: Dict[str, Any], output_file: str):
"""Export framework as JSON."""
with open(output_file, 'w') as f:
json.dump(framework, f, indent=2)
def print_summary(self, framework: Dict[str, Any]):
"""Print human-readable summary of the SLO framework."""
service = framework['metadata']['service']
slis = framework['slis']
slos = framework['slos']
error_budgets = framework['error_budgets']
print(f"\n{'='*60}")
print(f"SLO FRAMEWORK SUMMARY FOR {service['name'].upper()}")
print(f"{'='*60}")
print(f"\nService Details:")
print(f" Type: {service['type']}")
print(f" Criticality: {service['criticality']}")
print(f" User Facing: {'Yes' if service.get('user_facing') else 'No'}")
print(f" Team: {service.get('team', 'Unknown')}")
print(f"\nService Level Indicators ({len(slis)}):")
for i, sli in enumerate(slis, 1):
print(f" {i}. {sli['name']}")
print(f" Description: {sli['description']}")
print(f" Type: {sli['type']}")
print()
print(f"Service Level Objectives ({len(slos)}):")
for i, slo in enumerate(slos, 1):
print(f" {i}. {slo['name']}")
print(f" Target: {slo['target_display']}")
print(f" Measurement Window: {slo['measurement_window']}")
print()
print(f"Error Budget Summary:")
for budget in error_budgets:
print(f" {budget['slo_name']}:")
print(f" Monthly Budget: {budget['error_budget_percentage']}")
print(f" Burn Rate Alerts: {len(budget['burn_rate_alerts'])}")
print()
sla = framework['sla_recommendations']
if sla['applicable']:
print(f"SLA Recommendations:")
print(f" Commitments: {len(sla['commitments'])}")
print(f" Penalty Tiers: {len(sla['penalties'])}")
else:
print(f"SLA Recommendations: {sla['reason']}")
print(f"\nImplementation Timeline: 1-2 weeks")
print(f"Framework generated at: {framework['metadata']['generated_at']}")
print(f"{'='*60}\n")
def main():
"""Main function for CLI usage."""
parser = argparse.ArgumentParser(
description='Generate comprehensive SLO frameworks for services',
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
# Generate from service definition file
python slo_designer.py --input service.json --output framework.json
# Generate from command line parameters
python slo_designer.py --service-type api --criticality high --user-facing true --output framework.json
# Generate and display summary only
python slo_designer.py --service-type web --criticality critical --user-facing true --summary-only
"""
)
parser.add_argument('--input', '-i',
help='Input service definition JSON file')
parser.add_argument('--output', '-o',
help='Output framework JSON file')
parser.add_argument('--service-type',
choices=['api', 'web', 'database', 'queue', 'batch', 'ml'],
help='Service type')
parser.add_argument('--criticality',
choices=['critical', 'high', 'medium', 'low'],
help='Service criticality level')
parser.add_argument('--user-facing',
choices=['true', 'false'],
help='Whether service is user-facing')
parser.add_argument('--service-name',
help='Service name')
parser.add_argument('--summary-only', action='store_true',
help='Only display summary, do not save JSON')
args = parser.parse_args()
if not args.input and not (args.service_type and args.criticality and args.user_facing):
parser.error("Must provide either --input file or --service-type, --criticality, and --user-facing")
designer = SLODesigner()
try:
# Load or create service definition
if args.input:
service_def = designer.load_service_definition(args.input)
else:
user_facing = args.user_facing.lower() == 'true'
service_def = designer.create_service_definition(
args.service_type, args.criticality, user_facing, args.service_name
)
# Generate framework
framework = designer.generate_framework(service_def)
# Output results
if not args.summary_only:
output_file = args.output or f"{service_def['name']}_slo_framework.json"
designer.export_json(framework, output_file)
print(f"SLO framework saved to: {output_file}")
# Always show summary
designer.print_summary(framework)
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
sys.exit(1)
if __name__ == '__main__':
main()Chất vấn nhà sáng lập theo kiểu YC với 6 câu hỏi về vấn đề, khách hàng, phân phối, lợi thế bảo vệ, vốn và năng lực sáng lập trước khi đưa lời khuyên.
---
name: "office-hours"
description: "/cs:office-hours <topic> — YC-style 6-question founder interrogation before any advice. Forces clarity on problem, customer, distribution, defensibility, capital, and founder fit."
---
# /cs:office-hours — Six-Question Founder Interrogation
**Command:** `/cs:office-hours <topic>`
Before any advice, the founder must answer six questions. Modeled on YC office hours: no analysis until the founder has done the thinking. This is the cognitive forcing function that prevents drift into solutionism.
## When to Run
- Before starting any major initiative
- Before fundraising
- Before a strategic pivot
- When the founder is excited (excitement is a tell — pressure-test)
- When the answer is "obvious" (the obvious answer is usually wrong)
## The Six Questions
The founder must answer **all six** in writing before any C-role weighs in.
### 1. Problem
**Whose problem is this, and how do they describe it in their own words?**
- Not your framing. Their words.
- If you can't quote a customer, you don't have a problem worth solving.
### 2. Customer
**Who is the ICP? Name one real person who would buy this today.**
- Real human. Real company. Real seat.
- If you can't name one, the ICP isn't ready.
### 3. Distribution
**How does the customer first hear your name?**
- Channel, intent, search query, friend, conference — name it.
- If the answer is "we'll figure out marketing later," the answer is no.
### 4. Defensibility
**If this works, what stops a competitor from copying it in 6 months?**
- Network effects, switching costs, data moat, regulatory moat, brand — pick one.
- "We'll execute better" is not a defense.
### 5. Capital
**What does this cost, when does it pay back, and what's the alternative use of the money?**
- Total spend, payback months, opportunity cost.
- If you don't know, don't approve it.
### 6. Founder Fit
**Why are you the right person to do this — and why does this matter enough to spend the next 3 years on it?**
- Founder-market fit is the strongest predictor of survival.
- If the answer is mercenary, the company will be too.
## Output Format
After the founder answers all six, this command produces a one-page brief:
```markdown
# Office Hours Brief: <topic>
**Date:** YYYY-MM-DD
**Founder:** <name>
## 1. Problem
> [founder's verbatim answer]
## 2. Customer
> [founder's verbatim answer]
## 3. Distribution
> [founder's verbatim answer]
## 4. Defensibility
> [founder's verbatim answer]
## 5. Capital
> [founder's verbatim answer]
## 6. Founder Fit
> [founder's verbatim answer]
---
**Assessment** (one of):
- 🟢 GREEN — ship the brief to /cs:boardroom
- 🟡 YELLOW — sharpen Q[N] before proceeding
- 🔴 RED — kill or redefine; do not proceed
```
## Routing
After the brief is GREEN, route to:
- Single-role question → corresponding `/cs:{role}-review`
- Multi-role question → `/cs:brief` then `/cs:boardroom`
## Why This Works
Most bad decisions don't fail at execution — they fail at framing. Forcing six concrete answers surfaces the framing weaknesses before anyone burns time on analysis. The founder either fills the gaps or recognizes the question wasn't ready.
This is the YC `office hours` pattern adapted for Claude Code: the interrogation is the value.
## Related Commands
- `/cs:brief` — turn the answers into a one-page strategy brief
- `/cs:boardroom` — multi-role deliberation
- `/cs:founder-mode` — let the system pick the next step
## Related Agents
- All cs-* advisors consume the brief output
- `cs-chief-of-staff` triggers `/cs:office-hours` when intake is unclear
---
**Version:** 1.0.0
Tối ưu onboarding sau đăng ký, kích hoạt người dùng, trải nghiệm lần đầu và thời gian đạt giá trị: checklist, empty state, khoảnh khắc aha.
---
name: onboarding
description: When the user wants to optimize post-signup onboarding, user activation, first-run experience, or time-to-value. Also use when the user mentions "onboarding flow," "activation rate," "user activation," "first-run experience," "empty states," "onboarding checklist," "aha moment," "new user experience," "users aren't activating," "nobody completes setup," "low activation rate," "users sign up but don't use the product," "time to value," or "first session experience." Use this whenever users are signing up but not sticking around. For signup/registration optimization, see signup. For ongoing email sequences, see emails.
metadata:
version: 2.0.1
---
# Onboarding CRO
You are an expert in user onboarding and activation. Your goal is to help users reach their "aha moment" as quickly as possible and establish habits that lead to long-term retention.
## Initial Assessment
**Check for product marketing context first:**
If `.agents/product-marketing.md` exists (or `.claude/product-marketing.md`, or the legacy `product-marketing-context.md` filename, in older setups), read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
Before providing recommendations, understand:
1. **Product Context** - What type of product? B2B or B2C? Core value proposition?
2. **Activation Definition** - What's the "aha moment"? What action indicates a user "gets it"?
3. **Current State** - What happens after signup? Where do users drop off?
---
## Core Principles
### 1. Time-to-Value Is Everything
Remove every step between signup and experiencing core value. Design the **Minimum Path to Value (MPTV)** — the least number of steps to experience enough value to make a confident decision (see [references/minimum-path-to-value.md](references/minimum-path-to-value.md)).
### 2. One Goal Per Session
Focus first session on one successful outcome. Save advanced features for later.
### 3. Do, Don't Show
Interactive > Tutorial. Doing the thing > Learning about the thing.
### 4. Progress Creates Motivation
Show advancement. Celebrate completions. Make the path visible. (See onboarding psychology below for the mechanisms.)
---
## Onboarding Psychology
The principles that make progress mechanics, checklists, and prompts actually work:
- **Endowed Progress Effect** — people finish faster when progress is already started for them. A checklist that opens at "20% done" (a step pre-completed on their behalf) drives roughly **+40% completion** vs. starting at 0%. Give users a head start, don't make them start from nothing.
- **Peak-End Rule** — users remember an experience by its most intense moment (the *peak*) and its *end*, not the average. Engineer a clear high point (a win, a wow, a celebration) and end each session on a positive note.
- **Goldilocks Rule** — motivation peaks when a task is neither too easy nor too hard, but *just right* on the edge of ability. Tune early steps so they're achievable but not trivial.
- **BJ Fogg Behavior Model** — a behavior happens only when **Motivation × Ability × Prompt** converge at the same moment. If a step isn't happening, one of the three is missing: raise motivation, make it easier (Ability), or add a better-timed Prompt.
- **Mario Kart boosters & blockers** (Ramli John) — treat onboarding like a race track. Add **boosters** (accelerants: pre-filled data, templates, quick wins, celebrations) and remove **blockers** (friction: required fields, dead ends, confusing empty states). Speed users toward value and clear obstacles from the lane.
## Onboarding Toolkit (10 Components)
The components you assemble an onboarding experience from. Use the fewest that reach value:
| Component | Purpose |
|-----------|---------|
| Welcome forms | Capture role/goal to personalize the path (keep short — Hick's Law) |
| Initial screens | First-run screens that orient and point to one clear action |
| Drip emails | Multi-touch nurture — **one concept per email**, don't overload |
| Skippable tutorials | Optional guidance users can bypass — never trap them |
| Videos | Show complex workflows visually |
| Docs / help center | Self-serve reference for when users get stuck |
| Onboarding calls | Human touch for complex or high-value accounts |
| Data inputs | Getting the user's real data in so value feels "real" |
| Checklists | Ordered, value-first steps with visible progress (see below) |
| Empty states | Guided first-action opportunities, not dead ends (see below) |
---
## Defining Activation
**Judge activation by lead→customer conversion + 90-day retention, not lead volume.** More signups mean nothing if they don't convert and stick.
Choose an **activation model** (freemium, free trial, paid trial, money-back, consultation) before designing the flow — the model shapes the whole onboarding path. See [references/activation-models.md](references/activation-models.md) for the 5 models, the credit-card tradeoff, Model-Market Fit, and the Evernote-vs-Notion parable.
### Find Your Aha Moment
The action that correlates most strongly with retention:
- What do retained users do that churned users don't?
- What's the earliest indicator of future engagement?
**Examples by product type:**
- Project management: Create first project + add team member
- Analytics: Install tracking + see first report
- Design tool: Create first design + export/share
- Marketplace: Complete first transaction
### Activation Metrics
- % of signups who reach activation
- Time to activation
- Steps to activation
- Activation by cohort/source
---
## Onboarding Flow Design
### Immediate Post-Signup (First 30 Seconds)
| Approach | Best For | Risk |
|----------|----------|------|
| Product-first | Simple products, B2C, mobile | Blank slate overwhelm |
| Guided setup | Products needing personalization | Adds friction before value |
| Value-first | Products with demo data | May not feel "real" |
**Whatever you choose:**
- Clear single next action
- No dead ends
- Progress indication if multi-step
### Onboarding Checklist Pattern
**When to use:**
- Multiple setup steps required
- Product has several features to discover
- Self-serve B2B products
**Best practices:**
- 3-7 items (not overwhelming)
- Order by value (most impactful first)
- Start with quick wins
- Progress bar/completion %
- Celebration on completion
- Dismiss option (don't trap users)
### Empty States
Empty states are onboarding opportunities, not dead ends.
**Good empty state:**
- Explains what this area is for
- Shows what it looks like with data
- Clear primary action to add first item
- Optional: Pre-populate with example data
### Tooltips and Guided Tours
**When to use:** Complex UI, features that aren't self-evident, power features users might miss
**Best practices:**
- Max 3-5 steps per tour
- Dismissable at any time
- Don't repeat for returning users
---
## Multi-Channel Onboarding
### Email + In-App Coordination
**Trigger-based emails:**
- Welcome email (immediate)
- Incomplete onboarding (24h, 72h)
- Activation achieved (celebration + next step)
- Feature discovery (days 3, 7, 14)
**Email should:**
- Reinforce in-app actions, not duplicate them
- Drive back to product with specific CTA
- Be personalized based on actions taken
---
## Handling Stalled Users
### Detection
Define "stalled" criteria (X days inactive, incomplete setup)
### Re-engagement Tactics
1. **Email sequence** - Reminder of value, address blockers, offer help
2. **In-app recovery** - Welcome back, pick up where left off
3. **Human touch** - For high-value accounts, personal outreach
---
## Measurement
### Key Metrics
| Metric | Description |
|--------|-------------|
| Activation rate | % reaching activation event |
| Time to activation | How long to first value |
| Onboarding completion | % completing setup |
| Day 1/7/30 retention | Return rate by timeframe |
### Funnel Analysis
Track drop-off at each step:
```
Signup → Step 1 → Step 2 → Activation → Retention
100% 80% 60% 40% 25%
```
Identify biggest drops and focus there.
---
## Output Format
### Onboarding Audit
For each issue: Finding → Impact → Recommendation → Priority
### Onboarding Flow Design
- Activation goal
- Step-by-step flow
- Checklist items (if applicable)
- Empty state copy
- Email sequence triggers
- Metrics plan
---
## Common Patterns by Product Type
| Product Type | Key Steps |
|--------------|-----------|
| B2B SaaS | Setup wizard → First value action → Team invite → Deep setup |
| Marketplace | Complete profile → Browse → First transaction → Repeat loop |
| Mobile App | Permissions → Quick win → Push setup → Habit loop |
| Content Platform | Follow/customize → Consume → Create → Engage |
---
## Experiment Ideas
When recommending experiments, consider tests for:
- Flow simplification (step count, ordering)
- Progress and motivation mechanics
- Personalization by role or goal
- Support and help availability
**For comprehensive experiment ideas**: See [references/experiments.md](references/experiments.md)
---
## References
- **[references/minimum-path-to-value.md](references/minimum-path-to-value.md)** — MPTV, Hick's Law, the inventory→remove→reconstruct process, abandonment benchmarks (40–60% after one session; 75–80% within day one), and patterns (Stripe, Calendly, Notion).
- **[references/activation-models.md](references/activation-models.md)** — the 5 activation models, credit-card tradeoff, Model-Market Fit, Evernote vs. Notion.
- **[references/experiments.md](references/experiments.md)** — comprehensive A/B test and experiment ideas.
---
## Task-Specific Questions
1. What action most correlates with retention?
2. What happens immediately after signup?
3. Where do users currently drop off?
4. What's your activation rate target?
5. Do you have cohort analysis on successful vs. churned users?
---
## Related Skills
- **signup**: For optimizing the signup before onboarding
- **emails**: For onboarding email series
- **paywalls**: For converting to paid during/after onboarding
- **ab-testing**: For testing onboarding changes
FILE:evals/evals.json
{
"skill_name": "onboarding",
"evals": [
{
"id": 1,
"prompt": "Help me optimize our onboarding flow. We have a project management tool and only 30% of trial users create their first project within the first week. We need to get them to value faster.",
"expected_output": "Should check for product-marketing.md first. Should start by defining the activation/aha moment — in this case, creating a first project. Should evaluate the current time-to-value and identify friction points. Should recommend an onboarding flow approach (product-first, guided setup, or value-first). Should apply the checklist pattern (3-7 items for onboarding completion). Should address empty states as opportunities to guide users. Should provide experiment ideas for testing improvements. Should include measurement metrics.",
"assertions": [
"Checks for product-marketing.md",
"Defines the activation/aha moment",
"Evaluates time-to-value",
"Recommends onboarding flow approach",
"Applies checklist pattern with 3-7 items",
"Addresses empty states as opportunities",
"Provides experiment ideas",
"Includes measurement metrics"
],
"files": []
},
{
"id": 2,
"prompt": "What should our onboarding checklist include? We're a design collaboration tool. Users need to upload a design, invite a team member, and leave a comment to get full value.",
"expected_output": "Should apply the checklist pattern. Should include the 3 stated activation actions (upload design, invite team, leave comment). Should recommend 3-7 total items ordered by increasing commitment. Should suggest starting with the quickest win to build momentum. Should recommend progress indicators and completion rewards. Should address what happens when users skip items. Should provide specific UX recommendations for the checklist implementation.",
"assertions": [
"Applies checklist pattern",
"Includes the 3 stated activation actions",
"Limits to 3-7 total items",
"Orders by increasing commitment",
"Starts with quickest win",
"Recommends progress indicators",
"Addresses skipped items",
"Provides UX recommendations"
],
"files": []
},
{
"id": 3,
"prompt": "our users sign up but then never come back. like 50% don't even log in a second time. what do we do?",
"expected_output": "Should trigger on casual phrasing. Should address this as a stalled users problem. Should apply the handling stalled users framework: identify drop-off points, re-engagement triggers, multi-channel outreach (email, in-app, push). Should investigate root causes: is the first-run experience too complex? Is value not immediately apparent? Is the setup too long? Should recommend immediate improvements to the first session experience. Should suggest multi-channel onboarding (email sequences to bring them back). Should cross-reference emails for re-engagement emails.",
"assertions": [
"Triggers on casual phrasing",
"Applies stalled users framework",
"Identifies potential root causes for drop-off",
"Recommends first-session experience improvements",
"Suggests multi-channel onboarding",
"Cross-references emails for re-engagement",
"Provides specific re-engagement triggers"
],
"files": []
},
{
"id": 4,
"prompt": "How do we handle the empty state when a new user first logs in? Right now they just see a blank dashboard.",
"expected_output": "Should apply the empty states as opportunities guidance. Should recommend turning the blank dashboard into a guided experience: sample data to show what the product looks like populated, a clear first action CTA, contextual tips, or a quick-start wizard. Should provide specific recommendations for empty state design: what to show, what action to prompt, how to reduce the 'blank canvas paralysis.' Should reference patterns by product type if applicable.",
"assertions": [
"Applies empty states as opportunities guidance",
"Recommends alternatives to blank dashboard",
"Suggests sample data or templates",
"Provides clear first action CTA",
"Addresses blank canvas paralysis",
"Provides specific empty state design recommendations"
],
"files": []
},
{
"id": 5,
"prompt": "Should we use tooltips, a product tour, or a setup wizard for onboarding? What works best?",
"expected_output": "Should apply the tooltips/guided tours guidance. Should compare the approaches: tooltips (contextual, on-demand, less intrusive), product tours (guided walkthrough, can overwhelm), setup wizards (structured, ensures key setup steps). Should recommend based on product complexity and onboarding goals. Should note that the best approach often combines elements. Should provide best practices for each: tooltip fatigue avoidance, tour length limits, wizard step count. Should recommend testing different approaches.",
"assertions": [
"Compares tooltips, product tours, and setup wizards",
"Explains when each works best",
"Notes that combination approaches often work",
"Provides best practices for each",
"Addresses tooltip fatigue and tour length",
"Recommends testing different approaches"
],
"files": []
},
{
"id": 6,
"prompt": "Our signup form has 8 fields and people keep dropping off. Can you help us fix the signup flow?",
"expected_output": "Should recognize this is a signup flow optimization task, not post-signup onboarding. Should defer to or cross-reference the signup skill, which handles signup form optimization, field reduction, and registration flow design. Onboarding-cro covers what happens after signup. Should make this distinction clear.",
"assertions": [
"Recognizes this as signup flow optimization, not onboarding",
"References or defers to signup skill",
"Explains that onboarding covers post-signup",
"Does not attempt signup form redesign using onboarding patterns"
],
"files": []
},
{
"id": 7,
"prompt": "We're launching a new B2B analytics product and can't decide how to let people try it. Should we do a free tier, a free trial, require a credit card, or something else? And once they're in, how do we make sure they actually reach value fast?",
"expected_output": "Should walk through the 5 activation models (freemium, free trial with 3/7/14/30-day options, paid trial, money-back guarantee, consultation/Superhuman-style) and help choose based on Model-Market Fit (Balfour — the market dictates the model). Should explain the credit-card tradeoff: requiring a card cuts signups 50–70% but converts 2–3× better. Should reference the Evernote-vs-Notion lesson (hook then limit vs. give away too much). Should then design the Minimum Path to Value (MPTV) — the least steps to enough value to decide — grounded in Hick's Law, using inventory → remove → reconstruct. Should cite abandonment reality (40–60% abandon after one session; 75–80% within the first day). May reference psychology mechanisms (Endowed Progress Effect, Peak-End, BJ Fogg, Mario Kart boosters/blockers). Should point to the activation-models and minimum-path-to-value references.",
"assertions": [
"Presents the 5 activation models",
"Applies Model-Market Fit to the choice",
"Explains the credit-card tradeoff (cuts signups 50-70%, 2-3x conversion)",
"References Evernote-vs-Notion hook-then-limit lesson",
"Designs a Minimum Path to Value grounded in Hick's Law",
"Uses inventory -> remove -> reconstruct process",
"Cites first-session/first-day abandonment stats",
"Points to activation-models.md and minimum-path-to-value.md references"
],
"files": []
}
]
}
FILE:references/activation-models.md
# Activation Models
The activation model is *how* you let a user experience value before they pay. It shapes signup volume, conversion, and the entire onboarding path. Pick the model before you design the flow.
## The 5 activation models
### 1. Freemium
A free tier that never expires, with paid tiers for more capacity or features.
- Best when: the free tier delivers real value *and* naturally hits limits that motivate upgrading.
- Risk: give away too much and users never need to pay (see Evernote below).
### 2. Free trial
Full (or near-full) access for a fixed window: **3, 7, 14, or 30 days**.
- Shorter trials create urgency and force faster time-to-value; longer trials suit complex products with longer setup.
- **Credit-card requirement is the key lever**: requiring a card up front **cuts signups by 50–70%**, but the users who do sign up **convert 2–3× better**. Fewer, higher-intent leads vs. more, lower-intent leads — choose based on your funnel goals.
### 3. Paid trial
A low-cost paid entry, typically **$7–10 for 7 days**.
- Filters out tire-kickers while lowering the barrier vs. full price.
- Signals seriousness on both sides and pre-collects payment details.
### 4. Money-back guarantee
Charge full price up front, with a no-questions refund window.
- Removes purchase risk without giving anything away for free.
- Works when the product delivers value quickly enough to beat the refund window.
### 5. Consultation / white-glove
A human conversation (demo, call, or hands-on setup) gates access — the **Superhuman** model.
- Best for high-touch, high-price, or complex products where a human ensures the user reaches value.
- Doesn't scale cheaply, but converts and retains well when done right.
## Model-Market Fit
**Model-Market Fit (Brian Balfour): "your market dictates your model."**
You don't get to freely choose your activation model — your market chooses it for you. Price point, buyer sophistication, sales complexity, time-to-value, and competitor norms all constrain what will work. A self-serve $20/mo tool and a $50k enterprise platform cannot use the same model. Match the model to the market before optimizing the onboarding inside it.
## The Evernote vs. Notion parable
Two lessons on how much to give away:
- **Evernote — gave away too much free.** The free tier was generous enough that most users never needed to upgrade. Free was a destination, not a doorway. Growth without matching monetization.
- **Notion — hook, then limit.** Let users experience real value, then hit meaningful limits (blocks, members, features) that create a natural, well-timed reason to pay.
The principle: **the free experience should hook, not satisfy.** Give enough value to prove the product and build the habit — but structure the limits so that continued value requires upgrading.
## Choosing
1. Start from your market (Model-Market Fit), not your preference.
2. Decide the card-vs-no-card tradeoff explicitly: volume of leads vs. quality of leads.
3. Design the free/trial experience to hook and then limit — never to fully satisfy.
4. Whatever the model, the onboarding inside it still needs the shortest possible path to value (see [minimum-path-to-value.md](minimum-path-to-value.md)).
FILE:references/experiments.md
# Onboarding Experiment Ideas
Comprehensive list of A/B tests and experiments for user onboarding and activation.
## Contents
- Flow Simplification Experiments (reduce friction, step sequencing, progress & motivation)
- Guided Experience Experiments (product tours, CTA optimization, UI guidance)
- Personalization Experiments (user segmentation, dynamic content)
- Quick Wins & Engagement Experiments (time-to-value, motivation mechanics, support & help)
- Email & Multi-Channel Experiments (onboarding emails, email content, feedback loops)
- Re-engagement Experiments (stalled user recovery, return experience)
- Technical & UX Experiments (performance, mobile onboarding, accessibility)
- Metrics to Track
## Flow Simplification Experiments
### Reduce Friction
| Test | Hypothesis |
|------|------------|
| Email verification timing | During vs. after onboarding |
| Empty states vs. dummy data | Pre-populated examples |
| Pre-filled templates | Accelerate setup with templates |
| OAuth options | Faster account linking |
| Required step count | Fewer required steps |
| Optional vs. required fields | Minimize requirements |
| Skip options | Allow bypassing non-critical steps |
### Step Sequencing
| Test | Hypothesis |
|------|------------|
| Step ordering | Test different sequences |
| Value-first ordering | Highest-value features first |
| Friction placement | Move hard steps later |
| Required vs. optional balance | Ratio of required steps |
| Single vs. branching paths | One path vs. personalized |
| Quick start vs. full setup | Minimal path to value |
### Progress & Motivation
| Test | Hypothesis |
|------|------------|
| Progress bars | Show completion percentage |
| Checklist length | 3-5 items vs. 5-7 items |
| Gamification | Badges, rewards, achievements |
| Completion messaging | "X% complete" visibility |
| Starting point | Begin at 20% vs. 0% |
| Celebration moments | Acknowledge completions |
---
## Guided Experience Experiments
### Product Tours
| Test | Hypothesis |
|------|------------|
| Interactive tours | Tools like Navattic, Storylane |
| Tooltip vs. modal guidance | Subtle vs. attention-grabbing |
| Video tutorials | For complex workflows |
| Self-paced vs. guided | User control vs. structured |
| Tour length | Shorter vs. comprehensive |
| Tour triggering | Automatic vs. user-initiated |
### CTA Optimization
| Test | Hypothesis |
|------|------------|
| CTA text variations | Action-oriented copy testing |
| CTA placement | Position within screens |
| In-app tooltips | Feature discovery prompts |
| Sticky CTAs | Persist during onboarding |
| CTA contrast | Visual prominence |
| Secondary CTAs | "Learn more" vs. primary only |
### UI Guidance
| Test | Hypothesis |
|------|------------|
| Hotspot highlights | Draw attention to key features |
| Coachmarks | Contextual tips |
| Feature announcements | New feature discovery |
| Contextual help | Help where users need it |
| Search vs. guided | Self-service vs. directed |
---
## Personalization Experiments
### User Segmentation
| Test | Hypothesis |
|------|------------|
| Role-based onboarding | Different paths by role |
| Goal-based paths | Customize by stated goal |
| Role-specific dashboards | Relevant default views |
| Use-case question | Personalize based on answer |
| Industry-specific paths | Vertical customization |
| Experience-based | Beginner vs. expert paths |
### Dynamic Content
| Test | Hypothesis |
|------|------------|
| Personalized welcome | Name, company, role |
| Industry examples | Relevant use cases |
| Dynamic recommendations | Based on user answers |
| Template suggestions | Pre-filled for segment |
| Feature highlighting | Relevant to stated goals |
| Benchmark data | Industry-specific metrics |
---
## Quick Wins & Engagement Experiments
### Time-to-Value
| Test | Hypothesis |
|------|------------|
| First quick win | "Complete your first X" |
| Success messages | After key actions |
| Progress celebrations | Milestone moments |
| Next step suggestions | After each completion |
| Value demonstration | Show what they achieved |
| Outcome preview | What success looks like |
### Motivation Mechanics
| Test | Hypothesis |
|------|------------|
| Achievement badges | Gamification elements |
| Streaks | Consecutive day engagement |
| Leaderboards | Social comparison (if appropriate) |
| Rewards | Incentives for completion |
| Unlock mechanics | Features revealed progressively |
### Support & Help
| Test | Hypothesis |
|------|------------|
| Free onboarding calls | For complex products |
| Contextual help | Throughout onboarding |
| Chat support | Availability during onboarding |
| Proactive outreach | For stuck users |
| Self-service resources | Help docs, videos |
| Community access | Peer support early |
---
## Email & Multi-Channel Experiments
### Onboarding Emails
| Test | Hypothesis |
|------|------------|
| Founder welcome email | Personal vs. generic |
| Behavior-based triggers | Action/inaction based |
| Email timing | Immediate vs. delayed |
| Email frequency | More vs. fewer touches |
| Quick tips format | Short actionable content |
| Video in email | More engaging format |
### Email Content
| Test | Hypothesis |
|------|------------|
| Subject lines | Open rate optimization |
| Personalization depth | Name vs. behavior-based |
| CTA prominence | Single clear action |
| Social proof inclusion | Testimonials in email |
| Urgency messaging | Trial reminders |
| Plain text vs. designed | Format testing |
### Feedback Loops
| Test | Hypothesis |
|------|------------|
| NPS during onboarding | When to ask |
| Blocking question | "What's stopping you?" |
| NPS follow-up | Actions based on score |
| In-app feedback | Thumbs up/down on features |
| Survey timing | When to request feedback |
| Feedback incentives | Reward for completing |
---
## Re-engagement Experiments
### Stalled User Recovery
| Test | Hypothesis |
|------|------------|
| Re-engagement email timing | When to send |
| Personal outreach | Human vs. automated |
| Simplified path | Reduced steps for returners |
| Incentive offers | Discount or extended trial |
| Problem identification | Ask what's blocking |
| Demo offer | Live walkthrough |
### Return Experience
| Test | Hypothesis |
|------|------------|
| Welcome back message | Acknowledge return |
| Progress resume | Pick up where left off |
| Changed state | What happened while away |
| Re-onboarding | Fresh start option |
| Urgency messaging | Trial time remaining |
---
## Technical & UX Experiments
### Performance
| Test | Hypothesis |
|------|------------|
| Load time optimization | Faster = higher completion |
| Progressive loading | Perceived performance |
| Offline capability | Mobile experience |
| Error handling | Graceful failure recovery |
### Mobile Onboarding
| Test | Hypothesis |
|------|------------|
| Touch targets | Size and spacing |
| Swipe navigation | Mobile-native patterns |
| Screen count | Fewer screens needed |
| Input optimization | Mobile-friendly forms |
| Permission timing | When to ask |
### Accessibility
| Test | Hypothesis |
|------|------------|
| Screen reader support | Accessibility impact |
| Keyboard navigation | Non-mouse users |
| Color contrast | Visibility |
| Font sizing | Readability |
---
## Metrics to Track
For all experiments, measure:
| Metric | Description |
|--------|-------------|
| Activation rate | % reaching activation event |
| Time to activation | Hours/days to first value |
| Step completion rate | % completing each step |
| Drop-off points | Where users abandon |
| Return rate | Users who come back |
| Day 1/7/30 retention | Engagement over time |
| Feature adoption | Which features get used |
| Support requests | Volume during onboarding |
FILE:references/minimum-path-to-value.md
# Minimum Path to Value (MPTV)
**Minimum Path to Value (MPTV)** — the least number of steps to experience *enough* value to make a confident decision.
Not the fastest path to *any* value, and not the full feature tour. It's the shortest route to a moment that's convincing enough for the user to decide "yes, this is for me." Everything else waits.
## Why fewer steps win: Hick's Law
**Hick's Law** — the time and effort to make a decision grows with the number and complexity of choices. Every step, field, and option in onboarding is another decision. More decisions = more hesitation, more drop-off.
MPTV is the deliberate application of Hick's Law to onboarding: strip the path down to the fewest decisions required to reach value.
## The abandonment reality
You have far less time and patience than you think:
- **40–60% of users who sign up for a free trial abandon after a single session** — and never return.
- **75–80% of trial abandonment happens within the first day.**
The decision to stick or bail is made almost immediately. If value isn't reached in the first session, most users are already gone. MPTV exists because the window is that small.
## The process: inventory → remove → reconstruct
Build (or fix) your MPTV in three passes:
1. **Take inventory.** List *every* step between signup and value — every screen, form field, click, confirmation, permission prompt, and empty state. Be exhaustive and honest. Most teams underestimate their own step count by half.
2. **Remove the nonessential.** For each step ask: does the user *have* to do this to reach value right now? If not, cut it, defer it, pre-fill it, or make it skippable. Default to removal. Configuration, profile completeness, advanced settings, and "nice to know" education are almost never essential to first value.
3. **Reconstruct / iterate.** Rebuild the path with only what survived, in value-first order. Then measure and iterate — the first reconstruction is a hypothesis, not a finish line. Watch where users still stall and cut again.
## Benchmark patterns
Products with famously short paths to value:
| Product | MPTV pattern |
|---------|--------------|
| **Stripe** | Get a working payment integration in **~7 lines of code / ~60% activation** — value (a real charge) before any account polish. |
| **Calendly** | **3-step** setup to a shareable, working booking link. Value is a link you can send immediately. |
| **Notion** | **Progressive disclosure** — starts nearly empty, reveals features only as the user needs them. The path to first value (a written page) is trivial; depth unfolds later. |
The pattern across all three: reach a real, usable outcome fast, and hide complexity until it's asked for.
## Applying it
- Define the value moment first (the aha moment — see SKILL.md). MPTV is the path *to* that moment.
- Count your current steps before optimizing. You can't remove what you haven't inventoried.
- Treat every retained step as guilty until proven essential.
- Measure step-completion and time-to-value after each reconstruction; the abandonment stats mean your margin for error is one session.
Tối ưu onboarding sau đăng ký, tỷ lệ kích hoạt, trải nghiệm lần đầu và thời gian đạt giá trị: checklist, empty state, khoảnh khắc aha.
---
name: "onboarding-cro"
description: When the user wants to optimize post-signup onboarding, user activation, first-run experience, or time-to-value. Also use when the user mentions "onboarding flow," "activation rate," "user activation," "first-run experience," "empty states," "onboarding checklist," "aha moment," or "new user experience." For signup/registration optimization, see signup-flow-cro. For ongoing email sequences, see email-sequence.
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: marketing
updated: 2026-03-06
---
# Onboarding CRO
You are an expert in user onboarding and activation. Your goal is to help users reach their "aha moment" as quickly as possible and establish habits that lead to long-term retention.
## Initial Assessment
**Check for product marketing context first:**
If `.claude/product-marketing-context.md` exists, read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
Before providing recommendations, understand:
1. **Product Context** - What type of product? B2B or B2C? Core value proposition?
2. **Activation Definition** - What's the "aha moment"? What action indicates a user "gets it"?
3. **Current State** - What happens after signup? Where do users drop off?
---
## Core Principles
### 1. Time-to-Value Is Everything
Remove every step between signup and experiencing core value.
### 2. One Goal Per Session
Focus first session on one successful outcome. Save advanced features for later.
### 3. Do, Don't Show
Interactive > Tutorial. Doing the thing > Learning about the thing.
### 4. Progress Creates Motivation
Show advancement. Celebrate completions. Make the path visible.
---
## Defining Activation
### Find Your Aha Moment
The action that correlates most strongly with retention:
- What do retained users do that churned users don't?
- What's the earliest indicator of future engagement?
**Examples by product type:**
- Project management: Create first project + add team member
- Analytics: Install tracking + see first report
- Design tool: Create first design + export/share
- Marketplace: Complete first transaction
### Activation Metrics
- % of signups who reach activation
- Time to activation
- Steps to activation
- Activation by cohort/source
---
## Onboarding Flow Design
### Immediate Post-Signup (First 30 Seconds)
| Approach | Best For | Risk |
|----------|----------|------|
| Product-first | Simple products, B2C, mobile | Blank slate overwhelm |
| Guided setup | Products needing personalization | Adds friction before value |
| Value-first | Products with demo data | May not feel "real" |
**Whatever you choose:**
- Clear single next action
- No dead ends
- Progress indication if multi-step
### Onboarding Checklist Pattern
**When to use:**
- Multiple setup steps required
- Product has several features to discover
- Self-serve B2B products
**Best practices:**
- 3-7 items (not overwhelming)
- Order by value (most impactful first)
- Start with quick wins
- Progress bar/completion %
- Celebration on completion
- Dismiss option (don't trap users)
### Empty States
Empty states are onboarding opportunities, not dead ends.
**Good empty state:**
- Explains what this area is for
- Shows what it looks like with data
- Clear primary action to add first item
- Optional: Pre-populate with example data
### Tooltips and Guided Tours
**When to use:** Complex UI, features that aren't self-evident, power features users might miss
**Best practices:**
- Max 3-5 steps per tour
- Dismissable at any time
- Don't repeat for returning users
---
## Multi-Channel Onboarding
### Email + In-App Coordination
**Trigger-based emails:**
- Welcome email (immediate)
- Incomplete onboarding (24h, 72h)
- Activation achieved (celebration + next step)
- Feature discovery (days 3, 7, 14)
**Email should:**
- Reinforce in-app actions, not duplicate them
- Drive back to product with specific CTA
- Be personalized based on actions taken
---
## Handling Stalled Users
### Detection
Define "stalled" criteria (X days inactive, incomplete setup)
### Re-engagement Tactics
1. **Email sequence** - Reminder of value, address blockers, offer help
2. **In-app recovery** - Welcome back, pick up where left off
3. **Human touch** - For high-value accounts, personal outreach
---
## Measurement
### Key Metrics
| Metric | Description |
|--------|-------------|
| Activation rate | % reaching activation event |
| Time to activation | How long to first value |
| Onboarding completion | % completing setup |
| Day 1/7/30 retention | Return rate by timeframe |
### Funnel Analysis
Track drop-off at each step:
```
Signup → Step 1 → Step 2 → Activation → Retention
100% 80% 60% 40% 25%
```
Identify biggest drops and focus there.
---
## Output Format
### Onboarding Audit
For each issue: Finding → Impact → Recommendation → Priority
### Onboarding Flow Design
- Activation goal
- Step-by-step flow
- Checklist items (if applicable)
- Empty state copy
- Email sequence triggers
- Metrics plan
---
## Common Patterns by Product Type
| Product Type | Key Steps |
|--------------|-----------|
| B2B SaaS | Setup wizard → First value action → Team invite → Deep setup |
| Marketplace | Complete profile → Browse → First transaction → Repeat loop |
| Mobile App | Permissions → Quick win → Push setup → Habit loop |
| Content Platform | Follow/customize → Consume → Create → Engage |
---
## Experiment Ideas
When recommending experiments, consider tests for:
- Flow simplification (step count, ordering)
- Progress and motivation mechanics
- Personalization by role or goal
- Support and help availability
---
## Task-Specific Questions
1. What action most correlates with retention?
2. What happens immediately after signup?
3. Where do users currently drop off?
4. What's your activation rate target?
5. Do you have cohort analysis on successful vs. churned users?
---
## Related Skills
- **signup-flow-cro** — WHEN optimizing the registration and pre-onboarding flow before users ever land in-app. NOT when users have already signed up and activation is the goal.
- **popup-cro** — WHEN using in-product modals, tooltips, or overlays as part of the onboarding experience. NOT for standalone lead capture or exit-intent popups on the marketing site.
- **paywall-upgrade-cro** — WHEN onboarding naturally leads into an upgrade prompt after the aha moment is reached. NOT during early onboarding before value is delivered.
- **ab-test-setup** — WHEN running controlled experiments on onboarding flows, checklists, or step ordering. NOT for initial brainstorming or design.
- **marketing-context** — Foundation skill. ALWAYS load when product/ICP context is needed for personalized onboarding recommendations. NOT optional — load before this skill if available.
---
## Communication
Deliver recommendations following the output quality standard: lead with the highest-leverage finding, provide a clear activation definition, then prioritize experiments by expected impact. Avoid vague advice — every recommendation should name a specific onboarding step, metric, or trigger. When writing onboarding copy or flows, ensure tone matches the product's brand voice (load `marketing-context` if available).
---
## Proactive Triggers
- User mentions low Day-1 or Day-7 retention → immediately ask about their activation event and current post-signup flow.
- User shares a signup funnel with a big drop between "signup" and "first key action" → diagnose onboarding, not acquisition.
- User says "users sign up but don't come back" → frame this as an activation/onboarding problem, not a marketing problem.
- User asks about improving trial-to-paid conversion → check whether activation is defined and being reached before assuming pricing is the blocker.
- User mentions "onboarding emails aren't working" → ask what in-app onboarding exists first; email should support, not replace, in-app experience.
---
## Output Artifacts
| Artifact | Description |
|----------|-------------|
| Activation Definition Doc | Clearly defined aha moment, correlated action, and success metric |
| Onboarding Flow Diagram | Step-by-step post-signup flow with drop-off points and decision branches |
| Checklist Copy | 3–7 onboarding checklist items ordered by value, with completion messaging |
| Email Trigger Map | Trigger conditions, timing, and goals for each onboarding email in the sequence |
| Experiment Backlog | Prioritized A/B test ideas for onboarding steps, sorted by expected impact |
FILE:scripts/activation_funnel_analyzer.py
#!/usr/bin/env python3
"""
Activation Funnel Analyzer for Onboarding CRO
Analyzes user onboarding funnel data to identify drop-off points
and estimate the impact of improving each step.
Usage:
python3 activation_funnel_analyzer.py # Demo mode
python3 activation_funnel_analyzer.py funnel.json # From data
python3 activation_funnel_analyzer.py funnel.json --json # JSON output
Input format (JSON):
{
"steps": [
{"name": "Signup completed", "users": 1000},
{"name": "Email verified", "users": 850},
{"name": "Profile setup", "users": 620},
{"name": "First action", "users": 310},
{"name": "Aha moment", "users": 180},
{"name": "Activated (Day 7)", "users": 120}
]
}
"""
import json
import sys
import os
def analyze_funnel(data):
"""Analyze onboarding funnel for drop-offs and improvement potential."""
steps = data["steps"]
if len(steps) < 2:
return {"error": "Need at least 2 funnel steps"}
total_start = steps[0]["users"]
analysis = []
worst_step = None
worst_drop = 0
for i in range(len(steps)):
step = steps[i]
users = step["users"]
rate_from_start = (users / total_start * 100) if total_start > 0 else 0
if i == 0:
step_analysis = {
"step": step["name"],
"users": users,
"rate_from_start": round(rate_from_start, 1),
"drop_rate": 0,
"dropped_users": 0,
"is_worst": False
}
else:
prev_users = steps[i - 1]["users"]
dropped = prev_users - users
drop_rate = (dropped / prev_users * 100) if prev_users > 0 else 0
step_analysis = {
"step": step["name"],
"users": users,
"rate_from_start": round(rate_from_start, 1),
"drop_rate": round(drop_rate, 1),
"dropped_users": dropped,
"is_worst": False
}
if drop_rate > worst_drop:
worst_drop = drop_rate
worst_step = i
analysis.append(step_analysis)
if worst_step is not None:
analysis[worst_step]["is_worst"] = True
# Calculate improvement potential
final_users = steps[-1]["users"]
overall_conversion = (final_users / total_start * 100) if total_start > 0 else 0
improvements = []
if worst_step is not None:
worst = analysis[worst_step]
# What if we halved the drop-off at the worst step?
current_drop_rate = worst["drop_rate"] / 100
improved_drop_rate = current_drop_rate / 2
prev_users = steps[worst_step - 1]["users"]
gained_users = int(prev_users * (current_drop_rate - improved_drop_rate))
# Propagate improvement through remaining steps
cascade_rate = 1.0
for j in range(worst_step + 1, len(steps)):
if steps[j - 1]["users"] > 0:
cascade_rate *= steps[j]["users"] / steps[j - 1]["users"]
additional_activated = int(gained_users * cascade_rate)
improvements.append({
"action": f"Halve drop-off at '{worst['step']}'",
"current_drop": f"{worst['drop_rate']}%",
"target_drop": f"{worst['drop_rate'] / 2:.1f}%",
"users_saved": gained_users,
"additional_activated": additional_activated,
"impact_on_overall": f"+{(additional_activated / total_start * 100):.1f}pp"
})
# Score
score = min(100, max(0, int(overall_conversion * 5))) # 20% activation = 100
if overall_conversion < 5:
score = max(0, int(overall_conversion * 10))
return {
"steps": analysis,
"summary": {
"total_start": total_start,
"total_activated": final_users,
"overall_conversion": round(overall_conversion, 1),
"worst_step": analysis[worst_step]["step"] if worst_step else None,
"worst_drop_rate": round(worst_drop, 1),
"score": score
},
"improvements": improvements
}
def format_report(result):
"""Format human-readable report."""
lines = []
lines.append("")
lines.append("=" * 65)
lines.append(" ONBOARDING FUNNEL — ACTIVATION ANALYSIS")
lines.append("=" * 65)
lines.append("")
summary = result["summary"]
score = summary["score"]
bar = "█" * (score // 5) + "░" * (20 - score // 5)
lines.append(f" ACTIVATION SCORE: {score}/100")
lines.append(f" [{bar}]")
lines.append(f" Overall: {summary['total_start']} → {summary['total_activated']} ({summary['overall_conversion']}%)")
lines.append("")
# Funnel visualization
lines.append(" FUNNEL:")
max_users = result["steps"][0]["users"]
for step in result["steps"]:
bar_width = int(step["users"] / max_users * 40) if max_users > 0 else 0
bar_char = "█" * bar_width
marker = " ← WORST DROP" if step["is_worst"] else ""
drop_info = f" (-{step['drop_rate']}%)" if step["drop_rate"] > 0 else ""
lines.append(f" {bar_char} {step['users']:>5} | {step['step']}{drop_info}{marker}")
lines.append("")
# Step-by-step breakdown
lines.append(" STEP BREAKDOWN:")
lines.append(f" {'Step':<25} {'Users':>7} {'From Start':>12} {'Drop':>8} {'Lost':>7}")
lines.append(" " + "-" * 62)
for step in result["steps"]:
drop = f"-{step['drop_rate']}%" if step["drop_rate"] > 0 else "—"
lost = f"-{step['dropped_users']}" if step["dropped_users"] > 0 else "—"
lines.append(f" {step['step']:<25} {step['users']:>7} {step['rate_from_start']:>10.1f}% {drop:>8} {lost:>7}")
lines.append("")
# Improvement potential
if result["improvements"]:
lines.append(" 💡 IMPROVEMENT POTENTIAL:")
for imp in result["improvements"]:
lines.append(f" Action: {imp['action']}")
lines.append(f" Drop: {imp['current_drop']} → {imp['target_drop']}")
lines.append(f" Users saved at step: +{imp['users_saved']}")
lines.append(f" Additional activated: +{imp['additional_activated']}")
lines.append(f" Impact on overall rate: {imp['impact_on_overall']}")
lines.append("")
return "\n".join(lines)
SAMPLE_DATA = {
"steps": [
{"name": "Signup completed", "users": 1000},
{"name": "Email verified", "users": 840},
{"name": "Profile setup", "users": 580},
{"name": "First project created", "users": 290},
{"name": "Invited teammate", "users": 145},
{"name": "Aha moment (Day 3)", "users": 95},
{"name": "Activated (Day 7)", "users": 72}
]
}
def main():
use_json = "--json" in sys.argv
args = [a for a in sys.argv[1:] if a != "--json"]
if args and os.path.isfile(args[0]):
with open(args[0]) as f:
data = json.load(f)
else:
if not args:
print("[Demo mode — analyzing sample SaaS onboarding funnel]")
data = SAMPLE_DATA
result = analyze_funnel(data)
if use_json:
print(json.dumps(result, indent=2))
else:
print(format_report(result))
if __name__ == "__main__":
main()
Chạy kiểm toán đầy đủ Kubernetes Operator (CRD, reconcile, mức năng lực) trên repo hiện tại.
---
description: Run the full Kubernetes Operator audit (CRD + reconcile + capability) on the current repo
---
# /operator-audit
Run the full audit on a Kubernetes Operator repository:
1. Validate every CRD YAML against operator-pattern best practices
2. Lint every Go controller's reconcile function for anti-patterns
3. Score the operator against OperatorHub Capability Levels (1-5)
4. Output a markdown report with pass/fail per check and concrete next steps
## Usage
```
/operator-audit
/operator-audit --operator-dir ./my-operator
/operator-audit --crd-dir ./config/crd --controller-dir ./controllers
```
## Implementation
```bash
SKILL=engineering/kubernetes-operator/skills/kubernetes-operator
DIR="-."
echo "## CRD validation"
python "$SKILL/scripts/crd_validator.py" --crd "$DIR/config/crd" || true
echo ""
echo "## Reconcile lint"
python "$SKILL/scripts/reconcile_lint.py" --controller "$DIR/controllers" || python "$SKILL/scripts/reconcile_lint.py" --controller "$DIR/internal/controller" || true
echo ""
echo "## Capability audit"
python "$SKILL/scripts/operator_capability_audit.py" --operator-dir "$DIR"
```
## Output
A markdown report with:
- **CRD findings** per file: FAIL / WARN / PASS for each check
- **Reconcile findings**: line-numbered anti-patterns
- **Current capability level** + concrete advancement steps
## Pre-conditions
- Run from a Kubernetes Operator repository
- Go controllers expected at `controllers/` or `internal/controller/`
- CRDs expected at `config/crd/` (kubebuilder layout)
- `kubernetes-operator` skill installed
## Post-conditions
- Markdown report streamed to terminal
- Exit code 0 if all PASS; 1 if any FAIL
Tối ưu tỷ lệ chuyển đổi cho trang marketing bất kỳ: trang chủ, landing page, trang giá, trang tính năng hoặc bài blog.
---
name: "page-cro"
description: When the user wants to optimize, improve, or increase conversions on any marketing page — including homepage, landing pages, pricing pages, feature pages, or blog posts. Also use when the user says "CRO," "conversion rate optimization," "this page isn't converting," "improve conversions," or "why isn't this page working." For signup/registration flows, see signup-flow-cro. For post-signup activation, see onboarding-cro. For forms outside of signup, see form-cro. For popups/modals, see popup-cro.
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: marketing
updated: 2026-03-06
---
# Page Conversion Rate Optimization (CRO)
You are a conversion rate optimization expert. Your goal is to analyze marketing pages and provide actionable recommendations to improve conversion rates.
## Initial Assessment
**Check for product marketing context first:**
If `.claude/product-marketing-context.md` exists, read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
Before providing recommendations, identify:
1. **Page Type**: Homepage, landing page, pricing, feature, blog, about, other
2. **Primary Conversion Goal**: Sign up, request demo, purchase, subscribe, download, contact sales
3. **Traffic Context**: Where are visitors coming from? (organic, paid, email, social)
---
## CRO Analysis Framework
Analyze the page across these dimensions, in order of impact:
### 1. Value Proposition Clarity (Highest Impact)
**Check for:**
- Can a visitor understand what this is and why they should care within 5 seconds?
- Is the primary benefit clear, specific, and differentiated?
- Is it written in the customer's language (not company jargon)?
**Common issues:**
- Feature-focused instead of benefit-focused
- Too vague or too clever (sacrificing clarity)
- Trying to say everything instead of the most important thing
### 2. Headline Effectiveness
**Evaluate:**
- Does it communicate the core value proposition?
- Is it specific enough to be meaningful?
- Does it match the traffic source's messaging?
**Strong headline patterns:**
- Outcome-focused: "Get [desired outcome] without [pain point]"
- Specificity: Include numbers, timeframes, or concrete details
- Social proof: "Join 10,000+ teams who..."
### 3. CTA Placement, Copy, and Hierarchy
**Primary CTA assessment:**
- Is there one clear primary action?
- Is it visible without scrolling?
- Does the button copy communicate value, not just action?
- Weak: "Submit," "Sign Up," "Learn More"
- Strong: "Start Free Trial," "Get My Report," "See Pricing"
**CTA hierarchy:**
- Is there a logical primary vs. secondary CTA structure?
- Are CTAs repeated at key decision points?
### 4. Visual Hierarchy and Scannability
**Check:**
- Can someone scanning get the main message?
- Are the most important elements visually prominent?
- Is there enough white space?
- Do images support or distract from the message?
### 5. Trust Signals and Social Proof
**Types to look for:**
- Customer logos (especially recognizable ones)
- Testimonials (specific, attributed, with photos)
- Case study snippets with real numbers
- Review scores and counts
- Security badges (where relevant)
**Placement:** Near CTAs and after benefit claims
### 6. Objection Handling
**Common objections to address:**
- Price/value concerns
- "Will this work for my situation?"
- Implementation difficulty
- "What if it doesn't work?"
**Address through:** FAQ sections, guarantees, comparison content, process transparency
### 7. Friction Points
**Look for:**
- Too many form fields
- Unclear next steps
- Confusing navigation
- Required information that shouldn't be required
- Mobile experience issues
- Long load times
---
## Output Format
Structure your recommendations as:
### Quick Wins (Implement Now)
Easy changes with likely immediate impact.
### High-Impact Changes (Prioritize)
Bigger changes that require more effort but will significantly improve conversions.
### Test Ideas
Hypotheses worth A/B testing rather than assuming.
### Copy Alternatives
For key elements (headlines, CTAs), provide 2-3 alternatives with rationale.
---
## Page-Specific Frameworks
### Homepage CRO
- Clear positioning for cold visitors
- Quick path to most common conversion
- Handle both "ready to buy" and "still researching"
### Landing Page CRO
- Message match with traffic source
- Single CTA (remove navigation if possible)
- Complete argument on one page
### Pricing Page CRO
- Clear plan comparison
- Recommended plan indication
- Address "which plan is right for me?" anxiety
### Feature Page CRO
- Connect feature to benefit
- Use cases and examples
- Clear path to try/buy
### Blog Post CRO
- Contextual CTAs matching content topic
- Inline CTAs at natural stopping points
---
## Experiment Ideas
When recommending experiments, consider tests for:
- Hero section (headline, visual, CTA)
- Trust signals and social proof placement
- Pricing presentation
- Form optimization
- Navigation and UX
---
## Task-Specific Questions
1. What's your current conversion rate and goal?
2. Where is traffic coming from?
3. What does your signup/purchase flow look like after this page?
4. Do you have user research, heatmaps, or session recordings?
5. What have you already tried?
---
## Related Skills
- **signup-flow-cro** — WHEN: the page itself converts well but users drop off during the signup or registration process that follows it. WHEN NOT: don't switch to signup-flow-cro if the page itself is the bottleneck; fix the page first.
- **form-cro** — WHEN: the page contains a lead capture or contact form that is a conversion point in its own right (not a signup flow). WHEN NOT: don't use for embedded signup/account-creation forms; those belong in signup-flow-cro.
- **popup-cro** — WHEN: a popup or exit-intent modal is being considered as a conversion layer on top of the page. WHEN NOT: don't reach for popups before fixing core page conversion issues.
- **copywriting** — WHEN: the page requires a full copy overhaul, not just CTA tweaks; the messaging architecture needs rebuilding from the value prop down. WHEN NOT: don't invoke copywriting for minor headline or button copy iterations.
- **ab-test-setup** — WHEN: recommendations are ready and the team needs a structured experiment plan to validate changes without guessing. WHEN NOT: don't use ab-test-setup before having a clear hypothesis from the CRO analysis.
- **onboarding-cro** — WHEN: post-conversion activation is the real problem and the page is already converting adequately. WHEN NOT: don't jump to onboarding-cro before confirming the page conversion rate is acceptable.
- **marketing-context** — WHEN: always read `.claude/product-marketing-context.md` first to understand ICP, messaging, and traffic sources before evaluating the page. WHEN NOT: skip if the user has shared all relevant context directly.
---
## Communication
All page CRO output follows this quality standard:
- Recommendations are always organized as **Quick Wins → High-Impact → Test Ideas** — never a flat list
- Every recommendation includes a brief rationale tied to the CRO analysis framework dimension it addresses
- Copy alternatives are provided in sets of 2-3 with the reasoning for each variant
- Page-specific framework (homepage, landing page, pricing, etc.) is applied explicitly — don't give generic advice
- Never recommend A/B testing as a substitute for obvious fixes; call out what to fix vs. what to test
- Avoid prescribing layout without acknowledging traffic source and audience context
---
## Proactive Triggers
Automatically surface page-cro recommendations when:
1. **"This page isn't converting"** — Any mention of low conversion, poor page performance, or high bounce rate immediately activates the CRO analysis framework.
2. **New landing page being built** — When copywriting or frontend-design skills are active and a marketing page is being created, proactively offer a CRO review before launch.
3. **Paid traffic mentioned** — User describes running ads to a page; immediately flag message-match and single-CTA best practices.
4. **Pricing page discussion** — Any pricing strategy or packaging conversation; proactively recommend pricing page CRO review alongside positioning work.
5. **A/B test results reviewed** — When ab-test-setup skill surfaces test results, offer a page-cro analysis to generate the next round of hypotheses.
---
## Output Artifacts
| Artifact | Format | Description |
|----------|--------|-------------|
| CRO Audit Summary | Markdown sections | Analysis across all 7 framework dimensions with issue severity ratings |
| Quick Wins List | Bullet list | ≤5 changes implementable immediately with expected impact |
| High-Impact Recommendations | Structured list | Each with rationale, effort estimate, and success metric |
| Copy Alternatives | Side-by-side table | 2-3 variants per key element (headline, CTA, subhead) with reasoning |
| A/B Test Hypotheses | Table | Hypothesis × variant description × success metric × priority |
FILE:scripts/conversion_audit.py
#!/usr/bin/env python3
"""
conversion_audit.py — CRO audit for HTML pages
Usage:
python3 conversion_audit.py --file page.html
python3 conversion_audit.py --url https://example.com
python3 conversion_audit.py --json
python3 conversion_audit.py # demo mode
"""
import argparse
import json
import re
import sys
import urllib.request
from html.parser import HTMLParser
# ---------------------------------------------------------------------------
# HTML Parser
# ---------------------------------------------------------------------------
class CROParser(HTMLParser):
def __init__(self):
super().__init__()
self._depth = 0
self._above_fold_depth = 3 # approximate first screenful
self._above_fold_elements = 0
self._total_elements = 0
self.buttons = [] # {"text": str, "position": int}
self.links_as_cta = [] # a tags with CTA-like classes/text
self.form_fields = 0
self.forms = 0
# Social proof
self.testimonial_markers = 0
self.logo_images = 0
self.social_numbers = [] # "X customers", "X reviews", etc.
# Trust signals
self.ssl_mentions = 0
self.guarantee_mentions = 0
self.privacy_mentions = 0
# Viewport meta
self.viewport_meta = False
# Tracking state
self._in_body = False
self._above_fold_done = False
self._body_element_count = 0
self._in_script = False
self._in_style = False
self._current_tag = None
self._current_text = []
self._element_position = 0 # rough position counter
# Full text (for regex scans)
self.full_text = []
def handle_starttag(self, tag, attrs):
attrs_dict = dict(attrs)
tag_lower = tag.lower()
if tag_lower == "script":
self._in_script = True
return
if tag_lower == "style":
self._in_style = True
return
if tag_lower == "body":
self._in_body = True
return
if tag_lower == "meta":
if attrs_dict.get("name", "").lower() == "viewport":
self.viewport_meta = True
if not self._in_body:
return
self._element_position += 1
# Buttons
if tag_lower == "button":
self._current_tag = "button"
self._current_text = []
elif tag_lower == "input":
input_type = attrs_dict.get("type", "text").lower()
if input_type == "submit":
val = attrs_dict.get("value", "Submit")
self.buttons.append({"text": val, "position": self._element_position})
elif input_type not in ("hidden", "submit"):
self.form_fields += 1
elif tag_lower == "textarea" or tag_lower == "select":
self.form_fields += 1
elif tag_lower == "form":
self.forms += 1
elif tag_lower == "a":
cls = attrs_dict.get("class", "").lower()
href = attrs_dict.get("href", "")
cta_classes = {"btn", "button", "cta", "call-to-action", "signup", "register"}
if any(c in cls for c in cta_classes):
self._current_tag = "a_cta"
self._current_text = []
elif tag_lower == "img":
src = attrs_dict.get("src", "").lower()
alt = attrs_dict.get("alt", "").lower()
cls = attrs_dict.get("class", "").lower()
if any(kw in src or kw in alt or kw in cls
for kw in ("logo", "partner", "client", "badge", "seal", "award", "cert")):
self.logo_images += 1
def handle_endtag(self, tag):
tag_lower = tag.lower()
if tag_lower == "script":
self._in_script = False
elif tag_lower == "style":
self._in_style = False
elif tag_lower == "button" and self._current_tag == "button":
text = " ".join(self._current_text).strip()
self.buttons.append({"text": text, "position": self._element_position})
self._current_tag = None
self._current_text = []
elif tag_lower == "a" and self._current_tag == "a_cta":
text = " ".join(self._current_text).strip()
self.links_as_cta.append({"text": text, "position": self._element_position})
self._current_tag = None
self._current_text = []
def handle_data(self, data):
if self._in_script or self._in_style:
return
text = data.strip()
if not text:
return
if self._current_tag in ("button", "a_cta"):
self._current_text.append(text)
if self._in_body:
self.full_text.append(text)
# ---------------------------------------------------------------------------
# Text-based signal detection
# ---------------------------------------------------------------------------
TESTIMONIAL_PATTERNS = [
r'\b(testimonial|review|quote|said|says|told us|customer story)\b',
r'[""][^""]{20,}[""]', # quoted text
r'\b\d[\d,]+ (reviews?|customers?|users?|clients?|companies)\b',
r'\bstar[s]?\b.{0,10}\b(rating|review)\b',
r'\b(trustpilot|g2|capterra|clutch)\b',
]
TRUST_PATTERNS = {
"ssl": [r'\b(ssl|https|secure|encrypted|tls|256.bit)\b'],
"guarantee": [r'\b(guarantee|guaranteed|money.back|refund|risk.free|no.risk)\b'],
"privacy": [r'\b(privacy|gdpr|data protection|we never share|no spam|unsubscribe)\b'],
}
CTA_TEXT_PATTERNS = [
r'\b(get started|sign up|try free|start free|buy now|order now|get access|'
r'download|schedule|book|claim|join|subscribe|register|contact us|learn more|'
r'get quote|request demo|start trial|get demo)\b',
]
def scan_text_signals(full_text: str) -> dict:
text_lower = full_text.lower()
testimonials = sum(
len(re.findall(p, text_lower, re.IGNORECASE))
for p in TESTIMONIAL_PATTERNS
)
trust = {}
for key, patterns in TRUST_PATTERNS.items():
trust[key] = sum(len(re.findall(p, text_lower, re.IGNORECASE)) for p in patterns)
cta_text_count = sum(
len(re.findall(p, text_lower, re.IGNORECASE))
for p in CTA_TEXT_PATTERNS
)
return {
"testimonial_signals": min(testimonials, 20),
"trust": trust,
"cta_text_count": cta_text_count,
}
# ---------------------------------------------------------------------------
# Scoring
# ---------------------------------------------------------------------------
def score_category(value, thresholds: list) -> int:
"""thresholds: [(min_value, score), ...] sorted asc. Returns score for first match."""
for min_val, score in sorted(thresholds, reverse=True):
if value >= min_val:
return score
return 0
def audit(html: str) -> dict:
parser = CROParser()
parser.feed(html)
full_text = " ".join(parser.full_text)
text_signals = scan_text_signals(full_text)
all_ctas = parser.buttons + parser.links_as_cta
total_cta_count = len(all_ctas) + text_signals["cta_text_count"]
# --- CTA ---
cta_score = score_category(total_cta_count, [(0, 0), (1, 50), (2, 75), (3, 90), (5, 100)])
cta_above_fold = len([c for c in all_ctas if c["position"] <= 5])
if cta_above_fold >= 1:
cta_score = min(100, cta_score + 10)
# --- Forms ---
if parser.forms == 0:
form_score = 60 # not all pages need forms
form_note = "No form detected (OK if not a lead gen page)"
elif parser.form_fields <= 3:
form_score = 100
form_note = f"{parser.form_fields} field(s) — minimal friction"
elif parser.form_fields <= 5:
form_score = 70
form_note = f"{parser.form_fields} field(s) — consider trimming"
else:
form_score = max(10, 100 - (parser.form_fields - 3) * 10)
form_note = f"{parser.form_fields} field(s) — too many, high friction"
# --- Social proof ---
social_signals = text_signals["testimonial_signals"] + parser.logo_images
social_score = score_category(social_signals, [(0, 0), (1, 40), (2, 65), (4, 85), (6, 100)])
# --- Trust signals ---
trust = text_signals["trust"]
trust_total = sum(min(1, v) for v in trust.values()) # 0-3
trust_score = score_category(trust_total, [(0, 20), (1, 60), (2, 80), (3, 100)])
# --- Viewport meta ---
viewport_score = 100 if parser.viewport_meta else 0
# --- Overall ---
weights = {
"cta": 0.30,
"social_proof": 0.25,
"trust_signals": 0.20,
"forms": 0.15,
"viewport_mobile": 0.10,
}
scores = {
"cta": cta_score,
"social_proof": social_score,
"trust_signals": trust_score,
"forms": form_score,
"viewport_mobile": viewport_score,
}
overall = round(sum(scores[k] * weights[k] for k in weights))
return {
"overall_score": overall,
"categories": {
"cta_buttons": {
"score": cta_score,
"button_count": len(parser.buttons),
"cta_link_count": len(parser.links_as_cta),
"cta_text_count": text_signals["cta_text_count"],
"above_fold_ctas": cta_above_fold,
"weight": "30%",
},
"social_proof": {
"score": social_score,
"testimonial_signals": text_signals["testimonial_signals"],
"logo_badge_images": parser.logo_images,
"total_signals": social_signals,
"weight": "25%",
},
"trust_signals": {
"score": trust_score,
"ssl_mentions": trust["ssl"],
"guarantee_mentions": trust["guarantee"],
"privacy_mentions": trust["privacy"],
"weight": "20%",
},
"forms": {
"score": form_score,
"form_count": parser.forms,
"field_count": parser.form_fields,
"note": form_note,
"weight": "15%",
},
"viewport_mobile": {
"score": viewport_score,
"viewport_meta_present": parser.viewport_meta,
"weight": "10%",
},
},
}
# ---------------------------------------------------------------------------
# Demo HTML
# ---------------------------------------------------------------------------
DEMO_HTML = """<!DOCTYPE html>
<html>
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Get Your Free Marketing Audit</title>
</head>
<body>
<header>
<img src="logo.png" alt="Acme Corp logo" class="logo">
<a href="#form" class="btn cta">Get Free Audit</a>
</header>
<section class="hero">
<h1>Stop Wasting Your Ad Budget</h1>
<p>Join 12,400 marketers who cut wasted spend by 35% in 30 days.</p>
<button>Start Free Trial</button>
</section>
<section class="social-proof">
<h2>What Our Customers Say</h2>
<blockquote>"This tool saved us $50,000 in the first quarter." — Sarah M., CMO</blockquote>
<blockquote>"Best investment we made in 2023." — James T., Head of Growth</blockquote>
<p>Rated 4.9/5 on G2 with 2,400+ reviews</p>
<p>Trusted by 500+ companies worldwide</p>
<img src="google-partner.png" alt="Google Partner badge" class="badge">
<img src="trustpilot.png" alt="Trustpilot certified" class="badge">
</section>
<section id="form">
<h2>Get Your Free Audit</h2>
<form>
<input type="text" name="name" placeholder="Your name">
<input type="email" name="email" placeholder="Work email">
<button type="submit">Get My Free Audit</button>
</form>
<p>🔒 SSL secured. We never share your data. Unsubscribe anytime.</p>
<p>30-day money-back guarantee. No risk.</p>
</section>
</body>
</html>"""
# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
def main():
parser = argparse.ArgumentParser(
description="CRO audit — analyzes an HTML page for conversion signals."
)
parser.add_argument("--file", help="Path to HTML file")
parser.add_argument("--url", help="URL to fetch and analyze")
parser.add_argument("--json", action="store_true", help="Output as JSON")
args = parser.parse_args()
if args.file:
with open(args.file, "r", encoding="utf-8", errors="replace") as f:
html = f.read()
elif args.url:
with urllib.request.urlopen(args.url, timeout=10) as resp:
html = resp.read().decode("utf-8", errors="replace")
else:
html = DEMO_HTML
if not args.json:
print("No input provided — running in demo mode.\n")
result = audit(html)
if args.json:
print(json.dumps(result, indent=2))
return
cats = result["categories"]
overall = result["overall_score"]
print("=" * 62)
print(f" CRO AUDIT RESULTS Overall Score: {overall}/100")
print("=" * 62)
rows = [
("CTA Buttons", "cta_buttons"),
("Social Proof", "social_proof"),
("Trust Signals", "trust_signals"),
("Forms", "forms"),
("Mobile Viewport", "viewport_mobile"),
]
for label, key in rows:
c = cats[key]
score = c["score"]
weight = c["weight"]
bar_len = round(score / 10)
bar = "█" * bar_len + "░" * (10 - bar_len)
icon = "✅" if score >= 70 else ("⚠️ " if score >= 40 else "❌")
print(f" {icon} {label:<18} [{bar}] {score:>3}/100 (weight {weight})")
print()
# Detail callouts
cta = cats["cta_buttons"]
print(f" CTAs: {cta['button_count']} buttons, {cta['cta_link_count']} CTA links, "
f"{cta['cta_text_count']} CTA text phrases, {cta['above_fold_ctas']} above fold")
sp = cats["social_proof"]
print(f" Social Proof: {sp['testimonial_signals']} testimonial signals, "
f"{sp['logo_badge_images']} logos/badges")
ts = cats["trust_signals"]
print(f" Trust: SSL({ts['ssl_mentions']}) Guarantee({ts['guarantee_mentions']}) "
f"Privacy({ts['privacy_mentions']})")
fm = cats["forms"]
print(f" Forms: {fm['form_count']} form(s), {fm['field_count']} field(s) — {fm['note']}")
print()
grade = "A" if overall >= 85 else "B" if overall >= 70 else "C" if overall >= 55 else "D" if overall >= 40 else "F"
print("=" * 62)
print(f" Grade: {grade} Score: {overall}/100")
print("=" * 62)
if __name__ == "__main__":
main()
Hỗ trợ chiến dịch quảng cáo trả phí trên Google Ads, Meta, LinkedIn, Twitter/X: nội dung quảng cáo, ROAS, CPA, retargeting và nhắm đối tượng.
---
name: "paid-ads"
description: "When the user wants help with paid advertising campaigns on Google Ads, Meta (Facebook/Instagram), LinkedIn, Twitter/X, or other ad platforms. Also use when the user mentions 'PPC,' 'paid media,' 'ad copy,' 'ad creative,' 'ROAS,' 'CPA,' 'ad campaign,' 'retargeting,' or 'audience targeting.' This skill covers campaign strategy, ad creation, audience targeting, and optimization."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: marketing
updated: 2026-03-06
---
# Paid Ads
You are an expert performance marketer with direct access to ad platform accounts. Your goal is to help create, optimize, and scale paid advertising campaigns that drive efficient customer acquisition.
## Before Starting
**Check for product marketing context first:**
If `.claude/product-marketing-context.md` exists, read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
Gather this context (ask if not provided):
### 1. Campaign Goals
- What's the primary objective? (Awareness, traffic, leads, sales, app installs)
- What's the target CPA or ROAS?
- What's the monthly/weekly budget?
- Any constraints? (Brand guidelines, compliance, geographic)
### 2. Product & Offer
- What are you promoting? (Product, free trial, lead magnet, demo)
- What's the landing page URL?
- What makes this offer compelling?
### 3. Audience
- Who is the ideal customer?
- What problem does your product solve for them?
- What are they searching for or interested in?
- Do you have existing customer data for lookalikes?
### 4. Current State
- Have you run ads before? What worked/didn't?
- Do you have existing pixel/conversion data?
- What's your current funnel conversion rate?
---
## Platform Selection Guide
| Platform | Best For | Use When |
|----------|----------|----------|
| **Google Ads** | High-intent search traffic | People actively search for your solution |
| **Meta** | Demand generation, visual products | Creating demand, strong creative assets |
| **LinkedIn** | B2B, decision-makers | Job title/company targeting matters, higher price points |
| **Twitter/X** | Tech audiences, thought leadership | Audience is active on X, timely content |
| **TikTok** | Younger demographics, viral creative | Audience skews 18-34, video capacity |
---
## Campaign Structure Best Practices
### Account Organization
```
Account
├── Campaign 1: [Objective] - [Audience/Product]
│ ├── Ad Set 1: [Targeting variation]
│ │ ├── Ad 1: [Creative variation A]
│ │ ├── Ad 2: [Creative variation B]
│ │ └── Ad 3: [Creative variation C]
│ └── Ad Set 2: [Targeting variation]
└── Campaign 2...
```
### Naming Conventions
```
[Platform]_[Objective]_[Audience]_[Offer]_[Date]
Examples:
META_Conv_Lookalike-Customers_FreeTrial_2024Q1
GOOG_Search_Brand_Demo_Ongoing
LI_LeadGen_CMOs-SaaS_Whitepaper_Mar24
```
### Budget Allocation
**Testing phase (first 2-4 weeks):**
- 70% to proven/safe campaigns
- 30% to testing new audiences/creative
**Scaling phase:**
- Consolidate budget into winning combinations
- Increase budgets 20-30% at a time
- Wait 3-5 days between increases for algorithm learning
---
## Ad Copy Frameworks
### Key Formulas
**Problem-Agitate-Solve (PAS):**
> [Problem] → [Agitate the pain] → [Introduce solution] → [CTA]
**Before-After-Bridge (BAB):**
> [Current painful state] → [Desired future state] → [Your product as bridge]
**Social Proof Lead:**
> [Impressive stat or testimonial] → [What you do] → [CTA]
**For detailed templates and headline formulas**: See [references/ad-copy-templates.md](references/ad-copy-templates.md)
---
## Audience Targeting Overview
### Platform Strengths
| Platform | Key Targeting | Best Signals |
|----------|---------------|--------------|
| Google | Keywords, search intent | What they're searching |
| Meta | Interests, behaviors, lookalikes | Engagement patterns |
| LinkedIn | Job titles, companies, industries | Professional identity |
### Key Concepts
- **Lookalikes**: Base on best customers (by LTV), not all customers
- **Retargeting**: Segment by funnel stage (visitors vs. cart abandoners)
- **Exclusions**: Always exclude existing customers and recent converters
**For detailed targeting strategies by platform**: See [references/audience-targeting.md](references/audience-targeting.md)
---
## Creative Best Practices
### Image Ads
- Clear product screenshots showing UI
- Before/after comparisons
- Stats and numbers as focal point
- Human faces (real, not stock)
- Bold, readable text overlay (keep under 20%)
### Video Ads Structure (15-30 sec)
1. Hook (0-3 sec): Pattern interrupt, question, or bold statement
2. Problem (3-8 sec): Relatable pain point
3. Solution (8-20 sec): Show product/benefit
4. CTA (20-30 sec): Clear next step
**Production tips:**
- Captions always (85% watch without sound)
- Vertical for Stories/Reels, square for feed
- Native feel outperforms polished
- First 3 seconds determine if they watch
### Creative Testing Hierarchy
1. Concept/angle (biggest impact)
2. Hook/headline
3. Visual style
4. Body copy
5. CTA
---
## Campaign Optimization
### Key Metrics by Objective
| Objective | Primary Metrics |
|-----------|-----------------|
| Awareness | CPM, Reach, Video view rate |
| Consideration | CTR, CPC, Time on site |
| Conversion | CPA, ROAS, Conversion rate |
### Optimization Levers
**If CPA is too high:**
1. Check landing page (is the problem post-click?)
2. Tighten audience targeting
3. Test new creative angles
4. Improve ad relevance/quality score
5. Adjust bid strategy
**If CTR is low:**
- Creative isn't resonating → test new hooks/angles
- Audience mismatch → refine targeting
- Ad fatigue → refresh creative
**If CPM is high:**
- Audience too narrow → expand targeting
- High competition → try different placements
- Low relevance score → improve creative fit
### Bid Strategy Progression
1. Start with manual or cost caps
2. Gather conversion data (50+ conversions)
3. Switch to automated with targets based on historical data
4. Monitor and adjust targets based on results
---
## Retargeting Strategies
### Funnel-Based Approach
| Funnel Stage | Audience | Message | Goal |
|--------------|----------|---------|------|
| Top | Blog readers, video viewers | Educational, social proof | Move to consideration |
| Middle | Pricing/feature page visitors | Case studies, demos | Move to decision |
| Bottom | Cart abandoners, trial users | Urgency, objection handling | Convert |
### Retargeting Windows
| Stage | Window | Frequency Cap |
|-------|--------|---------------|
| Hot (cart/trial) | 1-7 days | Higher OK |
| Warm (key pages) | 7-30 days | 3-5x/week |
| Cold (any visit) | 30-90 days | 1-2x/week |
### Exclusions to Set Up
- Existing customers (unless upsell)
- Recent converters (7-14 day window)
- Bounced visitors (<10 sec)
- Irrelevant pages (careers, support)
---
## Reporting & Analysis
### Weekly Review
- Spend vs. budget pacing
- CPA/ROAS vs. targets
- Top and bottom performing ads
- Audience performance breakdown
- Frequency check (fatigue risk)
- Landing page conversion rate
### Attribution Considerations
- Platform attribution is inflated
- Use UTM parameters consistently
- Compare platform data to GA4
- Look at blended CAC, not just platform CPA
---
## Platform Setup
Before launching campaigns, ensure proper tracking and account setup.
**For complete setup checklists by platform**: See [references/platform-setup-checklists.md](references/platform-setup-checklists.md)
### Universal Pre-Launch Checklist
- [ ] Conversion tracking tested with real conversion
- [ ] Landing page loads fast (<3 sec)
- [ ] Landing page mobile-friendly
- [ ] UTM parameters working
- [ ] Budget set correctly
- [ ] Targeting matches intended audience
---
## Common Mistakes to Avoid
### Strategy
- Launching without conversion tracking
- Too many campaigns (fragmenting budget)
- Not giving algorithms enough learning time
- Optimizing for wrong metric
### Targeting
- Audiences too narrow or too broad
- Not excluding existing customers
- Overlapping audiences competing
### Creative
- Only one ad per ad set
- Not refreshing creative (fatigue)
- Mismatch between ad and landing page
### Budget
- Spreading too thin across campaigns
- Making big budget changes (disrupts learning)
- Stopping campaigns during learning phase
---
## Task-Specific Questions
1. What platform(s) are you currently running or want to start with?
2. What's your monthly ad budget?
3. What does a successful conversion look like (and what's it worth)?
4. Do you have existing creative assets or need to create them?
5. What landing page will ads point to?
6. Do you have pixel/conversion tracking set up?
---
## Tool Integrations
For implementation, see the [tools registry](../../tools/REGISTRY.md). Key advertising platforms:
| Platform | Best For | MCP | Guide |
|----------|----------|:---:|-------|
| **Google Ads** | Search intent, high-intent traffic | ✓ | [google-ads.md](../../tools/integrations/google-ads.md) |
| **Meta Ads** | Demand gen, visual products, B2C | - | [meta-ads.md](../../tools/integrations/meta-ads.md) |
| **LinkedIn Ads** | B2B, job title targeting | - | [linkedin-ads.md](../../tools/integrations/linkedin-ads.md) |
| **TikTok Ads** | Younger demographics, video | - | [tiktok-ads.md](../../tools/integrations/tiktok-ads.md) |
For tracking, see also: [ga4.md](../../tools/integrations/ga4.md), [segment.md](../../tools/integrations/segment.md)
---
## Related Skills
- **ad-creative** — WHEN you need deep creative direction for ad visuals, video scripts, or creative concepting beyond basic image/copy guidelines. NOT for campaign strategy, targeting, or bidding decisions.
- **analytics-tracking** — WHEN setting up conversion tracking pixels, UTM parameters, and attribution models before or during campaign launch. NOT for campaign creation or creative work.
- **campaign-analytics** — WHEN analyzing campaign performance data, diagnosing underperforming campaigns, or building reporting dashboards. NOT for initial campaign setup or creative production.
- **copywriting** — WHEN landing pages linked from ads need copy optimization to match ad messaging and improve post-click conversion. NOT for the ad copy itself.
- **marketing-context** — Foundation skill for ICP, positioning, and messaging alignment. ALWAYS load before writing ad copy or selecting targeting to ensure message-market fit.
---
## Communication
Always confirm conversion tracking is in place before recommending creative or targeting changes — a campaign without proper attribution is guesswork. When recommending budget allocation, state the rationale (testing vs. scaling phase). Deliver ad copy as complete, ready-to-launch sets: headline variants, body copy, and CTA. Proactively flag when a landing page mismatch (ad promise ≠ page promise) is the likely conversion bottleneck. Load `marketing-context` for ICP and positioning before writing any copy.
---
## Proactive Triggers
- User asks why ROAS is dropping → check creative fatigue and ad frequency before adjusting targeting or bids.
- User wants to launch their first paid campaign → run through the pre-launch checklist (conversion tracking, landing page speed, UTMs) before touching creative.
- User mentions high CTR but low conversions → diagnose landing page, not the ad; redirect to `page-cro` or `copywriting` skill.
- User is scaling budget aggressively → warn about algorithm learning phase disruption; recommend 20-30% incremental increases with 3-5 day stabilization windows.
- User asks about B2B lead generation via ads → recommend LinkedIn for job-title targeting and flag that CPL will be higher but lead quality better than Meta for high-ACV products.
---
## Output Artifacts
| Artifact | Description |
|----------|-------------|
| Campaign Architecture | Full account structure with campaign names, ad set targeting, naming conventions, and budget allocation |
| Ad Copy Set | 3 headline variants, body copy, and CTA for each ad format and platform, ready to launch |
| Audience Targeting Brief | Primary audiences, lookalike seeds, retargeting segments, and exclusion lists per platform |
| Pre-Launch Checklist | Platform-specific tracking verification, landing page audit, and UTM parameter setup |
| Weekly Optimization Report Template | Metrics dashboard structure with CPA/ROAS targets, fatigue signals, and decision triggers |
FILE:references/ad-copy-templates.md
# Ad Copy Templates Reference
Detailed formulas and templates for writing high-converting ad copy.
## Primary Text Formulas
### Problem-Agitate-Solve (PAS)
```
[Problem statement]
[Agitate the pain]
[Introduce solution]
[CTA]
```
**Example:**
> Spending hours on manual reporting every week?
> While you're buried in spreadsheets, your competitors are making decisions.
> [Product] automates your reports in minutes.
> Start your free trial →
---
### Before-After-Bridge (BAB)
```
[Current painful state]
[Desired future state]
[Your product as the bridge]
```
**Example:**
> Before: Chasing down approvals across email, Slack, and spreadsheets.
> After: Every approval tracked, automated, and on time.
> [Product] connects your tools and keeps projects moving.
---
### Social Proof Lead
```
[Impressive stat or testimonial]
[What you do]
[CTA]
```
**Example:**
> "We cut our reporting time by 75%." — Sarah K., Marketing Director
> [Product] automates the reports you hate building.
> See how it works →
---
### Feature-Benefit Bridge
```
[Feature]
[So that...]
[Which means...]
```
**Example:**
> Real-time collaboration on documents
> So your team always works from the latest version
> Which means no more version confusion or lost work
---
### Direct Response
```
[Bold claim/outcome]
[Proof point]
[CTA with urgency if genuine]
```
**Example:**
> Cut your reporting time by 80%
> Join 5,000+ marketing teams already using [Product]
> Start free → First month 50% off
---
## Headline Formulas
### For Search Ads
| Formula | Example |
|---------|---------|
| [Keyword] + [Benefit] | "Project Management That Teams Actually Use" |
| [Action] + [Outcome] | "Automate Reports \| Save 10 Hours Weekly" |
| [Question] | "Tired of Manual Data Entry?" |
| [Number] + [Benefit] | "500+ Teams Trust [Product] for [Outcome]" |
| [Keyword] + [Differentiator] | "CRM Built for Small Teams" |
| [Price/Offer] + [Keyword] | "Free Project Management \| No Credit Card" |
### For Social Ads
| Type | Example |
|------|---------|
| Outcome hook | "How we 3x'd our conversion rate" |
| Curiosity hook | "The reporting hack no one talks about" |
| Contrarian hook | "Why we stopped using [common tool]" |
| Specificity hook | "The exact template we use for..." |
| Question hook | "What if you could cut your admin time in half?" |
| Number hook | "7 ways to improve your workflow today" |
| Story hook | "We almost gave up. Then we found..." |
---
## CTA Variations
### Soft CTAs (awareness/consideration)
Best for: Top of funnel, cold audiences, complex products
- Learn More
- See How It Works
- Watch Demo
- Get the Guide
- Explore Features
- See Examples
- Read the Case Study
### Hard CTAs (conversion)
Best for: Bottom of funnel, warm audiences, clear offers
- Start Free Trial
- Get Started Free
- Book a Demo
- Claim Your Discount
- Buy Now
- Sign Up Free
- Get Instant Access
### Urgency CTAs (use when genuine)
Best for: Limited-time offers, scarcity situations
- Limited Time: 30% Off
- Offer Ends [Date]
- Only X Spots Left
- Last Chance
- Early Bird Pricing Ends Soon
### Action-Oriented CTAs
Best for: Active voice, clear next step
- Start Saving Time Today
- Get Your Free Report
- See Your Score
- Calculate Your ROI
- Build Your First Project
---
## Platform-Specific Copy Guidelines
### Google Search Ads
- **Headline limits:** 30 characters each (up to 15 headlines)
- **Description limits:** 90 characters each (up to 4 descriptions)
- Include keywords naturally
- Use all available headline slots
- Include numbers and stats when possible
- Test dynamic keyword insertion
### Meta Ads (Facebook/Instagram)
- **Primary text:** 125 characters visible (can be longer, gets truncated)
- **Headline:** 40 characters recommended
- Front-load the hook (first line matters most)
- Emojis can work but test
- Questions perform well
- Keep image text under 20%
### LinkedIn Ads
- **Intro text:** 600 characters max (150 recommended)
- **Headline:** 200 characters max (70 recommended)
- Professional tone (but not boring)
- Specific job outcomes resonate
- Stats and social proof important
- Avoid consumer-style hype
---
## Copy Testing Priority
When testing ad copy, focus on these elements in order of impact:
1. **Hook/angle** (biggest impact on performance)
2. **Headline**
3. **Primary benefit**
4. **CTA**
5. **Supporting proof points**
Test one element at a time for clean data.
FILE:references/audience-targeting.md
# Audience Targeting Reference
Detailed targeting strategies for each major ad platform.
## Google Ads Audiences
### Search Campaign Targeting
**Keywords:**
- Exact match: [keyword] — most precise, lower volume
- Phrase match: "keyword" — moderate precision and volume
- Broad match: keyword — highest volume, use with smart bidding
**Audience layering:**
- Add audiences in "observation" mode first
- Analyze performance by audience
- Switch to "targeting" mode for high performers
**RLSA (Remarketing Lists for Search Ads):**
- Bid higher on past visitors searching your terms
- Show different ads to returning searchers
- Exclude converters from prospecting campaigns
### Display/YouTube Targeting
**Custom intent audiences:**
- Based on recent search behavior
- Create from your converting keywords
- High intent, good for prospecting
**In-market audiences:**
- People actively researching solutions
- Pre-built by Google
- Layer with demographics for precision
**Affinity audiences:**
- Based on interests and habits
- Better for awareness
- Broad but can exclude irrelevant
**Customer match:**
- Upload email lists
- Retarget existing customers
- Create lookalikes from best customers
**Similar/lookalike audiences:**
- Based on your customer match lists
- Expand reach while maintaining relevance
- Best when source list is high-quality customers
---
## Meta Audiences
### Core Audiences (Interest/Demographic)
**Interest targeting tips:**
- Layer interests with AND logic for precision
- Use Audience Insights to research interests
- Start broad, let algorithm optimize
- Exclude existing customers always
**Demographic targeting:**
- Age and gender (if product-specific)
- Location (down to zip/postal code)
- Language
- Education and work (limited data now)
**Behavior targeting:**
- Purchase behavior
- Device usage
- Travel patterns
- Life events
### Custom Audiences
**Website visitors:**
- All visitors (last 180 days max)
- Specific page visitors
- Time on site thresholds
- Frequency (visited X times)
**Customer list:**
- Upload emails/phone numbers
- Match rate typically 30-70%
- Refresh regularly for accuracy
**Engagement audiences:**
- Video viewers (25%, 50%, 75%, 95%)
- Page/profile engagers
- Form openers
- Instagram engagers
**App activity:**
- App installers
- In-app events
- Purchase events
### Lookalike Audiences
**Source audience quality matters:**
- Use high-LTV customers, not all customers
- Purchasers > leads > all visitors
- Minimum 100 source users, ideally 1,000+
**Size recommendations:**
- 1% — most similar, smallest reach
- 1-3% — good balance for most
- 3-5% — broader, good for scale
- 5-10% — very broad, awareness only
**Layering strategies:**
- Lookalike + interest = more precision early
- Test lookalike-only as you scale
- Exclude the source audience
---
## LinkedIn Audiences
### Job-Based Targeting
**Job titles:**
- Be specific (CMO vs. "Marketing")
- LinkedIn normalizes titles, but verify
- Stack related titles
- Exclude irrelevant titles
**Job functions:**
- Broader than titles
- Combine with seniority level
- Good for awareness campaigns
**Seniority levels:**
- Entry, Senior, Manager, Director, VP, CXO, Partner
- Layer with function for precision
**Skills:**
- Self-reported, less reliable
- Good for technical roles
- Use as expansion layer
### Company-Based Targeting
**Company size:**
- 1-10, 11-50, 51-200, 201-500, 501-1000, 1001-5000, 5000+
- Key filter for B2B
**Industry:**
- Based on company classification
- Can be broad, layer with other criteria
**Company names (ABM):**
- Upload target account list
- Minimum 300 companies recommended
- Match rate varies
**Company growth rate:**
- Hiring rapidly = budget available
- Good signal for timing
### High-Performing Combinations
| Use Case | Targeting Combination |
|----------|----------------------|
| Enterprise sales | Company size 1000+ + VP/CXO + Industry |
| SMB sales | Company size 11-200 + Manager/Director + Function |
| Developer tools | Skills + Job function + Company type |
| ABM campaigns | Company list + Decision-maker titles |
| Broad awareness | Industry + Seniority + Geography |
---
## Twitter/X Audiences
### Targeting options:
- Follower lookalikes (accounts similar to followers of X)
- Interest categories
- Keywords (in tweets)
- Conversation topics
- Events
- Tailored audiences (your lists)
### Best practices:
- Follower lookalikes of relevant accounts work well
- Keyword targeting catches active conversations
- Lower CPMs than LinkedIn/Meta
- Less precise, better for awareness
---
## TikTok Audiences
### Targeting options:
- Demographics (age, gender, location)
- Interests (TikTok's categories)
- Behaviors (video interactions)
- Device (iOS/Android, connection type)
- Custom audiences (pixel, customer file)
- Lookalike audiences
### Best practices:
- Younger skew (18-34 primarily)
- Interest targeting is broad
- Creative matters more than targeting
- Let algorithm optimize with broad targeting
---
## Audience Size Guidelines
| Platform | Minimum Recommended | Ideal Range |
|----------|-------------------|-------------|
| Google Search | 1,000+ searches/mo | 5,000-50,000 |
| Google Display | 100,000+ | 500K-5M |
| Meta | 100,000+ | 500K-10M |
| LinkedIn | 50,000+ | 100K-500K |
| Twitter/X | 50,000+ | 100K-1M |
| TikTok | 100,000+ | 1M+ |
Too narrow = expensive, slow learning
Too broad = wasted spend, poor relevance
---
## Exclusion Strategy
Always exclude:
- Existing customers (unless upsell)
- Recent converters (7-14 days)
- Bounced visitors (<10 sec)
- Employees (by company or email list)
- Irrelevant page visitors (careers, support)
- Competitors (if identifiable)
FILE:references/copy-frameworks.md
# Ad Copy Frameworks
Reference for selecting the right copy framework based on product type and campaign goal. Each framework includes a structure template and platform-specific length constraints.
## Framework selection matrix
| Product type | Pain-point heavy? | Transformation story? | Feature-led? | Recommended framework |
|---|---|---|---|---|
| SaaS / B2B | ✅ | | | PAS (Problem-Agitate-Solve) |
| Coaching / courses | | ✅ | | BAB (Before-After-Bridge) |
| Ecommerce / physical | | | ✅ | FAB (Features-Advantages-Benefits) |
| Content / info product | ✅ | ✅ | | AIDA (Attention-Interest-Desire-Action) |
| App / tool launch | | | ✅ | 4P (Promise-Picture-Proof-Push) |
| Services / consulting | ✅ | ✅ | | Star-Story-Solution |
## The 6 frameworks
### PAS — Problem → Agitate → Solve
**Best for:** Pain-point products, SaaS solving specific frustrations
```
Problem: Name the exact pain (1 sentence)
Agitate: Twist the knife — what happens if they don't fix it (1-2 sentences)
Solve: Your product is the answer (1 sentence + CTA)
```
### BAB — Before → After → Bridge
**Best for:** Transformation products, coaching, courses
```
Before: Current painful state (1 sentence)
After: Desired state they'll achieve (1 sentence)
Bridge: Your product connects the two (1 sentence + CTA)
```
### AIDA — Attention → Interest → Desire → Action
**Best for:** Content marketing, info products, broad audiences
```
Attention: Hook with a surprising stat or question
Interest: Explain why this matters to them
Desire: Show social proof or specific outcomes
Action: Clear CTA with urgency
```
### FAB — Features → Advantages → Benefits
**Best for:** Product-led, ecommerce, feature-rich offerings
```
Feature: What it has (spec/capability)
Advantage: Why that matters vs alternatives
Benefit: What the user gains (outcome)
```
### 4P — Promise → Picture → Proof → Push
**Best for:** App launches, tools, direct response
```
Promise: Bold claim (1 headline)
Picture: Vivid scenario of life with the product
Proof: Social proof, stats, testimonial
Push: Strong CTA with urgency/scarcity
```
### Star-Story-Solution
**Best for:** Personal brands, services, consulting
```
Star: Introduce the hero (the customer, not you)
Story: Their struggle (relatable narrative)
Solution: How your service transforms their situation
```
## Platform-specific constraints
| Platform | Headline | Body | CTA |
|---|---|---|---|
| Google RSA | 30 chars × 15 headlines | 90 chars × 4 descriptions | Auto from list |
| Meta Feed | 40 chars (before truncation) | 125 chars primary text (before "See more") | Button from list |
| Meta Stories | 40 chars overlay | Minimal — visual-first | Swipe up / button |
| LinkedIn Sponsored | 70 chars intro text visible | 150 chars before truncation | Button from list |
| TikTok | Overlay text in video | Caption 100 chars | Button from list |
| Microsoft | 30 chars × 15 headlines | 90 chars × 4 descriptions | Auto from list |
## Brand DNA extraction (7 voice axes)
Before writing ad copy, extract the brand's voice profile on these 7 axes:
```json
{
"formal_casual": 0.7, // 0 = corporate formal, 1 = casual/friendly
"bold_subtle": 0.6, // 0 = understated, 1 = bold/provocative
"technical_human": 0.4, // 0 = jargon-heavy, 1 = plain language
"serious_playful": 0.5, // 0 = gravitas, 1 = humor/wit
"traditional_innovative": 0.8, // 0 = established, 1 = cutting-edge
"exclusive_inclusive": 0.6, // 0 = luxury/elite, 1 = accessible/everyone
"data_emotional": 0.5 // 0 = stats-driven, 1 = story-driven
}
```
Save as `brand-profile.json` for reuse across campaigns. Each axis is 0.0-1.0.
FILE:references/platform-setup-checklists.md
# Platform Setup Checklists
Complete setup checklists for major ad platforms.
## Google Ads Setup
### Account Foundation
- [ ] Google Ads account created and verified
- [ ] Billing information added
- [ ] Time zone and currency set correctly
- [ ] Account access granted to team members
### Conversion Tracking
- [ ] Google tag installed on all pages
- [ ] Conversion actions created (purchase, lead, signup)
- [ ] Conversion values assigned (if applicable)
- [ ] Enhanced conversions enabled
- [ ] Test conversions firing correctly
- [ ] Import conversions from GA4 (optional)
### Analytics Integration
- [ ] Google Analytics 4 linked
- [ ] Auto-tagging enabled
- [ ] GA4 audiences available in Google Ads
- [ ] Cross-domain tracking set up (if multiple domains)
### Audience Setup
- [ ] Remarketing tag verified
- [ ] Website visitor audiences created:
- All visitors (180 days)
- Key page visitors (pricing, demo, features)
- Converters (for exclusion)
- [ ] Customer match lists uploaded
- [ ] Similar audiences enabled
### Campaign Readiness
- [ ] Negative keyword lists created:
- Universal negatives (free, jobs, careers, reviews, complaints)
- Competitor negatives (if needed)
- Irrelevant industry terms
- [ ] Location targeting set (include/exclude)
- [ ] Language targeting set
- [ ] Ad schedule configured (if B2B, business hours)
- [ ] Device bid adjustments considered
### Ad Extensions
- [ ] Sitelinks (4-6 relevant pages)
- [ ] Callouts (key benefits, offers)
- [ ] Structured snippets (features, types, services)
- [ ] Call extension (if phone leads valuable)
- [ ] Lead form extension (if using)
- [ ] Price extensions (if applicable)
- [ ] Image extensions (where available)
### Brand Protection
- [ ] Brand campaign running (protect branded terms)
- [ ] Competitor campaigns considered
- [ ] Brand terms in negative lists for non-brand campaigns
---
## Meta Ads Setup
### Business Manager Foundation
- [ ] Business Manager created
- [ ] Business verified (if running certain ad types)
- [ ] Ad account created within Business Manager
- [ ] Payment method added
- [ ] Team access configured with proper roles
### Pixel & Tracking
- [ ] Meta Pixel installed on all pages
- [ ] Standard events configured:
- PageView (automatic)
- ViewContent (product/feature pages)
- Lead (form submissions)
- Purchase (conversions)
- AddToCart (if e-commerce)
- InitiateCheckout (if e-commerce)
- [ ] Conversions API (CAPI) set up for server-side tracking
- [ ] Event Match Quality score > 6
- [ ] Test events in Events Manager
### Domain & Aggregated Events
- [ ] Domain verified in Business Manager
- [ ] Aggregated Event Measurement configured
- [ ] Top 8 events prioritized in order of importance
- [ ] Web events prioritized for iOS 14+ tracking
### Audience Setup
- [ ] Custom audiences created:
- Website visitors (all, 30/60/90/180 days)
- Key page visitors
- Video viewers (25%, 50%, 75%, 95%)
- Page/Instagram engagers
- Customer list uploaded
- [ ] Lookalike audiences created (1%, 1-3%)
- [ ] Saved audiences for common targeting
### Catalog (E-commerce)
- [ ] Product catalog connected
- [ ] Product feed updating correctly
- [ ] Catalog sales campaigns enabled
- [ ] Dynamic product ads configured
### Creative Assets
- [ ] Images in correct sizes:
- Feed: 1080x1080 (1:1)
- Stories/Reels: 1080x1920 (9:16)
- Landscape: 1200x628 (1.91:1)
- [ ] Videos in correct formats
- [ ] Ad copy variations ready
- [ ] UTM parameters in all destination URLs
### Compliance
- [ ] Special Ad Categories declared (if housing, credit, employment, politics)
- [ ] Landing page complies with Meta policies
- [ ] No prohibited content in ads
---
## LinkedIn Ads Setup
### Campaign Manager Foundation
- [ ] Campaign Manager account created
- [ ] Company Page connected
- [ ] Billing information added
- [ ] Team access configured
### Insight Tag & Tracking
- [ ] LinkedIn Insight Tag installed on all pages
- [ ] Tag verified and firing
- [ ] Conversion tracking configured:
- URL-based conversions
- Event-specific conversions
- [ ] Conversion values set (if applicable)
### Audience Setup
- [ ] Matched Audiences created:
- Website retargeting audiences
- Company list uploaded (for ABM)
- Contact list uploaded
- [ ] Lookalike audiences created
- [ ] Saved audiences for common targeting
### Lead Gen Forms (if using)
- [ ] Lead gen form templates created
- [ ] Form fields selected (minimize for conversion)
- [ ] Privacy policy URL added
- [ ] Thank you message configured
- [ ] CRM integration set up (or CSV export process)
### Document Ads (if using)
- [ ] Documents uploaded (PDF, PowerPoint)
- [ ] Gating configured (full gate or preview)
- [ ] Lead gen form connected
### Creative Assets
- [ ] Single image ads: 1200x627 (1.91:1) or 1080x1080 (1:1)
- [ ] Carousel images ready
- [ ] Video specs met (if using)
- [ ] Ad copy within character limits:
- Intro text: 600 max, 150 recommended
- Headline: 200 max, 70 recommended
### Budget Considerations
- [ ] Budget realistic for LinkedIn CPCs ($8-15+ typical)
- [ ] Audience size validated (50K+ recommended)
- [ ] Daily vs. lifetime budget decided
- [ ] Bid strategy selected
---
## Twitter/X Ads Setup
### Account Foundation
- [ ] Ads account created
- [ ] Payment method added
- [ ] Account verified (if required)
### Tracking
- [ ] Twitter Pixel installed
- [ ] Conversion events created
- [ ] Website tag verified
### Audience Setup
- [ ] Tailored audiences created:
- Website visitors
- Customer lists
- [ ] Follower lookalikes identified
- [ ] Interest and keyword targets researched
### Creative
- [ ] Tweet copy within 280 characters
- [ ] Images: 1200x675 (1.91:1) or 1200x1200 (1:1)
- [ ] Video specs met (if using)
- [ ] Cards configured (website, app, etc.)
---
## TikTok Ads Setup
### Account Foundation
- [ ] TikTok Ads Manager account created
- [ ] Business verification completed
- [ ] Payment method added
### Pixel & Tracking
- [ ] TikTok Pixel installed
- [ ] Events configured (ViewContent, Purchase, etc.)
- [ ] Events API set up (recommended)
### Audience Setup
- [ ] Custom audiences created
- [ ] Lookalike audiences created
- [ ] Interest categories identified
### Creative
- [ ] Vertical video (9:16) ready
- [ ] Native-feeling content (not too polished)
- [ ] First 3 seconds are compelling hooks
- [ ] Captions added (most watch without sound)
- [ ] Music/sounds selected (licensed if needed)
---
## Universal Pre-Launch Checklist
Before launching any campaign:
- [ ] Conversion tracking tested with real conversion
- [ ] Landing page loads fast (<3 sec)
- [ ] Landing page mobile-friendly
- [ ] UTM parameters working
- [ ] Budget set correctly (daily vs. lifetime)
- [ ] Start/end dates correct
- [ ] Targeting matches intended audience
- [ ] Ad creative approved
- [ ] Team notified of launch
- [ ] Reporting dashboard ready
FILE:references/scoring-system.md
# Ad Account Scoring System
Reference for `ad_health_scorer.py`. Defines the weighted scoring algorithm, severity multipliers, and platform-specific category weights.
## Scoring formula
```
Category_Score = Σ(Check_Result × Severity_Multiplier) / Σ(Severity_Multiplier) × 100
Platform_Score = Σ(Category_Score × Category_Weight)
Aggregate_Score = Σ(Platform_Score × Budget_Share)
```
## Severity multipliers
| Severity | Multiplier | Meaning | SLA |
|---|---|---|---|
| Critical | 5.0x | Blocks revenue or burns budget | Fix immediately |
| High | 3.0x | Significant performance impact | Fix within 1 week |
| Medium | 1.5x | Optimization opportunity | Fix within 1 month |
| Low | 0.5x | Polish / best practice | Backlog |
Critical issues dominate the score. A single critical failure drops the category score significantly, which is the correct behavior — a missing conversion tag invalidates everything downstream.
## Platform category weights
### Google Ads
| Category | Weight | Key checks |
|---|---|---|
| Conversion Tracking | 25% | Tag installed, Enhanced Conversions, attribution model, conversion window |
| Wasted Spend | 20% | Negative keywords, search terms review, broad match rules, 3× CPA kill rule |
| Account Structure | 15% | Naming conventions, ad group size, campaign types |
| Keywords | 15% | Quality Score, duplicates, match types, search intent alignment |
| Ads | 15% | RSA headlines count, extensions, A/B testing |
| Settings | 10% | Location targeting, schedules, networks, bidding strategy |
### Meta (Facebook/Instagram)
| Category | Weight | Key checks |
|---|---|---|
| Pixel & CAPI | 30% | Pixel installed, CAPI active, event deduplication, domain verification |
| Creative | 30% | Format diversity, fatigue detection, safe zones, copy length |
| Structure | 20% | CBO, campaign naming, advantage+ settings |
| Audience | 20% | Lookalike seed size, exclusions, overlap, custom audiences |
### LinkedIn
| Category | Weight | Key checks |
|---|---|---|
| Technical | 25% | Insight tag, conversion events, matched audiences |
| Targeting | 25% | Audience size, job function vs title, company lists |
| Creative | 25% | Format mix, single-image vs carousel vs video, CTA alignment |
| Budget | 25% | Daily budget sufficiency, bid strategy, pacing |
### TikTok
| Category | Weight | Key checks |
|---|---|---|
| Pixel | 25% | Pixel installed, events configured, match quality |
| Creative | 30% | Native-feel content, format mix, hook rate (3s), UGC ratio |
| Targeting | 25% | Interest vs behavior, custom audiences, lookalikes |
| Budget | 20% | Learning phase budget (50× target CPA), pacing |
## Grade bands
| Grade | Score | Meaning |
|---|---|---|
| A | 90-100 | Excellent — maintain and scale |
| B | 75-89 | Good — address high-priority items |
| C | 60-74 | Needs work — systematic improvements needed |
| D | 40-59 | Poor — significant issues blocking performance |
| F | <40 | Critical — account needs fundamental restructuring |
Bands are calibrated wider than SEO scoring because ad accounts typically have more actionable but non-critical issues (e.g., missing extensions, suboptimal ad copy).
## Quick Wins formula
```
Quick Win = severity ∈ {critical, high} AND result = "warn" (not full fail)
```
Quick wins are issues that are important (high severity) but partially working (warn, not fail) — meaning the fix is usually small: enable a toggle, add a few negative keywords, activate an extension.
## Hard rules (quality gates)
These combinations should NEVER be recommended together:
- Broad Match + Manual CPC (wastes budget without smart bidding control)
- CPA target below $5 with < $50/day budget (can't exit learning phase)
- Conversion action = page view as primary (inflates numbers, misleads bidding)
The scorer doesn't enforce these directly but the SKILL.md workflow should flag them as critical failures.
FILE:scripts/ad_health_scorer.py
#!/usr/bin/env python3
"""
ad_health_scorer.py — Weighted 0-100 ad account health score with multi-platform support.
Scores ad accounts across platform-specific categories with severity multipliers
and budget-weighted cross-platform aggregation.
Severity multipliers:
critical = 5x weight (blocks revenue or burns budget)
high = 3x weight (significant impact)
medium = 1.5x weight (optimization opportunity)
low = 0.5x weight (backlog polish)
Platform category weights:
Google: Conversion Tracking 25%, Wasted Spend 20%, Structure 15%, Keywords 15%, Ads 15%, Settings 10%
Meta: Pixel/CAPI 30%, Creative 30%, Structure 20%, Audience 20%
LinkedIn: Technical 25%, Targeting 25%, Creative 25%, Budget 25%
TikTok: Pixel 25%, Creative 30%, Targeting 25%, Budget 20%
Cross-platform aggregation:
Aggregate Score = Σ(Platform_Score × Platform_Budget_Share)
Grade bands (calibrated wider — ad accounts naturally score lower):
A = 90-100, B = 75-89, C = 60-74, D = 40-59, F = <40
Usage:
python ad_health_scorer.py --checks checks.json
python ad_health_scorer.py --checks checks.json --platform google --budget 5000
python ad_health_scorer.py --multi platforms.json # multi-platform aggregation
python ad_health_scorer.py --demo
python ad_health_scorer.py --demo --json
"""
from __future__ import annotations
import argparse
import json
import sys
from collections import defaultdict
from pathlib import Path
SEVERITY_MULTIPLIER = {"critical": 5.0, "high": 3.0, "medium": 1.5, "low": 0.5}
PLATFORM_WEIGHTS = {
"google": {
"conversion_tracking": 0.25,
"wasted_spend": 0.20,
"account_structure": 0.15,
"keywords": 0.15,
"ads": 0.15,
"settings": 0.10,
},
"meta": {
"pixel_capi": 0.30,
"creative": 0.30,
"structure": 0.20,
"audience": 0.20,
},
"linkedin": {
"technical": 0.25,
"targeting": 0.25,
"creative": 0.25,
"budget": 0.25,
},
"tiktok": {
"pixel": 0.25,
"creative": 0.30,
"targeting": 0.25,
"budget": 0.20,
},
}
DEMO_CHECKS = {
"google": [
{"category": "conversion_tracking", "check": "Google Ads conversion tag installed", "result": "pass", "severity": "critical"},
{"category": "conversion_tracking", "check": "Enhanced Conversions enabled", "result": "fail", "severity": "critical", "detail": "Missing enhanced conversions — losing 15-30% attribution"},
{"category": "conversion_tracking", "check": "Conversion window appropriate", "result": "pass", "severity": "medium"},
{"category": "wasted_spend", "check": "Negative keyword coverage", "result": "warn", "severity": "high", "detail": "Only 12 negative keywords — review search terms report"},
{"category": "wasted_spend", "check": "No broad match + manual CPC", "result": "pass", "severity": "critical"},
{"category": "wasted_spend", "check": "Search terms review (last 30d)", "result": "fail", "severity": "high", "detail": "23% of spend on irrelevant terms"},
{"category": "account_structure", "check": "Campaign naming convention", "result": "pass", "severity": "low"},
{"category": "account_structure", "check": "Ad groups ≤ 20 keywords each", "result": "warn", "severity": "medium", "detail": "2 ad groups with 30+ keywords"},
{"category": "keywords", "check": "No duplicate keywords across campaigns", "result": "pass", "severity": "high"},
{"category": "keywords", "check": "Quality Score ≥ 6 on top spenders", "result": "warn", "severity": "high", "detail": "3 keywords with QS 4-5"},
{"category": "ads", "check": "RSA with ≥ 3 headlines", "result": "pass", "severity": "medium"},
{"category": "ads", "check": "Ad extensions active (sitelinks, callouts)", "result": "fail", "severity": "medium", "detail": "No callout extensions"},
{"category": "settings", "check": "Location targeting correct", "result": "pass", "severity": "high"},
{"category": "settings", "check": "Ad schedule aligned with business hours", "result": "pass", "severity": "low"},
],
"meta": [
{"category": "pixel_capi", "check": "Meta Pixel installed", "result": "pass", "severity": "critical"},
{"category": "pixel_capi", "check": "Conversions API (CAPI) active", "result": "fail", "severity": "critical", "detail": "No server-side events — degraded attribution post-iOS14"},
{"category": "creative", "check": "Creative diversity (≥ 3 formats)", "result": "warn", "severity": "high", "detail": "Only static images — add video and carousel"},
{"category": "creative", "check": "No creative fatigue (CTR stable)", "result": "pass", "severity": "high"},
{"category": "structure", "check": "CBO enabled", "result": "pass", "severity": "medium"},
{"category": "audience", "check": "Lookalike seed ≥ 1000 users", "result": "pass", "severity": "medium"},
],
}
def score_platform(checks, platform):
weights = PLATFORM_WEIGHTS.get(platform, {})
by_category = defaultdict(list)
for c in checks:
by_category[c.get("category", "other")].append(c)
category_scores = {}
findings = []
quick_wins = []
for cat, cat_checks in by_category.items():
weighted_pass = 0.0
weighted_total = 0.0
for check in cat_checks:
result = check.get("result", "fail")
severity = check.get("severity", "medium")
mult = SEVERITY_MULTIPLIER.get(severity, 1.0)
score = {"pass": 1.0, "warn": 0.5, "fail": 0.0}.get(result, 0.0)
weighted_pass += score * mult
weighted_total += mult
if result != "pass":
finding = {
"platform": platform,
"category": cat,
"check": check.get("check", ""),
"result": result,
"severity": severity,
"detail": check.get("detail", ""),
}
findings.append(finding)
# Quick win: high/critical severity + warn (not full fail)
if severity in ("critical", "high") and result == "warn":
quick_wins.append(finding)
cat_score = (weighted_pass / weighted_total * 100) if weighted_total > 0 else 100
category_scores[cat] = round(cat_score, 1)
# Weighted overall
overall = 0.0
total_weight = 0.0
for cat, weight in weights.items():
if cat in category_scores:
overall += category_scores[cat] * weight
total_weight += weight
overall = (overall / total_weight) if total_weight > 0 else 0.0
if overall >= 90:
grade = "A"
elif overall >= 75:
grade = "B"
elif overall >= 60:
grade = "C"
elif overall >= 40:
grade = "D"
else:
grade = "F"
findings.sort(key=lambda f: {"critical": 0, "high": 1, "medium": 2, "low": 3}.get(f["severity"], 99))
return {
"platform": platform,
"overall_score": round(overall, 1),
"grade": grade,
"category_scores": category_scores,
"total_checks": len(checks),
"passed": sum(1 for c in checks if c.get("result") == "pass"),
"warnings": sum(1 for c in checks if c.get("result") == "warn"),
"failures": sum(1 for c in checks if c.get("result") == "fail"),
"findings": findings,
"quick_wins": quick_wins,
}
def aggregate_platforms(platform_results, budgets=None):
if not budgets:
# Equal weight
budgets = {p["platform"]: 1.0 / len(platform_results) for p in platform_results}
total_budget = sum(budgets.values())
shares = {k: v / total_budget for k, v in budgets.items()}
aggregate = 0.0
for pr in platform_results:
share = shares.get(pr["platform"], 0)
aggregate += pr["overall_score"] * share
return {
"aggregate_score": round(aggregate, 1),
"budget_shares": {k: round(v, 2) for k, v in shares.items()},
"platform_scores": {pr["platform"]: pr["overall_score"] for pr in platform_results},
}
def print_report(result):
print(f"Ad Health Score ({result['platform'].upper()}): {result['overall_score']}/100 (Grade: {result['grade']})")
print(f"Checks: {result['total_checks']} — {result['passed']} pass, {result['warnings']} warn, {result['failures']} fail")
print()
print("Category Breakdown:")
for cat, score in sorted(result["category_scores"].items()):
bar = "█" * int(score / 5) + "░" * (20 - int(score / 5))
print(f" {cat:25s} {bar} {score:5.1f}/100")
print()
if result["quick_wins"]:
print(f"Quick Wins ({len(result['quick_wins'])}):")
for f in result["quick_wins"]:
print(f" ⚡ [{f['severity'].upper()}] {f['check']}: {f['detail']}")
print()
if result["findings"]:
print(f"Findings ({len(result['findings'])}):")
for f in result["findings"]:
detail = f" — {f['detail']}" if f["detail"] else ""
print(f" [{f['severity'].upper()}/{f['result'].upper()}] {f['check']}{detail}")
def main():
p = argparse.ArgumentParser(
description="Compute weighted 0-100 ad account health score with severity multipliers.",
epilog="Supports Google, Meta, LinkedIn, TikTok. Run with --demo for a sample report.",
)
p.add_argument("--checks", help="Path to checks JSON file (array of check objects)")
p.add_argument("--platform", choices=list(PLATFORM_WEIGHTS.keys()), default="google")
p.add_argument("--budget", type=float, default=None, help="Monthly budget (for multi-platform weighting)")
p.add_argument("--multi", help="Path to multi-platform JSON {platform: {checks: [...], budget: N}}")
p.add_argument("--json", action="store_true", help="JSON output")
p.add_argument("--demo", action="store_true", help="Run with demo data")
args = p.parse_args()
if args.demo:
results = []
for platform, checks in DEMO_CHECKS.items():
results.append(score_platform(checks, platform))
agg = aggregate_platforms(results, {"google": 3000, "meta": 2000})
if args.json:
print(json.dumps({"platforms": results, "aggregate": agg}, indent=2))
else:
for r in results:
print_report(r)
print()
print(f"Cross-Platform Aggregate: {agg['aggregate_score']}/100")
print(f"Budget shares: {agg['budget_shares']}")
return
if args.multi:
data = json.loads(Path(args.multi).read_text())
results = []
budgets = {}
for platform, pdata in data.items():
results.append(score_platform(pdata["checks"], platform))
budgets[platform] = pdata.get("budget", 1000)
agg = aggregate_platforms(results, budgets)
if args.json:
print(json.dumps({"platforms": results, "aggregate": agg}, indent=2))
else:
for r in results:
print_report(r)
print()
print(f"Cross-Platform Aggregate: {agg['aggregate_score']}/100")
return
if args.checks:
checks = json.loads(Path(args.checks).read_text())
result = score_platform(checks, args.platform)
if args.json:
print(json.dumps(result, indent=2))
else:
print_report(result)
return
p.print_help()
if __name__ == "__main__":
main()
FILE:scripts/roas_calculator.py
#!/usr/bin/env python3
"""
roas_calculator.py — ROAS and paid-ads metrics calculator
Usage:
python3 roas_calculator.py --spend 5000 --revenue 18000 --conversions 120 --leads 400 --margin 40
python3 roas_calculator.py --file campaign.json
python3 roas_calculator.py --json # demo + JSON output
python3 roas_calculator.py # demo mode
"""
import argparse
import json
import sys
# ---------------------------------------------------------------------------
# Calculation core
# ---------------------------------------------------------------------------
def calculate(spend: float, revenue: float = 0.0, conversions: int = 0,
leads: int = 0, margin_pct: float = 0.0,
impressions: int = 0, clicks: int = 0) -> dict:
results = {
"inputs": {
"ad_spend": spend,
"revenue": revenue,
"conversions": conversions,
"leads": leads,
"margin_pct": margin_pct,
"impressions": impressions,
"clicks": clicks,
}
}
metrics = {}
# --- ROAS ---
if revenue > 0 and spend > 0:
roas = revenue / spend
metrics["roas"] = {
"value": round(roas, 2),
"formula": "revenue / ad_spend",
"interpretation": _roas_label(roas),
}
# --- Break-even ROAS ---
if margin_pct > 0:
be_roas = 100 / margin_pct
metrics["break_even_roas"] = {
"value": round(be_roas, 2),
"formula": "100 / margin_%",
"note": f"Need {be_roas:.1f}x ROAS to cover ad costs at {margin_pct}% margin",
}
if revenue > 0:
actual_roas = revenue / spend
profitable = actual_roas >= be_roas
metrics["profitability"] = {
"is_profitable": profitable,
"gap": round(actual_roas - be_roas, 2),
"note": "Profitable ✅" if profitable else f"Unprofitable ❌ — need +{be_roas - actual_roas:.2f}x ROAS",
}
# --- CPA ---
if conversions > 0 and spend > 0:
cpa = spend / conversions
metrics["cpa"] = {
"value": round(cpa, 2),
"formula": "ad_spend / conversions",
"unit": "cost per acquisition",
}
if revenue > 0:
rev_per_conversion = revenue / conversions
metrics["revenue_per_conversion"] = {
"value": round(rev_per_conversion, 2),
"roi_per_conversion": round((rev_per_conversion - cpa) / cpa * 100, 1),
}
# --- CPL ---
if leads > 0 and spend > 0:
cpl = spend / leads
metrics["cpl"] = {
"value": round(cpl, 2),
"formula": "ad_spend / leads",
"unit": "cost per lead",
}
if conversions > 0:
lead_to_conv_rate = conversions / leads * 100
metrics["lead_to_conversion_rate"] = {
"value": round(lead_to_conv_rate, 1),
"unit": "%",
}
# --- Conversion rate ---
if clicks > 0 and conversions > 0:
cvr = conversions / clicks * 100
metrics["conversion_rate"] = {
"value": round(cvr, 2),
"unit": "%",
"benchmark": "2-5% typical for paid search",
}
if clicks > 0 and leads > 0:
lcr = leads / clicks * 100
metrics["lead_capture_rate"] = {
"value": round(lcr, 2),
"unit": "%",
}
# --- CTR ---
if impressions > 0 and clicks > 0:
ctr = clicks / impressions * 100
metrics["ctr"] = {
"value": round(ctr, 2),
"unit": "%",
"benchmark": "2-5% for search, 0.1-0.5% for display",
}
cpm = spend / impressions * 1000
metrics["cpm"] = {
"value": round(cpm, 2),
"unit": "cost per 1000 impressions",
}
cpc = spend / clicks
metrics["cpc"] = {
"value": round(cpc, 2),
"unit": "cost per click",
}
results["metrics"] = metrics
results["recommendations"] = _recommendations(metrics, spend, margin_pct)
return results
def _roas_label(roas: float) -> str:
if roas >= 8:
return "Excellent (8x+)"
if roas >= 5:
return "Strong (5-8x)"
if roas >= 3:
return "Good (3-5x)"
if roas >= 2:
return "Acceptable (2-3x) — check margins"
if roas >= 1:
return "Below target (<2x) — likely unprofitable"
return "Losing money (<1x)"
def _recommendations(metrics: dict, spend: float, margin_pct: float) -> list:
recs = []
roas = metrics.get("roas", {}).get("value")
be_roas = metrics.get("break_even_roas", {}).get("value")
if roas and be_roas:
if roas < be_roas:
shortfall = round((be_roas - roas) * spend, 2)
recs.append(f"⚠️ Losing ,.2f/period — pause or restructure campaign immediately")
elif roas < be_roas * 1.5:
recs.append("⚠️ Marginally profitable — optimize creatives and targeting before scaling")
else:
recs.append("✅ Profitable — consider increasing budget or duplicating campaign")
cpa = metrics.get("cpa", {}).get("value")
cpl = metrics.get("cpl", {}).get("value")
cvr = metrics.get("conversion_rate", {}).get("value")
if cvr and cvr < 2:
recs.append(f"⚠️ CVR {cvr}% is low — test new landing pages, headlines, and CTAs")
elif cvr and cvr >= 5:
recs.append(f"✅ Strong CVR {cvr}% — maximize traffic to this funnel")
if cpa and cpl:
l2c = metrics.get("lead_to_conversion_rate", {}).get("value", 0)
if l2c < 10:
recs.append(f"⚠️ Lead-to-close rate {l2c}% is low — review sales qualification or nurture sequence")
ctr = metrics.get("ctr", {}).get("value")
if ctr:
if ctr < 1:
recs.append(f"⚠️ CTR {ctr}% is low — refresh ad copy and audience targeting")
elif ctr >= 5:
recs.append(f"✅ High CTR {ctr}% — strong creative, ensure LP matches ad message")
if not recs:
recs.append("Add more data (margin %, impressions, leads) for actionable recommendations")
return recs
# ---------------------------------------------------------------------------
# Demo data
# ---------------------------------------------------------------------------
DEMO_DATA = {
"spend": 8500,
"revenue": 34200,
"conversions": 142,
"leads": 680,
"margin_pct": 35,
"impressions": 185000,
"clicks": 3700,
}
# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
def main():
parser = argparse.ArgumentParser(
description="ROAS calculator — paid ads performance metrics and recommendations."
)
parser.add_argument("--spend", type=float, help="Total ad spend ($)")
parser.add_argument("--revenue", type=float, default=0, help="Total attributed revenue ($)")
parser.add_argument("--conversions", type=int, default=0, help="Number of purchases/conversions")
parser.add_argument("--leads", type=int, default=0, help="Number of leads generated")
parser.add_argument("--margin", type=float, default=0, help="Gross margin %% (e.g. 40)")
parser.add_argument("--impressions", type=int, default=0, help="Total impressions")
parser.add_argument("--clicks", type=int, default=0, help="Total clicks")
parser.add_argument("--file", help="JSON file with campaign data")
parser.add_argument("--json", action="store_true", help="Output as JSON")
args = parser.parse_args()
if args.file:
with open(args.file, "r") as f:
data = json.load(f)
elif args.spend:
data = {
"spend": args.spend,
"revenue": args.revenue,
"conversions": args.conversions,
"leads": args.leads,
"margin_pct": args.margin,
"impressions": args.impressions,
"clicks": args.clicks,
}
else:
data = DEMO_DATA
if not args.json:
print("No input provided — running in demo mode.\n")
result = calculate(
spend=data.get("spend", 0),
revenue=data.get("revenue", 0),
conversions=data.get("conversions", 0),
leads=data.get("leads", 0),
margin_pct=data.get("margin_pct", 0),
impressions=data.get("impressions", 0),
clicks=data.get("clicks", 0),
)
if args.json:
print(json.dumps(result, indent=2))
return
inp = result["inputs"]
metrics = result["metrics"]
recs = result["recommendations"]
print("=" * 62)
print(" PAID ADS PERFORMANCE REPORT")
print("=" * 62)
print(f" Spend: >10,.2f")
if inp["revenue"]: print(f" Revenue: >10,.2f")
if inp["conversions"]:print(f" Conversions:{inp['conversions']:>10}")
if inp["leads"]: print(f" Leads: {inp['leads']:>10}")
if inp["impressions"]:print(f" Impressions:{inp['impressions']:>10,}")
if inp["clicks"]: print(f" Clicks: {inp['clicks']:>10,}")
print()
print(" METRICS")
print(" " + "─" * 58)
metric_labels = [
("roas", "ROAS", lambda m: f"{m['value']}x — {m['interpretation']}"),
("break_even_roas", "Break-even ROAS", lambda m: f"{m['value']}x — {m['note']}"),
("profitability", "Profitability", lambda m: m['note']),
("cpa", "CPA", lambda m: f",.2f / {m['unit']}"),
("revenue_per_conversion", "Rev/Conversion", lambda m: f",.2f (ROI {m['roi_per_conversion']}%)"),
("cpl", "CPL", lambda m: f",.2f / {m['unit']}"),
("lead_to_conversion_rate","Lead→Conv Rate", lambda m: f"{m['value']}%"),
("conversion_rate", "Conversion Rate", lambda m: f"{m['value']}% ({m['benchmark']})"),
("ctr", "CTR", lambda m: f"{m['value']}%"),
("cpc", "CPC", lambda m: f",.2f"),
("cpm", "CPM", lambda m: f",.2f"),
]
for key, label, fmt in metric_labels:
if key in metrics:
try:
detail = fmt(metrics[key])
print(f" {label:<24} {detail}")
except Exception:
pass
print()
print(" RECOMMENDATIONS")
print(" " + "─" * 58)
for rec in recs:
print(f" {rec}")
print("=" * 62)
if __name__ == "__main__":
main()
Tạo persona người dùng dựa trên dữ liệu cho nghiên cứu UX và thiết kế sản phẩm.
--- name: persona description: Generate data-driven user personas for UX research and product design. Usage: /persona generate [options] --- # /persona Generate structured user personas with demographics, goals, pain points, and behavioral patterns. ## Usage ``` /persona generate Generate persona (interactive) /persona generate json Generate persona as JSON ``` ## Input Format Interactive mode prompts for product context. Alternatively, provide context inline: ``` /persona generate > Product: B2B project management tool > Target: Engineering managers at mid-size companies > Key problem: Cross-team visibility ``` ## Examples ``` /persona generate /persona generate json /persona generate json > persona-eng-manager.json ``` ## Scripts - `product-team/ux-researcher-designer/scripts/persona_generator.py` — Persona generator (positional `json` arg for JSON output) ## Skill Reference > `product-team/ux-researcher-designer/SKILL.md`
Nhận diện stack công nghệ và sinh cấu hình pipeline CI/CD.
--- name: pipeline description: Detect stack and generate CI/CD pipeline configs. Usage: /pipeline <detect|generate> [options] --- # /pipeline Detect project stack and generate CI/CD pipeline configurations for GitHub Actions or GitLab CI. ## Usage ``` /pipeline detect [--repo <project-dir>] Detect stack, tools, and services /pipeline generate --platform github|gitlab [--repo <project-dir>] Generate pipeline YAML ``` ## Examples ``` /pipeline detect --repo ./my-project /pipeline generate --platform github --repo . /pipeline generate --platform gitlab --repo . ``` ## Scripts - `engineering/ci-cd-pipeline-builder/scripts/stack_detector.py` — Detect stack and tooling (`--repo <path>`, `--format text|json`) - `engineering/ci-cd-pipeline-builder/scripts/pipeline_generator.py` — Generate pipeline YAML (`--platform github|gitlab`, `--repo <path>`, `--input <stack.json>`, `--output <file>`) ## Skill Reference → `engineering/ci-cd-pipeline-builder/SKILL.md`
Hỗ trợ lập kế hoạch và chia nhỏ công việc thành các bước thực hiện.
# Planner Agent ## Vai trò Tác nhân lên kế hoạch — chịu trách nhiệm tổ chức, ưu tiên và phân bổ nguồn lực cho công việc cá nhân. ## Nhiệm vụ chính - Tiếp nhận danh sách công việc từ người dùng - Phân loại và ưu tiên theo ma trận Eisenhower - Lập timeline thực tế có buffer - Gợi ý khung giờ làm việc phù hợp ## Đầu vào - Danh sách việc cần làm - Deadline - Mức năng lượng dự kiến ## Đầu ra - Kế hoạch ngày/tuần có thứ tự ưu tiên - 3 việc quan trọng nhất (MIT) - Danh sách việc nên hoãn hoặc bỏ - Định dạng markdown có thể lưu vào Drive ## Tiêu chí đánh giá - Kế hoạch có thực tế không? (không nhồi nhét) - Có tính đến phục hồi năng lượng không? - Có xử lý được nếu phát sinh thêm việc không? ## Phối hợp Sau khi lên kế hoạch, chuyển sang QA Reviewer để kiểm tra tính khả thi.
Bộ 6 skill quản lý dự án: PM cấp cao, scrum master, chuyên gia Jira (JQL), Confluence, quản trị Atlassian, tạo template, tích hợp MCP với Jira/Confluence.
--- name: "pm-skills" description: "6 project management agent skills and plugins for Claude Code, Codex, Gemini CLI, Cursor, OpenClaw. Senior PM, scrum master, Jira expert (JQL), Confluence expert, Atlassian admin, template creator. MCP integration for live Jira/Confluence automation." version: 2.9.0 author: Alireza Rezvani license: MIT tags: - project-management - jira - confluence - atlassian - scrum - agile agents: - claude-code - codex-cli - openclaw --- # Project Management Skills 6 production-ready project management skills with Atlassian MCP integration. ## Quick Start ### Claude Code ``` /read project-management/jira-expert/SKILL.md ``` ### Codex CLI ```bash npx agent-skills-cli add alirezarezvani/claude-skills/project-management ``` ## Skills Overview | Skill | Folder | Focus | |-------|--------|-------| | Senior PM | `senior-pm/` | Portfolio management, risk analysis, resource planning | | Scrum Master | `scrum-master/` | Velocity forecasting, sprint health, retrospectives | | Jira Expert | `jira-expert/` | JQL queries, workflows, automation, dashboards | | Confluence Expert | `confluence-expert/` | Knowledge bases, page layouts, macros | | Atlassian Admin | `atlassian-admin/` | User management, permissions, integrations | | Atlassian Templates | `atlassian-templates/` | Blueprints, custom layouts, reusable content | ## Python Tools 6 scripts, all stdlib-only: ```bash python3 senior-pm/scripts/project_health_dashboard.py --help python3 scrum-master/scripts/velocity_analyzer.py --help ``` ## Rules - Load only the specific skill SKILL.md you need - Use MCP tools for live Jira/Confluence operations when available
Rà soát chi tiêu SaaS hằng năm, phân loại chi tiêu theo danh mục, phân tích chu kỳ mua và hợp nhất nhà cung cấp cân bằng rủi ro.
---
name: procurement-optimizer
description: Use when running an annual SaaS audit, doing category-level spend review, or rationalizing the supplier base — when the user needs to do a spend audit, spend categorization (UNSPSC-aligned), purchasing-cycle analysis, or risk-balanced supplier consolidation. Triggers on "spend audit", "SaaS audit", "spend categorization", "supplier rationalization", "supplier consolidation", "purchasing cycle", "procurement review", "category strategy", "duplicate SaaS", "renewal cluster". Ships 3 stdlib-only Python tools (UNSPSC-aligned spend categorizer with Pareto breakdown and industry profiles, purchasing-cycle analyzer that surfaces bottleneck categories per Goldratt's Theory of Constraints, supplier-consolidation planner that refuses single-source recommendations for tier-1 categories without a documented break-glass plan), 3 reference docs each citing 7+ authoritative sources (A.T. Kearney / Hackett / Spend Matters / UNSPSC / Productiv / Vendr / Tropic / IACCM / ISM / BCG), and a 20-minute spend-intake template. Distinct from sibling vendor-management (performance scoring of vendors you keep paying), finance/financial-analysis (close + report, not category strategy), and c-level-advisor/general-counsel-advisor (contract law, not category rationalization).
version: 2.8.0
author: claude-code-skills
license: MIT
tags: [bizops, procurement, spend-categorization, supplier-consolidation, unspsc, saas-audit, purchasing-cycle]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# Procurement Optimizer — Spend Categorization + Supplier Rationalization
You are a Head of Procurement / Head of BizOps / VP Finance operator running the annual category review. Your job is **what to buy, from whom, on what cadence** — not how the vendor you already chose is performing (that's `vendor-management`). You categorize spend along a UNSPSC-aligned taxonomy, find the Pareto-20% of categories driving 80% of cost, surface purchasing-cycle bottlenecks, and produce a **risk-balanced** supplier-consolidation plan that refuses to collapse tier-1 categories to single-source without a documented contingency.
## Purpose
A typical mid-stage company has:
- Software spend up 40% YoY with no single owner who can name the top growth categories.
- 3 monitoring tools, 2 expense platforms, 4 email-marketing tools — duplicate-function clusters that nobody consolidated because no one had the data to defend the recommendation.
- A purchasing cycle where some categories close in 5 days and others take 90, but the "average" hides the constraint.
- Renewal dates clustered in the same month, destroying negotiation leverage.
This skill produces a deterministic, defensible artifact for each problem: categorized spend with Pareto, cycle-time scorecard by category, and a consolidation plan with explicit risk flags.
## When to use
- Annual SaaS audit and category-level spend review.
- A category owner wants to know which 5 categories drove this year's spend growth.
- Finance flags that software spend is up 40% YoY and needs a Pareto by category, not by vendor.
- BizOps suspects duplicate-function tools (monitoring, expense, email-marketing) and needs a defensible consolidation plan.
- The CFO wants tighter approval thresholds and needs cycle-time data per category to justify it.
- Post-acquisition, two procurement teams need to merge category taxonomies and dedupe the supplier base.
## When NOT to use
- Scoring or auditing an individual vendor you've already decided to keep paying → sibling `vendor-management`.
- Financial close, monthly reporting, or P&L analysis → `finance/financial-analysis`.
- Drafting or negotiating contract terms → `c-level-advisor/general-counsel-advisor`.
- Building outbound sales proposals → `business-growth/contract-and-proposal-writer`.
## Workflow
### Step 1 — Intake spend
Have the user fill out `assets/spend_intake_template.md` (20 minutes for a typical mid-stage company). The skeleton expects line items with `{supplier, description, category_hint, annual_spend, frequency, currency}`. If prior-year spend is available, include it for YoY analysis.
### Step 2 — Categorize and find the Pareto
Run `scripts/spend_categorizer.py --input spend.json --profile <profile> --output categorized.md`.
The categorizer maps each line item to a UNSPSC-aligned Class → Family → Segment (built-in map of ~30 categories tuned for tech-startup spend: Software/SaaS, Hardware, Cloud Infrastructure, Professional Services, Marketing Services, Legal, Recruiting, Travel, Office, Insurance, Benefits, etc. — NOT the full 100k UNSPSC database). Output includes:
- Categorized line items
- Pareto: which 20% of categories drive 80% of spend?
- Top-10 YoY growth categories (when prior-year provided)
Profiles re-prioritize the category map: `tech-startup` (heavy SaaS / cloud), `scaleup` (sales tools / recruiting heavy), `enterprise` (professional services / facilities heavy), `services`, `manufacturing`.
### Step 3 — Analyze the purchasing cycle
Run `scripts/purchasing_cycle_analyzer.py --input pos.json --output cycle.md`.
For each PO record `{category, request_date, approval_date, po_issued_date, goods_received_date, payment_date, approver_hops}`, the analyzer computes per-category:
- Cycle time T-request → T-PO (median, P90)
- T-PO → T-pay (median, P90)
- Approver-hop count (median)
It then flags categories with cycle time > 2× the cross-category median as **bottleneck** categories. This is Goldratt's Theory of Constraints applied to procurement: the system throughput is set by the slowest step, and the slowest step is almost always one specific category (legal review on services contracts, security review on tier-1 SaaS).
### Step 4 — Plan supplier consolidation with risk balancing
Run `scripts/supplier_consolidation.py --input suppliers.json --profile <profile> --output consolidation_plan.md`.
The planner identifies **duplicate-function clusters** (e.g., 3 monitoring tools, 2 expense platforms). For each cluster:
- Picks a recommended consolidation winner (highest criticality tier survives, OR lowest switching-cost winner if the cluster is tier-3, depending on cluster type).
- **Flags risk:** does NOT recommend collapse to single-source for any tier-1 criticality category unless the input explicitly flags a documented break-glass plan. The output says explicitly: "DO NOT CONSOLIDATE — tier-1 cluster, no break-glass on record. Add a 72-hour contingency plan first."
- Estimates savings: current cluster spend − winner spend − migration cost (sum of switching-cost estimates of losers).
- Renewal-date clustering analysis: flags categories where ≥ 3 contracts renew within the same calendar month (no leverage).
### Step 5 — Synthesize the procurement review
Combine the 3 artifacts into a BizOps-ready digest:
- Top 5 categories driving YoY spend growth (categorizer)
- Top 3 bottleneck categories blocking throughput (cycle analyzer)
- Top 5 consolidation opportunities with estimated savings and risk flags (consolidation planner)
- All renewal clusters destroying leverage
- Tier-1 single-source exposure points needing break-glass plans before any consolidation
## Scripts
| Script | Purpose |
|---|---|
| `scripts/spend_categorizer.py` | UNSPSC-aligned categorization + Pareto + YoY growth |
| `scripts/purchasing_cycle_analyzer.py` | Per-category cycle time + Goldratt bottleneck flag |
| `scripts/supplier_consolidation.py` | Duplicate-function clustering + risk-flagged consolidation plan |
All three accept `--input` (JSON), `--output` (markdown path), `--sample` (run with built-in sample data), and `--help`. The two with industry-specific category priorities accept `--profile {tech-startup,scaleup,enterprise,services,manufacturing}`.
## References
- `references/spend_management_canon.md` — A.T. Kearney *Spend Management*, Procurement Leaders, Gartner Procurement, BCG Procurement value creation, Hackett benchmarks, Pierre Mitchell / Spend Matters, UNSPSC official taxonomy.
- `references/saas_management_canon.md` — Productiv / Zylo / Vendr / Tropic SaaS sprawl reports, BetterCloud SaaS Operations, Gartner SMP Magic Quadrant, Bain SaaS spend, Forrester SaaS portfolio management, Tomasz Tunguz on SaaS sprawl, Patrick Campbell / ProfitWell on SaaS unit economics.
- `references/procurement_anti_patterns.md` — A.T. Kearney maverick-spend, IACCM/WorldCC, McKinsey on category-strategy mistakes, Hackett purchasing-cycle research, BCG on supplier-consolidation risks, Spend Matters failed-rationalization analyses, ISM lessons learned.
## Assumptions
1. The user has access to AP / expense / SaaS-management exports, or can hand-assemble a spend list of the top 100-200 line items (the Pareto holds — top 20% of suppliers will be most of the spend).
2. Prior-year spend is preferred (for YoY) but optional; the categorizer degrades gracefully if absent.
3. Purchasing-cycle data is preferred but optional; if absent, the user gets categorization + consolidation only.
4. Supplier criticality (`tier-1/2/3`) is a **judgment call by the user**, not derived from spend alone. Tier-1 = revenue-blocking if the supplier disappears. The tool refuses to infer this — the user must mark it.
5. The output artifacts (categorized markdown, cycle scorecard, consolidation plan) are **inputs to a human decision**, not the decision itself.
## Anti-patterns
- **Consolidate to single-source for tier-1 critical category without a break-glass plan.** Cost savings buy nothing if the consolidated supplier disappears. See `references/procurement_anti_patterns.md`.
- **Categorize by vendor name, not by what's purchased.** Workday could be "HR Software" OR "Finance Software" depending on which modules are licensed. The line-item `description` and `category_hint` drive categorization, not the supplier name.
- **Ignore renewal-date clustering.** Twelve tier-2 contracts that all renew in March mean zero negotiation leverage on any of them. Spread them.
- **Approve-by-default for sub-$5K spend.** This is the death-by-a-thousand-SaaS pattern. The categorizer surfaces "small-spend, many-supplier" clusters explicitly.
- **No quarterly renewal review.** Annual is too coarse for SaaS, which renews continuously across the year.
- **Rationalize without measuring switching cost.** Consolidating 3 tools to save $50k when migration costs $200k is not a savings.
- **Consolidate based on price alone, ignoring integration debt.** The cheap tool that doesn't integrate with your data warehouse is more expensive than the expensive one that does.
- **Treat shadow IT spend as marketing's problem.** It is procurement's problem. Marketing-tool sprawl is the #1 driver of SaaS-spend growth in scaleups.
## Distinct from
- **Sibling `vendor-management`** — that's performance scoring (uptime, SLA, third-party risk) for vendors you've already decided to keep paying. This is **spend rationalization + supplier consolidation** — deciding WHICH vendors to keep.
- **`finance/financial-analysis`** — that's financial close, P&L, reporting, DCF. This is operational procurement: category strategy and supplier rationalization, not financial reporting.
- **`c-level-advisor/general-counsel-advisor`** — that's contract law (indemnity, IP, liquidated damages). This is category-level spend strategy. Once you've decided which 3 monitoring tools to consolidate to 1, GC reviews the contract terms of the survivor.
- **`business-growth/contract-and-proposal-writer`** — that's outbound proposals to win customers. This is inbound supplier rationalization.
- **`finance/budgeting`** — that's annual budget planning. This is the inside view: where the budget is actually leaking.
## Forcing-question library (Matt Pocock grill discipline)
Walked one at a time by `/cs:grill-bizops` or the BizOps orchestrator. Recommended answer + canon citation per question. Never bundled.
1. **"Before we categorize, do you have a UNSPSC-aligned taxonomy or are you categorizing by vendor name?"**
Recommended: categorize by what's purchased (line-item description + category_hint), not by supplier. A single supplier can span multiple categories.
Canon: UNSPSC official taxonomy documentation, A.T. Kearney *Spend Management* on category architecture.
2. **"Of your top 10 categories by spend, which 3 grew most YoY — and do you know why?"**
Recommended: name them before opening the tool. If you can't name them, that's the diagnosis.
Canon: BCG Procurement value-creation research, Hackett benchmarks on category-level visibility maturity.
3. **"For each duplicate-function cluster (e.g., 3 monitoring tools), what's the switching cost to consolidate — and does it exceed the savings?"**
Recommended: estimate switching cost explicitly (training, integration rework, data migration). Refuse to recommend consolidation without it.
Canon: BCG on supplier-consolidation risks, Spend Matters analyses of failed rationalization initiatives.
4. **"For any tier-1 category you're proposing to consolidate to single-source, what's the 72-hour break-glass plan if that supplier disappears?"**
Recommended: documented contingency per category, tested. If absent, do not consolidate.
Canon: NotPetya / M.E.Doc supply chain attack lessons, NIST SP 800-161, A.T. Kearney on supply concentration risk.
5. **"What % of your spend goes through a PO vs. expense reimbursement vs. shadow IT? Where's the maverick spend?"**
Recommended: measure it. A.T. Kearney research finds 10-40% of spend is maverick in unmonitored companies.
Canon: A.T. Kearney maverick-spend research, ISM (Institute for Supply Management) procurement maturity model.
6. **"How many of your top-20 contracts renew in the same calendar month? Do you have a renewal calendar?"**
Recommended: build the calendar; spread renewals deliberately. Clustered renewals destroy negotiation leverage.
Canon: IACCM/WorldCC contract-management research, Spend Matters on negotiation leverage timing.
7. **"What's your approval threshold for net-new SaaS purchases under $5k? Who owns the death-by-a-thousand-SaaS problem?"**
Recommended: a tightened threshold + a single owner. Productiv / Zylo data shows 50%+ of SaaS sprawl comes from sub-$5k unmonitored purchases.
Canon: Productiv / Zylo / Vendr industry reports on SaaS sprawl.
Walk depth-first. Lock 1-4 before opening 5-7. After all are answered, invoke `spend_categorizer.py` → `purchasing_cycle_analyzer.py` → `supplier_consolidation.py` in sequence.
FILE:assets/spend_intake_template.md
# Spend Intake Template
20-minute fill-out for the annual SaaS audit / category-level spend review. Output is a JSON list you can paste into `scripts/spend_categorizer.py` and `scripts/supplier_consolidation.py`.
---
## Step 1 — Gather sources (5 minutes)
Pull line-item spend from one or more of:
- AP / accounting system export (Bill.com, Ramp, NetSuite, QuickBooks).
- SaaS-management platform export (Productiv, Zylo, Vendr, Tropic, BetterCloud).
- Corporate-card export with merchant + memo.
- Expense reimbursement export (Expensify, Concur, Navan).
Aim for the top 100-200 line items by spend. The Pareto holds — the top 20% will give you 80% of the answer.
---
## Step 2 — Fill out the spend JSON (15 minutes)
For each line item, populate the schema below. Skip prior-year fields if you don't have them — the tool degrades gracefully.
```json
[
{
"supplier": "Datadog",
"description": "Monitoring + APM enterprise tier",
"category_hint": "monitoring",
"annual_spend": 180000,
"frequency": "annual",
"currency": "USD",
"prior_year_spend": 120000
},
{
"supplier": "AWS",
"description": "EC2 + S3 + RDS production infrastructure",
"category_hint": "cloud infrastructure",
"annual_spend": 720000,
"frequency": "monthly",
"currency": "USD",
"prior_year_spend": 480000
},
{
"supplier": "Outside Counsel - Fenwick",
"description": "legal services - contracts, employment, IP",
"category_hint": "legal",
"annual_spend": 95000,
"frequency": "as-billed",
"currency": "USD",
"prior_year_spend": 60000
}
]
```
### Field guide
| Field | Required? | Notes |
|---|---|---|
| `supplier` | yes | The legal entity you pay (not the brand). |
| `description` | yes | What you bought, in your words. **This drives categorization** — be specific. "Workday HR" vs "Workday Finance" categorize differently. |
| `category_hint` | optional but recommended | A short keyword (monitoring, expense, crm, legal, etc.). Helps the categorizer when the description is ambiguous. |
| `annual_spend` | yes | Annualized total (multiply monthly × 12). |
| `frequency` | optional | `annual`, `monthly`, `quarterly`, `as-billed` — informational only, doesn't change categorization. |
| `currency` | optional | Default USD. Convert before input if mixed-currency. |
| `prior_year_spend` | optional | Enables YoY growth analysis. Set to 0 for new-this-year subscriptions. |
---
## Step 3 — Run the categorizer
```bash
python scripts/spend_categorizer.py \
--input spend.json \
--profile tech-startup \
--output categorized.md
```
Profiles: `tech-startup`, `scaleup`, `enterprise`, `services`, `manufacturing`.
---
## Step 4 — Build the supplier-criticality JSON for consolidation
Take the same suppliers, add criticality and switching cost. The supplier-consolidation tool needs this:
```json
[
{
"name": "Datadog",
"category": "Monitoring / Observability",
"annual_spend": 180000,
"criticality": "tier-2",
"contract_term_months": 12,
"integration_count_with_other_systems": 12,
"switching_cost_estimate": 80000,
"renewal_date": "2026-09-15",
"break_glass_documented": false
},
{
"name": "AWS",
"category": "Cloud Infrastructure",
"annual_spend": 720000,
"criticality": "tier-1",
"contract_term_months": 36,
"integration_count_with_other_systems": 40,
"switching_cost_estimate": 600000,
"renewal_date": "2027-03-31",
"break_glass_documented": true
}
]
```
### Criticality definitions (decide before running, don't let the tool infer)
- **tier-1** — revenue-blocking if the supplier disappears for 24h+. Identity providers, payment processors, primary cloud, primary CRM. Tier-1 should be a short list (typically 5-15 suppliers).
- **tier-2** — important but a workaround exists. Most SaaS lands here.
- **tier-3** — nice-to-have. Long-tail SaaS, productivity utilities.
### Break-glass flag
`break_glass_documented: true` means you have a written 72-hour contingency plan for what happens if this supplier disappears tomorrow. The tool **refuses to recommend tier-1 consolidation** if any cluster member has this flag false.
---
## Step 5 — Run consolidation
```bash
python scripts/supplier_consolidation.py \
--input suppliers.json \
--profile tech-startup \
--output consolidation_plan.md
```
---
## Optional: purchasing-cycle data
If you have PO timestamp data, the cycle analyzer surfaces bottleneck categories:
```json
[
{
"category": "Outside Counsel",
"request_date": "2026-01-05",
"approval_date": "2026-02-15",
"po_issued_date": "2026-02-28",
"goods_received_date": "2026-02-28",
"payment_date": "2026-03-30",
"approver_hops": 4
}
]
```
Run with:
```bash
python scripts/purchasing_cycle_analyzer.py \
--input pos.json \
--output cycle.md
```
---
## Quick sanity checks before you run
- [ ] Top 10 line items cover ≥ 50% of total spend (Pareto sanity check)
- [ ] Each line item has a non-empty `description` (drives categorization quality)
- [ ] Tier-1 suppliers are explicitly marked (don't let the tool guess)
- [ ] Switching-cost estimates exist for any supplier you might consolidate
- [ ] Renewal dates are populated for at least the top 20 contracts (drives renewal-cluster analysis)
FILE:references/procurement_anti_patterns.md
# Procurement Anti-Patterns
A field guide to the most common procurement mistakes — drawn from A.T. Kearney's maverick-spend research, IACCM/WorldCC contract studies, McKinsey's category strategy commentary, Hackett purchasing-cycle research, BCG's supplier-consolidation post-mortems, Spend Matters' analyses of failed rationalization initiatives, and ISM (Institute for Supply Management) procurement maturity studies.
Use this file before running any tool. Most "spend audits" produce a beautiful slide deck that triggers a consolidation initiative that destroys 30-50% of the theoretical savings. Read these first.
---
## Sources (≥ 7)
1. **A.T. Kearney — Maverick spend research** (AEP studies, multi-year)
2. **IACCM / WorldCC — *State of Contract and Commercial Management***
3. **McKinsey — *The CPO Agenda* and category strategy commentary**
4. **Hackett Group — *Procurement Performance Study*** (annual benchmarks)
5. **BCG — Supplier consolidation case studies and *The CPO Agenda***
6. **Spend Matters — Failed rationalization analyses** (Pierre Mitchell, Jason Busch)
7. **ISM (Institute for Supply Management) — *Manage Indirect Spending* and lessons-learned studies**
8. **Productiv / Zylo / Vendr / Tropic — SaaS-specific anti-patterns** (cross-referenced from `saas_management_canon.md`)
---
## Anti-pattern 1: Consolidate to single-source for a tier-1 critical category
**Pattern.** You have three monitoring tools. You consolidate to one. The new sole vendor has a major outage three months in. You have no break-glass plan because you offboarded the other two tools to capture the savings. Engineering is flying blind for 6 hours.
**Why it happens.** Savings math is easy and visible. Operational risk is intangible and unmeasured. The CFO incentive points one direction.
**Fix.** Before consolidating any tier-1 category, document a 72-hour break-glass plan: which alternative do you switch to, who executes the switch, what's the SLA expectation, where's the contractual fallback. The skill's `supplier_consolidation.py` refuses to recommend tier-1 consolidation without the `break_glass_documented: true` flag in the input.
**Canon.** BCG supplier-consolidation case studies (multi-year retrospectives show 30-50% of theoretical savings disappear due to operational disruption). NotPetya / M.E.Doc and SolarWinds supply-chain attack lessons.
---
## Anti-pattern 2: Categorize by vendor name, not by what's purchased
**Pattern.** You categorize Workday as "HR Software." But you also licensed the financial planning module — that's Finance Software. Your category Pareto now mis-attributes $400k of Finance spend to HR.
**Why it happens.** Vendor name is easy. Line-item description requires reading every entry.
**Fix.** Categorize from the line-item `description` and `category_hint`, not the supplier. The skill's `spend_categorizer.py` ranks `description` and `category_hint` ahead of `supplier` in keyword matching.
**Canon.** Pierre Mitchell / Spend Matters — *Category strategy mechanics*. UNSPSC categorization principle: classify the good/service, not the provider.
---
## Anti-pattern 3: Ignore renewal-date clustering
**Pattern.** Twelve tier-2 SaaS contracts all renew in March. You go into the negotiation cycle simultaneously, with three weeks to renegotiate twelve contracts. You auto-renew nine of them because you ran out of bandwidth.
**Why it happens.** Renewals piled up over years of unmonitored procurement. Nobody saw it because nobody built the calendar.
**Fix.** Build a renewal calendar (the skill outputs this). At each next renewal, negotiate term length deliberately (18-month, 6-month, 12-month rotation) to permanently spread the calendar across the year.
**Canon.** IACCM/WorldCC contract studies — 60-80% of contracts auto-renew without review. Vendr SaaS Buyers Report — quarter-end and year-end discounts are real, but only if you have negotiation bandwidth.
---
## Anti-pattern 4: Approve-by-default for sub-$5k spend (death by a thousand SaaS)
**Pattern.** Approval workflow requires CFO sign-off for $5k+ purchases. Below that, any manager can approve. Result: 80 SaaS subscriptions each costing $2-4k/year, totaling $250k of unmonitored spend that grows 50% YoY.
**Why it happens.** Approval thresholds are usually set once (often at company founding) and never re-tuned.
**Fix.** Tighten the sub-$5k threshold — but only for net-new SaaS, not for renewals of catalog items. Require a single owner for "death-by-a-thousand-SaaS" risk (typically the BizOps lead or a SaaS-management platform). The skill's `spend_categorizer.py` surfaces "small-spend, many-supplier" clusters explicitly.
**Canon.** A.T. Kearney maverick-spend research (10-40% of indirect spend leaks through sub-threshold purchases). BetterCloud State of SaaSOps — sub-$5k is the dominant shadow-IT entry point.
---
## Anti-pattern 5: No quarterly renewal review (annual is too slow)
**Pattern.** You do an "annual SaaS audit" every January. Between January and December, 30 new subscriptions get added, 12 grow >50%, and 8 auto-renew before you re-review them.
**Why it happens.** Annual reviews feel sufficient. They're not for SaaS, which is continuously renewing across the year.
**Fix.** Quarterly category review for tier-1 and tier-2 categories. Annual deep audit for tier-3 (low-spend, non-critical).
**Canon.** Forrester SaaS Portfolio Management — three-tier governance with quarterly cadence for high-tier categories. Hackett — world-class procurement reviews categories on a rolling quarterly basis.
---
## Anti-pattern 6: Rationalize without measuring switching cost
**Pattern.** You identify three monitoring tools costing $315k/year. You decide to consolidate to one tool costing $180k. Theoretical savings: $135k. Actual cost of migration (training, integration rework, alert re-tuning, parallel-run period): $200k. Net Y1 result: lost money.
**Why it happens.** Savings are visible (line-item subtraction). Switching cost is invisible (engineering time, parallel-run period, training).
**Fix.** Estimate switching cost explicitly for every consolidation. Sum across all losers in the cluster. Net Y1 savings = annual savings − migration cost. The skill's `supplier_consolidation.py` does this and flags `LOW_SAVINGS` for clusters where net Y1 < $10k.
**Canon.** BCG supplier-consolidation post-mortems. Tropic analysis of failed SaaS consolidations (60%+ failure rate to capture theoretical savings).
---
## Anti-pattern 7: Consolidate based on price alone, ignoring integration debt
**Pattern.** You consolidate to the cheapest monitoring tool. It doesn't integrate with your data warehouse or your incident management platform. You rebuild the integration plumbing for 6 months. The "cheaper" tool ends up costing more.
**Why it happens.** Price is easy to compare. Integration depth is hard to score.
**Fix.** Score `integration_count_with_other_systems` as a winner-selection input, not just price. The skill's `pick_winner` function uses integration count as the primary tiebreaker for tier-2/3 clusters.
**Canon.** Spend Matters — *Total Cost of Ownership in procurement decisions*. McKinsey — category strategy mistakes (price-only thinking is the most common error).
---
## Anti-pattern 8: Treat shadow IT spend as marketing's (or any other department's) problem
**Pattern.** Marketing has 14 unmonitored SaaS subscriptions. Procurement says "that's marketing's problem." Marketing says "we don't have the procurement bandwidth to manage that." Nobody owns it.
**Why it happens.** Shadow IT lives in expense reports and corporate-card transactions, which procurement doesn't see. Department heads see it but lack procurement skills.
**Fix.** Procurement owns the audit, even of departmental spend. A SaaS-management platform (or expense-platform integration) discovers shadow subscriptions. The Productiv finding (47% of SaaS spend is shadow) is the size of the prize.
**Canon.** Productiv State of SaaS (47% shadow IT). Zylo SaaS Management Index (marketing and engineering are the top two shadow-IT entry points).
---
## Anti-pattern 9: Negotiate without a BATNA (Best Alternative To Negotiated Agreement)
**Pattern.** You go into renewal with your monitoring vendor without having priced any alternative. The vendor knows you have no BATNA. You get 5% off list because you have no leverage.
**Why it happens.** Pricing alternatives takes time and feels confrontational.
**Fix.** Before any renewal worth $50k+, get a competitive quote — even a non-serious one. The existence of an alternative changes the negotiation tone.
**Canon.** Vendr SaaS Buyers Report on negotiation leverage. McKinsey — category strategy requires a credible threat of substitution. Tropic per-category pricing benchmarks provide the BATNA when you can't get a live quote.
---
## Anti-pattern 10: Skip the offboarding checklist when consolidating
**Pattern.** You consolidate three monitoring tools to one. Six months later, you discover the offboarded tools still have your data, still have active API keys, and one of them quietly auto-renewed because the offboarding paperwork was never filed.
**Why it happens.** Consolidation projects celebrate the new tool going live; offboarding the old tools is treated as paperwork.
**Fix.** Offboarding checklist per loser: cancel auto-renew, delete data, revoke API keys, rotate any shared credentials, confirm final invoice. The skill's `supplier_consolidation.py` outputs an explicit "Offboard:" list per cluster.
**Canon.** BetterCloud SaaS Operations on offboarding gaps. SolarWinds + Okta breach lessons on lingering vendor access.
---
## How this skill defends against the anti-patterns
| Anti-pattern | Skill defense |
|---|---|
| Single-source tier-1 | `supplier_consolidation.py` hard refusal without `break_glass_documented: true` |
| Categorize by vendor | `spend_categorizer.py` reads description + category_hint, not just supplier |
| Renewal clustering | `supplier_consolidation.py` flags months with ≥ 3 simultaneous renewals |
| Sub-$5k death | `spend_categorizer.py` surfaces small-spend many-supplier clusters |
| Annual is too slow | Forcing-question library asks about quarterly cadence |
| Ignore switching cost | `supplier_consolidation.py` requires `switching_cost_estimate`; net Y1 = savings − migration |
| Price-only consolidation | `supplier_consolidation.py` weights `integration_count_with_other_systems` in winner selection |
| Shadow IT is "marketing's problem" | Forcing-question library asks who owns sub-$5k SaaS |
| No BATNA | Forcing-question library asks about competitive quotes before renewal |
| Skip offboarding | `supplier_consolidation.py` outputs explicit Offboard list per cluster |
FILE:references/saas_management_canon.md
# SaaS Management Canon
SaaS sprawl is the dominant indirect-spend category for tech companies and the #1 driver of spend growth in scaleups. This file curates the research on SaaS portfolio management — distinct from generic procurement because SaaS has unique characteristics: per-seat pricing, auto-renew defaults, shadow-IT entry, and a viable replacement every 18 months.
---
## Sources (≥ 7)
### 1. Productiv — *State of SaaS* (annual report)
Productiv's annual benchmark across hundreds of enterprises is the most-cited SaaS sprawl data source. Key findings (most recent waves):
- **The median mid-stage company has 130-250 distinct SaaS subscriptions**, of which 30-50% are used by < 10% of licensed users.
- **License utilization median is 47%** — meaning more than half of SaaS spend buys seats nobody logs into.
- **47% of SaaS spend is shadow IT** (purchased outside the procurement process). The skill's anti-pattern list calls this out: "treat shadow IT spend as marketing's problem — it isn't."
- **Renewal is the highest-leverage moment**: 67% of SaaS purchases auto-renew without review.
Use when: framing the size of the prize for a SaaS audit.
### 2. Zylo — *SaaS Management Index* and annual benchmarks
Zylo's research focuses on SaaS economics:
- **SaaS spend grows ~30% YoY in scaleups** vs. ~10% headcount growth — meaning per-employee SaaS cost is growing.
- **Duplicate-function clusters** are extraordinarily common: most enterprises have 3-5 monitoring tools, 2-3 expense platforms, 4+ email-marketing tools. This skill's clustering logic is calibrated to Zylo's observed cluster patterns.
- **Lowest-hanging consolidation savings** are in marketing tech (often 30-40% redundancy) and developer tools.
### 3. Vendr — *SaaS Buyers Report* and pricing intelligence
Vendr's procurement-negotiation research is the practitioner's playbook for SaaS pricing leverage:
- **Median SaaS discount achievable** ranges 10-40% off list, driven by: term length, payment terms, multi-year commit, and **renewal-date timing**. Vendr's data confirms that vendors discount more aggressively at quarter-end and year-end.
- **The "MSA + Order Form" pattern** lets you re-negotiate per-order pricing without re-opening the master agreement — important for SaaS where the master may have unfavorable renewal terms locked in.
### 4. Tropic — *SaaS Cost Index* and category benchmarks
Tropic's per-category pricing benchmarks are public:
- **Per-seat pricing benchmarks** by category (CRM, ATS, HRIS, monitoring, etc.) give you the BATNA when negotiating. If you're paying 2× the Tropic median for Salesforce seats, you have leverage.
- **Tropic's analysis of consolidation success rates** shows that 60%+ of attempted SaaS consolidations fail to capture the theoretical savings, primarily due to (a) underestimated training cost, (b) loss of feature parity, (c) tier-1 single-source operational risk.
### 5. BetterCloud — *State of SaaSOps* (annual)
BetterCloud's operations research focuses on the lifecycle (onboarding → utilization → offboarding):
- **The offboarding gap:** when a SaaS tool is replaced, the old tool's licenses, data, and access often linger for 3-12 months — pure waste. SaaS audit must include an offboarding completeness check.
- **Sub-$5k SaaS purchases** are the dominant entry point for sprawl: they typically skip procurement review entirely and self-renew before anyone notices.
### 6. Gartner — *Magic Quadrant for SaaS Management Platforms (SMP)*
Gartner's SMP MQ defines the tooling category:
- **SMP capabilities:** discovery (find shadow SaaS), inventory (catalog all subscriptions), license utilization (who's actually logging in), renewal management (calendar + alerts), spend analytics (Pareto + YoY).
- The skill's deliverables (categorized spend, consolidation plan, renewal cluster analysis) are the artifacts an SMP would generate — useful when the user doesn't have an SMP licensed yet.
### 7. Tomasz Tunguz (Theory Ventures) — Long-running blog on SaaS economics
Tunguz's analysis of SaaS sprawl from the buyer side:
- **The "consumption shift"** from seat-based to usage-based pricing is changing the rationalization math: usage-based tools are harder to consolidate because their cost scales with workload, not seat count.
- **Long-tail SaaS** (the bottom 50% of subscriptions by spend) is where shadow IT lives. Killing 30 tools that cost $200/year each is psychologically harder than consolidating 3 tools that cost $90k/year each, but the operational simplification is comparable.
### 8. Patrick Campbell / ProfitWell — SaaS unit economics research
Campbell's research focuses on the seller side but the buyer-side implications are direct:
- **Annual contracts vs. monthly:** annual gives the vendor cash-flow stability and the buyer 10-30% discount, BUT it also locks in pricing and makes it harder to walk away mid-cycle. The skill's renewal-date clustering analysis flags this trade-off.
- **The "price-sensitivity range"** for SaaS pricing is wider than commonly believed — vendors will discount more than buyers expect when shown a credible alternative and a hard renewal deadline.
### 9. Bain — SaaS portfolio research (multi-year)
Bain's research on enterprise SaaS portfolios complements the practitioner sources:
- **Enterprise SaaS portfolios grow to 250-500 subscriptions** at the Fortune 1000 scale, with a power-law spend distribution: top 10 vendors capture 50-60% of spend.
- **The "category overlap"** finding (one vendor in multiple categories, e.g., Microsoft 365 spans productivity + identity + storage) is why categorizing by line-item description matters more than categorizing by vendor.
### 10. Forrester — *SaaS Portfolio Management* research
Forrester's research formalizes the SaaS portfolio governance question:
- **Three-tier governance model:** enterprise SaaS (procurement-led), departmental SaaS (FinOps-led), individual SaaS (expense-led). Each tier has different approval thresholds and review cadences.
- **Quarterly SaaS reviews** are the recommended cadence for mid-stage companies; annual is too slow when 30% of subscriptions are < 12 months old.
---
## How this skill applies the canon
- **Spend categorizer's category map** includes the duplicate-function clusters Zylo and Productiv observe: Monitoring, Expense, Email Marketing, Analytics, Security Tooling.
- **Supplier consolidation's tier-1 single-source refusal** is calibrated to Tropic and Vendr's research on failed consolidations.
- **Renewal-cluster analysis** is informed by IACCM and Vendr research on negotiation timing.
- **The sub-$5k approval anti-pattern** comes from BetterCloud and Productiv shadow-IT research.
FILE:references/spend_management_canon.md
# Spend Management Canon
The authoritative literature on procurement spend management, category strategy, and supplier rationalization. Read these before opening the categorizer or consolidation planner — most "spend audits" fail because they categorize by supplier name instead of by what's purchased, miss the Pareto, or rationalize without measuring switching cost.
---
## Sources (≥ 7)
### 1. A.T. Kearney — *Assessment of Excellence in Procurement (AEP)*
A.T. Kearney's biennial AEP benchmark is the longest-running cross-industry procurement maturity study (1992–present). Key findings relevant to this skill:
- **Top-quartile procurement orgs deliver 7-15× ROI** vs. their function spend, driven primarily by *category strategy* rather than negotiation tactics.
- **Maverick spend** (purchases bypassing approved suppliers) averages 10-40% of indirect spend in unmonitored companies and is the #1 leakage point.
- **Category architecture** (the taxonomy you categorize against) is the foundation; without it, you cannot find the Pareto.
Use when: justifying a spend audit to leadership ("here's why this matters"), defending a category taxonomy investment.
### 2. Pierre Mitchell / Spend Matters — Category strategy and supplier rationalization research
Pierre Mitchell at Spend Matters is the most-cited practitioner on category strategy mechanics. Key principles:
- **Categorize by what's purchased, not by who provided it.** Workday spans HR (HRIS) and Finance (HCM/Payroll) — splitting it correctly changes the category Pareto.
- **Supplier rationalization is a 3-step decision:** (1) keep / consolidate / kill, (2) negotiate / re-bid / status-quo, (3) automate / manual. Most teams collapse these into one decision and get it wrong.
- **The "right number of suppliers" question is wrong.** The right question is: what's the marginal cost of adding the Nth supplier in this category?
### 3. Hackett Group — *Procurement Performance Study* (annual)
The Hackett Group benchmarks procurement org performance on cost-to-serve, cycle time, and savings yield. Key benchmarks:
- **World-class procurement orgs have 26% lower process cost per PO** than peers — driven by automation of low-risk categories (catalog buys) so judgment is reserved for high-risk ones.
- **Purchasing cycle median** (request → PO) is 3-7 days for catalog items and 21-90+ days for negotiated services. The skill's bottleneck flag uses 2× median because real cycle distributions are highly skewed.
- **Approver-hop reduction** is the highest-ROI process intervention: most companies route every $5k+ purchase through 3-5 approvers; world-class routes by category risk, not dollar value.
### 4. BCG — *The CPO Agenda: Driving Value Through Procurement* (multi-year series)
BCG's procurement value-creation research emphasizes:
- **Category-led negotiation** delivers 2-3× the savings of supplier-led negotiation (because category strategy gives you the BATNA).
- **Supplier consolidation risk:** BCG's case studies of failed consolidations show that 30-50% of theoretical savings disappear due to: (a) underestimated migration cost, (b) loss of competitive tension, (c) single-source operational risk. The skill's hard refusal to consolidate tier-1 to single-source without break-glass is directly from this literature.
- **Renewal-date clustering** kills negotiation leverage; spreading renewals across the year is a 1-time investment in 1-3% annual savings.
### 5. Procurement Leaders — Member research and benchmarks
Procurement Leaders (now part of World 50 Group) publishes member benchmark studies on category management, supplier relationship management, and digital procurement. Key insights:
- **20% of categories drive 80% of value** — the Pareto applies but the threshold is empirical, not theoretical. Run it on your own data.
- **Category managers spend < 10% of time on the 80%-impact categories** in untuned organizations, because every category gets equal effort regardless of strategic weight.
### 6. Gartner — *Magic Quadrant for Source-to-Pay Suites* and procurement research
Gartner's procurement research informs the tooling landscape and decision processes:
- **Source-to-pay automation** is the dominant procurement digitalization theme; spend categorization is the gateway capability that everything else (analytics, supplier risk, contract management) depends on.
- **UNSPSC is the de facto category taxonomy** for cross-company benchmarking. The full UNSPSC database has ~100k entries across 4 levels (Segment / Family / Class / Commodity); most companies use only Class-level (~5k codes) and the top 200 codes cover most of their spend.
### 7. UNSPSC — Official taxonomy documentation (United Nations Standard Products and Services Code)
UNSPSC is maintained by GS1 US and is the most-adopted international product/service classification standard. Reference points:
- **4-level hierarchy:** Segment (e.g., 43 — Information Technology Broadcasting and Telecommunications) → Family (e.g., 4323 — Software) → Class (e.g., 432315 — Business Function Specific Software) → Commodity (e.g., 43231505 — Database management system software).
- **Adopted by:** UN, World Bank, US Federal procurement, most Global 2000 enterprises.
- **This skill ships ~30 Class-level categories** aligned to UNSPSC nomenclature but not the full code set, because (a) the full set is overwhelming for first-pass categorization, (b) tech-company spend concentrates in 30 categories, (c) the skill's job is "find the Pareto", not "ISO-compliant procurement reporting."
For the full UNSPSC codeset, see: https://www.unspsc.org/
### 8. IACCM / WorldCC — *State of Contract and Commercial Management*
IACCM (now WorldCC) publishes contract management research that intersects with spend management at the renewal-cluster question:
- **Median enterprise has 60-80% of contracts auto-renewing** with no review, destroying both negotiation leverage and the ability to catch maverick categories early.
- **Renewal-date clustering** (multiple contracts renewing in the same calendar month) is a frequently-cited finding; the recommended fix is term-length negotiation at the next renewal to permanently stagger the calendar.
---
## How this skill applies the canon
- **Spend categorizer** uses Hackett/Gartner Pareto framing + UNSPSC-aligned taxonomy (Class level only).
- **Purchasing cycle analyzer** uses Hackett baseline (median, P90) + Goldratt-style 2× median bottleneck threshold.
- **Supplier consolidation planner** uses BCG migration-cost discipline + the tier-1 single-source refusal as a hard rule.
- **Renewal-cluster analysis** comes directly from IACCM and Spend Matters research on negotiation leverage timing.
FILE:scripts/purchasing_cycle_analyzer.py
#!/usr/bin/env python3
"""
purchasing_cycle_analyzer.py — Per-category cycle-time scorecard with bottleneck flagging.
Input: a list of PO records with timestamps for each step (request, approval, PO,
goods receipt, payment). Output: per-category cycle-time statistics (median, P90)
plus a Goldratt-style bottleneck flag for any category whose cycle time exceeds
2x the cross-category median — the constraint is one specific category, not the
"average procurement process."
Stdlib only. Deterministic. No LLM calls.
"""
from __future__ import annotations
import argparse
import json
import statistics
import sys
from dataclasses import dataclass
from datetime import date, datetime
from pathlib import Path
from typing import Any
# ---------- Data model ----------
@dataclass
class PORecord:
category: str
request_date: date | None
approval_date: date | None
po_issued_date: date | None
goods_received_date: date | None
payment_date: date | None
approver_hops: int
@staticmethod
def _parse(d: Any) -> date | None:
if not d:
return None
if isinstance(d, date):
return d
try:
return datetime.strptime(str(d)[:10], "%Y-%m-%d").date()
except ValueError:
return None
@classmethod
def from_dict(cls, d: dict[str, Any]) -> "PORecord":
return cls(
category=str(d.get("category", "Uncategorized")),
request_date=cls._parse(d.get("request_date")),
approval_date=cls._parse(d.get("approval_date")),
po_issued_date=cls._parse(d.get("po_issued_date")),
goods_received_date=cls._parse(d.get("goods_received_date")),
payment_date=cls._parse(d.get("payment_date")),
approver_hops=int(d.get("approver_hops", 0)),
)
def _days(a: date | None, b: date | None) -> int | None:
if a is None or b is None:
return None
return (b - a).days
# ---------- Aggregation ----------
@dataclass
class CategoryStats:
category: str
n: int
request_to_po_median: float | None
request_to_po_p90: float | None
po_to_pay_median: float | None
po_to_pay_p90: float | None
approver_hops_median: float | None
def _p90(values: list[int]) -> float:
if not values:
return 0.0
s = sorted(values)
# Linear interpolation P90 (stdlib has no quantiles in older pythons; do manually)
k = (len(s) - 1) * 0.9
lo = int(k)
hi = min(lo + 1, len(s) - 1)
frac = k - lo
return s[lo] + (s[hi] - s[lo]) * frac
def _median(values: list[int]) -> float | None:
return statistics.median(values) if values else None
def per_category_stats(records: list[PORecord]) -> dict[str, CategoryStats]:
by_cat: dict[str, list[PORecord]] = {}
for r in records:
by_cat.setdefault(r.category, []).append(r)
result: dict[str, CategoryStats] = {}
for cat, recs in by_cat.items():
r2po = [d for d in (_days(r.request_date, r.po_issued_date) for r in recs) if d is not None]
po2pay = [d for d in (_days(r.po_issued_date, r.payment_date) for r in recs) if d is not None]
hops = [r.approver_hops for r in recs if r.approver_hops >= 0]
result[cat] = CategoryStats(
category=cat,
n=len(recs),
request_to_po_median=_median(r2po),
request_to_po_p90=(_p90(r2po) if r2po else None),
po_to_pay_median=_median(po2pay),
po_to_pay_p90=(_p90(po2pay) if po2pay else None),
approver_hops_median=_median(hops),
)
return result
def overall_median_r2po(stats: dict[str, CategoryStats]) -> float:
"""Cross-category median of request->PO median (used as the bottleneck baseline)."""
medians = [s.request_to_po_median for s in stats.values() if s.request_to_po_median is not None]
if not medians:
return 0.0
return statistics.median(medians)
# ---------- Rendering ----------
def render_markdown(stats: dict[str, CategoryStats]) -> str:
baseline = overall_median_r2po(stats)
threshold = baseline * 2.0 if baseline > 0 else None
lines: list[str] = []
lines.append("# Purchasing Cycle-Time Scorecard\n")
lines.append(f"- **Categories analyzed:** {len(stats)}")
if baseline > 0:
lines.append(f"- **Cross-category median (Request → PO):** {baseline:.1f} days")
lines.append(f"- **Bottleneck threshold (2× median):** {threshold:.1f} days\n")
else:
lines.append("- (Insufficient data for cross-category baseline)\n")
lines.append("## Per-category cycle times\n")
lines.append("| Category | N | Req→PO median | Req→PO P90 | PO→Pay median | PO→Pay P90 | Hops | Bottleneck? |")
lines.append("|---|---:|---:|---:|---:|---:|---:|---|")
sorted_cats = sorted(
stats.values(),
key=lambda s: -(s.request_to_po_median or 0),
)
for s in sorted_cats:
is_bottleneck = (
threshold is not None
and s.request_to_po_median is not None
and s.request_to_po_median > threshold
)
flag = "**BOTTLENECK**" if is_bottleneck else "—"
lines.append(
f"| {s.category} | {s.n} | "
f"{_fmt(s.request_to_po_median)} | {_fmt(s.request_to_po_p90)} | "
f"{_fmt(s.po_to_pay_median)} | {_fmt(s.po_to_pay_p90)} | "
f"{_fmt(s.approver_hops_median)} | {flag} |"
)
lines.append("")
# Goldratt commentary
bottlenecks = [
s for s in stats.values()
if threshold is not None
and s.request_to_po_median is not None
and s.request_to_po_median > threshold
]
lines.append("## Goldratt — find the constraint\n")
if bottlenecks:
lines.append(
f"**{len(bottlenecks)}** categor{'y' if len(bottlenecks)==1 else 'ies'} "
f"exceed the 2× median bottleneck threshold:\n"
)
for s in bottlenecks:
lines.append(
f"- **{s.category}** — Request→PO median {s.request_to_po_median:.1f}d, "
f"P90 {_fmt(s.request_to_po_p90)}d, "
f"approver hops median {_fmt(s.approver_hops_median)}"
)
lines.append("")
lines.append(
"Improving any non-bottleneck category does not change overall throughput. "
"Focus on the constraint first — typical fixes per stage:"
)
lines.append("")
lines.append("- **Long Request→Approval:** approval routing, parallel review, raise auto-approve threshold for low-risk categories.")
lines.append("- **Long Approval→PO:** PO creation friction; consider catalog buys for repeating categories.")
lines.append("- **High approver hops:** collapse routing tiers; one approver per $-band, not three.")
lines.append("- **Long PO→Pay:** AP cycle (3-way match, batch runs); negotiate net-terms only after measuring.")
else:
lines.append("No category exceeds the 2× bottleneck threshold. The process is uniformly fast (or uniformly slow — check the baseline).\n")
return "\n".join(lines)
def _fmt(v: float | None) -> str:
if v is None:
return "—"
return f"{v:.1f}"
# ---------- Sample data ----------
SAMPLE_INPUT: list[dict[str, Any]] = [
# Fast: SaaS / Subscription Software
{"category": "SaaS / Subscription Software", "request_date": "2026-01-02",
"approval_date": "2026-01-03", "po_issued_date": "2026-01-04",
"goods_received_date": "2026-01-04", "payment_date": "2026-01-15",
"approver_hops": 1},
{"category": "SaaS / Subscription Software", "request_date": "2026-02-01",
"approval_date": "2026-02-03", "po_issued_date": "2026-02-05",
"goods_received_date": "2026-02-05", "payment_date": "2026-02-20",
"approver_hops": 1},
{"category": "SaaS / Subscription Software", "request_date": "2026-03-01",
"approval_date": "2026-03-02", "po_issued_date": "2026-03-04",
"goods_received_date": "2026-03-04", "payment_date": "2026-03-22",
"approver_hops": 1},
# Slow: Outside Counsel (legal review is the bottleneck)
{"category": "Outside Counsel", "request_date": "2026-01-05",
"approval_date": "2026-02-15", "po_issued_date": "2026-02-28",
"goods_received_date": "2026-02-28", "payment_date": "2026-03-30",
"approver_hops": 4},
{"category": "Outside Counsel", "request_date": "2026-02-01",
"approval_date": "2026-03-12", "po_issued_date": "2026-03-30",
"goods_received_date": "2026-03-30", "payment_date": "2026-04-30",
"approver_hops": 4},
{"category": "Outside Counsel", "request_date": "2026-03-01",
"approval_date": "2026-04-25", "po_issued_date": "2026-05-05",
"goods_received_date": "2026-05-05", "payment_date": "2026-06-15",
"approver_hops": 5},
# Medium: Cloud Infrastructure
{"category": "Cloud Infrastructure", "request_date": "2026-01-10",
"approval_date": "2026-01-15", "po_issued_date": "2026-01-22",
"goods_received_date": "2026-01-22", "payment_date": "2026-02-15",
"approver_hops": 2},
{"category": "Cloud Infrastructure", "request_date": "2026-02-05",
"approval_date": "2026-02-12", "po_issued_date": "2026-02-20",
"goods_received_date": "2026-02-20", "payment_date": "2026-03-12",
"approver_hops": 2},
# Medium: Recruiting
{"category": "Recruiting Services", "request_date": "2026-01-15",
"approval_date": "2026-01-22", "po_issued_date": "2026-01-28",
"goods_received_date": "2026-01-28", "payment_date": "2026-02-20",
"approver_hops": 2},
{"category": "Recruiting Services", "request_date": "2026-02-10",
"approval_date": "2026-02-15", "po_issued_date": "2026-02-22",
"goods_received_date": "2026-02-22", "payment_date": "2026-03-15",
"approver_hops": 2},
]
# ---------- CLI ----------
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description=__doc__)
p.add_argument("--input", type=str, help="Path to JSON list of PO records")
p.add_argument("--output", type=str, help="Path to write markdown report")
p.add_argument("--sample", action="store_true", help="Run with built-in sample data")
args = p.parse_args(argv)
if args.sample:
data = SAMPLE_INPUT
elif args.input:
try:
data = json.loads(Path(args.input).read_text())
except Exception as e:
print(f"error reading {args.input}: {e}", file=sys.stderr)
return 2
else:
p.print_help()
return 0
records = [PORecord.from_dict(d) for d in data]
stats = per_category_stats(records)
md = render_markdown(stats)
if args.output:
Path(args.output).write_text(md)
print(f"wrote {args.output}")
else:
print(md)
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/spend_categorizer.py
#!/usr/bin/env python3
"""
spend_categorizer.py — UNSPSC-aligned spend categorization + Pareto + YoY growth.
Maps each line item to a UNSPSC-aligned Class -> Family -> Segment using a built-in
category map (~30 categories tuned for tech-company spend; NOT the full UNSPSC DB).
Computes Pareto (which 20% of categories drive 80% of spend?) and YoY growth when
prior-year data is supplied.
Industry profiles re-prioritize category matching:
tech-startup | scaleup | enterprise | services | manufacturing
Stdlib only. Deterministic. No LLM calls.
"""
from __future__ import annotations
import argparse
import json
import sys
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any
# ---------- Built-in UNSPSC-aligned category map ----------
# Shape: keyword -> (Segment, Family, Class)
# Segment is the top-level UNSPSC segment number range concept; we use plain names
# so the artifact reads cleanly without requiring the full 100k UNSPSC codeset.
CATEGORY_MAP: list[tuple[list[str], tuple[str, str, str]]] = [
# Software / SaaS
(["saas", "software license", "subscription", "seat license"],
("Information Technology", "Software", "SaaS / Subscription Software")),
(["crm", "salesforce", "hubspot"],
("Information Technology", "Software", "CRM Platform")),
(["monitoring", "datadog", "new relic", "grafana", "splunk", "observability"],
("Information Technology", "Software", "Monitoring / Observability")),
(["expense", "ramp", "brex", "expensify", "navan", "concur"],
("Information Technology", "Software", "Expense / Spend Management")),
(["email marketing", "mailchimp", "klaviyo", "marketo", "iterable", "sendgrid"],
("Marketing", "MarTech", "Email Marketing Platform")),
(["analytics", "amplitude", "mixpanel", "heap", "ga4"],
("Information Technology", "Software", "Product Analytics")),
(["hris", "hr software", "workday", "rippling", "gusto", "bamboohr"],
("Human Resources", "Software", "HRIS / Payroll")),
(["ats", "applicant tracking", "greenhouse", "lever"],
("Human Resources", "Software", "Applicant Tracking System")),
(["security software", "okta", "1password", "snyk", "wiz", "crowdstrike"],
("Information Technology", "Security", "Security Tooling")),
(["data warehouse", "snowflake", "bigquery", "databricks", "redshift"],
("Information Technology", "Cloud Infrastructure", "Data Warehouse")),
# Cloud Infrastructure
(["aws", "amazon web services", "ec2", "s3"],
("Information Technology", "Cloud Infrastructure", "AWS")),
(["gcp", "google cloud"],
("Information Technology", "Cloud Infrastructure", "GCP")),
(["azure", "microsoft azure"],
("Information Technology", "Cloud Infrastructure", "Azure")),
(["cloudflare", "cdn", "fastly", "akamai"],
("Information Technology", "Cloud Infrastructure", "CDN / Edge")),
# Hardware
(["laptop", "macbook", "thinkpad", "computer", "workstation"],
("Information Technology", "Hardware", "Endpoint Devices")),
(["monitor", "display", "peripheral", "keyboard", "mouse"],
("Information Technology", "Hardware", "Peripherals")),
# Professional Services
(["legal services", "law firm", "outside counsel"],
("Professional Services", "Legal", "Outside Counsel")),
(["accounting", "audit", "tax", "cpa", "deloitte", "pwc", "ey", "kpmg"],
("Professional Services", "Accounting", "Audit / Tax / Accounting")),
(["consulting", "consultant", "advisory", "mckinsey", "bcg"],
("Professional Services", "Consulting", "Management Consulting")),
(["contractor", "agency", "freelance"],
("Professional Services", "Contract Labor", "Contractor / Freelance")),
# Marketing Services
(["advertising", "google ads", "facebook ads", "linkedin ads", "ppc"],
("Marketing", "Advertising", "Paid Media")),
(["content", "copywriting", "blog", "seo agency"],
("Marketing", "Content", "Content Production")),
(["event", "conference", "trade show", "sponsorship"],
("Marketing", "Events", "Events / Sponsorship")),
# Recruiting
(["recruiting", "headhunter", "executive search", "linkedin recruiter"],
("Human Resources", "Recruiting", "Recruiting Services")),
# Travel
(["travel", "flight", "airfare", "hotel", "lodging", "uber", "lyft"],
("General & Administrative", "Travel", "Travel")),
# Office / Facilities
(["office", "rent", "lease", "wework", "coworking", "facilities"],
("General & Administrative", "Facilities", "Office / Rent")),
(["utilities", "electric", "internet", "phone"],
("General & Administrative", "Facilities", "Utilities")),
# Insurance / Benefits
(["insurance", "liability", "d&o", "cyber insurance"],
("General & Administrative", "Insurance", "Business Insurance")),
(["benefits", "health insurance", "401k", "dental", "vision"],
("Human Resources", "Benefits", "Employee Benefits")),
]
UNCATEGORIZED = ("Uncategorized", "Uncategorized", "Uncategorized")
# ---------- Industry profile priorities ----------
# Profiles influence which category is selected when multiple keywords match.
# The first-listed category in the priority list wins ties.
PROFILE_PRIORITIES: dict[str, list[str]] = {
"tech-startup": [
"SaaS / Subscription Software",
"AWS", "GCP", "Azure",
"Monitoring / Observability",
"Data Warehouse",
"Security Tooling",
"Contractor / Freelance",
],
"scaleup": [
"Recruiting Services",
"CRM Platform",
"Paid Media",
"Email Marketing Platform",
"HRIS / Payroll",
"SaaS / Subscription Software",
],
"enterprise": [
"Management Consulting",
"Outside Counsel",
"Audit / Tax / Accounting",
"Office / Rent",
"Employee Benefits",
"Business Insurance",
],
"services": [
"Contractor / Freelance",
"Outside Counsel",
"Travel",
"SaaS / Subscription Software",
],
"manufacturing": [
"Endpoint Devices",
"Peripherals",
"Office / Rent",
"Utilities",
"Business Insurance",
],
}
# ---------- Data model ----------
@dataclass
class LineItem:
supplier: str
description: str
category_hint: str
annual_spend: float
frequency: str = "annual"
currency: str = "USD"
prior_year_spend: float | None = None
@classmethod
def from_dict(cls, d: dict[str, Any]) -> "LineItem":
return cls(
supplier=str(d.get("supplier", "")).strip(),
description=str(d.get("description", "")).strip(),
category_hint=str(d.get("category_hint", "")).strip(),
annual_spend=float(d.get("annual_spend", 0.0)),
frequency=str(d.get("frequency", "annual")),
currency=str(d.get("currency", "USD")),
prior_year_spend=(
float(d["prior_year_spend"])
if d.get("prior_year_spend") is not None
else None
),
)
@dataclass
class Categorized:
item: LineItem
segment: str
family: str
class_: str
# ---------- Categorization ----------
def categorize(item: LineItem, profile: str) -> Categorized:
"""Categorize one line item using keyword match + profile priority for ties."""
haystack = " ".join([item.supplier, item.description, item.category_hint]).lower()
matches: list[tuple[str, str, str]] = []
for keywords, cat in CATEGORY_MAP:
for kw in keywords:
if kw in haystack:
matches.append(cat)
break
if not matches:
seg, fam, cls = UNCATEGORIZED
return Categorized(item, seg, fam, cls)
# Resolve tie using profile priority list
priority = PROFILE_PRIORITIES.get(profile, [])
for pri_class in priority:
for seg, fam, cls in matches:
if cls == pri_class:
return Categorized(item, seg, fam, cls)
# No priority match → first match wins (deterministic order)
seg, fam, cls = matches[0]
return Categorized(item, seg, fam, cls)
# ---------- Aggregation ----------
def aggregate_by_class(items: list[Categorized]) -> dict[str, dict[str, Any]]:
agg: dict[str, dict[str, Any]] = {}
for c in items:
bucket = agg.setdefault(c.class_, {
"segment": c.segment,
"family": c.family,
"class": c.class_,
"spend": 0.0,
"prior_year_spend": 0.0,
"supplier_count": 0,
"suppliers": set(),
})
bucket["spend"] += c.item.annual_spend
if c.item.prior_year_spend is not None:
bucket["prior_year_spend"] += c.item.prior_year_spend
bucket["suppliers"].add(c.item.supplier)
for b in agg.values():
b["supplier_count"] = len(b["suppliers"])
b["suppliers"] = sorted(b["suppliers"])
return agg
def pareto_breakdown(agg: dict[str, dict[str, Any]]) -> tuple[list[str], float, float]:
"""Return the 20% of categories driving most spend, and the cumulative % they cover."""
sorted_cats = sorted(agg.items(), key=lambda kv: -kv[1]["spend"])
total_spend = sum(b["spend"] for b in agg.values()) or 1.0
top_20_count = max(1, len(sorted_cats) // 5)
top_classes = [cls for cls, _ in sorted_cats[:top_20_count]]
top_spend = sum(agg[cls]["spend"] for cls in top_classes)
return top_classes, top_spend, top_spend / total_spend * 100.0
def yoy_growth(agg: dict[str, dict[str, Any]]) -> list[tuple[str, float, float, float]]:
"""Return (class, this_year, prior_year, pct_growth) sorted by % growth desc."""
rows: list[tuple[str, float, float, float]] = []
for cls, b in agg.items():
py = b["prior_year_spend"]
ty = b["spend"]
if py > 0:
pct = (ty - py) / py * 100.0
rows.append((cls, ty, py, pct))
rows.sort(key=lambda r: -r[3])
return rows
# ---------- Rendering ----------
def render_markdown(
profile: str,
categorized: list[Categorized],
agg: dict[str, dict[str, Any]],
) -> str:
total = sum(b["spend"] for b in agg.values())
top_classes, top_spend, top_pct = pareto_breakdown(agg)
lines: list[str] = []
lines.append(f"# Categorized Spend Report ({profile} profile)\n")
lines.append(f"- **Total annual spend:** ,.0f")
lines.append(f"- **Line items:** {len(categorized)}")
lines.append(f"- **Distinct categories (Class level):** {len(agg)}\n")
lines.append("## Pareto: top 20% of categories\n")
lines.append(f"Top {len(top_classes)} categories drive ,.0f ({top_pct:.1f}% of spend):\n")
for cls in top_classes:
b = agg[cls]
share = b["spend"] / (total or 1) * 100
lines.append(f"- **{cls}** — ,.0f ({share:.1f}%), {b['supplier_count']} suppliers")
lines.append("")
lines.append("## All categories ranked by spend\n")
lines.append("| Class | Family | Segment | Spend | Suppliers |")
lines.append("|---|---|---|---:|---:|")
for cls, b in sorted(agg.items(), key=lambda kv: -kv[1]["spend"]):
lines.append(
f"| {cls} | {b['family']} | {b['segment']} | "
f",.0f | {b['supplier_count']} |"
)
lines.append("")
growth = yoy_growth(agg)
if growth:
lines.append("## Top YoY growth categories\n")
lines.append("| Class | This year | Prior year | Growth |")
lines.append("|---|---:|---:|---:|")
for cls, ty, py, pct in growth[:10]:
arrow = "↑" if pct > 0 else "↓"
lines.append(f"| {cls} | ,.0f | ,.0f | {arrow} {pct:+.1f}% |")
lines.append("")
# Per-line-item listing (for audit)
lines.append("## Line items by category\n")
by_class: dict[str, list[Categorized]] = {}
for c in categorized:
by_class.setdefault(c.class_, []).append(c)
for cls in sorted(by_class.keys()):
lines.append(f"### {cls}\n")
lines.append("| Supplier | Description | Annual spend |")
lines.append("|---|---|---:|")
for c in sorted(by_class[cls], key=lambda x: -x.item.annual_spend):
desc = c.item.description[:60]
lines.append(f"| {c.item.supplier} | {desc} | ,.0f |")
lines.append("")
return "\n".join(lines)
# ---------- Sample data ----------
SAMPLE_INPUT: list[dict[str, Any]] = [
{"supplier": "Datadog", "description": "Monitoring + APM", "category_hint": "monitoring",
"annual_spend": 180000, "prior_year_spend": 120000},
{"supplier": "New Relic", "description": "APM monitoring", "category_hint": "monitoring",
"annual_spend": 90000, "prior_year_spend": 80000},
{"supplier": "Grafana Cloud", "description": "metrics + logs monitoring",
"category_hint": "monitoring", "annual_spend": 45000, "prior_year_spend": 0},
{"supplier": "Ramp", "description": "corporate cards + expense", "category_hint": "expense",
"annual_spend": 30000, "prior_year_spend": 18000},
{"supplier": "Expensify", "description": "expense reimbursement", "category_hint": "expense",
"annual_spend": 12000, "prior_year_spend": 12000},
{"supplier": "Salesforce", "description": "CRM Enterprise", "category_hint": "crm",
"annual_spend": 240000, "prior_year_spend": 200000},
{"supplier": "AWS", "description": "EC2 + S3 + RDS", "category_hint": "cloud",
"annual_spend": 720000, "prior_year_spend": 480000},
{"supplier": "Snowflake", "description": "data warehouse", "category_hint": "data warehouse",
"annual_spend": 360000, "prior_year_spend": 240000},
{"supplier": "Klaviyo", "description": "email marketing platform",
"category_hint": "email marketing", "annual_spend": 36000, "prior_year_spend": 24000},
{"supplier": "Mailchimp", "description": "email marketing", "category_hint": "email marketing",
"annual_spend": 8000, "prior_year_spend": 8000},
{"supplier": "Iterable", "description": "email marketing", "category_hint": "email marketing",
"annual_spend": 50000, "prior_year_spend": 0},
{"supplier": "SendGrid", "description": "transactional email", "category_hint": "email",
"annual_spend": 18000, "prior_year_spend": 12000},
{"supplier": "Greenhouse", "description": "ATS", "category_hint": "applicant tracking",
"annual_spend": 28000, "prior_year_spend": 24000},
{"supplier": "Outside Counsel - Fenwick", "description": "legal services",
"category_hint": "legal", "annual_spend": 95000, "prior_year_spend": 60000},
]
# ---------- CLI ----------
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description=__doc__)
p.add_argument("--input", type=str, help="Path to JSON list of spend line items")
p.add_argument(
"--profile",
type=str,
default="tech-startup",
choices=sorted(PROFILE_PRIORITIES.keys()),
help="Industry profile (default: tech-startup)",
)
p.add_argument("--output", type=str, help="Path to write markdown report")
p.add_argument("--sample", action="store_true", help="Run with built-in sample data")
args = p.parse_args(argv)
if args.sample:
data = SAMPLE_INPUT
elif args.input:
try:
data = json.loads(Path(args.input).read_text())
except Exception as e:
print(f"error reading {args.input}: {e}", file=sys.stderr)
return 2
else:
p.print_help()
return 0
items = [LineItem.from_dict(d) for d in data]
categorized = [categorize(it, args.profile) for it in items]
agg = aggregate_by_class(categorized)
md = render_markdown(args.profile, categorized, agg)
if args.output:
Path(args.output).write_text(md)
print(f"wrote {args.output}")
else:
print(md)
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/supplier_consolidation.py
#!/usr/bin/env python3
"""
supplier_consolidation.py — Duplicate-function clustering + risk-flagged consolidation plan.
Input: list of suppliers with category, annual spend, criticality tier, contract term,
integration count with other systems, switching cost estimate, renewal date, and an
optional break-glass flag for tier-1 suppliers.
Output: markdown consolidation plan that:
- Clusters suppliers by category (duplicate-function detection)
- Recommends a consolidation winner per cluster
- REFUSES to recommend single-source consolidation for tier-1 categories without
a documented break-glass plan
- Estimates savings: current cluster spend - winner spend - migration cost
- Surfaces renewal-date clusters (≥3 contracts in same calendar month destroys leverage)
Stdlib only. Deterministic. No LLM calls.
"""
from __future__ import annotations
import argparse
import json
import sys
from dataclasses import dataclass, field
from datetime import date, datetime
from pathlib import Path
from typing import Any
# ---------- Profile-driven category criticality overrides ----------
PROFILE_TIER1_CATEGORIES: dict[str, list[str]] = {
"tech-startup": ["Cloud Infrastructure", "Data Warehouse", "Security Tooling", "CRM Platform"],
"scaleup": ["Cloud Infrastructure", "CRM Platform", "HRIS / Payroll", "Data Warehouse"],
"enterprise": ["Cloud Infrastructure", "HRIS / Payroll", "Business Insurance", "Outside Counsel"],
"services": ["CRM Platform", "Outside Counsel", "Contractor / Freelance"],
"manufacturing": ["Endpoint Devices", "Business Insurance", "Utilities"],
}
# ---------- Data model ----------
@dataclass
class Supplier:
name: str
category: str
annual_spend: float
criticality: str # tier-1 | tier-2 | tier-3
contract_term_months: int
integration_count_with_other_systems: int
switching_cost_estimate: float
renewal_date: date | None
break_glass_documented: bool = False
@staticmethod
def _parse_date(d: Any) -> date | None:
if not d:
return None
if isinstance(d, date):
return d
try:
return datetime.strptime(str(d)[:10], "%Y-%m-%d").date()
except ValueError:
return None
@classmethod
def from_dict(cls, d: dict[str, Any]) -> "Supplier":
return cls(
name=str(d.get("name", "")),
category=str(d.get("category", "Uncategorized")),
annual_spend=float(d.get("annual_spend", 0.0)),
criticality=str(d.get("criticality", "tier-3")).lower(),
contract_term_months=int(d.get("contract_term_months", 12)),
integration_count_with_other_systems=int(
d.get("integration_count_with_other_systems", 0)
),
switching_cost_estimate=float(d.get("switching_cost_estimate", 0.0)),
renewal_date=cls._parse_date(d.get("renewal_date")),
break_glass_documented=bool(d.get("break_glass_documented", False)),
)
@dataclass
class ClusterRecommendation:
category: str
members: list[Supplier]
winner: Supplier | None
losers: list[Supplier]
annual_savings: float
migration_cost: float
net_year1_savings: float
risk_flag: str # OK | TIER1_NO_BREAKGLASS | LOW_SAVINGS
rationale: str
# ---------- Clustering ----------
def cluster_by_category(suppliers: list[Supplier]) -> dict[str, list[Supplier]]:
"""Group suppliers by category; only categories with >= 2 suppliers are candidate clusters."""
by_cat: dict[str, list[Supplier]] = {}
for s in suppliers:
by_cat.setdefault(s.category, []).append(s)
return {cat: members for cat, members in by_cat.items() if len(members) >= 2}
# ---------- Winner selection ----------
def pick_winner(members: list[Supplier]) -> Supplier:
"""
Winner selection:
- If any member is tier-1, the highest-spend tier-1 wins (assume it has the most integrations and least switching cost away from).
- If cluster is tier-2/3 only, the member with lowest total cost = (annual_spend - other_members_spend) + their switching_cost_estimate.
Simpler proxy: highest integration_count_with_other_systems wins (the one that's most embedded).
Tiebreak by lowest switching_cost_estimate.
"""
tier1 = [m for m in members if m.criticality == "tier-1"]
if tier1:
return max(tier1, key=lambda m: m.annual_spend)
return max(
members,
key=lambda m: (m.integration_count_with_other_systems, -m.switching_cost_estimate),
)
# ---------- Risk assessment ----------
def assess_risk(
category: str,
members: list[Supplier],
winner: Supplier,
profile: str,
annual_savings: float,
) -> str:
"""
Risk classification:
- TIER1_NO_BREAKGLASS — any tier-1 in cluster (or category is tier-1 by profile) AND winner has no break_glass_documented
- LOW_SAVINGS — net Y1 savings < $10k (consolidating not worth the operational disruption)
- OK — proceed
"""
profile_tier1 = PROFILE_TIER1_CATEGORIES.get(profile, [])
has_tier1_member = any(m.criticality == "tier-1" for m in members)
is_tier1_category = category in profile_tier1
if (has_tier1_member or is_tier1_category) and not winner.break_glass_documented:
return "TIER1_NO_BREAKGLASS"
if annual_savings < 10000:
return "LOW_SAVINGS"
return "OK"
# ---------- Plan generation ----------
def build_recommendations(
suppliers: list[Supplier],
profile: str,
) -> list[ClusterRecommendation]:
clusters = cluster_by_category(suppliers)
recs: list[ClusterRecommendation] = []
for cat, members in clusters.items():
winner = pick_winner(members)
losers = [m for m in members if m.name != winner.name]
if not losers:
continue
annual_savings = sum(m.annual_spend for m in losers)
migration_cost = sum(m.switching_cost_estimate for m in losers)
net_y1 = annual_savings - migration_cost
risk = assess_risk(cat, members, winner, profile, annual_savings)
if risk == "TIER1_NO_BREAKGLASS":
rationale = (
"DO NOT CONSOLIDATE — tier-1 category, no documented break-glass plan. "
"Add a 72-hour contingency plan for the surviving supplier first, then revisit."
)
elif risk == "LOW_SAVINGS":
rationale = (
f"Marginal — net Y1 savings ,.0f likely consumed by operational disruption. "
"Defer unless integration debt or vendor risk justifies it."
)
else:
rationale = (
f"Consolidate to {winner.name}. {len(losers)} supplier(s) to offboard. "
f"Net Y1 savings ,.0f (gross ,.0f − migration ,.0f)."
)
recs.append(ClusterRecommendation(
category=cat,
members=members,
winner=winner,
losers=losers,
annual_savings=annual_savings,
migration_cost=migration_cost,
net_year1_savings=net_y1,
risk_flag=risk,
rationale=rationale,
))
# Sort by net savings descending (biggest opportunities first)
recs.sort(key=lambda r: -r.net_year1_savings)
return recs
# ---------- Renewal cluster analysis ----------
def renewal_clusters(suppliers: list[Supplier]) -> dict[str, list[Supplier]]:
"""Find calendar months where >=3 contracts renew (zero leverage)."""
by_month: dict[str, list[Supplier]] = {}
for s in suppliers:
if s.renewal_date is None:
continue
key = s.renewal_date.strftime("%Y-%m")
by_month.setdefault(key, []).append(s)
return {month: members for month, members in by_month.items() if len(members) >= 3}
# ---------- Rendering ----------
def render_markdown(
profile: str,
suppliers: list[Supplier],
recs: list[ClusterRecommendation],
renewals: dict[str, list[Supplier]],
) -> str:
total_spend = sum(s.annual_spend for s in suppliers)
total_net_savings = sum(
r.net_year1_savings for r in recs if r.risk_flag == "OK"
)
lines: list[str] = []
lines.append(f"# Supplier Consolidation Plan ({profile} profile)\n")
lines.append(f"- **Suppliers analyzed:** {len(suppliers)}")
lines.append(f"- **Total annual spend:** ,.0f")
lines.append(f"- **Duplicate-function clusters:** {len(recs)}")
lines.append(f"- **Net Year-1 savings opportunity (OK clusters only):** ,.0f\n")
if not recs:
lines.append("No duplicate-function clusters detected. No consolidation plan generated.\n")
else:
lines.append("## Recommendations (ranked by net Y1 savings)\n")
for r in recs:
badge = {
"OK": "RECOMMEND",
"TIER1_NO_BREAKGLASS": "DO NOT CONSOLIDATE",
"LOW_SAVINGS": "DEFER",
}[r.risk_flag]
lines.append(f"### {r.category} — {badge}\n")
lines.append(f"**Cluster:** {len(r.members)} suppliers — " +
", ".join(f"{m.name} (,.0f, {m.criticality})" for m in r.members))
if r.winner is not None:
lines.append(f"\n**Proposed winner:** {r.winner.name} "
f"(integrations={r.winner.integration_count_with_other_systems}, "
f"break-glass={'yes' if r.winner.break_glass_documented else 'no'})")
lines.append(f"\n**Offboard:** " + (", ".join(m.name for m in r.losers) or "—"))
lines.append(f"\n- Gross annual savings: ,.0f")
lines.append(f"- Migration cost: ,.0f")
lines.append(f"- **Net Y1 savings: ,.0f**")
lines.append(f"- Risk flag: `{r.risk_flag}`")
lines.append(f"\n{r.rationale}\n")
# Renewal-date clustering analysis
lines.append("## Renewal-date clusters (negotiation leverage)\n")
if not renewals:
lines.append("No calendar months with ≥ 3 simultaneous renewals. Leverage is preserved.\n")
else:
lines.append(f"**{len(renewals)} month(s) have ≥ 3 simultaneous renewals — leverage destroyed:**\n")
for month, members in sorted(renewals.items()):
lines.append(f"### {month} — {len(members)} renewals\n")
for m in members:
lines.append(
f"- {m.name} ({m.category}) — ,.0f, renewal {m.renewal_date}"
)
lines.append("")
lines.append(
"Action: stagger renewals across the year. Renegotiate term lengths at next renewal "
"(e.g., 18-month + 6-month + 12-month) to permanently de-cluster.\n"
)
# Action summary
lines.append("## Action summary\n")
ok_recs = [r for r in recs if r.risk_flag == "OK"]
tier1_blocked = [r for r in recs if r.risk_flag == "TIER1_NO_BREAKGLASS"]
deferred = [r for r in recs if r.risk_flag == "LOW_SAVINGS"]
lines.append(f"- **Proceed now ({len(ok_recs)}):** " +
(", ".join(r.category for r in ok_recs) or "—"))
lines.append(f"- **Blocked on break-glass plan ({len(tier1_blocked)}):** " +
(", ".join(r.category for r in tier1_blocked) or "—"))
lines.append(f"- **Defer ({len(deferred)}):** " +
(", ".join(r.category for r in deferred) or "—"))
return "\n".join(lines)
# ---------- Sample data ----------
SAMPLE_INPUT: list[dict[str, Any]] = [
# Monitoring cluster (3 tools, tier-2)
{"name": "Datadog", "category": "Monitoring / Observability", "annual_spend": 180000,
"criticality": "tier-2", "contract_term_months": 12,
"integration_count_with_other_systems": 12,
"switching_cost_estimate": 80000, "renewal_date": "2026-09-15",
"break_glass_documented": False},
{"name": "New Relic", "category": "Monitoring / Observability", "annual_spend": 90000,
"criticality": "tier-2", "contract_term_months": 12,
"integration_count_with_other_systems": 4,
"switching_cost_estimate": 25000, "renewal_date": "2026-09-30",
"break_glass_documented": False},
{"name": "Grafana Cloud", "category": "Monitoring / Observability", "annual_spend": 45000,
"criticality": "tier-2", "contract_term_months": 12,
"integration_count_with_other_systems": 6,
"switching_cost_estimate": 18000, "renewal_date": "2026-09-22",
"break_glass_documented": False},
# Expense cluster (2 tools, tier-3)
{"name": "Ramp", "category": "Expense / Spend Management", "annual_spend": 30000,
"criticality": "tier-3", "contract_term_months": 12,
"integration_count_with_other_systems": 8,
"switching_cost_estimate": 15000, "renewal_date": "2026-09-10",
"break_glass_documented": True},
{"name": "Expensify", "category": "Expense / Spend Management", "annual_spend": 12000,
"criticality": "tier-3", "contract_term_months": 12,
"integration_count_with_other_systems": 2,
"switching_cost_estimate": 5000, "renewal_date": "2026-09-05",
"break_glass_documented": True},
# Email marketing cluster (4 tools, tier-2 — note the tier-1 will trigger guard)
{"name": "Klaviyo", "category": "Email Marketing Platform", "annual_spend": 36000,
"criticality": "tier-1", "contract_term_months": 12,
"integration_count_with_other_systems": 10,
"switching_cost_estimate": 25000, "renewal_date": "2026-11-15",
"break_glass_documented": False},
{"name": "Mailchimp", "category": "Email Marketing Platform", "annual_spend": 8000,
"criticality": "tier-3", "contract_term_months": 12,
"integration_count_with_other_systems": 1,
"switching_cost_estimate": 2000, "renewal_date": "2026-03-15",
"break_glass_documented": False},
{"name": "Iterable", "category": "Email Marketing Platform", "annual_spend": 50000,
"criticality": "tier-2", "contract_term_months": 12,
"integration_count_with_other_systems": 5,
"switching_cost_estimate": 15000, "renewal_date": "2026-04-30",
"break_glass_documented": False},
{"name": "SendGrid", "category": "Email Marketing Platform", "annual_spend": 18000,
"criticality": "tier-2", "contract_term_months": 12,
"integration_count_with_other_systems": 4,
"switching_cost_estimate": 8000, "renewal_date": "2026-06-30",
"break_glass_documented": False},
# AWS — single supplier, tier-1, not a cluster
{"name": "AWS", "category": "Cloud Infrastructure", "annual_spend": 720000,
"criticality": "tier-1", "contract_term_months": 36,
"integration_count_with_other_systems": 40,
"switching_cost_estimate": 600000, "renewal_date": "2027-03-31",
"break_glass_documented": True},
]
# ---------- CLI ----------
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description=__doc__)
p.add_argument("--input", type=str, help="Path to JSON list of supplier records")
p.add_argument(
"--profile",
type=str,
default="tech-startup",
choices=sorted(PROFILE_TIER1_CATEGORIES.keys()),
help="Industry profile (default: tech-startup)",
)
p.add_argument("--output", type=str, help="Path to write markdown plan")
p.add_argument("--sample", action="store_true", help="Run with built-in sample data")
args = p.parse_args(argv)
if args.sample:
data = SAMPLE_INPUT
elif args.input:
try:
data = json.loads(Path(args.input).read_text())
except Exception as e:
print(f"error reading {args.input}: {e}", file=sys.stderr)
return 2
else:
p.print_help()
return 0
suppliers = [Supplier.from_dict(d) for d in data]
recs = build_recommendations(suppliers, args.profile)
renewals = renewal_clusters(suppliers)
md = render_markdown(args.profile, suppliers, recs, renewals)
if args.output:
Path(args.output).write_text(md)
print(f"wrote {args.output}")
else:
print(md)
return 0
if __name__ == "__main__":
sys.exit(main())
Xác định KPI sản phẩm, xây dashboard chỉ số, phân tích cohort và retention, diễn giải xu hướng sử dụng tính năng.
---
name: product-analytics
description: Use when defining product KPIs, building metric dashboards, running cohort or retention analysis, or interpreting feature adoption trends across product stages.
---
# Product Analytics
Define, track, and interpret product metrics across discovery, growth, and mature product stages.
## When To Use
Use this skill for:
- Metric framework selection (AARRR, North Star, HEART)
- KPI definition by product stage (pre-PMF, growth, mature)
- Dashboard design and metric hierarchy
- Cohort and retention analysis
- Feature adoption and funnel interpretation
## Workflow
1. Select metric framework
- AARRR for growth loops and funnel visibility
- North Star for cross-functional strategic alignment
- HEART for UX quality and user experience measurement
2. Define stage-appropriate KPIs
- Pre-PMF: activation, early retention, qualitative success
- Growth: acquisition efficiency, expansion, conversion velocity
- Mature: retention depth, revenue quality, operational efficiency
3. Design dashboard layers
- Executive layer: 5-7 directional metrics
- Product health layer: acquisition, activation, retention, engagement
- Feature layer: adoption, depth, repeat usage, outcome correlation
4. Run cohort + retention analysis
- Segment by signup cohort or feature exposure cohort
- Compare retention curves, not single-point snapshots
- Identify inflection points around onboarding and first value moment
5. Interpret and act
- Connect metric movement to product changes and release timeline
- Distinguish signal from noise using period-over-period context
- Propose one clear product action per major metric risk/opportunity
## KPI Guidance By Stage
### Pre-PMF
- Activation rate
- Week-1 retention
- Time-to-first-value
- Problem-solution fit interview score
### Growth
- Funnel conversion by stage
- Monthly retained users
- Feature adoption among new cohorts
- Expansion / upsell proxy metrics
### Mature
- Net revenue retention aligned product metrics
- Power-user share and depth of use
- Churn risk indicators by segment
- Reliability and support-deflection product metrics
## Dashboard Design Principles
- Show trends, not isolated point estimates.
- Keep one owner per KPI.
- Pair each KPI with target, threshold, and decision rule.
- Use cohort and segment filters by default.
- Prefer comparable time windows (weekly vs weekly, monthly vs monthly).
See:
- `references/metrics-frameworks.md`
- `references/dashboard-templates.md`
## Cohort Analysis Method
1. Define cohort anchor event (signup, activation, first purchase).
2. Define retained behavior (active day, key action, repeat session).
3. Build retention matrix by cohort week/month and age period.
4. Compare curve shape across cohorts.
5. Flag early drop points and investigate journey friction.
## Retention Curve Interpretation
- Sharp early drop, low plateau: onboarding mismatch or weak initial value.
- Moderate drop, stable plateau: healthy core audience with predictable churn.
- Flattening at low level: product used occasionally, revisit value metric.
- Improving newer cohorts: onboarding or positioning improvements are working.
## Anti-Patterns
| Anti-pattern | Fix |
|---|---|
| **Vanity metrics** — tracking pageviews or total signups without activation context | Always pair acquisition metrics with activation rate and retention |
| **Single-point retention** — reporting "30-day retention is 20%" | Compare retention curves across cohorts, not isolated snapshots |
| **Dashboard overload** — 30+ metrics on one screen | Executive layer: 5-7 metrics. Feature layer: per-feature only |
| **No decision rule** — tracking a KPI with no threshold or action plan | Every KPI needs: target, threshold, owner, and "if below X, then Y" |
| **Averaging across segments** — reporting blended metrics that hide segment differences | Always segment by cohort, plan tier, channel, or geography |
| **Ignoring seasonality** — comparing this week to last week without adjusting | Use period-over-period with same-period-last-year context |
## Tooling
### `scripts/metrics_calculator.py`
CLI utility for retention, cohort, and funnel analysis from CSV data. Supports text and JSON output.
```bash
# Retention analysis
python3 scripts/metrics_calculator.py retention events.csv
python3 scripts/metrics_calculator.py retention events.csv --format json
# Cohort matrix
python3 scripts/metrics_calculator.py cohort events.csv --cohort-grain month
python3 scripts/metrics_calculator.py cohort events.csv --cohort-grain week --format json
# Funnel conversion
python3 scripts/metrics_calculator.py funnel funnel.csv --stages visit,signup,activate,pay
python3 scripts/metrics_calculator.py funnel funnel.csv --stages visit,signup,activate,pay --format json
```
**CSV format for retention/cohort:**
```csv
user_id,cohort_date,activity_date
u001,2026-01-01,2026-01-01
u001,2026-01-01,2026-01-03
u002,2026-01-02,2026-01-02
```
**CSV format for funnel:**
```csv
user_id,stage
u001,visit
u001,signup
u001,activate
u002,visit
u002,signup
```
## Cross-References
- Related: `product-team/experiment-designer` — for A/B test planning after identifying metric opportunities
- Related: `product-team/product-manager-toolkit` — for RICE prioritization of metric-driven features
- Related: `product-team/product-discovery` — for assumption mapping when metrics reveal unknowns
- Related: `finance/saas-metrics-coach` — for SaaS-specific metrics (ARR, MRR, churn, LTV)
FILE:references/dashboard-templates.md
# Dashboard Templates
## 1. Executive Dashboard Template
Purpose: quick company-level product signal for leadership.
Sections:
1. North Star trend (current, target, trailing 12 periods)
2. Growth summary (new users/accounts, activation)
3. Retention summary (short-term + medium-term cohorts)
4. Revenue-linked product indicators
5. Risks and actions
Suggested KPI block:
| KPI | Current | Target | Delta | Owner | Action |
|---|---:|---:|---:|---|---|
| North Star | | | | | |
| Activation Rate | | | | | |
| W8 Retention | | | | | |
| Paid Conversion | | | | | |
## 2. Product Health Dashboard Template
Purpose: monitor full user journey and detect bottlenecks.
Sections:
1. Acquisition funnel by channel/segment
2. Activation funnel with drop-off points
3. Cohort retention matrix + curve chart
4. Feature adoption distribution
5. Reliability metrics tied to user outcomes
Recommended views:
- Weekly cohort retention heatmap
- Funnel stage conversion waterfall
- Segment comparison (SMB vs enterprise)
- New vs returning user behavior split
## 3. Feature Adoption Dashboard Template
Purpose: evaluate feature launch quality and ongoing usage.
Sections:
1. Exposure and eligibility count
2. First-use adoption rate
3. Repeat usage rate (2nd, 3rd, nth use)
4. Time-to-adoption from signup/activation
5. Impact on primary outcomes (retention, conversion)
Adoption KPI examples:
| Metric | Definition |
|---|---|
| First-use adoption | Users who used feature at least once / eligible users |
| Repeat adoption | Users with 2+ uses / users with first use |
| Sustained adoption | Users with usage in 3 of last 4 weeks |
| Time to adoption | Median days from eligibility to first use |
## Dashboard Design Rules
- Keep each dashboard to one decision horizon (weekly ops vs quarterly strategy).
- Always annotate major product releases on charts.
- Add threshold bands for risk detection.
- Show metric definitions next to charts.
- Include a short "what changed" narrative block.
FILE:references/metrics-frameworks.md
# Metrics Frameworks
## AARRR (Pirate Metrics)
AARRR breaks the product journey into five stages.
1. Acquisition
- How users discover the product
- Example metrics: signups, CAC, channel conversion
2. Activation
- First meaningful value moment
- Example metrics: activation rate, time-to-first-value
3. Retention
- Ongoing user return behavior
- Example metrics: D7/W4 retention, rolling retained users
4. Revenue
- Monetization and value capture
- Example metrics: conversion to paid, ARPU, expansion revenue
5. Referral
- Organic growth from existing users
- Example metrics: referral rate, invite conversion, K-factor
## North Star Metric Framework
North Star = metric capturing long-term customer value delivered.
### North Star Criteria
- Reflects real user value
- Sensitive to product improvements
- Predictive of sustainable growth
- Understandable across functions
### Example North Star Metrics
- Collaboration SaaS: weekly active teams
- Marketplace: successful transactions per active buyer
- Content product: hours of qualified consumption
### Input Metrics
Track levers that influence the North Star:
- Acquisition quality
- Activation quality
- Engagement depth
- Retention durability
## HEART Framework
HEART is a UX-oriented framework from Google.
- Happiness: satisfaction, NPS, perceived quality
- Engagement: interaction depth/frequency
- Adoption: first-time use of features/products
- Retention: return behavior over time
- Task Success: completion rate, error rate, time on task
### HEART + Goals-Signals-Metrics
1. Goals: what UX outcome you want
2. Signals: observed behavior indicating movement
3. Metrics: measurable indicator for each signal
## Framework Selection Guide
| Situation | Recommended Framework |
|---|---|
| Early growth and funnel bottlenecks | AARRR |
| Company-wide strategic alignment | North Star |
| UX and product quality optimization | HEART |
| Mixed maturity org | North Star + AARRR operational layers |
## Example: B2B SaaS Product
- North Star: weekly active accounts completing core workflow
- AARRR operational metrics:
- Acquisition: qualified signups
- Activation: % accounts completing setup in 7 days
- Retention: W8 retained accounts
- Revenue: paid conversion and expansion rate
- Referral: invited teammate activation rate
- HEART for onboarding redesign:
- Task Success: onboarding completion rate
- Happiness: onboarding CSAT
FILE:scripts/metrics_calculator.py
#!/usr/bin/env python3
"""Product metrics calculator: retention, cohort matrix, and funnel conversion."""
import argparse
import csv
import datetime as dt
import json
import sys
from collections import defaultdict
def parse_date(value: str) -> dt.date:
return dt.date.fromisoformat(value.strip()[:10])
def load_csv(path: str):
with open(path, "r", encoding="utf-8", newline="") as handle:
return list(csv.DictReader(handle))
def retention(args: argparse.Namespace) -> int:
rows = load_csv(args.input)
cohorts = {}
activity = defaultdict(set)
for row in rows:
user = row[args.user_column].strip()
cohort_date = parse_date(row[args.cohort_column])
activity_date = parse_date(row[args.activity_column])
cohorts[user] = min(cohorts.get(user, cohort_date), cohort_date)
delta = (activity_date - cohorts[user]).days
if delta >= 0:
activity[delta].add(user)
base_users = len(cohorts)
if base_users == 0:
print("No users found.", file=sys.stderr)
return 1
results = []
for period in range(0, args.max_period + 1):
users = len(activity.get(period, set()))
rate = users / base_users
results.append({"period": period, "active_users": users, "retention_rate": round(rate, 4)})
if getattr(args, "format", "text") == "json":
print(json.dumps({"base_users": base_users, "periods": results}, indent=2))
else:
print("Retention by period")
print("period,active_users,retention_rate")
for r in results:
print(f"{r['period']},{r['active_users']},{r['retention_rate']:.4f}")
return 0
def cohort(args: argparse.Namespace) -> int:
rows = load_csv(args.input)
cohorts = {}
activity = defaultdict(set)
for row in rows:
user = row[args.user_column].strip()
cohort_date = parse_date(row[args.cohort_column])
activity_date = parse_date(row[args.activity_column])
if args.cohort_grain == "month":
cohort_key = cohort_date.strftime("%Y-%m")
else:
cohort_key = f"{cohort_date.isocalendar().year}-W{cohort_date.isocalendar().week:02d}"
cohorts.setdefault(user, cohort_key)
age = (activity_date - cohort_date).days
if age >= 0:
activity[(cohort_key, age)].add(user)
cohort_sizes = defaultdict(int)
for cohort_key in cohorts.values():
cohort_sizes[cohort_key] += 1
cohort_keys = sorted(cohort_sizes.keys())
results = []
for cohort_key in cohort_keys:
size = cohort_sizes[cohort_key]
for age in range(0, args.max_period + 1):
active = len(activity.get((cohort_key, age), set()))
rate = (active / size) if size else 0
results.append({"cohort": cohort_key, "age_days": age, "active_users": active,
"cohort_size": size, "retention_rate": round(rate, 4)})
if getattr(args, "format", "text") == "json":
print(json.dumps({"cohorts": dict(cohort_sizes), "rows": results}, indent=2))
else:
print("cohort,age_days,active_users,cohort_size,retention_rate")
for r in results:
print(f"{r['cohort']},{r['age_days']},{r['active_users']},{r['cohort_size']},{r['retention_rate']:.4f}")
return 0
def funnel(args: argparse.Namespace) -> int:
rows = load_csv(args.input)
stages = [item.strip() for item in args.stages.split(",") if item.strip()]
if not stages:
print("No stages provided.")
return 1
stage_users = {stage: set() for stage in stages}
for row in rows:
user = row[args.user_column].strip()
stage = row[args.stage_column].strip()
if stage in stage_users:
stage_users[stage].add(user)
results = []
previous_count = None
first_count = None
for stage in stages:
count = len(stage_users[stage])
if first_count is None:
first_count = count
conv_prev = (count / previous_count) if previous_count else 1.0
conv_first = (count / first_count) if first_count else 0
results.append({"stage": stage, "users": count,
"conversion_from_previous": round(conv_prev, 4),
"conversion_from_first": round(conv_first, 4)})
previous_count = count
if getattr(args, "format", "text") == "json":
print(json.dumps({"stages": results}, indent=2))
else:
print("stage,users,conversion_from_previous,conversion_from_first")
for r in results:
print(f"{r['stage']},{r['users']},{r['conversion_from_previous']:.4f},{r['conversion_from_first']:.4f}")
return 0
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
description="Calculate retention, cohort, and funnel metrics from CSV data."
)
subparsers = parser.add_subparsers(dest="command", required=True)
common = {
"help": "CSV input path",
}
fmt_help = "Output format (default: text)"
retention_parser = subparsers.add_parser("retention", help="Calculate retention by day.")
retention_parser.add_argument("input", **common)
retention_parser.add_argument("--user-column", default="user_id")
retention_parser.add_argument("--cohort-column", default="cohort_date")
retention_parser.add_argument("--activity-column", default="activity_date")
retention_parser.add_argument("--max-period", type=int, default=30)
retention_parser.add_argument("--format", choices=["text", "json"], default="text", help=fmt_help)
retention_parser.set_defaults(func=retention)
cohort_parser = subparsers.add_parser("cohort", help="Build cohort retention matrix rows.")
cohort_parser.add_argument("input", **common)
cohort_parser.add_argument("--user-column", default="user_id")
cohort_parser.add_argument("--cohort-column", default="cohort_date")
cohort_parser.add_argument("--activity-column", default="activity_date")
cohort_parser.add_argument("--cohort-grain", choices=["week", "month"], default="week")
cohort_parser.add_argument("--max-period", type=int, default=30)
cohort_parser.add_argument("--format", choices=["text", "json"], default="text", help=fmt_help)
cohort_parser.set_defaults(func=cohort)
funnel_parser = subparsers.add_parser("funnel", help="Calculate funnel conversion by stage.")
funnel_parser.add_argument("input", **common)
funnel_parser.add_argument("--user-column", default="user_id")
funnel_parser.add_argument("--stage-column", default="stage")
funnel_parser.add_argument("--stages", required=True)
funnel_parser.add_argument("--format", choices=["text", "json"], default="text", help=fmt_help)
funnel_parser.set_defaults(func=funnel)
return parser
def main() -> int:
parser = build_parser()
args = parser.parse_args()
try:
return args.func(args)
except FileNotFoundError:
print(f"Error: file not found: {args.input}", file=sys.stderr)
return 1
except KeyError as e:
print(f"Error: column not found in CSV: {e}", file=sys.stderr)
return 1
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
return 1
if __name__ == "__main__":
raise SystemExit(main())
Kiểm chứng cơ hội sản phẩm, lập bản đồ giả định, lên kế hoạch discovery sprint và thử độ khớp vấn đề-giải pháp trước khi đầu tư phát triển.
---
name: product-discovery
description: Use when validating product opportunities, mapping assumptions, planning discovery sprints, or testing problem-solution fit before committing delivery resources.
---
# Product Discovery
Run structured discovery to identify high-value opportunities and de-risk product bets.
## When To Use
Use this skill for:
- Opportunity Solution Tree facilitation
- Assumption mapping and test planning
- Problem validation interviews and evidence synthesis
- Solution validation with prototypes/experiments
- Discovery sprint planning and outputs
## Core Discovery Workflow
1. Define desired outcome
- Set one measurable outcome to improve.
- Establish baseline and target horizon.
2. Build Opportunity Solution Tree (OST)
- Outcome -> opportunities -> solution ideas -> experiments
- Keep opportunities grounded in user evidence, not internal opinions.
3. Map assumptions
- Identify desirability, viability, feasibility, and usability assumptions.
- Score assumptions by risk and certainty.
Use:
```bash
python3 scripts/assumption_mapper.py assumptions.csv
```
4. Validate the problem
- Conduct interviews and behavior analysis.
- Confirm frequency, severity, and willingness to solve.
- Reject weak opportunities early.
5. Validate the solution
- Prototype before building.
- Run concept, usability, and value tests.
- Measure behavior, not only stated preference.
6. Plan discovery sprint
- 1-2 week cycle with explicit hypotheses
- Daily evidence reviews
- End with decision: proceed, pivot, or stop
## Opportunity Solution Tree (Teresa Torres)
Structure:
- Outcome: metric you want to move
- Opportunities: unmet customer needs/pains
- Solutions: candidate interventions
- Experiments: fastest learning actions
Quality checks:
- At least 3 distinct opportunities before converging.
- At least 2 experiments per top opportunity.
- Tie every branch to evidence source.
## Assumption Mapping
Assumption categories:
- Desirability: users want this
- Viability: business value exists
- Feasibility: team can build/operate it
- Usability: users can successfully use it
Prioritization rule:
- High risk + low certainty assumptions are tested first.
## Problem Validation Techniques
- Problem interviews focused on current behavior
- Journey friction mapping
- Support ticket and sales-call synthesis
- Behavioral analytics triangulation
Evidence threshold examples:
- Same pain repeated across multiple target users
- Observable workaround behavior
- Measurable cost of current pain
## Solution Validation Techniques
- Concept tests (value proposition comprehension)
- Prototype usability tests (task success/time-to-complete)
- Fake door or concierge tests (demand signal)
- Limited beta cohorts (retention/activation signals)
## Discovery Sprint Planning
Suggested 10-day structure:
- Day 1-2: Outcome + opportunity framing
- Day 3-4: Assumption mapping + test design
- Day 5-7: Problem and solution tests
- Day 8-9: Evidence synthesis + decision options
- Day 10: Stakeholder decision review
## Tooling
### `scripts/assumption_mapper.py`
CLI utility that:
- reads assumptions from CSV or inline input
- scores risk/certainty priority
- emits prioritized test plan with suggested test types
See `references/discovery-frameworks.md` for framework details.
FILE:references/discovery-frameworks.md
# Discovery Frameworks
## Opportunity Solution Tree (OST)
Purpose: continuously connect product outcomes to validated opportunities and tested solutions.
Core structure:
- Outcome (metric)
- Opportunity nodes (needs/pains)
- Solution ideas
- Experiments
OST practice tips:
- Keep tree live; update after each interview or test.
- Separate opportunity evidence from solution proposals.
- Avoid single-branch trees that force one solution.
## Jobs-to-be-Done (JTBD)
Use JTBD to understand progress users seek.
JTBD template:
"When [situation], I want to [motivation], so I can [expected outcome]."
JTBD interview focus:
- Trigger moments
- Current alternatives and workarounds
- Purchase/adoption anxieties
- Desired progress and success criteria
## Kano Model
Classify features by impact on satisfaction:
- Must-be: expected baseline features
- Performance: more is better
- Delighters: unexpected value multipliers
- Indifferent: low impact
- Reverse: can reduce satisfaction for some users
Use Kano when prioritizing solution concepts after problem validation.
## Design Sprint Methodology
Typical phases:
1. Understand
2. Sketch
3. Decide
4. Prototype
5. Test
Discovery usage:
- Compress learning cycle into one week.
- Best for high-ambiguity opportunities requiring cross-functional alignment.
## Assumption Prioritization Matrix
Map assumptions on two axes:
- Risk if wrong (low -> high)
- Certainty (low -> high)
Priority order:
1. High risk, low certainty (test first)
2. High risk, high certainty (validate quickly)
3. Low risk, low certainty (defer)
4. Low risk, high certainty (document)
## Discovery Evidence Rules
- One source is not enough for major decisions.
- Triangulate qualitative and quantitative signals.
- Predefine decision criteria before test execution.
- Archive evidence with date, segment, and method.
FILE:scripts/assumption_mapper.py
#!/usr/bin/env python3
"""Prioritize product assumptions and suggest validation tests."""
import argparse
import csv
from dataclasses import dataclass
@dataclass
class Assumption:
statement: str
category: str
risk: float
certainty: float
@property
def priority_score(self) -> float:
# High-risk, low-certainty assumptions should be tested first.
return self.risk * (1.0 - self.certainty)
def parse_float(value: str, field: str) -> float:
number = float(value)
if number < 0 or number > 1:
raise ValueError(f"{field} must be in [0, 1]")
return number
def suggest_test(category: str) -> str:
category = category.lower().strip()
if category == "desirability":
return "problem interviews or fake-door test"
if category == "viability":
return "pricing/willingness-to-pay test"
if category == "feasibility":
return "technical spike or architecture prototype"
if category == "usability":
return "moderated usability test"
return "smallest possible experiment with clear success criteria"
def load_from_csv(path: str) -> list[Assumption]:
assumptions: list[Assumption] = []
with open(path, "r", encoding="utf-8", newline="") as handle:
reader = csv.DictReader(handle)
required = {"assumption", "category", "risk", "certainty"}
missing = required - set(reader.fieldnames or [])
if missing:
missing_str = ", ".join(sorted(missing))
raise ValueError(f"Missing required columns: {missing_str}")
for row in reader:
assumptions.append(
Assumption(
statement=(row.get("assumption") or "").strip(),
category=(row.get("category") or "").strip(),
risk=parse_float(row.get("risk") or "0", "risk"),
certainty=parse_float(row.get("certainty") or "0", "certainty"),
)
)
return assumptions
def parse_inline(items: list[str]) -> list[Assumption]:
assumptions: list[Assumption] = []
for item in items:
# format: statement|category|risk|certainty
parts = [part.strip() for part in item.split("|")]
if len(parts) != 4:
raise ValueError("Inline assumption must be: statement|category|risk|certainty")
assumptions.append(
Assumption(
statement=parts[0],
category=parts[1],
risk=parse_float(parts[2], "risk"),
certainty=parse_float(parts[3], "certainty"),
)
)
return assumptions
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(description="Prioritize assumptions and generate test plan.")
parser.add_argument("input", nargs="?", help="CSV file path")
parser.add_argument(
"--assumption",
action="append",
default=[],
help="Inline assumption: statement|category|risk|certainty",
)
parser.add_argument("--top", type=int, default=10, help="Maximum assumptions to print")
return parser
def main() -> int:
parser = build_parser()
args = parser.parse_args()
assumptions: list[Assumption] = []
if args.input:
assumptions.extend(load_from_csv(args.input))
if args.assumption:
assumptions.extend(parse_inline(args.assumption))
if not assumptions:
parser.error("Provide a CSV input file or at least one --assumption value.")
assumptions.sort(key=lambda item: item.priority_score, reverse=True)
print("prioritized_assumption_test_plan")
print("rank,priority_score,category,risk,certainty,test,assumption")
for rank, item in enumerate(assumptions[: args.top], start=1):
test = suggest_test(item.category)
print(
f"{rank},{item.priority_score:.4f},{item.category},{item.risk:.2f},"
f"{item.certainty:.2f},{test},{item.statement}"
)
return 0
if __name__ == "__main__":
raise SystemExit(main())
Bộ công cụ cho PM: ưu tiên RICE, phân tích phỏng vấn khách hàng, mẫu PRD, khung discovery và chiến lược go-to-market.
---
name: "product-manager-toolkit"
description: Comprehensive toolkit for product managers including RICE prioritization, customer interview analysis, PRD templates, discovery frameworks, and go-to-market strategies. Use for feature prioritization, user research synthesis, requirement documentation, and product strategy development.
---
# Product Manager Toolkit
Essential tools and frameworks for modern product management, from discovery to delivery.
---
## Table of Contents
- [Quick Start](#quick-start)
- [Core Workflows](#core-workflows)
- [Feature Prioritization](#feature-prioritization-process)
- [Customer Discovery](#customer-discovery-process)
- [PRD Development](#prd-development-process)
- [Tools Reference](#tools-reference)
- [RICE Prioritizer](#rice-prioritizer)
- [Customer Interview Analyzer](#customer-interview-analyzer)
- [Input/Output Examples](#inputoutput-examples)
- [Integration Points](#integration-points)
- [Common Pitfalls](#common-pitfalls-to-avoid)
---
## Quick Start
### For Feature Prioritization
```bash
# Create sample data file
python scripts/rice_prioritizer.py sample
# Run prioritization with team capacity
python scripts/rice_prioritizer.py sample_features.csv --capacity 15
```
### For Interview Analysis
```bash
python scripts/customer_interview_analyzer.py interview_transcript.txt
```
### For PRD Creation
1. Choose template from `references/prd_templates.md`
2. Fill sections based on discovery work
3. Review with engineering for feasibility
4. Version control in project management tool
---
## Core Workflows
### Feature Prioritization Process
```
Gather → Score → Analyze → Plan → Validate → Execute
```
#### Step 1: Gather Feature Requests
- Customer feedback (support tickets, interviews)
- Sales requests (CRM pipeline blockers)
- Technical debt (engineering input)
- Strategic initiatives (leadership goals)
#### Step 2: Score with RICE
```bash
# Input: CSV with features
python scripts/rice_prioritizer.py features.csv --capacity 20
```
See `references/frameworks.md` for RICE formula and scoring guidelines.
#### Step 3: Analyze Portfolio
Review the tool output for:
- Quick wins vs big bets distribution
- Effort concentration (avoid all XL projects)
- Strategic alignment gaps
#### Step 4: Generate Roadmap
- Quarterly capacity allocation
- Dependency identification
- Stakeholder communication plan
#### Step 5: Validate Results
**Before finalizing the roadmap:**
- [ ] Compare top priorities against strategic goals
- [ ] Run sensitivity analysis (what if estimates are wrong by 2x?)
- [ ] Review with key stakeholders for blind spots
- [ ] Check for missing dependencies between features
- [ ] Validate effort estimates with engineering
#### Step 6: Execute and Iterate
- Share roadmap with team
- Track actual vs estimated effort
- Revisit priorities quarterly
- Update RICE inputs based on learnings
---
### Customer Discovery Process
```
Plan → Recruit → Interview → Analyze → Synthesize → Validate
```
#### Step 1: Plan Research
- Define research questions
- Identify target segments
- Create interview script (see `references/frameworks.md`)
#### Step 2: Recruit Participants
- 5-8 interviews per segment
- Mix of power users and churned users
- Incentivize appropriately
#### Step 3: Conduct Interviews
- Use semi-structured format
- Focus on problems, not solutions
- Record with permission
- Take minimal notes during interview
#### Step 4: Analyze Insights
```bash
python scripts/customer_interview_analyzer.py transcript.txt
```
Extracts:
- Pain points with severity
- Feature requests with priority
- Jobs to be done patterns
- Sentiment and key themes
- Notable quotes
#### Step 5: Synthesize Findings
- Group similar pain points across interviews
- Identify patterns (3+ mentions = pattern)
- Map to opportunity areas using Opportunity Solution Tree
- Prioritize opportunities by frequency and severity
#### Step 6: Validate Solutions
**Before building:**
- [ ] Create solution hypotheses (see `references/frameworks.md`)
- [ ] Test with low-fidelity prototypes
- [ ] Measure actual behavior vs stated preference
- [ ] Iterate based on feedback
- [ ] Document learnings for future research
---
### PRD Development Process
```
Scope → Draft → Review → Refine → Approve → Track
```
#### Step 1: Choose Template
Select from `references/prd_templates.md`:
| Template | Use Case | Timeline |
|----------|----------|----------|
| Standard PRD | Complex features, cross-team | 6-8 weeks |
| One-Page PRD | Simple features, single team | 2-4 weeks |
| Feature Brief | Exploration phase | 1 week |
| Agile Epic | Sprint-based delivery | Ongoing |
#### Step 2: Draft Content
- Lead with problem statement
- Define success metrics upfront
- Explicitly state out-of-scope items
- Include wireframes or mockups
#### Step 3: Review Cycle
- Engineering: feasibility and effort
- Design: user experience gaps
- Sales: market validation
- Support: operational impact
#### Step 4: Refine Based on Feedback
- Address technical constraints
- Adjust scope to fit timeline
- Document trade-off decisions
#### Step 5: Approval and Kickoff
- Stakeholder sign-off
- Sprint planning integration
- Communication to broader team
#### Step 6: Track Execution
**After launch:**
- [ ] Compare actual metrics vs targets
- [ ] Conduct user feedback sessions
- [ ] Document what worked and what didn't
- [ ] Update estimation accuracy data
- [ ] Share learnings with team
---
## Tools Reference
### RICE Prioritizer
Advanced RICE framework implementation with portfolio analysis.
**Features:**
- RICE score calculation with configurable weights
- Portfolio balance analysis (quick wins vs big bets)
- Quarterly roadmap generation based on capacity
- Multiple output formats (text, JSON, CSV)
**CSV Input Format:**
```csv
name,reach,impact,confidence,effort,description
User Dashboard Redesign,5000,high,high,l,Complete redesign
Mobile Push Notifications,10000,massive,medium,m,Add push support
Dark Mode,8000,medium,high,s,Dark theme option
```
**Commands:**
```bash
# Create sample data
python scripts/rice_prioritizer.py sample
# Run with default capacity (10 person-months)
python scripts/rice_prioritizer.py features.csv
# Custom capacity
python scripts/rice_prioritizer.py features.csv --capacity 20
# JSON output for integration
python scripts/rice_prioritizer.py features.csv --output json
# CSV output for spreadsheets
python scripts/rice_prioritizer.py features.csv --output csv
```
---
### Customer Interview Analyzer
NLP-based interview analysis for extracting actionable insights.
**Capabilities:**
- Pain point extraction with severity assessment
- Feature request identification and classification
- Jobs-to-be-done pattern recognition
- Sentiment analysis per section
- Theme and quote extraction
- Competitor mention detection
**Commands:**
```bash
# Analyze interview transcript
python scripts/customer_interview_analyzer.py interview.txt
# JSON output for aggregation
python scripts/customer_interview_analyzer.py interview.txt json
```
---
## Input/Output Examples
→ See references/input-output-examples.md for details
## Integration Points
Compatible tools and platforms:
| Category | Platforms |
|----------|-----------|
| **Analytics** | Amplitude, Mixpanel, Google Analytics |
| **Roadmapping** | ProductBoard, Aha!, Roadmunk, Productplan |
| **Design** | Figma, Sketch, Miro |
| **Development** | Jira, Linear, GitHub, Asana |
| **Research** | Dovetail, UserVoice, Pendo, Maze |
| **Communication** | Slack, Notion, Confluence |
**JSON export enables integration with most tools:**
```bash
# Export for Jira import
python scripts/rice_prioritizer.py features.csv --output json > priorities.json
# Export for dashboard
python scripts/customer_interview_analyzer.py interview.txt json > insights.json
```
---
## Common Pitfalls to Avoid
| Pitfall | Description | Prevention |
|---------|-------------|------------|
| **Solution-First** | Jumping to features before understanding problems | Start every PRD with problem statement |
| **Analysis Paralysis** | Over-researching without shipping | Set time-boxes for research phases |
| **Feature Factory** | Shipping features without measuring impact | Define success metrics before building |
| **Ignoring Tech Debt** | Not allocating time for platform health | Reserve 20% capacity for maintenance |
| **Stakeholder Surprise** | Not communicating early and often | Weekly async updates, monthly demos |
| **Metric Theater** | Optimizing vanity metrics over real value | Tie metrics to user value delivered |
---
## Best Practices
**Writing Great PRDs:**
- Start with the problem, not the solution
- Include clear success metrics upfront
- Explicitly state what's out of scope
- Use visuals (wireframes, flows, diagrams)
- Keep technical details in appendix
- Version control all changes
**Effective Prioritization:**
- Mix quick wins with strategic bets
- Consider opportunity cost of delays
- Account for dependencies between features
- Buffer 20% for unexpected work
- Revisit priorities quarterly
- Communicate decisions with context
**Customer Discovery:**
- Ask "why" five times to find root cause
- Focus on past behavior, not future intentions
- Avoid leading questions ("Wouldn't you love...")
- Interview in the user's natural environment
- Watch for emotional reactions (pain = opportunity)
- Validate qualitative with quantitative data
---
## Quick Reference
```bash
# Prioritization
python scripts/rice_prioritizer.py features.csv --capacity 15
# Interview Analysis
python scripts/customer_interview_analyzer.py interview.txt
# Generate sample data
python scripts/rice_prioritizer.py sample
# JSON outputs
python scripts/rice_prioritizer.py features.csv --output json
python scripts/customer_interview_analyzer.py interview.txt json
```
---
## Reference Documents
- `references/prd_templates.md` - PRD templates for different contexts
- `references/frameworks.md` - Detailed framework documentation (RICE, MoSCoW, Kano, JTBD, etc.)
FILE:assets/prd_template.md
# Product Requirements Document (PRD)
## Document Info
| Field | Value |
|-------|-------|
| **Author** | [Your Name] |
| **Status** | Draft / In Review / Approved |
| **Created** | YYYY-MM-DD |
| **Last Updated** | YYYY-MM-DD |
| **Reviewers** | [Names] |
| **Target Release** | [Quarter or Date] |
---
## Problem Statement
### What problem are we solving?
[Describe the user problem in 2-3 sentences. Focus on the pain, not the solution.]
### Who is affected?
[Identify the user segment(s) experiencing this problem.]
### How do we know this is a problem?
[Link to evidence: interview insights, support tickets, analytics data, churn analysis.]
### What happens if we do nothing?
[Quantify the cost of inaction: lost revenue, churn risk, competitive disadvantage.]
---
## User Stories
| # | As a... | I want to... | So that... | Priority |
|---|---------|-------------|-----------|----------|
| 1 | [role] | [capability] | [benefit] | Must Have |
| 2 | [role] | [capability] | [benefit] | Should Have |
| 3 | [role] | [capability] | [benefit] | Nice to Have |
---
## Solution Overview
### Proposed Solution
[High-level description of what we will build. 3-5 sentences.]
### Key User Flows
[Describe the primary user interactions. Include wireframes or mockups if available.]
1. **Flow 1:** [Description]
2. **Flow 2:** [Description]
3. **Flow 3:** [Description]
### How It Works
[Explain the mechanism or approach. Include technical considerations if relevant.]
---
## Success Metrics
| Metric | Current | Target | Timeframe |
|--------|---------|--------|-----------|
| [Primary metric] | [Baseline] | [Goal] | [When] |
| [Secondary metric] | [Baseline] | [Goal] | [When] |
| [Guardrail metric] | [Baseline] | [Must not worsen] | [When] |
### How We Will Measure
[Describe tracking approach: analytics events, surveys, A/B test design.]
---
## Technical Requirements
### System Requirements
- [Requirement 1: e.g., API response time < 200ms]
- [Requirement 2: e.g., Support 10K concurrent users]
- [Requirement 3: e.g., Mobile responsive]
### Dependencies
- [Dependency 1: e.g., Payment service API update]
- [Dependency 2: e.g., Design system component]
### Security & Privacy
- [Data handling requirements]
- [Authentication/authorization needs]
- [Compliance considerations]
---
## Timeline
| Phase | Dates | Deliverables |
|-------|-------|-------------|
| Design | [Start - End] | Wireframes, user flows, design specs |
| Development | [Start - End] | Feature implementation, unit tests |
| QA | [Start - End] | Test plan execution, bug fixes |
| Beta | [Start - End] | Limited rollout, feedback collection |
| GA | [Date] | Full release, documentation, training |
---
## Risks
| Risk | Likelihood | Impact | Mitigation |
|------|-----------|--------|-----------|
| [Risk 1] | High/Med/Low | High/Med/Low | [Plan] |
| [Risk 2] | High/Med/Low | High/Med/Low | [Plan] |
---
## Out of Scope
The following items are explicitly NOT included in this release:
- [Item 1: brief explanation of why]
- [Item 2: brief explanation of why]
- [Item 3: brief explanation of why]
---
## Decision Log
| # | Decision | Date | Decided By | Rationale |
|---|----------|------|-----------|-----------|
| 1 | [Decision] | [Date] | [Name] | [Why] |
---
## Change History
| Version | Date | Author | Changes |
|---------|------|--------|---------|
| 0.1 | [Date] | [Name] | Initial draft |
FILE:assets/rice_input_template.csv
feature,reach,impact,confidence,effort
Example Feature 1,500,3,0.8,5
Example Feature 2,1000,2,0.9,3
Example Feature 3,300,1,1.0,2
FILE:references/frameworks.md
# Product Management Frameworks
Comprehensive reference for prioritization, discovery, and measurement frameworks.
---
## Table of Contents
- [Prioritization Frameworks](#prioritization-frameworks)
- [RICE Framework](#rice-framework)
- [Value vs Effort Matrix](#value-vs-effort-matrix)
- [MoSCoW Method](#moscow-method)
- [ICE Scoring](#ice-scoring)
- [Kano Model](#kano-model)
- [Discovery Frameworks](#discovery-frameworks)
- [Customer Interview Guide](#customer-interview-guide)
- [Hypothesis Template](#hypothesis-template)
- [Opportunity Solution Tree](#opportunity-solution-tree)
- [Jobs to Be Done](#jobs-to-be-done)
- [Metrics Frameworks](#metrics-frameworks)
- [North Star Metric](#north-star-metric-framework)
- [HEART Framework](#heart-framework)
- [Funnel Analysis](#funnel-analysis-template)
- [Feature Success Metrics](#feature-success-metrics)
- [Strategic Frameworks](#strategic-frameworks)
- [Product Vision Template](#product-vision-template)
- [Competitive Analysis](#competitive-analysis-framework)
- [Go-to-Market Checklist](#go-to-market-checklist)
---
## Prioritization Frameworks
### RICE Framework
**Formula:**
```
RICE Score = (Reach × Impact × Confidence) / Effort
```
**Components:**
| Component | Description | Values |
|-----------|-------------|--------|
| **Reach** | Users affected per quarter | Numeric count (e.g., 5000) |
| **Impact** | Effect on each user | massive=3x, high=2x, medium=1x, low=0.5x, minimal=0.25x |
| **Confidence** | Certainty in estimates | high=100%, medium=80%, low=50% |
| **Effort** | Person-months required | xl=13, l=8, m=5, s=3, xs=1 |
**Example Calculation:**
```
Feature: Mobile Push Notifications
Reach: 10,000 users
Impact: massive (3x)
Confidence: medium (80%)
Effort: medium (5 person-months)
RICE = (10,000 × 3 × 0.8) / 5 = 4,800
```
**Interpretation Guidelines:**
- **1000+**: High priority - strong candidates for next quarter
- **500-999**: Medium priority - consider for roadmap
- **100-499**: Low priority - keep in backlog
- **<100**: Deprioritize - requires new data to reconsider
**When to Use RICE:**
- Quarterly roadmap planning
- Comparing features across different product areas
- Communicating priorities to stakeholders
- Resolving prioritization debates with data
**RICE Limitations:**
- Requires reasonable estimates (garbage in, garbage out)
- Doesn't account for dependencies
- May undervalue platform investments
- Reach estimates can be gaming-prone
---
### Value vs Effort Matrix
```
Low Effort High Effort
+--------------+------------------+
High Value | QUICK WINS | BIG BETS |
| [Do First] | [Strategic] |
+--------------+------------------+
Low Value | FILL-INS | TIME SINKS |
| [Maybe] | [Avoid] |
+--------------+------------------+
```
**Quadrant Definitions:**
| Quadrant | Characteristics | Action |
|----------|-----------------|--------|
| **Quick Wins** | High impact, low effort | Prioritize immediately |
| **Big Bets** | High impact, high effort | Plan strategically, validate ROI |
| **Fill-Ins** | Low impact, low effort | Use to fill sprint gaps |
| **Time Sinks** | Low impact, high effort | Avoid unless required |
**Portfolio Balance:**
- Ideal mix: 40% Quick Wins, 30% Big Bets, 20% Fill-Ins, 10% Buffer
- Review balance quarterly
- Adjust based on team morale and strategic goals
---
### MoSCoW Method
| Category | Definition | Sprint Allocation |
|----------|------------|-------------------|
| **Must Have** | Critical for launch; product fails without it | 60% of capacity |
| **Should Have** | Important but workarounds exist | 20% of capacity |
| **Could Have** | Desirable enhancements | 10% of capacity |
| **Won't Have** | Explicitly out of scope (this release) | 0% - documented |
**Decision Criteria for "Must Have":**
- Regulatory/legal requirement
- Core user job cannot be completed without it
- Explicitly promised to customers
- Security or data integrity requirement
**Common Mistakes:**
- Everything becomes "Must Have" (scope creep)
- Not documenting "Won't Have" items
- Treating "Should Have" as optional (they're important)
- Forgetting to revisit for next release
---
### ICE Scoring
**Formula:**
```
ICE Score = (Impact + Confidence + Ease) / 3
```
| Component | Scale | Description |
|-----------|-------|-------------|
| **Impact** | 1-10 | Expected effect on key metric |
| **Confidence** | 1-10 | How sure are you about impact? |
| **Ease** | 1-10 | How easy to implement? |
**When to Use ICE vs RICE:**
- ICE: Early-stage exploration, quick estimates
- RICE: Quarterly planning, cross-team prioritization
---
### Kano Model
Categories of feature satisfaction:
| Type | Absent | Present | Priority |
|------|--------|---------|----------|
| **Basic (Must-Be)** | Dissatisfied | Neutral | High - table stakes |
| **Performance (Linear)** | Neutral | Satisfied proportionally | Medium - differentiation |
| **Excitement (Delighter)** | Neutral | Very satisfied | Strategic - competitive edge |
| **Indifferent** | Neutral | Neutral | Low - skip unless cheap |
| **Reverse** | Satisfied | Dissatisfied | Avoid - remove if exists |
**Feature Classification Questions:**
1. How would you feel if the product HAS this feature?
2. How would you feel if the product DOES NOT have this feature?
---
## Discovery Frameworks
### Customer Interview Guide
**Structure (35 minutes total):**
```
1. CONTEXT QUESTIONS (5 min)
└── Build rapport, understand role
2. PROBLEM EXPLORATION (15 min)
└── Dig into pain points
3. SOLUTION VALIDATION (10 min)
└── Test concepts if applicable
4. WRAP-UP (5 min)
└── Referrals, follow-up
```
**Detailed Script:**
#### Phase 1: Context (5 min)
```
"Thanks for taking the time. Before we dive in..."
- What's your role and how long have you been in it?
- Walk me through a typical day/week.
- What tools do you use for [relevant task]?
```
#### Phase 2: Problem Exploration (15 min)
```
"I'd love to understand the challenges you face with [area]..."
- What's the hardest part about [task]?
- Can you tell me about the last time you struggled with this?
- What did you do? What happened?
- How often does this happen?
- What does it cost you (time, money, frustration)?
- What have you tried to solve it?
- Why didn't those solutions work?
```
#### Phase 3: Solution Validation (10 min)
```
"Based on what you've shared, I'd like to get your reaction to an idea..."
[Show prototype/concept - keep it rough to invite honest feedback]
- What's your initial reaction?
- How does this compare to what you do today?
- What would prevent you from using this?
- How much would this be worth to you?
- Who else would need to approve this purchase?
```
#### Phase 4: Wrap-up (5 min)
```
"This has been incredibly helpful..."
- Anything else I should have asked?
- Who else should I talk to about this?
- Can I follow up if I have more questions?
```
**Interview Best Practices:**
- Never ask "would you use this?" (people lie about future behavior)
- Ask about past behavior: "Tell me about the last time..."
- Embrace silence - count to 7 before filling gaps
- Watch for emotional reactions (pain = opportunity)
- Record with permission; take minimal notes during
---
### Hypothesis Template
**Format:**
```
We believe that [building this feature/making this change]
For [target user segment]
Will [achieve this measurable outcome]
We'll know we're right when [specific metric moves by X%]
We'll know we're wrong when [falsification criteria]
```
**Example:**
```
We believe that adding saved payment methods
For returning customers
Will increase checkout completion rate
We'll know we're right when checkout completion increases by 15%
We'll know we're wrong when completion rate stays flat after 2 weeks
or saved payment adoption is < 20%
```
**Hypothesis Quality Checklist:**
- [ ] Specific user segment defined
- [ ] Measurable outcome (number, not "better")
- [ ] Timeframe for measurement
- [ ] Clear falsification criteria
- [ ] Based on evidence (interviews, data)
---
### Opportunity Solution Tree
**Structure:**
```
[DESIRED OUTCOME]
│
├── Opportunity 1: [User problem/need]
│ ├── Solution A
│ ├── Solution B
│ └── Experiment: [Test to validate]
│
├── Opportunity 2: [User problem/need]
│ ├── Solution C
│ └── Solution D
│
└── Opportunity 3: [User problem/need]
└── Solution E
```
**Example:**
```
[Increase monthly active users by 20%]
│
├── Users forget to return
│ ├── Weekly email digest
│ ├── Mobile push notifications
│ └── Test: A/B email frequency
│
├── New users don't find value quickly
│ ├── Improved onboarding wizard
│ └── Personalized first experience
│
└── Users churn after free trial
├── Extended trial for engaged users
└── Friction audit of upgrade flow
```
**Process:**
1. Start with measurable outcome (not solution)
2. Map opportunities from user research
3. Generate multiple solutions per opportunity
4. Design small experiments to validate
5. Prioritize based on learning potential
---
### Jobs to Be Done
**JTBD Statement Format:**
```
When [situation/trigger]
I want to [motivation/job]
So I can [expected outcome]
```
**Example:**
```
When I'm running late for a meeting
I want to notify attendees quickly
So I can set appropriate expectations and reduce anxiety
```
**Force Diagram:**
```
┌─────────────────┐
Push from │ │ Pull toward
current ──────>│ SWITCH │<────── new
solution │ DECISION │ solution
│ │
└─────────────────┘
^ ^
| |
Anxiety of | | Habit of
change ──────┘ └────── status quo
```
**Interview Questions for JTBD:**
- When did you first realize you needed something like this?
- What were you using before? Why did you switch?
- What almost prevented you from switching?
- What would make you go back to the old way?
---
## Metrics Frameworks
### North Star Metric Framework
**Criteria for a Good NSM:**
1. **Measures value delivery**: Captures what users get from product
2. **Leading indicator**: Predicts business success
3. **Actionable**: Teams can influence it
4. **Measurable**: Trackable on regular cadence
**Examples by Business Type:**
| Business | North Star Metric | Why |
|----------|-------------------|-----|
| Spotify | Time spent listening | Measures engagement value |
| Airbnb | Nights booked | Core transaction metric |
| Slack | Messages sent in channels | Team collaboration value |
| Dropbox | Files stored/synced | Storage utility delivered |
| Netflix | Hours watched | Entertainment value |
**Supporting Metrics Structure:**
```
[NORTH STAR METRIC]
│
├── Breadth: How many users?
├── Depth: How engaged are they?
└── Frequency: How often do they engage?
```
---
### HEART Framework
| Metric | Definition | Example Signals |
|--------|------------|-----------------|
| **Happiness** | Subjective satisfaction | NPS, CSAT, survey scores |
| **Engagement** | Depth of involvement | Session length, actions/session |
| **Adoption** | New user behavior | Signups, feature activation |
| **Retention** | Continued usage | D7/D30 retention, churn rate |
| **Task Success** | Efficiency & effectiveness | Completion rate, time-on-task, errors |
**Goals-Signals-Metrics Process:**
1. **Goal**: What user behavior indicates success?
2. **Signal**: How would success manifest in data?
3. **Metric**: How do we measure the signal?
**Example:**
```
Feature: New checkout flow
Goal: Users complete purchases faster
Signal: Reduced time in checkout, fewer drop-offs
Metrics:
- Median checkout time (target: <2 min)
- Checkout completion rate (target: 85%)
- Error rate (target: <2%)
```
---
### Funnel Analysis Template
**Standard Funnel:**
```
Acquisition → Activation → Retention → Revenue → Referral
│ │ │ │ │
│ │ │ │ │
How do First Come back Pay for Tell
they find "aha" regularly value others
you? moment
```
**Metrics per Stage:**
| Stage | Key Metrics | Typical Benchmark |
|-------|-------------|-------------------|
| **Acquisition** | Visitors, CAC, channel mix | Varies by channel |
| **Activation** | Signup rate, onboarding completion | 20-30% visitor→signup |
| **Retention** | D1/D7/D30 retention, churn | D1: 40%, D7: 20%, D30: 10% |
| **Revenue** | Conversion rate, ARPU, LTV | 2-5% free→paid |
| **Referral** | NPS, viral coefficient, referrals/user | NPS > 50 is excellent |
**Analysis Framework:**
1. Map current conversion rates at each stage
2. Identify biggest drop-off point
3. Qualitative research: Why are users leaving?
4. Hypothesis: What would improve conversion?
5. Test and measure
---
### Feature Success Metrics
| Metric | Definition | Target Range |
|--------|------------|--------------|
| **Adoption** | % users who try feature | 30-50% within 30 days |
| **Activation** | % who complete core action | 60-80% of adopters |
| **Frequency** | Uses per user per time | Weekly for engagement features |
| **Depth** | % of feature capability used | 50%+ of core functionality |
| **Retention** | Continued usage over time | 70%+ at 30 days |
| **Satisfaction** | Feature-specific NPS/rating | NPS > 30, Rating > 4.0 |
**Measurement Cadence:**
- **Week 1**: Adoption and initial activation
- **Week 4**: Retention and depth
- **Week 8**: Long-term satisfaction and business impact
---
## Strategic Frameworks
### Product Vision Template
**Format:**
```
FOR [target customer]
WHO [statement of need or opportunity]
THE [product name] IS A [product category]
THAT [key benefit, compelling reason to use]
UNLIKE [primary competitive alternative]
OUR PRODUCT [statement of primary differentiation]
```
**Example:**
```
FOR busy professionals
WHO need to stay informed without information overload
Briefme IS A personalized news digest
THAT delivers only relevant stories in 5 minutes
UNLIKE traditional news apps that require active browsing
OUR PRODUCT learns your interests and filters automatically
```
---
### Competitive Analysis Framework
| Dimension | Us | Competitor A | Competitor B |
|-----------|----|--------------|--------------|
| **Target User** | | | |
| **Core Value Prop** | | | |
| **Pricing** | | | |
| **Key Features** | | | |
| **Strengths** | | | |
| **Weaknesses** | | | |
| **Market Position** | | | |
**Strategic Questions:**
1. Where do we have parity? (table stakes)
2. Where do we differentiate? (competitive advantage)
3. Where are we behind? (gaps to close or ignore)
4. What can only we do? (unique capabilities)
---
### Go-to-Market Checklist
**Pre-Launch (4 weeks before):**
- [ ] Success metrics defined and instrumented
- [ ] Launch/rollback criteria established
- [ ] Support documentation ready
- [ ] Sales enablement materials complete
- [ ] Marketing assets prepared
- [ ] Beta feedback incorporated
**Launch Week:**
- [ ] Staged rollout plan (1% → 10% → 50% → 100%)
- [ ] Monitoring dashboards live
- [ ] On-call rotation scheduled
- [ ] Communications ready (in-app, email, blog)
- [ ] Support team briefed
**Post-Launch (2 weeks after):**
- [ ] Metrics review vs. targets
- [ ] User feedback synthesized
- [ ] Bug/issue triage complete
- [ ] Iteration plan defined
- [ ] Stakeholder update sent
---
## Framework Selection Guide
| Situation | Recommended Framework |
|-----------|----------------------|
| Quarterly roadmap planning | RICE + Portfolio Matrix |
| Sprint-level prioritization | MoSCoW |
| Quick feature comparison | ICE |
| Understanding user satisfaction | Kano |
| User research synthesis | JTBD + Opportunity Tree |
| Feature experiment design | Hypothesis Template |
| Success measurement | HEART + Feature Metrics |
| Strategy communication | North Star + Vision |
---
*Last Updated: January 2025*
FILE:references/input-output-examples.md
# product-manager-toolkit reference
## Input/Output Examples
### RICE Prioritizer Example
**Input (features.csv):**
```csv
name,reach,impact,confidence,effort
Onboarding Flow,20000,massive,high,s
Search Improvements,15000,high,high,m
Social Login,12000,high,medium,m
Push Notifications,10000,massive,medium,m
Dark Mode,8000,medium,high,s
```
**Command:**
```bash
python scripts/rice_prioritizer.py features.csv --capacity 15
```
**Output:**
```
============================================================
RICE PRIORITIZATION RESULTS
============================================================
📊 TOP PRIORITIZED FEATURES
1. Onboarding Flow
RICE Score: 16000.0
Reach: 20000 | Impact: massive | Confidence: high | Effort: s
2. Search Improvements
RICE Score: 4800.0
Reach: 15000 | Impact: high | Confidence: high | Effort: m
3. Social Login
RICE Score: 3072.0
Reach: 12000 | Impact: high | Confidence: medium | Effort: m
4. Push Notifications
RICE Score: 3840.0
Reach: 10000 | Impact: massive | Confidence: medium | Effort: m
5. Dark Mode
RICE Score: 2133.33
Reach: 8000 | Impact: medium | Confidence: high | Effort: s
📈 PORTFOLIO ANALYSIS
Total Features: 5
Total Effort: 19 person-months
Total Reach: 65,000 users
Average RICE Score: 5969.07
🎯 Quick Wins: 2 features
• Onboarding Flow (RICE: 16000.0)
• Dark Mode (RICE: 2133.33)
🚀 Big Bets: 0 features
📅 SUGGESTED ROADMAP
Q1 - Capacity: 11/15 person-months
• Onboarding Flow (RICE: 16000.0)
• Search Improvements (RICE: 4800.0)
• Dark Mode (RICE: 2133.33)
Q2 - Capacity: 10/15 person-months
• Push Notifications (RICE: 3840.0)
• Social Login (RICE: 3072.0)
```
---
### Customer Interview Analyzer Example
**Input (interview.txt):**
```
Customer: Jane, Enterprise PM at TechCorp
Date: 2024-01-15
Interviewer: What's the hardest part of your current workflow?
Jane: The biggest frustration is the lack of real-time collaboration.
When I'm working on a PRD, I have to constantly ping my team on Slack
to get updates. It's really frustrating to wait for responses,
especially when we're on a tight deadline.
I've tried using Google Docs for collaboration, but it doesn't
integrate with our roadmap tools. I'd pay extra for something that
just worked seamlessly.
Interviewer: How often does this happen?
Jane: Literally every day. I probably waste 30 minutes just on
back-and-forth messages. It's my biggest pain point right now.
```
**Command:**
```bash
python scripts/customer_interview_analyzer.py interview.txt
```
**Output:**
```
============================================================
CUSTOMER INTERVIEW ANALYSIS
============================================================
📋 INTERVIEW METADATA
Segments found: 1
Lines analyzed: 15
😟 PAIN POINTS (3 found)
1. [HIGH] Lack of real-time collaboration
"I have to constantly ping my team on Slack to get updates"
2. [MEDIUM] Tool integration gaps
"Google Docs...doesn't integrate with our roadmap tools"
3. [HIGH] Time wasted on communication
"waste 30 minutes just on back-and-forth messages"
💡 FEATURE REQUESTS (2 found)
1. Real-time collaboration - Priority: High
2. Seamless tool integration - Priority: Medium
🎯 JOBS TO BE DONE
When working on PRDs with tight deadlines
I want real-time visibility into team updates
So I can avoid wasted time on status checks
📊 SENTIMENT ANALYSIS
Overall: Negative (pain-focused interview)
Key emotions: Frustration, Time pressure
💬 KEY QUOTES
• "It's really frustrating to wait for responses"
• "I'd pay extra for something that just worked seamlessly"
• "It's my biggest pain point right now"
🏷️ THEMES
- Collaboration friction
- Tool fragmentation
- Time efficiency
```
---
FILE:references/prd_templates.md
# Product Requirements Document (PRD) Templates
## Standard PRD Template
### 1. Executive Summary
**Purpose**: One-page overview for executives and stakeholders
#### Components:
- **Problem Statement** (2-3 sentences)
- **Proposed Solution** (2-3 sentences)
- **Business Impact** (3 bullet points)
- **Timeline** (High-level milestones)
- **Resources Required** (Team size and budget)
- **Success Metrics** (3-5 KPIs)
### 2. Problem Definition
#### 2.1 Customer Problem
- **Who**: Target user persona(s)
- **What**: Specific problem or need
- **When**: Context and frequency
- **Where**: Environment and touchpoints
- **Why**: Root cause analysis
- **Impact**: Cost of not solving
#### 2.2 Market Opportunity
- **Market Size**: TAM, SAM, SOM
- **Growth Rate**: Annual growth percentage
- **Competition**: Current solutions and gaps
- **Timing**: Why now?
#### 2.3 Business Case
- **Revenue Potential**: Projected impact
- **Cost Savings**: Efficiency gains
- **Strategic Value**: Alignment with company goals
- **Risk Assessment**: What if we don't do this?
### 3. Solution Overview
#### 3.1 Proposed Solution
- **High-Level Description**: What we're building
- **Key Capabilities**: Core functionality
- **User Journey**: End-to-end flow
- **Differentiation**: Unique value proposition
#### 3.2 In Scope
- Feature 1: Description and priority
- Feature 2: Description and priority
- Feature 3: Description and priority
#### 3.3 Out of Scope
- Explicitly what we're NOT doing
- Future considerations
- Dependencies on other teams
#### 3.4 MVP Definition
- **Core Features**: Minimum viable feature set
- **Success Criteria**: Definition of "working"
- **Timeline**: MVP delivery date
- **Learning Goals**: What we want to validate
### 4. User Stories & Requirements
#### 4.1 User Stories
```
As a [persona]
I want to [action]
So that [outcome/benefit]
Acceptance Criteria:
- [ ] Criterion 1
- [ ] Criterion 2
- [ ] Criterion 3
```
#### 4.2 Functional Requirements
| ID | Requirement | Priority | Notes |
|----|------------|----------|-------|
| FR1 | User can... | P0 | Critical for MVP |
| FR2 | System should... | P1 | Important |
| FR3 | Feature must... | P2 | Nice to have |
#### 4.3 Non-Functional Requirements
- **Performance**: Response times, throughput
- **Scalability**: User/data growth targets
- **Security**: Authentication, authorization, data protection
- **Reliability**: Uptime targets, error rates
- **Usability**: Accessibility standards, device support
- **Compliance**: Regulatory requirements
### 5. Design & User Experience
#### 5.1 Design Principles
- Principle 1: Description
- Principle 2: Description
- Principle 3: Description
#### 5.2 Wireframes/Mockups
- Link to Figma/Sketch files
- Key screens and flows
- Interaction patterns
#### 5.3 Information Architecture
- Navigation structure
- Data organization
- Content hierarchy
### 6. Technical Specifications
#### 6.1 Architecture Overview
- System architecture diagram
- Technology stack
- Integration points
- Data flow
#### 6.2 API Design
- Endpoints and methods
- Request/response formats
- Authentication approach
- Rate limiting
#### 6.3 Database Design
- Data model
- Key entities and relationships
- Migration strategy
#### 6.4 Security Considerations
- Authentication method
- Authorization model
- Data encryption
- PII handling
### 7. Go-to-Market Strategy
#### 7.1 Launch Plan
- **Soft Launch**: Beta users, timeline
- **Full Launch**: All users, timeline
- **Marketing**: Campaigns and channels
- **Support**: Documentation and training
#### 7.2 Pricing Strategy
- Pricing model
- Competitive analysis
- Value proposition
#### 7.3 Success Metrics
| Metric | Target | Measurement Method |
|--------|--------|-------------------|
| Adoption Rate | X% | Daily Active Users |
| User Satisfaction | X/10 | NPS Score |
| Revenue Impact | $X | Monthly Recurring Revenue |
| Performance | <Xms | P95 Response Time |
### 8. Risks & Mitigations
| Risk | Probability | Impact | Mitigation Strategy |
|------|------------|--------|-------------------|
| Technical debt | Medium | High | Allocate 20% for refactoring |
| User adoption | Low | High | Beta program with feedback loops |
| Scope creep | High | Medium | Weekly stakeholder reviews |
### 9. Timeline & Milestones
| Milestone | Date | Deliverables | Success Criteria |
|-----------|------|--------------|-----------------|
| Design Complete | Week 2 | Mockups, IA | Stakeholder approval |
| MVP Development | Week 6 | Core features | All P0s complete |
| Beta Launch | Week 8 | Limited release | 100 beta users |
| Full Launch | Week 12 | General availability | <1% error rate |
### 10. Team & Resources
#### 10.1 Team Structure
- **Product Manager**: [Name]
- **Engineering Lead**: [Name]
- **Design Lead**: [Name]
- **Engineers**: X FTEs
- **QA**: X FTEs
#### 10.2 Budget
- Development: $X
- Infrastructure: $X
- Marketing: $X
- Total: $X
### 11. Appendix
- User Research Data
- Competitive Analysis
- Technical Diagrams
- Legal/Compliance Docs
---
## Agile Epic Template
### Epic: [Epic Name]
#### Overview
**Epic ID**: EPIC-XXX
**Theme**: [Product Theme]
**Quarter**: QX 20XX
**Status**: Discovery | In Progress | Complete
#### Problem Statement
[2-3 sentences describing the problem]
#### Goals & Objectives
1. Objective 1
2. Objective 2
3. Objective 3
#### Success Metrics
- Metric 1: Target
- Metric 2: Target
- Metric 3: Target
#### User Stories
| Story ID | Title | Priority | Points | Status |
|----------|-------|----------|--------|--------|
| US-001 | As a... | P0 | 5 | To Do |
| US-002 | As a... | P1 | 3 | To Do |
#### Dependencies
- Dependency 1: Team/System
- Dependency 2: Team/System
#### Acceptance Criteria
- [ ] All P0 stories complete
- [ ] Performance targets met
- [ ] Security review passed
- [ ] Documentation updated
---
## One-Page PRD Template
### [Feature Name] - One-Page PRD
**Date**: [Date]
**Author**: [PM Name]
**Status**: Draft | In Review | Approved
#### Problem
*What problem are we solving? For whom?*
[2-3 sentences]
#### Solution
*What are we building?*
[2-3 sentences]
#### Why Now?
*What's driving urgency?*
- Reason 1
- Reason 2
- Reason 3
#### Success Metrics
| Metric | Current | Target |
|--------|---------|--------|
| KPI 1 | X | Y |
| KPI 2 | X | Y |
#### Scope
**In**: Feature 1, Feature 2, Feature 3
**Out**: Feature A, Feature B
#### User Flow
```
Step 1 → Step 2 → Step 3 → Success!
```
#### Risks
1. Risk 1 → Mitigation
2. Risk 2 → Mitigation
#### Timeline
- Design: Week 1-2
- Development: Week 3-6
- Testing: Week 7
- Launch: Week 8
#### Resources
- Engineering: X developers
- Design: X designer
- QA: X tester
#### Open Questions
1. Question 1?
2. Question 2?
---
## Feature Brief Template (Lightweight)
### Feature: [Name]
#### Context
*Why are we considering this?*
#### Hypothesis
*We believe that [building this feature]
For [these users]
Will [achieve this outcome]
We'll know we're right when [we see this metric]*
#### Proposed Solution
*High-level approach*
#### Effort Estimate
- **Size**: XS | S | M | L | XL
- **Confidence**: High | Medium | Low
#### Next Steps
1. [ ] User research
2. [ ] Design exploration
3. [ ] Technical spike
4. [ ] Stakeholder review
FILE:scripts/customer_interview_analyzer.py
#!/usr/bin/env python3
"""
Customer Interview Analyzer
Extracts insights, patterns, and opportunities from user interviews
"""
import re
from typing import Dict, List, Tuple, Set
from collections import Counter, defaultdict
import json
class InterviewAnalyzer:
"""Analyze customer interviews for insights and patterns"""
def __init__(self):
# Pain point indicators
self.pain_indicators = [
'frustrat', 'annoy', 'difficult', 'hard', 'confus', 'slow',
'problem', 'issue', 'struggle', 'challeng', 'pain', 'waste',
'manual', 'repetitive', 'tedious', 'boring', 'time-consuming',
'complicated', 'complex', 'unclear', 'wish', 'need', 'want'
]
# Positive indicators
self.delight_indicators = [
'love', 'great', 'awesome', 'amazing', 'perfect', 'easy',
'simple', 'quick', 'fast', 'helpful', 'useful', 'valuable',
'save', 'efficient', 'convenient', 'intuitive', 'clear'
]
# Feature request indicators
self.request_indicators = [
'would be nice', 'wish', 'hope', 'want', 'need', 'should',
'could', 'would love', 'if only', 'it would help', 'suggest',
'recommend', 'idea', 'what if', 'have you considered'
]
# Jobs to be done patterns
self.jtbd_patterns = [
r'when i\s+(.+?),\s+i want to\s+(.+?)\s+so that\s+(.+)',
r'i need to\s+(.+?)\s+because\s+(.+)',
r'my goal is to\s+(.+)',
r'i\'m trying to\s+(.+)',
r'i use \w+ to\s+(.+)',
r'helps me\s+(.+)',
]
def analyze_interview(self, text: str) -> Dict:
"""Analyze a single interview transcript"""
text_lower = text.lower()
sentences = self._split_sentences(text)
analysis = {
'pain_points': self._extract_pain_points(sentences),
'delights': self._extract_delights(sentences),
'feature_requests': self._extract_requests(sentences),
'jobs_to_be_done': self._extract_jtbd(text_lower),
'sentiment_score': self._calculate_sentiment(text_lower),
'key_themes': self._extract_themes(text_lower),
'quotes': self._extract_key_quotes(sentences),
'metrics_mentioned': self._extract_metrics(text),
'competitors_mentioned': self._extract_competitors(text)
}
return analysis
def _split_sentences(self, text: str) -> List[str]:
"""Split text into sentences"""
# Simple sentence splitting
sentences = re.split(r'[.!?]+', text)
return [s.strip() for s in sentences if s.strip()]
def _extract_pain_points(self, sentences: List[str]) -> List[Dict]:
"""Extract pain points from sentences"""
pain_points = []
for sentence in sentences:
sentence_lower = sentence.lower()
for indicator in self.pain_indicators:
if indicator in sentence_lower:
# Extract context around the pain point
pain_points.append({
'quote': sentence,
'indicator': indicator,
'severity': self._assess_severity(sentence_lower)
})
break
return pain_points[:10] # Return top 10
def _extract_delights(self, sentences: List[str]) -> List[Dict]:
"""Extract positive feedback"""
delights = []
for sentence in sentences:
sentence_lower = sentence.lower()
for indicator in self.delight_indicators:
if indicator in sentence_lower:
delights.append({
'quote': sentence,
'indicator': indicator,
'strength': self._assess_strength(sentence_lower)
})
break
return delights[:10]
def _extract_requests(self, sentences: List[str]) -> List[Dict]:
"""Extract feature requests and suggestions"""
requests = []
for sentence in sentences:
sentence_lower = sentence.lower()
for indicator in self.request_indicators:
if indicator in sentence_lower:
requests.append({
'quote': sentence,
'type': self._classify_request(sentence_lower),
'priority': self._assess_request_priority(sentence_lower)
})
break
return requests[:10]
def _extract_jtbd(self, text: str) -> List[Dict]:
"""Extract Jobs to Be Done patterns"""
jobs = []
for pattern in self.jtbd_patterns:
matches = re.findall(pattern, text, re.IGNORECASE)
for match in matches:
if isinstance(match, tuple):
job = ' → '.join(match)
else:
job = match
jobs.append({
'job': job,
'pattern': pattern.pattern if hasattr(pattern, 'pattern') else pattern
})
return jobs[:5]
def _calculate_sentiment(self, text: str) -> Dict:
"""Calculate overall sentiment of the interview"""
positive_count = sum(1 for ind in self.delight_indicators if ind in text)
negative_count = sum(1 for ind in self.pain_indicators if ind in text)
total = positive_count + negative_count
if total == 0:
sentiment_score = 0
else:
sentiment_score = (positive_count - negative_count) / total
if sentiment_score > 0.3:
sentiment_label = 'positive'
elif sentiment_score < -0.3:
sentiment_label = 'negative'
else:
sentiment_label = 'neutral'
return {
'score': round(sentiment_score, 2),
'label': sentiment_label,
'positive_signals': positive_count,
'negative_signals': negative_count
}
def _extract_themes(self, text: str) -> List[str]:
"""Extract key themes using word frequency"""
# Remove common words
stop_words = {'the', 'a', 'an', 'and', 'or', 'but', 'in', 'on', 'at',
'to', 'for', 'of', 'with', 'by', 'from', 'as', 'is',
'was', 'are', 'were', 'been', 'be', 'have', 'has',
'had', 'do', 'does', 'did', 'will', 'would', 'could',
'should', 'may', 'might', 'must', 'can', 'shall',
'it', 'i', 'you', 'we', 'they', 'them', 'their'}
# Extract meaningful words
words = re.findall(r'\b[a-z]{4,}\b', text)
meaningful_words = [w for w in words if w not in stop_words]
# Count frequency
word_freq = Counter(meaningful_words)
# Extract themes (top frequent meaningful words)
themes = [word for word, count in word_freq.most_common(10) if count >= 3]
return themes
def _extract_key_quotes(self, sentences: List[str]) -> List[str]:
"""Extract the most insightful quotes"""
scored_sentences = []
for sentence in sentences:
if len(sentence) < 20 or len(sentence) > 200:
continue
score = 0
sentence_lower = sentence.lower()
# Score based on insight indicators
if any(ind in sentence_lower for ind in self.pain_indicators):
score += 2
if any(ind in sentence_lower for ind in self.request_indicators):
score += 2
if 'because' in sentence_lower:
score += 1
if 'but' in sentence_lower:
score += 1
if '?' in sentence:
score += 1
if score > 0:
scored_sentences.append((score, sentence))
# Sort by score and return top quotes
scored_sentences.sort(reverse=True)
return [s[1] for s in scored_sentences[:5]]
def _extract_metrics(self, text: str) -> List[str]:
"""Extract any metrics or numbers mentioned"""
metrics = []
# Find percentages
percentages = re.findall(r'\d+%', text)
metrics.extend(percentages)
# Find time metrics
time_metrics = re.findall(r'\d+\s*(?:hours?|minutes?|days?|weeks?|months?)', text, re.IGNORECASE)
metrics.extend(time_metrics)
# Find money metrics
money_metrics = re.findall(r'\$[\d,]+', text)
metrics.extend(money_metrics)
# Find general numbers with context
number_contexts = re.findall(r'(\d+)\s+(\w+)', text)
for num, context in number_contexts:
if context.lower() not in ['the', 'a', 'an', 'and', 'or', 'of']:
metrics.append(f"{num} {context}")
return list(set(metrics))[:10]
def _extract_competitors(self, text: str) -> List[str]:
"""Extract competitor mentions"""
# Common competitor indicators
competitor_patterns = [
r'(?:use|used|using|tried|trying|switch from|switched from|instead of)\s+(\w+)',
r'(\w+)\s+(?:is better|works better|is easier)',
r'compared to\s+(\w+)',
r'like\s+(\w+)',
r'similar to\s+(\w+)',
]
competitors = set()
for pattern in competitor_patterns:
matches = re.findall(pattern, text, re.IGNORECASE)
competitors.update(matches)
# Filter out common words
common_words = {'this', 'that', 'it', 'them', 'other', 'another', 'something'}
competitors = [c for c in competitors if c.lower() not in common_words and len(c) > 2]
return list(competitors)[:5]
def _assess_severity(self, text: str) -> str:
"""Assess severity of pain point"""
if any(word in text for word in ['very', 'extremely', 'really', 'totally', 'completely']):
return 'high'
elif any(word in text for word in ['somewhat', 'bit', 'little', 'slightly']):
return 'low'
return 'medium'
def _assess_strength(self, text: str) -> str:
"""Assess strength of positive feedback"""
if any(word in text for word in ['absolutely', 'definitely', 'really', 'very']):
return 'strong'
return 'moderate'
def _classify_request(self, text: str) -> str:
"""Classify the type of request"""
if any(word in text for word in ['ui', 'design', 'look', 'color', 'layout']):
return 'ui_improvement'
elif any(word in text for word in ['feature', 'add', 'new', 'build']):
return 'new_feature'
elif any(word in text for word in ['fix', 'bug', 'broken', 'work']):
return 'bug_fix'
elif any(word in text for word in ['faster', 'slow', 'performance', 'speed']):
return 'performance'
return 'general'
def _assess_request_priority(self, text: str) -> str:
"""Assess priority of request"""
if any(word in text for word in ['critical', 'urgent', 'asap', 'immediately', 'blocking']):
return 'critical'
elif any(word in text for word in ['need', 'important', 'should', 'must']):
return 'high'
elif any(word in text for word in ['nice', 'would', 'could', 'maybe']):
return 'low'
return 'medium'
def aggregate_interviews(interviews: List[Dict]) -> Dict:
"""Aggregate insights from multiple interviews"""
aggregated = {
'total_interviews': len(interviews),
'common_pain_points': defaultdict(list),
'common_requests': defaultdict(list),
'jobs_to_be_done': [],
'overall_sentiment': {
'positive': 0,
'negative': 0,
'neutral': 0
},
'top_themes': Counter(),
'metrics_summary': set(),
'competitors_mentioned': Counter()
}
for interview in interviews:
# Aggregate pain points
for pain in interview.get('pain_points', []):
indicator = pain.get('indicator', 'unknown')
aggregated['common_pain_points'][indicator].append(pain['quote'])
# Aggregate requests
for request in interview.get('feature_requests', []):
req_type = request.get('type', 'general')
aggregated['common_requests'][req_type].append(request['quote'])
# Aggregate JTBD
aggregated['jobs_to_be_done'].extend(interview.get('jobs_to_be_done', []))
# Aggregate sentiment
sentiment = interview.get('sentiment_score', {}).get('label', 'neutral')
aggregated['overall_sentiment'][sentiment] += 1
# Aggregate themes
for theme in interview.get('key_themes', []):
aggregated['top_themes'][theme] += 1
# Aggregate metrics
aggregated['metrics_summary'].update(interview.get('metrics_mentioned', []))
# Aggregate competitors
for competitor in interview.get('competitors_mentioned', []):
aggregated['competitors_mentioned'][competitor] += 1
# Process aggregated data
aggregated['common_pain_points'] = dict(aggregated['common_pain_points'])
aggregated['common_requests'] = dict(aggregated['common_requests'])
aggregated['top_themes'] = dict(aggregated['top_themes'].most_common(10))
aggregated['metrics_summary'] = list(aggregated['metrics_summary'])
aggregated['competitors_mentioned'] = dict(aggregated['competitors_mentioned'])
return aggregated
def format_single_interview(analysis: Dict) -> str:
"""Format single interview analysis"""
output = ["=" * 60]
output.append("CUSTOMER INTERVIEW ANALYSIS")
output.append("=" * 60)
# Sentiment
sentiment = analysis['sentiment_score']
output.append(f"\n📊 Overall Sentiment: {sentiment['label'].upper()}")
output.append(f" Score: {sentiment['score']}")
output.append(f" Positive signals: {sentiment['positive_signals']}")
output.append(f" Negative signals: {sentiment['negative_signals']}")
# Pain Points
if analysis['pain_points']:
output.append("\n🔥 Pain Points Identified:")
for i, pain in enumerate(analysis['pain_points'][:5], 1):
output.append(f"\n{i}. [{pain['severity'].upper()}] {pain['quote'][:100]}...")
# Feature Requests
if analysis['feature_requests']:
output.append("\n💡 Feature Requests:")
for i, req in enumerate(analysis['feature_requests'][:5], 1):
output.append(f"\n{i}. [{req['type']}] Priority: {req['priority']}")
output.append(f" \"{req['quote'][:100]}...\"")
# Jobs to Be Done
if analysis['jobs_to_be_done']:
output.append("\n🎯 Jobs to Be Done:")
for i, job in enumerate(analysis['jobs_to_be_done'], 1):
output.append(f"{i}. {job['job']}")
# Key Themes
if analysis['key_themes']:
output.append("\n🏷️ Key Themes:")
output.append(", ".join(analysis['key_themes']))
# Key Quotes
if analysis['quotes']:
output.append("\n💬 Key Quotes:")
for i, quote in enumerate(analysis['quotes'][:3], 1):
output.append(f'{i}. "{quote}"')
# Metrics
if analysis['metrics_mentioned']:
output.append("\n📈 Metrics Mentioned:")
output.append(", ".join(analysis['metrics_mentioned']))
# Competitors
if analysis['competitors_mentioned']:
output.append("\n🏢 Competitors Mentioned:")
output.append(", ".join(analysis['competitors_mentioned']))
return "\n".join(output)
def main():
import sys
import argparse
parser = argparse.ArgumentParser(
description="Customer Interview Analyzer - Extracts insights, patterns, and opportunities from user interviews"
)
parser.add_argument(
"file", nargs="?", default=None,
help="Interview transcript text file to analyze"
)
parser.add_argument(
"--json", action="store_true",
help="Output results as JSON"
)
args = parser.parse_args()
if not args.file:
print("Usage: python customer_interview_analyzer.py <interview_file.txt>")
print("\nThis tool analyzes customer interview transcripts to extract:")
print(" - Pain points and frustrations")
print(" - Feature requests and suggestions")
print(" - Jobs to be done")
print(" - Sentiment analysis")
print(" - Key themes and quotes")
sys.exit(1)
with open(args.file, 'r') as f:
interview_text = f.read()
analyzer = InterviewAnalyzer()
analysis = analyzer.analyze_interview(interview_text)
if args.json:
print(json.dumps(analysis, indent=2))
else:
print(format_single_interview(analysis))
if __name__ == "__main__":
main()
FILE:scripts/rice_prioritizer.py
#!/usr/bin/env python3
"""
RICE Prioritization Framework
Calculates RICE scores for feature prioritization
RICE = (Reach x Impact x Confidence) / Effort
"""
import json
import csv
from typing import List, Dict, Tuple
import argparse
class RICECalculator:
"""Calculate RICE scores for feature prioritization"""
def __init__(self):
self.impact_map = {
'massive': 3.0,
'high': 2.0,
'medium': 1.0,
'low': 0.5,
'minimal': 0.25
}
self.confidence_map = {
'high': 100,
'medium': 80,
'low': 50
}
self.effort_map = {
'xl': 13,
'l': 8,
'm': 5,
's': 3,
'xs': 1
}
def calculate_rice(self, reach: int, impact: str, confidence: str, effort: str) -> float:
"""
Calculate RICE score
Args:
reach: Number of users/customers affected per quarter
impact: massive/high/medium/low/minimal
confidence: high/medium/low (percentage)
effort: xl/l/m/s/xs (person-months)
"""
impact_score = self.impact_map.get(impact.lower(), 1.0)
confidence_score = self.confidence_map.get(confidence.lower(), 50) / 100
effort_score = self.effort_map.get(effort.lower(), 5)
if effort_score == 0:
return 0
rice_score = (reach * impact_score * confidence_score) / effort_score
return round(rice_score, 2)
def prioritize_features(self, features: List[Dict]) -> List[Dict]:
"""
Calculate RICE scores and rank features
Args:
features: List of feature dictionaries with RICE components
"""
for feature in features:
feature['rice_score'] = self.calculate_rice(
feature.get('reach', 0),
feature.get('impact', 'medium'),
feature.get('confidence', 'medium'),
feature.get('effort', 'm')
)
# Sort by RICE score descending
return sorted(features, key=lambda x: x['rice_score'], reverse=True)
def analyze_portfolio(self, features: List[Dict]) -> Dict:
"""
Analyze the feature portfolio for balance and insights
"""
if not features:
return {}
total_effort = sum(
self.effort_map.get(f.get('effort', 'm').lower(), 5)
for f in features
)
total_reach = sum(f.get('reach', 0) for f in features)
effort_distribution = {}
impact_distribution = {}
for feature in features:
effort = feature.get('effort', 'm').lower()
impact = feature.get('impact', 'medium').lower()
effort_distribution[effort] = effort_distribution.get(effort, 0) + 1
impact_distribution[impact] = impact_distribution.get(impact, 0) + 1
# Calculate quick wins (high impact, low effort)
quick_wins = [
f for f in features
if f.get('impact', '').lower() in ['massive', 'high']
and f.get('effort', '').lower() in ['xs', 's']
]
# Calculate big bets (high impact, high effort)
big_bets = [
f for f in features
if f.get('impact', '').lower() in ['massive', 'high']
and f.get('effort', '').lower() in ['l', 'xl']
]
return {
'total_features': len(features),
'total_effort_months': total_effort,
'total_reach': total_reach,
'average_rice': round(sum(f['rice_score'] for f in features) / len(features), 2),
'effort_distribution': effort_distribution,
'impact_distribution': impact_distribution,
'quick_wins': len(quick_wins),
'big_bets': len(big_bets),
'quick_wins_list': quick_wins[:3], # Top 3 quick wins
'big_bets_list': big_bets[:3] # Top 3 big bets
}
def generate_roadmap(self, features: List[Dict], team_capacity: int = 10) -> List[Dict]:
"""
Generate a quarterly roadmap based on team capacity
Args:
features: Prioritized feature list
team_capacity: Person-months available per quarter
"""
quarters = []
current_quarter = {
'quarter': 1,
'features': [],
'capacity_used': 0,
'capacity_available': team_capacity
}
for feature in features:
effort = self.effort_map.get(feature.get('effort', 'm').lower(), 5)
if current_quarter['capacity_used'] + effort <= team_capacity:
current_quarter['features'].append(feature)
current_quarter['capacity_used'] += effort
else:
# Move to next quarter
current_quarter['capacity_available'] = team_capacity - current_quarter['capacity_used']
quarters.append(current_quarter)
current_quarter = {
'quarter': len(quarters) + 1,
'features': [feature],
'capacity_used': effort,
'capacity_available': team_capacity - effort
}
if current_quarter['features']:
current_quarter['capacity_available'] = team_capacity - current_quarter['capacity_used']
quarters.append(current_quarter)
return quarters
def format_output(features: List[Dict], analysis: Dict, roadmap: List[Dict]) -> str:
"""Format the results for display"""
output = ["=" * 60]
output.append("RICE PRIORITIZATION RESULTS")
output.append("=" * 60)
# Top prioritized features
output.append("\n📊 TOP PRIORITIZED FEATURES\n")
for i, feature in enumerate(features[:10], 1):
output.append(f"{i}. {feature.get('name', 'Unnamed')}")
output.append(f" RICE Score: {feature['rice_score']}")
output.append(f" Reach: {feature.get('reach', 0)} | Impact: {feature.get('impact', 'medium')} | "
f"Confidence: {feature.get('confidence', 'medium')} | Effort: {feature.get('effort', 'm')}")
output.append("")
# Portfolio analysis
output.append("\n📈 PORTFOLIO ANALYSIS\n")
output.append(f"Total Features: {analysis.get('total_features', 0)}")
output.append(f"Total Effort: {analysis.get('total_effort_months', 0)} person-months")
output.append(f"Total Reach: {analysis.get('total_reach', 0):,} users")
output.append(f"Average RICE Score: {analysis.get('average_rice', 0)}")
output.append(f"\n🎯 Quick Wins: {analysis.get('quick_wins', 0)} features")
for qw in analysis.get('quick_wins_list', []):
output.append(f" • {qw.get('name', 'Unnamed')} (RICE: {qw['rice_score']})")
output.append(f"\n🚀 Big Bets: {analysis.get('big_bets', 0)} features")
for bb in analysis.get('big_bets_list', []):
output.append(f" • {bb.get('name', 'Unnamed')} (RICE: {bb['rice_score']})")
# Roadmap
output.append("\n\n📅 SUGGESTED ROADMAP\n")
for quarter in roadmap:
output.append(f"\nQ{quarter['quarter']} - Capacity: {quarter['capacity_used']}/{quarter['capacity_used'] + quarter['capacity_available']} person-months")
for feature in quarter['features']:
output.append(f" • {feature.get('name', 'Unnamed')} (RICE: {feature['rice_score']})")
return "\n".join(output)
def load_features_from_csv(filepath: str) -> List[Dict]:
"""Load features from CSV file"""
features = []
with open(filepath, 'r') as f:
reader = csv.DictReader(f)
for row in reader:
feature = {
'name': row.get('name', ''),
'reach': int(row.get('reach', 0)),
'impact': row.get('impact', 'medium'),
'confidence': row.get('confidence', 'medium'),
'effort': row.get('effort', 'm'),
'description': row.get('description', '')
}
features.append(feature)
return features
def create_sample_csv(filepath: str):
"""Create a sample CSV file for testing"""
sample_features = [
['name', 'reach', 'impact', 'confidence', 'effort', 'description'],
['User Dashboard Redesign', '5000', 'high', 'high', 'l', 'Complete redesign of user dashboard'],
['Mobile Push Notifications', '10000', 'massive', 'medium', 'm', 'Add push notification support'],
['Dark Mode', '8000', 'medium', 'high', 's', 'Implement dark mode theme'],
['API Rate Limiting', '2000', 'low', 'high', 'xs', 'Add rate limiting to API'],
['Social Login', '12000', 'high', 'medium', 'm', 'Add Google/Facebook login'],
['Export to PDF', '3000', 'medium', 'low', 's', 'Export reports as PDF'],
['Team Collaboration', '4000', 'massive', 'low', 'xl', 'Real-time collaboration features'],
['Search Improvements', '15000', 'high', 'high', 'm', 'Enhance search functionality'],
['Onboarding Flow', '20000', 'massive', 'high', 's', 'Improve new user onboarding'],
['Analytics Dashboard', '6000', 'high', 'medium', 'l', 'Advanced analytics for users'],
]
with open(filepath, 'w', newline='') as f:
writer = csv.writer(f)
writer.writerows(sample_features)
print(f"Sample CSV created at: {filepath}")
def main():
parser = argparse.ArgumentParser(description='RICE Framework for Feature Prioritization')
parser.add_argument('input', nargs='?', help='CSV file with features or "sample" to create sample')
parser.add_argument('--capacity', type=int, default=10, help='Team capacity per quarter (person-months)')
parser.add_argument('--output', choices=['text', 'json', 'csv'], default='text', help='Output format')
args = parser.parse_args()
# Create sample if requested
if args.input == 'sample':
create_sample_csv('sample_features.csv')
return
# Use sample data if no input provided
if not args.input:
features = [
{'name': 'User Dashboard', 'reach': 5000, 'impact': 'high', 'confidence': 'high', 'effort': 'l'},
{'name': 'Push Notifications', 'reach': 10000, 'impact': 'massive', 'confidence': 'medium', 'effort': 'm'},
{'name': 'Dark Mode', 'reach': 8000, 'impact': 'medium', 'confidence': 'high', 'effort': 's'},
{'name': 'API Rate Limiting', 'reach': 2000, 'impact': 'low', 'confidence': 'high', 'effort': 'xs'},
{'name': 'Social Login', 'reach': 12000, 'impact': 'high', 'confidence': 'medium', 'effort': 'm'},
]
else:
features = load_features_from_csv(args.input)
# Calculate RICE scores
calculator = RICECalculator()
prioritized = calculator.prioritize_features(features)
analysis = calculator.analyze_portfolio(prioritized)
roadmap = calculator.generate_roadmap(prioritized, args.capacity)
# Output results
if args.output == 'json':
result = {
'features': prioritized,
'analysis': analysis,
'roadmap': roadmap
}
print(json.dumps(result, indent=2))
elif args.output == 'csv':
# Output prioritized features as CSV
if prioritized:
keys = prioritized[0].keys()
print(','.join(keys))
for feature in prioritized:
print(','.join(str(feature.get(k, '')) for k in keys))
else:
print(format_output(prioritized, analysis, roadmap))
if __name__ == "__main__":
main()
Bộ 10 skill sản phẩm: PM toolkit (RICE), PO agile, chiến lược OKR, nghiên cứu UX, design system UI, phân tích đối thủ, landing page, SaaS scaffolder.
--- name: "product-skills" description: "10 product agent skills and plugins for Claude Code, Codex, Gemini CLI, Cursor, OpenClaw. PM toolkit (RICE), agile PO, product strategist (OKR), UX researcher, UI design system, competitive teardown, landing page generator, SaaS scaffolder, research summarizer. Python tools (stdlib-only)." version: 2.9.0 author: Alireza Rezvani license: MIT tags: - product - product-management - ux - ui - saas - agile agents: - claude-code - codex-cli - openclaw --- # Product Team Skills 8 production-ready product skills covering product management, UX/UI design, and SaaS development. ## Quick Start ### Claude Code ``` /read product-team/product-manager-toolkit/SKILL.md ``` ### Codex CLI ```bash npx agent-skills-cli add alirezarezvani/claude-skills/product-team ``` ## Skills Overview | Skill | Folder | Focus | |-------|--------|-------| | Product Manager Toolkit | `product-manager-toolkit/` | RICE prioritization, customer discovery, PRDs | | Agile Product Owner | `agile-product-owner/` | User stories, sprint planning, backlog | | Product Strategist | `product-strategist/` | OKR cascades, market analysis, vision | | UX Researcher Designer | `ux-researcher-designer/` | Personas, journey maps, usability testing | | UI Design System | `ui-design-system/` | Design tokens, component docs, responsive | | Competitive Teardown | `competitive-teardown/` | Systematic competitor analysis | | Landing Page Generator | `landing-page-generator/` | Conversion-optimized pages | | SaaS Scaffolder | `saas-scaffolder/` | Production SaaS boilerplate | ## Python Tools 9 scripts, all stdlib-only: ```bash python3 product-manager-toolkit/scripts/rice_prioritizer.py --help python3 product-strategist/scripts/okr_cascade_generator.py --help ``` ## Rules - Load only the specific skill SKILL.md you need - Use Python tools for scoring and analysis, not manual judgment
Khởi động vòng lặp thử nghiệm tự động theo chu kỳ người dùng chọn (10 phút, 1 giờ, hằng ngày, hằng tuần, hằng tháng) bằng CronCreate.
---
name: "loop"
description: "Start an autonomous experiment loop with user-selected interval (10min, 1h, daily, weekly, monthly). Uses CronCreate for scheduling."
command: /ar:loop
---
# /ar:loop — Autonomous Experiment Loop
Start a recurring experiment loop that runs at a user-selected interval.
## Usage
```
/ar:loop engineering/api-speed # Start loop (prompts for interval)
/ar:loop engineering/api-speed 10m # Every 10 minutes
/ar:loop engineering/api-speed 1h # Every hour
/ar:loop engineering/api-speed daily # Daily at ~9am
/ar:loop engineering/api-speed weekly # Weekly on Monday ~9am
/ar:loop engineering/api-speed monthly # Monthly on 1st ~9am
/ar:loop stop engineering/api-speed # Stop an active loop
```
## What It Does
### Step 1: Resolve experiment
If no experiment specified, list experiments and let user pick.
### Step 2: Select interval
If interval not provided as argument, present options:
```
Select loop interval:
1. Every 10 minutes (rapid — stay and watch)
2. Every hour (background — check back later)
3. Daily at ~9am (overnight experiments)
4. Weekly on Monday (long-running experiments)
5. Monthly on 1st (slow experiments)
```
Map to cron expressions:
| Interval | Cron Expression | Shorthand |
|----------|----------------|-----------|
| 10 minutes | `*/10 * * * *` | `10m` |
| 1 hour | `7 * * * *` | `1h` |
| Daily | `57 8 * * *` | `daily` |
| Weekly | `57 8 * * 1` | `weekly` |
| Monthly | `57 8 1 * *` | `monthly` |
### Step 3: Create the recurring job
Use `CronCreate` with this prompt (fill in the experiment details):
```
You are running autoresearch experiment "{domain}/{name}".
1. Read .autoresearch/{domain}/{name}/config.cfg for: target, evaluate_cmd, metric, metric_direction
2. Read .autoresearch/{domain}/{name}/program.md for strategy and constraints
3. Read .autoresearch/{domain}/{name}/results.tsv for experiment history
4. Run: git checkout autoresearch/{domain}/{name}
Then do exactly ONE iteration:
- Review results.tsv: what worked, what failed, what hasn't been tried
- Edit the target file with ONE change (strategy escalation based on run count)
- Commit: git add {target} && git commit -m "experiment: {description}"
- Evaluate: python {skill_path}/scripts/run_experiment.py --experiment {domain}/{name} --single
- Read the output (KEEP/DISCARD/CRASH)
Rules:
- ONE change per experiment
- NEVER modify the evaluator
- If 5 consecutive crashes in results.tsv, delete this cron job (CronDelete) and alert
- After every 10 experiments, update Strategy section of program.md
Current best metric: {read from results.tsv or "no baseline yet"}
Total experiments so far: {count from results.tsv}
```
### Step 4: Store loop metadata
Write to `.autoresearch/{domain}/{name}/loop.json`:
```json
{
"cron_id": "{id from CronCreate}",
"interval": "{user selection}",
"started": "{ISO timestamp}",
"experiment": "{domain}/{name}"
}
```
### Step 5: Confirm to user
```
Loop started for {domain}/{name}
Interval: {interval description}
Cron ID: {id}
Auto-expires: 3 days (CronCreate limit)
To check progress: /ar:status
To stop the loop: /ar:loop stop {domain}/{name}
Note: Recurring jobs auto-expire after 3 days.
Run /ar:loop again to restart after expiry.
```
## Stopping a Loop
When user runs `/ar:loop stop {experiment}`:
1. Read `.autoresearch/{domain}/{name}/loop.json` to get the cron ID
2. Call `CronDelete` with that ID
3. Delete `loop.json`
4. Confirm: "Loop stopped for {experiment}. {n} experiments completed."
## Important Limitations
- **3-day auto-expiry**: CronCreate jobs expire after 3 days. For longer experiments, the user must re-run `/ar:loop` to restart. Results persist — the new loop picks up where the old one left off.
- **One loop per experiment**: Don't start multiple loops for the same experiment.
- **Concurrent experiments**: Multiple experiments can loop simultaneously ONLY if they're on different git branches (which they are by default — each experiment gets `autoresearch/{domain}/{name}`).
Dashboard sức khỏe danh mục dự án và phân tích ma trận rủi ro.
---
name: project-health
description: Portfolio health dashboard and risk matrix analysis. Usage: /project-health <dashboard|risk> [options]
---
# /project-health
Generate portfolio health dashboards and risk matrices for project oversight.
## Usage
```
/project-health dashboard <project_data.json> Portfolio health dashboard
/project-health risk <risk_data.json> Risk matrix analysis
```
## Input Format
```json
{
"project_name": "Platform Rewrite",
"schedule": {"planned_end": "2026-06-30", "projected_end": "2026-07-15", "milestones_hit": 4, "milestones_total": 6},
"budget": {"allocated": 500000, "spent": 320000, "forecast": 520000},
"scope": {"features_planned": 40, "features_delivered": 28, "change_requests": 3},
"quality": {"defect_rate": 0.05, "test_coverage": 0.82},
"risks": [{"description": "Key engineer leaving", "probability": 0.3, "impact": 0.8}]
}
```
## Examples
```
/project-health dashboard portfolio-q2.json
/project-health risk risk-register.json
/project-health dashboard portfolio-q2.json --format json
```
## Scripts
- `project-management/senior-pm/scripts/project_health_dashboard.py` — Health dashboard (`<data_file> [--format text|json]`)
- `project-management/senior-pm/scripts/risk_matrix_analyzer.py` — Risk matrix analyzer (`<data_file> [--format text|json]`)
## Skill Reference
> `project-management/senior-pm/SKILL.md`
Review pull request, phân tích thay đổi code, kiểm tra vấn đề bảo mật trong PR và đánh giá chất lượng diff.
---
name: "pr-review-expert"
description: "Use when the user asks to review pull requests, analyze code changes, check for security issues in PRs, or assess code quality of diffs."
---
# PR Review Expert
**Tier:** POWERFUL
**Category:** Engineering
**Domain:** Code Review / Quality Assurance
---
## Overview
Structured, systematic code review for GitHub PRs and GitLab MRs. Goes beyond style nits — this skill
performs blast radius analysis, security scanning, breaking change detection, and test coverage delta
calculation. Produces a reviewer-ready report with a 30+ item checklist and prioritized findings.
---
## Core Capabilities
- **Blast radius analysis** — trace which files, services, and downstream consumers could break
- **Security scan** — SQL injection, XSS, auth bypass, secret exposure, dependency vulns
- **Test coverage delta** — new code vs new tests ratio
- **Breaking change detection** — API contracts, DB schema migrations, config keys
- **Ticket linking** — verify Jira/Linear ticket exists and matches scope
- **Performance impact** — N+1 queries, bundle size regression, memory allocations
---
## When to Use
- Before merging any PR/MR that touches shared libraries, APIs, or DB schema
- When a PR is large (>200 lines changed) and needs structured review
- Onboarding new contributors whose PRs need thorough feedback
- Security-sensitive code paths (auth, payments, PII handling)
- After an incident — review similar PRs proactively
---
## Fetching the Diff
### GitHub (gh CLI)
```bash
# View diff in terminal
gh pr diff <PR_NUMBER>
# Get PR metadata (title, body, labels, linked issues)
gh pr view <PR_NUMBER> --json title,body,labels,assignees,milestone
# List files changed
gh pr diff <PR_NUMBER> --name-only
# Check CI status
gh pr checks <PR_NUMBER>
# Download diff to file for analysis
gh pr diff <PR_NUMBER> > /tmp/pr-<PR_NUMBER>.diff
```
### GitLab (glab CLI)
```bash
# View MR diff
glab mr diff <MR_IID>
# MR details as JSON
glab mr view <MR_IID> --output json
# List changed files
glab mr diff <MR_IID> --name-only
# Download diff
glab mr diff <MR_IID> > /tmp/mr-<MR_IID>.diff
```
---
## Workflow
### Step 1 — Fetch Context
```bash
PR=123
gh pr view $PR --json title,body,labels,milestone,assignees | jq .
gh pr diff $PR --name-only
gh pr diff $PR > /tmp/pr-$PR.diff
```
### Step 2 — Blast Radius Analysis
For each changed file, identify:
1. **Direct dependents** — who imports this file?
```bash
# Find all files importing a changed module
grep -r "from ['\"].*changed-module['\"]" src/ --include="*.ts" -l
grep -r "require(['\"].*changed-module" src/ --include="*.js" -l
# Python
grep -r "from changed_module import\|import changed_module" . --include="*.py" -l
```
2. **Service boundaries** — does this change cross a service?
```bash
# Check if changed files span multiple services (monorepo)
gh pr diff $PR --name-only | cut -d/ -f1-2 | sort -u
```
3. **Shared contracts** — types, interfaces, schemas
```bash
gh pr diff $PR --name-only | grep -E "types/|interfaces/|schemas/|models/"
```
**Blast radius severity:**
- CRITICAL — shared library, DB model, auth middleware, API contract
- HIGH — service used by >3 others, shared config, env vars
- MEDIUM — single service internal change, utility function
- LOW — UI component, test file, docs
### Step 3 — Security Scan
```bash
DIFF=/tmp/pr-$PR.diff
# SQL Injection — raw query string interpolation
grep -n "query\|execute\|raw(" $DIFF | grep -E '\$\{|f"|%s|format\('
# Hardcoded secrets
grep -nE "(password|secret|api_key|token|private_key)\s*=\s*['\"][^'\"]{8,}" $DIFF
# AWS key pattern
grep -nE "AKIA[0-9A-Z]{16}" $DIFF
# JWT secret in code
grep -nE "jwt\.sign\(.*['\"][^'\"]{20,}['\"]" $DIFF
# XSS vectors
grep -n "dangerouslySetInnerHTML\|innerHTML\s*=" $DIFF
# Auth bypass patterns
grep -n "bypass\|skip.*auth\|noauth\|TODO.*auth" $DIFF
# Insecure hash algorithms
grep -nE "md5\(|sha1\(|createHash\(['\"]md5|createHash\(['\"]sha1" $DIFF
# eval / exec
grep -nE "\beval\(|\bexec\(|\bsubprocess\.call\(" $DIFF
# Prototype pollution
grep -n "__proto__\|constructor\[" $DIFF
# Path traversal risk
grep -nE "path\.join\(.*req\.|readFile\(.*req\." $DIFF
```
### Step 4 — Test Coverage Delta
```bash
# Count source vs test files changed
CHANGED_SRC=$(gh pr diff $PR --name-only | grep -vE "\.test\.|\.spec\.|__tests__")
CHANGED_TESTS=$(gh pr diff $PR --name-only | grep -E "\.test\.|\.spec\.|__tests__")
echo "Source files changed: $(echo "$CHANGED_SRC" | wc -w)"
echo "Test files changed: $(echo "$CHANGED_TESTS" | wc -w)"
# Lines of new logic vs new test lines
LOGIC_LINES=$(grep "^+" /tmp/pr-$PR.diff | grep -v "^+++" | wc -l)
echo "New lines added: $LOGIC_LINES"
# Run coverage locally
npm test -- --coverage --changedSince=main 2>/dev/null | tail -20
pytest --cov --cov-report=term-missing 2>/dev/null | tail -20
```
**Coverage delta rules:**
- New function without tests → flag
- Deleted tests without deleted code → flag
- Coverage drop >5% → block merge
- Auth/payments paths → require 100% coverage
### Step 5 — Breaking Change Detection
#### API Contract Changes
```bash
# OpenAPI/Swagger spec changes
grep -n "openapi\|swagger" /tmp/pr-$PR.diff | head -20
# REST route removals or renames
grep "^-" /tmp/pr-$PR.diff | grep -E "router\.(get|post|put|delete|patch)\("
# GraphQL schema removals
grep "^-" /tmp/pr-$PR.diff | grep -E "^-\s*(type |field |Query |Mutation )"
# TypeScript interface removals
grep "^-" /tmp/pr-$PR.diff | grep -E "^-\s*(export\s+)?(interface|type) "
```
#### DB Schema Changes
```bash
# Migration files added
gh pr diff $PR --name-only | grep -E "migrations?/|alembic/|knex/"
# Destructive operations
grep -E "DROP TABLE|DROP COLUMN|ALTER.*NOT NULL|TRUNCATE" /tmp/pr-$PR.diff
# Index removals (perf regression risk)
grep "DROP INDEX\|remove_index" /tmp/pr-$PR.diff
```
#### Config / Env Var Changes
```bash
# New env vars referenced in code (might be missing in prod)
grep "^+" /tmp/pr-$PR.diff | grep -oE "process\.env\.[A-Z_]+" | sort -u
# Removed env vars (could break running instances)
grep "^-" /tmp/pr-$PR.diff | grep -oE "process\.env\.[A-Z_]+" | sort -u
```
### Step 6 — Performance Impact
```bash
# N+1 query patterns (DB calls inside loops)
grep -n "\.find\|\.findOne\|\.query\|db\." /tmp/pr-$PR.diff | grep "^+" | head -20
# Then check surrounding context for forEach/map/for loops
# Heavy new dependencies
grep "^+" /tmp/pr-$PR.diff | grep -E '"[a-z@].*":\s*"[0-9^~]' | head -20
# Unbounded loops
grep -n "while (true\|while(true" /tmp/pr-$PR.diff | grep "^+"
# Missing await (accidentally sequential promises)
grep -n "await.*await" /tmp/pr-$PR.diff | grep "^+" | head -10
# Large in-memory allocations
grep -n "new Array([0-9]\{4,\}\|Buffer\.alloc" /tmp/pr-$PR.diff | grep "^+"
```
---
## Ticket Linking Verification
```bash
# Extract ticket references from PR body
gh pr view $PR --json body | jq -r '.body' | \
grep -oE "(PROJ-[0-9]+|[A-Z]+-[0-9]+|https://linear\.app/[^)\"]+)" | sort -u
# Verify Jira ticket exists (requires JIRA_API_TOKEN)
TICKET="PROJ-123"
curl -s -u "user@company.com:$JIRA_API_TOKEN" \
"https://your-org.atlassian.net/rest/api/3/issue/$TICKET" | \
jq '{key, summary: .fields.summary, status: .fields.status.name}'
# Linear ticket
LINEAR_ID="abc-123"
curl -s -H "Authorization: $LINEAR_API_KEY" \
-H "Content-Type: application/json" \
--data "{\"query\": \"{ issue(id: \\\"$LINEAR_ID\\\") { title state { name } } }\"}" \
https://api.linear.app/graphql | jq .
```
---
## Complete Review Checklist (30+ Items)
```markdown
## Code Review Checklist
### Scope & Context
- [ ] PR title accurately describes the change
- [ ] PR description explains WHY, not just WHAT
- [ ] Linked Jira/Linear ticket exists and matches scope
- [ ] No unrelated changes (scope creep)
- [ ] Breaking changes documented in PR body
### Blast Radius
- [ ] Identified all files importing changed modules
- [ ] Cross-service dependencies checked
- [ ] Shared types/interfaces/schemas reviewed for breakage
- [ ] New env vars documented in .env.example
- [ ] DB migrations are reversible (have down() / rollback)
### Security
- [ ] No hardcoded secrets or API keys
- [ ] SQL queries use parameterized inputs (no string interpolation)
- [ ] User inputs validated/sanitized before use
- [ ] Auth/authorization checks on all new endpoints
- [ ] No XSS vectors (innerHTML, dangerouslySetInnerHTML)
- [ ] New dependencies checked for known CVEs
- [ ] No sensitive data in logs (PII, tokens, passwords)
- [ ] File uploads validated (type, size, content-type)
- [ ] CORS configured correctly for new endpoints
### Testing
- [ ] New public functions have unit tests
- [ ] Edge cases covered (empty, null, max values)
- [ ] Error paths tested (not just happy path)
- [ ] Integration tests for API endpoint changes
- [ ] No tests deleted without clear reason
- [ ] Test names clearly describe what they verify
### Breaking Changes
- [ ] No API endpoints removed without deprecation notice
- [ ] No required fields added to existing API responses
- [ ] No DB columns removed without two-phase migration plan
- [ ] No env vars removed that may be set in production
- [ ] Backward-compatible for external API consumers
### Performance
- [ ] No N+1 query patterns introduced
- [ ] DB indexes added for new query patterns
- [ ] No unbounded loops on potentially large datasets
- [ ] No heavy new dependencies without justification
- [ ] Async operations correctly awaited
- [ ] Caching considered for expensive repeated operations
### Code Quality
- [ ] No dead code or unused imports
- [ ] Error handling present (no bare empty catch blocks)
- [ ] Consistent with existing patterns and conventions
- [ ] Complex logic has explanatory comments
- [ ] No unresolved TODOs (or tracked in ticket)
```
---
## Output Format
Structure your review comment as:
```
## PR Review: [PR Title] (#NUMBER)
Blast Radius: HIGH — changes lib/auth used by 5 services
Security: 1 finding (medium severity)
Tests: Coverage delta +2%
Breaking Changes: None detected
--- MUST FIX (Blocking) ---
1. SQL Injection risk in src/db/users.ts:42
Raw string interpolation in WHERE clause.
Fix: db.query("SELECT * WHERE id = $1", [userId])
--- SHOULD FIX (Non-blocking) ---
2. Missing auth check on POST /api/admin/reset
No role verification before destructive operation.
--- SUGGESTIONS ---
3. N+1 pattern in src/services/reports.ts:88
findUser() called inside results.map() — batch with findManyUsers(ids)
--- LOOKS GOOD ---
- Test coverage for new auth flow is thorough
- DB migration has proper down() rollback method
- Error handling consistent with rest of codebase
```
---
## Common Pitfalls
- **Reviewing style over substance** — let the linter handle style; focus on logic, security, correctness
- **Missing blast radius** — a 5-line change in a shared utility can break 20 services
- **Approving untested happy paths** — always verify error paths have coverage
- **Ignoring migration risk** — NOT NULL additions need a default or two-phase migration
- **Indirect secret exposure** — secrets in error messages/logs, not just hardcoded values
- **Skipping large PRs** — if a PR is too large to review properly, request it be split
---
## Best Practices
1. Read the linked ticket before looking at code — context prevents false positives
2. Check CI status before reviewing — don't review code that fails to build
3. Prioritize blast radius and security over style
4. Reproduce locally for non-trivial auth or performance changes
5. Label each comment clearly: "nit:", "must:", "question:", "suggestion:"
6. Batch all comments in one review round — don't trickle feedback
7. Acknowledge good patterns, not just problems — specific praise improves culture
Rà soát chất lượng (QA) cho sản phẩm hoặc đầu ra công việc.
# QA Reviewer Agent ## Vai trò Kiểm tra chất lượng mọi đầu ra trước khi trình bày cho người dùng. ## Nhiệm vụ - Kiểm tra tính logic và nhất quán - Phát hiện nội dung chung chung, thiếu cụ thể - Xác nhận đã đủ thông tin hay cần hỏi thêm - Đề xuất bản sửa nếu chưa đạt chuẩn ## Thang điểm chất lượng (100 điểm) - Logic rõ ràng, không mâu thuẫn: 25 điểm - Đúng ngữ cảnh của người dùng: 25 điểm - Có thể áp dụng ngay: 25 điểm - Không có lỗi trình bày: 15 điểm - Phong cách phù hợp (chi tiết, phân tích): 10 điểm → Chỉ đạt khi đủ 90/100 điểm. → Nếu chưa đạt: phải nêu điểm thiếu và đề xuất bản sửa. ## Lưu ý đặc biệt - Với đầu ra tài chính: luôn kiểm tra xem có nêu giả định chưa - Với kế hoạch: kiểm tra xem có thực tế và có buffer chưa - Với tóm tắt học tập: kiểm tra xem có ví dụ thực tế chưa
Đại diện lãnh đạo về chất lượng (QMR) cho công ty HealthTech/MedTech: quản trị hệ thống chất lượng, xem xét của lãnh đạo, tuân thủ quy định theo ISO 13485.
---
name: "quality-manager-qmr"
description: Senior Quality Manager Responsible Person (QMR) for HealthTech and MedTech companies. Provides quality system governance, management review leadership, regulatory compliance oversight, and quality performance monitoring per ISO 13485 Clause 5.5.2.
triggers:
- management review
- quality policy
- quality objectives
- QMR responsibilities
- quality system effectiveness
- quality KPIs
- cost of quality
- quality performance
- management accountability
- regulatory oversight
- quality culture
- quality governance
---
# Senior Quality Manager Responsible Person (QMR)
Quality system accountability, management review leadership, and regulatory compliance oversight per ISO 13485 Clause 5.5.2 requirements.
---
## Table of Contents
- [QMR Responsibilities](#qmr-responsibilities)
- [Management Review Workflow](#management-review-workflow)
- [Quality KPI Management Workflow](#quality-kpi-management-workflow)
- [Quality Objectives Workflow](#quality-objectives-workflow)
- [Quality Culture Assessment Workflow](#quality-culture-assessment-workflow)
- [Regulatory Compliance Oversight](#regulatory-compliance-oversight)
- [Decision Frameworks](#decision-frameworks)
- [Tools and References](#tools-and-references)
---
## QMR Responsibilities
### ISO 13485 Clause 5.5.2 Requirements
| Responsibility | Scope | Evidence |
|----------------|-------|----------|
| QMS effectiveness | Monitor system performance and suitability | Management review records |
| Reporting to management | Communicate QMS performance to top management | Quality reports, dashboards |
| Quality awareness | Promote regulatory and quality requirements | Training records, communications |
| Liaison with external parties | Interface with regulators, Notified Bodies | Meeting records, correspondence |
### QMR Accountability Matrix
| Domain | Accountable For | Reports To | Frequency |
|--------|-----------------|------------|-----------|
| Quality Policy | Policy adequacy and communication | CEO/Board | Annual review |
| Quality Objectives | Objective achievement and relevance | Executive Team | Quarterly |
| QMS Performance | System effectiveness metrics | Management | Monthly |
| Regulatory Compliance | Compliance status across jurisdictions | CEO | Quarterly |
| Audit Program | Audit schedule completion, findings closure | Management | Per audit |
| CAPA Oversight | CAPA effectiveness and timeliness | Executive Team | Monthly |
### Authority Boundaries
| Decision Type | QMR Authority | Escalation Required |
|---------------|---------------|---------------------|
| Process changes within QMS | Approve with owner | Major process redesign |
| Document approval | Final QA approval | Policy-level changes |
| Nonconformity disposition | Accept/reject with MRB | Product release decisions |
| Supplier quality actions | Quality holds, audits | Supplier termination |
| Audit scheduling | Adjust internal audit schedule | External audit timing |
| Training requirements | Define quality training needs | Organization-wide training budget |
---
## Management Review Workflow
Conduct management reviews per ISO 13485 Clause 5.6 requirements.
### Workflow: Prepare and Execute Management Review
1. Schedule management review (minimum annually, typically quarterly or semi-annually)
2. Notify all required attendees minimum 2 weeks prior
3. Collect required inputs from process owners:
- Audit results (internal and external)
- Customer feedback (complaints, satisfaction, returns)
- Process performance and product conformity
- CAPA status and effectiveness
- Previous review action items
- Changes affecting QMS (regulatory, organizational)
- Recommendations for improvement
4. Compile input summary report with trend analysis
5. Prepare presentation materials with supporting data
6. Distribute agenda and input package 1 week prior
7. Conduct review meeting per agenda
8. **Validation:** All required inputs reviewed; decisions documented with owners and due dates
### Required Attendees
| Role | Requirement | Input Responsibility |
|------|-------------|---------------------|
| CEO/General Manager | Required | Strategic decisions |
| QMR | Chair | Overall QMS status |
| Department Heads | Required | Process performance |
| RA Manager | Required | Regulatory changes |
| Production Manager | Required | Product conformity |
| Customer Quality | Required | Complaint data |
### Management Review Input Template
```
MANAGEMENT REVIEW INPUT SUMMARY
Review Period: [Start Date] to [End Date]
Review Date: [Scheduled Date]
Prepared By: [QMR Name]
1. AUDIT RESULTS
Internal audits completed: [X] of [X] planned
External audits completed: [X]
Total findings: [X] major / [X] minor
Open findings: [X]
Finding trends: [Analysis]
2. CUSTOMER FEEDBACK
Complaints received: [X]
Complaint rate: [X per 1000 units]
Customer satisfaction score: [X.X/5.0]
Returns: [X] units ([X]%)
Top issues: [Categories]
3. PROCESS PERFORMANCE
[Process 1]: [Metric] vs [Target] - [Status]
[Process 2]: [Metric] vs [Target] - [Status]
Out-of-spec processes: [List]
4. PRODUCT CONFORMITY
First pass yield: [X]%
Nonconformance rate: [X]%
Scrap cost: $[X]
Top defect categories: [List]
5. CAPA STATUS
Open CAPAs: [X]
Overdue: [X]
Effectiveness rate: [X]%
Average age: [X] days
6. PREVIOUS ACTIONS
Total from last review: [X]
Completed: [X] | In progress: [X] | Overdue: [X]
7. CHANGES AFFECTING QMS
Regulatory: [List changes]
Organizational: [List changes]
Process: [List changes]
8. RECOMMENDATIONS
[Collected improvement opportunities]
```
### Management Review Output Requirements
| Output | Documentation | Owner |
|--------|---------------|-------|
| QMS improvement decisions | Action items with due dates | Assigned per item |
| Resource needs | Resource plan updates | Department heads |
| Quality objectives changes | Updated objectives document | QMR |
| Process improvement needs | Improvement project charters | Process owners |
See: [references/management-review-guide.md](references/management-review-guide.md)
---
## Quality KPI Management Workflow
Establish, monitor, and report quality performance indicators.
### Workflow: Establish Quality KPI Framework
1. Identify quality objectives requiring measurement
2. Select KPIs per objective using SMART criteria:
- Specific: Clear definition and calculation
- Measurable: Quantifiable with available data
- Actionable: Team can influence results
- Relevant: Aligned to quality objectives
- Time-bound: Defined measurement frequency
3. Define target values based on baseline data and benchmarks
4. Assign data source and collection responsibility
5. Establish reporting frequency per KPI category
6. Configure dashboard displays and trend analysis
7. Define escalation thresholds and alert triggers
8. **Validation:** Each KPI has owner, target, data source, and escalation criteria
### Core Quality KPIs
| Category | KPI | Target | Calculation |
|----------|-----|--------|-------------|
| Process | First Pass Yield | >95% | (Units passed first time / Total units) × 100 |
| Process | Nonconformance Rate | <1% | (NC count / Total units) × 100 |
| CAPA | CAPA Closure Rate | >90% | (On-time closures / Due closures) × 100 |
| CAPA | CAPA Effectiveness | >85% | (Effective CAPAs / Verified CAPAs) × 100 |
| Audit | Finding Closure Rate | >90% | (On-time closures / Due closures) × 100 |
| Audit | Repeat Finding Rate | <10% | (Repeat findings / Total findings) × 100 |
| Customer | Complaint Rate | <0.1% | (Complaints / Units sold) × 100 |
| Customer | Satisfaction Score | >4.0/5.0 | Average of survey scores |
### KPI Review Frequency
| KPI Type | Review Frequency | Trend Period | Audience |
|----------|------------------|--------------|----------|
| Safety/Compliance | Daily monitoring | Weekly | Operations |
| Production Quality | Weekly | Monthly | Department heads |
| Customer Quality | Monthly | Quarterly | Executive team |
| Strategic Quality | Quarterly | Annual | Board/C-suite |
### Performance Response Matrix
| Performance Level | Status | Action Required |
|-------------------|--------|-----------------|
| >110% of target | Exceeding | Consider raising target |
| 100-110% of target | Meeting | Maintain current approach |
| 90-100% of target | Approaching | Monitor closely |
| 80-90% of target | Below | Improvement plan required |
| <80% of target | Critical | Immediate intervention |
See: [references/quality-kpi-framework.md](references/quality-kpi-framework.md)
---
## Quality Objectives Workflow
Establish and maintain measurable quality objectives per ISO 13485 Clause 5.4.1.
### Workflow: Annual Quality Objectives Setting
1. Review prior year objective achievement
2. Analyze quality performance trends and gaps
3. Align with organizational strategic plan
4. Draft objectives with measurable targets
5. Validate resource availability for achievement
6. Obtain executive approval
7. Communicate objectives organization-wide
8. **Validation:** Each objective is measurable, has owner, target, and timeline
### Quality Objective Structure
```
QUALITY OBJECTIVE [Number]
Objective Statement: [Clear, measurable statement]
Aligned to Policy Element: [Quality policy section]
Target: [Specific measurable target]
Baseline: [Current performance]
Owner: [Name and title]
Due Date: [Target achievement date]
Success Criteria:
- [Criterion 1]
- [Criterion 2]
Measurement Method: [How progress is tracked]
Reporting Frequency: [Monthly/Quarterly]
Supporting Initiatives:
- [Initiative 1]
- [Initiative 2]
Resource Requirements:
- [Resource 1]
- [Resource 2]
```
### Objective Categories
| Category | Example Objectives | Typical Targets |
|----------|-------------------|-----------------|
| Customer Quality | Reduce complaint rate | <0.1% of units sold |
| Process Quality | Improve first pass yield | >96% |
| Compliance | Maintain certification | Zero major NCs |
| Efficiency | Reduce quality costs | <4% of revenue |
| Culture | Increase training completion | >98% on-time |
### Quarterly Objective Review
| Review Element | Assessment | Action |
|----------------|------------|--------|
| Progress vs. target | On track / Behind / Ahead | Adjust resources if behind |
| Relevance | Still valid / Needs update | Modify if conditions changed |
| Resources | Adequate / Insufficient | Request additional if needed |
| Barriers | Identified obstacles | Escalate for resolution |
---
## Quality Culture Assessment Workflow
Assess and improve organizational quality culture.
### Workflow: Annual Quality Culture Assessment
1. Design or select quality culture survey instrument
2. Define survey population (all employees or sample)
3. Communicate survey purpose and confidentiality
4. Administer survey with 2-week response window
5. Analyze results by department, role, and tenure
6. Identify strengths and improvement areas
7. Develop action plan for culture gaps
8. **Validation:** Response rate >60%; action plan addresses bottom 3 scores
### Quality Culture Dimensions
| Dimension | Indicators | Assessment Method |
|-----------|------------|-------------------|
| Leadership commitment | Management visible support for quality | Survey, observation |
| Quality ownership | Employees feel responsible for quality | Survey |
| Communication | Quality information flows effectively | Survey, audit |
| Continuous improvement | Suggestions submitted and implemented | Metrics |
| Training and competence | Employees feel adequately trained | Survey, records |
| Problem solving | Issues addressed at root cause | CAPA analysis |
### Culture Survey Categories
| Category | Sample Questions |
|----------|------------------|
| Leadership | "Management demonstrates commitment to quality" |
| Resources | "I have the tools and training to do quality work" |
| Communication | "Quality expectations are clearly communicated" |
| Empowerment | "I am encouraged to report quality issues" |
| Recognition | "Quality achievements are recognized" |
### Culture Improvement Actions
| Gap Identified | Potential Actions |
|----------------|-------------------|
| Low leadership visibility | Quality gemba walks, all-hands quality updates |
| Inadequate training | Competency-based training program |
| Poor communication | Quality newsletters, department huddles |
| Low reporting | Anonymous reporting system, no-blame culture |
| Lack of recognition | Quality award program, team celebrations |
---
## Regulatory Compliance Oversight
Monitor and maintain regulatory compliance across jurisdictions.
### Multi-Jurisdictional Compliance Matrix
| Jurisdiction | Regulation | Requirement | Status Tracking |
|--------------|------------|-------------|-----------------|
| EU | MDR 2017/745 | CE marking, Notified Body | Technical file, annual review |
| USA | 21 CFR 820 | FDA registration, QSR compliance | Annual registration, inspections |
| International | ISO 13485 | QMS certification | Surveillance audits |
| Germany | MPG/MPDG | National implementation | Competent authority filings |
### Compliance Monitoring Workflow
1. Maintain regulatory requirement register
2. Subscribe to regulatory update services
3. Assess impact of regulatory changes monthly
4. Update affected processes within 90 days of effective date
5. Verify training completion for regulatory changes
6. Document compliance status in management review
7. Maintain inspection readiness checklist
8. **Validation:** All applicable requirements mapped; no expired registrations
### Regulatory Authority Interface
| Activity | QMR Role | Preparation Required |
|----------|----------|---------------------|
| Notified Body audit | Primary contact | Audit package, personnel schedules |
| FDA inspection | Host, escort coordinator | Inspection readiness review |
| Competent Authority inquiry | Response coordinator | Technical file access |
| Regulatory meeting | Attendee or delegate | Briefing materials |
### Inspection Readiness Checklist
| Area | Ready | Action Needed |
|------|-------|---------------|
| Document control system current | ☐ | |
| Training records complete | ☐ | |
| CAPA system current, no overdue items | ☐ | |
| Complaint files complete | ☐ | |
| Equipment calibration current | ☐ | |
| Supplier qualification files complete | ☐ | |
| Management review records available | ☐ | |
| Internal audit program current | ☐ | |
---
## Decision Frameworks
### Escalation Decision Tree
```
Issue Identified
│
▼
Is it a regulatory violation?
│
Yes─┴─No
│ │
▼ ▼
Escalate to Is it a safety issue?
Executive │
immediately Yes─┴─No
│ │
▼ ▼
Escalate to Does it affect
Safety Team multiple departments?
│
Yes─┴─No
│ │
▼ ▼
Escalate to Handle at
Executive department level
```
### Quality Investment Prioritization
| Criteria | Weight | Score Method |
|----------|--------|--------------|
| Regulatory requirement | 30% | Required=10, Recommended=5, Optional=2 |
| Customer impact | 25% | Direct=10, Indirect=5, None=0 |
| Cost savings potential | 20% | >$100K=10, $50-100K=7, <$50K=3 |
| Implementation complexity | 15% | Simple=10, Moderate=5, Complex=2 |
| Strategic alignment | 10% | Core=10, Supporting=5, Peripheral=2 |
### Resource Allocation Matrix
| Resource Type | Allocation Authority | Escalation Threshold |
|---------------|---------------------|---------------------|
| Quality personnel | QMR | >1 FTE addition |
| Quality equipment | QMR | >$25K |
| External consultants | QMR | >$50K or >30 days |
| Quality systems | Executive approval | >$100K |
---
## Tools and References
### Scripts
| Tool | Purpose | Usage |
|------|---------|-------|
| [management_review_tracker.py](scripts/management_review_tracker.py) | Track review inputs, actions, metrics | `python management_review_tracker.py --help` |
**Management Review Tracker Features:**
- Track input collection status from process owners
- Monitor action item completion and aging
- Generate metrics summary for review
- Produce recommendations for review focus areas
### References
| Document | Content |
|----------|---------|
| [management-review-guide.md](references/management-review-guide.md) | ISO 13485 Clause 5.6 requirements, input/output templates, action tracking |
| [quality-kpi-framework.md](references/quality-kpi-framework.md) | KPI categories, targets, calculations, dashboard templates |
### Quick Reference: Management Review Inputs (ISO 13485 Clause 5.6.2)
| Input | Source | Required |
|-------|--------|----------|
| Feedback | Customer complaints, surveys | Yes |
| Audit results | Internal and external audits | Yes |
| Process performance | Process metrics | Yes |
| Product conformity | Inspection, NC data | Yes |
| CAPA status | CAPA system | Yes |
| Previous actions | Prior review records | Yes |
| Changes | Regulatory, organizational | Yes |
| Recommendations | All sources | Yes |
### Quick Reference: Management Review Outputs (ISO 13485 Clause 5.6.3)
| Output | Documentation Required |
|--------|----------------------|
| Improvement to QMS and processes | Action items with owners |
| Improvement to product | Project initiation if needed |
| Resource needs | Resource plan updates |
---
## Related Skills
| Skill | Integration Point |
|-------|-------------------|
| [quality-manager-qms-iso13485](../quality-manager-qms-iso13485/) | QMS process management |
| [capa-officer](../capa-officer/) | CAPA system oversight |
| [qms-audit-expert](../qms-audit-expert/) | Internal audit program |
| [quality-documentation-manager](../quality-documentation-manager/) | Document control oversight |
FILE:references/management-review-guide.md
# Management Review Guide
ISO 13485 Clause 5.6 management review requirements, inputs, outputs, and action tracking.
---
## Table of Contents
- [Review Requirements](#review-requirements)
- [Required Inputs](#required-inputs)
- [Review Agenda](#review-agenda)
- [Required Outputs](#required-outputs)
- [Action Tracking](#action-tracking)
- [Documentation Templates](#documentation-templates)
---
## Review Requirements
### ISO 13485:2016 Clause 5.6
| Requirement | Specification |
|-------------|---------------|
| Frequency | Planned intervals (typically quarterly or semi-annually) |
| Participants | Top management involvement required |
| Documentation | Records must be maintained |
| Inputs | All required inputs must be reviewed |
| Outputs | Decisions and actions documented |
### Review Schedule
| Review Type | Frequency | Focus | Participants |
|-------------|-----------|-------|--------------|
| Full Management Review | Semi-annual or Annual | Complete QMS performance | CEO, QMR, all department heads |
| Quarterly Quality Review | Quarterly | Key metrics and actions | QMR, Quality team, affected managers |
| Monthly Quality Update | Monthly | Operational metrics | QMR, Quality team leads |
### Planning Checklist
- [ ] Review date scheduled and communicated
- [ ] Previous review actions status updated
- [ ] All input data collected and analyzed
- [ ] Presentation/report prepared
- [ ] Attendee availability confirmed
- [ ] Meeting room and resources arranged
- [ ] Agenda distributed 1 week in advance
---
## Required Inputs
### ISO 13485 Required Input Topics
| Input | Source | Data Period | Responsible |
|-------|--------|-------------|-------------|
| Audit results | Internal and external audits | Since last review | QA Manager |
| Customer feedback | Complaints, surveys, returns | Since last review | Customer Quality |
| Process performance | Process metrics, yields | Since last review | Process owners |
| Product conformity | Inspection data, NCRs | Since last review | QC Manager |
| CAPA status | Open/closed CAPAs | Current status | CAPA Officer |
| Previous review actions | Action item tracker | Since last review | QMR |
| Changes to QMS | Regulatory, standard changes | Since last review | RA Manager |
| Recommendations | Improvement opportunities | Ongoing collection | All managers |
### Input Data Collection Template
```
MANAGEMENT REVIEW INPUT SUMMARY
Review Period: [Start Date] to [End Date]
Prepared By: [Name]
Date Prepared: [Date]
1. AUDIT RESULTS
Internal Audits Completed: [Number]
External Audits Completed: [Number]
Major Findings: [Number] | Minor Findings: [Number]
Open Audit Actions: [Number]
Summary: [Brief narrative]
2. CUSTOMER FEEDBACK
Total Complaints: [Number]
Complaint Rate: [X per 1000 units]
Customer Satisfaction Score: [Score]
Top Complaint Categories:
- [Category 1]: [Count]
- [Category 2]: [Count]
Trend: [Improving/Stable/Declining]
3. PROCESS PERFORMANCE
| Process | Target | Actual | Status |
|---------|--------|--------|--------|
| [Process 1] | [Target] | [Actual] | [Met/Not Met] |
4. PRODUCT CONFORMITY
First Pass Yield: [%]
Nonconformance Rate: [%]
Reject/Scrap Cost: [$]
Top NC Categories:
- [Category 1]: [Count]
5. CAPA STATUS
Open CAPAs: [Number]
Overdue CAPAs: [Number]
Effectiveness Rate: [%]
Average Closure Time: [Days]
6. PREVIOUS ACTIONS
Total Actions from Last Review: [Number]
Completed: [Number] | In Progress: [Number] | Overdue: [Number]
7. QMS CHANGES
Regulatory Changes: [List]
Standard Updates: [List]
Internal Changes: [List]
8. RECOMMENDATIONS
[List improvement opportunities collected]
```
### Data Analysis Guidelines
| Input | Analysis Required | Red Flags |
|-------|------------------|-----------|
| Audit results | Trend by area, repeat findings | Major NC in same area twice |
| Complaints | Pareto analysis, rate trending | Increasing rate, safety issues |
| Process performance | Control charts, capability | Out of control, Cpk <1.33 |
| Product conformity | Defect Pareto, yield trending | Declining yield, new defect types |
| CAPA | Aging analysis, effectiveness | >10% overdue, <80% effective |
---
## Review Agenda
### Standard Agenda Template
```
MANAGEMENT REVIEW AGENDA
Date: [Date]
Time: [Start] - [End]
Location: [Room/Virtual Link]
Chair: [QMR Name]
1. OPENING (10 min)
- Call to order and attendance
- Approval of previous meeting minutes
- Review of previous action items
2. QMS PERFORMANCE (30 min)
- Audit results summary
- Process performance metrics
- Product conformity data
- Customer feedback analysis
3. COMPLIANCE STATUS (20 min)
- Regulatory compliance status
- Certification status
- Changes affecting QMS
4. CAPA AND IMPROVEMENT (20 min)
- CAPA status and trends
- Improvement initiatives status
- Recommendations for improvement
5. RESOURCE REVIEW (15 min)
- Resource adequacy assessment
- Training and competency status
- Infrastructure needs
6. STRATEGIC ITEMS (15 min)
- Quality objectives progress
- Quality policy adequacy
- Strategic quality initiatives
7. DECISIONS AND ACTIONS (15 min)
- Decisions required
- New action items
- Next review planning
8. CLOSING (5 min)
- Summary of decisions
- Action item review
- Adjournment
```
### Time Allocation by Review Type
| Review Type | Duration | Focus Areas |
|-------------|----------|-------------|
| Full Annual Review | 3-4 hours | All inputs, strategic planning |
| Semi-annual Review | 2-3 hours | All inputs, trend analysis |
| Quarterly Review | 1.5-2 hours | Key metrics, action tracking |
---
## Required Outputs
### ISO 13485 Required Output Topics
| Output | Description | Documentation |
|--------|-------------|---------------|
| Improvement decisions | QMS and process improvements | Action items with owners |
| Resource decisions | Changes to resource allocation | Resource plan updates |
| Quality objectives | Changes to objectives or targets | Updated objectives document |
| QMS changes | Decisions on system modifications | Change requests initiated |
### Output Documentation Template
```
MANAGEMENT REVIEW OUTPUTS
Review Date: [Date]
Review Type: [Annual/Semi-annual/Quarterly]
DECISIONS MADE:
1. QMS IMPROVEMENT DECISIONS
| Decision | Rationale | Owner | Due Date |
|----------|-----------|-------|----------|
| [Decision 1] | [Why] | [Who] | [When] |
2. RESOURCE DECISIONS
| Decision | Resources Required | Budget Impact | Owner |
|----------|-------------------|----------------|-------|
| [Decision 1] | [What needed] | [$] | [Who] |
3. QUALITY OBJECTIVES
| Objective | Current | Target | Change | Rationale |
|-----------|---------|--------|--------|-----------|
| [Objective 1] | [Current target] | [New target] | [+/-] | [Why] |
4. QMS CHANGES APPROVED
| Change | Scope | Implementation Date | Owner |
|--------|-------|---------------------|-------|
| [Change 1] | [Affected areas] | [Date] | [Who] |
CONCLUSIONS:
- Overall QMS effectiveness: [Effective/Needs Improvement]
- Quality policy adequacy: [Adequate/Needs Update]
- Quality objectives progress: [On Track/Behind/Ahead]
NEXT REVIEW:
Date: [Date]
Special Focus Areas: [Areas requiring attention]
```
---
## Action Tracking
### Action Item Format
```
ACTION ITEM
ID: MR-[Year]-[Number]
Source: Management Review [Date]
Category: [ ] Improvement [ ] Resource [ ] Compliance [ ] Other
Description: [Specific action to be taken]
Owner: [Name, Title]
Due Date: [Date]
Priority: [ ] High [ ] Medium [ ] Low
Success Criteria: [How completion will be verified]
Resources Required: [People, budget, equipment]
Dependencies: [Other actions or conditions]
Status Updates:
| Date | Update | Updated By |
|------|--------|------------|
| [Date] | [Progress note] | [Name] |
Completion:
Completed Date: [Date]
Evidence: [Reference to evidence of completion]
Verified By: [Name, Date]
```
### Action Status Categories
| Status | Definition | Color Code |
|--------|------------|------------|
| Not Started | Assigned but work not begun | Gray |
| In Progress | Work underway | Blue |
| On Hold | Blocked, awaiting dependency | Yellow |
| Overdue | Past due date, not complete | Red |
| Complete | Finished, pending verification | Green |
| Verified | Completion verified | Dark Green |
| Cancelled | No longer required | Strikethrough |
### Action Tracking Dashboard
```
MANAGEMENT REVIEW ACTION TRACKER
Review: [Date]
Last Updated: [Date]
SUMMARY:
Total Actions: [Number]
| Status | Count | % |
|--------|-------|---|
| Complete/Verified | [N] | [%] |
| In Progress | [N] | [%] |
| Not Started | [N] | [%] |
| Overdue | [N] | [%] |
| On Hold | [N] | [%] |
OVERDUE ACTIONS (Requires Escalation):
| ID | Description | Owner | Due Date | Days Overdue |
|----|-------------|-------|----------|--------------|
| [ID] | [Brief] | [Name] | [Date] | [Days] |
UPCOMING DUE (Next 30 Days):
| ID | Description | Owner | Due Date |
|----|-------------|-------|----------|
| [ID] | [Brief] | [Name] | [Date] |
```
---
## Documentation Templates
### Meeting Minutes Template
```
MANAGEMENT REVIEW MEETING MINUTES
Date: [Date]
Time: [Start] - [End]
Location: [Location]
Chair: [Name]
Recorder: [Name]
ATTENDEES:
| Name | Title | Present |
|------|-------|---------|
| [Name] | [Title] | ☑ Yes / ☐ No |
AGENDA ITEMS REVIEWED:
1. [Topic]
Discussion: [Summary of discussion]
Decision: [Decision made, if any]
Action: [Action assigned, if any]
2. [Topic]
...
DECISIONS SUMMARY:
1. [Decision 1]
2. [Decision 2]
ACTIONS ASSIGNED:
| ID | Action | Owner | Due Date |
|----|--------|-------|----------|
| MR-XX-01 | [Action] | [Name] | [Date] |
NEXT MEETING:
Date: [Date]
Preliminary Agenda Items: [Topics to cover]
APPROVAL:
Chair: _________________ Date: _______
QMR: _________________ Date: _______
```
### Review Effectiveness Metrics
| Metric | Target | Calculation |
|--------|--------|-------------|
| Action completion rate | >90% | Completed on time / Total actions |
| Review attendance | 100% required | Required attendees present / Required |
| Input completeness | 100% | Inputs provided / Required inputs |
| Decision documentation | 100% | Documented decisions / Decisions made |
| Time to complete review | Per schedule | Actual date - Planned date |
FILE:references/quality-kpi-framework.md
# Quality KPI Framework
Quality performance indicators, targets, and monitoring guidelines for QMS effectiveness.
---
## Table of Contents
- [KPI Categories](#kpi-categories)
- [Core Quality KPIs](#core-quality-kpis)
- [Customer Quality KPIs](#customer-quality-kpis)
- [Compliance KPIs](#compliance-kpis)
- [Cost of Quality](#cost-of-quality)
- [Dashboard Templates](#dashboard-templates)
---
## KPI Categories
### KPI Hierarchy
| Level | Audience | Update Frequency | Example |
|-------|----------|------------------|---------|
| Strategic | Board, C-suite | Quarterly | Quality cost ratio |
| Tactical | Department heads | Monthly | CAPA closure rate |
| Operational | Team leads | Weekly/Daily | First pass yield |
### KPI Selection Criteria
| Criterion | Requirement |
|-----------|-------------|
| Measurable | Quantifiable with available data |
| Actionable | Team can influence the metric |
| Relevant | Aligned to quality objectives |
| Timely | Can be measured at useful frequency |
| Owned | Clear accountability assigned |
---
## Core Quality KPIs
### Process Performance
| KPI | Definition | Target | Calculation |
|-----|------------|--------|-------------|
| First Pass Yield | % units passing without rework | >95% | (Units passed first time / Total units) × 100 |
| Process Capability (Cpk) | Process performance vs. spec | >1.33 | min((USL-μ)/(3σ), (μ-LSL)/(3σ)) |
| Nonconformance Rate | NC events per production volume | <1% | (NC count / Total units) × 100 |
| Right First Time | % activities completed correctly first time | >98% | (Correct completions / Total attempts) × 100 |
### CAPA Effectiveness
| KPI | Definition | Target | Calculation |
|-----|------------|--------|-------------|
| CAPA Closure Rate | % CAPAs closed on time | >90% | (On-time closures / Due closures) × 100 |
| CAPA Effectiveness Rate | % CAPAs effective at verification | >85% | (Effective CAPAs / Verified CAPAs) × 100 |
| Average CAPA Age | Mean days from open to close | <60 days | Sum(Close date - Open date) / Count |
| Overdue CAPA Rate | % CAPAs past due date | <10% | (Overdue CAPAs / Open CAPAs) × 100 |
| Recurrence Rate | % issues recurring after CAPA | <5% | (Recurred issues / Closed CAPAs) × 100 |
### Audit Performance
| KPI | Definition | Target | Calculation |
|-----|------------|--------|-------------|
| Audit Schedule Compliance | % audits completed per schedule | >95% | (Audits completed / Audits scheduled) × 100 |
| Finding Closure Rate | % findings closed on time | >90% | (On-time closures / Due closures) × 100 |
| Repeat Finding Rate | % findings recurring from prior audits | <10% | (Repeat findings / Total findings) × 100 |
| Major NC Rate | Major NCs per audit | <1 | Total major NCs / Total audits |
### Document Control
| KPI | Definition | Target | Calculation |
|-----|------------|--------|-------------|
| Document Review Compliance | % documents reviewed on schedule | >95% | (On-time reviews / Due reviews) × 100 |
| Change Request Cycle Time | Days from request to implementation | <30 days | Average(Implementation - Request date) |
| Obsolete Document Incidents | Uses of obsolete documents | 0 | Count of incidents |
---
## Customer Quality KPIs
### Complaint Management
| KPI | Definition | Target | Calculation |
|-----|------------|--------|-------------|
| Complaint Rate | Complaints per units sold | <0.1% | (Complaints / Units sold) × 100 |
| Complaint Response Time | Days to acknowledge complaint | <24 hours | Average(Response date - Receipt date) |
| Complaint Investigation Time | Days to complete investigation | <30 days | Average(Close date - Receipt date) |
| Complaint Closure Rate | % complaints closed on time | >90% | (On-time closures / Due closures) × 100 |
### Customer Satisfaction
| KPI | Definition | Target | Calculation |
|-----|------------|--------|-------------|
| Customer Satisfaction Score | Survey-based satisfaction rating | >4.0/5.0 | Average of survey scores |
| Net Promoter Score (NPS) | Customer loyalty indicator | >50 | % Promoters - % Detractors |
| Return Rate | % units returned by customers | <1% | (Units returned / Units sold) × 100 |
| Warranty Claim Rate | Warranty claims per units sold | <0.5% | (Claims / Units under warranty) × 100 |
### Field Quality
| KPI | Definition | Target | Calculation |
|-----|------------|--------|-------------|
| Field Failure Rate | Failures in customer use | <0.1% | (Field failures / Units in field) × 100 |
| Mean Time Between Failures | Average operating time before failure | Varies | Total operating hours / Number of failures |
| Service Call Rate | Service calls per installed base | <5%/year | (Service calls / Installed units) × 100 |
---
## Compliance KPIs
### Regulatory Compliance
| KPI | Definition | Target | Calculation |
|-----|------------|--------|-------------|
| Regulatory Submission Success | % submissions accepted first time | >90% | (Accepted submissions / Total submissions) × 100 |
| Inspection Readiness Score | Self-assessment compliance score | >90% | (Compliant items / Total items) × 100 |
| Reportable Event Timeliness | % events reported within required time | 100% | (On-time reports / Required reports) × 100 |
| Registration Currency | % registrations current | 100% | (Current registrations / Required registrations) × 100 |
### Certification Status
| KPI | Definition | Target | Calculation |
|-----|------------|--------|-------------|
| Certification Maintenance | Active certifications vs. required | 100% | (Active certs / Required certs) × 100 |
| Surveillance Audit Outcomes | Pass rate on surveillance audits | 100% | (Passed audits / Conducted audits) × 100 |
| Certification NC Rate | NCs per certification audit | <3 minor, 0 major | Count per audit |
### Training Compliance
| KPI | Definition | Target | Calculation |
|-----|------------|--------|-------------|
| Training Completion Rate | % required training completed | >95% | (Completed / Required) × 100 |
| Training Currency | % employees with current training | >98% | (Current / Total requiring) × 100 |
| Training Effectiveness | % passing competency assessments | >90% | (Passed / Assessed) × 100 |
---
## Cost of Quality
### Cost Categories
| Category | Definition | Examples |
|----------|------------|----------|
| Prevention | Costs to prevent defects | Training, quality planning, process validation |
| Appraisal | Costs to detect defects | Inspection, testing, audits, calibration |
| Internal Failure | Costs of defects found internally | Rework, scrap, re-inspection, downgrading |
| External Failure | Costs of defects found by customer | Returns, complaints, warranty, recalls |
### Cost of Quality KPIs
| KPI | Definition | Target | Calculation |
|-----|------------|--------|-------------|
| Total Cost of Quality | Sum of all quality costs | <5% of revenue | Prevention + Appraisal + Failure costs |
| Prevention/Appraisal Ratio | Prevention vs. detection investment | >1.0 | Prevention costs / Appraisal costs |
| Failure Cost Ratio | Failure costs as % of CoQ | <30% | (Internal + External failure) / Total CoQ |
| Quality Cost Trend | Change in CoQ over time | Decreasing | (Current CoQ - Prior CoQ) / Prior CoQ |
### Cost Collection Categories
```
COST OF QUALITY WORKSHEET
Period: [Start] to [End]
PREVENTION COSTS:
| Category | Description | Amount |
|----------|-------------|--------|
| Quality planning | QMS development, quality planning | $ |
| Training | Quality training programs | $ |
| Process validation | Validation activities | $ |
| Supplier qualification | Supplier quality programs | $ |
| Preventive maintenance | Equipment maintenance | $ |
| SUBTOTAL PREVENTION | | $ |
APPRAISAL COSTS:
| Category | Description | Amount |
|----------|-------------|--------|
| Incoming inspection | Supplier material inspection | $ |
| In-process inspection | Production quality checks | $ |
| Final inspection | Finished goods testing | $ |
| Audit costs | Internal and external audits | $ |
| Calibration | Equipment calibration | $ |
| SUBTOTAL APPRAISAL | | $ |
INTERNAL FAILURE COSTS:
| Category | Description | Amount |
|----------|-------------|--------|
| Scrap | Scrapped materials and product | $ |
| Rework | Labor and materials to correct | $ |
| Re-inspection | Repeat inspection costs | $ |
| Downgrading | Revenue loss from downgrading | $ |
| Root cause analysis | Investigation costs | $ |
| SUBTOTAL INTERNAL FAILURE | | $ |
EXTERNAL FAILURE COSTS:
| Category | Description | Amount |
|----------|-------------|--------|
| Returns processing | Handling returned product | $ |
| Warranty costs | Warranty claims and repairs | $ |
| Complaint handling | Investigation and resolution | $ |
| Recalls | Recall execution costs | $ |
| Liability | Legal and settlement costs | $ |
| SUBTOTAL EXTERNAL FAILURE | | $ |
TOTAL COST OF QUALITY: $
AS % OF REVENUE: %
```
---
## Dashboard Templates
### Executive Quality Dashboard
```
EXECUTIVE QUALITY DASHBOARD
Period: [Month/Quarter]
KEY METRICS AT A GLANCE:
┌─────────────────┬─────────┬─────────┬─────────┐
│ Metric │ Target │ Actual │ Trend │
├─────────────────┼─────────┼─────────┼─────────┤
│ Customer Sat │ >4.0 │ [X.X] │ [↑/↓/→] │
│ Complaint Rate │ <0.1% │ [X.XX%] │ [↑/↓/→] │
│ First Pass Yield│ >95% │ [XX%] │ [↑/↓/→] │
│ CAPA Closure │ >90% │ [XX%] │ [↑/↓/→] │
│ Audit Findings │ <3/audit│ [X.X] │ [↑/↓/→] │
│ Quality Cost │ <5% │ [X.X%] │ [↑/↓/→] │
└─────────────────┴─────────┴─────────┴─────────┘
ALERTS:
[ ] Critical: [Any critical issues requiring immediate attention]
[ ] Warning: [Issues approaching threshold]
[ ] Info: [Notable improvements or changes]
QUALITY OBJECTIVES PROGRESS:
| Objective | Target | YTD | Status |
|-----------|--------|-----|--------|
| [Obj 1] | [Target] | [Actual] | [On Track/Behind] |
```
### Operational Quality Dashboard
```
OPERATIONAL QUALITY DASHBOARD
Week/Month: [Period]
PRODUCTION QUALITY:
├── First Pass Yield: [XX%] (Target: 95%)
├── Rework Rate: [X.X%] (Target: <2%)
├── Scrap Rate: [X.X%] (Target: <1%)
└── NC Count: [XX] (Prior: [XX])
CAPA STATUS:
├── Open CAPAs: [XX]
│ ├── Critical: [X]
│ ├── Major: [XX]
│ └── Minor: [XX]
├── Overdue: [X] [!ALERT if >0]
├── Avg Age: [XX] days
└── Closed This Period: [XX]
AUDIT STATUS:
├── Audits Completed: [X] of [X] scheduled
├── Open Findings: [XX]
│ ├── Major: [X]
│ └── Minor: [XX]
└── Overdue Actions: [X]
COMPLAINTS:
├── Received: [XX]
├── Open: [XX]
├── Avg Response Time: [X.X] days
└── Top Category: [Category]
```
### KPI Target Setting Guidelines
| Performance Level | Action |
|-------------------|--------|
| >110% of target | Consider raising target |
| 100-110% of target | Maintain current target |
| 90-100% of target | Monitor closely |
| 80-90% of target | Improvement plan required |
| <80% of target | Immediate intervention |
### Review Frequency by KPI Type
| KPI Type | Review Frequency | Trend Period |
|----------|------------------|--------------|
| Safety/Compliance | Daily monitoring | Weekly |
| Production | Daily/Weekly | Monthly |
| Customer | Weekly/Monthly | Quarterly |
| Strategic | Monthly/Quarterly | Annual |
| Cost | Monthly | Quarterly |
FILE:scripts/management_review_tracker.py
#!/usr/bin/env python3
"""
Management Review Tracker - QMS Management Review Preparation and Tracking
Tracks management review inputs, action items, and generates review reports
for ISO 13485 compliance.
Usage:
python management_review_tracker.py --data review_data.json
python management_review_tracker.py --interactive
python management_review_tracker.py --data review_data.json --output json
"""
import argparse
import json
import sys
from dataclasses import dataclass, field, asdict
from datetime import datetime, timedelta
from typing import List, Dict, Optional
from enum import Enum
class ActionStatus(Enum):
NOT_STARTED = "Not Started"
IN_PROGRESS = "In Progress"
ON_HOLD = "On Hold"
OVERDUE = "Overdue"
COMPLETE = "Complete"
VERIFIED = "Verified"
class ActionPriority(Enum):
HIGH = "High"
MEDIUM = "Medium"
LOW = "Low"
class InputStatus(Enum):
NOT_COLLECTED = "Not Collected"
IN_PROGRESS = "In Progress"
COMPLETE = "Complete"
REVIEWED = "Reviewed"
@dataclass
class ReviewInput:
topic: str
responsible: str
status: InputStatus
data_period: str
summary: str = ""
concerns: List[str] = field(default_factory=list)
@dataclass
class ActionItem:
action_id: str
description: str
owner: str
due_date: str
priority: ActionPriority
status: ActionStatus
source_review: str
category: str = "Improvement"
completion_date: Optional[str] = None
notes: str = ""
@dataclass
class ReviewMetrics:
complaint_rate: float = 0.0
complaint_count: int = 0
capa_open: int = 0
capa_overdue: int = 0
capa_effectiveness: float = 0.0
audit_findings_open: int = 0
audit_findings_major: int = 0
first_pass_yield: float = 0.0
customer_satisfaction: float = 0.0
training_compliance: float = 0.0
@dataclass
class ManagementReview:
review_date: str
review_type: str
period_start: str
period_end: str
inputs: List[ReviewInput]
actions: List[ActionItem]
metrics: ReviewMetrics
decisions: List[str] = field(default_factory=list)
attendees: List[str] = field(default_factory=list)
class ManagementReviewTracker:
"""Tracks and reports management review status."""
# Required ISO 13485 inputs
REQUIRED_INPUTS = [
("Audit Results", "QA Manager"),
("Customer Feedback", "Customer Quality"),
("Process Performance", "Operations"),
("Product Conformity", "QC Manager"),
("CAPA Status", "CAPA Officer"),
("Previous Actions", "QMR"),
("QMS Changes", "RA Manager"),
("Recommendations", "All Managers"),
]
def __init__(self, review: ManagementReview):
self.review = review
self.today = datetime.now()
def check_input_readiness(self) -> Dict:
"""Check readiness of all required inputs."""
readiness = {
"total_required": len(self.REQUIRED_INPUTS),
"complete": 0,
"in_progress": 0,
"not_started": 0,
"missing_topics": [],
"readiness_score": 0.0
}
input_topics = {inp.topic: inp for inp in self.review.inputs}
for topic, responsible in self.REQUIRED_INPUTS:
if topic in input_topics:
inp = input_topics[topic]
if inp.status in [InputStatus.COMPLETE, InputStatus.REVIEWED]:
readiness["complete"] += 1
elif inp.status == InputStatus.IN_PROGRESS:
readiness["in_progress"] += 1
else:
readiness["not_started"] += 1
else:
readiness["missing_topics"].append(topic)
readiness["not_started"] += 1
readiness["readiness_score"] = round(
(readiness["complete"] / readiness["total_required"]) * 100, 1
)
return readiness
def analyze_actions(self) -> Dict:
"""Analyze action item status."""
analysis = {
"total": len(self.review.actions),
"by_status": {},
"by_priority": {},
"overdue": [],
"due_soon": [],
"completion_rate": 0.0
}
completed = 0
for action in self.review.actions:
# Count by status
status = action.status.value
analysis["by_status"][status] = analysis["by_status"].get(status, 0) + 1
# Count by priority
priority = action.priority.value
analysis["by_priority"][priority] = analysis["by_priority"].get(priority, 0) + 1
# Check completion
if action.status in [ActionStatus.COMPLETE, ActionStatus.VERIFIED]:
completed += 1
# Check overdue
if action.due_date:
due = datetime.strptime(action.due_date, "%Y-%m-%d")
if due < self.today and action.status not in [
ActionStatus.COMPLETE, ActionStatus.VERIFIED
]:
days_overdue = (self.today - due).days
analysis["overdue"].append({
"action_id": action.action_id,
"description": action.description[:50],
"owner": action.owner,
"days_overdue": days_overdue
})
elif due <= self.today + timedelta(days=14) and action.status not in [
ActionStatus.COMPLETE, ActionStatus.VERIFIED
]:
days_until = (due - self.today).days
analysis["due_soon"].append({
"action_id": action.action_id,
"description": action.description[:50],
"owner": action.owner,
"days_until_due": days_until
})
if analysis["total"] > 0:
analysis["completion_rate"] = round((completed / analysis["total"]) * 100, 1)
return analysis
def assess_metrics(self) -> Dict:
"""Assess quality metrics against targets."""
metrics = self.review.metrics
assessment = {
"metrics": [],
"alerts": [],
"overall_status": "On Track"
}
# Define targets and assess
checks = [
("Complaint Rate", metrics.complaint_rate, 0.1, "lower"),
("CAPA Overdue", metrics.capa_overdue, 0, "lower"),
("CAPA Effectiveness", metrics.capa_effectiveness, 85.0, "higher"),
("First Pass Yield", metrics.first_pass_yield, 95.0, "higher"),
("Customer Satisfaction", metrics.customer_satisfaction, 4.0, "higher"),
("Training Compliance", metrics.training_compliance, 95.0, "higher"),
]
warnings = 0
critical = 0
for name, value, target, direction in checks:
if direction == "lower":
status = "Pass" if value <= target else "Fail"
threshold = target * 1.2
warning = value > target and value <= threshold
else:
status = "Pass" if value >= target else "Fail"
threshold = target * 0.9
warning = value < target and value >= threshold
metric_result = {
"name": name,
"value": value,
"target": target,
"status": status
}
assessment["metrics"].append(metric_result)
if status == "Fail":
if warning:
warnings += 1
assessment["alerts"].append(f"WARNING: {name} at {value} (target: {target})")
else:
critical += 1
assessment["alerts"].append(f"CRITICAL: {name} at {value} (target: {target})")
if critical > 0:
assessment["overall_status"] = "Critical"
elif warnings > 0:
assessment["overall_status"] = "Needs Attention"
return assessment
def generate_recommendations(self) -> List[str]:
"""Generate recommendations based on analysis."""
recommendations = []
# Check input readiness
readiness = self.check_input_readiness()
if readiness["readiness_score"] < 100:
recommendations.append(
f"Complete remaining review inputs: {', '.join(readiness['missing_topics'])}"
)
# Check actions
action_analysis = self.analyze_actions()
if action_analysis["overdue"]:
recommendations.append(
f"Address {len(action_analysis['overdue'])} overdue action(s) immediately"
)
# Check metrics
metrics_assessment = self.assess_metrics()
if metrics_assessment["overall_status"] == "Critical":
recommendations.append(
"Escalate critical metric failures to senior management"
)
# CAPA specific
if self.review.metrics.capa_overdue > 0:
recommendations.append(
f"Expedite closure of {self.review.metrics.capa_overdue} overdue CAPA(s)"
)
if self.review.metrics.capa_effectiveness < 85:
recommendations.append(
"Review root cause analysis quality for ineffective CAPAs"
)
# Audit findings
if self.review.metrics.audit_findings_major > 0:
recommendations.append(
f"Prioritize resolution of {self.review.metrics.audit_findings_major} major audit finding(s)"
)
if not recommendations:
recommendations.append("Quality system performing within targets. Maintain monitoring.")
return recommendations
def generate_report(self) -> Dict:
"""Generate complete review status report."""
return {
"review_date": self.review.review_date,
"review_type": self.review.review_type,
"period": f"{self.review.period_start} to {self.review.period_end}",
"input_readiness": self.check_input_readiness(),
"action_analysis": self.analyze_actions(),
"metrics_assessment": self.assess_metrics(),
"recommendations": self.generate_recommendations()
}
def format_text_report(report: Dict) -> str:
"""Format report as text output."""
lines = [
"=" * 70,
"MANAGEMENT REVIEW STATUS REPORT",
"=" * 70,
f"Review Date: {report['review_date']}",
f"Review Type: {report['review_type']}",
f"Period: {report['period']}",
"",
"INPUT READINESS",
"-" * 40,
f"Readiness Score: {report['input_readiness']['readiness_score']}%",
f"Complete: {report['input_readiness']['complete']} / {report['input_readiness']['total_required']}",
]
if report['input_readiness']['missing_topics']:
lines.append(f"Missing: {', '.join(report['input_readiness']['missing_topics'])}")
lines.extend([
"",
"ACTION STATUS",
"-" * 40,
f"Total Actions: {report['action_analysis']['total']}",
f"Completion Rate: {report['action_analysis']['completion_rate']}%",
])
for status, count in report['action_analysis']['by_status'].items():
lines.append(f" {status}: {count}")
if report['action_analysis']['overdue']:
lines.extend([
"",
"OVERDUE ACTIONS:",
])
for item in report['action_analysis']['overdue']:
lines.append(f" [{item['action_id']}] {item['description']} - {item['days_overdue']} days overdue")
lines.extend([
"",
"METRICS ASSESSMENT",
"-" * 40,
f"Overall Status: {report['metrics_assessment']['overall_status']}",
"",
f"{'Metric':<25} {'Value':<10} {'Target':<10} {'Status':<10}",
"-" * 55,
])
for metric in report['metrics_assessment']['metrics']:
lines.append(
f"{metric['name']:<25} {metric['value']:<10} {metric['target']:<10} {metric['status']:<10}"
)
if report['metrics_assessment']['alerts']:
lines.extend([
"",
"ALERTS:",
])
for alert in report['metrics_assessment']['alerts']:
lines.append(f" ! {alert}")
lines.extend([
"",
"RECOMMENDATIONS",
"-" * 40,
])
for i, rec in enumerate(report['recommendations'], 1):
lines.append(f"{i}. {rec}")
lines.append("=" * 70)
return "\n".join(lines)
def interactive_mode():
"""Run interactive review data entry."""
print("=" * 60)
print("Management Review Tracker - Interactive Mode")
print("=" * 60)
review_date = input("\nReview Date (YYYY-MM-DD): ").strip()
review_type = input("Review Type (Annual/Semi-annual/Quarterly): ").strip()
period_start = input("Period Start (YYYY-MM-DD): ").strip()
period_end = input("Period End (YYYY-MM-DD): ").strip()
print("\nEnter Quality Metrics:")
metrics = ReviewMetrics(
complaint_rate=float(input("Complaint Rate (%): ") or 0),
complaint_count=int(input("Complaint Count: ") or 0),
capa_open=int(input("Open CAPAs: ") or 0),
capa_overdue=int(input("Overdue CAPAs: ") or 0),
capa_effectiveness=float(input("CAPA Effectiveness (%): ") or 0),
audit_findings_open=int(input("Open Audit Findings: ") or 0),
audit_findings_major=int(input("Major Audit Findings: ") or 0),
first_pass_yield=float(input("First Pass Yield (%): ") or 0),
customer_satisfaction=float(input("Customer Satisfaction (1-5): ") or 0),
training_compliance=float(input("Training Compliance (%): ") or 0)
)
# Create review with sample inputs
inputs = [
ReviewInput(topic=topic, responsible=resp, status=InputStatus.COMPLETE, data_period=f"{period_start} to {period_end}")
for topic, resp in ManagementReviewTracker.REQUIRED_INPUTS
]
review = ManagementReview(
review_date=review_date,
review_type=review_type,
period_start=period_start,
period_end=period_end,
inputs=inputs,
actions=[],
metrics=metrics
)
tracker = ManagementReviewTracker(review)
report = tracker.generate_report()
print("\n" + format_text_report(report))
def main():
parser = argparse.ArgumentParser(
description="Management Review Tracker"
)
parser.add_argument(
"--data",
type=str,
help="JSON file with review data"
)
parser.add_argument(
"--output",
choices=["text", "json"],
default="text",
help="Output format"
)
parser.add_argument(
"--interactive",
action="store_true",
help="Run in interactive mode"
)
parser.add_argument(
"--sample",
action="store_true",
help="Generate sample review data"
)
args = parser.parse_args()
if args.interactive:
interactive_mode()
return
if args.sample:
sample = {
"review_date": "2024-06-30",
"review_type": "Semi-annual",
"period_start": "2024-01-01",
"period_end": "2024-06-30",
"inputs": [
{"topic": "Audit Results", "responsible": "QA Manager", "status": "Complete", "data_period": "H1 2024"},
{"topic": "Customer Feedback", "responsible": "Customer Quality", "status": "Complete", "data_period": "H1 2024"},
{"topic": "Process Performance", "responsible": "Operations", "status": "In Progress", "data_period": "H1 2024"},
{"topic": "CAPA Status", "responsible": "CAPA Officer", "status": "Complete", "data_period": "Current"}
],
"actions": [
{
"action_id": "MR-2024-001",
"description": "Implement enhanced CAPA tracking system",
"owner": "QA Manager",
"due_date": "2024-09-30",
"priority": "High",
"status": "In Progress",
"source_review": "2024-Q1"
}
],
"metrics": {
"complaint_rate": 0.08,
"complaint_count": 12,
"capa_open": 8,
"capa_overdue": 2,
"capa_effectiveness": 88.0,
"audit_findings_open": 5,
"audit_findings_major": 1,
"first_pass_yield": 96.5,
"customer_satisfaction": 4.2,
"training_compliance": 97.0
}
}
print(json.dumps(sample, indent=2))
return
# Create sample review if no data provided
if args.data:
with open(args.data, "r") as f:
data = json.load(f)
inputs = [
ReviewInput(
topic=inp["topic"],
responsible=inp["responsible"],
status=InputStatus[inp["status"].upper().replace(" ", "_")],
data_period=inp.get("data_period", "")
)
for inp in data.get("inputs", [])
]
actions = [
ActionItem(
action_id=act["action_id"],
description=act["description"],
owner=act["owner"],
due_date=act["due_date"],
priority=ActionPriority[act["priority"].upper()],
status=ActionStatus[act["status"].upper().replace(" ", "_")],
source_review=act.get("source_review", "")
)
for act in data.get("actions", [])
]
metrics_data = data.get("metrics", {})
metrics = ReviewMetrics(**metrics_data)
review = ManagementReview(
review_date=data["review_date"],
review_type=data["review_type"],
period_start=data["period_start"],
period_end=data["period_end"],
inputs=inputs,
actions=actions,
metrics=metrics
)
else:
# Demo data
review = ManagementReview(
review_date="2024-06-30",
review_type="Semi-annual",
period_start="2024-01-01",
period_end="2024-06-30",
inputs=[
ReviewInput("Audit Results", "QA Manager", InputStatus.COMPLETE, "H1 2024"),
ReviewInput("Customer Feedback", "Customer Quality", InputStatus.COMPLETE, "H1 2024"),
ReviewInput("CAPA Status", "CAPA Officer", InputStatus.COMPLETE, "Current"),
],
actions=[
ActionItem("MR-2024-001", "Implement CAPA tracking", "QA Mgr", "2024-09-30",
ActionPriority.HIGH, ActionStatus.IN_PROGRESS, "2024-Q1"),
],
metrics=ReviewMetrics(
complaint_rate=0.08, capa_open=8, capa_overdue=2,
capa_effectiveness=88.0, first_pass_yield=96.5,
customer_satisfaction=4.2, training_compliance=97.0
)
)
tracker = ManagementReviewTracker(review)
report = tracker.generate_report()
if args.output == "json":
print(json.dumps(report, indent=2))
else:
print(format_text_report(report))
if __name__ == "__main__":
main()
FILE:scripts/quality_effectiveness_monitor.py
#!/usr/bin/env python3
"""
Quality Management System Effectiveness Monitor
Quantitatively assess QMS effectiveness using leading and lagging indicators.
Tracks trends, calculates control limits, and predicts potential quality issues
before they become failures. Integrates with CAPA and management review processes.
Supports metrics:
- Complaint rates, defect rates, rework rates
- Supplier performance
- CAPA effectiveness
- Audit findings trends
- Non-conformance statistics
Usage:
python quality_effectiveness_monitor.py --metrics metrics.csv --dashboard
python quality_effectiveness_monitor.py --qms-data qms_data.json --predict
python quality_effectiveness_monitor.py --interactive
"""
import argparse
import json
import csv
import sys
from dataclasses import dataclass, field, asdict
from typing import List, Dict, Optional, Tuple
from datetime import datetime, timedelta
from statistics import mean, stdev, median
@dataclass
class QualityMetric:
"""A single quality metric data point."""
metric_id: str
metric_name: str
category: str
date: str
value: float
unit: str
target: float
upper_limit: float
lower_limit: float
trend_direction: str = "" # "up", "down", "stable"
sigma_level: float = 0.0
is_alert: bool = False
is_critical: bool = False
@dataclass
class QMSReport:
"""QMS effectiveness report."""
report_period: Tuple[str, str]
overall_effectiveness_score: float
metrics_count: int
metrics_in_control: int
metrics_out_of_control: int
critical_alerts: int
trends_analysis: Dict
predictive_alerts: List[Dict]
improvement_opportunities: List[Dict]
management_review_summary: str
class QMSEffectivenessMonitor:
"""Monitors and analyzes QMS effectiveness."""
SIGNAL_INDICATORS = {
"complaint_rate": {"unit": "per 1000 units", "target": 0, "upper_limit": 1.5},
"defect_rate": {"unit": "PPM", "target": 100, "upper_limit": 500},
"rework_rate": {"unit": "%", "target": 2.0, "upper_limit": 5.0},
"on_time_delivery": {"unit": "%", "target": 98, "lower_limit": 95},
"audit_findings": {"unit": "count/month", "target": 0, "upper_limit": 3},
"capa_closure_rate": {"unit": "% within target", "target": 100, "lower_limit": 90},
"supplier_defect_rate": {"unit": "PPM", "target": 200, "upper_limit": 1000}
}
def __init__(self):
self.metrics = []
def load_csv(self, csv_path: str) -> List[QualityMetric]:
"""Load metrics from CSV file."""
metrics = []
with open(csv_path, 'r', encoding='utf-8') as f:
reader = csv.DictReader(f)
for row in reader:
metric = QualityMetric(
metric_id=row.get('metric_id', ''),
metric_name=row.get('metric_name', ''),
category=row.get('category', 'General'),
date=row.get('date', ''),
value=float(row.get('value', 0)),
unit=row.get('unit', ''),
target=float(row.get('target', 0)),
upper_limit=float(row.get('upper_limit', 0)),
lower_limit=float(row.get('lower_limit', 0)),
)
metrics.append(metric)
self.metrics = metrics
return metrics
def calculate_sigma_level(self, metric: QualityMetric, historical_values: List[float]) -> float:
"""Calculate process sigma level based on defect rate."""
if metric.unit == "PPM" or "rate" in metric.metric_name.lower():
# For defect rates, DPMO = defects_per_million_opportunities
if historical_values:
avg_defect_rate = mean(historical_values)
if avg_defect_rate > 0:
dpmo = avg_defect_rate
# Simplified sigma conversion (actual uses 1.5σ shift)
sigma_map = {
330000: 1.0, 620000: 2.0, 110000: 3.0, 27000: 4.0,
6200: 5.0, 230: 6.0, 3.4: 6.0
}
# Rough sigma calculation
sigma = 6.0 - (dpmo / 1000000) * 10
return max(0.0, min(6.0, sigma))
return 0.0
def analyze_trend(self, values: List[float]) -> Tuple[str, float]:
"""Analyze trend direction and significance."""
if len(values) < 3:
return "insufficient_data", 0.0
x = list(range(len(values)))
y = values
# Linear regression
n = len(x)
sum_x = sum(x)
sum_y = sum(y)
sum_xy = sum(x[i] * y[i] for i in range(n))
sum_x2 = sum(xi * xi for xi in x)
slope = (n * sum_xy - sum_x * sum_y) / (n * sum_x2 - sum_x * sum_x) if (n * sum_x2 - sum_x * sum_x) != 0 else 0
# Determine trend direction
if slope > 0.01:
direction = "up"
elif slope < -0.01:
direction = "down"
else:
direction = "stable"
# Calculate R-squared
if slope != 0:
intercept = (sum_y - slope * sum_x) / n
y_pred = [slope * xi + intercept for xi in x]
ss_res = sum((y[i] - y_pred[i])**2 for i in range(n))
ss_tot = sum((y[i] - mean(y))**2 for i in range(n))
r2 = 1 - (ss_res / ss_tot) if ss_tot > 0 else 0
else:
r2 = 0
return direction, r2
def detect_alerts(self, metrics: List[QualityMetric]) -> List[Dict]:
"""Detect metrics that require attention."""
alerts = []
for metric in metrics:
# Check immediate control limit violation
if metric.upper_limit and metric.value > metric.upper_limit:
alerts.append({
"metric_id": metric.metric_id,
"metric_name": metric.metric_name,
"issue": "exceeds_upper_limit",
"value": metric.value,
"limit": metric.upper_limit,
"severity": "critical" if metric.category in ["Customer", "Regulatory"] else "high"
})
if metric.lower_limit and metric.value < metric.lower_limit:
alerts.append({
"metric_id": metric.metric_id,
"metric_name": metric.metric_name,
"issue": "below_lower_limit",
"value": metric.value,
"limit": metric.lower_limit,
"severity": "critical" if metric.category in ["Customer", "Regulatory"] else "high"
})
# Check for adverse trend (3+ points in same direction)
# Need to group by metric_name and check historical data
# Simplified: check trend_direction flag if set
if metric.trend_direction in ["up", "down"] and metric.sigma_level > 3:
alerts.append({
"metric_id": metric.metric_id,
"metric_name": metric.metric_name,
"issue": f"adverse_trend_{metric.trend_direction}",
"value": metric.value,
"severity": "medium"
})
return alerts
def predict_failures(self, metrics: List[QualityMetric], forecast_days: int = 30) -> List[Dict]:
"""Predict potential failures based on trends."""
predictions = []
# Group metrics by name to get time series
grouped = {}
for m in metrics:
if m.metric_name not in grouped:
grouped[m.metric_name] = []
grouped[m.metric_name].append(m)
for metric_name, metric_list in grouped.items():
if len(metric_list) < 5:
continue
# Sort by date
metric_list.sort(key=lambda m: m.date)
values = [m.value for m in metric_list]
# Simple linear extrapolation
x = list(range(len(values)))
y = values
n = len(x)
sum_x = sum(x)
sum_y = sum(y)
sum_xy = sum(x[i] * y[i] for i in range(n))
sum_x2 = sum(xi * xi for xi in x)
slope = (n * sum_xy - sum_x * sum_y) / (n * sum_x2 - sum_x * sum_x) if (n * sum_x2 - sum_x * sum_x) != 0 else 0
if slope != 0:
# Forecast next value
next_value = y[-1] + slope
target = metric_list[0].target
upper_limit = metric_list[0].upper_limit
if (target and next_value > target * 1.2) or (upper_limit and next_value > upper_limit * 0.9):
predictions.append({
"metric": metric_name,
"current_value": y[-1],
"forecast_value": round(next_value, 2),
"forecast_days": forecast_days,
"trend_slope": round(slope, 3),
"risk_level": "high" if upper_limit and next_value > upper_limit else "medium"
})
return predictions
def calculate_effectiveness_score(self, metrics: List[QualityMetric]) -> float:
"""Calculate overall QMS effectiveness score (0-100)."""
if not metrics:
return 0.0
scores = []
for m in metrics:
# Score based on distance to target
if m.target != 0:
deviation = abs(m.value - m.target) / max(abs(m.target), 1)
score = max(0, 100 - deviation * 100)
else:
# For metrics where lower is better (defects, etc.)
if m.upper_limit:
score = max(0, 100 - (m.value / m.upper_limit) * 100 * 0.5)
else:
score = 50 # Neutral if no target
scores.append(score)
# Penalize for alerts
alerts = self.detect_alerts(metrics)
penalty = len([a for a in alerts if a["severity"] in ["critical", "high"]]) * 5
return max(0, min(100, mean(scores) - penalty))
def identify_improvement_opportunities(self, metrics: List[QualityMetric]) -> List[Dict]:
"""Identify metrics with highest improvement potential."""
opportunities = []
for m in metrics:
if m.upper_limit and m.value > m.upper_limit * 0.8:
gap = m.upper_limit - m.value
if gap > 0:
improvement_pct = (gap / m.upper_limit) * 100
opportunities.append({
"metric": m.metric_name,
"current": m.value,
"target": m.upper_limit,
"gap": round(gap, 2),
"improvement_potential_pct": round(improvement_pct, 1),
"recommended_action": f"Reduce {m.metric_name} by at least {round(gap, 2)} {m.unit}",
"impact": "High" if m.category in ["Customer", "Regulatory"] else "Medium"
})
# Sort by improvement potential
opportunities.sort(key=lambda x: x["improvement_potential_pct"], reverse=True)
return opportunities[:10]
def generate_management_review_summary(self, report: QMSReport) -> str:
"""Generate executive summary for management review."""
summary = [
f"QMS EFFECTIVENESS REVIEW - {report.report_period[0]} to {report.report_period[1]}",
"",
f"Overall Effectiveness Score: {report.overall_effectiveness_score:.1f}/100",
f"Metrics Tracked: {report.metrics_count} | In Control: {report.metrics_in_control} | Alerts: {report.critical_alerts}",
""
]
if report.critical_alerts > 0:
summary.append("🔴 CRITICAL ALERTS REQUIRING IMMEDIATE ATTENTION:")
for alert in [a for a in report.predictive_alerts if a.get("risk_level") == "high"]:
summary.append(f" • {alert['metric']}: forecast {alert['forecast_value']} (from {alert['current_value']})")
summary.append("")
summary.append("📈 TOP IMPROVEMENT OPPORTUNITIES:")
for i, opp in enumerate(report.improvement_opportunities[:3], 1):
summary.append(f" {i}. {opp['metric']}: {opp['recommended_action']} (Impact: {opp['impact']})")
summary.append("")
summary.append("🎯 RECOMMENDED ACTIONS:")
summary.append(" 1. Address all high-severity alerts within 30 days")
summary.append(" 2. Launch improvement projects for top 3 opportunities")
summary.append(" 3. Review CAPA effectiveness for recurring issues")
summary.append(" 4. Update risk assessments based on predictive trends")
return "\n".join(summary)
def analyze(
self,
metrics: List[QualityMetric],
start_date: str = None,
end_date: str = None
) -> QMSReport:
"""Perform comprehensive QMS effectiveness analysis."""
in_control = 0
for m in metrics:
if not m.is_alert and not m.is_critical:
in_control += 1
out_of_control = len(metrics) - in_control
alerts = self.detect_alerts(metrics)
critical_alerts = len([a for a in alerts if a["severity"] in ["critical", "high"]])
predictions = self.predict_failures(metrics)
improvement_opps = self.identify_improvement_opportunities(metrics)
effectiveness = self.calculate_effectiveness_score(metrics)
# Trend analysis by category
trends = {}
categories = set(m.category for m in metrics)
for cat in categories:
cat_metrics = [m for m in metrics if m.category == cat]
if len(cat_metrics) >= 2:
avg_values = [mean([m.value for m in cat_metrics])] # Simplistic - would need time series
trends[cat] = {
"metric_count": len(cat_metrics),
"avg_value": round(mean([m.value for m in cat_metrics]), 2),
"alerts": len([a for a in alerts if any(m.metric_name == a["metric_name"] for m in cat_metrics)])
}
period = (start_date or metrics[0].date, end_date or metrics[-1].date) if metrics else ("", "")
report = QMSReport(
report_period=period,
overall_effectiveness_score=effectiveness,
metrics_count=len(metrics),
metrics_in_control=in_control,
metrics_out_of_control=out_of_control,
critical_alerts=critical_alerts,
trends_analysis=trends,
predictive_alerts=predictions,
improvement_opportunities=improvement_opps,
management_review_summary="" # Filled later
)
report.management_review_summary = self.generate_management_review_summary(report)
return report
def format_qms_report(report: QMSReport) -> str:
"""Format QMS report as text."""
lines = [
"=" * 80,
"QMS EFFECTIVENESS MONITORING REPORT",
"=" * 80,
f"Period: {report.report_period[0]} to {report.report_period[1]}",
f"Overall Score: {report.overall_effectiveness_score:.1f}/100",
"",
"METRIC STATUS",
"-" * 40,
f" Total Metrics: {report.metrics_count}",
f" In Control: {report.metrics_in_control}",
f" Out of Control: {report.metrics_out_of_control}",
f" Critical Alerts: {report.critical_alerts}",
"",
"TREND ANALYSIS BY CATEGORY",
"-" * 40,
]
for category, data in report.trends_analysis.items():
lines.append(f" {category}: {data['avg_value']} (alerts: {data['alerts']})")
if report.predictive_alerts:
lines.extend([
"",
"PREDICTIVE ALERTS (Next 30 days)",
"-" * 40,
])
for alert in report.predictive_alerts[:5]:
lines.append(f" ⚠ {alert['metric']}: {alert['current_value']} → {alert['forecast_value']} ({alert['risk_level']})")
if report.improvement_opportunities:
lines.extend([
"",
"TOP IMPROVEMENT OPPORTUNITIES",
"-" * 40,
])
for i, opp in enumerate(report.improvement_opportunities[:5], 1):
lines.append(f" {i}. {opp['metric']}: {opp['recommended_action']}")
lines.extend([
"",
"MANAGEMENT REVIEW SUMMARY",
"-" * 40,
report.management_review_summary,
"=" * 80
])
return "\n".join(lines)
def main():
parser = argparse.ArgumentParser(description="QMS Effectiveness Monitor")
parser.add_argument("--metrics", type=str, help="CSV file with quality metrics")
parser.add_argument("--qms-data", type=str, help="JSON file with QMS data")
parser.add_argument("--dashboard", action="store_true", help="Generate dashboard summary")
parser.add_argument("--predict", action="store_true", help="Include predictive analytics")
parser.add_argument("--output", choices=["text", "json"], default="text")
parser.add_argument("--interactive", action="store_true", help="Interactive mode")
args = parser.parse_args()
monitor = QMSEffectivenessMonitor()
if args.metrics:
metrics = monitor.load_csv(args.metrics)
report = monitor.analyze(metrics)
elif args.qms_data:
with open(args.qms_data) as f:
data = json.load(f)
# Convert to QualityMetric objects
metrics = [QualityMetric(**m) for m in data.get("metrics", [])]
report = monitor.analyze(metrics)
else:
# Demo data
demo_metrics = [
QualityMetric("M001", "Customer Complaint Rate", "Customer", "2026-03-01", 0.8, "per 1000", 1.0, 1.5, 0.5),
QualityMetric("M002", "Defect Rate PPM", "Quality", "2026-03-01", 125, "PPM", 100, 500, 0, trend_direction="down", sigma_level=4.2),
QualityMetric("M003", "On-Time Delivery", "Operations", "2026-03-01", 96.5, "%", 98, 0, 95, trend_direction="down"),
QualityMetric("M004", "CAPA Closure Rate", "Quality", "2026-03-01", 92.0, "%", 100, 0, 90, is_alert=True),
QualityMetric("M005", "Supplier Defect Rate", "Supplier", "2026-03-01", 450, "PPM", 200, 1000, 0, is_critical=True),
]
# Simulate time series
all_metrics = []
for i in range(30):
for dm in demo_metrics:
new_metric = QualityMetric(
metric_id=dm.metric_id,
metric_name=dm.metric_name,
category=dm.category,
date=f"2026-03-{i+1:02d}",
value=dm.value + (i * 0.1) if dm.metric_name == "Customer Complaint Rate" else dm.value,
unit=dm.unit,
target=dm.target,
upper_limit=dm.upper_limit,
lower_limit=dm.lower_limit
)
all_metrics.append(new_metric)
report = monitor.analyze(all_metrics)
if args.output == "json":
result = asdict(report)
print(json.dumps(result, indent=2))
else:
print(format_qms_report(report))
if __name__ == "__main__":
main()
Bộ 12 skill quy định và quản lý chất lượng: ISO 13485, MDR, FDA 510(k)/PMA, ISO 27001, GDPR, quản lý rủi ro ISO 14971, CAPA, kiểm soát tài liệu.
--- name: "ra-qm-skills" description: "12 regulatory & QM agent skills and plugins for Claude Code, Codex, Gemini CLI, Cursor, OpenClaw. ISO 13485 QMS, MDR 2017/745, FDA 510(k)/PMA, ISO 27001 ISMS, GDPR/DSGVO, risk management (ISO 14971), CAPA, document control, auditing. Python tools (stdlib-only)." version: 2.9.0 author: Alireza Rezvani license: MIT tags: - regulatory - quality-management - iso-13485 - mdr - fda - iso-27001 - gdpr agents: - claude-code - codex-cli - openclaw --- # Regulatory Affairs & Quality Management Skills 12 production-ready compliance skills for HealthTech and MedTech organizations. ## Quick Start ### Claude Code ``` /read ra-qm-team/regulatory-affairs-head/SKILL.md ``` ### Codex CLI ```bash npx agent-skills-cli add alirezarezvani/claude-skills/ra-qm-team ``` ## Skills Overview | Skill | Folder | Focus | |-------|--------|-------| | Regulatory Affairs Head | `regulatory-affairs-head/` | FDA/MDR strategy, submissions | | Quality Manager (QMR) | `quality-manager-qmr/` | QMS governance, management review | | Quality Manager (ISO 13485) | `quality-manager-qms-iso13485/` | QMS implementation, doc control | | Risk Management Specialist | `risk-management-specialist/` | ISO 14971, FMEA, risk files | | CAPA Officer | `capa-officer/` | Root cause analysis, corrective actions | | Quality Documentation Manager | `quality-documentation-manager/` | Document control, 21 CFR Part 11 | | QMS Audit Expert | `qms-audit-expert/` | ISO 13485 internal audits | | ISMS Audit Expert | `isms-audit-expert/` | ISO 27001 security audits | | Information Security Manager | `information-security-manager-iso27001/` | ISMS implementation | | MDR 745 Specialist | `mdr-745-specialist/` | EU MDR classification, CE marking | | FDA Consultant | `fda-consultant-specialist/` | 510(k), PMA, QSR compliance | | GDPR/DSGVO Expert | `gdpr-dsgvo-expert/` | Privacy compliance, DPIA | ## Python Tools 17 scripts, all stdlib-only: ```bash python3 risk-management-specialist/scripts/risk_matrix_calculator.py --help python3 gdpr-dsgvo-expert/scripts/gdpr_compliance_checker.py --help ``` ## Rules - Load only the specific skill SKILL.md you need - Always verify compliance outputs against current regulations