Quét, sửa và xác minh tuân thủ WCAG 2.2 mức A và AA cho React, Next.js, Vue, Angular, Svelte và HTML thuần.
---
name: "a11y-audit"
description: "Accessibility audit skill for scanning, fixing, and verifying WCAG 2.2 Level A and AA compliance across React, Next.js, Vue, Angular, Svelte, and plain HTML codebases. Use when auditing accessibility, fixing a11y violations, checking color contrast, generating compliance reports, or integrating accessibility checks into CI/CD pipelines."
---
# Accessibility Audit
WCAG 2.2 Accessibility Audit and Remediation Skill
## Description
The a11y-audit skill provides a complete accessibility audit pipeline for modern web applications. It implements a three-phase workflow -- Scan, Fix, Verify -- that identifies WCAG 2.2 Level A and AA violations, generates exact fix code per framework, and produces stakeholder-ready compliance reports.
For every violation it finds, it provides the precise before/after code fix tailored to your framework (React, Next.js, Vue, Angular, Svelte, or plain HTML).
**What this skill does:**
1. **Scans** your codebase for every WCAG 2.2 Level A and AA violation, categorized by severity (Critical, Major, Minor)
2. **Fixes** each violation with framework-specific before/after code patterns
3. **Verifies** that fixes resolve the original violations and introduces no regressions
4. **Reports** findings in a structured format suitable for developers, PMs, and compliance stakeholders
5. **Integrates** into CI/CD pipelines to prevent accessibility regressions
## Features
| Feature | Description |
|---------|-------------|
| **Full WCAG 2.2 Scan** | Checks all Level A and AA success criteria across your codebase |
| **Framework Detection** | Auto-detects React, Next.js, Vue, Angular, Svelte, or plain HTML |
| **Severity Classification** | Categorizes each violation as Critical, Major, or Minor |
| **Fix Code Generation** | Produces before/after code diffs for every issue |
| **Color Contrast Checker** | Validates foreground/background pairs against AA and AAA ratios |
| **Compliance Reporting** | Generates stakeholder reports with pass/fail summaries |
| **CI/CD Integration** | GitHub Actions, GitLab CI, Azure DevOps pipeline configs |
| **Keyboard Navigation Audit** | Detects missing focus management and tab order issues |
| **ARIA Validation** | Checks for incorrect, redundant, or missing ARIA attributes |
### Severity Definitions
| Severity | Definition | Example | SLA |
|----------|-----------|---------|-----|
| **Critical** | Blocks access for entire user groups | Missing alt text, no keyboard access to navigation | Fix before release |
| **Major** | Significant barrier that degrades experience | Insufficient color contrast, missing form labels | Fix within current sprint |
| **Minor** | Usability issue that causes friction | Redundant ARIA roles, suboptimal heading hierarchy | Fix within next 2 sprints |
## Usage
### Quick Start
```bash
# Scan entire project
python scripts/a11y_scanner.py /path/to/project
# Scan with JSON output for tooling
python scripts/a11y_scanner.py /path/to/project --json
# Check color contrast for specific values
python scripts/contrast_checker.py --fg "#777777" --bg "#ffffff"
# Check contrast across a CSS/Tailwind file
python scripts/contrast_checker.py --file /path/to/styles.css
```
### Slash Command
```
/a11y-audit # Audit current project
/a11y-audit --scope src/ # Audit specific directory
/a11y-audit --fix # Audit and auto-apply fixes
/a11y-audit --report # Generate stakeholder report
/a11y-audit --ci # Output CI-compatible results
```
### Three-Phase Workflow
**Phase 1: Scan** -- Walk the source tree, detect framework, apply rule set.
```bash
python scripts/a11y_scanner.py /path/to/project --format table
```
**Phase 2: Fix** -- Apply framework-specific fixes for each violation.
> See [references/framework-a11y-patterns.md](references/framework-a11y-patterns.md) for the complete fix patterns catalog.
**Phase 3: Verify** -- Re-run the scanner to confirm fixes and check for regressions.
```bash
python scripts/a11y_scanner.py /path/to/project --baseline audit-baseline.json
```
## Example: React Component Audit
```tsx
// BEFORE: src/components/ProductCard.tsx
function ProductCard({ product }) {
return (
<div onClick={() => navigate(`/product/product.id`)}>
<img src={product.image} />
<div style={{ color: '#aaa', fontSize: '12px' }}>{product.name}</div>
<span style={{ color: '#999' }}>product.price</span>
</div>
);
}
```
| # | WCAG | Severity | Issue |
|---|------|----------|-------|
| 1 | 1.1.1 | Critical | `<img>` missing `alt` attribute |
| 2 | 2.1.1 | Critical | `<div onClick>` not keyboard accessible |
| 3 | 1.4.3 | Major | Color `#aaa` on white fails contrast (2.32:1, needs 4.5:1) |
| 4 | 1.4.3 | Major | Color `#999` on white fails contrast (2.85:1, needs 4.5:1) |
| 5 | 4.1.2 | Major | Interactive element missing role and accessible name |
```tsx
// AFTER: src/components/ProductCard.tsx
function ProductCard({ product }) {
return (
<a href={`/product/product.id`} className="product-card"
aria-label={`View product.name - $product.price`}>
<img src={product.image} alt={product.imageAlt || product.name} />
<div style={{ color: '#595959', fontSize: '12px' }}>{product.name}</div>
<span style={{ color: '#767676' }}>product.price</span>
</a>
);
}
```
> See [references/examples-by-framework.md](references/examples-by-framework.md) for Vue, Angular, Next.js, and Svelte examples.
## Tools Reference
### a11y_scanner.py
```
Usage: python scripts/a11y_scanner.py <path> [options]
Options:
--json Output results as JSON
--format {table,csv} Output format (default: table)
--severity {critical,major,minor} Filter by minimum severity
--framework {react,vue,angular,svelte,html,auto} Force framework (default: auto)
--baseline FILE Compare against previous scan results
--report Generate stakeholder report
--output FILE Write results to file
--quiet Suppress output, exit code only
--ci CI mode: non-zero exit on critical issues
```
### contrast_checker.py
```
Usage: python scripts/contrast_checker.py [options]
Options:
--fg COLOR Foreground color (hex)
--bg COLOR Background color (hex)
--file FILE Scan CSS file for color pairs
--tailwind DIR Scan directory for Tailwind color classes
--json Output results as JSON
--suggest Suggest accessible alternatives for failures
--level {aa,aaa} Target conformance level (default: aa)
```
## Common Pitfalls
| Pitfall | Correct Approach |
|---------|------------------|
| `role="button"` on a `<div>` | Use native `<button>` -- includes keyboard handling for free |
| `tabindex="0"` on everything | Only interactive elements need focus; use native elements |
| `aria-label` on non-interactive elements | Use `aria-labelledby` pointing to visible text |
| `display: none` for screen reader hiding | Use `.sr-only` class instead |
| Color alone to convey meaning | Add icons, text labels, or patterns alongside color |
| Placeholder as only label | Always provide a visible `<label>` |
| `outline: none` without replacement | Always provide a visible focus indicator via `focus-visible` |
| Empty `alt=""` on informational images | Informational images need descriptive alt text |
| Skipping heading levels (h1 -> h3) | Heading levels must be sequential |
| `onClick` without `onKeyDown` | Add keyboard support or prefer native elements |
| Ignoring `prefers-reduced-motion` | Wrap animations in `@media (prefers-reduced-motion: no-preference)` |
## Related Skills
| Skill | Relationship |
|-------|-------------|
| **senior-frontend** | Frontend patterns used in a11y fixes |
| **code-reviewer** | Include a11y checks in code review workflows |
| **senior-qa** | Integration of a11y testing into QA processes |
| **playwright-pro** | Automated browser testing with accessibility assertions |
| **epic-design** | WCAG 2.1 AA compliant animations and scroll storytelling |
| **tdd-guide** | Test-driven development patterns for a11y test cases |
## Reference Documentation
| Reference | Description |
|-----------|-------------|
| [wcag-quick-ref.md](references/wcag-quick-ref.md) | WCAG 2.2 Level A & AA criteria quick reference |
| [wcag-22-new-criteria.md](references/wcag-22-new-criteria.md) | New WCAG 2.2 success criteria (Focus Appearance, Target Size, etc.) |
| [aria-patterns.md](references/aria-patterns.md) | ARIA patterns, keyboard interaction, and live regions |
| [framework-a11y-patterns.md](references/framework-a11y-patterns.md) | Framework-specific fix patterns (React, Vue, Angular, Svelte, HTML) |
| [color-contrast-guide.md](references/color-contrast-guide.md) | Color contrast checker details, Tailwind palette mapping, sr-only class |
| [ci-cd-integration.md](references/ci-cd-integration.md) | GitHub Actions, GitLab CI, Azure DevOps, pre-commit hook configs |
| [audit-report-template.md](references/audit-report-template.md) | Stakeholder-ready audit report template |
| [testing-checklist.md](references/testing-checklist.md) | Manual testing checklist (keyboard, screen reader, visual, forms) |
| [examples-by-framework.md](references/examples-by-framework.md) | Full audit examples for Vue, Angular, Next.js, and Svelte |
## Resources
- [WCAG 2.2 Specification](https://www.w3.org/TR/WCAG22/)
- [WAI-ARIA Authoring Practices 1.2](https://www.w3.org/WAI/ARIA/apg/)
- [Deque axe-core Rules](https://github.com/dequelabs/axe-core/blob/develop/doc/rule-descriptions.md)
- [eslint-plugin-jsx-a11y](https://github.com/jsx-eslint/eslint-plugin-jsx-a11y)
FILE:assets/sample-component.tsx
// Sample React component with intentional a11y issues for testing
import React from 'react';
export function UserCard({ user, onEdit, onDelete }) {
return (
<div className="card" onClick={() => onEdit(user.id)}>
<img src={user.avatar} />
<div className="name">{user.name}</div>
<div className="email">{user.email}</div>
<div className="actions">
<div onClick={() => onDelete(user.id)} style={{ color: '#aaa', cursor: 'pointer' }}>
Delete
</div>
<a href="#">Edit</a>
</div>
<input placeholder="Add note" />
</div>
);
}
export function SearchBar() {
return (
<div>
<input type="text" placeholder="Search..." />
<div onClick={() => alert('searching')} tabIndex={5}>
🔍
</div>
</div>
);
}
export function DataTable({ rows }) {
return (
<table>
<tr>
<td><b>Name</b></td>
<td><b>Email</b></td>
<td><b>Status</b></td>
</tr>
{rows.map((row) => (
<tr key={row.id}>
<td>{row.name}</td>
<td>{row.email}</td>
<td style={{ color: row.active ? 'green' : 'red' }}>
{row.active ? '●' : '●'}
</td>
</tr>
))}
</table>
);
}
FILE:expected_outputs/sample-contrast-output.txt
Contrast Check: #777777 on #ffffff
Foreground: #777777 (r=119, g=119, b=119)
Background: #ffffff (r=255, g=255, b=255)
Contrast Ratio: 4.48:1
Normal text (4.5:1 required):
AA: FAIL (4.48 < 4.5)
AAA: FAIL (4.48 < 7.0)
Large text (3.0:1 required):
AA: PASS (4.48 >= 3.0)
AAA: FAIL (4.48 < 4.5)
UI components (3.0:1 required):
AA: PASS (4.48 >= 3.0)
Verdict: FAIL — does not meet AA for normal text
---
Contrast Check: #1a1a2e on #ffffff
Foreground: #1a1a2e (r=26, g=26, b=46)
Background: #ffffff (r=255, g=255, b=255)
Contrast Ratio: 17.06:1
Normal text (4.5:1 required):
AA: PASS
AAA: PASS
Large text (3.0:1 required):
AA: PASS
AAA: PASS
UI components (3.0:1 required):
AA: PASS
Verdict: PASS — meets AAA for all categories
FILE:expected_outputs/sample-scan-output.json
{
"summary": {
"files_scanned": 1,
"files_with_issues": 1,
"total_issues": 9,
"critical": 3,
"serious": 4,
"moderate": 2,
"minor": 0,
"verdict": "FAIL"
},
"findings": [
{
"severity": "critical",
"category": "IMG-ALT",
"file": "sample-component.tsx",
"line": 7,
"code": "<img src={user.avatar} />",
"wcag": "1.1.1",
"message": "Image missing alt attribute",
"fix": "Add alt text: alt=\"description of image\""
},
{
"severity": "critical",
"category": "KB-CLICK",
"file": "sample-component.tsx",
"line": 5,
"code": "<div className=\"card\" onClick={() => onEdit(user.id)}>",
"wcag": "2.1.1",
"message": "Click handler on non-interactive element without keyboard support",
"fix": "Use <button> or add role=\"button\", tabIndex={0}, onKeyDown"
}
]
}
FILE:expected_outputs/sample-scan-report.md
# A11y Audit Report — sample-component.tsx
**Scanned:** 1 file | **Issues:** 9 | **Status:** FAIL
## Critical (3)
### 1. Missing alt text on image
- **File:** sample-component.tsx:7
- **Code:** `<img src={user.avatar} />`
- **WCAG:** 1.1.1 Non-text Content (Level A)
- **Fix:** Add descriptive alt text: `<img src={user.avatar} alt={`user.name's avatar`} />`
### 2. Click handler without keyboard support
- **File:** sample-component.tsx:5
- **Code:** `<div className="card" onClick={() => onEdit(user.id)}>`
- **WCAG:** 2.1.1 Keyboard (Level A)
- **Fix:** Use `<button>` or add `role="button"`, `tabIndex={0}`, and `onKeyDown`
### 3. Click handler without keyboard support
- **File:** sample-component.tsx:11
- **Code:** `<div onClick={() => onDelete(user.id)} ...>`
- **WCAG:** 2.1.1 Keyboard (Level A)
- **Fix:** Replace `<div>` with `<button>`
## Serious (4)
### 4. Missing form label
- **File:** sample-component.tsx:15
- **Code:** `<input placeholder="Add note" />`
- **WCAG:** 3.3.2 Labels or Instructions (Level A)
- **Fix:** Add `<label>` or `aria-label="Add note"`
### 5. Empty link
- **File:** sample-component.tsx:14
- **Code:** `<a href="#">Edit</a>`
- **WCAG:** 2.4.4 Link Purpose (Level A)
- **Fix:** Use a real href or replace with `<button>`
### 6. tabindex greater than 0
- **File:** sample-component.tsx:24
- **Code:** `tabIndex={5}`
- **WCAG:** 2.4.3 Focus Order (Level A)
- **Fix:** Use `tabIndex={0}` — positive values disrupt natural tab order
### 7. Missing table headers
- **File:** sample-component.tsx:30
- **Code:** `<td><b>Name</b></td>` (using td+b instead of th)
- **WCAG:** 1.3.1 Info and Relationships (Level A)
- **Fix:** Use `<th scope="col">Name</th>`
## Moderate (2)
### 8. Missing form label
- **File:** sample-component.tsx:22
- **Code:** `<input type="text" placeholder="Search..." />`
- **WCAG:** 3.3.2 Labels or Instructions (Level A)
- **Fix:** Add `aria-label="Search"` or visible label
### 9. Color as sole indicator
- **File:** sample-component.tsx:38
- **Code:** `style={{ color: row.active ? 'green' : 'red' }}`
- **WCAG:** 1.4.1 Use of Color (Level A)
- **Fix:** Add text or icon alongside color: `{row.active ? '✓ Active' : '✗ Inactive'}`
FILE:references/aria-patterns.md
# ARIA Patterns & Keyboard Interaction Reference
## Landmark Roles
Every page should have these landmarks:
```html
<header role="banner"> <!-- Site header — once per page -->
<nav role="navigation"> <!-- Navigation — can have multiple with aria-label -->
<main role="main"> <!-- Main content — once per page -->
<aside role="complementary"> <!-- Sidebar — related but not essential -->
<footer role="contentinfo"> <!-- Site footer — once per page -->
<form role="search"> <!-- Search form -->
```
**Semantic HTML equivalents:** `<header>`, `<nav>`, `<main>`, `<aside>`, `<footer>` provide implicit roles — no need to double up with explicit `role` attributes.
## Live Regions
### When to Use
| Pattern | Attribute | Use Case |
|---------|-----------|----------|
| Polite | `aria-live="polite"` | Toast notifications, status updates, search result counts |
| Assertive | `aria-live="assertive"` | Error messages, urgent alerts, form validation errors |
| Status | `role="status"` | Loading indicators, progress updates |
| Alert | `role="alert"` | Error dialogs, time-sensitive warnings |
| Log | `role="log"` | Chat messages, activity feeds |
| Timer | `role="timer"` | Countdown timers |
### Implementation
```html
<!-- Toast notifications -->
<div aria-live="polite" aria-atomic="true">
<!-- Inject toast content here dynamically -->
</div>
<!-- Form validation errors -->
<div aria-live="assertive" role="alert">
<p>Please enter a valid email address.</p>
</div>
<!-- Loading state -->
<div role="status" aria-live="polite">
Loading results...
</div>
```
**Key rule:** The live region container must exist in the DOM *before* content is injected. Adding `aria-live` to a newly created element won't announce it.
## Focus Management
### Focus Trap (Modals)
```javascript
// Trap focus inside modal
const modal = document.querySelector('[role="dialog"]');
const focusable = modal.querySelectorAll(
'a[href], button, textarea, input, select, [tabindex]:not([tabindex="-1"])'
);
const first = focusable[0];
const last = focusable[focusable.length - 1];
modal.addEventListener('keydown', (e) => {
if (e.key === 'Tab') {
if (e.shiftKey && document.activeElement === first) {
e.preventDefault();
last.focus();
} else if (!e.shiftKey && document.activeElement === last) {
e.preventDefault();
first.focus();
}
}
if (e.key === 'Escape') closeModal();
});
```
### Focus Restoration
```javascript
// Save focus before opening modal
const trigger = document.activeElement;
openModal();
// Restore focus on close
function closeModal() {
modal.hidden = true;
trigger.focus();
}
```
### Skip Link
```html
<a href="#main-content" class="skip-link">Skip to main content</a>
<!-- ... navigation ... -->
<main id="main-content" tabindex="-1">
```
```css
.skip-link {
position: absolute;
left: -9999px;
z-index: 999;
}
.skip-link:focus {
left: 10px;
top: 10px;
background: #000;
color: #fff;
padding: 8px 16px;
}
```
## Keyboard Interaction Patterns
### Tabs
```
Tab → Move to tab list, then to tab panel
Arrow Left/Right → Switch between tabs
Home → First tab
End → Last tab
```
```html
<div role="tablist" aria-label="Settings">
<button role="tab" aria-selected="true" aria-controls="panel-1" id="tab-1">General</button>
<button role="tab" aria-selected="false" aria-controls="panel-2" id="tab-2" tabindex="-1">Security</button>
</div>
<div role="tabpanel" id="panel-1" aria-labelledby="tab-1">...</div>
<div role="tabpanel" id="panel-2" aria-labelledby="tab-2" hidden>...</div>
```
### Combobox / Autocomplete
```
Arrow Down → Open list / next option
Arrow Up → Previous option
Enter → Select option
Escape → Close list
Type → Filter options
```
### Menu
```
Enter/Space → Activate item
Arrow Down → Next item
Arrow Up → Previous item
Arrow Right → Open submenu
Arrow Left → Close submenu
Escape → Close menu
```
### Accordion
```
Enter/Space → Toggle section
Arrow Down → Next header
Arrow Up → Previous header
Home → First header
End → Last header
```
## Framework-Specific ARIA
### React
```jsx
// Announce route changes (SPA)
<div aria-live="polite" className="sr-only">
{`Navigated to pageTitle`}
</div>
// Error boundary with accessible error
<div role="alert">
<h2>Something went wrong</h2>
<p>{error.message}</p>
</div>
```
### Vue
```vue
<!-- Announce dynamic content -->
<div aria-live="polite">
<p v-if="results.length">{{ results.length }} results found</p>
</div>
<!-- Accessible toggle -->
<button
:aria-expanded="isOpen"
:aria-controls="panelId"
@click="toggle"
>
{{ isOpen ? 'Collapse' : 'Expand' }}
</button>
```
### Angular
```html
<!-- cdkTrapFocus for modals -->
<div cdkTrapFocus cdkTrapFocusAutoCapture role="dialog" aria-labelledby="dialog-title">
<h2 id="dialog-title">Confirm Action</h2>
</div>
<!-- LiveAnnouncer service -->
<!-- In component: this.liveAnnouncer.announce('Item added to cart'); -->
```
## Common ARIA Mistakes
| Mistake | Why It's Wrong | Fix |
|---------|---------------|-----|
| `<div role="button">` without keyboard | Div doesn't get keyboard events | Use `<button>` or add `tabindex="0"` + `onkeydown` |
| `aria-hidden="true"` on focusable element | Screen reader skips it but keyboard reaches it | Remove from tab order too: `tabindex="-1"` |
| `aria-label` overriding visible text | Confusing for sighted screen reader users | Use `aria-labelledby` pointing to visible text |
| Redundant ARIA on semantic HTML | `<nav role="navigation">` is redundant | Drop the `role` — `<nav>` implies it |
| `aria-live` on container that already has content | Initial content gets announced on load | Add `aria-live` to empty container, inject content after |
| Missing `aria-expanded` on toggles | Screen reader can't tell if section is open | Add `aria-expanded="true/false"` |
FILE:references/audit-report-template.md
# Audit Report Template
The scanner generates a stakeholder-ready report when run with the `--report` flag:
```bash
python scripts/a11y_scanner.py /path/to/project --report --output audit-report.md
```
## Generated Report Structure
```markdown
# Accessibility Audit Report
**Project:** Acme Dashboard
**Date:** 2026-03-18
**Standard:** WCAG 2.2 Level AA
**Tool:** a11y-audit v2.1.2
## Executive Summary
- Files Scanned: 127
- Total Violations: 14
- Critical: 3 | Major: 7 | Minor: 4
- Estimated Remediation: 8-12 hours
- Compliance Score: 72% (Target: 100%)
## Violations by Category
| Category | Count | Severity Breakdown |
|----------|-------|--------------------|
| Missing Alt Text | 3 | 2 Critical, 1 Minor |
| Keyboard Access | 4 | 2 Critical, 2 Major |
| Color Contrast | 3 | 3 Major |
| Form Labels | 2 | 2 Major |
| ARIA Usage | 2 | 2 Minor |
## Detailed Findings
[Per-violation details with file, line, WCAG criterion, and fix]
## Remediation Priority
1. Fix all Critical issues (blocks release)
2. Fix Major issues in current sprint
3. Schedule Minor issues for next sprint
## Recommendations
- Add a11y linting to CI pipeline (eslint-plugin-jsx-a11y)
- Include keyboard testing in QA checklist
- Schedule quarterly manual audit with assistive technology
```
FILE:references/ci-cd-integration.md
# CI/CD Integration for Accessibility Auditing
## GitHub Actions
```yaml
# .github/workflows/a11y-audit.yml
name: Accessibility Audit
on:
pull_request:
paths:
- 'src/**/*.tsx'
- 'src/**/*.vue'
- 'src/**/*.html'
- 'src/**/*.svelte'
jobs:
a11y-audit:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: '3.11'
- name: Run A11y Scanner
run: |
python scripts/a11y_scanner.py ./src --json > a11y-results.json
- name: Check for Critical Issues
run: |
python -c "
import json, sys
with open('a11y-results.json') as f:
data = json.load(f)
critical = [v for v in data.get('violations', []) if v['severity'] == 'critical']
if critical:
print(f'FAILED: {len(critical)} critical a11y violations found')
for v in critical:
print(f\" [{v['wcag']}] {v['file']}:{v['line']} - {v['message']}\")
sys.exit(1)
print('PASSED: No critical a11y violations')
"
- name: Upload Audit Report
if: always()
uses: actions/upload-artifact@v4
with:
name: a11y-audit-report
path: a11y-results.json
- name: Comment on PR
if: failure()
uses: marocchino/sticky-pull-request-comment@v2
with:
header: a11y-audit
message: |
## Accessibility Audit Failed
Critical WCAG 2.2 violations were found. See the uploaded artifact for details.
Run `python scripts/a11y_scanner.py ./src` locally to view and fix issues.
```
## GitLab CI
```yaml
# .gitlab-ci.yml
a11y-audit:
stage: test
image: python:3.11-slim
script:
- python scripts/a11y_scanner.py ./src --json > a11y-results.json
- python -c "
import json, sys;
data = json.load(open('a11y-results.json'));
critical = [v for v in data.get('violations', []) if v['severity'] == 'critical'];
sys.exit(1) if critical else print('A11y audit passed')
"
artifacts:
paths:
- a11y-results.json
when: always
rules:
- changes:
- "src/**/*.{tsx,vue,html,svelte}"
```
## Azure DevOps
```yaml
# azure-pipelines.yml
- task: PythonScript@0
displayName: 'Run A11y Audit'
inputs:
scriptSource: 'filePath'
scriptPath: 'scripts/a11y_scanner.py'
arguments: './src --json --output $(Build.ArtifactStagingDirectory)/a11y-results.json'
- task: PublishBuildArtifacts@1
condition: always()
inputs:
PathtoPublish: '$(Build.ArtifactStagingDirectory)/a11y-results.json'
ArtifactName: 'a11y-audit-report'
```
## Pre-Commit Hook
```bash
#!/bin/bash
# .git/hooks/pre-commit
# Run a11y scan on staged files only
STAGED_FILES=$(git diff --cached --name-only --diff-filter=ACM | grep -E '\.(tsx|vue|html|svelte|jsx)$')
if [ -n "$STAGED_FILES" ]; then
echo "Running accessibility audit on staged files..."
for file in $STAGED_FILES; do
python scripts/a11y_scanner.py "$file" --severity critical --quiet
if [ $? -ne 0 ]; then
echo "A11y audit FAILED for $file. Fix critical issues before committing."
exit 1
fi
done
echo "A11y audit passed."
fi
```
FILE:references/color-contrast-guide.md
# Color Contrast Guide
## Contrast Checker Usage
The `contrast_checker.py` script validates color pairs against WCAG 2.2 contrast requirements.
```bash
# Check a single color pair
python scripts/contrast_checker.py --fg "#777777" --bg "#ffffff"
# Output:
# Foreground: #777777 | Background: #ffffff
# Contrast Ratio: 4.48:1
# AA Normal Text (4.5:1): FAIL
# AA Large Text (3.0:1): PASS
# AAA Normal Text (7.0:1): FAIL
# Suggested alternative: #767676 (4.54:1 - passes AA)
# Scan a CSS file for all color pairs
python scripts/contrast_checker.py --file src/styles/globals.css
# Scan Tailwind classes in components
python scripts/contrast_checker.py --tailwind src/components/
```
## Common Contrast Fixes
| Original Color | Contrast on White | Fix | New Contrast |
|----------------|------------------|-----|--------------|
| `#aaaaaa` | 2.32:1 | `#767676` | 4.54:1 (AA) |
| `#999999` | 2.85:1 | `#767676` | 4.54:1 (AA) |
| `#888888` | 3.54:1 | `#767676` | 4.54:1 (AA) |
| `#777777` | 4.48:1 | `#757575` | 4.60:1 (AA) |
| `#66bb6a` | 3.06:1 | `#2e7d32` | 5.87:1 (AA) |
| `#42a5f5` | 2.81:1 | `#1565c0` | 6.08:1 (AA) |
| `#ef5350` | 3.13:1 | `#c62828` | 5.57:1 (AA) |
## Tailwind CSS Accessible Palette Mapping
| Inaccessible Class | Contrast on White | Accessible Alternative | Contrast |
|---------------------|------------------|----------------------|----------|
| `text-gray-400` | 2.68:1 | `text-gray-600` | 5.74:1 |
| `text-blue-400` | 2.81:1 | `text-blue-700` | 5.96:1 |
| `text-green-400` | 2.12:1 | `text-green-700` | 5.18:1 |
| `text-red-400` | 3.04:1 | `text-red-700` | 6.05:1 |
| `text-yellow-500` | 1.47:1 | `text-yellow-800` | 7.34:1 |
## Screen Reader Utility Class
Every project should include this utility class for visually hiding content while keeping it accessible to screen readers:
```css
/* Visually hidden but accessible to screen readers */
.sr-only {
position: absolute;
width: 1px;
height: 1px;
padding: 0;
margin: -1px;
overflow: hidden;
clip: rect(0, 0, 0, 0);
white-space: nowrap;
border-width: 0;
}
/* Allow the element to be focusable when navigated to via keyboard */
.sr-only-focusable:focus,
.sr-only-focusable:active {
position: static;
width: auto;
height: auto;
padding: inherit;
margin: inherit;
overflow: visible;
clip: auto;
white-space: inherit;
}
```
Tailwind CSS includes this as `sr-only` by default. For other frameworks:
- **Angular**: Add to `styles.scss`
- **Vue**: Add to `assets/global.css`
- **Svelte**: Add to `app.css`
FILE:references/examples-by-framework.md
# Accessibility Audit Examples by Framework
## Example 1: Vue SFC Form Audit
```vue
<!-- BEFORE: src/components/LoginForm.vue -->
<template>
<form @submit="handleLogin">
<input type="text" placeholder="Email" v-model="email" />
<input type="password" placeholder="Password" v-model="password" />
<div v-if="error" style="color: red">{{ error }}</div>
<div @click="handleLogin">Sign In</div>
</form>
</template>
```
**Violations detected:**
| # | WCAG | Severity | Issue |
|---|------|----------|-------|
| 1 | 1.3.1 | Critical | Inputs missing associated `<label>` elements |
| 2 | 3.3.2 | Major | Placeholder text used as only label (disappears on input) |
| 3 | 2.1.1 | Critical | `<div @click>` not keyboard accessible |
| 4 | 4.1.3 | Major | Error message not announced to screen readers |
| 5 | 3.3.1 | Major | Error not programmatically associated with input |
```vue
<!-- AFTER: src/components/LoginForm.vue -->
<template>
<form @submit.prevent="handleLogin" aria-label="Sign in to your account">
<div class="field">
<label for="login-email">Email</label>
<input
id="login-email"
type="email"
v-model="email"
autocomplete="email"
required
:aria-describedby="emailError ? 'email-error' : undefined"
:aria-invalid="!!emailError"
/>
<span v-if="emailError" id="email-error" role="alert">
{{ emailError }}
</span>
</div>
<div class="field">
<label for="login-password">Password</label>
<input
id="login-password"
type="password"
v-model="password"
autocomplete="current-password"
required
:aria-describedby="passwordError ? 'password-error' : undefined"
:aria-invalid="!!passwordError"
/>
<span v-if="passwordError" id="password-error" role="alert">
{{ passwordError }}
</span>
</div>
<div v-if="error" role="alert" aria-live="assertive" class="form-error">
{{ error }}
</div>
<button type="submit">Sign In</button>
</form>
</template>
```
## Example 2: Angular Template Audit
```html
<!-- BEFORE: src/app/dashboard/dashboard.component.html -->
<div class="tabs">
<div *ngFor="let tab of tabs"
(click)="selectTab(tab)"
[class.active]="tab.active">
{{ tab.label }}
</div>
</div>
<div class="tab-content">
<div *ngIf="selectedTab">{{ selectedTab.content }}</div>
</div>
```
**Violations detected:**
| # | WCAG | Severity | Issue |
|---|------|----------|-------|
| 1 | 4.1.2 | Critical | Tab widget missing ARIA roles (`tablist`, `tab`, `tabpanel`) |
| 2 | 2.1.1 | Critical | Tabs not keyboard navigable (arrow keys, Home, End) |
| 3 | 2.4.11 | Major | No visible focus indicator on active tab |
```html
<!-- AFTER: src/app/dashboard/dashboard.component.html -->
<div class="tabs" role="tablist" aria-label="Dashboard sections">
<button
*ngFor="let tab of tabs; let i = index"
role="tab"
[id]="'tab-' + tab.id"
[attr.aria-selected]="tab.active"
[attr.aria-controls]="'panel-' + tab.id"
[attr.tabindex]="tab.active ? 0 : -1"
(click)="selectTab(tab)"
(keydown)="handleTabKeydown($event, i)"
class="tab-button"
[class.active]="tab.active">
{{ tab.label }}
</button>
</div>
<div
*ngIf="selectedTab"
role="tabpanel"
[id]="'panel-' + selectedTab.id"
[attr.aria-labelledby]="'tab-' + selectedTab.id"
tabindex="0"
class="tab-content">
{{ selectedTab.content }}
</div>
```
**Supporting TypeScript for keyboard navigation:**
```typescript
// dashboard.component.ts
handleTabKeydown(event: KeyboardEvent, index: number): void {
const tabCount = this.tabs.length;
let newIndex = index;
switch (event.key) {
case 'ArrowRight':
newIndex = (index + 1) % tabCount;
break;
case 'ArrowLeft':
newIndex = (index - 1 + tabCount) % tabCount;
break;
case 'Home':
newIndex = 0;
break;
case 'End':
newIndex = tabCount - 1;
break;
default:
return;
}
event.preventDefault();
this.selectTab(this.tabs[newIndex]);
// Move focus to the new tab button
const tabElement = document.getElementById(`tab-this.tabs[newIndex].id`);
tabElement?.focus();
}
```
## Example 3: Next.js Page-Level Audit
```tsx
// BEFORE: src/app/page.tsx
export default function Home() {
return (
<main>
<div className="text-4xl font-bold">Welcome to Acme</div>
<div className="mt-4">
Build better products with our platform.
</div>
<div className="mt-8 bg-blue-600 text-white px-6 py-3 rounded cursor-pointer"
onClick={() => router.push('/signup')}>
Get Started
</div>
</main>
);
}
```
**Violations detected:**
| # | WCAG | Severity | Issue |
|---|------|----------|-------|
| 1 | 1.3.1 | Major | Heading uses `<div>` instead of `<h1>` -- no semantic structure |
| 2 | 2.4.2 | Major | Page missing `<title>` (Next.js metadata) |
| 3 | 2.1.1 | Critical | CTA uses `<div onClick>` -- not keyboard accessible |
| 4 | 3.1.1 | Minor | `<html>` missing `lang` attribute (check `layout.tsx`) |
```tsx
// AFTER: src/app/page.tsx
import type { Metadata } from 'next';
import Link from 'next/link';
export const metadata: Metadata = {
title: 'Acme - Build Better Products',
description: 'Build better products with the Acme platform.',
};
export default function Home() {
return (
<main>
<h1 className="text-4xl font-bold">Welcome to Acme</h1>
<p className="mt-4">
Build better products with our platform.
</p>
<Link
href="/signup"
className="mt-8 inline-block bg-blue-600 text-white px-6 py-3 rounded
hover:bg-blue-700 focus-visible:outline focus-visible:outline-2
focus-visible:outline-offset-2 focus-visible:outline-blue-600"
>
Get Started
</Link>
</main>
);
}
```
```tsx
// Also fix: src/app/layout.tsx
export default function RootLayout({ children }: { children: React.ReactNode }) {
return (
<html lang="en">
<body>{children}</body>
</html>
);
}
```
## Example 4: Svelte Component Audit
```svelte
<!-- BEFORE: src/lib/components/Accordion.svelte -->
<script>
let openIndex = -1;
function toggle(i) { openIndex = openIndex === i ? -1 : i; }
</script>
{#each items as item, i}
<div class="header" on:click={() => toggle(i)}>
{item.title}
</div>
{#if openIndex === i}
<div class="content">{item.body}</div>
{/if}
{/each}
```
**Violations detected:**
| # | WCAG | Severity | Issue |
|---|------|----------|-------|
| 1 | 4.1.2 | Critical | Accordion missing ARIA roles and properties |
| 2 | 2.1.1 | Critical | Headers not keyboard accessible |
| 3 | 2.5.8 | Minor | Click targets may be smaller than 24x24px (NEW in WCAG 2.2) |
```svelte
<!-- AFTER: src/lib/components/Accordion.svelte -->
<script>
export let items = [];
let openIndex = -1;
function toggle(i) {
openIndex = openIndex === i ? -1 : i;
}
function handleKeydown(event, i) {
if (event.key === 'Enter' || event.key === ' ') {
event.preventDefault();
toggle(i);
}
}
</script>
<div class="accordion">
{#each items as item, i}
<h3>
<button
class="accordion-header"
aria-expanded={openIndex === i}
aria-controls="panel-{i}"
id="header-{i}"
on:click={() => toggle(i)}
on:keydown={(e) => handleKeydown(e, i)}
>
{item.title}
<span class="icon" aria-hidden="true">
{openIndex === i ? '−' : '+'}
</span>
</button>
</h3>
<div
id="panel-{i}"
role="region"
aria-labelledby="header-{i}"
class="accordion-content"
class:open={openIndex === i}
hidden={openIndex !== i}
>
{item.body}
</div>
{/each}
</div>
<style>
.accordion-header {
min-height: 44px; /* WCAG 2.5.8 Target Size */
width: 100%;
padding: 12px 16px;
cursor: pointer;
text-align: left;
}
.accordion-header:focus-visible {
outline: 2px solid #005fcc;
outline-offset: 2px;
}
</style>
```
FILE:references/framework-a11y-patterns.md
# Framework-Specific Accessibility Patterns
## React / Next.js
### Common Issues and Fixes
**Image alt text:**
```jsx
// ❌ Bad
<img src="/hero.jpg" />
<Image src="/hero.jpg" width={800} height={400} />
// ✅ Good
<img src="/hero.jpg" alt="Team collaborating in office" />
<Image src="/hero.jpg" width={800} height={400} alt="Team collaborating in office" />
// ✅ Decorative image
<img src="/divider.svg" alt="" role="presentation" />
```
**Form labels:**
```jsx
// ❌ Bad — placeholder as label
<input placeholder="Email" type="email" />
// ✅ Good — explicit label
<label htmlFor="email">Email</label>
<input id="email" type="email" placeholder="user@example.com" />
// ✅ Good — aria-label for icon-only inputs
<input type="search" aria-label="Search products" />
```
**Click handlers on divs:**
```jsx
// ❌ Bad — not keyboard accessible
<div onClick={handleClick}>Click me</div>
// ✅ Good — use button
<button onClick={handleClick}>Click me</button>
// ✅ If div is required — add keyboard support
<div
role="button"
tabIndex={0}
onClick={handleClick}
onKeyDown={(e) => { if (e.key === 'Enter' || e.key === ' ') handleClick(); }}
>
Click me
</div>
```
**SPA route announcements (Next.js App Router):**
```jsx
// Layout component — announce page changes
'use client';
import { usePathname } from 'next/navigation';
import { useEffect, useState } from 'react';
export function RouteAnnouncer() {
const pathname = usePathname();
const [announcement, setAnnouncement] = useState('');
useEffect(() => {
const title = document.title;
setAnnouncement(`Navigated to title`);
}, [pathname]);
return (
<div aria-live="assertive" role="status" className="sr-only">
{announcement}
</div>
);
}
```
**Focus management after dynamic content:**
```jsx
// After adding item to list, announce it
const [items, setItems] = useState([]);
const statusRef = useRef(null);
const addItem = (item) => {
setItems([...items, item]);
// Announce to screen readers
statusRef.current.textContent = `item.name added to list`;
};
return (
<>
<div ref={statusRef} aria-live="polite" className="sr-only" />
{/* list content */}
</>
);
```
### React-Specific Libraries
- `@radix-ui/*` — accessible primitives (Dialog, Tabs, Select, etc.)
- `@headlessui/react` — unstyled accessible components
- `react-aria` — Adobe's accessibility hooks
- `eslint-plugin-jsx-a11y` — lint rules for JSX accessibility
## Vue 3
### Common Issues and Fixes
**Dynamic content announcements:**
```vue
<template>
<div aria-live="polite" class="sr-only">
{{ announcement }}
</div>
<button @click="search">Search</button>
<ul v-if="results.length">
<li v-for="r in results" :key="r.id">{{ r.name }}</li>
</ul>
</template>
<script setup>
import { ref } from 'vue';
const results = ref([]);
const announcement = ref('');
async function search() {
results.value = await fetchResults();
announcement.value = `results.value.length results found`;
}
</script>
```
**Conditional rendering with focus:**
```vue
<template>
<button @click="showForm = true">Add Item</button>
<form v-if="showForm" ref="formRef">
<label for="name">Name</label>
<input id="name" ref="nameInput" />
</form>
</template>
<script setup>
import { ref, nextTick } from 'vue';
const showForm = ref(false);
const nameInput = ref(null);
watch(showForm, async (val) => {
if (val) {
await nextTick();
nameInput.value?.focus();
}
});
</script>
```
### Vue-Specific Libraries
- `vue-announcer` — route change announcements
- `@headlessui/vue` — accessible components
- `eslint-plugin-vuejs-accessibility` — lint rules
## Angular
### Common Issues and Fixes
**CDK accessibility utilities:**
```typescript
import { LiveAnnouncer } from '@angular/cdk/a11y';
import { FocusTrapFactory } from '@angular/cdk/a11y';
@Component({...})
export class MyComponent {
constructor(
private liveAnnouncer: LiveAnnouncer,
private focusTrapFactory: FocusTrapFactory
) {}
addItem(item: Item) {
this.items.push(item);
this.liveAnnouncer.announce(`item.name added`);
}
openDialog(element: HTMLElement) {
const focusTrap = this.focusTrapFactory.create(element);
focusTrap.focusInitialElement();
}
}
```
**Template-driven forms:**
```html
<!-- ❌ Bad -->
<input [formControl]="email" placeholder="Email" />
<!-- ✅ Good -->
<label for="email">Email address</label>
<input id="email" [formControl]="email"
[attr.aria-invalid]="email.invalid && email.touched"
[attr.aria-describedby]="email.invalid ? 'email-error' : null" />
<div id="email-error" *ngIf="email.invalid && email.touched" role="alert">
Please enter a valid email address.
</div>
```
### Angular-Specific Tools
- `@angular/cdk/a11y` — `FocusTrap`, `LiveAnnouncer`, `FocusMonitor`
- `codelyzer` — a11y lint rules for Angular templates
## Svelte / SvelteKit
### Common Issues and Fixes
```svelte
<!-- ❌ Bad — on:click without keyboard -->
<div on:click={handleClick}>Action</div>
<!-- ✅ Good — Svelte a11y warning built-in -->
<button on:click={handleClick}>Action</button>
<!-- ✅ Accessible toggle -->
<button
on:click={() => isOpen = !isOpen}
aria-expanded={isOpen}
aria-controls="panel"
>
{isOpen ? 'Close' : 'Open'} Details
</button>
{#if isOpen}
<div id="panel" role="region" aria-labelledby="toggle-btn">
Panel content
</div>
{/if}
```
**Note:** Svelte has built-in a11y warnings in the compiler — it flags missing alt text, click-without-keyboard, and other common issues at build time.
## Plain HTML
### Checklist for Static Sites
```html
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Descriptive Page Title</title>
</head>
<body>
<!-- Skip link -->
<a href="#main" class="skip-link">Skip to main content</a>
<header>
<nav aria-label="Main navigation">
<ul>
<li><a href="/">Home</a></li>
<li><a href="/about" aria-current="page">About</a></li>
</ul>
</nav>
</header>
<main id="main" tabindex="-1">
<h1>Page Heading</h1>
<!-- Only one h1 per page -->
<!-- Heading levels don't skip (h1 → h2 → h3, never h1 → h3) -->
</main>
<footer>
<p>© 2026 Company Name</p>
</footer>
</body>
</html>
```
## CSS Accessibility Patterns
### Focus Indicators
```css
/* ❌ Bad — removes focus indicator entirely */
:focus { outline: none; }
/* ✅ Good — custom focus indicator */
:focus-visible {
outline: 2px solid #005fcc;
outline-offset: 2px;
}
/* ✅ Good — enhanced for high contrast mode */
@media (forced-colors: active) {
:focus-visible {
outline: 2px solid ButtonText;
}
}
```
### Reduced Motion
```css
/* ✅ Respect prefers-reduced-motion */
@media (prefers-reduced-motion: reduce) {
*, *::before, *::after {
animation-duration: 0.01ms !important;
animation-iteration-count: 1 !important;
transition-duration: 0.01ms !important;
}
}
```
### Screen Reader Only
```css
.sr-only {
position: absolute;
width: 1px;
height: 1px;
padding: 0;
margin: -1px;
overflow: hidden;
clip: rect(0, 0, 0, 0);
white-space: nowrap;
border-width: 0;
}
```
## Fix Patterns Catalog
### React / Next.js Fix Patterns
#### Missing Alt Text (1.1.1)
```tsx
// BEFORE
<img src={hero} />
// AFTER - Informational image
<img src={hero} alt="Team collaborating around a whiteboard" />
// AFTER - Decorative image
<img src={divider} alt="" role="presentation" />
```
#### Non-Interactive Element with Click Handler (2.1.1)
```tsx
// BEFORE
<div onClick={handleClick}>Click me</div>
// AFTER - If it navigates
<Link href="/destination">Click me</Link>
// AFTER - If it performs an action
<button type="button" onClick={handleClick}>Click me</button>
```
#### Missing Focus Management in Modals (2.4.3)
```tsx
// BEFORE
function Modal({ isOpen, onClose, children }) {
if (!isOpen) return null;
return <div className="modal-overlay">{children}</div>;
}
// AFTER
import { useEffect, useRef } from 'react';
function Modal({ isOpen, onClose, children, title }) {
const modalRef = useRef(null);
const previousFocus = useRef(null);
useEffect(() => {
if (isOpen) {
previousFocus.current = document.activeElement;
modalRef.current?.focus();
} else {
previousFocus.current?.focus();
}
}, [isOpen]);
useEffect(() => {
if (!isOpen) return;
const handleKeydown = (e) => {
if (e.key === 'Escape') onClose();
if (e.key === 'Tab') {
const focusable = modalRef.current?.querySelectorAll(
'button, [href], input, select, textarea, [tabindex]:not([tabindex="-1"])'
);
if (!focusable?.length) return;
const first = focusable[0];
const last = focusable[focusable.length - 1];
if (e.shiftKey && document.activeElement === first) {
e.preventDefault();
last.focus();
} else if (!e.shiftKey && document.activeElement === last) {
e.preventDefault();
first.focus();
}
}
};
document.addEventListener('keydown', handleKeydown);
return () => document.removeEventListener('keydown', handleKeydown);
}, [isOpen, onClose]);
if (!isOpen) return null;
return (
<div className="modal-overlay" onClick={onClose} aria-hidden="true">
<div
ref={modalRef}
role="dialog"
aria-modal="true"
aria-label={title}
tabIndex={-1}
onClick={(e) => e.stopPropagation()}
>
<button
onClick={onClose}
aria-label="Close dialog"
className="modal-close"
>
×
</button>
{children}
</div>
</div>
);
}
```
#### Focus Appearance (2.4.11 -- NEW in WCAG 2.2)
```css
/* BEFORE */
button:focus {
outline: none; /* Removes default focus indicator */
}
/* AFTER - Meets WCAG 2.2 Focus Appearance */
button:focus-visible {
outline: 2px solid #005fcc;
outline-offset: 2px;
}
```
```tsx
// Tailwind CSS pattern
<button className="focus-visible:outline focus-visible:outline-2 focus-visible:outline-offset-2 focus-visible:outline-blue-600">
Submit
</button>
```
### Vue Fix Patterns
#### Missing Form Labels (1.3.1)
```vue
<!-- BEFORE -->
<input type="text" v-model="name" placeholder="Name" />
<!-- AFTER -->
<label for="user-name">Name</label>
<input id="user-name" type="text" v-model="name" autocomplete="name" />
```
#### Dynamic Content Without Live Region (4.1.3)
```vue
<!-- BEFORE -->
<div v-if="status">{{ statusMessage }}</div>
<!-- AFTER -->
<div aria-live="polite" aria-atomic="true">
<p v-if="status">{{ statusMessage }}</p>
</div>
```
#### Vue Router Navigation Announcements (2.4.2)
```typescript
// router/index.ts
router.afterEach((to) => {
const title = to.meta.title || 'Page';
document.title = `title | My App`;
// Announce route change to screen readers
const announcer = document.getElementById('route-announcer');
if (announcer) {
announcer.textContent = `Navigated to title`;
}
});
```
```vue
<!-- App.vue - Add announcer element -->
<div
id="route-announcer"
role="status"
aria-live="assertive"
aria-atomic="true"
class="sr-only"
></div>
```
### Angular Fix Patterns
#### Missing ARIA on Custom Components (4.1.2)
```typescript
// BEFORE
@Component({
selector: 'app-dropdown',
template: `
<div (click)="toggle()">{{ selected }}</div>
<div *ngIf="isOpen">
<div *ngFor="let opt of options" (click)="select(opt)">{{ opt }}</div>
</div>
`
})
// AFTER
@Component({
selector: 'app-dropdown',
template: `
<button
role="combobox"
[attr.aria-expanded]="isOpen"
aria-haspopup="listbox"
[attr.aria-label]="label"
(click)="toggle()"
(keydown)="handleKeydown($event)"
>
{{ selected }}
</button>
<ul *ngIf="isOpen" role="listbox" [attr.aria-label]="label + ' options'">
<li
*ngFor="let opt of options; let i = index"
role="option"
[attr.aria-selected]="opt === selected"
[attr.id]="'option-' + i"
(click)="select(opt)"
(keydown)="handleOptionKeydown($event, opt, i)"
tabindex="-1"
>
{{ opt }}
</li>
</ul>
`
})
```
#### Angular CDK A11y Module Integration
```typescript
// Use Angular CDK for focus trap in dialogs
import { A11yModule } from '@angular/cdk/a11y';
@Component({
template: `
<div cdkTrapFocus cdkTrapFocusAutoCapture>
<h2 id="dialog-title">Edit Profile</h2>
<!-- dialog content -->
</div>
`
})
```
### Svelte Fix Patterns
#### Accessible Announcements (4.1.3)
```svelte
<!-- BEFORE -->
{#if message}
<p class="toast">{message}</p>
{/if}
<!-- AFTER -->
<div aria-live="polite" class="sr-only">
{#if message}
<p>{message}</p>
{/if}
</div>
<div class="toast" aria-hidden="true">
{#if message}
<p>{message}</p>
{/if}
</div>
```
#### SvelteKit Page Titles (2.4.2)
```svelte
<!-- +page.svelte -->
<svelte:head>
<title>Dashboard | My App</title>
</svelte:head>
```
### Plain HTML Fix Patterns
#### Skip Navigation Link (2.4.1)
```html
<!-- BEFORE -->
<body>
<nav><!-- long navigation --></nav>
<main><!-- content --></main>
</body>
<!-- AFTER -->
<body>
<a href="#main-content" class="skip-link">Skip to main content</a>
<nav aria-label="Main navigation"><!-- long navigation --></nav>
<main id="main-content" tabindex="-1"><!-- content --></main>
</body>
```
```css
.skip-link {
position: absolute;
top: -40px;
left: 0;
padding: 8px 16px;
background: #005fcc;
color: #fff;
z-index: 1000;
transition: top 0.2s;
}
.skip-link:focus {
top: 0;
}
```
#### Accessible Data Table (1.3.1)
```html
<!-- BEFORE -->
<table>
<tr><td>Name</td><td>Email</td><td>Role</td></tr>
<tr><td>Alice</td><td>alice@co.com</td><td>Admin</td></tr>
</table>
<!-- AFTER -->
<table aria-label="Team members">
<caption class="sr-only">List of team members and their roles</caption>
<thead>
<tr>
<th scope="col">Name</th>
<th scope="col">Email</th>
<th scope="col">Role</th>
</tr>
</thead>
<tbody>
<tr>
<th scope="row">Alice</th>
<td>alice@co.com</td>
<td>Admin</td>
</tr>
</tbody>
</table>
```
FILE:references/testing-checklist.md
# Accessibility Testing Checklist
Use this checklist after applying fixes to verify accessibility manually.
## Keyboard Navigation
- [ ] All interactive elements reachable via Tab key
- [ ] Tab order follows visual/logical reading order
- [ ] Focus indicator visible on every focusable element (2px+ outline)
- [ ] Modals trap focus and return focus on close
- [ ] Escape key closes modals, dropdowns, and popups
- [ ] Arrow keys navigate within composite widgets (tabs, menus, listboxes)
- [ ] No keyboard traps (user can always Tab away)
## Screen Reader
- [ ] All images have appropriate alt text (or `alt=""` for decorative)
- [ ] Headings create logical document outline (h1 -> h2 -> h3)
- [ ] Form inputs have associated labels
- [ ] Error messages announced via `aria-live` or `role="alert"`
- [ ] Page title updates on navigation (SPA)
- [ ] Dynamic content changes announced appropriately
## Visual
- [ ] Text contrast meets 4.5:1 for normal text, 3:1 for large text
- [ ] UI component contrast meets 3:1 against background
- [ ] Content reflows without horizontal scrolling at 320px width
- [ ] Text resizable to 200% without loss of content
- [ ] No information conveyed by color alone
- [ ] Focus indicators meet 2.4.11 Focus Appearance criteria
## Motion and Media
- [ ] Animations respect `prefers-reduced-motion`
- [ ] No auto-playing media with audio
- [ ] No content flashing more than 3 times per second
- [ ] Video has captions; audio has transcripts
## Forms
- [ ] All inputs have visible labels
- [ ] Required fields indicated (not by color alone)
- [ ] Error messages specific and associated with input via `aria-describedby`
- [ ] Autocomplete attributes present on common fields (name, email, etc.)
- [ ] No CAPTCHA without alternative method (WCAG 2.2 3.3.8)
FILE:references/wcag-22-new-criteria.md
# WCAG 2.2 New Success Criteria Reference
These criteria were added in WCAG 2.2 and are commonly missed.
## 2.4.11 Focus Appearance (Level AA)
The focus indicator must have a minimum area of a 2px perimeter around the component and a contrast ratio of at least 3:1 against adjacent colors.
**Pattern:**
```css
:focus-visible {
outline: 2px solid #005fcc;
outline-offset: 2px;
}
```
## 2.5.7 Dragging Movements (Level AA)
Any functionality that uses dragging must have a single-pointer alternative (click, tap).
**Pattern:**
```tsx
// Sortable list: support both drag and button-based reorder
<li draggable onDragStart={handleDrag}>
{item.name}
<button onClick={() => moveUp(index)} aria-label={`Move item.name up`}>
Move Up
</button>
<button onClick={() => moveDown(index)} aria-label={`Move item.name down`}>
Move Down
</button>
</li>
```
## 2.5.8 Target Size (Level AA)
Interactive targets must be at least 24x24 CSS pixels, with exceptions for inline text links and elements where the spacing provides equivalent clearance.
**Pattern:**
```css
button, a, input, select, textarea {
min-height: 24px;
min-width: 24px;
}
/* Recommended: 44x44px for touch targets */
@media (pointer: coarse) {
button, a, input[type="checkbox"], input[type="radio"] {
min-height: 44px;
min-width: 44px;
}
}
```
## 3.3.7 Redundant Entry (Level A)
Information previously entered by the user must be auto-populated or available for selection when needed again in the same process.
**Pattern:**
```tsx
// Multi-step form: persist data across steps
const [formData, setFormData] = useState({});
// Step 2 pre-fills shipping address from billing
<input
defaultValue={formData.billingAddress || ''}
autoComplete="shipping street-address"
/>
```
## 3.3.8 Accessible Authentication (Level AA)
Authentication must not require cognitive function tests (e.g., remembering a password, solving a puzzle) unless an alternative is provided.
**Pattern:**
- Support password managers (`autocomplete="current-password"`)
- Offer passkey / biometric authentication
- Allow copy-paste in password fields (never block paste)
- Provide email/SMS OTP as alternative to CAPTCHA
FILE:references/wcag-quick-ref.md
# WCAG 2.2 Quick Reference — Level A & AA
## Perceivable
### 1.1 Text Alternatives
| Criterion | Level | Requirement | Common Violation |
|-----------|-------|-------------|------------------|
| 1.1.1 Non-text Content | A | All images have `alt` text; decorative images use `alt=""` or `role="presentation"` | `<img src="logo.png">` without alt |
### 1.2 Time-Based Media
| Criterion | Level | Requirement | Common Violation |
|-----------|-------|-------------|------------------|
| 1.2.1 Audio-only / Video-only | A | Provide transcript or audio description | Video without captions |
| 1.2.2 Captions | A | Captions for all prerecorded audio in video | Missing `<track kind="captions">` |
| 1.2.3 Audio Description | A | Audio description for prerecorded video | No descriptive track |
| 1.2.5 Audio Description (Prerecorded) | AA | Audio description for all prerecorded video | Same as 1.2.3 but stricter |
### 1.3 Adaptable
| Criterion | Level | Requirement | Common Violation |
|-----------|-------|-------------|------------------|
| 1.3.1 Info and Relationships | A | Semantic markup conveys structure | Using `<div>` instead of `<nav>`, `<main>`, `<header>` |
| 1.3.2 Meaningful Sequence | A | Reading order matches visual order | CSS flex/grid reordering without DOM reorder |
| 1.3.3 Sensory Characteristics | A | Don't rely solely on color, shape, position | "Click the red button" |
| 1.3.4 Orientation | AA | Content not restricted to portrait/landscape | CSS `orientation: portrait` lock |
| 1.3.5 Identify Input Purpose | AA | Input purpose identifiable via `autocomplete` | Missing `autocomplete="email"` on email inputs |
### 1.4 Distinguishable
| Criterion | Level | Requirement | Ratio |
|-----------|-------|-------------|-------|
| 1.4.1 Use of Color | A | Color not sole means of conveying info | Red-only error indicators |
| 1.4.2 Audio Control | A | Auto-playing audio has pause/stop | `autoplay` without `controls` |
| 1.4.3 Contrast (Minimum) | AA | Text: 4.5:1, Large text: 3:1 | Light gray text on white |
| 1.4.4 Resize Text | AA | Text resizable to 200% without loss | Fixed `px` font sizes |
| 1.4.5 Images of Text | AA | Use real text, not text in images | Logo text as PNG |
| 1.4.10 Reflow | AA | Content reflows at 320px width | Horizontal scrolling at mobile widths |
| 1.4.11 Non-text Contrast | AA | UI components and graphics: 3:1 | Low-contrast borders, icons |
| 1.4.12 Text Spacing | AA | No loss of content when spacing adjusted | Fixed-height containers clipping |
| 1.4.13 Content on Hover/Focus | AA | Dismissible, hoverable, persistent | Tooltips that disappear on mouse move |
## Operable
### 2.1 Keyboard Accessible
| Criterion | Level | Requirement | Common Violation |
|-----------|-------|-------------|------------------|
| 2.1.1 Keyboard | A | All functionality via keyboard | `onClick` without `onKeyDown` |
| 2.1.2 No Keyboard Trap | A | Focus can move away from any component | Modal without focus trap escape |
| 2.1.4 Character Key Shortcuts | A | Single-key shortcuts can be turned off | `accesskey` conflicts |
### 2.4 Navigable
| Criterion | Level | Requirement | Common Violation |
|-----------|-------|-------------|------------------|
| 2.4.1 Bypass Blocks | A | Skip navigation link | No "Skip to content" link |
| 2.4.2 Page Titled | A | Descriptive `<title>` | `<title>Untitled</title>` |
| 2.4.3 Focus Order | A | Logical tab order | `tabindex` > 0 |
| 2.4.4 Link Purpose | A | Link text describes destination | "Click here", "Read more" |
| 2.4.6 Headings and Labels | AA | Descriptive headings | Generic headings |
| 2.4.7 Focus Visible | AA | Visible focus indicator | `outline: none` without replacement |
| 2.4.11 Focus Not Obscured | AA | Focused element not hidden by sticky header | Fixed header covering focused element |
### 2.5 Input Modalities
| Criterion | Level | Requirement | Common Violation |
|-----------|-------|-------------|------------------|
| 2.5.1 Pointer Gestures | A | Multi-point gestures have single-point alternative | Pinch-to-zoom only |
| 2.5.2 Pointer Cancellation | A | Down-event doesn't trigger action | `mousedown` instead of `click` |
| 2.5.3 Label in Name | A | Visible label is in accessible name | Button shows "Submit" but `aria-label="btn1"` |
| 2.5.4 Motion Actuation | A | Motion-triggered actions have alternative | Shake-to-undo only |
| 2.5.7 Dragging Movements | AA | Drag has single-pointer alternative | Drag-and-drop only reordering |
| 2.5.8 Target Size | AA | Touch targets minimum 24x24 CSS pixels | Tiny mobile buttons |
## Understandable
### 3.1 Readable
| Criterion | Level | Requirement | Common Violation |
|-----------|-------|-------------|------------------|
| 3.1.1 Language of Page | A | `<html lang="en">` | Missing `lang` attribute |
| 3.1.2 Language of Parts | AA | `lang` on foreign-language spans | Mixed-language content without `lang` |
### 3.2 Predictable
| Criterion | Level | Requirement | Common Violation |
|-----------|-------|-------------|------------------|
| 3.2.1 On Focus | A | Focus doesn't trigger unexpected change | Auto-submitting on focus |
| 3.2.2 On Input | A | Input doesn't trigger unexpected change | Auto-navigating on select change |
| 3.2.3 Consistent Navigation | AA | Navigation consistent across pages | Menu order changes per page |
| 3.2.4 Consistent Identification | AA | Same function = same label | "Search" vs "Find" for same action |
### 3.3 Input Assistance
| Criterion | Level | Requirement | Common Violation |
|-----------|-------|-------------|------------------|
| 3.3.1 Error Identification | A | Errors described in text | Red border only, no message |
| 3.3.2 Labels or Instructions | A | Labels for required input | Placeholder as only label |
| 3.3.3 Error Suggestion | AA | Suggest corrections | "Invalid input" without guidance |
| 3.3.4 Error Prevention | AA | Reversible submissions for legal/financial | No confirmation for payment |
| 3.3.7 Redundant Entry | A | Don't ask for same info twice | Re-entering address in checkout |
| 3.3.8 Accessible Authentication | AA | No cognitive function test for login | CAPTCHA without audio alternative |
## Robust
### 4.1 Compatible
| Criterion | Level | Requirement | Common Violation |
|-----------|-------|-------------|------------------|
| 4.1.2 Name, Role, Value | A | Custom controls have accessible name and role | Custom dropdown without ARIA |
| 4.1.3 Status Messages | AA | Status updates announced without focus change | Toast without `aria-live` |
FILE:scripts/a11y_scanner.py
#!/usr/bin/env python3
"""WCAG 2.2 Accessibility Scanner for Frontend Codebases.
Scans HTML, JSX, TSX, Vue, Svelte, and CSS files for accessibility
violations across 10 categories: images, forms, headings, landmarks,
keyboard, ARIA, color/contrast, links, tables, and media.
Usage:
python a11y_scanner.py /path/to/project
python a11y_scanner.py /path/to/project --json
python a11y_scanner.py /path/to/project --severity critical,serious
python a11y_scanner.py /path/to/project --format json
"""
import argparse
import json
import os
import re
import sys
from dataclasses import dataclass, asdict
from typing import List, Optional
@dataclass
class Finding:
"""A single accessibility finding."""
rule_id: str
category: str
severity: str
message: str
file: str
line: int
snippet: str
wcag_criterion: str
fix: str
# ---------------------------------------------------------------------------
# Rule definitions: each returns a list of Finding from a single file
# ---------------------------------------------------------------------------
VALID_ARIA_ATTRS = {
"aria-activedescendant", "aria-atomic", "aria-autocomplete", "aria-busy",
"aria-checked", "aria-colcount", "aria-colindex", "aria-colspan",
"aria-controls", "aria-current", "aria-describedby", "aria-details",
"aria-disabled", "aria-dropeffect", "aria-errormessage", "aria-expanded",
"aria-flowto", "aria-grabbed", "aria-haspopup", "aria-hidden",
"aria-invalid", "aria-keyshortcuts", "aria-label", "aria-labelledby",
"aria-level", "aria-live", "aria-modal", "aria-multiline",
"aria-multiselectable", "aria-orientation", "aria-owns", "aria-placeholder",
"aria-posinset", "aria-pressed", "aria-readonly", "aria-relevant",
"aria-required", "aria-roledescription", "aria-rowcount", "aria-rowindex",
"aria-rowspan", "aria-selected", "aria-setsize", "aria-sort",
"aria-valuemax", "aria-valuemin", "aria-valuenow", "aria-valuetext",
"aria-braillelabel", "aria-brailleroledescription", "aria-description",
}
BAD_LINK_TEXT = re.compile(
r">\s*(click here|here|read more|more|link|this)\s*<", re.IGNORECASE
)
TAG_RE = re.compile(r"<(\w[\w-]*)\b([^>]*)(/?)>", re.DOTALL)
ATTR_RE = re.compile(r"""([\w:.-]+)\s*=\s*(?:"([^"]*)"|'([^']*)'|(\S+))""")
ATTR_BOOL_RE = re.compile(r"\b([\w:.-]+)(?=\s|/?>|$)")
INLINE_COLOR_RE = re.compile(
r'style\s*=\s*["\'][^"\']*\bcolor\s*:', re.IGNORECASE
)
ARIA_ATTR_RE = re.compile(r"\baria-[\w-]+")
def _attrs(attr_str: str) -> dict:
"""Parse HTML/JSX attribute string into a dict."""
result = {}
for m in ATTR_RE.finditer(attr_str):
result[m.group(1)] = m.group(2) or m.group(3) or m.group(4) or ""
# boolean attrs
cleaned = ATTR_RE.sub("", attr_str)
for m in ATTR_BOOL_RE.finditer(cleaned):
name = m.group(1)
if name not in result and not name.startswith("/"):
result[name] = True
return result
def _snippet(line_text: str) -> str:
"""Trim a line for display as a code snippet."""
s = line_text.rstrip("\n\r")
return s[:120] + "..." if len(s) > 120 else s
def _find(rule_id, cat, sev, msg, fp, ln, snip, wcag, fix):
return Finding(rule_id, cat, sev, msg, fp, ln, snip, wcag, fix)
# ---------- Images ----------------------------------------------------------
def check_img_missing_alt(tag, attrs, fp, ln, snip):
if tag == "img" and "alt" not in attrs:
return _find("img-alt-missing", "images", "critical",
"<img> missing alt attribute",
fp, ln, snip, "1.1.1 Non-text Content",
"Add alt=\"description\" or alt=\"\" for decorative images.")
def check_img_empty_alt_informative(tag, attrs, fp, ln, snip):
if tag == "img" and attrs.get("alt") == "" and attrs.get("src", ""):
src = attrs.get("src", "")
if not any(kw in src.lower() for kw in ("spacer", "border", "decorat", "bg")):
return _find("img-alt-empty-informative", "images", "serious",
"<img> has empty alt but may be informative",
fp, ln, snip, "1.1.1 Non-text Content",
"If image conveys information, add descriptive alt text.")
def check_img_decorative_has_alt(tag, attrs, fp, ln, snip):
if tag == "img" and attrs.get("role") == "presentation" and attrs.get("alt", "") != "":
return _find("img-decorative-alt", "images", "moderate",
"Decorative image (role=presentation) should have alt=\"\"",
fp, ln, snip, "1.1.1 Non-text Content",
"Set alt=\"\" on decorative images with role=presentation.")
# ---------- Forms -----------------------------------------------------------
def check_input_missing_label(tag, attrs, fp, ln, snip):
input_types = {"text", "email", "password", "search", "tel", "url", "number", "date"}
if tag == "input" and attrs.get("type", "text") in input_types:
if "aria-label" not in attrs and "aria-labelledby" not in attrs and "id" not in attrs:
return _find("form-input-no-label", "forms", "critical",
"<input> has no id, aria-label, or aria-labelledby",
fp, ln, snip, "1.3.1 Info and Relationships",
"Add id + <label for>, or aria-label attribute.")
def check_input_no_aria_label(tag, attrs, fp, ln, snip):
if tag in ("select", "textarea"):
if "aria-label" not in attrs and "aria-labelledby" not in attrs and "id" not in attrs:
return _find("form-select-no-label", "forms", "critical",
f"<{tag}> has no accessible name",
fp, ln, snip, "4.1.2 Name, Role, Value",
f"Add aria-label or id + <label for> to <{tag}>.")
def check_orphan_label(lines, fp):
"""Labels whose 'for' points to a non-existent id."""
findings = []
ids = set()
label_fors = []
for ln, line in enumerate(lines, 1):
for m in re.finditer(r'\bid\s*=\s*["\']([^"\']+)["\']', line):
ids.add(m.group(1))
for m in re.finditer(r'<label[^>]*\bfor\s*=\s*["\']([^"\']+)["\']', line):
label_fors.append((ln, m.group(1), line))
for ln, for_val, line in label_fors:
if for_val not in ids:
findings.append(_find("form-orphan-label", "forms", "serious",
f"<label for=\"{for_val}\"> references non-existent id",
fp, ln, _snippet(line), "1.3.1 Info and Relationships",
f"Ensure an element with id=\"{for_val}\" exists."))
return findings
def check_fieldset_legend(lines, fp):
"""Radio/checkbox groups without fieldset."""
findings = []
radio_lines = []
has_fieldset = any("fieldset" in l.lower() for l in lines)
for ln, line in enumerate(lines, 1):
if re.search(r'type\s*=\s*["\'](?:radio|checkbox)["\']', line, re.I):
radio_lines.append((ln, line))
if radio_lines and not has_fieldset:
ln, line = radio_lines[0]
findings.append(_find("form-missing-fieldset", "forms", "serious",
"Radio/checkbox group without <fieldset>/<legend>",
fp, ln, _snippet(line), "1.3.1 Info and Relationships",
"Wrap related radio/checkbox inputs in <fieldset> with <legend>."))
return findings
# ---------- Headings --------------------------------------------------------
def check_headings(lines, fp):
findings = []
heading_levels = []
for ln, line in enumerate(lines, 1):
for m in re.finditer(r"<[hH]([1-6])\b", line):
heading_levels.append((int(m.group(1)), ln, line))
if not heading_levels:
return findings
# Missing h1
levels_seen = {h[0] for h in heading_levels}
if 1 not in levels_seen and any(l <= 3 for l in levels_seen):
findings.append(_find("heading-missing-h1", "headings", "serious",
"Page has headings but no <h1>",
fp, heading_levels[0][1], _snippet(heading_levels[0][2]),
"1.3.1 Info and Relationships",
"Add a single <h1> as the main page heading."))
# Multiple h1s
h1_lines = [(ln, line) for lvl, ln, line in heading_levels if lvl == 1]
if len(h1_lines) > 1:
findings.append(_find("heading-multiple-h1", "headings", "moderate",
f"Page has {len(h1_lines)} <h1> elements",
fp, h1_lines[1][0], _snippet(h1_lines[1][1]),
"1.3.1 Info and Relationships",
"Use a single <h1> per page. Demote others to <h2>+."))
# Skipped levels
prev_level = 0
for lvl, ln, line in heading_levels:
if prev_level > 0 and lvl > prev_level + 1:
findings.append(_find("heading-skipped", "headings", "moderate",
f"Heading level skips from h{prev_level} to h{lvl}",
fp, ln, _snippet(line),
"1.3.1 Info and Relationships",
f"Use <h{prev_level + 1}> instead of <h{lvl}>."))
prev_level = lvl
return findings
# ---------- Landmarks -------------------------------------------------------
def check_landmarks(lines, fp):
findings = []
content = "\n".join(lines)
# Missing main landmark
if not re.search(r'<main\b|role\s*=\s*["\']main["\']', content, re.I):
findings.append(_find("landmark-no-main", "landmarks", "serious",
"Page missing <main> landmark",
fp, 1, "", "1.3.1 Info and Relationships",
"Add a <main> element to wrap primary content."))
# Missing nav
if not re.search(r'<nav\b|role\s*=\s*["\']navigation["\']', content, re.I):
findings.append(_find("landmark-no-nav", "landmarks", "moderate",
"Page missing <nav> landmark",
fp, 1, "", "1.3.1 Info and Relationships",
"Add <nav> for primary navigation blocks."))
# Missing skip link
if not re.search(r'skip.{0,10}(nav|main|content)', content, re.I):
findings.append(_find("landmark-no-skip-link", "landmarks", "serious",
"Page missing skip navigation link",
fp, 1, "", "2.4.1 Bypass Blocks",
"Add <a href=\"#main\">Skip to main content</a> as first focusable element."))
return findings
# ---------- Keyboard --------------------------------------------------------
def check_tabindex_positive(tag, attrs, fp, ln, snip):
ti = attrs.get("tabindex", "")
if isinstance(ti, str) and ti.lstrip("-").isdigit() and int(ti) > 0:
return _find("keyboard-tabindex-positive", "keyboard", "serious",
f"tabindex={ti} creates unexpected tab order",
fp, ln, snip, "2.4.3 Focus Order",
"Use tabindex=\"0\" or tabindex=\"-1\" instead of positive values.")
def check_click_no_keyboard(tag, attrs, fp, ln, snip):
has_click = "onClick" in attrs or "onclick" in attrs or "@click" in attrs or "on:click" in attrs
has_key = any(k for k in attrs if "keydown" in k.lower() or "keyup" in k.lower() or "keypress" in k.lower())
if tag in ("div", "span", "td", "li", "p", "section") and has_click and not has_key:
if attrs.get("role") not in ("button", "link", "tab", "menuitem"):
return _find("keyboard-click-no-key", "keyboard", "critical",
f"<{tag}> has click handler but no keyboard handler",
fp, ln, snip, "2.1.1 Keyboard",
f"Add onKeyDown handler or use <button> instead of <{tag}>.")
def check_autofocus_misuse(tag, attrs, fp, ln, snip):
if "autofocus" in attrs or "autoFocus" in attrs:
if tag not in ("input", "textarea", "select"):
return _find("keyboard-autofocus", "keyboard", "moderate",
f"autofocus on <{tag}> can disorient screen reader users",
fp, ln, snip, "3.2.1 On Focus",
"Avoid autofocus on non-input elements. Use focus management instead.")
# ---------- ARIA ------------------------------------------------------------
def check_invalid_aria(tag, attrs, fp, ln, snip):
findings = []
for key in attrs:
if key.startswith("aria-") and key.lower() not in VALID_ARIA_ATTRS:
findings.append(_find("aria-invalid-attr", "aria", "serious",
f"Invalid ARIA attribute: {key}",
fp, ln, snip, "4.1.2 Name, Role, Value",
f"Remove or replace \"{key}\" with a valid ARIA attribute."))
return findings
def check_aria_hidden_focusable(tag, attrs, fp, ln, snip):
if attrs.get("aria-hidden") in ("true", True):
focusable_tags = {"a", "button", "input", "select", "textarea"}
if tag in focusable_tags or (isinstance(attrs.get("tabindex", ""), str) and
attrs.get("tabindex", "-1") != "-1"):
return _find("aria-hidden-focusable", "aria", "critical",
f"aria-hidden=\"true\" on focusable <{tag}>",
fp, ln, snip, "4.1.2 Name, Role, Value",
"Remove aria-hidden or make element non-focusable (tabindex=\"-1\").")
def check_aria_live_missing(lines, fp):
"""Alert/status roles or live regions without aria-live."""
findings = []
for ln, line in enumerate(lines, 1):
if re.search(r'role\s*=\s*["\'](?:alert|status)["\']', line, re.I):
if "aria-live" not in line:
findings.append(_find("aria-live-missing", "aria", "serious",
"role=alert/status without explicit aria-live",
fp, ln, _snippet(line),
"4.1.3 Status Messages",
"Add aria-live=\"assertive\" (alert) or aria-live=\"polite\" (status)."))
return findings
# ---------- Color/Contrast --------------------------------------------------
def check_inline_color(tag, attrs, fp, ln, snip):
style = attrs.get("style", "")
if isinstance(style, str) and re.search(r"\bcolor\s*:", style, re.I):
if not re.search(r"background", style, re.I):
return _find("color-inline-no-bg", "color", "moderate",
"Inline color set without background — contrast may be insufficient",
fp, ln, snip, "1.4.3 Contrast (Minimum)",
"Ensure foreground and background colors meet 4.5:1 contrast ratio.")
def check_text_over_image(lines, fp):
"""Detects patterns where text is positioned over background images without overlay."""
findings = []
for ln, line in enumerate(lines, 1):
if re.search(r"background-image\s*:", line, re.I):
if not re.search(r"(overlay|rgba|linear-gradient)", line, re.I):
findings.append(_find("color-text-over-image", "color", "serious",
"Background image without contrast overlay for text",
fp, ln, _snippet(line),
"1.4.3 Contrast (Minimum)",
"Add a semi-transparent overlay or ensure text contrast."))
return findings
# ---------- Links -----------------------------------------------------------
def check_empty_link(tag, attrs, fp, ln, snip):
if tag == "a" and not attrs.get("aria-label") and not attrs.get("aria-labelledby"):
return None # handled by line-level check below
def check_empty_links_line(lines, fp):
findings = []
for ln, line in enumerate(lines, 1):
# <a ...></a> or <a ...> </a>
if re.search(r"<a\b[^>]*>\s*</a>", line, re.I):
if "aria-label" not in line and "aria-labelledby" not in line:
findings.append(_find("link-empty", "links", "critical",
"Empty link — no text or accessible name",
fp, ln, _snippet(line), "2.4.4 Link Purpose",
"Add link text or aria-label."))
# Bad link text
if BAD_LINK_TEXT.search(line):
findings.append(_find("link-bad-text", "links", "serious",
"Link uses vague text like 'click here'",
fp, ln, _snippet(line), "2.4.4 Link Purpose",
"Use descriptive link text that makes sense out of context."))
return findings
def check_same_page_link(tag, attrs, fp, ln, snip):
href = attrs.get("href", "")
if tag == "a" and isinstance(href, str) and href == "#":
return _find("link-empty-fragment", "links", "moderate",
"Link with href=\"#\" — use a button or valid fragment",
fp, ln, snip, "2.4.4 Link Purpose",
"Use <button> for actions or href=\"#section-id\" for anchors.")
# ---------- Tables ----------------------------------------------------------
def check_table_headers(lines, fp):
findings = []
in_table = False
table_start = 0
has_th = False
has_caption = False
has_aria_label = False
for ln, line in enumerate(lines, 1):
if re.search(r"<table\b", line, re.I):
in_table = True
table_start = ln
has_th = False
has_caption = False
has_aria_label = "aria-label" in line
if in_table:
if "<th" in line.lower():
has_th = True
if "<caption" in line.lower():
has_caption = True
if re.search(r"</table>", line, re.I):
if not has_th:
findings.append(_find("table-no-headers", "tables", "serious",
"<table> has no <th> header cells",
fp, table_start, _snippet(lines[table_start - 1]),
"1.3.1 Info and Relationships",
"Add <th> elements to identify column/row headers."))
if not has_caption and not has_aria_label:
findings.append(_find("table-no-caption", "tables", "moderate",
"<table> missing <caption> or aria-label",
fp, table_start, _snippet(lines[table_start - 1]),
"1.3.1 Info and Relationships",
"Add <caption> or aria-label to describe the table."))
in_table = False
return findings
# ---------- Media -----------------------------------------------------------
def check_media_captions(tag, attrs, fp, ln, snip):
if tag == "video":
return None # handled at block level
def check_media_captions_block(lines, fp):
findings = []
in_video = False
video_start = 0
has_track = False
has_controls = False
has_autoplay = False
for ln, line in enumerate(lines, 1):
if re.search(r"<video\b", line, re.I):
in_video = True
video_start = ln
has_track = False
has_controls = "controls" in line.lower()
has_autoplay = "autoplay" in line.lower()
if in_video:
if re.search(r'<track\b[^>]*kind\s*=\s*["\']captions["\']', line, re.I):
has_track = True
if "controls" in line.lower():
has_controls = True
if re.search(r"</video>", line, re.I) or (re.search(r"<video\b", line, re.I) and "/>" in line):
if not has_track:
findings.append(_find("media-no-captions", "media", "critical",
"<video> missing captions track",
fp, video_start, _snippet(lines[video_start - 1]),
"1.2.2 Captions (Prerecorded)",
"Add <track kind=\"captions\" src=\"...\" srclang=\"en\">."))
if has_autoplay and not has_controls:
findings.append(_find("media-autoplay-no-controls", "media", "serious",
"<video> has autoplay without controls",
fp, video_start, _snippet(lines[video_start - 1]),
"1.4.2 Audio Control",
"Add the controls attribute so users can pause/stop."))
in_video = False
# Single-line video tags
for ln, line in enumerate(lines, 1):
if re.search(r"<audio\b", line, re.I):
if "autoplay" in line.lower() and "controls" not in line.lower():
findings.append(_find("media-audio-autoplay", "media", "serious",
"<audio> has autoplay without controls",
fp, ln, _snippet(line), "1.4.2 Audio Control",
"Add the controls attribute to <audio>."))
return findings
# ---------------------------------------------------------------------------
# Scanner engine
# ---------------------------------------------------------------------------
SUPPORTED_EXTENSIONS = {".html", ".htm", ".jsx", ".tsx", ".vue", ".svelte", ".css"}
TAG_LEVEL_CHECKS = [
check_img_missing_alt,
check_img_empty_alt_informative,
check_img_decorative_has_alt,
check_input_missing_label,
check_input_no_aria_label,
check_tabindex_positive,
check_click_no_keyboard,
check_autofocus_misuse,
check_aria_hidden_focusable,
check_inline_color,
check_same_page_link,
]
TAG_LEVEL_MULTI_CHECKS = [
check_invalid_aria,
]
def scan_file(filepath: str) -> List[Finding]:
"""Scan a single file and return all findings."""
findings: List[Finding] = []
try:
with open(filepath, "r", encoding="utf-8", errors="replace") as f:
lines = f.readlines()
except (OSError, IOError):
return findings
# Tag-level checks
for ln, line in enumerate(lines, 1):
for m in TAG_RE.finditer(line):
tag = m.group(1).lower()
attr_str = m.group(2)
attrs = _attrs(attr_str)
snip = _snippet(line)
for check in TAG_LEVEL_CHECKS:
result = check(tag, attrs, filepath, ln, snip)
if result:
findings.append(result)
for check in TAG_LEVEL_MULTI_CHECKS:
results = check(tag, attrs, filepath, ln, snip)
if results:
findings.extend(results)
# File-level / multi-line checks
findings.extend(check_orphan_label(lines, filepath))
findings.extend(check_fieldset_legend(lines, filepath))
findings.extend(check_headings(lines, filepath))
findings.extend(check_landmarks(lines, filepath))
findings.extend(check_aria_live_missing(lines, filepath))
findings.extend(check_text_over_image(lines, filepath))
findings.extend(check_empty_links_line(lines, filepath))
findings.extend(check_table_headers(lines, filepath))
findings.extend(check_media_captions_block(lines, filepath))
return findings
def collect_files(path: str) -> List[str]:
"""Recursively collect scannable files under path."""
files = []
if os.path.isfile(path):
_, ext = os.path.splitext(path)
if ext.lower() in SUPPORTED_EXTENSIONS:
files.append(path)
return files
for root, dirs, filenames in os.walk(path):
# Skip common non-source directories
dirs[:] = [d for d in dirs if d not in (
"node_modules", ".git", "dist", "build", "__pycache__",
".next", ".nuxt", "vendor", "coverage"
)]
for fname in filenames:
_, ext = os.path.splitext(fname)
if ext.lower() in SUPPORTED_EXTENSIONS:
files.append(os.path.join(root, fname))
files.sort()
return files
# ---------------------------------------------------------------------------
# Output formatting
# ---------------------------------------------------------------------------
SEVERITY_ORDER = {"critical": 0, "serious": 1, "moderate": 2, "minor": 3}
def format_human(findings: List[Finding], files_scanned: int) -> str:
"""Format findings as human-readable text report."""
if not findings:
return (f"Scanned {files_scanned} file(s) -- no accessibility issues found.\n"
"All checks passed.")
lines = []
lines.append(f"WCAG 2.2 Accessibility Scan Results")
lines.append(f"{'=' * 50}")
lines.append(f"Files scanned: {files_scanned}")
lines.append(f"Issues found: {len(findings)}")
# Summary by severity
severity_counts = {}
for f in findings:
severity_counts[f.severity] = severity_counts.get(f.severity, 0) + 1
for sev in ("critical", "serious", "moderate", "minor"):
if sev in severity_counts:
lines.append(f" {sev.upper():10s}: {severity_counts[sev]}")
lines.append("")
# Summary by category
cat_counts = {}
for f in findings:
cat_counts[f.category] = cat_counts.get(f.category, 0) + 1
lines.append("By category:")
for cat in sorted(cat_counts, key=lambda c: -cat_counts[c]):
lines.append(f" {cat:20s}: {cat_counts[cat]}")
lines.append("")
# Detailed findings sorted by severity then file
sorted_findings = sorted(findings, key=lambda f: (SEVERITY_ORDER.get(f.severity, 9), f.file, f.line))
for i, f in enumerate(sorted_findings, 1):
lines.append(f"[{f.severity.upper()}] {f.rule_id}")
lines.append(f" File: {f.file}:{f.line}")
lines.append(f" WCAG: {f.wcag_criterion}")
lines.append(f" Issue: {f.message}")
if f.snippet:
lines.append(f" Code: {f.snippet}")
lines.append(f" Fix: {f.fix}")
lines.append("")
return "\n".join(lines)
def format_json(findings: List[Finding], files_scanned: int) -> str:
"""Format findings as JSON."""
severity_counts = {}
for f in findings:
severity_counts[f.severity] = severity_counts.get(f.severity, 0) + 1
report = {
"summary": {
"files_scanned": files_scanned,
"total_issues": len(findings),
"by_severity": severity_counts,
},
"findings": [asdict(f) for f in findings],
}
return json.dumps(report, indent=2)
# ---------------------------------------------------------------------------
# CLI
# ---------------------------------------------------------------------------
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
prog="a11y_scanner",
description="Scan frontend codebases for WCAG 2.2 accessibility violations.",
epilog=(
"Supported file types: .html, .htm, .jsx, .tsx, .vue, .svelte, .css\n"
"Exit codes: 0 = pass, 1 = critical/serious found, 2 = moderate/minor only"
),
formatter_class=argparse.RawDescriptionHelpFormatter,
)
parser.add_argument(
"path",
help="File or directory to scan",
)
parser.add_argument(
"--json", dest="json_flag", action="store_true",
help="Output results as JSON (shorthand for --format json)",
)
parser.add_argument(
"--format", dest="output_format", choices=["text", "json"],
default="text",
help="Output format: text (default) or json",
)
parser.add_argument(
"--severity", dest="severity",
default=None,
help="Comma-separated severity filter (e.g. critical,serious)",
)
return parser
def main():
parser = build_parser()
args = parser.parse_args()
path = os.path.abspath(args.path)
if not os.path.exists(path):
print(f"Error: path does not exist: {path}", file=sys.stderr)
sys.exit(1)
use_json = args.json_flag or args.output_format == "json"
# Collect and scan files
files = collect_files(path)
if not files:
print(f"No scannable files found in: {path}", file=sys.stderr)
sys.exit(0)
all_findings: List[Finding] = []
for fpath in files:
all_findings.extend(scan_file(fpath))
# Filter by severity if requested
if args.severity:
allowed = {s.strip().lower() for s in args.severity.split(",")}
all_findings = [f for f in all_findings if f.severity in allowed]
# Output
if use_json:
print(format_json(all_findings, len(files)))
else:
print(format_human(all_findings, len(files)))
# Exit code
severities = {f.severity for f in all_findings}
if severities & {"critical", "serious"}:
sys.exit(1)
elif severities & {"moderate", "minor"}:
sys.exit(2)
else:
sys.exit(0)
if __name__ == "__main__":
main()
FILE:scripts/contrast_checker.py
#!/usr/bin/env python3
"""WCAG 2.2 Color Contrast Checker.
Checks foreground/background color pairs against WCAG 2.2 contrast ratio
thresholds for normal text, large text, and UI components. Supports hex,
rgb(), and named CSS colors.
Usage:
python contrast_checker.py "#ffffff" "#000000"
python contrast_checker.py --suggest "#336699"
python contrast_checker.py --batch styles.css
python contrast_checker.py --demo
"""
import argparse
import json
import re
import sys
# ---------------------------------------------------------------------------
# Named CSS colors (25 common ones)
# ---------------------------------------------------------------------------
NAMED_COLORS = {
"black": (0, 0, 0),
"white": (255, 255, 255),
"red": (255, 0, 0),
"green": (0, 128, 0),
"blue": (0, 0, 255),
"yellow": (255, 255, 0),
"cyan": (0, 255, 255),
"magenta": (255, 0, 255),
"gray": (128, 128, 128),
"grey": (128, 128, 128),
"orange": (255, 165, 0),
"purple": (128, 0, 128),
"pink": (255, 192, 203),
"brown": (165, 42, 42),
"navy": (0, 0, 128),
"teal": (0, 128, 128),
"olive": (128, 128, 0),
"maroon": (128, 0, 0),
"lime": (0, 255, 0),
"aqua": (0, 255, 255),
"silver": (192, 192, 192),
"gold": (255, 215, 0),
"coral": (255, 127, 80),
"salmon": (250, 128, 114),
"tomato": (255, 99, 71),
}
# WCAG thresholds: (label, required_ratio)
WCAG_THRESHOLDS = [
("AA Normal Text", 4.5),
("AA Large Text", 3.0),
("AA UI Components", 3.0),
("AAA Normal Text", 7.0),
("AAA Large Text", 4.5),
]
# ---------------------------------------------------------------------------
# Color parsing
# ---------------------------------------------------------------------------
def parse_color(color_str: str) -> tuple:
"""Parse a color string into an (R, G, B) tuple.
Accepts:
- #RRGGBB or #RGB hex
- rgb(r, g, b) with values 0-255
- Named CSS colors
"""
s = color_str.strip().lower()
# Named color
if s in NAMED_COLORS:
return NAMED_COLORS[s]
# Hex: #RGB or #RRGGBB
hex_match = re.match(r"^#([0-9a-f]{3}|[0-9a-f]{6})$", s)
if hex_match:
h = hex_match.group(1)
if len(h) == 3:
r, g, b = int(h[0] * 2, 16), int(h[1] * 2, 16), int(h[2] * 2, 16)
else:
r, g, b = int(h[0:2], 16), int(h[2:4], 16), int(h[4:6], 16)
return (r, g, b)
# rgb(r, g, b)
rgb_match = re.match(r"^rgb\(\s*(\d{1,3})\s*,\s*(\d{1,3})\s*,\s*(\d{1,3})\s*\)$", s)
if rgb_match:
r, g, b = int(rgb_match.group(1)), int(rgb_match.group(2)), int(rgb_match.group(3))
if not all(0 <= c <= 255 for c in (r, g, b)):
raise ValueError(f"RGB values must be 0-255, got rgb({r},{g},{b})")
return (r, g, b)
raise ValueError(
f"Invalid color format: '{color_str}'. "
"Use #RRGGBB, #RGB, rgb(r,g,b), or a named color."
)
def color_to_hex(rgb: tuple) -> str:
"""Convert an (R, G, B) tuple to #RRGGBB."""
return f"#{rgb[0]:02x}{rgb[1]:02x}{rgb[2]:02x}"
# ---------------------------------------------------------------------------
# WCAG luminance and contrast
# ---------------------------------------------------------------------------
def relative_luminance(rgb: tuple) -> float:
"""Calculate relative luminance per WCAG 2.2 (sRGB).
https://www.w3.org/TR/WCAG22/#dfn-relative-luminance
"""
channels = []
for c in rgb:
s = c / 255.0
channels.append(s / 12.92 if s <= 0.04045 else ((s + 0.055) / 1.055) ** 2.4)
return 0.2126 * channels[0] + 0.7152 * channels[1] + 0.0722 * channels[2]
def contrast_ratio(rgb1: tuple, rgb2: tuple) -> float:
"""Return the WCAG contrast ratio between two colors (>= 1.0)."""
l1 = relative_luminance(rgb1)
l2 = relative_luminance(rgb2)
lighter = max(l1, l2)
darker = min(l1, l2)
return (lighter + 0.05) / (darker + 0.05)
def evaluate_contrast(ratio: float) -> list:
"""Return pass/fail results for each WCAG threshold."""
results = []
for label, threshold in WCAG_THRESHOLDS:
results.append({
"level": label,
"required": threshold,
"ratio": round(ratio, 2),
"pass": ratio >= threshold,
})
return results
# ---------------------------------------------------------------------------
# Suggest accessible backgrounds
# ---------------------------------------------------------------------------
def suggest_backgrounds(fg_rgb: tuple, target_ratio: float = 4.5, count: int = 8) -> list:
"""Given a foreground color, suggest background colors passing AA normal text.
Strategy: walk luminance in both directions (lighter / darker) from the
foreground and collect the first colors that meet the target ratio.
"""
suggestions = []
# Try a spread of grays and tinted variants
candidates = []
for v in range(0, 256, 1):
candidates.append((v, v, v)) # grays
# Also try tinted versions toward the complement
fr, fg, fb = fg_rgb
for v in range(0, 256, 2):
candidates.append((v, min(255, v + 20), min(255, v + 40)))
candidates.append((min(255, v + 40), v, min(255, v + 20)))
candidates.append((min(255, v + 20), min(255, v + 40), v))
seen = set()
scored = []
for c in candidates:
cr = contrast_ratio(fg_rgb, c)
if cr >= target_ratio and c not in seen:
seen.add(c)
scored.append((cr, c))
# Sort by ratio closest to target (prefer minimal-change backgrounds)
scored.sort(key=lambda x: x[0])
for cr, c in scored[:count]:
suggestions.append({"hex": color_to_hex(c), "rgb": list(c), "ratio": round(cr, 2)})
return suggestions
# ---------------------------------------------------------------------------
# Batch CSS parsing
# ---------------------------------------------------------------------------
_COLOR_RE = re.compile(
r"(#[0-9a-fA-F]{3,6}|rgb\(\s*\d{1,3}\s*,\s*\d{1,3}\s*,\s*\d{1,3}\s*\))"
)
def extract_css_pairs(css_text: str) -> list:
"""Extract color / background-color pairs from CSS declarations.
Returns a list of dicts with selector, foreground, and background strings.
"""
pairs = []
# Split into rule blocks
block_re = re.compile(r"([^{}]+)\{([^}]+)\}", re.DOTALL)
for m in block_re.finditer(css_text):
selector = m.group(1).strip()
body = m.group(2)
fg = bg = None
# Match color: ... (but not background-color)
fg_match = re.search(
r"(?<![-])color\s*:\s*([^;]+);", body, re.IGNORECASE
)
bg_match = re.search(
r"background(?:-color)?\s*:\s*([^;]+);", body, re.IGNORECASE
)
if fg_match:
val = fg_match.group(1).strip()
c = _COLOR_RE.search(val)
if c:
fg = c.group(1)
elif val.lower() in NAMED_COLORS:
fg = val.lower()
if bg_match:
val = bg_match.group(1).strip()
c = _COLOR_RE.search(val)
if c:
bg = c.group(1)
elif val.lower() in NAMED_COLORS:
bg = val.lower()
if fg and bg:
pairs.append({"selector": selector, "foreground": fg, "background": bg})
return pairs
# ---------------------------------------------------------------------------
# Output formatting
# ---------------------------------------------------------------------------
def format_result_human(fg_str: str, bg_str: str, ratio: float, results: list) -> str:
"""Format a contrast check result for the terminal."""
lines = [
f"Foreground : {fg_str}",
f"Background : {bg_str}",
f"Contrast : {ratio:.2f}:1",
"",
]
for r in results:
status = "PASS" if r["pass"] else "FAIL"
lines.append(f" [{status}] {r['level']:20s} (requires {r['required']}:1)")
return "\n".join(lines)
def format_suggestions_human(fg_str: str, suggestions: list) -> str:
"""Format suggested backgrounds for the terminal."""
lines = [f"Foreground: {fg_str}", "Suggested accessible backgrounds (AA Normal Text):"]
if not suggestions:
lines.append(" No suggestions found.")
for s in suggestions:
lines.append(f" {s['hex']} ratio={s['ratio']}:1")
return "\n".join(lines)
# ---------------------------------------------------------------------------
# Demo
# ---------------------------------------------------------------------------
DEMO_PAIRS = [
("#ffffff", "#000000"),
("#336699", "#ffffff"),
("#ff6600", "#ffffff"),
("navy", "white"),
("rgb(100,100,100)", "#eeeeee"),
]
def run_demo(as_json: bool) -> None:
"""Run demo checks and print results."""
all_results = []
for fg_str, bg_str in DEMO_PAIRS:
fg_rgb = parse_color(fg_str)
bg_rgb = parse_color(bg_str)
ratio = contrast_ratio(fg_rgb, bg_rgb)
results = evaluate_contrast(ratio)
entry = {
"foreground": fg_str,
"background": bg_str,
"foreground_hex": color_to_hex(fg_rgb),
"background_hex": color_to_hex(bg_rgb),
"ratio": round(ratio, 2),
"results": results,
}
all_results.append(entry)
if as_json:
print(json.dumps({"demo": True, "checks": all_results}, indent=2))
else:
print("=" * 60)
print("WCAG 2.2 Contrast Checker - Demo")
print("=" * 60)
for entry in all_results:
print()
print(
format_result_human(
entry["foreground"], entry["background"],
entry["ratio"], entry["results"],
)
)
print()
print("-" * 60)
print("Suggestion demo for foreground #336699:")
suggestions = suggest_backgrounds(parse_color("#336699"))
print(format_suggestions_human("#336699", suggestions))
# ---------------------------------------------------------------------------
# CLI
# ---------------------------------------------------------------------------
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
description="WCAG 2.2 Color Contrast Checker. "
"Checks foreground/background pairs against AA and AAA thresholds.",
epilog="Examples:\n"
" %(prog)s '#ffffff' '#000000'\n"
" %(prog)s --suggest '#336699'\n"
" %(prog)s --batch styles.css\n"
" %(prog)s --demo\n",
formatter_class=argparse.RawDescriptionHelpFormatter,
)
parser.add_argument(
"foreground",
nargs="?",
help="Foreground (text) color: #RRGGBB, #RGB, rgb(r,g,b), or named color",
)
parser.add_argument(
"background",
nargs="?",
help="Background color: #RRGGBB, #RGB, rgb(r,g,b), or named color",
)
parser.add_argument(
"--suggest",
metavar="COLOR",
help="Suggest accessible background colors for the given foreground color",
)
parser.add_argument(
"--batch",
metavar="CSS_FILE",
help="Extract color pairs from a CSS file and check each",
)
parser.add_argument(
"--json",
action="store_true",
dest="json_output",
help="Output results as JSON",
)
parser.add_argument(
"--demo",
action="store_true",
help="Show example output with sample color pairs",
)
return parser
def main() -> int:
parser = build_parser()
args = parser.parse_args()
# --demo mode
if args.demo:
run_demo(args.json_output)
return 0
# --suggest mode
if args.suggest:
try:
fg_rgb = parse_color(args.suggest)
except ValueError as exc:
print(f"Error: {exc}", file=sys.stderr)
return 1
suggestions = suggest_backgrounds(fg_rgb)
if args.json_output:
print(json.dumps({
"foreground": args.suggest,
"foreground_hex": color_to_hex(fg_rgb),
"suggestions": suggestions,
}, indent=2))
else:
print(format_suggestions_human(args.suggest, suggestions))
return 0
# --batch mode
if args.batch:
try:
with open(args.batch, "r", encoding="utf-8") as fh:
css_text = fh.read()
except FileNotFoundError:
print(f"Error: file not found: {args.batch}", file=sys.stderr)
return 1
except OSError as exc:
print(f"Error reading file: {exc}", file=sys.stderr)
return 1
pairs = extract_css_pairs(css_text)
if not pairs:
msg = "No color/background-color pairs found in the CSS file."
if args.json_output:
print(json.dumps({"batch": args.batch, "pairs": [], "message": msg}, indent=2))
else:
print(msg)
return 0
all_results = []
has_failure = False
for pair in pairs:
try:
fg_rgb = parse_color(pair["foreground"])
bg_rgb = parse_color(pair["background"])
except ValueError as exc:
entry = {
"selector": pair["selector"],
"foreground": pair["foreground"],
"background": pair["background"],
"error": str(exc),
}
all_results.append(entry)
continue
ratio = contrast_ratio(fg_rgb, bg_rgb)
results = evaluate_contrast(ratio)
if not results[0]["pass"]: # AA Normal Text
has_failure = True
entry = {
"selector": pair["selector"],
"foreground": pair["foreground"],
"background": pair["background"],
"foreground_hex": color_to_hex(fg_rgb),
"background_hex": color_to_hex(bg_rgb),
"ratio": round(ratio, 2),
"results": results,
}
all_results.append(entry)
if args.json_output:
print(json.dumps({"batch": args.batch, "pairs": all_results}, indent=2))
else:
print(f"Batch check: {args.batch}")
print("=" * 60)
for entry in all_results:
print(f"\nSelector: {entry['selector']}")
if "error" in entry:
print(f" Error: {entry['error']}")
else:
print(
format_result_human(
entry["foreground"], entry["background"],
entry["ratio"], entry["results"],
)
)
print()
summary_pass = sum(1 for e in all_results if "ratio" in e and e["results"][0]["pass"])
summary_total = sum(1 for e in all_results if "ratio" in e)
print(f"Summary: {summary_pass}/{summary_total} pairs pass AA Normal Text")
return 1 if has_failure else 0
# Default: check a single pair
if not args.foreground or not args.background:
parser.error(
"Provide foreground and background colors, or use --suggest, --batch, or --demo."
)
try:
fg_rgb = parse_color(args.foreground)
except ValueError as exc:
print(f"Error (foreground): {exc}", file=sys.stderr)
return 1
try:
bg_rgb = parse_color(args.background)
except ValueError as exc:
print(f"Error (background): {exc}", file=sys.stderr)
return 1
ratio = contrast_ratio(fg_rgb, bg_rgb)
results = evaluate_contrast(ratio)
if args.json_output:
print(json.dumps({
"foreground": args.foreground,
"background": args.background,
"foreground_hex": color_to_hex(fg_rgb),
"background_hex": color_to_hex(bg_rgb),
"ratio": round(ratio, 2),
"results": results,
}, indent=2))
else:
print(format_result_human(args.foreground, args.background, ratio, results))
return 0 if results[0]["pass"] else 1
if __name__ == "__main__":
sys.exit(main())
Phỏng vấn 6 câu hỏi để đánh giá nội bộ hệ thống quản lý AI theo ISO/IEC 42001 trước chứng nhận hoặc kiểm toán.
--- name: "aims-audit" description: "/cs:aims-audit <scope> — ISO/IEC 42001 AIMS internal-audit 6-question forcing interrogation. Use before certification stage 1, before annual internal audit cycles, or when onboarding a new AI system into an existing AIMS." --- # /cs:aims-audit — AIMS ISO 42001 Forcing Questions **Command:** `/cs:aims-audit <scope>` The ISO 42001 AIMS specialist pressure-tests any AI Management System work. Six questions before any certification commitment, internal audit cycle, or new-system onboarding. ## When to Run - Before stage 1 ISO 42001 certification audit - Before annual internal audit cycle (Clause 9.2) - When onboarding a new AI system into existing AIMS scope - When AI risk register hasn't been refreshed in > 6 months - After material model change (re-evaluate risks per Clause 6.1.2) - When audit findings hint at AIMS / ISMS / QMS duplication ## The Six AIMS Questions ### 1. Does the AIMS scope statement name every AI system? **Scope omission = certification finding.** - Including: embedded models, third-party AI services, "experimental" production systems - Run `aims_gap_analyzer.py` to verify Clause 4.3 evidence - "AI features added by SaaS vendors we use" = in scope if they affect the company's services ### 2. Does the AI policy commit to lawful use AND beneficial purpose AND human oversight AND continual improvement? **Missing any of the four = critical nonconformity at stage 1.** - AI policy is NOT info-sec policy — it has separate substantive content - Reference ISO 42001 Annex A.2.2 + Clause 5.2 - Marketing-copy "AI ethics" doesn't pass ### 3. What's the risk register coverage, and which Annex A controls treat each risk? **Risk identification without control mapping = Clause 6.1.3 fails.** - Run `ai_risk_register_builder.py` per ISO 23894 methodology - Every high/critical risk must link to ≥ 1 Annex A control - "Residual verdict: additional_treatment_required" must be closed before stage 1 ### 4. Has the AI risk assessment been re-run since the last material model change? **Concept drift is not a one-time event.** - Article 9 EU AI Act + ISO 42001 Clause 6.1.2 both require iterative risk assessment - Material change = retraining on new data, fine-tuning, architecture change, deployment context change - If "we did it 18 months ago and haven't touched it," the AIMS is broken ### 5. What's the Clause 9.2 internal audit plan, and is auditor independence respected? **Without 9.2 plan, the AIMS is incomplete.** - Run `aims_audit_scheduler.py` with scope + auditors + prior findings - Audit every clause + applicable Annex A control over rolling 3-year cycle - Same auditor cannot audit own work - Cross-check with cs-quality-regulatory if integrated with 13485 audit programme ### 6. Has the AIMS been integrated with existing ISMS / QMS, or built in parallel? **Parallel systems = 5x ongoing maintenance cost.** - 60% of Clauses 4-10 evidence reuses ISO 27001 / 13485 with AI scope appended - CAPA loop should be ONE loop with AI-tagged nonconformities, not separate - Reference `cross_framework_mapping_ai.md` for the reuse map - Cross-check with cs-ciso-advisor on ISO 27001 alignment ## Workflow ```bash # 1. AIMS gap analysis python ../../ra-qm-team/skills/iso42001-specialist/scripts/aims_gap_analyzer.py evidence.json # 2. AI risk register python ../../ra-qm-team/skills/iso42001-specialist/scripts/ai_risk_register_builder.py risks.json # 3. Internal audit plan python ../../ra-qm-team/skills/iso42001-specialist/scripts/aims_audit_scheduler.py audit_scope.json # 4. Cross-framework reuse map (via compliance-os) python ../../skills/compliance-os/scripts/cross_framework_mapper.py program.json ``` ## Output Format ```markdown # AIMS Audit: <scope> **Date:** YYYY-MM-DD ## The Decision Being Made [gap-closure | risk-treatment | audit-scope | new-system-onboarding] ## Gap Analysis (Clauses 4-10) - Weighted coverage: X% - Critical gaps: N - Major gaps: M - Certification readiness: ready | stage_2_candidate | not_ready ## AI Risk Register - Total risks: N - By severity: critical=X, high=Y, medium=Z, low=W - Requires additional treatment: K - Top risk requiring action: <description> ## Clause 9.2 Audit Plan - 12-month coverage: clauses=X, controls=Y - Auditor independence: clean | issues - Prior-year follow-up: scheduled in Q1 ## Cross-Framework Reuse - ISO 27001 evidence reused: % of AIMS Clauses 4-10 - 13485 evidence reused: % (if applicable) - Net-new for AIMS: % (mostly Annex A) ## Verdict 🟢 STAGE-1-READY | 🟡 CLOSE-CRITICALS-FIRST | 🔴 NOT-READY ## Top 3 Actions [3 concrete next steps with owner + date] ``` ## Routing - `/cs:compliance-readiness` — for multi-framework view - `/cs:ai-act-readiness` — if EU AI Act also applies - `/cs:caio-review` — for executive AI strategy decisions - `/cs:ciso-review` — for ISO 27001 cross-framework alignment - `/cs:decide` — to log the verdict - `/cs:freeze 30` — on certification commitments ## Related - Agent: [`cs-aims-iso42001`](../../agents/cs-aims-iso42001.md) - Skill: [`iso42001-specialist`](../../../ra-qm-team/skills/iso42001-specialist/SKILL.md) - Adjacent: `../../skills/compliance-os/`, `../ai-act-readiness/`, `../compliance-readiness/` --- **Version:** 1.0.0
Quản trị Jira, Confluence, Bitbucket, Trello: người dùng, phân quyền, bảo mật, tích hợp và cấu hình hệ thống.
---
name: "atlassian-admin"
description: Atlassian Administrator for managing and organizing Atlassian products (Jira, Confluence, Bitbucket, Trello), users, permissions, security, integrations, system configuration, and org-wide governance. Use when asked to add users to Jira, change Confluence permissions, configure access control, update admin settings, manage Atlassian groups, set up SSO, install marketplace apps, review security policies, or handle any org-wide Atlassian administration task.
---
# Atlassian Administrator Expert
## Workflows
### User Provisioning
1. Create user account: `admin.atlassian.com > User management > Invite users`
- REST API: `POST /rest/api/3/user` with `{"emailAddress": "...", "displayName": "...","products": [...]}`
2. Add to appropriate groups: `admin.atlassian.com > User management > Groups > [group] > Add members`
3. Assign product access (Jira, Confluence) via `admin.atlassian.com > Products > [product] > Access`
4. Configure default permissions per group scheme
5. Send welcome email with onboarding info
6. **NOTIFY**: Relevant team leads of new member
7. **VERIFY**: Confirm user appears active at `admin.atlassian.com/o/{orgId}/users` and can log in
### User Deprovisioning
1. **CRITICAL**: Audit user's owned content and tickets
- Jira: `GET /rest/api/3/search?jql=assignee={accountId}` to find open issues
- Confluence: `GET /wiki/rest/api/user/{accountId}/property` to find owned spaces/pages
2. Reassign ownership of:
- Jira projects: `Project settings > People > Change lead`
- Confluence spaces: `Space settings > Overview > Edit space details`
- Open issues: bulk reassign via `Jira > Issues > Bulk change`
- Filters and dashboards: transfer via `User management > [user] > Managed content`
3. Remove from all groups: `admin.atlassian.com > User management > [user] > Groups`
4. Revoke product access
5. Deactivate account: `admin.atlassian.com > User management > [user] > Deactivate`
- REST API: `DELETE /rest/api/3/user?accountId={accountId}`
6. **VERIFY**: Confirm `GET /rest/api/3/user?accountId={accountId}` returns `"active": false`
7. Document deprovisioning in audit log
8. **USE**: Jira Expert to reassign any remaining issues
### Group Management
1. Create groups: `admin.atlassian.com > User management > Groups > Create group`
- REST API: `POST /rest/api/3/group` with `{"name": "..."}`
- Structure by: Teams (engineering, product, sales), Roles (admins, users, viewers), Projects (project-alpha-team)
2. Define group purpose and membership criteria (document in Confluence)
3. Assign default permissions per group
4. Add users to appropriate groups
5. **VERIFY**: Confirm group members via `GET /rest/api/3/group/member?groupName={name}`
6. Regular review and cleanup (quarterly)
7. **USE**: Confluence Expert to document group structure
### Permission Scheme Design
**Jira Permission Schemes** (`Jira Settings > Issues > Permission Schemes`):
- **Public Project**: All users can view, members can edit
- **Team Project**: Team members full access, stakeholders view
- **Restricted Project**: Named individuals only
- **Admin Project**: Admins only
**Confluence Permission Schemes** (`Confluence Admin > Space permissions`):
- **Public Space**: All users view, space members edit
- **Team Space**: Team-specific access
- **Personal Space**: Individual user only
- **Restricted Space**: Named individuals and groups
**Best Practices**:
- Use groups, not individual permissions
- Principle of least privilege
- Regular permission audits
- Document permission rationale
### SSO Configuration
1. Choose identity provider (Okta, Azure AD, Google)
2. Configure SAML settings: `admin.atlassian.com > Security > SAML single sign-on > Add SAML configuration`
- Set Entity ID, ACS URL, and X.509 certificate from IdP
3. Test SSO with admin account (keep password login active during test)
4. Test with regular user account
5. Enable SSO for organization
6. Enforce SSO: `admin.atlassian.com > Security > Authentication policies > Enforce SSO`
7. Configure SCIM for auto-provisioning: `admin.atlassian.com > User provisioning > [IdP] > Enable SCIM`
8. **VERIFY**: Confirm SSO flow succeeds and audit logs show `saml.login.success` events
9. Monitor SSO logs: `admin.atlassian.com > Security > Audit log > filter: SSO`
### Marketplace App Management
1. Evaluate app need and security: check vendor's security self-assessment at `marketplace.atlassian.com`
2. Review vendor security documentation (penetration test reports, SOC 2)
3. Test app in sandbox environment
4. Purchase or request trial: `admin.atlassian.com > Billing > Manage subscriptions`
5. Install app: `admin.atlassian.com > Products > [product] > Apps > Find new apps`
6. Configure app settings per vendor documentation
7. Train users on app usage
8. **VERIFY**: Confirm app appears in `GET /rest/plugins/1.0/` and health check passes
9. Monitor app performance and usage; review annually for continued need
### System Performance Optimization
**Jira** (`Jira Settings > System`):
- Archive old projects: `Project settings > Archive project`
- Reindex: `Jira Settings > System > Indexing > Full re-index`
- Clean up unused workflows and schemes: `Jira Settings > Issues > Workflows`
- Monitor queue/thread counts: `Jira Settings > System > System info`
**Confluence** (`Confluence Admin > Configuration`):
- Archive inactive spaces: `Space tools > Overview > Archive space`
- Remove orphaned pages: `Confluence Admin > Orphaned pages`
- Monitor index and cache: `Confluence Admin > Cache management`
**Monitoring Cadence**:
- Daily health checks: `admin.atlassian.com > Products > [product] > Health`
- Weekly performance reports
- Monthly capacity planning
- Quarterly optimization reviews
### Integration Setup
**Common Integrations**:
- **Slack**: `Jira Settings > Apps > Slack integration` — notifications for Jira and Confluence
- **GitHub/Bitbucket**: `Jira Settings > Apps > DVCS accounts` — link commits to issues
- **Microsoft Teams**: `admin.atlassian.com > Apps > Microsoft Teams`
- **Zoom**: Available via Marketplace app `zoom-for-jira`
- **Salesforce**: Via Marketplace app `salesforce-connector`
**Configuration Steps**:
1. Review integration requirements and OAuth scopes needed
2. Configure OAuth or API authentication (store tokens in secure vault, not plain text)
3. Map fields and data flows
4. Test integration thoroughly with sample data
5. Document configuration in Confluence runbook
6. Train users on integration features
7. **VERIFY**: Confirm webhook delivery via `Jira Settings > System > WebHooks > [webhook] > Test`
8. Monitor integration health via app-specific dashboards
## Global Configuration
### Jira Global Settings (`Jira Settings > Issues`)
**Issue Types**: Create and manage org-wide issue types; define issue type schemes; standardize across projects
**Workflows**: Create global workflow templates via `Workflows > Add workflow`; manage workflow schemes
**Custom Fields**: Create org-wide custom fields at `Custom fields > Add custom field`; manage field configurations and context
**Notification Schemes**: Configure default notification rules; create custom notification schemes; manage email templates
### Confluence Global Settings (`Confluence Admin`)
**Blueprints & Templates**: Create org-wide templates at `Configuration > Global Templates and Blueprints`; manage blueprint availability
**Themes & Appearance**: Configure org branding at `Configuration > Themes`; customize logos and colors
**Macros**: Enable/disable macros at `Configuration > Macro usage`; configure macro permissions
### Security Settings (`admin.atlassian.com > Security`)
**Authentication**:
- Password policies: `Security > Authentication policies > Edit`
- Session timeout: `Security > Session duration`
- API token management: `Security > API token controls`
**Data Residency**: Configure data location at `admin.atlassian.com > Data residency > Pin products`
**Audit Logs**: `admin.atlassian.com > Security > Audit log`
- Enable comprehensive logging; export via `GET /admin/v1/orgs/{orgId}/audit-log`
- Retain per policy (minimum 7 years for SOC 2/GDPR compliance)
## Governance & Policies
### Access Governance
- Quarterly review of all user access: `admin.atlassian.com > User management > Export users`
- Verify user roles and permissions; remove inactive users
- Limit org admins to 2–3 individuals; audit admin actions monthly
- Require MFA for all admins: `Security > Authentication policies > Require 2FA`
### Naming Conventions
**Jira**: Project keys 3–4 uppercase letters (PROJ, WEB); issue types Title Case; custom fields prefixed (CF: Story Points)
**Confluence**: Spaces use Team/Project prefix (TEAM: Engineering); pages descriptive and consistent; labels lowercase, hyphen-separated
### Change Management
**Major Changes**: Announce 2 weeks in advance; test in sandbox; create rollback plan; execute during off-peak; post-implementation review
**Minor Changes**: Announce 48 hours in advance; document in change log; monitor for issues
## Disaster Recovery
### Backup Strategy
**Jira & Confluence**: Daily automated backups; weekly manual verification; 30-day retention; offsite storage
- Trigger manual backup: `Jira Settings > System > Backup system` / `Confluence Admin > Backup and Restore`
**Recovery Testing**: Quarterly recovery drills; document procedures; measure RTO and RPO
### Incident Response
**Severity Levels**:
- **P1 (Critical)**: System down — respond in 15 min
- **P2 (High)**: Major feature broken — respond in 1 hour
- **P3 (Medium)**: Minor issue — respond in 4 hours
- **P4 (Low)**: Enhancement — respond in 24 hours
**Response Steps**:
1. Acknowledge and log incident
2. Assess impact and severity
3. Communicate status to stakeholders
4. Investigate root cause (check `admin.atlassian.com > Products > [product] > Health` and Atlassian Status Page)
5. Implement fix
6. **VERIFY**: Confirm resolution via affected user test and health check
7. Post-mortem and lessons learned
## Metrics & Reporting
**System Health**: Active users (daily/weekly/monthly), storage utilization, API rate limits, integration health, response times
- Export via: `GET /admin/v1/orgs/{orgId}/users` for user counts; product-specific analytics dashboards
**Usage Analytics**: Most active projects/spaces, content creation trends, user engagement, search patterns
**Compliance Metrics**: User access review completion, security audit findings, failed login attempts, API token usage
## Decision Framework & Handoff Protocols
**Escalate to Atlassian Support**: System outage, performance degradation org-wide, data loss/corruption, license/billing issues, complex migrations
**Delegate to Product Experts**:
- Jira Expert: Project-specific configuration
- Confluence Expert: Space-specific settings
- Scrum Master: Team workflow needs
- Senior PM: Strategic planning input
**Involve Security Team**: Security incidents, unusual access patterns, compliance audit preparation, new integration security review
**TO Jira Expert**: New global workflows, custom fields, permission schemes, or automation capabilities available
**TO Confluence Expert**: New global templates, space permission schemes, blueprints, or macros configured
**TO Senior PM**: Usage analytics, capacity planning insights, cost optimization, security compliance status
**TO Scrum Master**: Team access provisioned, board configuration options, automation rules, integrations enabled
**FROM All Roles**: User access requests, permission changes, app installation requests, configuration support, incident reports
## Atlassian MCP Integration
**Primary Tools**: Jira MCP, Confluence MCP
**Admin Operations**:
- User and group management via API
- Bulk permission updates
- Configuration audits
- Usage reporting
- System health monitoring
- Automated compliance checks
**Integration Points**:
- Support all roles with admin capabilities
- Enable Jira Expert with global configurations
- Provide Confluence Expert with template management
- Ensure Senior PM has visibility into org health
- Enable Scrum Master with team provisioning
FILE:assets/permission_scheme_template.json
{
"permissionScheme": {
"name": "Standard Project Permission Scheme",
"description": "Default permission scheme for standard projects. Assigns permissions based on project roles.",
"version": "1.0",
"lastUpdated": "YYYY-MM-DD",
"owner": "IT Admin Team"
},
"roles": {
"projectAdmin": {
"description": "Full project administration including configuration and user management",
"typicalGroups": ["project-leads", "engineering-managers"]
},
"developer": {
"description": "Create and manage issues, transitions, and attachments",
"typicalGroups": ["dept-engineering", "dept-product"]
},
"user": {
"description": "View issues, add comments, and create basic issues",
"typicalGroups": ["org-all-employees"]
},
"viewer": {
"description": "Read-only access to project issues and boards",
"typicalGroups": ["stakeholders", "external-contractors"]
}
},
"permissions": {
"project": {
"ADMINISTER_PROJECTS": {
"description": "Manage project settings, roles, and permissions",
"grantedTo": ["projectAdmin"]
},
"BROWSE_PROJECTS": {
"description": "View the project and its issues",
"grantedTo": ["projectAdmin", "developer", "user", "viewer"]
},
"VIEW_DEV_TOOLS": {
"description": "View development panel (commits, branches, PRs)",
"grantedTo": ["projectAdmin", "developer"]
},
"VIEW_READONLY_WORKFLOW": {
"description": "View read-only workflow",
"grantedTo": ["projectAdmin", "developer", "user", "viewer"]
}
},
"issues": {
"CREATE_ISSUES": {
"description": "Create new issues in the project",
"grantedTo": ["projectAdmin", "developer", "user"]
},
"EDIT_ISSUES": {
"description": "Edit issue fields",
"grantedTo": ["projectAdmin", "developer"]
},
"DELETE_ISSUES": {
"description": "Delete issues permanently",
"grantedTo": ["projectAdmin"]
},
"ASSIGN_ISSUES": {
"description": "Assign issues to team members",
"grantedTo": ["projectAdmin", "developer"]
},
"ASSIGNABLE_USER": {
"description": "Be assigned to issues",
"grantedTo": ["projectAdmin", "developer"]
},
"CLOSE_ISSUES": {
"description": "Close/resolve issues",
"grantedTo": ["projectAdmin", "developer"]
},
"RESOLVE_ISSUES": {
"description": "Set issue resolution",
"grantedTo": ["projectAdmin", "developer"]
},
"TRANSITION_ISSUES": {
"description": "Transition issues through workflow",
"grantedTo": ["projectAdmin", "developer", "user"]
},
"LINK_ISSUES": {
"description": "Create and remove issue links",
"grantedTo": ["projectAdmin", "developer"]
},
"MOVE_ISSUES": {
"description": "Move issues between projects",
"grantedTo": ["projectAdmin"]
},
"SCHEDULE_ISSUES": {
"description": "Set due dates on issues",
"grantedTo": ["projectAdmin", "developer"]
},
"SET_ISSUE_SECURITY": {
"description": "Set security level on issues",
"grantedTo": ["projectAdmin"]
}
},
"comments": {
"ADD_COMMENTS": {
"description": "Add comments to issues",
"grantedTo": ["projectAdmin", "developer", "user"]
},
"EDIT_ALL_COMMENTS": {
"description": "Edit any comment",
"grantedTo": ["projectAdmin"]
},
"EDIT_OWN_COMMENTS": {
"description": "Edit own comments",
"grantedTo": ["projectAdmin", "developer", "user"]
},
"DELETE_ALL_COMMENTS": {
"description": "Delete any comment",
"grantedTo": ["projectAdmin"]
},
"DELETE_OWN_COMMENTS": {
"description": "Delete own comments",
"grantedTo": ["projectAdmin", "developer", "user"]
}
},
"attachments": {
"CREATE_ATTACHMENTS": {
"description": "Attach files to issues",
"grantedTo": ["projectAdmin", "developer", "user"]
},
"DELETE_ALL_ATTACHMENTS": {
"description": "Delete any attachment",
"grantedTo": ["projectAdmin"]
},
"DELETE_OWN_ATTACHMENTS": {
"description": "Delete own attachments",
"grantedTo": ["projectAdmin", "developer", "user"]
}
},
"worklogs": {
"WORK_ON_ISSUES": {
"description": "Log work on issues",
"grantedTo": ["projectAdmin", "developer"]
},
"EDIT_ALL_WORKLOGS": {
"description": "Edit any worklog",
"grantedTo": ["projectAdmin"]
},
"EDIT_OWN_WORKLOGS": {
"description": "Edit own worklogs",
"grantedTo": ["projectAdmin", "developer"]
},
"DELETE_ALL_WORKLOGS": {
"description": "Delete any worklog",
"grantedTo": ["projectAdmin"]
},
"DELETE_OWN_WORKLOGS": {
"description": "Delete own worklogs",
"grantedTo": ["projectAdmin", "developer"]
}
}
},
"projectMappings": [
{
"projectKey": "EXAMPLE",
"projectName": "Example Project",
"scheme": "Standard Project Permission Scheme",
"roleAssignments": {
"projectAdmin": ["project-leads"],
"developer": ["team-example-devs"],
"user": ["org-all-employees"],
"viewer": ["stakeholders-example"]
}
}
],
"notes": {
"usage": "Copy this template and customize role assignments per project. Use group names that match your Atlassian groups.",
"review": "Review permission scheme assignments quarterly as part of access review.",
"changes": "Any changes to permission schemes should be documented and approved by IT Admin."
}
}
FILE:references/security-hardening-guide.md
# Atlassian Cloud Security Hardening Guide
## Overview
This guide provides a comprehensive security hardening checklist for Atlassian Cloud products (Jira, Confluence, Bitbucket). It covers identity management, access controls, data protection, and monitoring practices aligned with enterprise security standards.
## Identity & Authentication
### SSO / SAML Setup
**Implementation Steps:**
1. Verify your domain in Atlassian Admin (admin.atlassian.com)
2. Claim all company email accounts
3. Configure SAML SSO with your identity provider (Okta, Azure AD, Google Workspace)
4. Set authentication policy to enforce SSO for all managed accounts
5. Test with a pilot group before full rollout
6. Disable password-based login for managed accounts
**Configuration Checklist:**
- [ ] Domain verified and accounts claimed
- [ ] SAML IdP configured with correct entity ID and SSO URL
- [ ] Attribute mapping: email, displayName, groups
- [ ] Single Logout (SLO) configured
- [ ] Authentication policy enforcing SSO
- [ ] Fallback access configured for emergency admin accounts
- [ ] SCIM provisioning enabled for automatic user sync
### Two-Factor Authentication (2FA)
**Enforcement Policy:**
- [ ] 2FA required for all managed accounts
- [ ] Enforce via authentication policy (not just recommended)
- [ ] Hardware security keys (FIDO2/WebAuthn) preferred for admin accounts
- [ ] TOTP (authenticator app) as minimum for all users
- [ ] SMS-based 2FA disabled (SIM swap vulnerability)
- [ ] Recovery codes generated and stored securely
### Session Management
- [ ] Session timeout set to 8 hours of inactivity (maximum)
- [ ] Absolute session timeout: 24 hours
- [ ] Require re-authentication for sensitive operations
- [ ] Monitor concurrent sessions per user
- [ ] Enforce session termination on password change
## Access Controls
### IP Allowlisting
**Configuration:**
- [ ] Enable IP allowlisting for organization
- [ ] Add corporate office IP ranges
- [ ] Add VPN exit node IP addresses
- [ ] Add CI/CD server IPs for API access
- [ ] Test access from all approved locations
- [ ] Document approved IP ranges with justification
- [ ] Review IP allowlist quarterly
**Exceptions:**
- Mobile access may require VPN or MDM solution
- Remote workers need VPN or conditional access policies
- API integrations need stable IP ranges
### API Token Management
**Policies:**
- [ ] Inventory all API tokens in use
- [ ] Set maximum token lifetime (90 days recommended)
- [ ] Require token rotation on schedule
- [ ] Use service accounts for integrations (not personal tokens)
- [ ] Monitor API token usage patterns
- [ ] Revoke tokens immediately on employee departure
- [ ] Document purpose and owner for each token
**Best Practices:**
- Use OAuth 2.0 (3LO) for user-context integrations
- Use API tokens only for service-to-service
- Store tokens in secrets management (never in code)
- Implement least-privilege scopes for OAuth apps
### Permission Model
- [ ] Review global permissions quarterly
- [ ] Use groups for permission assignment (not individual users)
- [ ] Implement role-based access for Jira projects
- [ ] Restrict Confluence space admin to designated owners
- [ ] Limit Jira system admin to 2-3 people
- [ ] Audit "anyone" or "logged in users" permissions
- [ ] Remove direct user permissions where groups exist
## Audit & Monitoring
### Audit Log Configuration
**What to Monitor:**
- User authentication events (login, logout, failed attempts)
- Permission changes (project, space, global)
- User account changes (creation, deactivation, group changes)
- API token creation and revocation
- App installations and updates
- Data export operations
- Admin configuration changes
**Setup Steps:**
- [ ] Enable organization audit log
- [ ] Configure audit log retention (minimum 1 year)
- [ ] Set up automated export to SIEM (Splunk, Datadog, etc.)
- [ ] Create alerts for suspicious patterns
- [ ] Schedule monthly audit log review
- [ ] Document incident response procedures for alerts
### Alerting Rules
**Critical Alerts (Immediate Response):**
- Multiple failed login attempts (>5 in 10 minutes)
- Admin permission grants to unexpected users
- API token created by non-service accounts
- Bulk data export or deletion
- New third-party app installed with broad permissions
**Warning Alerts (Same-Day Review):**
- New admin users added
- Permission scheme changes
- Authentication policy modifications
- IP allowlist changes
- User deactivation (verify it is expected)
## Data Protection
### Data Residency
- [ ] Configure data residency realm (US, EU, AU, etc.)
- [ ] Verify product data pinned to selected region
- [ ] Document data residency for compliance audits
- [ ] Review data residency coverage (some metadata may be global)
- [ ] Monitor for new residency options from Atlassian
### Encryption
- [ ] Verify encryption at rest (AES-256, managed by Atlassian)
- [ ] Verify encryption in transit (TLS 1.2+)
- [ ] Review Atlassian's encryption key management practices
- [ ] Consider BYOK (Bring Your Own Key) for Atlassian Guard Premium
### Data Loss Prevention
- [ ] Configure content restrictions for sensitive pages/issues
- [ ] Implement classification labels (public, internal, confidential)
- [ ] Restrict file attachment types if needed
- [ ] Monitor bulk exports and downloads
- [ ] Set up DLP rules for sensitive data patterns (PII, credentials)
## Mobile Device Management
### Mobile Access Controls
- [ ] Require MDM enrollment for mobile Atlassian apps
- [ ] Enforce device encryption
- [ ] Require screen lock with biometrics or PIN
- [ ] Enable remote wipe capability
- [ ] Block rooted/jailbroken devices
- [ ] Restrict copy/paste to managed apps
- [ ] Set app-level PIN for Atlassian apps
### Mobile Policies
- [ ] Define approved mobile devices/OS versions
- [ ] Enforce automatic app updates
- [ ] Configure offline data access limits
- [ ] Set maximum offline cache duration
- [ ] Review mobile access logs monthly
## Third-Party App Security
### App Review Process
- [ ] Maintain approved app list (whitelist)
- [ ] Review app permissions before installation
- [ ] Verify app is Atlassian Marketplace certified
- [ ] Check app vendor security certifications
- [ ] Assess data access scope (read-only vs read-write)
- [ ] Review app privacy policy
- [ ] Document app owner and business justification
### App Governance
- [ ] Audit installed apps quarterly
- [ ] Remove unused apps (no usage in 90 days)
- [ ] Monitor app permission changes
- [ ] Restrict app installation to admins only
- [ ] Review Atlassian Guard app access policies
- [ ] Set up alerts for new app installations
## Compliance Documentation
### Required Documentation
- [ ] Security policy for Atlassian Cloud usage
- [ ] Access control matrix (roles, permissions, justification)
- [ ] Incident response plan for Atlassian security events
- [ ] Data classification policy applied to Atlassian content
- [ ] Third-party app risk assessments
- [ ] Annual security review report
### Compliance Frameworks
- **SOC 2:** Map Atlassian controls to Trust Service Criteria
- **ISO 27001:** Align with Annex A controls for cloud services
- **GDPR:** Configure data residency, right to deletion, DPAs
- **HIPAA:** Review BAA availability, encryption, access controls
## Hardening Schedule
| Task | Frequency | Owner |
|------|-----------|-------|
| Permission audit | Quarterly | IT Admin |
| API token rotation | Every 90 days | Integration owners |
| App review | Quarterly | IT Admin |
| Audit log review | Monthly | Security team |
| IP allowlist review | Quarterly | IT Admin |
| Authentication policy review | Semi-annually | Security team |
| Full security assessment | Annually | Security team |
| User access review | Quarterly | Managers + IT Admin |
| Data residency verification | Annually | Compliance |
| Mobile device audit | Quarterly | IT Admin |
FILE:references/user-provisioning-checklist.md
# User Provisioning & Lifecycle Management Checklist
## Overview
This checklist covers the complete user lifecycle in Atlassian Cloud products, from onboarding through offboarding. Consistent provisioning ensures security, compliance, and a smooth user experience.
## Onboarding Steps
### Pre-Provisioning
- [ ] Receive approved access request (ticket or HR system trigger)
- [ ] Verify employee record in HR system
- [ ] Determine role-based access level (see Role Templates below)
- [ ] Identify required Atlassian products (Jira, Confluence, Bitbucket)
- [ ] Identify required project/space access
### Account Creation
- [ ] User account auto-provisioned via SCIM (preferred) or manually created
- [ ] Email domain matches verified organization domain
- [ ] SSO authentication verified (user can log in via IdP)
- [ ] 2FA enrollment confirmed
- [ ] Correct product access assigned (Jira, Confluence, Bitbucket)
### Group Membership
- [ ] Add to organization-level groups (e.g., `all-employees`)
- [ ] Add to department group (e.g., `engineering`, `product`, `marketing`)
- [ ] Add to team-specific groups (e.g., `team-platform`, `team-mobile`)
- [ ] Add to project groups as needed (e.g., `project-alpha-members`)
- [ ] Verify group membership grants correct permissions
### Product Configuration
- [ ] **Jira:** Add to correct project roles (Developer, User, Admin)
- [ ] **Jira:** Assign to correct board(s)
- [ ] **Jira:** Set default dashboard if applicable
- [ ] **Confluence:** Grant access to relevant spaces
- [ ] **Confluence:** Add to space groups with appropriate permission level
- [ ] **Bitbucket:** Grant repository access per team
- [ ] **Bitbucket:** Configure branch permissions
### Welcome & Training
- [ ] Send welcome email with access details and key links
- [ ] Share Confluence onboarding page (getting started guide)
- [ ] Assign onboarding buddy for Atlassian tool questions
- [ ] Schedule optional training session for new users
- [ ] Provide link to internal Atlassian usage guidelines
## Role-Based Access Templates
### Developer
- **Jira:** Project Developer role (create, edit, transition issues)
- **Confluence:** Team space editor, documentation spaces viewer
- **Bitbucket:** Repository write access for team repos
### Product Manager
- **Jira:** Project Admin role (manage boards, workflows, components)
- **Confluence:** Product spaces editor, all team spaces viewer
- **Bitbucket:** Repository read access (optional)
### Designer
- **Jira:** Project User role (view, comment, transition)
- **Confluence:** Design space editor, product spaces editor
- **Bitbucket:** No access (unless needed)
### Engineering Manager
- **Jira:** Project Admin for managed projects, viewer for others
- **Confluence:** Team space admin, all spaces viewer
- **Bitbucket:** Repository admin for team repos
### Executive / Stakeholder
- **Jira:** Viewer role on strategic projects, dashboard access
- **Confluence:** Viewer on relevant spaces
- **Bitbucket:** No access
### Contractor / External
- **Jira:** Project User role, limited to specific projects
- **Confluence:** Viewer on specific spaces only (no edit)
- **Bitbucket:** Repository read access, specific repos only
- **Additional:** Set account expiration date, restrict IP access
## Group Membership Standards
### Naming Convention
```
org-{company} # Organization-wide groups
dept-{department} # Department groups
team-{team-name} # Team-specific groups
project-{project} # Project-scoped groups
role-{role} # Role-based groups (role-admin, role-viewer)
```
### Standard Groups
| Group | Purpose | Products |
|-------|---------|----------|
| `org-all-employees` | All full-time employees | Jira, Confluence |
| `dept-engineering` | All engineers | Jira, Confluence, Bitbucket |
| `dept-product` | All product team | Jira, Confluence |
| `dept-marketing` | All marketing team | Confluence |
| `role-jira-admins` | Jira administrators | Jira |
| `role-confluence-admins` | Confluence administrators | Confluence |
| `role-org-admins` | Organization administrators | All |
## Offboarding Procedure
### Immediate Actions (Day of Departure)
- [ ] Deactivate user account in Atlassian (or via IdP/SCIM)
- [ ] Revoke all API tokens associated with the user
- [ ] Revoke all OAuth app authorizations
- [ ] Transfer ownership of critical Confluence pages
- [ ] Reassign Jira issues (open/in-progress items)
- [ ] Remove from all groups
- [ ] Document access removal in offboarding ticket
### Within 24 Hours
- [ ] Verify account is fully deactivated (cannot log in)
- [ ] Check for shared credentials or service accounts
- [ ] Review audit log for recent activity
- [ ] Transfer Confluence space ownership if applicable
- [ ] Update Jira project leads/component leads if applicable
- [ ] Remove from any Atlassian Marketplace vendor accounts
### Within 7 Days
- [ ] Verify no lingering sessions or cached access
- [ ] Review integrations the user may have set up
- [ ] Check for automation rules owned by the user
- [ ] Update team dashboards and filters
- [ ] Confirm with manager that all transfers are complete
### Data Retention
- [ ] User content (pages, issues, comments) retained per policy
- [ ] Personal spaces archived or transferred
- [ ] Account marked as deactivated (not deleted) for audit trail
- [ ] Data deletion request processed if required (GDPR)
## Quarterly Access Reviews
### Review Process
1. Generate user access report from Atlassian Admin
2. Distribute to managers for team verification
3. Managers confirm or flag each user's access level
4. IT Admin processes approved changes
5. Document review completion for compliance
### Review Checklist
- [ ] All active accounts match current employee list
- [ ] No accounts for departed employees
- [ ] Group memberships align with current roles
- [ ] Admin access limited to approved administrators
- [ ] External/contractor accounts have valid expiration dates
- [ ] Service accounts documented with current owners
- [ ] Unused accounts (no login in 90 days) flagged for review
### Compliance Documentation
- [ ] Access review completion date recorded
- [ ] Manager sign-off captured (email or ticket)
- [ ] Changes made during review documented
- [ ] Exceptions documented with justification and approval
- [ ] Report filed for audit purposes
- [ ] Next review date scheduled
## Automation Opportunities
### SCIM Provisioning
- Automatically create/deactivate accounts based on IdP changes
- Sync group membership from IdP groups
- Reduce manual provisioning errors
- Ensure immediate deactivation on termination
### Workflow Automation
- Trigger onboarding checklist from HR system event
- Auto-assign to groups based on department/role attributes
- Send welcome messages via Confluence automation
- Schedule access reviews via Jira recurring tickets
### Monitoring
- Alert on accounts without 2FA after 7 days
- Alert on admin group changes
- Weekly report of new and deactivated accounts
- Monthly stale account report (no login in 90 days)
FILE:scripts/permission_audit_tool.py
#!/usr/bin/env python3
"""
Permission Audit Tool
Analyzes Atlassian permission schemes for security issues. Checks for
over-permissioned groups, direct user permissions, missing restrictions on
sensitive actions, inconsistencies across projects, and compliance gaps.
Usage:
python permission_audit_tool.py permissions.json
python permission_audit_tool.py permissions.json --format json
"""
import argparse
import json
import sys
from typing import Any, Dict, List, Optional, Set
# ---------------------------------------------------------------------------
# Audit Configuration
# ---------------------------------------------------------------------------
SENSITIVE_PERMISSIONS = {
"administer_project",
"administer_jira",
"delete_issues",
"delete_all_comments",
"delete_all_attachments",
"manage_watchers",
"modify_reporter",
"bulk_change",
"system_admin",
"manage_group_filter_subscriptions",
}
RECOMMENDED_GROUP_ONLY_PERMISSIONS = {
"browse_projects",
"create_issues",
"edit_issues",
"transition_issues",
"assign_issues",
"resolve_issues",
"close_issues",
"add_comments",
"edit_all_comments",
}
SEVERITY_WEIGHTS = {
"critical": 25,
"high": 15,
"medium": 8,
"low": 3,
"info": 1,
}
# ---------------------------------------------------------------------------
# Audit Checks
# ---------------------------------------------------------------------------
def check_over_permissioned_groups(
schemes: List[Dict[str, Any]],
) -> List[Dict[str, str]]:
"""Check for groups with overly broad admin access."""
findings = []
for scheme in schemes:
scheme_name = scheme.get("name", "Unknown Scheme")
grants = scheme.get("grants", [])
group_permissions = {}
for grant in grants:
group = grant.get("group", "")
permission = grant.get("permission", "").lower()
if group:
if group not in group_permissions:
group_permissions[group] = set()
group_permissions[group].add(permission)
for group, perms in group_permissions.items():
admin_perms = perms & SENSITIVE_PERMISSIONS
if len(admin_perms) >= 3:
findings.append({
"rule": "over_permissioned_group",
"severity": "high",
"scheme": scheme_name,
"group": group,
"message": f"Group '{group}' has {len(admin_perms)} sensitive permissions "
f"in scheme '{scheme_name}': {', '.join(sorted(admin_perms))}. "
f"Review if all are necessary.",
})
if "system_admin" in perms or "administer_jira" in perms:
findings.append({
"rule": "admin_access_warning",
"severity": "critical",
"scheme": scheme_name,
"group": group,
"message": f"Group '{group}' has system/Jira admin access in '{scheme_name}'. "
f"Ensure this is strictly necessary and membership is limited.",
})
return findings
def check_direct_user_permissions(
schemes: List[Dict[str, Any]],
) -> List[Dict[str, str]]:
"""Check for permissions granted directly to users instead of groups."""
findings = []
for scheme in schemes:
scheme_name = scheme.get("name", "Unknown Scheme")
grants = scheme.get("grants", [])
for grant in grants:
user = grant.get("user", "")
permission = grant.get("permission", "")
if user and not grant.get("group"):
severity = "high" if permission.lower() in SENSITIVE_PERMISSIONS else "medium"
findings.append({
"rule": "direct_user_permission",
"severity": severity,
"scheme": scheme_name,
"user": user,
"message": f"User '{user}' has direct permission '{permission}' in '{scheme_name}'. "
f"Use groups instead for maintainability and audit clarity.",
})
return findings
def check_missing_restrictions(
schemes: List[Dict[str, Any]],
) -> List[Dict[str, str]]:
"""Check for missing restrictions on sensitive actions."""
findings = []
for scheme in schemes:
scheme_name = scheme.get("name", "Unknown Scheme")
grants = scheme.get("grants", [])
granted_permissions = set()
for grant in grants:
granted_permissions.add(grant.get("permission", "").lower())
# Check if delete permissions are unrestricted
delete_perms = {"delete_issues", "delete_all_comments", "delete_all_attachments"}
unrestricted_deletes = delete_perms & granted_permissions
for grant in grants:
perm = grant.get("permission", "").lower()
group = grant.get("group", "")
if perm in delete_perms and group:
# Check if granted to broad groups
broad_groups = {"users", "everyone", "all-users", "jira-users", "jira-software-users"}
if group.lower() in broad_groups:
findings.append({
"rule": "unrestricted_delete",
"severity": "critical",
"scheme": scheme_name,
"message": f"Delete permission '{perm}' granted to broad group '{group}' "
f"in '{scheme_name}'. Restrict to admins or leads only.",
})
# Check if admin permissions exist
admin_perms = {"administer_project", "administer_jira", "system_admin"}
if not (admin_perms & granted_permissions):
findings.append({
"rule": "no_admin_defined",
"severity": "medium",
"scheme": scheme_name,
"message": f"No explicit admin permission defined in '{scheme_name}'. "
f"Ensure project administration is properly assigned.",
})
return findings
def check_scheme_consistency(
schemes: List[Dict[str, Any]],
) -> List[Dict[str, str]]:
"""Check for inconsistencies across permission schemes."""
findings = []
if len(schemes) < 2:
return findings
# Compare permission sets across schemes
scheme_perms = {}
for scheme in schemes:
name = scheme.get("name", "Unknown")
perms = set()
for grant in scheme.get("grants", []):
perms.add(grant.get("permission", "").lower())
scheme_perms[name] = perms
# Find schemes with significantly different permission sets
all_perms = set()
for perms in scheme_perms.values():
all_perms |= perms
scheme_names = list(scheme_perms.keys())
for i in range(len(scheme_names)):
for j in range(i + 1, len(scheme_names)):
name_a = scheme_names[i]
name_b = scheme_names[j]
diff = scheme_perms[name_a].symmetric_difference(scheme_perms[name_b])
if len(diff) > 5:
findings.append({
"rule": "scheme_inconsistency",
"severity": "medium",
"message": f"Schemes '{name_a}' and '{name_b}' differ significantly "
f"({len(diff)} different permissions). Review for intentional differences.",
})
return findings
def check_compliance_gaps(
schemes: List[Dict[str, Any]],
) -> List[Dict[str, str]]:
"""Check for common compliance gaps."""
findings = []
for scheme in schemes:
scheme_name = scheme.get("name", "Unknown Scheme")
grants = scheme.get("grants", [])
groups_used = set()
users_used = set()
for grant in grants:
if grant.get("group"):
groups_used.add(grant["group"])
if grant.get("user"):
users_used.add(grant["user"])
# Check for separation of duties
admin_groups = set()
for grant in grants:
if grant.get("permission", "").lower() in SENSITIVE_PERMISSIONS and grant.get("group"):
admin_groups.add(grant["group"])
if len(admin_groups) == 1 and len(groups_used) > 1:
findings.append({
"rule": "separation_of_duties",
"severity": "info",
"scheme": scheme_name,
"message": f"Only one group ('{next(iter(admin_groups))}') holds all sensitive permissions "
f"in '{scheme_name}'. Consider separating duties across multiple groups.",
})
# Check user count
if len(users_used) > 5:
findings.append({
"rule": "too_many_direct_users",
"severity": "high",
"scheme": scheme_name,
"message": f"Scheme '{scheme_name}' has {len(users_used)} direct user grants. "
f"Migrate to group-based permissions for better governance.",
})
return findings
# ---------------------------------------------------------------------------
# Main Analysis
# ---------------------------------------------------------------------------
def audit_permissions(data: Dict[str, Any]) -> Dict[str, Any]:
"""Run full permission audit."""
schemes = data.get("schemes", [])
if not schemes:
# Try treating the entire input as a single scheme
if data.get("grants") or data.get("name"):
schemes = [data]
else:
return {
"risk_score": 0,
"grade": "invalid",
"error": "No permission schemes found in input",
"findings": [],
"summary": {},
}
all_findings = []
all_findings.extend(check_over_permissioned_groups(schemes))
all_findings.extend(check_direct_user_permissions(schemes))
all_findings.extend(check_missing_restrictions(schemes))
all_findings.extend(check_scheme_consistency(schemes))
all_findings.extend(check_compliance_gaps(schemes))
# Calculate risk score (higher = more risk)
summary = {"critical": 0, "high": 0, "medium": 0, "low": 0, "info": 0}
total_penalty = 0
for finding in all_findings:
severity = finding["severity"]
summary[severity] = summary.get(severity, 0) + 1
total_penalty += SEVERITY_WEIGHTS.get(severity, 0)
risk_score = min(100, total_penalty)
health_score = max(0, 100 - risk_score)
if health_score >= 85:
grade = "excellent"
elif health_score >= 70:
grade = "good"
elif health_score >= 50:
grade = "fair"
else:
grade = "poor"
# Generate remediation recommendations
remediations = _generate_remediations(all_findings)
return {
"risk_score": risk_score,
"health_score": health_score,
"grade": grade,
"schemes_analyzed": len(schemes),
"findings": all_findings,
"summary": summary,
"remediations": remediations,
}
def _generate_remediations(findings: List[Dict[str, str]]) -> List[str]:
"""Generate remediation recommendations."""
remediations = []
rules_seen = set()
for finding in findings:
rule = finding["rule"]
if rule in rules_seen:
continue
rules_seen.add(rule)
if rule == "over_permissioned_group":
remediations.append("Review and reduce sensitive permissions for over-permissioned groups. Apply principle of least privilege.")
elif rule == "admin_access_warning":
remediations.append("Audit admin group membership. Limit system/Jira admin access to essential personnel only.")
elif rule == "direct_user_permission":
remediations.append("Migrate direct user permissions to group-based grants. Create functional groups for common permission sets.")
elif rule == "unrestricted_delete":
remediations.append("Restrict delete permissions to project admins or leads. Remove from broad user groups.")
elif rule == "scheme_inconsistency":
remediations.append("Standardize permission schemes across projects. Document intentional differences.")
elif rule == "too_many_direct_users":
remediations.append("Create groups for users with direct permissions. This simplifies onboarding/offboarding.")
elif rule == "separation_of_duties":
remediations.append("Consider splitting admin responsibilities across multiple groups for better separation of duties.")
elif rule == "no_admin_defined":
remediations.append("Define explicit admin permissions in each scheme to ensure proper project governance.")
return remediations
# ---------------------------------------------------------------------------
# Output Formatting
# ---------------------------------------------------------------------------
def format_text_output(result: Dict[str, Any]) -> str:
"""Format results as readable text report."""
lines = []
lines.append("=" * 60)
lines.append("PERMISSION AUDIT REPORT")
lines.append("=" * 60)
lines.append("")
if "error" in result:
lines.append(f"ERROR: {result['error']}")
return "\n".join(lines)
lines.append("AUDIT SUMMARY")
lines.append("-" * 30)
lines.append(f"Risk Score: {result['risk_score']}/100 (lower is better)")
lines.append(f"Health Score: {result['health_score']}/100")
lines.append(f"Grade: {result['grade'].title()}")
lines.append(f"Schemes Analyzed: {result['schemes_analyzed']}")
lines.append("")
summary = result.get("summary", {})
lines.append("FINDINGS BY SEVERITY")
lines.append("-" * 30)
lines.append(f"Critical: {summary.get('critical', 0)}")
lines.append(f"High: {summary.get('high', 0)}")
lines.append(f"Medium: {summary.get('medium', 0)}")
lines.append(f"Low: {summary.get('low', 0)}")
lines.append(f"Info: {summary.get('info', 0)}")
lines.append("")
findings = result.get("findings", [])
if findings:
lines.append("DETAILED FINDINGS")
lines.append("-" * 30)
for i, finding in enumerate(findings, 1):
severity = finding["severity"].upper()
lines.append(f"{i}. [{severity}] {finding['message']}")
lines.append(f" Rule: {finding['rule']}")
if finding.get("scheme"):
lines.append(f" Scheme: {finding['scheme']}")
lines.append("")
remediations = result.get("remediations", [])
if remediations:
lines.append("REMEDIATION RECOMMENDATIONS")
lines.append("-" * 30)
for i, rem in enumerate(remediations, 1):
lines.append(f"{i}. {rem}")
return "\n".join(lines)
def format_json_output(result: Dict[str, Any]) -> Dict[str, Any]:
"""Format results as JSON."""
return result
# ---------------------------------------------------------------------------
# CLI Interface
# ---------------------------------------------------------------------------
def main() -> int:
"""Main CLI entry point."""
parser = argparse.ArgumentParser(
description="Audit Atlassian permission schemes for security issues"
)
parser.add_argument(
"permissions_file",
help="JSON file with permission scheme data",
)
parser.add_argument(
"--format",
choices=["text", "json"],
default="text",
help="Output format (default: text)",
)
args = parser.parse_args()
try:
with open(args.permissions_file, "r") as f:
data = json.load(f)
result = audit_permissions(data)
if args.format == "json":
print(json.dumps(format_json_output(result), indent=2))
else:
print(format_text_output(result))
return 0
except FileNotFoundError:
print(f"Error: File '{args.permissions_file}' not found", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON in '{args.permissions_file}': {e}", file=sys.stderr)
return 1
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
return 1
if __name__ == "__main__":
sys.exit(main())
Thảo luận 6 giai đoạn giữa các vai trò C-suite với cách ly độc lập, phản biện và tổng hợp, đầu ra là biên bản HĐQT.
---
name: "boardroom"
description: "/cs:boardroom <brief> — 6-phase multi-role deliberation across the C-suite with Phase 2 isolation, critic pre-screen, and synthesis. Outputs a board memo."
---
# /cs:boardroom — Multi-Role Boardroom Deliberation
**Command:** `/cs:boardroom <brief-path>`
Runs the `board-meeting` skill protocol across the C-suite for a single strategy brief. This is the **heart of the plugin** — the multi-role deliberation that gstack's review chain only approximates.
## Pipeline Position
```
/cs:office-hours → /cs:brief → /cs:boardroom → /cs:decide → /cs:execute → /cs:post-mortem
↑ you are here
```
## The 6 Phases (from board-meeting skill)
### Phase 1 — Briefing
- Chief of Staff distributes the brief to all advisors marked in **Affected Roles**.
- Each advisor reads company-context.md + the brief.
- No discussion yet.
### Phase 2 — Independent Thinking (ISOLATION)
- **Critical:** each advisor produces their position **independently**, without seeing others' positions.
- This prevents groupthink and surfaces dissent.
- Each writes: their voice's opening, recommendation, top 3 concerns, top 3 supports.
### Phase 3 — Cross-Examination
- Positions revealed simultaneously.
- Each advisor critiques the others' positions on the dimensions they own:
- cs-cfo-advisor critiques the math
- cs-ciso-advisor critiques the risk
- cs-cpo-advisor critiques the JTBD
- cs-cmo-advisor critiques the positioning
- cs-cro-advisor critiques the revenue math
- etc.
### Phase 4 — Devil's Advocate Pass
- `executive-mentor/devils-advocate` agent runs `/em:challenge` on the leading option.
- Surfaces three concerns with severity ratings.
### Phase 5 — Synthesis
- Chief of Staff synthesizes: which option commands majority, what are unresolved dissents.
- Produces the **board memo** with recommendation + dissent.
### Phase 6 — Decision Hand-off
- Memo is presented to the founder.
- Founder accepts, modifies, or rejects.
- Approved memo routes to `/cs:decide` for logging.
## Output: Board Memo
Saved to `~/.claude/boardroom/YYYY-MM-DD-<slug>.md`:
```markdown
# Board Memo: <topic>
**Date:** YYYY-MM-DD
**Brief:** <link to /cs:brief file>
**Status:** AWAITING FOUNDER DECISION | APPROVED | REJECTED
## Question
[One sentence from the brief]
## Recommended Option
**<Option name>** — chosen because <synthesis reasoning>
## Vote Tally
| Advisor | Vote | One-Sentence Reason |
|---|---|---|
| cs-ceo-advisor | A | <reason> |
| cs-cfo-advisor | A | <reason> |
| cs-cto-advisor | B | <reason> |
| ... | | |
## Dissent
- **<dissenter>:** <unresolved concern>
## Devil's Advocate Concerns
1. **CRITICAL** — <concern> — Mitigation: <plan>
2. **HIGH** — <concern> — Mitigation: <plan>
3. **MEDIUM** — <concern> — Mitigation: <plan>
## Success & Kill Criteria
[Copied from brief, refined by the panel]
## Recommended Decision Path
- `/cs:decide` → log the decision
- `/cs:execute` → 90-day plan
- `/cs:cross-eval` → multi-model sanity check (optional, high-stakes)
- `/cs:freeze N` → cooldown lock (optional, irreversible)
```
## Why Phase 2 Isolation Matters
If advisors see each other's positions before forming their own, they anchor. Phase 2 isolation is the single highest-leverage practice in the board-meeting protocol — it surfaces the dissents that sycophancy would have suppressed.
## Why This Beats gstack's Review Chain
| | gstack `/autoplan` | `/cs:boardroom` |
|---|---|---|
| Roles | CEO → design → eng (3) | Up to 10 C-roles |
| Order | Sequential | Phase 2 isolation, then simultaneous |
| Dissent capture | Implicit | Explicit dissent column |
| Adversarial pass | No | Phase 4 devil's advocate |
| Output | Reviewed plan | Voted memo with dissent + kill criteria |
## Workflow
1. Read brief from `~/.claude/briefs/<file>`
2. Identify affected roles
3. Invoke each cs-* advisor independently (Phase 2)
4. Collect positions
5. Run cross-examination round (Phase 3)
6. Run `/em:challenge` on leading option (Phase 4)
7. Synthesize memo (Phase 5)
8. Hand off to founder (Phase 6)
## Routing
- `/cs:decide` — log approved memo
- `/cs:cross-eval` — high-stakes second opinion
- `/cs:freeze` — cooldown lock
## Related
- Agent: [`cs-chief-of-staff`](../../agents/cs-chief-of-staff.md)
- Skills: [`board-meeting`](../../../skills/board-meeting/SKILL.md), [`executive-mentor`](../../../executive-mentor/)
---
**Version:** 1.0.0
Tạo bản tóm tắt chiến lược một trang từ buổi office hours, bước đầu của quy trình sprint chiến lược.
---
name: "brief"
description: "/cs:brief <topic> — Generate a one-page strategy brief from an office-hours intake. First step in the strategic sprint pipeline."
---
# /cs:brief — One-Page Strategy Brief
**Command:** `/cs:brief <topic>` or `/cs:brief <office-hours-output>`
Turns intake (raw question or office-hours output) into a one-page strategy brief that the boardroom can deliberate on. This is **Step 1** of the strategic sprint pipeline.
## Pipeline Position
```
/cs:office-hours → /cs:brief → /cs:boardroom → /cs:decide → /cs:execute → /cs:post-mortem
↑ you are here
```
## Inputs
- A topic string, **or**
- An office-hours brief (preferred — more rigor)
- `~/.claude/company-context.md` (loaded automatically)
## Output
A single Markdown file under `~/.claude/briefs/YYYY-MM-DD-<slug>.md` with this structure:
```markdown
# Strategy Brief: <topic>
**Date:** YYYY-MM-DD
**Author:** cs-chief-of-staff
**Status:** DRAFT | UNDER REVIEW | APPROVED | RETIRED
## Context
[1-2 paragraphs: where the company sits today on this topic — pulled from company-context.md]
## Question
[The one sentence question the boardroom must answer]
## Options
1. **Option A:** <name> — <one-sentence summary>
2. **Option B:** <name> — <one-sentence summary>
3. **Option C:** <name> — <one-sentence summary>
(Minimum 2 options. "Do nothing" is always an option.)
## Assumptions
- <assumption 1 — explicit>
- <assumption 2>
- <assumption 3>
## Constraints
- Time: <by when must this decide>
- Money: <budget envelope>
- People: <who can / can't be reallocated>
- Reversibility: <one-way door | two-way door>
## Affected Roles
[Which cs-* advisors should weigh in. Used to route to /cs:boardroom panel composition.]
- [ ] cs-ceo-advisor
- [ ] cs-cfo-advisor
- [ ] cs-cto-advisor
- [ ] cs-cmo-advisor
- [ ] cs-cro-advisor
- [ ] cs-cpo-advisor
- [ ] cs-coo-advisor
- [ ] cs-chro-advisor
- [ ] cs-ciso-advisor
- [ ] cs-chief-of-staff
## Success Criteria
[Measurable outcomes that define success — set BEFORE the decision]
- <metric 1, threshold, timeframe>
- <metric 2, threshold, timeframe>
## Kill Criteria
[What signal would tell you in 90 days that this was the wrong call]
- <metric, threshold, action if missed>
```
## Workflow
1. Load company-context.md via context-engine
2. If input is office-hours output, parse the 6 answers
3. If input is a raw topic, prompt the founder for the missing pieces
4. Draft 2-3 options (never just one — every brief needs a counterfactual)
5. Make assumptions and constraints explicit
6. Identify affected roles → drives panel composition for `/cs:boardroom`
7. Write success + kill criteria BEFORE the decision (this is the rigor moment)
8. Save to `~/.claude/briefs/`
## Why This Step Exists
The biggest decision-making failure is debating implementation before agreeing on the question. The brief locks the question, options, and success criteria so the boardroom can deliberate without scope creep.
This is also the **artifact handoff** — the next command consumes this file, not your memory.
## Routing
- `/cs:boardroom <brief>` — multi-role deliberation
- `/cs:cross-eval <brief>` — multi-model sanity check before boardroom (for high-stakes)
- `/cs:freeze <brief>` — cooldown lock for irreversible decisions
## Related
- Agent: [`cs-chief-of-staff`](../../agents/cs-chief-of-staff.md)
- Skills: [`context-engine`](../../../skills/context-engine/SKILL.md), [`board-meeting`](../../../skills/board-meeting/SKILL.md)
---
**Version:** 1.0.0
Tự động hóa tác vụ trình duyệt: thu thập web, điền biểu mẫu, chụp màn hình và trích xuất dữ liệu có cấu trúc.
---
name: "browser-automation"
description: "Use when the user asks to automate browser tasks, scrape websites, fill forms, capture screenshots, extract structured data from web pages, or build web automation workflows. NOT for testing — use playwright-pro for that."
---
# Browser Automation - POWERFUL
## Overview
The Browser Automation skill provides comprehensive tools and knowledge for building production-grade web automation workflows using Playwright. This skill covers data extraction, form filling, screenshot capture, session management, and anti-detection patterns for reliable browser automation at scale.
**When to use this skill:**
- Scraping structured data from websites (tables, listings, search results)
- Automating multi-step browser workflows (login, fill forms, download files)
- Capturing screenshots or PDFs of web pages
- Extracting data from SPAs and JavaScript-heavy sites
- Building repeatable browser-based data pipelines
**When NOT to use this skill:**
- Writing browser tests or E2E test suites — use **playwright-pro** instead
- Testing API endpoints — use **api-test-suite-builder** instead
- Load testing or performance benchmarking — use **performance-profiler** instead
**Why Playwright over Selenium or Puppeteer:**
- **Auto-wait built in** — no explicit `sleep()` or `waitForElement()` needed for most actions
- **Multi-browser from one API** — Chromium, Firefox, WebKit with zero config changes
- **Network interception** — block ads, mock responses, capture API calls natively
- **Browser contexts** — isolated sessions without spinning up new browser instances
- **Codegen** — `playwright codegen` records your actions and generates scripts
- **Async-first** — Python async/await for high-throughput scraping
## Core Competencies
### 1. Web Scraping Patterns
**Selector priority (most to least reliable):**
1. `data-testid`, `data-id`, or custom data attributes — stable across redesigns
2. `#id` selectors — unique but may change between deploys
3. Semantic selectors: `article`, `nav`, `main`, `section` — resilient to CSS changes
4. Class-based: `.product-card`, `.price` — brittle if classes are generated (e.g., CSS modules)
5. Positional: `nth-child()`, `nth-of-type()` — last resort, breaks on layout changes
Use XPath only when CSS cannot express the relationship (e.g., ancestor traversal, text-based selection).
**Pagination strategies:** next-button, URL-based (`?page=N`), infinite scroll, load-more button. See [data_extraction_recipes.md](references/data_extraction_recipes.md) for complete pagination handlers and scroll patterns.
### 2. Form Filling & Multi-Step Workflows
Break multi-step forms into discrete functions per step. Each function fills fields, clicks "Next"/"Continue", and waits for the next step to load (URL change or DOM element).
Key patterns: login flows, multi-page forms, file uploads (including drag-and-drop zones), native and custom dropdown handling. See [playwright_browser_api.md](references/playwright_browser_api.md) for complete API reference on `fill()`, `select_option()`, `set_input_files()`, and `expect_file_chooser()`.
### 3. Screenshot & PDF Capture
- **Full page:** `await page.screenshot(path="full.png", full_page=True)`
- **Element:** `await page.locator("div.chart").screenshot(path="chart.png")`
- **PDF (Chromium only):** `await page.pdf(path="out.pdf", format="A4", print_background=True)`
- **Visual regression:** Take screenshots at known states, store baselines in version control with naming: `{page}_{viewport}_{state}.png`
See [playwright_browser_api.md](references/playwright_browser_api.md) for full screenshot/PDF options.
### 4. Structured Data Extraction
Core extraction patterns:
- **Tables to JSON** — Extract `<thead>` headers and `<tbody>` rows into dictionaries
- **Listings to arrays** — Map repeating card elements using a field-selector map (supports `::attr()` for attributes)
- **Nested/threaded data** — Recursive extraction for comments with replies, category trees
See [data_extraction_recipes.md](references/data_extraction_recipes.md) for complete extraction functions, price parsing, data cleaning utilities, and output format helpers (JSON, CSV, JSONL).
### 5. Cookie & Session Management
- **Save/restore cookies:** `context.cookies()` and `context.add_cookies()`
- **Full storage state** (cookies + localStorage): `context.storage_state(path="state.json")` to save, `browser.new_context(storage_state="state.json")` to restore
**Best practice:** Save state after login, reuse across scraping sessions. Check session validity before starting a long job — make a lightweight request to a protected page and verify you are not redirected to login. See [playwright_browser_api.md](references/playwright_browser_api.md) for cookie and storage state API details.
### 6. Anti-Detection Patterns
Modern websites detect automation through multiple vectors. Apply these in priority order:
1. **WebDriver flag removal** — Remove `navigator.webdriver = true` via init script (critical)
2. **Custom user agent** — Rotate through real browser UAs; never use the default headless UA
3. **Realistic viewport** — Set 1920x1080 or similar real-world dimensions (default 800x600 is a red flag)
4. **Request throttling** — Add `random.uniform()` delays between actions
5. **Proxy support** — Per-browser or per-context proxy configuration
See [anti_detection_patterns.md](references/anti_detection_patterns.md) for the complete stealth stack: navigator property hardening, WebGL/canvas fingerprint evasion, behavioral simulation (mouse movement, typing speed, scroll patterns), proxy rotation strategies, and detection self-test URLs.
### 7. Dynamic Content Handling
- **SPA rendering:** Wait for content selectors (`wait_for_selector`), not the page load event
- **AJAX/Fetch waiting:** Use `page.expect_response("**/api/data*")` to intercept and wait for specific API calls
- **Shadow DOM:** Playwright pierces open Shadow DOM with `>>` operator: `page.locator("custom-element >> .inner-class")`
- **Lazy-loaded images:** Scroll elements into view with `scroll_into_view_if_needed()` to trigger loading
See [playwright_browser_api.md](references/playwright_browser_api.md) for wait strategies, network interception, and Shadow DOM details.
### 8. Error Handling & Retry Logic
- **Retry with backoff:** Wrap page interactions in retry logic with exponential backoff (e.g., 1s, 2s, 4s)
- **Fallback selectors:** On `TimeoutError`, try alternative selectors before failing
- **Error-state screenshots:** Capture `page.screenshot(path="error-state.png")` on unexpected failures for debugging
- **Rate limit detection:** Check for HTTP 429 responses and respect `Retry-After` headers
See [anti_detection_patterns.md](references/anti_detection_patterns.md) for the complete exponential backoff implementation and rate limiter class.
## Workflows
### Workflow 1: Single-Page Data Extraction
**Scenario:** Extract product data from a single page with JavaScript-rendered content.
**Steps:**
1. Launch browser in headed mode during development (`headless=False`), switch to headless for production
2. Navigate to URL and wait for content selector
3. Extract data using `query_selector_all` with field mapping
4. Validate extracted data (check for nulls, expected types)
5. Output as JSON
```python
async def extract_single_page(url, selectors):
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
context = await browser.new_context(
viewport={"width": 1920, "height": 1080},
user_agent="Mozilla/5.0 ..."
)
page = await context.new_page()
await page.goto(url, wait_until="networkidle")
data = await extract_listings(page, selectors["container"], selectors["fields"])
await browser.close()
return data
```
### Workflow 2: Multi-Page Scraping with Pagination
**Scenario:** Scrape search results across 50+ pages.
**Steps:**
1. Launch browser with anti-detection settings
2. Navigate to first page
3. Extract data from current page
4. Check if "Next" button exists and is enabled
5. Click next, wait for new content to load (not just navigation)
6. Repeat until no next page or max pages reached
7. Deduplicate results by unique key
8. Write output incrementally (don't hold everything in memory)
```python
async def scrape_paginated(base_url, selectors, max_pages=100):
all_data = []
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
page = await (await browser.new_context()).new_page()
await page.goto(base_url)
for page_num in range(max_pages):
items = await extract_listings(page, selectors["container"], selectors["fields"])
all_data.extend(items)
next_btn = page.locator(selectors["next_button"])
if await next_btn.count() == 0 or await next_btn.is_disabled():
break
await next_btn.click()
await page.wait_for_selector(selectors["container"])
await human_delay(800, 2000)
await browser.close()
return all_data
```
### Workflow 3: Authenticated Workflow Automation
**Scenario:** Log into a portal, navigate a multi-step form, download a report.
**Steps:**
1. Check for existing session state file
2. If no session, perform login and save state
3. Navigate to target page using saved session
4. Fill multi-step form with provided data
5. Wait for download to trigger
6. Save downloaded file to target directory
```python
async def authenticated_workflow(credentials, form_data, download_dir):
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
state_file = "session_state.json"
# Restore or create session
if os.path.exists(state_file):
context = await browser.new_context(storage_state=state_file)
else:
context = await browser.new_context()
page = await context.new_page()
await login(page, credentials["url"], credentials["user"], credentials["pass"])
await context.storage_state(path=state_file)
page = await context.new_page()
await page.goto(form_data["target_url"])
# Fill form steps
for step_fn in [fill_step_1, fill_step_2]:
await step_fn(page, form_data)
# Handle download
async with page.expect_download() as dl_info:
await page.click("button:has-text('Download Report')")
download = await dl_info.value
await download.save_as(os.path.join(download_dir, download.suggested_filename))
await browser.close()
```
## Tools Reference
| Script | Purpose | Key Flags | Output |
|--------|---------|-----------|--------|
| `scraping_toolkit.py` | Generate Playwright scraping script skeleton | `--url`, `--selectors`, `--paginate`, `--output` | Python script or JSON config |
| `form_automation_builder.py` | Generate form-fill automation script from field spec | `--fields`, `--url`, `--output` | Python automation script |
| `anti_detection_checker.py` | Audit a Playwright script for detection vectors | `--file`, `--verbose` | Risk report with score |
All scripts are stdlib-only. Run `python3 <script> --help` for full usage.
## Anti-Patterns
### Hardcoded Waits
**Bad:** `await page.wait_for_timeout(5000)` before every action.
**Good:** Use `wait_for_selector`, `wait_for_url`, `expect_response`, or `wait_for_load_state`. Hardcoded waits are flaky and slow.
### No Error Recovery
**Bad:** Linear script that crashes on first failure.
**Good:** Wrap each page interaction in try/except. Take error-state screenshots. Implement retry with exponential backoff.
### Ignoring robots.txt
**Bad:** Scraping without checking robots.txt directives.
**Good:** Fetch and parse robots.txt before scraping. Respect `Crawl-delay`. Skip disallowed paths. Add your bot name to User-Agent if running at scale.
### Storing Credentials in Scripts
**Bad:** Hardcoding usernames and passwords in Python files.
**Good:** Use environment variables, `.env` files (gitignored), or a secrets manager. Pass credentials via CLI arguments.
### No Rate Limiting
**Bad:** Hammering a site with 100 requests/second.
**Good:** Add random delays between requests (1-3s for polite scraping). Monitor for 429 responses. Implement exponential backoff.
### Selector Fragility
**Bad:** Relying on auto-generated class names (`.css-1a2b3c`) or deep nesting (`div > div > div > span:nth-child(3)`).
**Good:** Use data attributes, semantic HTML, or text-based locators. Test selectors in browser DevTools first.
### Not Cleaning Up Browser Instances
**Bad:** Launching browsers without closing them, leading to resource leaks.
**Good:** Always use `try/finally` or async context managers to ensure `browser.close()` is called.
### Running Headed in Production
**Bad:** Using `headless=False` in production/CI.
**Good:** Develop with headed mode for debugging, deploy with `headless=True`. Use environment variable to toggle: `headless = os.environ.get("HEADLESS", "true") == "true"`.
## Cross-References
- **playwright-pro** — Browser testing skill. Use for E2E tests, test assertions, test fixtures. Browser Automation is for data extraction and workflow automation, not testing.
- **api-test-suite-builder** — When the website has a public API, hit the API directly instead of scraping the rendered page. Faster, more reliable, less detectable.
- **performance-profiler** — If your automation scripts are slow, profile the bottlenecks before adding concurrency.
- **env-secrets-manager** — For securely managing credentials used in authenticated automation workflows.
FILE:references/anti_detection_patterns.md
# Anti-Detection Patterns for Browser Automation
This reference covers techniques to make Playwright automation less detectable by anti-bot services. These are defense-in-depth measures — no single technique is sufficient, but combining them significantly reduces detection risk.
## Detection Vectors
Anti-bot systems detect automation through multiple signals. Understanding what they check helps you counter effectively.
### Tier 1: Trivial Detection (Every Site Checks These)
1. **navigator.webdriver** — Set to `true` by all automation frameworks
2. **User-Agent string** — Default headless UA contains "HeadlessChrome"
3. **WebGL renderer** — Headless Chrome reports "SwiftShader" or "Google SwiftShader"
### Tier 2: Common Detection (Most Anti-Bot Services)
4. **Viewport/screen dimensions** — Unusual sizes flag automation
5. **Plugins array** — Empty in headless mode, populated in real browsers
6. **Languages** — Missing or mismatched locale
7. **Request timing** — Machine-speed interactions
8. **Mouse movement** — No mouse events between clicks
### Tier 3: Advanced Detection (Cloudflare, DataDome, PerimeterX)
9. **Canvas fingerprint** — Headless renders differently
10. **WebGL fingerprint** — GPU-specific rendering variations
11. **Audio fingerprint** — AudioContext processing differences
12. **Font enumeration** — Different available fonts in headless
13. **Behavioral analysis** — Scroll patterns, click patterns, reading time
## Stealth Techniques
### 1. WebDriver Flag Removal
The most critical fix. Every anti-bot check starts here.
```python
await page.add_init_script("""
// Remove webdriver flag
Object.defineProperty(navigator, 'webdriver', {
get: () => undefined,
});
// Remove Playwright-specific properties
delete window.__playwright;
delete window.__pw_manual;
""")
```
### 2. User Agent Configuration
Match the user agent to the browser you are launching. A Chrome UA with Firefox-specific headers is a red flag.
```python
# Chrome 120 on Windows 10 (most common configuration globally)
CHROME_WIN = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
# Chrome 120 on macOS
CHROME_MAC = "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
# Chrome 120 on Linux
CHROME_LINUX = "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
# Firefox 121 on Windows
FIREFOX_WIN = "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:121.0) Gecko/20100101 Firefox/121.0"
```
**Rules:**
- Update UAs every 2-3 months as browser versions increment
- Match UA platform to `navigator.platform` override
- If using Chromium, use Chrome UAs. If Firefox, use Firefox UAs.
- Never use obviously fake or ancient UAs
### 3. Viewport and Screen Properties
Common real-world screen resolutions (from analytics data):
| Resolution | Market Share | Use For |
|-----------|-------------|---------|
| 1920x1080 | ~23% | Default choice |
| 1366x768 | ~14% | Laptop simulation |
| 1536x864 | ~9% | Scaled laptop |
| 1440x900 | ~7% | MacBook |
| 2560x1440 | ~5% | High-end desktop |
```python
import random
VIEWPORTS = [
{"width": 1920, "height": 1080},
{"width": 1366, "height": 768},
{"width": 1536, "height": 864},
{"width": 1440, "height": 900},
]
viewport = random.choice(VIEWPORTS)
context = await browser.new_context(
viewport=viewport,
screen=viewport, # screen should match viewport
)
```
### 4. Navigator Properties Hardening
```python
STEALTH_INIT = """
// Plugins (headless Chrome has 0 plugins, real Chrome has 3-5)
Object.defineProperty(navigator, 'plugins', {
get: () => {
const plugins = [
{ name: 'Chrome PDF Plugin', filename: 'internal-pdf-viewer' },
{ name: 'Chrome PDF Viewer', filename: 'mhjfbmdgcfjbbpaeojofohoefgiehjai' },
{ name: 'Native Client', filename: 'internal-nacl-plugin' },
];
plugins.length = 3;
return plugins;
},
});
// Languages
Object.defineProperty(navigator, 'languages', {
get: () => ['en-US', 'en'],
});
// Platform (match to user agent)
Object.defineProperty(navigator, 'platform', {
get: () => 'Win32', // or 'MacIntel' for macOS UA
});
// Hardware concurrency (real browsers report CPU cores)
Object.defineProperty(navigator, 'hardwareConcurrency', {
get: () => 8,
});
// Device memory (Chrome-specific)
Object.defineProperty(navigator, 'deviceMemory', {
get: () => 8,
});
// Connection info
Object.defineProperty(navigator, 'connection', {
get: () => ({
effectiveType: '4g',
rtt: 50,
downlink: 10,
saveData: false,
}),
});
"""
await context.add_init_script(STEALTH_INIT)
```
### 5. WebGL Fingerprint Evasion
Headless Chrome uses SwiftShader for WebGL, which anti-bot services detect.
```python
# Option A: Launch with a real GPU (headed mode on a machine with GPU)
browser = await p.chromium.launch(headless=False)
# Option B: Override WebGL renderer info
await page.add_init_script("""
const getParameter = WebGLRenderingContext.prototype.getParameter;
WebGLRenderingContext.prototype.getParameter = function(parameter) {
if (parameter === 37445) {
return 'Intel Inc.'; // UNMASKED_VENDOR_WEBGL
}
if (parameter === 37446) {
return 'Intel(R) Iris(TM) Plus Graphics 640'; // UNMASKED_RENDERER_WEBGL
}
return getParameter.call(this, parameter);
};
""")
```
### 6. Canvas Fingerprint Noise
Anti-bot services render text/shapes to a canvas and hash the output. Headless Chrome produces a different hash.
```python
await page.add_init_script("""
const originalToDataURL = HTMLCanvasElement.prototype.toDataURL;
HTMLCanvasElement.prototype.toDataURL = function(type) {
if (type === 'image/png' || type === undefined) {
// Add minimal noise to the canvas to change fingerprint
const ctx = this.getContext('2d');
if (ctx) {
const imageData = ctx.getImageData(0, 0, this.width, this.height);
for (let i = 0; i < imageData.data.length; i += 4) {
// Shift one channel by +/- 1 (imperceptible)
imageData.data[i] = imageData.data[i] ^ 1;
}
ctx.putImageData(imageData, 0, 0);
}
}
return originalToDataURL.apply(this, arguments);
};
""")
```
## Request Throttling Patterns
### Human-Like Delays
Real users do not click at machine speed. Add realistic delays between actions.
```python
import random
import asyncio
async def human_delay(action_type="browse"):
"""Add realistic delay based on action type."""
delays = {
"browse": (1.0, 3.0), # Browsing between pages
"read": (2.0, 8.0), # Reading content
"fill": (0.3, 0.8), # Between form fields
"click": (0.1, 0.5), # Before clicking
"scroll": (0.5, 1.5), # Between scroll actions
}
min_s, max_s = delays.get(action_type, (0.5, 2.0))
await asyncio.sleep(random.uniform(min_s, max_s))
```
### Request Rate Limiting
```python
import time
class RateLimiter:
"""Enforce minimum delay between requests."""
def __init__(self, min_interval_seconds=1.0):
self.min_interval = min_interval_seconds
self.last_request_time = 0
async def wait(self):
elapsed = time.time() - self.last_request_time
if elapsed < self.min_interval:
await asyncio.sleep(self.min_interval - elapsed)
self.last_request_time = time.time()
# Usage
limiter = RateLimiter(min_interval_seconds=2.0)
for url in urls:
await limiter.wait()
await page.goto(url)
```
### Exponential Backoff on Errors
```python
async def with_backoff(coro_factory, max_retries=5, base_delay=1.0):
for attempt in range(max_retries):
try:
return await coro_factory()
except Exception as e:
if attempt == max_retries - 1:
raise
delay = base_delay * (2 ** attempt) + random.uniform(0, 1)
print(f"Attempt {attempt + 1} failed: {e}. Retrying in {delay:.1f}s...")
await asyncio.sleep(delay)
```
## Proxy Rotation Strategies
### Single Proxy
```python
browser = await p.chromium.launch(
proxy={"server": "http://proxy.example.com:8080"}
)
```
### Authenticated Proxy
```python
context = await browser.new_context(
proxy={
"server": "http://proxy.example.com:8080",
"username": "user",
"password": "pass",
}
)
```
### Rotating Proxy Pool
```python
PROXIES = [
"http://proxy1.example.com:8080",
"http://proxy2.example.com:8080",
"http://proxy3.example.com:8080",
]
async def create_context_with_proxy(browser):
proxy = random.choice(PROXIES)
return await browser.new_context(
proxy={"server": proxy}
)
```
### Per-Request Proxy (via Context Rotation)
Playwright does not support per-request proxy switching. Achieve it by creating a new context for each request or batch:
```python
async def scrape_url(browser, url, proxy):
context = await browser.new_context(proxy={"server": proxy})
page = await context.new_page()
try:
await page.goto(url)
data = await extract_data(page)
return data
finally:
await context.close()
```
### SOCKS5 Proxy
```python
browser = await p.chromium.launch(
proxy={"server": "socks5://proxy.example.com:1080"}
)
```
## Headless Detection Avoidance
### Running Chrome Channel Instead of Chromium
The bundled Chromium binary has different properties than a real Chrome install. Using the Chrome channel makes the browser indistinguishable from a normal install.
```python
# Use installed Chrome instead of bundled Chromium
browser = await p.chromium.launch(channel="chrome", headless=True)
```
**Requirements:** Chrome must be installed on the system.
### New Headless Mode (Chrome 112+)
Chrome's "new headless" mode is harder to detect than the old one:
```python
browser = await p.chromium.launch(
args=["--headless=new"],
)
```
### Avoiding Common Flags
Do NOT pass these flags — they are headless-detection signals:
- `--disable-gpu` (old headless workaround, not needed)
- `--no-sandbox` (security risk, detectable)
- `--disable-setuid-sandbox` (same as above)
## Behavioral Evasion
### Mouse Movement Simulation
Anti-bot services track mouse events. A click without preceding mouse movement is suspicious.
```python
async def human_click(page, selector):
"""Click with preceding mouse movement."""
element = await page.query_selector(selector)
box = await element.bounding_box()
if box:
# Move to element with slight offset
x = box["x"] + box["width"] / 2 + random.uniform(-5, 5)
y = box["y"] + box["height"] / 2 + random.uniform(-5, 5)
await page.mouse.move(x, y, steps=random.randint(5, 15))
await asyncio.sleep(random.uniform(0.05, 0.2))
await page.mouse.click(x, y)
```
### Typing Speed Variation
```python
async def human_type(page, selector, text):
"""Type with variable speed like a human."""
await page.click(selector)
for char in text:
await page.keyboard.type(char)
# Faster for common keys, slower for special characters
if char in "aeiou tnrs":
await asyncio.sleep(random.uniform(0.03, 0.08))
else:
await asyncio.sleep(random.uniform(0.08, 0.20))
```
### Scroll Behavior
Real users scroll gradually, not in instant jumps.
```python
async def human_scroll(page, distance=None):
"""Scroll down gradually like a human."""
if distance is None:
distance = random.randint(300, 800)
current = 0
while current < distance:
step = random.randint(50, 150)
await page.mouse.wheel(0, step)
current += step
await asyncio.sleep(random.uniform(0.05, 0.15))
```
## Detection Testing
### Self-Check Script
Navigate to these URLs to test your stealth configuration:
- `https://bot.sannysoft.com/` — Comprehensive bot detection test
- `https://abrahamjuliot.github.io/creepjs/` — Advanced fingerprint analysis
- `https://browserleaks.com/webgl` — WebGL fingerprint details
- `https://browserleaks.com/canvas` — Canvas fingerprint details
### Quick Test Pattern
```python
async def test_stealth(page):
"""Navigate to detection test page and report results."""
await page.goto("https://bot.sannysoft.com/")
await page.wait_for_timeout(3000)
# Check for failed tests
failed = await page.eval_on_selector_all(
"td.failed",
"els => els.map(e => e.parentElement.querySelector('td').textContent)"
)
if failed:
print(f"FAILED checks: {failed}")
else:
print("All checks passed.")
await page.screenshot(path="stealth_test.png", full_page=True)
```
## Recommended Stealth Stack
For most automation tasks, apply these in order of priority:
1. **WebDriver flag removal** — Critical, takes 2 lines
2. **Custom user agent** — Critical, takes 1 line
3. **Viewport configuration** — High priority, takes 1 line
4. **Request delays** — High priority, add random.uniform() calls
5. **Navigator properties** — Medium priority, init script block
6. **Chrome channel** — Medium priority, one launch option
7. **WebGL override** — Low priority unless hitting advanced anti-bot
8. **Canvas noise** — Low priority unless hitting advanced anti-bot
9. **Proxy rotation** — Only for high-volume or repeated scraping
10. **Behavioral simulation** — Only for sites with behavioral analysis
FILE:references/data_extraction_recipes.md
# Data Extraction Recipes
Practical patterns for extracting structured data from web pages using Playwright. Each recipe is a self-contained pattern you can adapt to your target site.
## CSS Selector Patterns for Common Structures
### E-Commerce Product Listings
```python
PRODUCT_SELECTORS = {
"container": "div.product-card, article.product, li.product-item",
"fields": {
"title": "h2.product-title, h3.product-name, [data-testid='product-title']",
"price": "span.price, .product-price, [data-testid='price']",
"original_price": "span.original-price, .was-price, del",
"rating": "span.rating, .star-rating, [data-rating]",
"review_count": "span.review-count, .num-reviews",
"image_url": "img.product-image::attr(src), img::attr(data-src)",
"product_url": "a.product-link::attr(href), h2 a::attr(href)",
"availability": "span.stock-status, .availability",
}
}
```
### News/Blog Article Listings
```python
ARTICLE_SELECTORS = {
"container": "article, div.post, div.article-card",
"fields": {
"headline": "h2 a, h3 a, .article-title",
"summary": "p.excerpt, .article-summary, .post-excerpt",
"author": "span.author, .byline, [rel='author']",
"date": "time, span.date, .published-date",
"category": "span.category, a.tag, .article-category",
"url": "h2 a::attr(href), .article-title a::attr(href)",
"image_url": "img.thumbnail::attr(src), .article-image img::attr(src)",
}
}
```
### Job Listings
```python
JOB_SELECTORS = {
"container": "div.job-card, li.job-listing, article.job",
"fields": {
"title": "h2.job-title, a.job-link, [data-testid='job-title']",
"company": "span.company-name, .employer, [data-testid='company']",
"location": "span.location, .job-location, [data-testid='location']",
"salary": "span.salary, .compensation, [data-testid='salary']",
"job_type": "span.job-type, .employment-type",
"posted_date": "time, span.posted, .date-posted",
"url": "a.job-link::attr(href), h2 a::attr(href)",
}
}
```
### Search Engine Results
```python
SERP_SELECTORS = {
"container": "div.g, .search-result, li.result",
"fields": {
"title": "h3, .result-title",
"url": "a::attr(href), cite",
"snippet": "div.VwiC3b, .result-snippet, .search-description",
"displayed_url": "cite, .result-url",
}
}
```
## Table Extraction Recipes
### Simple HTML Table to JSON
The most common extraction pattern. Works for any standard `<table>` with `<thead>` and `<tbody>`.
```python
async def extract_table(page, table_selector="table"):
"""Extract an HTML table into a list of dictionaries."""
data = await page.evaluate(f"""
(selector) => {{
const table = document.querySelector(selector);
if (!table) return null;
// Get headers
const headers = Array.from(table.querySelectorAll('thead th, thead td'))
.map(th => th.textContent.trim());
// If no thead, use first row as headers
if (headers.length === 0) {{
const firstRow = table.querySelector('tr');
if (firstRow) {{
headers.push(...Array.from(firstRow.querySelectorAll('th, td'))
.map(cell => cell.textContent.trim()));
}}
}}
// Get data rows
const rows = Array.from(table.querySelectorAll('tbody tr'));
return rows.map(row => {{
const cells = Array.from(row.querySelectorAll('td'));
const obj = {{}};
cells.forEach((cell, i) => {{
if (i < headers.length) {{
obj[headers[i]] = cell.textContent.trim();
}}
}});
return obj;
}});
}}
""", table_selector)
return data or []
```
### Table with Links and Attributes
When table cells contain links or data attributes, not just text:
```python
async def extract_rich_table(page, table_selector="table"):
"""Extract table including links and data attributes."""
return await page.evaluate(f"""
(selector) => {{
const table = document.querySelector(selector);
if (!table) return [];
const headers = Array.from(table.querySelectorAll('thead th'))
.map(th => th.textContent.trim());
return Array.from(table.querySelectorAll('tbody tr')).map(row => {{
const obj = {{}};
Array.from(row.querySelectorAll('td')).forEach((cell, i) => {{
const key = headers[i] || `col_{i}`;
obj[key] = cell.textContent.trim();
// Extract link if present
const link = cell.querySelector('a');
if (link) {{
obj[key + '_url'] = link.href;
}}
// Extract data attributes
for (const attr of cell.attributes) {{
if (attr.name.startsWith('data-')) {{
obj[key + '_' + attr.name] = attr.value;
}}
}}
}});
return obj;
}});
}}
""", table_selector)
```
### Multi-Page Table (Paginated)
```python
async def extract_paginated_table(page, table_selector, next_selector, max_pages=50):
"""Extract data from a table that spans multiple pages."""
all_rows = []
headers = None
for page_num in range(max_pages):
# Extract current page
page_data = await page.evaluate(f"""
(selector) => {{
const table = document.querySelector(selector);
if (!table) return {{ headers: [], rows: [] }};
const hs = Array.from(table.querySelectorAll('thead th'))
.map(th => th.textContent.trim());
const rs = Array.from(table.querySelectorAll('tbody tr')).map(row =>
Array.from(row.querySelectorAll('td')).map(td => td.textContent.trim())
);
return {{ headers: hs, rows: rs }};
}}
""", table_selector)
if headers is None and page_data["headers"]:
headers = page_data["headers"]
for row in page_data["rows"]:
all_rows.append(dict(zip(headers or [], row)))
# Check for next page
next_btn = page.locator(next_selector)
if await next_btn.count() == 0 or await next_btn.is_disabled():
break
await next_btn.click()
await page.wait_for_load_state("networkidle")
await page.wait_for_timeout(random.randint(800, 2000))
return all_rows
```
## Product Listing Extraction
### Generic Listing Extractor
Works for any repeating card/list pattern:
```python
async def extract_listings(page, container_sel, field_map):
"""
Extract data from repeating elements.
field_map: dict mapping field names to CSS selectors.
Special suffixes:
::attr(name) — extract attribute instead of text
::html — extract innerHTML
"""
items = []
cards = await page.query_selector_all(container_sel)
for card in cards:
item = {}
for field_name, selector in field_map.items():
try:
if "::attr(" in selector:
sel, attr = selector.split("::attr(")
attr = attr.rstrip(")")
el = await card.query_selector(sel)
item[field_name] = await el.get_attribute(attr) if el else None
elif selector.endswith("::html"):
sel = selector.replace("::html", "")
el = await card.query_selector(sel)
item[field_name] = await el.inner_html() if el else None
else:
el = await card.query_selector(selector)
item[field_name] = (await el.text_content()).strip() if el else None
except Exception:
item[field_name] = None
items.append(item)
return items
```
### With Price Parsing
```python
import re
def parse_price(text):
"""Extract numeric price from text like '$1,234.56' or '1.234,56 EUR'."""
if not text:
return None
# Remove currency symbols and whitespace
cleaned = re.sub(r'[^\d.,]', '', text.strip())
if not cleaned:
return None
# Handle European format (1.234,56)
if ',' in cleaned and '.' in cleaned:
if cleaned.rindex(',') > cleaned.rindex('.'):
cleaned = cleaned.replace('.', '').replace(',', '.')
else:
cleaned = cleaned.replace(',', '')
elif ',' in cleaned:
# Could be 1,234 or 1,23 — check decimal places
parts = cleaned.split(',')
if len(parts[-1]) <= 2:
cleaned = cleaned.replace(',', '.')
else:
cleaned = cleaned.replace(',', '')
try:
return float(cleaned)
except ValueError:
return None
async def extract_products_with_prices(page, container_sel, field_map, price_field="price"):
"""Extract listings and parse prices into floats."""
items = await extract_listings(page, container_sel, field_map)
for item in items:
if price_field in item and item[price_field]:
item[f"{price_field}_raw"] = item[price_field]
item[price_field] = parse_price(item[price_field])
return items
```
## Pagination Handling
### Next-Button Pagination
The most common pattern. Click "Next" until the button disappears or is disabled.
```python
async def paginate_via_next_button(page, next_selector, content_selector, max_pages=100):
"""
Yield page objects as you paginate through results.
next_selector: CSS selector for the "Next" button/link
content_selector: CSS selector to wait for after navigation (confirms new page loaded)
"""
pages_scraped = 0
while pages_scraped < max_pages:
yield page # Caller extracts data from current page
pages_scraped += 1
next_btn = page.locator(next_selector)
if await next_btn.count() == 0:
break
try:
is_disabled = await next_btn.is_disabled()
except Exception:
is_disabled = True
if is_disabled:
break
await next_btn.click()
await page.wait_for_selector(content_selector, state="attached")
await page.wait_for_timeout(random.randint(500, 1500))
```
### URL-Based Pagination
When pages follow a predictable URL pattern:
```python
async def paginate_via_url(page, url_template, start=1, max_pages=100):
"""
Navigate through pages using URL parameters.
url_template: URL with {page} placeholder, e.g., "https://example.com/search?page={page}"
"""
for page_num in range(start, start + max_pages):
url = url_template.format(page=page_num)
response = await page.goto(url, wait_until="networkidle")
if response and response.status == 404:
break
yield page, page_num
await page.wait_for_timeout(random.randint(800, 2500))
```
### Infinite Scroll
For sites that load content as you scroll:
```python
async def paginate_via_scroll(page, item_selector, max_scrolls=100, no_change_limit=3):
"""
Scroll to load more content until no new items appear.
item_selector: CSS selector for individual items (used to count progress)
no_change_limit: Stop after N scrolls with no new items
"""
previous_count = 0
no_change_streak = 0
for scroll_num in range(max_scrolls):
# Count current items
current_count = await page.locator(item_selector).count()
if current_count == previous_count:
no_change_streak += 1
if no_change_streak >= no_change_limit:
break
else:
no_change_streak = 0
previous_count = current_count
# Scroll to bottom
await page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
await page.wait_for_timeout(random.randint(1000, 2500))
# Check for "Load More" button that might appear
load_more = page.locator("button:has-text('Load More'), button:has-text('Show More')")
if await load_more.count() > 0 and await load_more.is_visible():
await load_more.click()
await page.wait_for_timeout(random.randint(1000, 2000))
return current_count
```
### Load-More Button
Simpler variant of infinite scroll where content loads via a button:
```python
async def paginate_via_load_more(page, button_selector, item_selector, max_clicks=50):
"""Click a 'Load More' button repeatedly until it disappears."""
for click_num in range(max_clicks):
btn = page.locator(button_selector)
if await btn.count() == 0 or not await btn.is_visible():
break
count_before = await page.locator(item_selector).count()
await btn.click()
# Wait for new items to appear
try:
await page.wait_for_function(
f"document.querySelectorAll('{item_selector}').length > {count_before}",
timeout=10000,
)
except Exception:
break # No new items loaded
await page.wait_for_timeout(random.randint(500, 1500))
return await page.locator(item_selector).count()
```
## Nested Data Extraction
### Comments with Replies (Threaded)
```python
async def extract_threaded_comments(page, parent_selector=".comments"):
"""Recursively extract threaded comments."""
return await page.evaluate(f"""
(parentSelector) => {{
function extractThread(container) {{
const comments = [];
const directChildren = container.querySelectorAll(':scope > .comment');
for (const comment of directChildren) {{
const authorEl = comment.querySelector('.author, .username');
const textEl = comment.querySelector('.comment-text, .comment-body');
const dateEl = comment.querySelector('time, .date');
const repliesContainer = comment.querySelector('.replies, .children');
comments.push({{
author: authorEl ? authorEl.textContent.trim() : null,
text: textEl ? textEl.textContent.trim() : null,
date: dateEl ? (dateEl.getAttribute('datetime') || dateEl.textContent.trim()) : null,
replies: repliesContainer ? extractThread(repliesContainer) : [],
}});
}}
return comments;
}}
const root = document.querySelector(parentSelector);
return root ? extractThread(root) : [];
}}
""", parent_selector)
```
### Nested Categories (Sidebar/Menu)
```python
async def extract_category_tree(page, root_selector="nav.categories"):
"""Extract nested category structure from a sidebar or menu."""
return await page.evaluate(f"""
(rootSelector) => {{
function extractLevel(container) {{
const items = [];
const directItems = container.querySelectorAll(':scope > li, :scope > div.category');
for (const item of directItems) {{
const link = item.querySelector(':scope > a');
const subMenu = item.querySelector(':scope > ul, :scope > div.sub-categories');
items.push({{
name: link ? link.textContent.trim() : item.textContent.trim().split('\\n')[0],
url: link ? link.href : null,
children: subMenu ? extractLevel(subMenu) : [],
}});
}}
return items;
}}
const root = document.querySelector(rootSelector);
return root ? extractLevel(root.querySelector('ul') || root) : [];
}}
""", root_selector)
```
### Accordion/Expandable Content
Some content is hidden behind accordion/expand toggles. Click to reveal, then extract.
```python
async def extract_accordion(page, toggle_selector, content_selector):
"""Expand all accordion items and extract their content."""
items = []
toggles = await page.query_selector_all(toggle_selector)
for toggle in toggles:
title = (await toggle.text_content()).strip()
# Click to expand
await toggle.click()
await page.wait_for_timeout(300)
# Find the associated content panel
content = await toggle.evaluate_handle(
f"el => el.closest('.accordion-item, .faq-item')?.querySelector('{content_selector}')"
)
body = None
if content:
body = (await content.text_content())
if body:
body = body.strip()
items.append({"title": title, "content": body})
return items
```
## Data Cleaning Utilities
### Post-Extraction Cleaning
```python
import re
def clean_text(text):
"""Normalize whitespace, remove zero-width characters."""
if not text:
return None
# Remove zero-width characters
text = re.sub(r'[\u200b\u200c\u200d\ufeff]', '', text)
# Normalize whitespace
text = re.sub(r'\s+', ' ', text).strip()
return text if text else None
def clean_url(url, base_url=None):
"""Convert relative URLs to absolute."""
if not url:
return None
url = url.strip()
if url.startswith("//"):
return "https:" + url
if url.startswith("/") and base_url:
return base_url.rstrip("/") + url
return url
def deduplicate(items, key_field):
"""Remove duplicate items based on a key field."""
seen = set()
unique = []
for item in items:
key = item.get(key_field)
if key and key not in seen:
seen.add(key)
unique.append(item)
return unique
```
### Output Formats
```python
import json
import csv
import io
def to_jsonl(items, file_path):
"""Write items as JSON Lines (one JSON object per line)."""
with open(file_path, "w") as f:
for item in items:
f.write(json.dumps(item, ensure_ascii=False) + "\n")
def to_csv(items, file_path):
"""Write items as CSV."""
if not items:
return
headers = list(items[0].keys())
with open(file_path, "w", newline="") as f:
writer = csv.DictWriter(f, fieldnames=headers)
writer.writeheader()
writer.writerows(items)
def to_json(items, file_path, indent=2):
"""Write items as a JSON array."""
with open(file_path, "w") as f:
json.dump(items, f, indent=indent, ensure_ascii=False)
```
FILE:references/playwright_browser_api.md
# Playwright Browser API Reference (Automation Focus)
This reference covers Playwright's Python async API for browser automation tasks — NOT testing. For test-specific APIs (assertions, fixtures, test runners), see playwright-pro.
## Browser Launch & Context
### Launching the Browser
```python
from playwright.async_api import async_playwright
async with async_playwright() as p:
# Chromium (recommended for most automation)
browser = await p.chromium.launch(headless=True)
# Firefox (better for some anti-detection scenarios)
browser = await p.firefox.launch(headless=True)
# WebKit (Safari engine — useful for Apple-specific sites)
browser = await p.webkit.launch(headless=True)
```
**Launch options:**
| Option | Type | Default | Purpose |
|--------|------|---------|---------|
| `headless` | bool | True | Run without visible window |
| `slow_mo` | int | 0 | Milliseconds to slow each operation (debugging) |
| `proxy` | dict | None | Proxy server configuration |
| `args` | list | [] | Additional Chromium flags |
| `downloads_path` | str | None | Directory for downloads |
| `channel` | str | None | Browser channel: "chrome", "msedge" |
### Browser Contexts (Session Isolation)
Browser contexts are isolated environments within a single browser instance. Each context has its own cookies, localStorage, and cache. Use them instead of launching multiple browsers.
```python
# Create isolated context
context = await browser.new_context(
viewport={"width": 1920, "height": 1080},
user_agent="Mozilla/5.0 ...",
locale="en-US",
timezone_id="America/New_York",
geolocation={"latitude": 40.7128, "longitude": -74.0060},
permissions=["geolocation"],
)
# Multiple contexts share one browser (resource efficient)
context_a = await browser.new_context() # User A session
context_b = await browser.new_context() # User B session
```
### Storage State (Session Persistence)
```python
# Save state after login (cookies + localStorage)
await context.storage_state(path="auth_state.json")
# Restore state in new context
context = await browser.new_context(storage_state="auth_state.json")
```
## Page Navigation
### Basic Navigation
```python
page = await context.new_page()
# Navigate with different wait strategies
await page.goto("https://example.com") # Default: "load"
await page.goto("https://example.com", wait_until="domcontentloaded") # Faster
await page.goto("https://example.com", wait_until="networkidle") # Wait for network quiet
await page.goto("https://example.com", timeout=30000) # Custom timeout (ms)
```
**`wait_until` options:**
- `"load"` — wait for the `load` event (all resources loaded)
- `"domcontentloaded"` — DOM is ready, images/styles may still load
- `"networkidle"` — no network requests for 500ms (best for SPAs)
- `"commit"` — response received, before any rendering
### Wait Strategies
```python
# Wait for a specific element to appear
await page.wait_for_selector("div.content", state="visible")
await page.wait_for_selector("div.loading", state="hidden") # Wait for loading to finish
await page.wait_for_selector("table tbody tr", state="attached") # In DOM but maybe not visible
# Wait for URL change
await page.wait_for_url("**/dashboard**")
await page.wait_for_url(re.compile(r"/dashboard/\d+"))
# Wait for specific network response
async with page.expect_response("**/api/data*") as resp_info:
await page.click("button.load")
response = await resp_info.value
json_data = await response.json()
# Wait for page load state
await page.wait_for_load_state("networkidle")
# Fixed wait (use sparingly — prefer the methods above)
await page.wait_for_timeout(1000) # milliseconds
```
### Navigation History
```python
await page.go_back()
await page.go_forward()
await page.reload()
```
## Element Interaction
### Finding Elements
```python
# Single element (returns first match)
element = await page.query_selector("css=div.product")
element = await page.query_selector("xpath=//div[@class='product']")
# Multiple elements
elements = await page.query_selector_all("div.product")
# Locator API (recommended — auto-waits, re-queries on each action)
locator = page.locator("div.product")
count = await locator.count()
first = locator.first
nth = locator.nth(2)
```
**Locator vs query_selector:**
- `query_selector` — returns an ElementHandle at a point in time. Can go stale if DOM changes.
- `locator` — returns a Locator that re-queries each time you interact with it. Preferred for reliability.
### Clicking
```python
await page.click("button.submit")
await page.click("a:has-text('Next')")
await page.dblclick("div.editable")
await page.click("button", position={"x": 10, "y": 10}) # Click at offset
await page.click("button", force=True) # Skip actionability checks
await page.click("button", modifiers=["Shift"]) # With modifier key
```
### Text Input
```python
# Fill (clears existing content first)
await page.fill("input#email", "user@example.com")
# Type (simulates keystroke-by-keystroke input — slower, more realistic)
await page.type("input#search", "query text", delay=50) # 50ms between keys
# Press specific keys
await page.press("input#search", "Enter")
await page.press("body", "Control+a")
```
### Dropdowns & Select
```python
# Native <select> element
await page.select_option("select#country", value="US")
await page.select_option("select#country", label="United States")
await page.select_option("select#tags", value=["tag1", "tag2"]) # Multi-select
# Custom dropdown (non-native)
await page.click("div.dropdown-trigger")
await page.click("li.option:has-text('United States')")
```
### Checkboxes & Radio Buttons
```python
await page.check("input#agree")
await page.uncheck("input#newsletter")
is_checked = await page.is_checked("input#agree")
```
### File Upload
```python
# Standard file input
await page.set_input_files("input[type='file']", "/path/to/file.pdf")
await page.set_input_files("input[type='file']", ["/path/a.pdf", "/path/b.pdf"])
# Clear file selection
await page.set_input_files("input[type='file']", [])
# Non-standard upload (drag-and-drop zones)
async with page.expect_file_chooser() as fc_info:
await page.click("div.upload-zone")
file_chooser = await fc_info.value
await file_chooser.set_files("/path/to/file.pdf")
```
### Hover & Focus
```python
await page.hover("div.menu-item")
await page.focus("input#search")
```
## Data Extraction
### Text Content
```python
# Get text content of an element
text = await page.text_content("h1.title")
inner_text = await page.inner_text("div.description") # Visible text only
inner_html = await page.inner_html("div.content") # HTML markup
# Get attribute
href = await page.get_attribute("a.link", "href")
src = await page.get_attribute("img.photo", "src")
```
### JavaScript Evaluation
```python
# Evaluate in page context
title = await page.evaluate("document.title")
scroll_height = await page.evaluate("document.body.scrollHeight")
# Evaluate on a specific element
text = await page.eval_on_selector("h1", "el => el.textContent")
texts = await page.eval_on_selector_all("li", "els => els.map(e => e.textContent.trim())")
# Complex extraction
data = await page.evaluate("""
() => {
const rows = document.querySelectorAll('table tbody tr');
return Array.from(rows).map(row => {
const cells = row.querySelectorAll('td');
return {
name: cells[0]?.textContent.trim(),
value: cells[1]?.textContent.trim(),
};
});
}
""")
```
### Screenshots & PDF
```python
# Full page screenshot
await page.screenshot(path="page.png", full_page=True)
# Viewport screenshot
await page.screenshot(path="viewport.png")
# Element screenshot
await page.locator("div.chart").screenshot(path="chart.png")
# PDF (Chromium only)
await page.pdf(path="page.pdf", format="A4", print_background=True)
# Screenshot as bytes (for processing without saving)
buffer = await page.screenshot()
```
## Network Interception
### Monitoring Requests
```python
# Listen for all responses
page.on("response", lambda response: print(f"{response.status} {response.url}"))
# Wait for a specific API call
async with page.expect_response("**/api/products*") as resp:
await page.click("button.load")
response = await resp.value
data = await response.json()
```
### Blocking Resources (Speed Up Scraping)
```python
# Block images, fonts, and CSS to speed up scraping
await page.route("**/*.{png,jpg,jpeg,gif,svg,woff,woff2,ttf}", lambda route: route.abort())
await page.route("**/*.css", lambda route: route.abort())
# Block specific domains (ads, analytics)
await page.route("**/google-analytics.com/**", lambda route: route.abort())
await page.route("**/facebook.com/**", lambda route: route.abort())
```
### Modifying Requests
```python
# Add custom headers
await page.route("**/*", lambda route: route.continue_(headers={
**route.request.headers,
"X-Custom-Header": "value"
}))
# Mock API responses
await page.route("**/api/data", lambda route: route.fulfill(
status=200,
content_type="application/json",
body=json.dumps({"items": []}),
))
```
## Dialog Handling
```python
# Auto-accept all dialogs
page.on("dialog", lambda dialog: dialog.accept())
# Handle specific dialog types
async def handle_dialog(dialog):
if dialog.type == "confirm":
await dialog.accept()
elif dialog.type == "prompt":
await dialog.accept("my input")
elif dialog.type == "alert":
await dialog.dismiss()
page.on("dialog", handle_dialog)
```
## File Downloads
```python
# Wait for download to start
async with page.expect_download() as dl_info:
await page.click("a.download-link")
download = await dl_info.value
# Save to specific path
await download.save_as("/path/to/downloads/" + download.suggested_filename)
# Get download as bytes
path = await download.path() # Temp file path
# Set download behavior at context level
context = await browser.new_context(accept_downloads=True)
```
## Frames & Iframes
```python
# Access iframe by selector
frame = page.frame_locator("iframe#content")
await frame.locator("button.submit").click()
# Access frame by name
frame = page.frame(name="editor")
# Access all frames
for frame in page.frames:
print(frame.url)
```
## Cookie Management
```python
# Get all cookies
cookies = await context.cookies()
# Get cookies for specific URL
cookies = await context.cookies(["https://example.com"])
# Add cookies
await context.add_cookies([{
"name": "session",
"value": "abc123",
"domain": "example.com",
"path": "/",
"httpOnly": True,
"secure": True,
}])
# Clear cookies
await context.clear_cookies()
```
## Concurrency Patterns
### Multiple Pages in One Context
```python
# Open multiple tabs in the same session
pages = []
for url in urls:
page = await context.new_page()
await page.goto(url)
pages.append(page)
# Process all pages
for page in pages:
data = await extract_data(page)
await page.close()
```
### Multiple Contexts for Parallel Sessions
```python
import asyncio
async def scrape_with_context(browser, url):
context = await browser.new_context(user_agent=random.choice(USER_AGENTS))
page = await context.new_page()
await page.goto(url)
data = await extract_data(page)
await context.close()
return data
# Run 5 concurrent scraping tasks
tasks = [scrape_with_context(browser, url) for url in urls[:5]]
results = await asyncio.gather(*tasks)
```
## Init Scripts (Stealth)
Init scripts run before any page script, in every new page/context.
```python
# Remove webdriver flag
await context.add_init_script("""
Object.defineProperty(navigator, 'webdriver', {get: () => undefined});
""")
# Override plugins (headless Chrome has empty plugins)
await context.add_init_script("""
Object.defineProperty(navigator, 'plugins', {
get: () => [1, 2, 3, 4, 5],
});
""")
# Override languages
await context.add_init_script("""
Object.defineProperty(navigator, 'languages', {
get: () => ['en-US', 'en'],
});
""")
# From file
await context.add_init_script(path="stealth.js")
```
## Common Automation Patterns
### Scrolling
```python
# Scroll to bottom
await page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
# Scroll element into view
await page.locator("div.target").scroll_into_view_if_needed()
# Smooth scroll simulation
await page.evaluate("""
async () => {
const delay = ms => new Promise(r => setTimeout(r, ms));
for (let i = 0; i < document.body.scrollHeight; i += 300) {
window.scrollTo(0, i);
await delay(100);
}
}
""")
```
### Clipboard Operations
```python
# Copy text
await page.evaluate("navigator.clipboard.writeText('hello')")
# Paste via keyboard
await page.keyboard.press("Control+v")
```
### Shadow DOM
```python
# Playwright pierces open shadow DOM with >> operator
await page.locator("my-component >> .inner-button").click()
# Or use the css= engine with >> for chained piercing
await page.locator("css=host-element >> css=.shadow-child").click()
```
FILE:scripts/anti_detection_checker.py
#!/usr/bin/env python3
"""
Anti-Detection Checker - Audits Playwright scripts for common bot detection vectors.
Analyzes a Playwright automation script and identifies patterns that make the
browser detectable as a bot. Produces a risk score (0-100) with specific
recommendations for each issue found.
Detection vectors checked:
- Headless mode usage
- Default/missing user agent configuration
- Viewport size (default 800x600 is a red flag)
- WebDriver flag (navigator.webdriver)
- Navigator property overrides
- Request throttling / human-like delays
- Cookie/session management
- Proxy configuration
- Error handling patterns
No external dependencies - uses only Python standard library.
"""
import argparse
import json
import os
import re
import sys
from dataclasses import dataclass, asdict
from typing import List, Optional
@dataclass
class Finding:
"""A single detection risk finding."""
category: str
severity: str # "critical", "high", "medium", "low", "info"
description: str
line: Optional[int]
recommendation: str
weight: int # Points added to risk score (0-15)
SEVERITY_WEIGHTS = {
"critical": 15,
"high": 10,
"medium": 5,
"low": 2,
"info": 0,
}
class AntiDetectionChecker:
"""Analyzes Playwright scripts for bot detection vulnerabilities."""
def __init__(self, script_content: str, file_path: str = "<stdin>"):
self.content = script_content
self.lines = script_content.split("\n")
self.file_path = file_path
self.findings: List[Finding] = []
def check_all(self) -> List[Finding]:
"""Run all detection checks."""
self._check_headless_mode()
self._check_user_agent()
self._check_viewport()
self._check_webdriver_flag()
self._check_navigator_properties()
self._check_request_delays()
self._check_error_handling()
self._check_proxy()
self._check_session_management()
self._check_browser_close()
self._check_stealth_imports()
return self.findings
def _find_line(self, pattern: str) -> Optional[int]:
"""Find the first line number matching a regex pattern."""
for i, line in enumerate(self.lines, 1):
if re.search(pattern, line):
return i
return None
def _has_pattern(self, pattern: str) -> bool:
"""Check if pattern exists anywhere in the script."""
return bool(re.search(pattern, self.content))
def _check_headless_mode(self):
"""Check if headless mode is properly configured."""
if self._has_pattern(r"headless\s*=\s*False"):
self.findings.append(Finding(
category="Headless Mode",
severity="high",
description="Browser launched in headed mode (headless=False). This is fine for development but should be headless=True in production.",
line=self._find_line(r"headless\s*=\s*False"),
recommendation="Use headless=True for production. Toggle via environment variable: headless=os.environ.get('HEADLESS', 'true') == 'true'",
weight=SEVERITY_WEIGHTS["high"],
))
elif not self._has_pattern(r"headless"):
# Default is headless=True in Playwright, which is correct
self.findings.append(Finding(
category="Headless Mode",
severity="info",
description="Using default headless mode (True). Good for production.",
line=None,
recommendation="No action needed. Default headless=True is correct.",
weight=SEVERITY_WEIGHTS["info"],
))
def _check_user_agent(self):
"""Check if a custom user agent is set."""
has_ua = self._has_pattern(r"user_agent\s*=") or self._has_pattern(r"userAgent")
has_ua_list = self._has_pattern(r"USER_AGENTS?\s*=\s*\[")
has_random_ua = self._has_pattern(r"random\.choice.*(?:USER_AGENT|user_agent|ua)")
if not has_ua:
self.findings.append(Finding(
category="User Agent",
severity="critical",
description="No custom user agent configured. Playwright's default user agent contains 'HeadlessChrome' which is trivially detected.",
line=None,
recommendation="Set a realistic user agent: context = await browser.new_context(user_agent='Mozilla/5.0 ...')",
weight=SEVERITY_WEIGHTS["critical"],
))
elif has_ua_list and has_random_ua:
self.findings.append(Finding(
category="User Agent",
severity="info",
description="User agent rotation detected. Good anti-detection practice.",
line=self._find_line(r"USER_AGENTS?\s*=\s*\["),
recommendation="Ensure user agents are recent and match the browser being launched (e.g., Chrome UA for Chromium).",
weight=SEVERITY_WEIGHTS["info"],
))
elif has_ua:
self.findings.append(Finding(
category="User Agent",
severity="low",
description="Custom user agent set but no rotation detected. Single user agent is fingerprint-able at scale.",
line=self._find_line(r"user_agent\s*="),
recommendation="Rotate through 5-10 recent user agents using random.choice().",
weight=SEVERITY_WEIGHTS["low"],
))
def _check_viewport(self):
"""Check viewport configuration."""
has_viewport = self._has_pattern(r"viewport\s*=\s*\{") or self._has_pattern(r"viewport.*width")
if not has_viewport:
self.findings.append(Finding(
category="Viewport Size",
severity="high",
description="No viewport configured. Default Playwright viewport (1280x720) is common among bots. Sites may flag unusual viewport distributions.",
line=None,
recommendation="Set a common desktop viewport: viewport={'width': 1920, 'height': 1080}. Vary across runs.",
weight=SEVERITY_WEIGHTS["high"],
))
else:
# Check for suspiciously small viewports
match = re.search(r"width['\"]?\s*[:=]\s*(\d+)", self.content)
if match:
width = int(match.group(1))
if width < 1024:
self.findings.append(Finding(
category="Viewport Size",
severity="medium",
description=f"Viewport width {width}px is unusually small. Most desktop browsers are 1366px+ wide.",
line=self._find_line(r"width.*" + str(width)),
recommendation="Use 1366x768 (most common) or 1920x1080. Avoid unusual sizes like 800x600.",
weight=SEVERITY_WEIGHTS["medium"],
))
else:
self.findings.append(Finding(
category="Viewport Size",
severity="info",
description=f"Viewport width {width}px is reasonable.",
line=self._find_line(r"width.*" + str(width)),
recommendation="No action needed.",
weight=SEVERITY_WEIGHTS["info"],
))
def _check_webdriver_flag(self):
"""Check if navigator.webdriver is being removed."""
has_webdriver_override = (
self._has_pattern(r"navigator.*webdriver") or
self._has_pattern(r"webdriver.*undefined") or
self._has_pattern(r"add_init_script.*webdriver")
)
if not has_webdriver_override:
self.findings.append(Finding(
category="WebDriver Flag",
severity="critical",
description="navigator.webdriver is not overridden. This is the most common bot detection check. Every major anti-bot service tests this property.",
line=None,
recommendation=(
"Add init script to remove the flag:\n"
" await page.add_init_script(\"Object.defineProperty(navigator, 'webdriver', {get: () => undefined});\")"
),
weight=SEVERITY_WEIGHTS["critical"],
))
else:
self.findings.append(Finding(
category="WebDriver Flag",
severity="info",
description="navigator.webdriver override detected.",
line=self._find_line(r"webdriver"),
recommendation="No action needed.",
weight=SEVERITY_WEIGHTS["info"],
))
def _check_navigator_properties(self):
"""Check for additional navigator property hardening."""
checks = {
"plugins": (r"navigator.*plugins", "navigator.plugins is empty in headless mode. Real browsers report installed plugins."),
"languages": (r"navigator.*languages", "navigator.languages should be set to match the user agent locale."),
"platform": (r"navigator.*platform", "navigator.platform should match the user agent OS."),
}
overridden_count = 0
for prop, (pattern, desc) in checks.items():
if self._has_pattern(pattern):
overridden_count += 1
if overridden_count == 0:
self.findings.append(Finding(
category="Navigator Properties",
severity="medium",
description="No navigator property hardening detected. Advanced anti-bot services check plugins, languages, and platform properties.",
line=None,
recommendation="Override navigator.plugins, navigator.languages, and navigator.platform via add_init_script() to match realistic browser fingerprints.",
weight=SEVERITY_WEIGHTS["medium"],
))
elif overridden_count < 3:
self.findings.append(Finding(
category="Navigator Properties",
severity="low",
description=f"Partial navigator hardening ({overridden_count}/3 properties). Consider covering all three: plugins, languages, platform.",
line=None,
recommendation="Add overrides for any missing properties among: plugins, languages, platform.",
weight=SEVERITY_WEIGHTS["low"],
))
def _check_request_delays(self):
"""Check for human-like request delays."""
has_sleep = self._has_pattern(r"asyncio\.sleep") or self._has_pattern(r"wait_for_timeout")
has_random_delay = (
self._has_pattern(r"random\.(uniform|randint|random)") and has_sleep
)
if not has_sleep:
self.findings.append(Finding(
category="Request Timing",
severity="high",
description="No delays between actions detected. Machine-speed interactions are the easiest behavior-based detection signal.",
line=None,
recommendation="Add random delays between page interactions: await asyncio.sleep(random.uniform(0.5, 2.0))",
weight=SEVERITY_WEIGHTS["high"],
))
elif not has_random_delay:
self.findings.append(Finding(
category="Request Timing",
severity="medium",
description="Fixed delays detected but no randomization. Constant timing intervals are detectable patterns.",
line=self._find_line(r"(asyncio\.sleep|wait_for_timeout)"),
recommendation="Use random delays: random.uniform(min_seconds, max_seconds) instead of fixed values.",
weight=SEVERITY_WEIGHTS["medium"],
))
else:
self.findings.append(Finding(
category="Request Timing",
severity="info",
description="Randomized delays detected between actions.",
line=self._find_line(r"random\.(uniform|randint)"),
recommendation="No action needed. Ensure delays are realistic (0.5-3s for browsing, 1-5s for reading).",
weight=SEVERITY_WEIGHTS["info"],
))
def _check_error_handling(self):
"""Check for error handling patterns."""
has_try_except = self._has_pattern(r"try\s*:") and self._has_pattern(r"except")
has_retry = self._has_pattern(r"retr(y|ies)") or self._has_pattern(r"max_retries|max_attempts")
if not has_try_except:
self.findings.append(Finding(
category="Error Handling",
severity="medium",
description="No try/except blocks found. Unhandled errors will crash the automation and leave browser instances running.",
line=None,
recommendation="Wrap page interactions in try/except. Handle TimeoutError, network errors, and element-not-found gracefully.",
weight=SEVERITY_WEIGHTS["medium"],
))
elif not has_retry:
self.findings.append(Finding(
category="Error Handling",
severity="low",
description="Error handling present but no retry logic detected. Transient failures (network blips, slow loads) will cause data loss.",
line=None,
recommendation="Add retry with exponential backoff for network operations and element interactions.",
weight=SEVERITY_WEIGHTS["low"],
))
def _check_proxy(self):
"""Check for proxy configuration."""
has_proxy = self._has_pattern(r"proxy\s*=\s*\{") or self._has_pattern(r"proxy.*server")
if not has_proxy:
self.findings.append(Finding(
category="Proxy",
severity="low",
description="No proxy configuration detected. Running from a single IP address is fine for small jobs but will trigger rate limits at scale.",
line=None,
recommendation="For high-volume scraping, use rotating proxies: proxy={'server': 'http://proxy:port'}",
weight=SEVERITY_WEIGHTS["low"],
))
def _check_session_management(self):
"""Check for session/cookie management."""
has_storage_state = self._has_pattern(r"storage_state")
has_cookies = self._has_pattern(r"cookies\(\)") or self._has_pattern(r"add_cookies")
if not has_storage_state and not has_cookies:
self.findings.append(Finding(
category="Session Management",
severity="low",
description="No session persistence detected. Each run will start fresh, requiring re-authentication.",
line=None,
recommendation="Use storage_state() to save/restore sessions across runs. This avoids repeated logins that may trigger security alerts.",
weight=SEVERITY_WEIGHTS["low"],
))
def _check_browser_close(self):
"""Check if browser is properly closed."""
has_close = self._has_pattern(r"browser\.close\(\)") or self._has_pattern(r"await.*close")
has_context_manager = self._has_pattern(r"async\s+with\s+async_playwright")
if not has_close and not has_context_manager:
self.findings.append(Finding(
category="Resource Cleanup",
severity="medium",
description="No browser.close() or context manager detected. Browser processes will leak on failure.",
line=None,
recommendation="Use 'async with async_playwright() as p:' or ensure browser.close() is in a finally block.",
weight=SEVERITY_WEIGHTS["medium"],
))
def _check_stealth_imports(self):
"""Check for stealth/anti-detection library usage."""
has_stealth = self._has_pattern(r"playwright_stealth|stealth_async|undetected")
if has_stealth:
self.findings.append(Finding(
category="Stealth Library",
severity="info",
description="Third-party stealth library detected. These provide additional fingerprint evasion but add dependencies.",
line=self._find_line(r"playwright_stealth|stealth_async|undetected"),
recommendation="Stealth libraries are helpful but not a silver bullet. Still implement manual checks for user agent, viewport, and timing.",
weight=SEVERITY_WEIGHTS["info"],
))
def get_risk_score(self) -> int:
"""Calculate overall risk score (0-100). Higher = more detectable."""
raw_score = sum(f.weight for f in self.findings)
# Cap at 100
return min(raw_score, 100)
def get_risk_level(self) -> str:
"""Get human-readable risk level."""
score = self.get_risk_score()
if score <= 10:
return "LOW"
elif score <= 30:
return "MODERATE"
elif score <= 50:
return "HIGH"
else:
return "CRITICAL"
def get_summary(self) -> dict:
"""Get a summary of the analysis."""
severity_counts = {"critical": 0, "high": 0, "medium": 0, "low": 0, "info": 0}
for f in self.findings:
severity_counts[f.severity] += 1
return {
"file": self.file_path,
"risk_score": self.get_risk_score(),
"risk_level": self.get_risk_level(),
"total_findings": len(self.findings),
"severity_counts": severity_counts,
"actionable_findings": len([f for f in self.findings if f.severity != "info"]),
}
def format_text_report(checker: AntiDetectionChecker, verbose: bool = False) -> str:
"""Format findings as human-readable text."""
lines = []
summary = checker.get_summary()
lines.append("=" * 60)
lines.append(" ANTI-DETECTION AUDIT REPORT")
lines.append("=" * 60)
lines.append(f"File: {summary['file']}")
lines.append(f"Risk Score: {summary['risk_score']}/100 ({summary['risk_level']})")
lines.append(f"Total Issues: {summary['actionable_findings']} actionable, {summary['severity_counts']['info']} info")
lines.append("")
# Severity breakdown
for sev in ["critical", "high", "medium", "low"]:
count = summary["severity_counts"][sev]
if count > 0:
lines.append(f" {sev.upper():10s} {count}")
lines.append("")
# Findings grouped by severity
severity_order = ["critical", "high", "medium", "low"]
if verbose:
severity_order.append("info")
for sev in severity_order:
sev_findings = [f for f in checker.findings if f.severity == sev]
if not sev_findings:
continue
lines.append(f"--- {sev.upper()} ---")
for f in sev_findings:
line_info = f" (line {f.line})" if f.line else ""
lines.append(f" [{f.category}]{line_info}")
lines.append(f" {f.description}")
lines.append(f" Fix: {f.recommendation}")
lines.append("")
# Exit code guidance
lines.append("-" * 60)
score = summary["risk_score"]
if score <= 10:
lines.append("Result: PASS - Low detection risk.")
elif score <= 30:
lines.append("Result: PASS with warnings - Address medium/high issues for production use.")
else:
lines.append("Result: FAIL - High detection risk. Fix critical and high issues before deploying.")
lines.append("")
return "\n".join(lines)
def main():
parser = argparse.ArgumentParser(
description="Audit a Playwright script for common bot detection vectors.",
epilog=(
"Examples:\n"
" %(prog)s --file scraper.py\n"
" %(prog)s --file scraper.py --verbose\n"
" %(prog)s --file scraper.py --json\n"
"\n"
"Exit codes:\n"
" 0 - Low risk (score 0-10)\n"
" 1 - Moderate to high risk (score 11-50)\n"
" 2 - Critical risk (score 51+)\n"
),
formatter_class=argparse.RawDescriptionHelpFormatter,
)
parser.add_argument(
"--file",
required=True,
help="Path to the Playwright script to audit",
)
parser.add_argument(
"--json",
action="store_true",
dest="json_output",
default=False,
help="Output results as JSON",
)
parser.add_argument(
"--verbose",
action="store_true",
default=False,
help="Include informational (non-actionable) findings in output",
)
args = parser.parse_args()
file_path = os.path.abspath(args.file)
if not os.path.isfile(file_path):
print(f"Error: File not found: {file_path}", file=sys.stderr)
sys.exit(2)
try:
with open(file_path, "r", encoding="utf-8") as f:
content = f.read()
except Exception as e:
print(f"Error reading file: {e}", file=sys.stderr)
sys.exit(2)
if not content.strip():
print("Error: File is empty.", file=sys.stderr)
sys.exit(2)
checker = AntiDetectionChecker(content, file_path)
checker.check_all()
if args.json_output:
output = checker.get_summary()
output["findings"] = [asdict(f) for f in checker.findings]
if not args.verbose:
output["findings"] = [f for f in output["findings"] if f["severity"] != "info"]
print(json.dumps(output, indent=2))
else:
print(format_text_report(checker, verbose=args.verbose))
# Exit code based on risk
score = checker.get_risk_score()
if score <= 10:
sys.exit(0)
elif score <= 50:
sys.exit(1)
else:
sys.exit(2)
if __name__ == "__main__":
main()
FILE:scripts/form_automation_builder.py
#!/usr/bin/env python3
"""
Form Automation Builder - Generates Playwright form-fill automation scripts.
Takes a JSON field specification and target URL, then produces a ready-to-run
Playwright script that fills forms, handles multi-step flows, and manages
file uploads.
No external dependencies - uses only Python standard library.
"""
import argparse
import json
import os
import sys
import textwrap
from datetime import datetime
SUPPORTED_FIELD_TYPES = {
"text": "page.fill('{selector}', '{value}')",
"password": "page.fill('{selector}', '{value}')",
"email": "page.fill('{selector}', '{value}')",
"textarea": "page.fill('{selector}', '{value}')",
"select": "page.select_option('{selector}', value='{value}')",
"checkbox": "page.check('{selector}')" if True else "page.uncheck('{selector}')",
"radio": "page.check('{selector}')",
"file": "page.set_input_files('{selector}', '{value}')",
"click": "page.click('{selector}')",
}
def validate_fields(fields):
"""Validate the field specification format. Returns list of issues."""
issues = []
if not isinstance(fields, list):
issues.append("Top-level structure must be a JSON array of field objects.")
return issues
for i, field in enumerate(fields):
if not isinstance(field, dict):
issues.append(f"Field {i}: must be a JSON object.")
continue
if "selector" not in field:
issues.append(f"Field {i}: missing required 'selector' key.")
if "type" not in field:
issues.append(f"Field {i}: missing required 'type' key.")
elif field["type"] not in SUPPORTED_FIELD_TYPES:
issues.append(
f"Field {i}: unsupported type '{field['type']}'. "
f"Supported: {', '.join(sorted(SUPPORTED_FIELD_TYPES.keys()))}"
)
if field.get("type") not in ("checkbox", "radio", "click") and "value" not in field:
issues.append(f"Field {i}: missing 'value' for type '{field.get('type', '?')}'.")
return issues
def generate_field_action(field, indent=8):
"""Generate the Playwright action line for a single field."""
ftype = field["type"]
selector = field["selector"]
value = field.get("value", "")
label = field.get("label", selector)
prefix = " " * indent
lines = []
lines.append(f'{prefix}# {label}')
if ftype == "checkbox":
if field.get("value", "true").lower() in ("true", "yes", "1", "on"):
lines.append(f'{prefix}await page.check("{selector}")')
else:
lines.append(f'{prefix}await page.uncheck("{selector}")')
elif ftype == "radio":
lines.append(f'{prefix}await page.check("{selector}")')
elif ftype == "click":
lines.append(f'{prefix}await page.click("{selector}")')
elif ftype == "select":
lines.append(f'{prefix}await page.select_option("{selector}", value="{value}")')
elif ftype == "file":
lines.append(f'{prefix}await page.set_input_files("{selector}", "{value}")')
else:
# text, password, email, textarea
lines.append(f'{prefix}await page.fill("{selector}", "{value}")')
# Add optional wait_after
wait_after = field.get("wait_after")
if wait_after:
lines.append(f'{prefix}await page.wait_for_selector("{wait_after}")')
return "\n".join(lines)
def build_form_script(url, fields, output_format="script"):
"""Build a Playwright form automation script from the field specification."""
issues = validate_fields(fields)
if issues:
return None, issues
if output_format == "json":
config = {
"url": url,
"fields": fields,
"field_count": len(fields),
"field_types": list(set(f["type"] for f in fields)),
"has_file_upload": any(f["type"] == "file" for f in fields),
"generated_at": datetime.now().isoformat(),
}
return config, None
# Group fields into steps if step markers are present
steps = {}
for field in fields:
step = field.get("step", 1)
if step not in steps:
steps[step] = []
steps[step].append(field)
multi_step = len(steps) > 1
# Generate step functions
step_functions = []
for step_num in sorted(steps.keys()):
step_fields = steps[step_num]
actions = "\n".join(generate_field_action(f) for f in step_fields)
if multi_step:
fn = textwrap.dedent(f"""\
async def fill_step_{step_num}(page):
\"\"\"Fill form step {step_num} ({len(step_fields)} fields).\"\"\"
print(f"Filling step {step_num}...")
{actions}
print(f"Step {step_num} complete.")
""")
else:
fn = textwrap.dedent(f"""\
async def fill_form(page):
\"\"\"Fill form ({len(step_fields)} fields).\"\"\"
print("Filling form...")
{actions}
print("Form filled.")
""")
step_functions.append(fn)
step_functions_str = "\n\n".join(step_functions)
# Generate main() call sequence
if multi_step:
step_calls = "\n".join(
f" await fill_step_{n}(page)" for n in sorted(steps.keys())
)
else:
step_calls = " await fill_form(page)"
submit_selector = None
for field in fields:
if field.get("type") == "click" and field.get("is_submit"):
submit_selector = field["selector"]
break
submit_block = ""
if submit_selector:
submit_block = textwrap.dedent(f"""\
# Submit
await page.click("{submit_selector}")
await page.wait_for_load_state("networkidle")
print("Form submitted.")
""")
script = textwrap.dedent(f'''\
#!/usr/bin/env python3
"""
Auto-generated Playwright form automation script.
Target: {url}
Fields: {len(fields)}
Steps: {len(steps)}
Generated: {datetime.now().isoformat()}
Requirements:
pip install playwright
playwright install chromium
"""
import asyncio
import random
from playwright.async_api import async_playwright
URL = "{url}"
USER_AGENTS = [
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
]
{step_functions_str}
async def main():
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
context = await browser.new_context(
viewport={{"width": 1920, "height": 1080}},
user_agent=random.choice(USER_AGENTS),
)
page = await context.new_page()
await page.add_init_script(
"Object.defineProperty(navigator, \'webdriver\', {{get: () => undefined}});"
)
print(f"Navigating to {{URL}}...")
await page.goto(URL, wait_until="networkidle")
{step_calls}
{submit_block}
print("Automation complete.")
await browser.close()
if __name__ == "__main__":
asyncio.run(main())
''')
return script, None
def main():
parser = argparse.ArgumentParser(
description="Generate Playwright form-fill automation scripts from a JSON field specification.",
epilog=textwrap.dedent("""\
Examples:
%(prog)s --url https://example.com/signup --fields fields.json
%(prog)s --url https://example.com/signup --fields fields.json --output fill_form.py
%(prog)s --url https://example.com/signup --fields fields.json --json
Field specification format (fields.json):
[
{"selector": "#email", "type": "email", "value": "user@example.com", "label": "Email"},
{"selector": "#password", "type": "password", "value": "s3cret"},
{"selector": "#country", "type": "select", "value": "US"},
{"selector": "#terms", "type": "checkbox", "value": "true"},
{"selector": "#avatar", "type": "file", "value": "/path/to/photo.jpg"},
{"selector": "button[type='submit']", "type": "click", "is_submit": true}
]
Supported field types: text, password, email, textarea, select, checkbox, radio, file, click
Multi-step forms: Add "step": N to each field to group into steps.
"""),
formatter_class=argparse.RawDescriptionHelpFormatter,
)
parser.add_argument(
"--url",
required=True,
help="Target form URL",
)
parser.add_argument(
"--fields",
required=True,
help="Path to JSON file containing field specifications",
)
parser.add_argument(
"--output",
help="Output file path (default: stdout)",
)
parser.add_argument(
"--json",
action="store_true",
dest="json_output",
default=False,
help="Output JSON configuration instead of Python script",
)
args = parser.parse_args()
# Load fields
fields_path = os.path.abspath(args.fields)
if not os.path.isfile(fields_path):
print(f"Error: Fields file not found: {fields_path}", file=sys.stderr)
sys.exit(2)
try:
with open(fields_path, "r") as f:
fields = json.load(f)
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON in {fields_path}: {e}", file=sys.stderr)
sys.exit(2)
output_format = "json" if args.json_output else "script"
result, errors = build_form_script(
url=args.url,
fields=fields,
output_format=output_format,
)
if errors:
print("Validation errors:", file=sys.stderr)
for err in errors:
print(f" - {err}", file=sys.stderr)
sys.exit(2)
if args.json_output:
output_text = json.dumps(result, indent=2)
else:
output_text = result
if args.output:
output_path = os.path.abspath(args.output)
with open(output_path, "w") as f:
f.write(output_text)
if not args.json_output:
os.chmod(output_path, 0o755)
print(f"Written to {output_path}", file=sys.stderr)
sys.exit(0)
else:
print(output_text)
sys.exit(0)
if __name__ == "__main__":
main()
FILE:scripts/scraping_toolkit.py
#!/usr/bin/env python3
"""
Scraping Toolkit - Generates Playwright scraping script skeletons.
Takes a URL pattern and CSS selectors as input and produces a ready-to-run
Playwright scraping script with pagination support, error handling, and
anti-detection patterns baked in.
No external dependencies - uses only Python standard library.
"""
import argparse
import json
import os
import sys
import textwrap
from datetime import datetime
def build_scraping_script(url, selectors, paginate=False, output_format="script"):
"""Build a Playwright scraping script from the given parameters."""
selector_list = [s.strip() for s in selectors.split(",") if s.strip()]
if not selector_list:
return None, "No valid selectors provided."
field_names = []
for sel in selector_list:
# Derive field name from selector: .product-title -> product_title
name = sel.strip("#.[]()>:+~ ")
name = name.replace("-", "_").replace(" ", "_").replace(".", "_")
# Remove non-alphanumeric
name = "".join(c if c.isalnum() or c == "_" else "" for c in name)
if not name:
name = f"field_{len(field_names)}"
field_names.append(name)
field_map = dict(zip(field_names, selector_list))
if output_format == "json":
config = {
"url": url,
"selectors": field_map,
"pagination": {
"enabled": paginate,
"next_selector": "a:has-text('Next'), button:has-text('Next')",
"max_pages": 50,
},
"anti_detection": {
"random_delay_ms": [800, 2500],
"user_agent_rotation": True,
"viewport": {"width": 1920, "height": 1080},
},
"output": {
"format": "jsonl",
"deduplicate_by": field_names[0] if field_names else None,
},
"generated_at": datetime.now().isoformat(),
}
return config, None
# Build Python script
fields_dict_str = "{\n"
for name, sel in field_map.items():
fields_dict_str += f' "{name}": "{sel}",\n'
fields_dict_str += " }"
pagination_block = ""
if paginate:
pagination_block = textwrap.dedent("""\
# --- Pagination ---
async def scrape_all_pages(page, container, fields, next_sel, max_pages=50):
all_items = []
for page_num in range(max_pages):
print(f"Scraping page {page_num + 1}...")
items = await extract_items(page, container, fields)
all_items.extend(items)
next_btn = page.locator(next_sel)
if await next_btn.count() == 0:
break
try:
is_disabled = await next_btn.is_disabled()
except Exception:
is_disabled = True
if is_disabled:
break
await next_btn.click()
await page.wait_for_load_state("networkidle")
await asyncio.sleep(random.uniform(0.8, 2.5))
return all_items
""")
main_call = "scrape_all_pages(page, CONTAINER, FIELDS, NEXT_SELECTOR)" if paginate else "extract_items(page, CONTAINER, FIELDS)"
script = textwrap.dedent(f'''\
#!/usr/bin/env python3
"""
Auto-generated Playwright scraping script.
Target: {url}
Generated: {datetime.now().isoformat()}
Requirements:
pip install playwright
playwright install chromium
"""
import asyncio
import json
import random
from playwright.async_api import async_playwright
# --- Configuration ---
URL = "{url}"
CONTAINER = "body" # Adjust to the repeating item container selector
FIELDS = {fields_dict_str}
NEXT_SELECTOR = "a:has-text('Next'), button:has-text('Next')"
USER_AGENTS = [
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
"Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
]
async def extract_items(page, container_selector, field_map):
"""Extract structured data from repeating elements."""
items = []
cards = await page.query_selector_all(container_selector)
for card in cards:
item = {{}}
for name, selector in field_map.items():
el = await card.query_selector(selector)
if el:
item[name] = (await el.text_content() or "").strip()
else:
item[name] = None
items.append(item)
return items
{pagination_block}
async def main():
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
context = await browser.new_context(
viewport={{"width": 1920, "height": 1080}},
user_agent=random.choice(USER_AGENTS),
)
page = await context.new_page()
# Remove WebDriver flag
await page.add_init_script(
"Object.defineProperty(navigator, \'webdriver\', {{get: () => undefined}});"
)
print(f"Navigating to {{URL}}...")
await page.goto(URL, wait_until="networkidle")
data = await {main_call}
print(json.dumps(data, indent=2, ensure_ascii=False))
await browser.close()
if __name__ == "__main__":
asyncio.run(main())
''')
return script, None
def main():
parser = argparse.ArgumentParser(
description="Generate Playwright scraping script skeletons from URL and selectors.",
epilog=(
"Examples:\n"
" %(prog)s --url https://example.com/products --selectors '.title,.price,.rating'\n"
" %(prog)s --url https://example.com/search --selectors '.name,.desc' --paginate\n"
" %(prog)s --url https://example.com --selectors '.item' --json\n"
" %(prog)s --url https://example.com --selectors '.item' --output scraper.py\n"
),
formatter_class=argparse.RawDescriptionHelpFormatter,
)
parser.add_argument(
"--url",
required=True,
help="Target URL to scrape",
)
parser.add_argument(
"--selectors",
required=True,
help="Comma-separated CSS selectors for data fields (e.g. '.title,.price,.rating')",
)
parser.add_argument(
"--paginate",
action="store_true",
default=False,
help="Include pagination handling in generated script",
)
parser.add_argument(
"--output",
help="Output file path (default: stdout)",
)
parser.add_argument(
"--json",
action="store_true",
dest="json_output",
default=False,
help="Output JSON configuration instead of Python script",
)
args = parser.parse_args()
output_format = "json" if args.json_output else "script"
result, error = build_scraping_script(
url=args.url,
selectors=args.selectors,
paginate=args.paginate,
output_format=output_format,
)
if error:
print(f"Error: {error}", file=sys.stderr)
sys.exit(2)
if args.json_output:
output_text = json.dumps(result, indent=2)
else:
output_text = result
if args.output:
output_path = os.path.abspath(args.output)
with open(output_path, "w") as f:
f.write(output_text)
if not args.json_output:
os.chmod(output_path, 0o755)
print(f"Written to {output_path}", file=sys.stderr)
sys.exit(0)
else:
print(output_text)
sys.exit(0)
if __name__ == "__main__":
main()
Chạy kiểm thử trên BrowserStack: kiểm thử đa trình duyệt, đám mây và tương thích trình duyệt.
---
name: "browserstack"
description: >-
Run tests on BrowserStack. Use when user mentions "browserstack",
"cross-browser", "cloud testing", "browser matrix", "test on safari",
"test on firefox", or "browser compatibility".
---
# BrowserStack Integration
Run Playwright tests on BrowserStack's cloud grid for cross-browser and cross-device testing.
## Prerequisites
Environment variables must be set:
- `BROWSERSTACK_USERNAME` — your BrowserStack username
- `BROWSERSTACK_ACCESS_KEY` — your access key
If not set, inform the user how to get them from [browserstack.com/accounts/settings](https://www.browserstack.com/accounts/settings) and stop.
## Capabilities
### 1. Configure for BrowserStack
```
/pw:browserstack setup
```
Steps:
1. Check current `playwright.config.ts`
2. Add BrowserStack connect options:
```typescript
// Add to playwright.config.ts
import { defineConfig } from '@playwright/test';
const isBS = !!process.env.BROWSERSTACK_USERNAME;
export default defineConfig({
// ... existing config
projects: isBS ? [
{
name: "chromelatestwindows-11",
use: {
connectOptions: {
wsEndpoint: `wss://cdp.browserstack.com/playwright?caps='chrome',
'browser_version': 'latest',
'os': 'Windows',
'os_version': '11',
'browserstack.username': process.env.BROWSERSTACK_USERNAME,
'browserstack.accessKey': process.env.BROWSERSTACK_ACCESS_KEY,))}`,
},
},
},
{
name: "firefoxlatestwindows-11",
use: {
connectOptions: {
wsEndpoint: `wss://cdp.browserstack.com/playwright?caps='playwright-firefox',
'browser_version': 'latest',
'os': 'Windows',
'os_version': '11',
'browserstack.username': process.env.BROWSERSTACK_USERNAME,
'browserstack.accessKey': process.env.BROWSERSTACK_ACCESS_KEY,))}`,
},
},
},
{
name: "webkitlatestos-x-ventura",
use: {
connectOptions: {
wsEndpoint: `wss://cdp.browserstack.com/playwright?caps='playwright-webkit',
'browser_version': 'latest',
'os': 'OS X',
'os_version': 'Ventura',
'browserstack.username': process.env.BROWSERSTACK_USERNAME,
'browserstack.accessKey': process.env.BROWSERSTACK_ACCESS_KEY,))}`,
},
},
},
] : [
// ... local projects fallback
],
});
```
3. Add npm script: `"test:e2e:cloud": "npx playwright test --project='chrome@*' --project='firefox@*' --project='webkit@*'"`
### 2. Run Tests on BrowserStack
```
/pw:browserstack run
```
Steps:
1. Verify credentials are set
2. Run tests with BrowserStack projects:
```bash
BROWSERSTACK_USERNAME=$BROWSERSTACK_USERNAME \
BROWSERSTACK_ACCESS_KEY=$BROWSERSTACK_ACCESS_KEY \
npx playwright test --project='chrome@*' --project='firefox@*'
```
3. Monitor execution
4. Report results per browser
### 3. Get Build Results
```
/pw:browserstack results
```
Steps:
1. Call `browserstack_get_builds` MCP tool
2. Get latest build's sessions
3. For each session:
- Status (pass/fail)
- Browser and OS
- Duration
- Video URL
- Log URLs
4. Format as summary table
### 4. Check Available Browsers
```
/pw:browserstack browsers
```
Steps:
1. Call `browserstack_get_browsers` MCP tool
2. Filter for Playwright-compatible browsers
3. Display available browser/OS combinations
### 5. Local Testing
```
/pw:browserstack local
```
For testing localhost or staging behind firewall:
1. Install BrowserStack Local: `npm install -D browserstack-local`
2. Add local tunnel to config
3. Provide setup instructions
## MCP Tools Used
| Tool | When |
|---|---|
| `browserstack_get_plan` | Check account limits |
| `browserstack_get_browsers` | List available browsers |
| `browserstack_get_builds` | List recent builds |
| `browserstack_get_sessions` | Get sessions in a build |
| `browserstack_get_session` | Get session details (video, logs) |
| `browserstack_update_session` | Mark pass/fail |
| `browserstack_get_logs` | Get text/network logs |
## Output
- Cross-browser test results table
- Per-browser pass/fail status
- Links to BrowserStack dashboard for video/screenshots
- Any browser-specific failures highlighted
Quản lý CAPA cho QMS thiết bị y tế: phân tích nguyên nhân gốc, hành động khắc phục và kiểm chứng hiệu quả.
---
name: "capa-officer"
description: CAPA system management for medical device QMS. Covers root cause analysis, corrective action planning, effectiveness verification, and CAPA metrics. Use for CAPA investigations, 5-Why analysis, fishbone diagrams, root cause determination, corrective action tracking, effectiveness verification, or CAPA program optimization.
triggers:
- CAPA investigation
- root cause analysis
- 5 Why analysis
- fishbone diagram
- corrective action
- preventive action
- effectiveness verification
- CAPA metrics
- nonconformance investigation
- quality issue investigation
- CAPA tracking
- audit finding CAPA
---
# CAPA Officer
Corrective and Preventive Action (CAPA) management within Quality Management Systems, focusing on systematic root cause analysis, action implementation, and effectiveness verification.
---
## Table of Contents
- [CAPA Investigation Workflow](#capa-investigation-workflow)
- [Root Cause Analysis](#root-cause-analysis)
- [Corrective Action Planning](#corrective-action-planning)
- [Effectiveness Verification](#effectiveness-verification)
- [CAPA Metrics and Reporting](#capa-metrics-and-reporting)
- [Reference Documentation](#reference-documentation)
- [Tools](#tools)
---
## CAPA Investigation Workflow
Conduct systematic CAPA investigation from initiation through closure:
1. Document trigger event with objective evidence
2. Assess significance and determine CAPA necessity
3. Form investigation team with relevant expertise
4. Collect data and evidence systematically
5. Select and apply appropriate RCA methodology
6. Identify root cause(s) with supporting evidence
7. Develop corrective and preventive actions
8. **Validation:** Root cause explains all symptoms; if eliminated, problem would not recur
### CAPA Necessity Determination
| Trigger Type | CAPA Required | Criteria |
|--------------|---------------|----------|
| Customer complaint (safety) | Yes | Any complaint involving patient/user safety |
| Customer complaint (quality) | Evaluate | Based on severity and frequency |
| Internal audit finding (Major) | Yes | Systematic failure or absence of element |
| Internal audit finding (Minor) | Recommended | Isolated lapse or partial implementation |
| Nonconformance (recurring) | Yes | Same NC type occurring 3+ times |
| Nonconformance (isolated) | Evaluate | Based on severity and risk |
| External audit finding | Yes | All Major and Minor findings |
| Trend analysis | Evaluate | Based on trend significance |
### Investigation Team Composition
| CAPA Severity | Required Team Members |
|---------------|----------------------|
| Critical | CAPA Officer, Process Owner, QA Manager, Subject Matter Expert, Management Rep |
| Major | CAPA Officer, Process Owner, Subject Matter Expert |
| Minor | CAPA Officer, Process Owner |
### Evidence Collection Checklist
- [ ] Problem description with specific details (what, where, when, who, how much)
- [ ] Timeline of events leading to issue
- [ ] Relevant records and documentation
- [ ] Interview notes from involved personnel
- [ ] Photos or physical evidence (if applicable)
- [ ] Related complaints, NCs, or previous CAPAs
- [ ] Process parameters and specifications
---
## Root Cause Analysis
Select and apply appropriate RCA methodology based on problem characteristics.
### RCA Method Selection Decision Tree
```
Is the issue safety-critical or involves system reliability?
├── Yes → Use FAULT TREE ANALYSIS
└── No → Is human error the suspected primary cause?
├── Yes → Use HUMAN FACTORS ANALYSIS
└── No → How many potential contributing factors?
├── 1-2 factors (linear causation) → Use 5 WHY ANALYSIS
├── 3-6 factors (complex, systemic) → Use FISHBONE DIAGRAM
└── Unknown/proactive assessment → Use FMEA
```
### 5 Why Analysis
Use when: Single-cause issues with linear causation, process deviations with clear failure point.
**Template:**
```
PROBLEM: [Clear, specific statement]
WHY 1: Why did [problem] occur?
BECAUSE: [First-level cause]
EVIDENCE: [Supporting data]
WHY 2: Why did [first-level cause] occur?
BECAUSE: [Second-level cause]
EVIDENCE: [Supporting data]
WHY 3: Why did [second-level cause] occur?
BECAUSE: [Third-level cause]
EVIDENCE: [Supporting data]
WHY 4: Why did [third-level cause] occur?
BECAUSE: [Fourth-level cause]
EVIDENCE: [Supporting data]
WHY 5: Why did [fourth-level cause] occur?
BECAUSE: [Root cause]
EVIDENCE: [Supporting data]
```
**Example - Calibration Overdue:**
```
PROBLEM: pH meter (EQ-042) found 2 months overdue for calibration
WHY 1: Why was calibration overdue?
BECAUSE: Equipment was not on calibration schedule
EVIDENCE: Calibration schedule reviewed, EQ-042 not listed
WHY 2: Why was it not on the schedule?
BECAUSE: Schedule not updated when equipment was purchased
EVIDENCE: Purchase date 2023-06-15, schedule dated 2023-01-01
WHY 3: Why was the schedule not updated?
BECAUSE: No process requires schedule update at equipment purchase
EVIDENCE: SOP-EQ-001 reviewed, no such requirement
WHY 4: Why is there no such requirement?
BECAUSE: Procedure written before equipment tracking was centralized
EVIDENCE: SOP last revised 2019, equipment system implemented 2021
WHY 5: Why has procedure not been updated?
BECAUSE: Periodic review did not assess compatibility with new systems
EVIDENCE: No review against new equipment system documented
ROOT CAUSE: Procedure review process does not assess compatibility
with organizational systems implemented after original procedure creation.
```
### Fishbone Diagram Categories (6M)
| Category | Focus Areas | Typical Causes |
|----------|-------------|----------------|
| Man (People) | Training, competency, workload | Skill gaps, fatigue, communication |
| Machine (Equipment) | Calibration, maintenance, age | Wear, malfunction, inadequate capacity |
| Method (Process) | Procedures, work instructions | Unclear steps, missing controls |
| Material | Specifications, suppliers, storage | Out-of-spec, degradation, contamination |
| Measurement | Calibration, methods, interpretation | Instrument error, wrong method |
| Mother Nature | Temperature, humidity, cleanliness | Environmental excursions |
See `references/rca-methodologies.md` for complete method details and templates.
### Root Cause Validation
Before proceeding to action planning, validate root cause:
- [ ] Root cause can be verified with objective evidence
- [ ] If root cause is eliminated, problem would not recur
- [ ] Root cause is within organizational control
- [ ] Root cause explains all observed symptoms
- [ ] No other significant causes remain unaddressed
---
## Corrective Action Planning
Develop effective actions addressing identified root causes:
1. Define immediate containment actions
2. Develop corrective actions targeting root cause
3. Identify preventive actions for similar processes
4. Assign responsibilities and resources
5. Establish timeline with milestones
6. Define success criteria and verification method
7. Document in CAPA action plan
8. **Validation:** Actions directly address root cause; success criteria are measurable
### Action Types
| Type | Purpose | Timeline | Example |
|------|---------|----------|---------|
| Containment | Stop immediate impact | 24-72 hours | Quarantine affected product |
| Correction | Fix the specific occurrence | 1-2 weeks | Rework or replace affected items |
| Corrective | Eliminate root cause | 30-90 days | Revise procedure, add controls |
| Preventive | Prevent in other areas | 60-120 days | Extend solution to similar processes |
### Action Plan Components
```
ACTION PLAN TEMPLATE
CAPA Number: [CAPA-XXXX]
Root Cause: [Identified root cause]
ACTION 1: [Specific action description]
- Type: [ ] Containment [ ] Correction [ ] Corrective [ ] Preventive
- Responsible: [Name, Title]
- Due Date: [YYYY-MM-DD]
- Resources: [Required resources]
- Success Criteria: [Measurable outcome]
- Verification Method: [How success will be verified]
ACTION 2: [Specific action description]
...
IMPLEMENTATION TIMELINE:
Week 1: [Milestone]
Week 2: [Milestone]
Week 4: [Milestone]
Week 8: [Milestone]
APPROVAL:
CAPA Owner: _____________ Date: _______
Process Owner: _____________ Date: _______
QA Manager: _____________ Date: _______
```
### Action Effectiveness Indicators
| Indicator | Target | Red Flag |
|-----------|--------|----------|
| Action scope | Addresses root cause completely | Treats only symptoms |
| Specificity | Measurable deliverables | Vague commitments |
| Timeline | Aggressive but achievable | No due dates or unrealistic |
| Resources | Identified and allocated | Not specified |
| Sustainability | Permanent solution | Temporary fix |
---
## Effectiveness Verification
Verify corrective actions achieved intended results:
1. Allow adequate implementation period (minimum 30-90 days)
2. Collect post-implementation data
3. Compare to pre-implementation baseline
4. Evaluate against success criteria
5. Verify no recurrence during verification period
6. Document verification evidence
7. Determine CAPA effectiveness
8. **Validation:** All criteria met with objective evidence; no recurrence observed
### Verification Timeline Guidelines
| CAPA Severity | Wait Period | Verification Window |
|---------------|-------------|---------------------|
| Critical | 30 days | 30-90 days post-implementation |
| Major | 60 days | 60-180 days post-implementation |
| Minor | 90 days | 90-365 days post-implementation |
### Verification Methods
| Method | Use When | Evidence Required |
|--------|----------|-------------------|
| Data trend analysis | Quantifiable issues | Pre/post comparison, trend charts |
| Process audit | Procedure compliance issues | Audit checklist, interview notes |
| Record review | Documentation issues | Sample records, compliance rate |
| Testing/inspection | Product quality issues | Test results, pass/fail data |
| Interview/observation | Training issues | Interview notes, observation records |
### Effectiveness Determination
```
Did recurrence occur during verification period?
├── Yes → CAPA INEFFECTIVE (re-investigate root cause)
└── No → Were all effectiveness criteria met?
├── Yes → CAPA EFFECTIVE (proceed to closure)
└── No → Extent of gap?
├── Minor gap → Extend verification or accept with justification
└── Significant gap → CAPA INEFFECTIVE (revise actions)
```
See `references/effectiveness-verification-guide.md` for detailed procedures.
---
## CAPA Metrics and Reporting
Monitor CAPA program performance through key indicators.
### Key Performance Indicators
| Metric | Target | Calculation |
|--------|--------|-------------|
| CAPA cycle time | <60 days average | (Close Date - Open Date) / Number of CAPAs |
| Overdue rate | <10% | Overdue CAPAs / Total Open CAPAs |
| First-time effectiveness | >90% | Effective on first verification / Total verified |
| Recurrence rate | <5% | Recurred issues / Total closed CAPAs |
| Investigation quality | 100% root cause validated | Root causes validated / Total CAPAs |
### Aging Analysis Categories
| Age Bucket | Status | Action Required |
|------------|--------|-----------------|
| 0-30 days | On track | Monitor progress |
| 31-60 days | Monitor | Review for delays |
| 61-90 days | Warning | Escalate to management |
| >90 days | Critical | Management intervention required |
### Management Review Inputs
Monthly CAPA status report includes:
- Open CAPA count by severity and status
- Overdue CAPA list with owners
- Cycle time trends
- Effectiveness rate trends
- Source analysis (complaints, audits, NCs)
- Recommendations for improvement
---
## Reference Documentation
### Root Cause Analysis Methodologies
`references/rca-methodologies.md` contains:
- Method selection decision tree
- 5 Why analysis template and example
- Fishbone diagram categories and template
- Fault Tree Analysis for safety-critical issues
- Human Factors Analysis for people-related causes
- FMEA for proactive risk assessment
- Hybrid approach guidance
### Effectiveness Verification Guide
`references/effectiveness-verification-guide.md` contains:
- Verification planning requirements
- Verification method selection
- Effectiveness criteria definition (SMART)
- Closure requirements by severity
- Ineffective CAPA process
- Documentation templates
---
## Tools
### CAPA Tracker
```bash
# Generate CAPA status report
python scripts/capa_tracker.py --capas capas.json
# Interactive mode for manual entry
python scripts/capa_tracker.py --interactive
# JSON output for integration
python scripts/capa_tracker.py --capas capas.json --output json
# Generate sample data file
python scripts/capa_tracker.py --sample > sample_capas.json
```
Calculates and reports:
- Summary metrics (open, closed, overdue, cycle time, effectiveness)
- Status distribution
- Severity and source analysis
- Aging report by time bucket
- Overdue CAPA list
- Actionable recommendations
### Sample CAPA Input
```json
{
"capas": [
{
"capa_number": "CAPA-2024-001",
"title": "Calibration overdue for pH meter",
"description": "pH meter EQ-042 found 2 months overdue",
"source": "AUDIT",
"severity": "MAJOR",
"status": "VERIFICATION",
"open_date": "2024-06-15",
"target_date": "2024-08-15",
"owner": "J. Smith",
"root_cause": "Procedure review gap",
"corrective_action": "Updated SOP-EQ-001"
}
]
}
```
---
## Regulatory Requirements
### ISO 13485:2016 Clause 8.5
| Sub-clause | Requirement | Key Activities |
|------------|-------------|----------------|
| 8.5.2 Corrective Action | Eliminate cause of nonconformity | NC review, cause determination, action evaluation, implementation, effectiveness review |
| 8.5.3 Preventive Action | Eliminate potential nonconformity | Trend analysis, cause determination, action evaluation, implementation, effectiveness review |
### FDA 21 CFR 820.100
Required CAPA elements:
- Procedures for implementing corrective and preventive action
- Analyzing quality data sources (complaints, NCs, audits, service records)
- Investigating cause of nonconformities
- Identifying actions needed to correct and prevent recurrence
- Verifying actions are effective and do not adversely affect device
- Submitting relevant information for management review
### Common FDA 483 Observations
| Observation | Root Cause Pattern |
|-------------|-------------------|
| CAPA not initiated for recurring issue | Trend analysis not performed |
| Root cause analysis superficial | Inadequate investigation training |
| Effectiveness not verified | No verification procedure |
| Actions do not address root cause | Symptom treatment vs. cause elimination |
FILE:references/effectiveness-verification-guide.md
# Effectiveness Verification Guide
CAPA effectiveness assessment procedures, verification methods, and closure criteria.
---
## Table of Contents
- [Verification Planning](#verification-planning)
- [Verification Methods](#verification-methods)
- [Effectiveness Criteria](#effectiveness-criteria)
- [Closure Requirements](#closure-requirements)
- [Ineffective CAPA Process](#ineffective-capa-process)
- [Documentation Templates](#documentation-templates)
---
## Verification Planning
### When to Plan Verification
Verification planning must occur BEFORE corrective action implementation:
| Stage | Planning Activity | Owner |
|-------|-------------------|-------|
| CAPA Initiation | Define preliminary verification approach | CAPA Owner |
| Root Cause Analysis | Refine criteria based on root cause | Investigation Team |
| Action Planning | Finalize verification method and timeline | CAPA Owner |
| Implementation | Schedule verification activities | Quality Assurance |
### Verification Timeline Guidelines
| CAPA Severity | Minimum Wait Period | Verification Window |
|---------------|---------------------|---------------------|
| Critical (Safety) | 30 days | 30-90 days post-implementation |
| Major | 60 days | 60-180 days post-implementation |
| Minor | 90 days | 90-365 days post-implementation |
**Rationale**: Waiting period ensures sufficient data collection and accounts for process variation.
### Verification Plan Components
```
VERIFICATION PLAN TEMPLATE
CAPA Number: [CAPA-XXXX]
Problem Statement: [Original issue]
Root Cause: [Identified root cause]
Corrective Action: [Implemented action]
VERIFICATION METHOD:
[ ] Data Trend Analysis
[ ] Process Audit
[ ] Record Review
[ ] Testing/Inspection
[ ] Interview/Observation
[ ] Multiple Methods (specify)
EFFECTIVENESS CRITERIA:
1. [Measurable criterion 1]
2. [Measurable criterion 2]
3. [Measurable criterion 3]
SUCCESS THRESHOLD:
- [Quantitative threshold, e.g., "Zero recurrence for 90 days"]
- [Qualitative threshold, e.g., "Procedure followed correctly 100%"]
DATA COLLECTION:
- Source: [Where data will come from]
- Sample Size: [Number of records/instances to review]
- Time Period: [Start and end dates]
- Responsible: [Who collects data]
VERIFICATION SCHEDULE:
- Implementation Complete: [Date]
- Waiting Period Ends: [Date]
- Verification Start: [Date]
- Verification Complete: [Date]
- Report Due: [Date]
APPROVAL:
CAPA Owner: _____________ Date: _______
Quality Assurance: _____________ Date: _______
```
---
## Verification Methods
### 1. Data Trend Analysis
**Best for:** Quantifiable issues with measurable outcomes (defect rates, cycle times, complaint trends)
**Procedure:**
1. Collect post-implementation data for defined period
2. Compare to pre-implementation baseline
3. Apply statistical analysis if sample size permits
4. Document trend direction and magnitude
**Example Criteria:**
- Defect rate reduced by ≥50% from baseline
- Zero recurrence of specific failure mode
- Process capability (Cpk) improved to ≥1.33
**Evidence Required:**
- Pre-implementation baseline data
- Post-implementation trend data
- Statistical analysis (if applicable)
- Trend charts with annotation
### 2. Process Audit
**Best for:** Procedure compliance issues, process control failures, systemic problems
**Procedure:**
1. Develop audit checklist based on corrective action
2. Conduct unannounced process audit
3. Interview operators and supervisors
4. Review records generated since implementation
5. Document compliance percentage
**Example Criteria:**
- 100% compliance with revised procedure
- All operators demonstrate competency
- No deviations observed during audit
**Evidence Required:**
- Audit checklist completed
- Interview notes
- Record samples reviewed
- Photos/observations (if applicable)
### 3. Record Review
**Best for:** Documentation issues, completeness problems, traceability failures
**Procedure:**
1. Define sample size based on volume (minimum 10 or 10%, whichever greater)
2. Review records generated post-implementation
3. Evaluate against specified requirements
4. Calculate compliance rate
**Example Criteria:**
- 100% of records meet completeness requirements
- All required signatures present
- Traceability maintained throughout
**Evidence Required:**
- List of records reviewed
- Compliance checklist results
- Non-compliance summary (if any)
### 4. Testing/Inspection
**Best for:** Product quality issues, equipment failures, specification non-conformances
**Procedure:**
1. Define test protocol based on corrective action
2. Conduct testing on post-implementation units
3. Compare results to acceptance criteria
4. Document pass/fail rates
**Example Criteria:**
- 100% of units pass revised inspection criteria
- All test results within specification
- Zero failures of targeted parameter
**Evidence Required:**
- Test protocol/method
- Test results data
- Pass/fail summary
- Comparison to pre-implementation results
### 5. Interview/Observation
**Best for:** Training issues, communication problems, human factors causes
**Procedure:**
1. Develop structured interview questions
2. Interview representative sample of affected personnel
3. Observe process execution in real-time
4. Document responses and observations
**Example Criteria:**
- All interviewed personnel demonstrate knowledge
- Observed practices match documented procedure
- No unsafe acts or workarounds observed
**Evidence Required:**
- Interview questions and responses
- Observation notes
- Training records (supporting)
---
## Effectiveness Criteria
### Defining Good Criteria
Criteria must be **SMART**:
| Element | Requirement | Example |
|---------|-------------|---------|
| **S**pecific | Clearly defined what to measure | "Calibration overdue rate" not "equipment issues" |
| **M**easurable | Quantifiable or objectively verifiable | "<2% overdue rate" not "improved timeliness" |
| **A**chievable | Realistic given the corrective action | Within capability of implemented solution |
| **R**elevant | Directly related to root cause | Addresses the actual problem |
| **T**ime-bound | Specified evaluation period | "For 90 consecutive days" |
### Criteria by Issue Type
| Issue Type | Typical Criteria | Threshold |
|------------|------------------|-----------|
| Nonconformance | Recurrence rate | Zero recurrence |
| Process deviation | Compliance rate | ≥95% compliance |
| Complaint | Complaint trend | ≥50% reduction |
| Calibration | Overdue rate | <2% overdue |
| Training | Competency pass rate | 100% pass |
| Documentation | Completeness rate | 100% complete |
| Supplier | Incoming reject rate | ≤1% reject rate |
### Sample Size Guidelines
| Population Size | Minimum Sample |
|-----------------|----------------|
| <10 | All (100%) |
| 10-50 | 10 |
| 51-100 | 15 |
| 101-500 | 20 |
| >500 | 25 or 10%, whichever less |
---
## Closure Requirements
### Closure Checklist
**CAPA Closure Prerequisites:**
- [ ] All corrective actions implemented
- [ ] Implementation evidence documented
- [ ] Verification waiting period complete
- [ ] Verification activities performed
- [ ] All effectiveness criteria met
- [ ] Verification evidence documented
- [ ] No recurrence during verification period
- [ ] CAPA owner review complete
- [ ] Quality Assurance review complete
- [ ] Documentation complete and filed
### Effectiveness Status Determination
```
EFFECTIVENESS DECISION TREE:
Did recurrence occur during verification period?
├── Yes → CAPA INEFFECTIVE (escalate per ineffective process)
└── No → Were all effectiveness criteria met?
├── Yes → Were any related issues identified?
│ ├── Yes → Open new CAPA if needed, close original
│ └── No → CAPA EFFECTIVE - proceed to closure
└── No → How many criteria missed?
├── Minor gap (1 criterion, marginal miss) →
│ Extend verification period OR accept with justification
└── Significant gap → CAPA INEFFECTIVE
EFFECTIVENESS DETERMINATION:
[ ] EFFECTIVE - All criteria met, no recurrence
[ ] EFFECTIVE WITH CONDITIONS - Minor gap, justified acceptance
[ ] INEFFECTIVE - Significant gaps or recurrence
```
### Closure Documentation
```
EFFECTIVENESS VERIFICATION REPORT
CAPA Number: [CAPA-XXXX]
Verification Complete Date: [Date]
Verified By: [Name, Title]
VERIFICATION SUMMARY:
| Criterion | Target | Actual | Status |
|-----------|--------|--------|--------|
| [Criterion 1] | [Target] | [Result] | ☑ Met / ☐ Not Met |
| [Criterion 2] | [Target] | [Result] | ☑ Met / ☐ Not Met |
| [Criterion 3] | [Target] | [Result] | ☑ Met / ☐ Not Met |
RECURRENCE CHECK:
- Recurrence during verification period: [ ] Yes [ ] No
- Related issues identified: [ ] Yes [ ] No
- If yes, describe: [Description]
EVIDENCE SUMMARY:
[List of evidence documents, record numbers, data sources]
EFFECTIVENESS DETERMINATION:
[ ] EFFECTIVE
[ ] EFFECTIVE WITH CONDITIONS: [Justification]
[ ] INEFFECTIVE: [Reason]
RECOMMENDED ACTION:
[ ] Close CAPA
[ ] Extend verification period to [Date]
[ ] Open new CAPA [CAPA-XXXX] for [Issue]
[ ] Re-investigate (return to root cause analysis)
APPROVALS:
CAPA Owner: _____________ Date: _______
Quality Assurance: _____________ Date: _______
Management (if Major/Critical): _____________ Date: _______
```
---
## Ineffective CAPA Process
### Definition of Ineffective
CAPA is ineffective when:
1. Original problem recurs during or after verification period
2. Effectiveness criteria not met
3. Root cause still present
4. Corrective action created new problems
### Ineffective CAPA Workflow
```
INEFFECTIVE CAPA DETECTED
│
├── 1. Immediate Actions
│ ├── Reopen CAPA (do not close as effective)
│ ├── Implement containment for recurrence
│ └── Notify CAPA owner and management
│
├── 2. Root Cause Re-evaluation
│ ├── Was original root cause correct?
│ │ ├── No → Conduct new root cause analysis
│ │ └── Yes → Was corrective action appropriate?
│ │ ├── No → Develop new corrective action
│ │ └── Yes → Was implementation adequate?
│ │ ├── No → Re-implement with improvements
│ │ └── Yes → Escalate (systemic issue)
│
├── 3. Escalation Criteria
│ ├── Second ineffective attempt → Management review required
│ ├── Safety-related recurrence → Immediate escalation
│ └── Pattern across multiple CAPAs → Systemic CAPA
│
└── 4. Documentation
├── Document ineffective status with evidence
├── Record re-investigation results
├── Update CAPA metrics/trending
└── Include in management review
```
### Preventing Ineffective CAPAs
| Common Cause | Prevention |
|--------------|------------|
| Superficial root cause | Validate root cause before action |
| Action addresses symptom not cause | Ensure action targets root cause |
| Implementation incomplete | Verify implementation before verification |
| Insufficient verification period | Allow adequate time for data collection |
| Wrong verification method | Match method to issue type |
| Unclear success criteria | Define SMART criteria upfront |
---
## Documentation Templates
### Verification Evidence Log
```
VERIFICATION EVIDENCE LOG
CAPA Number: [CAPA-XXXX]
| Doc/Record # | Description | Date | Reviewed By | Finding |
|--------------|-------------|------|-------------|---------|
| [Number] | [Description] | [Date] | [Reviewer] | [Compliant/Finding] |
| [Number] | [Description] | [Date] | [Reviewer] | [Compliant/Finding] |
SUMMARY:
- Total records reviewed: [Number]
- Compliant: [Number] ([Percentage]%)
- Non-compliant: [Number] ([Percentage]%)
CONCLUSION:
[Statement on whether evidence supports effectiveness]
```
### Trend Analysis Summary
```
TREND ANALYSIS FOR CAPA VERIFICATION
CAPA Number: [CAPA-XXXX]
Metric: [What is being measured]
BASELINE (Pre-Implementation):
- Period: [Start] to [End]
- Value: [Baseline value]
- Data points: [Number]
POST-IMPLEMENTATION:
- Period: [Start] to [End]
- Value: [Current value]
- Data points: [Number]
CHANGE:
- Absolute change: [Value]
- Percentage change: [Percentage]%
- Target: [Target value/change]
- Status: [ ] Met [ ] Not Met
TREND CHART:
[Include or reference trend chart showing before/after comparison]
STATISTICAL SIGNIFICANCE (if applicable):
- Method: [t-test, chi-square, etc.]
- p-value: [Value]
- Conclusion: [Statistically significant / Not significant]
```
### Interview Summary Template
```
VERIFICATION INTERVIEW SUMMARY
CAPA Number: [CAPA-XXXX]
Interviewer: [Name]
Date: [Date]
INTERVIEWEE:
- Name: [Name]
- Role: [Job title]
- Department: [Department]
- Experience: [Years in role]
QUESTIONS AND RESPONSES:
Q1: [Question about awareness of change]
A1: [Response summary]
Knowledge demonstrated: [ ] Yes [ ] Partial [ ] No
Q2: [Question about implementation of change]
A2: [Response summary]
Compliance demonstrated: [ ] Yes [ ] Partial [ ] No
Q3: [Question about understanding rationale]
A3: [Response summary]
Understanding demonstrated: [ ] Yes [ ] Partial [ ] No
OBSERVATION NOTES:
[Any relevant observations during interview]
CONCLUSION:
[ ] Interviewee demonstrates full knowledge and compliance
[ ] Interviewee demonstrates partial knowledge (specify gaps)
[ ] Interviewee does not demonstrate required knowledge
```
FILE:references/rca-methodologies.md
# Root Cause Analysis Methodologies
Decision criteria, templates, and implementation guidance for RCA techniques.
---
## Table of Contents
- [Method Selection Matrix](#method-selection-matrix)
- [5 Why Analysis](#5-why-analysis)
- [Fishbone Diagram](#fishbone-diagram)
- [Fault Tree Analysis](#fault-tree-analysis)
- [Human Factors Analysis](#human-factors-analysis)
- [Failure Mode and Effects Analysis](#failure-mode-and-effects-analysis)
- [Selecting the Right Method](#selecting-the-right-method)
---
## Method Selection Matrix
### When to Use Each Method
| Method | Use When | Problem Type | Team Size | Time Required |
|--------|----------|--------------|-----------|---------------|
| 5 Why | Single-cause issues, process deviations | Linear causation | 1-3 people | 30-60 min |
| Fishbone | Multi-factor problems, 3-6 contributing factors | Complex, systemic | 3-8 people | 2-4 hours |
| Fault Tree | Safety-critical failures, reliability issues | System failures | 2-5 people | 4-8 hours |
| Human Factors | Procedure/training-related issues | Human error | 3-6 people | 2-4 hours |
| FMEA | Systematic risk assessment, design review | Potential failures | 4-10 people | 8-16 hours |
### Quick Selection Decision Tree
```
Is the issue safety-critical or involves system reliability?
├── Yes → Use FAULT TREE ANALYSIS
└── No → Is human error the suspected primary cause?
├── Yes → Use HUMAN FACTORS ANALYSIS
└── No → How many potential contributing factors?
├── 1-2 factors → Use 5 WHY ANALYSIS
├── 3-6 factors → Use FISHBONE DIAGRAM
└── Unknown/Many → Use FMEA (proactive) or Fishbone (reactive)
```
---
## 5 Why Analysis
### Overview
Simple, iterative technique asking "why" repeatedly (typically 5 times) to drill from symptoms to root cause.
### When to Use
- Single-cause issues with linear causation
- Process deviations with clear failure point
- Quick investigations requiring rapid resolution
- Problems where symptoms clearly link to cause
### When NOT to Use
- Complex multi-factor problems
- Safety-critical incidents requiring comprehensive analysis
- Issues with multiple interacting causes
- When systemic factors are suspected
### 5 Why Template
```
PROBLEM STATEMENT:
[Clear, specific description of what happened, when, where, and impact]
WHY 1: Why did [problem] occur?
BECAUSE: [First-level cause]
EVIDENCE: [Data/observation supporting this cause]
WHY 2: Why did [first-level cause] occur?
BECAUSE: [Second-level cause]
EVIDENCE: [Data/observation supporting this cause]
WHY 3: Why did [second-level cause] occur?
BECAUSE: [Third-level cause]
EVIDENCE: [Data/observation supporting this cause]
WHY 4: Why did [third-level cause] occur?
BECAUSE: [Fourth-level cause]
EVIDENCE: [Data/observation supporting this cause]
WHY 5: Why did [fourth-level cause] occur?
BECAUSE: [Root cause - typically systemic or management system failure]
EVIDENCE: [Data/observation supporting this cause]
ROOT CAUSE VALIDATION:
- [ ] Can the root cause be verified with evidence?
- [ ] If root cause is eliminated, would problem recur?
- [ ] Is the root cause within organizational control?
- [ ] Does the root cause explain all symptoms?
```
### Example: Calibration Overdue
```
PROBLEM: pH meter (EQ-042) found 2 months overdue for calibration
WHY 1: Why was calibration overdue?
BECAUSE: The equipment was not on the calibration schedule
EVIDENCE: Calibration schedule reviewed, EQ-042 not listed
WHY 2: Why was it not on the calibration schedule?
BECAUSE: The schedule was not updated when equipment was purchased
EVIDENCE: Purchase date 2023-06-15, schedule dated 2023-01-01
WHY 3: Why was the schedule not updated?
BECAUSE: No process requires schedule update at equipment purchase
EVIDENCE: Equipment procedure SOP-EQ-001 reviewed, no such requirement
WHY 4: Why is there no requirement to update the schedule?
BECAUSE: The procedure was written before equipment tracking was centralized
EVIDENCE: SOP-EQ-001 last revised 2019, equipment system implemented 2021
WHY 5: Why has the procedure not been updated?
BECAUSE: Periodic procedure review did not assess compatibility with new systems
EVIDENCE: No documented review of SOP-EQ-001 against new equipment system
ROOT CAUSE: Procedure review process does not assess compatibility
with organizational systems implemented after original procedure creation
```
---
## Fishbone Diagram
### Overview
Also called Ishikawa or cause-and-effect diagram. Organizes potential causes into categories branching from the problem statement.
### Standard Categories (6M)
| Category | Focus Areas | Typical Causes |
|----------|-------------|----------------|
| **Man** (People) | Training, competency, workload | Skill gaps, fatigue, communication |
| **Machine** (Equipment) | Calibration, maintenance, age | Wear, malfunction, inadequate capacity |
| **Method** (Process) | Procedures, work instructions | Unclear steps, missing controls |
| **Material** | Specifications, suppliers, storage | Out-of-spec, degradation, contamination |
| **Measurement** | Calibration, methods, interpretation | Instrument error, wrong method |
| **Mother Nature** (Environment) | Temperature, humidity, cleanliness | Environmental excursions |
### Fishbone Template
```
PROBLEM STATEMENT: [Effect being investigated]
┌── Man ────────────────┐
│ ├─ [Cause 1] │
│ ├─ [Cause 2] │
│ └─ [Cause 3] │
│ │
┌── Machine ────────┤ ├── Method ──────────┐
│ ├─ [Cause 1] │ │ ├─ [Cause 1] │
│ ├─ [Cause 2] │ PROBLEM │ ├─ [Cause 2] │
│ └─ [Cause 3] ├───────────────────────┤ └─ [Cause 3] │
│ │ │ │
├── Material ───────┤ ├── Measurement ─────┤
│ ├─ [Cause 1] │ │ ├─ [Cause 1] │
│ ├─ [Cause 2] │ │ ├─ [Cause 2] │
│ └─ [Cause 3] │ │ └─ [Cause 3] │
│ │
└── Environment ────────┘
├─ [Cause 1]
├─ [Cause 2]
└─ [Cause 3]
CAUSE PRIORITIZATION:
| Cause | Category | Likelihood | Evidence | Priority |
|-------|----------|------------|----------|----------|
| [Cause A] | Method | High | [Evidence] | 1 |
| [Cause B] | Man | Medium | [Evidence] | 2 |
ROOT CAUSES IDENTIFIED:
1. [Primary root cause with supporting evidence]
2. [Contributing cause with supporting evidence]
```
### Facilitation Guidelines
1. Assemble cross-functional team (3-8 people)
2. Define problem statement clearly before starting
3. Brainstorm causes without judgment first
4. Organize into categories after brainstorming
5. Drill down on each major cause (sub-causes)
6. Prioritize based on evidence and likelihood
7. Validate top causes with data
---
## Fault Tree Analysis
### Overview
Top-down, deductive analysis starting with undesired event and systematically identifying all potential causes using Boolean logic (AND/OR gates).
### When to Use
- Safety-critical system failures
- Complex system reliability analysis
- Events with multiple failure pathways
- Regulatory-required investigations (FDA, MDR)
### FTA Symbols
| Symbol | Name | Meaning |
|--------|------|---------|
| Rectangle | Top Event / Intermediate Event | Undesired event or intermediate fault |
| Circle | Basic Event | Primary fault requiring no further analysis |
| Diamond | Undeveloped Event | Event not fully analyzed (data limitation) |
| AND Gate | Requires all inputs | All child events must occur for parent |
| OR Gate | Requires any input | Any child event causes parent |
### FTA Template
```
TOP EVENT: [Undesired event under investigation]
LEVEL 1 (Immediate Causes):
[Top Event]
│
└── OR GATE ──┬── [Cause 1.1]
├── [Cause 1.2]
└── [Cause 1.3]
LEVEL 2 (Contributing Causes):
[Cause 1.1]
│
└── AND GATE ──┬── [Cause 2.1]
└── [Cause 2.2]
MINIMAL CUT SETS:
(Combinations of basic events that cause top event)
1. {Basic Event A, Basic Event B} ← Both required (AND)
2. {Basic Event C} ← Single point failure (OR)
3. {Basic Event D, Basic Event E} ← Both required (AND)
CRITICAL PATH ANALYSIS:
Most likely failure pathway: [Description]
Single points of failure: [List]
RECOMMENDATIONS:
- Address single points of failure first
- Add redundancy where AND gates show vulnerability
- Prioritize controls on highest probability paths
```
### Cut Set Analysis
Minimal cut sets identify the smallest combination of basic events causing the top event:
- **Single-element cut sets**: Single points of failure (highest priority)
- **Two-element cut sets**: Dual failure scenarios
- **Probability calculation**: P(Top Event) = Union of P(Cut Sets)
---
## Human Factors Analysis
### Overview
Systematic analysis of human error focusing on cognitive, physical, and organizational factors contributing to performance failures.
### HFACS Categories
Human Factors Analysis and Classification System:
| Level | Category | Examples |
|-------|----------|----------|
| **Unsafe Acts** | Errors, violations | Skill-based, decision, perceptual errors |
| **Preconditions** | Conditions for unsafe acts | Fatigue, mental state, CRM, physical environment |
| **Unsafe Supervision** | Supervisory failures | Inadequate supervision, planned inappropriate ops |
| **Organizational Influences** | Organizational failures | Resource management, organizational climate |
### Human Error Types
| Type | Description | Example | Mitigation |
|------|-------------|---------|------------|
| Slip | Execution error in routine task | Wrong button pressed | Error-proofing, forcing functions |
| Lapse | Memory failure | Forgot step in procedure | Checklists, reminders |
| Mistake | Planning/decision error | Wrong procedure selected | Training, decision aids |
| Violation | Intentional deviation | Skipped step to save time | Culture change, supervision |
### Human Factors Investigation Template
```
INCIDENT DESCRIPTION:
[What happened, who was involved, when, where]
UNSAFE ACTS ANALYSIS:
Type of Error: [ ] Slip [ ] Lapse [ ] Mistake [ ] Violation
Description: [Specific action or inaction]
Task Being Performed: [Activity at time of error]
Experience Level: [Novice/Intermediate/Expert]
PRECONDITIONS FOR UNSAFE ACTS:
Cognitive Factors:
- [ ] Task complexity exceeded capability
- [ ] Time pressure
- [ ] Distraction/interruption
- [ ] Mental fatigue
Physical Factors:
- [ ] Physical fatigue
- [ ] Inadequate lighting
- [ ] Noise interference
- [ ] Workspace ergonomics
Team Factors:
- [ ] Communication breakdown
- [ ] Coordination failure
- [ ] Inadequate leadership
SUPERVISORY FACTORS:
- [ ] Inadequate supervision
- [ ] Failed to correct known problem
- [ ] Inappropriate staffing
- [ ] Authorized unnecessary risk
ORGANIZATIONAL FACTORS:
- [ ] Resource management deficiency
- [ ] Organizational process issue
- [ ] Organizational culture/climate
ROOT CAUSE(S):
[Human factors root causes identified]
CORRECTIVE ACTIONS:
| Action | Target Factor | Priority |
|--------|---------------|----------|
| [Action 1] | [Factor addressed] | High |
| [Action 2] | [Factor addressed] | Medium |
```
---
## Failure Mode and Effects Analysis
### Overview
Proactive, systematic technique identifying potential failure modes, their causes, and effects before failures occur.
### FMEA Types
| Type | Application | Scope |
|------|-------------|-------|
| Design FMEA (DFMEA) | Product design | Component and system design failures |
| Process FMEA (PFMEA) | Manufacturing process | Process step failures |
| System FMEA | System-level analysis | System interaction failures |
### Risk Priority Number (RPN)
RPN = Severity (S) × Occurrence (O) × Detection (D)
**Severity Scale (1-10):**
| Rating | Effect | Criteria |
|--------|--------|----------|
| 10 | Hazardous | Failure affects safe operation, no warning |
| 8-9 | Very High | Primary function lost, high impact |
| 6-7 | High | Performance degraded, customer dissatisfied |
| 4-5 | Moderate | Some performance loss, moderate impact |
| 2-3 | Low | Minor effect, slight inconvenience |
| 1 | None | No discernible effect |
**Occurrence Scale (1-10):**
| Rating | Likelihood | Failure Rate |
|--------|------------|--------------|
| 10 | Very High | >1 in 10 |
| 7-9 | High | 1 in 20 - 1 in 100 |
| 4-6 | Moderate | 1 in 400 - 1 in 2,000 |
| 2-3 | Low | 1 in 15,000 - 1 in 150,000 |
| 1 | Remote | <1 in 1,500,000 |
**Detection Scale (1-10):**
| Rating | Detection | Criteria |
|--------|-----------|----------|
| 10 | Absolute Uncertainty | No inspection/control, defect will reach customer |
| 7-9 | Very Remote to Remote | Controls unlikely to detect |
| 4-6 | Moderate | Controls may detect |
| 2-3 | High | Controls likely to detect |
| 1 | Almost Certain | Controls will almost certainly detect |
### FMEA Template
```
PROCESS/PRODUCT: [Name]
FMEA TEAM: [Members]
DATE: [Date]
| Item/Step | Failure Mode | Effect | S | Cause | O | Controls | D | RPN | Action |
|-----------|--------------|--------|---|-------|---|----------|---|-----|--------|
| [Item 1] | [How it fails] | [Impact] | 8 | [Why] | 4 | [Current] | 6 | 192 | [Action] |
| [Item 2] | [How it fails] | [Impact] | 6 | [Why] | 3 | [Current] | 4 | 72 | [Action] |
RPN THRESHOLD: Actions required for RPN > [threshold]
HIGH SEVERITY RULE: Actions required for S >= 9 regardless of RPN
ACTION PRIORITIZATION:
1. Address all items with S >= 9 first
2. Address items with highest RPN
3. Focus on reducing Occurrence (prevention)
4. Then improve Detection (inspection)
```
---
## Selecting the Right Method
### Decision Flowchart
```
START: Investigation Required
│
├── Is this a proactive assessment (no failure yet)?
│ └── Yes → Use FMEA
│
├── Is the issue safety-critical?
│ └── Yes → Use FAULT TREE ANALYSIS
│
├── Is human error the primary concern?
│ └── Yes → Use HUMAN FACTORS ANALYSIS
│
├── Are there multiple contributing factors (3+)?
│ ├── Yes → Use FISHBONE DIAGRAM
│ └── No → Use 5 WHY ANALYSIS
│
└── Uncertain? → Start with 5 WHY, escalate to FISHBONE if needed
```
### Hybrid Approach
For complex investigations, combine methods:
1. **Initial screening**: 5 Why for quick cause identification
2. **Detailed analysis**: Fishbone to explore all categories
3. **Validation**: Fault Tree for critical failure paths
4. **Systemic factors**: Human Factors for people-related causes
5. **Prevention**: FMEA for future risk mitigation
### Documentation Requirements
| Method | Required Outputs | Retention |
|--------|------------------|-----------|
| 5 Why | Completed template with evidence | CAPA record |
| Fishbone | Diagram + prioritized causes | CAPA record |
| Fault Tree | FTA diagram + cut set analysis | DHF/CAPA record |
| Human Factors | HFACS analysis + actions | CAPA record |
| FMEA | FMEA worksheet + action tracking | Design file |
FILE:scripts/capa_tracker.py
#!/usr/bin/env python3
"""
CAPA Tracker - Corrective and Preventive Action Management Tool
Tracks CAPA status, calculates metrics, identifies overdue items,
and generates reports for management review.
Usage:
python capa_tracker.py --capas capas.json
python capa_tracker.py --interactive
python capa_tracker.py --capas capas.json --output json
"""
import argparse
import json
import sys
from dataclasses import dataclass, field, asdict
from datetime import datetime, timedelta
from typing import List, Dict, Optional
from enum import Enum
class CAPAStatus(Enum):
OPEN = "Open"
INVESTIGATION = "Investigation"
ACTION_PLANNING = "Action Planning"
IMPLEMENTATION = "Implementation"
VERIFICATION = "Verification"
CLOSED_EFFECTIVE = "Closed - Effective"
CLOSED_INEFFECTIVE = "Closed - Ineffective"
class CAPASeverity(Enum):
CRITICAL = "Critical"
MAJOR = "Major"
MINOR = "Minor"
class CAPASource(Enum):
COMPLAINT = "Customer Complaint"
AUDIT = "Internal Audit"
EXTERNAL_AUDIT = "External Audit"
NONCONFORMANCE = "Nonconformance"
MANAGEMENT_REVIEW = "Management Review"
TREND_ANALYSIS = "Trend Analysis"
REGULATORY = "Regulatory Feedback"
OTHER = "Other"
@dataclass
class CAPA:
capa_number: str
title: str
description: str
source: CAPASource
severity: CAPASeverity
status: CAPAStatus
open_date: str
target_date: str
owner: str
root_cause: str = ""
corrective_action: str = ""
verification_date: Optional[str] = None
close_date: Optional[str] = None
days_open: int = 0
is_overdue: bool = False
@dataclass
class CAPAMetrics:
total_capas: int
open_capas: int
closed_capas: int
overdue_capas: int
avg_cycle_time: float
effectiveness_rate: float
by_status: Dict[str, int]
by_severity: Dict[str, int]
by_source: Dict[str, int]
overdue_list: List[Dict]
recommendations: List[str]
class CAPATracker:
"""CAPA tracking and metrics calculator."""
# Target cycle times by severity (days)
TARGET_CYCLE_TIMES = {
CAPASeverity.CRITICAL: 30,
CAPASeverity.MAJOR: 60,
CAPASeverity.MINOR: 90,
}
def __init__(self, capas: List[CAPA]):
self.capas = capas
self.today = datetime.now()
self._calculate_derived_fields()
def _calculate_derived_fields(self):
"""Calculate days open and overdue status."""
for capa in self.capas:
open_date = datetime.strptime(capa.open_date, "%Y-%m-%d")
if capa.close_date:
close_date = datetime.strptime(capa.close_date, "%Y-%m-%d")
capa.days_open = (close_date - open_date).days
else:
capa.days_open = (self.today - open_date).days
target_date = datetime.strptime(capa.target_date, "%Y-%m-%d")
if not capa.close_date and self.today > target_date:
capa.is_overdue = True
def calculate_metrics(self) -> CAPAMetrics:
"""Calculate comprehensive CAPA metrics."""
total = len(self.capas)
# Status counts
closed_statuses = [CAPAStatus.CLOSED_EFFECTIVE, CAPAStatus.CLOSED_INEFFECTIVE]
open_capas = [c for c in self.capas if c.status not in closed_statuses]
closed_capas = [c for c in self.capas if c.status in closed_statuses]
overdue_capas = [c for c in self.capas if c.is_overdue]
# Average cycle time (closed CAPAs only)
if closed_capas:
avg_cycle = sum(c.days_open for c in closed_capas) / len(closed_capas)
else:
avg_cycle = 0.0
# Effectiveness rate
effective = [c for c in self.capas if c.status == CAPAStatus.CLOSED_EFFECTIVE]
ineffective = [c for c in self.capas if c.status == CAPAStatus.CLOSED_INEFFECTIVE]
if effective or ineffective:
effectiveness = len(effective) / (len(effective) + len(ineffective)) * 100
else:
effectiveness = 0.0
# Counts by category
by_status = {}
for status in CAPAStatus:
count = len([c for c in self.capas if c.status == status])
if count > 0:
by_status[status.value] = count
by_severity = {}
for severity in CAPASeverity:
count = len([c for c in self.capas if c.severity == severity])
if count > 0:
by_severity[severity.value] = count
by_source = {}
for source in CAPASource:
count = len([c for c in self.capas if c.source == source])
if count > 0:
by_source[source.value] = count
# Overdue list
overdue_list = []
for capa in sorted(overdue_capas, key=lambda c: c.days_open, reverse=True):
target = datetime.strptime(capa.target_date, "%Y-%m-%d")
days_overdue = (self.today - target).days
overdue_list.append({
"capa_number": capa.capa_number,
"title": capa.title,
"severity": capa.severity.value,
"status": capa.status.value,
"days_overdue": days_overdue,
"owner": capa.owner
})
# Generate recommendations
recommendations = self._generate_recommendations(
open_capas, overdue_capas, effectiveness, avg_cycle
)
return CAPAMetrics(
total_capas=total,
open_capas=len(open_capas),
closed_capas=len(closed_capas),
overdue_capas=len(overdue_capas),
avg_cycle_time=round(avg_cycle, 1),
effectiveness_rate=round(effectiveness, 1),
by_status=by_status,
by_severity=by_severity,
by_source=by_source,
overdue_list=overdue_list,
recommendations=recommendations
)
def _generate_recommendations(
self,
open_capas: List[CAPA],
overdue_capas: List[CAPA],
effectiveness: float,
avg_cycle: float
) -> List[str]:
"""Generate actionable recommendations."""
recommendations = []
# Overdue CAPAs
if overdue_capas:
critical_overdue = [c for c in overdue_capas if c.severity == CAPASeverity.CRITICAL]
if critical_overdue:
recommendations.append(
f"URGENT: {len(critical_overdue)} critical CAPA(s) overdue. "
"Escalate to management immediately."
)
else:
recommendations.append(
f"ACTION: {len(overdue_capas)} CAPA(s) overdue. "
"Review and update target dates or expedite closure."
)
# Effectiveness rate
if effectiveness < 80 and effectiveness > 0:
recommendations.append(
f"CONCERN: Effectiveness rate at {effectiveness:.0f}%. "
"Review root cause analysis quality and corrective action adequacy."
)
# Cycle time
if avg_cycle > 60:
recommendations.append(
f"IMPROVEMENT: Average cycle time is {avg_cycle:.0f} days. "
"Target is 60 days. Review investigation and approval bottlenecks."
)
# Investigation backlog
in_investigation = [c for c in open_capas if c.status == CAPAStatus.INVESTIGATION]
if len(in_investigation) > 5:
recommendations.append(
f"WORKLOAD: {len(in_investigation)} CAPAs in investigation phase. "
"Consider additional resources or prioritization."
)
# Stuck in verification
in_verification = [c for c in open_capas if c.status == CAPAStatus.VERIFICATION]
old_verification = [c for c in in_verification if c.days_open > 120]
if old_verification:
recommendations.append(
f"STALLED: {len(old_verification)} CAPA(s) in verification >120 days. "
"Complete effectiveness checks or extend with justification."
)
# Source patterns
complaint_capas = [c for c in self.capas if c.source == CAPASource.COMPLAINT]
if len(complaint_capas) > len(self.capas) * 0.4:
recommendations.append(
"TREND: >40% of CAPAs from customer complaints. "
"Review preventive action effectiveness and quality controls."
)
if not recommendations:
recommendations.append(
"CAPA program operating within targets. "
"Continue monitoring key metrics."
)
return recommendations
def get_aging_report(self) -> Dict:
"""Generate aging analysis of open CAPAs."""
open_statuses = [
CAPAStatus.OPEN, CAPAStatus.INVESTIGATION,
CAPAStatus.ACTION_PLANNING, CAPAStatus.IMPLEMENTATION,
CAPAStatus.VERIFICATION
]
open_capas = [c for c in self.capas if c.status in open_statuses]
aging_buckets = {
"0-30 days": [],
"31-60 days": [],
"61-90 days": [],
"91-120 days": [],
">120 days": []
}
for capa in open_capas:
days = capa.days_open
if days <= 30:
bucket = "0-30 days"
elif days <= 60:
bucket = "31-60 days"
elif days <= 90:
bucket = "61-90 days"
elif days <= 120:
bucket = "91-120 days"
else:
bucket = ">120 days"
aging_buckets[bucket].append({
"capa_number": capa.capa_number,
"title": capa.title,
"days_open": days,
"status": capa.status.value,
"severity": capa.severity.value
})
return aging_buckets
def format_text_output(metrics: CAPAMetrics, aging: Dict) -> str:
"""Format metrics as text report."""
lines = [
"=" * 70,
"CAPA STATUS REPORT",
"=" * 70,
f"Generated: {datetime.now().strftime('%Y-%m-%d %H:%M')}",
"",
"SUMMARY METRICS",
"-" * 40,
f"Total CAPAs: {metrics.total_capas}",
f"Open CAPAs: {metrics.open_capas}",
f"Closed CAPAs: {metrics.closed_capas}",
f"Overdue CAPAs: {metrics.overdue_capas}",
f"Avg Cycle Time: {metrics.avg_cycle_time} days",
f"Effectiveness Rate: {metrics.effectiveness_rate}%",
"",
"STATUS DISTRIBUTION",
"-" * 40,
]
for status, count in metrics.by_status.items():
bar = "█" * min(count, 20)
lines.append(f" {status:<25} {bar} {count}")
lines.extend([
"",
"SEVERITY DISTRIBUTION",
"-" * 40,
])
for severity, count in metrics.by_severity.items():
bar = "█" * min(count, 20)
lines.append(f" {severity:<25} {bar} {count}")
lines.extend([
"",
"SOURCE DISTRIBUTION",
"-" * 40,
])
for source, count in metrics.by_source.items():
bar = "█" * min(count, 20)
lines.append(f" {source:<25} {bar} {count}")
lines.extend([
"",
"AGING ANALYSIS",
"-" * 40,
])
for bucket, capas in aging.items():
lines.append(f" {bucket}: {len(capas)} CAPA(s)")
if metrics.overdue_list:
lines.extend([
"",
"OVERDUE CAPAs",
"-" * 40,
f"{'CAPA #':<12} {'Title':<25} {'Days':<6} {'Owner':<15}",
"-" * 60,
])
for item in metrics.overdue_list[:10]:
title = item["title"][:24] if len(item["title"]) > 24 else item["title"]
lines.append(
f"{item['capa_number']:<12} {title:<25} "
f"{item['days_overdue']:<6} {item['owner']:<15}"
)
if len(metrics.overdue_list) > 10:
lines.append(f"... and {len(metrics.overdue_list) - 10} more")
lines.extend([
"",
"RECOMMENDATIONS",
"-" * 40,
])
for i, rec in enumerate(metrics.recommendations, 1):
lines.append(f"{i}. {rec}")
lines.append("=" * 70)
return "\n".join(lines)
def interactive_mode():
"""Run interactive CAPA entry mode."""
print("=" * 60)
print("CAPA Tracker - Interactive Mode")
print("=" * 60)
capas = []
print("\nEnter CAPAs (blank CAPA number to finish):\n")
while True:
capa_num = input("CAPA Number (e.g., CAPA-2024-001): ").strip()
if not capa_num:
break
title = input("Title: ").strip()
description = input("Description: ").strip()
print("Source options: C=Complaint, A=Audit, N=Nonconformance, M=Management Review, T=Trend, O=Other")
source_input = input("Source [C/A/N/M/T/O]: ").strip().upper()
source_map = {
"C": CAPASource.COMPLAINT,
"A": CAPASource.AUDIT,
"N": CAPASource.NONCONFORMANCE,
"M": CAPASource.MANAGEMENT_REVIEW,
"T": CAPASource.TREND_ANALYSIS,
"O": CAPASource.OTHER
}
source = source_map.get(source_input, CAPASource.OTHER)
print("Severity: C=Critical, M=Major, I=Minor")
severity_input = input("Severity [C/M/I]: ").strip().upper()
severity_map = {
"C": CAPASeverity.CRITICAL,
"M": CAPASeverity.MAJOR,
"I": CAPASeverity.MINOR
}
severity = severity_map.get(severity_input, CAPASeverity.MINOR)
print("Status: O=Open, I=Investigation, P=Action Planning, M=Implementation, V=Verification, E=Closed Effective, N=Closed Ineffective")
status_input = input("Status [O/I/P/M/V/E/N]: ").strip().upper()
status_map = {
"O": CAPAStatus.OPEN,
"I": CAPAStatus.INVESTIGATION,
"P": CAPAStatus.ACTION_PLANNING,
"M": CAPAStatus.IMPLEMENTATION,
"V": CAPAStatus.VERIFICATION,
"E": CAPAStatus.CLOSED_EFFECTIVE,
"N": CAPAStatus.CLOSED_INEFFECTIVE
}
status = status_map.get(status_input, CAPAStatus.OPEN)
open_date = input("Open Date (YYYY-MM-DD): ").strip()
target_date = input("Target Date (YYYY-MM-DD): ").strip()
owner = input("Owner: ").strip()
close_date = None
if status in [CAPAStatus.CLOSED_EFFECTIVE, CAPAStatus.CLOSED_INEFFECTIVE]:
close_date = input("Close Date (YYYY-MM-DD): ").strip()
capas.append(CAPA(
capa_number=capa_num,
title=title,
description=description,
source=source,
severity=severity,
status=status,
open_date=open_date,
target_date=target_date,
owner=owner,
close_date=close_date if close_date else None
))
print(f"\nAdded: {capa_num}\n")
if not capas:
print("No CAPAs entered. Exiting.")
return
tracker = CAPATracker(capas)
metrics = tracker.calculate_metrics()
aging = tracker.get_aging_report()
print("\n" + format_text_output(metrics, aging))
def main():
parser = argparse.ArgumentParser(
description="CAPA Tracking and Metrics Tool"
)
parser.add_argument(
"--capas",
type=str,
help="JSON file with CAPA data"
)
parser.add_argument(
"--output",
choices=["text", "json"],
default="text",
help="Output format"
)
parser.add_argument(
"--interactive",
action="store_true",
help="Run in interactive mode"
)
parser.add_argument(
"--sample",
action="store_true",
help="Generate sample CAPA data file"
)
args = parser.parse_args()
if args.interactive:
interactive_mode()
return
if args.sample:
sample_data = {
"capas": [
{
"capa_number": "CAPA-2024-001",
"title": "Calibration overdue for pH meter",
"description": "pH meter EQ-042 found 2 months overdue",
"source": "AUDIT",
"severity": "MAJOR",
"status": "VERIFICATION",
"open_date": "2024-06-15",
"target_date": "2024-08-15",
"owner": "J. Smith",
"root_cause": "No trigger for schedule update at equipment purchase",
"corrective_action": "Updated SOP-EQ-001 to require schedule update"
},
{
"capa_number": "CAPA-2024-002",
"title": "Customer complaint - labeling error",
"description": "Wrong lot number on product label",
"source": "COMPLAINT",
"severity": "CRITICAL",
"status": "INVESTIGATION",
"open_date": "2024-09-01",
"target_date": "2024-10-01",
"owner": "M. Jones"
},
{
"capa_number": "CAPA-2024-003",
"title": "Training records incomplete",
"description": "Missing effectiveness verification for 3 operators",
"source": "AUDIT",
"severity": "MINOR",
"status": "CLOSED_EFFECTIVE",
"open_date": "2024-03-10",
"target_date": "2024-06-10",
"owner": "A. Brown",
"close_date": "2024-05-20"
}
]
}
print(json.dumps(sample_data, indent=2))
return
if args.capas:
with open(args.capas, "r") as f:
data = json.load(f)
capas = []
for c in data.get("capas", []):
try:
source = CAPASource[c.get("source", "OTHER").upper()]
except KeyError:
source = CAPASource.OTHER
try:
severity = CAPASeverity[c.get("severity", "MINOR").upper()]
except KeyError:
severity = CAPASeverity.MINOR
try:
status = CAPAStatus[c.get("status", "OPEN").upper()]
except KeyError:
status = CAPAStatus.OPEN
capas.append(CAPA(
capa_number=c["capa_number"],
title=c.get("title", ""),
description=c.get("description", ""),
source=source,
severity=severity,
status=status,
open_date=c["open_date"],
target_date=c["target_date"],
owner=c.get("owner", ""),
root_cause=c.get("root_cause", ""),
corrective_action=c.get("corrective_action", ""),
verification_date=c.get("verification_date"),
close_date=c.get("close_date")
))
else:
# Demo data if no file provided
capas = [
CAPA(
capa_number="CAPA-2024-001",
title="Calibration overdue",
description="pH meter overdue",
source=CAPASource.AUDIT,
severity=CAPASeverity.MAJOR,
status=CAPAStatus.VERIFICATION,
open_date="2024-06-15",
target_date="2024-08-15",
owner="J. Smith"
),
CAPA(
capa_number="CAPA-2024-002",
title="Labeling error complaint",
description="Wrong lot number",
source=CAPASource.COMPLAINT,
severity=CAPASeverity.CRITICAL,
status=CAPAStatus.INVESTIGATION,
open_date="2024-09-01",
target_date="2024-10-01",
owner="M. Jones"
),
CAPA(
capa_number="CAPA-2024-003",
title="Training records incomplete",
description="Missing effectiveness verification",
source=CAPASource.AUDIT,
severity=CAPASeverity.MINOR,
status=CAPAStatus.CLOSED_EFFECTIVE,
open_date="2024-03-10",
target_date="2024-06-10",
owner="A. Brown",
close_date="2024-05-20"
)
]
tracker = CAPATracker(capas)
metrics = tracker.calculate_metrics()
aging = tracker.get_aging_report()
if args.output == "json":
output = {
"metrics": asdict(metrics),
"aging": aging
}
print(json.dumps(output, indent=2))
else:
print(format_text_output(metrics, aging))
if __name__ == "__main__":
main()
FILE:scripts/root_cause_analyzer.py
#!/usr/bin/env python3
"""
Root Cause Analyzer - Structured root cause analysis for CAPA investigations.
Supports multiple analysis methodologies:
- 5-Why Analysis
- Fishbone (Ishikawa) Diagram
- Fault Tree Analysis
- Kepner-Tregoe Problem Analysis
Generates structured root cause reports and CAPA recommendations.
Usage:
python root_cause_analyzer.py --method 5why --problem "High defect rate in assembly line"
python root_cause_analyzer.py --interactive
python root_cause_analyzer.py --data investigation.json --output json
"""
import argparse
import json
import sys
from dataclasses import dataclass, field, asdict
from typing import List, Dict, Optional
from enum import Enum
from datetime import datetime
class AnalysisMethod(Enum):
FIVE_WHY = "5-Why"
FISHBONE = "Fishbone"
FAULT_TREE = "Fault Tree"
KEPNER_TREGOE = "Kepner-Tregoe"
class RootCauseCategory(Enum):
MAN = "Man (People)"
MACHINE = "Machine (Equipment)"
MATERIAL = "Material"
METHOD = "Method (Process)"
MEASUREMENT = "Measurement"
ENVIRONMENT = "Environment"
MANAGEMENT = "Management (Policy)"
SOFTWARE = "Software/Data"
class SeverityLevel(Enum):
LOW = "Low"
MEDIUM = "Medium"
HIGH = "High"
CRITICAL = "Critical"
@dataclass
class WhyStep:
"""A single step in 5-Why analysis."""
level: int
question: str
answer: str
evidence: str = ""
verified: bool = False
@dataclass
class FishboneCause:
"""A cause in fishbone analysis."""
category: str
cause: str
sub_causes: List[str] = field(default_factory=list)
is_root: bool = False
evidence: str = ""
@dataclass
class FaultEvent:
"""An event in fault tree analysis."""
event_id: str
description: str
is_basic: bool = True # Basic events have no children
gate_type: str = "OR" # OR, AND
children: List[str] = field(default_factory=list)
probability: Optional[float] = None
@dataclass
class RootCauseFinding:
"""Identified root cause with evidence."""
cause_id: str
description: str
category: str
evidence: List[str] = field(default_factory=list)
contributing_factors: List[str] = field(default_factory=list)
systemic: bool = False # Whether it's a systemic vs. local issue
@dataclass
class CAPARecommendation:
"""Corrective or preventive action recommendation."""
action_id: str
action_type: str # "Corrective" or "Preventive"
description: str
addresses_cause: str # cause_id
priority: str
estimated_effort: str
responsible_role: str
effectiveness_criteria: List[str] = field(default_factory=list)
@dataclass
class RootCauseAnalysis:
"""Complete root cause analysis result."""
investigation_id: str
problem_statement: str
analysis_method: str
root_causes: List[RootCauseFinding]
recommendations: List[CAPARecommendation]
analysis_details: Dict
confidence_level: float
investigator_notes: List[str] = field(default_factory=list)
class RootCauseAnalyzer:
"""Performs structured root cause analysis."""
def __init__(self):
self.analysis_steps = []
self.findings = []
def analyze_5why(self, problem: str, whys: List[Dict] = None) -> Dict:
"""Perform 5-Why analysis."""
steps = []
if whys:
for i, w in enumerate(whys, 1):
steps.append(WhyStep(
level=i,
question=w.get("question", f"Why did this occur? (Level {i})"),
answer=w.get("answer", ""),
evidence=w.get("evidence", ""),
verified=w.get("verified", False)
))
# Analyze depth and quality
depth = len(steps)
has_root = any(
s.answer and ("system" in s.answer.lower() or "policy" in s.answer.lower() or "process" in s.answer.lower())
for s in steps
)
return {
"method": "5-Why Analysis",
"steps": [asdict(s) for s in steps],
"depth": depth,
"reached_systemic_cause": has_root,
"quality_score": min(100, depth * 20 + (20 if has_root else 0))
}
def analyze_fishbone(self, problem: str, causes: List[Dict] = None) -> Dict:
"""Perform fishbone (Ishikawa) analysis."""
categories = {}
fishbone_causes = []
if causes:
for c in causes:
cat = c.get("category", "Method")
cause = c.get("cause", "")
sub = c.get("sub_causes", [])
if cat not in categories:
categories[cat] = []
categories[cat].append({
"cause": cause,
"sub_causes": sub,
"is_root": c.get("is_root", False),
"evidence": c.get("evidence", "")
})
fishbone_causes.append(FishboneCause(
category=cat,
cause=cause,
sub_causes=sub,
is_root=c.get("is_root", False),
evidence=c.get("evidence", "")
))
root_causes = [fc for fc in fishbone_causes if fc.is_root]
return {
"method": "Fishbone (Ishikawa) Analysis",
"problem": problem,
"categories": categories,
"total_causes": len(fishbone_causes),
"root_causes_identified": len(root_causes),
"categories_covered": list(categories.keys()),
"recommended_categories": [c.value for c in RootCauseCategory],
"missing_categories": [c.value for c in RootCauseCategory if c.value.split(" (")[0] not in categories]
}
def analyze_fault_tree(self, top_event: str, events: List[Dict] = None) -> Dict:
"""Perform fault tree analysis."""
fault_events = {}
if events:
for e in events:
fault_events[e["event_id"]] = FaultEvent(
event_id=e["event_id"],
description=e.get("description", ""),
is_basic=e.get("is_basic", True),
gate_type=e.get("gate_type", "OR"),
children=e.get("children", []),
probability=e.get("probability")
)
# Find basic events (root causes)
basic_events = {eid: ev for eid, ev in fault_events.items() if ev.is_basic}
intermediate_events = {eid: ev for eid, ev in fault_events.items() if not ev.is_basic}
return {
"method": "Fault Tree Analysis",
"top_event": top_event,
"total_events": len(fault_events),
"basic_events": len(basic_events),
"intermediate_events": len(intermediate_events),
"basic_event_details": [asdict(e) for e in basic_events.values()],
"cut_sets": self._find_cut_sets(fault_events)
}
def _find_cut_sets(self, events: Dict[str, FaultEvent]) -> List[List[str]]:
"""Find minimal cut sets (combinations of basic events that cause top event)."""
# Simplified cut set analysis
cut_sets = []
for eid, event in events.items():
if not event.is_basic and event.gate_type == "AND":
cut_sets.append(event.children)
return cut_sets[:5] # Return top 5
def generate_recommendations(
self,
root_causes: List[RootCauseFinding],
problem: str
) -> List[CAPARecommendation]:
"""Generate CAPA recommendations based on root causes."""
recommendations = []
for i, cause in enumerate(root_causes, 1):
# Corrective action (fix the immediate cause)
recommendations.append(CAPARecommendation(
action_id=f"CA-{i:03d}",
action_type="Corrective",
description=f"Address immediate cause: {cause.description}",
addresses_cause=cause.cause_id,
priority=self._assess_priority(cause),
estimated_effort=self._estimate_effort(cause),
responsible_role=self._suggest_responsible(cause),
effectiveness_criteria=[
f"Elimination of {cause.description} confirmed by audit",
"No recurrence within 90 days",
"Metrics return to acceptable range"
]
))
# Preventive action (prevent recurrence in other areas)
if cause.systemic:
recommendations.append(CAPARecommendation(
action_id=f"PA-{i:03d}",
action_type="Preventive",
description=f"Systemic prevention: Update process/procedure to prevent similar issues",
addresses_cause=cause.cause_id,
priority="Medium",
estimated_effort="2-4 weeks",
responsible_role="Quality Manager",
effectiveness_criteria=[
"Updated procedure approved and implemented",
"Training completed for affected personnel",
"No similar issues in related processes within 6 months"
]
))
return recommendations
def _assess_priority(self, cause: RootCauseFinding) -> str:
if cause.systemic or "safety" in cause.description.lower():
return "High"
elif "quality" in cause.description.lower():
return "Medium"
return "Low"
def _estimate_effort(self, cause: RootCauseFinding) -> str:
if cause.systemic:
return "4-8 weeks"
elif len(cause.contributing_factors) > 3:
return "2-4 weeks"
return "1-2 weeks"
def _suggest_responsible(self, cause: RootCauseFinding) -> str:
category_roles = {
"Man": "Training Manager",
"Machine": "Engineering Manager",
"Material": "Supply Chain Manager",
"Method": "Process Owner",
"Measurement": "Quality Engineer",
"Environment": "Facilities Manager",
"Management": "Department Head",
"Software": "IT/Software Manager"
}
cat_key = cause.category.split(" (")[0] if "(" in cause.category else cause.category
return category_roles.get(cat_key, "Quality Manager")
def full_analysis(
self,
problem: str,
method: str = "5-Why",
analysis_data: Dict = None
) -> RootCauseAnalysis:
"""Perform complete root cause analysis."""
investigation_id = f"RCA-{datetime.now().strftime('%Y%m%d-%H%M')}"
analysis_details = {}
root_causes = []
if method == "5-Why" and analysis_data:
analysis_details = self.analyze_5why(problem, analysis_data.get("whys", []))
# Extract root cause from deepest why
steps = analysis_details.get("steps", [])
if steps:
last_step = steps[-1]
root_causes.append(RootCauseFinding(
cause_id="RC-001",
description=last_step.get("answer", "Unknown"),
category="Systemic",
evidence=[s.get("evidence", "") for s in steps if s.get("evidence")],
systemic=analysis_details.get("reached_systemic_cause", False)
))
elif method == "Fishbone" and analysis_data:
analysis_details = self.analyze_fishbone(problem, analysis_data.get("causes", []))
for i, cat in enumerate(analysis_data.get("causes", [])):
if cat.get("is_root"):
root_causes.append(RootCauseFinding(
cause_id=f"RC-{i+1:03d}",
description=cat.get("cause", ""),
category=cat.get("category", ""),
evidence=[cat.get("evidence", "")] if cat.get("evidence") else [],
sub_causes=cat.get("sub_causes", []),
systemic=True
))
recommendations = self.generate_recommendations(root_causes, problem)
# Confidence based on evidence and method
confidence = 0.7
if root_causes and any(rc.evidence for rc in root_causes):
confidence = 0.85
if len(root_causes) > 1:
confidence = min(0.95, confidence + 0.05)
return RootCauseAnalysis(
investigation_id=investigation_id,
problem_statement=problem,
analysis_method=method,
root_causes=root_causes,
recommendations=recommendations,
analysis_details=analysis_details,
confidence_level=confidence
)
def format_rca_text(rca: RootCauseAnalysis) -> str:
"""Format RCA report as text."""
lines = [
"=" * 70,
"ROOT CAUSE ANALYSIS REPORT",
"=" * 70,
f"Investigation ID: {rca.investigation_id}",
f"Analysis Method: {rca.analysis_method}",
f"Confidence Level: {rca.confidence_level:.0%}",
"",
"PROBLEM STATEMENT",
"-" * 40,
f" {rca.problem_statement}",
"",
"ROOT CAUSES IDENTIFIED",
"-" * 40,
]
for rc in rca.root_causes:
lines.extend([
f"",
f" [{rc.cause_id}] {rc.description}",
f" Category: {rc.category}",
f" Systemic: {'Yes' if rc.systemic else 'No'}",
])
if rc.evidence:
lines.append(f" Evidence:")
for ev in rc.evidence:
if ev:
lines.append(f" • {ev}")
if rc.contributing_factors:
lines.append(f" Contributing Factors:")
for cf in rc.contributing_factors:
lines.append(f" - {cf}")
lines.extend([
"",
"RECOMMENDED ACTIONS",
"-" * 40,
])
for rec in rca.recommendations:
lines.extend([
f"",
f" [{rec.action_id}] {rec.action_type}: {rec.description}",
f" Priority: {rec.priority} | Effort: {rec.estimated_effort}",
f" Responsible: {rec.responsible_role}",
f" Effectiveness Criteria:",
])
for ec in rec.effectiveness_criteria:
lines.append(f" ✓ {ec}")
if "steps" in rca.analysis_details:
lines.extend([
"",
"5-WHY CHAIN",
"-" * 40,
])
for step in rca.analysis_details["steps"]:
lines.extend([
f"",
f" Why {step['level']}: {step['question']}",
f" → {step['answer']}",
])
if step.get("evidence"):
lines.append(f" Evidence: {step['evidence']}")
lines.append("=" * 70)
return "\n".join(lines)
def main():
parser = argparse.ArgumentParser(description="Root Cause Analyzer for CAPA Investigations")
parser.add_argument("--problem", type=str, help="Problem statement")
parser.add_argument("--method", choices=["5why", "fishbone", "fault-tree", "kt"],
default="5why", help="Analysis method")
parser.add_argument("--data", type=str, help="JSON file with analysis data")
parser.add_argument("--output", choices=["text", "json"], default="text", help="Output format")
parser.add_argument("--interactive", action="store_true", help="Interactive mode")
args = parser.parse_args()
analyzer = RootCauseAnalyzer()
if args.data:
with open(args.data) as f:
data = json.load(f)
problem = data.get("problem", "Unknown problem")
method = data.get("method", "5-Why")
rca = analyzer.full_analysis(problem, method, data)
elif args.problem:
method_map = {"5why": "5-Why", "fishbone": "Fishbone", "fault-tree": "Fault Tree", "kt": "Kepner-Tregoe"}
rca = analyzer.full_analysis(args.problem, method_map.get(args.method, "5-Why"))
else:
# Demo
demo_data = {
"method": "5-Why",
"whys": [
{"question": "Why did the product fail inspection?", "answer": "Surface defect detected on 15% of units", "evidence": "QC inspection records"},
{"question": "Why did surface defects occur?", "answer": "Injection molding temperature was outside spec", "evidence": "Process monitoring data"},
{"question": "Why was temperature outside spec?", "answer": "Temperature controller calibration drift", "evidence": "Calibration log"},
{"question": "Why did calibration drift go undetected?", "answer": "No automated alert for drift, manual checks missed it", "evidence": "SOP review"},
{"question": "Why was there no automated alert?", "answer": "Process monitoring system lacks drift detection capability - systemic gap", "evidence": "System requirements review"}
]
}
rca = analyzer.full_analysis("High defect rate in injection molding process", "5-Why", demo_data)
if args.output == "json":
result = {
"investigation_id": rca.investigation_id,
"problem": rca.problem_statement,
"method": rca.analysis_method,
"root_causes": [asdict(rc) for rc in rca.root_causes],
"recommendations": [asdict(rec) for rec in rca.recommendations],
"analysis_details": rca.analysis_details,
"confidence": rca.confidence_level
}
print(json.dumps(result, indent=2, default=str))
else:
print(format_rca_text(rca))
if __name__ == "__main__":
main()
Giả định kế hoạch thất bại sau 12 tháng rồi lần ngược để tìm điểm yếu, giả định và rủi ro thực thi.
--- name: "challenge" description: "Pre-mortem plan analysis. Imagine the plan failed 12 months from now and work backwards to find the weaknesses. Surfaces assumptions, dependencies, and execution risks before committing resources. Use when before significant resource commitment, before presenting to a board or investors, when feedback has been one-sidedly positive, or when there is pressure to move fast and figure it out later." --- # /em:challenge — Pre-Mortem Plan Analysis **Command:** `/em:challenge <plan>` Systematically finds weaknesses in any plan before reality does. Not to kill the plan — to make it survive contact with reality. --- ## The Core Idea Most plans fail for predictable reasons. Not bad luck — bad assumptions. Overestimated demand. Underestimated complexity. Dependencies nobody questioned. Timing that made sense in a spreadsheet but not in the real world. The pre-mortem technique: **imagine it's 12 months from now and this plan failed spectacularly. Now work backwards. Why?** That's not pessimism. It's how you build something that doesn't collapse. --- ## When to Run a Challenge - Before committing significant resources to a plan - Before presenting to the board or investors - When you notice you're only hearing positive feedback about the plan - When the plan requires multiple external dependencies to align - When there's pressure to move fast and "figure it out later" - When you feel excited about the plan (excitement is a signal to scrutinize harder) --- ## The Challenge Framework ### Step 1: Extract Core Assumptions Before you can test a plan, you need to surface everything it assumes to be true. For each section of the plan, ask: - What has to be true for this to work? - What are we assuming about customer behavior? - What are we assuming about competitor response? - What are we assuming about our own execution capability? - What external factors does this depend on? **Common assumption categories:** - **Market assumptions** — size, growth rate, customer willingness to pay, buying cycle - **Execution assumptions** — team capacity, velocity, no major hires needed - **Customer assumptions** — they have the problem, they know they have it, they'll pay to solve it - **Competitive assumptions** — incumbents won't respond, no new entrant, moat holds - **Financial assumptions** — burn rate, revenue timing, CAC, LTV ratios - **Dependency assumptions** — partner will deliver, API won't change, regulations won't shift ### Step 2: Rate Each Assumption For every assumption extracted, rate it on two dimensions: **Confidence level (how sure are you this is true):** - **High** — verified with data, customer conversations, market research - **Medium** — directionally right but not validated - **Low** — plausible but untested - **Unknown** — we simply don't know **Impact if wrong (what happens if this assumption fails):** - **Critical** — plan fails entirely - **High** — major delay or cost overrun - **Medium** — significant rework required - **Low** — manageable adjustment ### Step 3: Map Vulnerabilities The matrix of Low/Unknown confidence × Critical/High impact = your highest-risk assumptions. **Vulnerability = Low confidence + High impact** These are not problems to ignore. They're the bets you're making. The question is: are you making them consciously? ### Step 4: Find the Dependency Chain Many plans fail not because any single assumption is wrong, but because multiple assumptions have to be right simultaneously. Map the chain: - Does assumption B depend on assumption A being true first? - If the first thing goes wrong, how many downstream things break? - What's the critical path? What has zero slack? ### Step 5: Test the Reversibility For each critical vulnerability: if this assumption turns out to be wrong at month 3, what do you do? - Can you pivot? - Can you cut scope? - Is money already spent? - Are commitments already made? The less reversible, the more rigorously you need to validate before committing. --- ## Output Format **Challenge Report: [Plan Name]** ``` CORE ASSUMPTIONS (extracted) 1. [Assumption] — Confidence: [H/M/L/?] — Impact if wrong: [Critical/High/Medium/Low] 2. ... VULNERABILITY MAP Critical risks (act before proceeding): • [#N] [Assumption] — WHY it might be wrong — WHAT breaks if it is High risks (validate before scaling): • ... DEPENDENCY CHAIN [Assumption A] → depends on → [Assumption B] → which enables → [Assumption C] Weakest link: [X] — if this breaks, [Y] and [Z] also fail REVERSIBILITY ASSESSMENT • Reversible bets: [list] • Irreversible commitments: [list — treat with extreme care] KILL SWITCHES What would have to be true at [30/60/90 days] to continue vs. kill/pivot? • Continue if: ... • Kill/pivot if: ... HARDENING ACTIONS 1. [Specific validation to do before proceeding] 2. [Alternative approach to consider] 3. [Contingency to build into the plan] ``` --- ## Challenge Patterns by Plan Type ### Product Roadmap - Are we building what customers will pay for, or what they said they wanted? - Does the velocity estimate account for real team capacity (not theoretical)? - What happens if the anchor feature takes 3× longer than estimated? - Who owns decisions when requirements conflict? ### Go-to-Market Plan - What's the actual ICP conversion rate, not the hoped-for one? - How many touches to close, and do you have the sales capacity for that? - What happens if the first 10 deals take 3 months instead of 1? - Is "land and expand" a real motion or a hope? ### Hiring Plan - What happens if the key hire takes 4 months to find, not 6 weeks? - Is the plan dependent on retaining specific people who might leave? - Does the plan account for ramp time (usually 3–6 months before full productivity)? - What's the burn impact if headcount leads revenue by 6 months? ### Fundraising Plan - What's your fallback if the lead investor passes? - Have you modeled the timeline if it takes 6 months, not 3? - What's your runway at current burn if the round closes at the low end? - What assumptions break if you raise 50% of the target amount? --- ## The Hardest Questions These are the ones people skip: - "What's the bear case, not the base case?" - "If this exact plan was run by a team we don't trust, would it work?" - "What are we not saying out loud because it's uncomfortable?" - "Who has incentives to make this plan sound better than it is?" - "What would an enemy of this plan attack first?" --- ## Deliverable The output of `/em:challenge` is not permission to stop. It's a vulnerability map. Now you can make conscious decisions: validate the risky assumptions, hedge the critical ones, or accept the bets you're making knowingly. Unknown risks are dangerous. Known risks are manageable.
Tự động đánh giá mã nhiều ngôn ngữ: phân tích PR, độ phức tạp, vi phạm SOLID, mã có mùi và tạo báo cáo.
---
name: "code-reviewer"
description: Code review automation for TypeScript, JavaScript, Python, Go, Swift, Kotlin, C#, .NET, Java, C, C++, Rust, Ruby, PHP, and Dart/Flutter. Analyzes PRs for complexity and risk, checks code quality for SOLID violations and code smells, generates review reports. Use when reviewing pull requests, analyzing code quality, identifying issues, generating review checklists.
---
# Code Reviewer
Automated code review tools for analyzing pull requests, detecting code quality issues, and generating review reports.
---
## How This Skill Is Organized
```
code-reviewer/
SKILL.md ← you are here (tools + dispatch table)
rules/
universal.md ← security, async, resources, exceptions, performance — all languages
languages/
python.md ← Python-specific rules + idioms
typescript.md ← TypeScript / JavaScript-specific rules + idioms
go.md ← Go-specific rules + idioms
swift.md ← Swift-specific rules + idioms
kotlin.md ← Kotlin-specific rules + idioms
csharp.md ← C# / .NET-specific rules + idioms
java.md ← Java-specific rules + idioms
c.md ← C -specific rules + idioms
cpp.md ← C++ -specific rules + idioms
rust.md ← Rust -specific rules + idioms
ruby.md ← Ruby -specific rules + idioms
php.md ← PHP-specific rules + idioms
dart.md ← Dart / Flutter-specific rules + idioms
```
### Loading order for every review
1. This file (`SKILL.md`) — tools and thresholds
2. `rules/universal.md` — always, for every language
3. The matching `languages/*.md` — one file based on the extension table below
That is always exactly **2 additional files**, regardless of scope.
| Extension(s) | Load |
|---|---|
| `.py` | `languages/python.md` |
| `.ts`, `.tsx`, `.js`, `.jsx`, `.mjs` | `languages/typescript.md` |
| `.go` | `languages/go.md` |
| `.swift` | `languages/swift.md` |
| `.kt`, `.kts` | `languages/kotlin.md` |
| `.cs`, `.csx`, `.razor`, `.cshtml` | `languages/csharp.md` |
| `.java` | `languages/java.md` |
| `.c`, `.h` | `languages/c.md` |
| `.cpp`, `.cc`, `.cxx`, `.hpp`, `.hh`, `.hxx` | `languages/cpp.md` |
| `.rs` | `languages/rust.md` |
| `.rb`, `.rake`, `.gemspec`, `.ru` | `languages/ruby.md` |
| `.php`, `.phtml` | `languages/php.md` |
| `.dart` | `languages/dart.md` |
---
## Tools
### PR Analyzer
Analyzes git diff between branches to assess review complexity and identify risks.
```bash
# Analyze current branch against main
python scripts/pr_analyzer.py /path/to/repo
# Compare specific branches
python scripts/pr_analyzer.py . --base main --head feature-branch
# JSON output for integration
python scripts/pr_analyzer.py /path/to/repo --json
```
**What it detects (universal — see also language file for language-specific signals):**
- Hardcoded secrets (passwords, API keys, tokens, connection strings)
- SQL / query injection patterns
- Debug statements left in production code
- Lint / analyzer suppression annotations
- TODO/FIXME comments
**Language-specific detections** are defined in each `languages/*.md` file.
**Output includes:**
- Complexity score (1-10)
- Risk categorization (critical, high, medium, low)
- File prioritization for review order
- Commit message validation
---
### Code Quality Checker
Analyzes source code for structural issues, code smells, and SOLID violations.
```bash
# Analyze a directory
python scripts/code_quality_checker.py /path/to/code
# Analyze specific language
# Valid values: python, typescript, javascript, go, swift, kotlin, csharp, java, c, cpp, rust, ruby, php, dart
python scripts/code_quality_checker.py . --language java
# JSON output
python scripts/code_quality_checker.py /path/to/code --json
```
**Universal thresholds:**
| Issue | Threshold |
|-------|-----------|
| Long function | >50 lines |
| Large file | >500 lines |
| God class | >20 methods |
| Too many params | >5 |
| Deep nesting | >4 levels |
| High complexity | >10 branches |
Language-specific checks are defined in each `languages/*.md` file.
---
### Review Report Generator
Combines PR analysis and code quality findings into structured review reports.
```bash
# Generate report for current repo
python scripts/review_report_generator.py /path/to/repo
# Markdown output
python scripts/review_report_generator.py . --format markdown --output review.md
# Use pre-computed analyses
python scripts/review_report_generator.py . \
--pr-analysis pr_results.json \
--quality-analysis quality_results.json
```
**Verdicts:**
| Score | Verdict |
|-------|---------|
| 90+ with no high issues | Approve |
| 75+ with ≤2 high issues | Approve with suggestions |
| 50-74 | Request changes |
| <50 or critical issues | Block |
---
## Adding a New Language
**Reviewer guidance (required):**
1. Create `languages/<name>.md` using any existing language file as a template — it must have sections: PR Analyzer Signals, Code Quality Checks, Security, Async, Resource Management, Exception Handling, Performance, Idioms.
2. Add the extension row to the dispatch table above.
That is all the agent-driven review needs.
**Deterministic analyzer support (optional, recommended):** the bundled scripts
only flag a language they explicitly know. To make `code_quality_checker.py`
score the new language:
3. Add the extensions to `LANGUAGE_EXTENSIONS` in `scripts/code_quality_checker.py` (this also adds the `--language` choice).
4. Add `function` / `class` / `method` regex entries for the language in the same file; otherwise it falls back to the Python patterns.
5. Optionally add a `check_<name>_specific_smells(...)` detector (see the C#, Java, and C ones) and call it from `analyze_file`.
6. Add `assets/sample_<name>_smells.<ext>` + `_clean` fixtures and commit the expected `--json` output under `expected_outputs/` as a regression guard.
---
## Regression Fixtures
Labelled fixtures live in `assets/` with their committed `--json` output in
`expected_outputs/` (C#, Java, and C). Drift from the committed JSON signals a
behaviour change in the analyzer:
```bash
python scripts/code_quality_checker.py assets/sample_java_smells.java --json \
| diff - expected_outputs/sample_java_smells_quality.json
```
FILE:assets/sample_csharp_clean.cs
// Sample C# file showing the fixed version of sample_csharp_smells.cs.
// Same shape, but every smell has been resolved per the patterns documented
// in rules/universal.md and languages/csharp.md.
//
// Run:
// python scripts/code_quality_checker.py assets/sample_csharp_clean.cs
//
// Expected: no HIGH C#-specific smells flagged.
using System;
using System.Net.Http;
using System.Threading.Tasks;
using System.Data.SqlClient;
using Microsoft.Extensions.Logging;
using Microsoft.Extensions.Options;
namespace Sample
{
public class DbOptions
{
// FIX: connection string from configuration, never inlined.
public string ConnectionString { get; init; } = "";
}
public class UserService
{
private readonly string _connectionString;
private readonly HttpClient _httpClient;
private readonly ILogger<UserService> _logger;
// FIX: IHttpClientFactory + IOptions, no hardcoded secrets, no `new HttpClient()`.
public UserService(
IHttpClientFactory httpClientFactory,
IOptions<DbOptions> dbOptions,
ILogger<UserService> logger)
{
_httpClient = httpClientFactory.CreateClient("api");
_connectionString = dbOptions.Value.ConnectionString;
_logger = logger;
}
// FIX: async Task (not async void) so callers can await and observe exceptions.
public async Task HandleClickAsync()
{
// FIX: await the Task instead of blocking on it.
var data = await FetchAsync().ConfigureAwait(false);
_logger.LogInformation("Fetched {Length} bytes", data.Length);
}
public async Task<string> FetchAsync()
{
try
{
// FIX: await the async call — Task is no longer discarded.
await FireAndForgetAsync().ConfigureAwait(false);
// FIX: real null check, no `!`.
var user = await GetCurrentUserAsync().ConfigureAwait(false);
if (user is null)
{
throw new InvalidOperationException("No current user");
}
_ = user.Name;
return await _httpClient
.GetStringAsync("https://api.example/data")
.ConfigureAwait(false);
}
catch (HttpRequestException ex)
{
// FIX: catch specific exception, log with context, rethrow.
_logger.LogError(ex, "Upstream fetch failed");
throw;
}
}
// FIX: real type, not `dynamic`.
public User? CurrentUser { get; private set; }
// FIX: `unsafe` removed — none of the logic actually needed pointers.
public int FirstValue(int[] values) => values.Length > 0 ? values[0] : 0;
// FIX: no #pragma / [SuppressMessage] — root cause fixed instead.
public string GetName(int id)
{
// FIX: `using var` disposes connection + command deterministically.
using var conn = new SqlConnection(_connectionString);
// FIX: parameterized query, no string concatenation.
using var cmd = new SqlCommand("SELECT name FROM users WHERE id = @id", conn);
cmd.Parameters.AddWithValue("@id", id);
conn.Open();
return (string)cmd.ExecuteScalar();
}
private Task<User?> GetCurrentUserAsync() => Task.FromResult<User?>(null);
private Task FireAndForgetAsync() => Task.CompletedTask;
}
public record User(int Id, string Name);
}
FILE:assets/sample_csharp_smells.cs
// Sample C# file demonstrating every C#-specific pattern the code-reviewer
// skill detects. Each smell is labelled inline. This file is NOT meant to
// compile cleanly — it is a fixture for code_quality_checker.py and
// pr_analyzer.py.
//
// Run:
// python scripts/code_quality_checker.py assets/sample_csharp_smells.cs
//
// Expected output: see expected_outputs/sample_csharp_smells_quality.json
using System;
using System.Net.Http;
using System.Threading.Tasks;
using System.Data.SqlClient;
using System.Diagnostics.CodeAnalysis;
namespace Sample
{
public class UserService
{
// [hardcoded_secrets] hardcoded connection string with password
public string ConnectionString = "Server=prod;Database=app;Password=hunter2;";
// [csharp_async_void] async void on a non-event-handler signature
public async void HandleClick(object sender, EventArgs e)
{
// [csharp_blocking_async] .Result blocks on Task in a sync context
var data = FetchAsync().Result;
// [console_log] Debug.WriteLine output statement
Debug.WriteLine(data);
}
public async Task<string> FetchAsync()
{
// [csharp_new_httpclient] new HttpClient() in method body
// [csharp_undisposed_idisposable] HttpClient not in `using`
var client = new HttpClient();
try
{
// [csharp_missing_await] FireAndForgetAsync() returns Task, never awaited
FireAndForgetAsync();
// [csharp_null_forgiving] `user!.Name` forces null-forgiving
var name = user!.Name;
return await client.GetStringAsync("https://api.example/data");
}
catch (Exception)
{
// [csharp_swallowed_exception] empty catch (Exception)
}
return null!;
}
// [loose_type] C# `dynamic` overuse
public dynamic Untyped = null;
// [csharp_unsafe_block] `unsafe` modifier on a method
public unsafe void Pointers()
{
int x = 0;
int* p = &x;
}
// [analyzer_disable] #pragma warning disable
#pragma warning disable CS0168
// [analyzer_disable] [SuppressMessage] attribute
[SuppressMessage("Style", "IDE0060")]
public string GetName(SqlConnection conn, int id)
{
// [csharp_undisposed_idisposable] SqlCommand without `using`
// [sql_concatenation] string concatenation builds SQL with user input
var cmd = new SqlCommand("SELECT name FROM users WHERE id = " + id, conn);
return cmd.ExecuteScalar().ToString();
}
}
}
FILE:assets/sample_c_clean.c
/*
* sample_c_clean.c — sample_c_smells.c refactored per
* rules/universal.md + languages/c.md. Same surface area, zero
* detector hits.
*/
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
void safe_input(void) {
char buf[64];
/* fgets is bounds-aware */
if (fgets(buf, sizeof(buf), stdin) == NULL) {
return;
}
char dest[10];
/* strncpy with explicit bound + manual null-terminate */
strncpy(dest, buf, sizeof(dest) - 1);
dest[sizeof(dest) - 1] = '\0';
/* strncat with remaining-space bound */
size_t room = sizeof(dest) - strlen(dest) - 1;
strncat(dest, "world", room);
char msg[100];
/* snprintf is bounds-aware */
snprintf(msg, sizeof(msg), "%s says hello", buf);
/* Format string is a literal; buf is an argument */
printf("%s\n", buf);
char name[32];
/* %s with explicit width prevents overflow */
scanf("%31s", name);
}
void checked_alloc(int n) {
/* malloc result is NULL-checked before any dereference */
char *buf = malloc(n);
if (buf == NULL) {
return;
}
buf[0] = 'x';
buf[1] = 'y';
strncpy(buf, "ok", n - 1);
free(buf);
buf = NULL;
printf("done\n");
}
void run_safe_cmd(void) {
/* system() with a string literal — no command-injection surface */
system("ls -la");
}
int main(int argc, char *argv[]) {
(void)argc;
(void)argv;
safe_input();
checked_alloc(100);
run_safe_cmd();
return 0;
}
FILE:assets/sample_c_smells.c
/*
* sample_c_smells.c — labelled instances of every C-specific pattern
* the code-reviewer skill flags. Every smell is annotated inline with
* its CWE and the rule from languages/c.md.
*
* Refactored counterpart: sample_c_clean.c
*/
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
void unsafe_input(void) {
char buf[64];
/* RULE: banned function gets() — CWE-242, no bounds check */
gets(buf);
char dest[10];
/* RULE: banned function strcpy() — no bounds check */
strcpy(dest, buf);
/* RULE: banned function strcat() — no bounds check */
strcat(dest, "world");
char msg[100];
/* RULE: banned function sprintf() — no bounds check */
sprintf(msg, "%s says hello", buf);
/* RULE: format-string vulnerability — CWE-134, buf controls format */
printf(buf);
char name[32];
/* RULE: unbounded scanf — %s without width, CWE-120 */
scanf("%s", name);
}
void leaky_alloc(int n) {
/* RULE: malloc result not NULL-checked within 5 lines — CWE-690 */
char *buf = malloc(n);
buf[0] = 'x';
buf[1] = 'y';
strcpy(buf, "leak");
/* RULE: free without zeroing pointer — CWE-416 dangling */
free(buf);
printf("done\n");
}
void run_user_cmd(const char *cmd_from_user) {
/* RULE: system() with non-literal argument — CWE-78 command injection */
system(cmd_from_user);
}
int main(int argc, char *argv[]) {
if (argc > 1) {
unsafe_input();
leaky_alloc(100);
run_user_cmd(argv[1]);
}
return 0;
}
FILE:assets/sample_java_clean.java
// Sample Java file showing the fixed version of sample_java_smells.java.
// Same shape, but every smell has been resolved per the patterns documented
// in rules/universal.md and languages/java.md.
//
// Run:
// python scripts/code_quality_checker.py assets/sample_java_clean.java
//
// Expected: no HIGH Java-specific smells flagged.
package sample;
import java.io.FileInputStream;
import java.io.InputStream;
import java.sql.Connection;
import java.sql.PreparedStatement;
import java.sql.ResultSet;
import com.fasterxml.jackson.databind.ObjectMapper;
public class UserService {
// FIX: heavy object shared as a singleton instead of constructed per call.
private static final ObjectMapper MAPPER = new ObjectMapper();
// FIX: connection string injected from configuration, never inlined.
private final String connectionString;
public UserService(String connectionString) {
this.connectionString = connectionString;
}
public String getName(Connection conn, int id) {
// FIX: try-with-resources guarantees the stream and statement close.
try (InputStream config = new FileInputStream("/etc/config");
// FIX: parameterized query, no string concatenation.
PreparedStatement stmt =
conn.prepareStatement("SELECT name FROM users WHERE id = ?")) {
stmt.setInt(1, id);
try (ResultSet rs = stmt.executeQuery()) {
return rs.next() ? rs.getString("name") : null;
}
} catch (Exception e) {
// FIX: rethrow with context instead of swallowing.
throw new IllegalStateException("Failed to load user " + id, e);
}
}
public void process() {
try {
Thread.sleep(1000);
} catch (InterruptedException e) {
// FIX: restore the interrupt flag so cancellation still propagates.
Thread.currentThread().interrupt();
}
}
}
FILE:assets/sample_java_smells.java
// Sample Java file demonstrating the Java-specific patterns the code-reviewer
// skill detects. Each smell is labelled inline. This file is NOT meant to
// compile cleanly — it is a fixture for code_quality_checker.py and
// pr_analyzer.py.
//
// Run:
// python scripts/code_quality_checker.py assets/sample_java_smells.java
//
// Expected output: see expected_outputs/sample_java_smells_quality.json
package sample;
import java.io.FileInputStream;
import java.sql.Connection;
import java.sql.Statement;
import com.fasterxml.jackson.databind.ObjectMapper;
public class UserService {
// [hardcoded_secrets] hardcoded JDBC URL with password
public String connectionString = "jdbc:postgresql://prod/app?user=app&password=hunter2";
// [analyzer_disable] @SuppressWarnings without justification
@SuppressWarnings("unchecked")
public String getName(Connection conn, int id) throws Exception {
// [java_unclosed_resource] FileInputStream not in try-with-resources
FileInputStream fis = new FileInputStream("/etc/config");
// [java_per_use_heavy_object] new ObjectMapper() constructed per call
ObjectMapper mapper = new ObjectMapper();
try {
Statement stmt = conn.createStatement();
// [sql_concatenation] string concatenation builds SQL with user input
return stmt.executeQuery("SELECT name FROM users WHERE id = " + id).toString();
} catch (Exception e) {
// [java_empty_catch] empty catch swallows the exception
}
return null;
}
public void process() {
try {
Thread.sleep(1000);
} catch (InterruptedException e) {
// [java_swallowed_interrupt] interrupt flag not restored
// [console_log] printStackTrace used as error handling
e.printStackTrace();
}
}
public void log(String message) {
// [console_log] System.out.println left in production code
System.out.println(message);
}
}
FILE:expected_outputs/sample_csharp_clean_quality.json
{
"file": "/home/user/claude-skills/engineering-team/skills/code-reviewer/assets/sample_csharp_clean.cs",
"language": "csharp",
"metrics": {
"lines": {
"total": 101,
"code": 67,
"blank": 14,
"comment": 20
},
"functions": 13,
"classes": 3,
"avg_complexity": 1.3
},
"quality_score": 98,
"grade": "A",
"smells": [
{
"type": "csharp_unused_using",
"severity": "low",
"message": "'using System;' appears unused",
"location": "System"
},
{
"type": "csharp_unused_using",
"severity": "low",
"message": "'using System.Net.Http;' appears unused",
"location": "System.Net.Http"
},
{
"type": "csharp_unused_using",
"severity": "low",
"message": "'using System.Threading.Tasks;' appears unused",
"location": "System.Threading.Tasks"
},
{
"type": "csharp_unused_using",
"severity": "low",
"message": "'using System.Data.SqlClient;' appears unused",
"location": "System.Data.SqlClient"
},
{
"type": "csharp_unused_using",
"severity": "low",
"message": "'using Microsoft.Extensions.Logging;' appears unused",
"location": "Microsoft.Extensions.Logging"
},
{
"type": "csharp_unused_using",
"severity": "low",
"message": "'using Microsoft.Extensions.Options;' appears unused",
"location": "Microsoft.Extensions.Options"
}
],
"solid_violations": [],
"function_details": [
{
"name": "HttpClient",
"parameters": 0,
"lines": 2,
"complexity": 1
},
{
"name": "UserService",
"parameters": 3,
"lines": 8,
"complexity": 1
},
{
"name": "Task",
"parameters": 1,
"lines": 2,
"complexity": 2
},
{
"name": "HandleClickAsync",
"parameters": 0,
"lines": 8,
"complexity": 1
},
{
"name": "FetchAsync",
"parameters": 0,
"lines": 12,
"complexity": 2
},
{
"name": "InvalidOperationException",
"parameters": 1,
"lines": 21,
"complexity": 3
},
{
"name": "FirstValue",
"parameters": 1,
"lines": 4,
"complexity": 1
},
{
"name": "GetName",
"parameters": 1,
"lines": 4,
"complexity": 1
},
{
"name": "SqlConnection",
"parameters": 1,
"lines": 3,
"complexity": 1
},
{
"name": "SqlCommand",
"parameters": 2,
"lines": 7,
"complexity": 1
}
],
"class_details": [
{
"name": "DbOptions",
"methods": 0,
"lines": 7
},
{
"name": "UserService",
"methods": 8,
"lines": 75
},
{
"name": "User",
"methods": 0,
"lines": 3
}
]
}
FILE:expected_outputs/sample_csharp_smells_quality.json
{
"file": "/home/user/claude-skills/engineering-team/skills/code-reviewer/assets/sample_csharp_smells.cs",
"language": "csharp",
"metrics": {
"lines": {
"total": 79,
"code": 43,
"blank": 11,
"comment": 25
},
"functions": 7,
"classes": 1,
"avg_complexity": 1.3
},
"quality_score": 45,
"grade": "F",
"smells": [
{
"type": "csharp_async_void",
"severity": "high",
"message": "'async void HandleClick' \u2014 only safe for event handlers; prefer 'async Task'",
"location": "HandleClick"
},
{
"type": "csharp_blocking_async",
"severity": "high",
"message": "Blocking call on async operation ('.Result' / '.Wait()' / '.GetAwaiter().GetResult()') \u2014 can deadlock in ASP.NET contexts",
"location": "offset 430"
},
{
"type": "csharp_swallowed_exception",
"severity": "high",
"message": "Empty catch block swallows exceptions silently",
"location": "offset 853"
},
{
"type": "csharp_undisposed_idisposable",
"severity": "medium",
"message": "'HttpClient' looks like IDisposable but is not wrapped in 'using' / 'using var'",
"location": "offset 555"
},
{
"type": "csharp_undisposed_idisposable",
"severity": "medium",
"message": "'SqlCommand' looks like IDisposable but is not wrapped in 'using' / 'using var'",
"location": "offset 1289"
},
{
"type": "csharp_new_httpclient",
"severity": "medium",
"message": "'new HttpClient()' \u2014 prefer IHttpClientFactory or a long-lived static instance to avoid socket exhaustion",
"location": "offset 606"
},
{
"type": "csharp_missing_await",
"severity": "medium",
"message": "Async method called without 'await' \u2014 Task is discarded",
"location": "line 42"
},
{
"type": "csharp_unused_using",
"severity": "low",
"message": "'using System;' appears unused",
"location": "System"
},
{
"type": "csharp_unused_using",
"severity": "low",
"message": "'using System.Net.Http;' appears unused",
"location": "System.Net.Http"
},
{
"type": "csharp_unused_using",
"severity": "low",
"message": "'using System.Threading.Tasks;' appears unused",
"location": "System.Threading.Tasks"
},
{
"type": "csharp_unused_using",
"severity": "low",
"message": "'using System.Data.SqlClient;' appears unused",
"location": "System.Data.SqlClient"
},
{
"type": "csharp_unused_using",
"severity": "low",
"message": "'using System.Diagnostics.CodeAnalysis;' appears unused",
"location": "System.Diagnostics.CodeAnalysis"
}
],
"solid_violations": [],
"function_details": [
{
"name": "HandleClick",
"parameters": 2,
"lines": 9,
"complexity": 1
},
{
"name": "FetchAsync",
"parameters": 0,
"lines": 3,
"complexity": 1
},
{
"name": "HttpClient",
"parameters": 0,
"lines": 3,
"complexity": 1
},
{
"name": "HttpClient",
"parameters": 0,
"lines": 24,
"complexity": 3
},
{
"name": "Pointers",
"parameters": 0,
"lines": 11,
"complexity": 1
},
{
"name": "GetName",
"parameters": 2,
"lines": 5,
"complexity": 1
},
{
"name": "SqlCommand",
"parameters": 2,
"lines": 6,
"complexity": 1
}
],
"class_details": [
{
"name": "UserService",
"methods": 4,
"lines": 61
}
]
}
FILE:expected_outputs/sample_c_clean_quality.json
{
"file": "/home/user/claude-skills/engineering-team/skills/code-reviewer/assets/sample_c_clean.c",
"language": "c",
"metrics": {
"lines": {
"total": 72,
"code": 43,
"blank": 17,
"comment": 12
},
"functions": 4,
"classes": 0,
"avg_complexity": 1.8
},
"quality_score": 100,
"grade": "A",
"smells": [
{
"type": "long_function",
"severity": "medium",
"message": "Function 'safe_input' has 61 lines (max: 50)",
"location": "safe_input"
},
{
"type": "magic_number",
"severity": "low",
"message": "Magic number 100 should be a named constant",
"location": "line 29"
},
{
"type": "magic_number",
"severity": "low",
"message": "Magic number 100 should be a named constant",
"location": "line 68"
}
],
"solid_violations": [],
"function_details": [
{
"name": "safe_input",
"parameters": 1,
"lines": 61,
"complexity": 3
},
{
"name": "checked_alloc",
"parameters": 1,
"lines": 31,
"complexity": 2
},
{
"name": "run_safe_cmd",
"parameters": 1,
"lines": 15,
"complexity": 1
},
{
"name": "main",
"parameters": 2,
"lines": 9,
"complexity": 1
}
],
"class_details": []
}
FILE:expected_outputs/sample_c_smells_quality.json
{
"file": "/home/user/claude-skills/engineering-team/skills/code-reviewer/assets/sample_c_smells.c",
"language": "c",
"metrics": {
"lines": {
"total": 67,
"code": 37,
"blank": 17,
"comment": 13
},
"functions": 4,
"classes": 0,
"avg_complexity": 2.0
},
"quality_score": 4,
"grade": "F",
"smells": [
{
"type": "long_function",
"severity": "medium",
"message": "Function 'unsafe_input' has 54 lines (max: 50)",
"location": "unsafe_input"
},
{
"type": "magic_number",
"severity": "low",
"message": "Magic number 242 should be a named constant",
"location": "line 17"
},
{
"type": "magic_number",
"severity": "low",
"message": "Magic number 100 should be a named constant",
"location": "line 27"
},
{
"type": "magic_number",
"severity": "low",
"message": "Magic number 134 should be a named constant",
"location": "line 31"
},
{
"type": "magic_number",
"severity": "low",
"message": "Magic number 120 should be a named constant",
"location": "line 35"
},
{
"type": "magic_number",
"severity": "low",
"message": "Magic number 690 should be a named constant",
"location": "line 41"
},
{
"type": "magic_number",
"severity": "low",
"message": "Magic number 416 should be a named constant",
"location": "line 47"
},
{
"type": "magic_number",
"severity": "low",
"message": "Magic number 100 should be a named constant",
"location": "line 62"
},
{
"type": "c_banned_gets",
"severity": "high",
"message": "'gets()' is unsafe: no bounds check, removed from C11 (CWE-242)",
"location": "offset 117"
},
{
"type": "c_banned_strcpy",
"severity": "high",
"message": "'strcpy()' is unsafe: no bounds check \u2014 prefer strncpy or strlcpy",
"location": "offset 157"
},
{
"type": "c_banned_strcpy",
"severity": "high",
"message": "'strcpy()' is unsafe: no bounds check \u2014 prefer strncpy or strlcpy",
"location": "offset 447"
},
{
"type": "c_banned_strcat",
"severity": "high",
"message": "'strcat()' is unsafe: no bounds check \u2014 prefer strncat or strlcat",
"location": "offset 186"
},
{
"type": "c_banned_sprintf",
"severity": "high",
"message": "'sprintf()' is unsafe: no bounds check \u2014 prefer snprintf",
"location": "offset 238"
},
{
"type": "c_format_string",
"severity": "high",
"message": "'printf(buf)' uses a non-literal format string \u2014 CWE-134 format string vulnerability",
"location": "offset 284"
},
{
"type": "c_unbounded_scanf",
"severity": "high",
"message": "scanf '%s' without a width specifier \u2014 unbounded read can overflow the destination buffer",
"location": "offset 326"
},
{
"type": "c_malloc_unchecked",
"severity": "medium",
"message": "'buf' from malloc/calloc/realloc is not NULL-checked within 5 lines \u2014 dereferencing NULL is UB (CWE-690)",
"location": "line 36"
},
{
"type": "c_free_without_null",
"severity": "low",
"message": "'free(buf)' not followed by 'buf = NULL;' \u2014 dangling pointer can be reused (CWE-416)",
"location": "line 42"
},
{
"type": "c_system_non_literal",
"severity": "high",
"message": "'system(cmd_from_user)' with a non-literal argument \u2014 command injection (CWE-78); use execve with validated args",
"location": "offset 571"
}
],
"solid_violations": [],
"function_details": [
{
"name": "unsafe_input",
"parameters": 1,
"lines": 54,
"complexity": 2
},
{
"name": "leaky_alloc",
"parameters": 1,
"lines": 28,
"complexity": 2
},
{
"name": "run_user_cmd",
"parameters": 1,
"lines": 15,
"complexity": 2
},
{
"name": "main",
"parameters": 2,
"lines": 9,
"complexity": 2
}
],
"class_details": []
}
FILE:expected_outputs/sample_java_clean_quality.json
{
"file": "/home/user/claude-skills/engineering-team/skills/code-reviewer/assets/sample_java_clean.java",
"language": "java",
"metrics": {
"lines": {
"total": 56,
"code": 33,
"blank": 9,
"comment": 14
},
"functions": 3,
"classes": 1,
"avg_complexity": 2.0
},
"quality_score": 100,
"grade": "A",
"smells": [
{
"type": "magic_number",
"severity": "low",
"message": "Magic number 1000 should be a named constant",
"location": "line 49"
}
],
"solid_violations": [],
"function_details": [
{
"name": "UserService",
"parameters": 1,
"lines": 5,
"complexity": 1
},
{
"name": "getName",
"parameters": 2,
"lines": 17,
"complexity": 3
},
{
"name": "process",
"parameters": 0,
"lines": 10,
"complexity": 2
}
],
"class_details": [
{
"name": "UserService",
"methods": 3,
"lines": 38
}
]
}
FILE:expected_outputs/sample_java_smells_quality.json
{
"file": "/home/user/claude-skills/engineering-team/skills/code-reviewer/assets/sample_java_smells.java",
"language": "java",
"metrics": {
"lines": {
"total": 57,
"code": 29,
"blank": 10,
"comment": 18
},
"functions": 3,
"classes": 1,
"avg_complexity": 2.0
},
"quality_score": 68,
"grade": "D",
"smells": [
{
"type": "magic_number",
"severity": "low",
"message": "Magic number 1000 should be a named constant",
"location": "line 44"
},
{
"type": "java_empty_catch",
"severity": "high",
"message": "Empty catch block swallows exceptions silently",
"location": "offset 684"
},
{
"type": "java_print_stack_trace",
"severity": "medium",
"message": "'printStackTrace()' is not real error handling \u2014 log via a proper logger or rethrow with context",
"location": "offset 913"
},
{
"type": "java_swallowed_interrupt",
"severity": "high",
"message": "InterruptedException caught without 'Thread.currentThread().interrupt()' \u2014 breaks cooperative cancellation",
"location": "offset 841"
},
{
"type": "java_unclosed_resource",
"severity": "medium",
"message": "'FileInputStream' looks like an AutoCloseable but is not in a try-with-resources statement",
"location": "offset 366"
},
{
"type": "java_per_use_heavy_object",
"severity": "medium",
"message": "'new ObjectMapper()' is expensive \u2014 share a singleton instance instead of constructing per call",
"location": "offset 451"
}
],
"solid_violations": [],
"function_details": [
{
"name": "getName",
"parameters": 2,
"lines": 18,
"complexity": 3
},
{
"name": "process",
"parameters": 0,
"lines": 11,
"complexity": 2
},
{
"name": "log",
"parameters": 1,
"lines": 6,
"complexity": 1
}
],
"class_details": [
{
"name": "UserService",
"methods": 3,
"lines": 40
}
]
}
FILE:languages/c.md
---
language: c
extensions: [".c", ".h"]
---
# C — Language-Specific Review Notes
Load this file alongside `rules/universal.md`. Universal rules are not repeated here — only C-specific rules and idioms.
---
## PR Analyzer — C Risk Signals
- `printf` / debug `fprintf(stderr, ...)` statements left in production code
- `// TODO` / `// FIXME` comments near memory management code — high risk
- Disabled compiler warnings (`#pragma GCC diagnostic ignore`, `-w` flags in Makefile)
- Hardcoded credentials or keys in source
- Use of banned functions: `gets`, `strcpy`, `strcat`, `sprintf`, `scanf` without width limits
---
## Code Quality — C Checks
- Functions longer than 50 lines — C functions tend to grow organically and become hard to reason about
- Missing `NULL` check after `malloc` / `calloc` / `realloc`
- Return value of functions ignored without explicit `(void)` cast
- Global mutable state used across translation units without clear ownership
- Magic numbers without `#define` or `const` — especially sizes and offsets
- Mixed `malloc`/`free` ownership — unclear which caller is responsible for freeing
---
## Security
- Flag `gets()` — no bounds checking, always a buffer overflow; replace with `fgets()`
- Flag `strcpy()` / `strcat()` — use `strncpy()` / `strncat()` with explicit size, or `strlcpy()` / `strlcat()`
- Flag `sprintf()` — use `snprintf()` with explicit buffer size
- Flag `scanf("%s", buf)` without a width specifier — unbounded read
- Flag `strlen()` result used as a signed integer — potential truncation on 64-bit
- Flag user-controlled data used as a format string (`printf(user_input)`) — format string attack
- Flag integer arithmetic used as array index without bounds check
- Flag signed integer overflow — undefined behavior in C
---
## Async / Concurrency
- Flag shared global or `static` variables accessed from multiple threads without a mutex or `_Atomic`
- Flag `pthread_mutex_t` / `sem_t` not initialized before use
- Flag signal handlers that call non-async-signal-safe functions (`malloc`, `printf`, etc.)
- Flag `volatile` used as a substitute for proper synchronization — it is not sufficient
- Flag lock acquisition order inconsistency across call sites — deadlock risk
---
## Resource Management
- Flag every `malloc` / `calloc` / `realloc` path — verify a matching `free` exists on all exit paths
- Flag `fopen` without a matching `fclose` on all paths including error paths
- Flag `dup` / `socket` / `open` file descriptors not closed on all paths
- Flag stack-allocated VLAs (variable-length arrays) of unbounded size — stack overflow risk
- Flag `realloc` return value assigned directly to the source pointer — leaks on failure
---
## Exception Handling
- Flag ignored return values from `malloc`, `fopen`, `read`, `write`, `close` — all can fail
- Flag `errno` checked after a function that doesn't set it, or not checked immediately after one that does
- Flag `perror` / `strerror` as the sole error handling in library code — propagate errors to callers
- Flag functions that return `-1` on error without documenting which `errno` values are possible
- Flag `assert()` used for runtime error handling — disabled by `NDEBUG` in production builds
---
## Performance
- Flag `strlen()` called repeatedly on the same string in a loop — cache the result
- Flag unnecessary copies of large structs passed by value — pass by pointer
- Flag `memcpy` / `memset` on overlapping regions — use `memmove` for overlapping
- Flag repeated heap allocations in a tight loop — consider a pool or stack allocation
- Flag `volatile` on variables not accessed by hardware or signal handlers — prevents optimization
---
## Idioms and Best Practices
### Memory Safety
- Every pointer must have a clear owner responsible for freeing it — document ownership in comments
- Set pointers to `NULL` immediately after `free` to catch use-after-free early
- Prefer `calloc` over `malloc` + `memset` for zero-initialized allocations
- Use `const` on pointer parameters that the function does not modify
### Defensive Coding
- Always check `NULL` returns from allocation functions
- Use `size_t` for sizes and counts — never `int`
- Prefer `snprintf` and `fgets` over any unbounded string function
- Compile with `-Wall -Wextra -Werror` and treat warnings as errors
### Portability
- Do not assume pointer size equals `int` size — use `intptr_t` / `uintptr_t`
- Do not rely on undefined behavior for performance — use compiler intrinsics instead
- Use `stdint.h` types (`uint32_t`, `int64_t`) for fixed-width requirements
FILE:languages/cpp.md
---
language: cpp
extensions: [".cpp", ".cc", ".cxx", ".hpp", ".hh", ".hxx"]
---
# C++ — Language-Specific Review Notes
Load this file alongside `rules/universal.md`. Universal rules are not repeated here — only C++-specific rules and idioms.
---
## PR Analyzer — C++ Risk Signals
- Raw `new` / `delete` outside of smart pointer wrappers
- `reinterpret_cast` — almost always a red flag; require justification
- Disabled compiler warnings (`#pragma warning(disable:...)`, `-w`)
- `// TODO` / `// FIXME` near ownership or lifetime code
- Hardcoded credentials or keys in source
- Use of deprecated C-style functions: `strcpy`, `sprintf`, `gets`
---
## Code Quality — C++ Checks
- Raw owning pointers (`T*`) used where `unique_ptr` / `shared_ptr` would express ownership
- `shared_ptr` overused where `unique_ptr` suffices — implies shared ownership unnecessarily
- `std::endl` used in hot paths — flushes the buffer every call; prefer `'\n'`
- Implicit conversions between signed and unsigned integers
- Virtual destructor missing on base classes with virtual methods
- `catch (...)` swallowing all exceptions without logging or re-throwing
---
## Security
- Flag `reinterpret_cast` on user-controlled data — potential type confusion
- Flag raw array indexing without bounds check — use `.at()` or assert bounds
- Flag `std::string` data passed to C APIs without null-termination guarantee — use `.c_str()`
- Flag hardcoded buffer sizes — derive from `sizeof` or use `std::array<T, N>`
- Flag `sscanf` / `sprintf` — use `std::istringstream` or `std::format` (C++20)
- Flag user-controlled data used as a format string
---
## Async / Concurrency
- Flag `std::shared_ptr` accessed from multiple threads — the pointer itself is not thread-safe for write; use `std::atomic<std::shared_ptr<T>>` (C++20) or external locking
- Flag `std::vector` / `std::map` mutated from multiple threads without a mutex
- Flag `std::mutex` locked twice in the same thread without `std::recursive_mutex` — deadlock
- Flag detached threads (`std::thread::detach`) with no lifetime coordination
- Flag `volatile` used instead of `std::atomic` for inter-thread communication
---
## Resource Management
- Flag raw `new` returning an owning pointer — wrap immediately in `std::make_unique` or `std::make_shared`
- Flag `delete` called manually outside of a destructor or smart pointer — ownership confusion
- Flag RAII violations — resources acquired in constructor but not released via destructor
- Flag `std::ifstream` / `std::ofstream` not checked for open failure before use
- Flag exceptions thrown from destructors — causes `std::terminate` if thrown during stack unwinding
---
## Exception Handling
- Flag `catch (...)` that swallows exceptions without logging or re-throwing
- Flag exceptions thrown from destructors — wrap in `try/catch` inside the destructor
- Flag `noexcept` on functions that can actually throw — causes `std::terminate`
- Flag exception specifications (`throw(...)`) — deprecated since C++11, removed in C++17
- Flag using exceptions for control flow in performance-critical paths
---
## Performance
- Flag pass-by-value for non-trivial types where pass-by-const-reference suffices
- Flag `std::vector::push_back` in a loop without `reserve` when size is known — repeated reallocations
- Flag `std::map` used where `std::unordered_map` would give O(1) lookup
- Flag `std::endl` in loops — prefer `'\n'` to avoid repeated buffer flushes
- Flag unnecessary copies from missing `std::move` on local temporaries being returned or passed
---
## Idioms and Best Practices
### Ownership and Lifetime
- Prefer `std::unique_ptr` for sole ownership, `std::shared_ptr` only for shared ownership
- Prefer `std::make_unique` / `std::make_shared` over `new` — exception-safe
- Use `std::weak_ptr` to break `shared_ptr` cycles
- Never use raw owning pointers in new code — they are for non-owning observation only
### Modern C++ (17/20)
- Prefer `std::optional<T>` over sentinel values or nullable pointers for optional returns
- Prefer `std::variant` over tagged unions
- Prefer `std::string_view` over `const std::string&` for read-only string parameters
- Prefer range-based `for` loops over index loops where the index isn't needed
- Prefer `if constexpr` over `#ifdef` for compile-time branching
### Type Safety
- Prefer `static_cast` over C-style casts — explicit and auditable
- Avoid `reinterpret_cast` except in low-level I/O or FFI code with a comment
- Use `enum class` over plain `enum` to avoid implicit integer conversions
FILE:languages/csharp.md
---
language: csharp
extensions: [".cs", ".csx", ".razor", ".cshtml"]
---
# C# / .NET — Language-Specific Review Notes
Load this file alongside `rules/universal.md`. Universal rules are not repeated here — only C#-specific rules and idioms.
---
## PR Analyzer — C# Risk Signals
- `#pragma warning disable` and `[SuppressMessage]` — verify they are justified
- `unsafe { }` blocks — require explicit sign-off
- Null-forgiving operator (`!`) used broadly without justification
- `dynamic` used outside of interop scenarios
- Hardcoded connection strings in source files
---
## Code Quality — C# Checks
- `async void` methods (except event handlers)
- `Task` returned but not awaited
- `IDisposable` objects not in `using` / `using var`
- Bare `catch { }` or `catch (Exception e) { }` swallowing silently
- Nullable reference types feature disabled at project level
---
## Security
- Flag raw string interpolation in SQL queries — require parameterized queries (`SqlCommand`) or EF Core
- Flag missing `[ValidateAntiForgeryToken]` on state-changing controller actions
- Flag user-controlled data passed to `Process.Start()` or `File` APIs without validation
- Flag hardcoded connection strings — require `appsettings.json` + secrets management
- Flag `[AllowAnonymous]` on endpoints that should be protected
---
## Async / Await
- Flag `async void` methods outside of event handlers — cannot be awaited and swallow exceptions
- Flag `.Result`, `.Wait()`, or `.GetAwaiter().GetResult()` on `Task` — causes deadlocks in ASP.NET contexts
- Flag missing `ConfigureAwait(false)` in library (non-application) code
- Flag `Task.Run()` wrapping synchronous code inside ASP.NET request handlers unnecessarily
- Flag `CancellationToken` not threaded through to downstream async calls
---
## Resource Management
- Flag `IDisposable` objects (`SqlConnection`, `HttpClient`, `FileStream`, etc.) not wrapped in `using` / `using var`
- Flag `HttpClient` instantiated with `new` inside a method — use `IHttpClientFactory` or a shared static instance to avoid socket exhaustion
- Flag `DbContext` registered as a singleton in DI — it must be scoped
- Flag `MemoryStream` / `MemoryCache` growing unboundedly without eviction policy
---
## Exception Handling
- Flag `catch { }` or `catch (Exception) { }` with no logging or re-throw — silent swallow
- Flag `catch (Exception e) { throw e; }` — resets the stack trace; use `throw;` instead
- Flag catching `Exception` when a specific type (`IOException`, `HttpRequestException`) is appropriate
- Flag exception filters (`when`) used for side effects that suppress the exception
- Flag exceptions used for control flow in hot paths — use `Try*` pattern methods instead
---
## Performance
- Flag `.ToList()` / `.ToArray()` on `IQueryable` before filtering — forces all rows into memory; filter server-side first
- Flag `string` concatenation in loops — use `StringBuilder`
- Flag `Enumerable.Count()` on `IQueryable` when only an existence check is needed — use `Any()`
- Flag `await` in a loop where `Task.WhenAll()` would parallelize the work
- Flag synchronous file or network I/O in an `async` method — use the async overload
---
## Idioms and Best Practices
### Null Safety
- Ensure `<Nullable>enable</Nullable>` is set in the project file
- Flag excessive use of `!` (null-forgiving) without a comment explaining why
- Prefer `is null` / `is not null` over `== null` for null checks
### LINQ
- Flag `First()` where `FirstOrDefault()` is safer
- Flag complex LINQ chains that would be clearer as explicit loops
### Modern C# (10+)
- Prefer `record` types for immutable data carriers
- Prefer `switch` expressions over `switch` statements where a value is returned
- Prefer primary constructors (C# 12) for simple dependency injection
- Prefer file-scoped namespaces (`namespace Foo;`) over block-scoped
- Prefer `is` pattern matching over explicit casts
FILE:languages/dart.md
---
language: dart
extensions: [".dart"]
---
# Dart / Flutter — Language-Specific Review Notes
Load this file alongside `rules/universal.md`. Universal rules are not repeated here — only Dart and Flutter-specific rules and idioms.
---
## PR Analyzer — Dart / Flutter Risk Signals
- `print()` statements left in production code — use a logging package
- `// ignore:` lint suppression comments — verify they are justified
- `!` null assertion operator used broadly without justification
- Hardcoded API keys, tokens, or URLs in Dart source — use environment variables or a secrets package
- `TODO` / `FIXME` near widget lifecycle or state management code
---
## Code Quality — Dart Checks
- `dynamic` used where a concrete type is known — defeats static analysis
- `!` (null assertion) used broadly — prefer null-safe patterns
- `StatefulWidget` used where `StatelessWidget` suffices — prefer stateless
- `setState` called with heavy computation inside — offload before calling
- `BuildContext` used across async gaps without checking `mounted`
- Missing `const` constructor on widgets that could be constant
---
## Security
- Flag API keys or secrets hardcoded in Dart source or `pubspec.yaml` — use `--dart-define` or a secrets manager
- Flag `http` package used without certificate validation disabled intentionally
- Flag `SharedPreferences` used to store sensitive data — use `flutter_secure_storage`
- Flag user-controlled input used in `dart:io` file path operations without sanitization
- Flag `WebView` loading arbitrary user-supplied URLs without validation
- Flag deep link / URL scheme handlers that don't validate the incoming URL before acting on it
---
## Async / Concurrency
- Flag `BuildContext` used after an `await` without checking `if (!mounted) return` — context may be invalid
- Flag `Future` returned but not `await`-ed and without `.catchError()` or `unawaited()` — floating future
- Flag `Isolate.spawn` without a clear message-passing protocol
- Flag heavy computation on the main isolate — offload with `compute()` or `Isolate.run()`
- Flag `StreamController` not closed when the owning widget is disposed — memory leak
- Flag `async*` / `yield*` generators with no error handling on the stream consumer side
---
## Resource Management
- Flag `StreamController` not closed in `dispose()`
- Flag `AnimationController` not disposed in `dispose()`
- Flag `TextEditingController` / `FocusNode` / `ScrollController` not disposed in `dispose()`
- Flag `Timer` not cancelled in `dispose()`
- Flag listeners added to `ChangeNotifier` / `ValueNotifier` without a corresponding `removeListener`
---
## Exception Handling
- Flag empty `catch` blocks — swallowed errors
- Flag `catchError` with no handler body — silent failure
- Flag `Future.error` not surfaced to the UI — show an error state
- Flag `FlutterError.onError` overridden without calling the original handler
- Prefer typed `on ExceptionType catch (e)` over generic `catch (e)` where the exception type is known
---
## Performance
- Flag `setState` called for changes that only affect a small subtree — use `ValueNotifier` / `provider` / `Riverpod` to scope rebuilds
- Flag expensive computation inside `build()` — move to `initState`, a controller, or a `FutureBuilder`
- Flag `ListView` without `ListView.builder` for long or infinite lists — builds all children at once
- Flag missing `const` on widgets that never change — prevents unnecessary rebuilds
- Flag `Image.network` without a caching package in a list — re-downloads on every scroll
- Flag `RepaintBoundary` missing around frequently-repainted widgets (animations, counters)
---
## Idioms and Best Practices
### Null Safety
- Prefer `?.` safe navigation and `??` null coalescing over `!` assertions
- Use `late` only when initialization is guaranteed before first access — document why
- Prefer early returns over deeply nested null checks
### Flutter Widget Patterns
- Prefer `StatelessWidget` + external state management over `StatefulWidget` for business logic
- Keep `build()` methods pure — no side effects, no heavy computation
- Extract repeated widget subtrees into named widget classes, not just methods, for better rebuild granularity
- Use `const` constructors wherever possible — compile-time constant widgets skip rebuilds entirely
### State Management
- Do not mix multiple state management approaches in the same feature
- Flag business logic inside `build()` — it belongs in a ViewModel, Notifier, or BLoC
- Prefer `Riverpod` / `provider` / `BLoC` over raw `setState` for anything beyond local UI state
### Modern Dart (3.x)
- Prefer `sealed` classes for exhaustive pattern matching on domain types
- Use records (`(int, String)`) for lightweight multi-value returns instead of ad hoc classes
- Use `switch` expressions with pattern matching instead of long `if/else` chains
- Prefer `final` for local variables — immutability by default
FILE:languages/go.md
---
language: go
extensions: [".go"]
---
# Go — Language-Specific Review Notes
Load this file alongside `rules/universal.md`. Universal rules are not repeated here — only Go-specific rules and idioms.
---
## PR Analyzer — Go Risk Signals
- `fmt.Println` / `log.Println` debug statements left in production code
- `//nolint` comments — verify they are justified
- `unsafe` package imports — require explicit sign-off
- Hardcoded credentials or tokens in source
---
## Code Quality — Go Checks
- Errors returned but not checked (`_ = someFunc()`)
- `panic()` used outside of package initialization
- Goroutines started without a clear lifetime or cancellation path
- `interface{}` / `any` used where a concrete type or typed interface would work
- Missing context propagation (`context.Context` not threaded through call chains)
---
## Security
- Flag `database/sql` queries built with `fmt.Sprintf` — require `?` / `$N` placeholders
- Flag `os/exec` calls with user-controlled arguments without sanitization
- Flag `html/template` bypassed in favor of `text/template` for HTML output
- Flag `http.ListenAndServeTLS` with `InsecureSkipVerify: true`
---
## Async / Concurrency
- Flag goroutines started with no clear lifetime or cancellation path — always pass `context.Context`
- Flag goroutines that write to a channel with no receiver and no `select` default — causes a leak
- Flag `time.Sleep()` used inside a goroutine as a synchronization mechanism
- Flag `sync.WaitGroup.Add()` called inside the goroutine it tracks — race condition
- Flag `sync.Mutex` copied by value — must always be used as a pointer or embedded in a struct
---
## Resource Management
- Flag `http.Response.Body` not closed after reading — even on error paths (`defer resp.Body.Close()`)
- Flag `os.File` not closed — use `defer f.Close()` immediately after opening
- Flag `rows.Close()` missing after `sql.Query()` — leaks the DB connection
- Flag `context.WithCancel` / `context.WithTimeout` cancel function not called — context and resources leak
---
## Exception Handling
- Flag errors assigned to `_` without a comment explaining why it is safe to ignore
- Flag errors not wrapped with `fmt.Errorf("...: %w", err)` — loses stack context
- Flag `errors.New` / `fmt.Errorf` strings starting with a capital letter or ending in punctuation — violates Go conventions
- Flag `panic()` used for expected runtime errors — reserve for programming errors and unrecoverable states
- Flag `recover()` used to silently swallow panics without logging
---
## Performance
- Flag `fmt.Sprintf` used for simple string concatenation — use `strings.Builder` or `+` for small cases
- Flag `append()` in a tight loop without pre-allocating slice capacity — use `make([]T, 0, n)`
- Flag `json.Marshal` / `json.Unmarshal` on large structs in hot paths — consider `json.Encoder` / streaming
- Flag goroutines spawned per-request without a worker pool for CPU-bound tasks
---
## Idioms and Best Practices
### Error Handling
- All returned errors must be checked — never assign to `_` without a comment
- Prefer wrapping with `fmt.Errorf("...: %w", err)` for stack context
- Use `errors.Is` / `errors.As` for error inspection — never string comparison
### Concurrency
- Every goroutine must have an owner responsible for its lifetime
- Always pass `context.Context` as the first argument to functions that do I/O or block
- Prefer `sync.WaitGroup` or `errgroup` over ad-hoc channel coordination
### Modern Go (1.18+)
- Prefer generics over `interface{}` for container types and utility functions
- Use `any` (alias for `interface{}`) in new code for readability
FILE:languages/java.md
---
language: java
extensions: [".java"]
---
# Java — Language-Specific Review Notes
Load this file alongside `rules/universal.md`. Universal rules are not repeated here — only Java-specific rules and idioms.
---
## PR Analyzer — Java Risk Signals
- `System.out.println` / `e.printStackTrace()` left in production code
- `@SuppressWarnings` annotations — verify they are justified
- Hardcoded JDBC URLs or credentials in source
- Raw type usage (`List`, `Map` without generics)
---
## Code Quality — Java Checks
- Empty `catch` blocks swallowing exceptions silently
- Checked exceptions caught and not re-thrown with context
- `Closeable` / `AutoCloseable` resources not in try-with-resources
- Raw type usage — defeats generics type safety
- Missing `@Override` on overriding methods
- `InterruptedException` caught without calling `Thread.currentThread().interrupt()`
---
## Security
- Flag JPQL / HQL or native SQL string concatenation — require named parameters or `CriteriaBuilder`
- Flag `@RequestMapping` without explicit HTTP method restriction on state-changing endpoints
- Flag user-controlled input passed to `Runtime.exec()` or `ProcessBuilder` without validation
- Flag `ObjectInputStream.readObject()` on untrusted data — unsafe deserialization
- Flag hardcoded JDBC URLs or credentials — require environment variables or a vault
---
## Async / Concurrency
- Flag `ExecutorService.submit()` return value ignored — exceptions are swallowed
- Flag `Thread.sleep()` used as a synchronization mechanism — use `CountDownLatch`, `CompletableFuture`, or `await()`
- Flag `CompletableFuture` chains with no `.exceptionally()` or `.handle()` terminal handler
- Flag `InterruptedException` caught without calling `Thread.currentThread().interrupt()`
- Flag `synchronized` on a non-final field — the lock object can be replaced
- Flag `HashMap` used in multi-threaded context — use `ConcurrentHashMap`
---
## Resource Management
- Flag `InputStream`, `OutputStream`, `Connection`, `ResultSet`, `PreparedStatement` not wrapped in try-with-resources
- Flag manual `finally { resource.close() }` — replace with try-with-resources
- Flag `HttpURLConnection` not disconnected after use
- Flag JDBC `Connection` obtained from a pool and not returned (missing `close()`) on all paths
- Flag `static` `HttpClient` or `Connection` fields shared across threads without connection pool management
---
## Exception Handling
- Flag empty `catch` blocks — `catch (Exception e) {}`
- Flag `InterruptedException` caught without `Thread.currentThread().interrupt()` — breaks cooperative cancellation
- Flag checked exceptions swallowed in a `catch` and not re-thrown or logged with context
- Flag `throw new RuntimeException(e)` without a descriptive message — loses context
- Flag `printStackTrace()` as the sole error handling — use a proper logger
---
## Performance
- Flag `String` concatenation in loops — use `StringBuilder`
- Flag `List.contains()` / `Map.get()` in a loop on large collections — review data structure choice
- Flag N+1 JPA / Hibernate queries — use `JOIN FETCH` or `@BatchSize`
- Flag `new ObjectMapper()` / `new Gson()` instantiated per-request — share a singleton
- Flag `ResultSet` fully iterated when only the first result is needed — use `LIMIT 1` in the query
---
## Idioms and Best Practices
### Null Safety
- Prefer returning `Optional<T>` over `null` from methods
- Flag unchecked dereferences without a prior null guard
- Do not catch `NullPointerException` — fix the root cause instead
### Collections and Streams
- Flag `==` used to compare `String` or boxed types — use `.equals()`
- Flag `.collect(Collectors.toList())` where `.toList()` (Java 16+) suffices
- Flag premature `.stream().collect()` round-trips that could be a single-pass operation
### Generics
- Flag raw types in any new code — always parameterize (`List<String>`, not `List`)
- Flag unchecked cast warnings suppressed without explanation
### Modern Java (11+)
- Prefer `var` for local variables where the type is obvious from the right-hand side
- Prefer records for pure data carriers over manual POJOs with getters/setters
- Prefer `instanceof` pattern matching (`if (obj instanceof String s)`) over explicit casts
- Prefer `switch` expressions over `switch` statements where a value is returned
FILE:languages/kotlin.md
---
language: kotlin
extensions: [".kt", ".kts"]
---
# Kotlin — Language-Specific Review Notes
Load this file alongside `rules/universal.md`. Universal rules are not repeated here — only Kotlin-specific rules and idioms.
---
## PR Analyzer — Kotlin Risk Signals
- `println()` statements left in production code
- `@Suppress` annotations — verify they are justified
- `!!` (not-null assertion) used broadly without justification
- Hardcoded credentials or API keys in source
---
## Code Quality — Kotlin Checks
- `!!` used broadly — prefer `?.let`, `?:`, or `requireNotNull()`
- `lateinit var` accessed before initialization
- Coroutines launched with `GlobalScope` — prefer scoped coroutines
- `runBlocking` used outside of tests or top-level entry points
---
## Security
- Flag Room / SQLite queries built with string concatenation — require parameterized queries
- Flag `WebView.loadUrl()` with user-controlled input without validation
- Flag credentials stored in `SharedPreferences` — require `EncryptedSharedPreferences` or Keychain
---
## Async / Coroutines
- Flag `GlobalScope.launch` / `GlobalScope.async` in production code — use a structured scope
- Flag `runBlocking` outside of tests or top-level main functions
- Flag `launch` / `async` without a `CoroutineExceptionHandler` or `supervisorScope` where individual failures should not cancel siblings
- Flag `Dispatchers.Main` used for CPU-bound work — use `Dispatchers.Default`
- Flag coroutine cancellation not respected — long loops should check `isActive` or call `yield()`
---
## Resource Management
- Flag `Closeable` / `AutoCloseable` not wrapped in `.use { }` (Kotlin's try-with-resources equivalent)
- Flag `OkHttpClient` / `Retrofit` instantiated per-request — share a singleton
- Flag `BroadcastReceiver` registered without a corresponding `unregisterReceiver` — memory / battery leak
- Flag coroutines that hold a resource across a `suspend` point without structured cleanup in `finally`
---
## Exception Handling
- Flag `runCatching { }.getOrNull()` used broadly — silently swallows all exceptions
- Flag `catch (e: Exception)` in coroutines without re-throwing `CancellationException` — breaks structured concurrency
- Flag empty `catch` blocks
- Flag `throw RuntimeException(e)` without a descriptive message
- Prefer typed `sealed class` error hierarchies over raw exceptions for domain errors in coroutine flows
---
## Performance
- Flag `buildString` / `StringBuilder` not used for multi-step string construction in loops
- Flag `List` used for frequent `contains` checks — prefer `Set`
- Flag `flow.collect {}` re-subscribing on every recomposition in Jetpack Compose — use `collectAsStateWithLifecycle`
- Flag `Dispatchers.IO` used for CPU-bound work — use `Dispatchers.Default`
- Flag `suspend` functions calling non-suspend blocking APIs directly — wrap with `withContext(Dispatchers.IO)`
---
## Idioms and Best Practices
### Null Safety
- Prefer safe call (`?.`) and Elvis operator (`?:`) over `!!`
- Use `requireNotNull()` / `checkNotNull()` with a descriptive message when null means a programming error
- Prefer `val` over `var` — immutability by default
### Modern Kotlin
- Prefer `data class` for value carriers
- Prefer `sealed class` / `sealed interface` for exhaustive `when` expressions
- Prefer extension functions over utility classes
- Prefer `object` declarations for singletons
FILE:languages/php.md
---
language: php
extensions: [".php", ".phtml", ".php3", ".php4", ".php5", ".phps"]
---
# PHP — Language-Specific Review Notes
Load this file alongside `rules/universal.md`. Universal rules are not repeated here — only PHP-specific rules and idioms.
---
## PR Analyzer — PHP Risk Signals
- `var_dump` / `print_r` / `echo` debug statements left in production code
- `@` error suppression operator — masks real errors; verify it is justified
- `// phpcs:ignore` / `// phpstan-ignore` comments — verify they are justified
- Hardcoded credentials, database passwords, or API keys in source
- `eval()` anywhere — almost always a security issue
- `$_GET` / `$_POST` / `$_REQUEST` / `$_COOKIE` used without sanitization
---
## Code Quality — PHP Checks
- Missing type declarations on function parameters and return types
- `mixed` return type used broadly — tighten to specific types
- Global variables (`global $var`) — pass dependencies explicitly
- Long functions (>50 lines) — PHP functions tend to accumulate logic
- `isset()` / `empty()` used to mask type errors instead of fixing the root cause
- Missing `strict_types=1` declaration at the top of the file
---
## Security
- Flag `$_GET` / `$_POST` / `$_REQUEST` used directly in SQL queries — require PDO prepared statements
- Flag `mysqli_query($conn, "SELECT ... WHERE id = " . $_GET['id'])` — SQL injection
- Flag `echo $_GET['name']` or any unescaped output — XSS; use `htmlspecialchars()` with `ENT_QUOTES`
- Flag `include` / `require` with user-controlled paths — local/remote file inclusion
- Flag `eval()` — remote code execution risk; no legitimate use in application code
- Flag `shell_exec` / `exec` / `system` / `passthru` with user-controlled input — command injection
- Flag `unserialize()` on untrusted data — arbitrary object instantiation and code execution
- Flag `move_uploaded_file` without MIME type validation and extension whitelist — file upload attack
- Flag `header("Location: " . $_GET['url'])` without validation — open redirect
- Flag missing CSRF token validation on state-changing form endpoints
---
## Async / Concurrency
- Flag long-running synchronous operations in a request cycle — offload to a queue (Laravel Queue, RabbitMQ)
- Flag `sleep()` used inside a request handler — blocks the PHP-FPM worker
- Flag shared mutable state in `static` properties accessed across requests in long-running processes (Swoole, RoadRunner)
- Flag missing idempotency in queued jobs — jobs can be retried on failure
---
## Resource Management
- Flag database connections not closed or returned to the pool (`$pdo = null` or `$conn->close()`)
- Flag `fopen` / `fwrite` without a matching `fclose` on all paths
- Flag `curl_init` without `curl_close` — leaks the curl handle
- Flag unbounded file uploads with no size or type restriction
- Flag sessions not explicitly closed (`session_write_close()`) before long operations — session locking blocks other requests
---
## Exception Handling
- Flag empty `catch` blocks — swallowed exceptions
- Flag `catch (Exception $e) {}` without logging — silent failure
- Flag `die()` / `exit()` used for error handling in library code — use exceptions
- Flag `@` operator used to suppress errors from functions that can fail — check return values instead
- Flag `trigger_error` used in new code — prefer exceptions
---
## Performance
- Flag N+1 Eloquent / Doctrine queries — use eager loading (`with()`, `load()`, `join`)
- Flag `count($array)` called repeatedly in a loop condition — cache the result
- Flag `array_push($arr, $val)` — use `$arr[] = $val` which is faster
- Flag `in_array` on large arrays without the strict third argument — use `isset` on a flipped array for O(1) lookup
- Flag `file_get_contents` on remote URLs in a request cycle — use an HTTP client with timeout and async where possible
- Flag Eloquent `all()` without pagination — loads entire table into memory
---
## Idioms and Best Practices
### Type Safety
- Always declare `declare(strict_types=1)` at the top of every file
- Use union types (`int|string`) and nullable types (`?string`) rather than `mixed`
- Use typed properties on classes — avoid untyped `public $foo`
- Use constructor promotion for simple value objects
### Modern PHP (8.x)
- Prefer `match` expressions over `switch` — strict comparison, no fall-through
- Use named arguments for functions with many optional parameters
- Use `enum` for fixed sets of values instead of class constants
- Use `readonly` properties for immutable data
- Use nullsafe operator (`?->`) instead of nested `isset` checks
- Use `first-class callable syntax` (`strlen(...)`) instead of string references
### Laravel / Symfony Specific
- Keep controllers thin — logic belongs in service classes or action classes
- Use form requests for validation — never validate in the controller directly
- Prefer Eloquent relationships over manual joins for readability
- Flag raw queries where the ORM can express the same intent safely
FILE:languages/python.md
---
language: python
extensions: [".py"]
---
# Python — Language-Specific Review Notes
Load this file alongside `rules/universal.md`. Universal rules are not repeated here — only Python-specific rules and idioms.
---
## PR Analyzer — Python Risk Signals
- `print()` statements left in production code
- `# noqa` and `# type: ignore` comments — verify they are justified
- `eval()` / `exec()` with any user-controlled input
- `pickle` used to deserialize untrusted data
- Hardcoded credentials or tokens in source
---
## Code Quality — Python Checks
- Bare `except:` or `except Exception:` swallowing silently
- Mutable default arguments (`def foo(items=[])`) — shared across calls
- `import *` — pollutes namespace and hides dependencies
- Missing type hints on public functions and methods
- `assert` used for runtime validation — stripped by `-O` flag
---
## Security
- Flag `eval()` / `exec()` with any user-controlled input
- Flag `pickle.loads()` on untrusted data — use `json` or `msgpack`
- Flag `subprocess` calls with `shell=True` and user input
- Flag `flask.render_template_string()` with user data (SSTI)
- Flag `SECRET_KEY` / `DEBUG = True` committed to source
---
## Async
- Flag `asyncio.get_event_loop().run_until_complete()` inside an already-running loop
- Flag mixing `threading` and `asyncio` without a clear bridge (`run_in_executor`)
- Flag CPU-bound work inside an `async def` without offloading to `ProcessPoolExecutor`
- Flag `time.sleep()` inside async functions — use `await asyncio.sleep()`
---
## Resource Management
- Flag `open()` not used as a context manager (`with open(...) as f`)
- Flag `requests.Session` created per-request instead of shared/reused
- Flag database connections not closed or returned to a pool on all paths
- Flag large files read entirely into memory with `.read()` — prefer streaming / chunked reads
---
## Exception Handling
- Flag bare `except:` — catches `BaseException` including `KeyboardInterrupt` and `SystemExit`
- Flag `except Exception: pass` — silently swallows errors
- Flag re-raising with `raise e` instead of `raise` — loses the original traceback
- Flag `except` clause too broad when the `try` block covers multiple operations with different failure modes — split them
---
## Performance
- Flag `+` string concatenation in loops — use `"".join()`
- Flag repeated `re.compile()` inside a loop — compile once at module level
- Flag `list.append()` in a loop where a list comprehension would be more efficient
- Flag `in` membership tests on `list` where the collection is large — use `set`
- Flag loading entire large files into memory — prefer streaming or chunked reads
---
## Idioms and Best Practices
### Type Safety
- All public functions and methods should have type annotations
- Prefer `X | None` (Python 3.10+) over `Optional[X]`
- Use `TypedDict` or `dataclass` over plain `dict` for structured data
### Modern Python (3.10+)
- Prefer `match` statements over long `if/elif` chains
- Prefer `dataclass` or `NamedTuple` over plain classes for data carriers
- Prefer `pathlib.Path` over `os.path` for file operations
- Prefer f-strings over `.format()` or `%` formatting
### None Safety
- Prefer explicit `if x is None` over falsy checks when `0` or `""` are valid values
- Flag functions returning `None` implicitly — make it explicit or raise
FILE:languages/ruby.md
---
language: ruby
extensions: [".rb", ".rake", ".gemspec", ".ru"]
---
# Ruby — Language-Specific Review Notes
Load this file alongside `rules/universal.md`. Universal rules are not repeated here — only Ruby-specific rules and idioms.
---
## PR Analyzer — Ruby Risk Signals
- `puts` / `p` / `pp` debug statements left in production code
- `# rubocop:disable` comments — verify they are justified
- `eval` / `instance_eval` / `class_eval` with user-controlled input
- Hardcoded credentials, tokens, or `SECRET_KEY_BASE` in source
- `binding.pry` / `byebug` / `debugger` left in code
---
## Code Quality — Ruby Checks
- Methods longer than 15 lines — Ruby idioms favor very small methods
- Classes with more than 10 public methods — possible god object
- `rescue Exception` — catches `SignalException` and `SystemExit`; use `rescue StandardError` or more specific types
- `method_missing` implemented without `respond_to_missing?`
- Deeply nested blocks (>3 levels) — extract to methods
- String interpolation used where a symbol would suffice (hash keys, etc.)
---
## Security
- Flag `eval` / `instance_eval` with user-controlled strings — remote code execution
- Flag `system()` / `exec()` / backtick calls with user-controlled input — shell injection
- Flag `YAML.load` on untrusted data — use `YAML.safe_load`
- Flag `Marshal.load` on untrusted data — arbitrary code execution
- Flag raw SQL string interpolation in ActiveRecord — use parameterized queries (`where("name = ?", name)`)
- Flag `params` passed directly to `redirect_to` without validation — open redirect
- Flag `render inline:` with user data — XSS via ERB
- Flag missing `strong_parameters` in Rails controllers — mass assignment vulnerability
---
## Async / Concurrency
- Flag shared mutable state accessed from multiple threads without a `Mutex`
- Flag `Thread.new` without storing the thread reference — exceptions are silently swallowed
- Flag `sleep` used as a synchronization mechanism in threaded code
- Flag `@@class_variables` mutated in multi-threaded contexts — not thread-safe
- Flag Sidekiq / ActiveJob workers that are not idempotent — jobs can be retried
---
## Resource Management
- Flag `File.open` without a block form — the block form guarantees `close`
- Flag database connections or HTTP clients not released in `ensure` blocks
- Flag `ActiveRecord` queries inside loops — N+1 pattern; use `includes` / `preload` / `eager_load`
- Flag `ObjectSpace` usage in production — memory and performance impact
---
## Exception Handling
- Flag `rescue Exception` — use `rescue StandardError` or a specific exception class
- Flag empty `rescue` blocks — swallowed errors
- Flag `rescue` used for control flow (e.g. rescuing `ActiveRecord::RecordNotFound` instead of using `find_by`)
- Flag re-raising with `raise e` instead of bare `raise` — loses the original backtrace
- Flag `ensure` blocks that can raise — masks the original exception
---
## Performance
- Flag N+1 ActiveRecord queries — use `includes`, `preload`, or `eager_load`
- Flag `Array#each` with string concatenation — use `map` + `join`
- Flag `select` + `map` that could be a single `filter_map`
- Flag `.count` on an ActiveRecord relation inside a view or loop — triggers a query each time
- Flag `require` inside a method body — constant overhead on every call
- Flag `Hash#merge` in a loop — use `merge!` or `each_with_object`
---
## Idioms and Best Practices
### Ruby Style
- Prefer `map` / `select` / `reject` / `reduce` over manual `each` + accumulator
- Prefer `&method(:name)` over `{ |x| some_method(x) }` for method reference blocks
- Prefer `freeze` on string constants to avoid repeated object allocation
- Use `attr_reader` / `attr_writer` / `attr_accessor` instead of manual getter/setter methods
- Prefer `Symbol#to_proc` (`&:method_name`) for simple single-method blocks
### Rails-Specific
- Keep controllers thin — logic belongs in service objects, models, or concerns
- Use `before_action` for authentication/authorization checks — never inline
- Prefer `find_by` over `where(...).first` — more intent-revealing
- Flag `after_commit` callbacks with side effects that should be in a service object
- Prefer `respond_to` blocks over separate controller actions for format variants
### Modern Ruby (3.x)
- Prefer pattern matching (`case/in`) for complex data destructuring
- Use numbered block parameters (`_1`, `_2`) only for very short, obvious blocks
- Prefer `Data.define` for simple immutable value objects (Ruby 3.2+)
FILE:languages/rust.md
---
language: rust
extensions: [".rs"]
---
# Rust — Language-Specific Review Notes
Load this file alongside `rules/universal.md`. Universal rules are not repeated here — only Rust-specific rules and idioms.
---
## PR Analyzer — Rust Risk Signals
- `unsafe { }` blocks — require explicit justification and sign-off
- `#[allow(...)]` attributes suppressing lints — verify they are justified
- `.unwrap()` / `.expect("")` on `Option` or `Result` outside of tests or prototypes
- Hardcoded credentials or tokens in source
- `TODO` / `FIXME` comments near `unsafe` or ownership code
---
## Code Quality — Rust Checks
- `.unwrap()` used broadly in production code — prefer `?`, `if let`, or `match`
- `clone()` called excessively — may indicate ownership design issues
- `Arc<Mutex<T>>` used where a simpler ownership model would work
- `Box<dyn Trait>` used where generics (`impl Trait`) would avoid heap allocation
- `pub` fields on structs that should enforce invariants — use accessor methods
---
## Security
- Flag `unsafe` blocks accessing raw pointers without clear safety invariant documented in a comment
- Flag `std::mem::transmute` — almost always a logic error or undefined behavior; require strong justification
- Flag `from_utf8_unchecked` on user-controlled data — use `from_utf8` with error handling
- Flag `unwrap()` on user-supplied input parsing — panics are a denial-of-service vector in server code
- Flag hardcoded secrets — use environment variables or a secrets crate
---
## Async / Concurrency
- Flag `std::sync::Mutex` used in async code — use `tokio::sync::Mutex` to avoid blocking the async runtime
- Flag `.await` inside a `std::sync::MutexGuard` scope — holds the lock across an await point, blocking other tasks
- Flag `spawn` without storing the `JoinHandle` — panics in the spawned task are silently ignored
- Flag `Arc<Mutex<T>>` cloned excessively — consider message passing via channels instead
- Flag blocking I/O calls (`std::fs`, `std::net`) inside async functions — use async equivalents
---
## Resource Management
- Flag manual `drop` called explicitly where the natural scope boundary suffices
- Flag `Rc<T>` used in multi-threaded code — use `Arc<T>`; the compiler catches this but flag in review for architecture discussion
- Flag `Vec` or `String` with large pre-allocated capacity never trimmed — call `.shrink_to_fit()` if long-lived
- Flag `impl Drop` that can panic — causes `abort` during stack unwinding
---
## Exception Handling
- Flag `.unwrap()` in production code outside of tests — use `?` to propagate or handle explicitly
- Flag `.expect("todo")` or `.expect("")` — messages must explain the invariant that guarantees safety
- Flag `panic!` used for recoverable errors — use `Result<T, E>`
- Flag `unwrap_or_default()` where the default silently masks a real error
- Prefer typed error enums (`thiserror`) over `Box<dyn Error>` for library crates
- Prefer `anyhow` for application-level error context; `thiserror` for library error types
---
## Performance
- Flag `.clone()` on large types in hot paths — review whether a reference or `Cow<T>` would work
- Flag `format!` used only to create a `String` from a literal — use `.to_string()` or `String::from`
- Flag `collect::<Vec<_>>()` followed immediately by `.iter()` — chain iterators instead
- Flag `Box<T>` for small types where stack allocation is fine
- Flag `Mutex` contention on a hot path — consider `RwLock` for read-heavy workloads or sharding
---
## Idioms and Best Practices
### Ownership
- Prefer borrowing (`&T`, `&mut T`) over cloning wherever the lifetime allows
- Use `Cow<'_, str>` for functions that sometimes need to own and sometimes borrow
- Prefer `impl Trait` in function signatures over `Box<dyn Trait>` for static dispatch
### Error Handling
- Use `?` operator to propagate errors — avoid manual `match Err(e) => return Err(e)`
- Define domain error types with `thiserror` in libraries; use `anyhow` in binaries
- Never use `.unwrap()` in library code — it panics the caller's thread
### Modern Rust
- Prefer `if let` / `while let` for single-variant matches over full `match`
- Prefer `?` over `unwrap` everywhere errors are recoverable
- Use `#[derive(Debug, Clone, PartialEq)]` consistently on data types
- Prefer `iter()` chains over manual loops — they compose and optimize well
- Use `clippy` and treat its lints as required — flag any `#[allow(clippy::...)]` in review
FILE:languages/swift.md
---
language: swift
extensions: [".swift"]
---
# Swift — Language-Specific Review Notes
Load this file alongside `rules/universal.md`. Universal rules are not repeated here — only Swift-specific rules and idioms.
---
## PR Analyzer — Swift Risk Signals
- `print()` statements left in production code
- Force unwrap (`!`) on optionals outside of tests or justified init
- Force cast (`as!`) without a safe fallback
- Hardcoded credentials or API keys in source
---
## Code Quality — Swift Checks
- Force unwrap (`!`) used broadly — prefer `guard let` or `if let`
- `try!` used outside of guaranteed-safe contexts
- Retain cycles in closures — missing `[weak self]` or `[unowned self]`
- `@objc` / `dynamic` used without an Objective-C interop reason
---
## Security
- Flag credentials stored in `UserDefaults` — require Keychain
- Flag `URLSession` requests over plain HTTP in production
- Flag `WKWebView` loading arbitrary user-supplied URLs without validation
---
## Async / Concurrency
- Flag `DispatchQueue.main.sync` called from the main thread — deadlock
- Flag `@escaping` closures capturing `self` strongly in reference cycles — use `[weak self]`
- Flag mixing `async/await` and `DispatchQueue` for the same operation without clear reasoning
- Flag `Task { }` (unstructured) where a structured `async let` or `TaskGroup` would maintain structure
- Flag data races — shared mutable state accessed from multiple tasks without an actor
---
## Resource Management
- Flag `URLSessionDataTask` started with no cancellation handle stored — cannot be cancelled if the view disappears
- Flag `NotificationCenter` observers added without a corresponding `removeObserver` — memory leak
- Flag `CLLocationManager` / `AVCaptureSession` not stopped when the owning view controller is dismissed
---
## Exception Handling
- Flag `try!` outside of guaranteed-safe contexts (test fixtures, constants) — crashes on failure
- Flag `try?` discarding errors where the failure mode matters to the caller
- Flag error types conforming to `Error` with no associated values or message — makes debugging hard
- Flag throwing functions calling `fatalError()` as a fallback — choose one error strategy
---
## Performance
- Flag `UIImage(named:)` called repeatedly for the same asset without caching
- Flag synchronous network calls on the main thread
- Flag `Array` used for frequent membership tests — prefer `Set`
- Flag `String` interpolation inside tight loops where a pre-built string would avoid allocations
---
## Idioms and Best Practices
### Optionals
- Prefer `guard let` for early exit; `if let` for local scope
- Prefer optional chaining (`?.`) over force unwrap
- Flag implicitly unwrapped optionals (`var x: String!`) outside of `@IBOutlet`
### Memory Management
- Flag closures capturing `self` strongly in reference cycles — use `[weak self]`
- Prefer `struct` over `class` for value semantics unless identity or inheritance is needed
- Use `unowned` only when the lifetime is guaranteed — otherwise `weak`
### Concurrency (Swift 5.5+)
- Prefer `async/await` over completion handlers in new code
- Flag `DispatchQueue.main.async` where `@MainActor` or `await MainActor.run` is more appropriate
FILE:languages/typescript.md
---
language: typescript
extensions: [".ts", ".tsx", ".js", ".jsx", ".mjs"]
---
# TypeScript / JavaScript — Language-Specific Review Notes
Load this file alongside `rules/universal.md`. Universal rules are not repeated here — only TypeScript/JavaScript-specific rules and idioms.
---
## PR Analyzer — TypeScript / JavaScript Risk Signals
- `console.log` / `debugger` statements left in production code
- `// eslint-disable` comments — verify they are justified
- `any` type annotations — require explicit justification
- `@ts-ignore` / `@ts-expect-error` — verify they are justified
- `eval()` with any dynamic or user-controlled input
- Hardcoded API keys or tokens in source
---
## Code Quality — TypeScript / JavaScript Checks
- `any` used broadly instead of proper typing
- Non-null assertion (`!`) used without justification
- `var` declarations — prefer `const` / `let`
- Missing `await` on async function calls
- Floating promises (no `.catch()` and no `await`)
- `==` used instead of `===`
---
## Security
- Flag `innerHTML`, `outerHTML`, `document.write()` with user-controlled data — use `textContent` or a sanitizer
- Flag `dangerouslySetInnerHTML` in React without a sanitizer
- Flag `eval()` / `new Function()` with dynamic input
- Flag JWT decoded without signature verification
- Flag missing `httpOnly` / `secure` flags on cookies
---
## Async / Promises
- Flag floating promises — async calls not `await`-ed and without `.catch()`
- Flag `Promise.all()` where `Promise.allSettled()` is safer (one failure should not cancel siblings)
- Flag `async` functions inside `forEach` — `forEach` does not await; use `for...of` or `Promise.all()`
- Flag unhandled promise rejection (no global `unhandledRejection` handler in Node.js services)
---
## Resource Management
- Flag `fs.createReadStream` / `fs.createWriteStream` with no `close` or `destroy` on error
- Flag `EventEmitter` listeners added in a loop without removal — memory leak
- Flag `setInterval` / `setTimeout` handles not cleared when the owning component unmounts or exits
- Flag database clients / pools not released after use in Node.js
---
## Exception Handling
- Flag `catch (e) {}` (empty catch) — swallowed error
- Flag `catch (e)` where `e` is used as `any` without narrowing — type the error properly
- Flag `Promise` rejection not handled — `.catch()` or `try/await/catch` required
- Flag re-throwing a new `Error` without wrapping the original — loses stack context
- Use `Error` subclasses for domain errors rather than plain strings or object literals
---
## Performance
- Flag `Array.prototype.find` / `filter` / `map` chained multiple times over the same array — combine into one pass
- Flag DOM queries (`document.querySelector`) inside loops — cache the result
- Flag `JSON.parse` / `JSON.stringify` in a hot path on large objects — consider streaming or partial parsing
- Flag `async` functions called sequentially in a loop where `Promise.all()` would parallelize them
---
## Idioms and Best Practices
### Type Safety (TypeScript)
- Prefer `unknown` over `any` for truly unknown values — forces a type guard before use
- Prefer type narrowing (`typeof`, `instanceof`, discriminated unions) over casting
- Enable `strict` mode in `tsconfig.json`
- Prefer `interface` for object shapes that may be extended; `type` for unions and aliases
### Modern JavaScript / TypeScript
- Prefer `const` by default; `let` only when reassignment is needed
- Prefer optional chaining (`?.`) and nullish coalescing (`??`) over manual null guards
- Prefer `structuredClone()` over manual deep-copy patterns
- Prefer named exports over default exports for better refactoring support
### Null / Undefined Safety
- Distinguish between `null` (intentional absence) and `undefined` (not set) — be consistent
- Flag `== null` checks that accidentally include `undefined` when only one is intended
FILE:README.md
# code-reviewer
Code review automation for TypeScript, JavaScript, Python, Go, Swift, Kotlin, C#, .NET, Java, C, C++, Rust, Ruby, PHP, and Dart/Flutter. Analyzes PRs for complexity and risk, checks code quality for SOLID violations and code smells, and generates review reports.
The full skill spec is [`SKILL.md`](./SKILL.md). This README is a quick reference for the 3 bundled scripts.
---
## How to use
### Quick install check
```bash
python scripts/pr_analyzer.py --help
python scripts/code_quality_checker.py --help
python scripts/review_report_generator.py --help
```
All three scripts are stdlib-only — no `pip install` required.
### Example 1 — review a pull request
```bash
# From inside the repo you want to analyze:
python /path/to/skills/code-reviewer/scripts/pr_analyzer.py . --base main --head HEAD
```
Outputs: complexity score (1-10), risk categorization (critical / high / medium / low), prioritized review order, commit-message validation.
### Example 2 — score a directory's code quality
```bash
python scripts/code_quality_checker.py /path/to/code
# Filter by language
python scripts/code_quality_checker.py /path/to/code --language csharp
# Machine-readable
python scripts/code_quality_checker.py /path/to/code --json
```
Outputs: quality score (0-100), letter grade, detected code smells, SOLID violations.
### Example 3 — combine into a review report
```bash
python scripts/review_report_generator.py /path/to/repo --format markdown --output review.md
```
Outputs: review verdict (approve / request changes / block), score, prioritized action items.
---
## Examples bundled with the skill
| File | Purpose |
|------|---------|
| [`assets/sample_csharp_smells.cs`](./assets/sample_csharp_smells.cs) | C# file with every C#-specific pattern this skill detects, labelled inline |
| [`assets/sample_csharp_clean.cs`](./assets/sample_csharp_clean.cs) | Same code refactored per `rules/universal.md` + `languages/csharp.md` |
| [`assets/sample_java_smells.java`](./assets/sample_java_smells.java) | Java file with every Java-specific pattern this skill detects, labelled inline |
| [`assets/sample_java_clean.java`](./assets/sample_java_clean.java) | Same code refactored per `rules/universal.md` + `languages/java.md` |
| [`assets/sample_c_smells.c`](./assets/sample_c_smells.c) | C file with every C-specific pattern this skill detects, labelled inline |
| [`assets/sample_c_clean.c`](./assets/sample_c_clean.c) | Same code refactored per `rules/universal.md` + `languages/c.md` |
| [`expected_outputs/*.json`](./expected_outputs/) | Expected `code_quality_checker.py --json` output for each fixture |
Use them as a regression-detection harness:
```bash
python scripts/code_quality_checker.py assets/sample_java_smells.java --json > /tmp/check.json
diff /tmp/check.json expected_outputs/sample_java_smells_quality.json
# silence means the detector still behaves as documented
```
---
## What it detects
See [`SKILL.md`](./SKILL.md) for the full pattern list, severity tiers, and references. Quick summary:
- **PR Analyzer** (`scripts/pr_analyzer.py`): hardcoded secrets / connection strings, SQL injection, debug statements (`console.*` / `System.out` / `printStackTrace`), analyzer suppressions (ESLint / Roslyn / `@SuppressWarnings`), `any` / `dynamic` overuse, TODO/FIXME, `unsafe` blocks, null-forgiving `!`, `async void`, blocking on `Task`.
- **Code Quality Checker** (`scripts/code_quality_checker.py`): long methods, large files, god classes, deep nesting, too many parameters, high cyclomatic complexity, swallowed exceptions, missing `await`, undisposed `IDisposable`, `new HttpClient()` in method body, unused `using` directives. Language-specific smell packs for C# (`async void`, blocking on `Task`), Java (empty catch, `printStackTrace`, swallowed `InterruptedException`, unclosed resources, per-call `ObjectMapper` / `Gson`), and C (banned functions `gets`/`strcpy`/`strcat`/`sprintf`/`vsprintf`, format-string vulnerability `printf(var)`, unbounded `scanf("%s")`, malloc-without-NULL-check, free-without-zeroing, `system()` with non-literal argument).
- **Review Report Generator** (`scripts/review_report_generator.py`): combines the above into a single markdown or JSON verdict.
---
## Review rules
Rules are split so every review loads exactly two files — the cross-language
baseline plus one language guide (see the dispatch table in [`SKILL.md`](./SKILL.md)):
- [`rules/universal.md`](./rules/universal.md) — cross-language rules: security, async/concurrency, resource management, exception handling, performance
- [`languages/`](./languages/) — one self-contained guide per language (`python`, `typescript`, `go`, `swift`, `kotlin`, `csharp`, `java`, `c`, `cpp`, `rust`, `ruby`, `php`, `dart`), each with Security / Async / Resource Management / Exception Handling / Performance / Idioms sections
FILE:rules/universal.md
# Universal Rules — All Languages
These rules apply regardless of language. Load this file for every review, alongside the relevant `languages/*.md` file.
---
## Security
- Flag any string interpolation or concatenation used to build SQL, shell, or LDAP queries — require parameterized queries or a safe API
- Flag hardcoded credentials, API keys, tokens, or secrets anywhere in source — require environment variables or a secrets manager
- Flag user-controlled input passed to file system, process execution, or URL redirect APIs without validation
- Flag overly broad CORS or CSP policies
---
## Async / Concurrency
- Flag shared mutable state accessed from multiple threads/coroutines/tasks without synchronization
- Flag fire-and-forget async operations with no error handling path
- Flag timeouts missing on any network or I/O call
- Flag unbounded queues or thread pools with no backpressure mechanism
---
## Resource Management
- Flag any resource (file, socket, DB connection, HTTP connection) acquired without a guaranteed release path
- Flag connection pools not returned to the pool on all code paths (including exceptions)
- Flag unbounded collections that grow without eviction — potential memory leak
- Flag resources held open longer than the operation they serve
---
## Exception Handling
- Flag empty catch/except blocks — swallowed exceptions hide bugs silently
- Flag catching the broadest possible exception type (`Exception`, `Throwable`, `error`) where a specific type is appropriate
- Flag exceptions used for normal control flow (signaling "not found", etc.) — use return values or `Optional`
- Flag error context lost when re-throwing — always wrap with the original cause
---
## Performance
- Flag N+1 query patterns — loading a collection then querying for each item individually
- Flag unbounded queries or API calls with no pagination or limit
- Flag synchronous I/O on a thread or event loop that serves concurrent requests
- Flag large objects serialized/deserialized repeatedly when they could be cached
- Flag string concatenation in tight loops — use a builder or join
FILE:scripts/code_quality_checker.py
#!/usr/bin/env python3
"""
Code Quality Checker
Analyzes source code for quality issues, code smells, complexity metrics,
and SOLID principle violations.
Usage:
python code_quality_checker.py /path/to/file.py
python code_quality_checker.py /path/to/directory --recursive
python code_quality_checker.py . --language typescript --json
"""
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Dict, List, Optional
# Language-specific file extensions.
# `c` is declared before `cpp` so plain `.h` resolves to C, matching the
# dispatch table in SKILL.md. C++ headers use `.hpp` / `.hh` / `.hxx`.
LANGUAGE_EXTENSIONS = {
"python": [".py"],
"typescript": [".ts", ".tsx"],
"javascript": [".js", ".jsx", ".mjs"],
"go": [".go"],
"swift": [".swift"],
"kotlin": [".kt", ".kts"],
"csharp": [".cs", ".csx", ".razor", ".cshtml"],
"java": [".java"],
"c": [".c", ".h"],
"cpp": [".cpp", ".cc", ".cxx", ".hpp", ".hh", ".hxx"],
"rust": [".rs"],
"ruby": [".rb", ".rake", ".gemspec", ".ru"],
"php": [".php", ".phtml"],
"dart": [".dart"],
}
# Code smell thresholds
THRESHOLDS = {
"long_function_lines": 50,
"too_many_parameters": 5,
"high_complexity": 10,
"god_class_methods": 20,
"max_imports": 15
}
def get_file_extension(filepath: Path) -> str:
"""Get file extension."""
return filepath.suffix.lower()
def detect_language(filepath: Path) -> Optional[str]:
"""Detect programming language from file extension."""
ext = get_file_extension(filepath)
for lang, extensions in LANGUAGE_EXTENSIONS.items():
if ext in extensions:
return lang
return None
def read_file_content(filepath: Path) -> str:
"""Read file content safely."""
try:
with open(filepath, "r", encoding="utf-8", errors="ignore") as f:
return f.read()
except Exception:
return ""
def calculate_cyclomatic_complexity(content: str) -> int:
"""
Estimate cyclomatic complexity based on control flow keywords.
"""
complexity = 1 # Base complexity
# Control flow patterns that increase complexity
patterns = [
r"\bif\b",
r"\belif\b",
r"\belse\b",
r"\bfor\b",
r"\bwhile\b",
r"\bcase\b",
r"\bcatch\b",
r"\bexcept\b",
r"\band\b",
r"\bor\b",
r"\|\|",
r"&&"
]
for pattern in patterns:
matches = re.findall(pattern, content, re.IGNORECASE)
complexity += len(matches)
return complexity
def count_lines(content: str) -> Dict[str, int]:
"""Count different types of lines in code."""
lines = content.split("\n")
total = len(lines)
blank = sum(1 for line in lines if not line.strip())
comment = 0
for line in lines:
stripped = line.strip()
if stripped.startswith("#") or stripped.startswith("//"):
comment += 1
elif stripped.startswith("/*") or stripped.startswith("'''") or stripped.startswith('"""'):
comment += 1
code = total - blank - comment
return {
"total": total,
"code": code,
"blank": blank,
"comment": comment
}
def find_functions(content: str, language: str) -> List[Dict]:
"""Find function definitions and their metrics."""
functions = []
# Language-specific function patterns
patterns = {
"python": r"def\s+(\w+)\s*\(([^)]*)\)",
"typescript": r"(?:function\s+(\w+)|(?:const|let|var)\s+(\w+)\s*=\s*(?:async\s+)?\([^)]*\)\s*=>)",
"javascript": r"(?:function\s+(\w+)|(?:const|let|var)\s+(\w+)\s*=\s*(?:async\s+)?\([^)]*\)\s*=>)",
"go": r"func\s+(?:\([^)]+\)\s+)?(\w+)\s*\(([^)]*)\)",
"swift": r"func\s+(\w+)\s*\(([^)]*)\)",
"kotlin": r"fun\s+(\w+)\s*\(([^)]*)\)",
# C#: require at least one method modifier (public/private/etc. or static/async/...)
# to distinguish declarations from invocations.
"csharp": (
r"(?:(?:public|private|protected|internal|static|async|virtual|"
r"override|sealed|abstract|partial|new|readonly|extern)\s+)+"
r"(?:[\w<>?,\s\[\]\.]+?\s+)?(\w+)\s*\(([^)]*)\)"
),
# Java: require at least one method modifier to distinguish
# declarations from invocations (mirrors the C# approach).
"java": (
r"(?:(?:public|private|protected|static|final|abstract|"
r"synchronized|native|default|strictfp)\s+)+"
r"(?:[\w<>?,\s\[\]\.]+?\s+)?(\w+)\s*\(([^)]*)\)"
),
# C: require an opening brace after the parens so prototypes and
# call sites don't get matched. Return type / qualifiers come first.
# Skip C control-flow keywords that look like function calls.
"c": (
r"^(?:static\s+|inline\s+|extern\s+|const\s+|unsigned\s+|"
r"signed\s+|volatile\s+|register\s+)*"
r"(?:[\w\*]+\s+\**)+"
r"(?!(?:if|while|for|switch|return|sizeof)\b)"
r"(\w+)\s*\(([^)]*)\)\s*\{"
),
# C++: like C but also catches `ClassName::method(...)` definitions
# and template return types like `std::vector<int>`.
"cpp": (
r"^(?:static\s+|inline\s+|extern\s+|const\s+|virtual\s+|"
r"explicit\s+|constexpr\s+|noexcept\s+)*"
r"(?:[\w:\*<>&,\s]+\s+\**)+"
r"(?!(?:if|while|for|switch|return|sizeof)\b)"
r"(\w+)(?:::\w+)?\s*\(([^)]*)\)\s*(?:const\s*)?"
r"(?:noexcept\s*)?(?:override\s*)?(?:final\s*)?\{"
),
# Rust: `fn` keyword is always present and unambiguous.
"rust": (
r"(?:pub(?:\([^)]+\))?\s+)?(?:async\s+)?(?:unsafe\s+)?"
r"(?:extern\s+\"[^\"]+\"\s+)?fn\s+(\w+)\s*"
r"(?:<[^>]+>)?\s*\(([^)]*)\)"
),
# Ruby: `def` keyword; params may be parenthesised or bare.
"ruby": (
r"def\s+(?:self\.)?(\w+[?!=]?)(?:\s*\(([^)]*)\)|\s*$|\s+\w)"
),
# PHP: `function` keyword is always present.
"php": (
r"(?:(?:public|private|protected|static|abstract|final)\s+)*"
r"function\s+(\w+)\s*\(([^)]*)\)"
),
# Dart: typed return followed by name and parens. Constructors
# (where name matches enclosing class) are not specially handled.
"dart": (
r"^\s*(?:static\s+|external\s+)*"
r"(?:Future<[^>]*>|Stream<[^>]*>|void|[\w<>?,\s]+?)\s+"
r"(\w+)\s*\(([^)]*)\)\s*(?:async\*?\s*|sync\*?\s*)?\{"
),
}
pattern = patterns.get(language, patterns["python"])
matches = re.finditer(pattern, content, re.MULTILINE)
for match in matches:
name = next((g for g in match.groups() if g), "anonymous")
params_str = match.group(2) if len(match.groups()) > 1 and match.group(2) else ""
# Count parameters
params = [p.strip() for p in params_str.split(",") if p.strip()]
param_count = len(params)
# Estimate function length
start_pos = match.end()
remaining = content[start_pos:]
next_func = re.search(pattern, remaining)
if next_func:
func_body = remaining[:next_func.start()]
else:
func_body = remaining[:min(2000, len(remaining))]
line_count = len(func_body.split("\n"))
complexity = calculate_cyclomatic_complexity(func_body)
functions.append({
"name": name,
"parameters": param_count,
"lines": line_count,
"complexity": complexity
})
return functions
def find_classes(content: str, language: str) -> List[Dict]:
"""Find class definitions and their metrics."""
classes = []
patterns = {
"python": r"class\s+(\w+)",
"typescript": r"class\s+(\w+)",
"javascript": r"class\s+(\w+)",
"go": r"type\s+(\w+)\s+struct",
"swift": r"class\s+(\w+)",
"kotlin": r"class\s+(\w+)",
"csharp": r"(?:class|struct|record|interface)\s+(\w+)",
"java": r"(?:class|interface|enum|record)\s+(\w+)",
# C has no classes; `struct` and `typedef struct` are the closest.
"c": r"(?:typedef\s+)?struct\s+(\w+)",
"cpp": r"(?:class|struct)\s+(\w+)",
# Rust uses `struct`, `enum`, `trait`, `union` for type definitions.
# `impl` blocks attach methods but are not type defs themselves.
"rust": r"(?:pub(?:\([^)]+\))?\s+)?(?:struct|enum|trait|union)\s+(\w+)",
"ruby": r"(?:class|module)\s+(\w+)",
"php": (
r"(?:abstract\s+|final\s+)?"
r"(?:class|interface|trait|enum)\s+(\w+)"
),
# Dart 3 class modifiers: final / interface / base / sealed / mixin.
"dart": (
r"(?:abstract\s+|sealed\s+|final\s+|base\s+|interface\s+)?"
r"(?:class|mixin|enum|extension)\s+(\w+)"
),
}
pattern = patterns.get(language, patterns["python"])
matches = re.finditer(pattern, content)
for match in matches:
name = match.group(1)
start_pos = match.end()
remaining = content[start_pos:]
next_class = re.search(pattern, remaining)
if next_class:
class_body = remaining[:next_class.start()]
else:
class_body = remaining
# Count methods
method_patterns = {
"python": r"def\s+\w+\s*\(",
"typescript": r"(?:public|private|protected)?\s*\w+\s*\([^)]*\)\s*[:{]",
"javascript": r"\w+\s*\([^)]*\)\s*\{",
"go": r"func\s+\(",
"swift": r"func\s+\w+",
"kotlin": r"fun\s+\w+",
"csharp": (
r"(?:(?:public|private|protected|internal|static|async|virtual|"
r"override|sealed|abstract|partial)\s+)+"
r"(?:[\w<>?,\s\[\]\.]+?\s+)?\w+\s*\("
),
"java": (
r"(?:(?:public|private|protected|static|final|abstract|"
r"synchronized|native|default|strictfp)\s+)+"
r"(?:[\w<>?,\s\[\]\.]+?\s+)?\w+\s*\("
),
# C has no classes; struct members are typically function pointers
# rather than methods. Use the function definition pattern.
"c": (
r"^(?:static\s+|inline\s+)*(?:[\w\*]+\s+\**)+"
r"(?!(?:if|while|for|switch|return|sizeof)\b)"
r"\w+\s*\([^)]*\)\s*\{"
),
"cpp": (
r"^(?:static\s+|inline\s+|virtual\s+|explicit\s+|"
r"constexpr\s+)*(?:[\w:\*<>&,\s]+\s+\**)+"
r"(?!(?:if|while|for|switch|return|sizeof)\b)"
r"\w+(?:::\w+)?\s*\([^)]*\)"
),
"rust": (
r"(?:pub(?:\([^)]+\))?\s+)?(?:async\s+)?(?:unsafe\s+)?"
r"fn\s+\w+"
),
"ruby": r"def\s+(?:self\.)?\w+[?!=]?",
"php": (
r"(?:(?:public|private|protected|static|abstract|final)\s+)*"
r"function\s+\w+\s*\("
),
"dart": (
r"^\s*(?:static\s+|external\s+)*"
r"(?:Future<[^>]*>|Stream<[^>]*>|void|[\w<>?,\s]+?)\s+"
r"\w+\s*\([^)]*\)\s*(?:async\*?\s*|sync\*?\s*)?\{"
),
}
method_pattern = method_patterns.get(language, method_patterns["python"])
methods = len(re.findall(method_pattern, class_body))
classes.append({
"name": name,
"methods": methods,
"lines": len(class_body.split("\n"))
})
return classes
def check_code_smells(content: str, functions: List[Dict], classes: List[Dict]) -> List[Dict]:
"""Check for code smells in the content."""
smells = []
# Long functions
for func in functions:
if func["lines"] > THRESHOLDS["long_function_lines"]:
smells.append({
"type": "long_function",
"severity": "medium",
"message": f"Function '{func['name']}' has {func['lines']} lines (max: {THRESHOLDS['long_function_lines']})",
"location": func["name"]
})
# Too many parameters
for func in functions:
if func["parameters"] > THRESHOLDS["too_many_parameters"]:
smells.append({
"type": "too_many_parameters",
"severity": "low",
"message": f"Function '{func['name']}' has {func['parameters']} parameters (max: {THRESHOLDS['too_many_parameters']})",
"location": func["name"]
})
# High complexity
for func in functions:
if func["complexity"] > THRESHOLDS["high_complexity"]:
severity = "high" if func["complexity"] > 20 else "medium"
smells.append({
"type": "high_complexity",
"severity": severity,
"message": f"Function '{func['name']}' has complexity {func['complexity']} (max: {THRESHOLDS['high_complexity']})",
"location": func["name"]
})
# God classes
for cls in classes:
if cls["methods"] > THRESHOLDS["god_class_methods"]:
smells.append({
"type": "god_class",
"severity": "high",
"message": f"Class '{cls['name']}' has {cls['methods']} methods (max: {THRESHOLDS['god_class_methods']})",
"location": cls["name"]
})
# Magic numbers
magic_pattern = r"\b(?<![.\"\'])\d{3,}\b(?!\.\d)"
for i, line in enumerate(content.split("\n"), 1):
if line.strip().startswith(("#", "//", "import", "from")):
continue
matches = re.findall(magic_pattern, line)
for match in matches[:1]: # One per line
smells.append({
"type": "magic_number",
"severity": "low",
"message": f"Magic number {match} should be a named constant",
"location": f"line {i}"
})
# Commented code patterns
commented_code_pattern = r"^\s*[#//]+\s*(if|for|while|def|function|class|const|let|var)\s"
for i, line in enumerate(content.split("\n"), 1):
if re.match(commented_code_pattern, line, re.IGNORECASE):
smells.append({
"type": "commented_code",
"severity": "low",
"message": "Commented-out code should be removed",
"location": f"line {i}"
})
return smells
def _strip_csharp_comments(content: str) -> str:
"""Remove // line comments and /* */ block comments so regex detectors
don't match keywords inside prose."""
no_block = re.sub(r"/\*.*?\*/", "", content, flags=re.DOTALL)
no_line = re.sub(r"//[^\n]*", "", no_block)
return no_line
def check_csharp_specific_smells(content: str) -> List[Dict]:
"""C# / .NET-specific code smells documented in SKILL.md."""
smells: List[Dict] = []
content = _strip_csharp_comments(content)
# async void (event handler exception only — caller must justify)
for match in re.finditer(r"\basync\s+void\s+(\w+)\s*\(", content):
smells.append({
"type": "csharp_async_void",
"severity": "high",
"message": (
f"'async void {match.group(1)}' — only safe for event handlers; "
"prefer 'async Task'"
),
"location": match.group(1),
})
# Blocking on async: .Result, .Wait(), .GetAwaiter().GetResult()
for match in re.finditer(
r"\.(?:Result\b|Wait\(\)|GetAwaiter\(\)\.GetResult\(\))", content
):
smells.append({
"type": "csharp_blocking_async",
"severity": "high",
"message": (
"Blocking call on async operation ('.Result' / '.Wait()' / "
"'.GetAwaiter().GetResult()') — can deadlock in ASP.NET contexts"
),
"location": f"offset {match.start()}",
})
# Bare catch / catch (Exception) that swallows
swallow_pattern = re.compile(
r"catch\s*(?:\(\s*(?:System\.)?Exception(?:\s+\w+)?\s*\))?\s*\{\s*\}"
)
for match in swallow_pattern.finditer(content):
smells.append({
"type": "csharp_swallowed_exception",
"severity": "high",
"message": "Empty catch block swallows exceptions silently",
"location": f"offset {match.start()}",
})
# IDisposable instantiated but not in `using` — heuristic: `new SomethingClient(`
# / `new SomethingStream(` / `new SqlConnection(` outside a `using` line.
disposable_hint = re.compile(
r"^(?!\s*using\b)\s*(?:var|[\w<>]+)\s+\w+\s*=\s*new\s+"
r"(\w*(?:Stream|Connection|Reader|Writer|Client|Context|Command))\s*\(",
re.MULTILINE,
)
for match in disposable_hint.finditer(content):
smells.append({
"type": "csharp_undisposed_idisposable",
"severity": "medium",
"message": (
f"'{match.group(1)}' looks like IDisposable but is not wrapped in "
"'using' / 'using var'"
),
"location": f"offset {match.start()}",
})
# HttpClient instantiated with `new` inside a method body (socket exhaustion)
httpclient_inline = re.compile(r"new\s+HttpClient\s*\(\s*\)")
for match in httpclient_inline.finditer(content):
smells.append({
"type": "csharp_new_httpclient",
"severity": "medium",
"message": (
"'new HttpClient()' — prefer IHttpClientFactory or a long-lived "
"static instance to avoid socket exhaustion"
),
"location": f"offset {match.start()}",
})
# Missing await: `Task.Run(` / async method call assigned but never awaited.
# Heuristic: a statement ending in `Async()` or `Async(...)` followed by `;`
# with no `await` keyword on the same line.
for line_no, line in enumerate(content.split("\n"), 1):
stripped = line.strip()
if not stripped or stripped.startswith(("//", "/*", "*")):
continue
if re.search(r"\b\w+Async\s*\([^)]*\)\s*;\s*$", stripped) and "await " not in stripped:
# Skip `return ...Async();` (forwarding the Task is legitimate)
if stripped.startswith("return "):
continue
smells.append({
"type": "csharp_missing_await",
"severity": "medium",
"message": "Async method called without 'await' — Task is discarded",
"location": f"line {line_no}",
})
# Unnecessary `using` directives — heuristic: `using` directive whose
# namespace tail isn't referenced anywhere else in the file.
using_directives = re.findall(
r"^using\s+(?:static\s+)?([A-Z]\w*(?:\.\w+)*)\s*;", content, re.MULTILINE
)
body = re.sub(r"^using\s+[^;]+;\s*$", "", content, flags=re.MULTILINE)
for ns in using_directives:
tail = ns.split(".")[-1]
if not re.search(rf"\b{re.escape(tail)}\b", body):
smells.append({
"type": "csharp_unused_using",
"severity": "low",
"message": f"'using {ns};' appears unused",
"location": ns,
})
return smells
def check_java_specific_smells(content: str) -> List[Dict]:
"""Java-specific code smells documented in languages/java.md."""
smells: List[Dict] = []
# Java comment syntax matches C#, so the same stripper applies.
content = _strip_csharp_comments(content)
# Empty catch block — swallows the exception silently.
for match in re.finditer(r"catch\s*\([^)]*\)\s*\{\s*\}", content):
smells.append({
"type": "java_empty_catch",
"severity": "high",
"message": "Empty catch block swallows exceptions silently",
"location": f"offset {match.start()}",
})
# printStackTrace() as error handling — use a logger instead.
for match in re.finditer(r"\.printStackTrace\s*\(\s*\)", content):
smells.append({
"type": "java_print_stack_trace",
"severity": "medium",
"message": (
"'printStackTrace()' is not real error handling — log via a "
"proper logger or rethrow with context"
),
"location": f"offset {match.start()}",
})
# InterruptedException caught without restoring the interrupt flag.
for match in re.finditer(
r"catch\s*\(\s*InterruptedException\s+(\w+)\s*\)\s*\{(.*?)\}",
content,
re.DOTALL,
):
if "interrupt()" not in match.group(2):
smells.append({
"type": "java_swallowed_interrupt",
"severity": "high",
"message": (
"InterruptedException caught without "
"'Thread.currentThread().interrupt()' — breaks cooperative "
"cancellation"
),
"location": f"offset {match.start()}",
})
# Closeable resource instantiated outside try-with-resources (leak heuristic).
resource_hint = re.compile(
r"^(?!\s*try\b)\s*(?:final\s+)?[\w<>\[\]]+\s+\w+\s*=\s*new\s+"
r"(\w*(?:InputStream|OutputStream|Reader|Writer|Stream|Connection))\s*\(",
re.MULTILINE,
)
for match in resource_hint.finditer(content):
smells.append({
"type": "java_unclosed_resource",
"severity": "medium",
"message": (
f"'{match.group(1)}' looks like an AutoCloseable but is not in a "
"try-with-resources statement"
),
"location": f"offset {match.start()}",
})
# Heavy object built per use instead of shared as a singleton.
# A `static` field assignment is the recommended singleton form — skip it.
heavy_object = re.compile(
r"^(?!.*\bstatic\b).*\bnew\s+(ObjectMapper|Gson)\s*\(\s*\)",
re.MULTILINE,
)
for match in heavy_object.finditer(content):
smells.append({
"type": "java_per_use_heavy_object",
"severity": "medium",
"message": (
f"'new {match.group(1)}()' is expensive — share a singleton "
"instance instead of constructing per call"
),
"location": f"offset {match.start()}",
})
return smells
def check_c_specific_smells(content: str) -> List[Dict]:
"""C-specific code smells documented in languages/c.md.
Focuses on memory-safety and command/format-string patterns that the
CERT C Coding Standard and the CWE catalogue rank as the
highest-impact footguns. C uses the same line/block comment syntax as
C# and Java, so the existing comment stripper applies.
"""
smells: List[Dict] = []
content = _strip_csharp_comments(content)
# 1. Banned functions — no bounds check on any of them.
banned = {
"gets": "no bounds check, removed from C11 (CWE-242)",
"strcpy": "no bounds check — prefer strncpy or strlcpy",
"strcat": "no bounds check — prefer strncat or strlcat",
"sprintf": "no bounds check — prefer snprintf",
"vsprintf": "no bounds check — prefer vsnprintf",
}
for fn, reason in banned.items():
for m in re.finditer(rf"\b{fn}\s*\(", content):
smells.append({
"type": f"c_banned_{fn}",
"severity": "high",
"message": f"'{fn}()' is unsafe: {reason}",
"location": f"offset {m.start()}",
})
# 2. Format-string vulnerability — printf/syslog called with a bare
# identifier as the format argument (CWE-134). Skip when the first
# arg is a string literal.
for fn in ("printf", "syslog"):
pattern = rf"\b{fn}\s*\(\s*(?!\")(\w+)\s*[,\)]"
for m in re.finditer(pattern, content):
smells.append({
"type": "c_format_string",
"severity": "high",
"message": (
f"'{fn}({m.group(1)})' uses a non-literal format string "
"— CWE-134 format string vulnerability"
),
"location": f"offset {m.start()}",
})
# 3. Unbounded scanf — `%s` without a width specifier invites overflow.
scanf_call = re.compile(
r"\b(?:scanf|fscanf|sscanf)\s*\(\s*[^)]*?\"([^\"]*)\""
)
for m in scanf_call.finditer(content):
fmt = m.group(1)
if "%s" in fmt and not re.search(r"%\d+s", fmt):
smells.append({
"type": "c_unbounded_scanf",
"severity": "high",
"message": (
"scanf '%s' without a width specifier — unbounded read "
"can overflow the destination buffer"
),
"location": f"offset {m.start()}",
})
# 4. malloc / calloc / realloc result dereferenced without a NULL check
# within 5 lines. CWE-690.
lines = content.split("\n")
malloc_assign = re.compile(
r"^\s*(?:[\w\*]+\s+)?\*?(\w+)\s*=\s*\(?[\w\s\*]*\)?\s*"
r"(?:m|c|re)alloc\s*\("
)
for i, line in enumerate(lines):
m = malloc_assign.match(line)
if not m:
continue
var = m.group(1)
window = "\n".join(lines[i + 1 : i + 6])
null_check = re.compile(
rf"\bif\s*\([^)]*(?:{re.escape(var)}\s*==\s*NULL"
rf"|NULL\s*==\s*{re.escape(var)}"
rf"|!\s*{re.escape(var)}\b"
rf"|{re.escape(var)}\s*!=\s*NULL)"
)
if not null_check.search(window):
smells.append({
"type": "c_malloc_unchecked",
"severity": "medium",
"message": (
f"'{var}' from malloc/calloc/realloc is not NULL-checked "
"within 5 lines — dereferencing NULL is UB (CWE-690)"
),
"location": f"line {i + 1}",
})
# 5. free(p) without setting p to NULL on the next real line.
# CWE-416 use-after-free guardrail.
free_call = re.compile(r"^\s*free\s*\(\s*(\w+)\s*\)\s*;")
for i, line in enumerate(lines):
m = free_call.match(line)
if not m:
continue
var = m.group(1)
for j in range(i + 1, min(i + 3, len(lines))):
nxt = lines[j].strip()
if not nxt:
continue
if re.match(rf"^{re.escape(var)}\s*=\s*NULL\s*;", nxt):
break
smells.append({
"type": "c_free_without_null",
"severity": "low",
"message": (
f"'free({var})' not followed by '{var} = NULL;' — "
"dangling pointer can be reused (CWE-416)"
),
"location": f"line {i + 1}",
})
break
# 6. system() with a non-string-literal argument — command injection.
system_pattern = re.compile(r"\bsystem\s*\(\s*(?!\"|NULL\b)(\w+)\s*\)")
for m in system_pattern.finditer(content):
smells.append({
"type": "c_system_non_literal",
"severity": "high",
"message": (
f"'system({m.group(1)})' with a non-literal argument — "
"command injection (CWE-78); use execve with validated args"
),
"location": f"offset {m.start()}",
})
return smells
def check_solid_violations(content: str) -> List[Dict]:
"""Check for potential SOLID principle violations."""
violations = []
# OCP: Type checking instead of polymorphism
type_checks = len(re.findall(r"isinstance\(|type\(.*\)\s*==|typeof\s+\w+\s*===", content))
if type_checks > 2:
violations.append({
"principle": "OCP",
"name": "Open/Closed Principle",
"severity": "medium",
"message": f"Found {type_checks} type checks - consider using polymorphism"
})
# LSP/ISP: NotImplementedError
not_impl = len(re.findall(r"raise\s+NotImplementedError|not\s+implemented", content, re.IGNORECASE))
if not_impl:
violations.append({
"principle": "LSP/ISP",
"name": "Liskov/Interface Segregation",
"severity": "low",
"message": f"Found {not_impl} unimplemented methods - may indicate oversized interface"
})
# DIP: Too many direct imports
imports = len(re.findall(r"^(?:import|from)\s+", content, re.MULTILINE))
if imports > THRESHOLDS["max_imports"]:
violations.append({
"principle": "DIP",
"name": "Dependency Inversion Principle",
"severity": "low",
"message": f"File has {imports} imports - consider dependency injection"
})
return violations
def calculate_quality_score(
line_metrics: Dict,
functions: List[Dict],
classes: List[Dict],
smells: List[Dict],
violations: List[Dict]
) -> int:
"""Calculate overall quality score (0-100)."""
score = 100
# Deduct for code smells
for smell in smells:
if smell["severity"] == "high":
score -= 10
elif smell["severity"] == "medium":
score -= 5
elif smell["severity"] == "low":
score -= 2
# Deduct for SOLID violations
for violation in violations:
if violation["severity"] == "high":
score -= 8
elif violation["severity"] == "medium":
score -= 4
elif violation["severity"] == "low":
score -= 2
# Bonus for good comment ratio (10-30%)
if line_metrics["total"] > 0:
comment_ratio = line_metrics["comment"] / line_metrics["total"]
if 0.1 <= comment_ratio <= 0.3:
score += 5
# Bonus for reasonable function sizes
if functions:
avg_lines = sum(f["lines"] for f in functions) / len(functions)
if avg_lines < 30:
score += 5
return max(0, min(100, score))
def get_grade(score: int) -> str:
"""Convert score to letter grade."""
if score >= 90:
return "A"
elif score >= 80:
return "B"
elif score >= 70:
return "C"
elif score >= 60:
return "D"
else:
return "F"
def analyze_file(filepath: Path) -> Dict:
"""Analyze a single file for code quality."""
language = detect_language(filepath)
if not language:
return {"error": f"Unsupported file type: {filepath.suffix}"}
content = read_file_content(filepath)
if not content:
return {"error": f"Could not read file: {filepath}"}
line_metrics = count_lines(content)
functions = find_functions(content, language)
classes = find_classes(content, language)
smells = check_code_smells(content, functions, classes)
if language == "csharp":
smells.extend(check_csharp_specific_smells(content))
if language == "java":
smells.extend(check_java_specific_smells(content))
if language == "c":
smells.extend(check_c_specific_smells(content))
violations = check_solid_violations(content)
score = calculate_quality_score(line_metrics, functions, classes, smells, violations)
return {
"file": str(filepath),
"language": language,
"metrics": {
"lines": line_metrics,
"functions": len(functions),
"classes": len(classes),
"avg_complexity": round(sum(f["complexity"] for f in functions) / max(1, len(functions)), 1)
},
"quality_score": score,
"grade": get_grade(score),
"smells": smells,
"solid_violations": violations,
"function_details": functions[:10],
"class_details": classes[:10]
}
def analyze_directory(
dir_path: Path,
recursive: bool = True,
language: Optional[str] = None
) -> Dict:
"""Analyze all files in a directory."""
results = []
extensions = []
if language:
extensions = LANGUAGE_EXTENSIONS.get(language, [])
else:
for exts in LANGUAGE_EXTENSIONS.values():
extensions.extend(exts)
pattern = "**/*" if recursive else "*"
for ext in extensions:
for filepath in dir_path.glob(f"{pattern}{ext}"):
if "node_modules" in str(filepath) or ".git" in str(filepath):
continue
result = analyze_file(filepath)
if "error" not in result:
results.append(result)
if not results:
return {"error": "No supported files found"}
total_score = sum(r["quality_score"] for r in results)
avg_score = total_score / len(results)
total_smells = sum(len(r["smells"]) for r in results)
total_violations = sum(len(r["solid_violations"]) for r in results)
return {
"directory": str(dir_path),
"files_analyzed": len(results),
"average_score": round(avg_score, 1),
"overall_grade": get_grade(int(avg_score)),
"total_code_smells": total_smells,
"total_solid_violations": total_violations,
"files": sorted(results, key=lambda x: x["quality_score"])
}
def print_report(analysis: Dict) -> None:
"""Print human-readable analysis report."""
if "error" in analysis:
print(f"Error: {analysis['error']}")
return
print("=" * 60)
print("CODE QUALITY REPORT")
print("=" * 60)
if "file" in analysis:
print(f"\nFile: {analysis['file']}")
print(f"Language: {analysis['language']}")
print(f"Quality Score: {analysis['quality_score']}/100 ({analysis['grade']})")
metrics = analysis["metrics"]
print(f"\nLines: {metrics['lines']['total']} ({metrics['lines']['code']} code, {metrics['lines']['comment']} comments)")
print(f"Functions: {metrics['functions']}")
print(f"Classes: {metrics['classes']}")
print(f"Avg Complexity: {metrics['avg_complexity']}")
if analysis["smells"]:
print("\n--- CODE SMELLS ---")
for smell in analysis["smells"][:10]:
print(f" [{smell['severity'].upper()}] {smell['message']} ({smell['location']})")
if analysis["solid_violations"]:
print("\n--- SOLID VIOLATIONS ---")
for v in analysis["solid_violations"]:
print(f" [{v['principle']}] {v['message']}")
else:
print(f"\nDirectory: {analysis['directory']}")
print(f"Files Analyzed: {analysis['files_analyzed']}")
print(f"Average Score: {analysis['average_score']}/100 ({analysis['overall_grade']})")
print(f"Total Code Smells: {analysis['total_code_smells']}")
print(f"Total SOLID Violations: {analysis['total_solid_violations']}")
print("\n--- FILES BY QUALITY ---")
for f in analysis["files"][:10]:
print(f" {f['quality_score']:3d}/100 [{f['grade']}] {f['file']}")
print("\n" + "=" * 60)
def main():
parser = argparse.ArgumentParser(
description="Analyze code quality, smells, and SOLID violations"
)
parser.add_argument(
"path",
help="File or directory to analyze"
)
parser.add_argument(
"--recursive", "-r",
action="store_true",
default=True,
help="Recursively analyze directories (default: true)"
)
parser.add_argument(
"--language", "-l",
choices=list(LANGUAGE_EXTENSIONS.keys()),
help="Filter by programming language"
)
parser.add_argument(
"--json",
action="store_true",
help="Output in JSON format"
)
parser.add_argument(
"--output", "-o",
help="Write output to file"
)
args = parser.parse_args()
target = Path(args.path).resolve()
if not target.exists():
print(f"Error: Path does not exist: {target}", file=sys.stderr)
sys.exit(1)
if target.is_file():
analysis = analyze_file(target)
else:
analysis = analyze_directory(target, args.recursive, args.language)
if args.json:
output = json.dumps(analysis, indent=2, default=str)
if args.output:
with open(args.output, "w") as f:
f.write(output)
print(f"Results written to {args.output}")
else:
print(output)
else:
print_report(analysis)
if __name__ == "__main__":
main()
FILE:scripts/pr_analyzer.py
#!/usr/bin/env python3
"""
PR Analyzer
Analyzes pull request changes for review complexity, risk assessment,
and generates review priorities.
Usage:
python pr_analyzer.py /path/to/repo
python pr_analyzer.py . --base main --head feature-branch
python pr_analyzer.py /path/to/repo --json
"""
import argparse
import json
import os
import re
import subprocess
import sys
from pathlib import Path
from typing import Dict, List, Optional, Tuple
# File categories for review prioritization
FILE_CATEGORIES = {
"critical": {
"patterns": [
r"auth", r"security", r"password", r"token", r"secret",
r"payment", r"billing", r"crypto", r"encrypt"
],
"weight": 5,
"description": "Security-sensitive files requiring careful review"
},
"high": {
"patterns": [
r"api", r"database", r"migration", r"schema", r"model",
r"config", r"env", r"middleware"
],
"weight": 4,
"description": "Core infrastructure files"
},
"medium": {
"patterns": [
r"service", r"controller", r"handler", r"util", r"helper"
],
"weight": 3,
"description": "Business logic files"
},
"low": {
"patterns": [
r"test", r"spec", r"mock", r"fixture", r"story",
r"readme", r"docs", r"\.md$"
],
"weight": 1,
"description": "Tests and documentation"
}
}
# Risky patterns to flag
RISK_PATTERNS = [
{
"name": "hardcoded_secrets",
"pattern": r"(password|secret|api_key|token|connection_?string)\s*[=:]\s*['\"][^'\"]+['\"]",
"severity": "critical",
"message": "Potential hardcoded secret or connection string detected"
},
{
"name": "todo_fixme",
"pattern": r"(TODO|FIXME|HACK|XXX):",
"severity": "low",
"message": "TODO/FIXME comment found"
},
{
"name": "console_log",
"pattern": (
r"console\.(log|debug|info|warn|error)\(|\bDebug\.WriteLine\(|"
r"\bSystem\.out\.print(?:ln)?\(|\.printStackTrace\("
),
"severity": "medium",
"message": (
"Debug output statement found "
"(console.* / Debug.WriteLine / System.out / printStackTrace)"
)
},
{
"name": "debugger",
"pattern": r"\bdebugger\b",
"severity": "high",
"message": "Debugger statement found"
},
{
"name": "analyzer_disable",
"pattern": (
r"eslint-disable|#pragma\s+warning\s+disable|\[SuppressMessage|"
r"@SuppressWarnings"
),
"severity": "medium",
"message": (
"Static-analyzer rule disabled "
"(ESLint / Roslyn / SuppressMessage / @SuppressWarnings)"
)
},
{
"name": "loose_type",
"pattern": r":\s*any\b|\bdynamic\s+\w+\s*[=;]",
"severity": "medium",
"message": "Loose type used (TypeScript 'any' or C# 'dynamic')"
},
{
"name": "sql_concatenation",
"pattern": r"(SELECT|INSERT|UPDATE|DELETE).*\+.*['\"]|(?:FromSql|ExecuteSql)\w*\([^)]*\$\"",
"severity": "critical",
"message": "Potential SQL injection (string concatenation or interpolation in query)"
},
{
"name": "csharp_unsafe_block",
"pattern": (
r"\bunsafe\s+(?:\{|public|private|protected|internal|static|sealed|"
r"partial|class|struct|void|int|string|long|short|byte|double|float|"
r"bool|char|ref|out|fixed)\b"
),
"severity": "high",
"message": "C# 'unsafe' code — requires memory-safety review"
},
{
"name": "csharp_null_forgiving",
"pattern": r"(?:\)\s*!\.|\w+!\.\w+)",
"severity": "medium",
"message": "Null-forgiving operator (!) used — verify the value is truly non-null"
},
{
"name": "csharp_async_void",
"pattern": r"\basync\s+void\s+\w+\s*\(",
"severity": "high",
"message": "'async void' method — use only for event handlers"
},
{
"name": "csharp_blocking_async",
"pattern": r"\.(?:Result\b|Wait\(\)|GetAwaiter\(\)\.GetResult\(\))",
"severity": "high",
"message": "Blocking call on async operation — can deadlock in ASP.NET contexts"
}
]
def run_git_command(cmd: List[str], cwd: Path) -> Tuple[bool, str]:
"""Run a git command and return success status and output."""
try:
result = subprocess.run(
cmd,
cwd=cwd,
capture_output=True,
text=True,
timeout=30
)
return result.returncode == 0, result.stdout.strip()
except subprocess.TimeoutExpired:
return False, "Command timed out"
except Exception as e:
return False, str(e)
def get_changed_files(repo_path: Path, base: str, head: str) -> List[Dict]:
"""Get list of changed files between two refs."""
success, output = run_git_command(
["git", "diff", "--name-status", f"{base}...{head}"],
repo_path
)
if not success:
# Try without the triple dot (for uncommitted changes)
success, output = run_git_command(
["git", "diff", "--name-status", base, head],
repo_path
)
if not success or not output:
# Fall back to staged changes
success, output = run_git_command(
["git", "diff", "--name-status", "--cached"],
repo_path
)
files = []
for line in output.split("\n"):
if not line.strip():
continue
parts = line.split("\t")
if len(parts) >= 2:
status = parts[0][0] # First character of status
filepath = parts[-1] # Handle renames (R100\told\tnew)
status_map = {
"A": "added",
"M": "modified",
"D": "deleted",
"R": "renamed",
"C": "copied"
}
files.append({
"path": filepath,
"status": status_map.get(status, "modified")
})
return files
def get_file_diff(repo_path: Path, filepath: str, base: str, head: str) -> str:
"""Get diff content for a specific file."""
success, output = run_git_command(
["git", "diff", f"{base}...{head}", "--", filepath],
repo_path
)
if not success:
success, output = run_git_command(
["git", "diff", "--cached", "--", filepath],
repo_path
)
return output if success else ""
def categorize_file(filepath: str) -> Tuple[str, int]:
"""Categorize a file based on its path and name."""
filepath_lower = filepath.lower()
for category, info in FILE_CATEGORIES.items():
for pattern in info["patterns"]:
if re.search(pattern, filepath_lower):
return category, info["weight"]
return "medium", 2 # Default category
def analyze_diff_for_risks(diff_content: str, filepath: str) -> List[Dict]:
"""Analyze diff content for risky patterns."""
risks = []
# Only analyze added lines (starting with +)
added_lines = [
line[1:] for line in diff_content.split("\n")
if line.startswith("+") and not line.startswith("+++")
]
content = "\n".join(added_lines)
for risk in RISK_PATTERNS:
matches = re.findall(risk["pattern"], content, re.IGNORECASE)
if matches:
risks.append({
"name": risk["name"],
"severity": risk["severity"],
"message": risk["message"],
"file": filepath,
"count": len(matches)
})
return risks
def count_changes(diff_content: str) -> Dict[str, int]:
"""Count additions and deletions in diff."""
additions = 0
deletions = 0
for line in diff_content.split("\n"):
if line.startswith("+") and not line.startswith("+++"):
additions += 1
elif line.startswith("-") and not line.startswith("---"):
deletions += 1
return {"additions": additions, "deletions": deletions}
def calculate_complexity_score(files: List[Dict], all_risks: List[Dict]) -> int:
"""Calculate overall PR complexity score (1-10)."""
score = 0
# File count contribution (max 3 points)
file_count = len(files)
if file_count > 20:
score += 3
elif file_count > 10:
score += 2
elif file_count > 5:
score += 1
# Total changes contribution (max 3 points)
total_changes = sum(f.get("additions", 0) + f.get("deletions", 0) for f in files)
if total_changes > 500:
score += 3
elif total_changes > 200:
score += 2
elif total_changes > 50:
score += 1
# Risk severity contribution (max 4 points)
critical_risks = sum(1 for r in all_risks if r["severity"] == "critical")
high_risks = sum(1 for r in all_risks if r["severity"] == "high")
score += min(2, critical_risks)
score += min(2, high_risks)
return min(10, max(1, score))
def analyze_commit_messages(repo_path: Path, base: str, head: str) -> Dict:
"""Analyze commit messages in the PR."""
success, output = run_git_command(
["git", "log", "--oneline", f"{base}...{head}"],
repo_path
)
if not success or not output:
return {"commits": 0, "issues": []}
commits = output.strip().split("\n")
issues = []
for commit in commits:
if len(commit) < 10:
continue
# Check for conventional commit format
message = commit[8:] if len(commit) > 8 else commit # Skip hash
if not re.match(r"^(feat|fix|docs|style|refactor|test|chore|perf|ci|build|revert)(\(.+\))?:", message):
issues.append({
"commit": commit[:7],
"issue": "Does not follow conventional commit format"
})
if len(message) > 72:
issues.append({
"commit": commit[:7],
"issue": "Commit message exceeds 72 characters"
})
return {
"commits": len(commits),
"issues": issues
}
def analyze_pr(
repo_path: Path,
base: str = "main",
head: str = "HEAD"
) -> Dict:
"""Perform complete PR analysis."""
# Get changed files
changed_files = get_changed_files(repo_path, base, head)
if not changed_files:
return {
"status": "no_changes",
"message": "No changes detected between branches"
}
# Analyze each file
all_risks = []
file_analyses = []
for file_info in changed_files:
filepath = file_info["path"]
category, weight = categorize_file(filepath)
# Get diff for the file
diff = get_file_diff(repo_path, filepath, base, head)
changes = count_changes(diff)
risks = analyze_diff_for_risks(diff, filepath)
all_risks.extend(risks)
file_analyses.append({
"path": filepath,
"status": file_info["status"],
"category": category,
"priority_weight": weight,
"additions": changes["additions"],
"deletions": changes["deletions"],
"risks": risks
})
# Sort by priority (highest first)
file_analyses.sort(key=lambda x: (-x["priority_weight"], x["path"]))
# Analyze commits
commit_analysis = analyze_commit_messages(repo_path, base, head)
# Calculate metrics
complexity = calculate_complexity_score(file_analyses, all_risks)
total_additions = sum(f["additions"] for f in file_analyses)
total_deletions = sum(f["deletions"] for f in file_analyses)
return {
"status": "analyzed",
"summary": {
"files_changed": len(file_analyses),
"total_additions": total_additions,
"total_deletions": total_deletions,
"complexity_score": complexity,
"complexity_label": get_complexity_label(complexity),
"commits": commit_analysis["commits"]
},
"risks": {
"critical": [r for r in all_risks if r["severity"] == "critical"],
"high": [r for r in all_risks if r["severity"] == "high"],
"medium": [r for r in all_risks if r["severity"] == "medium"],
"low": [r for r in all_risks if r["severity"] == "low"]
},
"files": file_analyses,
"commit_issues": commit_analysis["issues"],
"review_order": [f["path"] for f in file_analyses[:10]] # Top 10 priority files
}
def get_complexity_label(score: int) -> str:
"""Get human-readable complexity label."""
if score <= 2:
return "Simple"
elif score <= 4:
return "Moderate"
elif score <= 6:
return "Complex"
elif score <= 8:
return "Very Complex"
else:
return "Critical"
def print_report(analysis: Dict) -> None:
"""Print human-readable analysis report."""
if analysis["status"] == "no_changes":
print("No changes detected.")
return
summary = analysis["summary"]
risks = analysis["risks"]
print("=" * 60)
print("PR ANALYSIS REPORT")
print("=" * 60)
print(f"\nComplexity: {summary['complexity_score']}/10 ({summary['complexity_label']})")
print(f"Files Changed: {summary['files_changed']}")
print(f"Lines: +{summary['total_additions']} / -{summary['total_deletions']}")
print(f"Commits: {summary['commits']}")
# Risk summary
print("\n--- RISK SUMMARY ---")
print(f"Critical: {len(risks['critical'])}")
print(f"High: {len(risks['high'])}")
print(f"Medium: {len(risks['medium'])}")
print(f"Low: {len(risks['low'])}")
# Critical and high risks details
if risks["critical"]:
print("\n--- CRITICAL RISKS ---")
for risk in risks["critical"]:
print(f" [{risk['file']}] {risk['message']} (x{risk['count']})")
if risks["high"]:
print("\n--- HIGH RISKS ---")
for risk in risks["high"]:
print(f" [{risk['file']}] {risk['message']} (x{risk['count']})")
# Commit message issues
if analysis["commit_issues"]:
print("\n--- COMMIT MESSAGE ISSUES ---")
for issue in analysis["commit_issues"][:5]:
print(f" {issue['commit']}: {issue['issue']}")
# Review order
print("\n--- SUGGESTED REVIEW ORDER ---")
for i, filepath in enumerate(analysis["review_order"], 1):
file_info = next(f for f in analysis["files"] if f["path"] == filepath)
print(f" {i}. [{file_info['category'].upper()}] {filepath}")
print("\n" + "=" * 60)
def main():
parser = argparse.ArgumentParser(
description="Analyze pull request for review complexity and risks"
)
parser.add_argument(
"repo_path",
nargs="?",
default=".",
help="Path to git repository (default: current directory)"
)
parser.add_argument(
"--base", "-b",
default="main",
help="Base branch for comparison (default: main)"
)
parser.add_argument(
"--head",
default="HEAD",
help="Head branch/commit for comparison (default: HEAD)"
)
parser.add_argument(
"--json",
action="store_true",
help="Output in JSON format"
)
parser.add_argument(
"--output", "-o",
help="Write output to file"
)
args = parser.parse_args()
repo_path = Path(args.repo_path).resolve()
if not (repo_path / ".git").exists():
print(f"Error: {repo_path} is not a git repository", file=sys.stderr)
sys.exit(1)
analysis = analyze_pr(repo_path, args.base, args.head)
if args.json:
output = json.dumps(analysis, indent=2)
if args.output:
with open(args.output, "w") as f:
f.write(output)
print(f"Results written to {args.output}")
else:
print(output)
else:
print_report(analysis)
if __name__ == "__main__":
main()
FILE:scripts/review_report_generator.py
#!/usr/bin/env python3
"""
Review Report Generator
Generates comprehensive code review reports by combining PR analysis
and code quality findings into structured, actionable reports.
Usage:
python review_report_generator.py /path/to/repo
python review_report_generator.py . --pr-analysis pr_results.json --quality-analysis quality_results.json
python review_report_generator.py /path/to/repo --format markdown --output review.md
"""
import argparse
import json
import os
import subprocess
import sys
from datetime import datetime
from pathlib import Path
from typing import Dict, List, Optional, Tuple
# Severity weights for prioritization
SEVERITY_WEIGHTS = {
"critical": 100,
"high": 75,
"medium": 50,
"low": 25,
"info": 10
}
# Review verdict thresholds
VERDICT_THRESHOLDS = {
"approve": {"max_critical": 0, "max_high": 0, "max_score": 100},
"approve_with_suggestions": {"max_critical": 0, "max_high": 2, "max_score": 85},
"request_changes": {"max_critical": 0, "max_high": 5, "max_score": 70},
"block": {"max_critical": float("inf"), "max_high": float("inf"), "max_score": 0}
}
def load_json_file(filepath: str) -> Optional[Dict]:
"""Load JSON file if it exists."""
try:
with open(filepath, "r") as f:
return json.load(f)
except (FileNotFoundError, json.JSONDecodeError):
return None
def run_pr_analyzer(repo_path: Path) -> Dict:
"""Run pr_analyzer.py and return results."""
script_path = Path(__file__).parent / "pr_analyzer.py"
if not script_path.exists():
return {"status": "error", "message": "pr_analyzer.py not found"}
try:
result = subprocess.run(
[sys.executable, str(script_path), str(repo_path), "--json"],
capture_output=True,
text=True,
timeout=120
)
if result.returncode == 0:
return json.loads(result.stdout)
return {"status": "error", "message": result.stderr}
except Exception as e:
return {"status": "error", "message": str(e)}
def run_quality_checker(repo_path: Path) -> Dict:
"""Run code_quality_checker.py and return results."""
script_path = Path(__file__).parent / "code_quality_checker.py"
if not script_path.exists():
return {"status": "error", "message": "code_quality_checker.py not found"}
try:
result = subprocess.run(
[sys.executable, str(script_path), str(repo_path), "--json"],
capture_output=True,
text=True,
timeout=300
)
if result.returncode == 0:
return json.loads(result.stdout)
return {"status": "error", "message": result.stderr}
except Exception as e:
return {"status": "error", "message": str(e)}
def calculate_review_score(pr_analysis: Dict, quality_analysis: Dict) -> int:
"""Calculate overall review score (0-100)."""
score = 100
# Deduct for PR risks
if "risks" in pr_analysis:
risks = pr_analysis["risks"]
score -= len(risks.get("critical", [])) * 15
score -= len(risks.get("high", [])) * 10
score -= len(risks.get("medium", [])) * 5
score -= len(risks.get("low", [])) * 2
# Deduct for code quality issues
if "issues" in quality_analysis:
issues = quality_analysis["issues"]
score -= len([i for i in issues if i.get("severity") == "critical"]) * 12
score -= len([i for i in issues if i.get("severity") == "high"]) * 8
score -= len([i for i in issues if i.get("severity") == "medium"]) * 4
score -= len([i for i in issues if i.get("severity") == "low"]) * 1
# Deduct for complexity
if "summary" in pr_analysis:
complexity = pr_analysis["summary"].get("complexity_score", 0)
if complexity > 7:
score -= 10
elif complexity > 5:
score -= 5
return max(0, min(100, score))
def determine_verdict(score: int, critical_count: int, high_count: int) -> Tuple[str, str]:
"""Determine review verdict based on score and issue counts."""
if critical_count > 0:
return "block", "Critical issues must be resolved before merge"
if score >= 90 and high_count == 0:
return "approve", "Code meets quality standards"
if score >= 75 and high_count <= 2:
return "approve_with_suggestions", "Minor improvements recommended"
if score >= 50:
return "request_changes", "Several issues need to be addressed"
return "block", "Significant issues prevent approval"
def generate_findings_list(pr_analysis: Dict, quality_analysis: Dict) -> List[Dict]:
"""Combine and prioritize all findings."""
findings = []
# Add PR risk findings
if "risks" in pr_analysis:
for severity, items in pr_analysis["risks"].items():
for item in items:
findings.append({
"source": "pr_analysis",
"severity": severity,
"category": item.get("name", "unknown"),
"message": item.get("message", ""),
"file": item.get("file", ""),
"count": item.get("count", 1)
})
# Add code quality findings
if "issues" in quality_analysis:
for issue in quality_analysis["issues"]:
findings.append({
"source": "quality_analysis",
"severity": issue.get("severity", "medium"),
"category": issue.get("type", "unknown"),
"message": issue.get("message", ""),
"file": issue.get("file", ""),
"line": issue.get("line", 0)
})
# Sort by severity weight
findings.sort(
key=lambda x: -SEVERITY_WEIGHTS.get(x["severity"], 0)
)
return findings
def generate_action_items(findings: List[Dict]) -> List[Dict]:
"""Generate prioritized action items from findings."""
action_items = []
seen_categories = set()
for finding in findings:
category = finding["category"]
severity = finding["severity"]
# Group similar issues
if category in seen_categories and severity not in ["critical", "high"]:
continue
action = {
"priority": "P0" if severity == "critical" else "P1" if severity == "high" else "P2",
"action": get_action_for_category(category, finding),
"severity": severity,
"files_affected": [finding["file"]] if finding.get("file") else []
}
action_items.append(action)
seen_categories.add(category)
return action_items[:15] # Top 15 actions
def get_action_for_category(category: str, finding: Dict) -> str:
"""Get actionable recommendation for issue category."""
actions = {
"hardcoded_secrets": "Remove hardcoded credentials and use environment variables or a secrets manager",
"sql_concatenation": "Use parameterized queries to prevent SQL injection",
"debugger": "Remove debugger statements before merging",
"console_log": "Remove or replace console statements with proper logging",
"todo_fixme": "Address TODO/FIXME comments or create tracking issues",
"disable_eslint": "Address the underlying issue instead of disabling lint rules",
"any_type": "Replace 'any' types with proper type definitions",
"long_function": "Break down function into smaller, focused units",
"god_class": "Split class into smaller, single-responsibility classes",
"too_many_params": "Use parameter objects or builder pattern",
"deep_nesting": "Refactor using early returns, guard clauses, or extraction",
"high_complexity": "Reduce cyclomatic complexity through refactoring",
"missing_error_handling": "Add proper error handling and recovery logic",
"duplicate_code": "Extract duplicate code into shared functions",
"magic_numbers": "Replace magic numbers with named constants",
"large_file": "Consider splitting into multiple smaller modules"
}
return actions.get(category, f"Review and address: {finding.get('message', category)}")
def format_markdown_report(report: Dict) -> str:
"""Generate markdown-formatted report."""
lines = []
# Header
lines.append("# Code Review Report")
lines.append("")
lines.append(f"**Generated:** {report['metadata']['generated_at']}")
lines.append(f"**Repository:** {report['metadata']['repository']}")
lines.append("")
# Executive Summary
lines.append("## Executive Summary")
lines.append("")
summary = report["summary"]
verdict = summary["verdict"]
verdict_emoji = {
"approve": "✅",
"approve_with_suggestions": "✅",
"request_changes": "⚠️",
"block": "❌"
}.get(verdict, "❓")
lines.append(f"**Verdict:** {verdict_emoji} {verdict.upper().replace('_', ' ')}")
lines.append(f"**Score:** {summary['score']}/100")
lines.append(f"**Rationale:** {summary['rationale']}")
lines.append("")
# Issue Counts
lines.append("### Issue Summary")
lines.append("")
lines.append("| Severity | Count |")
lines.append("|----------|-------|")
for severity in ["critical", "high", "medium", "low"]:
count = summary["issue_counts"].get(severity, 0)
lines.append(f"| {severity.capitalize()} | {count} |")
lines.append("")
# PR Statistics (if available)
if "pr_summary" in report:
pr = report["pr_summary"]
lines.append("### Change Statistics")
lines.append("")
lines.append(f"- **Files Changed:** {pr.get('files_changed', 'N/A')}")
lines.append(f"- **Lines Added:** +{pr.get('total_additions', 0)}")
lines.append(f"- **Lines Removed:** -{pr.get('total_deletions', 0)}")
lines.append(f"- **Complexity:** {pr.get('complexity_label', 'N/A')}")
lines.append("")
# Action Items
if report.get("action_items"):
lines.append("## Action Items")
lines.append("")
for i, item in enumerate(report["action_items"], 1):
priority = item["priority"]
emoji = "🔴" if priority == "P0" else "🟠" if priority == "P1" else "🟡"
lines.append(f"{i}. {emoji} **[{priority}]** {item['action']}")
if item.get("files_affected"):
lines.append(f" - Files: {', '.join(item['files_affected'][:3])}")
lines.append("")
# Critical Findings
critical_findings = [f for f in report.get("findings", []) if f["severity"] == "critical"]
if critical_findings:
lines.append("## Critical Issues (Must Fix)")
lines.append("")
for finding in critical_findings:
lines.append(f"- **{finding['category']}** in `{finding.get('file', 'unknown')}`")
lines.append(f" - {finding['message']}")
lines.append("")
# High Priority Findings
high_findings = [f for f in report.get("findings", []) if f["severity"] == "high"]
if high_findings:
lines.append("## High Priority Issues")
lines.append("")
for finding in high_findings[:10]:
lines.append(f"- **{finding['category']}** in `{finding.get('file', 'unknown')}`")
lines.append(f" - {finding['message']}")
lines.append("")
# Review Order (if available)
if "review_order" in report:
lines.append("## Suggested Review Order")
lines.append("")
for i, filepath in enumerate(report["review_order"][:10], 1):
lines.append(f"{i}. `{filepath}`")
lines.append("")
# Footer
lines.append("---")
lines.append("*Generated by Code Reviewer*")
return "\n".join(lines)
def format_text_report(report: Dict) -> str:
"""Generate plain text report."""
lines = []
lines.append("=" * 60)
lines.append("CODE REVIEW REPORT")
lines.append("=" * 60)
lines.append("")
lines.append(f"Generated: {report['metadata']['generated_at']}")
lines.append(f"Repository: {report['metadata']['repository']}")
lines.append("")
summary = report["summary"]
verdict = summary["verdict"].upper().replace("_", " ")
lines.append(f"VERDICT: {verdict}")
lines.append(f"SCORE: {summary['score']}/100")
lines.append(f"RATIONALE: {summary['rationale']}")
lines.append("")
lines.append("--- ISSUE SUMMARY ---")
for severity in ["critical", "high", "medium", "low"]:
count = summary["issue_counts"].get(severity, 0)
lines.append(f" {severity.capitalize()}: {count}")
lines.append("")
if report.get("action_items"):
lines.append("--- ACTION ITEMS ---")
for i, item in enumerate(report["action_items"][:10], 1):
lines.append(f" {i}. [{item['priority']}] {item['action']}")
lines.append("")
critical = [f for f in report.get("findings", []) if f["severity"] == "critical"]
if critical:
lines.append("--- CRITICAL ISSUES ---")
for f in critical:
lines.append(f" [{f.get('file', 'unknown')}] {f['message']}")
lines.append("")
lines.append("=" * 60)
return "\n".join(lines)
def generate_report(
repo_path: Path,
pr_analysis: Optional[Dict] = None,
quality_analysis: Optional[Dict] = None
) -> Dict:
"""Generate comprehensive review report."""
# Run analyses if not provided
if pr_analysis is None:
pr_analysis = run_pr_analyzer(repo_path)
if quality_analysis is None:
quality_analysis = run_quality_checker(repo_path)
# Generate findings
findings = generate_findings_list(pr_analysis, quality_analysis)
# Count issues by severity
issue_counts = {
"critical": len([f for f in findings if f["severity"] == "critical"]),
"high": len([f for f in findings if f["severity"] == "high"]),
"medium": len([f for f in findings if f["severity"] == "medium"]),
"low": len([f for f in findings if f["severity"] == "low"])
}
# Calculate score and verdict
score = calculate_review_score(pr_analysis, quality_analysis)
verdict, rationale = determine_verdict(
score,
issue_counts["critical"],
issue_counts["high"]
)
# Generate action items
action_items = generate_action_items(findings)
# Build report
report = {
"metadata": {
"generated_at": datetime.now().isoformat(),
"repository": str(repo_path),
"version": "1.0.0"
},
"summary": {
"score": score,
"verdict": verdict,
"rationale": rationale,
"issue_counts": issue_counts
},
"findings": findings,
"action_items": action_items
}
# Add PR summary if available
if pr_analysis.get("status") == "analyzed":
report["pr_summary"] = pr_analysis.get("summary", {})
report["review_order"] = pr_analysis.get("review_order", [])
# Add quality summary if available
if quality_analysis.get("status") == "analyzed":
report["quality_summary"] = quality_analysis.get("summary", {})
return report
def main():
parser = argparse.ArgumentParser(
description="Generate comprehensive code review reports"
)
parser.add_argument(
"repo_path",
nargs="?",
default=".",
help="Path to repository (default: current directory)"
)
parser.add_argument(
"--pr-analysis",
help="Path to pre-computed PR analysis JSON"
)
parser.add_argument(
"--quality-analysis",
help="Path to pre-computed quality analysis JSON"
)
parser.add_argument(
"--format", "-f",
choices=["text", "markdown", "json"],
default="text",
help="Output format (default: text)"
)
parser.add_argument(
"--output", "-o",
help="Write output to file"
)
parser.add_argument(
"--json",
action="store_true",
help="Output as JSON (shortcut for --format json)"
)
args = parser.parse_args()
repo_path = Path(args.repo_path).resolve()
if not repo_path.exists():
print(f"Error: Path does not exist: {repo_path}", file=sys.stderr)
sys.exit(1)
# Load pre-computed analyses if provided
pr_analysis = None
quality_analysis = None
if args.pr_analysis:
pr_analysis = load_json_file(args.pr_analysis)
if not pr_analysis:
print(f"Warning: Could not load PR analysis from {args.pr_analysis}")
if args.quality_analysis:
quality_analysis = load_json_file(args.quality_analysis)
if not quality_analysis:
print(f"Warning: Could not load quality analysis from {args.quality_analysis}")
# Generate report
report = generate_report(repo_path, pr_analysis, quality_analysis)
# Format output
output_format = "json" if args.json else args.format
if output_format == "json":
output = json.dumps(report, indent=2)
elif output_format == "markdown":
output = format_markdown_report(report)
else:
output = format_text_report(report)
# Write or print output
if args.output:
with open(args.output, "w") as f:
f.write(output)
print(f"Report written to {args.output}")
else:
print(output)
if __name__ == "__main__":
main()
Gợi ý và chọn đúng lệnh, agent, skill trong Claude Code khi chưa rõ nên dùng công cụ nào.
---
name: "command-guide"
description: >
Claude Code Command Selection Guide - Automatically recommend and select the right
commands, agents, and skills in Claude Code.
Use when: (1) user is unsure which command or tool to use, (2) needs to decide which
agent/skill best fits the current task, (3) querying usage scenarios for /plan, /tdd,
/compact, /loop and other commands, (4) understanding when to invoke planner,
code-reviewer, build-error-resolver and other agents, (5) needs command cheat sheet
or decision flowchart.
Triggers: "which command to use", "which agent", "command selection", "how to use /plan",
"when to use /compact", "agent selection guide", "command cheat sheet", "skill recommendation".
---
# Claude Code Command Selection Guide
This skill helps you choose the most appropriate command, agent, or skill for different scenarios.
## Quick Decision Flowchart
```mermaid
graph TD
A[User Request] --> B{Request Type?}
B -->|New Feature| C[/plan]
B -->|Bug Fix| D[/tdd or build-error-resolver]
B -->|Code Review| E[/code-review or code-reviewer agent]
B -->|Testing| F[/e2e or tdd-guide agent]
B -->|Context Too Long| G[/compact]
B -->|Documentation| H[/docs or docs-lookup agent]
B -->|Looping Task| I[/loop]
B -->|Security Review| J[security-reviewer agent]
C --> K[planner agent]
D --> L{Build Failed?}
L -->|Yes| M[build-error-resolver]
L -->|No| N[tdd-guide]
E --> O[code-reviewer]
F --> P[e2e-runner]
```
## 1. Built-in Slash Commands
### Session Management Commands
| Command | Use Case | Example |
|---------|----------|---------|
| `/compact` | Context too long (>150K tokens), slow response, task phase transition | `/compact` or auto-trigger |
| `/clear` | Start fresh conversation, clear history | `/clear` |
| `/loop` | Periodic task execution, automated looping work | `/loop 5m check build status` |
| `/help` | View help, learn commands | `/help` |
| `/fast` | Need faster response (Opus 4.6 only) | `/fast` |
| `/model` | Switch model | `/model sonnet` |
### Development Workflow Commands
| Command | Use Case | Activation Timing |
|---------|----------|-------------------|
| `/plan` | Start new feature, architecture refactor, complex tasks | **Enter Plan Mode** |
| `/tdd` | Write tests, TDD development workflow | When test guidance needed |
| `/e2e` | E2E testing, critical user flow verification | When browser testing needed |
| `/code-review` | Code quality review | After writing code |
| `/build-fix` | Build failure, type errors | When build fails |
| `/learn` | Extract patterns from session, learning | Before session ends |
| `/skill-create` | Create new skill from git history | When repeating patterns found |
### Documentation & Query Commands
| Command | Use Case | Example |
|---------|----------|---------|
| `/docs` | Update project documentation | `/docs` |
| `/update-codemaps` | Update code maps | `/update-codemaps` |
| `/remember` | Save memory to memory system | `/remember user prefers concise output` |
| `/tasks` | View task list | `/tasks` |
---
## 2. Agents Selection
### Development Workflow Agents
| Agent | Trigger Condition | Purpose |
|-------|-------------------|---------|
| `planner` | Complex feature request, architectural decision | Create implementation plan |
| `architect` | System design, tech stack selection | Architecture analysis and decisions |
| `tdd-guide` | New feature, bug fix | TDD workflow guidance |
| `code-reviewer` | **Invoke immediately after writing code** | Code quality review |
| `security-reviewer` | Handling auth, API, sensitive data | Security vulnerability detection |
### Problem Solving Agents
| Agent | Trigger Condition | Purpose |
|-------|-------------------|---------|
| `build-error-resolver` | **Invoke immediately when build fails** | Fix build/type errors |
| `e2e-runner` | Critical user flows, before PR | E2E test execution |
| `refactor-cleaner` | Code maintenance, dead code cleanup | Dead code detection and cleanup |
| `doc-updater` | Update docs, codemaps | Documentation sync |
### Research & Exploration Agents
| Agent | Trigger Condition | Purpose |
|-------|-------------------|---------|
| `Explore` | Codebase exploration, file finding | Quick codebase exploration |
| `general-purpose` | Complex multi-step tasks | General task handling |
| `docs-lookup` | Query library/framework docs | Get latest API documentation |
---
## 3. Skills Selection
### Workflow Skills
| Skill | Trigger Timing | Purpose |
|-------|----------------|---------|
| `tdd-workflow` | Developing new feature/fixing bug | Complete TDD workflow guidance |
| `verification-loop` | After feature completion, before PR | Comprehensive verification (build/test/lint/security) |
| `strategic-compact` | Long session, context pressure | Guide when to manually `/compact` |
### Architecture & Pattern Skills
| Skill | Trigger Timing | Purpose |
|-------|----------------|---------|
| `frontend-patterns` | Frontend development | React/Next.js/Vue best practices |
| `backend-patterns` | Backend development | API/service architecture patterns |
| `api-design` | API design | RESTful/API design standards |
| `mcp-server-patterns` | MCP server development | MCP configuration and patterns |
### Testing Skills
| Skill | Trigger Timing | Purpose |
|-------|----------------|---------|
| `e2e-testing` | E2E testing needs | Playwright test generation |
| `security-review` | Security review needs | OWASP Top 10 detection |
### Research Skills
| Skill | Trigger Timing | Purpose |
|-------|----------------|---------|
| `deep-research` | Need deep research | Multi-round search and research |
| `exa-search` | Need web search | Web content search |
| `documentation-lookup` | Query library docs | Context7 documentation query |
---
## 4. Scenario Decision Matrix
### By Task Phase
| Phase | Recommended Tool Combination | Reason |
|-------|------------------------------|--------|
| **Requirements Analysis** | `planner` + `Explore` | Plan first, explore later |
| **Architecture Design** | `architect` + `api-design` skill | Professional architecture guidance |
| **Pre-Development** | `tdd-guide` + `tdd-workflow` skill | Test first |
| **During Development** | Direct edit + quick iteration | Stay in flow |
| **Post-Development** | `code-reviewer` + `verification-loop` | Quality gate |
| **Testing Phase** | `e2e-runner` + `e2e-testing` skill | Complete test coverage |
| **Before PR** | `security-reviewer` + `verification-loop` | Final verification |
| **Build Failure** | `build-error-resolver` | Focused fix |
### By Problem Type
| Problem | Invoke Immediately | Note |
|---------|--------------------|------|
| Build failure | `build-error-resolver` | Minimal changes, quick fix |
| Type error | `build-error-resolver` | TypeScript specialist |
| Bug fix | `tdd-guide` | Write test then fix |
| Security vulnerability | `security-reviewer` | OWASP detection |
| Poor code quality | `code-reviewer` | Immediate review |
| Missing documentation | `doc-updater` | Auto update |
| Dead code | `refactor-cleaner` | Safe cleanup |
### By Development Type
| Development Type | Skills Combination |
|------------------|--------------------|
| Frontend feature | `frontend-patterns` + `tdd-workflow` |
| Backend API | `backend-patterns` + `api-design` + `tdd-workflow` |
| MCP server | `mcp-server-patterns` + `tdd-workflow` |
| Database | `database-reviewer` agent |
| Security feature | `security-reviewer` + `security-review` skill |
---
## 5. Parallel Execution Strategy
### Parallelizable Scenarios
Recommended: Launch multiple independent tasks simultaneously
Scenario: Preparing PR after code completion
- Agent 1: code-reviewer (code quality)
- Agent 2: security-reviewer (security review)
- Agent 3: e2e-runner (E2E tests)
Scenario: Large refactor analysis
- Agent 1: architect (architecture analysis)
- Agent 2: Explore (code exploration)
- Agent 3: refactor-cleaner (dead code detection)
### Sequential Execution Required
Cannot parallelize: Dependencies exist
Scenario: Fixing build error
- Sequence: build-error-resolver -> test verification -> code-reviewer
Scenario: New feature development
- Sequence: planner -> tdd-guide (write tests) -> implementation -> code-reviewer
---
## 6. Auto-Trigger Rules
### Invoke Without User Request
| Situation | Auto Action |
|-----------|-------------|
| Code written/modified | **Immediately invoke** `code-reviewer` |
| Build fails | **Immediately invoke** `build-error-resolver` |
| Complex feature request | **Immediately invoke** `planner` |
| Handling auth/sensitive data | **Immediately invoke** `security-reviewer` |
| New feature/bug fix | **Immediately invoke** `tdd-guide` |
| Architectural decision | **Immediately invoke** `architect` |
---
## 7. Context Management Timing
| Indicator | Trigger `/compact` |
|-----------|-------------------|
| Token > 150K | Immediately compact |
| Slow response | Suggest compact |
| Task phase switch | Compact at boundary |
| Major milestone completed | Compact then continue |
| Debugging ends -> new task | Clear debug traces |
**Best Practices**:
- Compact after research, before implementation (preserve plan)
- Compact after milestone completion (clear intermediate state)
- Don't compact mid-implementation (lose variables/paths)
---
## 8. Command Cheat Sheet
```
Development Workflow:
/plan -> Enter planning mode (complex tasks)
/tdd -> TDD workflow
/e2e -> E2E testing
/code-review -> Code review
/build-fix -> Fix build
Session Management:
/compact -> Compact context
/clear -> Clear session
/loop -> Looping task
/fast -> Fast mode
Documentation & Memory:
/docs -> Update docs
/remember -> Save memory
/tasks -> View tasks
Help:
/help -> View all commands
```
---
## 9. Usage Examples
### Example 1: New Feature Development
User: Add user authentication feature
Workflow:
1. /plan -> planner agent creates plan
2. tdd-guide -> write tests
3. Implementation -> edit code
4. code-reviewer -> code review
5. security-reviewer -> security review (auth sensitive)
6. e2e-runner -> E2E tests
7. /compact -> compact after milestone completion
### Example 2: Build Failure
User: npm run build failed
Workflow:
1. build-error-resolver -> analyze error, minimal fix
2. Verify build success
3. code-reviewer -> check fix quality
### Example 3: Code Refactoring
User: Refactor authentication module
Workflow:
1. architect -> architecture analysis
2. planner -> implementation plan
3. refactor-cleaner -> dead code detection
4. tdd-guide -> ensure test coverage
5. Implementation -> refactor code
6. verification-loop -> comprehensive verification
---
**Core Principles**:
1. **Plan first, implement later** - Use `/plan` for complex tasks
2. **Test first** - Use `tdd-guide` for new features
3. **Review immediately after coding** - Use `code-reviewer` when code complete
4. **Fix build immediately when failed** - Use `build-error-resolver`
5. **Review sensitive code** - Use `security-reviewer` for auth/API
6. **Verify comprehensively before PR** - Use `verification-loop`
Phân tích sản phẩm đối thủ từ trang giá, đánh giá ứng dụng, tin tuyển dụng, SEO, tạo ma trận tính năng và SWOT.
---
name: "competitive-teardown"
description: "Analyzes competitor products and companies by synthesizing data from pricing pages, app store reviews, job postings, SEO signals, and social media into structured competitive intelligence. Produces feature comparison matrices scored across 12 dimensions, SWOT analyses, positioning maps, UX audits, pricing model breakdowns, action item roadmaps, and stakeholder presentation templates. Use when conducting competitor analysis, comparing products against competitors, researching the competitive landscape, building battle cards for sales, preparing for a product strategy or roadmap session, responding to a competitor's new feature or pricing change, or performing a quarterly competitive review."
---
# Competitive Teardown
**Tier:** POWERFUL
**Category:** Product Team
**Domain:** Competitive Intelligence, Product Strategy, Market Analysis
---
## When to Use
- Before a product strategy or roadmap session
- When a competitor launches a major feature or pricing change
- Quarterly competitive review
- Before a sales pitch where you need battle card data
- When entering a new market segment
---
## Teardown Workflow
Follow these steps in sequence to produce a complete teardown:
1. **Define competitors** — List 2–4 competitors to analyze. Confirm which is the primary focus.
2. **Collect data** — Use `references/data-collection-guide.md` to gather raw signals from at least 3 sources per competitor (website, reviews, job postings, SEO, social).
_Validation checkpoint: Before proceeding, confirm you have pricing data, at least 20 reviews, and job posting counts for each competitor._
3. **Score using rubric** — Apply the 12-dimension rubric below to produce a numeric scorecard for each competitor and your own product.
_Validation checkpoint: Every dimension should have a score and at least one supporting evidence note._
4. **Generate outputs** — Populate the templates in `references/analysis-templates.md` (Feature Matrix, Pricing Analysis, SWOT, Positioning Map, UX Audit).
5. **Build action plan** — Translate findings into the Action Items template (quick wins / medium-term / strategic).
6. **Package for stakeholders** — Assemble the Stakeholder Presentation using outputs from steps 3–5.
---
## Data Collection Guide
> Full executable scripts for each source are in `references/data-collection-guide.md`. Summaries of what to capture are below.
### 1. Website Analysis
Key things to capture:
- Pricing tiers and price points
- Feature lists per tier
- Primary CTA and messaging
- Case studies / customer logos (signals ICP)
- Integration logos
- Trust signals (certifications, compliance badges)
### 2. App Store Reviews
Review sentiment categories:
- **Praise** → what users love (defend / strengthen these)
- **Feature requests** → unmet needs (opportunity gaps)
- **Bugs** → quality signals
- **UX complaints** → friction points you can beat them on
**Sample App Store query (iTunes Search API):**
```
GET https://itunes.apple.com/search?term=<competitor_name>&entity=software&limit=1
# Extract trackId, then:
GET https://itunes.apple.com/rss/customerreviews/id=<trackId>/sortBy=mostRecent/json?l=en&limit=50
```
Parse `entry[].content.label` for review text and `entry[].im:rating.label` for star rating.
### 3. Job Postings (Team Size & Tech Stack Signals)
Signals from job postings:
- **Engineering volume** → scaling vs. consolidating
- **Specific tech mentions** → stack (React/Vue, Postgres/Mongo, AWS/GCP)
- **Sales/CS ratio** → product-led vs. sales-led motion
- **Data/ML roles** → upcoming AI features
- **Compliance roles** → regulatory expansion
### 4. SEO Analysis
SEO signals to capture:
- Top 20 organic keywords (intent: informational / navigational / commercial)
- Domain Authority / backlink count
- Blog publishing cadence and topics
- Which pages rank (product pages vs. blog vs. docs)
### 5. Social Media Sentiment
Capture recent mentions via Twitter/X API v2, Reddit, or LinkedIn. Look for recurring praise, complaints, and feature requests. See `references/data-collection-guide.md` for API query examples.
---
## Scoring Rubric (12 Dimensions, 1-5)
| # | Dimension | 1 (Weak) | 3 (Average) | 5 (Best-in-class) |
|---|-----------|----------|-------------|-------------------|
| 1 | **Features** | Core only, many gaps | Solid coverage | Comprehensive + unique |
| 2 | **Pricing** | Confusing / overpriced | Market-rate, clear | Transparent, flexible, fair |
| 3 | **UX** | Confusing, high friction | Functional | Delightful, minimal friction |
| 4 | **Performance** | Slow, unreliable | Acceptable | Fast, high uptime |
| 5 | **Docs** | Sparse, outdated | Decent coverage | Comprehensive, searchable |
| 6 | **Support** | Email only, slow | Chat + email | 24/7, great response |
| 7 | **Integrations** | 0-5 integrations | 6-25 | 26+ or deep ecosystem |
| 8 | **Security** | No mentions | SOC2 claimed | SOC2 Type II, ISO 27001 |
| 9 | **Scalability** | No enterprise tier | Mid-market ready | Enterprise-grade |
| 10 | **Brand** | Generic, unmemorable | Decent positioning | Strong, differentiated |
| 11 | **Community** | None | Forum / Slack | Active, vibrant community |
| 12 | **Innovation** | No recent releases | Quarterly | Frequent, meaningful |
**Example completed row** (Competitor: Acme Corp, Dimension 3 – UX):
| Dimension | Acme Corp Score | Evidence |
|-----------|----------------|---------|
| UX | 2 | App Store reviews cite "confusing navigation" (38 mentions); onboarding requires 7 steps before TTFV; no onboarding wizard; CC required at signup. |
Apply this pattern to all 12 dimensions for each competitor.
---
## Templates
> Full template markdown is in `references/analysis-templates.md`. Abbreviated reference below.
### Feature Comparison Matrix
Rows: core features, pricing tiers, platform capabilities (web, iOS, Android, API).
Columns: your product + up to 3 competitors.
Score each cell 1–5. Sum to get total out of 60.
**Score legend:** 5=Best-in-class, 4=Strong, 3=Average, 2=Below average, 1=Weak/Missing
### Pricing Analysis
Capture per competitor: model type (per-seat / usage-based / flat rate / freemium), entry/mid/enterprise price points, free trial length.
Summarize: price leader, value leader, premium positioning, your position, and 2–3 pricing opportunity bullets.
### SWOT Analysis
For each competitor: 3–5 bullets per quadrant (Strengths, Weaknesses, Opportunities for us, Threats to us). Anchor every bullet to a data signal (review quote, job posting count, pricing page, etc.).
### Positioning Map
2x2 axes (e.g., Simple ↔ Complex / Low Value ↔ High Value). Place each competitor and your product. Bubble size = market share or funding. See `references/analysis-templates.md` for ASCII and editable versions.
### UX Audit Checklist
Onboarding: TTFV (minutes), steps to activation, CC-required, onboarding wizard quality.
Key workflows: steps, friction points, comparative score (yours vs. theirs).
Mobile: iOS/Android ratings, feature parity, top complaint and praise.
Navigation: global search, keyboard shortcuts, in-app help.
### Action Items
| Horizon | Effort | Examples |
|---------|--------|---------|
| Quick wins (0–4 wks) | Low | Add review badges, publish comparison landing page |
| Medium-term (1–3 mo) | Moderate | Launch free tier, improve onboarding TTFV, add top-requested integration |
| Strategic (3–12 mo) | High | Enter new market, build API v2, achieve SOC2 Type II |
### Stakeholder Presentation (7 slides)
1. **Executive Summary** — Threat level (LOW/MEDIUM/HIGH/CRITICAL), top strength, top opportunity, recommended action
2. **Market Position** — 2x2 positioning map
3. **Feature Scorecard** — 12-dimension radar or table, total scores
4. **Pricing Analysis** — Comparison table + key insight
5. **UX Highlights** — What they do better (3 bullets) vs. where we win (3 bullets)
6. **Voice of Customer** — Top 3 review complaints (quoted or paraphrased)
7. **Our Action Plan** — Quick wins, medium-term, strategic priorities; Appendix with raw data
## Related Skills
- **Product Strategist** (`product-team/product-strategist/`) — Competitive insights feed OKR and strategy planning
- **Landing Page Generator** (`product-team/landing-page-generator/`) — Competitive positioning informs landing page messaging
FILE:references/analysis-templates.md
# Competitive Analysis Templates
## 1. SWOT Analysis Template
### Company/Product: [Competitor Name]
**Date:** [Analysis Date] | **Analyst:** [Name] | **Version:** [1.0]
#### Strengths (Internal Advantages)
| # | Strength | Evidence | Impact |
|---|----------|----------|--------|
| 1 | [e.g., Strong brand recognition] | [Source/data point] | High/Med/Low |
| 2 | | | |
| 3 | | | |
#### Weaknesses (Internal Limitations)
| # | Weakness | Evidence | Exploitability |
|---|----------|----------|---------------|
| 1 | [e.g., Limited API capabilities] | [Source/data point] | High/Med/Low |
| 2 | | | |
| 3 | | | |
#### Opportunities (External Favorable)
| # | Opportunity | Timeframe | Our Advantage |
|---|------------|-----------|---------------|
| 1 | [e.g., Competitor slow to adopt AI] | Short/Med/Long | [How we capitalize] |
| 2 | | | |
| 3 | | | |
#### Threats (External Unfavorable)
| # | Threat | Likelihood | Mitigation |
|---|--------|-----------|-----------|
| 1 | [e.g., Competitor acquired by larger company] | High/Med/Low | [Our response plan] |
| 2 | | | |
| 3 | | | |
---
## 2. Porter's Five Forces (Product Application)
### Market: [Your Product Category]
#### Force 1: Competitive Rivalry (Intensity: High/Med/Low)
- Number of direct competitors: ___
- Market growth rate: ___% annually
- Product differentiation level: High/Med/Low
- Switching costs for customers: High/Med/Low
- Exit barriers: High/Med/Low
- **Assessment:** [Summary of competitive rivalry intensity]
#### Force 2: Threat of New Entrants (Intensity: High/Med/Low)
- Capital requirements: High/Med/Low
- Technology barriers: High/Med/Low
- Network effects strength: Strong/Moderate/Weak
- Regulatory barriers: High/Med/Low
- Brand loyalty in market: Strong/Moderate/Weak
- **Assessment:** [Summary of new entrant threat]
#### Force 3: Threat of Substitutes (Intensity: High/Med/Low)
- Alternative solutions: [List substitutes]
- Price-performance of substitutes: Better/Same/Worse
- Switching costs to substitutes: High/Med/Low
- Customer propensity to switch: High/Med/Low
- **Assessment:** [Summary of substitute threat]
#### Force 4: Bargaining Power of Buyers (Power: High/Med/Low)
- Buyer concentration: Concentrated/Fragmented
- Price sensitivity: High/Med/Low
- Information availability: Full/Partial/Limited
- Switching costs: High/Med/Low
- Volume of purchases: High/Med/Low
- **Assessment:** [Summary of buyer power]
#### Force 5: Bargaining Power of Suppliers (Power: High/Med/Low)
- Key technology dependencies: [List]
- Cloud provider lock-in: High/Med/Low
- Talent market tightness: Tight/Balanced/Loose
- Data source dependencies: Critical/Important/Optional
- **Assessment:** [Summary of supplier power]
#### Overall Industry Attractiveness: [Score 1-10]
---
## 3. Competitive Positioning Map
### Axis Definitions
- **X-Axis:** [e.g., Ease of Use] (Low to High)
- **Y-Axis:** [e.g., Feature Completeness] (Low to High)
### Competitor Positions
| Competitor | X Score (1-10) | Y Score (1-10) | Quadrant |
|-----------|---------------|---------------|----------|
| Your Product | ___ | ___ | ___ |
| Competitor A | ___ | ___ | ___ |
| Competitor B | ___ | ___ | ___ |
| Competitor C | ___ | ___ | ___ |
| Competitor D | ___ | ___ | ___ |
### Quadrant Definitions
- **Top-Right (Leaders):** High on both axes - market leaders
- **Top-Left (Feature-Rich):** High features, lower ease of use - complex tools
- **Bottom-Right (Simple):** Easy to use, fewer features - niche players
- **Bottom-Left (Laggards):** Low on both axes - disruption candidates
### Positioning Insights
- **White space opportunities:** [Areas with no competitor presence]
- **Crowded areas:** [Where competition is fiercest]
- **Our trajectory:** [Direction we're moving on the map]
---
## 4. Win/Loss Analysis Template
### Deal: [Opportunity Name]
**Date:** [Close Date] | **Result:** Won / Lost | **Competitor:** [Name]
#### Deal Context
- **Deal Size:** $___
- **Sales Cycle:** ___ days
- **Segment:** SMB / Mid-Market / Enterprise
- **Industry:** ___
- **Decision Makers:** [Roles involved]
- **Evaluation Criteria:** [What mattered most to buyer]
#### Competitive Comparison (Buyer Perspective)
| Factor | Us (Score 1-5) | Competitor (Score 1-5) | Decisive? |
|--------|---------------|----------------------|-----------|
| Product Fit | | | Yes/No |
| Pricing | | | Yes/No |
| Ease of Use | | | Yes/No |
| Support Quality | | | Yes/No |
| Integration | | | Yes/No |
| Brand/Trust | | | Yes/No |
| Implementation | | | Yes/No |
#### Win/Loss Factors
- **Primary reason for outcome:** [Single most important factor]
- **Secondary factors:** [Supporting reasons]
- **Buyer quotes:** ["Direct quotes from debrief"]
#### Action Items
| # | Action | Owner | Due Date |
|---|--------|-------|----------|
| 1 | [e.g., Improve onboarding flow] | [Name] | [Date] |
| 2 | | | |
---
## 5. Battle Card Template
### Competitor: [Name]
**Last Updated:** [Date] | **Confidence:** High/Med/Low
#### Quick Facts
- **Founded:** ___
- **Funding:** $___
- **Employees:** ___
- **Customers:** ___
- **HQ:** ___
#### Elevator Pitch (Their Positioning)
> [How the competitor describes themselves in one sentence]
#### Our Positioning Against Them
> [How we differentiate - our one-liner against this competitor]
#### Where They Win
| Strength | Our Counter |
|----------|------------|
| [e.g., Lower price point] | [e.g., Emphasize TCO including implementation costs] |
| [e.g., Larger integration marketplace] | [e.g., Highlight quality over quantity, key integrations] |
| | |
#### Where We Win
| Our Strength | Evidence |
|-------------|----------|
| [e.g., Superior onboarding experience] | [Metric or customer quote] |
| [e.g., Better enterprise security] | [Certification or feature] |
| | |
#### Landmines to Set
Questions to ask prospects that expose competitor weaknesses:
1. "Have you evaluated how [specific capability] scales beyond [threshold]?"
2. "What's their approach to [area where competitor is weak]?"
3. "Can you share their uptime SLA and historical performance?"
#### Objection Handling
| Objection | Response |
|-----------|----------|
| "[Competitor] is cheaper" | [Value-based response] |
| "[Competitor] has more features" | [Quality/relevance response] |
| "We already use [Competitor]" | [Migration/coexistence story] |
#### Trap Questions They Set
Questions competitors ask about us, and how to respond:
1. **Q:** "[Our known weakness]?" **A:** [Honest, redirect response]
2. **Q:** "[Feature gap]?" **A:** [Roadmap or alternative approach]
#### Recent Intel
- [Date]: [Notable change - pricing, feature, hire, funding]
- [Date]: [Notable change]
FILE:references/competitive-analysis-frameworks.md
# Competitive Analysis Frameworks
This reference provides practical frameworks for evaluating competitors and positioning decisions.
## Porter's Five Forces
Assess the competitive intensity of your market:
1. Threat of new entrants
- Barriers to entry (capital, regulation, network effects)
- Speed of competitor replication
2. Bargaining power of suppliers
- Dependency on core infrastructure vendors
- Concentration of key technical providers
3. Bargaining power of buyers
- Customer switching costs
- Procurement complexity and contract leverage
4. Threat of substitutes
- Adjacent alternatives solving the same job
- DIY and internal build options
5. Rivalry among existing competitors
- Number and similarity of competitors
- Price competition and differentiation pressure
### Five Forces Template
| Force | Current Pressure (Low/Med/High) | Evidence | Strategic Response |
|---|---|---|---|
| New Entrants | | | |
| Supplier Power | | | |
| Buyer Power | | | |
| Substitutes | | | |
| Rivalry | | | |
## SWOT Analysis
Use SWOT to map internal and external context quickly.
### SWOT Template
| Strengths (Internal) | Weaknesses (Internal) |
|---|---|
| What we do better than alternatives | Where competitors outperform us |
| Unique capabilities or assets | Known product or go-to-market gaps |
| Opportunities (External) | Threats (External) |
|---|---|
| Market trends we can exploit | Competitor moves or macro risks |
| Unserved segments and use cases | Regulatory, platform, or pricing pressure |
### SWOT Quality Checklist
- Base every point on evidence, not assumptions.
- Separate observations from conclusions.
- Prioritize top 3 items per quadrant.
## Feature Comparison Matrix
Compare products on meaningful buying criteria, not vanity features.
### Feature Matrix Template
| Dimension | Weight | Your Product | Competitor A | Competitor B | Notes |
|---|---:|---:|---:|---:|---|
| Core workflow coverage | 25% | | | | |
| Ease of implementation | 15% | | | | |
| Performance / reliability | 15% | | | | |
| Integrations / ecosystem | 15% | | | | |
| Security / compliance | 15% | | | | |
| Pricing / TCO | 15% | | | | |
Scoring scale recommendation: 1-5 (weak to strong).
## Competitive Positioning Map
Create a 2-axis map showing market whitespace and crowding.
### Positioning Map Steps
1. Select two high-signal dimensions customers care about.
2. Place each competitor based on evidence (pricing pages, reviews, demos).
3. Mark clusters where products are undifferentiated.
4. Identify white space where demand exists but options are weak.
Example axes:
- X-axis: Ease of use
- Y-axis: Enterprise readiness
## Blue Ocean Strategy Canvas
Use a strategy canvas to decide where to raise, reduce, eliminate, or create factors.
### ERRC Grid (Eliminate-Reduce-Raise-Create)
| Eliminate | Reduce | Raise | Create |
|---|---|---|---|
| Commodity table-stakes not valued by target users | Costly features with weak adoption | Differentiators tied to target job-to-be-done | New value dimensions competitors ignore |
### Strategy Canvas Checklist
- Compare value curves between your product and top competitors.
- Ensure target segment is explicit.
- Tie every strategic choice to measurable outcome.
FILE:references/data-collection-guide.md
# Competitive Data Collection Guide
## Overview
This guide outlines systematic approaches for gathering competitive intelligence from publicly available sources. All methods described here are ethical and rely on information that competitors have made publicly accessible.
## Public Data Sources
### Review Platforms
- **G2**: Enterprise software reviews, feature comparisons, satisfaction scores
- **Capterra**: SMB-focused reviews, pricing transparency, deployment details
- **TrustRadius**: In-depth reviews with verified users, TrustMaps
- **Product Hunt**: Launch positioning, early adopter sentiment, feature highlights
- **App Store / Google Play**: Mobile app ratings, review themes, update frequency
### Company Publications
- **Pricing Pages**: Tier structure, feature gating, enterprise vs self-serve
- **Changelogs / Release Notes**: Development velocity, feature priorities, tech direction
- **Blog Posts**: Strategic messaging, thought leadership topics, market positioning
- **Case Studies**: Target customer profiles, value propositions, success metrics
- **Help Documentation**: Feature depth, API capabilities, integration ecosystem
### Talent & Organization Signals
- **Job Postings**: Technology stack, team growth areas, strategic initiatives
- **LinkedIn**: Team size, org structure, key hires, department ratios
- **Glassdoor**: Company culture, internal challenges, growth trajectory
### Financial & Legal
- **Patent Filings**: Innovation direction, defensive IP, technology differentiation
- **SEC Filings (public companies)**: Revenue, growth rate, customer count, churn
- **Crunchbase / PitchBook**: Funding rounds, investors, valuation trends
### Technical Intelligence
- **BuiltWith / Wappalyzer**: Technology stack detection
- **GitHub**: Open-source contributions, SDK quality, developer engagement
- **API Documentation**: Integration capabilities, rate limits, data models
- **Status Pages**: Uptime history, incident frequency, infrastructure maturity
## Data Points to Collect Per Competitor
### Product
- Core features and capabilities (feature-by-feature matrix)
- Unique differentiators and proprietary technology
- Platform support (web, mobile, desktop, API)
- Integration ecosystem (number and quality of integrations)
- Performance benchmarks (if available from reviews)
### Business
- Pricing tiers and per-seat/usage costs
- Target customer segments (SMB, mid-market, enterprise)
- Estimated customer count and notable logos
- Geographic focus and localization
- Go-to-market model (PLG, sales-led, hybrid)
### Team & Technology
- Estimated team size and engineering ratio
- Technology stack and infrastructure choices
- Development velocity (release frequency)
- Open-source involvement and developer relations
### Market Position
- Market share estimates
- Brand perception and NPS (from reviews)
- Analyst coverage (Gartner, Forrester positioning)
- Partnership and channel strategy
## Ethical Guidelines
1. **Use only public information** - Never access private systems, NDA-protected content, or internal documents
2. **No deception** - Do not misrepresent yourself to obtain information (e.g., fake sales inquiries)
3. **Respect terms of service** - Follow scraping policies and API usage terms
4. **Attribute sources** - Document where each data point came from for verification
5. **No employee poaching for intelligence** - Hiring decisions should be talent-driven, not intelligence-driven
6. **Legal compliance** - Ensure data collection complies with local regulations
## Update Cadence Recommendations
| Data Type | Frequency | Trigger Events |
|-----------|-----------|---------------|
| Pricing | Monthly | Competitor pricing page changes |
| Features | Bi-weekly | Changelog updates, product launches |
| Reviews | Monthly | Batch review analysis |
| Job Postings | Monthly | Hiring surge detection |
| Financials | Quarterly | Earnings reports, funding rounds |
| Tech Stack | Quarterly | Major platform changes |
| Full Teardown | Quarterly | Strategic planning cycles |
## Collection Workflow
1. **Set up monitoring** - Google Alerts, competitor RSS feeds, social listening
2. **Schedule regular sweeps** - Calendar recurring data collection tasks
3. **Centralize data** - Use a shared competitive intelligence database or spreadsheet
4. **Validate findings** - Cross-reference multiple sources for accuracy
5. **Tag and categorize** - Apply consistent taxonomy for easy retrieval
6. **Share insights** - Distribute relevant findings to product, sales, and marketing teams
7. **Archive versions** - Maintain historical snapshots for trend analysis
## Tools for Automation
- **Google Alerts**: Free monitoring for competitor mentions
- **Visualping**: Website change detection (pricing pages, feature pages)
- **Feedly**: RSS aggregation for competitor blogs and news
- **SimilarWeb**: Traffic estimates and audience overlap
- **SEMrush / Ahrefs**: SEO positioning and content strategy analysis
FILE:references/scoring-rubric.md
# Competitive Scoring Rubric
## Overview
This rubric provides a standardized framework for evaluating competitors across key dimensions. Consistent scoring enables meaningful comparisons and tracks competitive position changes over time.
## Scoring Scale (1-10)
| Score | Label | Definition |
|-------|-------|-----------|
| 1-2 | Poor | Significant gaps, major usability issues, or missing capability |
| 3-4 | Below Average | Basic functionality with notable limitations |
| 5-6 | Average | Meets market expectations, no standout qualities |
| 7-8 | Above Average | Strong execution with clear advantages |
| 9-10 | Exceptional | Industry-leading, sets the standard for others |
## Dimension Categories
### 1. User Experience (UX) - Weight: 20%
- **Onboarding**: Time to first value, setup complexity, guided flows
- **Navigation**: Information architecture, discoverability, consistency
- **Visual Design**: Modern aesthetics, brand coherence, accessibility
- **Performance**: Page load times, responsiveness, offline capability
- **Mobile Experience**: Native app quality, responsive design, feature parity
### 2. Feature Completeness - Weight: 25%
- **Core Features**: Coverage of essential use cases
- **Advanced Features**: Power user capabilities, automation, customization
- **Workflow Support**: End-to-end process coverage without workarounds
- **API & Extensibility**: API coverage, webhook support, SDK quality
- **Innovation**: Unique capabilities not found in competitors
### 3. Pricing & Value - Weight: 15%
- **Transparency**: Clear pricing without hidden costs
- **Flexibility**: Plan options matching different customer sizes
- **Value-to-Cost Ratio**: Feature access relative to price point
- **Free Tier / Trial**: Quality of free offering for evaluation
- **Contract Terms**: Lock-in requirements, cancellation ease
### 4. Integrations - Weight: 10%
- **Native Integrations**: Number and quality of built-in connectors
- **Marketplace**: Third-party app ecosystem breadth
- **API Quality**: Documentation, reliability, rate limits
- **Data Import/Export**: Migration ease, format support
- **Workflow Automation**: Zapier, Make, native automation support
### 5. Support & Documentation - Weight: 10%
- **Documentation Quality**: Completeness, searchability, freshness
- **Support Channels**: Chat, email, phone, community availability
- **Response Time**: SLA adherence, resolution speed
- **Self-Service**: Knowledge base, video tutorials, community forums
- **Onboarding Support**: Dedicated CSM, implementation assistance
### 6. Performance & Reliability - Weight: 10%
- **Uptime**: Historical availability, SLA commitments
- **Speed**: Application responsiveness under normal load
- **Scalability**: Performance at high volume, enterprise readiness
- **Data Handling**: Large dataset support, bulk operations
- **Global Performance**: CDN, regional deployments, latency
### 7. Security & Compliance - Weight: 10%
- **Authentication**: SSO, MFA, RBAC granularity
- **Data Protection**: Encryption at rest and in transit, data residency
- **Certifications**: SOC 2, ISO 27001, GDPR, HIPAA compliance
- **Audit Trail**: Activity logging, access monitoring
- **Privacy Controls**: Data retention policies, right to deletion
## Weighting Guidelines
Default weights above suit most B2B SaaS evaluations. Adjust based on:
- **Enterprise buyers**: Increase Security (15%), Support (15%), reduce Pricing (10%)
- **Developer tools**: Increase Integrations (20%), Features (30%), reduce UX (10%)
- **SMB products**: Increase Pricing (25%), UX (25%), reduce Security (5%)
- **Regulated industries**: Increase Security (25%), reduce Features (15%)
## Calibration Process
1. **Anchor scoring** - Score your own product first to establish baseline
2. **Multiple scorers** - Have 2-3 team members score independently
3. **Discuss outliers** - Reconcile scores that differ by more than 2 points
4. **Document evidence** - Record specific examples justifying each score
5. **Normalize quarterly** - Re-calibrate as market expectations evolve
## Bias Mitigation
- **Avoid halo effect** - Score each dimension independently, not influenced by overall impression
- **Use evidence, not feelings** - Every score must link to observable data points
- **Include competitor strengths** - Resist tendency to under-score competitors
- **Rotate scorers** - Different team members bring fresh perspectives
- **Blind scoring** - When possible, evaluate features without knowing which competitor
- **Customer validation** - Compare internal scores against user review sentiment
## Composite Score Calculation
```
Weighted Score = SUM(Dimension Score x Dimension Weight)
Example:
UX(8) x 0.20 = 1.60
Features(7) x 0.25 = 1.75
Pricing(6) x 0.15 = 0.90
Integrations(8) x 0.10 = 0.80
Support(7) x 0.10 = 0.70
Performance(9) x 0.10 = 0.90
Security(8) x 0.10 = 0.80
---
Total = 7.45 / 10
```
## Output Format
Present results as a comparison matrix with color coding:
- Green (8-10): Competitive advantage
- Yellow (5-7): Market parity
- Red (1-4): Competitive gap
FILE:scripts/competitive_matrix_builder.py
#!/usr/bin/env python3
"""Competitive Matrix Builder — Analyze and score competitors across feature dimensions.
Generates weighted competitive matrices, gap analysis, and positioning insights
from structured competitor data.
Usage:
python competitive_matrix_builder.py competitors.json --format json
python competitive_matrix_builder.py competitors.json --format text
python competitive_matrix_builder.py competitors.json --format text --weights pricing=2,ux=1.5
"""
import argparse
import json
import sys
from typing import Dict, List, Any, Optional
from datetime import datetime
from statistics import mean, stdev
def load_competitors(path: str) -> Dict[str, Any]:
"""Load competitor data from JSON file."""
with open(path, "r") as f:
return json.load(f)
def normalize_score(value: float, min_val: float = 1.0, max_val: float = 10.0) -> float:
"""Normalize a score to 0-100 scale."""
return max(0.0, min(100.0, ((value - min_val) / (max_val - min_val)) * 100))
def calculate_weighted_scores(
competitors: List[Dict[str, Any]],
dimensions: List[str],
weights: Optional[Dict[str, float]] = None
) -> List[Dict[str, Any]]:
"""Calculate weighted scores for each competitor across dimensions."""
if weights is None:
weights = {d: 1.0 for d in dimensions}
results = []
for comp in competitors:
scores = comp.get("scores", {})
weighted_total = 0.0
weight_sum = 0.0
dimension_results = {}
for dim in dimensions:
raw = scores.get(dim, 0)
w = weights.get(dim, 1.0)
normalized = normalize_score(raw)
weighted = normalized * w
weighted_total += weighted
weight_sum += w
dimension_results[dim] = {
"raw": raw,
"normalized": round(normalized, 1),
"weight": w,
"weighted": round(weighted, 1)
}
overall = round(weighted_total / weight_sum, 1) if weight_sum > 0 else 0
results.append({
"name": comp["name"],
"overall_score": overall,
"dimensions": dimension_results,
"tier": classify_tier(overall),
"pricing": comp.get("pricing", {}),
"strengths": comp.get("strengths", []),
"weaknesses": comp.get("weaknesses", [])
})
results.sort(key=lambda x: x["overall_score"], reverse=True)
return results
def classify_tier(score: float) -> str:
"""Classify competitor into tier based on overall score."""
if score >= 80:
return "Leader"
elif score >= 60:
return "Strong Competitor"
elif score >= 40:
return "Viable Alternative"
elif score >= 20:
return "Niche Player"
else:
return "Weak"
def gap_analysis(
your_scores: Dict[str, float],
competitor_scores: List[Dict[str, Any]],
dimensions: List[str]
) -> Dict[str, Any]:
"""Identify gaps between your product and competitors."""
gaps = {}
for dim in dimensions:
your_val = your_scores.get(dim, 0)
comp_vals = [c["dimensions"][dim]["raw"] for c in competitor_scores if dim in c.get("dimensions", {})]
if not comp_vals:
continue
avg_comp = mean(comp_vals)
best_comp = max(comp_vals)
gap_to_avg = round(your_val - avg_comp, 1)
gap_to_best = round(your_val - best_comp, 1)
gaps[dim] = {
"your_score": your_val,
"competitor_avg": round(avg_comp, 1),
"competitor_best": best_comp,
"gap_to_avg": gap_to_avg,
"gap_to_best": gap_to_best,
"status": "ahead" if gap_to_avg > 0.5 else ("behind" if gap_to_avg < -0.5 else "parity"),
"priority": "high" if gap_to_best < -2 else ("medium" if gap_to_best < -1 else "low")
}
return {
"gaps": gaps,
"biggest_opportunities": sorted(
[{"dimension": k, **v} for k, v in gaps.items() if v["status"] == "behind"],
key=lambda x: x["gap_to_best"]
)[:5],
"competitive_advantages": sorted(
[{"dimension": k, **v} for k, v in gaps.items() if v["status"] == "ahead"],
key=lambda x: -x["gap_to_avg"]
)[:5]
}
def positioning_analysis(scored: List[Dict[str, Any]]) -> Dict[str, Any]:
"""Generate positioning insights from scored competitors."""
scores = [c["overall_score"] for c in scored]
return {
"market_leaders": [c["name"] for c in scored if c["tier"] == "Leader"],
"your_rank": next((i + 1 for i, c in enumerate(scored) if c.get("is_you")), None),
"total_competitors": len(scored),
"score_distribution": {
"mean": round(mean(scores), 1) if scores else 0,
"stdev": round(stdev(scores), 1) if len(scores) > 1 else 0,
"min": round(min(scores), 1) if scores else 0,
"max": round(max(scores), 1) if scores else 0
},
"tier_distribution": {
tier: len([c for c in scored if c["tier"] == tier])
for tier in ["Leader", "Strong Competitor", "Viable Alternative", "Niche Player", "Weak"]
}
}
def format_text(result: Dict[str, Any]) -> str:
"""Format results as human-readable text."""
lines = []
lines.append("=" * 70)
lines.append("COMPETITIVE MATRIX ANALYSIS")
lines.append(f"Generated: {result['generated_at']}")
lines.append("=" * 70)
# Ranking table
lines.append("\n## COMPETITIVE RANKING\n")
lines.append(f"{'Rank':<6}{'Competitor':<25}{'Score':<10}{'Tier':<20}")
lines.append("-" * 61)
for i, c in enumerate(result["scored_competitors"], 1):
marker = " ← YOU" if c.get("is_you") else ""
lines.append(f"{i:<6}{c['name']:<25}{c['overall_score']:<10}{c['tier']:<20}{marker}")
# Dimension breakdown
lines.append("\n## DIMENSION BREAKDOWN\n")
dims = result["dimensions"]
header = f"{'Dimension':<20}" + "".join(f"{c['name'][:12]:<14}" for c in result["scored_competitors"])
lines.append(header)
lines.append("-" * len(header))
for dim in dims:
row = f"{dim:<20}"
for c in result["scored_competitors"]:
val = c["dimensions"].get(dim, {}).get("raw", "N/A")
row += f"{val:<14}"
lines.append(row)
# Gap analysis
if result.get("gap_analysis"):
ga = result["gap_analysis"]
if ga["biggest_opportunities"]:
lines.append("\n## BIGGEST OPPORTUNITIES (where you're behind)\n")
for opp in ga["biggest_opportunities"]:
lines.append(f" • {opp['dimension']}: You={opp['your_score']}, "
f"Best={opp['competitor_best']}, Gap={opp['gap_to_best']} "
f"[{opp['priority'].upper()} priority]")
if ga["competitive_advantages"]:
lines.append("\n## COMPETITIVE ADVANTAGES (where you lead)\n")
for adv in ga["competitive_advantages"]:
lines.append(f" • {adv['dimension']}: You={adv['your_score']}, "
f"Avg={adv['competitor_avg']}, Lead=+{adv['gap_to_avg']}")
# Positioning
pos = result.get("positioning", {})
if pos:
lines.append("\n## MARKET POSITIONING\n")
lines.append(f" Market Leaders: {', '.join(pos.get('market_leaders', ['None']))}")
if pos.get("your_rank"):
lines.append(f" Your Rank: #{pos['your_rank']} of {pos['total_competitors']}")
dist = pos.get("score_distribution", {})
lines.append(f" Score Range: {dist.get('min', 0)} - {dist.get('max', 0)} "
f"(avg: {dist.get('mean', 0)}, stdev: {dist.get('stdev', 0)})")
lines.append("\n" + "=" * 70)
return "\n".join(lines)
def build_matrix(data: Dict[str, Any], weight_overrides: Optional[Dict[str, float]] = None) -> Dict[str, Any]:
"""Main entry: build competitive matrix from input data."""
competitors = data.get("competitors", [])
dimensions = data.get("dimensions", [])
your_product = data.get("your_product", {})
if not competitors:
return {"error": "No competitors provided"}
if not dimensions:
# Auto-detect from first competitor's scores
dimensions = list(competitors[0].get("scores", {}).keys())
weights = data.get("weights", {})
if weight_overrides:
weights.update(weight_overrides)
# Include your product in scoring if provided
all_entries = list(competitors)
if your_product:
your_product["is_you"] = True
all_entries.insert(0, your_product)
scored = calculate_weighted_scores(all_entries, dimensions, weights)
# Mark your product
for s in scored:
if any(c.get("is_you") and c["name"] == s["name"] for c in all_entries):
s["is_you"] = True
result = {
"generated_at": datetime.now().isoformat(),
"dimensions": dimensions,
"weights": weights if weights else {d: 1.0 for d in dimensions},
"scored_competitors": scored,
"positioning": positioning_analysis(scored)
}
if your_product:
result["gap_analysis"] = gap_analysis(
your_product.get("scores", {}), scored, dimensions
)
return result
def parse_weights(weight_str: str) -> Dict[str, float]:
"""Parse weight string like 'pricing=2,ux=1.5' into dict."""
weights = {}
for pair in weight_str.split(","):
if "=" in pair:
k, v = pair.split("=", 1)
weights[k.strip()] = float(v.strip())
return weights
def main():
parser = argparse.ArgumentParser(
description="Build competitive matrix with scoring and gap analysis"
)
parser.add_argument("input", help="Path to competitors JSON file")
parser.add_argument("--format", choices=["json", "text"], default="text",
help="Output format (default: text)")
parser.add_argument("--weights", type=str, default=None,
help="Weight overrides: 'dim1=2.0,dim2=1.5'")
parser.add_argument("--output", type=str, default=None,
help="Output file path (default: stdout)")
args = parser.parse_args()
data = load_competitors(args.input)
weight_overrides = parse_weights(args.weights) if args.weights else None
result = build_matrix(data, weight_overrides)
if args.format == "json":
output = json.dumps(result, indent=2)
else:
output = format_text(result)
if args.output:
with open(args.output, "w") as f:
f.write(output)
print(f"Output written to {args.output}")
else:
print(output)
if __name__ == "__main__":
main()
Lấy đồng thuận từ nhiều mô hình AI (Claude, Codex, Gemini) khi rà soát bản ghi nhớ hội đồng hoặc tài liệu chiến lược.
---
name: "cross-eval"
description: "/cs:cross-eval <memo> — Multi-model consensus on a board memo or strategy brief. Claude + Codex + Gemini cross-review with graceful degradation."
---
# /cs:cross-eval — Multi-Model Consensus
**Command:** `/cs:cross-eval <memo-or-brief>`
Runs the same memo through multiple model providers and reconciles divergences. Use for **high-stakes, irreversible decisions** where single-model bias is too costly: M&A, major fundraises, layoffs, strategic pivots, regulatory commitments.
Adapted from gstack's `/codex` cross-review pattern, generalized to **business memos** instead of code PRs.
## When to Run
- Before signing a term sheet
- Before announcing a layoff
- Before committing to a regulated market
- Before any decision where reversing costs > 6 months of company time
- When the boardroom vote was split or had a CRITICAL dissent
## Models Used (graceful degradation)
The command tries to invoke each available model in order:
1. **Claude** (primary, always available) — the boardroom's native voice
2. **Codex / OpenAI** (if `OPENAI_API_KEY` or `codex` CLI available)
3. **Gemini** (if `GEMINI_API_KEY` or `gemini` CLI available)
If only Claude is available, the command runs **Claude-only with adversarial mode** — same model, different prompt seeds — and clearly labels the output as single-model.
## Workflow
1. Read the memo / brief
2. Probe environment for available model CLIs / API keys
3. For each available model:
- Send the memo with this prompt prefix:
> "You are an independent C-suite reviewer. The following is a board memo from another company's boardroom. Identify the top 3 concerns, the top 3 supports, and your vote (APPROVE / REJECT / DEFER). Do not deferentially agree — assume the memo's reasoning is flawed until proven otherwise."
4. Collect three independent reviews
5. Reconcile: where do they agree? Where do they diverge?
6. Surface the divergences as questions for the founder
## Output Format
Saved to `~/.claude/cross-eval/YYYY-MM-DD-<slug>.md`:
```markdown
# Cross-Eval: <memo title>
**Date:** YYYY-MM-DD
**Memo reviewed:** <link>
**Models invoked:** Claude / Codex / Gemini (or noted fallbacks)
## Vote Tally
| Model | Vote | Confidence |
|---|---|---|
| Claude | APPROVE | High |
| Codex | DEFER | Med |
| Gemini | APPROVE | Low |
## Consensus Concerns (≥2 models flagged)
1. <concern> — flagged by Claude + Codex
2. <concern> — flagged by all 3
## Divergent Concerns (1 model flagged)
- <Codex only:> <concern> — worth a second look
- <Gemini only:> <concern> — likely noise, but check
## Consensus Supports (≥2 models endorsed)
1. <support>
2. <support>
## Recommendation
- 🟢 GO if 2+ models APPROVE and no CRITICAL concerns from any model
- 🟡 PAUSE if any model is DEFER or any concern is CRITICAL
- 🔴 STOP if 2+ models REJECT
## Open Questions for Founder
1. <question raised by divergence>
2. <question raised by divergence>
```
## Why This Matters
Single-model recommendations have systematic biases. Claude trends helpful and may under-weight risk. Codex (OpenAI) trends more cautious on emerging-market and regulatory topics. Gemini trends more cautious on technical scale claims. Disagreement is signal, not noise.
This is the **safety net before irreversibility** — not a replacement for outside counsel or a real board.
## Graceful Degradation
If only Claude is available:
```markdown
**Models available:** Claude only
**Mode:** ADVERSARIAL — running 3 independent Claude passes with different system prompts:
1. Standard reviewer
2. Devil's advocate (must find 3 critical concerns)
3. Steelman (must find 3 strongest reasons to approve)
This is weaker than true multi-model. Treat the result as suggestive, not conclusive.
```
## Routing
- `/cs:decide` — if consensus is GO
- `/cs:freeze` — if consensus is PAUSE
- `/cs:boardroom` (re-run) — if consensus is STOP
## Related
- Skills: [`board-meeting`](../../../skills/board-meeting/SKILL.md), [`executive-mentor`](../../../executive-mentor/)
- Inspiration: gstack's `/codex` cross-review pattern (adapted to business memos)
---
**Version:** 1.0.0
Lên kế hoạch webinar từ mục tiêu kinh doanh, cứu webinar kém hiệu quả hoặc biến webinar cũ thành nguồn tạo lead theo yêu cầu.
---
name: "cs-webinar"
description: "/cs:webinar — Webinar & virtual-event marketing workflow. Plan a webinar from scratch (sized backward from the business goal), rescue one whose numbers disappointed (score the funnel, fix the broken stage), or turn a past webinar into an evergreen on-demand lead engine. Covers the full funnel: registration, promotion runway, show-up, live engagement, live-to-close, and segmented follow-up. Treats a webinar as a funnel, not an event."
---
# /cs:webinar — Webinar & Virtual Event Marketing
**Command:** `/cs:webinar [mode] [args]`
The `cs-webinar` command is the **entry point for webinar workflows**: plan → promote → run → follow up, or diagnose → fix → re-run.
## When To Run
- Planning a webinar, virtual event, live demo, workshop, masterclass, fireside chat, or virtual summit from scratch
- Rescuing a webinar whose numbers disappointed — low registrations, low show-up, or attendees who don't convert
- Turning a one-time webinar into an always-on evergreen / on-demand engine
- Scoring an existing funnel to find the stage that's actually broken
## When NOT To Run
- Full product launch (not just a webinar) → use `/cs:launch` / launch-strategy
- Generic lifecycle nurture email unrelated to an event → use the `emails` skill
- In-person field-event logistics (venue, catering, booth) → out of scope
## Modes
### `plan` — Design the whole motion from scratch
```bash
/cs:webinar plan
```
Walks the intake, locks the promise + format, sizes the funnel backward from the business goal,
builds the promotion runway, and designs show-up + live-to-close + follow-up. Delivers a full plan
using `templates/webinar-plan-template.md`.
### `rescue` — Diagnose and fix an underperforming webinar
```bash
/cs:webinar rescue --input funnel.json
```
Scores the funnel, names the weakest stage, and returns ranked fixes targeting the actual bottleneck
— not a reflexive landing-page rewrite.
### `evergreen` — Convert a past webinar to on-demand
```bash
/cs:webinar evergreen
```
Maps the on-demand registration → watch → follow-up automation, with honest live-vs-simulated framing.
### `score` — Run the funnel scorer directly
```bash
/cs:webinar score --input funnel.json
/cs:webinar score # embedded sample data
```
## Minimal Intake (3 Questions)
| Q | Asks | When |
|---|---|---|
| Q1 | Which mode — plan / rescue / evergreen? | Always |
| Q2 | Business goal + conversion action (leads, pipeline, adoption, retention, brand)? | Always (drives the backward funnel math) |
| Q3 | Audience temperature (customers / warm / owned_cold / paid_cold)? | Always (selects benchmarks) |
Read `marketing-context.md` first if it exists — it covers brand voice, personas, and customer language,
so you only ask for what's specific to this event.
## Workflow
```bash
# Mode: rescue / score — find the broken stage first
python3 marketing-skill/skills/webinar-marketing/scripts/webinar_funnel_scorer.py funnel.json
# → overall 0-100 score + per-stage rate vs. benchmark + named bottleneck
# Pipe JSON via stdin
cat funnel.json | python3 marketing-skill/skills/webinar-marketing/scripts/webinar_funnel_scorer.py -
# Demo on embedded sample data (no --help flag — run with no args)
python3 marketing-skill/skills/webinar-marketing/scripts/webinar_funnel_scorer.py
```
Input JSON (`registrations` + `attended_live` required; rest optional):
```json
{
"invited": 5000, "page_visits": 1800, "registrations": 620,
"attended_live": 180, "cta_clicks": 40, "conversions": 14,
"audience": "owned_cold", "runtime_min": 45, "avg_watch_min": 26
}
```
## The Funnel Math (Plan Backward)
Always size from the business goal backward — this stops anyone celebrating 800 registrations while 6 people buy:
```
Business goal: 20 sales-qualified opportunities
÷ attendee→SQO rate (~10%) → need 200 engaged attendees
÷ register→attend (~35% live) → need ~570 registrations
÷ landing-page CVR (~40%) → need ~1,425 landing-page visits
→ promotion must drive ~1,425 qualified visits
```
If the required visits exceed the reachable audience, fix the goal, format, or promotion budget *now*.
## Audience Benchmarks
The scorer calibrates per audience temperature (warmer audiences convert better at every stage):
| Audience | Page→Reg | Reg→Attend | Attend→CTA | Attend→Convert |
|---|---|---|---|---|
| `customers` | 40% | 50% | 25% | 12% |
| `warm` | 35% | 42% | 22% | 10% |
| `owned_cold` | 25% | 35% | 18% | 7% |
| `paid_cold` | 18% | 28% | 15% | 5% |
## Anti-Patterns Rejected
- Celebrating registrations while show-up or conversion quietly fails
- Rewriting the landing page when the broken stage is show-up or live-to-close
- Promoting before sizing the funnel backward from the business goal
- Obvious fake-live framing that erodes audience trust
- Treating a webinar as an event instead of a funnel
## Trigger Phrases
- "plan a webinar" / "webinar strategy"
- "my webinar isn't converting" / "low show-up rate"
- "webinar promotion" / "webinar follow-up"
- "virtual event" / "live demo" / "masterclass" / "fireside chat" / "virtual summit"
- "evergreen webinar" / "on-demand webinar"
- "registration funnel" / "attendance rate"
## Related
- Agent: [`cs-webinar-marketer`](../agents/marketing/cs-webinar-marketer.md)
- Skill: [`webinar-marketing`](../marketing-skill/skills/webinar-marketing/SKILL.md)
- Companion: `/cs:aeo` (get supporting content cited by AI search), launch-strategy (full launches)
---
**Version:** 2.9.0
**License:** MIT
Tạo video demo, hướng dẫn sản phẩm, giới thiệu tính năng hoặc GIF từ ảnh chụp màn hình hay mô tả cảnh bằng playwright, ffmpeg, edge-tts.
--- name: "demo-video" description: "Use when the user asks to create a demo video, product walkthrough, feature showcase, animated presentation, marketing video, or GIF from screenshots or scene descriptions. Orchestrates playwright, ffmpeg, and edge-tts MCPs to produce polished video content." --- # Demo Video You are a video producer. Not a slideshow maker. Every frame has a job. Every second earns the next. ## Overview Create polished demo videos by orchestrating browser rendering, text-to-speech, and video compositing. Think like a video producer — story arc, pacing, emotion, visual hierarchy. Turns screenshots and scene descriptions into shareable product demos. ## When to Use This Skill - User asks to create a demo video, product walkthrough, or feature showcase - User wants an animated presentation, marketing video, or product teaser - User wants to turn screenshots or UI captures into a polished video or GIF - User says "make a video", "create a demo", "record a demo", "promo video" ## Core Workflow ### 1. Choose a rendering mode Before starting, verify available tools: - **playwright MCP available?** — needed for automated screenshots. Fallback: ask user to screenshot the HTML files manually. - **edge-tts available?** — needed for narration audio. Fallback: output narration text files for user to record or use any TTS tool. - **ffmpeg available?** — needed for compositing. Fallback: output individual scene images + audio files with manual ffmpeg commands the user can run. If none are available, produce HTML scene files + `scenes.json` manifest + narration scripts. The user can composite manually or use any video editor. | Mode | How | When | |------|-----|------| | **MCP Orchestration** | HTML → playwright screenshots → edge-tts audio → ffmpeg composite | Use when playwright + edge-tts + ffmpeg MCPs are all connected | | **Manual** | Write HTML scene files, provide ffmpeg commands for user to run | Use when MCPs are not available | ### 2. Pick a story structure **The Classic Demo (30-60s):** Hook (3s) -> Problem (5s) -> Magic Moment (5s) -> Proof (15s) -> Social Proof (4s) -> Invite (4s) **The Problem-Solution (20-40s):** Before (6s) -> After (6s) -> How (10s) -> CTA (4s) **The 15-Second Teaser:** Hook (2s) -> Demo (8s) -> Logo (3s) -> Tagline (2s) ### 3. Design scenes **If no screenshots are provided:** - For CLI/terminal tools: generate HTML scenes with terminal-style dark background, monospace font, and animated typing effect - For conceptual demos: use text-heavy scenes with the color language and typography system - Ask the user for screenshots only if the product is visual and descriptions are insufficient Every scene has exactly ONE primary focus: - Title scenes: product name - Problem scenes: the pain (red, chaotic) - Solution scenes: the result (green, spacious) - Feature scenes: the highlighted screenshot region - End scenes: URL / CTA button ### 4. Write narration - One idea per scene. If you need "and" you need two scenes. - Lead with the verb. "Organize your tabs" not "Tab organization is provided." - No jargon. "Your tabs organize themselves" not "AI-powered tab categorization." - Use contrast. "24 tabs. One click. 5 groups." ## Output Artifacts For each video, produce these files in a `demo-output/` directory: 1. `scenes/` — one HTML file per scene (1920x1080 viewport) 2. `narration/` — one `.txt` file per scene (for edge-tts input) 3. `scenes.json` — manifest listing scenes in order with durations and narration text 4. `build.sh` — shell script that runs the full pipeline: - `playwright screenshot` each HTML scene → `frames/` - `edge-tts` each narration file → `audio/` - `ffmpeg` concat with crossfade transitions → `output.mp4` If MCPs are unavailable, still produce items 1-3. Include the ffmpeg commands in `build.sh` for the user to run manually. ## Scene Design System See [references/scene-design-system.md](references/scene-design-system.md) for the full design system: color language, animation timing, typography, HTML layout, voice options, and pacing guide. ## Quality Checklist - [ ] Video has audio stream - [ ] Resolution is 1920x1080 - [ ] No black frames between scenes - [ ] First 3 seconds grab attention - [ ] Every scene has one focus point - [ ] End card has URL and CTA ## Anti-Patterns | Anti-pattern | Fix | |---|---| | **Slideshow pacing** — every scene same duration, no rhythm | Vary durations: hooks 3s, proof 8s, CTA 4s | | **Wall of text on screen** | Move info to narration, simplify visuals | | **Generic narration** — "This feature lets you..." | Use specific numbers and concrete verbs | | **No story arc** — just listing features | Use problem -> solution -> proof structure | | **Raw screenshots** | Always add rounded corners, shadows, dark background | | **Using `ease` or `linear` animations** | Use spring curve: `cubic-bezier(0.16, 1, 0.3, 1)` | ## Cross-References - Related: `engineering/browser-automation` — for playwright-based browser workflows - See also: [framecraft](https://github.com/vaddisrinivas/framecraft) — open-source scene rendering pipeline FILE:references/scene-design-system.md # Scene Design System Reference material for demo video scene design — colors, typography, animation timing, voice options, and pacing. ## Color Language | Color | Meaning | Use for | |-------|---------|---------| | `#c5d5ff` | Trust | Titles, logo | | `#7c6af5` | Premium | Subtitles, badges | | `#4ade80` | Success | "After" states | | `#f28b82` | Problem | "Before" states | | `#fbbf24` | Energy | Callouts | | `#0d0e12` | Background | Always dark mode | ## Animation Timing ``` Element entrance: 0.5-0.8s (cubic-bezier(0.16, 1, 0.3, 1)) Between elements: 0.2-0.4s gap Scene transition: 0.3-0.5s crossfade Hold after last anim: 1.0-2.0s ``` ## Typography ``` Title: 48-72px, weight 800 Subtitle: 24-32px, weight 400, muted Bullets: 18-22px, weight 600, pill background Font: Inter (Google Fonts) ``` ## HTML Scene Layout (1920x1080) ```html <body> <h1 class="title">...</h1> <!-- Top 15% --> <div class="hero">...</div> <!-- Middle 65% --> <div class="footer">...</div> <!-- Bottom 20% --> </body> ``` Background: dark with subtle purple-blue glow gradients. Screenshots: always `border-radius: 12px` with `box-shadow`. Easing: always `cubic-bezier(0.16, 1, 0.3, 1)` — never `ease` or `linear`. ## Voice Options (edge-tts) | Voice | Best for | |-------|----------| | `andrew` | Product demos, launches | | `jenny` | Tutorials, onboarding | | `davis` | Enterprise, security | | `emma` | Consumer products | ## Pacing Guide | Duration | Max words | Fill | |----------|-----------|------| | 3-4s | 8-12 | ~70% | | 5-6s | 15-22 | ~75% | | 7-8s | 22-30 | ~80% |
Xây hệ thống email giao dịch: mẫu React Email, tích hợp Resend/Postmark/SendGrid/SES, xem trước, đa ngôn ngữ, chế độ tối, chống spam.
---
name: "email-template-builder"
description: "Build complete transactional email systems: React Email templates, provider integration (Resend, Postmark, SendGrid, AWS SES), preview server, i18n support, dark mode, spam optimization, analytics tracking. Use when adding transactional email to a new product, migrating between email providers, refactoring legacy email templates for accessibility, or adding internationalization to existing templates."
---
# Email Template Builder
**Tier:** POWERFUL
**Category:** Engineering Team
**Domain:** Transactional Email / Communications Infrastructure
---
## Overview
Build complete transactional email systems: React Email templates, provider integration, preview server, i18n support, dark mode, spam optimization, and analytics tracking. Output production-ready code for Resend, Postmark, SendGrid, or AWS SES.
---
## Core Capabilities
- React Email templates (welcome, verification, password reset, invoice, notification, digest)
- MJML templates for maximum email client compatibility
- Multi-provider support with unified sending interface
- Local preview server with hot reload
- i18n/localization with typed translation keys
- Dark mode support using media queries
- Spam score optimization checklist
- Open/click tracking with UTM parameters
---
## When to Use
- Setting up transactional email for a new product
- Migrating from a legacy email system
- Adding new email types (invoice, digest, notification)
- Debugging email deliverability issues
- Implementing i18n for email templates
---
## Project Structure
```
emails/
├── components/
│ ├── layout/
│ │ ├── email-layout.tsx # Base layout with brand header/footer
│ │ └── email-button.tsx # CTA button component
│ ├── partials/
│ │ ├── header.tsx
│ │ └── footer.tsx
├── templates/
│ ├── welcome.tsx
│ ├── verify-email.tsx
│ ├── password-reset.tsx
│ ├── invoice.tsx
│ ├── notification.tsx
│ └── weekly-digest.tsx
├── lib/
│ ├── send.ts # Unified send function
│ ├── providers/
│ │ ├── resend.ts
│ │ ├── postmark.ts
│ │ └── ses.ts
│ └── tracking.ts # UTM + analytics
├── i18n/
│ ├── en.ts
│ └── de.ts
└── preview/ # Dev preview server
└── server.ts
```
---
## Base Email Layout
```tsx
// emails/components/layout/email-layout.tsx
import {
Body, Container, Head, Html, Img, Preview, Section, Text, Hr, Font
} from "@react-email/components"
interface EmailLayoutProps {
preview: string
children: React.ReactNode
}
export function EmailLayout({ preview, children }: EmailLayoutProps) {
return (
<Html lang="en">
<Head>
<Font
fontFamily="Inter"
fallbackFontFamily="Arial"
webFont={{ url: "https://fonts.gstatic.com/s/inter/v13/UcCO3FwrK3iLTeHuS_nVMrMxCp50SjIw2boKoduKmMEVuLyfAZ9hiJ-Ek-_EeA.woff2", format: "woff2" }}
fontWeight={400}
fontStyle="normal"
/>
{/* Dark mode styles */}
<style>{`
@media (prefers-color-scheme: dark) {
.email-body { background-color: #0f0f0f !important; }
.email-container { background-color: #1a1a1a !important; }
.email-text { color: #e5e5e5 !important; }
.email-heading { color: #ffffff !important; }
.email-divider { border-color: #333333 !important; }
}
`}</style>
</Head>
<Preview>{preview}</Preview>
<Body className="email-body" style={styles.body}>
<Container className="email-container" style={styles.container}>
{/* Header */}
<Section style={styles.header}>
<Img src="https://yourapp.com/logo.png" width={120} height={40} alt="MyApp" />
</Section>
{/* Content */}
<Section style={styles.content}>
{children}
</Section>
{/* Footer */}
<Hr style={styles.divider} />
<Section style={styles.footer}>
<Text style={styles.footerText}>
MyApp Inc. · 123 Main St · San Francisco, CA 94105
</Text>
<Text style={styles.footerText}>
<a href="{{unsubscribe_url}}" style={styles.link}>Unsubscribe</a>
{" · "}
<a href="https://yourapp.com/privacy" style={styles.link}>Privacy Policy</a>
</Text>
</Section>
</Container>
</Body>
</Html>
)
}
const styles = {
body: { backgroundColor: "#f5f5f5", fontFamily: "Inter, Arial, sans-serif" },
container: { maxWidth: "600px", margin: "0 auto", backgroundColor: "#ffffff", borderRadius: "8px", overflow: "hidden" },
header: { padding: "24px 32px", borderBottom: "1px solid #e5e5e5" },
content: { padding: "32px" },
divider: { borderColor: "#e5e5e5", margin: "0 32px" },
footer: { padding: "24px 32px" },
footerText: { fontSize: "12px", color: "#6b7280", textAlign: "center" as const, margin: "4px 0" },
link: { color: "#6b7280", textDecoration: "underline" },
}
```
---
## Welcome Email
```tsx
// emails/templates/welcome.tsx
import { Button, Heading, Text } from "@react-email/components"
import { EmailLayout } from "../components/layout/email-layout"
interface WelcomeEmailProps {
name: "string"
confirmUrl: string
trialDays?: number
}
export function WelcomeEmail({ name, confirmUrl, trialDays = 14 }: WelcomeEmailProps) {
return (
<EmailLayout preview={`Welcome to MyApp, name! Confirm your email to get started.`}>
<Heading style={styles.h1}>Welcome to MyApp, {name}!</Heading>
<Text style={styles.text}>
We're excited to have you on board. You've got {trialDays} days to explore everything MyApp has to offer — no credit card required.
</Text>
<Text style={styles.text}>
First, confirm your email address to activate your account:
</Text>
<Button href={confirmUrl} style={styles.button}>
Confirm Email Address
</Button>
<Text style={styles.hint}>
Button not working? Copy and paste this link into your browser:
<br />
<a href={confirmUrl} style={styles.link}>{confirmUrl}</a>
</Text>
<Text style={styles.text}>
Once confirmed, you can:
</Text>
<ul style={styles.list}>
<li>Connect your first project in 2 minutes</li>
<li>Invite your team (free for up to 3 members)</li>
<li>Set up Slack notifications</li>
</ul>
</EmailLayout>
)
}
export default WelcomeEmail
const styles = {
h1: { fontSize: "28px", fontWeight: "700", color: "#111827", margin: "0 0 16px" },
text: { fontSize: "16px", lineHeight: "1.6", color: "#374151", margin: "0 0 16px" },
button: { backgroundColor: "#4f46e5", color: "#ffffff", borderRadius: "6px", fontSize: "16px", fontWeight: "600", padding: "12px 24px", textDecoration: "none", display: "inline-block", margin: "8px 0 24px" },
hint: { fontSize: "13px", color: "#6b7280" },
link: { color: "#4f46e5" },
list: { fontSize: "16px", lineHeight: "1.8", color: "#374151", paddingLeft: "20px" },
}
```
---
## Invoice Email
```tsx
// emails/templates/invoice.tsx
import { Row, Column, Section, Heading, Text, Hr, Button } from "@react-email/components"
import { EmailLayout } from "../components/layout/email-layout"
interface InvoiceItem { description: string; amount: number }
interface InvoiceEmailProps {
name: "string"
invoiceNumber: string
invoiceDate: string
dueDate: string
items: InvoiceItem[]
total: number
currency: string
downloadUrl: string
}
export function InvoiceEmail({ name, invoiceNumber, invoiceDate, dueDate, items, total, currency = "USD", downloadUrl }: InvoiceEmailProps) {
const formatter = new Intl.NumberFormat("en-US", { style: "currency", currency })
return (
<EmailLayout preview={`Invoice invoiceNumber - formatter.format(total / 100)`}>
<Heading style={styles.h1}>Invoice #{invoiceNumber}</Heading>
<Text style={styles.text}>Hi {name},</Text>
<Text style={styles.text}>Here's your invoice from MyApp. Thank you for your continued support.</Text>
{/* Invoice Meta */}
<Section style={styles.metaBox}>
<Row>
<Column><Text style={styles.metaLabel}>Invoice Date</Text><Text style={styles.metaValue}>{invoiceDate}</Text></Column>
<Column><Text style={styles.metaLabel}>Due Date</Text><Text style={styles.metaValue}>{dueDate}</Text></Column>
<Column><Text style={styles.metaLabel}>Amount Due</Text><Text style={styles.metaValueLarge}>{formatter.format(total / 100)}</Text></Column>
</Row>
</Section>
{/* Line Items */}
<Section style={styles.table}>
<Row style={styles.tableHeader}>
<Column><Text style={styles.tableHeaderText}>Description</Text></Column>
<Column><Text style={{ ...styles.tableHeaderText, textAlign: "right" }}>Amount</Text></Column>
</Row>
{items.map((item, i) => (
<Row key={i} style={i % 2 === 0 ? styles.tableRowEven : styles.tableRowOdd}>
<Column><Text style={styles.tableCell}>{item.description}</Text></Column>
<Column><Text style={{ ...styles.tableCell, textAlign: "right" }}>{formatter.format(item.amount / 100)}</Text></Column>
</Row>
))}
<Hr style={styles.divider} />
<Row>
<Column><Text style={styles.totalLabel}>Total</Text></Column>
<Column><Text style={styles.totalValue}>{formatter.format(total / 100)}</Text></Column>
</Row>
</Section>
<Button href={downloadUrl} style={styles.button}>Download PDF Invoice</Button>
</EmailLayout>
)
}
export default InvoiceEmail
const styles = {
h1: { fontSize: "24px", fontWeight: "700", color: "#111827", margin: "0 0 16px" },
text: { fontSize: "15px", lineHeight: "1.6", color: "#374151", margin: "0 0 12px" },
metaBox: { backgroundColor: "#f9fafb", borderRadius: "8px", padding: "16px", margin: "16px 0" },
metaLabel: { fontSize: "12px", color: "#6b7280", fontWeight: "600", textTransform: "uppercase" as const, margin: "0 0 4px" },
metaValue: { fontSize: "14px", color: "#111827", margin: 0 },
metaValueLarge: { fontSize: "20px", fontWeight: "700", color: "#4f46e5", margin: 0 },
table: { width: "100%", margin: "16px 0" },
tableHeader: { backgroundColor: "#f3f4f6", borderRadius: "4px" },
tableHeaderText: { fontSize: "12px", fontWeight: "600", color: "#374151", padding: "8px 12px", textTransform: "uppercase" as const },
tableRowEven: { backgroundColor: "#ffffff" },
tableRowOdd: { backgroundColor: "#f9fafb" },
tableCell: { fontSize: "14px", color: "#374151", padding: "10px 12px" },
divider: { borderColor: "#e5e5e5", margin: "8px 0" },
totalLabel: { fontSize: "16px", fontWeight: "700", color: "#111827", padding: "8px 12px" },
totalValue: { fontSize: "16px", fontWeight: "700", color: "#111827", textAlign: "right" as const, padding: "8px 12px" },
button: { backgroundColor: "#4f46e5", color: "#fff", borderRadius: "6px", padding: "12px 24px", fontSize: "15px", fontWeight: "600", textDecoration: "none" },
}
```
---
## Unified Send Function
```typescript
// emails/lib/send.ts
import { Resend } from "resend"
import { render } from "@react-email/render"
import { WelcomeEmail } from "../templates/welcome"
import { InvoiceEmail } from "../templates/invoice"
import { addTrackingParams } from "./tracking"
const resend = new Resend(process.env.RESEND_API_KEY)
type EmailPayload =
| { type: "welcome"; props: Parameters<typeof WelcomeEmail>[0] }
| { type: "invoice"; props: Parameters<typeof InvoiceEmail>[0] }
export async function sendEmail(to: string, payload: EmailPayload) {
const templates = {
welcome: { component: WelcomeEmail, subject: "Welcome to MyApp — confirm your email" },
invoice: { component: InvoiceEmail, subject: `Invoice from MyApp` },
}
const template = templates[payload.type]
const html = render(template.component(payload.props as any))
const trackedHtml = addTrackingParams(html, { campaign: payload.type })
const result = await resend.emails.send({
from: "MyApp <hello@yourapp.com>",
to,
subject: template.subject,
html: trackedHtml,
tags: [{ name: "email-type", value: payload.type }],
})
return result
}
```
---
## Preview Server Setup
```typescript
// package.json scripts
{
"scripts": {
"email:dev": "email dev --dir emails/templates --port 3001",
"email:build": "email export --dir emails/templates --outDir emails/out"
}
}
// Run: npm run email:dev
// Opens: http://localhost:3001
// Shows all templates with live preview and hot reload
```
---
## i18n Support
```typescript
// emails/i18n/en.ts
export const en = {
welcome: {
preview: (name: "string-welcome-to-myapp-name"
heading: (name: "string-welcome-to-myapp-name"
body: (days: number) => `You've got days days to explore everything.`,
cta: "Confirm Email Address",
},
}
// emails/i18n/de.ts
export const de = {
welcome: {
preview: (name: "string-willkommen-bei-myapp-name"
heading: (name: "string-willkommen-bei-myapp-name"
body: (days: number) => `Du hast days Tage Zeit, alles zu erkunden.`,
cta: "E-Mail-Adresse bestätigen",
},
}
// Usage in template
import { en, de } from "../i18n"
const t = locale === "de" ? de : en
```
---
## Spam Score Optimization Checklist
- [ ] Sender domain has SPF, DKIM, and DMARC records configured
- [ ] From address uses your own domain (not gmail.com/hotmail.com)
- [ ] Subject line under 50 characters, no ALL CAPS, no "FREE!!!"
- [ ] Text-to-image ratio: at least 60% text
- [ ] Plain text version included alongside HTML
- [ ] Unsubscribe link in every marketing email (CAN-SPAM, GDPR)
- [ ] No URL shorteners — use full branded links
- [ ] No red-flag words: "guarantee", "no risk", "limited time offer" in subject
- [ ] Single CTA per email — no 5 different buttons
- [ ] Image alt text on every image
- [ ] HTML validates — no broken tags
- [ ] Test with Mail-Tester.com before first send (target: 9+/10)
---
## Analytics Tracking
```typescript
// emails/lib/tracking.ts
interface TrackingParams {
campaign: string
medium?: string
source?: string
}
export function addTrackingParams(html: string, params: TrackingParams): string {
const utmString = new URLSearchParams({
utm_source: params.source ?? "email",
utm_medium: params.medium ?? "transactional",
utm_campaign: params.campaign,
}).toString()
// Add UTM params to all links in the email
return html.replace(/href="(https?:\/\/[^"]+)"/g, (match, url) => {
const separator = url.includes("?") ? "&" : "?"
return `href="urlseparatorutmString"`
})
}
```
---
## Common Pitfalls
- **Inline styles required** — most email clients strip `<head>` styles; React Email handles this
- **Max width 600px** — anything wider breaks on Gmail mobile
- **No flexbox/grid** — use `<Row>` and `<Column>` from react-email, not CSS grid
- **Dark mode media queries** — must use `!important` to override inline styles
- **Missing plain text** — all major providers have a plain text field; always populate it
- **Transactional vs marketing** — use separate sending domains/IPs to protect deliverability
Biến mẫu đã kiểm chứng hoặc cách gỡ lỗi thành skill độc lập có thể tái sử dụng, gồm SKILL.md, tài liệu tham khảo và ví dụ.
---
name: "extract"
description: "Turn a proven pattern or debugging solution into a standalone reusable skill with SKILL.md, reference docs, and examples."
---
# /si:extract — Create Skills from Patterns
Transforms a recurring pattern or debugging solution into a standalone, portable skill that can be installed in any project.
## Usage
```
/si:extract <pattern description> # Interactive extraction
/si:extract <pattern> --name docker-m1-fixes # Specify skill name
/si:extract <pattern> --output ./skills/ # Custom output directory
/si:extract <pattern> --dry-run # Preview without creating files
```
## When to Extract
A learning qualifies for skill extraction when ANY of these are true:
| Criterion | Signal |
|---|---|
| **Recurring** | Same issue across 2+ projects |
| **Non-obvious** | Required real debugging to discover |
| **Broadly applicable** | Not tied to one specific codebase |
| **Complex solution** | Multi-step fix that's easy to forget |
| **User-flagged** | "Save this as a skill", "I want to reuse this" |
## Workflow
### Step 1: Identify the pattern
Read the user's description. Search auto-memory for related entries:
```bash
MEMORY_DIR="$HOME/.claude/projects/$(pwd | sed 's|/|%2F|g; s|%2F|/|; s|^/||')/memory"
grep -rni "<keywords>" "$MEMORY_DIR/"
```
If found in auto-memory, use those entries as source material. If not, use the user's description directly.
### Step 2: Determine skill scope
Ask (max 2 questions):
- "What problem does this solve?" (if not clear)
- "Should this include code examples?" (if applicable)
### Step 3: Generate skill name
Rules for naming:
- Lowercase, hyphens between words
- Descriptive but concise (2-4 words)
- Examples: `docker-m1-fixes`, `api-timeout-patterns`, `pnpm-workspace-setup`
**Reserved fragments — must NOT appear in the skill name:**
- `claude`
- `anthropic`
For skills about Claude Code itself, use the `cc-` prefix instead:
- ❌ `claude-code-settings` → ✅ `cc-settings`
- ❌ `claude-code-maintenance` → ✅ `cc-maintenance`
- ❌ `claude-mcp-tools` → ✅ `cc-mcp-tools`
- ❌ `claude-plugin-development` → ✅ `cc-plugin-development`
Before writing the skill directory, check the proposed name against this list.
If a reserved fragment is present, transform it (drop the fragment or replace
the `claude*`/`anthropic*` prefix with `cc-`) and confirm with the user.
### Step 4: Create the skill files
**Spawn the `skill-extractor` agent** for the actual file generation.
The agent creates:
```
<skill-name>/
├── SKILL.md # Main skill file with frontmatter
├── README.md # Human-readable overview
└── reference/ # (optional) Supporting documentation
└── examples.md # Concrete examples and edge cases
```
### Step 5: SKILL.md structure
The generated SKILL.md must follow this format:
```markdown
---
name: "skill-name"
description: "<one-line description>. Use when: <trigger conditions>."
---
# <Skill Title>
> One-line summary of what this skill solves.
## Quick Reference
| Problem | Solution |
|---------|----------|
| {{problem 1}} | {{solution 1}} |
| {{problem 2}} | {{solution 2}} |
## The Problem
{{2-3 sentences explaining what goes wrong and why it's non-obvious.}}
## Solutions
### Option 1: {{Name}} (Recommended)
{{Step-by-step with code examples.}}
### Option 2: {{Alternative}}
{{For when Option 1 doesn't apply.}}
## Trade-offs
| Approach | Pros | Cons |
|----------|------|------|
| Option 1 | {{pros}} | {{cons}} |
| Option 2 | {{pros}} | {{cons}} |
## Edge Cases
- {{edge case 1 and how to handle it}}
- {{edge case 2 and how to handle it}}
```
### Step 6: Quality gates
Before finalizing, verify:
- [ ] SKILL.md has valid YAML frontmatter with `name` and `description`
- [ ] `name` matches the folder name (lowercase, hyphens)
- [ ] `name` does NOT contain reserved fragments `claude` or `anthropic` (use `cc-` prefix for Claude Code skills)
- [ ] Description includes "Use when:" trigger conditions
- [ ] Solutions are self-contained (no external context needed)
- [ ] Code examples are complete and copy-pasteable
- [ ] No project-specific hardcoded values (paths, URLs, credentials)
- [ ] No unnecessary dependencies
### Step 7: Report
```
✅ Skill extracted: {{skill-name}}
Files created:
{{path}}/SKILL.md ({{lines}} lines)
{{path}}/README.md ({{lines}} lines)
{{path}}/reference/examples.md ({{lines}} lines)
Install: /plugin install (copy to your skills directory)
Publish: clawhub publish {{path}}
Source: MEMORY.md entries at lines {{n, m, ...}} (retained — the skill is portable, the memory is project-specific)
```
## Examples
### Extracting a debugging pattern
```
/si:extract "Fix for Docker builds failing on Apple Silicon with platform mismatch"
```
Creates `docker-m1-fixes/SKILL.md` with:
- The platform mismatch error message
- Three solutions (build flag, Dockerfile, docker-compose)
- Trade-offs table
- Performance note about Rosetta 2 emulation
### Extracting a workflow pattern
```
/si:extract "Always regenerate TypeScript API client after modifying OpenAPI spec"
```
Creates `api-client-regen/SKILL.md` with:
- Why manual regen is needed
- The exact command sequence
- CI integration snippet
- Common failure modes
## Tips
- Extract patterns that would save time in a *different* project
- Keep skills focused — one problem per skill
- Include the error messages people would search for
- Test the skill by reading it without the original context — does it make sense?
Tư vấn FDA cho công ty thiết bị y tế: lộ trình 510(k)/PMA/De Novo, tuân thủ QSR (21 CFR 820), HIPAA và an ninh mạng thiết bị.
---
name: "fda-consultant-specialist"
description: FDA regulatory consultant for medical device companies. Provides 510(k)/PMA/De Novo pathway guidance, QSR (21 CFR 820) compliance, HIPAA assessments, and device cybersecurity. Use when user mentions FDA submission, 510(k), PMA, De Novo, QSR, premarket, predicate device, substantial equivalence, HIPAA medical device, or FDA cybersecurity.
---
# FDA Consultant Specialist
FDA regulatory consulting for medical device manufacturers covering submission pathways, Quality System Regulation (QSR), HIPAA compliance, and device cybersecurity requirements.
## Table of Contents
- [FDA Pathway Selection](#fda-pathway-selection)
- [510(k) Submission Process](#510k-submission-process)
- [QSR Compliance](#qsr-compliance)
- [HIPAA for Medical Devices](#hipaa-for-medical-devices)
- [Device Cybersecurity](#device-cybersecurity)
- [Resources](#resources)
---
## FDA Pathway Selection
Determine the appropriate FDA regulatory pathway based on device classification and predicate availability.
### Decision Framework
```
Predicate device exists?
├── YES → Substantially equivalent?
│ ├── YES → 510(k) Pathway
│ │ ├── No design changes → Abbreviated 510(k)
│ │ ├── Manufacturing only → Special 510(k)
│ │ └── Design/performance → Traditional 510(k)
│ └── NO → PMA or De Novo
└── NO → Novel device?
├── Low-to-moderate risk → De Novo
└── High risk (Class III) → PMA
```
### Pathway Comparison
| Pathway | When to Use | Timeline | Cost |
|---------|-------------|----------|------|
| 510(k) Traditional | Predicate exists, design changes | 90 days | $21,760 |
| 510(k) Special | Manufacturing changes only | 30 days | $21,760 |
| 510(k) Abbreviated | Guidance/standard conformance | 30 days | $21,760 |
| De Novo | Novel, low-moderate risk | 150 days | $134,676 |
| PMA | Class III, no predicate | 180+ days | $425,000+ |
### Pre-Submission Strategy
1. Identify product code and classification
2. Search 510(k) database for predicates
3. Assess substantial equivalence feasibility
4. Prepare Q-Sub questions for FDA
5. Schedule Pre-Sub meeting if needed
**Reference:** See [fda_submission_guide.md](references/fda_submission_guide.md) for pathway decision matrices and submission requirements.
---
## 510(k) Submission Process
### Workflow
```
Phase 1: Planning
├── Step 1: Identify predicate device(s)
├── Step 2: Compare intended use and technology
├── Step 3: Determine testing requirements
└── Checkpoint: SE argument feasible?
Phase 2: Preparation
├── Step 4: Complete performance testing
├── Step 5: Prepare device description
├── Step 6: Document SE comparison
├── Step 7: Finalize labeling
└── Checkpoint: All required sections complete?
Phase 3: Submission
├── Step 8: Assemble submission package
├── Step 9: Submit via eSTAR
├── Step 10: Track acknowledgment
└── Checkpoint: Submission accepted?
Phase 4: Review
├── Step 11: Monitor review status
├── Step 12: Respond to AI requests
├── Step 13: Receive decision
└── Verification: SE letter received?
```
### Required Sections (21 CFR 807.87)
| Section | Content |
|---------|---------|
| Cover Letter | Submission type, device ID, contact info |
| Form 3514 | CDRH premarket review cover sheet |
| Device Description | Physical description, principles of operation |
| Indications for Use | Form 3881, patient population, use environment |
| SE Comparison | Side-by-side comparison with predicate |
| Performance Testing | Bench, biocompatibility, electrical safety |
| Software Documentation | Level of concern, hazard analysis (IEC 62304) |
| Labeling | IFU, package labels, warnings |
| 510(k) Summary | Public summary of submission |
### Common RTA Issues
| Issue | Prevention |
|-------|------------|
| Missing user fee | Verify payment before submission |
| Incomplete Form 3514 | Review all fields, ensure signature |
| No predicate identified | Confirm K-number in FDA database |
| Inadequate SE comparison | Address all technological characteristics |
---
## QSR Compliance
Quality System Regulation (21 CFR Part 820) requirements for medical device manufacturers.
### Key Subsystems
| Section | Title | Focus |
|---------|-------|-------|
| 820.20 | Management Responsibility | Quality policy, org structure, management review |
| 820.30 | Design Controls | Input, output, review, verification, validation |
| 820.40 | Document Controls | Approval, distribution, change control |
| 820.50 | Purchasing Controls | Supplier qualification, purchasing data |
| 820.70 | Production Controls | Process validation, environmental controls |
| 820.100 | CAPA | Root cause analysis, corrective actions |
| 820.181 | Device Master Record | Specifications, procedures, acceptance criteria |
### Design Controls Workflow (820.30)
```
Step 1: Design Input
└── Capture user needs, intended use, regulatory requirements
Verification: Inputs reviewed and approved?
Step 2: Design Output
└── Create specifications, drawings, software architecture
Verification: Outputs traceable to inputs?
Step 3: Design Review
└── Conduct reviews at each phase milestone
Verification: Review records with signatures?
Step 4: Design Verification
└── Perform testing against specifications
Verification: All tests pass acceptance criteria?
Step 5: Design Validation
└── Confirm device meets user needs in actual use conditions
Verification: Validation report approved?
Step 6: Design Transfer
└── Release to production with DMR complete
Verification: Transfer checklist complete?
```
### CAPA Process (820.100)
1. **Identify**: Document nonconformity or potential problem
2. **Investigate**: Perform root cause analysis (5 Whys, Fishbone)
3. **Plan**: Define corrective/preventive actions
4. **Implement**: Execute actions, update documentation
5. **Verify**: Confirm implementation complete
6. **Effectiveness**: Monitor for recurrence (30-90 days)
7. **Close**: Management approval and closure
**Reference:** See [qsr_compliance_requirements.md](references/qsr_compliance_requirements.md) for detailed QSR implementation guidance.
---
## HIPAA for Medical Devices
HIPAA requirements for devices that create, store, transmit, or access Protected Health Information (PHI).
### Applicability
| Device Type | HIPAA Applies |
|-------------|---------------|
| Standalone diagnostic (no data transmission) | No |
| Connected device transmitting patient data | Yes |
| Device with EHR integration | Yes |
| SaMD storing patient information | Yes |
| Wellness app (no diagnosis) | Only if stores PHI |
### Required Safeguards
```
Administrative (§164.308)
├── Security officer designation
├── Risk analysis and management
├── Workforce training
├── Incident response procedures
└── Business associate agreements
Physical (§164.310)
├── Facility access controls
├── Workstation security
└── Device disposal procedures
Technical (§164.312)
├── Access control (unique IDs, auto-logoff)
├── Audit controls (logging)
├── Integrity controls (checksums, hashes)
├── Authentication (MFA recommended)
└── Transmission security (TLS 1.2+)
```
### Risk Assessment Steps
1. Inventory all systems handling ePHI
2. Document data flows (collection, storage, transmission)
3. Identify threats and vulnerabilities
4. Assess likelihood and impact
5. Determine risk levels
6. Implement controls
7. Document residual risk
**Reference:** See [hipaa_compliance_framework.md](references/hipaa_compliance_framework.md) for implementation checklists and BAA templates.
---
## Device Cybersecurity
FDA cybersecurity requirements for connected medical devices.
### Premarket Requirements
| Element | Description |
|---------|-------------|
| Threat Model | STRIDE analysis, attack trees, trust boundaries |
| Security Controls | Authentication, encryption, access control |
| SBOM | Software Bill of Materials (CycloneDX or SPDX) |
| Security Testing | Penetration testing, vulnerability scanning |
| Vulnerability Plan | Disclosure process, patch management |
### Device Tier Classification
**Tier 1 (Higher Risk):**
- Connects to network/internet
- Cybersecurity incident could cause patient harm
**Tier 2 (Standard Risk):**
- All other connected devices
### Postmarket Obligations
1. Monitor NVD and ICS-CERT for vulnerabilities
2. Assess applicability to device components
3. Develop and test patches
4. Communicate with customers
5. Report to FDA per guidance
### Coordinated Vulnerability Disclosure
```
Researcher Report
↓
Acknowledgment (48 hours)
↓
Initial Assessment (5 days)
↓
Fix Development
↓
Coordinated Public Disclosure
```
**Reference:** See [device_cybersecurity_guidance.md](references/device_cybersecurity_guidance.md) for SBOM format examples and threat modeling templates.
---
## Resources
### scripts/
| Script | Purpose |
|--------|---------|
| `fda_submission_tracker.py` | Track 510(k)/PMA/De Novo submission milestones and timelines |
| `qsr_compliance_checker.py` | Assess 21 CFR 820 compliance against project documentation |
| `hipaa_risk_assessment.py` | Evaluate HIPAA safeguards in medical device software |
### references/
| File | Content |
|------|---------|
| `fda_submission_guide.md` | 510(k), De Novo, PMA submission requirements and checklists |
| `qsr_compliance_requirements.md` | 21 CFR 820 implementation guide with templates |
| `hipaa_compliance_framework.md` | HIPAA Security Rule safeguards and BAA requirements |
| `device_cybersecurity_guidance.md` | FDA cybersecurity requirements, SBOM, threat modeling |
| `fda_capa_requirements.md` | CAPA process, root cause analysis, effectiveness verification |
### Usage Examples
```bash
# Track FDA submission status
python scripts/fda_submission_tracker.py /path/to/project --type 510k
# Assess QSR compliance
python scripts/qsr_compliance_checker.py /path/to/project --section 820.30
# Run HIPAA risk assessment
python scripts/hipaa_risk_assessment.py /path/to/project --category technical
```
FILE:references/device_cybersecurity_guidance.md
# Medical Device Cybersecurity Guidance
Complete framework for FDA cybersecurity requirements based on FDA guidance documents and recognized consensus standards.
---
## Table of Contents
- [Regulatory Framework](#regulatory-framework)
- [Premarket Cybersecurity](#premarket-cybersecurity)
- [Postmarket Cybersecurity](#postmarket-cybersecurity)
- [Threat Modeling](#threat-modeling)
- [Security Controls](#security-controls)
- [Software Bill of Materials](#software-bill-of-materials)
- [Vulnerability Management](#vulnerability-management)
- [Documentation Requirements](#documentation-requirements)
---
## Regulatory Framework
### FDA Guidance Documents
| Document | Scope | Key Requirements |
|----------|-------|------------------|
| Premarket Cybersecurity (2023) | 510(k), PMA, De Novo | Security design, SBOM, threat modeling |
| Postmarket Management (2016) | All marketed devices | Vulnerability monitoring, patching |
| Content of Premarket Submissions | Submission format | Documentation structure |
### PATCH Act Requirements (2023)
**Cyber Device Definition:**
- Contains software
- Can connect to internet
- May be vulnerable to cybersecurity threats
**Manufacturer Obligations:**
1. Submit plan to monitor, identify, and address vulnerabilities
2. Design, develop, and maintain processes to ensure device security
3. Provide software bill of materials (SBOM)
4. Comply with other requirements under section 524B
### Recognized Consensus Standards
| Standard | Scope | FDA Recognition |
|----------|-------|-----------------|
| IEC 62443 | Industrial automation security | Recognized |
| NIST Cybersecurity Framework | Security framework | Referenced |
| UL 2900 | Software cybersecurity | Recognized |
| AAMI TIR57 | Medical device cybersecurity | Referenced |
| IEC 81001-5-1 | Health software security | Recognized |
---
## Premarket Cybersecurity
### Cybersecurity Documentation Requirements
```
Cybersecurity Documentation Package:
├── 1. Security Risk Assessment
│ ├── Threat model
│ ├── Vulnerability assessment
│ ├── Risk analysis
│ └── Risk mitigation
├── 2. Security Architecture
│ ├── System diagram
│ ├── Data flow diagram
│ ├── Trust boundaries
│ └── Security controls
├── 3. Cybersecurity Testing
│ ├── Penetration testing
│ ├── Vulnerability scanning
│ ├── Fuzz testing
│ └── Security code review
├── 4. SBOM
│ ├── Software components
│ ├── Versions
│ └── Known vulnerabilities
├── 5. Vulnerability Management Plan
│ ├── Monitoring process
│ ├── Disclosure process
│ └── Patch management
└── 6. Labeling
├── Security instructions
└── End-of-life plan
```
### Device Tier Classification
**Tier 1 - Higher Cybersecurity Risk:**
- Device can connect to another product or network
- A cybersecurity incident could directly result in patient harm
**Tier 2 - Standard Cybersecurity Risk:**
- Device NOT a Tier 1 device
- Still requires cybersecurity documentation
**Documentation Depth by Tier:**
| Element | Tier 1 | Tier 2 |
|---------|--------|--------|
| Threat model | Comprehensive | Basic |
| Penetration testing | Required | Recommended |
| SBOM | Required | Required |
| Security testing | Full suite | Core testing |
### Security by Design Principles
```markdown
## Secure Product Development Framework (SPDF)
### 1. Security Risk Management
- Integrate security into QMS
- Apply throughout product lifecycle
- Document security decisions
### 2. Security Architecture
- Defense in depth
- Least privilege
- Secure defaults
- Fail securely
### 3. Cybersecurity Testing
- Verify security controls
- Test for known vulnerabilities
- Validate threat mitigations
### 4. Cybersecurity Transparency
- SBOM provision
- Vulnerability disclosure
- Coordinated vulnerability disclosure
### 5. Cybersecurity Maintenance
- Monitor for vulnerabilities
- Provide timely updates
- Support throughout lifecycle
```
---
## Postmarket Cybersecurity
### Vulnerability Monitoring
**Sources to Monitor:**
- National Vulnerability Database (NVD)
- ICS-CERT advisories
- Third-party component vendors
- Security researcher reports
- Customer/user reports
**Monitoring Process:**
```
Daily/Weekly Monitoring:
├── NVD feed check
├── Vendor security bulletins
├── Security mailing lists
└── ISAC notifications
Monthly Review:
├── Component vulnerability analysis
├── Risk re-assessment
├── Patch status review
└── Trending threat analysis
Quarterly Assessment:
├── Comprehensive vulnerability scan
├── Third-party security audit
├── Update threat model
└── Security metrics review
```
### Vulnerability Assessment and Response
**CVSS-Based Triage:**
| CVSS Score | Severity | Response Timeframe |
|------------|----------|-------------------|
| 9.0-10.0 | Critical | 24-48 hours assessment |
| 7.0-8.9 | High | 1 week assessment |
| 4.0-6.9 | Medium | 30 days assessment |
| 0.1-3.9 | Low | Quarterly review |
**Exploitability Assessment:**
```markdown
## Vulnerability Exploitation Assessment
### Device-Specific Factors
- [ ] Is the vulnerability reachable in device configuration?
- [ ] Are mitigating controls in place?
- [ ] What is the attack surface exposure?
- [ ] What is the potential patient harm?
### Environment Factors
- [ ] Is exploit code publicly available?
- [ ] Is the vulnerability being actively exploited?
- [ ] What is the typical deployment environment?
### Risk Determination
Uncontrolled Risk = Exploitability × Impact × Exposure
| Risk Level | Action |
|------------|--------|
| Unacceptable | Immediate remediation |
| Elevated | Prioritized remediation |
| Acceptable | Monitor, routine update |
```
### Patch and Update Management
**Update Classification:**
| Type | Description | Regulatory Path |
|------|-------------|-----------------|
| Security patch | Addresses vulnerability only | May not require new submission |
| Software update | New features + security | Evaluate per guidance |
| Major upgrade | Significant changes | New 510(k) evaluation |
**FDA's Cybersecurity Policies:**
1. **Routine Updates:** Generally do not require premarket review
2. **Remediation of Vulnerabilities:** No premarket review if:
- No new risks introduced
- No changes to intended use
- Adequate design controls followed
---
## Threat Modeling
### STRIDE Methodology
| Threat | Description | Device Example |
|--------|-------------|----------------|
| **S**poofing | Pretending to be someone/something else | Fake device identity |
| **T**ampering | Modifying data or code | Altering dosage parameters |
| **R**epudiation | Denying actions | Hiding malicious commands |
| **I**nformation Disclosure | Exposing information | PHI data leak |
| **D**enial of Service | Making resource unavailable | Device becomes unresponsive |
| **E**levation of Privilege | Gaining unauthorized access | Admin access from user |
### Threat Model Template
```markdown
## Device Threat Model
### 1. System Description
Device Name: _____________________
Device Type: _____________________
Intended Use: ____________________
### 2. Architecture Diagram
[Include system diagram with trust boundaries]
### 3. Data Flow Diagram
[Document data flows and data types]
### 4. Entry Points
| Entry Point | Protocol | Authentication | Data Type |
|-------------|----------|----------------|-----------|
| USB port | USB HID | None | Config data |
| Network | HTTPS | Certificate | PHI |
| Bluetooth | BLE | Pairing | Commands |
### 5. Assets
| Asset | Sensitivity | Integrity | Availability |
|-------|-------------|-----------|--------------|
| Patient data | High | High | Medium |
| Device firmware | High | Critical | High |
| Configuration | Medium | High | Medium |
### 6. Threat Analysis
| Threat ID | STRIDE | Entry Point | Asset | Mitigation |
|-----------|--------|-------------|-------|------------|
| T-001 | Spoofing | Network | Auth | Mutual TLS |
| T-002 | Tampering | USB | Firmware | Secure boot |
| T-003 | Information | Network | PHI | Encryption |
### 7. Risk Assessment
| Threat | Likelihood | Impact | Risk | Accept/Mitigate |
|--------|------------|--------|------|-----------------|
| T-001 | Medium | High | High | Mitigate |
| T-002 | Low | Critical | High | Mitigate |
| T-003 | Medium | High | High | Mitigate |
```
### Attack Trees
**Example: Unauthorized Access to Device**
```
Goal: Gain Unauthorized Access
├── 1. Physical Access Attack
│ ├── 1.1 Steal device
│ ├── 1.2 Access debug port
│ └── 1.3 Extract storage media
├── 2. Network Attack
│ ├── 2.1 Exploit unpatched vulnerability
│ ├── 2.2 Man-in-the-middle attack
│ └── 2.3 Credential theft
├── 3. Social Engineering
│ ├── 3.1 Phishing for credentials
│ └── 3.2 Insider threat
└── 4. Supply Chain Attack
├── 4.1 Compromised component
└── 4.2 Malicious update
```
---
## Security Controls
### Authentication and Access Control
**Authentication Requirements:**
| Access Level | Authentication | Session Management |
|--------------|----------------|-------------------|
| Patient | PIN/biometric | Auto-logout |
| Clinician | Password + MFA | Timeout 15 min |
| Service | Certificate | Per-session |
| Admin | MFA + approval | Audit logged |
**Password Requirements:**
- Minimum 8 characters (12+ recommended)
- Complexity requirements
- Secure storage (hashed, salted)
- Account lockout after failed attempts
- Forced change on first use
### Encryption Requirements
**Data at Rest:**
- AES-256 for sensitive data
- Secure key storage (TPM, secure enclave)
- Key rotation procedures
**Data in Transit:**
- TLS 1.2 or higher
- Strong cipher suites
- Certificate validation
- Perfect forward secrecy
**Encryption Implementation Checklist:**
```markdown
## Encryption Controls
### Key Management
- [ ] Keys stored in hardware security module or equivalent
- [ ] Key generation uses cryptographically secure RNG
- [ ] Key rotation procedures documented
- [ ] Key revocation procedures documented
- [ ] Key escrow/recovery procedures (if applicable)
### Algorithm Selection
- [ ] AES-256 for symmetric encryption
- [ ] RSA-2048+ or ECDSA P-256+ for asymmetric
- [ ] SHA-256 or better for hashing
- [ ] No deprecated algorithms (MD5, SHA-1, DES)
### Implementation
- [ ] Using well-vetted cryptographic libraries
- [ ] Proper initialization vector handling
- [ ] Protection against timing attacks
- [ ] Secure key zeroing after use
```
### Secure Communications
**Network Security Controls:**
| Layer | Control | Implementation |
|-------|---------|----------------|
| Transport | TLS 1.2+ | Mutual authentication |
| Network | Firewall | Whitelist only |
| Application | API security | Rate limiting, validation |
| Data | Encryption | End-to-end |
### Code Integrity
**Secure Boot Chain:**
```
Root of Trust (Hardware)
↓
Bootloader (Signed)
↓
Operating System (Verified)
↓
Application (Authenticated)
↓
Configuration (Integrity-checked)
```
**Software Integrity Controls:**
- Code signing for all software
- Signature verification before execution
- Anti-rollback protection
- Secure update mechanism
---
## Software Bill of Materials
### SBOM Requirements
**NTIA Minimum Elements:**
1. Supplier name
2. Component name
3. Version of component
4. Other unique identifiers (PURL, CPE)
5. Dependency relationship
6. Author of SBOM data
7. Timestamp
### SBOM Formats
| Format | Standard | Use Case |
|--------|----------|----------|
| SPDX | ISO/IEC 5962:2021 | Comprehensive |
| CycloneDX | OWASP | Security-focused |
| SWID | ISO/IEC 19770-2 | Asset management |
### SBOM Template (CycloneDX)
```xml
<?xml version="1.0" encoding="UTF-8"?>
<bom xmlns="http://cyclonedx.org/schema/bom/1.4">
<metadata>
<timestamp>2024-01-15T00:00:00Z</timestamp>
<tools>
<tool>
<vendor>Manufacturer</vendor>
<name>SBOM Generator</name>
<version>1.0.0</version>
</tool>
</tools>
<component type="device">
<name>Medical Device XYZ</name>
<version>2.0.0</version>
<supplier>
<name>Device Manufacturer</name>
</supplier>
</component>
</metadata>
<components>
<component type="library">
<name>openssl</name>
<version>1.1.1k</version>
<purl>pkg:generic/openssl@1.1.1k</purl>
<licenses>
<license>
<id>Apache-2.0</id>
</license>
</licenses>
</component>
<!-- Additional components -->
</components>
<dependencies>
<dependency ref="device-xyz">
<dependency ref="openssl"/>
</dependency>
</dependencies>
</bom>
```
### SBOM Management Process
```
1. Initial SBOM Creation
└── During development, before submission
2. Vulnerability Monitoring
└── Continuous monitoring against NVD
3. SBOM Updates
└── With each software release
4. Customer Communication
└── SBOM provided on request
5. FDA Submission
└── Included in premarket submission
```
---
## Vulnerability Management
### Vulnerability Disclosure
**Coordinated Vulnerability Disclosure (CVD):**
```markdown
## Vulnerability Disclosure Policy
### Reporting
- Security contact: security@manufacturer.com
- PGP key available at: [URL]
- Bug bounty program: [if applicable]
### Response Timeline
- Acknowledgment: Within 48 hours
- Initial assessment: Within 5 business days
- Status updates: Every 30 days
- Target remediation: Per severity
### Public Disclosure
- Coordinated with reporter
- After remediation available
- Include mitigations if patch delayed
### Safe Harbor
[Statement on not pursuing legal action against good-faith reporters]
```
### Vulnerability Response Process
```
Discovery
↓
Triage (CVSS + Exploitability)
↓
Risk Assessment
↓
Remediation Development
↓
Testing and Validation
↓
Deployment/Communication
↓
Verification
↓
Closure
```
### Customer Communication
**Security Advisory Template:**
```markdown
## Security Advisory
### Advisory ID: [ID]
### Published: [Date]
### Severity: [Critical/High/Medium/Low]
### Affected Products
- Product A, versions 1.0-2.0
- Product B, versions 3.0-3.5
### Description
[Description of vulnerability without exploitation details]
### Impact
[What could happen if exploited]
### Mitigation
[Steps to reduce risk before patch available]
### Remediation
- Patch version: X.X.X
- Download: [URL]
- Installation instructions: [Link]
### Credits
[Acknowledge reporter if agreed]
### References
- CVE-XXXX-XXXX
- Manufacturer reference: [ID]
```
---
## Documentation Requirements
### Premarket Submission Checklist
```markdown
## Cybersecurity Documentation for Premarket Submission
### Device Description (Tier 1 and 2)
- [ ] Cybersecurity risk level justification
- [ ] Global system diagram
- [ ] Data flow diagram
### Security Risk Management (Tier 1 and 2)
- [ ] Threat model
- [ ] Security risk assessment
- [ ] Traceability matrix
### Security Architecture (Tier 1 and 2)
- [ ] Defense-in-depth description
- [ ] Security controls list
- [ ] Trust boundaries identified
### Testing Documentation
#### Tier 1
- [ ] Penetration test report
- [ ] Vulnerability scan results
- [ ] Fuzz testing results
- [ ] Static code analysis
- [ ] Third-party component testing
#### Tier 2
- [ ] Security testing summary
- [ ] Known vulnerability analysis
### SBOM (Tier 1 and 2)
- [ ] Complete component inventory
- [ ] Known vulnerability assessment
- [ ] Support and update plan
### Vulnerability Management (Tier 1 and 2)
- [ ] Vulnerability handling policy
- [ ] Coordinated disclosure process
- [ ] Security update plan
### Labeling (Tier 1 and 2)
- [ ] User security instructions
- [ ] End-of-support date
- [ ] Security contact information
```
### Recommended File Structure
```
Cybersecurity_Documentation/
├── 01_Executive_Summary.pdf
├── 02_Device_Description/
│ ├── System_Diagram.pdf
│ └── Data_Flow_Diagram.pdf
├── 03_Security_Risk_Assessment/
│ ├── Threat_Model.pdf
│ ├── Risk_Assessment.pdf
│ └── Traceability_Matrix.xlsx
├── 04_Security_Architecture/
│ ├── Architecture_Description.pdf
│ ├── Security_Controls.pdf
│ └── Trust_Boundary_Analysis.pdf
├── 05_Security_Testing/
│ ├── Penetration_Test_Report.pdf
│ ├── Vulnerability_Scan_Results.pdf
│ ├── Fuzz_Testing_Report.pdf
│ └── Code_Analysis_Report.pdf
├── 06_SBOM/
│ ├── SBOM.xml (CycloneDX)
│ └── Vulnerability_Analysis.pdf
├── 07_Vulnerability_Management/
│ ├── Vulnerability_Policy.pdf
│ └── Disclosure_Process.pdf
└── 08_Labeling/
└── Security_Instructions.pdf
```
---
## Quick Reference
### Common Cybersecurity Deficiencies
| Deficiency | Resolution |
|------------|------------|
| Incomplete threat model | Document all entry points, assets, threats |
| No SBOM provided | Generate using automated tools |
| Weak authentication | Implement MFA, strong passwords |
| Missing encryption | Add TLS 1.2+, AES-256 |
| No vulnerability management plan | Create monitoring and response procedures |
| Insufficient testing | Conduct penetration testing |
### Security Testing Requirements
| Test Type | Tier 1 | Tier 2 | Tools |
|-----------|--------|--------|-------|
| Penetration testing | Required | Recommended | Manual + automated |
| Vulnerability scanning | Required | Required | Nessus, OpenVAS |
| Fuzz testing | Required | Recommended | AFL, Peach |
| Static analysis | Required | Recommended | SonarQube, Coverity |
| Dynamic analysis | Required | Recommended | Burp Suite, ZAP |
### Recognized Standards Mapping
| FDA Requirement | IEC 62443 | NIST CSF |
|-----------------|-----------|----------|
| Threat modeling | SR 3 | ID.RA |
| Access control | SR 1, SR 2 | PR.AC |
| Encryption | SR 4 | PR.DS |
| Audit logging | SR 6 | PR.PT, DE.AE |
| Patch management | SR 7 | PR.MA |
| Incident response | SR 6 | RS.RP |
FILE:references/fda_capa_requirements.md
# FDA CAPA Requirements
Complete guide to Corrective and Preventive Action requirements per 21 CFR 820.100.
---
## Table of Contents
- [CAPA Regulation Overview](#capa-regulation-overview)
- [CAPA Sources](#capa-sources)
- [CAPA Process](#capa-process)
- [Root Cause Analysis](#root-cause-analysis)
- [Action Implementation](#action-implementation)
- [Effectiveness Verification](#effectiveness-verification)
- [Documentation Requirements](#documentation-requirements)
- [FDA Inspection Focus Areas](#fda-inspection-focus-areas)
---
## CAPA Regulation Overview
### 21 CFR 820.100 Requirements
```
§820.100 Corrective and preventive action
(a) Each manufacturer shall establish and maintain procedures for
implementing corrective and preventive action. The procedures shall
include requirements for:
(1) Analyzing processes, work operations, concessions, quality audit
reports, quality records, service records, complaints, returned
product, and other sources of quality data to identify existing
and potential causes of nonconforming product, or other quality
problems.
(2) Investigating the cause of nonconformities relating to product,
processes, and the quality system.
(3) Identifying the action(s) needed to correct and prevent recurrence
of nonconforming product and other quality problems.
(4) Verifying or validating the corrective and preventive action to
ensure that such action is effective and does not adversely affect
the finished device.
(5) Implementing and recording changes in methods and procedures needed
to correct and prevent identified quality problems.
(6) Ensuring that information related to quality problems or nonconforming
product is disseminated to those directly responsible for assuring
the quality of such product or the prevention of such problems.
(7) Submitting relevant information on identified quality problems, as
well as corrective and preventive actions, for management review.
```
### Definitions
| Term | Definition |
|------|------------|
| **Correction** | Action to eliminate a detected nonconformity |
| **Corrective Action** | Action to eliminate the cause of a detected nonconformity to prevent recurrence |
| **Preventive Action** | Action to eliminate the cause of a potential nonconformity to prevent occurrence |
| **Root Cause** | The fundamental reason for the occurrence of a problem |
| **Effectiveness** | Confirmation that actions achieved intended results |
### CAPA vs. Correction
```
Problem Detected
├── Correction (Immediate)
│ └── Fix the immediate issue
│ Example: Replace defective part
│
└── CAPA (Systemic)
└── Address root cause
Example: Fix process that caused defect
```
---
## CAPA Sources
### Data Sources for CAPA Input
**Internal Sources:**
- Nonconforming product reports (NCRs)
- Internal audit findings
- Process deviations
- Manufacturing data trends
- Equipment failures
- Employee observations
- Training deficiencies
**External Sources:**
- Customer complaints
- Service records
- Returned product
- Regulatory feedback (483s, warning letters)
- Adverse event reports (MDRs)
- Field safety corrective actions
### CAPA Threshold Criteria
**Mandatory CAPA Triggers:**
| Source | Threshold |
|--------|-----------|
| Audit findings | All major/critical findings |
| Customer complaints | Any safety-related |
| NCRs | Recurring (3+ occurrences) |
| Regulatory feedback | All observations |
| MDR/vigilance | All reportable events |
**Discretionary CAPA Evaluation:**
| Source | Consideration |
|--------|---------------|
| Trend data | Statistical significance |
| Process deviations | Impact assessment |
| Minor audit findings | Risk-based |
| Supplier issues | Frequency and severity |
### Trend Analysis
**Statistical Process Control:**
```markdown
## Monthly CAPA Trend Review
### Complaint Trending
- [ ] Complaints by product
- [ ] Complaints by failure mode
- [ ] Geographic distribution
- [ ] Customer type analysis
### NCR Trending
- [ ] NCRs by product/process
- [ ] NCRs by cause code
- [ ] NCRs by supplier
- [ ] Scrap/rework rates
### Threshold Monitoring
| Metric | Threshold | Current | Status |
|--------|-----------|---------|--------|
| Complaints/month | <10 | | |
| NCR rate | <2% | | |
| Recurring issues | 0 | | |
```
---
## CAPA Process
### CAPA Workflow
```
1. Initiation
├── Problem identification
├── Initial assessment
└── CAPA determination
2. Investigation
├── Data collection
├── Root cause analysis
└── Impact assessment
3. Action Planning
├── Correction (if applicable)
├── Corrective action
└── Preventive action
4. Implementation
├── Execute actions
├── Document changes
└── Train affected personnel
5. Verification
├── Verify implementation
├── Validate effectiveness
└── Monitor for recurrence
6. Closure
├── Management approval
├── Final documentation
└── Trend data update
```
### CAPA Form Template
```markdown
## CAPA Record
### Section 1: Identification
CAPA Number: ________________
Initiated By: ________________
Date Initiated: ______________
Priority: ☐ Critical ☐ Major ☐ Minor
Source:
☐ Audit Finding ☐ Complaint ☐ NCR
☐ Service Record ☐ MDR ☐ Trend Data
☐ Regulatory ☐ Other: ____________
### Section 2: Problem Description
Products Affected: _______________________
Processes Affected: _____________________
Quantity/Scope: _________________________
Problem Statement:
[Clear, specific description of the nonconformity or potential problem]
### Section 3: Immediate Correction
Correction Taken: _______________________
Date Completed: _________________________
Verified By: ____________________________
### Section 4: Investigation
Investigation Lead: _____________________
Investigation Start Date: _______________
Data Collected:
☐ Complaint records ☐ Production records
☐ Test data ☐ Training records
☐ Process documentation ☐ Supplier data
Root Cause Analysis Method:
☐ 5 Whys ☐ Fishbone ☐ Fault Tree ☐ Other
Root Cause Statement:
[Specific, factual statement of the root cause]
Contributing Factors:
1. _____________________________________
2. _____________________________________
### Section 5: Action Plan
#### Corrective Actions
| Action | Owner | Target Date | Status |
|--------|-------|-------------|--------|
| | | | |
#### Preventive Actions
| Action | Owner | Target Date | Status |
|--------|-------|-------------|--------|
| | | | |
### Section 6: Verification
Verification Method: ____________________
Verification Criteria: __________________
Verification Date: _____________________
Verified By: ___________________________
Verification Results:
☐ Actions implemented as planned
☐ No adverse effects identified
☐ Documentation updated
### Section 7: Effectiveness Review
Effectiveness Review Date: ______________
Review Period: ________________________
Reviewer: _____________________________
Effectiveness Criteria:
[Specific, measurable criteria for success]
Results:
☐ Effective - problem has not recurred
☐ Not Effective - additional action required
Evidence:
[Reference to data showing effectiveness]
### Section 8: Closure
Closure Date: _________________________
Approved By: __________________________
Management Review Submitted: ☐ Yes ☐ No
Date: ________________________________
```
---
## Root Cause Analysis
### 5 Whys Technique
**Example: Device Fails Final Test**
```
Problem: 5% of devices fail functional test at final inspection
Why 1: Component X is out of tolerance
Why 2: Component X was accepted at incoming inspection
Why 3: Incoming inspection sampling missed defective lot
Why 4: Sampling plan inadequate for component criticality
Why 5: Risk classification of component not updated after design change
Root Cause: Risk classification process did not include design change trigger
```
**5 Whys Template:**
```markdown
## 5 Whys Analysis
Problem Statement: _________________________________
Why 1: _____________________________________________
Evidence: __________________________________________
Why 2: _____________________________________________
Evidence: __________________________________________
Why 3: _____________________________________________
Evidence: __________________________________________
Why 4: _____________________________________________
Evidence: __________________________________________
Why 5: _____________________________________________
Evidence: __________________________________________
Root Cause: ________________________________________
Verification: How do we know this is the root cause?
________________________________________________
```
### Fishbone (Ishikawa) Diagram
**Categories for Medical Device Manufacturing:**
```
┌─────────────────────────────────────────┐
│ PROBLEM │
└─────────────────────────────────────────┘
▲
┌───────────────────────────┼───────────────────────────┐
│ │ │
┌──────┴──────┐ ┌──────┴──────┐ ┌──────┴──────┐
│ PERSONNEL │ │ METHODS │ │ MATERIALS │
│ │ │ │ │ │
│ • Training │ │ • SOP gaps │ │ • Supplier │
│ • Skills │ │ • Process │ │ • Specs │
│ • Attention │ │ • Sequence │ │ • Storage │
└─────────────┘ └─────────────┘ └─────────────┘
│ │ │
└───────────────────────────┼───────────────────────────┘
│
┌──────────────┐ ┌──────┴──────┐ ┌──────────────┐
│ MEASUREMENT │ │ EQUIPMENT │ │ ENVIRONMENT │
│ │ │ │ │ │
│ • Calibration│ │ • Maintenance│ │ • Temperature│
│ • Method │ │ • Capability │ │ • Humidity │
│ • Accuracy │ │ • Tooling │ │ • Cleanliness│
└──────────────┘ └─────────────┘ └──────────────┘
```
### Fault Tree Analysis
**For Complex Failures:**
```
Top Event: Device Failure
│
┌───────────────┼───────────────┐
│ │ │
AND/OR AND/OR AND/OR
│ │ │
┌─────┴─────┐ ┌─────┴─────┐ ┌─────┴─────┐
│ Component │ │ Software │ │ User │
│ Failure │ │ Failure │ │ Error │
└─────┬─────┘ └─────┬─────┘ └─────┬─────┘
│ │ │
Basic Events Basic Events Basic Events
```
### Root Cause Categories
| Category | Examples | Evidence Sources |
|----------|----------|------------------|
| Design | Specification error, tolerance stack-up | DHF, design review records |
| Process | Procedure inadequate, sequence error | Process validation, work instructions |
| Personnel | Training gap, human error | Training records, interviews |
| Equipment | Calibration drift, maintenance | Calibration records, logs |
| Material | Supplier quality, storage | Incoming inspection, COCs |
| Environment | Temperature, contamination | Environmental monitoring |
| Management | Resource allocation, priorities | Management review records |
---
## Action Implementation
### Corrective Action Requirements
**Effective Corrective Actions:**
1. Address identified root cause
2. Are specific and measurable
3. Have assigned ownership
4. Have realistic target dates
5. Consider impact on other processes
6. Include verification method
**Action Types:**
| Type | Description | Example |
|------|-------------|---------|
| Process change | Modify procedure or method | Update SOP with additional step |
| Design change | Modify product design | Add tolerance specification |
| Training | Improve personnel capability | Conduct retraining |
| Equipment | Modify or replace equipment | Upgrade inspection equipment |
| Supplier | Address supplier quality | Audit supplier, add requirements |
| Documentation | Improve or add documentation | Create work instruction |
### Change Control Integration
```
CAPA Action Identified
│
▼
Change Request Initiated
│
▼
Impact Assessment
├── Regulatory impact
├── Product impact
├── Process impact
└── Documentation impact
│
▼
Change Approved
│
▼
Implementation
├── Document updates
├── Training
├── Validation (if required)
└── Effective date
│
▼
CAPA Verification
```
### Training Requirements
**When Training is Required:**
- New or revised procedures
- New equipment or tools
- Process changes
- Findings related to personnel performance
**Training Documentation:**
```markdown
## CAPA-Related Training Record
CAPA Number: _______________
Training Subject: ___________
Training Date: ______________
Trainer: ___________________
Attendees:
| Name | Signature | Date |
|------|-----------|------|
| | | |
Training Content:
- [ ] Root cause explanation
- [ ] Process/procedure changes
- [ ] New requirements
- [ ] Competency verification
Competency Verified By: _______________
Date: _______________
```
---
## Effectiveness Verification
### Verification vs. Validation
| Verification | Validation |
|--------------|------------|
| Actions implemented correctly | Actions achieved intended results |
| Short-term check | Long-term monitoring |
| Process-focused | Outcome-focused |
### Effectiveness Criteria
**SMART Criteria:**
- **S**pecific: Clearly defined outcome
- **M**easurable: Quantifiable metrics
- **A**chievable: Realistic expectations
- **R**elevant: Related to root cause
- **T**ime-bound: Defined monitoring period
**Examples:**
| Problem | Root Cause | Action | Effectiveness Criteria |
|---------|------------|--------|----------------------|
| 5% test failures | Inadequate sampling | Increase sampling | <1% failure rate for 3 months |
| Customer complaints | Unclear instructions | Revise IFU | Zero complaints on topic for 6 months |
| NCRs from supplier | No incoming inspection | Add inspection | Zero supplier NCRs for 90 days |
### Effectiveness Review Template
```markdown
## CAPA Effectiveness Review
CAPA Number: _______________
Review Date: _______________
Reviewer: __________________
### Review Criteria
Original Problem: _________________
Effectiveness Metric: ______________
Success Threshold: ________________
Review Period: ____________________
### Data Analysis
| Period | Metric Value | Threshold | Pass/Fail |
|--------|--------------|-----------|-----------|
| Month 1 | | | |
| Month 2 | | | |
| Month 3 | | | |
### Conclusion
☐ Effective - Criteria met, CAPA may be closed
☐ Partially Effective - Additional monitoring required
☐ Not Effective - Additional actions required
### Evidence
[Reference to supporting data: complaint logs, NCR reports, audit results, etc.]
### Next Steps (if not effective)
___________________________________
___________________________________
### Approval
Reviewer Signature: _______________ Date: _______
Quality Approval: _________________ Date: _______
```
### Monitoring Period Guidelines
| CAPA Type | Minimum Monitoring |
|-----------|-------------------|
| Product quality | 3 production lots or 90 days |
| Process | 3 months of production |
| Complaints | 6 months |
| Audit findings | Until next audit |
| Supplier | 3 lots or 90 days |
---
## Documentation Requirements
### CAPA File Contents
```
CAPA File Structure:
├── CAPA Form (all sections completed)
├── Investigation Records
│ ├── Data collected
│ ├── Root cause analysis worksheets
│ └── Impact assessment
├── Action Documentation
│ ├── Action plans
│ ├── Change requests (if applicable)
│ └── Training records
├── Verification Evidence
│ ├── Implementation verification
│ ├── Effectiveness data
│ └── Trend analysis
└── Closure Documentation
├── Closure approval
└── Management review submission
```
### Record Retention
Per 21 CFR 820.180:
- Records shall be retained for the design and expected life of the device
- Minimum of 2 years from date of release for commercial distribution
**CAPA Record Retention:**
- Retain for lifetime of product + 2 years
- Include all supporting documentation
- Maintain audit trail for changes
### Traceability
**Required Traceability:**
- CAPA to source (complaint, NCR, audit finding)
- CAPA to affected products/lots
- CAPA to corrective actions taken
- CAPA to verification evidence
- CAPA to management review
---
## FDA Inspection Focus Areas
### Common 483 Observations
| Observation | Prevention |
|-------------|------------|
| CAPA not initiated when required | Define clear CAPA triggers |
| Root cause analysis inadequate | Use structured RCA methods |
| Actions don't address root cause | Verify action-cause linkage |
| Effectiveness not verified | Define measurable criteria |
| CAPA not timely | Set and track target dates |
| Trend analysis not performed | Implement monthly trending |
| Management review missing CAPA input | Include in management review agenda |
### Inspection Preparation
**CAPA Readiness Checklist:**
```markdown
## FDA Inspection CAPA Preparation
### Documentation Review
- [ ] All CAPAs have complete documentation
- [ ] No overdue CAPAs
- [ ] Root cause documented with evidence
- [ ] Effectiveness verified and documented
- [ ] All open CAPAs have current status
### Metrics Available
- [ ] CAPA by source
- [ ] CAPA cycle time
- [ ] Overdue CAPA trend
- [ ] Effectiveness rate
- [ ] Recurring issues
### Process Evidence
- [ ] CAPA procedure current
- [ ] Training records complete
- [ ] Trend analysis documented
- [ ] Management review records show CAPA input
### Common Questions Prepared
- How do you initiate a CAPA?
- How do you determine root cause?
- How do you verify effectiveness?
- Show me your overdue CAPAs
- Show me CAPAs from complaints
```
### CAPA Metrics Dashboard
| Metric | Target | Calculation |
|--------|--------|-------------|
| On-time initiation | 100% | CAPAs initiated within 30 days |
| On-time closure | >90% | CAPAs closed by target date |
| Effectiveness rate | >85% | Effective at first review / Total |
| Average cycle time | <90 days | Average days to closure |
| Overdue CAPAs | 0 | CAPAs past target date |
| Recurring issues | <5% | Repeat CAPAs / Total |
---
## Quick Reference
### CAPA Decision Tree
```
Quality Issue Identified
│
▼
Is it an isolated incident?
├── YES → Correction only (document, may not need CAPA)
│ Evaluate for trend
│
└── NO → Is it a systemic issue?
├── YES → Initiate CAPA
│ Determine if Corrective or Preventive
│
└── MAYBE → Investigate further
Monitor for recurrence
May escalate to CAPA
```
### Root Cause vs. Symptom
| Symptom (NOT root cause) | Root Cause (Address this) |
|--------------------------|---------------------------|
| "Operator made error" | Training inadequate for task |
| "Component was defective" | Incoming inspection ineffective |
| "SOP not followed" | SOP unclear or impractical |
| "Equipment malfunctioned" | Maintenance schedule inadequate |
| "Supplier shipped wrong part" | Purchasing requirements unclear |
### Action Effectiveness Verification
| Action Type | Verification Method | Timeframe |
|-------------|---------------------|-----------|
| Procedure change | Audit for compliance | 30-60 days |
| Training | Competency assessment | Immediate |
| Design change | Product testing | Per protocol |
| Supplier action | Incoming inspection data | 3 lots |
| Equipment | Calibration/performance | Per schedule |
### Integration with Other Systems
| System | CAPA Integration Point |
|--------|------------------------|
| Complaints | Trigger for CAPA, complaint closure after CAPA |
| NCR | Trend to CAPA, NCR references CAPA |
| Audit | Findings generate CAPA, CAPA closure audit |
| Design Control | Design change via CAPA, DHF update |
| Supplier | Supplier CAPA, supplier audit findings |
| Risk Management | Risk file update post-CAPA |
FILE:references/fda_submission_guide.md
# FDA Submission Guide
Complete framework for 510(k), De Novo, and PMA submissions to the FDA.
---
## Table of Contents
- [Submission Pathway Selection](#submission-pathway-selection)
- [510(k) Premarket Notification](#510k-premarket-notification)
- [De Novo Classification](#de-novo-classification)
- [PMA Premarket Approval](#pma-premarket-approval)
- [Pre-Submission Program](#pre-submission-program)
- [FDA Review Timeline](#fda-review-timeline)
---
## Submission Pathway Selection
### Decision Matrix
```
Is there a legally marketed predicate device?
├── YES → Is your device substantially equivalent?
│ ├── YES → 510(k) Pathway
│ │ ├── No changes from predicate → Abbreviated 510(k)
│ │ ├── Manufacturing changes only → Special 510(k)
│ │ └── Design/performance changes → Traditional 510(k)
│ └── NO → PMA or De Novo
└── NO → Is it a novel low-to-moderate risk device?
├── YES → De Novo Classification Request
└── NO → PMA Pathway (Class III)
```
### Classification Determination
| Class | Risk Level | Pathway | Examples |
|-------|------------|---------|----------|
| I | Low | Exempt or 510(k) | Bandages, stethoscopes |
| II | Moderate | 510(k) | Powered wheelchairs, pregnancy tests |
| III | High | PMA | Pacemakers, heart valves |
### Predicate Device Search
**Database Sources:**
1. FDA 510(k) Database: https://www.accessdata.fda.gov/scripts/cdrh/cfdocs/cfpmn/pmn.cfm
2. FDA Product Classification Database
3. FDA PMA Database
4. FDA De Novo Database
**Search Criteria:**
- Product code (3-letter code)
- Device name keywords
- Intended use similarity
- Technological characteristics
---
## 510(k) Premarket Notification
### Required Sections (21 CFR 807.87)
#### 1. Administrative Information
```
Cover Letter
├── Submission type (Traditional/Special/Abbreviated)
├── Device name and classification
├── Predicate device(s) identification
├── Contact information
└── Signature of authorized representative
CDRH Premarket Review Submission Cover Sheet (FDA Form 3514)
├── Section A: Applicant Information
├── Section B: Device Information
├── Section C: Submission Information
└── Section D: Truth and Accuracy Statement
```
#### 2. Device Description
| Element | Required Content |
|---------|------------------|
| Device Name | Trade name, common name, classification name |
| Intended Use | Disease/condition, patient population, use environment |
| Physical Description | Materials, dimensions, components |
| Principles of Operation | How the device achieves intended use |
| Accessories | Included items, optional components |
| Variants/Models | All versions included in submission |
#### 3. Substantial Equivalence Comparison
```
Comparison Table Format:
┌────────────────────┬─────────────────┬─────────────────┐
│ Characteristic │ Subject Device │ Predicate │
├────────────────────┼─────────────────┼─────────────────┤
│ Intended Use │ [Your device] │ [Predicate] │
│ Technological │ │ │
│ Characteristics │ │ │
│ Performance │ │ │
│ Safety │ │ │
└────────────────────┴─────────────────┴─────────────────┘
Substantial Equivalence Argument:
1. Same intended use? YES/NO
2. Same technological characteristics? YES/NO
3. If different technology, does it raise new safety/effectiveness questions? YES/NO
4. Performance data demonstrates equivalence? YES/NO
```
#### 4. Performance Testing
**Bench Testing:**
- Mechanical/structural testing
- Electrical safety (IEC 60601-1 if applicable)
- Biocompatibility (ISO 10993 series)
- Sterilization validation
- Shelf life/stability testing
- Software verification (IEC 62304 if applicable)
**Clinical Data (if required):**
- Clinical study summaries
- Literature review
- Adverse event data
#### 5. Labeling
**Required Elements:**
- Instructions for Use (IFU)
- Device labeling (package, carton)
- Indications for Use statement
- Contraindications, warnings, precautions
- Advertising materials (if applicable)
### 510(k) Acceptance Checklist
```markdown
## Pre-Submission Verification
- [ ] FDA Form 3514 complete and signed
- [ ] User fee payment ($21,760 for FY2024, small business exemptions available)
- [ ] Device description complete
- [ ] Predicate device identified with 510(k) number
- [ ] Substantial equivalence comparison table
- [ ] Indications for Use statement (FDA Form 3881)
- [ ] Performance data summary
- [ ] Labeling (IFU, device labels)
- [ ] 510(k) summary or statement
- [ ] Truthful and Accuracy statement signed
- [ ] Environmental assessment or categorical exclusion
```
---
## De Novo Classification
### Eligibility Criteria
1. Novel device with no legally marketed predicate
2. Low-to-moderate risk (would be Class I or II if predicate existed)
3. General controls alone (Class I) or with special controls (Class II) provide reasonable assurance of safety and effectiveness
### Required Content
#### Risk Assessment
```
Risk Analysis Requirements:
├── Hazard Identification
│ ├── Biological hazards
│ ├── Mechanical hazards
│ ├── Electrical hazards
│ ├── Use-related hazards
│ └── Cybersecurity hazards (if applicable)
├── Risk Estimation
│ ├── Probability of occurrence
│ ├── Severity of harm
│ └── Risk level (High/Medium/Low)
├── Risk Evaluation
│ ├── Acceptability criteria
│ └── Benefit-risk analysis
└── Risk Control Measures
├── Design controls
├── Protective measures
└── Information for safety
```
#### Proposed Classification
| Classification | Controls | Rationale |
|----------------|----------|-----------|
| Class I | General controls only | Low risk, general controls adequate |
| Class II | General + Special controls | Moderate risk, special controls needed |
#### Special Controls (for Class II)
Define specific controls such as:
- Performance testing requirements
- Labeling requirements
- Post-market surveillance
- Patient registry
- Design specifications
---
## PMA Premarket Approval
### PMA Application Contents
#### Technical Sections
1. **Device Description and Intended Use**
- Detailed design specifications
- Operating principles
- Complete indications for use
2. **Manufacturing Information**
- Manufacturing process description
- Quality system information
- Facility registration
3. **Nonclinical Laboratory Studies**
- Bench testing results
- Animal studies (if applicable)
- Biocompatibility testing
4. **Clinical Investigation**
- IDE number and approval date
- Clinical protocol
- Clinical study results
- Statistical analysis
- Adverse events
5. **Labeling**
- Complete labeling
- Patient labeling (if applicable)
#### Clinical Data Requirements
```
Clinical Study Design:
├── Study Objectives
│ ├── Primary endpoint(s)
│ └── Secondary endpoint(s)
├── Study Population
│ ├── Inclusion criteria
│ ├── Exclusion criteria
│ └── Sample size justification
├── Study Design
│ ├── Randomized controlled trial
│ ├── Single-arm study with OPC
│ └── Other design with justification
├── Statistical Analysis Plan
│ ├── Analysis populations
│ ├── Statistical methods
│ └── Handling of missing data
└── Safety Monitoring
├── Adverse event definitions
├── Stopping rules
└── DSMB oversight
```
### IDE (Investigational Device Exemption)
**When Required:**
- Significant risk device clinical studies
- Studies not exempt under 21 CFR 812.2
**IDE Application Content:**
- Investigational plan
- Manufacturing information
- Investigator agreements
- IRB approvals
- Informed consent forms
- Labeling
- Risk analysis
---
## Pre-Submission Program
### Q-Submission Types
| Type | Purpose | FDA Response |
|------|---------|--------------|
| Pre-Sub | Feedback on planned submission | Written feedback or meeting |
| Informational | Share information, no feedback | Acknowledgment only |
| Study Risk | Determination of study risk level | Risk determination |
| Agreement/Determination | Binding agreement on specific issue | Formal agreement |
### Pre-Sub Meeting Preparation
```
Pre-Submission Package:
1. Cover letter with meeting request
2. Device description
3. Regulatory history (if any)
4. Proposed submission pathway
5. Specific questions (maximum 5-6)
6. Supporting data/information
Meeting Types:
- Written response only (default)
- Teleconference (90 minutes)
- In-person meeting (90 minutes)
```
### Effective Question Formulation
**Good Question Format:**
```
Question: Does FDA agree that [specific proposal] is acceptable for [specific purpose]?
Background: [Brief context - 1-2 paragraphs]
Proposal: [Your specific proposal - detailed but concise]
Rationale: [Why you believe this is appropriate]
```
**Avoid:**
- Open-ended questions ("What should we do?")
- Multiple questions combined
- Questions already answered in guidance
---
## FDA Review Timeline
### Standard Review Times
| Submission Type | FDA Goal | Typical Range |
|----------------|----------|---------------|
| 510(k) Traditional | 90 days | 90-150 days |
| 510(k) Special | 30 days | 30-60 days |
| 510(k) Abbreviated | 30 days | 30-60 days |
| De Novo | 150 days | 150-300 days |
| PMA | 180 days | 12-24 months |
| Pre-Sub Response | 70-75 days | 60-90 days |
### Review Process Stages
```
510(k) Review Timeline:
Day 0: Submission received
Day 1-15: Acceptance review
├── Accept → Substantive review begins
└── Refuse to Accept (RTA) → 180 days to respond
Day 15-90: Substantive review
├── Additional Information (AI) request stops clock
├── Interactive review may occur
└── Decision by Day 90 goal
Decision:
├── Substantially Equivalent (SE) → Clearance letter
├── Not Substantially Equivalent (NSE) → Appeal or new submission
└── Withdrawn
```
### Additional Information Requests
**Response Best Practices:**
- Respond within 30-60 days
- Use FDA's question numbering
- Provide complete responses
- Include amended sections clearly marked
- Reference specific guidance documents
---
## Submission Best Practices
### Document Formatting
- Use PDF format (PDF/A preferred)
- Bookmarks for each section
- Hyperlinks to cross-references
- Table of contents with page numbers
- Consistent headers/footers
### eSTAR (Electronic Submission Template)
FDA's recommended electronic submission format for 510(k):
- Structured data entry
- Built-in validation
- Automatic formatting
- Reduced RTA rate
### Common Refuse to Accept (RTA) Issues
| Issue | Prevention |
|-------|------------|
| Missing user fee | Verify payment before submission |
| Incomplete Form 3514 | Review all fields, ensure signature |
| Missing predicate | Confirm predicate is legally marketed |
| Inadequate device description | Include all models, accessories |
| Missing Indications for Use | Use FDA Form 3881 |
| Incomplete SE comparison | Address all characteristics |
FILE:references/hipaa_compliance_framework.md
# HIPAA Compliance Framework for Medical Devices
Complete guide to HIPAA requirements for medical device manufacturers and software developers.
---
## Table of Contents
- [HIPAA Overview](#hipaa-overview)
- [Privacy Rule Requirements](#privacy-rule-requirements)
- [Security Rule Requirements](#security-rule-requirements)
- [Medical Device Considerations](#medical-device-considerations)
- [Risk Assessment](#risk-assessment)
- [Implementation Specifications](#implementation-specifications)
- [Business Associate Agreements](#business-associate-agreements)
- [Breach Notification](#breach-notification)
---
## HIPAA Overview
### Applicability to Medical Devices
| Entity Type | HIPAA Applicability |
|-------------|---------------------|
| Healthcare providers | Covered Entity (CE) |
| Health plans | Covered Entity (CE) |
| Healthcare clearinghouses | Covered Entity (CE) |
| Device manufacturers | Business Associate (BA) if handling PHI |
| SaMD developers | Business Associate (BA) if handling PHI |
| Cloud service providers | Business Associate (BA) |
### Protected Health Information (PHI)
**PHI Definition:** Individually identifiable health information transmitted or maintained in any form.
**18 HIPAA Identifiers:**
```
1. Names
2. Geographic data (smaller than state)
3. Dates (except year) related to individual
4. Phone numbers
5. Fax numbers
6. Email addresses
7. Social Security numbers
8. Medical record numbers
9. Health plan beneficiary numbers
10. Account numbers
11. Certificate/license numbers
12. Vehicle identifiers
13. Device identifiers and serial numbers
14. Web URLs
15. IP addresses
16. Biometric identifiers
17. Full face photos
18. Any other unique identifying number
```
### Electronic PHI (ePHI)
PHI that is created, stored, transmitted, or received in electronic form. Most relevant for:
- Connected medical devices
- Medical device software (SaMD)
- Mobile health applications
- Cloud-based healthcare systems
---
## Privacy Rule Requirements
### Minimum Necessary Standard
**Principle:** Limit PHI access, use, and disclosure to the minimum necessary to accomplish the intended purpose.
**Implementation:**
- Role-based access controls
- Access audit logging
- Data segmentation
- Need-to-know policies
### Patient Rights
| Right | Device Implication |
|-------|---------------------|
| Access | Provide mechanism to view/export data |
| Amendment | Allow corrections to patient data |
| Accounting of disclosures | Log all PHI disclosures |
| Restriction requests | Support data sharing restrictions |
| Confidential communications | Secure communication channels |
### Use and Disclosure
**Permitted Uses:**
- Treatment, Payment, Healthcare Operations (TPO)
- With patient authorization
- Public health activities
- Required by law
- Health oversight activities
**Medical Device Context:**
- Device data for treatment: Permitted
- Data analytics by manufacturer: Requires BAA or de-identification
- Research use: Requires authorization or IRB waiver
---
## Security Rule Requirements
### Administrative Safeguards
#### Security Management Process (§164.308(a)(1))
**Required Specifications:**
```markdown
## Security Management Process
### Risk Analysis
- [ ] Identify systems with ePHI
- [ ] Document potential threats and vulnerabilities
- [ ] Assess likelihood and impact
- [ ] Document current controls
- [ ] Determine risk levels
### Risk Management
- [ ] Implement security measures
- [ ] Document residual risk
- [ ] Management approval
### Sanction Policy
- [ ] Define workforce sanctions
- [ ] Document enforcement procedures
### Information System Activity Review
- [ ] Define audit procedures
- [ ] Review logs regularly
- [ ] Document findings
```
#### Workforce Security (§164.308(a)(3))
| Specification | Type | Implementation |
|---------------|------|----------------|
| Authorization/supervision | Addressable | Access approval process |
| Workforce clearance | Addressable | Background checks |
| Termination procedures | Addressable | Access revocation |
#### Information Access Management (§164.308(a)(4))
**Access Control Elements:**
- Access authorization
- Access establishment and modification
- Unique user identification
- Automatic logoff
#### Security Awareness and Training (§164.308(a)(5))
**Training Topics:**
- Security reminders
- Protection from malicious software
- Login monitoring
- Password management
#### Security Incident Procedures (§164.308(a)(6))
**Incident Response Requirements:**
1. Identify and document incidents
2. Report security incidents
3. Respond to mitigate harmful effects
4. Document outcomes
#### Contingency Plan (§164.308(a)(7))
```markdown
## Contingency Plan Components
### Data Backup Plan (Required)
- Backup frequency: _____
- Backup verification: _____
- Off-site storage: _____
### Disaster Recovery Plan (Required)
- Recovery time objective: _____
- Recovery point objective: _____
- Recovery procedures: _____
### Emergency Mode Operation (Required)
- Critical functions: _____
- Manual procedures: _____
- Communication plan: _____
### Testing and Revision (Addressable)
- Test frequency: _____
- Last test date: _____
- Revision history: _____
### Applications and Data Criticality (Addressable)
- Critical systems: _____
- Priority recovery order: _____
```
### Physical Safeguards
#### Facility Access Controls (§164.310(a)(1))
| Specification | Type | Implementation |
|---------------|------|----------------|
| Contingency operations | Addressable | Physical access during emergency |
| Facility security plan | Addressable | Physical access policies |
| Access control/validation | Addressable | Visitor management |
| Maintenance records | Addressable | Physical maintenance logs |
#### Workstation Use (§164.310(b))
**Requirements:**
- Policies for workstation use
- Physical environment considerations
- Secure positioning
- Screen privacy
#### Workstation Security (§164.310(c))
**Physical Safeguards:**
- Cable locks
- Restricted areas
- Surveillance
- Clean desk policy
#### Device and Media Controls (§164.310(d)(1))
**Critical for Medical Devices:**
```markdown
## Device and Media Controls
### Disposal (Required)
- [ ] Wipe procedures for devices with ePHI
- [ ] Certificate of destruction
- [ ] Media sanitization per NIST 800-88
### Media Re-use (Required)
- [ ] Sanitization before re-use
- [ ] Verification of removal
- [ ] Documentation
### Accountability (Addressable)
- [ ] Hardware inventory
- [ ] Movement tracking
- [ ] Responsibility assignment
### Data Backup and Storage (Addressable)
- [ ] Retrievable copies
- [ ] Secure storage location
- [ ] Access controls on backup media
```
### Technical Safeguards
#### Access Control (§164.312(a)(1))
| Specification | Type | Implementation |
|---------------|------|----------------|
| Unique user identification | Required | Individual accounts |
| Emergency access | Required | Break-glass procedures |
| Automatic logoff | Addressable | Session timeout |
| Encryption and decryption | Addressable | At-rest encryption |
#### Audit Controls (§164.312(b))
**Audit Log Contents:**
- User identification
- Event type
- Date and time
- Success/failure
- Affected data
**Medical Device Considerations:**
- Log all access to patient data
- Protect logs from tampering
- Retain logs per policy (minimum 6 years)
- Real-time alerting for critical events
#### Integrity (§164.312(c)(1))
**ePHI Integrity Controls:**
- Hash verification
- Digital signatures
- Version control
- Change detection
#### Person or Entity Authentication (§164.312(d))
**Authentication Methods:**
- Passwords (strong requirements)
- Biometrics
- Hardware tokens
- Multi-factor authentication (recommended)
#### Transmission Security (§164.312(e)(1))
| Specification | Type | Implementation |
|---------------|------|----------------|
| Integrity controls | Addressable | TLS, message authentication |
| Encryption | Addressable | TLS 1.2+, AES-256 |
---
## Medical Device Considerations
### Connected Medical Device Security
**Data Flow Analysis:**
```
Device → Local Network → Internet → Cloud → EHR
│ │ │ │ │
└─ ePHI at rest ePHI in transit ePHI at rest
Encrypt Encrypt TLS Encrypt + Access Control
```
### SaMD (Software as a Medical Device)
**HIPAA Requirements for SaMD:**
1. Encryption of stored patient data
2. Secure authentication
3. Audit logging
4. Access controls
5. Secure communication protocols
6. Backup and recovery
7. Incident response
### Mobile Medical Applications
**Additional Considerations:**
- Device loss/theft protection
- Remote wipe capability
- App sandboxing
- Secure data storage
- API security
### Cloud-Based Devices
**Cloud Provider Requirements:**
- BAA with cloud provider
- Data residency (US only for HIPAA)
- Encryption key management
- Audit log access
- Incident notification
---
## Risk Assessment
### HIPAA Risk Assessment Process
```
Step 1: Scope Definition
├── Identify systems with ePHI
├── Document data flows
└── Identify business associates
Step 2: Threat Identification
├── Natural threats (fire, flood)
├── Human threats (hackers, insiders)
├── Environmental threats (power, HVAC)
└── Technical threats (malware, system failure)
Step 3: Vulnerability Assessment
├── Administrative controls
├── Physical controls
├── Technical controls
└── Gap analysis
Step 4: Risk Analysis
├── Likelihood assessment
├── Impact assessment
├── Risk level determination
└── Risk prioritization
Step 5: Risk Treatment
├── Accept
├── Mitigate
├── Transfer
└── Avoid
Step 6: Documentation
├── Risk register
├── Risk management plan
└── Remediation tracking
```
### Risk Assessment Template
```markdown
## HIPAA Risk Assessment
### System Information
System Name: _____________________
System Owner: ____________________
Date: ___________________________
### Asset Inventory
| Asset | ePHI Type | Location | Classification |
|-------|-----------|----------|----------------|
| | | | |
### Threat Analysis
| Threat | Likelihood (1-5) | Impact (1-5) | Risk Score |
|--------|------------------|--------------|------------|
| | | | |
### Vulnerability Assessment
| Safeguard Category | Gap Identified | Severity | Remediation |
|--------------------|----------------|----------|-------------|
| Administrative | | | |
| Physical | | | |
| Technical | | | |
### Risk Treatment Plan
| Risk | Treatment | Owner | Timeline | Status |
|------|-----------|-------|----------|--------|
| | | | | |
### Approval
Risk Assessment Approved: _______________ Date: _______
Next Assessment Due: _______________
```
---
## Implementation Specifications
### Required vs. Addressable
**Required:** Must be implemented as specified
**Addressable:**
1. Implement as specified, OR
2. Implement alternative measure, OR
3. Not implement if not reasonable and appropriate (document rationale)
### Implementation Status Matrix
| Safeguard | Specification | Type | Status | Evidence |
|-----------|---------------|------|--------|----------|
| §164.308(a)(1)(ii)(A) | Risk analysis | R | ☐ | |
| §164.308(a)(1)(ii)(B) | Risk management | R | ☐ | |
| §164.308(a)(3)(ii)(A) | Authorization/supervision | A | ☐ | |
| §164.308(a)(5)(ii)(A) | Security reminders | A | ☐ | |
| §164.310(a)(2)(i) | Contingency operations | A | ☐ | |
| §164.310(d)(2)(i) | Disposal | R | ☐ | |
| §164.312(a)(2)(i) | Unique user ID | R | ☐ | |
| §164.312(a)(2)(ii) | Emergency access | R | ☐ | |
| §164.312(a)(2)(iv) | Encryption (at rest) | A | ☐ | |
| §164.312(e)(2)(ii) | Encryption (transit) | A | ☐ | |
---
## Business Associate Agreements
### When Required
BAA required when business associate:
- Creates, receives, maintains, or transmits PHI
- Provides services involving PHI use/disclosure
### BAA Requirements
**Required Provisions:**
1. Permitted and required uses of PHI
2. Subcontractor requirements
3. Appropriate safeguards
4. Breach notification
5. Termination provisions
6. Return or destruction of PHI
### Medical Device Manufacturer BAA Template
```markdown
## Business Associate Agreement
This Agreement is entered into as of [Date] between:
COVERED ENTITY: [Healthcare Provider/Plan Name]
BUSINESS ASSOCIATE: [Device Manufacturer Name]
### 1. Definitions
[Standard HIPAA definitions]
### 2. Obligations of Business Associate
Business Associate agrees to:
a) Not use or disclose PHI other than as permitted
b) Use appropriate safeguards to prevent improper use/disclosure
c) Report any security incident or breach
d) Ensure subcontractors agree to same restrictions
e) Make PHI available for individual access
f) Make PHI available for amendment
g) Document and make available disclosures
h) Make internal practices available to HHS
i) Return or destroy PHI at termination
### 3. Permitted Uses and Disclosures
Business Associate may:
a) Use PHI for device operation and maintenance
b) Use PHI for quality improvement
c) De-identify PHI per HIPAA standards
d) Create aggregate data
e) Report to FDA as required
### 4. Security Requirements
Business Associate shall implement:
a) Administrative safeguards per §164.308
b) Physical safeguards per §164.310
c) Technical safeguards per §164.312
### 5. Breach Notification
Business Associate shall:
a) Report breaches within [60 days/contractual period]
b) Provide information for breach notification
c) Mitigate harmful effects
### 6. Term and Termination
[Standard termination provisions]
### Signatures
COVERED ENTITY: _________________ Date: _______
BUSINESS ASSOCIATE: _____________ Date: _______
```
---
## Breach Notification
### Breach Definition
**Breach:** Acquisition, access, use, or disclosure of unsecured PHI in a manner not permitted that compromises security or privacy.
**Exceptions:**
1. Unintentional acquisition by workforce member acting in good faith
2. Inadvertent disclosure between authorized persons
3. Good faith belief that unauthorized person couldn't retain information
### Risk Assessment for Breach
**Factors to Consider:**
1. Nature and extent of PHI involved
2. Unauthorized person who received PHI
3. Whether PHI was actually acquired/viewed
4. Extent to which risk has been mitigated
### Notification Requirements
| Audience | Timing | Method |
|----------|--------|--------|
| Individuals | 60 days from discovery | First-class mail or email |
| HHS | 60 days (if >500) | HHS breach portal |
| HHS | Annual (if <500) | Annual report |
| Media | 60 days (if >500 in state) | Prominent media outlet |
### Breach Response Procedure
```markdown
## Breach Response Procedure
### Phase 1: Detection and Containment (Immediate)
- [ ] Identify scope of breach
- [ ] Contain breach (stop ongoing access)
- [ ] Preserve evidence
- [ ] Notify incident response team
- [ ] Document timeline
### Phase 2: Investigation (1-14 days)
- [ ] Determine what PHI was involved
- [ ] Identify affected individuals
- [ ] Assess risk of harm
- [ ] Document investigation findings
### Phase 3: Risk Assessment (15-30 days)
- [ ] Apply four-factor risk assessment
- [ ] Determine if notification required
- [ ] Document decision rationale
### Phase 4: Notification (Within 60 days)
- [ ] Prepare individual notification letters
- [ ] Submit to HHS (if required)
- [ ] Media notification (if required)
- [ ] Retain copies of notifications
### Phase 5: Remediation (Ongoing)
- [ ] Implement corrective actions
- [ ] Update policies and procedures
- [ ] Train workforce
- [ ] Monitor for additional impact
```
### Breach Notification Content
**Individual Notification Must Include:**
1. Description of what happened
2. Types of PHI involved
3. Steps individuals should take
4. What entity is doing to investigate
5. What entity is doing to prevent future breaches
6. Contact information for questions
---
## Compliance Checklist
### Administrative Safeguards Checklist
```markdown
## Administrative Safeguards
- [ ] Security Management Process
- [ ] Risk analysis completed and documented
- [ ] Risk management plan in place
- [ ] Sanction policy documented
- [ ] Information system activity review conducted
- [ ] Assigned Security Responsibility
- [ ] Security Officer designated
- [ ] Contact information documented
- [ ] Workforce Security
- [ ] Authorization procedures
- [ ] Background checks (if applicable)
- [ ] Termination procedures
- [ ] Information Access Management
- [ ] Access authorization policies
- [ ] Access establishment procedures
- [ ] Access modification procedures
- [ ] Security Awareness and Training
- [ ] Training program established
- [ ] Security reminders distributed
- [ ] Protection from malicious software training
- [ ] Password management training
- [ ] Security Incident Procedures
- [ ] Incident response plan
- [ ] Incident documentation procedures
- [ ] Reporting mechanisms
- [ ] Contingency Plan
- [ ] Data backup plan
- [ ] Disaster recovery plan
- [ ] Emergency mode operation plan
- [ ] Testing and revision procedures
```
### Technical Safeguards Checklist
```markdown
## Technical Safeguards
- [ ] Access Control
- [ ] Unique user identification
- [ ] Emergency access procedure
- [ ] Automatic logoff
- [ ] Encryption (at rest)
- [ ] Audit Controls
- [ ] Audit logging implemented
- [ ] Log review procedures
- [ ] Log retention policy
- [ ] Integrity
- [ ] Mechanism to authenticate ePHI
- [ ] Integrity controls in place
- [ ] Authentication
- [ ] Person/entity authentication
- [ ] Strong password policy
- [ ] Transmission Security
- [ ] Integrity controls (in transit)
- [ ] Encryption (TLS 1.2+)
```
---
## Quick Reference
### Common HIPAA Violations
| Violation | Prevention |
|-----------|------------|
| Unauthorized access | Role-based access, MFA |
| Lost/stolen devices | Encryption, remote wipe |
| Improper disposal | NIST 800-88 sanitization |
| Insufficient training | Annual training program |
| Missing BAAs | BA inventory and tracking |
| Insufficient audit logs | Comprehensive logging |
### Penalty Structure
| Tier | Knowledge | Per Violation | Annual Maximum |
|------|-----------|---------------|----------------|
| 1 | Unknown | $100-$50,000 | $1,500,000 |
| 2 | Reasonable cause | $1,000-$50,000 | $1,500,000 |
| 3 | Willful neglect (corrected) | $10,000-$50,000 | $1,500,000 |
| 4 | Willful neglect (not corrected) | $50,000 | $1,500,000 |
### FDA-HIPAA Intersection
| Device Scenario | FDA | HIPAA |
|-----------------|-----|-------|
| Standalone diagnostic | 510(k)/PMA | If transmits PHI |
| Connected insulin pump | Class III PMA | Yes (patient data) |
| Wellness app (no diagnosis) | Exempt | If stores PHI |
| EHR-integrated device | May apply | Yes |
| Research device | IDE | IRB may waive |
FILE:references/qsr_compliance_requirements.md
# Quality System Regulation (QSR) Compliance
Complete guide to 21 CFR Part 820 requirements for medical device manufacturers.
---
## Table of Contents
- [QSR Overview](#qsr-overview)
- [Management Responsibility (820.20)](#management-responsibility-82020)
- [Design Controls (820.30)](#design-controls-82030)
- [Document Controls (820.40)](#document-controls-82040)
- [Purchasing Controls (820.50)](#purchasing-controls-82050)
- [Production and Process Controls (820.70-75)](#production-and-process-controls-82070-75)
- [CAPA (820.100)](#capa-820100)
- [Device Master Record (820.181)](#device-master-record-820181)
- [FDA Inspection Readiness](#fda-inspection-readiness)
---
## QSR Overview
### Applicability
The QSR applies to:
- Finished device manufacturers
- Specification developers
- Initial distributors of imported devices
- Contract manufacturers
- Repackagers and relabelers
### Exemptions
| Device Class | Exemption Status |
|--------------|------------------|
| Class I (most) | Exempt from design controls (820.30) |
| Class I (listed) | Fully exempt from QSR |
| Class II | Full QSR compliance |
| Class III | Full QSR compliance |
### QSR Structure
```
21 CFR Part 820 Subparts:
├── A - General Provisions (820.1-5)
├── B - Quality System Requirements (820.20-25)
├── C - Design Controls (820.30)
├── D - Document Controls (820.40)
├── E - Purchasing Controls (820.50)
├── F - Identification and Traceability (820.60-65)
├── G - Production and Process Controls (820.70-75)
├── H - Acceptance Activities (820.80-86)
├── I - Nonconforming Product (820.90)
├── J - Corrective and Preventive Action (820.100)
├── K - Labeling and Packaging Control (820.120-130)
├── L - Handling, Storage, Distribution, Installation (820.140-170)
├── M - Records (820.180-198)
├── N - Servicing (820.200)
└── O - Statistical Techniques (820.250)
```
---
## Management Responsibility (820.20)
### Quality Policy
**Requirements:**
- Documented quality policy
- Objectives for quality
- Commitment to meeting requirements
- Communicated throughout organization
**Quality Policy Template:**
```markdown
## Quality Policy Statement
[Company Name] is committed to designing, manufacturing, and distributing
medical devices that meet customer requirements and applicable regulatory
standards. We achieve this through:
1. Maintaining an effective Quality Management System
2. Continuous improvement of our processes
3. Compliance with 21 CFR Part 820 and applicable standards
4. Training and empowering employees
5. Supplier quality management
Approved by: _______________ Date: _______________
Management Representative
```
### Organization
| Role | Responsibilities | Documentation |
|------|------------------|---------------|
| Management Representative | QMS oversight, FDA liaison | Org chart, job description |
| Quality Manager | Day-to-day QMS operations | Procedures, authority matrix |
| Design Authority | Design control decisions | DHF sign-offs |
| Production Manager | Manufacturing compliance | Process documentation |
### Management Review
**Frequency:** At least annually (more frequently recommended)
**Required Inputs:**
1. Audit results (internal and external)
2. Customer feedback and complaints
3. Process performance metrics
4. Product conformity data
5. CAPA status
6. Changes affecting QMS
7. Recommendations for improvement
**Required Outputs:**
- Decisions on improvement actions
- Resource needs
- Quality objectives updates
**Management Review Agenda Template:**
```markdown
## Management Review Meeting
Date: _______________
Attendees: _______________
### Agenda Items
1. Review of previous action items
2. Quality objectives and metrics
3. Internal audit results
4. Customer complaints summary
5. CAPA status report
6. Supplier quality performance
7. Regulatory updates
8. Resource requirements
9. Improvement opportunities
### Decisions and Actions
| Item | Decision | Owner | Due Date |
|------|----------|-------|----------|
| | | | |
### Next Review Date: _______________
```
---
## Design Controls (820.30)
### When Required
Design controls are required for:
- Class II devices (most)
- Class III devices (all)
- Class I devices with software
- Class I devices on exemption list exceptions
### Design Control Process Flow
```
Design Input (820.30c)
↓
Design Output (820.30d)
↓
Design Review (820.30e)
↓
Design Verification (820.30f)
↓
Design Validation (820.30g)
↓
Design Transfer (820.30h)
↓
Design Changes (820.30i)
↓
Design History File (820.30j)
```
### Design Input Requirements
**Must Include:**
- Intended use and user requirements
- Patient population
- Performance requirements
- Safety requirements
- Regulatory requirements
- Risk management requirements
**Verification Criteria:**
- Complete (all requirements captured)
- Unambiguous (clear interpretation)
- Not conflicting
- Verifiable or validatable
### Design Output Requirements
| Output Type | Examples | Verification Method |
|-------------|----------|---------------------|
| Device specifications | Drawings, BOMs | Inspection, testing |
| Manufacturing specs | Process parameters | Process validation |
| Software specs | Source code, architecture | Software V&V |
| Labeling | IFU, labels | Review against inputs |
**Essential Requirements:**
- Traceable to design inputs
- Contains acceptance criteria
- Identifies critical characteristics
### Design Review
**Review Stages:**
1. Concept review (feasibility)
2. Design input review (requirements complete)
3. Preliminary design review (architecture)
4. Critical design review (detailed design)
5. Final design review (transfer readiness)
**Participants:**
- Representative of each design function
- Other specialists as needed
- Independent reviewers (no direct design responsibility)
**Documentation:**
- Meeting minutes
- Issues identified
- Resolution actions
- Approval signatures
### Design Verification
**Methods:**
- Inspections and measurements
- Bench testing
- Analysis and calculations
- Simulations
- Comparisons to similar designs
**Verification Matrix Template:**
```markdown
| Req ID | Requirement | Verification Method | Pass Criteria | Result |
|--------|-------------|---------------------|---------------|--------|
| REQ-001 | Dimension tolerance | Measurement | ±0.5mm | |
| REQ-002 | Tensile strength | Testing per ASTM | >500 MPa | |
| REQ-003 | Software function | Unit testing | 100% pass | |
```
### Design Validation
**Definition:** Confirmation that device meets user needs and intended uses
**Validation Requirements:**
- Use initial production units (or equivalent)
- Simulated or actual use conditions
- Includes software validation
**Validation Types:**
1. **Bench validation** - Laboratory simulated use
2. **Clinical validation** - Human subjects (may require IDE)
3. **Usability validation** - Human factors testing
### Design Transfer
**Transfer Checklist:**
```markdown
## Design Transfer Verification
- [ ] DMR complete and approved
- [ ] Manufacturing processes validated
- [ ] Training completed
- [ ] Inspection procedures established
- [ ] Supplier qualifications complete
- [ ] Labeling approved
- [ ] Risk analysis updated
- [ ] Regulatory clearance/approval obtained
```
### Design History File (DHF)
**Contents:**
- Design and development plan
- Design input records
- Design output records
- Design review records
- Design verification records
- Design validation records
- Design transfer records
- Design change records
- Risk management file
---
## Document Controls (820.40)
### Document Approval and Distribution
**Requirements:**
- Documents reviewed and approved before use
- Approved documents available at point of use
- Obsolete documents removed or marked
- Changes reviewed and approved
### Document Control Matrix
| Document Type | Author | Reviewer | Approver | Distribution |
|---------------|--------|----------|----------|--------------|
| SOPs | Process owner | QA | Quality Manager | Controlled |
| Work Instructions | Supervisor | QA | Manager | Controlled |
| Forms | QA | QA | Quality Manager | Controlled |
| Drawings | Engineer | Peer | Design Authority | Controlled |
### Change Control
**Change Request Process:**
```
1. Initiate Change Request
└── Description, justification, impact assessment
2. Technical Review
└── Engineering, quality, regulatory assessment
3. Change Classification
├── Minor: No regulatory impact
├── Moderate: May affect compliance
└── Major: Regulatory submission required
4. Approval
└── Change Control Board (CCB) or designated authority
5. Implementation
└── Training, document updates, inventory actions
6. Verification
└── Confirm change implemented correctly
7. Close Change Request
└── Documentation complete
```
---
## Purchasing Controls (820.50)
### Supplier Qualification
**Qualification Criteria:**
- Quality system capability
- Product/service quality history
- Financial stability
- Regulatory compliance history
**Qualification Methods:**
| Method | When Used | Documentation |
|--------|-----------|---------------|
| On-site audit | Critical suppliers, high risk | Audit report |
| Questionnaire | Initial screening | Completed form |
| Certification review | ISO certified suppliers | Cert copies |
| Product qualification | Incoming inspection data | Test results |
### Approved Supplier List (ASL)
**ASL Requirements:**
- Supplier name and contact
- Products/services approved
- Qualification date and method
- Qualification status
- Re-evaluation schedule
### Purchasing Data
**Purchase Order Requirements:**
- Complete product specifications
- Quality requirements
- Applicable standards
- Inspection/acceptance requirements
- Right of access for verification
---
## Production and Process Controls (820.70-75)
### Process Validation (820.75)
**When Required:**
- Process output cannot be fully verified
- Deficiencies would only appear after use
- Examples: sterilization, welding, molding
**Validation Protocol Elements:**
```markdown
## Process Validation Protocol
### 1. Protocol Approval
Prepared by: _______________ Date: _______________
Approved by: _______________ Date: _______________
### 2. Process Description
[Describe process, equipment, materials, parameters]
### 3. Acceptance Criteria
| Parameter | Specification | Test Method |
|-----------|---------------|-------------|
| | | |
### 4. Equipment Qualification
- IQ (Installation Qualification): _______________
- OQ (Operational Qualification): _______________
- PQ (Performance Qualification): _______________
### 5. Validation Runs
Number of runs: _____ (minimum 3)
Lot sizes: _____
### 6. Results Summary
| Run | Date | Parameters | Results | Pass/Fail |
|-----|------|------------|---------|-----------|
| 1 | | | | |
| 2 | | | | |
| 3 | | | | |
### 7. Conclusion
Process validated: Yes / No
Revalidation triggers: _____
```
### Environmental Controls (820.70(c))
**Controlled Conditions:**
- Temperature and humidity
- Particulate contamination (cleanrooms)
- ESD (electrostatic discharge)
- Lighting levels
**Monitoring Requirements:**
- Continuous or periodic monitoring
- Documented limits
- Out-of-specification procedures
- Calibrated equipment
### Personnel (820.70(d))
**Training Requirements:**
- Job-specific training
- Competency verification
- Retraining for significant changes
- Training records maintained
**Training Record Template:**
```markdown
## Training Record
Employee: _______________ ID: _______________
Position: _______________
| Training Topic | Trainer | Date | Method | Competency Verified |
|----------------|---------|------|--------|---------------------|
| | | | | Signature: ________ |
```
### Equipment (820.70(g))
**Requirements:**
- Maintenance schedule
- Calibration program
- Adjustment limits documented
- Inspection before use
### Calibration (820.72)
**Calibration Program Elements:**
1. Equipment identification
2. Calibration frequency
3. Calibration procedures
4. Accuracy requirements
5. Traceability to NIST standards
6. Out-of-tolerance actions
---
## CAPA (820.100)
### CAPA Sources
- Customer complaints
- Nonconforming product
- Audit findings
- Process monitoring
- Returned products
- MDR/Vigilance reports
- Trend analysis
### CAPA Process
```
1. Identification
└── Problem statement, data collection
2. Investigation
└── Root cause analysis (5 Whys, Fishbone, etc.)
3. Action Determination
├── Correction: Immediate fix
└── Corrective/Preventive: Address root cause
4. Implementation
└── Action execution, documentation
5. Verification
└── Confirm actions completed
6. Effectiveness Review
└── Problem recurrence check (30-90 days)
7. Closure
└── Management approval
```
### Root Cause Analysis Tools
**5 Whys Example:**
```
Problem: Device failed during use
Why 1: Component failed
Why 2: Component was out of specification
Why 3: Incoming inspection did not detect
Why 4: Inspection procedure inadequate
Why 5: Procedure not updated for new component
Root Cause: Document control failure - procedure not updated
```
**Fishbone Categories:**
- Man (People)
- Machine (Equipment)
- Method (Process)
- Material
- Measurement
- Environment
### CAPA Metrics
| Metric | Target | Frequency |
|--------|--------|-----------|
| CAPA on-time closure | >90% | Monthly |
| Overdue CAPAs | <5 | Monthly |
| Effectiveness rate | >85% | Quarterly |
| Average days to closure | <60 | Monthly |
---
## Device Master Record (820.181)
### DMR Contents
```
Device Master Record
├── Device specifications
│ ├── Drawings
│ ├── Composition/formulation
│ └── Component specifications
├── Production process specifications
│ ├── Manufacturing procedures
│ ├── Assembly instructions
│ └── Process parameters
├── Quality assurance procedures
│ ├── Acceptance criteria
│ ├── Inspection procedures
│ └── Test methods
├── Packaging and labeling specifications
│ ├── Package drawings
│ ├── Label content
│ └── IFU content
├── Installation, maintenance, servicing procedures
└── Environmental requirements
```
### Device History Record (DHR) - 820.184
**DHR Contents:**
- Dates of manufacture
- Quantity manufactured
- Quantity released for distribution
- Acceptance records
- Primary identification label
- Device identification and control numbers
### Quality System Record (QSR) - 820.186
**QSR Contents:**
- Procedures and changes
- Calibration records
- Distribution records
- Complaint files
- CAPA records
- Audit reports
---
## FDA Inspection Readiness
### Pre-Inspection Preparation
**30-Day Readiness Checklist:**
```markdown
## FDA Inspection Readiness
### Documentation Review
- [ ] Quality manual current
- [ ] SOPs reviewed and approved
- [ ] Training records complete
- [ ] CAPA files complete
- [ ] Complaint files organized
- [ ] DMR/DHR accessible
- [ ] Management review records current
### Facility Review
- [ ] Controlled areas properly identified
- [ ] Equipment calibration current
- [ ] Environmental monitoring records available
- [ ] Storage conditions appropriate
- [ ] Quarantine areas clearly marked
### Personnel Preparation
- [ ] Escort team identified
- [ ] Subject matter experts briefed
- [ ] Front desk/reception notified
- [ ] Conference room reserved
- [ ] FDA credentials verification process
### Record Accessibility
- [ ] Electronic records accessible
- [ ] Backup copies available
- [ ] Audit trail functional
- [ ] Archive records retrievable
```
### During Inspection
**Escort Guidelines:**
1. One designated escort with investigator at all times
2. Answer questions truthfully and concisely
3. Don't volunteer information not requested
4. Request clarification if question unclear
5. Get help from SME for technical questions
6. Document all requests and commitments
**Record Request Tracking:**
| Request # | Date | Document Requested | Provided By | Date Provided |
|-----------|------|-------------------|-------------|---------------|
| | | | | |
### Post-Inspection
**FDA 483 Response:**
- Due within 15 business days
- Address each observation specifically
- Include corrective actions and timeline
- Provide evidence of completion where possible
**Response Format:**
```markdown
## Observation [Number]
### FDA Observation:
[Copy verbatim from Form 483]
### Company Response:
#### Understanding of Observation:
[Demonstrate understanding of the concern]
#### Immediate Correction:
[Actions already taken]
#### Root Cause Analysis:
[Investigation findings]
#### Corrective Actions:
| Action | Responsible | Target Date | Status |
|--------|-------------|-------------|--------|
| | | | |
#### Preventive Actions:
[Systemic improvements]
#### Verification:
[How effectiveness will be verified]
```
---
## Compliance Metrics Dashboard
### Key Performance Indicators
| Category | Metric | Target | Current |
|----------|--------|--------|---------|
| CAPA | On-time closure rate | >90% | |
| CAPA | Effectiveness rate | >85% | |
| Complaints | Response time (days) | <5 | |
| Training | Compliance rate | 100% | |
| Calibration | On-time rate | 100% | |
| Audit | Findings closure rate | >95% | |
| NCR | Recurring issues | <5% | |
| Supplier | Quality rate | >98% | |
### Trend Analysis
**Monthly Review Items:**
- Complaint trends by product/failure mode
- NCR trends by cause code
- CAPA effectiveness
- Supplier quality
- Production yields
- Customer feedback
---
## Quick Reference
### Common 483 Observations
| Observation | Prevention |
|-------------|------------|
| CAPA not effective | Verify effectiveness before closure |
| Training incomplete | Competency-based training records |
| Document control gaps | Regular procedure reviews |
| Complaint investigation | Thorough, documented investigations |
| Supplier controls weak | Robust qualification and monitoring |
| Validation inadequate | Follow IQ/OQ/PQ protocols |
### Regulatory Cross-References
| QSR Section | ISO 13485 Clause |
|-------------|------------------|
| 820.20 | 5.1, 5.5, 5.6 |
| 820.30 | 7.3 |
| 820.40 | 4.2.4 |
| 820.50 | 7.4 |
| 820.70 | 7.5.1 |
| 820.75 | 7.5.6 |
| 820.100 | 8.5.2, 8.5.3 |
FILE:scripts/fda_submission_tracker.py
#!/usr/bin/env python3
"""
FDA Submission Tracker
Tracks FDA submission status, calculates timelines, and monitors regulatory milestones
for 510(k), De Novo, and PMA submissions.
Usage:
python fda_submission_tracker.py <project_dir>
python fda_submission_tracker.py <project_dir> --type 510k
python fda_submission_tracker.py <project_dir> --json
"""
import argparse
import json
import os
import sys
from datetime import datetime, timedelta
from pathlib import Path
from typing import Dict, List, Optional, Any
# FDA review timeline targets (calendar days)
FDA_TIMELINES = {
"510k_traditional": {
"acceptance_review": 15,
"substantive_review": 90,
"total_goal": 90,
"ai_response": 180 # Days to respond to Additional Information
},
"510k_special": {
"acceptance_review": 15,
"substantive_review": 30,
"total_goal": 30,
"ai_response": 180
},
"510k_abbreviated": {
"acceptance_review": 15,
"substantive_review": 30,
"total_goal": 30,
"ai_response": 180
},
"de_novo": {
"acceptance_review": 60,
"substantive_review": 150,
"total_goal": 150,
"ai_response": 180
},
"pma": {
"acceptance_review": 45,
"substantive_review": 180,
"total_goal": 180,
"ai_response": 180
},
"pma_supplement": {
"acceptance_review": 15,
"substantive_review": 180,
"total_goal": 180,
"ai_response": 180
}
}
# Submission milestones by type
MILESTONES = {
"510k": [
{"id": "predicate_identified", "name": "Predicate Device Identified", "phase": "planning"},
{"id": "testing_complete", "name": "Performance Testing Complete", "phase": "preparation"},
{"id": "documentation_complete", "name": "Submission Documentation Complete", "phase": "preparation"},
{"id": "submission_sent", "name": "Submission Sent to FDA", "phase": "submission"},
{"id": "acknowledgment_received", "name": "FDA Acknowledgment Received", "phase": "review"},
{"id": "acceptance_decision", "name": "Acceptance Review Complete", "phase": "review"},
{"id": "ai_request", "name": "Additional Information Request", "phase": "review", "optional": True},
{"id": "ai_response", "name": "AI Response Submitted", "phase": "review", "optional": True},
{"id": "se_decision", "name": "Substantial Equivalence Decision", "phase": "decision"},
{"id": "clearance_letter", "name": "510(k) Clearance Letter Received", "phase": "decision"}
],
"de_novo": [
{"id": "classification_determined", "name": "Classification Determination", "phase": "planning"},
{"id": "special_controls_defined", "name": "Special Controls Defined", "phase": "preparation"},
{"id": "risk_assessment_complete", "name": "Risk Assessment Complete", "phase": "preparation"},
{"id": "testing_complete", "name": "Performance Testing Complete", "phase": "preparation"},
{"id": "submission_sent", "name": "Submission Sent to FDA", "phase": "submission"},
{"id": "acknowledgment_received", "name": "FDA Acknowledgment Received", "phase": "review"},
{"id": "acceptance_decision", "name": "Acceptance Review Complete", "phase": "review"},
{"id": "ai_request", "name": "Additional Information Request", "phase": "review", "optional": True},
{"id": "ai_response", "name": "AI Response Submitted", "phase": "review", "optional": True},
{"id": "classification_decision", "name": "De Novo Classification Decision", "phase": "decision"}
],
"pma": [
{"id": "ide_approved", "name": "IDE Approval (if required)", "phase": "planning", "optional": True},
{"id": "clinical_complete", "name": "Clinical Study Complete", "phase": "preparation"},
{"id": "clinical_report_complete", "name": "Clinical Study Report Complete", "phase": "preparation"},
{"id": "documentation_complete", "name": "PMA Documentation Complete", "phase": "preparation"},
{"id": "submission_sent", "name": "PMA Submission Sent to FDA", "phase": "submission"},
{"id": "acknowledgment_received", "name": "FDA Acknowledgment Received", "phase": "review"},
{"id": "filing_decision", "name": "Filing Decision", "phase": "review"},
{"id": "ai_request", "name": "Major Deficiency Letter", "phase": "review", "optional": True},
{"id": "ai_response", "name": "Deficiency Response Submitted", "phase": "review", "optional": True},
{"id": "panel_meeting", "name": "Advisory Committee Meeting", "phase": "review", "optional": True},
{"id": "approval_decision", "name": "PMA Approval Decision", "phase": "decision"}
]
}
def find_submission_config(project_dir: Path) -> Optional[Dict]:
"""Find and load submission configuration file."""
config_paths = [
project_dir / "fda_submission.json",
project_dir / "regulatory" / "fda_submission.json",
project_dir / ".fda" / "submission.json"
]
for config_path in config_paths:
if config_path.exists():
try:
with open(config_path) as f:
return json.load(f)
except json.JSONDecodeError:
continue
return None
def calculate_timeline_status(submission_type: str, milestones: Dict[str, str]) -> Dict:
"""Calculate timeline status based on submission type and milestone dates."""
timeline_config = FDA_TIMELINES.get(submission_type, FDA_TIMELINES["510k_traditional"])
result = {
"submission_type": submission_type,
"timeline_config": timeline_config,
"status": "not_started",
"days_elapsed": 0,
"days_remaining": None,
"projected_decision_date": None,
"on_track": None
}
# Check if submission has been sent
if "submission_sent" in milestones:
try:
submission_date = datetime.strptime(milestones["submission_sent"], "%Y-%m-%d")
today = datetime.now()
result["days_elapsed"] = (today - submission_date).days
# Check for AI hold
ai_hold_days = 0
if "ai_request" in milestones and "ai_response" in milestones:
ai_request_date = datetime.strptime(milestones["ai_request"], "%Y-%m-%d")
ai_response_date = datetime.strptime(milestones["ai_response"], "%Y-%m-%d")
ai_hold_days = (ai_response_date - ai_request_date).days
elif "ai_request" in milestones and "ai_response" not in milestones:
ai_request_date = datetime.strptime(milestones["ai_request"], "%Y-%m-%d")
ai_hold_days = (today - ai_request_date).days
result["status"] = "ai_hold"
# Calculate review days (excluding AI hold)
review_days = result["days_elapsed"] - ai_hold_days
# Determine status
if "se_decision" in milestones or "approval_decision" in milestones or "classification_decision" in milestones:
result["status"] = "complete"
elif "acceptance_decision" in milestones:
result["status"] = "substantive_review"
elif "acknowledgment_received" in milestones:
result["status"] = "acceptance_review"
else:
result["status"] = "submitted"
# Calculate projected decision date
if result["status"] not in ["complete", "ai_hold"]:
goal_days = timeline_config["total_goal"]
result["days_remaining"] = max(0, goal_days - review_days)
result["projected_decision_date"] = (submission_date + timedelta(days=goal_days + ai_hold_days)).strftime("%Y-%m-%d")
result["on_track"] = review_days <= goal_days
except ValueError:
pass
return result
def analyze_milestone_status(submission_type: str, completed_milestones: Dict[str, str]) -> List[Dict]:
"""Analyze milestone completion status."""
milestone_list = MILESTONES.get(submission_type.split("_")[0], MILESTONES["510k"])
results = []
for milestone in milestone_list:
status = {
"id": milestone["id"],
"name": milestone["name"],
"phase": milestone["phase"],
"optional": milestone.get("optional", False),
"completed": milestone["id"] in completed_milestones,
"completion_date": completed_milestones.get(milestone["id"])
}
results.append(status)
return results
def calculate_submission_readiness(project_dir: Path, submission_type: str) -> Dict:
"""Check submission readiness by looking for required documentation."""
required_docs = {
"510k": [
{"name": "Device Description", "patterns": ["device_description*", "device_desc*"]},
{"name": "Indications for Use", "patterns": ["indications*", "ifu*"]},
{"name": "Substantial Equivalence", "patterns": ["substantial_equiv*", "se_comparison*", "predicate*"]},
{"name": "Performance Testing", "patterns": ["performance*", "test_report*", "bench_test*"]},
{"name": "Biocompatibility", "patterns": ["biocompat*", "iso_10993*"]},
{"name": "Labeling", "patterns": ["label*", "ifu*", "instructions*"]},
{"name": "Software Documentation", "patterns": ["software*", "iec_62304*"], "optional": True},
{"name": "Sterilization Validation", "patterns": ["steriliz*", "sterility*"], "optional": True}
],
"de_novo": [
{"name": "Device Description", "patterns": ["device_description*", "device_desc*"]},
{"name": "Risk Assessment", "patterns": ["risk*", "hazard*"]},
{"name": "Special Controls", "patterns": ["special_control*"]},
{"name": "Performance Testing", "patterns": ["performance*", "test_report*"]},
{"name": "Labeling", "patterns": ["label*", "ifu*"]}
],
"pma": [
{"name": "Device Description", "patterns": ["device_description*"]},
{"name": "Manufacturing Information", "patterns": ["manufacturing*", "production*"]},
{"name": "Clinical Study Report", "patterns": ["clinical*", "csr*"]},
{"name": "Nonclinical Testing", "patterns": ["nonclinical*", "bench*", "preclinical*"]},
{"name": "Risk Analysis", "patterns": ["risk*", "fmea*"]},
{"name": "Labeling", "patterns": ["label*", "ifu*"]}
]
}
docs_to_check = required_docs.get(submission_type.split("_")[0], required_docs["510k"])
# Search common documentation directories
doc_dirs = [
project_dir / "regulatory",
project_dir / "regulatory" / "fda",
project_dir / "docs",
project_dir / "documentation",
project_dir / "dhf",
project_dir
]
results = []
for doc in docs_to_check:
found = False
found_path = None
for doc_dir in doc_dirs:
if not doc_dir.exists():
continue
for pattern in doc["patterns"]:
matches = list(doc_dir.glob(f"**/{pattern}"))
matches.extend(list(doc_dir.glob(f"**/{pattern.upper()}")))
if matches:
found = True
found_path = str(matches[0].relative_to(project_dir))
break
if found:
break
results.append({
"name": doc["name"],
"required": not doc.get("optional", False),
"found": found,
"path": found_path
})
required_found = sum(1 for r in results if r["required"] and r["found"])
required_total = sum(1 for r in results if r["required"])
return {
"documents": results,
"required_complete": required_found,
"required_total": required_total,
"readiness_percentage": round((required_found / required_total) * 100, 1) if required_total > 0 else 0
}
def generate_sample_config() -> Dict:
"""Generate sample submission configuration."""
return {
"submission_type": "510k_traditional",
"device_name": "Example Medical Device",
"product_code": "ABC",
"predicate_device": {
"name": "Predicate Device Name",
"k_number": "K123456"
},
"milestones": {
"predicate_identified": "2024-01-15",
"testing_complete": "2024-03-01",
"documentation_complete": "2024-03-15"
},
"contacts": {
"regulatory_lead": "Name",
"quality_lead": "Name"
},
"notes": "Add milestone dates as they are completed"
}
def print_text_report(result: Dict) -> None:
"""Print human-readable report."""
print("=" * 60)
print("FDA SUBMISSION TRACKER REPORT")
print("=" * 60)
if "error" in result:
print(f"\nError: {result['error']}")
print(f"\nTo create a configuration file, run with --init")
return
# Basic info
print(f"\nDevice: {result.get('device_name', 'Unknown')}")
print(f"Submission Type: {result['submission_type']}")
print(f"Product Code: {result.get('product_code', 'N/A')}")
# Timeline status
timeline = result["timeline_status"]
print(f"\n--- Timeline Status ---")
print(f"Status: {timeline['status'].upper()}")
print(f"Days Elapsed: {timeline['days_elapsed']}")
if timeline["days_remaining"] is not None:
print(f"Days Remaining (FDA goal): {timeline['days_remaining']}")
if timeline["projected_decision_date"]:
print(f"Projected Decision Date: {timeline['projected_decision_date']}")
if timeline["on_track"] is not None:
status = "ON TRACK" if timeline["on_track"] else "BEHIND SCHEDULE"
print(f"Timeline Status: {status}")
# Milestones
print(f"\n--- Milestones ---")
for ms in result["milestones"]:
status = "[X]" if ms["completed"] else "[ ]"
optional = " (optional)" if ms["optional"] else ""
date = f" - {ms['completion_date']}" if ms["completion_date"] else ""
print(f" {status} {ms['name']}{optional}{date}")
# Readiness
if "readiness" in result:
print(f"\n--- Submission Readiness ---")
readiness = result["readiness"]
print(f"Readiness: {readiness['readiness_percentage']}% ({readiness['required_complete']}/{readiness['required_total']} required docs)")
print("\n Documents:")
for doc in readiness["documents"]:
status = "[X]" if doc["found"] else "[ ]"
req = "(required)" if doc["required"] else "(optional)"
path = f" - {doc['path']}" if doc["path"] else ""
print(f" {status} {doc['name']} {req}{path}")
# Recommendations
if result.get("recommendations"):
print(f"\n--- Recommendations ---")
for i, rec in enumerate(result["recommendations"], 1):
print(f" {i}. {rec}")
print("\n" + "=" * 60)
def generate_recommendations(result: Dict) -> List[str]:
"""Generate actionable recommendations based on status."""
recommendations = []
timeline = result["timeline_status"]
# Timeline recommendations
if timeline["status"] == "ai_hold":
recommendations.append("Priority: Respond to FDA Additional Information request within 180 days")
elif timeline["on_track"] is False:
recommendations.append("Warning: Submission is behind FDA review schedule - consider contacting FDA")
# Milestone recommendations
completed_phases = set()
for ms in result["milestones"]:
if ms["completed"]:
completed_phases.add(ms["phase"])
if "submission" not in completed_phases and "preparation" in completed_phases:
recommendations.append("Ready for submission: Documentation complete, proceed with FDA submission")
# Readiness recommendations
if "readiness" in result:
missing_required = [d for d in result["readiness"]["documents"] if d["required"] and not d["found"]]
if missing_required:
docs = ", ".join(d["name"] for d in missing_required[:3])
recommendations.append(f"Missing required documentation: {docs}")
return recommendations
def analyze_submission(project_dir: Path, submission_type: Optional[str] = None) -> Dict:
"""Main analysis function."""
# Try to find existing configuration
config = find_submission_config(project_dir)
if config is None:
# No config found - do basic analysis
sub_type = submission_type or "510k_traditional"
result = {
"submission_type": sub_type,
"config_found": False,
"timeline_status": calculate_timeline_status(sub_type, {}),
"milestones": analyze_milestone_status(sub_type, {}),
"readiness": calculate_submission_readiness(project_dir, sub_type)
}
else:
# Config found - full analysis
sub_type = config.get("submission_type", submission_type or "510k_traditional")
milestones = config.get("milestones", {})
result = {
"submission_type": sub_type,
"device_name": config.get("device_name"),
"product_code": config.get("product_code"),
"predicate_device": config.get("predicate_device"),
"config_found": True,
"timeline_status": calculate_timeline_status(sub_type, milestones),
"milestones": analyze_milestone_status(sub_type, milestones),
"readiness": calculate_submission_readiness(project_dir, sub_type)
}
# Generate recommendations
result["recommendations"] = generate_recommendations(result)
return result
def main():
parser = argparse.ArgumentParser(
description="FDA Submission Tracker - Monitor 510(k), De Novo, and PMA submissions"
)
parser.add_argument(
"project_dir",
nargs="?",
default=".",
help="Project directory to analyze (default: current directory)"
)
parser.add_argument(
"--type",
choices=["510k", "510k_traditional", "510k_special", "510k_abbreviated",
"de_novo", "pma", "pma_supplement"],
help="Submission type (overrides config file)"
)
parser.add_argument(
"--json",
action="store_true",
help="Output in JSON format"
)
parser.add_argument(
"--init",
action="store_true",
help="Create sample configuration file"
)
args = parser.parse_args()
project_dir = Path(args.project_dir).resolve()
if not project_dir.exists():
print(f"Error: Directory not found: {project_dir}", file=sys.stderr)
sys.exit(1)
if args.init:
config_path = project_dir / "fda_submission.json"
if config_path.exists():
print(f"Configuration file already exists: {config_path}")
sys.exit(1)
sample = generate_sample_config()
if args.type:
sample["submission_type"] = args.type
with open(config_path, "w") as f:
json.dump(sample, f, indent=2)
print(f"Created sample configuration: {config_path}")
print("Edit this file with your submission details and milestone dates.")
return
result = analyze_submission(project_dir, args.type)
if args.json:
print(json.dumps(result, indent=2))
else:
print_text_report(result)
if __name__ == "__main__":
main()
FILE:scripts/hipaa_risk_assessment.py
#!/usr/bin/env python3
"""
HIPAA Risk Assessment Tool
Evaluates HIPAA compliance for medical device software and connected devices
by analyzing code and documentation for security safeguards.
Usage:
python hipaa_risk_assessment.py <project_dir>
python hipaa_risk_assessment.py <project_dir> --category technical
python hipaa_risk_assessment.py <project_dir> --json
"""
import argparse
import json
import os
import re
import sys
from datetime import datetime
from pathlib import Path
from typing import Dict, List, Optional, Any, Tuple
# HIPAA Security Rule safeguards
HIPAA_SAFEGUARDS = {
"administrative": {
"title": "Administrative Safeguards (§164.308)",
"controls": {
"security_management": {
"title": "Security Management Process",
"requirement": "Risk analysis, risk management, sanction policy",
"doc_patterns": ["risk_assessment*", "security_policy*", "sanction*"],
"code_patterns": [],
"weight": 10
},
"security_officer": {
"title": "Assigned Security Responsibility",
"requirement": "Designated security official",
"doc_patterns": ["security_officer*", "hipaa_officer*", "privacy_officer*"],
"code_patterns": [],
"weight": 5
},
"workforce_security": {
"title": "Workforce Security",
"requirement": "Authorization/supervision, clearance, termination procedures",
"doc_patterns": ["access_control*", "termination*", "hr_security*"],
"code_patterns": [],
"weight": 5
},
"access_management": {
"title": "Information Access Management",
"requirement": "Access authorization, establishment, modification",
"doc_patterns": ["access_management*", "role_definition*", "access_control*"],
"code_patterns": [r"role.*based", r"permission", r"authorization"],
"weight": 8
},
"security_training": {
"title": "Security Awareness and Training",
"requirement": "Training program, security reminders",
"doc_patterns": ["training*", "security_awareness*"],
"code_patterns": [],
"weight": 5
},
"incident_procedures": {
"title": "Security Incident Procedures",
"requirement": "Incident response and reporting",
"doc_patterns": ["incident*", "breach*", "security_event*"],
"code_patterns": [r"incident.*report", r"security.*alert", r"breach.*notify"],
"weight": 8
},
"contingency_plan": {
"title": "Contingency Plan",
"requirement": "Backup, disaster recovery, emergency mode",
"doc_patterns": ["contingency*", "disaster_recovery*", "backup*", "dr_plan*"],
"code_patterns": [r"backup", r"recovery", r"failover"],
"weight": 8
},
"evaluation": {
"title": "Evaluation",
"requirement": "Periodic security evaluations",
"doc_patterns": ["security_audit*", "hipaa_audit*", "compliance_review*"],
"code_patterns": [],
"weight": 5
},
"baa": {
"title": "Business Associate Contracts",
"requirement": "Written contracts with business associates",
"doc_patterns": ["baa*", "business_associate*", "vendor_agreement*"],
"code_patterns": [],
"weight": 5
}
}
},
"physical": {
"title": "Physical Safeguards (§164.310)",
"controls": {
"facility_access": {
"title": "Facility Access Controls",
"requirement": "Physical access procedures and controls",
"doc_patterns": ["facility_access*", "physical_security*", "access_control*"],
"code_patterns": [],
"weight": 5
},
"workstation_use": {
"title": "Workstation Use",
"requirement": "Policies for workstation use and security",
"doc_patterns": ["workstation*", "endpoint*", "device_policy*"],
"code_patterns": [],
"weight": 3
},
"device_media": {
"title": "Device and Media Controls",
"requirement": "Disposal, media re-use, accountability",
"doc_patterns": ["media_disposal*", "device_disposal*", "data_sanitization*"],
"code_patterns": [r"secure.*delete", r"wipe", r"sanitize"],
"weight": 5
}
}
},
"technical": {
"title": "Technical Safeguards (§164.312)",
"controls": {
"access_control": {
"title": "Access Control",
"requirement": "Unique user ID, emergency access, auto logoff, encryption",
"doc_patterns": ["access_control*", "authentication*", "session*"],
"code_patterns": [
r"authentication",
r"authorize",
r"session.*timeout",
r"auto.*logout",
r"unique.*id",
r"user.*id"
],
"weight": 10
},
"audit_controls": {
"title": "Audit Controls",
"requirement": "Record and examine activity in systems with ePHI",
"doc_patterns": ["audit_log*", "access_log*", "security_log*"],
"code_patterns": [
r"audit.*log",
r"access.*log",
r"log.*access",
r"security.*event",
r"logger"
],
"weight": 10
},
"integrity": {
"title": "Integrity Controls",
"requirement": "Mechanism to authenticate ePHI",
"doc_patterns": ["data_integrity*", "checksum*", "hash*"],
"code_patterns": [
r"checksum",
r"hash",
r"hmac",
r"integrity.*check",
r"digital.*signature"
],
"weight": 8
},
"authentication": {
"title": "Person or Entity Authentication",
"requirement": "Verify identity of person or entity seeking access",
"doc_patterns": ["authentication*", "identity*", "mfa*", "2fa*"],
"code_patterns": [
r"authenticate",
r"mfa",
r"two.*factor",
r"2fa",
r"multi.*factor",
r"oauth",
r"jwt"
],
"weight": 10
},
"transmission_security": {
"title": "Transmission Security",
"requirement": "Encryption during transmission",
"doc_patterns": ["encryption*", "tls*", "ssl*", "transport_security*"],
"code_patterns": [
r"https",
r"tls",
r"ssl",
r"encrypt.*transit",
r"secure.*connection"
],
"weight": 10
}
}
}
}
# PHI data patterns to detect in code
PHI_PATTERNS = [
(r"patient.*name", "Patient Name"),
(r"ssn|social.*security", "Social Security Number"),
(r"date.*of.*birth|dob", "Date of Birth"),
(r"medical.*record", "Medical Record Number"),
(r"health.*plan", "Health Plan ID"),
(r"diagnosis|icd.*code", "Diagnosis/ICD Code"),
(r"prescription|medication", "Medication/Prescription"),
(r"insurance", "Insurance Information"),
(r"phone.*number|telephone", "Phone Number"),
(r"email.*address", "Email Address"),
(r"address|street|city|zip", "Physical Address"),
(r"biometric", "Biometric Data")
]
# Security vulnerability patterns (dynamic code execution, hardcoded secrets)
VULNERABILITY_PATTERNS = [
(r"password.*=.*['\"]", "Hardcoded password"),
(r"api.*key.*=.*['\"]", "Hardcoded API key"),
(r"secret.*=.*['\"]", "Hardcoded secret"),
(r"http://(?!localhost)", "Unencrypted HTTP connection"),
(r"verify.*=.*False", "SSL verification disabled"),
(r"dynamic.*code.*execution", "Dynamic code execution risk"),
(r"disable.*ssl", "SSL disabled"),
(r"insecure", "Insecure configuration")
]
def scan_documentation(project_dir: Path, patterns: List[str]) -> List[str]:
"""Scan for documentation matching patterns."""
found = []
doc_dirs = [
project_dir / "docs",
project_dir / "documentation",
project_dir / "policies",
project_dir / "compliance",
project_dir / "hipaa",
project_dir
]
for doc_dir in doc_dirs:
if not doc_dir.exists():
continue
for pattern in patterns:
for ext in ["*.md", "*.pdf", "*.docx", "*.doc", "*.txt"]:
try:
for match in doc_dir.glob(f"**/{pattern}{ext}"):
rel_path = str(match.relative_to(project_dir))
if rel_path not in found:
found.append(rel_path)
except Exception:
continue
return found
def scan_code_patterns(project_dir: Path, patterns: List[str]) -> List[Dict]:
"""Scan source code for patterns."""
matches = []
code_extensions = ["*.py", "*.js", "*.ts", "*.java", "*.cs", "*.go", "*.rb"]
src_dirs = [
project_dir / "src",
project_dir / "app",
project_dir / "lib",
project_dir
]
for src_dir in src_dirs:
if not src_dir.exists():
continue
for ext in code_extensions:
try:
for file_path in src_dir.glob(f"**/{ext}"):
# Skip node_modules, venv, etc.
if any(skip in str(file_path) for skip in ["node_modules", "venv", ".venv", "__pycache__", ".git"]):
continue
try:
content = file_path.read_text(encoding='utf-8', errors='ignore')
for pattern in patterns:
if re.search(pattern, content, re.IGNORECASE):
rel_path = str(file_path.relative_to(project_dir))
matches.append({
"file": rel_path,
"pattern": pattern
})
break # One match per file per control is enough
except Exception:
continue
except Exception:
continue
return matches
def detect_phi_handling(project_dir: Path) -> Dict:
"""Detect potential PHI handling in code."""
phi_found = []
code_extensions = ["*.py", "*.js", "*.ts", "*.java", "*.cs", "*.go"]
for ext in code_extensions:
try:
for file_path in project_dir.glob(f"**/{ext}"):
if any(skip in str(file_path) for skip in ["node_modules", "venv", ".venv", "__pycache__", ".git"]):
continue
try:
content = file_path.read_text(encoding='utf-8', errors='ignore')
rel_path = str(file_path.relative_to(project_dir))
for pattern, phi_type in PHI_PATTERNS:
if re.search(pattern, content, re.IGNORECASE):
phi_found.append({
"file": rel_path,
"phi_type": phi_type
})
break
except Exception:
continue
except Exception:
continue
return {
"phi_detected": len(phi_found) > 0,
"files_with_phi": phi_found,
"phi_types": list(set(p["phi_type"] for p in phi_found))
}
def detect_security_vulnerabilities(project_dir: Path) -> List[Dict]:
"""Scan for security vulnerabilities."""
vulnerabilities = []
code_extensions = ["*.py", "*.js", "*.ts", "*.java", "*.cs", "*.go", "*.yaml", "*.yml", "*.json"]
for ext in code_extensions:
try:
for file_path in project_dir.glob(f"**/{ext}"):
if any(skip in str(file_path) for skip in ["node_modules", "venv", ".venv", "__pycache__", ".git"]):
continue
try:
content = file_path.read_text(encoding='utf-8', errors='ignore')
rel_path = str(file_path.relative_to(project_dir))
for pattern, vuln_type in VULNERABILITY_PATTERNS:
matches = re.findall(pattern, content, re.IGNORECASE)
if matches:
vulnerabilities.append({
"file": rel_path,
"vulnerability": vuln_type,
"count": len(matches)
})
except Exception:
continue
except Exception:
continue
return vulnerabilities
def assess_control(project_dir: Path, control_id: str, control_data: Dict) -> Dict:
"""Assess a single HIPAA control."""
doc_evidence = scan_documentation(project_dir, control_data["doc_patterns"])
code_evidence = scan_code_patterns(project_dir, control_data["code_patterns"]) if control_data["code_patterns"] else []
# Determine compliance status
has_docs = len(doc_evidence) > 0
has_code = len(code_evidence) > 0
if has_docs and (has_code or not control_data["code_patterns"]):
status = "implemented"
score = 100
elif has_docs or has_code:
status = "partial"
score = 50
else:
status = "gap"
score = 0
return {
"control_id": control_id,
"title": control_data["title"],
"requirement": control_data["requirement"],
"status": status,
"score": score,
"weight": control_data["weight"],
"weighted_score": (score * control_data["weight"]) / 100,
"documentation": doc_evidence,
"code_evidence": [e["file"] for e in code_evidence]
}
def assess_category(project_dir: Path, category_id: str, category_data: Dict) -> Dict:
"""Assess a HIPAA safeguard category."""
control_results = []
total_weight = 0
weighted_score = 0
for control_id, control_data in category_data["controls"].items():
result = assess_control(project_dir, control_id, control_data)
control_results.append(result)
total_weight += control_data["weight"]
weighted_score += result["weighted_score"]
category_score = round((weighted_score / total_weight) * 100, 1) if total_weight > 0 else 0
return {
"category": category_id,
"title": category_data["title"],
"score": category_score,
"controls": control_results,
"compliant": sum(1 for c in control_results if c["status"] == "implemented"),
"partial": sum(1 for c in control_results if c["status"] == "partial"),
"gaps": sum(1 for c in control_results if c["status"] == "gap")
}
def calculate_risk_level(overall_score: float, vulnerabilities: List[Dict], phi_data: Dict) -> Dict:
"""Calculate overall HIPAA risk level."""
# Base risk from compliance score
if overall_score >= 80:
base_risk = "LOW"
base_score = 1
elif overall_score >= 60:
base_risk = "MEDIUM"
base_score = 2
elif overall_score >= 40:
base_risk = "HIGH"
base_score = 3
else:
base_risk = "CRITICAL"
base_score = 4
# Adjust for vulnerabilities
critical_vulns = sum(1 for v in vulnerabilities if "password" in v["vulnerability"].lower() or "secret" in v["vulnerability"].lower())
if critical_vulns > 0:
base_score = min(4, base_score + 1)
# Adjust for PHI handling
if phi_data["phi_detected"] and base_score < 4:
base_score = min(4, base_score + 0.5)
# Map back to risk level
risk_levels = {1: "LOW", 2: "MEDIUM", 3: "HIGH", 4: "CRITICAL"}
final_risk = risk_levels.get(int(base_score), "HIGH")
return {
"risk_level": final_risk,
"compliance_score": overall_score,
"vulnerability_count": len(vulnerabilities),
"phi_handling_detected": phi_data["phi_detected"]
}
def generate_recommendations(assessment: Dict) -> List[str]:
"""Generate prioritized recommendations."""
recommendations = []
# Technical safeguards first (highest priority for software)
for cat in assessment["categories"]:
if cat["category"] == "technical":
for control in cat["controls"]:
if control["status"] == "gap":
recommendations.append(f"CRITICAL: Implement {control['title']} - {control['requirement']}")
elif control["status"] == "partial":
recommendations.append(f"HIGH: Complete {control['title']} implementation")
# Administrative safeguards
for cat in assessment["categories"]:
if cat["category"] == "administrative":
for control in cat["controls"]:
if control["status"] == "gap":
recommendations.append(f"MEDIUM: Document {control['title']} procedures")
# Vulnerabilities
for vuln in assessment.get("vulnerabilities", [])[:5]:
recommendations.append(f"SECURITY: Fix {vuln['vulnerability']} in {vuln['file']}")
return recommendations[:10] # Top 10
def print_text_report(result: Dict) -> None:
"""Print human-readable report."""
print("=" * 70)
print("HIPAA SECURITY RULE COMPLIANCE ASSESSMENT")
print("=" * 70)
# Risk summary
risk = result["risk_assessment"]
print(f"\nRISK LEVEL: {risk['risk_level']}")
print(f"Compliance Score: {risk['compliance_score']}%")
print(f"Vulnerabilities Found: {risk['vulnerability_count']}")
print(f"PHI Handling Detected: {'Yes' if risk['phi_handling_detected'] else 'No'}")
# Category scores
print("\n--- SAFEGUARD CATEGORIES ---")
for cat in result["categories"]:
status = "OK" if cat["score"] >= 70 else "NEEDS ATTENTION"
print(f" {cat['title']}: {cat['score']}% [{status}]")
print(f" Implemented: {cat['compliant']}, Partial: {cat['partial']}, Gaps: {cat['gaps']}")
# Gaps
print("\n--- COMPLIANCE GAPS ---")
gap_count = 0
for cat in result["categories"]:
for control in cat["controls"]:
if control["status"] == "gap":
gap_count += 1
print(f" [{cat['category'].upper()}] {control['title']}")
print(f" Requirement: {control['requirement']}")
if gap_count == 0:
print(" No critical gaps identified")
# PHI Detection
if result["phi_detection"]["phi_detected"]:
print("\n--- PHI HANDLING DETECTED ---")
print(f" PHI Types: {', '.join(result['phi_detection']['phi_types'])}")
print(f" Files: {len(result['phi_detection']['files_with_phi'])}")
# Vulnerabilities
if result["vulnerabilities"]:
print("\n--- SECURITY VULNERABILITIES ---")
for vuln in result["vulnerabilities"][:10]:
print(f" - {vuln['vulnerability']}: {vuln['file']}")
# Recommendations
if result["recommendations"]:
print("\n--- RECOMMENDATIONS ---")
for i, rec in enumerate(result["recommendations"], 1):
print(f" {i}. {rec}")
print("\n" + "=" * 70)
print(f"Assessment Date: {datetime.now().strftime('%Y-%m-%d %H:%M')}")
print("=" * 70)
def main():
parser = argparse.ArgumentParser(
description="HIPAA Risk Assessment Tool for Medical Device Software"
)
parser.add_argument(
"project_dir",
nargs="?",
default=".",
help="Project directory to analyze (default: current directory)"
)
parser.add_argument(
"--category",
choices=["administrative", "physical", "technical"],
help="Assess specific safeguard category only"
)
parser.add_argument(
"--json",
action="store_true",
help="Output in JSON format"
)
parser.add_argument(
"--detailed",
action="store_true",
help="Include detailed evidence in output"
)
args = parser.parse_args()
project_dir = Path(args.project_dir).resolve()
if not project_dir.exists():
print(f"Error: Directory not found: {project_dir}", file=sys.stderr)
sys.exit(1)
# Filter categories if specific one requested
categories_to_assess = HIPAA_SAFEGUARDS
if args.category:
categories_to_assess = {args.category: HIPAA_SAFEGUARDS[args.category]}
# Perform assessment
category_results = []
total_weight = 0
weighted_score = 0
for cat_id, cat_data in categories_to_assess.items():
cat_result = assess_category(project_dir, cat_id, cat_data)
category_results.append(cat_result)
# Calculate weighted average
cat_weight = sum(c["weight"] for c in cat_data["controls"].values())
total_weight += cat_weight
weighted_score += (cat_result["score"] * cat_weight) / 100
overall_score = round((weighted_score / total_weight) * 100, 1) if total_weight > 0 else 0
# Additional scans
phi_detection = detect_phi_handling(project_dir)
vulnerabilities = detect_security_vulnerabilities(project_dir)
# Risk assessment
risk_assessment = calculate_risk_level(overall_score, vulnerabilities, phi_detection)
result = {
"project_dir": str(project_dir),
"assessment_date": datetime.now().isoformat(),
"overall_score": overall_score,
"risk_assessment": risk_assessment,
"categories": category_results if args.detailed else [
{
"category": c["category"],
"title": c["title"],
"score": c["score"],
"compliant": c["compliant"],
"partial": c["partial"],
"gaps": c["gaps"]
}
for c in category_results
],
"phi_detection": phi_detection,
"vulnerabilities": vulnerabilities,
"recommendations": []
}
result["recommendations"] = generate_recommendations(result)
if args.json:
print(json.dumps(result, indent=2))
else:
print_text_report(result)
if __name__ == "__main__":
main()
FILE:scripts/qsr_compliance_checker.py
#!/usr/bin/env python3
"""
QSR Compliance Checker
Assesses compliance with 21 CFR Part 820 (Quality System Regulation) by analyzing
project documentation and identifying gaps.
Usage:
python qsr_compliance_checker.py <project_dir>
python qsr_compliance_checker.py <project_dir> --section 820.30
python qsr_compliance_checker.py <project_dir> --json
"""
import argparse
import json
import os
import re
import sys
from datetime import datetime
from pathlib import Path
from typing import Dict, List, Optional, Any
# QSR sections and requirements
QSR_REQUIREMENTS = {
"820.20": {
"title": "Management Responsibility",
"subsections": {
"820.20(a)": {
"title": "Quality Policy",
"required_evidence": ["quality_policy", "quality_manual", "quality_objectives"],
"doc_patterns": ["quality_policy*", "quality_manual*", "qms_manual*"],
"keywords": ["quality policy", "quality objectives", "management commitment"]
},
"820.20(b)": {
"title": "Organization",
"required_evidence": ["org_chart", "job_descriptions", "authority_matrix"],
"doc_patterns": ["org_chart*", "organization*", "job_desc*", "authority*"],
"keywords": ["organizational structure", "responsibility", "authority"]
},
"820.20(c)": {
"title": "Management Review",
"required_evidence": ["management_review_procedure", "management_review_records"],
"doc_patterns": ["management_review*", "mgmt_review*", "qmr*"],
"keywords": ["management review", "review meeting", "quality system effectiveness"]
}
}
},
"820.30": {
"title": "Design Controls",
"subsections": {
"820.30(a)": {
"title": "Design and Development Planning",
"required_evidence": ["design_plan", "development_plan"],
"doc_patterns": ["design_plan*", "dev_plan*", "development_plan*"],
"keywords": ["design planning", "development phases", "design milestones"]
},
"820.30(b)": {
"title": "Design Input",
"required_evidence": ["design_input", "requirements_specification"],
"doc_patterns": ["design_input*", "requirement*", "srs*", "prs*"],
"keywords": ["design input", "requirements", "user needs", "intended use"]
},
"820.30(c)": {
"title": "Design Output",
"required_evidence": ["design_output", "specifications", "drawings"],
"doc_patterns": ["design_output*", "specification*", "drawing*", "bom*"],
"keywords": ["design output", "specifications", "acceptance criteria"]
},
"820.30(d)": {
"title": "Design Review",
"required_evidence": ["design_review_procedure", "design_review_records"],
"doc_patterns": ["design_review*", "dr_record*", "dr_minutes*"],
"keywords": ["design review", "review meeting", "design evaluation"]
},
"820.30(e)": {
"title": "Design Verification",
"required_evidence": ["verification_plan", "verification_results"],
"doc_patterns": ["verification*", "test_report*", "dv_*"],
"keywords": ["verification", "testing", "design verification"]
},
"820.30(f)": {
"title": "Design Validation",
"required_evidence": ["validation_plan", "validation_results"],
"doc_patterns": ["validation*", "clinical*", "usability*", "val_*"],
"keywords": ["validation", "user needs", "intended use", "clinical evaluation"]
},
"820.30(g)": {
"title": "Design Transfer",
"required_evidence": ["transfer_checklist", "transfer_verification"],
"doc_patterns": ["transfer*", "production_release*"],
"keywords": ["design transfer", "manufacturing", "production"]
},
"820.30(h)": {
"title": "Design Changes",
"required_evidence": ["change_control_procedure", "change_records"],
"doc_patterns": ["change_control*", "ecn*", "eco*", "dcr*"],
"keywords": ["design change", "change control", "modification"]
},
"820.30(i)": {
"title": "Design History File",
"required_evidence": ["dhf_index", "dhf"],
"doc_patterns": ["dhf*", "design_history*"],
"keywords": ["design history file", "DHF", "design records"]
}
}
},
"820.40": {
"title": "Document Controls",
"subsections": {
"820.40(a)": {
"title": "Document Approval and Distribution",
"required_evidence": ["document_control_procedure"],
"doc_patterns": ["document_control*", "doc_control*", "sop_document*"],
"keywords": ["document approval", "document distribution", "controlled documents"]
},
"820.40(b)": {
"title": "Document Changes",
"required_evidence": ["document_change_procedure", "revision_history"],
"doc_patterns": ["revision_history*", "document_change*"],
"keywords": ["document change", "revision", "document modification"]
}
}
},
"820.50": {
"title": "Purchasing Controls",
"subsections": {
"820.50(a)": {
"title": "Evaluation of Suppliers",
"required_evidence": ["supplier_qualification_procedure", "approved_supplier_list"],
"doc_patterns": ["supplier*", "asl*", "vendor*"],
"keywords": ["supplier evaluation", "approved supplier", "vendor qualification"]
},
"820.50(b)": {
"title": "Purchasing Data",
"required_evidence": ["purchasing_procedure", "purchase_order_requirements"],
"doc_patterns": ["purchas*", "procurement*"],
"keywords": ["purchasing data", "specifications", "quality requirements"]
}
}
},
"820.70": {
"title": "Production and Process Controls",
"subsections": {
"820.70(a)": {
"title": "General Process Controls",
"required_evidence": ["manufacturing_procedures", "work_instructions"],
"doc_patterns": ["manufacturing*", "production*", "work_instruction*", "wi_*"],
"keywords": ["manufacturing process", "production", "process parameters"]
},
"820.70(b)": {
"title": "Production and Process Changes",
"required_evidence": ["process_change_procedure"],
"doc_patterns": ["process_change*", "manufacturing_change*"],
"keywords": ["process change", "production change", "change control"]
},
"820.70(c)": {
"title": "Environmental Control",
"required_evidence": ["environmental_control_procedure", "monitoring_records"],
"doc_patterns": ["environmental*", "cleanroom*", "env_monitoring*"],
"keywords": ["environmental control", "cleanroom", "contamination"]
},
"820.70(d)": {
"title": "Personnel",
"required_evidence": ["training_procedure", "training_records"],
"doc_patterns": ["training*", "personnel*", "competency*"],
"keywords": ["training", "personnel qualification", "competency"]
},
"820.70(e)": {
"title": "Contamination Control",
"required_evidence": ["contamination_control_procedure"],
"doc_patterns": ["contamination*", "cleaning*", "hygiene*"],
"keywords": ["contamination", "cleaning", "hygiene"]
},
"820.70(f)": {
"title": "Buildings",
"required_evidence": ["facility_requirements"],
"doc_patterns": ["facility*", "building*"],
"keywords": ["facility", "buildings", "manufacturing area"]
},
"820.70(g)": {
"title": "Equipment",
"required_evidence": ["equipment_maintenance_procedure", "maintenance_records"],
"doc_patterns": ["equipment*", "maintenance*", "preventive_maintenance*"],
"keywords": ["equipment", "maintenance", "calibration"]
},
"820.70(h)": {
"title": "Manufacturing Material",
"required_evidence": ["material_handling_procedure"],
"doc_patterns": ["material*", "handling*", "storage*"],
"keywords": ["manufacturing material", "handling", "storage"]
},
"820.70(i)": {
"title": "Automated Processes",
"required_evidence": ["software_validation", "automated_process_validation"],
"doc_patterns": ["software_val*", "csv*", "automation*"],
"keywords": ["software validation", "automated", "computer system"]
}
}
},
"820.72": {
"title": "Inspection, Measuring, and Test Equipment",
"subsections": {
"820.72(a)": {
"title": "Calibration",
"required_evidence": ["calibration_procedure", "calibration_records"],
"doc_patterns": ["calibration*", "cal_*"],
"keywords": ["calibration", "accuracy", "measurement"]
},
"820.72(b)": {
"title": "Calibration Standards",
"required_evidence": ["calibration_standards", "traceability_records"],
"doc_patterns": ["calibration_standard*", "nist*", "traceability*"],
"keywords": ["calibration standards", "NIST", "traceability"]
}
}
},
"820.75": {
"title": "Process Validation",
"subsections": {
"820.75(a)": {
"title": "Process Validation Requirements",
"required_evidence": ["process_validation_procedure", "validation_protocols"],
"doc_patterns": ["process_validation*", "pv_*", "validation_protocol*"],
"keywords": ["process validation", "IQ", "OQ", "PQ"]
},
"820.75(b)": {
"title": "Validation Monitoring",
"required_evidence": ["validation_monitoring", "revalidation_criteria"],
"doc_patterns": ["revalidation*", "validation_monitoring*"],
"keywords": ["monitoring", "revalidation", "process performance"]
}
}
},
"820.90": {
"title": "Nonconforming Product",
"subsections": {
"820.90(a)": {
"title": "Nonconforming Product Control",
"required_evidence": ["ncr_procedure", "nonconforming_records"],
"doc_patterns": ["ncr*", "nonconform*", "nc_*"],
"keywords": ["nonconforming", "NCR", "disposition"]
},
"820.90(b)": {
"title": "Nonconformance Review",
"required_evidence": ["ncr_review_procedure"],
"doc_patterns": ["ncr_review*", "mrb*"],
"keywords": ["review", "disposition", "concession"]
}
}
},
"820.100": {
"title": "Corrective and Preventive Action",
"subsections": {
"820.100(a)": {
"title": "CAPA Procedure",
"required_evidence": ["capa_procedure", "capa_records"],
"doc_patterns": ["capa*", "corrective*", "preventive*"],
"keywords": ["CAPA", "corrective action", "preventive action", "root cause"]
}
}
},
"820.120": {
"title": "Device Labeling",
"subsections": {
"820.120": {
"title": "Labeling Controls",
"required_evidence": ["labeling_procedure", "label_inspection"],
"doc_patterns": ["label*", "labeling*"],
"keywords": ["labeling", "label inspection", "UDI"]
}
}
},
"820.180": {
"title": "General Requirements - Records",
"subsections": {
"820.180": {
"title": "Records Requirements",
"required_evidence": ["records_management_procedure", "retention_schedule"],
"doc_patterns": ["record*", "retention*", "archive*"],
"keywords": ["records", "retention", "archive", "backup"]
}
}
},
"820.181": {
"title": "Device Master Record",
"subsections": {
"820.181": {
"title": "DMR Contents",
"required_evidence": ["dmr_index", "dmr"],
"doc_patterns": ["dmr*", "device_master*"],
"keywords": ["device master record", "DMR", "specifications"]
}
}
},
"820.184": {
"title": "Device History Record",
"subsections": {
"820.184": {
"title": "DHR Contents",
"required_evidence": ["dhr_template", "dhr_records"],
"doc_patterns": ["dhr*", "device_history*", "batch_record*"],
"keywords": ["device history record", "DHR", "production record"]
}
}
},
"820.198": {
"title": "Complaint Files",
"subsections": {
"820.198": {
"title": "Complaint Handling",
"required_evidence": ["complaint_procedure", "complaint_records"],
"doc_patterns": ["complaint*", "customer_feedback*"],
"keywords": ["complaint", "customer feedback", "MDR"]
}
}
}
}
def search_documentation(project_dir: Path, patterns: List[str], keywords: List[str]) -> Dict:
"""Search for documentation matching patterns and keywords."""
result = {
"documents_found": [],
"keyword_matches": [],
"evidence_strength": "none"
}
# Common documentation directories
doc_dirs = [
project_dir / "qms",
project_dir / "quality",
project_dir / "docs",
project_dir / "documentation",
project_dir / "procedures",
project_dir / "sops",
project_dir / "dhf",
project_dir / "dmr",
project_dir
]
# Search for document patterns
for doc_dir in doc_dirs:
if not doc_dir.exists():
continue
for pattern in patterns:
for ext in ["*.md", "*.pdf", "*.docx", "*.doc", "*.txt"]:
full_pattern = f"**/{pattern}{ext}" if not pattern.endswith("*") else f"**/{pattern[:-1]}{ext}"
try:
matches = list(doc_dir.glob(full_pattern))
for match in matches:
rel_path = str(match.relative_to(project_dir))
if rel_path not in result["documents_found"]:
result["documents_found"].append(rel_path)
except Exception:
continue
# Search for keywords in markdown and text files
for doc_dir in doc_dirs:
if not doc_dir.exists():
continue
for ext in ["*.md", "*.txt"]:
try:
for file_path in doc_dir.glob(f"**/{ext}"):
try:
content = file_path.read_text(encoding='utf-8', errors='ignore').lower()
for keyword in keywords:
if keyword.lower() in content:
rel_path = str(file_path.relative_to(project_dir))
if rel_path not in result["keyword_matches"]:
result["keyword_matches"].append(rel_path)
except Exception:
continue
except Exception:
continue
# Determine evidence strength
if result["documents_found"] and result["keyword_matches"]:
result["evidence_strength"] = "strong"
elif result["documents_found"] or result["keyword_matches"]:
result["evidence_strength"] = "partial"
else:
result["evidence_strength"] = "none"
return result
def assess_section(project_dir: Path, section_id: str, section_data: Dict) -> Dict:
"""Assess compliance for a QSR section."""
result = {
"section": section_id,
"title": section_data["title"],
"subsections": [],
"compliance_score": 0,
"total_subsections": len(section_data["subsections"]),
"compliant_subsections": 0
}
for subsection_id, subsection_data in section_data["subsections"].items():
evidence = search_documentation(
project_dir,
subsection_data["doc_patterns"],
subsection_data["keywords"]
)
subsection_result = {
"subsection": subsection_id,
"title": subsection_data["title"],
"required_evidence": subsection_data["required_evidence"],
"evidence_found": evidence,
"status": "gap" if evidence["evidence_strength"] == "none" else (
"partial" if evidence["evidence_strength"] == "partial" else "compliant"
)
}
if subsection_result["status"] == "compliant":
result["compliant_subsections"] += 1
elif subsection_result["status"] == "partial":
result["compliant_subsections"] += 0.5
result["subsections"].append(subsection_result)
if result["total_subsections"] > 0:
result["compliance_score"] = round(
(result["compliant_subsections"] / result["total_subsections"]) * 100, 1
)
return result
def generate_gap_report(assessment_results: List[Dict]) -> Dict:
"""Generate gap analysis report."""
gaps = []
recommendations = []
for section in assessment_results:
for subsection in section["subsections"]:
if subsection["status"] != "compliant":
gap = {
"section": subsection["subsection"],
"title": subsection["title"],
"status": subsection["status"],
"missing_evidence": subsection["required_evidence"]
}
gaps.append(gap)
if subsection["status"] == "gap":
recommendations.append(
f"{subsection['subsection']}: Create documentation for {subsection['title']}"
)
else:
recommendations.append(
f"{subsection['subsection']}: Enhance documentation for {subsection['title']}"
)
return {
"total_gaps": len([g for g in gaps if g["status"] == "gap"]),
"total_partial": len([g for g in gaps if g["status"] == "partial"]),
"gaps": gaps,
"priority_recommendations": recommendations[:10] # Top 10
}
def calculate_overall_compliance(assessment_results: List[Dict]) -> Dict:
"""Calculate overall QSR compliance score."""
total_subsections = 0
compliant_subsections = 0
section_scores = {}
for section in assessment_results:
total_subsections += section["total_subsections"]
compliant_subsections += section["compliant_subsections"]
section_scores[section["section"]] = section["compliance_score"]
overall_score = round((compliant_subsections / total_subsections) * 100, 1) if total_subsections > 0 else 0
# Determine compliance level
if overall_score >= 90:
level = "HIGH"
color = "green"
elif overall_score >= 70:
level = "MEDIUM"
color = "yellow"
elif overall_score >= 50:
level = "LOW"
color = "orange"
else:
level = "CRITICAL"
color = "red"
return {
"overall_score": overall_score,
"compliance_level": level,
"total_subsections": total_subsections,
"compliant_subsections": compliant_subsections,
"section_scores": section_scores
}
def print_text_report(result: Dict) -> None:
"""Print human-readable compliance report."""
print("=" * 70)
print("21 CFR PART 820 (QSR) COMPLIANCE ASSESSMENT")
print("=" * 70)
# Overall compliance
overall = result["overall_compliance"]
print(f"\nOVERALL COMPLIANCE: {overall['overall_score']}% ({overall['compliance_level']})")
print(f"Subsections Assessed: {overall['total_subsections']}")
print(f"Compliant/Partial: {overall['compliant_subsections']}")
# Section summary
print("\n--- SECTION SCORES ---")
for section in result["assessment"]:
status = "OK" if section["compliance_score"] >= 70 else "GAP"
print(f" {section['section']} {section['title']}: {section['compliance_score']}% [{status}]")
# Gap analysis
gap_report = result["gap_report"]
print(f"\n--- GAP ANALYSIS ---")
print(f"Critical Gaps: {gap_report['total_gaps']}")
print(f"Partial Compliance: {gap_report['total_partial']}")
if gap_report["gaps"]:
print("\n Gaps Identified:")
for gap in gap_report["gaps"][:15]: # Show top 15
status = "GAP" if gap["status"] == "gap" else "PARTIAL"
print(f" [{status}] {gap['section']}: {gap['title']}")
# Recommendations
if gap_report["priority_recommendations"]:
print("\n--- PRIORITY RECOMMENDATIONS ---")
for i, rec in enumerate(gap_report["priority_recommendations"], 1):
print(f" {i}. {rec}")
print("\n" + "=" * 70)
print(f"Assessment Date: {datetime.now().strftime('%Y-%m-%d %H:%M')}")
print("=" * 70)
def main():
parser = argparse.ArgumentParser(
description="QSR Compliance Checker - Assess 21 CFR 820 compliance"
)
parser.add_argument(
"project_dir",
nargs="?",
default=".",
help="Project directory to analyze (default: current directory)"
)
parser.add_argument(
"--section",
help="Analyze specific QSR section only (e.g., 820.30)"
)
parser.add_argument(
"--json",
action="store_true",
help="Output in JSON format"
)
parser.add_argument(
"--detailed",
action="store_true",
help="Include detailed evidence in output"
)
args = parser.parse_args()
project_dir = Path(args.project_dir).resolve()
if not project_dir.exists():
print(f"Error: Directory not found: {project_dir}", file=sys.stderr)
sys.exit(1)
# Filter sections if specific one requested
sections_to_assess = QSR_REQUIREMENTS
if args.section:
if args.section in QSR_REQUIREMENTS:
sections_to_assess = {args.section: QSR_REQUIREMENTS[args.section]}
else:
print(f"Error: Unknown section: {args.section}", file=sys.stderr)
print(f"Available sections: {', '.join(QSR_REQUIREMENTS.keys())}")
sys.exit(1)
# Perform assessment
assessment_results = []
for section_id, section_data in sections_to_assess.items():
section_result = assess_section(project_dir, section_id, section_data)
assessment_results.append(section_result)
# Generate reports
overall_compliance = calculate_overall_compliance(assessment_results)
gap_report = generate_gap_report(assessment_results)
result = {
"project_dir": str(project_dir),
"assessment_date": datetime.now().isoformat(),
"overall_compliance": overall_compliance,
"assessment": assessment_results if args.detailed else [
{
"section": s["section"],
"title": s["title"],
"compliance_score": s["compliance_score"],
"status": "compliant" if s["compliance_score"] >= 70 else "gap"
}
for s in assessment_results
],
"gap_report": gap_report
}
if args.json:
print(json.dumps(result, indent=2))
else:
print_text_report(result)
if __name__ == "__main__":
main()
Sửa các test Playwright bị lỗi hoặc chập chờn, gỡ lỗi test hỏng và lỗi xuất hiện ngẫu nhiên.
---
name: "fix"
description: >-
Fix failing or flaky Playwright tests. Use when user says "fix test",
"flaky test", "test failing", "debug test", "test broken", "test passes
sometimes", or "intermittent failure".
---
# Fix Failing or Flaky Tests
Diagnose and fix a Playwright test that fails or passes intermittently using a systematic taxonomy.
## Input
`$ARGUMENTS` contains:
- A test file path: `e2e/login.spec.ts`
- A test name: ""should redirect after login"`
- A description: `"the checkout test fails in CI but passes locally"`
## Steps
### 1. Reproduce the Failure
Run the test to capture the error:
```bash
npx playwright test <file> --reporter=list
```
If the test passes, it's likely flaky. Run burn-in:
```bash
npx playwright test <file> --repeat-each=10 --reporter=list
```
If it still passes, try with parallel workers:
```bash
npx playwright test --fully-parallel --workers=4 --repeat-each=5
```
### 2. Capture Trace
Run with full tracing:
```bash
npx playwright test <file> --trace=on --retries=0
```
Read the trace output. Use `/debug` to analyze trace files if available.
### 3. Categorize the Failure
Load `flaky-taxonomy.md` from this skill directory.
Every failing test falls into one of four categories:
| Category | Symptom | Diagnosis |
|---|---|---|
| **Timing/Async** | Fails intermittently everywhere | `--repeat-each=20` reproduces locally |
| **Test Isolation** | Fails in suite, passes alone | `--workers=1 --grep "test name"` passes |
| **Environment** | Fails in CI, passes locally | Compare CI vs local screenshots/traces |
| **Infrastructure** | Random, no pattern | Error references browser internals |
### 4. Apply Targeted Fix
**Timing/Async:**
- Replace `waitForTimeout()` with web-first assertions
- Add `await` to missing Playwright calls
- Wait for specific network responses before asserting
- Use `toBeVisible()` before interacting with elements
**Test Isolation:**
- Remove shared mutable state between tests
- Create test data per-test via API or fixtures
- Use unique identifiers (timestamps, random strings) for test data
- Check for database state leaks
**Environment:**
- Match viewport sizes between local and CI
- Account for font rendering differences in screenshots
- Use `docker` locally to match CI environment
- Check for timezone-dependent assertions
**Infrastructure:**
- Increase timeout for slow CI runners
- Add retries in CI config (`retries: 2`)
- Check for browser OOM (reduce parallel workers)
- Ensure browser dependencies are installed
### 5. Verify the Fix
Run the test 10 times to confirm stability:
```bash
npx playwright test <file> --repeat-each=10 --reporter=list
```
All 10 must pass. If any fail, go back to step 3.
### 6. Prevent Recurrence
Suggest:
- Add to CI with `retries: 2` if not already
- Enable `trace: 'on-first-retry'` in config
- Add the fix pattern to project's test conventions doc
## Output
- Root cause category and specific issue
- The fix applied (with diff)
- Verification result (10/10 passes)
- Prevention recommendation
FILE:flaky-taxonomy.md
# Flaky Test Taxonomy
## Decision Tree
```
Test is flaky
│
├── Fails locally with --repeat-each=20?
│ ├── YES → TIMING / ASYNC
│ │ ├── Missing await? → Add await
│ │ ├── waitForTimeout? → Replace with assertion
│ │ ├── Race condition? → Wait for specific event
│ │ └── Animation? → Wait for animation end or disable
│ │
│ └── NO → Continue...
│
├── Passes alone, fails in suite?
│ ├── YES → TEST ISOLATION
│ │ ├── Shared variable? → Make per-test
│ │ ├── Database state? → Reset per-test
│ │ ├── localStorage? → Clear in beforeEach
│ │ └── Cookie leak? → Use isolated contexts
│ │
│ └── NO → Continue...
│
├── Fails in CI, passes locally?
│ ├── YES → ENVIRONMENT
│ │ ├── Viewport? → Set explicit size
│ │ ├── Fonts? → Use Docker locally
│ │ ├── Timezone? → Use UTC everywhere
│ │ └── Network? → Mock external services
│ │
│ └── NO → INFRASTRUCTURE
│ ├── Browser crash? → Reduce workers
│ ├── OOM? → Limit parallel tests
│ ├── DNS? → Add retry config
│ └── File system? → Use unique temp dirs
```
## Common Fixes by Category
### Timing / Async
**Missing await:**
```typescript
// BAD — race condition
page.goto('/dashboard');
expect(page.getByText('Welcome')).toBeVisible();
// GOOD
await page.goto('/dashboard');
await expect(page.getByText('Welcome')).toBeVisible();
```
**Clicking before visible:**
```typescript
// BAD — element may not be ready
await page.getByRole('button', { name: 'Submit' }).click();
// GOOD — ensure visible first
const submitBtn = page.getByRole('button', { name: 'Submit' });
await expect(submitBtn).toBeVisible();
await submitBtn.click();
```
**Race with network:**
```typescript
// BAD — data might not be loaded
await page.goto('/users');
await expect(page.getByRole('table')).toBeVisible();
// GOOD — wait for API response
const responsePromise = page.waitForResponse('**/api/users');
await page.goto('/users');
await responsePromise;
await expect(page.getByRole('table')).toBeVisible();
```
### Test Isolation
**Shared state fix:**
```typescript
// BAD — tests share userId
let userId: string;
test('create', async () => { userId = '123'; });
test('read', async () => { /* uses userId */ });
// GOOD — each test is independent
test('read user', async ({ request }) => {
const response = await request.post('/api/users', { data: { name: 'Test' } });
const { id } = await response.json();
// Use id within this test
});
```
**localStorage cleanup:**
```typescript
test.beforeEach(async ({ page }) => {
await page.goto('/');
await page.evaluate(() => localStorage.clear());
});
```
### Environment
**Explicit viewport:**
```typescript
test.use({ viewport: { width: 1280, height: 720 } });
```
**Timezone-safe dates:**
```typescript
// BAD
expect(dateText).toBe('March 5, 2026');
// GOOD — timezone independent
expect(dateText).toMatch(/\d{1,2}\/\d{1,2}\/\d{4}/);
```
### Infrastructure
**Retry config:**
```typescript
// playwright.config.ts
export default defineConfig({
retries: process.env.CI ? 2 : 0,
workers: process.env.CI ? 2 : undefined,
});
```
**Increase timeout for CI:**
```typescript
test.setTimeout(60_000); // 60s for slow CI
```
Sửa và gỡ lỗi có hệ thống một tính năng hoặc module từ đầu đến cuối trên mọi tệp và phụ thuộc, không dùng cho lỗi đơn lẻ.
---
name: "focused-fix"
description: "Use when the user asks to fix, debug, or make a specific feature/module/area work end-to-end. Triggers: 'make X work', 'fix the Y feature', 'the Z module is broken', 'focus on [area]'. Not for quick single-bug fixes — this is for systematic deep-dive repair across all files and dependencies."
---
# Focused Fix — Deep-Dive Feature Repair
## When to Use
Activate when the user asks to fix, debug, or make a specific feature/module/area work. Key triggers:
- "make X work"
- "fix the Y feature"
- "the Z module is broken"
- "focus on [area]"
- "this feature needs to work properly"
This is NOT for quick single-bug fixes (use systematic-debugging for that). This is for when an entire feature or module needs systematic repair — tracing every dependency, reading logs, checking tests, mapping the full dependency graph.
```dot
digraph when_to_use {
"User reports feature broken" [shape=diamond];
"Single bug or symptom?" [shape=diamond];
"Use systematic-debugging" [shape=box];
"Entire feature/module needs repair?" [shape=diamond];
"Use focused-fix" [shape=box];
"Something else" [shape=box];
"User reports feature broken" -> "Single bug or symptom?";
"Single bug or symptom?" -> "Use systematic-debugging" [label="yes"];
"Single bug or symptom?" -> "Entire feature/module needs repair?" [label="no"];
"Entire feature/module needs repair?" -> "Use focused-fix" [label="yes"];
"Entire feature/module needs repair?" -> "Something else" [label="no"];
}
```
## The Iron Law
```
NO FIXES WITHOUT COMPLETING SCOPE → TRACE → DIAGNOSE FIRST
```
If you haven't finished Phase 3, you cannot propose fixes. Period.
**Violating the letter of these phases is violating the spirit of focused repair.**
## Protocol — STRICTLY follow these 5 phases IN ORDER
```dot
digraph phases {
rankdir=LR;
SCOPE [shape=box, label="Phase 1\nSCOPE"];
TRACE [shape=box, label="Phase 2\nTRACE"];
DIAGNOSE [shape=box, label="Phase 3\nDIAGNOSE"];
FIX [shape=box, label="Phase 4\nFIX"];
VERIFY [shape=box, label="Phase 5\nVERIFY"];
SCOPE -> TRACE -> DIAGNOSE -> FIX -> VERIFY;
FIX -> DIAGNOSE [label="fix broke\nsomething else"];
FIX -> ESCALATE [label="3+ fixes\ncreate new issues"];
ESCALATE [shape=doubleoctagon, label="STOP\nQuestion Architecture\nDiscuss with User"];
}
```
### Phase 1: SCOPE — Map the Feature Boundary
Before touching any code, understand the full scope of the feature.
1. Ask the user: "Which feature/folder should I focus on?" if not already clear
2. Identify the PRIMARY folder/files for this feature
3. Map EVERY file in that folder — read each one, understand its purpose
4. Create a feature manifest:
```
FEATURE SCOPE:
Primary path: src/features/auth/
Entry points: [files that are imported by other parts of the app]
Internal files: [files only used within this feature]
Total files: N
Total lines: N
```
### Phase 2: TRACE — Map All Dependencies (Inside AND Outside)
Trace every connection this feature has to the rest of the codebase.
**INBOUND (what this feature imports):**
1. For every import statement in every file in the feature folder:
- Trace it to its source
- Verify the source file exists
- Verify the imported entity (function, type, component) exists and is exported
- Check if the types/signatures match what the feature expects
2. Check for:
- Environment variables used (grep for process.env, import.meta.env, os.environ, etc.)
- Config files referenced
- Database models/schemas used
- API endpoints called
- Third-party packages imported
**OUTBOUND (what imports this feature):**
1. Search the entire codebase for imports from this feature folder
2. For each consumer:
- Verify they're importing entities that actually exist
- Check if they're using the correct API/interface
- Note if any consumers are using deprecated patterns
Output format:
```
DEPENDENCY MAP:
Inbound (this feature depends on):
src/lib/db.ts → used in auth/repository.ts (getUserById, createUser)
src/lib/jwt.ts → used in auth/service.ts (signToken, verifyToken)
@prisma/client → used in auth/repository.ts
process.env.JWT_SECRET → used in auth/service.ts
process.env.DATABASE_URL → used via prisma
Outbound (depends on this feature):
src/app/api/login/route.ts → imports { login } from auth/service
src/app/api/register/route.ts → imports { register } from auth/service
src/middleware.ts → imports { verifyToken } from auth/service
Env vars required: JWT_SECRET, DATABASE_URL
Config files: prisma/schema.prisma (User model)
```
### Phase 3: DIAGNOSE — Find Every Issue
Systematically check for problems. Run ALL of these checks:
**CODE QUALITY:**
- [ ] Every import resolves to a real file/export
- [ ] No circular dependencies within the feature
- [ ] Types are consistent across boundaries (no `any` at interfaces)
- [ ] Error handling exists for all async operations
- [ ] No TODO/FIXME/HACK comments indicating known issues
**RUNTIME:**
- [ ] All required environment variables are set (check .env)
- [ ] Database migrations are up to date (if applicable)
- [ ] API endpoints return expected shapes
- [ ] No hardcoded values that should be configurable
**TESTS:**
- [ ] Run ALL tests related to this feature: find them by searching for imports from the feature folder
- [ ] Record every failure with full error output
- [ ] Check test coverage — are there untested code paths?
**LOGS & ERRORS:**
- [ ] Search for any log files, error reports, or Sentry-style error tracking
- [ ] Check git log for recent changes to this feature: `git log --oneline -20 -- <feature-path>`
- [ ] Check if any recent commits might have broken something: `git log --oneline -5 --all -- <files that this feature depends on>`
**CONFIGURATION:**
- [ ] Verify all config files this feature depends on are valid
- [ ] Check for mismatches between development and production configs
- [ ] Verify third-party service credentials are valid (if testable)
**ROOT-CAUSE CONFIRMATION:**
For each CRITICAL issue found, confirm root cause before adding it to the fix list:
- State clearly: "I think X is the root cause because Y"
- Trace the data/control flow backward to verify — don't trust surface-level symptoms
- If the issue spans multiple components, add diagnostic logging at each boundary to identify which layer fails
- **REQUIRED SUB-SKILL:** For complex bugs found during diagnosis, apply `superpowers:systematic-debugging` Phase 1 (Root Cause Investigation) to confirm before proceeding
**RISK LABELING:**
Assign each issue a risk label:
| Risk | Criteria |
|---|---|
| HIGH | Public API surface / breaking interface contract / DB schema / auth or security logic / widely imported module (>3 callers) / git hotspot |
| MED | Internal module with tests / shared utility / config with runtime impact / internal callers of changed functions |
| LOW | Leaf module / isolated file / test-only change / single-purpose helper with no callers |
Output format:
```
DIAGNOSIS REPORT:
Issues found: N
CRITICAL:
1. [HIGH] [file:line] — description of issue. Root cause: [confirmed explanation]
2. [HIGH] [file:line] — description of issue. Root cause: [confirmed explanation]
WARNINGS:
1. [MED] [file:line] — description of issue
2. [LOW] [file:line] — description of issue
TESTS:
Ran: N tests
Passed: N
Failed: N
[list each failure with one-line summary]
```
### Phase 4: FIX — Repair Everything Systematically
Fix issues in this EXACT order:
1. **DEPENDENCIES FIRST** — fix broken imports, missing packages, wrong versions
2. **TYPES SECOND** — fix type mismatches at feature boundaries
3. **LOGIC THIRD** — fix actual business logic bugs
4. **TESTS FOURTH** — fix or create tests for each fix
5. **INTEGRATION LAST** — verify the feature works end-to-end with its consumers
Rules:
- Fix ONE issue at a time
- After each fix, run the related test to confirm it works
- If a fix breaks something else, STOP and re-evaluate (go back to DIAGNOSE)
- Keep a running log of every change made
- Never change code outside the feature folder without explicitly stating why
- Fix HIGH-risk issues before MED, MED before LOW
**ESCALATION RULE — 3-Strike Architecture Check:**
If 3+ fixes in this phase create NEW issues (not pre-existing ones), STOP immediately.
This pattern indicates an architectural problem, not a bug collection:
- Each fix reveals new shared state / coupling / problem in a different place
- Fixes require "massive refactoring" to implement
- Each fix creates new symptoms elsewhere
**Action:** Stop fixing. Tell the user: "3+ fixes have cascaded into new issues. This suggests the feature's architecture may need rethinking, not patching. Here's what I've found: [summary]. Should we continue fixing symptoms or discuss restructuring?"
Do NOT attempt fix #4 without this discussion.
Output after each fix:
```
FIX #1:
File: auth/service.ts:45
Issue: signToken called with wrong argument order
Change: swapped (expiresIn, payload) to (payload, expiresIn)
Test: auth.test.ts → PASSES
```
### Phase 5: VERIFY — Confirm Everything Works
After all fixes are applied:
1. Run ALL tests in the feature folder — every single one must pass
2. Run ALL tests in files that IMPORT from this feature — must pass
3. Run the full test suite if available — check for regressions
4. If the feature has a UI, describe how to manually verify it
5. Summarize all changes made
Final output:
```
FOCUSED FIX COMPLETE:
Feature: auth
Files changed: 4
Total fixes: 7
Tests: 23/23 passing
Regressions: 0
Changes:
1. auth/service.ts — fixed token signing argument order
2. auth/repository.ts — added null check for user lookup
3. auth/middleware.ts — fixed async error handling
4. auth/types.ts — aligned UserResponse type with actual DB schema
Consumers verified:
- src/app/api/login/route.ts ✅
- src/app/api/register/route.ts ✅
- src/middleware.ts ✅
```
## Red Flags — STOP and Return to Current Phase
If you catch yourself thinking any of these, you are skipping phases:
- "I can see the bug, let me just fix it" → STOP. You haven't traced dependencies yet.
- "Scoping is overkill, it's obviously just this file" → STOP. That's always wrong for feature-level fixes.
- "I'll map dependencies after I fix the obvious stuff" → STOP. You'll miss root causes.
- "The user said fix X, so I only need to look at X" → STOP. Features have dependencies.
- "Tests are passing so I'm done" → STOP. Did you run consumer tests too?
- "I don't need to check env vars for this" → STOP. Config issues masquerade as code bugs.
- "One more fix should do it" (after 2+ cascading failures) → STOP. Escalate.
- "I'll skip the diagnosis report, the fixes are obvious" → STOP. Write it down.
**ALL of these mean: Return to the phase you're supposed to be in.**
## Common Rationalizations
| Excuse | Reality |
|---|---|
| "The feature is small, I don't need all 5 phases" | Small features have dependencies too. Phases 1-2 take minutes for small features — do them. |
| "I already know this codebase" | Knowledge decays. Trace the actual imports, don't rely on memory. |
| "The user wants speed, not process" | Skipping phases causes rework. Systematic is faster than thrashing. |
| "Only one file is broken" | If only one file were broken, the user would say "fix this bug", not "make the feature work." |
| "I fixed the tests, so it works" | Tests can pass while consumers are broken. Verify Phase 5 fully. |
| "The dependency map is too big to trace" | Then the feature is too big to fix without tracing. That's exactly why you need it. |
| "Root cause is obvious, I don't need to confirm" | "Obvious" root causes are wrong 40% of the time. Confirm with evidence. |
| "3 cascading failures is normal for a big fix" | 3 cascading failures means you're patching symptoms of an architectural problem. |
## Anti-Patterns — NEVER do these
| Anti-Pattern | Why It's Wrong |
|---|---|
| Starting to fix code before mapping all dependencies | You'll miss root causes and create whack-a-mole fixes |
| Fixing only the file the user mentioned | Related files likely have issues too |
| Ignoring environment variables and configuration | Many "code bugs" are actually config issues |
| Skipping the test run phase | You can't verify fixes without running tests |
| Making changes outside the feature folder without explaining why | Unexpected side effects confuse the user |
| Fixing symptoms in consumer files instead of root cause in feature | Band-aids that break when the next consumer appears |
| Declaring "done" without running verification tests | Untested fixes are unverified fixes |
| Changing the public API without updating all consumers | Breaks everything that depends on the feature |
## Related Skills
- **`superpowers:systematic-debugging`** — Use within Phase 3 for root-cause tracing of individual complex bugs
- **`superpowers:verification-before-completion`** — Use within Phase 5 before claiming the feature is fixed
- **`scope`** — If you need to understand blast radius before starting, run scope first then focused-fix
## Quick Reference
| Phase | Key Action | Output |
|---|---|---|
| SCOPE | Read every file, map entry points | Feature manifest |
| TRACE | Map inbound + outbound dependencies | Dependency map |
| DIAGNOSE | Check code, runtime, tests, logs, config | Diagnosis report |
| FIX | Fix in order: deps → types → logic → tests → integration | Fix log per issue |
| VERIFY | Run all tests, check consumers, summarize | Completion report |
Tự động chuyển câu hỏi của nhà sáng lập đến cố vấn C-level phù hợp hoặc phiên họp hội đồng cho chủ đề đa vai trò.
---
name: "founder-mode"
description: "/cs:founder-mode <question> — Auto-routes any founder question to the right C-role advisor or to /cs:boardroom for multi-role topics. The single-command entry point."
---
# /cs:founder-mode — The Auto-Router
**Command:** `/cs:founder-mode <question>`
The single command a founder needs to remember. Routes the question to the right C-role automatically, or triggers `/cs:boardroom` if multi-role.
This is the **killer command** — the answer to "I don't know which slash command to use." Type the question; the system figures out the room.
## Routing Logic
The router (via `cs-chief-of-staff`) does keyword + intent matching:
| Signal in question | Route |
|---|---|
| burn, runway, fundraise, dilution, model, LTV, CAC | `cs-cfo-advisor` |
| pipeline, win rate, forecast, NRR, churn, ramp | `cs-cro-advisor` |
| positioning, ICP, message, brand, channel, campaign | `cs-cmo-advisor` |
| roadmap, PMF, JTBD, North Star, RICE, kill | `cs-cpo-advisor` |
| cadence, OKR, scorecard, DRI, operating system, rhythm | `cs-coo-advisor` |
| hiring, comp, ladder, level, attrition, eNPS, equity | `cs-chro-advisor` |
| security, threat, breach, compliance, audit, SOC 2 | `cs-ciso-advisor` |
| architecture, scaling, tech debt, SLO, latency | `cs-cto-advisor` |
| contract, IP, term sheet, regulator, license | `/cs:gc-review` |
| strategy, vision, board, M&A, raise, exit | `cs-ceo-advisor` |
| **2+ signals from different roles** | `/cs:boardroom` |
| **ambiguous** | `/cs:office-hours` first, then route |
## Workflow
1. Parse the question for role signals
2. If exactly one role: invoke that cs-* agent directly
3. If 2+ roles: build a brief via `/cs:brief` and trigger `/cs:boardroom`
4. If ambiguous / no signal match: trigger `/cs:office-hours` to force the founder to sharpen
5. Log the routing decision (raw layer) via `decision-logger`
## Output
The router emits one of three responses:
### Single-role route
```
**Routing:** cs-cfo-advisor
**Why:** Question hits burn rate and unit economics.
**Next:** Invoking cs-cfo-advisor with company-context loaded.
[Advisor's response follows]
```
### Multi-role route
```
**Routing:** /cs:boardroom
**Why:** Question touches CFO + CMO + CPO (pricing change has finance, positioning, and product implications).
**Next:** Building brief via /cs:brief, then running boardroom.
Brief saved: ~/.claude/briefs/2026-05-12-pricing-v3.md
Run: /cs:boardroom ~/.claude/briefs/2026-05-12-pricing-v3.md
```
### Ambiguous → office hours
```
**Routing:** /cs:office-hours
**Why:** Question is too broad ("should we grow faster?"). Need framing before any advisor can help.
**Next:** Six-question intake.
[Office hours questions follow]
```
## Why This Is the Killer Command
gstack requires the founder to know all 23 slash commands and pick the right one. That's a cognitive tax. `/cs:founder-mode` collapses that to one — the system picks. This is also where persistent memory pays off: with company-context.md + decision-logger, the router knows what's already been decided and won't re-litigate.
## Examples
```
/cs:founder-mode "should we raise a Series B now or wait 6 months?"
→ boardroom (CFO + CEO + CRO touched)
/cs:founder-mode "the win rate dropped 20% this month"
→ cs-cro-advisor
/cs:founder-mode "let's hire a VP Marketing"
→ boardroom (CHRO + CMO + CFO touched)
/cs:founder-mode "should we be growing faster?"
→ /cs:office-hours (too ambiguous)
```
## Related
- Agent: [`cs-chief-of-staff`](../../agents/cs-chief-of-staff.md) — does the routing
- Skill: [`chief-of-staff`](../../../skills/chief-of-staff/SKILL.md) — routing logic
- Skill: [`context-engine`](../../../skills/context-engine/SKILL.md) — loads context
---
**Version:** 1.0.0
Thiết kế kiến trúc GCP: triển khai GKE, Cloud Run, pipeline BigQuery, tối ưu chi phí và di chuyển lên Google Cloud.
---
name: "gcp-cloud-architect"
description: "Design GCP architectures for startups and enterprises. Use when asked to design Google Cloud infrastructure, deploy to GKE or Cloud Run, configure BigQuery pipelines, optimize GCP costs, or migrate to GCP. Covers Cloud Run, GKE, Cloud Functions, Cloud SQL, BigQuery, and cost optimization."
---
# GCP Cloud Architect
Design scalable, cost-effective Google Cloud architectures for startups and enterprises with infrastructure-as-code templates.
---
## Workflow
### Step 1: Gather Requirements
Collect application specifications:
```
- Application type (web app, mobile backend, data pipeline, SaaS)
- Expected users and requests per second
- Budget constraints (monthly spend limit)
- Team size and GCP experience level
- Compliance requirements (GDPR, HIPAA, SOC 2)
- Availability requirements (SLA, RPO/RTO)
```
### Step 2: Design Architecture
Run the architecture designer to get pattern recommendations:
```bash
python scripts/architecture_designer.py --input requirements.json
```
**Example output:**
```json
{
"recommended_pattern": "serverless_web",
"service_stack": ["Cloud Storage", "Cloud CDN", "Cloud Run", "Firestore", "Identity Platform"],
"estimated_monthly_cost_usd": 30,
"pros": ["Low ops overhead", "Pay-per-use", "Auto-scaling", "No cold starts on Cloud Run min instances"],
"cons": ["Vendor lock-in", "Regional limitations", "Eventual consistency with Firestore"]
}
```
Select from recommended patterns:
- **Serverless Web**: Cloud Storage + Cloud CDN + Cloud Run + Firestore
- **Microservices on GKE**: GKE Autopilot + Cloud SQL + Memorystore + Cloud Pub/Sub
- **Serverless Data Pipeline**: Pub/Sub + Dataflow + BigQuery + Looker
- **ML Platform**: Vertex AI + Cloud Storage + BigQuery + Cloud Functions
See `references/architecture_patterns.md` for detailed pattern specifications.
**Validation checkpoint:** Confirm the recommended pattern matches the team's operational maturity and compliance requirements before proceeding to Step 3.
### Step 3: Estimate Cost
Analyze estimated costs and optimization opportunities:
```bash
python scripts/cost_optimizer.py --resources current_setup.json --monthly-spend 2000
```
**Example output:**
```json
{
"current_monthly_usd": 2000,
"recommendations": [
{ "action": "Right-size Cloud SQL db-custom-4-16384 to db-custom-2-8192", "savings_usd": 380, "priority": "high" },
{ "action": "Purchase 1-yr committed use discount for GKE nodes", "savings_usd": 290, "priority": "high" },
{ "action": "Move Cloud Storage objects >90 days to Nearline", "savings_usd": 75, "priority": "medium" }
],
"total_potential_savings_usd": 745
}
```
Output includes:
- Monthly cost breakdown by service
- Right-sizing recommendations
- Committed use discount opportunities
- Sustained use discount analysis
- Potential monthly savings
Use the [GCP Pricing Calculator](https://cloud.google.com/products/calculator) for detailed estimates.
### Step 4: Generate IaC
Create infrastructure-as-code for the selected pattern:
```bash
python scripts/deployment_manager.py --app-name my-app --pattern serverless_web --region us-central1
```
**Example Terraform HCL output (Cloud Run + Firestore):**
```hcl
terraform {
required_providers {
google = {
source = "hashicorp/google"
version = "~> 5.0"
}
}
}
provider "google" {
project = var.project_id
region = var.region
}
variable "project_id" {
description = "GCP project ID"
type = string
}
variable "region" {
description = "GCP region"
type = string
default = "us-central1"
}
resource "google_cloud_run_v2_service" "api" {
name = "var.environment-var.app_name-api"
location = var.region
template {
containers {
image = "gcr.io/var.project_id/var.app_name:latest"
resources {
limits = {
cpu = "1000m"
memory = "512Mi"
}
}
env {
name = "FIRESTORE_PROJECT"
value = var.project_id
}
}
scaling {
min_instance_count = 0
max_instance_count = 10
}
}
}
resource "google_firestore_database" "default" {
project = var.project_id
name = "(default)"
location_id = var.region
type = "FIRESTORE_NATIVE"
}
```
**Example gcloud CLI deployment:**
```bash
# Deploy Cloud Run service
gcloud run deploy my-app-api \
--image gcr.io/$PROJECT_ID/my-app:latest \
--region us-central1 \
--platform managed \
--allow-unauthenticated \
--memory 512Mi \
--cpu 1 \
--min-instances 0 \
--max-instances 10
# Create Firestore database
gcloud firestore databases create --location=us-central1
```
> Full templates including Cloud CDN, Identity Platform, IAM, and Cloud Monitoring are generated by `deployment_manager.py` and also available in `references/architecture_patterns.md`.
### Step 5: Configure CI/CD
Set up automated deployment with Cloud Build or GitHub Actions:
```yaml
# cloudbuild.yaml
steps:
- name: 'gcr.io/cloud-builders/docker'
args: ['build', '-t', 'gcr.io/$PROJECT_ID/my-app:$COMMIT_SHA', '.']
- name: 'gcr.io/cloud-builders/docker'
args: ['push', 'gcr.io/$PROJECT_ID/my-app:$COMMIT_SHA']
- name: 'gcr.io/google.com/cloudsdktool/cloud-sdk'
entrypoint: gcloud
args:
- 'run'
- 'deploy'
- 'my-app-api'
- '--image=gcr.io/$PROJECT_ID/my-app:$COMMIT_SHA'
- '--region=us-central1'
- '--platform=managed'
images:
- 'gcr.io/$PROJECT_ID/my-app:$COMMIT_SHA'
```
```bash
# Connect repo and create trigger
gcloud builds triggers create github \
--repo-name=my-app \
--repo-owner=my-org \
--branch-pattern="^main$" \
--build-config=cloudbuild.yaml
```
### Step 6: Security Review
Verify security configuration:
```bash
# Review IAM bindings
gcloud projects get-iam-policy $PROJECT_ID --format=json
# Check service account permissions
gcloud iam service-accounts list --project=$PROJECT_ID
# Verify VPC Service Controls (if applicable)
gcloud access-context-manager perimeters list --policy=$POLICY_ID
```
**Security checklist:**
- IAM roles follow least privilege (prefer predefined roles over basic roles)
- Service accounts use Workload Identity for GKE
- VPC Service Controls configured for sensitive APIs
- Cloud KMS encryption keys for customer-managed encryption
- Cloud Audit Logs enabled for all admin activity
- Organization policies restrict public access
- Secret Manager used for all credentials
**If deployment fails:**
1. Check the failure reason:
```bash
gcloud run services describe my-app-api --region us-central1
gcloud logging read "resource.type=cloud_run_revision" --limit=20
```
2. Review Cloud Logging for application errors.
3. Fix the configuration or container image.
4. Redeploy:
```bash
gcloud run deploy my-app-api --image gcr.io/$PROJECT_ID/my-app:latest --region us-central1
```
**Common failure causes:**
- IAM permission errors -- verify service account roles and `--allow-unauthenticated` flag
- Quota exceeded -- request quota increase via IAM & Admin > Quotas
- Container startup failure -- check container logs and health check configuration
- Region not enabled -- enable the required APIs with `gcloud services enable`
---
## Tools
### architecture_designer.py
Recommends GCP services based on workload requirements.
```bash
python scripts/architecture_designer.py --input requirements.json --output design.json
```
**Input:** JSON with app type, scale, budget, compliance needs
**Output:** Recommended pattern, service stack, cost estimate, pros/cons
### cost_optimizer.py
Analyzes GCP resources for cost savings.
```bash
python scripts/cost_optimizer.py --resources inventory.json --monthly-spend 5000
```
**Output:** Recommendations for:
- Idle resource removal
- Machine type right-sizing
- Committed use discounts
- Storage class transitions
- Network egress optimization
### deployment_manager.py
Generates gcloud CLI deployment scripts and Terraform configurations.
```bash
python scripts/deployment_manager.py --app-name my-app --pattern serverless_web --region us-central1
```
**Output:** Production-ready deployment scripts with:
- Cloud Run or GKE deployment
- Firestore or Cloud SQL setup
- Identity Platform configuration
- IAM roles with least privilege
- Cloud Monitoring and Logging
---
## Quick Start
### Web App on Cloud Run (< $100/month)
```
Ask: "Design a serverless web backend for a mobile app with 1000 users"
Result:
- Cloud Run for API (auto-scaling, no cold start with min instances)
- Firestore for data (pay-per-operation)
- Identity Platform for authentication
- Cloud Storage + Cloud CDN for static assets
- Estimated: $15-40/month
```
### Microservices on GKE ($500-2000/month)
```
Ask: "Design a scalable architecture for a SaaS platform with 50k users"
Result:
- GKE Autopilot for containerized workloads
- Cloud SQL (PostgreSQL) with read replicas
- Memorystore (Redis) for session caching
- Cloud CDN for global delivery
- Cloud Build for CI/CD
- Multi-zone deployment
```
### Serverless Data Pipeline
```
Ask: "Design a real-time analytics pipeline for event data"
Result:
- Pub/Sub for event ingestion
- Dataflow (Apache Beam) for stream processing
- BigQuery for analytics and warehousing
- Looker for dashboards
- Cloud Functions for lightweight transforms
```
### ML Platform
```
Ask: "Design a machine learning platform for model training and serving"
Result:
- Vertex AI for training and prediction
- Cloud Storage for datasets and model artifacts
- BigQuery for feature store
- Cloud Functions for preprocessing triggers
- Cloud Monitoring for model drift detection
```
---
## Input Requirements
Provide these details for architecture design:
| Requirement | Description | Example |
|-------------|-------------|---------|
| Application type | What you're building | SaaS platform, mobile backend |
| Expected scale | Users, requests/sec | 10k users, 100 RPS |
| Budget | Monthly GCP limit | $500/month max |
| Team context | Size, GCP experience | 3 devs, intermediate |
| Compliance | Regulatory needs | HIPAA, GDPR, SOC 2 |
| Availability | Uptime requirements | 99.9% SLA, 1hr RPO |
**JSON Format:**
```json
{
"application_type": "saas_platform",
"expected_users": 10000,
"requests_per_second": 100,
"budget_monthly_usd": 500,
"team_size": 3,
"gcp_experience": "intermediate",
"compliance": ["SOC2"],
"availability_sla": "99.9%"
}
```
---
## Output Formats
### Architecture Design
- Pattern recommendation with rationale
- Service stack diagram (ASCII)
- Monthly cost estimate and trade-offs
### IaC Templates
- **Terraform HCL**: Production-ready Google provider configs
- **gcloud CLI**: Scripted deployment commands
- **Cloud Build YAML**: CI/CD pipeline definitions
### Cost Analysis
- Current spend breakdown with optimization recommendations
- Priority action list (high/medium/low) and implementation checklist
---
## Anti-Patterns
| Anti-Pattern | Why It Fails | Better Approach |
|---|---|---|
| Using default VPC for production | No isolation, shared firewall rules | Create custom VPC with private subnets |
| Over-provisioning GKE node pools | Wasted cost on idle capacity | Use GKE Autopilot or cluster autoscaler |
| Storing secrets in environment variables | Visible in Cloud Console, logs | Use Secret Manager with Workload Identity |
| Ignoring sustained use discounts | Missing 20-30% automatic savings | Right-size VMs for consistent baseline usage |
| Single-region deployment for SaaS | One region outage = full downtime | Multi-region with Cloud Load Balancing |
| BigQuery on-demand for heavy workloads | Unpredictable costs at scale | Use BigQuery slots (flat-rate) for consistent workloads |
| Running Cloud Functions for long tasks | 9-minute timeout, cold starts | Use Cloud Run for tasks > 60 seconds |
---
## Cross-References
| Skill | Relationship |
|-------|-------------|
| `engineering-team/aws-solution-architect` | AWS equivalent — same 6-step workflow, different services |
| `engineering-team/azure-cloud-architect` | Azure equivalent — completes the cloud trifecta |
| `engineering-team/senior-devops` | Broader DevOps scope — pipelines, monitoring, containerization |
| `engineering/terraform-patterns` | IaC implementation — use for Terraform modules targeting GCP |
| `engineering/ci-cd-pipeline-builder` | Pipeline construction — automates Cloud Build and deployment |
---
## Reference Documentation
| Document | Contents |
|----------|----------|
| `references/architecture_patterns.md` | 6 patterns: serverless, GKE microservices, three-tier, data pipeline, ML platform, multi-region |
| `references/service_selection.md` | Decision matrices for compute, database, storage, messaging |
| `references/best_practices.md` | Naming, labels, IAM, networking, monitoring, disaster recovery |
FILE:references/architecture_patterns.md
# GCP Architecture Patterns
Reference guide for selecting the right GCP architecture pattern based on application requirements.
---
## Table of Contents
- [Pattern Selection Matrix](#pattern-selection-matrix)
- [Pattern 1: Serverless Web Application](#pattern-1-serverless-web-application)
- [Pattern 2: Microservices on GKE](#pattern-2-microservices-on-gke)
- [Pattern 3: Three-Tier Application](#pattern-3-three-tier-application)
- [Pattern 4: Serverless Data Pipeline](#pattern-4-serverless-data-pipeline)
- [Pattern 5: ML Platform](#pattern-5-ml-platform)
- [Pattern 6: Multi-Region High Availability](#pattern-6-multi-region-high-availability)
---
## Pattern Selection Matrix
| Pattern | Best For | Users | Monthly Cost | Complexity |
|---------|----------|-------|--------------|------------|
| Serverless Web | MVP, SaaS, mobile backend | <50K | $30-400 | Low |
| Microservices on GKE | Complex services, enterprise | 10K-500K | $400-2500 | Medium |
| Three-Tier | Traditional web, e-commerce | 10K-200K | $300-1500 | Medium |
| Data Pipeline | Analytics, ETL, streaming | Any | $100-2000 | Medium-High |
| ML Platform | Training, serving, MLOps | Any | $200-5000 | High |
| Multi-Region HA | Global apps, DR | >100K | 2x single | High |
---
## Pattern 1: Serverless Web Application
### Use Case
SaaS platforms, mobile backends, low-traffic websites, MVPs
### Architecture Diagram
```
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ Cloud CDN │────▶│ Cloud │ │ Identity │
│ (CDN) │ │ Storage │ │ Platform │
└─────────────┘ │ (Static) │ │ (Auth) │
└─────────────┘ └──────┬──────┘
│
┌─────────────┐ ┌─────────────┐ ┌──────▼──────┐
│ Cloud DNS │────▶│ Cloud │────▶│ Cloud Run │
│ (DNS) │ │ Load Bal. │ │ (API) │
└─────────────┘ └─────────────┘ └──────┬──────┘
│
┌──────▼──────┐
│ Firestore │
│ (Database) │
└─────────────┘
```
### Service Stack
| Layer | Service | Configuration |
|-------|---------|---------------|
| Frontend | Cloud Storage + Cloud CDN | Static hosting with HTTPS |
| API | Cloud Run | Containerized API with auto-scaling |
| Database | Firestore | Native mode, pay-per-operation |
| Auth | Identity Platform | Multi-provider authentication |
| CI/CD | Cloud Build | Automated container deployments |
### Terraform Example
```hcl
resource "google_cloud_run_v2_service" "api" {
name = "my-app-api"
location = "us-central1"
template {
containers {
image = "gcr.io/my-project/my-app:latest"
resources {
limits = {
cpu = "1000m"
memory = "512Mi"
}
}
}
scaling {
min_instance_count = 0
max_instance_count = 10
}
}
}
```
### Cost Breakdown (10K users)
| Service | Monthly Cost |
|---------|-------------|
| Cloud Run | $5-25 |
| Firestore | $5-30 |
| Cloud CDN | $5-15 |
| Cloud Storage | $1-5 |
| Identity Platform | $0-10 |
| **Total** | **$16-85** |
### Pros and Cons
**Pros:**
- Scale-to-zero (pay nothing when idle)
- Container-based (no runtime restrictions)
- Built-in HTTPS and custom domains
- Auto-scaling with no configuration
**Cons:**
- Cold starts if min instances = 0
- Firestore query limitations vs SQL
- Vendor lock-in to GCP
---
## Pattern 2: Microservices on GKE
### Use Case
Complex business systems, enterprise applications, platform engineering
### Architecture Diagram
```
┌─────────────┐ ┌─────────────┐
│ Cloud CDN │────▶│ Global │
│ (CDN) │ │ Load Bal. │
└─────────────┘ └──────┬──────┘
│
┌──────▼──────┐
│ GKE │
│ Autopilot │
└──────┬──────┘
│
┌──────────────────┼──────────────────┐
│ │ │
┌──────▼──────┐ ┌──────▼──────┐ ┌──────▼──────┐
│ Cloud SQL │ │ Memorystore │ │ Pub/Sub │
│ (Postgres) │ │ (Redis) │ │ (Messaging) │
└─────────────┘ └─────────────┘ └─────────────┘
```
### Service Stack
| Layer | Service | Configuration |
|-------|---------|---------------|
| CDN | Cloud CDN | Edge caching, HTTPS |
| Load Balancer | Global Application LB | Backend services, health checks |
| Compute | GKE Autopilot | Managed node provisioning |
| Database | Cloud SQL PostgreSQL | Regional HA, read replicas |
| Cache | Memorystore Redis | Session, query caching |
| Messaging | Pub/Sub | Async service communication |
### GKE Autopilot Configuration
```yaml
# Deployment manifest
apiVersion: apps/v1
kind: Deployment
metadata:
name: api-service
spec:
replicas: 2
selector:
matchLabels:
app: api-service
template:
metadata:
labels:
app: api-service
spec:
serviceAccountName: api-workload-sa
containers:
- name: api
image: us-central1-docker.pkg.dev/my-project/my-app/api:latest
ports:
- containerPort: 8080
resources:
requests:
cpu: "500m"
memory: "512Mi"
limits:
cpu: "1000m"
memory: "1Gi"
env:
- name: DB_HOST
valueFrom:
secretKeyRef:
name: db-credentials
key: host
readinessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 5
periodSeconds: 10
---
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api-service-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: api-service
minReplicas: 2
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
```
### Cost Breakdown (50K users)
| Service | Monthly Cost |
|---------|-------------|
| GKE Autopilot | $150-400 |
| Cloud Load Balancing | $25-50 |
| Cloud SQL | $100-300 |
| Memorystore | $40-80 |
| Pub/Sub | $5-20 |
| **Total** | **$320-850** |
---
## Pattern 3: Three-Tier Application
### Use Case
Traditional web apps, e-commerce, CMS, applications with complex queries
### Architecture Diagram
```
┌─────────────┐ ┌─────────────┐
│ Cloud CDN │────▶│ Global │
│ (CDN) │ │ Load Bal. │
└─────────────┘ └──────┬──────┘
│
┌──────▼──────┐
│ Cloud Run │
│ (or MIG) │
└──────┬──────┘
│
┌──────────────────┼──────────────────┐
│ │ │
┌──────▼──────┐ ┌──────▼──────┐ ┌──────▼──────┐
│ Cloud SQL │ │ Memorystore │ │ Cloud │
│ (Database) │ │ (Redis) │ │ Storage │
└─────────────┘ └─────────────┘ └─────────────┘
```
### Service Stack
| Layer | Service | Configuration |
|-------|---------|---------------|
| CDN | Cloud CDN | Edge caching, compression |
| Load Balancer | External Application LB | SSL termination, health checks |
| Compute | Cloud Run or Managed Instance Group | Auto-scaling containers or VMs |
| Database | Cloud SQL (MySQL/PostgreSQL) | Regional HA, automated backups |
| Cache | Memorystore Redis | Session store, query cache |
| Storage | Cloud Storage | Uploads, static assets, backups |
### Cost Breakdown (50K users)
| Service | Monthly Cost |
|---------|-------------|
| Cloud Run / MIG | $80-200 |
| Cloud Load Balancing | $25-50 |
| Cloud SQL | $100-250 |
| Memorystore | $30-60 |
| Cloud Storage | $10-30 |
| **Total** | **$245-590** |
---
## Pattern 4: Serverless Data Pipeline
### Use Case
Analytics, IoT data ingestion, log processing, real-time streaming, ETL
### Architecture Diagram
```
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ Sources │────▶│ Pub/Sub │────▶│ Dataflow │
│ (Apps/IoT) │ │ (Ingest) │ │ (Process) │
└─────────────┘ └─────────────┘ └──────┬──────┘
│
┌─────────────┐ ┌─────────────┐ ┌──────▼──────┐
│ Looker │◀────│ BigQuery │◀────│ Cloud │
│ (Dashbd) │ │(Warehouse) │ │ Storage │
└─────────────┘ └─────────────┘ │ (Data Lake) │
└─────────────┘
```
### Service Stack
| Layer | Service | Purpose |
|-------|---------|---------|
| Ingestion | Pub/Sub | Real-time event capture |
| Processing | Dataflow (Apache Beam) | Stream/batch transforms |
| Warehouse | BigQuery | SQL analytics at scale |
| Storage | Cloud Storage | Raw data lake |
| Visualization | Looker / Looker Studio | Dashboards and reports |
| Orchestration | Cloud Composer (Airflow) | Pipeline scheduling |
### Dataflow Pipeline Example
```python
import apache_beam as beam
from apache_beam.options.pipeline_options import PipelineOptions
options = PipelineOptions([
'--runner=DataflowRunner',
'--project=my-project',
'--region=us-central1',
'--temp_location=gs://my-bucket/temp',
'--streaming'
])
with beam.Pipeline(options=options) as p:
(p
| 'ReadPubSub' >> beam.io.ReadFromPubSub(topic='projects/my-project/topics/events')
| 'ParseJSON' >> beam.Map(lambda x: json.loads(x))
| 'WindowInto' >> beam.WindowInto(beam.window.FixedWindows(60))
| 'WriteBQ' >> beam.io.WriteToBigQuery(
'my-project:analytics.events',
schema='event_id:STRING,event_type:STRING,timestamp:TIMESTAMP',
write_disposition=beam.io.BigQueryDisposition.WRITE_APPEND
))
```
### Cost Breakdown
| Service | Monthly Cost |
|---------|-------------|
| Pub/Sub | $5-30 |
| Dataflow | $30-200 |
| BigQuery (on-demand) | $10-100 |
| Cloud Storage | $5-30 |
| Looker Studio | $0 (free) |
| **Total** | **$50-360** |
---
## Pattern 5: ML Platform
### Use Case
Model training, serving, MLOps, feature engineering
### Architecture Diagram
```
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ BigQuery │────▶│ Vertex AI │────▶│ Vertex AI │
│ (Features) │ │ (Training) │ │ (Endpoints) │
└─────────────┘ └──────┬──────┘ └─────────────┘
│
┌─────────────┐ ┌──────▼──────┐ ┌─────────────┐
│ Cloud │◀────│ Cloud │────▶│ Vertex AI │
│ Functions │ │ Storage │ │ Pipelines │
│ (Triggers) │ │ (Artifacts) │ │ (MLOps) │
└─────────────┘ └─────────────┘ └─────────────┘
```
### Service Stack
| Layer | Service | Purpose |
|-------|---------|---------|
| Data | BigQuery | Feature engineering, exploration |
| Training | Vertex AI Training | Custom or AutoML training |
| Serving | Vertex AI Endpoints | Online/batch prediction |
| Storage | Cloud Storage | Datasets, model artifacts |
| Orchestration | Vertex AI Pipelines | ML workflow automation |
| Monitoring | Vertex AI Model Monitoring | Drift and skew detection |
### Vertex AI Training Example
```python
from google.cloud import aiplatform
aiplatform.init(project='my-project', location='us-central1')
job = aiplatform.CustomTrainingJob(
display_name='my-model-training',
script_path='train.py',
container_uri='us-docker.pkg.dev/vertex-ai/training/tf-gpu.2-12:latest',
requirements=['pandas', 'scikit-learn'],
)
model = job.run(
replica_count=1,
machine_type='n1-standard-8',
accelerator_type='NVIDIA_TESLA_T4',
accelerator_count=1,
)
endpoint = model.deploy(
deployed_model_display_name='my-model-v1',
machine_type='n1-standard-4',
min_replica_count=1,
max_replica_count=5,
)
```
### Cost Breakdown
| Service | Monthly Cost |
|---------|-------------|
| Vertex AI Training (T4 GPU) | $50-500 |
| Vertex AI Prediction | $30-200 |
| BigQuery | $10-50 |
| Cloud Storage | $5-30 |
| **Total** | **$95-780** |
---
## Pattern 6: Multi-Region High Availability
### Use Case
Global applications, disaster recovery, data sovereignty compliance
### Architecture Diagram
```
┌─────────────┐
│ Cloud DNS │
│(Geo routing)│
└──────┬──────┘
│
┌────────────────┼────────────────┐
│ │
┌──────▼──────┐ ┌──────▼──────┐
│us-central1 │ │europe-west1 │
│ Cloud Run │ │ Cloud Run │
└──────┬──────┘ └──────┬──────┘
│ │
┌──────▼──────┐ ┌──────▼──────┐
│Cloud Spanner│◀── Replication ──▶│Cloud Spanner│
│ (Region) │ │ (Region) │
└─────────────┘ └─────────────┘
```
### Service Stack
| Component | Service | Configuration |
|-----------|---------|---------------|
| DNS | Cloud DNS | Geolocation or latency routing |
| CDN | Cloud CDN | Multiple regional origins |
| Compute | Cloud Run (multi-region) | Deployed in each region |
| Database | Cloud Spanner (multi-region) | Strong global consistency |
| Storage | Cloud Storage (multi-region) | Automatic geo-redundancy |
### Cloud DNS Geolocation Policy
```bash
# Create geolocation routing policy
gcloud dns record-sets create api.example.com \
--zone=my-zone \
--type=A \
--routing-policy-type=GEO \
--routing-policy-data="us-central1=projects/my-project/regions/us-central1/addresses/api-us;europe-west1=projects/my-project/regions/europe-west1/addresses/api-eu"
```
### Cost Considerations
| Factor | Impact |
|--------|--------|
| Compute | 2x (each region) |
| Cloud Spanner | Multi-region 3x regional price |
| Data Transfer | Cross-region replication costs |
| Cloud DNS | Geolocation queries premium |
| **Total** | **2-3x single region** |
---
## Pattern Comparison Summary
### Latency
| Pattern | Typical Latency |
|---------|-----------------|
| Serverless Web | 30-150ms (Cloud Run) |
| GKE Microservices | 15-80ms |
| Three-Tier | 20-100ms |
| Multi-Region | <50ms (regional) |
### Scaling Characteristics
| Pattern | Scale Limit | Scale Speed |
|---------|-------------|-------------|
| Serverless Web | 1000 instances/service | Seconds |
| GKE Microservices | Cluster node limits | Minutes |
| Data Pipeline | Unlimited (Dataflow) | Seconds |
| Multi-Region | Regional limits | Seconds |
### Operational Complexity
| Pattern | Setup | Maintenance | Debugging |
|---------|-------|-------------|-----------|
| Serverless Web | Low | Low | Medium |
| GKE Microservices | Medium | Medium | Medium |
| Three-Tier | Medium | Medium | Low |
| Data Pipeline | High | Medium | High |
| ML Platform | High | High | High |
| Multi-Region | High | High | High |
FILE:references/best_practices.md
# GCP Best Practices
Production-ready practices for naming, labels, IAM, networking, monitoring, and disaster recovery.
---
## Table of Contents
- [Naming Conventions](#naming-conventions)
- [Labels and Organization](#labels-and-organization)
- [IAM and Security](#iam-and-security)
- [Networking](#networking)
- [Monitoring and Logging](#monitoring-and-logging)
- [Cost Optimization](#cost-optimization)
- [Disaster Recovery](#disaster-recovery)
- [Common Pitfalls](#common-pitfalls)
---
## Naming Conventions
### Resource Naming Pattern
```
{environment}-{project}-{resource-type}-{purpose}
Examples:
prod-myapp-gke-cluster
dev-myapp-sql-primary
staging-myapp-run-api
prod-myapp-gcs-uploads
```
### Project Naming
```
{org}-{team}-{environment}
Examples:
acme-platform-prod
acme-platform-dev
acme-data-prod
```
### Naming Rules
| Resource | Format | Max Length | Example |
|----------|--------|-----------|---------|
| Project ID | lowercase, hyphens | 30 chars | acme-platform-prod |
| GKE Cluster | lowercase, hyphens | 40 chars | prod-api-cluster |
| Cloud Run | lowercase, hyphens | 49 chars | prod-myapp-api |
| Cloud SQL | lowercase, hyphens | 84 chars | prod-myapp-sql-primary |
| GCS Bucket | lowercase, hyphens, dots | 63 chars | acme-prod-myapp-uploads |
| Service Account | lowercase, hyphens | 30 chars | myapp-run-sa |
---
## Labels and Organization
### Required Labels
Apply these labels to all resources:
```
labels:
environment: "prod" # dev, staging, prod
team: "platform" # team owning the resource
app: "myapp" # application name
cost-center: "eng-001" # billing allocation
managed-by: "terraform" # terraform, gcloud, console
```
### Label-Based Cost Reporting
```bash
# Export billing data to BigQuery with labels
# Then query by label:
SELECT
labels.value AS environment,
SUM(cost) AS total_cost
FROM `billing_export.gcp_billing_export_v1_*`
CROSS JOIN UNNEST(labels) AS labels
WHERE labels.key = 'environment'
GROUP BY environment
ORDER BY total_cost DESC
```
### Organization Hierarchy
```
Organization
├── Folder: Production
│ ├── Project: platform-prod
│ ├── Project: data-prod
│ └── Project: ml-prod
├── Folder: Non-Production
│ ├── Project: platform-dev
│ ├── Project: platform-staging
│ └── Project: data-dev
└── Folder: Shared Services
├── Project: shared-networking
├── Project: shared-security
└── Project: shared-monitoring
```
---
## IAM and Security
### Principle of Least Privilege
```bash
# BAD: Basic roles are too broad
gcloud projects add-iam-policy-binding my-project \
--member="user:dev@example.com" \
--role="roles/editor"
# GOOD: Use predefined roles
gcloud projects add-iam-policy-binding my-project \
--member="user:dev@example.com" \
--role="roles/run.developer"
```
### Service Account Best Practices
```bash
# 1. Create dedicated SA per workload
gcloud iam service-accounts create myapp-api-sa \
--display-name="MyApp API Service Account"
# 2. Grant only required roles
gcloud projects add-iam-policy-binding my-project \
--member="serviceAccount:myapp-api-sa@my-project.iam.gserviceaccount.com" \
--role="roles/datastore.user"
# 3. Use Workload Identity for GKE (no key files)
gcloud iam service-accounts add-iam-policy-binding \
myapp-api-sa@my-project.iam.gserviceaccount.com \
--role="roles/iam.workloadIdentityUser" \
--member="serviceAccount:my-project.svc.id.goog[default/myapp-api-ksa]"
# 4. NEVER download SA key files in production
# Instead, use attached service accounts or impersonation
```
### VPC Service Controls
```bash
# Create a service perimeter to restrict data exfiltration
gcloud access-context-manager perimeters create my-perimeter \
--title="Production Data Perimeter" \
--resources="projects/123456" \
--restricted-services="bigquery.googleapis.com,storage.googleapis.com" \
--policy=$POLICY_ID
```
### Organization Policies
```bash
# Restrict external IPs on VMs
gcloud resource-manager org-policies set-policy \
--project=my-project policy.yaml
# policy.yaml
constraint: compute.vmExternalIpAccess
listPolicy:
allValues: DENY
# Restrict public Cloud Storage
constraint: storage.publicAccessPrevention
booleanPolicy:
enforced: true
```
### Encryption
| Layer | Service | Default |
|-------|---------|---------|
| At rest | Google-managed keys | Always enabled |
| At rest | CMEK (Cloud KMS) | Optional, recommended |
| In transit | TLS 1.3 | Always enabled |
| Application | Cloud KMS | Encrypt sensitive fields |
```bash
# Create CMEK key for Cloud SQL
gcloud kms keys create myapp-sql-key \
--keyring=myapp-keyring \
--location=us-central1 \
--purpose=encryption
# Use CMEK with Cloud SQL
gcloud sql instances create myapp-db \
--disk-encryption-key=projects/my-project/locations/us-central1/keyRings/myapp-keyring/cryptoKeys/myapp-sql-key
```
---
## Networking
### VPC Design
```bash
# Create custom VPC (avoid default network)
gcloud compute networks create myapp-vpc \
--subnet-mode=custom
# Create subnets with secondary ranges for GKE
gcloud compute networks subnets create myapp-subnet \
--network=myapp-vpc \
--region=us-central1 \
--range=10.0.0.0/20 \
--secondary-range pods=10.4.0.0/14,services=10.8.0.0/20 \
--enable-private-google-access
```
### Shared VPC
Use Shared VPC for multi-project environments:
```
Host Project (shared-networking)
├── VPC: shared-vpc
│ ├── Subnet: prod-us-central1 → Service Project: platform-prod
│ ├── Subnet: prod-europe-west1 → Service Project: platform-prod
│ └── Subnet: dev-us-central1 → Service Project: platform-dev
```
### Firewall Rules
```bash
# Allow internal traffic
gcloud compute firewall-rules create allow-internal \
--network=myapp-vpc \
--allow=tcp,udp,icmp \
--source-ranges=10.0.0.0/8
# Allow health checks from Google load balancers
gcloud compute firewall-rules create allow-health-checks \
--network=myapp-vpc \
--allow=tcp:8080 \
--source-ranges=35.191.0.0/16,130.211.0.0/22 \
--target-tags=allow-health-check
# Deny all other ingress (implicit, but be explicit)
gcloud compute firewall-rules create deny-all-ingress \
--network=myapp-vpc \
--action=DENY \
--rules=all \
--direction=INGRESS \
--priority=65534
```
### Private Google Access
Always enable Private Google Access to reach GCP APIs without public IPs:
```bash
gcloud compute networks subnets update myapp-subnet \
--region=us-central1 \
--enable-private-google-access
```
---
## Monitoring and Logging
### Cloud Monitoring Setup
```bash
# Create uptime check
gcloud monitoring uptime create \
--display-name="API Health Check" \
--resource-type=cloud-run-revision \
--resource-labels="service_name=myapp-api,location=us-central1" \
--check-request-path="/health" \
--period=60s
# Create alerting policy
gcloud alpha monitoring policies create \
--display-name="High Error Rate" \
--condition-display-name="Cloud Run 5xx > 1%" \
--condition-filter='resource.type="cloud_run_revision" AND metric.type="run.googleapis.com/request_count" AND metric.labels.response_code_class="5xx"' \
--condition-threshold-value=1 \
--notification-channels="projects/my-project/notificationChannels/12345"
```
### Key Metrics to Monitor
| Service | Metric | Alert Threshold |
|---------|--------|-----------------|
| Cloud Run | request_latencies (p99) | >2s |
| Cloud Run | request_count (5xx) | >1% of total |
| Cloud SQL | cpu/utilization | >80% |
| Cloud SQL | disk/utilization | >85% |
| GKE | container/cpu/utilization | >80% |
| GKE | node/cpu/allocatable_utilization | >85% |
| Pub/Sub | subscription/oldest_unacked_message_age | >300s |
| BigQuery | query/execution_time | >60s |
### Log-Based Metrics
```bash
# Create a metric for application errors
gcloud logging metrics create app-errors \
--description="Application error count" \
--log-filter='resource.type="cloud_run_revision" AND severity>=ERROR'
# Create log sink to BigQuery for analysis
gcloud logging sinks create audit-logs-bq \
bigquery.googleapis.com/projects/my-project/datasets/audit_logs \
--log-filter='logName="projects/my-project/logs/cloudaudit.googleapis.com%2Factivity"'
```
### Log Exclusion (Cost Reduction)
```bash
# Exclude verbose debug logs to save on Cloud Logging costs
gcloud logging sinks create _Default \
--log-filter='NOT (severity="DEBUG" OR severity="DEFAULT")' \
--description="Exclude debug-level logs"
# Or create exclusion filters
gcloud logging exclusions create exclude-debug \
--log-filter='severity="DEBUG"' \
--description="Exclude debug logs to reduce costs"
```
---
## Cost Optimization
### Committed Use Discounts
| Term | Compute Discount | Memory Discount |
|------|-----------------|-----------------|
| 1 year | 37% | 37% |
| 3 years | 55% | 55% |
```bash
# Check recommendations
gcloud recommender recommendations list \
--project=my-project \
--location=us-central1 \
--recommender=google.compute.commitment.UsageCommitmentRecommender
```
### Sustained Use Discounts
Automatic discounts for resources running >25% of the month:
| Usage | Discount |
|-------|----------|
| 25-50% | 20% |
| 50-75% | 40% |
| 75-100% | 60% |
### BigQuery Cost Control
```sql
-- Use partitioning to limit data scanned
CREATE TABLE my_dataset.events
PARTITION BY DATE(timestamp)
CLUSTER BY event_type
AS SELECT * FROM raw_events;
-- Estimate query cost before running
-- Use --dry_run flag
bq query --dry_run --use_legacy_sql=false \
'SELECT * FROM my_dataset.events WHERE DATE(timestamp) = "2026-01-01"'
```
### Cloud Storage Optimization
```bash
# Enable Autoclass for automatic class management
gsutil mb -l us-central1 --autoclass gs://my-bucket/
# Set lifecycle policy
gsutil lifecycle set lifecycle.json gs://my-bucket/
```
---
## Disaster Recovery
### RPO/RTO Targets
| Tier | RPO | RTO | Strategy |
|------|-----|-----|----------|
| Tier 1 (Critical) | 0 | <1 hour | Multi-region active-active |
| Tier 2 (Important) | <1 hour | <4 hours | Regional HA + cross-region backup |
| Tier 3 (Standard) | <24 hours | <24 hours | Automated backups + restore |
### Backup Strategy
```bash
# Cloud SQL automated backups
gcloud sql instances patch myapp-db \
--backup-start-time=02:00 \
--enable-point-in-time-recovery
# Firestore scheduled exports
gcloud firestore export gs://myapp-backups/firestore/$(date +%Y%m%d)
# GKE cluster backup with Backup for GKE
gcloud beta container backup-restore backup-plans create myapp-plan \
--project=my-project \
--location=us-central1 \
--cluster=projects/my-project/locations/us-central1/clusters/myapp-cluster \
--all-namespaces \
--cron-schedule="0 2 * * *"
```
### Multi-Region Failover
```bash
# Cloud SQL cross-region replica for DR
gcloud sql instances create myapp-db-replica \
--master-instance-name=myapp-db \
--region=us-east1
# Promote replica during failover
gcloud sql instances promote-replica myapp-db-replica
```
---
## Common Pitfalls
### Technical Debt
| Pitfall | Solution |
|---------|----------|
| Using default VPC | Always create custom VPCs |
| Not enabling audit logs | Enable Cloud Audit Logs from day one |
| Single-region deployment | Plan for multi-zone at minimum |
| No IaC | Use Terraform from the start |
### Security Mistakes
| Mistake | Prevention |
|---------|------------|
| SA key files in code | Use Workload Identity, attached SAs |
| Public GCS buckets | Enable org policy for public access prevention |
| Basic roles (Owner/Editor) | Use predefined or custom roles |
| No encryption key management | Use CMEK for sensitive data |
| Default service account | Create dedicated SAs per workload |
### Performance Issues
| Issue | Solution |
|-------|----------|
| Cold starts on Cloud Run | Set min-instances=1 for latency-critical services |
| Slow BigQuery queries | Partition tables, use clustering, avoid SELECT * |
| GKE pod scheduling delays | Use PodDisruptionBudget, pre-provision with Autopilot |
| Firestore hotspots | Distribute writes across document IDs evenly |
### Cost Surprises
| Surprise | Prevention |
|----------|------------|
| Undeleted resources | Label everything, review weekly |
| Egress costs | Keep traffic in same region, use Private Google Access |
| Cloud NAT charges | Use Private Google Access for GCP service traffic |
| Log ingestion costs | Set exclusion filters for debug/verbose logs |
| BigQuery full scans | Always use partitioning and clustering |
| Idle GKE clusters | Delete dev clusters nightly, use Autopilot |
FILE:references/service_selection.md
# GCP Service Selection Guide
Quick reference for choosing the right GCP service based on requirements.
---
## Table of Contents
- [Compute Services](#compute-services)
- [Database Services](#database-services)
- [Storage Services](#storage-services)
- [Messaging and Events](#messaging-and-events)
- [API and Integration](#api-and-integration)
- [Networking](#networking)
- [Security and Identity](#security-and-identity)
---
## Compute Services
### Decision Matrix
| Requirement | Recommended Service |
|-------------|---------------------|
| HTTP-triggered containers, auto-scaling | Cloud Run |
| Event-driven, short tasks (<9 min) | Cloud Functions (2nd gen) |
| Kubernetes workloads, microservices | GKE Autopilot |
| Custom VMs, GPU/TPU | Compute Engine |
| Batch processing, HPC | Batch |
| Kubernetes with full control | GKE Standard |
### Cloud Run
**Best for:** Containerized HTTP services, APIs, web backends
```
Limits:
- vCPU: 1-8 per instance
- Memory: 128 MiB - 32 GiB
- Request timeout: 3600 seconds
- Concurrency: 1-1000 per instance
- Min instances: 0 (scale-to-zero)
- Max instances: 1000
Pricing: Per vCPU-second + GiB-second (free tier: 2M requests/month)
```
**Use when:**
- Containerized apps with HTTP endpoints
- Variable/unpredictable traffic
- Want scale-to-zero capability
- No Kubernetes expertise needed
**Avoid when:**
- Non-HTTP workloads (use Cloud Functions or GKE)
- Need GPU/TPU (use Compute Engine or GKE)
- Require persistent local storage
### Cloud Functions (2nd gen)
**Best for:** Event-driven functions, lightweight triggers, webhooks
```
Limits:
- Execution: 9 minutes max (2nd gen), 9 minutes (1st gen)
- Memory: 128 MB - 32 GB
- Concurrency: Up to 1000 per instance (2nd gen)
- Runtimes: Node.js, Python, Go, Java, .NET, Ruby, PHP
Pricing: $0.40 per million invocations + compute time
```
**Use when:**
- Event-driven processing (Pub/Sub, Cloud Storage, Firestore)
- Lightweight API endpoints
- Scheduled tasks (Cloud Scheduler triggers)
- Minimal infrastructure management
**Avoid when:**
- Long-running processes (>9 min)
- Complex multi-container apps
- Need fine-grained scaling control
### GKE Autopilot
**Best for:** Kubernetes workloads with managed node provisioning
```
Limits:
- Pod resources: 0.25-112 vCPU, 0.5-896 GiB memory
- GPU support: NVIDIA T4, L4, A100, H100
- Management fee: $0.10/hour per cluster ($74.40/month)
Pricing: Per pod vCPU-hour + GiB-hour (no node management)
```
**Use when:**
- Team has Kubernetes expertise
- Need pod-level resource control
- Multi-container services
- GPU workloads
### Compute Engine
**Best for:** Custom configurations, specialized hardware
```
Machine Types:
- General: e2, n2, n2d, c3
- Compute: c2, c2d
- Memory: m1, m2, m3
- Accelerator: a2 (GPU), a3 (GPU)
- Storage: z3
Pricing Options:
- On-demand, Spot (60-91% discount), Committed Use (37-55% discount)
```
**Use when:**
- Need GPU/TPU
- Windows workloads
- Specific hardware requirements
- Lift-and-shift migrations
---
## Database Services
### Decision Matrix
| Data Type | Query Pattern | Scale | Recommended |
|-----------|--------------|-------|-------------|
| Key-value, document | Simple lookups, real-time | Any | Firestore |
| Wide-column | High-throughput reads/writes | >1TB | Cloud Bigtable |
| Relational | Complex joins, ACID | Variable | Cloud SQL |
| Relational, global | Strong consistency, global | Large | Cloud Spanner |
| Time-series | Time-based queries | Any | Bigtable or BigQuery |
| Analytics, warehouse | SQL analytics | Petabytes | BigQuery |
### Firestore
**Best for:** Document data, mobile/web apps, real-time sync
```
Limits:
- Document size: 1 MiB max
- Field depth: 20 nested levels
- Write rate: 10,000 writes/sec per database
- Indexes: Automatic single-field, manual composite
Pricing:
- Reads: $0.036 per 100K reads
- Writes: $0.108 per 100K writes
- Storage: $0.108 per GiB/month
- Free tier: 50K reads, 20K writes, 1 GiB storage per day
```
**Use when:**
- Mobile/web apps needing offline sync
- Real-time data updates
- Flexible schema
- Serverless architecture
**Avoid when:**
- Complex SQL queries with joins
- Heavy analytics workloads
- Data >1 MiB per document
### Cloud SQL
**Best for:** Relational data with familiar SQL
| Engine | Version | Max Storage | Max Connections |
|--------|---------|-------------|-----------------|
| PostgreSQL | 15 | 64 TB | Instance-dependent |
| MySQL | 8.0 | 64 TB | Instance-dependent |
| SQL Server | 2022 | 64 TB | Instance-dependent |
```
Pricing:
- Machine type + storage + networking
- HA: 2x cost (regional instance)
- Read replicas: Per-replica pricing
```
**Use when:**
- Relational data with complex queries
- Existing SQL expertise
- Need ACID transactions
- Migration from on-premises databases
### Cloud Spanner
**Best for:** Globally distributed relational data
```
Limits:
- Storage: Unlimited
- Nodes: 1-100+ per instance
- Consistency: Strong global consistency
Pricing:
- Regional: $0.90/node-hour (~$657/month per node)
- Multi-region: $2.70/node-hour (~$1,971/month per node)
- Storage: $0.30/GiB/month
```
**Use when:**
- Global applications needing strong consistency
- Relational data at massive scale
- 99.999% availability requirement
- Horizontal scaling with SQL
### BigQuery
**Best for:** Analytics, data warehouse, SQL on massive datasets
```
Limits:
- Query: 6-hour timeout
- Concurrent queries: 100 default
- Streaming inserts: 100K rows/sec per table
Pricing:
- On-demand: $6.25 per TB queried (first 1 TB free/month)
- Editions: Autoscale slots starting at $0.04/slot-hour
- Storage: $0.02/GiB (active), $0.01/GiB (long-term)
```
### Firestore vs Cloud SQL vs Spanner
| Factor | Firestore | Cloud SQL | Cloud Spanner |
|--------|-----------|-----------|---------------|
| Query flexibility | Document-based | Full SQL | Full SQL |
| Scaling | Automatic | Vertical + read replicas | Horizontal |
| Consistency | Strong (single region) | ACID | Strong (global) |
| Cost model | Per-operation | Per-hour | Per-node-hour |
| Operational | Zero management | Managed (some ops) | Managed |
| Best for | Mobile/web apps | Traditional apps | Global scale |
---
## Storage Services
### Cloud Storage Classes
| Class | Access Pattern | Min Duration | Cost (GiB/mo) |
|-------|---------------|--------------|----------------|
| Standard | Frequent | None | $0.020 |
| Nearline | Monthly access | 30 days | $0.010 |
| Coldline | Quarterly access | 90 days | $0.004 |
| Archive | Annual access | 365 days | $0.0012 |
### Lifecycle Policy Example
```json
{
"lifecycle": {
"rule": [
{
"action": { "type": "SetStorageClass", "storageClass": "NEARLINE" },
"condition": { "age": 30, "matchesStorageClass": ["STANDARD"] }
},
{
"action": { "type": "SetStorageClass", "storageClass": "COLDLINE" },
"condition": { "age": 90, "matchesStorageClass": ["NEARLINE"] }
},
{
"action": { "type": "SetStorageClass", "storageClass": "ARCHIVE" },
"condition": { "age": 365, "matchesStorageClass": ["COLDLINE"] }
},
{
"action": { "type": "Delete" },
"condition": { "age": 2555 }
}
]
}
}
```
### Autoclass
Automatically transitions objects between storage classes based on access patterns. Recommended for mixed or unknown access patterns.
```bash
gsutil mb -l us-central1 --autoclass gs://my-bucket/
```
### Block and File Storage
| Service | Use Case | Access |
|---------|----------|--------|
| Persistent Disk | GCE/GKE block storage | Single instance (RW) or multi (RO) |
| Filestore | NFS shared file system | Multiple instances |
| Parallelstore | HPC parallel file system | High throughput |
| Cloud Storage FUSE | Mount GCS as filesystem | Any compute |
---
## Messaging and Events
### Decision Matrix
| Pattern | Service | Use Case |
|---------|---------|----------|
| Pub/sub messaging | Pub/Sub | Event streaming, microservice decoupling |
| Task queue | Cloud Tasks | Asynchronous task execution with retries |
| Workflow orchestration | Workflows | Multi-step service orchestration |
| Batch orchestration | Cloud Composer | Complex DAG-based pipelines (Airflow) |
| Event triggers | Eventarc | Route events to Cloud Run, GKE, Workflows |
### Pub/Sub
**Best for:** Event-driven architectures, stream processing
```
Limits:
- Message size: 10 MB max
- Throughput: Unlimited (auto-scaling)
- Retention: 7 days default (up to 31 days)
- Ordering: Per ordering key
Pricing: $40/TiB for message delivery
```
```python
# Pub/Sub publisher example
from google.cloud import pubsub_v1
import json
publisher = pubsub_v1.PublisherClient()
topic_path = publisher.topic_path('my-project', 'events')
def publish_event(event_type, payload):
data = json.dumps(payload).encode('utf-8')
future = publisher.publish(
topic_path,
data,
event_type=event_type
)
return future.result()
```
### Cloud Tasks
**Best for:** Asynchronous task execution with delivery guarantees
```
Features:
- Configurable retry policies
- Rate limiting
- Scheduled delivery
- HTTP and App Engine targets
Pricing: $0.40 per million operations
```
### Eventarc
**Best for:** Routing cloud events to services
```python
# Eventarc routes events from 130+ Google Cloud sources
# to Cloud Run, GKE, or Workflows
# Example: Trigger Cloud Run on Cloud Storage upload
# gcloud eventarc triggers create my-trigger \
# --destination-run-service=my-service \
# --event-filters="type=google.cloud.storage.object.v1.finalized" \
# --event-filters="bucket=my-bucket"
```
---
## API and Integration
### API Gateway vs Cloud Endpoints vs Cloud Run
| Factor | API Gateway | Cloud Endpoints | Cloud Run (direct) |
|--------|-------------|-----------------|---------------------|
| Protocol | REST, gRPC | REST, gRPC | Any HTTP |
| Auth | API keys, JWT, Firebase | API keys, JWT | IAM, custom |
| Rate limiting | Built-in | Built-in | Manual |
| Cost | Per-call pricing | Per-call pricing | Per-request |
| Best for | External APIs | Internal APIs | Simple services |
### Cloud Endpoints Configuration
```yaml
# openapi.yaml
swagger: "2.0"
info:
title: "My API"
version: "1.0.0"
host: "my-api-xyz.apigateway.my-project.cloud.goog"
schemes:
- "https"
paths:
/users:
get:
summary: "List users"
operationId: "listUsers"
x-google-backend:
address: "https://my-app-api-xyz.a.run.app"
security:
- api_key: []
securityDefinitions:
api_key:
type: "apiKey"
name: "key"
in: "query"
```
### Workflows
**Best for:** Orchestrating multi-service processes
```yaml
# workflow.yaml
main:
steps:
- processOrder:
call: http.post
args:
url: https://orders-service.run.app/process
body:
orderId: args.orderId
result: orderResult
- checkInventory:
switch:
- condition: orderResult.body.inStock
next: shipOrder
next: backOrder
- shipOrder:
call: http.post
args:
url: https://shipping-service.run.app/ship
body:
orderId: args.orderId
result: shipResult
- backOrder:
call: http.post
args:
url: https://inventory-service.run.app/backorder
body:
orderId: args.orderId
```
---
## Networking
### VPC Components
| Component | Purpose |
|-----------|---------|
| VPC | Isolated network (global resource) |
| Subnet | Regional network segment |
| Cloud NAT | Outbound internet for private instances |
| Cloud Router | Dynamic routing (BGP) |
| Private Google Access | Access GCP APIs without public IP |
| VPC Peering | Connect two VPC networks |
| Shared VPC | Share VPC across projects |
### VPC Design Pattern
```
VPC: 10.0.0.0/16 (global)
Subnet us-central1:
10.0.0.0/20 (primary)
10.4.0.0/14 (pods - secondary)
10.8.0.0/20 (services - secondary)
- GKE cluster, Cloud Run (VPC connector)
Subnet us-east1:
10.0.16.0/20 (primary)
- Cloud SQL (private IP), Memorystore
Subnet europe-west1:
10.0.32.0/20 (primary)
- DR / multi-region workloads
```
### Private Google Access
```bash
# Enable Private Google Access on a subnet
gcloud compute networks subnets update my-subnet \
--region=us-central1 \
--enable-private-google-access
```
---
## Security and Identity
### IAM Best Practices
```bash
# Prefer predefined roles over basic roles
# BAD: roles/editor (too broad)
# GOOD: roles/run.invoker (specific)
# Grant role to service account
gcloud projects add-iam-policy-binding my-project \
--member="serviceAccount:my-sa@my-project.iam.gserviceaccount.com" \
--role="roles/datastore.user" \
--condition='expression=resource.name.startsWith("projects/my-project/databases/(default)/documents/users"),title=firestore-users-only'
```
### Service Account Best Practices
| Practice | Description |
|----------|-------------|
| One SA per service | Separate service accounts per workload |
| Workload Identity | Bind K8s SAs to GCP SAs in GKE |
| Short-lived tokens | Use impersonation instead of key files |
| No SA keys | Avoid downloading JSON key files |
### Secret Manager vs Environment Variables
| Factor | Secret Manager | Env Variables |
|--------|---------------|---------------|
| Rotation | Automatic versioning | Manual redeploy |
| Audit | Cloud Audit Logs | No audit trail |
| Access control | IAM per-secret | Per-service |
| Pricing | $0.06/10K access ops | Free |
| Use case | Credentials, API keys | Non-sensitive config |
### Secret Manager Usage
```python
from google.cloud import secretmanager
def get_secret(project_id, secret_id, version="latest"):
client = secretmanager.SecretManagerServiceClient()
name = f"projects/{project_id}/secrets/{secret_id}/versions/{version}"
response = client.access_secret_version(request={"name": name})
return response.payload.data.decode("UTF-8")
# Usage
db_password = get_secret("my-project", "db-password")
```
FILE:scripts/architecture_designer.py
"""
GCP architecture design and service recommendation module.
Generates architecture patterns based on application requirements.
"""
import argparse
import json
import sys
from typing import Dict, List, Any
from enum import Enum
class ApplicationType(Enum):
"""Types of applications supported."""
WEB_APP = "web_application"
MOBILE_BACKEND = "mobile_backend"
DATA_PIPELINE = "data_pipeline"
MICROSERVICES = "microservices"
SAAS_PLATFORM = "saas_platform"
ML_PLATFORM = "ml_platform"
class ArchitectureDesigner:
"""Design GCP architectures based on requirements."""
def __init__(self, requirements: Dict[str, Any]):
"""
Initialize with application requirements.
Args:
requirements: Dictionary containing app type, traffic, budget, etc.
"""
self.app_type = requirements.get('application_type', 'web_application')
self.expected_users = requirements.get('expected_users', 1000)
self.requests_per_second = requirements.get('requests_per_second', 10)
self.budget_monthly = requirements.get('budget_monthly_usd', 500)
self.team_size = requirements.get('team_size', 3)
self.gcp_experience = requirements.get('gcp_experience', 'beginner')
self.compliance_needs = requirements.get('compliance', [])
self.data_size_gb = requirements.get('data_size_gb', 10)
def recommend_architecture_pattern(self) -> Dict[str, Any]:
"""
Recommend architecture pattern based on requirements.
Returns:
Dictionary with recommended pattern and services
"""
if self.app_type in ['web_application', 'saas_platform']:
if self.expected_users < 10000:
return self._serverless_web_architecture()
elif self.expected_users < 100000:
return self._gke_microservices_architecture()
else:
return self._multi_region_architecture()
elif self.app_type == 'mobile_backend':
return self._serverless_mobile_backend()
elif self.app_type == 'data_pipeline':
return self._data_pipeline_architecture()
elif self.app_type == 'microservices':
return self._gke_microservices_architecture()
elif self.app_type == 'ml_platform':
return self._ml_platform_architecture()
else:
return self._serverless_web_architecture()
def _serverless_web_architecture(self) -> Dict[str, Any]:
"""Serverless web application pattern using Cloud Run."""
return {
'pattern_name': 'Serverless Web Application',
'description': 'Fully serverless architecture with Cloud Run and Firestore',
'use_case': 'SaaS platforms, low to medium traffic websites, MVPs',
'services': {
'frontend': {
'service': 'Cloud Storage + Cloud CDN',
'purpose': 'Static website hosting with global CDN',
'configuration': {
'bucket': 'Website bucket with public access',
'cdn': 'Cloud CDN with custom domain and HTTPS',
'caching': 'Cache-Control headers, edge caching'
}
},
'api': {
'service': 'Cloud Run',
'purpose': 'Containerized API backend with auto-scaling',
'configuration': {
'cpu': '1 vCPU',
'memory': '512 Mi',
'min_instances': '0 (scale to zero)',
'max_instances': '10',
'concurrency': '80 requests per instance',
'timeout': '300 seconds'
}
},
'database': {
'service': 'Firestore',
'purpose': 'NoSQL document database with real-time sync',
'configuration': {
'mode': 'Native mode',
'location': 'Regional or multi-region',
'security_rules': 'Firestore security rules',
'backup': 'Scheduled exports to Cloud Storage'
}
},
'authentication': {
'service': 'Identity Platform',
'purpose': 'User authentication and authorization',
'configuration': {
'providers': 'Email/password, Google, Apple, OIDC',
'mfa': 'SMS or TOTP multi-factor authentication',
'token_expiration': '1 hour access, 30 days refresh'
}
},
'cicd': {
'service': 'Cloud Build',
'purpose': 'Automated build and deployment from Git',
'configuration': {
'source': 'GitHub or Cloud Source Repositories',
'build': 'Automatic on commit',
'environments': 'dev, staging, production'
}
}
},
'estimated_cost': {
'monthly_usd': self._calculate_serverless_cost(),
'breakdown': {
'Cloud CDN': '5-20 USD',
'Cloud Run': '5-25 USD',
'Firestore': '5-30 USD',
'Identity Platform': '0-10 USD (free tier: 50k MAU)',
'Cloud Storage': '1-5 USD'
}
},
'pros': [
'No server management',
'Auto-scaling with scale-to-zero',
'Pay only for what you use',
'No cold starts with min instances',
'Container-based (no runtime restrictions)'
],
'cons': [
'Vendor lock-in to GCP',
'Regional availability considerations',
'Debugging distributed systems complex',
'Firestore query limitations vs SQL'
],
'scaling_characteristics': {
'users_supported': '1k - 100k',
'requests_per_second': '100 - 10,000',
'scaling_method': 'Automatic (Cloud Run auto-scaling)'
}
}
def _gke_microservices_architecture(self) -> Dict[str, Any]:
"""GKE-based microservices architecture."""
return {
'pattern_name': 'Microservices on GKE',
'description': 'Kubernetes-native architecture with managed services',
'use_case': 'SaaS platforms, complex microservices, enterprise applications',
'services': {
'load_balancer': {
'service': 'Cloud Load Balancing',
'purpose': 'Global HTTP(S) load balancing',
'configuration': {
'type': 'External Application Load Balancer',
'ssl': 'Google-managed SSL certificate',
'health_checks': '/health endpoint, 10s interval',
'cdn': 'Cloud CDN enabled for static content'
}
},
'compute': {
'service': 'GKE Autopilot',
'purpose': 'Managed Kubernetes for containerized workloads',
'configuration': {
'mode': 'Autopilot (fully managed node provisioning)',
'scaling': 'Horizontal Pod Autoscaler',
'networking': 'VPC-native with Alias IPs',
'workload_identity': 'Enabled for secure service account binding'
}
},
'database': {
'service': 'Cloud SQL (PostgreSQL)',
'purpose': 'Managed relational database',
'configuration': {
'tier': 'db-custom-2-8192 (2 vCPU, 8 GB RAM)',
'high_availability': 'Regional with automatic failover',
'read_replicas': '1-2 for read scaling',
'backup': 'Automated daily backups, 7-day retention',
'encryption': 'Customer-managed encryption key (CMEK)'
}
},
'cache': {
'service': 'Memorystore (Redis)',
'purpose': 'Session storage, application caching',
'configuration': {
'tier': 'Basic (1 GB) or Standard (HA)',
'version': 'Redis 7.0',
'eviction_policy': 'allkeys-lru'
}
},
'messaging': {
'service': 'Pub/Sub',
'purpose': 'Asynchronous messaging between services',
'configuration': {
'topics': 'Per-domain event topics',
'subscriptions': 'Pull or push delivery',
'dead_letter': 'Dead letter topic after 5 retries',
'ordering': 'Ordering keys for ordered delivery'
}
},
'storage': {
'service': 'Cloud Storage',
'purpose': 'User uploads, backups, logs',
'configuration': {
'storage_class': 'Standard with lifecycle policies',
'versioning': 'Enabled for important buckets',
'lifecycle': 'Transition to Nearline after 30 days'
}
}
},
'estimated_cost': {
'monthly_usd': self._calculate_gke_cost(),
'breakdown': {
'Cloud Load Balancing': '20-40 USD',
'GKE Autopilot': '75-250 USD',
'Cloud SQL': '80-250 USD',
'Memorystore': '30-80 USD',
'Pub/Sub': '5-20 USD',
'Cloud Storage': '5-20 USD'
}
},
'pros': [
'Kubernetes ecosystem compatibility',
'Fine-grained scaling control',
'Multi-cloud portability',
'Rich service mesh (Anthos Service Mesh)',
'Managed node provisioning with Autopilot'
],
'cons': [
'Higher baseline costs than serverless',
'Kubernetes learning curve',
'More operational complexity',
'GKE management fee ($74.40/month per cluster)'
],
'scaling_characteristics': {
'users_supported': '10k - 500k',
'requests_per_second': '1,000 - 50,000',
'scaling_method': 'HPA + Cluster Autoscaler'
}
}
def _serverless_mobile_backend(self) -> Dict[str, Any]:
"""Serverless mobile backend with Firebase."""
return {
'pattern_name': 'Serverless Mobile Backend',
'description': 'Mobile-first backend with Firebase and Cloud Functions',
'use_case': 'Mobile apps, real-time applications, offline-first apps',
'services': {
'api': {
'service': 'Cloud Functions (2nd gen)',
'purpose': 'Event-driven API handlers',
'configuration': {
'runtime': 'Node.js 20 or Python 3.12',
'memory': '256 MB - 1 GB',
'timeout': '60 seconds',
'concurrency': 'Up to 1000 concurrent'
}
},
'database': {
'service': 'Firestore',
'purpose': 'Real-time NoSQL database with offline sync',
'configuration': {
'mode': 'Native mode',
'multi_region': 'nam5 or eur3 for HA',
'security_rules': 'Client-side access control',
'indexes': 'Composite indexes for queries'
}
},
'file_storage': {
'service': 'Cloud Storage (Firebase)',
'purpose': 'User uploads (images, videos, documents)',
'configuration': {
'access': 'Firebase Security Rules',
'resumable_uploads': 'Enabled for large files',
'cdn': 'Automatic via Firebase Hosting CDN'
}
},
'authentication': {
'service': 'Firebase Authentication',
'purpose': 'User management and federation',
'configuration': {
'providers': 'Email, Google, Apple, Phone',
'anonymous_auth': 'Enabled for guest access',
'custom_claims': 'Role-based access control',
'multi_tenancy': 'Supported via Identity Platform'
}
},
'push_notifications': {
'service': 'Firebase Cloud Messaging (FCM)',
'purpose': 'Push notifications to mobile devices',
'configuration': {
'platforms': 'iOS (APNs), Android, Web',
'topics': 'Topic-based group messaging',
'analytics': 'Notification delivery tracking'
}
},
'analytics': {
'service': 'Google Analytics (Firebase)',
'purpose': 'User analytics and event tracking',
'configuration': {
'events': 'Custom and automatic events',
'audiences': 'User segmentation',
'bigquery_export': 'Raw event export to BigQuery'
}
}
},
'estimated_cost': {
'monthly_usd': 40 + (self.expected_users * 0.004),
'breakdown': {
'Cloud Functions': '5-30 USD',
'Firestore': '10-50 USD',
'Cloud Storage': '5-20 USD',
'Identity Platform': '0-15 USD',
'FCM': '0 USD (free)',
'Analytics': '0 USD (free)'
}
},
'pros': [
'Real-time data sync built-in',
'Offline-first support',
'Firebase SDKs for iOS/Android/Web',
'Free tier covers most MVPs',
'Rapid development with Firebase console'
],
'cons': [
'Firestore query limitations',
'Vendor lock-in to Firebase/GCP',
'Cost scaling can be unpredictable',
'Limited server-side control'
],
'scaling_characteristics': {
'users_supported': '1k - 1M',
'requests_per_second': '100 - 100,000',
'scaling_method': 'Automatic (Firebase managed)'
}
}
def _data_pipeline_architecture(self) -> Dict[str, Any]:
"""Serverless data pipeline with BigQuery."""
return {
'pattern_name': 'Serverless Data Pipeline',
'description': 'Scalable data ingestion, processing, and analytics',
'use_case': 'Analytics, IoT data, log processing, ETL, data warehousing',
'services': {
'ingestion': {
'service': 'Pub/Sub',
'purpose': 'Real-time event and data ingestion',
'configuration': {
'throughput': 'Unlimited (auto-scaling)',
'retention': '7 days (configurable to 31 days)',
'ordering': 'Ordering keys for ordered delivery',
'dead_letter': 'Dead letter topic for failed messages'
}
},
'processing': {
'service': 'Dataflow (Apache Beam)',
'purpose': 'Stream and batch data processing',
'configuration': {
'mode': 'Streaming or batch',
'autoscaling': 'Horizontal autoscaling',
'workers': f'{max(1, self.data_size_gb // 20)} initial workers',
'sdk': 'Python or Java Apache Beam SDK'
}
},
'warehouse': {
'service': 'BigQuery',
'purpose': 'Serverless data warehouse and analytics',
'configuration': {
'pricing': 'On-demand ($6.25/TB queried) or slots',
'partitioning': 'By ingestion time or custom field',
'clustering': 'Up to 4 clustering columns',
'streaming_insert': 'Real-time data availability'
}
},
'storage': {
'service': 'Cloud Storage (Data Lake)',
'purpose': 'Raw data lake and archival storage',
'configuration': {
'format': 'Parquet or Avro (columnar)',
'partitioning': 'By date (year/month/day)',
'lifecycle': 'Transition to Coldline after 90 days',
'catalog': 'Dataplex for data governance'
}
},
'visualization': {
'service': 'Looker / Looker Studio',
'purpose': 'Business intelligence dashboards',
'configuration': {
'source': 'BigQuery direct connection',
'refresh': 'Real-time or scheduled',
'sharing': 'Embedded or web dashboards'
}
},
'orchestration': {
'service': 'Cloud Composer (Airflow)',
'purpose': 'Workflow orchestration for batch pipelines',
'configuration': {
'environment': 'Cloud Composer 2 (auto-scaling)',
'dags': 'Python DAG definitions',
'scheduling': 'Cron-based scheduling'
}
}
},
'estimated_cost': {
'monthly_usd': self._calculate_data_pipeline_cost(),
'breakdown': {
'Pub/Sub': '5-30 USD',
'Dataflow': '20-150 USD',
'BigQuery': '10-100 USD (on-demand)',
'Cloud Storage': '5-30 USD',
'Looker Studio': '0 USD (free)',
'Cloud Composer': '300+ USD (if used)'
}
},
'pros': [
'Fully serverless data stack',
'BigQuery scales to petabytes',
'Real-time and batch in same pipeline',
'Cost-effective with on-demand pricing',
'ML integration via BigQuery ML'
],
'cons': [
'Dataflow has steep learning curve (Beam SDK)',
'BigQuery costs based on data scanned',
'Cloud Composer expensive for small workloads',
'Schema evolution requires planning'
],
'scaling_characteristics': {
'events_per_second': '1,000 - 10,000,000',
'data_volume': '1 GB - 1 PB per day',
'scaling_method': 'Automatic (all services auto-scale)'
}
}
def _ml_platform_architecture(self) -> Dict[str, Any]:
"""ML platform architecture with Vertex AI."""
return {
'pattern_name': 'ML Platform',
'description': 'End-to-end machine learning platform',
'use_case': 'Model training, serving, MLOps, feature engineering',
'services': {
'ml_platform': {
'service': 'Vertex AI',
'purpose': 'Training, tuning, and serving ML models',
'configuration': {
'training': 'Custom or AutoML training jobs',
'prediction': 'Online or batch prediction endpoints',
'pipelines': 'Vertex AI Pipelines for MLOps',
'feature_store': 'Vertex AI Feature Store'
}
},
'data': {
'service': 'BigQuery',
'purpose': 'Feature engineering and data exploration',
'configuration': {
'ml': 'BigQuery ML for in-warehouse models',
'export': 'Export to Cloud Storage for training',
'feature_engineering': 'SQL-based transformations'
}
},
'storage': {
'service': 'Cloud Storage',
'purpose': 'Datasets, model artifacts, experiment logs',
'configuration': {
'buckets': 'Separate buckets for data/models/logs',
'versioning': 'Enabled for model artifacts',
'lifecycle': 'Archive old experiment data'
}
},
'triggers': {
'service': 'Cloud Functions',
'purpose': 'Event-driven preprocessing and triggers',
'configuration': {
'triggers': 'Cloud Storage, Pub/Sub, Scheduler',
'preprocessing': 'Data validation and transforms',
'notifications': 'Training completion alerts'
}
},
'monitoring': {
'service': 'Vertex AI Model Monitoring',
'purpose': 'Detect data drift and model degradation',
'configuration': {
'skew_detection': 'Training-serving skew alerts',
'drift_detection': 'Feature drift monitoring',
'alerting': 'Cloud Monitoring integration'
}
}
},
'estimated_cost': {
'monthly_usd': 200 + (self.data_size_gb * 2),
'breakdown': {
'Vertex AI Training': '50-500 USD (GPU dependent)',
'Vertex AI Prediction': '30-200 USD',
'BigQuery': '20-100 USD',
'Cloud Storage': '10-50 USD',
'Cloud Functions': '5-20 USD'
}
},
'pros': [
'End-to-end ML lifecycle management',
'AutoML for rapid prototyping',
'Integrated with BigQuery and Cloud Storage',
'Managed model serving with autoscaling',
'Built-in experiment tracking'
],
'cons': [
'GPU costs can escalate quickly',
'Vertex AI pricing is complex',
'Limited customization vs self-managed',
'Vendor lock-in for model artifacts'
],
'scaling_characteristics': {
'training': 'Multi-GPU, distributed training',
'prediction': '1 - 1000+ replicas',
'scaling_method': 'Automatic endpoint scaling'
}
}
def _multi_region_architecture(self) -> Dict[str, Any]:
"""Multi-region high availability architecture."""
return {
'pattern_name': 'Multi-Region High Availability',
'description': 'Global deployment with disaster recovery',
'use_case': 'Global applications, 99.99% uptime, compliance',
'services': {
'dns': {
'service': 'Cloud DNS',
'purpose': 'Global DNS with health-checked routing',
'configuration': {
'routing_policy': 'Geolocation or weighted routing',
'health_checks': 'HTTP health checks per region',
'failover': 'Automatic DNS failover'
}
},
'cdn': {
'service': 'Cloud CDN',
'purpose': 'Edge caching and acceleration',
'configuration': {
'origins': 'Multiple regional backends',
'cache_modes': 'CACHE_ALL_STATIC or USE_ORIGIN_HEADERS',
'edge_locations': 'Global (100+ locations)'
}
},
'compute': {
'service': 'Multi-region GKE or Cloud Run',
'purpose': 'Active-active deployment across regions',
'configuration': {
'regions': 'us-central1 (primary), europe-west1 (secondary)',
'deployment': 'Cloud Deploy for multi-region rollout',
'traffic_split': 'Global Load Balancer with traffic management'
}
},
'database': {
'service': 'Cloud Spanner or Firestore multi-region',
'purpose': 'Globally consistent database',
'configuration': {
'spanner': 'Multi-region config (nam-eur-asia1)',
'firestore': 'Multi-region location (nam5, eur3)',
'consistency': 'Strong consistency (Spanner) or eventual (Firestore)',
'replication': 'Automatic cross-region replication'
}
},
'storage': {
'service': 'Cloud Storage (dual-region or multi-region)',
'purpose': 'Geo-redundant object storage',
'configuration': {
'location': 'Dual-region (us-central1+us-east1) or multi-region (US)',
'turbo_replication': '15-minute RPO with turbo replication',
'versioning': 'Enabled for critical data'
}
}
},
'estimated_cost': {
'monthly_usd': self._calculate_gke_cost() * 2.0,
'breakdown': {
'Cloud DNS': '5-15 USD',
'Cloud CDN': '20-100 USD',
'Compute (2 regions)': '150-500 USD',
'Cloud Spanner': '500-2000 USD (multi-region)',
'Data transfer (cross-region)': '50-200 USD'
}
},
'pros': [
'Global low latency',
'High availability (99.99%+)',
'Disaster recovery built-in',
'Data sovereignty compliance',
'Automatic failover'
],
'cons': [
'2x+ costs vs single region',
'Cloud Spanner is expensive',
'Complex deployment pipeline',
'Cross-region data transfer costs',
'Operational overhead'
],
'scaling_characteristics': {
'users_supported': '100k - 100M',
'requests_per_second': '10,000 - 10,000,000',
'scaling_method': 'Per-region auto-scaling + global load balancing'
}
}
def _calculate_serverless_cost(self) -> float:
"""Estimate serverless architecture cost."""
requests_per_month = self.requests_per_second * 2_592_000
cloud_run_cost = max(5, (requests_per_month / 1_000_000) * 0.40)
firestore_cost = max(5, self.data_size_gb * 0.18)
cdn_cost = max(5, self.expected_users * 0.008)
storage_cost = max(1, self.data_size_gb * 0.02)
total = cloud_run_cost + firestore_cost + cdn_cost + storage_cost
return min(total, self.budget_monthly)
def _calculate_gke_cost(self) -> float:
"""Estimate GKE microservices architecture cost."""
gke_management = 74.40 # Autopilot cluster fee
pod_cost = max(2, self.expected_users // 5000) * 35
cloud_sql_cost = 120 # db-custom-2-8192 baseline
memorystore_cost = 35 # Basic 1 GB
lb_cost = 25
total = gke_management + pod_cost + cloud_sql_cost + memorystore_cost + lb_cost
return min(total, self.budget_monthly)
def _calculate_data_pipeline_cost(self) -> float:
"""Estimate data pipeline cost."""
pubsub_cost = max(5, self.data_size_gb * 0.5)
dataflow_cost = max(20, self.data_size_gb * 1.5)
bigquery_cost = max(10, self.data_size_gb * 0.02 * 6.25)
storage_cost = self.data_size_gb * 0.02
total = pubsub_cost + dataflow_cost + bigquery_cost + storage_cost
return min(total, self.budget_monthly)
def generate_service_checklist(self) -> list:
"""Generate implementation checklist for recommended architecture."""
architecture = self.recommend_architecture_pattern()
checklist = [
{
'phase': 'Planning',
'tasks': [
'Review architecture pattern and services',
'Estimate costs using GCP Pricing Calculator',
'Define environment strategy (dev, staging, prod)',
'Set up GCP Organization and projects',
'Define labeling strategy for resources'
]
},
{
'phase': 'Foundation',
'tasks': [
'Create VPC with subnets (if using GKE/Compute)',
'Configure Cloud NAT for private resources',
'Set up IAM roles and service accounts',
'Enable Cloud Audit Logs',
'Configure Organization policies'
]
},
{
'phase': 'Core Services',
'tasks': [
f"Deploy {service['service']}"
for service in architecture['services'].values()
]
},
{
'phase': 'Security',
'tasks': [
'Configure firewall rules and VPC Service Controls',
'Enable encryption (Cloud KMS) for all services',
'Set up Cloud Armor WAF rules',
'Configure Secret Manager for credentials',
'Enable Security Command Center'
]
},
{
'phase': 'Monitoring',
'tasks': [
'Create Cloud Monitoring dashboards',
'Set up alerting policies for critical metrics',
'Configure notification channels (email, Slack, PagerDuty)',
'Enable Cloud Trace for distributed tracing',
'Set up log-based metrics and log sinks'
]
},
{
'phase': 'CI/CD',
'tasks': [
'Set up Cloud Build triggers',
'Configure automated testing',
'Implement canary or rolling deployments',
'Set up rollback procedures',
'Document deployment process'
]
}
]
return checklist
def main():
parser = argparse.ArgumentParser(
description='GCP Architecture Designer - Recommends GCP services based on workload requirements'
)
parser.add_argument(
'--input', '-i',
type=str,
help='Path to JSON file with application requirements'
)
parser.add_argument(
'--output', '-o',
type=str,
help='Path to write design output JSON'
)
parser.add_argument(
'--json',
action='store_true',
help='Output as JSON format'
)
parser.add_argument(
'--app-type',
type=str,
choices=['web_application', 'mobile_backend', 'data_pipeline',
'microservices', 'saas_platform', 'ml_platform'],
default='web_application',
help='Application type (default: web_application)'
)
parser.add_argument(
'--users',
type=int,
default=1000,
help='Expected number of users (default: 1000)'
)
parser.add_argument(
'--budget',
type=float,
default=500,
help='Monthly budget in USD (default: 500)'
)
args = parser.parse_args()
if args.input:
try:
with open(args.input, 'r') as f:
requirements = json.load(f)
except FileNotFoundError:
print(f"Error: File '{args.input}' not found.", file=sys.stderr)
sys.exit(1)
except json.JSONDecodeError:
print(f"Error: File '{args.input}' is not valid JSON.", file=sys.stderr)
sys.exit(1)
else:
requirements = {
'application_type': args.app_type,
'expected_users': args.users,
'budget_monthly_usd': args.budget
}
designer = ArchitectureDesigner(requirements)
result = designer.recommend_architecture_pattern()
checklist = designer.generate_service_checklist()
output = {
'architecture': result,
'implementation_checklist': checklist
}
if args.output:
with open(args.output, 'w') as f:
json.dump(output, f, indent=2)
print(f"Design written to {args.output}")
elif args.json:
print(json.dumps(output, indent=2))
else:
print(f"\nRecommended Pattern: {result['pattern_name']}")
print(f"Description: {result['description']}")
print(f"Use Case: {result['use_case']}")
print(f"\nServices:")
for name, svc in result['services'].items():
print(f" - {name}: {svc['service']} ({svc['purpose']})")
print(f"\nEstimated Monthly Cost: .2f")
print(f"\nPros: {', '.join(result['pros'])}")
print(f"Cons: {', '.join(result['cons'])}")
if __name__ == '__main__':
main()
FILE:scripts/cost_optimizer.py
"""
GCP cost optimization analyzer.
Provides cost-saving recommendations for GCP resources.
"""
import argparse
import json
import sys
from typing import Dict, List, Any
class CostOptimizer:
"""Analyze GCP costs and provide optimization recommendations."""
def __init__(self, current_resources: Dict[str, Any], monthly_spend: float):
"""
Initialize with current GCP resources and spending.
Args:
current_resources: Dictionary of current GCP resources
monthly_spend: Current monthly GCP spend in USD
"""
self.resources = current_resources
self.monthly_spend = monthly_spend
self.recommendations = []
def analyze_and_optimize(self) -> Dict[str, Any]:
"""
Analyze current setup and generate cost optimization recommendations.
Returns:
Dictionary with recommendations and potential savings
"""
self.recommendations = []
potential_savings = 0.0
compute_savings = self._analyze_compute()
potential_savings += compute_savings
storage_savings = self._analyze_storage()
potential_savings += storage_savings
database_savings = self._analyze_database()
potential_savings += database_savings
network_savings = self._analyze_networking()
potential_savings += network_savings
general_savings = self._analyze_general_optimizations()
potential_savings += general_savings
return {
'current_monthly_spend': self.monthly_spend,
'potential_monthly_savings': round(potential_savings, 2),
'optimized_monthly_spend': round(self.monthly_spend - potential_savings, 2),
'savings_percentage': round((potential_savings / self.monthly_spend) * 100, 2) if self.monthly_spend > 0 else 0,
'recommendations': self.recommendations,
'priority_actions': self._prioritize_recommendations()
}
def _analyze_compute(self) -> float:
"""Analyze compute resources (GCE, GKE, Cloud Run)."""
savings = 0.0
gce_instances = self.resources.get('gce_instances', [])
if gce_instances:
idle_count = sum(1 for inst in gce_instances if inst.get('cpu_utilization', 100) < 10)
if idle_count > 0:
idle_cost = idle_count * 50
savings += idle_cost
self.recommendations.append({
'service': 'Compute Engine',
'type': 'Idle Resources',
'issue': f'{idle_count} GCE instances with <10% CPU utilization',
'recommendation': 'Stop or delete idle instances, or downsize to smaller machine types',
'potential_savings': idle_cost,
'priority': 'high'
})
# Check for committed use discounts
on_demand_count = sum(1 for inst in gce_instances if inst.get('pricing', 'on-demand') == 'on-demand')
if on_demand_count >= 2:
cud_savings = on_demand_count * 50 * 0.37 # 37% savings with 1-yr CUD
savings += cud_savings
self.recommendations.append({
'service': 'Compute Engine',
'type': 'Committed Use Discounts',
'issue': f'{on_demand_count} instances on on-demand pricing',
'recommendation': 'Purchase 1-year committed use discounts for predictable workloads (37% savings) or 3-year (55% savings)',
'potential_savings': cud_savings,
'priority': 'medium'
})
# Check for sustained use discounts awareness
short_lived = sum(1 for inst in gce_instances if inst.get('uptime_hours_month', 730) < 200)
if short_lived > 0:
self.recommendations.append({
'service': 'Compute Engine',
'type': 'Scheduling',
'issue': f'{short_lived} instances running <200 hours/month',
'recommendation': 'Use Instance Scheduler to stop dev/test instances outside business hours',
'potential_savings': short_lived * 20,
'priority': 'medium'
})
savings += short_lived * 20
# GKE optimization
gke_clusters = self.resources.get('gke_clusters', [])
for cluster in gke_clusters:
if cluster.get('mode', 'standard') == 'standard':
node_utilization = cluster.get('avg_node_utilization', 100)
if node_utilization < 40:
autopilot_savings = cluster.get('monthly_cost', 500) * 0.30
savings += autopilot_savings
self.recommendations.append({
'service': 'GKE',
'type': 'Cluster Mode',
'issue': f'Standard GKE cluster with <40% node utilization',
'recommendation': 'Migrate to GKE Autopilot to pay only for pod resources, or enable cluster autoscaler',
'potential_savings': autopilot_savings,
'priority': 'high'
})
# Cloud Run optimization
cloud_run_services = self.resources.get('cloud_run_services', [])
for svc in cloud_run_services:
if svc.get('min_instances', 0) > 0 and svc.get('avg_rps', 100) < 1:
min_inst_savings = svc.get('min_instances', 1) * 15
savings += min_inst_savings
self.recommendations.append({
'service': 'Cloud Run',
'type': 'Min Instances',
'issue': f'Service {svc.get("name", "unknown")} has min instances but very low traffic',
'recommendation': 'Set min-instances to 0 for low-traffic services to enable scale-to-zero',
'potential_savings': min_inst_savings,
'priority': 'medium'
})
return savings
def _analyze_storage(self) -> float:
"""Analyze Cloud Storage resources."""
savings = 0.0
gcs_buckets = self.resources.get('gcs_buckets', [])
for bucket in gcs_buckets:
size_gb = bucket.get('size_gb', 0)
storage_class = bucket.get('storage_class', 'STANDARD')
if not bucket.get('has_lifecycle_policy', False) and size_gb > 100:
lifecycle_savings = size_gb * 0.012
savings += lifecycle_savings
self.recommendations.append({
'service': 'Cloud Storage',
'type': 'Lifecycle Policy',
'issue': f'Bucket {bucket.get("name", "unknown")} ({size_gb} GB) has no lifecycle policy',
'recommendation': 'Add lifecycle rule: Transition to Nearline after 30 days, Coldline after 90 days, Archive after 365 days',
'potential_savings': lifecycle_savings,
'priority': 'medium'
})
if storage_class == 'STANDARD' and size_gb > 500:
class_savings = size_gb * 0.006
savings += class_savings
self.recommendations.append({
'service': 'Cloud Storage',
'type': 'Storage Class',
'issue': f'Large bucket ({size_gb} GB) using Standard class',
'recommendation': 'Enable Autoclass for automatic storage class management based on access patterns',
'potential_savings': class_savings,
'priority': 'high'
})
return savings
def _analyze_database(self) -> float:
"""Analyze Cloud SQL, Firestore, and BigQuery costs."""
savings = 0.0
cloud_sql_instances = self.resources.get('cloud_sql_instances', [])
for db in cloud_sql_instances:
if db.get('connections_per_day', 1000) < 10:
db_cost = db.get('monthly_cost', 100)
savings += db_cost * 0.8
self.recommendations.append({
'service': 'Cloud SQL',
'type': 'Idle Resource',
'issue': f'Database {db.get("name", "unknown")} has <10 connections/day',
'recommendation': 'Stop database if not needed, or take a backup and delete',
'potential_savings': db_cost * 0.8,
'priority': 'high'
})
if db.get('utilization', 100) < 30 and not db.get('has_ha', False):
rightsize_savings = db.get('monthly_cost', 200) * 0.35
savings += rightsize_savings
self.recommendations.append({
'service': 'Cloud SQL',
'type': 'Right-sizing',
'issue': f'Cloud SQL instance {db.get("name", "unknown")} has low utilization (<30%)',
'recommendation': 'Downsize to a smaller machine type (e.g., db-custom-2-8192 to db-f1-micro for dev)',
'potential_savings': rightsize_savings,
'priority': 'medium'
})
# BigQuery optimization
bigquery_datasets = self.resources.get('bigquery_datasets', [])
for dataset in bigquery_datasets:
if dataset.get('pricing_model', 'on_demand') == 'on_demand':
monthly_tb_scanned = dataset.get('monthly_tb_scanned', 0)
if monthly_tb_scanned > 10:
slot_savings = (monthly_tb_scanned * 6.25) * 0.30
savings += slot_savings
self.recommendations.append({
'service': 'BigQuery',
'type': 'Pricing Model',
'issue': f'Scanning {monthly_tb_scanned} TB/month on on-demand pricing',
'recommendation': 'Switch to BigQuery editions with slots for predictable costs (30%+ savings at this volume)',
'potential_savings': slot_savings,
'priority': 'high'
})
if not dataset.get('has_partitioning', False):
partition_savings = dataset.get('monthly_query_cost', 50) * 0.50
savings += partition_savings
self.recommendations.append({
'service': 'BigQuery',
'type': 'Table Partitioning',
'issue': f'Tables in {dataset.get("name", "unknown")} lack partitioning',
'recommendation': 'Partition tables by date and add clustering columns to reduce bytes scanned',
'potential_savings': partition_savings,
'priority': 'medium'
})
return savings
def _analyze_networking(self) -> float:
"""Analyze networking costs (egress, Cloud NAT, etc.)."""
savings = 0.0
cloud_nat_gateways = self.resources.get('cloud_nat_gateways', [])
if len(cloud_nat_gateways) > 1:
extra_nats = len(cloud_nat_gateways) - 1
nat_savings = extra_nats * 45
savings += nat_savings
self.recommendations.append({
'service': 'Cloud NAT',
'type': 'Resource Consolidation',
'issue': f'{len(cloud_nat_gateways)} Cloud NAT gateways deployed',
'recommendation': 'Consolidate NAT gateways in dev/staging, or use Private Google Access for GCP services',
'potential_savings': nat_savings,
'priority': 'high'
})
egress_gb = self.resources.get('monthly_egress_gb', 0)
if egress_gb > 1000:
cdn_savings = egress_gb * 0.04 # CDN is cheaper than direct egress
savings += cdn_savings
self.recommendations.append({
'service': 'Networking',
'type': 'CDN Optimization',
'issue': f'High egress volume ({egress_gb} GB/month)',
'recommendation': 'Enable Cloud CDN to serve cached content at lower egress rates',
'potential_savings': cdn_savings,
'priority': 'medium'
})
return savings
def _analyze_general_optimizations(self) -> float:
"""General GCP cost optimizations."""
savings = 0.0
# Log retention
log_sinks = self.resources.get('log_sinks', [])
if not log_sinks:
log_volume_gb = self.resources.get('monthly_log_volume_gb', 0)
if log_volume_gb > 50:
log_savings = log_volume_gb * 0.50 * 0.6
savings += log_savings
self.recommendations.append({
'service': 'Cloud Logging',
'type': 'Log Exclusion',
'issue': f'{log_volume_gb} GB/month of logs without exclusion filters',
'recommendation': 'Create log exclusion filters for verbose/debug logs and route remaining to Cloud Storage via log sinks',
'potential_savings': log_savings,
'priority': 'medium'
})
# Unattached persistent disks
persistent_disks = self.resources.get('persistent_disks', [])
unattached = sum(1 for disk in persistent_disks if not disk.get('attached', True))
if unattached > 0:
disk_savings = unattached * 10 # ~$10/month per 100 GB disk
savings += disk_savings
self.recommendations.append({
'service': 'Compute Engine',
'type': 'Unused Resources',
'issue': f'{unattached} unattached persistent disks',
'recommendation': 'Snapshot and delete unused persistent disks',
'potential_savings': disk_savings,
'priority': 'high'
})
# Static external IPs
static_ips = self.resources.get('static_ips', [])
unused_ips = sum(1 for ip in static_ips if not ip.get('in_use', True))
if unused_ips > 0:
ip_savings = unused_ips * 7.30 # $0.01/hour = $7.30/month
savings += ip_savings
self.recommendations.append({
'service': 'Networking',
'type': 'Unused Resources',
'issue': f'{unused_ips} unused static external IP addresses',
'recommendation': 'Release unused static IPs to avoid hourly charges',
'potential_savings': ip_savings,
'priority': 'high'
})
# Budget alerts
if not self.resources.get('has_budget_alerts', False):
self.recommendations.append({
'service': 'Cloud Billing',
'type': 'Cost Monitoring',
'issue': 'No budget alerts configured',
'recommendation': 'Set up Cloud Billing budgets with alerts at 50%, 80%, 100% of monthly budget',
'potential_savings': 0,
'priority': 'high'
})
# Recommender API
if not self.resources.get('uses_recommender', False):
self.recommendations.append({
'service': 'Active Assist',
'type': 'Visibility',
'issue': 'GCP Recommender not reviewed',
'recommendation': 'Review Active Assist recommendations for right-sizing, idle resources, and committed use discounts',
'potential_savings': 0,
'priority': 'medium'
})
return savings
def _prioritize_recommendations(self) -> List[Dict[str, Any]]:
"""Get top priority recommendations."""
high_priority = [r for r in self.recommendations if r['priority'] == 'high']
high_priority.sort(key=lambda x: x.get('potential_savings', 0), reverse=True)
return high_priority[:5]
def generate_optimization_checklist(self) -> List[Dict[str, Any]]:
"""Generate actionable checklist for cost optimization."""
return [
{
'category': 'Immediate Actions (Today)',
'items': [
'Release unused static IPs',
'Delete unattached persistent disks',
'Stop idle Compute Engine instances',
'Set up billing budget alerts'
]
},
{
'category': 'This Week',
'items': [
'Add Cloud Storage lifecycle policies',
'Create log exclusion filters for verbose logs',
'Right-size Cloud SQL instances',
'Review Active Assist recommendations'
]
},
{
'category': 'This Month',
'items': [
'Evaluate committed use discounts',
'Migrate GKE Standard to Autopilot where applicable',
'Partition and cluster BigQuery tables',
'Enable Cloud CDN for high-egress services'
]
},
{
'category': 'Ongoing',
'items': [
'Review billing reports weekly',
'Label all resources for cost allocation',
'Monitor Active Assist recommendations monthly',
'Conduct quarterly cost optimization reviews'
]
}
]
def main():
parser = argparse.ArgumentParser(
description='GCP Cost Optimizer - Analyzes GCP resources and recommends cost savings'
)
parser.add_argument(
'--resources', '-r',
type=str,
help='Path to JSON file with current GCP resource inventory'
)
parser.add_argument(
'--monthly-spend', '-s',
type=float,
default=1000,
help='Current monthly GCP spend in USD (default: 1000)'
)
parser.add_argument(
'--output', '-o',
type=str,
help='Path to write optimization report JSON'
)
parser.add_argument(
'--json',
action='store_true',
help='Output as JSON format'
)
parser.add_argument(
'--checklist',
action='store_true',
help='Generate optimization checklist'
)
args = parser.parse_args()
if args.resources:
try:
with open(args.resources, 'r') as f:
resources = json.load(f)
except FileNotFoundError:
print(f"Error: File '{args.resources}' not found.", file=sys.stderr)
sys.exit(1)
except json.JSONDecodeError:
print(f"Error: File '{args.resources}' is not valid JSON.", file=sys.stderr)
sys.exit(1)
else:
resources = {}
optimizer = CostOptimizer(resources, args.monthly_spend)
result = optimizer.analyze_and_optimize()
if args.checklist:
result['checklist'] = optimizer.generate_optimization_checklist()
if args.output:
with open(args.output, 'w') as f:
json.dump(result, f, indent=2)
print(f"Report written to {args.output}")
elif args.json:
print(json.dumps(result, indent=2))
else:
print(f"\nGCP Cost Optimization Report")
print(f"{'=' * 40}")
print(f"Current Monthly Spend: .2f")
print(f"Potential Savings: .2f")
print(f"Optimized Spend: .2f")
print(f"Savings Percentage: {result['savings_percentage']}%")
print(f"\nTop Priority Actions:")
for i, action in enumerate(result['priority_actions'], 1):
print(f" {i}. [{action['service']}] {action['recommendation']}")
print(f" Savings: .2f/month")
print(f"\nTotal Recommendations: {len(result['recommendations'])}")
if __name__ == '__main__':
main()
FILE:scripts/deployment_manager.py
"""
GCP deployment script generator.
Creates gcloud CLI scripts and Terraform configurations for GCP architectures.
"""
import argparse
import json
import sys
from typing import Dict, Any
class DeploymentManager:
"""Generate GCP deployment scripts and IaC configurations."""
def __init__(self, app_name: str, requirements: Dict[str, Any]):
"""
Initialize with application requirements.
Args:
app_name: Application name (used for resource naming)
requirements: Dictionary with pattern, region, project requirements
"""
self.app_name = app_name.lower().replace(' ', '-')
self.requirements = requirements
self.region = requirements.get('region', 'us-central1')
self.project_id = requirements.get('project_id', 'my-project')
self.pattern = requirements.get('pattern', 'serverless_web')
def generate_gcloud_script(self) -> str:
"""
Generate gcloud CLI deployment script.
Returns:
Shell script as string
"""
if self.pattern == 'serverless_web':
return self._gcloud_serverless_web()
elif self.pattern == 'gke_microservices':
return self._gcloud_gke_microservices()
elif self.pattern == 'data_pipeline':
return self._gcloud_data_pipeline()
else:
return self._gcloud_serverless_web()
def _gcloud_serverless_web(self) -> str:
"""Generate gcloud script for serverless web pattern."""
return f"""#!/bin/bash
# GCP Serverless Web Deployment Script
# Application: {self.app_name}
# Region: {self.region}
# Pattern: Cloud Run + Firestore + Cloud Storage + Cloud CDN
set -euo pipefail
PROJECT_ID="{self.project_id}"
REGION="{self.region}"
APP_NAME="{self.app_name}"
ENVIRONMENT="-dev}"
echo "=== Deploying $APP_NAME to GCP ($ENVIRONMENT) ==="
# 1. Set project
gcloud config set project $PROJECT_ID
# 2. Enable required APIs
echo "Enabling required APIs..."
gcloud services enable \\
run.googleapis.com \\
firestore.googleapis.com \\
cloudbuild.googleapis.com \\
artifactregistry.googleapis.com \\
secretmanager.googleapis.com \\
compute.googleapis.com \\
monitoring.googleapis.com \\
logging.googleapis.com
# 3. Create Artifact Registry repository
echo "Creating Artifact Registry repository..."
gcloud artifacts repositories create $APP_NAME \\
--repository-format=docker \\
--location=$REGION \\
--description="Docker images for $APP_NAME" \\
|| echo "Repository already exists"
# 4. Build and push container image
echo "Building container image..."
gcloud builds submit \\
--tag $REGION-docker.pkg.dev/$PROJECT_ID/$APP_NAME/$APP_NAME:latest \\
.
# 5. Create Firestore database
echo "Creating Firestore database..."
gcloud firestore databases create \\
--location=$REGION \\
--type=firestore-native \\
|| echo "Firestore database already exists"
# 6. Create service account for Cloud Run
echo "Creating service account..."
SA_NAME="{APP_NAME}-run-sa"
gcloud iam service-accounts create $SA_NAME \\
--display-name="$APP_NAME Cloud Run Service Account" \\
|| echo "Service account already exists"
# Grant Firestore access
gcloud projects add-iam-policy-binding $PROJECT_ID \\
--member="serviceAccount:$SA_NAME@$PROJECT_ID.iam.gserviceaccount.com" \\
--role="roles/datastore.user" \\
--condition=None
# Grant Secret Manager access
gcloud projects add-iam-policy-binding $PROJECT_ID \\
--member="serviceAccount:$SA_NAME@$PROJECT_ID.iam.gserviceaccount.com" \\
--role="roles/secretmanager.secretAccessor" \\
--condition=None
# 7. Deploy Cloud Run service
echo "Deploying Cloud Run service..."
gcloud run deploy $APP_NAME-api \\
--image $REGION-docker.pkg.dev/$PROJECT_ID/$APP_NAME/$APP_NAME:latest \\
--region $REGION \\
--platform managed \\
--service-account $SA_NAME@$PROJECT_ID.iam.gserviceaccount.com \\
--memory 512Mi \\
--cpu 1 \\
--min-instances 0 \\
--max-instances 10 \\
--set-env-vars "PROJECT_ID=$PROJECT_ID,ENVIRONMENT=$ENVIRONMENT" \\
--allow-unauthenticated
# 8. Create Cloud Storage bucket for static assets
echo "Creating static assets bucket..."
BUCKET_NAME="{PROJECT_ID}-{APP_NAME}-static"
gsutil mb -l $REGION gs://$BUCKET_NAME/ || echo "Bucket already exists"
gsutil iam ch allUsers:objectViewer gs://$BUCKET_NAME
# 9. Set up Cloud Monitoring alerting
echo "Setting up monitoring..."
gcloud alpha monitoring policies create \\
--notification-channels="" \\
--display-name="$APP_NAME High Error Rate" \\
--condition-display-name="Cloud Run 5xx Error Rate" \\
--condition-filter='resource.type="cloud_run_revision" AND metric.type="run.googleapis.com/request_count" AND metric.labels.response_code_class="5xx"' \\
--condition-threshold-value=10 \\
--condition-threshold-duration=60s \\
|| echo "Alert policy creation requires additional configuration"
# 10. Output deployment info
echo ""
echo "=== Deployment Complete ==="
SERVICE_URL=$(gcloud run services describe $APP_NAME-api --region $REGION --format 'value(status.url)')
echo "Cloud Run URL: $SERVICE_URL"
echo "Static Bucket: gs://$BUCKET_NAME"
echo "Firestore: https://console.cloud.google.com/firestore?project=$PROJECT_ID"
echo "Monitoring: https://console.cloud.google.com/monitoring?project=$PROJECT_ID"
"""
def _gcloud_gke_microservices(self) -> str:
"""Generate gcloud script for GKE microservices pattern."""
return f"""#!/bin/bash
# GCP GKE Microservices Deployment Script
# Application: {self.app_name}
# Region: {self.region}
# Pattern: GKE Autopilot + Cloud SQL + Memorystore
set -euo pipefail
PROJECT_ID="{self.project_id}"
REGION="{self.region}"
APP_NAME="{self.app_name}"
ENVIRONMENT="-dev}"
CLUSTER_NAME="{APP_NAME}-cluster"
NETWORK_NAME="{APP_NAME}-vpc"
echo "=== Deploying $APP_NAME GKE Microservices ($ENVIRONMENT) ==="
# 1. Set project
gcloud config set project $PROJECT_ID
# 2. Enable required APIs
echo "Enabling required APIs..."
gcloud services enable \\
container.googleapis.com \\
sqladmin.googleapis.com \\
redis.googleapis.com \\
cloudbuild.googleapis.com \\
artifactregistry.googleapis.com \\
secretmanager.googleapis.com \\
servicenetworking.googleapis.com \\
compute.googleapis.com
# 3. Create VPC network
echo "Creating VPC network..."
gcloud compute networks create $NETWORK_NAME \\
--subnet-mode=auto \\
|| echo "Network already exists"
# Allocate IP range for private services
gcloud compute addresses create google-managed-services-$NETWORK_NAME \\
--global \\
--purpose=VPC_PEERING \\
--prefix-length=16 \\
--network=$NETWORK_NAME \\
|| echo "IP range already exists"
gcloud services vpc-peerings connect \\
--service=servicenetworking.googleapis.com \\
--ranges=google-managed-services-$NETWORK_NAME \\
--network=$NETWORK_NAME \\
|| echo "VPC peering already exists"
# 4. Create GKE Autopilot cluster
echo "Creating GKE Autopilot cluster..."
gcloud container clusters create-auto $CLUSTER_NAME \\
--region $REGION \\
--network $NETWORK_NAME \\
--release-channel regular \\
--enable-master-authorized-networks \\
--enable-private-nodes \\
|| echo "Cluster already exists"
# 5. Get cluster credentials
gcloud container clusters get-credentials $CLUSTER_NAME --region $REGION
# 6. Create Cloud SQL instance
echo "Creating Cloud SQL instance..."
gcloud sql instances create $APP_NAME-db \\
--database-version=POSTGRES_15 \\
--tier=db-custom-2-8192 \\
--region=$REGION \\
--network=$NETWORK_NAME \\
--no-assign-ip \\
--availability-type=regional \\
--backup-start-time=02:00 \\
--storage-auto-increase \\
|| echo "Cloud SQL instance already exists"
# Create database
gcloud sql databases create $APP_NAME \\
--instance=$APP_NAME-db \\
|| echo "Database already exists"
# 7. Create Memorystore Redis instance
echo "Creating Memorystore Redis instance..."
gcloud redis instances create $APP_NAME-cache \\
--size=1 \\
--region=$REGION \\
--redis-version=redis_7_0 \\
--network=$NETWORK_NAME \\
--tier=basic \\
|| echo "Redis instance already exists"
# 8. Configure Workload Identity
echo "Configuring Workload Identity..."
SA_NAME="{APP_NAME}-workload"
gcloud iam service-accounts create $SA_NAME \\
--display-name="$APP_NAME Workload Identity SA" \\
|| echo "Service account already exists"
gcloud projects add-iam-policy-binding $PROJECT_ID \\
--member="serviceAccount:$SA_NAME@$PROJECT_ID.iam.gserviceaccount.com" \\
--role="roles/cloudsql.client"
gcloud iam service-accounts add-iam-policy-binding \\
$SA_NAME@$PROJECT_ID.iam.gserviceaccount.com \\
--role="roles/iam.workloadIdentityUser" \\
--member="serviceAccount:$PROJECT_ID.svc.id.goog[default/$SA_NAME]"
echo ""
echo "=== GKE Cluster Ready ==="
echo "Cluster: $CLUSTER_NAME"
echo "Cloud SQL: $APP_NAME-db"
echo "Redis: $APP_NAME-cache"
echo ""
echo "Next: Apply Kubernetes manifests with 'kubectl apply -f k8s/'"
"""
def _gcloud_data_pipeline(self) -> str:
"""Generate gcloud script for data pipeline pattern."""
return f"""#!/bin/bash
# GCP Data Pipeline Deployment Script
# Application: {self.app_name}
# Region: {self.region}
# Pattern: Pub/Sub + Dataflow + BigQuery
set -euo pipefail
PROJECT_ID="{self.project_id}"
REGION="{self.region}"
APP_NAME="{self.app_name}"
echo "=== Deploying $APP_NAME Data Pipeline ==="
# 1. Set project
gcloud config set project $PROJECT_ID
# 2. Enable required APIs
echo "Enabling required APIs..."
gcloud services enable \\
pubsub.googleapis.com \\
dataflow.googleapis.com \\
bigquery.googleapis.com \\
storage.googleapis.com \\
monitoring.googleapis.com
# 3. Create Pub/Sub topic and subscription
echo "Creating Pub/Sub resources..."
gcloud pubsub topics create $APP_NAME-events \\
|| echo "Topic already exists"
gcloud pubsub subscriptions create $APP_NAME-events-sub \\
--topic=$APP_NAME-events \\
--ack-deadline=60 \\
--message-retention-duration=7d \\
|| echo "Subscription already exists"
# Dead letter topic
gcloud pubsub topics create $APP_NAME-events-dlq \\
|| echo "DLQ topic already exists"
gcloud pubsub subscriptions update $APP_NAME-events-sub \\
--dead-letter-topic=$APP_NAME-events-dlq \\
--max-delivery-attempts=5
# 4. Create BigQuery dataset and table
echo "Creating BigQuery resources..."
bq mk --dataset --location=$REGION $PROJECT_ID:{APP_NAME//-/_}_analytics \\
|| echo "Dataset already exists"
bq mk --table \\
$PROJECT_ID:{APP_NAME//-/_}_analytics.events \\
event_id:STRING,event_type:STRING,payload:STRING,timestamp:TIMESTAMP,processed_at:TIMESTAMP \\
--time_partitioning_type=DAY \\
--time_partitioning_field=timestamp \\
--clustering_fields=event_type \\
|| echo "Table already exists"
# 5. Create Cloud Storage bucket for Dataflow temp/staging
echo "Creating staging bucket..."
STAGING_BUCKET="{PROJECT_ID}-{APP_NAME}-dataflow"
gsutil mb -l $REGION gs://$STAGING_BUCKET/ || echo "Bucket already exists"
# 6. Create service account for Dataflow
echo "Creating Dataflow service account..."
SA_NAME="{APP_NAME}-dataflow-sa"
gcloud iam service-accounts create $SA_NAME \\
--display-name="$APP_NAME Dataflow Worker SA" \\
|| echo "Service account already exists"
for ROLE in roles/dataflow.worker roles/bigquery.dataEditor roles/pubsub.subscriber roles/storage.objectAdmin; do
gcloud projects add-iam-policy-binding $PROJECT_ID \\
--member="serviceAccount:$SA_NAME@$PROJECT_ID.iam.gserviceaccount.com" \\
--role="$ROLE" \\
--condition=None
done
echo ""
echo "=== Data Pipeline Infrastructure Ready ==="
echo "Pub/Sub Topic: $APP_NAME-events"
echo "BigQuery Dataset: {APP_NAME//-/_}_analytics"
echo "Staging Bucket: gs://$STAGING_BUCKET"
echo ""
echo "Next: Deploy Dataflow job with Apache Beam pipeline"
echo " python -m apache_beam.examples.streaming_wordcount \\\\"
echo " --runner DataflowRunner \\\\"
echo " --project $PROJECT_ID \\\\"
echo " --region $REGION \\\\"
echo " --temp_location gs://$STAGING_BUCKET/temp"
"""
def generate_terraform_configuration(self) -> str:
"""
Generate Terraform configuration for the selected pattern.
Returns:
Terraform HCL configuration as string
"""
if self.pattern == 'serverless_web':
return self._terraform_serverless_web()
elif self.pattern == 'gke_microservices':
return self._terraform_gke_microservices()
else:
return self._terraform_serverless_web()
def _terraform_serverless_web(self) -> str:
"""Generate Terraform for serverless web pattern."""
return f"""terraform {{
required_version = ">= 1.0"
required_providers {{
google = {{
source = "hashicorp/google"
version = "~> 5.0"
}}
}}
}}
provider "google" {{
project = var.project_id
region = var.region
}}
variable "project_id" {{
description = "GCP project ID"
type = string
}}
variable "region" {{
description = "GCP region"
type = string
default = "{self.region}"
}}
variable "environment" {{
description = "Environment name"
type = string
default = "dev"
}}
variable "app_name" {{
description = "Application name"
type = string
default = "{self.app_name}"
}}
# Enable required APIs
resource "google_project_service" "apis" {{
for_each = toset([
"run.googleapis.com",
"firestore.googleapis.com",
"secretmanager.googleapis.com",
"artifactregistry.googleapis.com",
"monitoring.googleapis.com",
])
project = var.project_id
service = each.value
}}
# Service Account for Cloud Run
resource "google_service_account" "cloud_run" {{
account_id = "{var.app_name}-run-sa"
display_name = "{var.app_name} Cloud Run Service Account"
}}
resource "google_project_iam_member" "firestore_user" {{
project = var.project_id
role = "roles/datastore.user"
member = "serviceAccount:{google_service_account.cloud_run.email}"
}}
resource "google_project_iam_member" "secret_accessor" {{
project = var.project_id
role = "roles/secretmanager.secretAccessor"
member = "serviceAccount:{google_service_account.cloud_run.email}"
}}
# Firestore Database
resource "google_firestore_database" "default" {{
project = var.project_id
name = "(default)"
location_id = var.region
type = "FIRESTORE_NATIVE"
depends_on = [google_project_service.apis["firestore.googleapis.com"]]
}}
# Cloud Run Service
resource "google_cloud_run_v2_service" "api" {{
name = "{var.environment}-{var.app_name}-api"
location = var.region
template {{
service_account = google_service_account.cloud_run.email
containers {{
image = "{var.region}-docker.pkg.dev/{var.project_id}/{var.app_name}/{var.app_name}:latest"
resources {{
limits = {{
cpu = "1000m"
memory = "512Mi"
}}
}}
env {{
name = "PROJECT_ID"
value = var.project_id
}}
env {{
name = "ENVIRONMENT"
value = var.environment
}}
}}
scaling {{
min_instance_count = 0
max_instance_count = 10
}}
}}
depends_on = [google_project_service.apis["run.googleapis.com"]]
labels = {{
environment = var.environment
app = var.app_name
}}
}}
# Allow unauthenticated access (public API)
resource "google_cloud_run_v2_service_iam_member" "public" {{
project = var.project_id
location = var.region
name = google_cloud_run_v2_service.api.name
role = "roles/run.invoker"
member = "allUsers"
}}
# Cloud Storage bucket for static assets
resource "google_storage_bucket" "static" {{
name = "{var.project_id}-{var.app_name}-static"
location = var.region
uniform_bucket_level_access = true
website {{
main_page_suffix = "index.html"
not_found_page = "404.html"
}}
lifecycle_rule {{
condition {{
age = 30
}}
action {{
type = "SetStorageClass"
storage_class = "NEARLINE"
}}
}}
labels = {{
environment = var.environment
app = var.app_name
}}
}}
# Outputs
output "cloud_run_url" {{
description = "Cloud Run service URL"
value = google_cloud_run_v2_service.api.uri
}}
output "static_bucket" {{
description = "Static assets bucket name"
value = google_storage_bucket.static.name
}}
output "service_account" {{
description = "Cloud Run service account email"
value = google_service_account.cloud_run.email
}}
"""
def _terraform_gke_microservices(self) -> str:
"""Generate Terraform for GKE microservices pattern."""
return f"""terraform {{
required_version = ">= 1.0"
required_providers {{
google = {{
source = "hashicorp/google"
version = "~> 5.0"
}}
}}
}}
provider "google" {{
project = var.project_id
region = var.region
}}
variable "project_id" {{
description = "GCP project ID"
type = string
}}
variable "region" {{
description = "GCP region"
type = string
default = "{self.region}"
}}
variable "environment" {{
description = "Environment name"
type = string
default = "dev"
}}
variable "app_name" {{
description = "Application name"
type = string
default = "{self.app_name}"
}}
# Enable required APIs
resource "google_project_service" "apis" {{
for_each = toset([
"container.googleapis.com",
"sqladmin.googleapis.com",
"redis.googleapis.com",
"servicenetworking.googleapis.com",
"secretmanager.googleapis.com",
])
project = var.project_id
service = each.value
}}
# VPC Network
resource "google_compute_network" "main" {{
name = "{var.app_name}-vpc"
auto_create_subnetworks = false
}}
resource "google_compute_subnetwork" "main" {{
name = "{var.app_name}-subnet"
ip_cidr_range = "10.0.0.0/20"
region = var.region
network = google_compute_network.main.id
secondary_ip_range {{
range_name = "pods"
ip_cidr_range = "10.4.0.0/14"
}}
secondary_ip_range {{
range_name = "services"
ip_cidr_range = "10.8.0.0/20"
}}
}}
# GKE Autopilot Cluster
resource "google_container_cluster" "main" {{
name = "{var.environment}-{var.app_name}-cluster"
location = var.region
enable_autopilot = true
network = google_compute_network.main.name
subnetwork = google_compute_subnetwork.main.name
ip_allocation_policy {{
cluster_secondary_range_name = "pods"
services_secondary_range_name = "services"
}}
release_channel {{
channel = "REGULAR"
}}
depends_on = [google_project_service.apis["container.googleapis.com"]]
}}
# Private Services Access for Cloud SQL
resource "google_compute_global_address" "private_ip" {{
name = "private-ip-range"
purpose = "VPC_PEERING"
address_type = "INTERNAL"
prefix_length = 16
network = google_compute_network.main.id
}}
resource "google_service_networking_connection" "private_vpc" {{
network = google_compute_network.main.id
service = "servicenetworking.googleapis.com"
reserved_peering_ranges = [google_compute_global_address.private_ip.name]
}}
# Cloud SQL PostgreSQL
resource "google_sql_database_instance" "main" {{
name = "{var.environment}-{var.app_name}-db"
database_version = "POSTGRES_15"
region = var.region
settings {{
tier = "db-custom-2-8192"
availability_type = "REGIONAL"
backup_configuration {{
enabled = true
start_time = "02:00"
point_in_time_recovery_enabled = true
}}
ip_configuration {{
ipv4_enabled = false
private_network = google_compute_network.main.id
}}
disk_autoresize = true
}}
depends_on = [google_service_networking_connection.private_vpc]
}}
resource "google_sql_database" "app" {{
name = var.app_name
instance = google_sql_database_instance.main.name
}}
# Memorystore Redis
resource "google_redis_instance" "cache" {{
name = "{var.environment}-{var.app_name}-cache"
tier = "BASIC"
memory_size_gb = 1
region = var.region
redis_version = "REDIS_7_0"
authorized_network = google_compute_network.main.id
depends_on = [google_project_service.apis["redis.googleapis.com"]]
labels = {{
environment = var.environment
app = var.app_name
}}
}}
# Outputs
output "cluster_name" {{
description = "GKE cluster name"
value = google_container_cluster.main.name
}}
output "cloud_sql_connection" {{
description = "Cloud SQL connection name"
value = google_sql_database_instance.main.connection_name
}}
output "redis_host" {{
description = "Memorystore Redis host"
value = google_redis_instance.cache.host
}}
"""
def main():
parser = argparse.ArgumentParser(
description='GCP Deployment Manager - Generates gcloud CLI scripts and Terraform configurations'
)
parser.add_argument(
'--app-name', '-a',
type=str,
required=True,
help='Application name'
)
parser.add_argument(
'--pattern', '-p',
type=str,
choices=['serverless_web', 'gke_microservices', 'data_pipeline'],
default='serverless_web',
help='Architecture pattern (default: serverless_web)'
)
parser.add_argument(
'--region', '-r',
type=str,
default='us-central1',
help='GCP region (default: us-central1)'
)
parser.add_argument(
'--project-id',
type=str,
default='my-project',
help='GCP project ID (default: my-project)'
)
parser.add_argument(
'--format', '-f',
type=str,
choices=['gcloud', 'terraform', 'both'],
default='both',
help='Output format (default: both)'
)
parser.add_argument(
'--output', '-o',
type=str,
help='Output directory for generated files'
)
parser.add_argument(
'--json',
action='store_true',
help='Output as JSON format'
)
args = parser.parse_args()
requirements = {
'pattern': args.pattern,
'region': args.region,
'project_id': args.project_id
}
manager = DeploymentManager(args.app_name, requirements)
if args.json:
output = {}
if args.format in ('gcloud', 'both'):
output['gcloud_script'] = manager.generate_gcloud_script()
if args.format in ('terraform', 'both'):
output['terraform_config'] = manager.generate_terraform_configuration()
print(json.dumps(output, indent=2))
elif args.output:
import os
os.makedirs(args.output, exist_ok=True)
if args.format in ('gcloud', 'both'):
gcloud_path = os.path.join(args.output, 'deploy.sh')
with open(gcloud_path, 'w') as f:
f.write(manager.generate_gcloud_script())
os.chmod(gcloud_path, 0o755)
print(f"gcloud script written to {gcloud_path}")
if args.format in ('terraform', 'both'):
tf_path = os.path.join(args.output, 'main.tf')
with open(tf_path, 'w') as f:
f.write(manager.generate_terraform_configuration())
print(f"Terraform config written to {tf_path}")
else:
if args.format in ('gcloud', 'both'):
print("# ===== gcloud CLI Script =====")
print(manager.generate_gcloud_script())
if args.format in ('terraform', 'both'):
print("# ===== Terraform Configuration =====")
print(manager.generate_terraform_configuration())
if __name__ == '__main__':
main()
Khung ứng phó sự cố từ phát hiện đến xử lý và rà soát sau sự cố: phân loại mức độ, dựng dòng thời gian, phân tích có cấu trúc.
---
name: "incident-commander"
description: "Comprehensive incident response framework from detection through resolution and post-incident review. Battle-tested SRE/DevOps practices: severity classification, timeline reconstruction, structured post-incident analysis. Use when declaring an incident, coordinating multi-team response during an outage, leading a post-mortem, or setting up on-call practices for a new service."
---
# Incident Commander Skill
**Category:** Engineering Team
**Tier:** POWERFUL
**Author:** Claude Skills Team
**Version:** 1.0.0
**Last Updated:** February 2026
## Overview
The Incident Commander skill provides a comprehensive incident response framework for managing technology incidents from detection through resolution and post-incident review. This skill implements battle-tested practices from SRE and DevOps teams at scale, providing structured tools for severity classification, timeline reconstruction, and thorough post-incident analysis.
## Key Features
- **Automated Severity Classification** - Intelligent incident triage based on impact and urgency metrics
- **Timeline Reconstruction** - Transform scattered logs and events into coherent incident narratives
- **Post-Incident Review Generation** - Structured PIRs with multiple RCA frameworks
- **Communication Templates** - Pre-built templates for stakeholder updates and escalations
- **Runbook Integration** - Generate actionable runbooks from incident patterns
## Skills Included
### Core Tools
1. **Incident Classifier** (`incident_classifier.py`)
- Analyzes incident descriptions and outputs severity levels
- Recommends response teams and initial actions
- Generates communication templates based on severity
2. **Timeline Reconstructor** (`timeline_reconstructor.py`)
- Processes timestamped events from multiple sources
- Reconstructs chronological incident timeline
- Identifies gaps and provides duration analysis
3. **PIR Generator** (`pir_generator.py`)
- Creates comprehensive Post-Incident Review documents
- Applies multiple RCA frameworks (5 Whys, Fishbone, Timeline)
- Generates actionable follow-up items
## Incident Response Framework
### Severity Classification System
#### SEV1 - Critical Outage
**Definition:** Complete service failure affecting all users or critical business functions
**Characteristics:**
- Customer-facing services completely unavailable
- Data loss or corruption affecting users
- Security breaches with customer data exposure
- Revenue-generating systems down
- SLA violations with financial penalties
**Response Requirements:**
- Immediate escalation to on-call engineer
- Incident Commander assigned within 5 minutes
- Executive notification within 15 minutes
- Public status page update within 15 minutes
- War room established
- All hands on deck if needed
**Communication Frequency:** Every 15 minutes until resolution
#### SEV2 - Major Impact
**Definition:** Significant degradation affecting subset of users or non-critical functions
**Characteristics:**
- Partial service degradation (>25% of users affected)
- Performance issues causing user frustration
- Non-critical features unavailable
- Internal tools impacting productivity
- Data inconsistencies not affecting user experience
**Response Requirements:**
- On-call engineer response within 15 minutes
- Incident Commander assigned within 30 minutes
- Status page update within 30 minutes
- Stakeholder notification within 1 hour
- Regular team updates
**Communication Frequency:** Every 30 minutes during active response
#### SEV3 - Minor Impact
**Definition:** Limited impact with workarounds available
**Characteristics:**
- Single feature or component affected
- <25% of users impacted
- Workarounds available
- Performance degradation not significantly impacting UX
- Non-urgent monitoring alerts
**Response Requirements:**
- Response within 2 hours during business hours
- Next business day response acceptable outside hours
- Internal team notification
- Optional status page update
**Communication Frequency:** At key milestones only
#### SEV4 - Low Impact
**Definition:** Minimal impact, cosmetic issues, or planned maintenance
**Characteristics:**
- Cosmetic bugs
- Documentation issues
- Logging or monitoring gaps
- Performance issues with no user impact
- Development/test environment issues
**Response Requirements:**
- Response within 1-2 business days
- Standard ticket/issue tracking
- No special escalation required
**Communication Frequency:** Standard development cycle updates
### Incident Commander Role
#### Primary Responsibilities
1. **Command and Control**
- Own the incident response process
- Make critical decisions about resource allocation
- Coordinate between technical teams and stakeholders
- Maintain situational awareness across all response streams
2. **Communication Hub**
- Provide regular updates to stakeholders
- Manage external communications (status pages, customer notifications)
- Facilitate effective communication between response teams
- Shield responders from external distractions
3. **Process Management**
- Ensure proper incident tracking and documentation
- Drive toward resolution while maintaining quality
- Coordinate handoffs between team members
- Plan and execute rollback strategies if needed
4. **Post-Incident Leadership**
- Ensure thorough post-incident reviews are conducted
- Drive implementation of preventive measures
- Share learnings with broader organization
#### Decision-Making Framework
**Emergency Decisions (SEV1/2):**
- Incident Commander has full authority
- Bias toward action over analysis
- Document decisions for later review
- Consult subject matter experts but don't get blocked
**Resource Allocation:**
- Can pull in any necessary team members
- Authority to escalate to senior leadership
- Can approve emergency spend for external resources
- Make call on communication channels and timing
**Technical Decisions:**
- Lean on technical leads for implementation details
- Make final calls on trade-offs between speed and risk
- Approve rollback vs. fix-forward strategies
- Coordinate testing and validation approaches
### Communication Templates
#### Initial Incident Notification (SEV1/2)
```
Subject: [SEV{severity}] {Service Name} - {Brief Description}
Incident Details:
- Start Time: {timestamp}
- Severity: SEV{level}
- Impact: {user impact description}
- Current Status: {investigating/mitigating/resolved}
Technical Details:
- Affected Services: {service list}
- Symptoms: {what users are experiencing}
- Initial Assessment: {suspected root cause if known}
Response Team:
- Incident Commander: {name}
- Technical Lead: {name}
- SMEs Engaged: {list}
Next Update: {timestamp}
Status Page: {link}
War Room: {bridge/chat link}
---
{Incident Commander Name}
{Contact Information}
```
#### Executive Summary (SEV1)
```
Subject: URGENT - Customer-Impacting Outage - {Service Name}
Executive Summary:
{2-3 sentence description of customer impact and business implications}
Key Metrics:
- Time to Detection: {X minutes}
- Time to Engagement: {X minutes}
- Estimated Customer Impact: {number/percentage}
- Current Status: {status}
- ETA to Resolution: {time or "investigating"}
Leadership Actions Required:
- [ ] Customer communication approval
- [ ] PR/Communications coordination
- [ ] Resource allocation decisions
- [ ] External vendor engagement
Incident Commander: {name} ({contact})
Next Update: {time}
---
This is an automated alert from our incident response system.
```
#### Customer Communication Template
```
We are currently experiencing {brief description of issue} affecting {scope of impact}.
Our engineering team was alerted at {time} and is actively working to resolve the issue. We will provide updates every {frequency} until resolved.
What we know:
- {factual statement of impact}
- {factual statement of scope}
- {brief status of response}
What we're doing:
- {primary response action}
- {secondary response action}
Workaround (if available):
{workaround steps or "No workaround currently available"}
We apologize for the inconvenience and will share more information as it becomes available.
Next update: {time}
Status page: {link}
```
### Stakeholder Management
#### Stakeholder Classification
**Internal Stakeholders:**
- **Engineering Leadership** - Technical decisions and resource allocation
- **Product Management** - Customer impact assessment and feature implications
- **Customer Support** - User communication and support ticket management
- **Sales/Account Management** - Customer relationship management for enterprise clients
- **Executive Team** - Business impact decisions and external communication approval
- **Legal/Compliance** - Regulatory reporting and liability assessment
**External Stakeholders:**
- **Customers** - Service availability and impact communication
- **Partners** - API availability and integration impacts
- **Vendors** - Third-party service dependencies and support escalation
- **Regulators** - Compliance reporting for regulated industries
- **Public/Media** - Transparency for public-facing outages
#### Communication Cadence by Stakeholder
| Stakeholder | SEV1 | SEV2 | SEV3 | SEV4 |
|-------------|------|------|------|------|
| Engineering Leadership | Real-time | 30min | 4hrs | Daily |
| Executive Team | 15min | 1hr | EOD | Weekly |
| Customer Support | Real-time | 30min | 2hrs | As needed |
| Customers | 15min | 1hr | Optional | None |
| Partners | 30min | 2hrs | Optional | None |
### Runbook Generation Framework
#### Dynamic Runbook Components
1. **Detection Playbooks**
- Monitoring alert definitions
- Triage decision trees
- Escalation trigger points
- Initial response actions
2. **Response Playbooks**
- Step-by-step mitigation procedures
- Rollback instructions
- Validation checkpoints
- Communication checkpoints
3. **Recovery Playbooks**
- Service restoration procedures
- Data consistency checks
- Performance validation
- User notification processes
#### Runbook Template Structure
```markdown
# {Service/Component} Incident Response Runbook
## Quick Reference
- **Severity Indicators:** {list of conditions for each severity level}
- **Key Contacts:** {on-call rotations and escalation paths}
- **Critical Commands:** {list of emergency commands with descriptions}
## Detection
### Monitoring Alerts
- {Alert name}: {description and thresholds}
- {Alert name}: {description and thresholds}
### Manual Detection Signs
- {Symptom}: {what to look for and where}
- {Symptom}: {what to look for and where}
## Initial Response (0-15 minutes)
1. **Assess Severity**
- [ ] Check {primary metric}
- [ ] Verify {secondary indicator}
- [ ] Classify as SEV{level} based on {criteria}
2. **Establish Command**
- [ ] Page Incident Commander if SEV1/2
- [ ] Create incident tracking ticket
- [ ] Join war room: {link/bridge info}
3. **Initial Investigation**
- [ ] Check recent deployments: {deployment log location}
- [ ] Review error logs: {log location and queries}
- [ ] Verify dependencies: {dependency check commands}
## Mitigation Strategies
### Strategy 1: {Name}
**Use when:** {conditions}
**Steps:**
1. {detailed step with commands}
2. {detailed step with expected outcomes}
3. {validation step}
**Rollback Plan:**
1. {rollback step}
2. {verification step}
### Strategy 2: {Name}
{similar structure}
## Recovery and Validation
1. **Service Restoration**
- [ ] {restoration step}
- [ ] Wait for {metric} to return to normal
- [ ] Validate end-to-end functionality
2. **Communication**
- [ ] Update status page
- [ ] Notify stakeholders
- [ ] Schedule PIR
## Common Pitfalls
- **{Pitfall}:** {description and how to avoid}
- **{Pitfall}:** {description and how to avoid}
## Reference Information
→ See references/reference-information.md for details
## Usage Examples
### Example 1: Database Connection Pool Exhaustion
```bash
# Classify the incident
echo '{"description": "Users reporting 500 errors, database connections timing out", "affected_users": "80%", "business_impact": "high"}' | python scripts/incident_classifier.py
# Reconstruct timeline from logs
python scripts/timeline_reconstructor.py --input assets/db_incident_events.json --output timeline.md
# Generate PIR after resolution
python scripts/pir_generator.py --incident assets/db_incident_data.json --timeline timeline.md --output pir.md
```
### Example 2: API Rate Limiting Incident
```bash
# Quick classification from stdin
echo "API rate limits causing customer API calls to fail" | python scripts/incident_classifier.py --format text
# Build timeline from multiple sources
python scripts/timeline_reconstructor.py --input assets/api_incident_logs.json --detect-phases --gap-analysis
# Generate comprehensive PIR
python scripts/pir_generator.py --incident assets/api_incident_summary.json --rca-method fishbone --action-items
```
## Best Practices
### During Incident Response
1. **Maintain Calm Leadership**
- Stay composed under pressure
- Make decisive calls with incomplete information
- Communicate confidence while acknowledging uncertainty
2. **Document Everything**
- All actions taken and their outcomes
- Decision rationale, especially for controversial calls
- Timeline of events as they happen
3. **Effective Communication**
- Use clear, jargon-free language
- Provide regular updates even when there's no new information
- Manage stakeholder expectations proactively
4. **Technical Excellence**
- Prefer rollbacks to risky fixes under pressure
- Validate fixes before declaring resolution
- Plan for secondary failures and cascading effects
### Post-Incident
1. **Blameless Culture**
- Focus on system failures, not individual mistakes
- Encourage honest reporting of what went wrong
- Celebrate learning and improvement opportunities
2. **Action Item Discipline**
- Assign specific owners and due dates
- Track progress publicly
- Prioritize based on risk and effort
3. **Knowledge Sharing**
- Share PIRs broadly within the organization
- Update runbooks based on lessons learned
- Conduct training sessions for common failure modes
4. **Continuous Improvement**
- Look for patterns across multiple incidents
- Invest in tooling and automation
- Regularly review and update processes
## Integration with Existing Tools
### Monitoring and Alerting
- PagerDuty/Opsgenie integration for escalation
- Datadog/Grafana for metrics and dashboards
- ELK/Splunk for log analysis and correlation
### Communication Platforms
- Slack/Teams for war room coordination
- Zoom/Meet for video bridges
- Status page providers (Statuspage.io, etc.)
### Documentation Systems
- Confluence/Notion for PIR storage
- GitHub/GitLab for runbook version control
- JIRA/Linear for action item tracking
### Change Management
- CI/CD pipeline integration
- Deployment tracking systems
- Feature flag platforms for quick rollbacks
## Conclusion
The Incident Commander skill provides a comprehensive framework for managing incidents from detection through post-incident review. By implementing structured processes, clear communication templates, and thorough analysis tools, teams can improve their incident response capabilities and build more resilient systems.
The key to successful incident management is preparation, practice, and continuous learning. Use this framework as a starting point, but adapt it to your organization's specific needs, culture, and technical environment.
Remember: The goal isn't to prevent all incidents (which is impossible), but to detect them quickly, respond effectively, communicate clearly, and learn continuously.
FILE:assets/incident_report_template.md
# Incident Report: [INC-YYYY-NNNN] [Title]
**Severity:** SEV[1-4]
**Status:** [Active | Mitigated | Resolved]
**Incident Commander:** [Name]
**Date:** [YYYY-MM-DD]
---
## Executive Summary
[2-3 sentence summary of the incident: what happened, impact scope, resolution status. Written for executive audience — no jargon, focus on business impact.]
---
## Impact Statement
| Metric | Value |
|--------|-------|
| **Duration** | [X hours Y minutes] |
| **Affected Users** | [number or percentage] |
| **Failed Transactions** | [number] |
| **Revenue Impact** | $[amount] |
| **Data Loss** | [Yes/No — if yes, detail below] |
| **SLA Impact** | [X.XX% availability for period] |
| **Affected Regions** | [list regions] |
| **Affected Services** | [list services] |
### Customer-Facing Impact
[Describe what customers experienced: error messages, degraded functionality, complete outage. Be specific about which user journeys were affected.]
---
## Timeline
| Time (UTC) | Phase | Event |
|------------|-------|-------|
| HH:MM | Detection | [First alert or report] |
| HH:MM | Declaration | [Incident declared, channel created] |
| HH:MM | Investigation | [Key investigation findings] |
| HH:MM | Mitigation | [Mitigation action taken] |
| HH:MM | Resolution | [Permanent fix applied] |
| HH:MM | Closure | [Incident closed, monitoring confirmed stable] |
### Key Decision Points
1. **[HH:MM] [Decision]** — [Rationale and outcome]
2. **[HH:MM] [Decision]** — [Rationale and outcome]
### Timeline Gaps
[Note any periods >15 minutes without logged events. These represent potential blind spots in the response.]
---
## Root Cause Analysis
### Root Cause
[Clear, specific statement of the root cause. Not "human error" — describe the systemic failure.]
### Contributing Factors
1. **[Factor Category: Process/Tooling/Human/Environment]** — [Description]
2. **[Factor Category]** — [Description]
3. **[Factor Category]** — [Description]
### 5-Whys Analysis
**Why did the service degrade?**
→ [Answer]
**Why did [answer above] happen?**
→ [Answer]
**Why did [answer above] happen?**
→ [Answer]
**Why did [answer above] happen?**
→ [Answer]
**Why did [answer above] happen?**
→ [Root systemic cause]
---
## Response Metrics
| Metric | Value | Target | Status |
|--------|-------|--------|--------|
| **MTTD** (Mean Time to Detect) | [X min] | <5 min | [Met/Missed] |
| **Time to Declare** | [X min] | <10 min | [Met/Missed] |
| **Time to Mitigate** | [X min] | <60 min (SEV1) | [Met/Missed] |
| **MTTR** (Mean Time to Resolve) | [X min] | <4 hr (SEV1) | [Met/Missed] |
| **Postmortem Timeliness** | [X hours] | <72 hr | [Met/Missed] |
---
## Action Items
| # | Priority | Action | Owner | Deadline | Type | Status |
|---|----------|--------|-------|----------|------|--------|
| 1 | P1 | [Action description] | [owner] | [date] | Detection | Open |
| 2 | P1 | [Action description] | [owner] | [date] | Prevention | Open |
| 3 | P2 | [Action description] | [owner] | [date] | Prevention | Open |
| 4 | P2 | [Action description] | [owner] | [date] | Process | Open |
### Action Item Types
- **Detection**: Improve ability to detect this class of issue faster
- **Prevention**: Prevent this class of issue from occurring
- **Mitigation**: Reduce impact when this class of issue occurs
- **Process**: Improve response process and coordination
---
## Lessons Learned
### What Went Well
- [Specific positive outcome from the response]
- [Specific positive outcome]
### What Didn't Go Well
- [Specific area for improvement]
- [Specific area for improvement]
### Where We Got Lucky
- [Things that could have made this worse but didn't]
---
## Communication Log
| Time (UTC) | Channel | Audience | Summary |
|------------|---------|----------|---------|
| HH:MM | Status Page | External | [Summary of update] |
| HH:MM | Slack #exec | Internal | [Summary of update] |
| HH:MM | Email | Customers | [Summary of notification] |
---
## Participants
| Name | Role |
|------|------|
| [Name] | Incident Commander |
| [Name] | Operations Lead |
| [Name] | Communications Lead |
| [Name] | Subject Matter Expert |
---
## Appendix
### Related Incidents
- [INC-YYYY-NNNN] — [Brief description of related incident]
### Reference Links
- [Link to monitoring dashboard]
- [Link to deployment logs]
- [Link to incident channel archive]
---
*This report follows the blameless postmortem principle. The goal is systemic improvement, not individual accountability. All contributing factors should trace to process, tooling, or environmental gaps that can be addressed with concrete action items.*
FILE:assets/runbook_template.md
# Runbook: [Service/Component Name]
**Owner:** [Team Name]
**Last Updated:** [YYYY-MM-DD]
**Reviewed By:** [Name]
**Review Cadence:** Quarterly
---
## Service Overview
| Property | Value |
|----------|-------|
| **Service** | [service-name] |
| **Repository** | [repo URL] |
| **Dashboard** | [monitoring dashboard URL] |
| **On-Call Rotation** | [PagerDuty/OpsGenie schedule URL] |
| **SLA Tier** | [Tier 1/2/3] |
| **Availability Target** | [99.9% / 99.95% / 99.99%] |
| **Dependencies** | [list upstream/downstream services] |
| **Owner Team** | [team name] |
| **Escalation Contact** | [name/email] |
### Architecture Summary
[2-3 sentence description of the service architecture. Include key components, data stores, and external dependencies.]
---
## Alert Response Decision Tree
### High Error Rate (>5%)
```
Error Rate Alert Fired
├── Check: Is this a deployment-related issue?
│ ├── YES → Go to "Recent Deployment Rollback" section
│ └── NO → Continue
├── Check: Is a downstream dependency failing?
│ ├── YES → Go to "Dependency Failure" section
│ └── NO → Continue
├── Check: Is there unusual traffic volume?
│ ├── YES → Go to "Traffic Spike" section
│ └── NO → Continue
└── Escalate: Engage on-call secondary + service owner
```
### High Latency (p99 > [threshold]ms)
```
Latency Alert Fired
├── Check: Database query latency elevated?
│ ├── YES → Go to "Database Performance" section
│ └── NO → Continue
├── Check: Connection pool utilization >80%?
│ ├── YES → Go to "Connection Pool Exhaustion" section
│ └── NO → Continue
├── Check: Memory/CPU pressure on service instances?
│ ├── YES → Go to "Resource Exhaustion" section
│ └── NO → Continue
└── Escalate: Engage on-call secondary + service owner
```
### Service Unavailable (Health Check Failing)
```
Health Check Alert Fired
├── Check: Are all instances down?
│ ├── YES → Go to "Complete Outage" section
│ └── NO → Continue
├── Check: Is only one AZ affected?
│ ├── YES → Go to "AZ Failure" section
│ └── NO → Continue
├── Check: Can instances be restarted?
│ ├── YES → Go to "Instance Restart" section
│ └── NO → Continue
└── Escalate: Declare incident, engage IC
```
---
## Common Scenarios
### Recent Deployment Rollback
**Symptoms:** Error rate spike or latency increase within 60 minutes of a deployment.
**Diagnosis:**
1. Check deployment history: `kubectl rollout history deployment/[service-name]`
2. Compare error rate timing with deployment timestamp
3. Review deployment diff for risky changes
**Mitigation:**
1. Initiate rollback: `kubectl rollout undo deployment/[service-name]`
2. Verify rollback: `kubectl rollout status deployment/[service-name]`
3. Confirm error rate returns to baseline (allow 5 minutes)
4. If rollback fails: escalate immediately
**Communication:** If customer-impacting, update status page within 5 minutes of confirming impact.
---
### Database Performance
**Symptoms:** Elevated query latency, connection pool saturation, timeout errors.
**Diagnosis:**
1. Check active queries: `SELECT * FROM pg_stat_activity WHERE state = 'active';`
2. Check for long-running queries: `SELECT pid, now() - pg_stat_activity.query_start AS duration, query FROM pg_stat_activity WHERE state != 'idle' ORDER BY duration DESC;`
3. Check connection count: `SELECT count(*) FROM pg_stat_activity;`
4. Check table bloat and vacuum status
**Mitigation:**
1. Kill long-running queries if identified: `SELECT pg_terminate_backend([pid]);`
2. If connection pool exhausted: increase pool size via config (requires restart)
3. If read replica available: redirect read traffic
4. If write-heavy: identify and defer non-critical writes
**Escalation Trigger:** If query latency >10s for >5 minutes, escalate to DBA on-call.
---
### Connection Pool Exhaustion
**Symptoms:** Connection timeout errors, pool utilization >90%, requests queuing.
**Diagnosis:**
1. Check pool metrics: current size, active connections, waiting requests
2. Check for connection leaks: connections held >30s without activity
3. Review recent config changes or deployments
**Mitigation:**
1. Increase pool size (if infrastructure allows): update config, rolling restart
2. Kill idle connections exceeding timeout
3. If caused by leak: identify and restart affected instances
4. Enable connection pool auto-scaling if available
**Prevention:** Pool utilization alerting at 70% (warning) and 85% (critical).
---
### Dependency Failure
**Symptoms:** Errors correlated with downstream service failures, circuit breakers tripping.
**Diagnosis:**
1. Check dependency status dashboards
2. Verify circuit breaker state: open/half-open/closed
3. Check for correlation with dependency deployments or incidents
4. Test dependency health endpoints directly
**Mitigation:**
1. If circuit breaker not tripping: verify timeout/threshold configuration
2. Enable graceful degradation (serve cached/default responses)
3. If critical path: engage dependency team via incident process
4. If non-critical path: disable feature flag for affected functionality
**Communication:** Coordinate with dependency team IC if both services have active incidents.
---
### Traffic Spike
**Symptoms:** Sudden traffic increase beyond normal patterns, resource saturation.
**Diagnosis:**
1. Check traffic source: organic growth vs. bot traffic vs. DDoS
2. Review rate limiting effectiveness
3. Check auto-scaling status and capacity
**Mitigation:**
1. If bot/DDoS: enable rate limiting, engage security team
2. If organic: trigger manual scale-up, increase auto-scaling limits
3. Enable request queuing or load shedding if at capacity
4. Consider feature flag toggles to reduce per-request cost
---
### Complete Outage
**Symptoms:** All instances unreachable, health checks failing across AZs.
**Diagnosis:**
1. Check infrastructure status (AWS/GCP status page)
2. Verify network connectivity and DNS resolution
3. Check for infrastructure-level incidents (region outage)
4. Review recent infrastructure changes (Terraform, network config)
**Mitigation:**
1. If infra provider issue: activate disaster recovery plan
2. If DNS issue: update DNS records, reduce TTL
3. If deployment corruption: redeploy last known good version
4. If data corruption: engage data recovery procedures
**Escalation:** Immediately declare SEV1 incident. Engage infrastructure team and management.
---
### Instance Restart
**Symptoms:** Individual instances unhealthy, OOM kills, process crashes.
**Diagnosis:**
1. Check instance logs for crash reason
2. Review memory/CPU usage patterns before crash
3. Check for memory leaks or resource exhaustion
4. Verify configuration consistency across instances
**Mitigation:**
1. Restart unhealthy instances: `kubectl delete pod [pod-name]`
2. If recurring: cordon node and migrate workloads
3. If memory leak: schedule immediate patch with increased memory limit
4. Monitor for recurrence after restart
---
### AZ Failure
**Symptoms:** All instances in one availability zone failing, others healthy.
**Diagnosis:**
1. Confirm AZ-specific failure vs. instance-specific issues
2. Check cloud provider AZ status
3. Verify load balancer is routing around failed AZ
**Mitigation:**
1. Ensure load balancer marks AZ instances as unhealthy
2. Scale up remaining AZs to handle redirected traffic
3. If auto-scaling: verify it's responding to increased load
4. Monitor remaining AZs for cascade effects
---
## Key Metrics & Dashboards
| Metric | Normal Range | Warning | Critical | Dashboard |
|--------|-------------|---------|----------|-----------|
| Error Rate | <0.1% | >1% | >5% | [link] |
| p99 Latency | <200ms | >500ms | >2000ms | [link] |
| CPU Usage | <60% | >75% | >90% | [link] |
| Memory Usage | <70% | >80% | >90% | [link] |
| DB Pool Usage | <50% | >70% | >85% | [link] |
| Request Rate | [baseline]±20% | ±50% | ±100% | [link] |
---
## Escalation Contacts
| Level | Contact | When |
|-------|---------|------|
| L1: On-Call Primary | [name/rotation] | First responder |
| L2: On-Call Secondary | [name/rotation] | Primary unavailable or needs help |
| L3: Service Owner | [name] | Complex issues, architectural decisions |
| L4: Engineering Manager | [name] | SEV1/SEV2, customer impact, resource needs |
| L5: VP Engineering | [name] | SEV1 >30 min, major customer/revenue impact |
---
## Maintenance Procedures
### Planned Maintenance Checklist
- [ ] Maintenance window scheduled and communicated (72 hours advance for Tier 1)
- [ ] Status page updated with planned maintenance notice
- [ ] Rollback plan documented and tested
- [ ] On-call notified of maintenance window
- [ ] Customer notification sent (if SLA-impacting)
- [ ] Post-maintenance verification plan ready
### Health Verification After Changes
1. Check all health endpoints return 200
2. Verify error rate returns to baseline within 5 minutes
3. Confirm latency within normal range
4. Run synthetic transaction test
5. Monitor for 15 minutes before declaring success
---
## Revision History
| Date | Author | Change |
|------|--------|--------|
| [YYYY-MM-DD] | [Name] | Initial version |
| [YYYY-MM-DD] | [Name] | [Description of update] |
---
*This runbook should be reviewed quarterly and updated after every incident that reveals missing procedures. The on-call engineer should be able to follow this document without prior context about the service. If any section requires tribal knowledge to execute, it needs to be expanded.*
FILE:assets/sample_incident_classification.json
{
"description": "Database connection timeouts causing 500 errors for payment processing API. Users unable to complete checkout. Error rate spiked from 0.1% to 45% starting at 14:30 UTC. Database monitoring shows connection pool exhaustion with 200/200 connections active.",
"service": "payment-api",
"affected_users": "80%",
"business_impact": "high",
"duration_minutes": 95,
"metadata": {
"error_rate": "45%",
"connection_pool_utilization": "100%",
"affected_regions": ["us-west", "us-east", "eu-west"],
"detection_method": "monitoring_alert",
"customer_escalations": 12
}
}
FILE:assets/sample_incident_data.json
{
"incident": {
"id": "INC-2024-0142",
"title": "Payment Service Degradation",
"severity": "SEV1",
"status": "resolved",
"declared_at": "2024-01-15T14:23:00Z",
"resolved_at": "2024-01-15T16:45:00Z",
"commander": "Jane Smith",
"service": "payment-gateway",
"affected_services": ["checkout", "subscription-billing"]
},
"events": [
{
"timestamp": "2024-01-15T14:15:00Z",
"type": "trigger",
"actor": "system",
"description": "Database connection pool utilization reaches 95% on payment-gateway primary",
"metadata": {"metric": "db_pool_utilization", "value": 95, "threshold": 90}
},
{
"timestamp": "2024-01-15T14:20:00Z",
"type": "detection",
"actor": "monitoring",
"description": "PagerDuty alert fired: payment-gateway error rate >5% (current: 8.2%)",
"metadata": {"alert_id": "PD-98765", "source": "datadog", "error_rate": 8.2}
},
{
"timestamp": "2024-01-15T14:21:00Z",
"type": "detection",
"actor": "monitoring",
"description": "Datadog alert: p99 latency on /api/payments exceeds 5000ms (current: 8500ms)",
"metadata": {"alert_id": "DD-54321", "source": "datadog", "latency_p99_ms": 8500}
},
{
"timestamp": "2024-01-15T14:23:00Z",
"type": "declaration",
"actor": "Jane Smith",
"description": "SEV1 declared. Incident channel #inc-20240115-payment-degradation created. Bridge call started.",
"metadata": {"channel": "#inc-20240115-payment-degradation", "severity": "SEV1"}
},
{
"timestamp": "2024-01-15T14:25:00Z",
"type": "investigation",
"actor": "Alice Chen",
"description": "Confirmed: database connection pool at 100% utilization. All new connections being rejected.",
"metadata": {"pool_size": 20, "active_connections": 20, "waiting_requests": 147}
},
{
"timestamp": "2024-01-15T14:28:00Z",
"type": "investigation",
"actor": "Carol Davis",
"description": "Identified recent deployment of user-api v2.4.1 at 13:45 UTC. New ORM version (3.2.0) changed connection handling behavior.",
"metadata": {"deployment": "user-api-v2.4.1", "deployed_at": "2024-01-15T13:45:00Z"}
},
{
"timestamp": "2024-01-15T14:30:00Z",
"type": "communication",
"actor": "Bob Kim",
"description": "Status page updated: Investigating - We are investigating increased error rates affecting payment processing.",
"metadata": {"channel": "status_page", "status": "investigating"}
},
{
"timestamp": "2024-01-15T14:35:00Z",
"type": "escalation",
"actor": "Jane Smith",
"description": "Escalated to VP Engineering. Customer impact confirmed: 12,500+ users affected, failed transactions accumulating.",
"metadata": {"escalated_to": "VP Engineering", "reason": "revenue_impact"}
},
{
"timestamp": "2024-01-15T14:40:00Z",
"type": "mitigation",
"actor": "Alice Chen",
"description": "Attempting mitigation: increasing connection pool size from 20 to 50 via config override.",
"metadata": {"action": "pool_resize", "old_value": 20, "new_value": 50}
},
{
"timestamp": "2024-01-15T14:45:00Z",
"type": "communication",
"actor": "Bob Kim",
"description": "Status page updated: Identified - The issue has been identified as a database configuration problem. We are implementing a fix.",
"metadata": {"channel": "status_page", "status": "identified"}
},
{
"timestamp": "2024-01-15T14:50:00Z",
"type": "investigation",
"actor": "Carol Davis",
"description": "Pool resize partially effective. Error rate dropped from 23% to 12%. ORM 3.2.0 opens 3x more connections per request than 3.1.2.",
"metadata": {"error_rate_before": 23.5, "error_rate_after": 12.1}
},
{
"timestamp": "2024-01-15T15:00:00Z",
"type": "mitigation",
"actor": "Alice Chen",
"description": "Decision: roll back ORM version to 3.1.2. Initiating rollback deployment of user-api v2.3.9.",
"metadata": {"action": "rollback", "target_version": "2.3.9", "rollback_reason": "orm_connection_leak"}
},
{
"timestamp": "2024-01-15T15:15:00Z",
"type": "mitigation",
"actor": "Alice Chen",
"description": "Rollback deployment complete. user-api v2.3.9 running in production. Connection pool utilization dropping.",
"metadata": {"deployment_duration_minutes": 15, "pool_utilization": 45}
},
{
"timestamp": "2024-01-15T15:20:00Z",
"type": "communication",
"actor": "Bob Kim",
"description": "Status page updated: Monitoring - A fix has been implemented and we are monitoring the results.",
"metadata": {"channel": "status_page", "status": "monitoring"}
},
{
"timestamp": "2024-01-15T15:30:00Z",
"type": "mitigation",
"actor": "Jane Smith",
"description": "Error rate back to baseline (<0.1%). Payment processing fully restored. Entering monitoring phase.",
"metadata": {"error_rate": 0.08, "pool_utilization": 32}
},
{
"timestamp": "2024-01-15T16:30:00Z",
"type": "investigation",
"actor": "Carol Davis",
"description": "Confirmed stable for 60 minutes. No degradation detected. Root cause documented: ORM 3.2.0 connection pooling incompatibility.",
"metadata": {"monitoring_duration_minutes": 60, "stable": true}
},
{
"timestamp": "2024-01-15T16:45:00Z",
"type": "resolution",
"actor": "Jane Smith",
"description": "Incident resolved. All services nominal. Postmortem scheduled for 2024-01-17 10:00 UTC.",
"metadata": {"postmortem_scheduled": "2024-01-17T10:00:00Z"}
},
{
"timestamp": "2024-01-15T16:50:00Z",
"type": "communication",
"actor": "Bob Kim",
"description": "Status page updated: Resolved - The issue has been resolved. Payment processing is operating normally.",
"metadata": {"channel": "status_page", "status": "resolved"}
}
],
"communications": [
{
"timestamp": "2024-01-15T14:30:00Z",
"channel": "status_page",
"audience": "external",
"message": "Investigating - We are investigating increased error rates affecting payment processing. Some transactions may fail. We will provide an update within 15 minutes."
},
{
"timestamp": "2024-01-15T14:35:00Z",
"channel": "slack_exec",
"audience": "internal",
"message": "SEV1 ACTIVE: Payment service degradation. ~12,500 users affected. Failed transactions accumulating. IC: Jane Smith. Bridge: [link]. ETA for mitigation: investigating."
},
{
"timestamp": "2024-01-15T14:45:00Z",
"channel": "status_page",
"audience": "external",
"message": "Identified - The issue has been identified as a database configuration problem following a recent deployment. We are implementing a fix. Next update in 15 minutes."
},
{
"timestamp": "2024-01-15T15:20:00Z",
"channel": "status_page",
"audience": "external",
"message": "Monitoring - A fix has been implemented and we are monitoring the results. Payment processing is recovering. We will provide a final update once we confirm stability."
},
{
"timestamp": "2024-01-15T16:50:00Z",
"channel": "status_page",
"audience": "external",
"message": "Resolved - The issue affecting payment processing has been resolved. All systems are operating normally. We will publish a full incident report within 48 hours."
}
],
"impact": {
"revenue_impact": "high",
"affected_users_percentage": 45,
"affected_regions": ["us-east-1", "eu-west-1"],
"data_integrity_risk": false,
"security_breach": false,
"customer_facing": true,
"degradation_type": "partial",
"workaround_available": false
},
"signals": {
"error_rate_percentage": 23.5,
"latency_p99_ms": 8500,
"affected_endpoints": ["/api/payments", "/api/checkout", "/api/subscriptions"],
"dependent_services": ["checkout", "subscription-billing", "order-service"],
"alert_count": 12,
"customer_reports": 8
},
"context": {
"recent_deployments": [
{
"service": "user-api",
"deployed_at": "2024-01-15T13:45:00Z",
"version": "2.4.1",
"changes": "Upgraded ORM from 3.1.2 to 3.2.0"
}
],
"ongoing_incidents": [],
"maintenance_windows": [],
"on_call": {
"primary": "alice@company.com",
"secondary": "bob@company.com",
"escalation_manager": "director-eng@company.com"
}
},
"resolution": {
"root_cause": "Database connection pool exhaustion caused by ORM 3.2.0 opening 3x more connections per request than previous version 3.1.2, exceeding the pool size of 20",
"contributing_factors": [
"Insufficient load testing of new ORM version under production-scale connection patterns",
"Connection pool monitoring alert threshold set too high (90%) with no warning at 70%",
"No canary deployment process for database configuration or ORM changes",
"Missing connection pool sizing documentation for service dependencies"
],
"mitigation_steps": [
"Increased connection pool size from 20 to 50 as temporary relief",
"Rolled back user-api from v2.4.1 (ORM 3.2.0) to v2.3.9 (ORM 3.1.2)"
],
"permanent_fix": "Load test ORM 3.2.0 with production connection patterns, update pool sizing, implement canary deployment for ORM changes",
"customer_impact": {
"affected_users": 12500,
"failed_transactions": 342,
"revenue_impact_usd": 28500,
"data_loss": false
}
},
"action_items": [
{
"title": "Add connection pool utilization alerting at 70% warning and 85% critical thresholds",
"owner": "alice@company.com",
"priority": "P1",
"deadline": "2024-01-22",
"type": "detection",
"status": "open"
},
{
"title": "Implement canary deployment pipeline for database configuration and ORM changes",
"owner": "bob@company.com",
"priority": "P1",
"deadline": "2024-02-01",
"type": "prevention",
"status": "open"
},
{
"title": "Load test ORM v3.2.0 with production-scale connection patterns before re-deployment",
"owner": "carol@company.com",
"priority": "P2",
"deadline": "2024-01-29",
"type": "prevention",
"status": "open"
},
{
"title": "Document connection pool sizing requirements for all services in runbook",
"owner": "alice@company.com",
"priority": "P2",
"deadline": "2024-02-05",
"type": "process",
"status": "open"
},
{
"title": "Add ORM connection behavior to integration test suite",
"owner": "carol@company.com",
"priority": "P3",
"deadline": "2024-02-15",
"type": "prevention",
"status": "open"
}
],
"participants": [
{"name": "Jane Smith", "role": "Incident Commander"},
{"name": "Alice Chen", "role": "Operations Lead"},
{"name": "Bob Kim", "role": "Communications Lead"},
{"name": "Carol Davis", "role": "Database SME"}
]
}
FILE:assets/sample_incident_pir_data.json
{
"incident_id": "INC-2024-0315-001",
"title": "Payment API Database Connection Pool Exhaustion",
"description": "Database connection pool exhaustion caused widespread 500 errors in payment processing API, preventing users from completing purchases. Root cause was an inefficient database query introduced in deployment v2.3.1.",
"severity": "sev2",
"start_time": "2024-03-15T14:30:00Z",
"end_time": "2024-03-15T15:35:00Z",
"duration": "1h 5m",
"affected_services": ["payment-api", "checkout-service", "subscription-billing"],
"customer_impact": "80% of users unable to complete payments or checkout. Approximately 2,400 failed payment attempts during the incident. Users experienced immediate 500 errors when attempting to pay.",
"business_impact": "Estimated revenue loss of $45,000 during outage period. No SLA breaches as resolution was within 2-hour window. 12 customer escalations through support channels.",
"incident_commander": "Mike Rodriguez",
"responders": [
"Sarah Chen - On-call Engineer, Primary Responder",
"Tom Wilson - Database Team Lead",
"Lisa Park - Database Engineer",
"Mike Rodriguez - Incident Commander",
"David Kumar - DevOps Engineer"
],
"status": "resolved",
"detection_details": {
"detection_method": "automated_monitoring",
"detection_time": "2024-03-15T14:30:00Z",
"alert_source": "Datadog error rate threshold",
"time_to_detection": "immediate"
},
"response_details": {
"time_to_response": "5 minutes",
"time_to_escalation": "10 minutes",
"time_to_resolution": "65 minutes",
"war_room_established": "2024-03-15T14:45:00Z",
"executives_notified": false,
"status_page_updated": true
},
"technical_details": {
"root_cause": "Inefficient database query introduced in deployment v2.3.1 caused each payment validation to take 15 seconds instead of normal 0.1 seconds, exhausting the 200-connection database pool",
"affected_regions": ["us-west", "us-east", "eu-west"],
"error_metrics": {
"peak_error_rate": "45%",
"normal_error_rate": "0.1%",
"connection_pool_max": 200,
"connections_exhausted_at": "100%"
},
"resolution_method": "rollback",
"rollback_target": "v2.2.9",
"rollback_duration": "7 minutes"
},
"communication_log": [
{
"timestamp": "2024-03-15T14:50:00Z",
"type": "status_page",
"message": "Investigating payment processing issues",
"audience": "customers"
},
{
"timestamp": "2024-03-15T15:35:00Z",
"type": "status_page",
"message": "Payment processing issues resolved",
"audience": "customers"
}
],
"lessons_learned_preview": [
"Deployment v2.3.1 code review missed performance implications of query change",
"Load testing didn't include realistic database query patterns",
"Connection pool monitoring could have provided earlier warning",
"Rollback procedure worked effectively - 7 minute rollback time"
],
"preliminary_action_items": [
"Fix inefficient query for v2.3.2 deployment",
"Add database query performance checks to CI pipeline",
"Improve load testing to include database performance scenarios",
"Add connection pool utilization alerts"
]
}
FILE:assets/sample_timeline_events.json
[
{
"timestamp": "2024-03-15T14:30:00Z",
"source": "datadog",
"type": "alert",
"message": "High error rate detected on payment-api: 45% error rate (threshold: 5%)",
"severity": "critical",
"actor": "monitoring-system",
"metadata": {
"alert_id": "ALT-001",
"metric_value": "45%",
"threshold": "5%"
}
},
{
"timestamp": "2024-03-15T14:32:00Z",
"source": "pagerduty",
"type": "escalation",
"message": "Paged on-call engineer Sarah Chen for payment-api alerts",
"severity": "high",
"actor": "pagerduty-system",
"metadata": {
"incident_id": "PD-12345",
"responder": "sarah.chen@company.com"
}
},
{
"timestamp": "2024-03-15T14:35:00Z",
"source": "slack",
"type": "communication",
"message": "Sarah Chen acknowledged the alert and is investigating payment-api issues",
"severity": "medium",
"actor": "sarah.chen",
"metadata": {
"channel": "#incidents",
"message_id": "1234567890.123456"
}
},
{
"timestamp": "2024-03-15T14:38:00Z",
"source": "application_logs",
"type": "log",
"message": "Database connection pool exhausted: 200/200 connections active, unable to acquire new connections",
"severity": "critical",
"actor": "payment-api",
"metadata": {
"log_level": "ERROR",
"component": "database_pool",
"connection_count": 200,
"max_connections": 200
}
},
{
"timestamp": "2024-03-15T14:40:00Z",
"source": "slack",
"type": "escalation",
"message": "Sarah Chen: Escalating to incident commander - database connection pool exhausted, need database team",
"severity": "high",
"actor": "sarah.chen",
"metadata": {
"channel": "#incidents",
"escalation_reason": "database_expertise_needed"
}
},
{
"timestamp": "2024-03-15T14:42:00Z",
"source": "pagerduty",
"type": "escalation",
"message": "Incident commander Mike Rodriguez assigned to incident PD-12345",
"severity": "high",
"actor": "pagerduty-system",
"metadata": {
"incident_commander": "mike.rodriguez@company.com",
"role": "incident_commander"
}
},
{
"timestamp": "2024-03-15T14:45:00Z",
"source": "slack",
"type": "communication",
"message": "Mike Rodriguez: War room established in #war-room-payment-api. Engaging database team.",
"severity": "high",
"actor": "mike.rodriguez",
"metadata": {
"channel": "#incidents",
"war_room": "#war-room-payment-api"
}
},
{
"timestamp": "2024-03-15T14:47:00Z",
"source": "pagerduty",
"type": "escalation",
"message": "Database team engineers paged: Tom Wilson, Lisa Park",
"severity": "medium",
"actor": "pagerduty-system",
"metadata": {
"team": "database-team",
"responders": ["tom.wilson@company.com", "lisa.park@company.com"]
}
},
{
"timestamp": "2024-03-15T14:50:00Z",
"source": "statuspage",
"type": "communication",
"message": "Status page updated: Investigating payment processing issues",
"severity": "medium",
"actor": "mike.rodriguez",
"metadata": {
"status": "investigating",
"affected_systems": ["payment-api"]
}
},
{
"timestamp": "2024-03-15T14:52:00Z",
"source": "slack",
"type": "communication",
"message": "Tom Wilson: Joining war room. Looking at database metrics now. Seeing unusual query patterns from recent deployment.",
"severity": "medium",
"actor": "tom.wilson",
"metadata": {
"channel": "#war-room-payment-api",
"investigation_focus": "database_metrics"
}
},
{
"timestamp": "2024-03-15T14:55:00Z",
"source": "database_monitoring",
"type": "log",
"message": "Identified slow query introduced in deployment v2.3.1: payment validation taking 15s per request",
"severity": "critical",
"actor": "database-monitor",
"metadata": {
"deployment_version": "v2.3.1",
"query_time": "15s",
"normal_query_time": "0.1s"
}
},
{
"timestamp": "2024-03-15T15:00:00Z",
"source": "slack",
"type": "communication",
"message": "Tom Wilson: Root cause identified - inefficient query in v2.3.1 deployment. Recommending immediate rollback.",
"severity": "high",
"actor": "tom.wilson",
"metadata": {
"channel": "#war-room-payment-api",
"root_cause": "inefficient_query",
"recommendation": "rollback"
}
},
{
"timestamp": "2024-03-15T15:02:00Z",
"source": "slack",
"type": "communication",
"message": "Mike Rodriguez: Approved rollback to v2.2.9. Sarah initiating rollback procedure.",
"severity": "high",
"actor": "mike.rodriguez",
"metadata": {
"channel": "#war-room-payment-api",
"decision": "rollback_approved",
"target_version": "v2.2.9"
}
},
{
"timestamp": "2024-03-15T15:05:00Z",
"source": "deployment_system",
"type": "action",
"message": "Rollback initiated: payment-api v2.3.1 → v2.2.9",
"severity": "medium",
"actor": "sarah.chen",
"metadata": {
"from_version": "v2.3.1",
"to_version": "v2.2.9",
"deployment_type": "rollback"
}
},
{
"timestamp": "2024-03-15T15:12:00Z",
"source": "deployment_system",
"type": "action",
"message": "Rollback completed successfully: payment-api now running v2.2.9 across all regions",
"severity": "medium",
"actor": "deployment-system",
"metadata": {
"deployment_status": "completed",
"regions": ["us-west", "us-east", "eu-west"]
}
},
{
"timestamp": "2024-03-15T15:15:00Z",
"source": "datadog",
"type": "log",
"message": "Error rate decreasing: payment-api error rate dropped to 8% and continuing to decline",
"severity": "medium",
"actor": "monitoring-system",
"metadata": {
"error_rate": "8%",
"trend": "decreasing"
}
},
{
"timestamp": "2024-03-15T15:18:00Z",
"source": "database_monitoring",
"type": "log",
"message": "Connection pool utilization normalizing: 45/200 connections active",
"severity": "low",
"actor": "database-monitor",
"metadata": {
"connection_count": 45,
"max_connections": 200,
"utilization": "22.5%"
}
},
{
"timestamp": "2024-03-15T15:25:00Z",
"source": "datadog",
"type": "log",
"message": "Error rate returned to normal: payment-api error rate now 0.2% (within normal range)",
"severity": "low",
"actor": "monitoring-system",
"metadata": {
"error_rate": "0.2%",
"status": "normal"
}
},
{
"timestamp": "2024-03-15T15:30:00Z",
"source": "slack",
"type": "communication",
"message": "Mike Rodriguez: All metrics returned to normal. Declaring incident resolved. Thanks to all responders.",
"severity": "low",
"actor": "mike.rodriguez",
"metadata": {
"channel": "#war-room-payment-api",
"status": "resolved"
}
},
{
"timestamp": "2024-03-15T15:35:00Z",
"source": "statuspage",
"type": "communication",
"message": "Status page updated: Payment processing issues resolved. All systems operational.",
"severity": "low",
"actor": "mike.rodriguez",
"metadata": {
"status": "resolved",
"duration": "65 minutes"
}
},
{
"timestamp": "2024-03-15T15:40:00Z",
"source": "slack",
"type": "communication",
"message": "Mike Rodriguez: PIR scheduled for tomorrow 10am. Action item: fix the inefficient query in v2.3.2",
"severity": "low",
"actor": "mike.rodriguez",
"metadata": {
"channel": "#incidents",
"pir_time": "2024-03-16T10:00:00Z",
"action_item": "fix_query_v2.3.2"
}
}
]
FILE:assets/simple_incident.json
{
"description": "Users reporting slow page loads on the main website",
"service": "web-frontend",
"affected_users": "25%",
"business_impact": "medium"
}
FILE:assets/simple_timeline_events.json
[
{
"timestamp": "2024-03-10T09:00:00Z",
"source": "monitoring",
"message": "High CPU utilization detected on web servers",
"severity": "medium",
"actor": "system"
},
{
"timestamp": "2024-03-10T09:05:00Z",
"source": "slack",
"message": "Engineer investigating high CPU alerts",
"severity": "medium",
"actor": "john.doe"
},
{
"timestamp": "2024-03-10T09:15:00Z",
"source": "deployment",
"message": "Deployed hotfix to reduce CPU usage",
"severity": "low",
"actor": "john.doe"
},
{
"timestamp": "2024-03-10T09:25:00Z",
"source": "monitoring",
"message": "CPU utilization returned to normal levels",
"severity": "low",
"actor": "system"
}
]
FILE:expected_outputs/incident_classification_text_output.txt
============================================================
INCIDENT CLASSIFICATION REPORT
============================================================
CLASSIFICATION:
Severity: SEV1
Confidence: 100.0%
Reasoning: Classified as SEV1 based on: keywords: timeout, 500 error; user impact: 80%
Timestamp: 2026-02-16T12:41:46.644096+00:00
RECOMMENDED RESPONSE:
Primary Team: Analytics Team
Supporting Teams: SRE, API Team, Backend Engineering, Finance Engineering, Payments Team, DevOps, Compliance Team, Database Team, Platform Team, Data Engineering
Response Time: 5 minutes
INITIAL ACTIONS:
1. Establish incident command (Priority 1)
Timeout: 5 minutes
Page incident commander and establish war room
2. Create incident ticket (Priority 1)
Timeout: 2 minutes
Create tracking ticket with all known details
3. Update status page (Priority 2)
Timeout: 15 minutes
Post initial status page update acknowledging incident
4. Notify executives (Priority 2)
Timeout: 15 minutes
Alert executive team of customer-impacting outage
5. Engage subject matter experts (Priority 3)
Timeout: 10 minutes
Page relevant SMEs based on affected systems
COMMUNICATION:
Subject: 🚨 [SEV1] payment-api - Database connection timeouts causing 500 errors fo...
Urgency: SEV1
Recipients: on-call, engineering-leadership, executives, customer-success
Channels: pager, phone, slack, email, status-page
Update Frequency: Every 15 minutes
============================================================
FILE:expected_outputs/pir_markdown_output.md
# Post-Incident Review: Payment API Database Connection Pool Exhaustion
## Executive Summary
On March 15, 2024, we experienced a sev2 incident affecting ['payment-api', 'checkout-service', 'subscription-billing']. The incident lasted 1h 5m and had the following impact: 80% of users unable to complete payments or checkout. Approximately 2,400 failed payment attempts during the incident. Users experienced immediate 500 errors when attempting to pay. The incident has been resolved and we have identified specific actions to prevent recurrence.
## Incident Overview
- **Incident ID:** INC-2024-0315-001
- **Date & Time:** 2024-03-15 14:30:00 UTC
- **Duration:** 1h 5m
- **Severity:** SEV2
- **Status:** Resolved
- **Incident Commander:** Mike Rodriguez
- **Responders:** Sarah Chen - On-call Engineer, Primary Responder, Tom Wilson - Database Team Lead, Lisa Park - Database Engineer, Mike Rodriguez - Incident Commander, David Kumar - DevOps Engineer
### Customer Impact
80% of users unable to complete payments or checkout. Approximately 2,400 failed payment attempts during the incident. Users experienced immediate 500 errors when attempting to pay.
### Business Impact
Estimated revenue loss of $45,000 during outage period. No SLA breaches as resolution was within 2-hour window. 12 customer escalations through support channels.
## Timeline
No detailed timeline available.
## Root Cause Analysis
### Analysis Method: 5 Whys Analysis
#### Why Analysis
**Why 1:** Why did Database connection pool exhaustion caused widespread 500 errors in payment processing API, preventing users from completing purchases. Root cause was an inefficient database query introduced in deployment v2.3.1.?
**Answer:** New deployment introduced a regression
**Why 2:** Why wasn't this detected earlier?
**Answer:** Code review process missed the issue
**Why 3:** Why didn't existing safeguards prevent this?
**Answer:** Testing environment didn't match production
**Why 4:** Why wasn't there a backup mechanism?
**Answer:** Further investigation needed
**Why 5:** Why wasn't this scenario anticipated?
**Answer:** Further investigation needed
## What Went Well
- The incident was successfully resolved
- Incident command was established
- Multiple team members collaborated on resolution
## What Didn't Go Well
- Analysis in progress
## Lessons Learned
Lessons learned to be documented following detailed analysis.
## Action Items
Action items to be defined.
## Follow-up and Prevention
### Prevention Measures
Based on the root cause analysis, the following preventive measures have been identified:
- Implement comprehensive testing for similar scenarios
- Improve monitoring and alerting coverage
- Enhance error handling and resilience patterns
### Follow-up Schedule
- 1 week: Review action item progress
- 1 month: Evaluate effectiveness of implemented changes
- 3 months: Conduct follow-up assessment and update preventive measures
## Appendix
### Additional Information
- Incident ID: INC-2024-0315-001
- Severity Classification: sev2
- Affected Services: payment-api, checkout-service, subscription-billing
### References
- Incident tracking ticket: [Link TBD]
- Monitoring dashboards: [Link TBD]
- Communication thread: [Link TBD]
---
*Generated on 2026-02-16 by PIR Generator*
FILE:expected_outputs/simple_incident_classification.txt
============================================================
INCIDENT CLASSIFICATION REPORT
============================================================
CLASSIFICATION:
Severity: SEV2
Confidence: 100.0%
Reasoning: Classified as SEV2 based on: keywords: slow; user impact: 25%
Timestamp: 2026-02-16T12:42:41.889774+00:00
RECOMMENDED RESPONSE:
Primary Team: UX Engineering
Supporting Teams: Product Engineering, Frontend Team
Response Time: 15 minutes
INITIAL ACTIONS:
1. Assign incident commander (Priority 1)
Timeout: 30 minutes
Assign IC and establish coordination channel
2. Create incident tracking (Priority 1)
Timeout: 5 minutes
Create incident ticket with details and timeline
3. Assess customer impact (Priority 2)
Timeout: 15 minutes
Determine scope and severity of user impact
4. Engage response team (Priority 2)
Timeout: 30 minutes
Page appropriate technical responders
5. Begin investigation (Priority 3)
Timeout: 15 minutes
Start technical analysis and debugging
COMMUNICATION:
Subject: ⚠️ [SEV2] web-frontend - Users reporting slow page loads on the main websit...
Urgency: SEV2
Recipients: on-call, engineering-leadership, product-team
Channels: pager, slack, email
Update Frequency: Every 30 minutes
============================================================
FILE:expected_outputs/timeline_reconstruction_text_output.txt
================================================================================
INCIDENT TIMELINE RECONSTRUCTION
================================================================================
OVERVIEW:
Time Range: 2024-03-15T14:30:00+00:00 to 2024-03-15T15:40:00+00:00
Total Duration: 70 minutes
Total Events: 21
Phases Detected: 12
PHASES:
DETECTION:
Start: 2024-03-15T14:30:00+00:00
Duration: 0.0 minutes
Events: 1
Description: Initial detection of the incident through monitoring or observation
ESCALATION:
Start: 2024-03-15T14:32:00+00:00
Duration: 0.0 minutes
Events: 1
Description: Escalation to additional resources or higher severity response
TRIAGE:
Start: 2024-03-15T14:35:00+00:00
Duration: 0.0 minutes
Events: 1
Description: Assessment and initial investigation of the incident
ESCALATION:
Start: 2024-03-15T14:38:00+00:00
Duration: 9.0 minutes
Events: 5
Description: Escalation to additional resources or higher severity response
TRIAGE:
Start: 2024-03-15T14:50:00+00:00
Duration: 0.0 minutes
Events: 1
Description: Assessment and initial investigation of the incident
ESCALATION:
Start: 2024-03-15T14:52:00+00:00
Duration: 10.0 minutes
Events: 4
Description: Escalation to additional resources or higher severity response
TRIAGE:
Start: 2024-03-15T15:05:00+00:00
Duration: 7.0 minutes
Events: 2
Description: Assessment and initial investigation of the incident
DETECTION:
Start: 2024-03-15T15:15:00+00:00
Duration: 0.0 minutes
Events: 1
Description: Initial detection of the incident through monitoring or observation
RESOLUTION:
Start: 2024-03-15T15:18:00+00:00
Duration: 0.0 minutes
Events: 1
Description: Confirmation that the incident has been resolved
DETECTION:
Start: 2024-03-15T15:25:00+00:00
Duration: 0.0 minutes
Events: 1
Description: Initial detection of the incident through monitoring or observation
RESOLUTION:
Start: 2024-03-15T15:30:00+00:00
Duration: 5.0 minutes
Events: 2
Description: Confirmation that the incident has been resolved
TRIAGE:
Start: 2024-03-15T15:40:00+00:00
Duration: 0.0 minutes
Events: 1
Description: Assessment and initial investigation of the incident
KEY METRICS:
Time to Mitigation: 0 minutes
Time to Resolution: 48.0 minutes
Events per Hour: 18.0
Unique Sources: 7
INCIDENT NARRATIVE:
Incident Timeline Summary:
The incident began at 2024-03-15 14:30:00 UTC and concluded at 2024-03-15 15:40:00 UTC, lasting approximately 70 minutes.
The incident progressed through 12 distinct phases: detection, escalation, triage, escalation, triage, escalation, triage, detection, resolution, detection, resolution, triage.
Key milestones:
- Detection: 14:30 (0 min)
- Escalation: 14:32 (0 min)
- Triage: 14:35 (0 min)
- Escalation: 14:38 (9 min)
- Triage: 14:50 (0 min)
- Escalation: 14:52 (10 min)
- Triage: 15:05 (7 min)
- Detection: 15:15 (0 min)
- Resolution: 15:18 (0 min)
- Detection: 15:25 (0 min)
- Resolution: 15:30 (5 min)
- Triage: 15:40 (0 min)
================================================================================
FILE:README.md
# Incident Commander Skill
A comprehensive incident response framework providing structured tools for managing technology incidents from detection through resolution and post-incident review.
## Overview
This skill implements battle-tested practices from SRE and DevOps teams at scale, providing:
- **Automated Severity Classification** - Intelligent incident triage
- **Timeline Reconstruction** - Transform scattered events into coherent narratives
- **Post-Incident Review Generation** - Structured PIRs with RCA frameworks
- **Communication Templates** - Pre-built stakeholder communication
- **Comprehensive Documentation** - Reference guides for incident response
## Quick Start
### Classify an Incident
```bash
# From JSON file
python scripts/incident_classifier.py --input incident.json --format text
# From stdin text
echo "Database is down affecting all users" | python scripts/incident_classifier.py --format text
# Interactive mode
python scripts/incident_classifier.py --interactive
```
### Reconstruct Timeline
```bash
# Analyze event timeline
python scripts/timeline_reconstructor.py --input events.json --format text
# With gap analysis
python scripts/timeline_reconstructor.py --input events.json --gap-analysis --format markdown
```
### Generate PIR Document
```bash
# Basic PIR
python scripts/pir_generator.py --incident incident.json --format markdown
# Comprehensive PIR with timeline
python scripts/pir_generator.py --incident incident.json --timeline timeline.json --rca-method fishbone
```
## Scripts
### incident_classifier.py
**Purpose:** Analyzes incident descriptions and provides severity classification, team recommendations, and response templates.
**Input:** JSON object with incident details or plain text description
**Output:** JSON + human-readable classification report
**Example Input:**
```json
{
"description": "Database connection timeouts causing 500 errors",
"service": "payment-api",
"affected_users": "80%",
"business_impact": "high"
}
```
**Key Features:**
- SEV1-4 severity classification
- Recommended response teams
- Initial action prioritization
- Communication templates
- Response timelines
### timeline_reconstructor.py
**Purpose:** Reconstructs incident timelines from timestamped events, identifies phases, and performs gap analysis.
**Input:** JSON array of timestamped events
**Output:** Formatted timeline with phase analysis and metrics
**Example Input:**
```json
[
{
"timestamp": "2024-01-01T12:00:00Z",
"source": "monitoring",
"message": "High error rate detected",
"severity": "critical",
"actor": "system"
}
]
```
**Key Features:**
- Phase detection (detection → triage → mitigation → resolution)
- Duration analysis
- Gap identification
- Communication effectiveness analysis
- Response metrics
### pir_generator.py
**Purpose:** Generates comprehensive Post-Incident Review documents with multiple RCA frameworks.
**Input:** Incident data JSON, optional timeline data
**Output:** Structured PIR document with RCA analysis
**Key Features:**
- Multiple RCA methods (5 Whys, Fishbone, Timeline, Bow Tie)
- Automated action item generation
- Lessons learned categorization
- Follow-up planning
- Completeness assessment
## Sample Data
The `assets/` directory contains sample data files for testing:
- `sample_incident_classification.json` - Database connection pool exhaustion incident
- `sample_timeline_events.json` - Complete timeline with 21 events across phases
- `sample_incident_pir_data.json` - Comprehensive incident data for PIR generation
- `simple_incident.json` - Minimal incident for basic testing
- `simple_timeline_events.json` - Simple 4-event timeline
## Expected Outputs
The `expected_outputs/` directory contains reference outputs showing what each script produces:
- `incident_classification_text_output.txt` - Detailed classification report
- `timeline_reconstruction_text_output.txt` - Complete timeline analysis
- `pir_markdown_output.md` - Full PIR document
- `simple_incident_classification.txt` - Basic classification example
## Reference Documentation
### references/incident_severity_matrix.md
Complete severity classification system with:
- SEV1-4 definitions and criteria
- Response requirements and timelines
- Escalation paths
- Communication requirements
- Decision trees and examples
### references/rca_frameworks_guide.md
Detailed guide for root cause analysis:
- 5 Whys methodology
- Fishbone (Ishikawa) diagram analysis
- Timeline analysis techniques
- Bow Tie analysis for high-risk incidents
- Framework selection guidelines
### references/communication_templates.md
Standardized communication templates:
- Severity-specific notification templates
- Stakeholder-specific messaging
- Escalation communications
- Resolution notifications
- Customer communication guidelines
## Usage Patterns
### End-to-End Incident Workflow
1. **Initial Classification**
```bash
echo "Payment API returning 500 errors for 70% of requests" | \
python scripts/incident_classifier.py --format text
```
2. **Timeline Reconstruction** (after collecting events)
```bash
python scripts/timeline_reconstructor.py \
--input events.json \
--gap-analysis \
--format markdown \
--output timeline.md
```
3. **PIR Generation** (after incident resolution)
```bash
python scripts/pir_generator.py \
--incident incident.json \
--timeline timeline.md \
--rca-method fishbone \
--output pir.md
```
### Integration Examples
**CI/CD Pipeline Integration:**
```bash
# Classify deployment issues
cat deployment_error.log | python scripts/incident_classifier.py --format json
```
**Monitoring Integration:**
```bash
# Process alert events
curl -s "monitoring-api/events" | python scripts/timeline_reconstructor.py --format text
```
**Runbook Generation:**
Use classification output to automatically select appropriate runbooks and escalation procedures.
## Quality Standards
- **Zero External Dependencies** - All scripts use only Python standard library
- **Dual Output Format** - Both JSON (machine-readable) and text (human-readable)
- **Robust Input Handling** - Graceful handling of missing or malformed data
- **Professional Defaults** - Opinionated, battle-tested configurations
- **Comprehensive Testing** - Sample data and expected outputs included
## Technical Requirements
- Python 3.6+
- No external dependencies required
- Works with standard Unix tools (pipes, redirection)
- Cross-platform compatible
## Severity Classification Reference
| Severity | Description | Response Time | Update Frequency |
|----------|-------------|---------------|------------------|
| **SEV1** | Complete outage | 5 minutes | Every 15 minutes |
| **SEV2** | Major degradation | 15 minutes | Every 30 minutes |
| **SEV3** | Minor impact | 2 hours | At milestones |
| **SEV4** | Low impact | 1-2 days | Weekly |
## Getting Help
Each script includes comprehensive help:
```bash
python scripts/incident_classifier.py --help
python scripts/timeline_reconstructor.py --help
python scripts/pir_generator.py --help
```
For methodology questions, refer to the reference documentation in the `references/` directory.
## Contributing
When adding new features:
1. Maintain zero external dependencies
2. Add comprehensive examples to `assets/`
3. Update expected outputs in `expected_outputs/`
4. Follow the established patterns for argument parsing and output formatting
## License
This skill is part of the claude-skills repository. See the main repository LICENSE for details.
FILE:references/communication_templates.md
# Incident Communication Templates
## Overview
This document provides standardized communication templates for incident response. These templates ensure consistent, clear communication across different severity levels and stakeholder groups.
## Template Usage Guidelines
### General Principles
1. **Be Clear and Concise** - Use simple language, avoid jargon
2. **Be Factual** - Only state what is known, avoid speculation
3. **Be Timely** - Send updates at committed intervals
4. **Be Actionable** - Include next steps and expected timelines
5. **Be Accountable** - Include contact information for follow-up
### Template Selection
- Choose templates based on incident severity and audience
- Customize templates with specific incident details
- Always include next update time and contact information
- Escalate template types as severity increases
---
## SEV1 Templates
### Initial Alert - Internal Teams
**Subject:** 🚨 [SEV1] CRITICAL: {Service} Complete Outage - Immediate Response Required
```
CRITICAL INCIDENT ALERT - IMMEDIATE ATTENTION REQUIRED
Incident Summary:
- Service: {Service Name}
- Status: Complete Outage
- Start Time: {Timestamp}
- Customer Impact: {Impact Description}
- Estimated Affected Users: {Number/Percentage}
Immediate Actions Needed:
✓ Incident Commander: {Name} - ASSIGNED
✓ War Room: {Bridge/Chat Link} - JOIN NOW
✓ On-Call Response: {Team} - PAGED
⏳ Executive Notification: In progress
⏳ Status Page Update: Within 15 minutes
Current Situation:
{Brief description of what we know}
What We're Doing:
{Immediate response actions being taken}
Next Update: {Timestamp - 15 minutes from now}
Incident Commander: {Name}
Contact: {Phone/Slack}
THIS IS A CUSTOMER-IMPACTING INCIDENT REQUIRING IMMEDIATE ATTENTION
```
### Executive Notification - SEV1
**Subject:** 🚨 URGENT: Customer-Impacting Outage - {Service}
```
EXECUTIVE ALERT: Critical customer-facing incident
Service: {Service Name}
Impact: {Customer impact description}
Duration: {Current duration} (started {start time})
Business Impact: {Revenue/SLA/compliance implications}
Customer Impact Summary:
- Affected Users: {Number/percentage}
- Revenue Impact: {$ amount if known}
- SLA Status: {Breach status}
- Customer Escalations: {Number if any}
Response Status:
- Incident Commander: {Name} ({contact})
- Response Team Size: {Number of engineers}
- Root Cause: {If known, otherwise "Under investigation"}
- ETA to Resolution: {If known, otherwise "Investigating"}
Executive Actions Required:
- [ ] Customer communication approval needed
- [ ] Legal/compliance notification: {If applicable}
- [ ] PR/Media response preparation: {If needed}
- [ ] Resource allocation decisions: {If escalation needed}
War Room: {Link}
Next Update: {15 minutes from now}
This incident meets SEV1 criteria and requires executive oversight.
{Incident Commander contact information}
```
### Customer Communication - SEV1
**Subject:** Service Disruption - Immediate Action Being Taken
```
We are currently experiencing a service disruption affecting {service description}.
What's Happening:
{Clear, customer-friendly description of the issue}
Impact:
{What customers are experiencing - be specific}
What We're Doing:
We detected this issue at {time} and immediately mobilized our engineering team. We are actively working to resolve this issue and will provide updates every 15 minutes.
Current Actions:
• {Action 1 - customer-friendly description}
• {Action 2 - customer-friendly description}
• {Action 3 - customer-friendly description}
Workaround:
{If available, provide clear steps}
{If not available: "We are working on alternative solutions and will share them as soon as available."}
Next Update: {Timestamp}
Status Page: {Link}
Support: {Contact information if different from usual}
We sincerely apologize for the inconvenience and are committed to resolving this as quickly as possible.
{Company Name} Team
```
### Status Page Update - SEV1
**Status:** Major Outage
```
{Timestamp} - Investigating
We are currently investigating reports of {service} being unavailable. Our team has been alerted and is actively investigating the cause.
Affected Services: {List of affected services}
Impact: {Customer-facing impact description}
We will provide an update within 15 minutes.
```
```
{Timestamp} - Identified
We have identified the cause of the {service} outage. Our engineering team is implementing a fix.
Root Cause: {Brief, customer-friendly explanation}
Expected Resolution: {Timeline if known}
Next update in 15 minutes.
```
```
{Timestamp} - Monitoring
The fix has been implemented and we are monitoring the service recovery.
Current Status: {Recovery progress}
Next Steps: {What we're monitoring}
We expect full service restoration within {timeframe}.
```
```
{Timestamp} - Resolved
{Service} is now fully operational. We have confirmed that all functionality is working as expected.
Total Duration: {Duration}
Root Cause: {Brief summary}
We apologize for the inconvenience. A full post-incident review will be conducted and shared within 24 hours.
```
---
## SEV2 Templates
### Team Notification - SEV2
**Subject:** ⚠️ [SEV2] {Service} Performance Issues - Response Team Mobilizing
```
SEV2 INCIDENT: Performance degradation requiring active response
Incident Details:
- Service: {Service Name}
- Issue: {Description of performance issue}
- Start Time: {Timestamp}
- Affected Users: {Percentage/description}
- Business Impact: {Impact on business operations}
Current Status:
{What we know about the issue}
Response Team:
- Incident Commander: {Name} ({contact})
- Primary Responder: {Name} ({team})
- Supporting Teams: {List of engaged teams}
Immediate Actions:
✓ {Action 1 - completed}
⏳ {Action 2 - in progress}
⏳ {Action 3 - next step}
Metrics:
- Error Rate: {Current vs normal}
- Response Time: {Current vs normal}
- Throughput: {Current vs normal}
Communication Plan:
- Internal Updates: Every 30 minutes
- Stakeholder Notification: {If needed}
- Status Page Update: {Planned/not needed}
Coordination Channel: {Slack channel}
Next Update: {30 minutes from now}
Incident Commander: {Name} | {Contact}
```
### Stakeholder Update - SEV2
**Subject:** [SEV2] Service Performance Update - {Service}
```
Service Performance Incident Update
Service: {Service Name}
Duration: {Current duration}
Impact: {Description of user impact}
Current Status:
{Brief status of the incident and response efforts}
What We Know:
• {Key finding 1}
• {Key finding 2}
• {Key finding 3}
What We're Doing:
• {Response action 1}
• {Response action 2}
• {Monitoring/verification steps}
Customer Impact:
{Realistic assessment of what users are experiencing}
Workaround:
{If available, provide steps}
Expected Resolution:
{Timeline if known, otherwise "Continuing investigation"}
Next Update: {30 minutes}
Contact: {Incident Commander information}
This incident is being actively managed and does not currently require escalation.
```
### Customer Communication - SEV2 (Optional)
**Subject:** Temporary Service Performance Issues
```
We are currently experiencing performance issues with {service name} that may affect your experience.
What You Might Notice:
{Specific symptoms users might experience}
What We're Doing:
Our team identified this issue at {time} and is actively working on a resolution. We expect to have this resolved within {timeframe}.
Workaround:
{If applicable, provide simple workaround steps}
We will update our status page at {link} with progress information.
Thank you for your patience as we work to resolve this issue quickly.
{Company Name} Support Team
```
---
## SEV3 Templates
### Team Assignment - SEV3
**Subject:** [SEV3] Issue Assignment - {Component} Issue
```
SEV3 Issue Assignment
Service/Component: {Affected component}
Issue: {Description}
Reported: {Timestamp}
Reporter: {Person/system that reported}
Issue Details:
{Detailed description of the problem}
Impact Assessment:
- Affected Users: {Scope}
- Business Impact: {Assessment}
- Urgency: {Business hours response appropriate}
Assignment:
- Primary: {Engineer name}
- Team: {Responsible team}
- Expected Response: {Within 2-4 hours}
Investigation Plan:
1. {Investigation step 1}
2. {Investigation step 2}
3. {Communication checkpoint}
Workaround:
{If known, otherwise "Investigating alternatives"}
This issue will be tracked in {ticket system} as {ticket number}.
Team Lead: {Name} | {Contact}
```
### Status Update - SEV3
**Subject:** [SEV3] Progress Update - {Component}
```
SEV3 Issue Progress Update
Issue: {Brief description}
Assigned to: {Engineer/Team}
Investigation Status: {Current progress}
Findings So Far:
{What has been discovered during investigation}
Next Steps:
{Planned actions and timeline}
Impact Update:
{Any changes to scope or urgency}
Expected Resolution:
{Timeline if known}
This issue continues to be tracked as SEV3 with no escalation required.
Contact: {Assigned engineer} | {Team lead}
```
---
## SEV4 Templates
### Issue Documentation - SEV4
**Subject:** [SEV4] Issue Documented - {Description}
```
SEV4 Issue Logged
Description: {Clear description of the issue}
Reporter: {Name/system}
Date: {Date reported}
Impact:
{Minimal impact description}
Priority Assessment:
This issue has been classified as SEV4 and will be addressed in the normal development cycle.
Assignment:
- Team: {Responsible team}
- Sprint: {Target sprint}
- Estimated Effort: {Story points/hours}
This issue is tracked as {ticket number} in {system}.
Product Owner: {Name}
```
---
## Escalation Templates
### Severity Escalation
**Subject:** ESCALATION: {Original Severity} → {New Severity} - {Service}
```
SEVERITY ESCALATION NOTIFICATION
Original Classification: {Original severity}
New Classification: {New severity}
Escalation Time: {Timestamp}
Escalated By: {Name and role}
Escalation Reasons:
• {Reason 1 - scope expansion/duration/impact}
• {Reason 2}
• {Reason 3}
Updated Impact:
{New assessment of customer/business impact}
Updated Response Requirements:
{New response team, communication frequency, etc.}
Previous Response Actions:
{Summary of actions taken under previous severity}
New Incident Commander: {If changed}
Updated Communication Plan: {New frequency/recipients}
All stakeholders should adjust response according to {new severity} protocols.
Incident Commander: {Name} | {Contact}
```
### Management Escalation
**Subject:** MANAGEMENT ESCALATION: Extended {Severity} Incident - {Service}
```
Management Escalation Required
Incident: {Service} {brief description}
Original Severity: {Severity}
Duration: {Current duration}
Escalation Trigger: {Duration threshold/scope change/customer escalation}
Current Status:
{Brief status of incident response}
Challenges Encountered:
• {Challenge 1}
• {Challenge 2}
• {Resource/expertise needs}
Business Impact:
{Updated assessment of business implications}
Management Decision Required:
• {Decision 1 - resource allocation/external expertise/communication}
• {Decision 2}
Recommended Actions:
{Incident Commander's recommendations}
This escalation follows standard procedures for {trigger type}.
Incident Commander: {Name}
Contact: {Phone/Slack}
War Room: {Link}
```
---
## Resolution Templates
### Resolution Confirmation - All Severities
**Subject:** RESOLVED: [{Severity}] {Service} Incident - {Brief Description}
```
INCIDENT RESOLVED
Service: {Service Name}
Issue: {Brief description}
Duration: {Total duration}
Resolution Time: {Timestamp}
Resolution Summary:
{Brief description of how the issue was resolved}
Root Cause:
{Brief explanation - detailed PIR to follow}
Impact Summary:
- Users Affected: {Final count/percentage}
- Business Impact: {Final assessment}
- Services Affected: {List}
Resolution Actions Taken:
• {Action 1}
• {Action 2}
• {Verification steps}
Monitoring:
We will continue monitoring {service} for {duration} to ensure stability.
Next Steps:
• Post-incident review scheduled for {date}
• Action items to be tracked in {system}
• Follow-up communication: {If needed}
Thank you to everyone who participated in the incident response.
Incident Commander: {Name}
```
### Customer Resolution Communication
**Subject:** Service Restored - Thank You for Your Patience
```
Service Update: Issue Resolved
We're pleased to report that the {service} issues have been fully resolved as of {timestamp}.
What Was Fixed:
{Customer-friendly explanation of the resolution}
Duration:
The issue lasted {duration} from {start time} to {end time}.
What We Learned:
{Brief, high-level takeaway}
Our Commitment:
We are conducting a thorough review of this incident and will implement improvements to prevent similar issues in the future. A summary of our findings and improvements will be shared {timeframe}.
We sincerely apologize for any inconvenience this may have caused and appreciate your patience while we worked to resolve the issue.
If you continue to experience any problems, please contact our support team at {contact information}.
Thank you,
{Company Name} Team
```
---
## Template Customization Guidelines
### Placeholders to Always Replace
- `{Service}` / `{Service Name}` - Specific service or component
- `{Timestamp}` - Specific date/time in consistent format
- `{Name}` / `{Contact}` - Actual names and contact information
- `{Duration}` - Actual time durations
- `{Link}` - Real URLs to war rooms, status pages, etc.
### Language Guidelines
- Use active voice ("We are investigating" not "The issue is being investigated")
- Be specific about timelines ("within 30 minutes" not "soon")
- Avoid technical jargon in customer communications
- Include empathy in customer-facing messages
- Use consistent terminology throughout incident lifecycle
### Timing Guidelines
| Severity | Initial Notification | Update Frequency | Resolution Notification |
|----------|---------------------|------------------|------------------------|
| SEV1 | Immediate (< 5 min) | Every 15 minutes | Immediate |
| SEV2 | Within 15 minutes | Every 30 minutes | Within 15 minutes |
| SEV3 | Within 2 hours | At milestones | Within 1 hour |
| SEV4 | Within 1 business day | Weekly | When resolved |
### Audience-Specific Considerations
#### Engineering Teams
- Include technical details
- Provide specific metrics and logs
- Include coordination channels
- List specific actions and owners
#### Executive/Business
- Focus on business impact
- Include customer and revenue implications
- Provide clear timeline and resource needs
- Highlight any external factors (PR, legal, compliance)
#### Customers
- Use plain language
- Focus on customer impact and workarounds
- Provide realistic timelines
- Include support contact information
- Show empathy and accountability
---
**Last Updated:** February 2026
**Next Review:** May 2026
**Owner:** Incident Management Team
FILE:references/incident-response-framework.md
# Incident Response Framework Reference
Production-grade incident management knowledge base synthesizing PagerDuty, Google SRE, and Atlassian methodologies into a unified, opinionated framework. This document is the source of truth for incident commanders operating under pressure.
---
## 1. Industry Framework Comparison
### PagerDuty Incident Response Model
PagerDuty's open-source incident response process defines four core roles and five process phases. The model prioritizes **speed of mobilization** over process perfection.
**Roles:**
- **Incident Commander (IC):** Owns the incident end-to-end. Does NOT perform technical investigation. Delegates, coordinates, and makes final escalation decisions. The IC is the single point of authority; conflicting opinions are resolved by the IC, not by committee.
- **Scribe:** Captures timestamped decisions, actions, and findings in the incident channel. The scribe never participates in technical work. A good scribe reduces postmortem preparation time by 70%.
- **Subject Matter Expert (SME):** Pulled in on-demand for specific subsystems. SMEs report findings to the IC, not to each other. Parallel SME investigations must be coordinated through the IC to avoid duplicated effort.
- **Customer Liaison:** Owns all outbound customer communication. Drafts status page updates for IC approval. Shields the technical team from inbound customer inquiries during active incidents.
**Process Phases:** Detect, Triage, Mobilize, Mitigate, Resolve, Postmortem.
**Communication Protocol:** PagerDuty mandates a dedicated Slack channel per incident, a bridge call for SEV1/SEV2, and status updates at fixed cadences (every 15 min for SEV1, every 30 min for SEV2). All decisions are announced in the channel, never in DMs or side threads.
### Google SRE: Managing Incidents (Chapter 14)
Google's SRE model, documented in *Site Reliability Engineering* (O'Reilly, 2016), emphasizes **role separation** and **clear handoffs** as the primary mechanisms for preventing incident chaos.
**Key Principles:**
- **Operational vs. Communication Tracks:** Google splits incident work into two parallel tracks. The operational track handles technical mitigation. The communication track handles stakeholder updates, executive briefings, and customer notifications. These tracks run independently with the IC bridging them.
- **Role Separation is Non-Negotiable:** The person debugging the system must never be the person updating stakeholders. Cognitive load from context-switching between technical work and communication degrades both outputs. Google measured a 40% increase in mean-time-to-resolution (MTTR) when a single person attempted both.
- **Clear Handoffs:** When an IC rotates out (recommended every 60-90 minutes for SEV1), the handoff includes: current status summary, active hypotheses, pending actions, and escalation state. Handoffs happen on the bridge call, not asynchronously.
- **Defined Command Post:** All communication flows through a single channel. Google uses the term "command post" -- a virtual or physical location where all incident participants converge.
### Atlassian Incident Management Model
Atlassian's model, published in their *Incident Management Handbook*, is **severity-driven** and **template-heavy**. It favors structured playbooks over improvisation.
**Key Characteristics:**
- **Severity Levels Drive Everything:** The assigned severity determines who gets paged, what communication templates are used, response time SLAs, and postmortem requirements. Severity is assigned at triage and reassessed every 30 minutes.
- **Handbook-Driven Approach:** Atlassian maintains runbooks for every known failure mode. During incidents, responders follow documented playbooks before improvising. This reduces MTTR for known issues by 50-60% but requires significant upfront investment in documentation.
- **Communication Templates:** Pre-written templates for status page updates, customer emails, and executive summaries. Templates include severity-specific language and are reviewed quarterly. This eliminates wordsmithing during active incidents.
- **Values-Based Decisions:** When runbooks do not cover the situation, Atlassian defaults to a decision hierarchy: (1) protect customer data, (2) restore service, (3) preserve evidence for root cause analysis.
### Framework Comparison Table
| Dimension | PagerDuty | Google SRE | Atlassian |
|-----------|-----------|------------|-----------|
| Primary strength | Speed of mobilization | Role separation discipline | Structured playbooks |
| IC authority model | IC has final say | IC coordinates, escalates to VP if blocked | IC follows handbook, escalates if off-script |
| Communication style | Dedicated channel + bridge | Command post with dual tracks | Template-driven status updates |
| Handoff protocol | Informal | Formal on-call handoff script | Rotation policy in handbook |
| Postmortem requirement | All SEV1/SEV2 | All incidents | SEV1/SEV2 mandatory, SEV3 optional |
| Best for | Fast-moving startups | Large-scale distributed systems | Regulated or process-heavy orgs |
| Weakness | Under-documented for edge cases | Heavyweight for small teams | Rigid, slow to adapt to novel failures |
### When to Use Which Framework
- **Teams under 20 engineers:** Start with PagerDuty's model. It is lightweight and prescriptive enough to work without heavy process investment. Add Atlassian-style runbooks as you identify recurring failure modes.
- **Teams running 50+ microservices:** Adopt Google SRE's dual-track model. The operational/communication split becomes critical when incidents span multiple teams and subsystems.
- **Regulated industries (finance, healthcare, government):** Use Atlassian's handbook-driven approach as the foundation. Regulatory auditors expect documented procedures, and templates satisfy compliance requirements for incident communication records.
- **Hybrid (recommended for most teams at scale):** Use PagerDuty's role definitions, Google's track separation, and Atlassian's template library. This is the approach codified in the rest of this document.
---
## 2. Severity Definitions
### Severity Classification Matrix
| Severity | Impact | Response Time | Update Cadence | Escalation Trigger | Example |
|----------|--------|---------------|----------------|---------------------|---------|
| **SEV1** | Total service outage or data breach affecting all users. Revenue loss exceeding $10K/hour. Security incident with active exfiltration. | Page IC + on-call within 5 min. All hands mobilized within 15 min. | Every 15 min to stakeholders. Continuous updates in incident channel. | Immediate executive notification. Board notification for data breaches. | Primary database cluster down. Payment processing system offline. Active ransomware attack. |
| **SEV2** | Major feature degraded for >30% of users. Revenue impact $1K-$10K/hour. Data integrity concerns without confirmed loss. | IC assigned within 15 min. Responders mobilized within 30 min. | Every 30 min to stakeholders. Every 15 min in incident channel. | Executive notification if unresolved after 1 hour. Upgrade to SEV1 if impact expands. | Search functionality returning errors for 40% of queries. Checkout flow failing intermittently. Authentication latency exceeding 10s. |
| **SEV3** | Minor feature degraded or non-critical service impaired. Workaround available. No direct revenue impact. | Acknowledged within 1 hour. Investigation started within 4 hours. | Every 2 hours to stakeholders if actively worked. Daily if deferred. | Escalate to SEV2 if workaround fails or user complaints exceed 50 in 1 hour. | Admin dashboard loading slowly. Email notifications delayed by 30+ minutes. Non-critical API endpoint returning 5xx for <5% of requests. |
| **SEV4** | Cosmetic issue, minor bug, or internal tooling degradation. No user-facing impact or negligible impact. | Acknowledged within 1 business day. Prioritized against backlog. | No scheduled updates. Tracked in issue tracker. | Escalate to SEV3 if internal productivity impact exceeds 2 hours/day across team. | Logging pipeline dropping non-critical debug logs. Internal metrics dashboard showing stale data. Minor UI alignment issue on one browser. |
### Customer-Facing Signals by Severity
**SEV1 Signals:** Support ticket volume spikes >500% of baseline within 15 minutes. Social media mentions of outage trend upward. Revenue dashboards show >95% drop in transaction volume. Multiple monitoring systems alarm simultaneously.
**SEV2 Signals:** Support ticket volume spikes 100-500% of baseline. Specific feature-related complaints cluster in support channels. Partial transaction failures visible in payment dashboards. Single monitoring system shows sustained alerting.
**SEV3 Signals:** Sporadic support tickets with a common pattern (under 20/hour). Users report intermittent issues with workarounds. Monitoring shows degraded but not critical metrics.
**SEV4 Signals:** Internal team notices issue during routine work. Occasional user mention with no pattern or urgency. Monitoring shows minor anomaly within acceptable thresholds.
### Severity Upgrade and Downgrade Criteria
**Upgrade from SEV2 to SEV1:** Impact expands to >80% of users, revenue impact confirmed above $10K/hour, data integrity compromise confirmed, or mitigation attempt fails after 45 minutes.
**Downgrade from SEV1 to SEV2:** Partial mitigation restores service for >70% of users, revenue impact drops below $10K/hour, and no ongoing data integrity concern.
**Downgrade from SEV2 to SEV3:** Workaround deployed and communicated, impact limited to <10% of users, and no revenue impact.
Severity changes must be announced by the IC in the incident channel with justification. The scribe logs the timestamp and rationale.
---
## 3. Role Definitions
### Incident Commander (IC)
The IC is the single decision-maker during an incident. This role exists to eliminate decision-by-committee, which adds 20-40 minutes to MTTR in measured studies.
**Responsibilities:**
- Assign severity level at triage (reassess every 30 minutes)
- Assign all other incident roles
- Approve status page updates before publication
- Make go/no-go decisions on mitigation strategies (rollback, feature flag, scaling)
- Decide when to escalate to executive leadership
- Declare incident resolved and initiate postmortem scheduling
**Decision Authority:** The IC can authorize rollbacks, page any team member regardless of org chart, approve customer communications, and override objections from individual contributors during active mitigation. The IC cannot approve financial expenditures above $50K or public press statements -- those require VP/C-level approval.
**What the IC Must NOT Do:** Debug code, write queries, SSH into production servers, or perform any hands-on technical work. The moment an IC starts debugging, incident coordination degrades. If the IC is the only person with domain expertise, they must hand off IC duties before engaging technically.
### Communications Lead
**Responsibilities:**
- Draft all status page updates using severity-appropriate templates
- Coordinate with Customer Liaison on outbound customer messaging
- Maintain the executive summary document (updated every 30 min for SEV1/SEV2)
- Manage the stakeholder notification list and delivery
- Post scheduled updates even when there is no new information ("We are continuing to investigate" is a valid update)
### Operations Lead
**Responsibilities:**
- Coordinate technical investigation across SMEs
- Maintain the running hypothesis list and assign investigation tasks
- Report technical findings to the IC in plain language
- Execute mitigation actions approved by the IC
- Track parallel workstreams and prevent duplicated effort
### Scribe
**Responsibilities:**
- Maintain a timestamped log of all decisions, actions, and findings
- Document who said what and when in the incident channel
- Capture rollback decisions, hypothesis changes, and escalation triggers
- Produce the initial postmortem timeline (saves 2-4 hours of postmortem prep)
### Subject Matter Experts (SMEs)
SMEs are paged on-demand by the IC for specific subsystems. They report findings to the Operations Lead, not directly to stakeholders. An SME who identifies a potential fix proposes it to the IC for approval before executing. SMEs are released from the incident explicitly by the IC when their subsystem is cleared.
### Customer Liaison
Owns the customer-facing voice during the incident. Monitors support channels for inbound customer reports. Drafts customer notification emails. Updates the public status page (after IC approval). Shields the technical team from direct customer inquiries during active mitigation.
---
## 4. Communication Protocols
### Incident Channel Naming Convention
Format: `#inc-YYYYMMDD-brief-desc`
Examples:
- `#inc-20260216-payment-api-timeout`
- `#inc-20260216-db-primary-failover`
- `#inc-20260216-auth-service-degraded`
Channel topic must include: severity, IC name, bridge call link, status page link.
Example topic: `SEV1 | IC: @jane.smith | Bridge: https://meet.example.com/inc-20260216 | Status: https://status.example.com`
### Internal Status Update Templates
**SEV1/SEV2 Update Template (posted in incident channel and executive Slack channel):**
```
INCIDENT UPDATE - [SEV1/SEV2] - [HH:MM UTC]
Status: [Investigating | Identified | Mitigating | Resolved]
Impact: [Specific user-facing impact in plain language]
Current Action: [What is actively being done right now]
Next Update: [HH:MM UTC]
IC: @[name]
```
**Executive Summary Template (for SEV1, updated every 30 min):**
```
EXECUTIVE SUMMARY - [Incident Title] - [HH:MM UTC]
Severity: SEV1
Duration: [X hours Y minutes]
Customer Impact: [Number of affected users/transactions]
Revenue Impact: [Estimated $ if known, "assessing" if not]
Current Status: [One sentence]
Mitigation ETA: [Estimated time or "unknown"]
Next Escalation Point: [What triggers executive action]
```
### Status Page Update Templates
**SEV1 Initial Post:**
```
Title: [Service Name] - Service Disruption
Body: We are currently experiencing a disruption affecting [service/feature].
Users may encounter [specific symptom: errors, timeouts, inability to access].
Our engineering team has been mobilized and is actively investigating.
We will provide an update within 15 minutes.
```
**SEV1 Update (mitigation in progress):**
```
Title: [Service Name] - Service Disruption (Update)
Body: We have identified the cause of the disruption affecting [service/feature]
and are implementing a fix. Some users may continue to experience [symptom].
We expect to have an update on resolution within [X] minutes.
```
**SEV1 Resolution:**
```
Title: [Service Name] - Resolved
Body: The disruption affecting [service/feature] has been resolved as of [HH:MM UTC].
Service has been restored to normal operation. Users should no longer experience
[symptom]. We will publish a full incident report within 48 hours.
We apologize for the inconvenience.
```
**SEV2 Initial Post:**
```
Title: [Service Name] - Degraded Performance
Body: We are investigating reports of degraded performance affecting [feature].
Some users may experience [specific symptom]. A workaround is [available/not yet available].
Our team is actively investigating and we will provide an update within 30 minutes.
```
### Bridge Call / War Room Etiquette
1. **Mute by default.** Unmute only when speaking to the IC or Operations Lead.
2. **Identify yourself before speaking.** "This is [name] from [team]." Every time.
3. **State findings, then recommendations.** "Database replication lag is 45 seconds and climbing. I recommend we fail over to the secondary cluster."
4. **IC confirms before action.** No unilateral action on production systems during an incident. The IC says "approved" or "hold" before anyone executes.
5. **No side conversations.** If two SMEs need to discuss a hypothesis, they take it to a breakout channel and report back findings to the main bridge.
6. **Time-box debugging.** The IC sets 15-minute timers for investigation threads. If a hypothesis is not confirmed or denied in 15 minutes, pivot to the next hypothesis or escalate.
### Customer Notification Templates
**SEV1 Customer Email (B2B, enterprise accounts):**
```
Subject: [Company Name] Service Incident - [Date]
Dear [Customer Name],
We are writing to inform you of a service incident affecting [product/service]
that began at [HH:MM UTC] on [date].
Impact: [Specific impact to this customer's usage]
Current Status: [Brief status]
Expected Resolution: [ETA if known, or "We are working to resolve this as quickly as possible"]
We will continue to provide updates every [15/30] minutes until resolution.
Your dedicated account team is available at [contact info] for any questions.
Sincerely,
[Name], [Title]
```
---
## 5. Escalation Matrix
### Escalation Tiers
**Tier 1 - Within Team (0-15 minutes):**
On-call engineer investigates. If the issue is within the team's domain and matches a known runbook, resolve without escalation. Page the IC if severity is SEV2 or higher, or if the issue is not resolved within 15 minutes.
**Tier 2 - Cross-Team (15-45 minutes):**
IC pages SMEs from adjacent teams. Common cross-team escalations: database team for replication issues, networking team for connectivity failures, security team for suspicious activity. Cross-team SMEs join the incident channel and bridge call.
**Tier 3 - Executive (45+ minutes or immediate for SEV1):**
VP of Engineering notified for all SEV1 incidents immediately. CTO notified if SEV1 exceeds 1 hour without mitigation progress. CEO notified if SEV1 involves data breach or regulatory implications. Executive involvement is for resource allocation and external communication decisions, not technical direction.
### Time-Based Escalation Triggers
| Elapsed Time | SEV1 Action | SEV2 Action |
|-------------|-------------|-------------|
| 0 min | Page IC + all on-call. Notify VP Eng. | Page IC + primary on-call. |
| 15 min | Confirm all roles staffed. Open bridge call. | IC assesses if additional SMEs needed. |
| 30 min | If no mitigation path identified, page backup on-call for all related services. | First stakeholder update. Reassess severity. |
| 45 min | Escalate to CTO if no progress. Consider customer notification. | If no progress, consider escalating to SEV1. |
| 60 min | CTO briefing. Initiate customer notification if not already done. | Notify VP Eng. Page cross-team SMEs. |
| 90 min | IC rotation (fresh IC takes over). Reassess all hypotheses. | IC rotation if needed. |
| 120 min | CEO briefing if data breach or regulatory risk. External PR team engaged. | Escalate to SEV1 if impact has not decreased. |
### Escalation Path Examples
**Database failover failure:**
On-call DBA (Tier 1, 0-15 min) -> IC + DBA team lead (Tier 2, 15 min) -> Infrastructure VP + cloud provider support (Tier 3, 45 min)
**Payment processing outage:**
On-call payments engineer (Tier 1, 0-5 min) -> IC + payments team lead + payment provider liaison (Tier 2, 5 min, immediate due to revenue impact) -> CFO + VP Eng (Tier 3, 15 min if provider-side issue confirmed)
**Security incident (suspected breach):**
Security on-call (Tier 1, 0-5 min) -> CISO + IC + legal counsel (Tier 2, immediate) -> CEO + external incident response firm (Tier 3, within 1 hour if breach confirmed)
### On-Call Rotation Best Practices
- **Primary + secondary on-call** for every critical service. Secondary is paged automatically if primary does not acknowledge within 5 minutes.
- **On-call shifts are 7 days maximum.** Longer rotations degrade alertness and response quality.
- **Handoff checklist:** Current open issues, recent deploys in the last 48 hours, known risks or maintenance windows, escalation contacts for dependent services.
- **On-call load budget:** No more than 2 pages per night on average, measured weekly. Exceeding this indicates systemic reliability issues that must be addressed with engineering investment, not heroic on-call effort.
---
## 6. Incident Lifecycle Phases
### Phase 1: Detection
Detection comes from three sources, in order of preference:
1. **Automated monitoring (preferred):** Alerting rules on latency (p99 > 2x baseline), error rates (5xx > 1% of requests), saturation (CPU > 85%, memory > 90%, disk > 80%), and business metrics (transaction volume drops > 20% from 15-minute rolling average). Alerts should fire within 60 seconds of threshold breach.
2. **Internal reports:** An engineer notices anomalous behavior during routine work. Internal detection typically adds 5-15 minutes to response time compared to automated monitoring.
3. **Customer reports:** Customers contact support about issues. This is the worst detection source. If customers detect incidents before monitoring, the monitoring coverage has a gap that must be closed in the postmortem.
**Detection SLA:** SEV1 incidents must be detected within 5 minutes of impact onset. If detection latency exceeds this, the postmortem must include a monitoring improvement action item.
### Phase 2: Triage
The first responder performs initial triage within 5 minutes of detection:
1. **Scope assessment:** How many users, services, or regions are affected? Check dashboards, not assumptions.
2. **Severity assignment:** Use the severity matrix in Section 2. When in doubt, assign higher severity. Downgrading is cheap; delayed escalation is expensive.
3. **IC assignment:** For SEV1/SEV2, page the on-call IC immediately. For SEV3, the first responder may self-assign IC duties.
4. **Initial hypothesis:** What changed in the last 2 hours? Check deploy logs, config changes, upstream dependency status, and traffic patterns. 70% of incidents correlate with a change deployed in the prior 2 hours.
### Phase 3: Mobilization
The IC executes mobilization within 10 minutes of assignment:
1. **Create incident channel:** `#inc-YYYYMMDD-brief-desc`. Set topic with severity, IC name, bridge link.
2. **Assign roles:** Communications Lead, Operations Lead, Scribe. For SEV3/SEV4, the IC may cover multiple roles.
3. **Open bridge call (SEV1/SEV2):** Share link in incident channel. All responders join within 5 minutes.
4. **Post initial summary:** Current understanding, affected services, assigned roles, first actions.
5. **Notify stakeholders:** Page dependent teams. Notify customer support leadership. For SEV1, notify executive chain per escalation matrix.
### Phase 4: Investigation
Investigation runs as parallel workstreams coordinated by the Operations Lead:
- **Workstream discipline:** Each SME investigates one hypothesis at a time. The Operations Lead tracks active hypotheses on a shared list. Completed investigations report: confirmed, denied, or inconclusive.
- **Hypothesis testing priority:** (1) Recent changes (deploys, configs, feature flags), (2) Upstream dependency failures, (3) Capacity exhaustion, (4) Data corruption, (5) Security compromise.
- **15-minute rule:** If a hypothesis is not confirmed or denied within 15 minutes, the IC decides whether to continue, pivot, or escalate. Unbounded investigation is the leading cause of extended MTTR.
- **Evidence collection:** Screenshots, log snippets, metric graphs, and query results are posted in the incident channel, not described verbally. The scribe tags evidence with timestamps.
### Phase 5: Mitigation
Mitigation prioritizes restoring service over finding root cause:
- **Rollback first:** If a deploy correlates with the incident, roll it back before investigating further. A 5-minute rollback beats a 45-minute investigation. Rollback authority rests with the IC.
- **Feature flags:** Disable the suspected feature via feature flag if available. This is faster and less risky than a full rollback.
- **Scaling:** If the issue is capacity-related, scale horizontally before investigating the traffic source.
- **Failover:** If a primary system is unrecoverable, fail over to the secondary. Test failover procedures quarterly so this is a routine, not a gamble.
- **Customer workaround:** If mitigation will take time, publish a workaround for customers (e.g., "Use the mobile app while we restore web access").
**Mitigation verification:** After applying mitigation, monitor key metrics for 15 minutes before declaring the issue mitigated. Premature declarations that the issue is mitigated followed by recurrence damage team credibility and customer trust.
### Phase 6: Resolution
Resolution is declared when the root cause is addressed and service is operating normally:
- **Verification checklist:** Error rates returned to baseline, latency returned to baseline, no ongoing customer reports, monitoring confirms stability for 30+ minutes.
- **Incident channel update:** IC posts final status with resolution summary, total duration, and next steps.
- **Status page update:** Post resolution notice within 15 minutes of declaring resolved.
- **Stand down:** IC explicitly releases all responders. SMEs return to normal work. Bridge call is closed.
### Phase 7: Postmortem
Postmortem is mandatory for SEV1 and SEV2. Optional but recommended for SEV3. Never conducted for SEV4.
- **Timeline:** Postmortem document drafted within 24 hours. Postmortem meeting held within 72 hours (3 business days). Action items assigned and tracked in the team's issue tracker.
- **Blameless standard:** The postmortem examines systems, processes, and tools -- not individual performance. "Why did the system allow this?" not "Why did [person] do this?"
- **Required sections:** Timeline (from scribe's log), root cause analysis (using 5 Whys or fault tree), impact summary (users, revenue, duration), what went well, what went poorly, action items with owners and due dates.
- **Action items and recurrence:** Every postmortem produces 3-7 concrete action items. Items without owners and due dates are not action items. Teams should close 80%+ within 30 days. If the same root cause appears in two postmortems within 6 months, escalate to engineering leadership as a systemic reliability investment area.
FILE:references/incident_severity_matrix.md
# Incident Severity Classification Matrix
## Overview
This document defines the severity classification system used for incident response. The classification determines response requirements, escalation paths, and communication frequency.
## Severity Levels
### SEV1 - Critical Outage
**Definition:** Complete service failure affecting all users or critical business functions
#### Impact Criteria
- Customer-facing services completely unavailable
- Data loss or corruption affecting users
- Security breaches with customer data exposure
- Revenue-generating systems down
- SLA violations with financial penalties
- > 75% of users affected
#### Response Requirements
| Metric | Requirement |
|--------|-------------|
| **Response Time** | Immediate (0-5 minutes) |
| **Incident Commander** | Assigned within 5 minutes |
| **War Room** | Established within 10 minutes |
| **Executive Notification** | Within 15 minutes |
| **Public Status Page** | Updated within 15 minutes |
| **Customer Communication** | Within 30 minutes |
#### Escalation Path
1. **Immediate**: On-call Engineer → Incident Commander
2. **15 minutes**: VP Engineering + Customer Success VP
3. **30 minutes**: CTO
4. **60 minutes**: CEO + Full Executive Team
#### Communication Requirements
- **Frequency**: Every 15 minutes until resolution
- **Channels**: PagerDuty, Phone, Slack, Email, Status Page
- **Recipients**: All engineering, executives, customer success
- **Template**: SEV1 Executive Alert Template
---
### SEV2 - Major Impact
**Definition:** Significant degradation affecting subset of users or non-critical functions
#### Impact Criteria
- Partial service degradation (25-75% of users affected)
- Performance issues causing user frustration
- Non-critical features unavailable
- Internal tools impacting productivity
- Data inconsistencies not affecting user experience
- API errors affecting integrations
#### Response Requirements
| Metric | Requirement |
|--------|-------------|
| **Response Time** | 15 minutes |
| **Incident Commander** | Assigned within 30 minutes |
| **Status Page Update** | Within 30 minutes |
| **Stakeholder Notification** | Within 1 hour |
| **Team Assembly** | Within 30 minutes |
#### Escalation Path
1. **Immediate**: On-call Engineer → Team Lead
2. **30 minutes**: Engineering Manager
3. **2 hours**: VP Engineering
4. **4 hours**: CTO (if unresolved)
#### Communication Requirements
- **Frequency**: Every 30 minutes during active response
- **Channels**: PagerDuty, Slack, Email
- **Recipients**: Engineering team, product team, relevant stakeholders
- **Template**: SEV2 Major Impact Template
---
### SEV3 - Minor Impact
**Definition:** Limited impact with workarounds available
#### Impact Criteria
- Single feature or component affected
- < 25% of users impacted
- Workarounds available
- Performance degradation not significantly impacting UX
- Non-urgent monitoring alerts
- Development/test environment issues
#### Response Requirements
| Metric | Requirement |
|--------|-------------|
| **Response Time** | 2 hours (business hours) |
| **After Hours Response** | Next business day |
| **Team Assignment** | Within 4 hours |
| **Status Page Update** | Optional |
| **Internal Notification** | Within 2 hours |
#### Escalation Path
1. **Immediate**: Assigned Engineer
2. **4 hours**: Team Lead
3. **1 business day**: Engineering Manager (if needed)
#### Communication Requirements
- **Frequency**: At key milestones only
- **Channels**: Slack, Email
- **Recipients**: Assigned team, team lead
- **Template**: SEV3 Minor Impact Template
---
### SEV4 - Low Impact
**Definition:** Minimal impact, cosmetic issues, or planned maintenance
#### Impact Criteria
- Cosmetic bugs
- Documentation issues
- Logging or monitoring gaps
- Performance issues with no user impact
- Development/test environment issues
- Feature requests or enhancements
#### Response Requirements
| Metric | Requirement |
|--------|-------------|
| **Response Time** | 1-2 business days |
| **Assignment** | Next sprint planning |
| **Tracking** | Standard ticket system |
| **Escalation** | None required |
#### Communication Requirements
- **Frequency**: Standard development cycle updates
- **Channels**: Ticket system
- **Recipients**: Product owner, assigned developer
- **Template**: Standard issue template
## Classification Guidelines
### User Impact Assessment
| Impact Scope | Description | Typical Severity |
|--------------|-------------|------------------|
| **All Users** | 100% of users affected | SEV1 |
| **Major Subset** | 50-75% of users affected | SEV1/SEV2 |
| **Significant Subset** | 25-50% of users affected | SEV2 |
| **Limited Users** | 5-25% of users affected | SEV2/SEV3 |
| **Few Users** | < 5% of users affected | SEV3/SEV4 |
| **No User Impact** | Internal only | SEV4 |
### Business Impact Assessment
| Business Impact | Description | Severity Boost |
|-----------------|-------------|----------------|
| **Revenue Loss** | Direct revenue impact | +1 severity level |
| **SLA Breach** | Contract violations | +1 severity level |
| **Regulatory** | Compliance implications | +1 severity level |
| **Brand Damage** | Public-facing issues | +1 severity level |
| **Security** | Data or system security | +2 severity levels |
### Duration Considerations
| Duration | Impact on Classification |
|----------|--------------------------|
| **< 15 minutes** | May reduce severity by 1 level |
| **15-60 minutes** | Standard classification |
| **1-4 hours** | May increase severity by 1 level |
| **> 4 hours** | Significant severity increase |
## Decision Tree
```
1. Is this a security incident with data exposure?
→ YES: SEV1 (regardless of user count)
→ NO: Continue to step 2
2. Are revenue-generating services completely down?
→ YES: SEV1
→ NO: Continue to step 3
3. What percentage of users are affected?
→ > 75%: SEV1
→ 25-75%: SEV2
→ 5-25%: SEV3
→ < 5%: SEV4
4. Apply business impact modifiers
5. Consider duration factors
6. When in doubt, err on higher severity
```
## Examples
### SEV1 Examples
- Payment processing system completely down
- All user authentication failing
- Database corruption causing data loss
- Security breach with customer data exposed
- Website returning 500 errors for all users
### SEV2 Examples
- Payment processing slow (30-second delays)
- Search functionality returning incomplete results
- API rate limits causing partner integration issues
- Dashboard displaying stale data (> 1 hour old)
- Mobile app crashing for 40% of users
### SEV3 Examples
- Single feature in admin panel not working
- Email notifications delayed by 1 hour
- Non-critical API endpoint returning errors
- Cosmetic UI bug in settings page
- Development environment deployment failing
### SEV4 Examples
- Typo in help documentation
- Log format change needed for analysis
- Non-critical performance optimization
- Internal tool enhancement request
- Test data cleanup needed
## Escalation Triggers
### Automatic Escalation
- SEV1 incidents automatically escalate every 30 minutes if unresolved
- SEV2 incidents escalate after 2 hours without significant progress
- Any incident with expanding scope increases severity
- Customer escalation to support triggers severity review
### Manual Escalation
- Incident Commander can escalate at any time
- Technical leads can request escalation
- Business stakeholders can request severity review
- External factors (media attention, regulatory) trigger escalation
## Communication Templates
### SEV1 Executive Alert
```
Subject: 🚨 CRITICAL INCIDENT - [Service] Complete Outage
URGENT: Customer-facing service outage requiring immediate attention
Service: [Service Name]
Start Time: [Timestamp]
Impact: [Description of customer impact]
Estimated Affected Users: [Number/Percentage]
Business Impact: [Revenue/SLA/Brand implications]
Incident Commander: [Name] ([Contact])
Response Team: [Team members engaged]
Current Status: [Brief status update]
Next Update: [Timestamp - 15 minutes from now]
War Room: [Bridge/Chat link]
This is a customer-impacting incident requiring executive awareness.
```
### SEV2 Major Impact
```
Subject: ⚠️ [SEV2] [Service] - Major Performance Impact
Major service degradation affecting user experience
Service: [Service Name]
Start Time: [Timestamp]
Impact: [Description of user impact]
Scope: [Affected functionality/users]
Response Team: [Team Lead] + [Team members]
Status: [Current mitigation efforts]
Workaround: [If available]
Next Update: 30 minutes
Status Page: [Link if updated]
```
## Review and Updates
This severity matrix should be reviewed quarterly and updated based on:
- Incident response learnings
- Business priority changes
- Service architecture evolution
- Regulatory requirement changes
- Customer feedback and SLA updates
**Last Updated:** February 2026
**Next Review:** May 2026
**Owner:** Engineering Leadership
FILE:references/rca_frameworks_guide.md
# Root Cause Analysis (RCA) Frameworks Guide
## Overview
This guide provides detailed instructions for applying various Root Cause Analysis frameworks during Post-Incident Reviews. Each framework offers a different perspective and approach to identifying underlying causes of incidents.
## Framework Selection Guidelines
| Incident Type | Recommended Framework | Why |
|---------------|----------------------|-----|
| **Process Failure** | 5 Whys | Simple, direct cause-effect chain |
| **Complex System Failure** | Fishbone + Timeline | Multiple contributing factors |
| **Human Error** | Fishbone | Systematic analysis of contributing factors |
| **Extended Incidents** | Timeline Analysis | Understanding decision points |
| **High-Risk Incidents** | Bow Tie | Comprehensive barrier analysis |
| **Recurring Issues** | 5 Whys + Fishbone | Deep dive into systemic issues |
---
## 5 Whys Analysis Framework
### Purpose
Iteratively drill down through cause-effect relationships to identify root causes.
### When to Use
- Simple, linear cause-effect chains
- Time-pressured analysis
- Process-related failures
- Individual component failures
### Process Steps
#### Step 1: Problem Statement
Write a clear, specific problem statement.
**Good Example:**
> "The payment API returned 500 errors for 2 hours on March 15, affecting 80% of checkout attempts."
**Poor Example:**
> "The system was broken."
#### Step 2: First Why
Ask why the problem occurred. Focus on immediate, observable causes.
**Example:**
- **Why 1:** Why did the payment API return 500 errors?
- **Answer:** The database connection pool was exhausted.
#### Step 3: Subsequent Whys
For each answer, ask "why" again. Continue until you reach a root cause.
**Example Chain:**
- **Why 2:** Why was the database connection pool exhausted?
- **Answer:** The application was creating more connections than usual.
- **Why 3:** Why was the application creating more connections?
- **Answer:** A new feature wasn't properly closing connections.
- **Why 4:** Why wasn't the feature properly closing connections?
- **Answer:** Code review missed the connection leak pattern.
- **Why 5:** Why did code review miss this pattern?
- **Answer:** We don't have automated checks for connection pooling best practices.
#### Step 4: Validation
Verify that addressing the root cause would prevent the original problem.
### Best Practices
1. **Ask at least 3 "whys"** - Surface causes are rarely root causes
2. **Focus on process failures, not people** - Avoid blame, focus on system improvements
3. **Use evidence** - Support each answer with data or observations
4. **Consider multiple paths** - Some problems have multiple root causes
5. **Test the logic** - Work backwards from root cause to problem
### Common Pitfalls
- **Stopping too early** - First few whys often reveal symptoms, not causes
- **Single-cause assumption** - Complex systems often have multiple contributing factors
- **Blame focus** - Focusing on individual mistakes rather than system failures
- **Vague answers** - Use specific, actionable answers
### 5 Whys Template
```markdown
## 5 Whys Analysis
**Problem Statement:** [Clear description of the incident]
**Why 1:** [First why question]
**Answer:** [Specific, evidence-based answer]
**Evidence:** [Supporting data, logs, observations]
**Why 2:** [Second why question]
**Answer:** [Specific answer based on Why 1]
**Evidence:** [Supporting evidence]
[Continue for 3-7 iterations]
**Root Cause(s) Identified:**
1. [Primary root cause]
2. [Secondary root cause if applicable]
**Validation:** [Confirm that addressing root causes would prevent recurrence]
```
---
## Fishbone (Ishikawa) Diagram Framework
### Purpose
Systematically analyze potential causes across multiple categories to identify contributing factors.
### When to Use
- Complex incidents with multiple potential causes
- When human factors are suspected
- Systemic or organizational issues
- When 5 Whys doesn't reveal clear root causes
### Categories
#### People (Human Factors)
- **Training and Skills**
- Insufficient training on new systems
- Lack of domain expertise
- Skill gaps in team
- Knowledge not shared across team
- **Communication**
- Poor communication between teams
- Unclear responsibilities
- Information not reaching right people
- Language/cultural barriers
- **Decision Making**
- Decisions made under pressure
- Insufficient information for decisions
- Risk assessment inadequate
- Approval processes bypassed
#### Process (Procedures and Workflows)
- **Documentation**
- Outdated procedures
- Missing runbooks
- Unclear instructions
- Process not documented
- **Change Management**
- Inadequate change review
- Rushed deployments
- Insufficient testing
- Rollback procedures unclear
- **Review and Approval**
- Code review gaps
- Architecture review skipped
- Security review insufficient
- Performance review missing
#### Technology (Systems and Tools)
- **Architecture**
- Single points of failure
- Insufficient redundancy
- Scalability limitations
- Tight coupling between systems
- **Monitoring and Alerting**
- Missing monitoring
- Alert fatigue
- Inadequate thresholds
- Poor alert routing
- **Tools and Automation**
- Manual processes prone to error
- Tool limitations
- Automation gaps
- Integration issues
#### Environment (External Factors)
- **Infrastructure**
- Hardware failures
- Network issues
- Capacity limitations
- Geographic dependencies
- **Dependencies**
- Third-party service failures
- External API changes
- Vendor issues
- Supply chain problems
- **External Pressure**
- Time pressure from business
- Resource constraints
- Regulatory changes
- Market conditions
### Process Steps
#### Step 1: Define the Problem
Place the incident at the "head" of the fishbone diagram.
#### Step 2: Brainstorm Causes
For each category, brainstorm potential contributing factors.
#### Step 3: Drill Down
For each factor, ask what caused that factor (sub-causes).
#### Step 4: Identify Primary Causes
Mark the most likely contributing factors based on evidence.
#### Step 5: Validate
Gather evidence to support or refute each suspected cause.
### Fishbone Template
```markdown
## Fishbone Analysis
**Problem:** [Incident description]
### People
**Training/Skills:**
- [Factor 1]: [Evidence/likelihood]
- [Factor 2]: [Evidence/likelihood]
**Communication:**
- [Factor 1]: [Evidence/likelihood]
**Decision Making:**
- [Factor 1]: [Evidence/likelihood]
### Process
**Documentation:**
- [Factor 1]: [Evidence/likelihood]
**Change Management:**
- [Factor 1]: [Evidence/likelihood]
**Review/Approval:**
- [Factor 1]: [Evidence/likelihood]
### Technology
**Architecture:**
- [Factor 1]: [Evidence/likelihood]
**Monitoring:**
- [Factor 1]: [Evidence/likelihood]
**Tools:**
- [Factor 1]: [Evidence/likelihood]
### Environment
**Infrastructure:**
- [Factor 1]: [Evidence/likelihood]
**Dependencies:**
- [Factor 1]: [Evidence/likelihood]
**External Factors:**
- [Factor 1]: [Evidence/likelihood]
### Primary Contributing Factors
1. [Factor with highest evidence/impact]
2. [Second most significant factor]
3. [Third most significant factor]
### Root Cause Hypothesis
[Synthesized explanation of how factors combined to cause incident]
```
---
## Timeline Analysis Framework
### Purpose
Analyze the chronological sequence of events to identify decision points, missed opportunities, and process gaps.
### When to Use
- Extended incidents (> 1 hour)
- Complex multi-phase incidents
- When response effectiveness is questioned
- Communication or coordination failures
### Analysis Dimensions
#### Detection Analysis
- **Time to Detection:** How long from onset to first alert?
- **Detection Method:** How was the incident first identified?
- **Alert Effectiveness:** Were the right people notified quickly?
- **False Negatives:** What signals were missed?
#### Response Analysis
- **Time to Response:** How long from detection to first response action?
- **Escalation Timing:** Were escalations timely and appropriate?
- **Resource Mobilization:** How quickly were the right people engaged?
- **Decision Points:** What key decisions were made and when?
#### Communication Analysis
- **Internal Communication:** How effective was team coordination?
- **External Communication:** Were stakeholders informed appropriately?
- **Communication Gaps:** Where did information flow break down?
- **Update Frequency:** Were updates provided at appropriate intervals?
#### Resolution Analysis
- **Mitigation Strategy:** Was the chosen approach optimal?
- **Alternative Paths:** What other options were considered?
- **Resource Allocation:** Were resources used effectively?
- **Verification:** How was resolution confirmed?
### Process Steps
#### Step 1: Event Reconstruction
Create comprehensive timeline with all available events.
#### Step 2: Phase Identification
Identify distinct phases (detection, triage, escalation, mitigation, resolution).
#### Step 3: Gap Analysis
Identify time gaps and analyze their causes.
#### Step 4: Decision Point Analysis
Examine key decision points and alternative paths.
#### Step 5: Effectiveness Assessment
Evaluate the overall effectiveness of the response.
### Timeline Template
```markdown
## Timeline Analysis
### Incident Phases
1. **Detection** ([start] - [end], [duration])
2. **Triage** ([start] - [end], [duration])
3. **Escalation** ([start] - [end], [duration])
4. **Mitigation** ([start] - [end], [duration])
5. **Resolution** ([start] - [end], [duration])
### Key Decision Points
**[Timestamp]:** [Decision made]
- **Context:** [Situation at time of decision]
- **Alternatives:** [Other options considered]
- **Outcome:** [Result of decision]
- **Assessment:** [Was this optimal?]
### Communication Timeline
**[Timestamp]:** [Communication event]
- **Channel:** [Slack/Email/Phone/etc.]
- **Audience:** [Who was informed]
- **Content:** [What was communicated]
- **Effectiveness:** [Assessment]
### Gaps and Delays
**[Time Period]:** [Description of gap]
- **Duration:** [Length of gap]
- **Cause:** [Why did gap occur]
- **Impact:** [Effect on incident response]
### Response Effectiveness
**Strengths:**
- [What went well]
- [Effective decisions/actions]
**Weaknesses:**
- [What could be improved]
- [Missed opportunities]
### Root Causes from Timeline
1. [Process-based root cause]
2. [Communication-based root cause]
3. [Decision-making root cause]
```
---
## Bow Tie Analysis Framework
### Purpose
Analyze both preventive measures (left side) and protective measures (right side) around an incident.
### When to Use
- High-severity incidents (SEV1)
- Security incidents
- Safety-critical systems
- When comprehensive barrier analysis is needed
### Components
#### Hazards
What conditions create the potential for incidents?
**Examples:**
- High traffic loads
- Software deployments
- Human interactions with critical systems
- Third-party dependencies
#### Top Event
What actually went wrong? This is the center of the bow tie.
**Examples:**
- "Database became unresponsive"
- "Payment processing failed"
- "User authentication service crashed"
#### Threats (Left Side)
What specific causes could lead to the top event?
**Examples:**
- Code defects in new deployment
- Database connection pool exhaustion
- Network connectivity issues
- DDoS attack
#### Consequences (Right Side)
What are the potential impacts of the top event?
**Examples:**
- Revenue loss
- Customer churn
- Regulatory violations
- Brand damage
- Data loss
#### Barriers
What controls exist (or could exist) to prevent threats or mitigate consequences?
**Preventive Barriers (Left Side):**
- Code reviews
- Automated testing
- Load testing
- Input validation
- Rate limiting
**Protective Barriers (Right Side):**
- Circuit breakers
- Failover systems
- Backup procedures
- Customer communication
- Rollback capabilities
### Process Steps
#### Step 1: Define the Top Event
Clearly state what went wrong.
#### Step 2: Identify Threats
Brainstorm all possible causes that could lead to the top event.
#### Step 3: Identify Consequences
List all potential impacts of the top event.
#### Step 4: Map Existing Barriers
Identify current controls for each threat and consequence.
#### Step 5: Assess Barrier Effectiveness
Evaluate how well each barrier worked (or failed).
#### Step 6: Recommend Additional Barriers
Identify new controls needed to prevent recurrence.
### Bow Tie Template
```markdown
## Bow Tie Analysis
**Top Event:** [What went wrong]
### Threats (Potential Causes)
1. **[Threat 1]**
- Likelihood: [High/Medium/Low]
- Current Barriers: [Preventive controls]
- Barrier Effectiveness: [Assessment]
2. **[Threat 2]**
- Likelihood: [High/Medium/Low]
- Current Barriers: [Preventive controls]
- Barrier Effectiveness: [Assessment]
### Consequences (Potential Impacts)
1. **[Consequence 1]**
- Severity: [High/Medium/Low]
- Current Barriers: [Protective controls]
- Barrier Effectiveness: [Assessment]
2. **[Consequence 2]**
- Severity: [High/Medium/Low]
- Current Barriers: [Protective controls]
- Barrier Effectiveness: [Assessment]
### Barrier Analysis
**Effective Barriers:**
- [Barrier that worked well]
- [Why it was effective]
**Failed Barriers:**
- [Barrier that failed]
- [Why it failed]
- [How to improve]
**Missing Barriers:**
- [Needed preventive control]
- [Needed protective control]
### Recommendations
**Preventive Measures:**
1. [New barrier to prevent threat]
2. [Improvement to existing barrier]
**Protective Measures:**
1. [New barrier to mitigate consequence]
2. [Improvement to existing barrier]
```
---
## Framework Comparison
| Framework | Time Required | Complexity | Best For | Output |
|-----------|---------------|------------|----------|---------|
| **5 Whys** | 30-60 minutes | Low | Simple, linear causes | Clear cause chain |
| **Fishbone** | 1-2 hours | Medium | Complex, multi-factor | Comprehensive factor map |
| **Timeline** | 2-3 hours | Medium | Extended incidents | Process improvements |
| **Bow Tie** | 2-4 hours | High | High-risk incidents | Barrier strategy |
## Combining Frameworks
### 5 Whys + Fishbone
Use 5 Whys for initial analysis, then Fishbone to explore contributing factors.
### Timeline + 5 Whys
Use Timeline to identify key decision points, then 5 Whys on critical failures.
### Fishbone + Bow Tie
Use Fishbone to identify causes, then Bow Tie to develop comprehensive prevention strategy.
## Quality Checklist
- [ ] Root causes address systemic issues, not symptoms
- [ ] Analysis is backed by evidence, not assumptions
- [ ] Multiple perspectives considered (technical, process, human)
- [ ] Recommendations are specific and actionable
- [ ] Analysis focuses on prevention, not blame
- [ ] Findings are validated against incident timeline
- [ ] Contributing factors are prioritized by impact
- [ ] Root causes link clearly to preventive actions
## Common Anti-Patterns
- **Human Error as Root Cause** - Dig deeper into why human error occurred
- **Single Root Cause** - Complex systems usually have multiple contributing factors
- **Technology-Only Focus** - Consider process and organizational factors
- **Blame Assignment** - Focus on system improvements, not individual fault
- **Generic Recommendations** - Provide specific, measurable actions
- **Surface-Level Analysis** - Ensure you've reached true root causes
---
**Last Updated:** February 2026
**Next Review:** August 2026
**Owner:** SRE Team + Engineering Leadership
FILE:references/reference-information.md
# incident-commander reference
## Reference Information
- **Architecture Diagram:** {link}
- **Monitoring Dashboard:** {link}
- **Related Runbooks:** {links to dependent service runbooks}
```
### Post-Incident Review (PIR) Framework
#### PIR Timeline and Ownership
**Timeline:**
- **24 hours:** Initial PIR draft completed by Incident Commander
- **3 business days:** Final PIR published with all stakeholder input
- **1 week:** Action items assigned with owners and due dates
- **4 weeks:** Follow-up review on action item progress
**Roles:**
- **PIR Owner:** Incident Commander (can delegate writing but owns completion)
- **Technical Contributors:** All engineers involved in response
- **Review Committee:** Engineering leadership, affected product teams
- **Action Item Owners:** Assigned based on expertise and capacity
#### Root Cause Analysis Frameworks
#### 1. Five Whys Method
The Five Whys technique involves asking "why" repeatedly to drill down to root causes:
**Example Application:**
- **Problem:** Database became unresponsive during peak traffic
- **Why 1:** Why did the database become unresponsive? → Connection pool was exhausted
- **Why 2:** Why was the connection pool exhausted? → Application was creating more connections than usual
- **Why 3:** Why was the application creating more connections? → New feature wasn't properly connection pooling
- **Why 4:** Why wasn't the feature properly connection pooling? → Code review missed this pattern
- **Why 5:** Why did code review miss this? → No automated checks for connection pooling patterns
**Best Practices:**
- Ask "why" at least 3 times, often need 5+ iterations
- Focus on process failures, not individual blame
- Each "why" should point to a actionable system improvement
- Consider multiple root cause paths, not just one linear chain
#### 2. Fishbone (Ishikawa) Diagram
Systematic analysis across multiple categories of potential causes:
**Categories:**
- **People:** Training, experience, communication, handoffs
- **Process:** Procedures, change management, review processes
- **Technology:** Architecture, tooling, monitoring, automation
- **Environment:** Infrastructure, dependencies, external factors
**Application Method:**
1. State the problem clearly at the "head" of the fishbone
2. For each category, brainstorm potential contributing factors
3. For each factor, ask what caused that factor (sub-causes)
4. Identify the factors most likely to be root causes
5. Validate root causes with evidence from the incident
#### 3. Timeline Analysis
Reconstruct the incident chronologically to identify decision points and missed opportunities:
**Timeline Elements:**
- **Detection:** When was the issue first observable? When was it first detected?
- **Notification:** How quickly were the right people informed?
- **Response:** What actions were taken and how effective were they?
- **Communication:** When were stakeholders updated?
- **Resolution:** What finally resolved the issue?
**Analysis Questions:**
- Where were there delays and what caused them?
- What decisions would we make differently with perfect information?
- Where did communication break down?
- What automation could have detected/resolved faster?
### Escalation Paths
#### Technical Escalation
**Level 1:** On-call engineer
- **Responsibility:** Initial response and common issue resolution
- **Escalation Trigger:** Issue not resolved within SLA timeframe
- **Timeframe:** 15 minutes (SEV1), 30 minutes (SEV2)
**Level 2:** Senior engineer/Team lead
- **Responsibility:** Complex technical issues requiring deeper expertise
- **Escalation Trigger:** Level 1 requests help or timeout occurs
- **Timeframe:** 30 minutes (SEV1), 1 hour (SEV2)
**Level 3:** Engineering Manager/Staff Engineer
- **Responsibility:** Cross-team coordination and architectural decisions
- **Escalation Trigger:** Issue spans multiple systems or teams
- **Timeframe:** 45 minutes (SEV1), 2 hours (SEV2)
**Level 4:** Director of Engineering/CTO
- **Responsibility:** Resource allocation and business impact decisions
- **Escalation Trigger:** Extended outage or significant business impact
- **Timeframe:** 1 hour (SEV1), 4 hours (SEV2)
#### Business Escalation
**Customer Impact Assessment:**
- **High:** Revenue loss, SLA breaches, customer churn risk
- **Medium:** User experience degradation, support ticket volume
- **Low:** Internal tools, development impact only
**Escalation Matrix:**
| Severity | Duration | Business Escalation |
|----------|----------|-------------------|
| SEV1 | Immediate | VP Engineering |
| SEV1 | 30 minutes | CTO + Customer Success VP |
| SEV1 | 1 hour | CEO + Full Executive Team |
| SEV2 | 2 hours | VP Engineering |
| SEV2 | 4 hours | CTO |
| SEV3 | 1 business day | Engineering Manager |
### Status Page Management
#### Update Principles
1. **Transparency:** Provide factual information without speculation
2. **Timeliness:** Update within committed timeframes
3. **Clarity:** Use customer-friendly language, avoid technical jargon
4. **Completeness:** Include impact scope, status, and next update time
#### Status Categories
- **Operational:** All systems functioning normally
- **Degraded Performance:** Some users may experience slowness
- **Partial Outage:** Subset of features unavailable
- **Major Outage:** Service unavailable for most/all users
- **Under Maintenance:** Planned maintenance window
#### Update Template
```
{Timestamp} - {Status Category}
{Brief description of current state}
Impact: {who is affected and how}
Cause: {root cause if known, "under investigation" if not}
Resolution: {what's being done to fix it}
Next update: {specific time}
We apologize for any inconvenience this may cause.
```
### Action Item Framework
#### Action Item Categories
1. **Immediate Fixes**
- Critical bugs discovered during incident
- Security vulnerabilities exposed
- Data integrity issues
2. **Process Improvements**
- Communication gaps
- Escalation procedure updates
- Runbook additions/updates
3. **Technical Debt**
- Architecture improvements
- Monitoring enhancements
- Automation opportunities
4. **Organizational Changes**
- Team structure adjustments
- Training requirements
- Tool/platform investments
#### Action Item Template
```
**Title:** {Concise description of the action}
**Priority:** {Critical/High/Medium/Low}
**Category:** {Fix/Process/Technical/Organizational}
**Owner:** {Assigned person}
**Due Date:** {Specific date}
**Success Criteria:** {How will we know this is complete}
**Dependencies:** {What needs to happen first}
**Related PIRs:** {Links to other incidents this addresses}
**Description:**
{Detailed description of what needs to be done and why}
**Implementation Plan:**
1. {Step 1}
2. {Step 2}
3. {Validation step}
**Progress Updates:**
- {Date}: {Progress update}
- {Date}: {Progress update}
```
FILE:references/sla-management-guide.md
# SLA Management Guide
> Comprehensive reference for Service Level Agreements, Objectives, and Indicators.
> Designed for incident commanders who must understand, protect, and communicate SLA status during and after incidents.
---
## 1. Definitions & Relationships
### Service Level Indicator (SLI)
An SLI is the quantitative measurement of a specific aspect of service quality. SLIs are the raw data that feed everything above them. They must be precisely defined, automatically collected, and unambiguous.
**Common SLI types by service:**
| Service Type | SLI | Measurement Method |
|---|---|---|
| Web Application | Request latency (p50, p95, p99) | Server-side histogram |
| Web Application | Availability (successful responses / total requests) | Load balancer logs |
| REST API | Error rate (5xx responses / total responses) | API gateway metrics |
| REST API | Throughput (requests per second) | Counter metric |
| Database | Query latency (p99) | Slow query log + APM |
| Database | Replication lag (seconds) | Replica monitoring |
| Message Queue | End-to-end delivery latency | Timestamp comparison |
| Message Queue | Message loss rate | Producer vs consumer counts |
| Storage | Durability (objects lost / objects stored) | Integrity checksums |
| CDN | Cache hit ratio | Edge server logs |
**SLI specification formula:**
```
SLI = (good events / total events) x 100
```
For availability: `SLI = (successful requests / total requests) x 100`
For latency: `SLI = (requests faster than threshold / total requests) x 100`
### Service Level Objective (SLO)
An SLO is the target value or range for an SLI. It defines the acceptable level of reliability. SLOs are internal goals that engineering teams commit to.
**Setting meaningful SLOs:**
1. Measure the current baseline over 30 days minimum
2. Subtract a safety margin (typically 0.05%-0.1% below actual performance)
3. Validate against user expectations and business requirements
4. Never set an SLO higher than what the system can sustain without heroics
**Common pitfall:** Setting 99.99% availability when 99.9% meets every user need. The jump from 99.9% to 99.99% is a 10x reduction in allowed downtime and typically requires 3-5x the engineering investment.
**SLO examples:**
- `99.9% of HTTP requests return a non-5xx response within each calendar month`
- `95% of API requests complete in under 200ms (p95 latency)`
- `99.95% of messages are delivered within 30 seconds of production`
### Service Level Agreement (SLA)
An SLA is a formal contract between a service provider and its customers that specifies consequences for failing to meet defined service levels. SLAs must always be looser than SLOs to provide a buffer zone.
**Rule of thumb:** If your SLO is 99.95%, your SLA should be 99.9% or lower. The gap between SLO and SLA is your safety margin.
### The Hierarchy
```
SLA (99.9%) ← Contract with customers, financial penalties
↑ backs
SLO (99.95%) ← Internal target, triggers error budget policy
↑ targets
SLI (measured) ← Raw metric: actual uptime = 99.97% this month
```
**Standard combinations by tier:**
| Tier | SLI (Metric) | SLO (Target) | SLA (Contract) | Allowed Downtime/Month |
|---|---|---|---|---|
| Critical (payments) | Availability | 99.99% | 99.95% | SLO: 4.38 min / SLA: 21.9 min |
| High (core API) | Availability | 99.95% | 99.9% | SLO: 21.9 min / SLA: 43.8 min |
| Standard (dashboard) | Availability | 99.9% | 99.5% | SLO: 43.8 min / SLA: 3.65 hrs |
| Low (internal tools) | Availability | 99.5% | 99.0% | SLO: 3.65 hrs / SLA: 7.3 hrs |
---
## 2. Error Budget Policy
### What Is an Error Budget
An error budget is the maximum amount of unreliability a service can have within a given period while still meeting its SLO. It is calculated as:
```
Error Budget = 1 - SLO target
```
For a 99.9% SLO over a 30-day month (43,200 minutes):
```
Error Budget = 1 - 0.999 = 0.001 = 0.1%
Allowed Downtime = 43,200 x 0.001 = 43.2 minutes
```
### Downtime Allowances by SLO
| SLO | Error Budget | Monthly Downtime | Quarterly Downtime | Annual Downtime |
|---|---|---|---|---|
| 99.0% | 1.0% | 7 hrs 18 min | 21 hrs 54 min | 3 days 15 hrs |
| 99.5% | 0.5% | 3 hrs 39 min | 10 hrs 57 min | 1 day 19 hrs |
| 99.9% | 0.1% | 43.8 min | 2 hrs 11 min | 8 hrs 46 min |
| 99.95% | 0.05% | 21.9 min | 1 hr 6 min | 4 hrs 23 min |
| 99.99% | 0.01% | 4.38 min | 13.1 min | 52.6 min |
| 99.999% | 0.001% | 26.3 sec | 78.9 sec | 5.26 min |
### Error Budget Consumption Tracking
Track budget consumption as a percentage of the total budget used so far in the current window:
```
Budget Consumed (%) = (actual bad minutes / allowed bad minutes) x 100
```
Example: SLO is 99.9% (43.8 min budget/month). On day 10, you have had 15 minutes of downtime.
```
Budget Consumed = (15 / 43.8) x 100 = 34.2%
Expected consumption at day 10 = (10/30) x 100 = 33.3%
Status: Slightly over pace (34.2% consumed at 33.3% of month elapsed)
```
### Burn Rate
Burn rate measures how fast the error budget is being consumed relative to the steady-state rate:
```
Burn Rate = (error rate observed / error rate allowed by SLO)
```
A burn rate of 1.0 means the budget will be exactly exhausted by the end of the window. A burn rate of 10 means the budget will be exhausted in 1/10th of the window.
**Burn rate to time-to-exhaustion (30-day month):**
| Burn Rate | Budget Exhausted In | Urgency |
|---|---|---|
| 1x | 30 days | On pace, monitoring only |
| 2x | 15 days | Elevated attention |
| 6x | 5 days | Active investigation required |
| 14.4x | 2.08 days (~50 hours) | Immediate page |
| 36x | 20 hours | Critical, all-hands |
| 720x | 1 hour | Total outage scenario |
### Error Budget Exhaustion Policy
When the error budget is consumed, the following actions trigger based on threshold:
**Tier 1 - Budget at 75% consumed (Yellow):**
- Notify service team lead via automated alert
- Freeze non-critical deployments to the affected service
- Conduct pre-emptive review of upcoming changes for risk
- Increase monitoring sensitivity (lower alert thresholds)
**Tier 2 - Budget at 100% consumed (Orange):**
- Hard feature freeze on the affected service
- Mandatory reliability sprint: all engineering effort redirected to reliability
- Daily status updates to engineering leadership
- Postmortem required for the incidents that consumed the budget
- Freeze lasts until budget replenishes to 50% or systemic fixes are verified
**Tier 3 - Budget at 150% consumed / SLA breach imminent (Red):**
- Escalation to VP Engineering and CTO
- Cross-team war room if dependencies are involved
- Customer communication prepared and staged
- Legal and finance teams briefed on potential SLA credit obligations
- Recovery plan with specific milestones required within 24 hours
### Error Budget Policy Template
```
SERVICE: [service-name]
SLO: [target]% availability over [rolling 30-day / calendar month] window
ERROR BUDGET: [calculated] minutes per window
BUDGET THRESHOLDS:
- 50% consumed: Team notification, increased vigilance
- 75% consumed: Feature freeze for this service, reliability focus
- 100% consumed: Full feature freeze, reliability sprint mandatory
- SLA threshold crossed: Executive escalation, customer communication
REVIEW CADENCE: Monthly budget review on [day], quarterly SLO adjustment
EXCEPTIONS: Planned maintenance windows excluded if communicated 72+ hours in advance
and within agreed maintenance allowance.
APPROVED BY: [Engineering Lead] / [Product Lead] / [Date]
```
---
## 3. SLA Breach Handling
### Detection Methods
**Automated detection (primary):**
- Real-time monitoring dashboards with SLA burn-rate alerts
- Automated SLA compliance calculations running every 5 minutes
- Threshold-based alerts when cumulative downtime approaches SLA limits
- Synthetic monitoring (external probes) for customer-perspective validation
**Manual review (secondary):**
- Monthly SLA compliance reports generated on the 1st of each month
- Customer-reported incidents cross-referenced with internal metrics
- Quarterly audits comparing measured SLIs against contracted SLAs
- Discrepancy review between internal metrics and customer-perceived availability
### Breach Classification
**Minor Breach:**
- SLA missed by less than 0.05 percentage points (e.g., 99.85% vs 99.9% SLA)
- Fewer than 3 discrete incidents contributed
- No single incident exceeded 30 minutes
- Customer impact was limited or partial degradation only
- Financial credit: typically 5-10% of monthly service fee
**Major Breach:**
- SLA missed by 0.05 to 0.5 percentage points
- Extended outage of 1-4 hours in a single incident, or multiple significant incidents
- Clear customer impact with support tickets generated
- Financial credit: typically 10-25% of monthly service fee
**Critical Breach:**
- SLA missed by more than 0.5 percentage points
- Total outage exceeding 4 hours, or repeated major incidents in same window
- Data loss, security incident, or compliance violation involved
- Financial credit: typically 25-100% of monthly service fee
- May trigger contract termination clauses
### Response Protocol
**For Minor Breach (within 3 business days):**
1. Generate SLA compliance report with exact metrics
2. Document contributing incidents with root causes
3. Send proactive notification to customer success manager
4. Issue service credits if contractually required (do not wait for customer to ask)
5. File internal improvement ticket with 30-day remediation target
**For Major Breach (within 24 hours):**
1. Incident commander confirms SLA impact calculation
2. Draft customer communication (see template below)
3. Executive sponsor reviews and approves communication
4. Issue service credits with detailed breakdown
5. Schedule root cause review with customer within 5 business days
6. Produce remediation plan with committed timelines
**For Critical Breach (immediate):**
1. Activate executive escalation chain
2. Legal team reviews contractual exposure
3. Finance team calculates credit obligations
4. Customer communication from VP or C-level within 4 hours
5. Dedicated remediation task force assigned
6. Weekly status updates to customer until remediation complete
7. Formal postmortem document shared with customer within 10 business days
### Customer Communication Template
```
Subject: Service Level Update - [Service Name] - [Month Year]
Dear [Customer Name],
We are writing to inform you that [Service Name] did not meet the committed
service level of [SLA target]% availability during [time period].
MEASURED PERFORMANCE: [actual]% availability
COMMITTED SLA: [SLA target]% availability
SHORTFALL: [delta] percentage points
CONTRIBUTING FACTORS:
- [Date/Time]: [Brief description of incident] ([duration] impact)
- [Date/Time]: [Brief description of incident] ([duration] impact)
SERVICE CREDIT: In accordance with our agreement, a credit of [amount/percentage]
will be applied to your next invoice.
REMEDIATION ACTIONS:
1. [Specific technical fix with completion date]
2. [Process improvement with implementation date]
3. [Monitoring enhancement with deployment date]
We take our service commitments seriously. [Name], [Title] is personally
overseeing the remediation and is available to discuss further at your convenience.
Sincerely,
[Name, Title]
```
### Legal and Compliance Considerations
- Maintain auditable records of all SLA measurements for the full contract term plus 2 years
- SLA calculations must use the measurement methodology defined in the contract, not internal approximations
- Force majeure clauses typically exclude natural disasters, but verify per contract
- Planned maintenance exclusions must match the exact notification procedures in the contract
- Multi-region SLAs may have separate calculations per region; verify aggregation method
---
## 4. Incident-to-SLA Mapping
### Downtime Calculation Methodologies
**Full outage:** Service completely unavailable. Every minute counts as a full minute of downtime.
```
Downtime = End Time - Start Time (in minutes)
```
**Partial degradation:** Service available but impaired. Apply a degradation factor:
```
Effective Downtime = Actual Duration x Degradation Factor
```
| Degradation Level | Factor | Description |
|---|---|---|
| Complete outage | 1.0 | Service fully unavailable |
| Severe degradation | 0.75 | >50% of requests failing or >10x latency |
| Moderate degradation | 0.5 | 10-50% of requests affected or 3-10x latency |
| Minor degradation | 0.25 | <10% of requests affected or <3x latency increase |
| Cosmetic / non-functional | 0.0 | No impact on core SLI metrics |
**Note:** The exact degradation factors must be agreed upon in the SLA contract. The above are industry-standard starting points.
### Planned vs Unplanned Downtime
Most SLAs exclude pre-announced maintenance windows from availability calculations, subject to conditions:
- Notification provided N hours/days in advance (commonly 72 hours)
- Maintenance occurs within an agreed window (e.g., Sunday 02:00-06:00 UTC)
- Total planned downtime does not exceed the monthly maintenance allowance (e.g., 4 hours/month)
- Any overrun beyond the planned window counts as unplanned downtime
```
SLA Availability = (Total Minutes - Excluded Maintenance - Unplanned Downtime) / (Total Minutes - Excluded Maintenance) x 100
```
### Multi-Service SLA Composition
When a customer-facing product depends on multiple services, composite SLA is calculated as:
**Serial dependency (all must be up):**
```
Composite SLA = SLA_A x SLA_B x SLA_C
Example: 99.9% x 99.95% x 99.99% = 99.84%
```
**Parallel / redundant (any one must be up):**
```
Composite Availability = 1 - ((1 - SLA_A) x (1 - SLA_B))
Example: 1 - ((1 - 0.999) x (1 - 0.999)) = 1 - 0.000001 = 99.9999%
```
This is critical during incidents: an outage in a shared dependency may breach SLAs for multiple customer-facing products simultaneously.
### Worked Examples
**Example 1: Simple outage**
- Service: Core API (SLA: 99.9%)
- Month: 30 days = 43,200 minutes
- Incident: Full outage from 14:23 to 14:38 UTC on the 12th (15 minutes)
- No other incidents this month
```
Availability = (43,200 - 15) / 43,200 x 100 = 99.965%
SLA Status: PASS (99.965% > 99.9%)
Error Budget Consumed: 15 / 43.2 = 34.7%
```
**Example 2: Partial degradation**
- Service: Payment Processing (SLA: 99.95%)
- Month: 30 days = 43,200 minutes
- Incident: 50% of transactions failing for 4 hours (240 minutes)
- Degradation factor: 0.5 (moderate - 50% of requests affected)
```
Effective Downtime = 240 x 0.5 = 120 minutes
Availability = (43,200 - 120) / 43,200 x 100 = 99.722%
SLA Status: FAIL (99.722% < 99.95%)
Shortfall: 0.228 percentage points → Major Breach
```
**Example 3: Multiple incidents**
- Service: Dashboard (SLA: 99.5%)
- Month: 31 days = 44,640 minutes
- Incident A: 45-minute full outage on the 5th
- Incident B: 2-hour severe degradation (factor 0.75) on the 18th
- Incident C: 30-minute full outage on the 25th
```
Total Effective Downtime = 45 + (120 x 0.75) + 30 = 45 + 90 + 30 = 165 minutes
Availability = (44,640 - 165) / 44,640 x 100 = 99.630%
SLA Status: PASS (99.630% > 99.5%)
Error Budget Consumed: 165 / 223.2 = 73.9% → Yellow threshold, feature freeze recommended
```
---
## 5. SLO Best Practices
### Start with User Journeys
Do not set SLOs based on infrastructure metrics. Start from what users experience:
1. Identify critical user journeys (e.g., "User completes checkout")
2. Map each journey to the services and dependencies involved
3. Define what "good" looks like for each journey (fast, error-free, complete)
4. Select the SLIs that most directly measure that user experience
5. Set SLO targets that reflect the minimum acceptable user experience
A database with 99.99% uptime is meaningless if the API in front of it has a bug causing 5% error rates.
### The Four Golden Signals as SLI Sources
From Google SRE, the four golden signals provide comprehensive service health:
| Signal | SLI Example | Typical SLO |
|---|---|---|
| Latency | p99 request duration < 500ms | 99% of requests under threshold |
| Traffic | Requests per second | N/A (capacity planning, not SLO) |
| Errors | 5xx rate as % of total requests | < 0.1% error rate over rolling window |
| Saturation | CPU/memory/queue depth | < 80% utilization (capacity SLI) |
For most services, latency and error rate are the two most important SLIs to back with SLOs.
### Setting SLO Targets
1. Collect 90 days of historical SLI data
2. Calculate the 5th percentile performance (worst 5% of days)
3. Set SLO slightly above that baseline (this ensures the SLO is achievable without heroics)
4. Validate: would a breach at this level actually impact users negatively?
5. Adjust upward only if user impact analysis demands it
**Never set SLOs by aspiration.** A 99.99% SLO on a service that has historically achieved 99.93% is a guaranteed source of perpetual firefighting with no reliability improvement.
### Review Cadence
- **Weekly:** Review current error budget burn rate, flag services approaching thresholds
- **Monthly:** Full SLO compliance review, adjust alert thresholds if needed
- **Quarterly:** Reassess SLO targets based on 90-day data, review SLA contract alignment
- **Annually:** Strategic SLO review tied to product roadmap and infrastructure investments
### Anti-Patterns
| Anti-Pattern | Problem | Fix |
|---|---|---|
| Vanity SLOs | Setting 99.99% to impress, then ignoring breaches | Set achievable targets, enforce budget policy |
| SLO Inflation | Ratcheting SLOs up whenever performance is good | Only increase SLOs when users demonstrably need it |
| Unmeasured SLAs | Committing contractual SLAs without actual SLI measurement | Instrument SLIs before signing SLA contracts |
| Copy-Paste SLOs | Same SLO for every service regardless of criticality | Tier services by business impact, set SLOs accordingly |
| Ignoring Dependencies | Setting aggressive SLOs without accounting for dependency reliability | Calculate composite SLA; your SLO cannot exceed dependency chain |
| Alert-Free SLOs | Having SLOs but no automated alerting on budget consumption | Every SLO must have corresponding burn rate alerts |
---
## 6. Monitoring & Alerting for SLAs
### Multi-Window Burn Rate Alerting
The Google SRE approach uses multiple time windows to balance speed of detection against alert noise. Each alert condition requires both a short window (for speed) and a long window (for confirmation):
**Alert configuration matrix:**
| Severity | Short Window | Short Threshold | Long Window | Long Threshold | Action |
|---|---|---|---|---|---|
| Critical (Page) | 1 hour | > 14.4x burn rate | 5 minutes | > 14.4x burn rate | Wake someone up |
| High (Page) | 6 hours | > 6x burn rate | 30 minutes | > 6x burn rate | Page on-call within 30 min |
| Medium (Ticket) | 3 days | > 1x burn rate | 6 hours | > 1x burn rate | Create ticket, next business day |
**Why these specific numbers:**
- 14.4x burn rate over 1 hour consumes 2% of monthly budget in that hour. At this rate, the entire 30-day budget is gone in ~50 hours. This demands immediate human attention.
- 6x burn rate over 6 hours consumes 5% of monthly budget. The budget will be exhausted in 5 days. Urgent but not wake-up-at-3am urgent.
- 1x burn rate over 3 days means you are on pace to exactly exhaust the budget. This needs investigation but is not an emergency.
### Burn Rate Alert Formulas
For a given time window, calculate the burn rate:
```
burn_rate = (error_count_in_window / request_count_in_window) / (1 - SLO_target)
```
Example for a 99.9% SLO, observing 50 errors out of 10,000 requests in a 1-hour window:
```
observed_error_rate = 50 / 10,000 = 0.005 (0.5%)
allowed_error_rate = 1 - 0.999 = 0.001 (0.1%)
burn_rate = 0.005 / 0.001 = 5.0
```
A burn rate of 5.0 means the error budget is being consumed 5 times faster than the sustainable rate.
### Alert Severity to SLA Risk Mapping
| Burn Rate | Budget Impact | SLA Risk | Response |
|---|---|---|---|
| < 1x | Under budget pace | None | Routine monitoring |
| 1x - 3x | On pace or slightly over | Low | Investigate next business day |
| 3x - 6x | Budget will exhaust in 5-10 days | Moderate | Investigate within 4 hours |
| 6x - 14.4x | Budget will exhaust in 2-5 days | High | Page on-call, respond in 30 min |
| > 14.4x | Budget will exhaust in < 2 days | Critical | Immediate page, incident declared |
| > 100x | Active major outage | SLA breach imminent | All-hands incident response |
### Dashboard Design for SLA Tracking
Every SLA-tracked service should have a dashboard with these panels:
**Row 1 - Current Status:**
- Current availability (real-time, rolling 5-minute window)
- Current error rate (real-time)
- Current p99 latency (real-time)
**Row 2 - Budget Status:**
- Error budget remaining (% of monthly budget, gauge visualization)
- Budget consumption timeline (line chart, actual vs expected burn)
- Budget burn rate (current 1h, 6h, and 3d burn rates)
**Row 3 - Historical Context:**
- 30-day availability trend (daily granularity)
- SLA compliance status for current and previous 3 months
- Incident markers overlaid on availability timeline
**Row 4 - Dependencies:**
- Upstream dependency availability (services this service depends on)
- Downstream impact scope (services that depend on this service)
- Composite SLA calculation for customer-facing products
### Alert Fatigue Prevention
Alert fatigue is the primary reason SLA monitoring fails in practice. Mitigation strategies:
1. **Require dual-window confirmation.** Never page on a single short window. Always require both the short window (for speed) and long window (for persistence) to fire simultaneously.
2. **Separate page-worthy from ticket-worthy.** Only two conditions should wake someone up: >14.4x burn rate sustained, or >6x burn rate sustained. Everything else is a ticket.
3. **Deduplicate aggressively.** If the same service triggers both a latency and error rate alert for the same underlying issue, group them into a single notification.
4. **Auto-resolve.** Alerts must auto-resolve when the burn rate drops below threshold. Never leave stale alerts open.
5. **Review alert quality monthly.** Track the ratio of actionable alerts to total alerts. Target >80% actionable rate. If an alert fires and no human action is needed, tune or remove it.
6. **Escalation, not repetition.** If an alert is not acknowledged within the response window, escalate to the next tier. Do not re-send the same alert every 5 minutes.
### Practical Monitoring Stack
| Layer | Tool Category | Purpose |
|---|---|---|
| Collection | Prometheus, OpenTelemetry, StatsD | Gather SLI metrics from services |
| Storage | Prometheus TSDB, Thanos, Mimir | Retain metrics for SLO window + 90 days |
| Calculation | Prometheus recording rules, Sloth | Pre-compute burn rates and budget consumption |
| Alerting | Alertmanager, PagerDuty, OpsGenie | Route alerts by severity and schedule |
| Visualization | Grafana, Datadog | Dashboards for real-time and historical SLA views |
| Reporting | Custom scripts, SLO generators | Monthly SLA compliance reports for customers |
**Retention requirement:** SLI data must be retained for at least the SLA reporting period (typically monthly or quarterly) plus a 90-day dispute window. Annual SLA reviews require 12 months of data at daily granularity minimum.
---
*Last updated: February 2026*
*For use with: incident-commander skill*
*Maintainer: Engineering Team*
FILE:scripts/incident_classifier.py
#!/usr/bin/env python3
"""
Incident Classifier
Analyzes incident descriptions and outputs severity levels, recommended response teams,
initial actions, and communication templates.
This tool uses pattern matching and keyword analysis to classify incidents according to
SEV1-4 criteria and provide structured response guidance.
Usage:
python incident_classifier.py --input incident.json
echo "Database is down" | python incident_classifier.py --format text
python incident_classifier.py --interactive
"""
import argparse
import json
import sys
import re
from datetime import datetime, timezone
from typing import Dict, List, Tuple, Optional, Any
class IncidentClassifier:
"""
Classifies incidents based on description, impact metrics, and business context.
Provides severity assessment, team recommendations, and response templates.
"""
def __init__(self):
"""Initialize the classifier with rules and templates."""
self.severity_rules = self._load_severity_rules()
self.team_mappings = self._load_team_mappings()
self.communication_templates = self._load_communication_templates()
self.action_templates = self._load_action_templates()
def _load_severity_rules(self) -> Dict[str, Dict]:
"""Load severity classification rules and keywords."""
return {
"sev1": {
"keywords": [
"down", "outage", "offline", "unavailable", "crashed", "failed",
"critical", "emergency", "dead", "broken", "timeout", "500 error",
"data loss", "corrupted", "breach", "security incident",
"revenue impact", "customer facing", "all users", "complete failure"
],
"impact_indicators": [
"100%", "all users", "entire service", "complete",
"revenue loss", "sla violation", "customer churn",
"security breach", "data corruption", "regulatory"
],
"duration_threshold": 0, # Immediate classification
"response_time": 300, # 5 minutes
"description": "Complete service failure affecting all users or critical business functions"
},
"sev2": {
"keywords": [
"degraded", "slow", "performance", "errors", "partial",
"intermittent", "high latency", "timeouts", "some users",
"feature broken", "api errors", "database slow"
],
"impact_indicators": [
"50%", "25-75%", "many users", "significant",
"performance degradation", "feature unavailable",
"support tickets", "user complaints"
],
"duration_threshold": 300, # 5 minutes
"response_time": 900, # 15 minutes
"description": "Significant degradation affecting subset of users or non-critical functions"
},
"sev3": {
"keywords": [
"minor", "cosmetic", "single feature", "workaround available",
"edge case", "rare issue", "non-critical", "internal tool",
"logging issue", "monitoring gap"
],
"impact_indicators": [
"<25%", "few users", "limited impact",
"workaround exists", "internal only",
"development environment"
],
"duration_threshold": 3600, # 1 hour
"response_time": 7200, # 2 hours
"description": "Limited impact with workarounds available"
},
"sev4": {
"keywords": [
"cosmetic", "documentation", "typo", "minor bug",
"enhancement", "nice to have", "low priority",
"test environment", "dev tools"
],
"impact_indicators": [
"no impact", "cosmetic only", "documentation",
"development", "testing", "non-production"
],
"duration_threshold": 86400, # 24 hours
"response_time": 172800, # 2 days
"description": "Minimal impact, cosmetic issues, or planned maintenance"
}
}
def _load_team_mappings(self) -> Dict[str, List[str]]:
"""Load team assignment rules based on service/component keywords."""
return {
"database": ["Database Team", "SRE", "Backend Engineering"],
"frontend": ["Frontend Team", "UX Engineering", "Product Engineering"],
"api": ["API Team", "Backend Engineering", "Platform Team"],
"infrastructure": ["SRE", "DevOps", "Platform Team"],
"security": ["Security Team", "SRE", "Compliance Team"],
"network": ["Network Engineering", "SRE", "Infrastructure Team"],
"authentication": ["Identity Team", "Security Team", "Backend Engineering"],
"payment": ["Payments Team", "Finance Engineering", "Compliance Team"],
"mobile": ["Mobile Team", "API Team", "QA Engineering"],
"monitoring": ["SRE", "Platform Team", "DevOps"],
"deployment": ["DevOps", "Release Engineering", "SRE"],
"data": ["Data Engineering", "Analytics Team", "Backend Engineering"]
}
def _load_communication_templates(self) -> Dict[str, Dict]:
"""Load communication templates for each severity level."""
return {
"sev1": {
"subject": "🚨 [SEV1] {service} - {brief_description}",
"body": """CRITICAL INCIDENT ALERT
Incident Details:
- Start Time: {timestamp}
- Severity: SEV1 - Critical Outage
- Service: {service}
- Impact: {impact_description}
- Current Status: Investigating
Customer Impact:
{customer_impact}
Response Team:
- Incident Commander: TBD (assigning now)
- Primary Responder: {primary_responder}
- SMEs Required: {subject_matter_experts}
Immediate Actions Taken:
{initial_actions}
War Room: {war_room_link}
Status Page: Will be updated within 15 minutes
Next Update: {next_update_time}
This is a customer-impacting incident requiring immediate attention.
{incident_commander_contact}"""
},
"sev2": {
"subject": "⚠️ [SEV2] {service} - {brief_description}",
"body": """MAJOR INCIDENT NOTIFICATION
Incident Details:
- Start Time: {timestamp}
- Severity: SEV2 - Major Impact
- Service: {service}
- Impact: {impact_description}
- Current Status: Investigating
User Impact:
{customer_impact}
Response Team:
- Primary Responder: {primary_responder}
- Supporting Team: {supporting_teams}
- Incident Commander: {incident_commander}
Initial Assessment:
{initial_assessment}
Next Steps:
{next_steps}
Updates will be provided every 30 minutes.
Status page: {status_page_link}
{contact_information}"""
},
"sev3": {
"subject": "ℹ️ [SEV3] {service} - {brief_description}",
"body": """MINOR INCIDENT NOTIFICATION
Incident Details:
- Start Time: {timestamp}
- Severity: SEV3 - Minor Impact
- Service: {service}
- Impact: {impact_description}
- Status: {current_status}
Details:
{incident_details}
Assigned Team: {assigned_team}
Estimated Resolution: {eta}
Workaround: {workaround}
This incident has limited customer impact and is being addressed during normal business hours.
{team_contact}"""
},
"sev4": {
"subject": "[SEV4] {service} - {brief_description}",
"body": """LOW PRIORITY ISSUE
Issue Details:
- Reported: {timestamp}
- Severity: SEV4 - Low Impact
- Component: {service}
- Description: {description}
This issue will be addressed in the normal development cycle.
Assigned to: {assigned_team}
Target Resolution: {target_date}
{standard_contact}"""
}
}
def _load_action_templates(self) -> Dict[str, List[Dict]]:
"""Load initial action templates for each severity level."""
return {
"sev1": [
{
"action": "Establish incident command",
"priority": 1,
"timeout_minutes": 5,
"description": "Page incident commander and establish war room"
},
{
"action": "Create incident ticket",
"priority": 1,
"timeout_minutes": 2,
"description": "Create tracking ticket with all known details"
},
{
"action": "Update status page",
"priority": 2,
"timeout_minutes": 15,
"description": "Post initial status page update acknowledging incident"
},
{
"action": "Notify executives",
"priority": 2,
"timeout_minutes": 15,
"description": "Alert executive team of customer-impacting outage"
},
{
"action": "Engage subject matter experts",
"priority": 3,
"timeout_minutes": 10,
"description": "Page relevant SMEs based on affected systems"
},
{
"action": "Begin technical investigation",
"priority": 3,
"timeout_minutes": 5,
"description": "Start technical diagnosis and mitigation efforts"
}
],
"sev2": [
{
"action": "Assign incident commander",
"priority": 1,
"timeout_minutes": 30,
"description": "Assign IC and establish coordination channel"
},
{
"action": "Create incident tracking",
"priority": 1,
"timeout_minutes": 5,
"description": "Create incident ticket with details and timeline"
},
{
"action": "Assess customer impact",
"priority": 2,
"timeout_minutes": 15,
"description": "Determine scope and severity of user impact"
},
{
"action": "Engage response team",
"priority": 2,
"timeout_minutes": 30,
"description": "Page appropriate technical responders"
},
{
"action": "Begin investigation",
"priority": 3,
"timeout_minutes": 15,
"description": "Start technical analysis and debugging"
},
{
"action": "Plan status communication",
"priority": 3,
"timeout_minutes": 30,
"description": "Determine if status page update is needed"
}
],
"sev3": [
{
"action": "Assign to appropriate team",
"priority": 1,
"timeout_minutes": 120,
"description": "Route to team with relevant expertise"
},
{
"action": "Create tracking ticket",
"priority": 1,
"timeout_minutes": 30,
"description": "Document issue in standard ticketing system"
},
{
"action": "Assess scope and impact",
"priority": 2,
"timeout_minutes": 60,
"description": "Understand full scope of the issue"
},
{
"action": "Identify workarounds",
"priority": 2,
"timeout_minutes": 60,
"description": "Find temporary solutions if possible"
},
{
"action": "Plan resolution approach",
"priority": 3,
"timeout_minutes": 120,
"description": "Develop plan for permanent fix"
}
],
"sev4": [
{
"action": "Create backlog item",
"priority": 1,
"timeout_minutes": 1440, # 24 hours
"description": "Add to team backlog for future sprint planning"
},
{
"action": "Triage and prioritize",
"priority": 2,
"timeout_minutes": 2880, # 2 days
"description": "Review and prioritize against other work"
},
{
"action": "Assign owner",
"priority": 3,
"timeout_minutes": 4320, # 3 days
"description": "Assign to appropriate developer when capacity allows"
}
]
}
def classify_incident(self, incident_data: Dict[str, Any]) -> Dict[str, Any]:
"""
Main classification method that analyzes incident data and returns
comprehensive response recommendations.
Args:
incident_data: Dictionary containing incident information
Returns:
Dictionary with classification results and recommendations
"""
# Extract key information from incident data
description = incident_data.get('description', '').lower()
affected_users = incident_data.get('affected_users', '0%')
business_impact = incident_data.get('business_impact', 'unknown')
service = incident_data.get('service', 'unknown service')
duration = incident_data.get('duration_minutes', 0)
# Classify severity
severity = self._classify_severity(description, affected_users, business_impact, duration)
# Determine response teams
response_teams = self._determine_teams(description, service)
# Generate initial actions
initial_actions = self._generate_initial_actions(severity, incident_data)
# Create communication template
communication = self._generate_communication(severity, incident_data)
# Calculate response timeline
timeline = self._generate_timeline(severity)
# Determine escalation path
escalation = self._determine_escalation(severity, business_impact)
return {
"classification": {
"severity": severity.upper(),
"confidence": self._calculate_confidence(description, affected_users, business_impact),
"reasoning": self._explain_classification(severity, description, affected_users),
"timestamp": datetime.now(timezone.utc).isoformat()
},
"response": {
"primary_team": response_teams[0] if response_teams else "General Engineering",
"supporting_teams": response_teams[1:] if len(response_teams) > 1 else [],
"all_teams": response_teams,
"response_time_minutes": self.severity_rules[severity]["response_time"] // 60
},
"initial_actions": initial_actions,
"communication": communication,
"timeline": timeline,
"escalation": escalation,
"incident_data": {
"service": service,
"description": incident_data.get('description', ''),
"affected_users": affected_users,
"business_impact": business_impact,
"duration_minutes": duration
}
}
def _classify_severity(self, description: str, affected_users: str,
business_impact: str, duration: int) -> str:
"""Classify incident severity based on multiple factors."""
scores = {"sev1": 0, "sev2": 0, "sev3": 0, "sev4": 0}
# Keyword analysis
for severity, rules in self.severity_rules.items():
for keyword in rules["keywords"]:
if keyword in description:
scores[severity] += 2
for indicator in rules["impact_indicators"]:
if indicator.lower() in description or indicator.lower() in affected_users.lower():
scores[severity] += 3
# Business impact weighting
if business_impact.lower() in ['critical', 'high', 'severe']:
scores["sev1"] += 5
scores["sev2"] += 3
elif business_impact.lower() in ['medium', 'moderate']:
scores["sev2"] += 3
scores["sev3"] += 2
elif business_impact.lower() in ['low', 'minimal']:
scores["sev3"] += 2
scores["sev4"] += 3
# User impact analysis
if '%' in affected_users:
try:
percentage = float(re.findall(r'\d+', affected_users)[0])
if percentage >= 75:
scores["sev1"] += 4
elif percentage >= 25:
scores["sev2"] += 4
elif percentage >= 5:
scores["sev3"] += 3
else:
scores["sev4"] += 2
except (IndexError, ValueError):
pass
# Duration consideration
if duration > 0:
if duration >= 3600: # 1 hour
scores["sev1"] += 2
scores["sev2"] += 1
elif duration >= 1800: # 30 minutes
scores["sev2"] += 2
scores["sev3"] += 1
# Return highest scoring severity
return max(scores, key=scores.get)
def _determine_teams(self, description: str, service: str) -> List[str]:
"""Determine which teams should respond based on affected systems."""
teams = set()
text_to_analyze = f"{description} {service}".lower()
for component, team_list in self.team_mappings.items():
if component in text_to_analyze:
teams.update(team_list)
# Default teams if no specific match
if not teams:
teams = {"General Engineering", "SRE"}
return list(teams)
def _generate_initial_actions(self, severity: str, incident_data: Dict) -> List[Dict]:
"""Generate prioritized initial actions based on severity."""
base_actions = self.action_templates[severity].copy()
# Customize actions based on incident details
for action in base_actions:
if severity in ["sev1", "sev2"]:
action["urgency"] = "immediate" if severity == "sev1" else "high"
else:
action["urgency"] = "normal" if severity == "sev3" else "low"
return base_actions
def _generate_communication(self, severity: str, incident_data: Dict) -> Dict:
"""Generate communication template filled with incident data."""
template = self.communication_templates[severity]
# Fill template with incident data
now = datetime.now(timezone.utc)
service = incident_data.get('service', 'Unknown Service')
description = incident_data.get('description', 'Incident detected')
communication = {
"subject": template["subject"].format(
service=service,
brief_description=description[:50] + "..." if len(description) > 50 else description
),
"body": template["body"],
"urgency": severity,
"recipients": self._determine_recipients(severity),
"channels": self._determine_channels(severity),
"frequency_minutes": self._get_update_frequency(severity)
}
return communication
def _generate_timeline(self, severity: str) -> Dict:
"""Generate expected response timeline."""
rules = self.severity_rules[severity]
now = datetime.now(timezone.utc)
milestones = []
if severity == "sev1":
milestones = [
{"milestone": "Incident Commander assigned", "minutes": 5},
{"milestone": "War room established", "minutes": 10},
{"milestone": "Initial status page update", "minutes": 15},
{"milestone": "Executive notification", "minutes": 15},
{"milestone": "First customer update", "minutes": 30}
]
elif severity == "sev2":
milestones = [
{"milestone": "Response team assembled", "minutes": 15},
{"milestone": "Initial assessment complete", "minutes": 30},
{"milestone": "Stakeholder notification", "minutes": 60},
{"milestone": "Status page update (if needed)", "minutes": 60}
]
elif severity == "sev3":
milestones = [
{"milestone": "Team assignment", "minutes": 120},
{"milestone": "Initial triage complete", "minutes": 240},
{"milestone": "Resolution plan created", "minutes": 480}
]
else: # sev4
milestones = [
{"milestone": "Backlog creation", "minutes": 1440},
{"milestone": "Priority assessment", "minutes": 2880}
]
return {
"response_time_minutes": rules["response_time"] // 60,
"milestones": milestones,
"update_frequency_minutes": self._get_update_frequency(severity)
}
def _determine_escalation(self, severity: str, business_impact: str) -> Dict:
"""Determine escalation requirements and triggers."""
escalation_rules = {
"sev1": {
"immediate": ["Incident Commander", "Engineering Manager"],
"15_minutes": ["VP Engineering", "Customer Success"],
"30_minutes": ["CTO"],
"60_minutes": ["CEO", "All C-Suite"],
"triggers": ["Extended outage", "Revenue impact", "Media attention"]
},
"sev2": {
"immediate": ["Team Lead", "On-call Engineer"],
"30_minutes": ["Engineering Manager"],
"120_minutes": ["VP Engineering"],
"triggers": ["No progress", "Expanding scope", "Customer escalation"]
},
"sev3": {
"immediate": ["Assigned Engineer"],
"240_minutes": ["Team Lead"],
"triggers": ["Issue complexity", "Multiple teams needed"]
},
"sev4": {
"immediate": ["Product Owner"],
"triggers": ["Customer request", "Stakeholder priority"]
}
}
return escalation_rules.get(severity, escalation_rules["sev4"])
def _determine_recipients(self, severity: str) -> List[str]:
"""Determine who should receive notifications."""
recipients = {
"sev1": ["on-call", "engineering-leadership", "executives", "customer-success"],
"sev2": ["on-call", "engineering-leadership", "product-team"],
"sev3": ["assigned-team", "team-lead"],
"sev4": ["assigned-engineer"]
}
return recipients.get(severity, recipients["sev4"])
def _determine_channels(self, severity: str) -> List[str]:
"""Determine communication channels to use."""
channels = {
"sev1": ["pager", "phone", "slack", "email", "status-page"],
"sev2": ["pager", "slack", "email"],
"sev3": ["slack", "email"],
"sev4": ["ticket-system"]
}
return channels.get(severity, channels["sev4"])
def _get_update_frequency(self, severity: str) -> int:
"""Get recommended update frequency in minutes."""
frequencies = {"sev1": 15, "sev2": 30, "sev3": 240, "sev4": 0}
return frequencies.get(severity, 0)
def _calculate_confidence(self, description: str, affected_users: str, business_impact: str) -> float:
"""Calculate confidence score for the classification."""
confidence = 0.5 # Base confidence
# Higher confidence with more specific information
if '%' in affected_users and any(char.isdigit() for char in affected_users):
confidence += 0.2
if business_impact.lower() in ['critical', 'high', 'medium', 'low']:
confidence += 0.15
if len(description.split()) > 5: # Detailed description
confidence += 0.15
return min(confidence, 1.0)
def _explain_classification(self, severity: str, description: str, affected_users: str) -> str:
"""Provide explanation for the classification decision."""
rules = self.severity_rules[severity]
matched_keywords = []
for keyword in rules["keywords"]:
if keyword in description.lower():
matched_keywords.append(keyword)
explanation = f"Classified as {severity.upper()} based on: "
reasons = []
if matched_keywords:
reasons.append(f"keywords: {', '.join(matched_keywords[:3])}")
if '%' in affected_users:
reasons.append(f"user impact: {affected_users}")
if not reasons:
reasons.append("default classification based on available information")
return explanation + "; ".join(reasons)
def format_json_output(result: Dict) -> str:
"""Format result as pretty JSON."""
return json.dumps(result, indent=2, ensure_ascii=False)
def format_text_output(result: Dict) -> str:
"""Format result as human-readable text."""
classification = result["classification"]
response = result["response"]
actions = result["initial_actions"]
communication = result["communication"]
output = []
output.append("=" * 60)
output.append("INCIDENT CLASSIFICATION REPORT")
output.append("=" * 60)
output.append("")
# Classification section
output.append("CLASSIFICATION:")
output.append(f" Severity: {classification['severity']}")
output.append(f" Confidence: {classification['confidence']:.1%}")
output.append(f" Reasoning: {classification['reasoning']}")
output.append(f" Timestamp: {classification['timestamp']}")
output.append("")
# Response section
output.append("RECOMMENDED RESPONSE:")
output.append(f" Primary Team: {response['primary_team']}")
if response['supporting_teams']:
output.append(f" Supporting Teams: {', '.join(response['supporting_teams'])}")
output.append(f" Response Time: {response['response_time_minutes']} minutes")
output.append("")
# Actions section
output.append("INITIAL ACTIONS:")
for i, action in enumerate(actions[:5], 1): # Show first 5 actions
output.append(f" {i}. {action['action']} (Priority {action['priority']})")
output.append(f" Timeout: {action['timeout_minutes']} minutes")
output.append(f" {action['description']}")
output.append("")
# Communication section
output.append("COMMUNICATION:")
output.append(f" Subject: {communication['subject']}")
output.append(f" Urgency: {communication['urgency'].upper()}")
output.append(f" Recipients: {', '.join(communication['recipients'])}")
output.append(f" Channels: {', '.join(communication['channels'])}")
if communication['frequency_minutes'] > 0:
output.append(f" Update Frequency: Every {communication['frequency_minutes']} minutes")
output.append("")
output.append("=" * 60)
return "\n".join(output)
def parse_input_text(text: str) -> Dict[str, Any]:
"""Parse free-form text input into structured incident data."""
# Basic parsing - in a real system, this would be more sophisticated
incident_data = {
"description": text.strip(),
"service": "unknown service",
"affected_users": "unknown",
"business_impact": "unknown"
}
# Try to extract service name
service_patterns = [
r'(?:service|api|database|server|application)\s+(\w+)',
r'(\w+)(?:\s+(?:is|has|service|api|database))',
r'(?:^|\s)(\w+)\s+(?:down|failed|broken)'
]
for pattern in service_patterns:
match = re.search(pattern, text.lower())
if match:
incident_data["service"] = match.group(1)
break
# Try to extract user impact
impact_patterns = [
r'(\d+%)\s+(?:of\s+)?(?:users?|customers?)',
r'(?:all|every|100%)\s+(?:users?|customers?)',
r'(?:some|many|several)\s+(?:users?|customers?)'
]
for pattern in impact_patterns:
match = re.search(pattern, text.lower())
if match:
incident_data["affected_users"] = match.group(1) if match.group(1) else match.group(0)
break
# Try to infer business impact
if any(word in text.lower() for word in ['critical', 'urgent', 'emergency', 'down', 'outage']):
incident_data["business_impact"] = "high"
elif any(word in text.lower() for word in ['slow', 'degraded', 'performance']):
incident_data["business_impact"] = "medium"
elif any(word in text.lower() for word in ['minor', 'cosmetic', 'small']):
incident_data["business_impact"] = "low"
return incident_data
def interactive_mode():
"""Run in interactive mode, prompting user for input."""
classifier = IncidentClassifier()
print("🚨 Incident Classifier - Interactive Mode")
print("=" * 50)
print("Enter incident details (or 'quit' to exit):")
print()
while True:
try:
description = input("Incident description: ").strip()
if description.lower() in ['quit', 'exit', 'q']:
break
if not description:
print("Please provide an incident description.")
continue
service = input("Affected service (optional): ").strip() or "unknown"
affected_users = input("Affected users (e.g., '50%', 'all users'): ").strip() or "unknown"
business_impact = input("Business impact (high/medium/low): ").strip() or "unknown"
incident_data = {
"description": description,
"service": service,
"affected_users": affected_users,
"business_impact": business_impact
}
result = classifier.classify_incident(incident_data)
print("\n" + "=" * 50)
print(format_text_output(result))
print("=" * 50)
print()
except KeyboardInterrupt:
print("\n\nExiting...")
break
except Exception as e:
print(f"Error: {e}")
def main():
"""Main function with argument parsing and execution."""
parser = argparse.ArgumentParser(
description="Classify incidents and provide response recommendations",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
python incident_classifier.py --input incident.json
echo "Database is down" | python incident_classifier.py --format text
python incident_classifier.py --interactive
Input JSON format:
{
"description": "Database connection timeouts",
"service": "user-service",
"affected_users": "80%",
"business_impact": "high"
}
"""
)
parser.add_argument(
"--input", "-i",
help="Input file path (JSON format) or '-' for stdin"
)
parser.add_argument(
"--format", "-f",
choices=["json", "text"],
default="json",
help="Output format (default: json)"
)
parser.add_argument(
"--interactive",
action="store_true",
help="Run in interactive mode"
)
parser.add_argument(
"--output", "-o",
help="Output file path (default: stdout)"
)
args = parser.parse_args()
# Interactive mode
if args.interactive:
interactive_mode()
return
classifier = IncidentClassifier()
try:
# Read input
if args.input == "-" or (not args.input and not sys.stdin.isatty()):
# Read from stdin
input_text = sys.stdin.read().strip()
if not input_text:
parser.error("No input provided")
# Try to parse as JSON first, then as text
try:
incident_data = json.loads(input_text)
except json.JSONDecodeError:
incident_data = parse_input_text(input_text)
elif args.input:
# Read from file
with open(args.input, 'r') as f:
incident_data = json.load(f)
else:
parser.error("No input specified. Use --input, --interactive, or pipe data to stdin.")
# Validate required fields
if not isinstance(incident_data, dict):
parser.error("Input must be a JSON object")
if "description" not in incident_data:
parser.error("Input must contain 'description' field")
# Classify incident
result = classifier.classify_incident(incident_data)
# Format output
if args.format == "json":
output = format_json_output(result)
else:
output = format_text_output(result)
# Write output
if args.output:
with open(args.output, 'w') as f:
f.write(output)
f.write('\n')
else:
print(output)
except FileNotFoundError as e:
print(f"Error: File not found - {e}", file=sys.stderr)
sys.exit(1)
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON - {e}", file=sys.stderr)
sys.exit(1)
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
sys.exit(1)
if __name__ == "__main__":
main()
FILE:scripts/incident_timeline_builder.py
#!/usr/bin/env python3
"""
Incident Timeline Builder
Builds structured incident timelines with automatic phase detection, gap analysis,
communication template generation, and response metrics calculation. Produces
professional reports suitable for post-incident review and stakeholder briefing.
Usage:
python incident_timeline_builder.py incident_data.json
python incident_timeline_builder.py incident_data.json --format json
python incident_timeline_builder.py incident_data.json --format markdown
cat incident_data.json | python incident_timeline_builder.py --format text
"""
import argparse
import json
import sys
from datetime import datetime, timedelta
from typing import Any, Dict, List, Optional, Tuple
# ---------------------------------------------------------------------------
# Configuration Constants
# ---------------------------------------------------------------------------
ISO_FORMAT = "%Y-%m-%dT%H:%M:%SZ"
EVENT_TYPES = [
"detection", "declaration", "escalation", "investigation",
"mitigation", "communication", "resolution", "action_item",
]
SEVERITY_LEVELS = {
"SEV1": {"label": "Critical", "rank": 1},
"SEV2": {"label": "Major", "rank": 2},
"SEV3": {"label": "Minor", "rank": 3},
"SEV4": {"label": "Low", "rank": 4},
}
PHASE_DEFINITIONS = [
{"name": "Detection", "trigger_types": ["detection"],
"description": "Issue detected via monitoring, alerting, or user report."},
{"name": "Triage", "trigger_types": ["declaration", "escalation"],
"description": "Incident declared, severity assessed, commander assigned."},
{"name": "Investigation", "trigger_types": ["investigation"],
"description": "Root cause analysis and impact assessment underway."},
{"name": "Mitigation", "trigger_types": ["mitigation"],
"description": "Active work to reduce or eliminate customer impact."},
{"name": "Resolution", "trigger_types": ["resolution"],
"description": "Service restored to normal operating parameters."},
]
GAP_THRESHOLD_MINUTES = 15
DECISION_EVENT_TYPES = {"escalation", "mitigation", "declaration", "resolution"}
# ---------------------------------------------------------------------------
# Data Model Classes
# ---------------------------------------------------------------------------
class IncidentEvent:
"""Represents a single event in the incident timeline."""
def __init__(self, data: Dict[str, Any]):
self.timestamp_raw: str = data.get("timestamp", "")
self.timestamp: Optional[datetime] = _parse_timestamp(self.timestamp_raw)
self.type: str = data.get("type", "unknown").lower().strip()
self.actor: str = data.get("actor", "unknown")
self.description: str = data.get("description", "")
self.metadata: Dict[str, Any] = data.get("metadata", {})
def to_dict(self) -> Dict[str, Any]:
result: Dict[str, Any] = {
"timestamp": self.timestamp_raw, "type": self.type,
"actor": self.actor, "description": self.description,
}
if self.metadata:
result["metadata"] = self.metadata
return result
@property
def is_decision_point(self) -> bool:
return self.type in DECISION_EVENT_TYPES
class IncidentPhase:
"""Represents a detected phase of the incident lifecycle."""
def __init__(self, name: str, description: str):
self.name: str = name
self.description: str = description
self.start_time: Optional[datetime] = None
self.end_time: Optional[datetime] = None
self.events: List[IncidentEvent] = []
@property
def duration_minutes(self) -> Optional[float]:
if self.start_time and self.end_time:
return (self.end_time - self.start_time).total_seconds() / 60.0
return None
def to_dict(self) -> Dict[str, Any]:
dur = self.duration_minutes
return {
"name": self.name, "description": self.description,
"start_time": self.start_time.strftime(ISO_FORMAT) if self.start_time else None,
"end_time": self.end_time.strftime(ISO_FORMAT) if self.end_time else None,
"duration_minutes": round(dur, 1) if dur is not None else None,
"event_count": len(self.events),
}
class CommunicationTemplate:
"""A generated communication message for a specific audience."""
def __init__(self, template_type: str, audience: str, subject: str, body: str):
self.template_type = template_type
self.audience = audience
self.subject = subject
self.body = body
def to_dict(self) -> Dict[str, Any]:
return {"template_type": self.template_type, "audience": self.audience,
"subject": self.subject, "body": self.body}
class TimelineGap:
"""Represents a gap in the timeline where no events were logged."""
def __init__(self, start: datetime, end: datetime, duration_minutes: float):
self.start = start
self.end = end
self.duration_minutes = duration_minutes
def to_dict(self) -> Dict[str, Any]:
return {"start": self.start.strftime(ISO_FORMAT),
"end": self.end.strftime(ISO_FORMAT),
"duration_minutes": round(self.duration_minutes, 1)}
class TimelineAnalysis:
"""Holds the complete analysis result for an incident timeline."""
def __init__(self):
self.incident_id: str = ""
self.incident_title: str = ""
self.severity: str = ""
self.status: str = ""
self.commander: str = ""
self.service: str = ""
self.affected_services: List[str] = []
self.declared_at: Optional[datetime] = None
self.resolved_at: Optional[datetime] = None
self.events: List[IncidentEvent] = []
self.phases: List[IncidentPhase] = []
self.gaps: List[TimelineGap] = []
self.decision_points: List[IncidentEvent] = []
self.metrics: Dict[str, Any] = {}
self.communications: List[CommunicationTemplate] = []
self.errors: List[str] = []
# ---------------------------------------------------------------------------
# Timestamp Helpers
# ---------------------------------------------------------------------------
def _parse_timestamp(raw: str) -> Optional[datetime]:
"""Parse an ISO-8601 timestamp string into a datetime object."""
if not raw:
return None
cleaned = raw.replace("Z", "+00:00") if raw.endswith("Z") else raw
try:
return datetime.fromisoformat(cleaned).replace(tzinfo=None)
except (ValueError, AttributeError):
pass
try:
return datetime.strptime(raw, ISO_FORMAT)
except ValueError:
return None
def _fmt_duration(minutes: Optional[float]) -> str:
"""Format a duration in minutes as a human-readable string."""
if minutes is None:
return "N/A"
if minutes < 1:
return f"{minutes * 60:.0f}s"
if minutes < 60:
return f"{minutes:.0f}m"
hours, remaining = int(minutes // 60), int(minutes % 60)
return f"{hours}h" if remaining == 0 else f"{hours}h {remaining}m"
def _fmt_ts(dt: Optional[datetime]) -> str:
"""Format a datetime as HH:MM:SS for display."""
return dt.strftime("%H:%M:%S") if dt else "??:??:??"
def _sev_label(sev: str) -> str:
"""Return the human label for a severity code."""
return SEVERITY_LEVELS.get(sev, {}).get("label", sev)
# ---------------------------------------------------------------------------
# Core Analysis Functions
# ---------------------------------------------------------------------------
def parse_incident_data(data: Dict[str, Any]) -> TimelineAnalysis:
"""Parse raw incident JSON into a TimelineAnalysis with populated fields."""
a = TimelineAnalysis()
inc = data.get("incident", {})
a.incident_id = inc.get("id", "UNKNOWN")
a.incident_title = inc.get("title", "Untitled Incident")
a.severity = inc.get("severity", "UNKNOWN").upper()
a.status = inc.get("status", "unknown").lower()
a.commander = inc.get("commander", "Unassigned")
a.service = inc.get("service", "unknown")
a.affected_services = inc.get("affected_services", [])
a.declared_at = _parse_timestamp(inc.get("declared_at", ""))
a.resolved_at = _parse_timestamp(inc.get("resolved_at", ""))
raw_events = data.get("events", [])
if not raw_events:
a.errors.append("No events found in incident data.")
return a
for raw in raw_events:
event = IncidentEvent(raw)
if event.timestamp is None:
a.errors.append(f"Skipping event with unparseable timestamp: {raw.get('timestamp', '')}")
continue
a.events.append(event)
a.events.sort(key=lambda e: e.timestamp) # type: ignore[arg-type]
return a
def detect_phases(analysis: TimelineAnalysis) -> None:
"""Detect incident lifecycle phases from the ordered event stream."""
if not analysis.events:
return
trigger_map: Dict[str, Dict[str, str]] = {}
for pdef in PHASE_DEFINITIONS:
for ttype in pdef["trigger_types"]:
trigger_map[ttype] = {"name": pdef["name"], "description": pdef["description"]}
phase_by_name: Dict[str, IncidentPhase] = {}
phase_order: List[str] = []
current: Optional[IncidentPhase] = None
for event in analysis.events:
pinfo = trigger_map.get(event.type)
if pinfo and pinfo["name"] not in phase_by_name:
if current is not None:
current.end_time = event.timestamp
phase = IncidentPhase(pinfo["name"], pinfo["description"])
phase.start_time = event.timestamp
phase_by_name[pinfo["name"]] = phase
phase_order.append(pinfo["name"])
current = phase
if current is not None:
current.events.append(event)
if current is not None:
current.end_time = analysis.resolved_at or analysis.events[-1].timestamp
analysis.phases = [phase_by_name[n] for n in phase_order]
def detect_gaps(analysis: TimelineAnalysis) -> None:
"""Identify gaps longer than GAP_THRESHOLD_MINUTES between consecutive events."""
for i in range(len(analysis.events) - 1):
ts_a, ts_b = analysis.events[i].timestamp, analysis.events[i + 1].timestamp
if ts_a is None or ts_b is None:
continue
delta = (ts_b - ts_a).total_seconds() / 60.0
if delta >= GAP_THRESHOLD_MINUTES:
analysis.gaps.append(TimelineGap(start=ts_a, end=ts_b, duration_minutes=delta))
def identify_decision_points(analysis: TimelineAnalysis) -> None:
"""Extract key decision-point events from the timeline."""
analysis.decision_points = [e for e in analysis.events if e.is_decision_point]
def calculate_metrics(analysis: TimelineAnalysis) -> None:
"""Calculate incident response metrics: MTTD, MTTR, phase durations."""
m: Dict[str, Any] = {}
det = [e for e in analysis.events if e.type == "detection"]
first_det = det[0].timestamp if det else None
first_ts = analysis.events[0].timestamp if analysis.events else None
# MTTD: first event to first detection.
if first_ts and first_det:
m["mttd_minutes"] = round((first_det - first_ts).total_seconds() / 60.0, 1)
else:
m["mttd_minutes"] = None
# MTTR: detection to resolution.
if first_det and analysis.resolved_at:
m["mttr_minutes"] = round((analysis.resolved_at - first_det).total_seconds() / 60.0, 1)
else:
m["mttr_minutes"] = None
# Total duration.
if analysis.declared_at and analysis.resolved_at:
m["total_duration_minutes"] = round(
(analysis.resolved_at - analysis.declared_at).total_seconds() / 60.0, 1)
else:
m["total_duration_minutes"] = None
# Phase durations.
m["phase_durations"] = {
p.name: (round(p.duration_minutes, 1) if p.duration_minutes is not None else None)
for p in analysis.phases
}
# Event counts by type.
tc: Dict[str, int] = {}
for e in analysis.events:
tc[e.type] = tc.get(e.type, 0) + 1
m["event_counts_by_type"] = tc
# Gap statistics.
m["gap_count"] = len(analysis.gaps)
if analysis.gaps:
gm = [g.duration_minutes for g in analysis.gaps]
m["longest_gap_minutes"] = round(max(gm), 1)
m["total_gap_minutes"] = round(sum(gm), 1)
else:
m["longest_gap_minutes"] = 0
m["total_gap_minutes"] = 0
m["total_events"] = len(analysis.events)
m["decision_point_count"] = len(analysis.decision_points)
m["phase_count"] = len(analysis.phases)
analysis.metrics = m
# ---------------------------------------------------------------------------
# Communication Template Generation
# ---------------------------------------------------------------------------
def generate_communications(analysis: TimelineAnalysis) -> None:
"""Generate four communication templates based on incident data."""
sev, sl = analysis.severity, _sev_label(analysis.severity)
title, svc = analysis.incident_title, analysis.service
affected = ", ".join(analysis.affected_services) or "none identified"
cmd, iid = analysis.commander, analysis.incident_id
decl = analysis.declared_at.strftime("%Y-%m-%d %H:%M UTC") if analysis.declared_at else "TBD"
resv = analysis.resolved_at.strftime("%Y-%m-%d %H:%M UTC") if analysis.resolved_at else "TBD"
dur = _fmt_duration(analysis.metrics.get("total_duration_minutes"))
resolved = analysis.status == "resolved"
# 1 -- Initial stakeholder notification
analysis.communications.append(CommunicationTemplate(
"initial_notification", "internal", f"[{sev}] Incident Declared: {title}",
f"An incident has been declared for {svc}.\n\n"
f"Incident ID: {iid}\nSeverity: {sev} ({sl})\nCommander: {cmd}\n"
f"Declared at: {decl}\nAffected services: {affected}\n\n"
f"The incident team is actively investigating. Updates will follow.",
))
# 2 -- Status page update
if resolved:
sp_subj = f"[Resolved] {title}"
sp_body = (f"The incident affecting {svc} has been resolved.\n\n"
f"Duration: {dur}\nAll affected services ({affected}) are restored. "
f"A post-incident review will be published within 48 hours.")
else:
sp_subj = f"[Investigating] {title}"
sp_body = (f"We are investigating degraded performance in {svc}. "
f"Affected services: {affected}.\n\n"
f"Our team is working to identify the root cause. Updates every 30 minutes.")
analysis.communications.append(CommunicationTemplate(
"status_page", "external", sp_subj, sp_body))
# 3 -- Executive summary
phase_lines = "\n".join(
f" - {p.name}: {_fmt_duration(p.duration_minutes)}" for p in analysis.phases
) or " No phase data available."
mttd = _fmt_duration(analysis.metrics.get("mttd_minutes"))
mttr = _fmt_duration(analysis.metrics.get("mttr_minutes"))
analysis.communications.append(CommunicationTemplate(
"executive_summary", "executive", f"Executive Summary: {iid} - {title}",
f"Incident: {iid} - {title}\nSeverity: {sev} ({sl})\n"
f"Service: {svc}\nCommander: {cmd}\nStatus: {analysis.status.capitalize()}\n"
f"Declared: {decl}\nResolved: {resv}\nDuration: {dur}\n\n"
f"Key Metrics:\n - MTTD: {mttd}\n - MTTR: {mttr}\n"
f" - Timeline Gaps: {analysis.metrics.get('gap_count', 0)}\n\n"
f"Phase Breakdown:\n{phase_lines}\n\nAffected Services: {affected}",
))
# 4 -- Customer notification
if resolved:
cust_body = (f"We experienced an issue affecting {svc} starting at {decl}.\n\n"
f"The issue was resolved at {resv} (duration: {dur}). "
f"We apologize for any inconvenience and are reviewing to prevent recurrence.")
else:
cust_body = (f"We are experiencing an issue affecting {svc} starting at {decl}.\n\n"
f"Our engineering team is actively working to resolve this. "
f"We will provide updates as the situation develops. We apologize for the inconvenience.")
analysis.communications.append(CommunicationTemplate(
"customer_notification", "external", f"Service Update: {title}", cust_body))
# ---------------------------------------------------------------------------
# Main Analysis Orchestrator
# ---------------------------------------------------------------------------
def build_timeline(data: Dict[str, Any]) -> TimelineAnalysis:
"""Run the full timeline analysis pipeline on raw incident data."""
analysis = parse_incident_data(data)
if analysis.errors and not analysis.events:
return analysis
detect_phases(analysis)
detect_gaps(analysis)
identify_decision_points(analysis)
calculate_metrics(analysis)
generate_communications(analysis)
return analysis
# ---------------------------------------------------------------------------
# Output Formatters
# ---------------------------------------------------------------------------
def format_text_output(analysis: TimelineAnalysis) -> str:
"""Format the analysis as a human-readable text report."""
L: List[str] = []
w = 64
L.append("=" * w)
L.append("INCIDENT TIMELINE REPORT")
L.append("=" * w)
L.append("")
if analysis.errors:
for err in analysis.errors:
L.append(f" WARNING: {err}")
L.append("")
if not analysis.events:
return "\n".join(L)
# Summary
L.append("INCIDENT SUMMARY")
L.append("-" * 32)
L.append(f" ID: {analysis.incident_id}")
L.append(f" Title: {analysis.incident_title}")
L.append(f" Severity: {analysis.severity}")
L.append(f" Status: {analysis.status.capitalize()}")
L.append(f" Commander: {analysis.commander}")
L.append(f" Service: {analysis.service}")
if analysis.affected_services:
L.append(f" Affected: {', '.join(analysis.affected_services)}")
L.append(f" Duration: {_fmt_duration(analysis.metrics.get('total_duration_minutes'))}")
L.append("")
# Key metrics
L.append("KEY METRICS")
L.append("-" * 32)
L.append(f" MTTD (Mean Time to Detect): {_fmt_duration(analysis.metrics.get('mttd_minutes'))}")
L.append(f" MTTR (Mean Time to Resolve): {_fmt_duration(analysis.metrics.get('mttr_minutes'))}")
L.append(f" Total Events: {analysis.metrics.get('total_events', 0)}")
L.append(f" Decision Points: {analysis.metrics.get('decision_point_count', 0)}")
L.append(f" Timeline Gaps (>{GAP_THRESHOLD_MINUTES}m): {analysis.metrics.get('gap_count', 0)}")
L.append("")
# Phases
L.append("INCIDENT PHASES")
L.append("-" * 32)
if analysis.phases:
for p in analysis.phases:
L.append(f" [{_fmt_ts(p.start_time)} - {_fmt_ts(p.end_time)}] {p.name} ({_fmt_duration(p.duration_minutes)})")
L.append(f" {p.description}")
L.append(f" Events: {len(p.events)}")
else:
L.append(" No phases detected.")
L.append("")
# Chronological timeline
L.append("CHRONOLOGICAL TIMELINE")
L.append("-" * 32)
for e in analysis.events:
marker = "*" if e.is_decision_point else " "
L.append(f" {_fmt_ts(e.timestamp)} {marker} [{e.type.upper():13s}] {e.actor}")
L.append(f" {e.description}")
L.append("")
L.append(" (* = key decision point)")
L.append("")
# Gap warnings
if analysis.gaps:
L.append("GAP ANALYSIS")
L.append("-" * 32)
for g in analysis.gaps:
L.append(f" WARNING: {_fmt_duration(g.duration_minutes)} gap between {_fmt_ts(g.start)} and {_fmt_ts(g.end)}")
L.append("")
# Decision points
if analysis.decision_points:
L.append("KEY DECISION POINTS")
L.append("-" * 32)
for dp in analysis.decision_points:
L.append(f" {_fmt_ts(dp.timestamp)} [{dp.type.upper()}] {dp.description}")
L.append("")
# Communications
if analysis.communications:
L.append("GENERATED COMMUNICATIONS")
L.append("-" * 32)
for c in analysis.communications:
L.append(f" Type: {c.template_type}")
L.append(f" Audience: {c.audience}")
L.append(f" Subject: {c.subject}")
L.append(" ---")
for bl in c.body.split("\n"):
L.append(f" {bl}")
L.append("")
L.append("=" * w)
L.append("END OF REPORT")
L.append("=" * w)
return "\n".join(L)
def format_json_output(analysis: TimelineAnalysis) -> Dict[str, Any]:
"""Format the analysis as a structured JSON-serializable dictionary."""
return {
"incident": {
"id": analysis.incident_id, "title": analysis.incident_title,
"severity": analysis.severity, "status": analysis.status,
"commander": analysis.commander, "service": analysis.service,
"affected_services": analysis.affected_services,
"declared_at": analysis.declared_at.strftime(ISO_FORMAT) if analysis.declared_at else None,
"resolved_at": analysis.resolved_at.strftime(ISO_FORMAT) if analysis.resolved_at else None,
},
"timeline": [e.to_dict() for e in analysis.events],
"phases": [p.to_dict() for p in analysis.phases],
"gaps": [g.to_dict() for g in analysis.gaps],
"decision_points": [e.to_dict() for e in analysis.decision_points],
"metrics": analysis.metrics,
"communications": [c.to_dict() for c in analysis.communications],
"errors": analysis.errors if analysis.errors else [],
}
def format_markdown_output(analysis: TimelineAnalysis) -> str:
"""Format the analysis as a professional Markdown report."""
L: List[str] = []
L.append(f"# Incident Timeline Report: {analysis.incident_id}")
L.append("")
if analysis.errors:
L.append("> **Warnings:**")
for err in analysis.errors:
L.append(f"> - {err}")
L.append("")
if not analysis.events:
return "\n".join(L)
# Summary table
L.append("## Incident Summary")
L.append("")
L.append("| Field | Value |")
L.append("|-------|-------|")
L.append(f"| **ID** | {analysis.incident_id} |")
L.append(f"| **Title** | {analysis.incident_title} |")
L.append(f"| **Severity** | {analysis.severity} ({_sev_label(analysis.severity)}) |")
L.append(f"| **Status** | {analysis.status.capitalize()} |")
L.append(f"| **Commander** | {analysis.commander} |")
L.append(f"| **Service** | {analysis.service} |")
if analysis.affected_services:
L.append(f"| **Affected Services** | {', '.join(analysis.affected_services)} |")
L.append(f"| **Duration** | {_fmt_duration(analysis.metrics.get('total_duration_minutes'))} |")
L.append("")
# Key metrics
L.append("## Key Metrics")
L.append("")
L.append(f"- **MTTD (Mean Time to Detect):** {_fmt_duration(analysis.metrics.get('mttd_minutes'))}")
L.append(f"- **MTTR (Mean Time to Resolve):** {_fmt_duration(analysis.metrics.get('mttr_minutes'))}")
L.append(f"- **Total Events:** {analysis.metrics.get('total_events', 0)}")
L.append(f"- **Decision Points:** {analysis.metrics.get('decision_point_count', 0)}")
L.append(f"- **Timeline Gaps (>{GAP_THRESHOLD_MINUTES}m):** {analysis.metrics.get('gap_count', 0)}")
if analysis.metrics.get("longest_gap_minutes", 0) > 0:
L.append(f"- **Longest Gap:** {_fmt_duration(analysis.metrics.get('longest_gap_minutes'))}")
L.append("")
# Phases table
L.append("## Incident Phases")
L.append("")
if analysis.phases:
L.append("| Phase | Start | End | Duration | Events |")
L.append("|-------|-------|-----|----------|--------|")
for p in analysis.phases:
L.append(f"| {p.name} | {_fmt_ts(p.start_time)} | {_fmt_ts(p.end_time)} | {_fmt_duration(p.duration_minutes)} | {len(p.events)} |")
L.append("")
# ASCII bar chart
max_dur = max((p.duration_minutes for p in analysis.phases if p.duration_minutes), default=0)
if max_dur and max_dur > 0:
L.append("### Phase Duration Distribution")
L.append("")
L.append("```")
for p in analysis.phases:
d = p.duration_minutes or 0
bar = "#" * int((d / max_dur) * 40)
L.append(f" {p.name:15s} |{bar} {_fmt_duration(d)}")
L.append("```")
L.append("")
else:
L.append("No phases detected.")
L.append("")
# Chronological timeline
L.append("## Chronological Timeline")
L.append("")
for e in analysis.events:
dm = " **[KEY DECISION]**" if e.is_decision_point else ""
L.append(f"- `{_fmt_ts(e.timestamp)}` **{e.type.upper()}** ({e.actor}){dm}")
L.append(f" - {e.description}")
L.append("")
# Gap analysis
if analysis.gaps:
L.append("## Gap Analysis")
L.append("")
L.append(f"> {len(analysis.gaps)} gap(s) of >{GAP_THRESHOLD_MINUTES} minutes detected. "
f"These may represent blind spots where important activity was not recorded.")
L.append("")
for g in analysis.gaps:
L.append(f"- **{_fmt_duration(g.duration_minutes)}** gap from `{_fmt_ts(g.start)}` to `{_fmt_ts(g.end)}`")
L.append("")
# Decision points
if analysis.decision_points:
L.append("## Key Decision Points")
L.append("")
for dp in analysis.decision_points:
L.append(f"1. `{_fmt_ts(dp.timestamp)}` **{dp.type.upper()}** - {dp.description}")
L.append("")
# Communications
if analysis.communications:
L.append("## Generated Communications")
L.append("")
for c in analysis.communications:
L.append(f"### {c.template_type.replace('_', ' ').title()} ({c.audience})")
L.append("")
L.append(f"**Subject:** {c.subject}")
L.append("")
for bl in c.body.split("\n"):
L.append(bl)
L.append("")
L.append("---")
L.append("")
# Event type breakdown
tc = analysis.metrics.get("event_counts_by_type", {})
if tc:
L.append("## Event Type Breakdown")
L.append("")
L.append("| Type | Count |")
L.append("|------|-------|")
for etype, count in sorted(tc.items(), key=lambda x: -x[1]):
L.append(f"| {etype} | {count} |")
L.append("")
L.append("---")
L.append(f"*Report generated for incident {analysis.incident_id}. All timestamps in UTC.*")
return "\n".join(L)
# ---------------------------------------------------------------------------
# CLI Interface
# ---------------------------------------------------------------------------
def main() -> int:
"""Main CLI entry point."""
parser = argparse.ArgumentParser(
description="Build structured incident timelines with phase detection and communication templates."
)
parser.add_argument(
"data_file", nargs="?", default=None,
help="JSON file with incident data (reads stdin if omitted)",
)
parser.add_argument(
"--format", choices=["text", "json", "markdown"], default="text",
help="Output format (default: text)",
)
args = parser.parse_args()
try:
if args.data_file:
try:
with open(args.data_file, "r") as f:
raw_data = json.load(f)
except FileNotFoundError:
print(f"Error: File '{args.data_file}' not found.", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON in '{args.data_file}': {e}", file=sys.stderr)
return 1
else:
if sys.stdin.isatty():
print("Error: No input file specified and stdin is a terminal. "
"Provide a file argument or pipe JSON to stdin.", file=sys.stderr)
return 1
try:
raw_data = json.load(sys.stdin)
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON on stdin: {e}", file=sys.stderr)
return 1
if not isinstance(raw_data, dict):
print("Error: Input must be a JSON object.", file=sys.stderr)
return 1
if "incident" not in raw_data and "events" not in raw_data:
print("Error: Input must contain at least 'incident' or 'events' keys.", file=sys.stderr)
return 1
analysis = build_timeline(raw_data)
if args.format == "json":
print(json.dumps(format_json_output(analysis), indent=2))
elif args.format == "markdown":
print(format_markdown_output(analysis))
else:
print(format_text_output(analysis))
return 0
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
return 1
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/pir_generator.py
#!/usr/bin/env python3
"""
PIR (Post-Incident Review) Generator
Generates comprehensive Post-Incident Review documents from incident data, timelines,
and actions taken. Applies multiple RCA frameworks including 5 Whys, Fishbone diagram,
and Timeline analysis.
This tool creates structured PIR documents with root cause analysis, lessons learned,
action items, and follow-up recommendations.
Usage:
python pir_generator.py --incident incident.json --timeline timeline.json --output pir.md
python pir_generator.py --incident incident.json --rca-method fishbone --action-items
cat incident.json | python pir_generator.py --format markdown
"""
import argparse
import json
import sys
import re
from datetime import datetime, timezone, timedelta
from typing import Dict, List, Optional, Any, Tuple
from collections import defaultdict, Counter
class PIRGenerator:
"""
Generates comprehensive Post-Incident Review documents with multiple
RCA frameworks, lessons learned, and actionable follow-up items.
"""
def __init__(self):
"""Initialize the PIR generator with templates and frameworks."""
self.rca_frameworks = self._load_rca_frameworks()
self.pir_templates = self._load_pir_templates()
self.severity_guidelines = self._load_severity_guidelines()
self.action_item_types = self._load_action_item_types()
self.lessons_learned_categories = self._load_lessons_learned_categories()
def _load_rca_frameworks(self) -> Dict[str, Dict]:
"""Load root cause analysis framework definitions."""
return {
"five_whys": {
"name": "5 Whys Analysis",
"description": "Iterative questioning technique to explore cause-and-effect relationships",
"steps": [
"State the problem clearly",
"Ask why the problem occurred",
"For each answer, ask why again",
"Continue until root cause is identified",
"Verify the root cause addresses the original problem"
],
"min_iterations": 3,
"max_iterations": 7
},
"fishbone": {
"name": "Fishbone (Ishikawa) Diagram",
"description": "Systematic analysis across multiple categories of potential causes",
"categories": [
{
"name": "People",
"description": "Human factors, training, communication, experience",
"examples": ["Training gaps", "Communication failures", "Skill deficits", "Staffing issues"]
},
{
"name": "Process",
"description": "Procedures, workflows, change management, review processes",
"examples": ["Missing procedures", "Inadequate reviews", "Change management gaps", "Documentation issues"]
},
{
"name": "Technology",
"description": "Systems, tools, architecture, automation",
"examples": ["Architecture limitations", "Tool deficiencies", "Automation gaps", "Infrastructure issues"]
},
{
"name": "Environment",
"description": "External factors, dependencies, infrastructure",
"examples": ["Third-party dependencies", "Network issues", "Hardware failures", "External service outages"]
}
]
},
"timeline": {
"name": "Timeline Analysis",
"description": "Chronological analysis of events to identify decision points and missed opportunities",
"focus_areas": [
"Detection timing and effectiveness",
"Response time and escalation paths",
"Decision points and alternative paths",
"Communication effectiveness",
"Mitigation strategy effectiveness"
]
},
"bow_tie": {
"name": "Bow Tie Analysis",
"description": "Analysis of both preventive and protective measures around an incident",
"components": [
"Hazards (what could go wrong)",
"Top events (what actually went wrong)",
"Threats (what caused it)",
"Consequences (what was the impact)",
"Barriers (what preventive/protective measures exist or could exist)"
]
}
}
def _load_pir_templates(self) -> Dict[str, str]:
"""Load PIR document templates for different severity levels."""
return {
"comprehensive": """# Post-Incident Review: {incident_title}
## Executive Summary
{executive_summary}
## Incident Overview
- **Incident ID:** {incident_id}
- **Date & Time:** {incident_date}
- **Duration:** {duration}
- **Severity:** {severity}
- **Status:** {status}
- **Incident Commander:** {incident_commander}
- **Responders:** {responders}
### Customer Impact
{customer_impact}
### Business Impact
{business_impact}
## Timeline
{timeline_section}
## Root Cause Analysis
{rca_section}
## What Went Well
{what_went_well}
## What Didn't Go Well
{what_went_wrong}
## Lessons Learned
{lessons_learned}
## Action Items
{action_items}
## Follow-up and Prevention
{prevention_measures}
## Appendix
{appendix_section}
---
*Generated on {generation_date} by PIR Generator*
""",
"standard": """# Post-Incident Review: {incident_title}
## Summary
{executive_summary}
## Incident Details
- **Date:** {incident_date}
- **Duration:** {duration}
- **Severity:** {severity}
- **Impact:** {customer_impact}
## Timeline
{timeline_section}
## Root Cause
{rca_section}
## Action Items
{action_items}
## Lessons Learned
{lessons_learned}
---
*Generated on {generation_date}*
""",
"brief": """# Incident Review: {incident_title}
**Date:** {incident_date} | **Duration:** {duration} | **Severity:** {severity}
## What Happened
{executive_summary}
## Root Cause
{rca_section}
## Actions
{action_items}
---
*{generation_date}*
"""
}
def _load_severity_guidelines(self) -> Dict[str, Dict]:
"""Load severity-specific PIR guidelines."""
return {
"sev1": {
"required_sections": ["executive_summary", "timeline", "rca", "action_items", "lessons_learned"],
"required_attendees": ["incident_commander", "technical_leads", "engineering_manager", "product_manager"],
"timeline_requirement": "Complete timeline with 15-minute intervals",
"rca_methods": ["five_whys", "fishbone", "timeline"],
"review_deadline_hours": 24,
"follow_up_weeks": 4
},
"sev2": {
"required_sections": ["summary", "timeline", "rca", "action_items"],
"required_attendees": ["incident_commander", "technical_leads", "team_lead"],
"timeline_requirement": "Key milestone timeline",
"rca_methods": ["five_whys", "timeline"],
"review_deadline_hours": 72,
"follow_up_weeks": 2
},
"sev3": {
"required_sections": ["summary", "rca", "action_items"],
"required_attendees": ["technical_lead", "team_member"],
"timeline_requirement": "Basic timeline",
"rca_methods": ["five_whys"],
"review_deadline_hours": 168, # 1 week
"follow_up_weeks": 1
},
"sev4": {
"required_sections": ["summary", "action_items"],
"required_attendees": ["assigned_engineer"],
"timeline_requirement": "Optional",
"rca_methods": ["brief_analysis"],
"review_deadline_hours": 336, # 2 weeks
"follow_up_weeks": 0
}
}
def _load_action_item_types(self) -> Dict[str, Dict]:
"""Load action item categorization and templates."""
return {
"immediate_fix": {
"priority": "P0",
"timeline": "24-48 hours",
"description": "Critical bugs or security issues that need immediate attention",
"template": "Fix {issue_description} to prevent recurrence of {incident_type}",
"owners": ["engineer", "team_lead"]
},
"process_improvement": {
"priority": "P1",
"timeline": "1-2 weeks",
"description": "Process gaps or communication issues identified",
"template": "Improve {process_area} to address {gap_description}",
"owners": ["team_lead", "process_owner"]
},
"monitoring_alerting": {
"priority": "P1",
"timeline": "1 week",
"description": "Missing monitoring or alerting capabilities",
"template": "Implement {monitoring_type} for {system_component}",
"owners": ["sre", "engineer"]
},
"documentation": {
"priority": "P2",
"timeline": "2-3 weeks",
"description": "Documentation gaps or runbook updates",
"template": "Update {documentation_type} to include {missing_information}",
"owners": ["technical_writer", "engineer"]
},
"training": {
"priority": "P2",
"timeline": "1 month",
"description": "Training needs or knowledge gaps",
"template": "Provide {training_type} training on {topic}",
"owners": ["training_coordinator", "subject_matter_expert"]
},
"architectural": {
"priority": "P1-P3",
"timeline": "1-3 months",
"description": "System design or architecture improvements",
"template": "Redesign {system_component} to improve {quality_attribute}",
"owners": ["architect", "engineering_manager"]
},
"tooling": {
"priority": "P2",
"timeline": "2-4 weeks",
"description": "Tool improvements or new tool requirements",
"template": "Implement {tool_type} to support {use_case}",
"owners": ["devops", "engineer"]
}
}
def _load_lessons_learned_categories(self) -> Dict[str, List[str]]:
"""Load categories for organizing lessons learned."""
return {
"detection_and_monitoring": [
"Monitoring gaps identified",
"Alert fatigue issues",
"Detection timing improvements",
"Observability enhancements"
],
"response_and_escalation": [
"Response time improvements",
"Escalation path optimization",
"Communication effectiveness",
"Resource allocation lessons"
],
"technical_systems": [
"Architecture resilience",
"Failure mode analysis",
"Performance bottlenecks",
"Dependency management"
],
"process_and_procedures": [
"Runbook effectiveness",
"Change management gaps",
"Review process improvements",
"Documentation quality"
],
"team_and_culture": [
"Training needs identified",
"Cross-team collaboration",
"Knowledge sharing gaps",
"Decision-making processes"
]
}
def generate_pir(self, incident_data: Dict[str, Any], timeline_data: Optional[Dict] = None,
rca_method: str = "five_whys", template_type: str = "comprehensive") -> Dict[str, Any]:
"""
Generate a comprehensive PIR document from incident data.
Args:
incident_data: Core incident information
timeline_data: Optional timeline reconstruction data
rca_method: RCA framework to use
template_type: PIR template type (comprehensive, standard, brief)
Returns:
Dictionary containing PIR document and metadata
"""
# Extract incident information
incident_info = self._extract_incident_info(incident_data)
# Generate root cause analysis
rca_results = self._perform_rca(incident_data, timeline_data, rca_method)
# Generate lessons learned
lessons_learned = self._generate_lessons_learned(incident_data, timeline_data, rca_results)
# Generate action items
action_items = self._generate_action_items(incident_data, rca_results, lessons_learned)
# Create timeline section
timeline_section = self._create_timeline_section(timeline_data, incident_info["severity"])
# Generate document sections
sections = self._generate_document_sections(
incident_info, rca_results, lessons_learned, action_items, timeline_section
)
# Build final document
template = self.pir_templates[template_type]
pir_document = template.format(**sections)
# Generate metadata
metadata = self._generate_metadata(incident_info, rca_results, action_items)
return {
"pir_document": pir_document,
"metadata": metadata,
"incident_info": incident_info,
"rca_results": rca_results,
"lessons_learned": lessons_learned,
"action_items": action_items,
"generation_timestamp": datetime.now(timezone.utc).isoformat()
}
def _extract_incident_info(self, incident_data: Dict) -> Dict[str, Any]:
"""Extract and normalize incident information."""
return {
"incident_id": incident_data.get("incident_id", "INC-" + datetime.now().strftime("%Y%m%d-%H%M")),
"title": incident_data.get("title", incident_data.get("description", "Incident")[:50]),
"description": incident_data.get("description", "No description provided"),
"severity": incident_data.get("severity", "unknown").lower(),
"start_time": self._parse_timestamp(incident_data.get("start_time", incident_data.get("timestamp", ""))),
"end_time": self._parse_timestamp(incident_data.get("end_time", "")),
"duration": self._calculate_duration(incident_data),
"affected_services": incident_data.get("affected_services", []),
"customer_impact": incident_data.get("customer_impact", "Unknown impact"),
"business_impact": incident_data.get("business_impact", "Unknown business impact"),
"incident_commander": incident_data.get("incident_commander", "TBD"),
"responders": incident_data.get("responders", []),
"status": incident_data.get("status", "resolved")
}
def _parse_timestamp(self, timestamp_str: str) -> Optional[datetime]:
"""Parse timestamp string to datetime object."""
if not timestamp_str:
return None
formats = [
"%Y-%m-%dT%H:%M:%S.%fZ",
"%Y-%m-%dT%H:%M:%SZ",
"%Y-%m-%d %H:%M:%S",
"%m/%d/%Y %H:%M:%S"
]
for fmt in formats:
try:
dt = datetime.strptime(timestamp_str, fmt)
if dt.tzinfo is None:
dt = dt.replace(tzinfo=timezone.utc)
return dt
except ValueError:
continue
return None
def _calculate_duration(self, incident_data: Dict) -> str:
"""Calculate incident duration in human-readable format."""
start_time = self._parse_timestamp(incident_data.get("start_time", ""))
end_time = self._parse_timestamp(incident_data.get("end_time", ""))
if start_time and end_time:
duration = end_time - start_time
total_minutes = int(duration.total_seconds() / 60)
if total_minutes < 60:
return f"{total_minutes} minutes"
elif total_minutes < 1440: # Less than 24 hours
hours = total_minutes // 60
minutes = total_minutes % 60
return f"{hours}h {minutes}m"
else:
days = total_minutes // 1440
hours = (total_minutes % 1440) // 60
return f"{days}d {hours}h"
return incident_data.get("duration", "Unknown duration")
def _perform_rca(self, incident_data: Dict, timeline_data: Optional[Dict], method: str) -> Dict[str, Any]:
"""Perform root cause analysis using specified method."""
if method == "five_whys":
return self._five_whys_analysis(incident_data, timeline_data)
elif method == "fishbone":
return self._fishbone_analysis(incident_data, timeline_data)
elif method == "timeline":
return self._timeline_analysis(incident_data, timeline_data)
elif method == "bow_tie":
return self._bow_tie_analysis(incident_data, timeline_data)
else:
return self._five_whys_analysis(incident_data, timeline_data) # Default
def _five_whys_analysis(self, incident_data: Dict, timeline_data: Optional[Dict]) -> Dict[str, Any]:
"""Perform 5 Whys root cause analysis."""
problem_statement = incident_data.get("description", "Incident occurred")
# Generate why questions based on incident data
whys = []
current_issue = problem_statement
# Generate systematic why questions
why_patterns = [
f"Why did {current_issue}?",
"Why wasn't this detected earlier?",
"Why didn't existing safeguards prevent this?",
"Why wasn't there a backup mechanism?",
"Why wasn't this scenario anticipated?"
]
# Try to infer answers from incident data
potential_answers = self._infer_why_answers(incident_data, timeline_data)
for i, why_question in enumerate(why_patterns):
answer = potential_answers[i] if i < len(potential_answers) else "Further investigation needed"
whys.append({
"question": why_question,
"answer": answer,
"evidence": self._find_supporting_evidence(answer, incident_data, timeline_data)
})
# Identify root causes from the analysis
root_causes = self._extract_root_causes(whys)
return {
"method": "five_whys",
"problem_statement": problem_statement,
"why_analysis": whys,
"root_causes": root_causes,
"confidence": self._calculate_rca_confidence(whys, incident_data)
}
def _fishbone_analysis(self, incident_data: Dict, timeline_data: Optional[Dict]) -> Dict[str, Any]:
"""Perform Fishbone (Ishikawa) diagram analysis."""
problem_statement = incident_data.get("description", "Incident occurred")
# Analyze each category
categories = {}
for category_info in self.rca_frameworks["fishbone"]["categories"]:
category_name = category_info["name"]
contributing_factors = self._identify_category_factors(
category_name, incident_data, timeline_data
)
categories[category_name] = {
"description": category_info["description"],
"factors": contributing_factors,
"examples": category_info["examples"]
}
# Identify primary contributing factors
primary_factors = self._identify_primary_factors(categories)
# Generate root cause hypothesis
root_causes = self._synthesize_fishbone_root_causes(categories, primary_factors)
return {
"method": "fishbone",
"problem_statement": problem_statement,
"categories": categories,
"primary_factors": primary_factors,
"root_causes": root_causes,
"confidence": self._calculate_rca_confidence(categories, incident_data)
}
def _timeline_analysis(self, incident_data: Dict, timeline_data: Optional[Dict]) -> Dict[str, Any]:
"""Perform timeline-based root cause analysis."""
if not timeline_data:
return {"method": "timeline", "error": "No timeline data provided"}
# Extract key decision points
decision_points = self._extract_decision_points(timeline_data)
# Identify missed opportunities
missed_opportunities = self._identify_missed_opportunities(timeline_data)
# Analyze response effectiveness
response_analysis = self._analyze_response_effectiveness(timeline_data)
# Generate timeline-based root causes
root_causes = self._extract_timeline_root_causes(
decision_points, missed_opportunities, response_analysis
)
return {
"method": "timeline",
"decision_points": decision_points,
"missed_opportunities": missed_opportunities,
"response_analysis": response_analysis,
"root_causes": root_causes,
"confidence": self._calculate_rca_confidence(timeline_data, incident_data)
}
def _bow_tie_analysis(self, incident_data: Dict, timeline_data: Optional[Dict]) -> Dict[str, Any]:
"""Perform Bow Tie analysis."""
# Identify the top event (what went wrong)
top_event = incident_data.get("description", "Service failure")
# Identify threats (what caused it)
threats = self._identify_threats(incident_data, timeline_data)
# Identify consequences (impact)
consequences = self._identify_consequences(incident_data)
# Identify existing barriers
existing_barriers = self._identify_existing_barriers(incident_data, timeline_data)
# Recommend additional barriers
recommended_barriers = self._recommend_additional_barriers(threats, consequences)
return {
"method": "bow_tie",
"top_event": top_event,
"threats": threats,
"consequences": consequences,
"existing_barriers": existing_barriers,
"recommended_barriers": recommended_barriers,
"confidence": self._calculate_rca_confidence(threats, incident_data)
}
def _infer_why_answers(self, incident_data: Dict, timeline_data: Optional[Dict]) -> List[str]:
"""Infer potential answers to why questions from available data."""
answers = []
# Look for clues in incident description
description = incident_data.get("description", "").lower()
# Common patterns and their inferred answers
if "database" in description and ("timeout" in description or "slow" in description):
answers.append("Database connection pool was exhausted")
answers.append("Connection pool configuration was insufficient for peak load")
answers.append("Load testing didn't include realistic database scenarios")
elif "deployment" in description or "release" in description:
answers.append("New deployment introduced a regression")
answers.append("Code review process missed the issue")
answers.append("Testing environment didn't match production")
elif "network" in description or "connectivity" in description:
answers.append("Network infrastructure had unexpected load")
answers.append("Network monitoring wasn't comprehensive enough")
answers.append("Redundancy mechanisms failed simultaneously")
else:
# Generic answers based on common root causes
answers.extend([
"System couldn't handle the load/request volume",
"Monitoring didn't detect the issue early enough",
"Error handling mechanisms were insufficient",
"Dependencies failed without proper circuit breakers",
"System lacked sufficient redundancy/resilience"
])
return answers[:5] # Return up to 5 answers
def _find_supporting_evidence(self, answer: str, incident_data: Dict, timeline_data: Optional[Dict]) -> List[str]:
"""Find supporting evidence for RCA answers."""
evidence = []
# Look for supporting information in incident data
if timeline_data and "timeline" in timeline_data:
events = timeline_data["timeline"].get("events", [])
for event in events:
event_message = event.get("message", "").lower()
if any(keyword in event_message for keyword in answer.lower().split()):
evidence.append(f"Timeline event: {event['message']}")
# Check incident metadata for supporting info
metadata = incident_data.get("metadata", {})
for key, value in metadata.items():
if isinstance(value, str) and any(keyword in value.lower() for keyword in answer.lower().split()):
evidence.append(f"Incident metadata: {key} = {value}")
return evidence[:3] # Return top 3 pieces of evidence
def _extract_root_causes(self, whys: List[Dict]) -> List[Dict]:
"""Extract root causes from 5 Whys analysis."""
root_causes = []
# The deepest "why" answers are typically closest to root causes
if len(whys) >= 3:
for i, why in enumerate(whys[-2:]): # Look at last 2 whys
if "further investigation needed" not in why["answer"].lower():
root_causes.append({
"cause": why["answer"],
"category": self._categorize_root_cause(why["answer"]),
"evidence": why["evidence"],
"confidence": "high" if len(why["evidence"]) > 1 else "medium"
})
return root_causes
def _categorize_root_cause(self, cause: str) -> str:
"""Categorize a root cause into standard categories."""
cause_lower = cause.lower()
if any(keyword in cause_lower for keyword in ["process", "procedure", "review", "change management"]):
return "Process"
elif any(keyword in cause_lower for keyword in ["training", "knowledge", "skill", "experience"]):
return "People"
elif any(keyword in cause_lower for keyword in ["system", "architecture", "code", "configuration"]):
return "Technology"
elif any(keyword in cause_lower for keyword in ["network", "infrastructure", "dependency", "third-party"]):
return "Environment"
else:
return "Unknown"
def _identify_category_factors(self, category: str, incident_data: Dict, timeline_data: Optional[Dict]) -> List[Dict]:
"""Identify contributing factors for a Fishbone category."""
factors = []
description = incident_data.get("description", "").lower()
if category == "People":
if "misconfigured" in description or "human error" in description:
factors.append({"factor": "Configuration error", "likelihood": "high"})
if timeline_data and self._has_delayed_response(timeline_data):
factors.append({"factor": "Delayed incident response", "likelihood": "medium"})
elif category == "Process":
if "deployment" in description:
factors.append({"factor": "Insufficient deployment validation", "likelihood": "high"})
if "code review" in incident_data.get("context", "").lower():
factors.append({"factor": "Code review process gaps", "likelihood": "medium"})
elif category == "Technology":
if "database" in description:
factors.append({"factor": "Database performance limitations", "likelihood": "high"})
if "timeout" in description or "latency" in description:
factors.append({"factor": "System performance bottlenecks", "likelihood": "high"})
elif category == "Environment":
if "network" in description:
factors.append({"factor": "Network infrastructure issues", "likelihood": "medium"})
if "third-party" in description or "external" in description:
factors.append({"factor": "External service dependencies", "likelihood": "medium"})
return factors
def _identify_primary_factors(self, categories: Dict) -> List[Dict]:
"""Identify primary contributing factors across all categories."""
primary_factors = []
for category_name, category_data in categories.items():
high_likelihood_factors = [
f for f in category_data["factors"]
if f.get("likelihood") == "high"
]
primary_factors.extend([
{**factor, "category": category_name}
for factor in high_likelihood_factors
])
return primary_factors
def _synthesize_fishbone_root_causes(self, categories: Dict, primary_factors: List[Dict]) -> List[Dict]:
"""Synthesize root causes from Fishbone analysis."""
root_causes = []
# Group primary factors by category
category_factors = defaultdict(list)
for factor in primary_factors:
category_factors[factor["category"]].append(factor)
# Create root causes from categories with multiple factors
for category, factors in category_factors.items():
if len(factors) > 1:
root_causes.append({
"cause": f"Multiple {category.lower()} issues contributed to the incident",
"category": category,
"contributing_factors": [f["factor"] for f in factors],
"confidence": "high"
})
elif len(factors) == 1:
root_causes.append({
"cause": factors[0]["factor"],
"category": category,
"confidence": "medium"
})
return root_causes
def _has_delayed_response(self, timeline_data: Dict) -> bool:
"""Check if timeline shows delayed response patterns."""
if not timeline_data or "gap_analysis" not in timeline_data:
return False
gaps = timeline_data["gap_analysis"].get("gaps", [])
return any(gap.get("type") == "phase_transition" for gap in gaps)
def _extract_decision_points(self, timeline_data: Dict) -> List[Dict]:
"""Extract key decision points from timeline."""
decision_points = []
if "timeline" in timeline_data and "phases" in timeline_data["timeline"]:
phases = timeline_data["timeline"]["phases"]
for i, phase in enumerate(phases):
if phase["name"] in ["escalation", "mitigation"]:
decision_points.append({
"timestamp": phase["start_time"],
"decision": f"Initiated {phase['name']} phase",
"phase": phase["name"],
"duration": phase["duration_minutes"]
})
return decision_points
def _identify_missed_opportunities(self, timeline_data: Dict) -> List[Dict]:
"""Identify missed opportunities from gap analysis."""
missed_opportunities = []
if "gap_analysis" in timeline_data:
gaps = timeline_data["gap_analysis"].get("gaps", [])
for gap in gaps:
if gap.get("severity") == "critical":
missed_opportunities.append({
"opportunity": f"Earlier {gap['type'].replace('_', ' ')}",
"gap_minutes": gap["gap_minutes"],
"potential_impact": "Could have reduced incident duration"
})
return missed_opportunities
def _analyze_response_effectiveness(self, timeline_data: Dict) -> Dict[str, Any]:
"""Analyze the effectiveness of incident response."""
effectiveness = {
"overall_rating": "unknown",
"strengths": [],
"weaknesses": [],
"metrics": {}
}
if "metrics" in timeline_data:
metrics = timeline_data["metrics"]
duration_metrics = metrics.get("duration_metrics", {})
# Analyze response times
time_to_mitigation = duration_metrics.get("time_to_mitigation_minutes", 0)
time_to_resolution = duration_metrics.get("time_to_resolution_minutes", 0)
if time_to_mitigation <= 30:
effectiveness["strengths"].append("Quick mitigation response")
else:
effectiveness["weaknesses"].append("Slow mitigation response")
if time_to_resolution <= 120:
effectiveness["strengths"].append("Fast resolution")
else:
effectiveness["weaknesses"].append("Extended resolution time")
effectiveness["metrics"] = {
"time_to_mitigation": time_to_mitigation,
"time_to_resolution": time_to_resolution
}
# Overall rating based on strengths vs weaknesses
if len(effectiveness["strengths"]) > len(effectiveness["weaknesses"]):
effectiveness["overall_rating"] = "effective"
elif len(effectiveness["weaknesses"]) > len(effectiveness["strengths"]):
effectiveness["overall_rating"] = "needs_improvement"
else:
effectiveness["overall_rating"] = "mixed"
return effectiveness
def _extract_timeline_root_causes(self, decision_points: List, missed_opportunities: List,
response_analysis: Dict) -> List[Dict]:
"""Extract root causes from timeline analysis."""
root_causes = []
# Root causes from missed opportunities
for opportunity in missed_opportunities:
if opportunity["gap_minutes"] > 60: # Significant gaps
root_causes.append({
"cause": f"Delayed response: {opportunity['opportunity']}",
"category": "Process",
"evidence": f"{opportunity['gap_minutes']} minute gap identified",
"confidence": "high"
})
# Root causes from response effectiveness
for weakness in response_analysis.get("weaknesses", []):
root_causes.append({
"cause": weakness,
"category": "Process",
"evidence": "Timeline analysis",
"confidence": "medium"
})
return root_causes
def _identify_threats(self, incident_data: Dict, timeline_data: Optional[Dict]) -> List[Dict]:
"""Identify threats for Bow Tie analysis."""
threats = []
description = incident_data.get("description", "").lower()
if "deployment" in description:
threats.append({"threat": "Defective code deployment", "likelihood": "medium"})
if "load" in description or "traffic" in description:
threats.append({"threat": "Unexpected load increase", "likelihood": "high"})
if "database" in description:
threats.append({"threat": "Database performance degradation", "likelihood": "medium"})
return threats
def _identify_consequences(self, incident_data: Dict) -> List[Dict]:
"""Identify consequences for Bow Tie analysis."""
consequences = []
customer_impact = incident_data.get("customer_impact", "").lower()
business_impact = incident_data.get("business_impact", "").lower()
if "all users" in customer_impact or "complete outage" in customer_impact:
consequences.append({"consequence": "Complete service unavailability", "severity": "critical"})
if "revenue" in business_impact:
consequences.append({"consequence": "Revenue loss", "severity": "high"})
return consequences
def _identify_existing_barriers(self, incident_data: Dict, timeline_data: Optional[Dict]) -> List[Dict]:
"""Identify existing preventive/protective barriers."""
barriers = []
# Look for evidence of existing controls
if timeline_data and "timeline" in timeline_data:
events = timeline_data["timeline"].get("events", [])
for event in events:
message = event.get("message", "").lower()
if "alert" in message or "monitoring" in message:
barriers.append({
"barrier": "Monitoring and alerting system",
"type": "detective",
"effectiveness": "partial"
})
elif "rollback" in message:
barriers.append({
"barrier": "Rollback capability",
"type": "corrective",
"effectiveness": "effective"
})
return barriers
def _recommend_additional_barriers(self, threats: List[Dict], consequences: List[Dict]) -> List[Dict]:
"""Recommend additional barriers based on threats and consequences."""
recommendations = []
for threat in threats:
if "deployment" in threat["threat"].lower():
recommendations.append({
"barrier": "Enhanced pre-deployment testing",
"type": "preventive",
"justification": "Prevent defective deployments reaching production"
})
elif "load" in threat["threat"].lower():
recommendations.append({
"barrier": "Auto-scaling and load shedding",
"type": "preventive",
"justification": "Handle unexpected load increases automatically"
})
return recommendations
def _calculate_rca_confidence(self, analysis_data: Any, incident_data: Dict) -> str:
"""Calculate confidence level for RCA results."""
# Simple heuristic based on available data
confidence_score = 0
# More detailed incident data increases confidence
if incident_data.get("description") and len(incident_data["description"]) > 50:
confidence_score += 1
if incident_data.get("timeline") or incident_data.get("events"):
confidence_score += 2
if incident_data.get("logs") or incident_data.get("monitoring_data"):
confidence_score += 2
# Analysis data completeness
if isinstance(analysis_data, list) and len(analysis_data) > 3:
confidence_score += 1
elif isinstance(analysis_data, dict) and len(analysis_data) > 5:
confidence_score += 1
if confidence_score >= 4:
return "high"
elif confidence_score >= 2:
return "medium"
else:
return "low"
def _generate_lessons_learned(self, incident_data: Dict, timeline_data: Optional[Dict],
rca_results: Dict) -> Dict[str, List[str]]:
"""Generate categorized lessons learned."""
lessons = defaultdict(list)
# Lessons from RCA
root_causes = rca_results.get("root_causes", [])
for root_cause in root_causes:
category = root_cause.get("category", "technical_systems").lower()
category_key = self._map_to_lessons_category(category)
lesson = f"Identified: {root_cause['cause']}"
lessons[category_key].append(lesson)
# Lessons from timeline analysis
if timeline_data and "gap_analysis" in timeline_data:
gaps = timeline_data["gap_analysis"].get("gaps", [])
for gap in gaps:
if gap.get("severity") == "critical":
lessons["response_and_escalation"].append(
f"Response time gap: {gap['type'].replace('_', ' ')} took {gap['gap_minutes']} minutes"
)
# Generic lessons based on incident characteristics
severity = incident_data.get("severity", "").lower()
if severity in ["sev1", "critical"]:
lessons["detection_and_monitoring"].append(
"Critical incidents require immediate detection and alerting"
)
return dict(lessons)
def _map_to_lessons_category(self, category: str) -> str:
"""Map RCA category to lessons learned category."""
mapping = {
"people": "team_and_culture",
"process": "process_and_procedures",
"technology": "technical_systems",
"environment": "technical_systems",
"unknown": "process_and_procedures"
}
return mapping.get(category, "technical_systems")
def _generate_action_items(self, incident_data: Dict, rca_results: Dict,
lessons_learned: Dict) -> List[Dict]:
"""Generate actionable follow-up items."""
action_items = []
# Actions from root causes
root_causes = rca_results.get("root_causes", [])
for root_cause in root_causes:
action_type = self._determine_action_type(root_cause)
action_template = self.action_item_types[action_type]
action_items.append({
"title": f"Address: {root_cause['cause'][:50]}...",
"description": root_cause["cause"],
"type": action_type,
"priority": action_template["priority"],
"timeline": action_template["timeline"],
"owner": "TBD",
"success_criteria": f"Prevent recurrence of {root_cause['cause'][:30]}...",
"related_root_cause": root_cause
})
# Actions from lessons learned
for category, lessons in lessons_learned.items():
if len(lessons) > 1: # Multiple lessons in same category indicate systematic issue
action_items.append({
"title": f"Improve {category.replace('_', ' ')}",
"description": f"Address multiple issues identified in {category}",
"type": "process_improvement",
"priority": "P1",
"timeline": "2-3 weeks",
"owner": "TBD",
"success_criteria": f"Comprehensive review and improvement of {category}"
})
# Standard actions based on severity
severity = incident_data.get("severity", "").lower()
if severity in ["sev1", "critical"]:
action_items.append({
"title": "Conduct comprehensive post-incident review",
"description": "Schedule PIR meeting with all stakeholders",
"type": "process_improvement",
"priority": "P0",
"timeline": "24-48 hours",
"owner": incident_data.get("incident_commander", "TBD"),
"success_criteria": "PIR completed and documented"
})
return action_items
def _determine_action_type(self, root_cause: Dict) -> str:
"""Determine action item type based on root cause."""
cause_text = root_cause.get("cause", "").lower()
category = root_cause.get("category", "").lower()
if any(keyword in cause_text for keyword in ["bug", "error", "failure", "crash"]):
return "immediate_fix"
elif any(keyword in cause_text for keyword in ["monitor", "alert", "detect"]):
return "monitoring_alerting"
elif any(keyword in cause_text for keyword in ["process", "procedure", "review"]):
return "process_improvement"
elif any(keyword in cause_text for keyword in ["document", "runbook", "knowledge"]):
return "documentation"
elif any(keyword in cause_text for keyword in ["training", "skill", "knowledge"]):
return "training"
elif any(keyword in cause_text for keyword in ["architecture", "design", "system"]):
return "architectural"
else:
return "process_improvement" # Default
def _create_timeline_section(self, timeline_data: Optional[Dict], severity: str) -> str:
"""Create timeline section for PIR document."""
if not timeline_data:
return "No detailed timeline available."
timeline_content = []
if "timeline" in timeline_data and "phases" in timeline_data["timeline"]:
timeline_content.append("### Phase Timeline")
timeline_content.append("")
phases = timeline_data["timeline"]["phases"]
for phase in phases:
timeline_content.append(f"**{phase['name'].title()} Phase**")
timeline_content.append(f"- Start: {phase['start_time']}")
timeline_content.append(f"- Duration: {phase['duration_minutes']} minutes")
timeline_content.append(f"- Events: {phase['event_count']}")
timeline_content.append("")
if "metrics" in timeline_data:
metrics = timeline_data["metrics"]
duration_metrics = metrics.get("duration_metrics", {})
timeline_content.append("### Key Metrics")
timeline_content.append("")
timeline_content.append(f"- Total Duration: {duration_metrics.get('total_duration_minutes', 'N/A')} minutes")
timeline_content.append(f"- Time to Mitigation: {duration_metrics.get('time_to_mitigation_minutes', 'N/A')} minutes")
timeline_content.append(f"- Time to Resolution: {duration_metrics.get('time_to_resolution_minutes', 'N/A')} minutes")
timeline_content.append("")
return "\n".join(timeline_content)
def _generate_document_sections(self, incident_info: Dict, rca_results: Dict,
lessons_learned: Dict, action_items: List[Dict],
timeline_section: str) -> Dict[str, str]:
"""Generate all document sections for PIR template."""
sections = {}
# Basic information
sections["incident_title"] = incident_info["title"]
sections["incident_id"] = incident_info["incident_id"]
sections["incident_date"] = incident_info["start_time"].strftime("%Y-%m-%d %H:%M:%S UTC") if incident_info["start_time"] else "Unknown"
sections["duration"] = incident_info["duration"]
sections["severity"] = incident_info["severity"].upper()
sections["status"] = incident_info["status"].title()
sections["incident_commander"] = incident_info["incident_commander"]
sections["responders"] = ", ".join(incident_info["responders"]) if incident_info["responders"] else "TBD"
sections["generation_date"] = datetime.now().strftime("%Y-%m-%d")
# Impact sections
sections["customer_impact"] = incident_info["customer_impact"]
sections["business_impact"] = incident_info["business_impact"]
# Executive summary
sections["executive_summary"] = self._create_executive_summary(incident_info, rca_results)
# Timeline
sections["timeline_section"] = timeline_section
# RCA section
sections["rca_section"] = self._create_rca_section(rca_results)
# What went well/wrong
sections["what_went_well"] = self._create_what_went_well_section(incident_info, rca_results)
sections["what_went_wrong"] = self._create_what_went_wrong_section(rca_results, lessons_learned)
# Lessons learned
sections["lessons_learned"] = self._create_lessons_learned_section(lessons_learned)
# Action items
sections["action_items"] = self._create_action_items_section(action_items)
# Prevention and appendix
sections["prevention_measures"] = self._create_prevention_section(rca_results, action_items)
sections["appendix_section"] = self._create_appendix_section(incident_info)
return sections
def _create_executive_summary(self, incident_info: Dict, rca_results: Dict) -> str:
"""Create executive summary section."""
summary_parts = []
# Incident description
summary_parts.append(f"On {incident_info['start_time'].strftime('%B %d, %Y') if incident_info['start_time'] else 'an unknown date'}, we experienced a {incident_info['severity']} incident affecting {incident_info.get('affected_services', ['our services'])}.")
# Duration and impact
summary_parts.append(f"The incident lasted {incident_info['duration']} and had the following impact: {incident_info['customer_impact']}")
# Root cause summary
root_causes = rca_results.get("root_causes", [])
if root_causes:
primary_cause = root_causes[0]["cause"]
summary_parts.append(f"Root cause analysis identified the primary issue as: {primary_cause}")
# Resolution
summary_parts.append(f"The incident has been {incident_info['status']} and we have identified specific actions to prevent recurrence.")
return " ".join(summary_parts)
def _create_rca_section(self, rca_results: Dict) -> str:
"""Create RCA section content."""
rca_content = []
method = rca_results.get("method", "unknown")
rca_content.append(f"### Analysis Method: {self.rca_frameworks.get(method, {}).get('name', method)}")
rca_content.append("")
if method == "five_whys" and "why_analysis" in rca_results:
rca_content.append("#### Why Analysis")
rca_content.append("")
for i, why in enumerate(rca_results["why_analysis"], 1):
rca_content.append(f"**Why {i}:** {why['question']}")
rca_content.append(f"**Answer:** {why['answer']}")
if why["evidence"]:
rca_content.append(f"**Evidence:** {', '.join(why['evidence'])}")
rca_content.append("")
elif method == "fishbone" and "categories" in rca_results:
rca_content.append("#### Contributing Factor Analysis")
rca_content.append("")
for category, data in rca_results["categories"].items():
if data["factors"]:
rca_content.append(f"**{category}:**")
for factor in data["factors"]:
rca_content.append(f"- {factor['factor']} (likelihood: {factor.get('likelihood', 'unknown')})")
rca_content.append("")
# Root causes summary
root_causes = rca_results.get("root_causes", [])
if root_causes:
rca_content.append("#### Identified Root Causes")
rca_content.append("")
for i, cause in enumerate(root_causes, 1):
rca_content.append(f"{i}. **{cause['cause']}**")
rca_content.append(f" - Category: {cause.get('category', 'Unknown')}")
rca_content.append(f" - Confidence: {cause.get('confidence', 'Unknown')}")
if cause.get("evidence"):
rca_content.append(f" - Evidence: {cause['evidence']}")
rca_content.append("")
return "\n".join(rca_content)
def _create_what_went_well_section(self, incident_info: Dict, rca_results: Dict) -> str:
"""Create what went well section."""
positives = []
# Generic positive aspects
if incident_info["status"] == "resolved":
positives.append("The incident was successfully resolved")
if incident_info["incident_commander"] != "TBD":
positives.append("Incident command was established")
if len(incident_info.get("responders", [])) > 1:
positives.append("Multiple team members collaborated on resolution")
# Analysis-specific positives
if rca_results.get("confidence") == "high":
positives.append("Root cause analysis provided clear insights")
if not positives:
positives.append("Incident response process was followed")
return "\n".join([f"- {positive}" for positive in positives])
def _create_what_went_wrong_section(self, rca_results: Dict, lessons_learned: Dict) -> str:
"""Create what went wrong section."""
issues = []
# Issues from RCA
root_causes = rca_results.get("root_causes", [])
for cause in root_causes[:3]: # Show top 3
issues.append(cause["cause"])
# Issues from lessons learned
for category, lessons in lessons_learned.items():
if lessons:
issues.append(f"{category.replace('_', ' ').title()}: {lessons[0]}")
if not issues:
issues.append("Analysis in progress")
return "\n".join([f"- {issue}" for issue in issues])
def _create_lessons_learned_section(self, lessons_learned: Dict) -> str:
"""Create lessons learned section."""
content = []
for category, lessons in lessons_learned.items():
if lessons:
content.append(f"### {category.replace('_', ' ').title()}")
content.append("")
for lesson in lessons:
content.append(f"- {lesson}")
content.append("")
if not content:
content.append("Lessons learned to be documented following detailed analysis.")
return "\n".join(content)
def _create_action_items_section(self, action_items: List[Dict]) -> str:
"""Create action items section."""
if not action_items:
return "Action items to be defined."
content = []
# Group by priority
priority_groups = defaultdict(list)
for item in action_items:
priority_groups[item.get("priority", "P3")].append(item)
for priority in ["P0", "P1", "P2", "P3"]:
items = priority_groups.get(priority, [])
if items:
content.append(f"### {priority} - {self._get_priority_description(priority)}")
content.append("")
for item in items:
content.append(f"**{item['title']}**")
content.append(f"- Owner: {item.get('owner', 'TBD')}")
content.append(f"- Timeline: {item.get('timeline', 'TBD')}")
content.append(f"- Success Criteria: {item.get('success_criteria', 'TBD')}")
content.append("")
return "\n".join(content)
def _get_priority_description(self, priority: str) -> str:
"""Get human-readable priority description."""
descriptions = {
"P0": "Critical - Immediate Action Required",
"P1": "High Priority - Complete Within 1-2 Weeks",
"P2": "Medium Priority - Complete Within 1 Month",
"P3": "Low Priority - Complete When Capacity Allows"
}
return descriptions.get(priority, "Unknown Priority")
def _create_prevention_section(self, rca_results: Dict, action_items: List[Dict]) -> str:
"""Create prevention and follow-up section."""
content = []
content.append("### Prevention Measures")
content.append("")
content.append("Based on the root cause analysis, the following preventive measures have been identified:")
content.append("")
# Extract prevention-focused action items
prevention_items = [item for item in action_items if "prevent" in item.get("description", "").lower()]
if prevention_items:
for item in prevention_items:
content.append(f"- {item['title']}: {item.get('description', '')}")
else:
content.append("- Implement comprehensive testing for similar scenarios")
content.append("- Improve monitoring and alerting coverage")
content.append("- Enhance error handling and resilience patterns")
content.append("")
content.append("### Follow-up Schedule")
content.append("")
content.append("- 1 week: Review action item progress")
content.append("- 1 month: Evaluate effectiveness of implemented changes")
content.append("- 3 months: Conduct follow-up assessment and update preventive measures")
return "\n".join(content)
def _create_appendix_section(self, incident_info: Dict) -> str:
"""Create appendix section."""
content = []
content.append("### Additional Information")
content.append("")
content.append(f"- Incident ID: {incident_info['incident_id']}")
content.append(f"- Severity Classification: {incident_info['severity']}")
if incident_info.get("affected_services"):
content.append(f"- Affected Services: {', '.join(incident_info['affected_services'])}")
content.append("")
content.append("### References")
content.append("")
content.append("- Incident tracking ticket: [Link TBD]")
content.append("- Monitoring dashboards: [Link TBD]")
content.append("- Communication thread: [Link TBD]")
return "\n".join(content)
def _generate_metadata(self, incident_info: Dict, rca_results: Dict, action_items: List[Dict]) -> Dict[str, Any]:
"""Generate PIR metadata for tracking and analysis."""
return {
"pir_id": f"PIR-{incident_info['incident_id']}",
"incident_severity": incident_info["severity"],
"rca_method": rca_results.get("method", "unknown"),
"rca_confidence": rca_results.get("confidence", "unknown"),
"total_action_items": len(action_items),
"critical_action_items": len([item for item in action_items if item.get("priority") == "P0"]),
"estimated_prevention_timeline": self._estimate_prevention_timeline(action_items),
"categories_affected": list(set(item.get("type", "unknown") for item in action_items)),
"review_completeness": self._assess_review_completeness(incident_info, rca_results, action_items)
}
def _estimate_prevention_timeline(self, action_items: List[Dict]) -> str:
"""Estimate timeline for implementing all prevention measures."""
if not action_items:
return "unknown"
# Find the longest timeline among action items
max_weeks = 0
for item in action_items:
timeline = item.get("timeline", "")
if "week" in timeline:
try:
weeks = int(re.findall(r'\d+', timeline)[0])
max_weeks = max(max_weeks, weeks)
except (IndexError, ValueError):
pass
elif "month" in timeline:
try:
months = int(re.findall(r'\d+', timeline)[0])
max_weeks = max(max_weeks, months * 4)
except (IndexError, ValueError):
pass
if max_weeks == 0:
return "1-2 weeks"
elif max_weeks <= 4:
return f"{max_weeks} weeks"
else:
return f"{max_weeks // 4} months"
def _assess_review_completeness(self, incident_info: Dict, rca_results: Dict, action_items: List[Dict]) -> float:
"""Assess completeness of the PIR (0-1 score)."""
score = 0.0
# Basic information completeness
if incident_info.get("description"):
score += 0.1
if incident_info.get("start_time"):
score += 0.1
if incident_info.get("customer_impact"):
score += 0.1
# RCA completeness
if rca_results.get("root_causes"):
score += 0.2
if rca_results.get("confidence") in ["medium", "high"]:
score += 0.1
# Action items completeness
if action_items:
score += 0.2
if any(item.get("owner") and item["owner"] != "TBD" for item in action_items):
score += 0.1
# Additional factors
if incident_info.get("incident_commander") != "TBD":
score += 0.1
if len(action_items) >= 3: # Multiple action items show thorough analysis
score += 0.1
return min(score, 1.0)
def format_json_output(result: Dict) -> str:
"""Format result as pretty JSON."""
return json.dumps(result, indent=2, ensure_ascii=False)
def format_markdown_output(result: Dict) -> str:
"""Format result as Markdown PIR document."""
return result.get("pir_document", "Error: No PIR document generated")
def format_text_output(result: Dict) -> str:
"""Format result as human-readable summary."""
if "error" in result:
return f"Error: {result['error']}"
metadata = result.get("metadata", {})
incident_info = result.get("incident_info", {})
rca_results = result.get("rca_results", {})
action_items = result.get("action_items", [])
output = []
output.append("=" * 60)
output.append("POST-INCIDENT REVIEW SUMMARY")
output.append("=" * 60)
output.append("")
# Basic info
output.append("INCIDENT INFORMATION:")
output.append(f" PIR ID: {metadata.get('pir_id', 'Unknown')}")
output.append(f" Severity: {incident_info.get('severity', 'Unknown').upper()}")
output.append(f" Duration: {incident_info.get('duration', 'Unknown')}")
output.append(f" Status: {incident_info.get('status', 'Unknown').title()}")
output.append("")
# RCA summary
output.append("ROOT CAUSE ANALYSIS:")
output.append(f" Method: {rca_results.get('method', 'Unknown')}")
output.append(f" Confidence: {rca_results.get('confidence', 'Unknown').title()}")
root_causes = rca_results.get("root_causes", [])
if root_causes:
output.append(f" Root Causes Identified: {len(root_causes)}")
for i, cause in enumerate(root_causes[:3], 1):
output.append(f" {i}. {cause.get('cause', 'Unknown')[:60]}...")
output.append("")
# Action items summary
output.append("ACTION ITEMS:")
output.append(f" Total Actions: {len(action_items)}")
output.append(f" Critical (P0): {metadata.get('critical_action_items', 0)}")
output.append(f" Prevention Timeline: {metadata.get('estimated_prevention_timeline', 'Unknown')}")
if action_items:
output.append(" Top Actions:")
for item in action_items[:3]:
output.append(f" - {item.get('title', 'Unknown')[:50]}...")
output.append("")
# Completeness
completeness = metadata.get("review_completeness", 0) * 100
output.append(f"REVIEW COMPLETENESS: {completeness:.0f}%")
output.append("")
output.append("=" * 60)
return "\n".join(output)
def main():
"""Main function with argument parsing and execution."""
parser = argparse.ArgumentParser(
description="Generate Post-Incident Review documents with RCA and action items",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
python pir_generator.py --incident incident.json --output pir.md
python pir_generator.py --incident incident.json --rca-method fishbone
cat incident.json | python pir_generator.py --format markdown
Incident JSON format:
{
"incident_id": "INC-2024-001",
"title": "Database performance degradation",
"description": "Users experiencing slow response times",
"severity": "sev2",
"start_time": "2024-01-01T12:00:00Z",
"end_time": "2024-01-01T14:30:00Z",
"customer_impact": "50% of users affected by slow page loads",
"business_impact": "Moderate user experience degradation",
"incident_commander": "Alice Smith",
"responders": ["Bob Jones", "Carol Johnson"]
}
"""
)
parser.add_argument(
"--incident", "-i",
help="Incident data file (JSON) or '-' for stdin"
)
parser.add_argument(
"--timeline", "-t",
help="Timeline reconstruction file (JSON)"
)
parser.add_argument(
"--output", "-o",
help="Output file path (default: stdout)"
)
parser.add_argument(
"--format", "-f",
choices=["json", "markdown", "text"],
default="markdown",
help="Output format (default: markdown)"
)
parser.add_argument(
"--rca-method",
choices=["five_whys", "fishbone", "timeline", "bow_tie"],
default="five_whys",
help="Root cause analysis method (default: five_whys)"
)
parser.add_argument(
"--template-type",
choices=["comprehensive", "standard", "brief"],
default="comprehensive",
help="PIR template type (default: comprehensive)"
)
parser.add_argument(
"--action-items",
action="store_true",
help="Generate detailed action items"
)
args = parser.parse_args()
generator = PIRGenerator()
try:
# Read incident data
if args.incident == "-" or (not args.incident and not sys.stdin.isatty()):
# Read from stdin
input_text = sys.stdin.read().strip()
if not input_text:
parser.error("No incident data provided")
incident_data = json.loads(input_text)
elif args.incident:
# Read from file
with open(args.incident, 'r') as f:
incident_data = json.load(f)
else:
parser.error("No incident data specified. Use --incident or pipe data to stdin.")
# Read timeline data if provided
timeline_data = None
if args.timeline:
with open(args.timeline, 'r') as f:
timeline_data = json.load(f)
# Validate incident data
if not isinstance(incident_data, dict):
parser.error("Incident data must be a JSON object")
if not incident_data.get("description") and not incident_data.get("title"):
parser.error("Incident data must contain 'description' or 'title'")
# Generate PIR
result = generator.generate_pir(
incident_data=incident_data,
timeline_data=timeline_data,
rca_method=args.rca_method,
template_type=args.template_type
)
# Format output
if args.format == "json":
output = format_json_output(result)
elif args.format == "markdown":
output = format_markdown_output(result)
else:
output = format_text_output(result)
# Write output
if args.output:
with open(args.output, 'w') as f:
f.write(output)
f.write('\n')
else:
print(output)
except FileNotFoundError as e:
print(f"Error: File not found - {e}", file=sys.stderr)
sys.exit(1)
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON - {e}", file=sys.stderr)
sys.exit(1)
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
sys.exit(1)
if __name__ == "__main__":
main()
FILE:scripts/postmortem_generator.py
#!/usr/bin/env python3
"""
Postmortem Generator - Generate structured postmortem reports with 5-Whys analysis.
Produces comprehensive incident postmortem documents from structured JSON input,
including root cause analysis, contributing factor classification, action item
validation, MTTD/MTTR metrics, and customer impact summaries.
Usage:
python postmortem_generator.py incident_data.json
python postmortem_generator.py incident_data.json --format markdown
python postmortem_generator.py incident_data.json --format json
cat incident_data.json | python postmortem_generator.py
Input:
JSON object with keys: incident, timeline, resolution, action_items, participants.
See SKILL.md for the full input schema.
"""
import argparse
import json
import sys
from datetime import datetime, timezone
from typing import Any, Dict, List, Optional, Tuple
# ---------- Constants and Configuration ----------
VERSION = "1.0.0"
SEVERITY_ORDER = {"SEV0": 0, "SEV1": 1, "SEV2": 2, "SEV3": 3, "SEV4": 4}
FACTOR_CATEGORIES = ("process", "tooling", "human", "environment", "external")
ACTION_TYPES = ("detection", "prevention", "mitigation", "process")
PRIORITY_ORDER = {"P0": 0, "P1": 1, "P2": 2, "P3": 3, "P4": 4}
POSTMORTEM_TARGET_HOURS = 72
# Industry benchmarks for incident response (minutes, except postmortem)
BENCHMARKS = {
"SEV0": {"mttd": 5, "mttr": 60, "mitigate": 30, "declare": 5},
"SEV1": {"mttd": 10, "mttr": 120, "mitigate": 60, "declare": 10},
"SEV2": {"mttd": 30, "mttr": 480, "mitigate": 120, "declare": 30},
"SEV3": {"mttd": 60, "mttr": 1440, "mitigate": 240, "declare": 60},
"SEV4": {"mttd": 120, "mttr": 2880, "mitigate": 480, "declare": 120},
}
CAT_TO_ACTION = {"process": "process", "tooling": "detection", "human": "prevention",
"environment": "mitigation", "external": "prevention"}
CAT_WEIGHT = {"process": 1.0, "tooling": 0.9, "human": 0.8, "environment": 0.7, "external": 0.6}
# Keywords used to classify contributing factors into categories
FACTOR_KEYWORDS = {
"process": ["process", "procedure", "workflow", "review", "approval", "checklist",
"runbook", "documentation", "policy", "standard", "protocol", "canary",
"deployment", "rollback", "change management"],
"tooling": ["tool", "monitor", "alert", "threshold", "automation", "test", "pipeline",
"ci/cd", "observability", "dashboard", "logging", "infrastructure",
"configuration", "config"],
"human": ["training", "knowledge", "experience", "communication", "handoff", "fatigue",
"oversight", "mistake", "error", "misunderstand", "assumption", "awareness"],
"environment": ["load", "traffic", "scale", "capacity", "resource", "network", "hardware",
"region", "latency", "timeout", "connection", "performance", "spike"],
"external": ["vendor", "third-party", "upstream", "downstream", "provider", "api",
"dependency", "partner", "dns", "cdn", "certificate"],
}
# 5-Whys templates per category (each list is 5 why->answer steps)
WHY_TEMPLATES = {
"process": [
"Why did this process gap exist? -> The existing process did not account for this scenario.",
"Why was the scenario not accounted for? -> It was not identified during the last process review.",
"Why was the process review incomplete? -> Reviews focus on known failure modes, not emerging risks.",
"Why are emerging risks not surfaced? -> No systematic mechanism to capture lessons from near-misses.",
"Why is there no near-miss capture mechanism? -> Incident learning is ad-hoc rather than systematic."],
"tooling": [
"Why did the tooling fail to catch this? -> The relevant metric was not monitored or the threshold was misconfigured.",
"Why was the threshold misconfigured? -> It was set during initial deployment and never revisited.",
"Why was it never revisited? -> There is no scheduled review of monitoring configurations.",
"Why is there no scheduled review? -> Monitoring ownership is diffuse across teams.",
"Why is ownership diffuse? -> No clear operational runbook assigns monitoring review responsibilities."],
"human": [
"Why did the human factor contribute? -> The individual lacked context needed to prevent the issue.",
"Why was context lacking? -> Knowledge was siloed and not documented accessibly.",
"Why was knowledge siloed? -> No structured onboarding or knowledge-sharing process for this area.",
"Why is there no knowledge-sharing process? -> Team capacity has been focused on feature delivery.",
"Why is capacity skewed toward features? -> Operational excellence is not weighted equally in planning."],
"environment": [
"Why did the environment cause this failure? -> System capacity was insufficient for the load pattern.",
"Why was capacity insufficient? -> Load projections did not account for this traffic pattern.",
"Why were projections inaccurate? -> Load testing does not replicate production-scale variability.",
"Why doesn't load testing replicate production? -> Test environments lack realistic traffic generators.",
"Why are traffic generators missing? -> Investment in production-like test infrastructure was deferred."],
"external": [
"Why did the external factor cause an incident? -> The system had a hard dependency with no fallback.",
"Why was there no fallback? -> The integration was assumed to be highly available.",
"Why was high availability assumed? -> SLA review of the external dependency was not performed.",
"Why was SLA review skipped? -> No standard checklist for evaluating third-party dependencies.",
"Why is there no evaluation checklist? -> Vendor management practices are informal and undocumented."],
}
THEME_RECS = {
"process": ["Establish a quarterly process review cadence covering change management and deployment procedures.",
"Implement a near-miss tracking system to surface latent risks before they become incidents.",
"Create pre-deployment checklists that require sign-off from the service owner."],
"tooling": ["Schedule quarterly reviews of alerting thresholds and monitoring coverage.",
"Assign explicit monitoring ownership per service in operational runbooks.",
"Invest in synthetic monitoring and canary analysis for critical paths."],
"human": ["Build structured onboarding that covers incident-prone areas and past postmortems.",
"Implement blameless knowledge-sharing sessions after each incident.",
"Balance operational excellence work alongside feature delivery in sprint planning."],
"environment": ["Conduct periodic capacity planning reviews using production traffic replays.",
"Invest in production-like load-testing infrastructure with realistic traffic profiles.",
"Implement auto-scaling policies with validated upper-bound thresholds."],
"external": ["Perform formal SLA reviews for all third-party dependencies annually.",
"Implement circuit breakers and fallbacks for external service integrations.",
"Maintain a dependency registry with risk ratings and contingency plans."],
}
MISSING_ACTION_TEMPLATES = {
"process": "Create or update runbook/checklist to prevent recurrence of this process gap",
"detection": "Add monitoring and alerting to detect this class of issue earlier",
"mitigation": "Implement auto-scaling or circuit-breaker to reduce blast radius",
"prevention": "Add automated safeguards (canary deploy, load test gate) to prevent recurrence",
}
# ---------- Data Model Classes ----------
class IncidentData:
"""Parsed incident metadata."""
def __init__(self, data: Dict[str, Any]) -> None:
self.id: str = data.get("id", "UNKNOWN")
self.title: str = data.get("title", "Untitled Incident")
self.severity: str = data.get("severity", "SEV3").upper()
self.commander: str = data.get("commander", "Unassigned")
self.service: str = data.get("service", "unknown-service")
self.affected_services: List[str] = data.get("affected_services", [])
def to_dict(self) -> Dict[str, Any]:
return {"id": self.id, "title": self.title, "severity": self.severity,
"commander": self.commander, "service": self.service,
"affected_services": self.affected_services}
class TimelineMetrics:
"""MTTD, MTTR, and other timing metrics computed from raw timestamps."""
def __init__(self, timeline: Dict[str, str], severity: str) -> None:
self.severity = severity
self.issue_started = self._parse(timeline.get("issue_started"))
self.detected_at = self._parse(timeline.get("detected_at"))
self.declared_at = self._parse(timeline.get("declared_at"))
self.mitigated_at = self._parse(timeline.get("mitigated_at"))
self.resolved_at = self._parse(timeline.get("resolved_at"))
self.postmortem_at = self._parse(timeline.get("postmortem_at"))
@staticmethod
def _parse(ts: Optional[str]) -> Optional[datetime]:
if ts is None:
return None
for fmt in ("%Y-%m-%dT%H:%M:%SZ", "%Y-%m-%dT%H:%M:%S%z", "%Y-%m-%dT%H:%M:%S"):
try:
dt = datetime.strptime(ts, fmt)
return dt if dt.tzinfo else dt.replace(tzinfo=timezone.utc)
except ValueError:
continue
return None
def _delta_min(self, start: Optional[datetime], end: Optional[datetime]) -> Optional[float]:
if start is None or end is None:
return None
return round((end - start).total_seconds() / 60.0, 1)
@property
def mttd(self) -> Optional[float]:
return self._delta_min(self.issue_started, self.detected_at)
@property
def mttr(self) -> Optional[float]:
return self._delta_min(self.detected_at, self.resolved_at)
@property
def time_to_mitigate(self) -> Optional[float]:
return self._delta_min(self.detected_at, self.mitigated_at)
@property
def time_to_declare(self) -> Optional[float]:
return self._delta_min(self.detected_at, self.declared_at)
@property
def postmortem_timeliness_hours(self) -> Optional[float]:
m = self._delta_min(self.resolved_at, self.postmortem_at)
return round(m / 60.0, 1) if m is not None else None
@property
def postmortem_on_time(self) -> Optional[bool]:
h = self.postmortem_timeliness_hours
return h <= POSTMORTEM_TARGET_HOURS if h is not None else None
def benchmark_comparison(self) -> Dict[str, Dict[str, Any]]:
bench = BENCHMARKS.get(self.severity, BENCHMARKS["SEV3"])
results: Dict[str, Dict[str, Any]] = {}
for name, actual, target in [("mttd", self.mttd, bench["mttd"]),
("mttr", self.mttr, bench["mttr"]),
("time_to_mitigate", self.time_to_mitigate, bench["mitigate"]),
("time_to_declare", self.time_to_declare, bench["declare"])]:
if actual is not None:
results[name] = {"actual_minutes": actual, "benchmark_minutes": target,
"met_benchmark": actual <= target,
"delta_minutes": round(actual - target, 1)}
h = self.postmortem_timeliness_hours
if h is not None:
results["postmortem_timeliness"] = {
"actual_hours": h, "target_hours": POSTMORTEM_TARGET_HOURS,
"met_target": self.postmortem_on_time, "delta_hours": round(h - POSTMORTEM_TARGET_HOURS, 1)}
return results
def to_dict(self) -> Dict[str, Any]:
return {"mttd_minutes": self.mttd, "mttr_minutes": self.mttr,
"time_to_mitigate_minutes": self.time_to_mitigate,
"time_to_declare_minutes": self.time_to_declare,
"postmortem_timeliness_hours": self.postmortem_timeliness_hours,
"postmortem_on_time": self.postmortem_on_time,
"benchmarks": self.benchmark_comparison()}
class ContributingFactor:
"""A classified contributing factor with weight and action-type mapping."""
def __init__(self, description: str, index: int) -> None:
self.description = description
self.index = index
self.category = self._classify()
self.weight = round(max(1.0 - index * 0.15, 0.3) * CAT_WEIGHT.get(self.category, 0.8), 2)
self.mapped_action_type = CAT_TO_ACTION.get(self.category, "process")
def _classify(self) -> str:
lower = self.description.lower()
scores = {cat: sum(1 for kw in kws if kw in lower) for cat, kws in FACTOR_KEYWORDS.items()}
best = max(scores, key=lambda k: scores[k])
return best if scores[best] > 0 else "process"
def to_dict(self) -> Dict[str, Any]:
return {"description": self.description, "category": self.category,
"weight": self.weight, "mapped_action_type": self.mapped_action_type}
class FiveWhysAnalysis:
"""Structured 5-Whys chain for a contributing factor."""
def __init__(self, factor: ContributingFactor) -> None:
self.factor = factor
self.systemic_theme: str = factor.category
self.chain: List[str] = [f"Why? {factor.description}"] + \
WHY_TEMPLATES.get(factor.category, WHY_TEMPLATES["process"])
def to_dict(self) -> Dict[str, Any]:
return {"factor": self.factor.description, "category": self.factor.category,
"chain": self.chain, "systemic_theme": self.systemic_theme}
class ActionItem:
"""Parsed and validated action item."""
def __init__(self, data: Dict[str, Any]) -> None:
self.title: str = data.get("title", "")
self.owner: str = data.get("owner", "")
self.priority: str = data.get("priority", "P3")
self.deadline: str = data.get("deadline", "")
self.type: str = data.get("type", "process")
self.status: str = data.get("status", "open")
self.validation_issues: List[str] = []
self.quality_score: int = 0
self._validate()
def _validate(self) -> None:
self.validation_issues = []
if not self.title:
self.validation_issues.append("Missing title")
if not self.owner:
self.validation_issues.append("Missing owner")
if not self.deadline:
self.validation_issues.append("Missing deadline")
if self.priority not in PRIORITY_ORDER:
self.validation_issues.append(f"Invalid priority: {self.priority}")
if self.type not in ACTION_TYPES:
self.validation_issues.append(f"Invalid type: {self.type}")
self.quality_score = self._score_quality()
def _score_quality(self) -> int:
"""Score 0-100: specific, measurable, achievable."""
s = 0
if len(self.title) > 10: s += 20
if self.owner: s += 20
if self.deadline: s += 20
if self.priority in PRIORITY_ORDER: s += 10
if self.type in ACTION_TYPES: s += 10
if any(kw in self.title.lower() for kw in ["%", "threshold", "within", "before",
"after", "less than", "greater than"]):
s += 10
if len(self.title.split()) >= 5: s += 10
return min(s, 100)
@property
def is_valid(self) -> bool:
return len(self.validation_issues) == 0
@property
def is_past_deadline(self) -> bool:
if not self.deadline or self.status != "open":
return False
try:
dl = datetime.strptime(self.deadline, "%Y-%m-%d").replace(tzinfo=timezone.utc)
return datetime.now(timezone.utc) > dl
except ValueError:
return False
def to_dict(self) -> Dict[str, Any]:
return {"title": self.title, "owner": self.owner, "priority": self.priority,
"deadline": self.deadline, "type": self.type, "status": self.status,
"is_valid": self.is_valid, "validation_issues": self.validation_issues,
"quality_score": self.quality_score, "is_past_deadline": self.is_past_deadline}
class PostmortemReport:
"""Complete postmortem document assembled from all analysis components."""
def __init__(self, raw: Dict[str, Any]) -> None:
self.raw = raw
self.incident = IncidentData(raw.get("incident", {}))
self.timeline = TimelineMetrics(raw.get("timeline", {}), self.incident.severity)
self.resolution: Dict[str, Any] = raw.get("resolution", {})
self.participants: List[Dict[str, str]] = raw.get("participants", [])
# Derived analysis
self.contributing_factors = [ContributingFactor(f, i)
for i, f in enumerate(self.resolution.get("contributing_factors", []))]
self.five_whys = [FiveWhysAnalysis(f) for f in self.contributing_factors]
self.action_items = [ActionItem(a) for a in raw.get("action_items", [])]
self.factor_distribution = self._compute_factor_distribution()
self.coverage_gaps = self._find_coverage_gaps()
self.suggested_actions = self._suggest_missing_actions()
self.theme_recommendations = self._build_theme_recommendations()
def _compute_factor_distribution(self) -> Dict[str, float]:
dist: Dict[str, float] = {c: 0.0 for c in FACTOR_CATEGORIES}
total = sum(f.weight for f in self.contributing_factors) or 1.0
for f in self.contributing_factors:
dist[f.category] += f.weight
return {k: round(v / total * 100, 1) for k, v in dist.items()}
def _find_coverage_gaps(self) -> List[str]:
factor_cats = {f.category for f in self.contributing_factors}
action_types = {a.type for a in self.action_items}
gaps = []
for cat in factor_cats:
expected = CAT_TO_ACTION.get(cat)
if expected and expected not in action_types:
gaps.append(f"No '{expected}' action item to address '{cat}' contributing factor")
return gaps
def _suggest_missing_actions(self) -> List[Dict[str, str]]:
factor_cats = {f.category for f in self.contributing_factors}
action_types = {a.type for a in self.action_items}
suggestions = []
for cat in factor_cats:
expected = CAT_TO_ACTION.get(cat)
if expected and expected not in action_types:
suggestions.append({
"type": expected,
"suggestion": MISSING_ACTION_TEMPLATES.get(expected, "Add an action item for this gap"),
"reason": f"No action item addresses the '{cat}' contributing factor"})
return suggestions
def _build_theme_recommendations(self) -> Dict[str, List[str]]:
seen: Dict[str, List[str]] = {}
for a in self.five_whys:
if a.systemic_theme not in seen:
seen[a.systemic_theme] = THEME_RECS.get(a.systemic_theme, [])
return seen
def customer_impact_summary(self) -> Dict[str, Any]:
impact = self.resolution.get("customer_impact", {})
affected = impact.get("affected_users", 0)
failed_tx = impact.get("failed_transactions", 0)
revenue = impact.get("revenue_impact_usd", 0)
data_loss = impact.get("data_loss", False)
comm_required = affected > 1000 or data_loss or revenue > 10000
sev = "high" if (affected > 10000 or revenue > 50000) else (
"medium" if (affected > 1000 or revenue > 5000) else "low")
return {"affected_users": affected, "failed_transactions": failed_tx,
"revenue_impact_usd": revenue, "data_loss": data_loss,
"data_integrity": "compromised" if data_loss else "intact",
"customer_communication_required": comm_required, "impact_severity": sev}
def executive_summary(self) -> str:
mttr = self.timeline.mttr
ci = self.customer_impact_summary()
mttr_str = f"{mttr:.0f} minutes" if mttr is not None else "unknown duration"
parts = [
f"On {self._fmt_date(self.timeline.issue_started)}, a {self.incident.severity} "
f"incident (\"{self.incident.title}\") impacted the {self.incident.service} service.",
f"The root cause was identified as: {self.resolution.get('root_cause', 'Unknown root cause')}.",
f"The incident was resolved in {mttr_str}, affecting approximately "
f"{ci['affected_users']:,} users with an estimated revenue impact of ,.2f.",
"Data loss was confirmed; affected customers must be notified." if ci["data_loss"]
else "No data loss occurred during this incident."]
return " ".join(parts)
@staticmethod
def _fmt_date(dt: Optional[datetime]) -> str:
return dt.strftime("%Y-%m-%d at %H:%M UTC") if dt else "an unknown date"
def overdue_p1_items(self) -> List[Dict[str, str]]:
return [{"title": a.title, "owner": a.owner, "deadline": a.deadline}
for a in self.action_items if a.priority in ("P0", "P1") and a.is_past_deadline]
def to_dict(self) -> Dict[str, Any]:
return {
"version": VERSION, "incident": self.incident.to_dict(),
"executive_summary": self.executive_summary(),
"timeline_metrics": self.timeline.to_dict(),
"customer_impact": self.customer_impact_summary(),
"root_cause": self.resolution.get("root_cause", ""),
"contributing_factors": [f.to_dict() for f in self.contributing_factors],
"factor_distribution": self.factor_distribution,
"five_whys_analysis": [a.to_dict() for a in self.five_whys],
"theme_recommendations": self.theme_recommendations,
"mitigation_steps": self.resolution.get("mitigation_steps", []),
"permanent_fix": self.resolution.get("permanent_fix", ""),
"action_items": [a.to_dict() for a in self.action_items],
"action_item_coverage_gaps": self.coverage_gaps,
"suggested_actions": self.suggested_actions,
"overdue_p1_items": self.overdue_p1_items(),
"participants": self.participants}
# ---------- Core Analysis Helpers ----------
def _bar(pct: float, width: int = 30) -> str:
"""Render a text-based horizontal bar chart segment."""
filled = int(round(pct / 100 * width))
return "[" + "#" * filled + "." * (width - filled) + "]"
def _generate_lessons(report: PostmortemReport) -> List[str]:
"""Derive lessons learned from the analysis."""
lessons: List[str] = []
bench = BENCHMARKS.get(report.incident.severity, BENCHMARKS["SEV3"])
mttd = report.timeline.mttd
if mttd is not None and mttd > bench["mttd"]:
lessons.append(
f"Detection took {mttd:.0f} minutes, exceeding the {bench['mttd']}-minute "
f"benchmark for {report.incident.severity}. Invest in earlier detection mechanisms.")
dist = report.factor_distribution
dominant = max(dist, key=lambda k: dist[k])
if dist[dominant] >= 50:
lessons.append(
f"The '{dominant}' category accounts for {dist[dominant]:.0f}% of contributing factors. "
f"Targeted improvements in this area will yield the highest return.")
if report.coverage_gaps:
lessons.append(
f"There are {len(report.coverage_gaps)} action item coverage gap(s). "
"Ensure every contributing factor category has a corresponding remediation action.")
avg_q = (sum(a.quality_score for a in report.action_items) / len(report.action_items)
if report.action_items else 0)
if avg_q < 70:
lessons.append(
f"Average action item quality score is {avg_q:.0f}/100. "
"Make action items more specific with measurable targets and clear ownership.")
if report.timeline.postmortem_on_time is False:
h = report.timeline.postmortem_timeliness_hours
lessons.append(
f"Postmortem was held {h:.0f} hours after resolution, exceeding the "
f"{POSTMORTEM_TARGET_HOURS}-hour target. Schedule postmortems sooner to capture context.")
if not lessons:
lessons.append("This incident was handled within benchmarks. Continue reinforcing "
"current practices and share this postmortem for organizational learning.")
return lessons
# ---------- Output Formatters ----------
def format_text(report: PostmortemReport) -> str:
"""Format the postmortem as plain text."""
L: List[str] = []
W = 72
def h1(title: str) -> None:
L.append(""); L.append("=" * W); L.append(f" {title}"); L.append("=" * W)
def h2(title: str) -> None:
L.append(""); L.append(f"--- {title} ---")
inc = report.incident
h1(f"POSTMORTEM: {inc.title}")
L.append(f" ID: {inc.id} | Severity: {inc.severity} | Service: {inc.service}")
L.append(f" Commander: {inc.commander}")
if inc.affected_services:
L.append(f" Affected services: {', '.join(inc.affected_services)}")
# Executive Summary
h1("EXECUTIVE SUMMARY")
L.append("")
for sentence in report.executive_summary().split(". "):
s = sentence.strip()
if s and not s.endswith("."): s += "."
if s: L.append(f" {s}")
# Timeline Metrics
h1("TIMELINE METRICS")
tm = report.timeline
L.append("")
for label, val, unit in [("MTTD (Time to Detect)", tm.mttd, "min"),
("MTTR (Time to Resolve)", tm.mttr, "min"),
("Time to Mitigate", tm.time_to_mitigate, "min"),
("Time to Declare", tm.time_to_declare, "min"),
("Postmortem Timeliness", tm.postmortem_timeliness_hours, "hrs")]:
L.append(f" {label:<30s} {f'{val:.1f} {unit}' if val is not None else 'N/A'}")
h2("Benchmark Comparison")
for name, d in tm.benchmark_comparison().items():
if "actual_minutes" in d:
st = "PASS" if d["met_benchmark"] else "FAIL"
L.append(f" {name:<25s} actual={d['actual_minutes']}min benchmark={d['benchmark_minutes']}min [{st}]")
elif "actual_hours" in d:
st = "PASS" if d["met_target"] else "FAIL"
L.append(f" {name:<25s} actual={d['actual_hours']}hrs target={d['target_hours']}hrs [{st}]")
# Customer Impact
h1("CUSTOMER IMPACT")
ci = report.customer_impact_summary()
L.append("")
L.append(f" Affected users: {ci['affected_users']:,}")
L.append(f" Failed transactions: {ci['failed_transactions']:,}")
L.append(f" Revenue impact: ,.2f")
L.append(f" Data integrity: {ci['data_integrity']}")
L.append(f" Impact severity: {ci['impact_severity']}")
L.append(f" Comms required: {'Yes' if ci['customer_communication_required'] else 'No'}")
# Root Cause
h1("ROOT CAUSE ANALYSIS")
L.append("")
L.append(f" {report.resolution.get('root_cause', 'Unknown')}")
h2("Contributing Factors")
for f in report.contributing_factors:
L.append(f" [{f.category.upper():<12s} w={f.weight:.2f}] {f.description}")
h2("Factor Distribution")
for cat, pct in sorted(report.factor_distribution.items(), key=lambda x: -x[1]):
if pct > 0:
L.append(f" {cat:<14s} {pct:5.1f}% {_bar(pct)}")
# 5-Whys
h1("5-WHYS ANALYSIS")
for analysis in report.five_whys:
L.append("")
L.append(f" Factor: {analysis.factor.description}")
L.append(f" Theme: {analysis.systemic_theme}")
for i, step in enumerate(analysis.chain):
L.append(f" {i}. {step}")
h2("Theme-Based Recommendations")
for theme, recs in report.theme_recommendations.items():
L.append(f" [{theme.upper()}]")
for rec in recs:
L.append(f" - {rec}")
# Mitigation & Fix
h1("MITIGATION AND RESOLUTION")
h2("Mitigation Steps Taken")
for step in report.resolution.get("mitigation_steps", []):
L.append(f" - {step}")
h2("Permanent Fix")
L.append(f" {report.resolution.get('permanent_fix', 'TBD')}")
# Action Items
h1("ACTION ITEMS")
L.append("")
hdr = f" {'Priority':<10s} {'Type':<14s} {'Owner':<25s} {'Deadline':<12s} {'Quality':<8s} Title"
L.append(hdr)
L.append(" " + "-" * (len(hdr) - 2))
for a in sorted(report.action_items, key=lambda x: PRIORITY_ORDER.get(x.priority, 99)):
flag = " *OVERDUE*" if a.is_past_deadline else ""
L.append(f" {a.priority:<10s} {a.type:<14s} {a.owner:<25s} {a.deadline:<12s} "
f"{a.quality_score:<8d} {a.title}{flag}")
if report.coverage_gaps:
h2("Coverage Gaps")
for gap in report.coverage_gaps:
L.append(f" WARNING: {gap}")
if report.suggested_actions:
h2("Suggested Additional Actions")
for s in report.suggested_actions:
L.append(f" [{s['type'].upper()}] {s['suggestion']}")
L.append(f" Reason: {s['reason']}")
overdue = report.overdue_p1_items()
if overdue:
h2("Overdue P0/P1 Items")
for item in overdue:
L.append(f" OVERDUE: {item['title']} (owner: {item['owner']}, deadline: {item['deadline']})")
# Participants
h1("PARTICIPANTS")
L.append("")
for p in report.participants:
L.append(f" {p.get('name', 'Unknown'):<25s} {p.get('role', '')}")
# Lessons Learned
h1("LESSONS LEARNED")
L.append("")
for i, lesson in enumerate(_generate_lessons(report), 1):
L.append(f" {i}. {lesson}")
L.append("")
L.append("=" * W)
L.append(f" Generated by postmortem_generator v{VERSION}")
L.append("=" * W)
L.append("")
return "\n".join(L)
def format_json(report: PostmortemReport) -> str:
"""Format the postmortem as JSON."""
data = report.to_dict()
data["lessons_learned"] = _generate_lessons(report)
return json.dumps(data, indent=2, default=str)
def format_markdown(report: PostmortemReport) -> str:
"""Format the postmortem as a Markdown document."""
L: List[str] = []
inc = report.incident
L.append(f"# Postmortem: {inc.title}")
L.append("")
L.append("| Field | Value |")
L.append("|-------|-------|")
L.append(f"| **ID** | {inc.id} |")
L.append(f"| **Severity** | {inc.severity} |")
L.append(f"| **Service** | {inc.service} |")
L.append(f"| **Commander** | {inc.commander} |")
if inc.affected_services:
L.append(f"| **Affected Services** | {', '.join(inc.affected_services)} |")
L.append("")
# Executive Summary
L.append("## Executive Summary\n")
L.append(report.executive_summary())
L.append("")
# Timeline Metrics
L.append("## Timeline Metrics\n")
L.append("| Metric | Value | Benchmark | Status |")
L.append("|--------|-------|-----------|--------|")
labels = {"mttd": "MTTD (Time to Detect)", "mttr": "MTTR (Time to Resolve)",
"time_to_mitigate": "Time to Mitigate", "time_to_declare": "Time to Declare",
"postmortem_timeliness": "Postmortem Timeliness"}
for key, label in labels.items():
b = report.timeline.benchmark_comparison().get(key)
if b and "actual_minutes" in b:
st = "PASS" if b["met_benchmark"] else "FAIL"
L.append(f"| {label} | {b['actual_minutes']} min | {b['benchmark_minutes']} min | {st} |")
elif b and "actual_hours" in b:
st = "PASS" if b["met_target"] else "FAIL"
L.append(f"| {label} | {b['actual_hours']} hrs | {b['target_hours']} hrs | {st} |")
L.append("")
# Customer Impact
L.append("## Customer Impact\n")
ci = report.customer_impact_summary()
L.append(f"- **Affected users:** {ci['affected_users']:,}")
L.append(f"- **Failed transactions:** {ci['failed_transactions']:,}")
L.append(f"- **Revenue impact:** ,.2f")
L.append(f"- **Data integrity:** {ci['data_integrity']}")
L.append(f"- **Impact severity:** {ci['impact_severity']}")
L.append(f"- **Customer communication required:** {'Yes' if ci['customer_communication_required'] else 'No'}")
L.append("")
# Root Cause Analysis
L.append("## Root Cause Analysis\n")
L.append(f"**Root cause:** {report.resolution.get('root_cause', 'Unknown')}")
L.append("")
L.append("### Contributing Factors\n")
L.append("| # | Category | Weight | Description |")
L.append("|---|----------|--------|-------------|")
for i, f in enumerate(report.contributing_factors, 1):
L.append(f"| {i} | {f.category} | {f.weight:.2f} | {f.description} |")
L.append("")
L.append("### Factor Distribution\n")
L.append("```")
for cat, pct in sorted(report.factor_distribution.items(), key=lambda x: -x[1]):
if pct > 0:
L.append(f" {cat:<14s} {pct:5.1f}% {_bar(pct, 25)}")
L.append("```")
L.append("")
# 5-Whys
L.append("## 5-Whys Analysis\n")
for analysis in report.five_whys:
L.append(f"### Factor: {analysis.factor.description}")
L.append(f"**Systemic theme:** {analysis.systemic_theme}\n")
for i, step in enumerate(analysis.chain):
L.append(f"{i}. {step}")
L.append("")
L.append("### Theme-Based Recommendations\n")
for theme, recs in report.theme_recommendations.items():
L.append(f"**{theme.capitalize()}:**")
for rec in recs:
L.append(f"- {rec}")
L.append("")
# Mitigation
L.append("## Mitigation and Resolution\n")
L.append("### Mitigation Steps Taken\n")
for step in report.resolution.get("mitigation_steps", []):
L.append(f"- {step}")
L.append("")
L.append("### Permanent Fix\n")
L.append(report.resolution.get("permanent_fix", "TBD"))
L.append("")
# Action Items
L.append("## Action Items\n")
L.append("| Priority | Type | Owner | Deadline | Quality | Title |")
L.append("|----------|------|-------|----------|---------|-------|")
for a in sorted(report.action_items, key=lambda x: PRIORITY_ORDER.get(x.priority, 99)):
flag = " **OVERDUE**" if a.is_past_deadline else ""
L.append(f"| {a.priority} | {a.type} | {a.owner} | {a.deadline} | {a.quality_score}/100 | {a.title}{flag} |")
L.append("")
if report.coverage_gaps:
L.append("### Coverage Gaps\n")
for gap in report.coverage_gaps:
L.append(f"> **WARNING:** {gap}")
L.append("")
if report.suggested_actions:
L.append("### Suggested Additional Actions\n")
for s in report.suggested_actions:
L.append(f"- **[{s['type'].upper()}]** {s['suggestion']}")
L.append(f" - _Reason: {s['reason']}_")
L.append("")
overdue = report.overdue_p1_items()
if overdue:
L.append("### Overdue P0/P1 Items\n")
for item in overdue:
L.append(f"- **{item['title']}** (owner: {item['owner']}, deadline: {item['deadline']})")
L.append("")
# Participants
L.append("## Participants\n")
L.append("| Name | Role |")
L.append("|------|------|")
for p in report.participants:
L.append(f"| {p.get('name', 'Unknown')} | {p.get('role', '')} |")
L.append("")
# Lessons Learned
L.append("## Lessons Learned\n")
for i, lesson in enumerate(_generate_lessons(report), 1):
L.append(f"{i}. {lesson}")
L.append("")
L.append("---")
L.append(f"_Generated by postmortem_generator v{VERSION}_")
L.append("")
return "\n".join(L)
# ---------- Input Loading ----------
def load_input(filepath: Optional[str]) -> Dict[str, Any]:
"""Load incident data from a file path or stdin."""
if filepath:
try:
with open(filepath, "r", encoding="utf-8") as fh:
return json.load(fh)
except FileNotFoundError:
print(f"Error: File not found: {filepath}", file=sys.stderr)
sys.exit(1)
except json.JSONDecodeError as exc:
print(f"Error: Invalid JSON in {filepath}: {exc}", file=sys.stderr)
sys.exit(1)
else:
if sys.stdin.isatty():
print("Error: No input file specified and no data on stdin.", file=sys.stderr)
print("Usage: postmortem_generator.py [data_file] or pipe JSON via stdin.", file=sys.stderr)
sys.exit(1)
try:
return json.load(sys.stdin)
except json.JSONDecodeError as exc:
print(f"Error: Invalid JSON on stdin: {exc}", file=sys.stderr)
sys.exit(1)
def validate_input(data: Dict[str, Any]) -> List[str]:
"""Return a list of validation warnings (non-fatal)."""
warnings: List[str] = []
for key in ("incident", "timeline", "resolution", "action_items"):
if key not in data:
warnings.append(f"Missing '{key}' section")
for ts in ("issue_started", "detected_at", "mitigated_at", "resolved_at"):
if ts not in data.get("timeline", {}):
warnings.append(f"Missing timeline field: {ts}")
res = data.get("resolution", {})
if "root_cause" not in res:
warnings.append("Missing 'root_cause' in resolution")
if not res.get("contributing_factors"):
warnings.append("No contributing factors provided")
return warnings
# ---------- CLI Entry Point ----------
def main() -> None:
"""CLI entry point for postmortem generation."""
parser = argparse.ArgumentParser(
description="Generate structured postmortem reports with 5-Whys analysis.",
epilog="Reads JSON from a file or stdin. Outputs text, JSON, or markdown.")
parser.add_argument("data_file", nargs="?", default=None,
help="JSON file with incident + resolution data (reads stdin if omitted)")
parser.add_argument("--format", choices=["text", "json", "markdown"], default="text",
dest="output_format", help="Output format (default: text)")
args = parser.parse_args()
data = load_input(args.data_file)
warnings = validate_input(data)
for w in warnings:
print(f"Warning: {w}", file=sys.stderr)
report = PostmortemReport(data)
formatters = {"text": format_text, "json": format_json, "markdown": format_markdown}
print(formatters[args.output_format](report))
if __name__ == "__main__":
main()
FILE:scripts/severity_classifier.py
#!/usr/bin/env python3
"""
Severity Classifier - Classify incident severity and generate escalation paths.
Analyses incident data across multiple dimensions (revenue impact, user scope,
data/security risk, service criticality, blast radius) to produce a weighted
severity score and map it to SEV1-SEV4. Generates escalation paths, on-call
routing, SLA impact assessments, and immediate action plans.
Table of Contents:
SeverityLevel - Enum-like severity definitions (SEV1-SEV4)
ImpactAssessment - Parsed impact data from incident input
SeverityScore - Multi-dimensional weighted scoring result
EscalationPath - Generated escalation routing and timelines
ActionPlan - Recommended immediate actions per severity
SLAImpact - SLA breach risk and error-budget assessment
parse_incident_data() - Validate and normalise raw JSON input
compute_dimension_scores() - Score each weighted dimension
classify_severity() - Map composite score to SEV1-SEV4
build_escalation_path() - Generate escalation routing
build_action_plan() - Generate immediate action checklist
assess_sla_impact() - SLA breach risk assessment
format_text() - Human-readable text output
format_json() - Machine-readable JSON output
format_markdown() - Markdown report output
main() - CLI entry point
Usage:
python severity_classifier.py incident.json
python severity_classifier.py incident.json --format json
python severity_classifier.py incident.json --format markdown
cat incident.json | python severity_classifier.py --format text
echo '{"incident":{...}}' | python severity_classifier.py
"""
import argparse
import json
import sys
from dataclasses import dataclass, field, asdict
from datetime import datetime, timezone
from typing import Any, Dict, List, Optional, Tuple
# ---------- Severity Level Definitions ----------------------------------------
class SeverityLevel:
"""Enum-like container for SEV1 through SEV4 definitions."""
SEV1 = "SEV1"
SEV2 = "SEV2"
SEV3 = "SEV3"
SEV4 = "SEV4"
DEFINITIONS: Dict[str, Dict[str, Any]] = {
"SEV1": {
"label": "Critical",
"description": (
"Complete service outage, confirmed data loss or corruption, "
"active security breach, or more than 50% of users affected."
),
"score_threshold": 0.75,
"response_time_minutes": 5,
"update_cadence_minutes": 15,
"executive_notify": True,
"war_room": True,
},
"SEV2": {
"label": "Major",
"description": (
"Significant service degradation, more than 25% of users "
"affected, no viable workaround, or high revenue impact."
),
"score_threshold": 0.50,
"response_time_minutes": 15,
"update_cadence_minutes": 30,
"executive_notify": False,
"war_room": True,
},
"SEV3": {
"label": "Moderate",
"description": (
"Partial degradation with workaround available, fewer than "
"25% of users affected, limited blast radius."
),
"score_threshold": 0.25,
"response_time_minutes": 30,
"update_cadence_minutes": 60,
"executive_notify": False,
"war_room": False,
},
"SEV4": {
"label": "Minor",
"description": (
"Cosmetic issue, low impact, minimal user effect, "
"informational or non-urgent."
),
"score_threshold": 0.0,
"response_time_minutes": 120,
"update_cadence_minutes": 240,
"executive_notify": False,
"war_room": False,
},
}
@classmethod
def from_score(cls, score: float) -> str:
"""Return the severity level string for a given composite score."""
for level in [cls.SEV1, cls.SEV2, cls.SEV3]:
if score >= cls.DEFINITIONS[level]["score_threshold"]:
return level
return cls.SEV4
@classmethod
def get_definition(cls, level: str) -> Dict[str, Any]:
return cls.DEFINITIONS.get(level, cls.DEFINITIONS[cls.SEV4])
# ---------- Configuration Constants -------------------------------------------
DIMENSION_WEIGHTS: Dict[str, float] = {
"revenue_impact": 0.25,
"user_impact_scope": 0.25,
"data_security_risk": 0.20,
"service_criticality": 0.15,
"blast_radius": 0.15,
}
REVENUE_IMPACT_SCORES: Dict[str, float] = {
"critical": 1.0,
"high": 0.8,
"medium": 0.5,
"low": 0.2,
"none": 0.0,
}
DEGRADATION_SCORES: Dict[str, float] = {
"complete": 1.0,
"major": 0.75,
"partial": 0.50,
"minor": 0.25,
"none": 0.0,
}
ERROR_RATE_THRESHOLDS: List[Tuple[float, float]] = [
(50.0, 1.0),
(25.0, 0.8),
(10.0, 0.6),
(5.0, 0.4),
(1.0, 0.2),
]
LATENCY_P99_THRESHOLDS_MS: List[Tuple[float, float]] = [
(10000, 1.0),
(5000, 0.8),
(2000, 0.6),
(1000, 0.4),
(500, 0.2),
]
SLA_TIERS: Dict[str, Dict[str, Any]] = {
"SEV1": {
"target_resolution_hours": 1,
"target_response_minutes": 5,
"sla_percentage": 99.95,
"monthly_error_budget_minutes": 21.6,
},
"SEV2": {
"target_resolution_hours": 4,
"target_response_minutes": 15,
"sla_percentage": 99.9,
"monthly_error_budget_minutes": 43.2,
},
"SEV3": {
"target_resolution_hours": 24,
"target_response_minutes": 60,
"sla_percentage": 99.5,
"monthly_error_budget_minutes": 216.0,
},
"SEV4": {
"target_resolution_hours": 72,
"target_response_minutes": 480,
"sla_percentage": 99.0,
"monthly_error_budget_minutes": 432.0,
},
}
ESCALATION_TEMPLATES: Dict[str, Dict[str, Any]] = {
"SEV1": {
"initial_notify": ["on-call-primary", "on-call-secondary", "engineering-manager"],
"escalate_after_minutes": 15,
"escalate_to": ["vp-engineering", "cto"],
"bridge_required": True,
"status_page_update": True,
"customer_comms": True,
},
"SEV2": {
"initial_notify": ["on-call-primary", "on-call-secondary"],
"escalate_after_minutes": 30,
"escalate_to": ["engineering-manager"],
"bridge_required": True,
"status_page_update": True,
"customer_comms": False,
},
"SEV3": {
"initial_notify": ["on-call-primary"],
"escalate_after_minutes": 120,
"escalate_to": ["on-call-secondary"],
"bridge_required": False,
"status_page_update": False,
"customer_comms": False,
},
"SEV4": {
"initial_notify": ["on-call-primary"],
"escalate_after_minutes": 480,
"escalate_to": [],
"bridge_required": False,
"status_page_update": False,
"customer_comms": False,
},
}
# ---------- Data Model Classes ------------------------------------------------
@dataclass
class ImpactAssessment:
"""Parsed and normalised impact data from incident input."""
revenue_impact: str = "none"
affected_users_percentage: float = 0.0
affected_regions: List[str] = field(default_factory=list)
data_integrity_risk: bool = False
security_breach: bool = False
customer_facing: bool = False
degradation_type: str = "none"
workaround_available: bool = True
@dataclass
class SeverityScore:
"""Multi-dimensional scoring result with per-dimension breakdown."""
composite_score: float = 0.0
severity_level: str = SeverityLevel.SEV4
dimensions: Dict[str, float] = field(default_factory=dict)
weighted_dimensions: Dict[str, float] = field(default_factory=dict)
contributing_factors: List[str] = field(default_factory=list)
auto_escalate_reasons: List[str] = field(default_factory=list)
@dataclass
class EscalationPath:
"""Generated escalation routing and notification schedule."""
severity_level: str = SeverityLevel.SEV4
immediate_notify: List[str] = field(default_factory=list)
escalation_chain: List[Dict[str, Any]] = field(default_factory=list)
cross_team_notify: List[str] = field(default_factory=list)
war_room_required: bool = False
bridge_link: str = ""
status_page_update: bool = False
customer_comms_required: bool = False
suggested_smes: List[str] = field(default_factory=list)
@dataclass
class ActionPlan:
"""Recommended immediate actions checklist for the incident."""
severity_level: str = SeverityLevel.SEV4
immediate_actions: List[str] = field(default_factory=list)
diagnostic_steps: List[str] = field(default_factory=list)
communication_actions: List[str] = field(default_factory=list)
rollback_assessment: Dict[str, Any] = field(default_factory=dict)
@dataclass
class SLAImpact:
"""SLA breach risk and error-budget assessment."""
severity_level: str = SeverityLevel.SEV4
sla_tier: Dict[str, Any] = field(default_factory=dict)
breach_risk: str = "low"
error_budget_impact_minutes: float = 0.0
remaining_budget_percentage: float = 100.0
estimated_time_to_breach_minutes: float = 0.0
recommendations: List[str] = field(default_factory=list)
# ---------- Input Parsing -----------------------------------------------------
def parse_incident_data(raw: Dict[str, Any]) -> Tuple[Dict, ImpactAssessment, Dict, Dict]:
"""
Validate and normalise raw JSON input into typed structures.
Returns:
(incident_info, impact_assessment, signals, context)
"""
incident = raw.get("incident", {})
if not incident:
raise ValueError("Input must contain an 'incident' key with title and description.")
impact_raw = raw.get("impact", {})
impact = ImpactAssessment(
revenue_impact=impact_raw.get("revenue_impact", "none"),
affected_users_percentage=float(impact_raw.get("affected_users_percentage", 0)),
affected_regions=impact_raw.get("affected_regions", []),
data_integrity_risk=bool(impact_raw.get("data_integrity_risk", False)),
security_breach=bool(impact_raw.get("security_breach", False)),
customer_facing=bool(impact_raw.get("customer_facing", False)),
degradation_type=impact_raw.get("degradation_type", "none"),
workaround_available=bool(impact_raw.get("workaround_available", True)),
)
signals = raw.get("signals", {})
context = raw.get("context", {})
return incident, impact, signals, context
# ---------- Core Scoring Engine -----------------------------------------------
def _score_revenue_impact(impact: ImpactAssessment) -> Tuple[float, List[str]]:
"""Score the revenue impact dimension (0.0 - 1.0)."""
factors: List[str] = []
score = REVENUE_IMPACT_SCORES.get(impact.revenue_impact, 0.0)
if impact.customer_facing and score >= 0.5:
score = min(1.0, score + 0.1)
factors.append("Customer-facing service with revenue exposure")
if not impact.workaround_available and score >= 0.5:
score = min(1.0, score + 0.1)
factors.append("No workaround available, prolonging revenue impact")
if score >= 0.8:
factors.append(f"Revenue impact rated '{impact.revenue_impact}'")
return score, factors
def _score_user_impact(impact: ImpactAssessment, signals: Dict) -> Tuple[float, List[str]]:
"""Score the user impact scope dimension (0.0 - 1.0)."""
factors: List[str] = []
pct = impact.affected_users_percentage
if pct >= 75:
score = 1.0
elif pct >= 50:
score = 0.85
elif pct >= 25:
score = 0.65
elif pct >= 10:
score = 0.45
elif pct >= 1:
score = 0.25
else:
score = 0.1
if pct > 0:
factors.append(f"{pct}% of users affected")
customer_reports = signals.get("customer_reports", 0)
if customer_reports > 20:
score = min(1.0, score + 0.15)
factors.append(f"{customer_reports} customer reports received")
elif customer_reports > 5:
score = min(1.0, score + 0.08)
factors.append(f"{customer_reports} customer reports received")
degradation_boost = DEGRADATION_SCORES.get(impact.degradation_type, 0.0) * 0.15
score = min(1.0, score + degradation_boost)
if impact.degradation_type in ("complete", "major"):
factors.append(f"Degradation type: {impact.degradation_type}")
return score, factors
def _score_data_security(impact: ImpactAssessment) -> Tuple[float, List[str]]:
"""Score the data/security risk dimension (0.0 - 1.0)."""
factors: List[str] = []
score = 0.0
if impact.security_breach:
score = 1.0
factors.append("Active security breach confirmed")
elif impact.data_integrity_risk:
score = 0.8
factors.append("Data integrity at risk")
if impact.customer_facing and impact.data_integrity_risk:
score = min(1.0, score + 0.1)
factors.append("Customer data potentially affected")
return score, factors
def _score_service_criticality(signals: Dict, context: Dict) -> Tuple[float, List[str]]:
"""Score service criticality based on signals and dependency graph."""
factors: List[str] = []
score = 0.0
dependent_services = signals.get("dependent_services", [])
dep_count = len(dependent_services)
if dep_count >= 5:
score = 1.0
factors.append(f"{dep_count} dependent services (critical hub)")
elif dep_count >= 3:
score = 0.75
factors.append(f"{dep_count} dependent services")
elif dep_count >= 1:
score = 0.5
factors.append(f"{dep_count} dependent service(s)")
else:
score = 0.2
affected_endpoints = signals.get("affected_endpoints", [])
if len(affected_endpoints) >= 5:
score = min(1.0, score + 0.15)
factors.append(f"{len(affected_endpoints)} endpoints affected")
elif len(affected_endpoints) >= 2:
score = min(1.0, score + 0.08)
factors.append(f"{len(affected_endpoints)} endpoints affected")
return score, factors
def _score_blast_radius(
impact: ImpactAssessment, signals: Dict
) -> Tuple[float, List[str]]:
"""Score blast radius from region spread, alert volume, and error rate."""
factors: List[str] = []
score = 0.0
region_count = len(impact.affected_regions)
if region_count >= 3:
score = 0.9
factors.append(f"Spanning {region_count} regions")
elif region_count == 2:
score = 0.6
factors.append(f"Spanning {region_count} regions")
elif region_count == 1:
score = 0.3
error_rate = signals.get("error_rate_percentage", 0.0)
for threshold, rate_score in ERROR_RATE_THRESHOLDS:
if error_rate >= threshold:
score = max(score, rate_score)
factors.append(f"Error rate at {error_rate}%")
break
latency = signals.get("latency_p99_ms", 0)
for threshold, lat_score in LATENCY_P99_THRESHOLDS_MS:
if latency >= threshold:
score = max(score, lat_score)
factors.append(f"P99 latency at {latency}ms")
break
alert_count = signals.get("alert_count", 0)
if alert_count >= 20:
score = min(1.0, score + 0.15)
factors.append(f"{alert_count} alerts firing")
elif alert_count >= 10:
score = min(1.0, score + 0.08)
factors.append(f"{alert_count} alerts firing")
return score, factors
def compute_dimension_scores(
impact: ImpactAssessment, signals: Dict, context: Dict
) -> SeverityScore:
"""Score each weighted dimension and produce a composite severity score."""
dimensions: Dict[str, float] = {}
weighted: Dict[str, float] = {}
all_factors: List[str] = []
auto_escalate: List[str] = []
# -- Revenue impact --
rev_score, rev_factors = _score_revenue_impact(impact)
dimensions["revenue_impact"] = round(rev_score, 3)
weighted["revenue_impact"] = round(rev_score * DIMENSION_WEIGHTS["revenue_impact"], 3)
all_factors.extend(rev_factors)
# -- User impact scope --
user_score, user_factors = _score_user_impact(impact, signals)
dimensions["user_impact_scope"] = round(user_score, 3)
weighted["user_impact_scope"] = round(user_score * DIMENSION_WEIGHTS["user_impact_scope"], 3)
all_factors.extend(user_factors)
# -- Data / security risk --
sec_score, sec_factors = _score_data_security(impact)
dimensions["data_security_risk"] = round(sec_score, 3)
weighted["data_security_risk"] = round(sec_score * DIMENSION_WEIGHTS["data_security_risk"], 3)
all_factors.extend(sec_factors)
# -- Service criticality --
svc_score, svc_factors = _score_service_criticality(signals, context)
dimensions["service_criticality"] = round(svc_score, 3)
weighted["service_criticality"] = round(svc_score * DIMENSION_WEIGHTS["service_criticality"], 3)
all_factors.extend(svc_factors)
# -- Blast radius --
blast_score, blast_factors = _score_blast_radius(impact, signals)
dimensions["blast_radius"] = round(blast_score, 3)
weighted["blast_radius"] = round(blast_score * DIMENSION_WEIGHTS["blast_radius"], 3)
all_factors.extend(blast_factors)
composite = sum(weighted.values())
# -- Auto-escalation overrides --
if impact.security_breach:
composite = max(composite, 0.85)
auto_escalate.append("Security breach triggers automatic SEV1 escalation")
if impact.data_integrity_risk and impact.customer_facing:
composite = max(composite, 0.76)
auto_escalate.append("Customer-facing data integrity risk triggers SEV1 floor")
if impact.affected_users_percentage >= 50 and impact.degradation_type == "complete":
composite = max(composite, 0.80)
auto_escalate.append("Complete outage affecting 50%+ users triggers SEV1 floor")
composite = min(1.0, round(composite, 3))
severity_level = SeverityLevel.from_score(composite)
return SeverityScore(
composite_score=composite,
severity_level=severity_level,
dimensions=dimensions,
weighted_dimensions=weighted,
contributing_factors=all_factors,
auto_escalate_reasons=auto_escalate,
)
# ---------- Classification Wrapper --------------------------------------------
def classify_severity(
incident: Dict, impact: ImpactAssessment, signals: Dict, context: Dict
) -> SeverityScore:
"""
Top-level classification: compute scores and return the final
SeverityScore including the resolved severity level.
"""
return compute_dimension_scores(impact, signals, context)
# ---------- Escalation Path Builder -------------------------------------------
def build_escalation_path(
severity_score: SeverityScore,
signals: Dict,
context: Dict,
) -> EscalationPath:
"""Generate the escalation routing based on severity and context."""
level = severity_score.severity_level
template = ESCALATION_TEMPLATES.get(level, ESCALATION_TEMPLATES["SEV4"])
on_call = context.get("on_call", {})
primary = on_call.get("primary", "on-call-primary@company.com")
secondary = on_call.get("secondary", "on-call-secondary@company.com")
immediate: List[str] = []
for role in template["initial_notify"]:
if role == "on-call-primary":
immediate.append(primary)
elif role == "on-call-secondary":
immediate.append(secondary)
else:
immediate.append(role)
chain: List[Dict[str, Any]] = []
if template["escalate_to"]:
chain.append({
"trigger_after_minutes": template["escalate_after_minutes"],
"notify": template["escalate_to"],
"reason": f"No resolution within {template['escalate_after_minutes']} minutes",
})
sev_def = SeverityLevel.get_definition(level)
if sev_def.get("executive_notify"):
chain.append({
"trigger_after_minutes": 15,
"notify": ["vp-engineering", "cto"],
"reason": "SEV1 executive notification policy",
})
cross_team: List[str] = []
dependent_services = signals.get("dependent_services", [])
for svc in dependent_services:
cross_team.append(f"{svc}-team")
suggested_smes: List[str] = []
affected_endpoints = signals.get("affected_endpoints", [])
if affected_endpoints:
suggested_smes.append(f"API owner for: {', '.join(affected_endpoints[:3])}")
if dependent_services:
suggested_smes.append(f"Service owners: {', '.join(dependent_services[:3])}")
ongoing = context.get("ongoing_incidents", [])
if ongoing:
suggested_smes.append("Incident coordinator (multiple active incidents)")
bridge_link = ""
if template["bridge_required"]:
bridge_link = f"https://bridge.company.com/incident-{level.lower()}"
return EscalationPath(
severity_level=level,
immediate_notify=immediate,
escalation_chain=chain,
cross_team_notify=cross_team,
war_room_required=template["bridge_required"],
bridge_link=bridge_link,
status_page_update=template["status_page_update"],
customer_comms_required=template.get("customer_comms", False),
suggested_smes=suggested_smes,
)
# ---------- Action Plan Builder -----------------------------------------------
def build_action_plan(
severity_score: SeverityScore,
incident: Dict,
impact: ImpactAssessment,
signals: Dict,
context: Dict,
) -> ActionPlan:
"""Generate the immediate action plan for the classified incident."""
level = severity_score.severity_level
sev_def = SeverityLevel.get_definition(level)
# -- Immediate actions --
immediate: List[str] = [
f"Acknowledge incident within {sev_def['response_time_minutes']} minutes",
"Join the war room / bridge call" if sev_def["war_room"] else "Open incident channel",
f"Post status update every {sev_def['update_cadence_minutes']} minutes",
]
if level in (SeverityLevel.SEV1, SeverityLevel.SEV2):
immediate.append("Page secondary on-call if primary unresponsive within 5 minutes")
immediate.append("Begin impact quantification for executive update")
if impact.security_breach:
immediate.insert(0, "CRITICAL: Initiate security incident response playbook")
immediate.append("Engage security team immediately")
immediate.append("Preserve forensic evidence -- do not restart services yet")
if impact.data_integrity_risk:
immediate.append("Halt writes to affected data stores if safe to do so")
immediate.append("Begin data integrity verification")
# -- Diagnostic steps --
diagnostics: List[str] = [
"Check service dashboards and recent metric trends",
"Review application logs for error spikes",
"Verify upstream and downstream dependency health",
]
error_rate = signals.get("error_rate_percentage", 0)
if error_rate > 10:
diagnostics.append(f"Investigate error rate spike ({error_rate}%)")
latency = signals.get("latency_p99_ms", 0)
if latency > 2000:
diagnostics.append(f"Investigate latency degradation (P99 = {latency}ms)")
affected_endpoints = signals.get("affected_endpoints", [])
if affected_endpoints:
diagnostics.append(
f"Trace requests to affected endpoints: {', '.join(affected_endpoints[:5])}"
)
dependent_services = signals.get("dependent_services", [])
if dependent_services:
diagnostics.append(
f"Check health of dependent services: {', '.join(dependent_services)}"
)
# -- Communication actions --
comms: List[str] = []
if sev_def.get("executive_notify"):
comms.append("Draft executive summary within 15 minutes")
if level in (SeverityLevel.SEV1, SeverityLevel.SEV2):
comms.append("Post initial status page update")
comms.append("Notify customer success team for proactive outreach")
comms.append(f"Schedule post-incident review within 48 hours")
# -- Rollback assessment --
recent_deploys = context.get("recent_deployments", [])
rollback: Dict[str, Any] = {"recent_deployment_detected": False, "recommendation": ""}
if recent_deploys:
latest = recent_deploys[0]
rollback["recent_deployment_detected"] = True
rollback["service"] = latest.get("service", "unknown")
rollback["version"] = latest.get("version", "unknown")
rollback["deployed_at"] = latest.get("deployed_at", "unknown")
detected_at = incident.get("detected_at", "")
deploy_time = latest.get("deployed_at", "")
if detected_at and deploy_time:
try:
det = datetime.fromisoformat(detected_at.replace("Z", "+00:00"))
dep = datetime.fromisoformat(deploy_time.replace("Z", "+00:00"))
delta_minutes = (det - dep).total_seconds() / 60
rollback["minutes_since_deploy"] = round(delta_minutes, 1)
if 0 < delta_minutes < 120:
rollback["recommendation"] = (
f"STRONG: Deployment of {latest.get('service')} v{latest.get('version')} "
f"occurred {round(delta_minutes)} minutes before detection. "
"Consider immediate rollback."
)
else:
rollback["recommendation"] = (
"Recent deployment is outside the typical correlation window. "
"Investigate other root causes first."
)
except (ValueError, TypeError):
rollback["recommendation"] = (
"Unable to parse timestamps. Manually assess deployment correlation."
)
else:
rollback["recommendation"] = (
"No recent deployments detected. Focus on infrastructure and dependency investigation."
)
return ActionPlan(
severity_level=level,
immediate_actions=immediate,
diagnostic_steps=diagnostics,
communication_actions=comms,
rollback_assessment=rollback,
)
# ---------- SLA Impact Assessment ---------------------------------------------
def assess_sla_impact(
severity_score: SeverityScore,
impact: ImpactAssessment,
signals: Dict,
) -> SLAImpact:
"""Calculate SLA breach risk and error-budget consumption."""
level = severity_score.severity_level
tier = SLA_TIERS.get(level, SLA_TIERS["SEV4"])
# Estimate ongoing burn rate (minutes of budget consumed per real minute)
user_pct = impact.affected_users_percentage / 100.0
degradation_factor = DEGRADATION_SCORES.get(impact.degradation_type, 0.25)
burn_rate = user_pct * degradation_factor
if burn_rate <= 0:
burn_rate = 0.01 # minimum if incident is open
monthly_budget = tier["monthly_error_budget_minutes"]
# Assume 30% of budget already consumed this month for conservative estimate
assumed_consumed_pct = 30.0
remaining_budget = monthly_budget * (1 - assumed_consumed_pct / 100.0)
if burn_rate > 0:
time_to_breach = remaining_budget / burn_rate
else:
time_to_breach = float("inf")
# Classify breach risk
if time_to_breach <= 30:
breach_risk = "critical"
elif time_to_breach <= 120:
breach_risk = "high"
elif time_to_breach <= 480:
breach_risk = "medium"
else:
breach_risk = "low"
budget_impact_per_hour = burn_rate * 60
error_budget_impact = round(budget_impact_per_hour, 2)
remaining_pct = round(
max(0.0, (remaining_budget / monthly_budget) * 100.0), 1
)
recommendations: List[str] = []
if breach_risk == "critical":
recommendations.append(
"SLA breach imminent. Prioritize resolution above all other work."
)
recommendations.append(
"Prepare customer communication about potential SLA credit."
)
elif breach_risk == "high":
recommendations.append(
"SLA breach likely within hours. Escalate to ensure rapid resolution."
)
elif breach_risk == "medium":
recommendations.append(
"Monitor error budget consumption. Resolve before end of business."
)
else:
recommendations.append(
"SLA impact is contained. Continue standard incident response."
)
recommendations.append(
f"Current burn rate: {round(burn_rate * 100, 1)}% of error budget per minute"
)
recommendations.append(
f"Estimated time to SLA breach: {round(time_to_breach, 0)} minutes "
f"({round(time_to_breach / 60, 1)} hours)"
)
return SLAImpact(
severity_level=level,
sla_tier=tier,
breach_risk=breach_risk,
error_budget_impact_minutes=error_budget_impact,
remaining_budget_percentage=remaining_pct,
estimated_time_to_breach_minutes=round(time_to_breach, 1),
recommendations=recommendations,
)
# ---------- Output Formatters -------------------------------------------------
def _header_line(char: str, width: int = 72) -> str:
return char * width
def format_text(
incident: Dict,
severity_score: SeverityScore,
escalation: EscalationPath,
action_plan: ActionPlan,
sla_impact: SLAImpact,
) -> str:
"""Render a human-readable text report."""
lines: List[str] = []
w = 72
lines.append(_header_line("=", w))
lines.append("INCIDENT SEVERITY CLASSIFICATION REPORT")
lines.append(_header_line("=", w))
lines.append("")
# -- Incident Summary --
lines.append(f"Title: {incident.get('title', 'N/A')}")
lines.append(f"Service: {incident.get('service', 'N/A')}")
lines.append(f"Detected: {incident.get('detected_at', 'N/A')}")
lines.append(f"Reporter: {incident.get('reporter', 'N/A')}")
lines.append("")
# -- Severity --
sev_def = SeverityLevel.get_definition(severity_score.severity_level)
lines.append(_header_line("-", w))
lines.append(f"SEVERITY: {severity_score.severity_level} ({sev_def['label']})")
lines.append(f"Composite Score: {severity_score.composite_score:.3f}")
lines.append(_header_line("-", w))
lines.append(f" {sev_def['description']}")
lines.append("")
# -- Dimension Breakdown --
lines.append("Dimension Scores:")
for dim, raw in severity_score.dimensions.items():
wt = severity_score.weighted_dimensions.get(dim, 0)
weight_cfg = DIMENSION_WEIGHTS.get(dim, 0)
label = dim.replace("_", " ").title()
lines.append(f" {label:<25s} raw={raw:.3f} weight={weight_cfg:.2f} weighted={wt:.3f}")
lines.append("")
if severity_score.contributing_factors:
lines.append("Contributing Factors:")
for f in severity_score.contributing_factors:
lines.append(f" - {f}")
lines.append("")
if severity_score.auto_escalate_reasons:
lines.append("Auto-Escalation Overrides:")
for r in severity_score.auto_escalate_reasons:
lines.append(f" * {r}")
lines.append("")
# -- Escalation Path --
lines.append(_header_line("-", w))
lines.append("ESCALATION PATH")
lines.append(_header_line("-", w))
lines.append(f"Immediate Notify: {', '.join(escalation.immediate_notify)}")
if escalation.war_room_required:
lines.append(f"War Room: Required ({escalation.bridge_link})")
else:
lines.append("War Room: Not required")
lines.append(f"Status Page: {'Update required' if escalation.status_page_update else 'No update needed'}")
lines.append(f"Customer Comms: {'Required' if escalation.customer_comms_required else 'Not required'}")
lines.append("")
if escalation.escalation_chain:
lines.append("Escalation Chain:")
for step in escalation.escalation_chain:
lines.append(
f" After {step['trigger_after_minutes']}min -> "
f"Notify: {', '.join(step['notify'])} ({step['reason']})"
)
lines.append("")
if escalation.cross_team_notify:
lines.append(f"Cross-Team Notify: {', '.join(escalation.cross_team_notify)}")
if escalation.suggested_smes:
lines.append("Suggested SMEs:")
for sme in escalation.suggested_smes:
lines.append(f" - {sme}")
lines.append("")
# -- Action Plan --
lines.append(_header_line("-", w))
lines.append("ACTION PLAN")
lines.append(_header_line("-", w))
lines.append("Immediate Actions:")
for i, action in enumerate(action_plan.immediate_actions, 1):
lines.append(f" {i}. {action}")
lines.append("")
lines.append("Diagnostic Steps:")
for i, step in enumerate(action_plan.diagnostic_steps, 1):
lines.append(f" {i}. {step}")
lines.append("")
lines.append("Communication Actions:")
for i, action in enumerate(action_plan.communication_actions, 1):
lines.append(f" {i}. {action}")
lines.append("")
rb = action_plan.rollback_assessment
lines.append("Rollback Assessment:")
if rb.get("recent_deployment_detected"):
lines.append(f" Recent Deploy: {rb.get('service', '?')} v{rb.get('version', '?')}")
lines.append(f" Deployed At: {rb.get('deployed_at', '?')}")
if "minutes_since_deploy" in rb:
lines.append(f" Minutes Before Detection: {rb['minutes_since_deploy']}")
lines.append(f" Recommendation: {rb.get('recommendation', 'N/A')}")
lines.append("")
# -- SLA Impact --
lines.append(_header_line("-", w))
lines.append("SLA IMPACT ASSESSMENT")
lines.append(_header_line("-", w))
lines.append(f"Breach Risk: {sla_impact.breach_risk.upper()}")
lines.append(f"Error Budget Impact: {sla_impact.error_budget_impact_minutes} min/hr")
lines.append(f"Remaining Budget: {sla_impact.remaining_budget_percentage}%")
lines.append(f"Est. Time to Breach: {sla_impact.estimated_time_to_breach_minutes} min")
tier = sla_impact.sla_tier
lines.append(f"Target Resolution: {tier.get('target_resolution_hours', '?')} hours")
lines.append(f"Target Response: {tier.get('target_response_minutes', '?')} minutes")
lines.append("")
if sla_impact.recommendations:
lines.append("SLA Recommendations:")
for rec in sla_impact.recommendations:
lines.append(f" - {rec}")
lines.append("")
lines.append(_header_line("=", w))
return "\n".join(lines)
def format_json(
incident: Dict,
severity_score: SeverityScore,
escalation: EscalationPath,
action_plan: ActionPlan,
sla_impact: SLAImpact,
) -> str:
"""Render a machine-readable JSON report."""
report = {
"classification_timestamp": datetime.now(timezone.utc).isoformat(),
"incident": incident,
"severity": asdict(severity_score),
"severity_definition": SeverityLevel.get_definition(severity_score.severity_level),
"escalation": asdict(escalation),
"action_plan": asdict(action_plan),
"sla_impact": asdict(sla_impact),
}
return json.dumps(report, indent=2, default=str)
def format_markdown(
incident: Dict,
severity_score: SeverityScore,
escalation: EscalationPath,
action_plan: ActionPlan,
sla_impact: SLAImpact,
) -> str:
"""Render a Markdown report suitable for incident tickets or wikis."""
lines: List[str] = []
sev_def = SeverityLevel.get_definition(severity_score.severity_level)
lines.append(f"# Incident Severity Classification: {severity_score.severity_level}")
lines.append("")
lines.append(f"**Classified:** {datetime.now(timezone.utc).strftime('%Y-%m-%d %H:%M UTC')}")
lines.append("")
lines.append("## Incident Summary")
lines.append("")
lines.append(f"| Field | Value |")
lines.append(f"|-------|-------|")
lines.append(f"| Title | {incident.get('title', 'N/A')} |")
lines.append(f"| Service | {incident.get('service', 'N/A')} |")
lines.append(f"| Detected | {incident.get('detected_at', 'N/A')} |")
lines.append(f"| Reporter | {incident.get('reporter', 'N/A')} |")
lines.append("")
lines.append("## Severity Classification")
lines.append("")
lines.append(
f"> **{severity_score.severity_level} -- {sev_def['label']}** "
f"(Score: {severity_score.composite_score:.3f})"
)
lines.append(f">")
lines.append(f"> {sev_def['description']}")
lines.append("")
lines.append("### Dimension Scores")
lines.append("")
lines.append("| Dimension | Raw | Weight | Weighted |")
lines.append("|-----------|-----|--------|----------|")
for dim, raw in severity_score.dimensions.items():
wt = severity_score.weighted_dimensions.get(dim, 0)
weight_cfg = DIMENSION_WEIGHTS.get(dim, 0)
label = dim.replace("_", " ").title()
lines.append(f"| {label} | {raw:.3f} | {weight_cfg:.2f} | {wt:.3f} |")
lines.append("")
if severity_score.contributing_factors:
lines.append("### Contributing Factors")
lines.append("")
for f in severity_score.contributing_factors:
lines.append(f"- {f}")
lines.append("")
if severity_score.auto_escalate_reasons:
lines.append("### Auto-Escalation Overrides")
lines.append("")
for r in severity_score.auto_escalate_reasons:
lines.append(f"- **{r}**")
lines.append("")
lines.append("## Escalation Path")
lines.append("")
lines.append(f"**Immediate Notify:** {', '.join(escalation.immediate_notify)}")
lines.append("")
if escalation.war_room_required:
lines.append(f"**War Room:** [Join Bridge]({escalation.bridge_link})")
else:
lines.append("**War Room:** Not required")
lines.append("")
if escalation.escalation_chain:
lines.append("### Escalation Chain")
lines.append("")
for step in escalation.escalation_chain:
lines.append(
f"- **After {step['trigger_after_minutes']} min:** "
f"Notify {', '.join(step['notify'])} -- {step['reason']}"
)
lines.append("")
if escalation.cross_team_notify:
lines.append(f"**Cross-Team:** {', '.join(escalation.cross_team_notify)}")
lines.append("")
if escalation.suggested_smes:
lines.append("### Suggested SMEs")
lines.append("")
for sme in escalation.suggested_smes:
lines.append(f"- {sme}")
lines.append("")
lines.append("## Action Plan")
lines.append("")
lines.append("### Immediate Actions")
lines.append("")
for i, action in enumerate(action_plan.immediate_actions, 1):
lines.append(f"{i}. {action}")
lines.append("")
lines.append("### Diagnostic Steps")
lines.append("")
for i, step in enumerate(action_plan.diagnostic_steps, 1):
lines.append(f"{i}. {step}")
lines.append("")
lines.append("### Communication")
lines.append("")
for i, action in enumerate(action_plan.communication_actions, 1):
lines.append(f"{i}. {action}")
lines.append("")
rb = action_plan.rollback_assessment
lines.append("### Rollback Assessment")
lines.append("")
if rb.get("recent_deployment_detected"):
lines.append(
f"| Deploy | {rb.get('service', '?')} v{rb.get('version', '?')} |"
)
lines.append(f"|--------|------|")
lines.append(f"| Deployed At | {rb.get('deployed_at', '?')} |")
if "minutes_since_deploy" in rb:
lines.append(f"| Minutes Before Detection | {rb['minutes_since_deploy']} |")
lines.append("")
lines.append(f"**Recommendation:** {rb.get('recommendation', 'N/A')}")
lines.append("")
lines.append("## SLA Impact")
lines.append("")
tier = sla_impact.sla_tier
lines.append(f"| Metric | Value |")
lines.append(f"|--------|-------|")
lines.append(f"| Breach Risk | **{sla_impact.breach_risk.upper()}** |")
lines.append(f"| Error Budget Impact | {sla_impact.error_budget_impact_minutes} min/hr |")
lines.append(f"| Remaining Budget | {sla_impact.remaining_budget_percentage}% |")
lines.append(f"| Est. Time to Breach | {sla_impact.estimated_time_to_breach_minutes} min |")
lines.append(f"| Target Resolution | {tier.get('target_resolution_hours', '?')} hours |")
lines.append(f"| Target Response | {tier.get('target_response_minutes', '?')} minutes |")
lines.append("")
if sla_impact.recommendations:
lines.append("### SLA Recommendations")
lines.append("")
for rec in sla_impact.recommendations:
lines.append(f"- {rec}")
lines.append("")
lines.append("---")
lines.append("*Generated by severity_classifier.py*")
return "\n".join(lines)
# ---------- CLI Entry Point ---------------------------------------------------
def main() -> None:
"""Parse arguments, read input, classify, and emit output."""
parser = argparse.ArgumentParser(
description="Classify incident severity and generate escalation paths.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""\
examples:
%(prog)s incident.json
%(prog)s incident.json --format json
%(prog)s incident.json --format markdown
cat incident.json | %(prog)s
cat incident.json | %(prog)s --format json
""",
)
parser.add_argument(
"data_file",
nargs="?",
default=None,
help="JSON file with incident data (reads stdin if omitted)",
)
parser.add_argument(
"--format",
choices=["text", "json", "markdown"],
default="text",
dest="output_format",
help="Output format (default: text)",
)
args = parser.parse_args()
# -- Read input --
try:
if args.data_file:
with open(args.data_file, "r", encoding="utf-8") as fh:
raw_data = json.load(fh)
else:
if sys.stdin.isatty():
parser.error("No input file provided and stdin is a terminal. Pipe JSON or pass a file.")
raw_data = json.load(sys.stdin)
except json.JSONDecodeError as exc:
print(f"Error: invalid JSON input -- {exc}", file=sys.stderr)
sys.exit(1)
except FileNotFoundError:
print(f"Error: file not found -- {args.data_file}", file=sys.stderr)
sys.exit(1)
except IOError as exc:
print(f"Error: could not read input -- {exc}", file=sys.stderr)
sys.exit(1)
# -- Parse and validate --
try:
incident, impact, signals, context = parse_incident_data(raw_data)
except ValueError as exc:
print(f"Error: {exc}", file=sys.stderr)
sys.exit(1)
# -- Classify --
severity_score = classify_severity(incident, impact, signals, context)
# -- Build outputs --
escalation = build_escalation_path(severity_score, signals, context)
action_plan = build_action_plan(severity_score, incident, impact, signals, context)
sla_impact = assess_sla_impact(severity_score, impact, signals)
# -- Format and print --
if args.output_format == "json":
output = format_json(incident, severity_score, escalation, action_plan, sla_impact)
elif args.output_format == "markdown":
output = format_markdown(incident, severity_score, escalation, action_plan, sla_impact)
else:
output = format_text(incident, severity_score, escalation, action_plan, sla_impact)
print(output)
# -- Exit code reflects severity --
if severity_score.severity_level == SeverityLevel.SEV1:
sys.exit(2)
elif severity_score.severity_level == SeverityLevel.SEV2:
sys.exit(1)
else:
sys.exit(0)
if __name__ == "__main__":
main()
FILE:scripts/timeline_reconstructor.py
#!/usr/bin/env python3
"""
Timeline Reconstructor
Reconstructs incident timelines from timestamped events (logs, alerts, Slack messages).
Identifies incident phases, calculates durations, and performs gap analysis.
This tool processes chronological event data and creates a coherent narrative
of how an incident progressed from detection through resolution.
Usage:
python timeline_reconstructor.py --input events.json --output timeline.md
python timeline_reconstructor.py --input events.json --detect-phases --gap-analysis
cat events.json | python timeline_reconstructor.py --format text
"""
import argparse
import json
import sys
import re
from datetime import datetime, timezone, timedelta
from typing import Dict, List, Optional, Any, Tuple
from collections import defaultdict, namedtuple
# Event data structure
Event = namedtuple('Event', ['timestamp', 'source', 'type', 'message', 'severity', 'actor', 'metadata'])
# Phase data structure
Phase = namedtuple('Phase', ['name', 'start_time', 'end_time', 'duration', 'events', 'description'])
class TimelineReconstructor:
"""
Reconstructs incident timelines from disparate event sources.
Identifies phases, calculates metrics, and performs gap analysis.
"""
def __init__(self):
"""Initialize the reconstructor with phase detection rules and templates."""
self.phase_patterns = self._load_phase_patterns()
self.event_types = self._load_event_types()
self.severity_mapping = self._load_severity_mapping()
self.gap_thresholds = self._load_gap_thresholds()
def _load_phase_patterns(self) -> Dict[str, Dict]:
"""Load patterns for identifying incident phases."""
return {
"detection": {
"keywords": [
"alert", "alarm", "triggered", "fired", "detected", "noticed",
"monitoring", "threshold exceeded", "anomaly", "spike",
"error rate", "latency increase", "timeout", "failure"
],
"event_types": ["alert", "monitoring", "notification"],
"priority": 1,
"description": "Initial detection of the incident through monitoring or observation"
},
"triage": {
"keywords": [
"investigating", "triaging", "assessing", "evaluating",
"checking", "looking into", "analyzing", "reviewing",
"diagnosis", "troubleshooting", "examining"
],
"event_types": ["investigation", "communication", "action"],
"priority": 2,
"description": "Assessment and initial investigation of the incident"
},
"escalation": {
"keywords": [
"escalating", "paging", "calling in", "requesting help",
"engaging", "involving", "notifying", "alerting team",
"incident commander", "war room", "all hands"
],
"event_types": ["escalation", "communication", "notification"],
"priority": 3,
"description": "Escalation to additional resources or higher severity response"
},
"mitigation": {
"keywords": [
"fixing", "patching", "deploying", "rolling back", "restarting",
"scaling", "rerouting", "bypassing", "workaround",
"implementing fix", "applying solution", "remediation"
],
"event_types": ["deployment", "action", "fix"],
"priority": 4,
"description": "Active mitigation efforts to resolve the incident"
},
"resolution": {
"keywords": [
"resolved", "fixed", "restored", "recovered", "back online",
"working", "normal", "stable", "healthy", "operational",
"incident closed", "service restored"
],
"event_types": ["resolution", "confirmation"],
"priority": 5,
"description": "Confirmation that the incident has been resolved"
},
"review": {
"keywords": [
"post-mortem", "retrospective", "review", "lessons learned",
"pir", "post-incident", "analysis", "follow-up",
"action items", "improvements"
],
"event_types": ["review", "documentation"],
"priority": 6,
"description": "Post-incident review and documentation activities"
}
}
def _load_event_types(self) -> Dict[str, Dict]:
"""Load event type classification rules."""
return {
"alert": {
"sources": ["monitoring", "nagios", "datadog", "newrelic", "prometheus"],
"indicators": ["alert", "alarm", "threshold", "metric"],
"severity_boost": 2
},
"log": {
"sources": ["application", "server", "container", "system"],
"indicators": ["error", "exception", "warn", "fail"],
"severity_boost": 1
},
"communication": {
"sources": ["slack", "teams", "email", "chat"],
"indicators": ["message", "notification", "update"],
"severity_boost": 0
},
"deployment": {
"sources": ["ci/cd", "jenkins", "github", "gitlab", "deploy"],
"indicators": ["deploy", "release", "build", "merge"],
"severity_boost": 3
},
"action": {
"sources": ["manual", "script", "automation", "operator"],
"indicators": ["executed", "ran", "performed", "applied"],
"severity_boost": 2
},
"escalation": {
"sources": ["pagerduty", "opsgenie", "oncall", "escalation"],
"indicators": ["paged", "escalated", "notified", "assigned"],
"severity_boost": 3
}
}
def _load_severity_mapping(self) -> Dict[str, int]:
"""Load severity level mappings."""
return {
"critical": 5, "crit": 5, "sev1": 5, "p1": 5,
"high": 4, "major": 4, "sev2": 4, "p2": 4,
"medium": 3, "moderate": 3, "sev3": 3, "p3": 3,
"low": 2, "minor": 2, "sev4": 2, "p4": 2,
"info": 1, "informational": 1, "debug": 1,
"unknown": 0
}
def _load_gap_thresholds(self) -> Dict[str, int]:
"""Load gap analysis thresholds in minutes."""
return {
"detection_to_triage": 15, # Should start investigating within 15 min
"triage_to_mitigation": 30, # Should start mitigation within 30 min
"mitigation_to_resolution": 120, # Should resolve within 2 hours
"communication_gap": 30, # Should communicate every 30 min
"action_gap": 60, # Should take actions every hour
"phase_transition": 45 # Should transition phases within 45 min
}
def reconstruct_timeline(self, events_data: List[Dict]) -> Dict[str, Any]:
"""
Main reconstruction method that processes events and builds timeline.
Args:
events_data: List of event dictionaries
Returns:
Dictionary with timeline analysis and metrics
"""
# Parse and normalize events
events = self._parse_events(events_data)
if not events:
return {"error": "No valid events found"}
# Sort events chronologically
events.sort(key=lambda e: e.timestamp)
# Detect phases
phases = self._detect_phases(events)
# Calculate metrics
metrics = self._calculate_metrics(events, phases)
# Perform gap analysis
gap_analysis = self._analyze_gaps(events, phases)
# Generate timeline narrative
narrative = self._generate_narrative(events, phases)
# Create summary statistics
summary = self._generate_summary(events, phases, metrics)
return {
"timeline": {
"total_events": len(events),
"time_range": {
"start": events[0].timestamp.isoformat(),
"end": events[-1].timestamp.isoformat(),
"duration_minutes": int((events[-1].timestamp - events[0].timestamp).total_seconds() / 60)
},
"phases": [self._phase_to_dict(phase) for phase in phases],
"events": [self._event_to_dict(event) for event in events]
},
"metrics": metrics,
"gap_analysis": gap_analysis,
"narrative": narrative,
"summary": summary,
"reconstruction_timestamp": datetime.now(timezone.utc).isoformat()
}
def _parse_events(self, events_data: List[Dict]) -> List[Event]:
"""Parse raw event data into normalized Event objects."""
events = []
for event_dict in events_data:
try:
# Parse timestamp
timestamp_str = event_dict.get("timestamp", event_dict.get("time", ""))
if not timestamp_str:
continue
timestamp = self._parse_timestamp(timestamp_str)
if not timestamp:
continue
# Extract other fields
source = event_dict.get("source", "unknown")
event_type = self._classify_event_type(event_dict)
message = event_dict.get("message", event_dict.get("description", ""))
severity = self._parse_severity(event_dict.get("severity", event_dict.get("level", "unknown")))
actor = event_dict.get("actor", event_dict.get("user", "system"))
# Extract metadata
metadata = {k: v for k, v in event_dict.items()
if k not in ["timestamp", "time", "source", "type", "message", "severity", "actor"]}
event = Event(
timestamp=timestamp,
source=source,
type=event_type,
message=message,
severity=severity,
actor=actor,
metadata=metadata
)
events.append(event)
except Exception as e:
# Skip invalid events but log them
continue
return events
def _parse_timestamp(self, timestamp_str: str) -> Optional[datetime]:
"""Parse various timestamp formats."""
# Common timestamp formats
formats = [
"%Y-%m-%dT%H:%M:%S.%fZ", # ISO with microseconds
"%Y-%m-%dT%H:%M:%SZ", # ISO without microseconds
"%Y-%m-%d %H:%M:%S", # Standard format
"%m/%d/%Y %H:%M:%S", # US format
"%d/%m/%Y %H:%M:%S", # EU format
"%Y-%m-%d %H:%M:%S.%f", # With microseconds
"%Y%m%d_%H%M%S", # Compact format
]
for fmt in formats:
try:
dt = datetime.strptime(timestamp_str, fmt)
# Ensure timezone awareness
if dt.tzinfo is None:
dt = dt.replace(tzinfo=timezone.utc)
return dt
except ValueError:
continue
# Try parsing as Unix timestamp
try:
timestamp_float = float(timestamp_str)
return datetime.fromtimestamp(timestamp_float, tz=timezone.utc)
except ValueError:
pass
return None
def _classify_event_type(self, event_dict: Dict) -> str:
"""Classify event type based on source and content."""
source = event_dict.get("source", "").lower()
message = event_dict.get("message", "").lower()
event_type = event_dict.get("type", "").lower()
# Check explicit type first
if event_type in self.event_types:
return event_type
# Classify based on source and content
for type_name, type_info in self.event_types.items():
# Check source patterns
if any(src in source for src in type_info["sources"]):
return type_name
# Check message indicators
if any(indicator in message for indicator in type_info["indicators"]):
return type_name
return "unknown"
def _parse_severity(self, severity_str: str) -> int:
"""Parse severity string to numeric value."""
severity_clean = str(severity_str).lower().strip()
return self.severity_mapping.get(severity_clean, 0)
def _detect_phases(self, events: List[Event]) -> List[Phase]:
"""Detect incident phases based on event patterns."""
phases = []
current_phase = None
phase_events = []
for event in events:
detected_phase = self._identify_phase(event)
if detected_phase != current_phase:
# End current phase if exists
if current_phase and phase_events:
phase_obj = Phase(
name=current_phase,
start_time=phase_events[0].timestamp,
end_time=phase_events[-1].timestamp,
duration=(phase_events[-1].timestamp - phase_events[0].timestamp).total_seconds() / 60,
events=phase_events.copy(),
description=self.phase_patterns[current_phase]["description"]
)
phases.append(phase_obj)
# Start new phase
current_phase = detected_phase
phase_events = [event]
else:
phase_events.append(event)
# Add final phase
if current_phase and phase_events:
phase_obj = Phase(
name=current_phase,
start_time=phase_events[0].timestamp,
end_time=phase_events[-1].timestamp,
duration=(phase_events[-1].timestamp - phase_events[0].timestamp).total_seconds() / 60,
events=phase_events,
description=self.phase_patterns[current_phase]["description"]
)
phases.append(phase_obj)
return self._merge_adjacent_phases(phases)
def _identify_phase(self, event: Event) -> str:
"""Identify which phase an event belongs to."""
message_lower = event.message.lower()
# Score each phase based on keywords and event type
phase_scores = {}
for phase_name, pattern_info in self.phase_patterns.items():
score = 0
# Keyword matching
for keyword in pattern_info["keywords"]:
if keyword in message_lower:
score += 2
# Event type matching
if event.type in pattern_info["event_types"]:
score += 3
# Severity boost for certain phases
if phase_name == "escalation" and event.severity >= 4:
score += 2
phase_scores[phase_name] = score
# Return highest scoring phase, default to triage
if phase_scores and max(phase_scores.values()) > 0:
return max(phase_scores, key=phase_scores.get)
return "triage" # Default phase
def _merge_adjacent_phases(self, phases: List[Phase]) -> List[Phase]:
"""Merge adjacent phases of the same type."""
if not phases:
return phases
merged = []
current_phase = phases[0]
for next_phase in phases[1:]:
if (next_phase.name == current_phase.name and
(next_phase.start_time - current_phase.end_time).total_seconds() < 300): # 5 min gap
# Merge phases
merged_events = current_phase.events + next_phase.events
current_phase = Phase(
name=current_phase.name,
start_time=current_phase.start_time,
end_time=next_phase.end_time,
duration=(next_phase.end_time - current_phase.start_time).total_seconds() / 60,
events=merged_events,
description=current_phase.description
)
else:
merged.append(current_phase)
current_phase = next_phase
merged.append(current_phase)
return merged
def _calculate_metrics(self, events: List[Event], phases: List[Phase]) -> Dict[str, Any]:
"""Calculate timeline metrics and KPIs."""
if not events or not phases:
return {}
start_time = events[0].timestamp
end_time = events[-1].timestamp
total_duration = (end_time - start_time).total_seconds() / 60
# Phase timing metrics
phase_durations = {phase.name: phase.duration for phase in phases}
# Detection metrics
detection_time = 0
if phases and phases[0].name == "detection":
detection_time = phases[0].duration
# Time to mitigation
mitigation_start = None
for phase in phases:
if phase.name == "mitigation":
mitigation_start = (phase.start_time - start_time).total_seconds() / 60
break
# Time to resolution
resolution_time = None
for phase in phases:
if phase.name == "resolution":
resolution_time = (phase.start_time - start_time).total_seconds() / 60
break
# Communication frequency
comm_events = [e for e in events if e.type == "communication"]
comm_frequency = len(comm_events) / (total_duration / 60) if total_duration > 0 else 0
# Action frequency
action_events = [e for e in events if e.type == "action"]
action_frequency = len(action_events) / (total_duration / 60) if total_duration > 0 else 0
# Event source distribution
source_counts = defaultdict(int)
for event in events:
source_counts[event.source] += 1
return {
"duration_metrics": {
"total_duration_minutes": round(total_duration, 1),
"detection_duration_minutes": round(detection_time, 1),
"time_to_mitigation_minutes": round(mitigation_start or 0, 1),
"time_to_resolution_minutes": round(resolution_time or 0, 1),
"phase_durations": {k: round(v, 1) for k, v in phase_durations.items()}
},
"activity_metrics": {
"total_events": len(events),
"events_per_hour": round((len(events) / (total_duration / 60)) if total_duration > 0 else 0, 1),
"communication_frequency": round(comm_frequency, 1),
"action_frequency": round(action_frequency, 1),
"unique_sources": len(source_counts),
"unique_actors": len(set(e.actor for e in events))
},
"phase_metrics": {
"total_phases": len(phases),
"phase_sequence": [p.name for p in phases],
"longest_phase": max(phases, key=lambda p: p.duration).name if phases else None,
"shortest_phase": min(phases, key=lambda p: p.duration).name if phases else None
},
"source_distribution": dict(source_counts)
}
def _analyze_gaps(self, events: List[Event], phases: List[Phase]) -> Dict[str, Any]:
"""Perform gap analysis to identify potential issues."""
gaps = []
warnings = []
# Check phase transition timing
for i in range(len(phases) - 1):
current_phase = phases[i]
next_phase = phases[i + 1]
transition_gap = (next_phase.start_time - current_phase.end_time).total_seconds() / 60
threshold_key = f"{current_phase.name}_to_{next_phase.name}"
threshold = self.gap_thresholds.get(threshold_key, self.gap_thresholds["phase_transition"])
if transition_gap > threshold:
gaps.append({
"type": "phase_transition",
"from_phase": current_phase.name,
"to_phase": next_phase.name,
"gap_minutes": round(transition_gap, 1),
"threshold_minutes": threshold,
"severity": "warning" if transition_gap < threshold * 2 else "critical"
})
# Check communication gaps
comm_events = [e for e in events if e.type == "communication"]
for i in range(len(comm_events) - 1):
gap_minutes = (comm_events[i+1].timestamp - comm_events[i].timestamp).total_seconds() / 60
if gap_minutes > self.gap_thresholds["communication_gap"]:
gaps.append({
"type": "communication_gap",
"gap_minutes": round(gap_minutes, 1),
"threshold_minutes": self.gap_thresholds["communication_gap"],
"severity": "warning" if gap_minutes < self.gap_thresholds["communication_gap"] * 2 else "critical"
})
# Check for missing phases
expected_phases = ["detection", "triage", "mitigation", "resolution"]
actual_phases = [p.name for p in phases]
missing_phases = [p for p in expected_phases if p not in actual_phases]
for missing_phase in missing_phases:
warnings.append({
"type": "missing_phase",
"phase": missing_phase,
"message": f"Expected phase '{missing_phase}' not detected in timeline"
})
# Check for unusually long phases
for phase in phases:
if phase.duration > 180: # 3 hours
warnings.append({
"type": "long_phase",
"phase": phase.name,
"duration_minutes": round(phase.duration, 1),
"message": f"Phase '{phase.name}' lasted {phase.duration:.0f} minutes, which is unusually long"
})
return {
"gaps": gaps,
"warnings": warnings,
"gap_summary": {
"total_gaps": len(gaps),
"critical_gaps": len([g for g in gaps if g.get("severity") == "critical"]),
"warning_gaps": len([g for g in gaps if g.get("severity") == "warning"]),
"missing_phases": len(missing_phases)
}
}
def _generate_narrative(self, events: List[Event], phases: List[Phase]) -> Dict[str, Any]:
"""Generate human-readable incident narrative."""
if not events or not phases:
return {"error": "Insufficient data for narrative generation"}
# Create phase-based narrative
phase_narratives = []
for phase in phases:
key_events = self._extract_key_events(phase.events)
narrative_text = self._create_phase_narrative(phase, key_events)
phase_narratives.append({
"phase": phase.name,
"start_time": phase.start_time.isoformat(),
"duration_minutes": round(phase.duration, 1),
"narrative": narrative_text,
"key_events": len(key_events),
"total_events": len(phase.events)
})
# Create overall summary
start_time = events[0].timestamp
end_time = events[-1].timestamp
total_duration = (end_time - start_time).total_seconds() / 60
summary = f"""Incident Timeline Summary:
The incident began at {start_time.strftime('%Y-%m-%d %H:%M:%S UTC')} and concluded at {end_time.strftime('%Y-%m-%d %H:%M:%S UTC')}, lasting approximately {total_duration:.0f} minutes.
The incident progressed through {len(phases)} distinct phases: {', '.join(p.name for p in phases)}.
Key milestones:"""
for phase in phases:
summary += f"\n- {phase.name.title()}: {phase.start_time.strftime('%H:%M')} ({phase.duration:.0f} min)"
return {
"summary": summary,
"phase_narratives": phase_narratives,
"timeline_type": self._classify_timeline_pattern(phases),
"complexity_score": self._calculate_complexity_score(events, phases)
}
def _extract_key_events(self, events: List[Event]) -> List[Event]:
"""Extract the most important events from a phase."""
# Sort by severity and timestamp
sorted_events = sorted(events, key=lambda e: (e.severity, e.timestamp), reverse=True)
# Take top events, but ensure chronological representation
key_events = []
# Always include first and last events
if events:
key_events.append(events[0])
if len(events) > 1:
key_events.append(events[-1])
# Add high-severity events
high_severity_events = [e for e in events if e.severity >= 4]
key_events.extend(high_severity_events[:3])
# Remove duplicates while preserving order
seen = set()
unique_events = []
for event in key_events:
event_key = (event.timestamp, event.message)
if event_key not in seen:
seen.add(event_key)
unique_events.append(event)
return sorted(unique_events, key=lambda e: e.timestamp)
def _create_phase_narrative(self, phase: Phase, key_events: List[Event]) -> str:
"""Create narrative text for a phase."""
phase_templates = {
"detection": "The incident was first detected when {first_event}. {additional_details}",
"triage": "Initial investigation began with {first_event}. The team {investigation_actions}",
"escalation": "The incident was escalated when {escalation_trigger}. {escalation_actions}",
"mitigation": "Mitigation efforts started with {first_action}. {mitigation_steps}",
"resolution": "The incident was resolved when {resolution_event}. {confirmation_steps}",
"review": "Post-incident review activities included {review_activities}"
}
template = phase_templates.get(phase.name, "During the {phase_name} phase, {activities}")
if not key_events:
return f"The {phase.name} phase lasted {phase.duration:.0f} minutes with {len(phase.events)} events."
first_event = key_events[0].message
# Customize based on phase
if phase.name == "detection":
return template.format(
first_event=first_event,
additional_details=f"This phase lasted {phase.duration:.0f} minutes with {len(phase.events)} total events."
)
elif phase.name == "triage":
actions = [e.message for e in key_events if "investigating" in e.message.lower() or "checking" in e.message.lower()]
investigation_text = "performed various diagnostic activities" if not actions else f"focused on {actions[0]}"
return template.format(
first_event=first_event,
investigation_actions=investigation_text
)
else:
return f"During the {phase.name} phase ({phase.duration:.0f} minutes), key activities included: {first_event}"
def _classify_timeline_pattern(self, phases: List[Phase]) -> str:
"""Classify the overall timeline pattern."""
phase_names = [p.name for p in phases]
if "escalation" in phase_names and phases[0].name == "detection":
return "standard_escalation"
elif len(phases) <= 3:
return "simple_resolution"
elif "review" in phase_names:
return "comprehensive_response"
else:
return "complex_incident"
def _calculate_complexity_score(self, events: List[Event], phases: List[Phase]) -> float:
"""Calculate incident complexity score (0-10)."""
score = 0.0
# Phase count contributes to complexity
score += min(len(phases) * 1.5, 6.0)
# Event count contributes to complexity
score += min(len(events) / 20, 2.0)
# Duration contributes to complexity
if events:
duration_hours = (events[-1].timestamp - events[0].timestamp).total_seconds() / 3600
score += min(duration_hours / 2, 2.0)
return min(score, 10.0)
def _generate_summary(self, events: List[Event], phases: List[Phase], metrics: Dict) -> Dict[str, Any]:
"""Generate comprehensive incident summary."""
if not events:
return {}
# Key statistics
start_time = events[0].timestamp
end_time = events[-1].timestamp
duration_minutes = metrics.get("duration_metrics", {}).get("total_duration_minutes", 0)
# Phase analysis
phase_analysis = {}
for phase in phases:
phase_analysis[phase.name] = {
"duration_minutes": round(phase.duration, 1),
"event_count": len(phase.events),
"start_time": phase.start_time.isoformat(),
"end_time": phase.end_time.isoformat()
}
# Actor involvement
actors = defaultdict(int)
for event in events:
actors[event.actor] += 1
return {
"incident_overview": {
"start_time": start_time.isoformat(),
"end_time": end_time.isoformat(),
"total_duration_minutes": round(duration_minutes, 1),
"total_events": len(events),
"phases_detected": len(phases)
},
"phase_analysis": phase_analysis,
"key_participants": dict(actors),
"event_sources": dict(defaultdict(int, {e.source: 1 for e in events})),
"complexity_indicators": {
"unique_sources": len(set(e.source for e in events)),
"unique_actors": len(set(e.actor for e in events)),
"high_severity_events": len([e for e in events if e.severity >= 4]),
"phase_transitions": len(phases) - 1 if phases else 0
}
}
def _event_to_dict(self, event: Event) -> Dict:
"""Convert Event namedtuple to dictionary."""
return {
"timestamp": event.timestamp.isoformat(),
"source": event.source,
"type": event.type,
"message": event.message,
"severity": event.severity,
"actor": event.actor,
"metadata": event.metadata
}
def _phase_to_dict(self, phase: Phase) -> Dict:
"""Convert Phase namedtuple to dictionary."""
return {
"name": phase.name,
"start_time": phase.start_time.isoformat(),
"end_time": phase.end_time.isoformat(),
"duration_minutes": round(phase.duration, 1),
"event_count": len(phase.events),
"description": phase.description
}
def format_json_output(result: Dict) -> str:
"""Format result as pretty JSON."""
return json.dumps(result, indent=2, ensure_ascii=False)
def format_text_output(result: Dict) -> str:
"""Format result as human-readable text."""
if "error" in result:
return f"Error: {result['error']}"
timeline = result["timeline"]
metrics = result["metrics"]
narrative = result["narrative"]
output = []
output.append("=" * 80)
output.append("INCIDENT TIMELINE RECONSTRUCTION")
output.append("=" * 80)
output.append("")
# Overview
time_range = timeline["time_range"]
output.append("OVERVIEW:")
output.append(f" Time Range: {time_range['start']} to {time_range['end']}")
output.append(f" Total Duration: {time_range['duration_minutes']} minutes")
output.append(f" Total Events: {timeline['total_events']}")
output.append(f" Phases Detected: {len(timeline['phases'])}")
output.append("")
# Phase summary
output.append("PHASES:")
for phase in timeline["phases"]:
output.append(f" {phase['name'].upper()}:")
output.append(f" Start: {phase['start_time']}")
output.append(f" Duration: {phase['duration_minutes']} minutes")
output.append(f" Events: {phase['event_count']}")
output.append(f" Description: {phase['description']}")
output.append("")
# Key metrics
if "duration_metrics" in metrics:
duration_metrics = metrics["duration_metrics"]
output.append("KEY METRICS:")
output.append(f" Time to Mitigation: {duration_metrics.get('time_to_mitigation_minutes', 'N/A')} minutes")
output.append(f" Time to Resolution: {duration_metrics.get('time_to_resolution_minutes', 'N/A')} minutes")
if "activity_metrics" in metrics:
activity = metrics["activity_metrics"]
output.append(f" Events per Hour: {activity.get('events_per_hour', 'N/A')}")
output.append(f" Unique Sources: {activity.get('unique_sources', 'N/A')}")
output.append("")
# Narrative
if "summary" in narrative:
output.append("INCIDENT NARRATIVE:")
output.append(narrative["summary"])
output.append("")
# Gap analysis
if "gap_analysis" in result and result["gap_analysis"]["gaps"]:
output.append("GAP ANALYSIS:")
for gap in result["gap_analysis"]["gaps"][:5]: # Show first 5 gaps
output.append(f" {gap['type'].replace('_', ' ').title()}: {gap['gap_minutes']} min gap (threshold: {gap['threshold_minutes']} min)")
output.append("")
output.append("=" * 80)
return "\n".join(output)
def format_markdown_output(result: Dict) -> str:
"""Format result as Markdown timeline."""
if "error" in result:
return f"# Error\n\n{result['error']}"
timeline = result["timeline"]
narrative = result.get("narrative", {})
output = []
output.append("# Incident Timeline")
output.append("")
# Overview
time_range = timeline["time_range"]
output.append("## Overview")
output.append("")
output.append(f"- **Duration:** {time_range['duration_minutes']} minutes")
output.append(f"- **Start Time:** {time_range['start']}")
output.append(f"- **End Time:** {time_range['end']}")
output.append(f"- **Total Events:** {timeline['total_events']}")
output.append("")
# Narrative summary
if "summary" in narrative:
output.append("## Summary")
output.append("")
output.append(narrative["summary"])
output.append("")
# Phase timeline
output.append("## Phase Timeline")
output.append("")
for phase in timeline["phases"]:
output.append(f"### {phase['name'].title()} Phase")
output.append("")
output.append(f"**Duration:** {phase['duration_minutes']} minutes ")
output.append(f"**Start:** {phase['start_time']} ")
output.append(f"**Events:** {phase['event_count']} ")
output.append("")
output.append(phase["description"])
output.append("")
# Detailed timeline
output.append("## Detailed Event Timeline")
output.append("")
for event in timeline["events"]:
timestamp = datetime.fromisoformat(event["timestamp"].replace('Z', '+00:00'))
output.append(f"**{timestamp.strftime('%H:%M:%S')}** [{event['source']}] {event['message']}")
output.append("")
return "\n".join(output)
def main():
"""Main function with argument parsing and execution."""
parser = argparse.ArgumentParser(
description="Reconstruct incident timeline from timestamped events",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
python timeline_reconstructor.py --input events.json --output timeline.md
python timeline_reconstructor.py --input events.json --detect-phases --gap-analysis
cat events.json | python timeline_reconstructor.py --format text
Input JSON format:
[
{
"timestamp": "2024-01-01T12:00:00Z",
"source": "monitoring",
"type": "alert",
"message": "High error rate detected",
"severity": "critical",
"actor": "system"
}
]
"""
)
parser.add_argument(
"--input", "-i",
help="Input file path (JSON format) or '-' for stdin"
)
parser.add_argument(
"--output", "-o",
help="Output file path (default: stdout)"
)
parser.add_argument(
"--format", "-f",
choices=["json", "text", "markdown"],
default="json",
help="Output format (default: json)"
)
parser.add_argument(
"--detect-phases",
action="store_true",
help="Enable advanced phase detection"
)
parser.add_argument(
"--gap-analysis",
action="store_true",
help="Perform gap analysis on timeline"
)
parser.add_argument(
"--min-events",
type=int,
default=1,
help="Minimum number of events required (default: 1)"
)
args = parser.parse_args()
reconstructor = TimelineReconstructor()
try:
# Read input
if args.input == "-" or (not args.input and not sys.stdin.isatty()):
# Read from stdin
input_text = sys.stdin.read().strip()
if not input_text:
parser.error("No input provided")
events_data = json.loads(input_text)
elif args.input:
# Read from file
with open(args.input, 'r') as f:
events_data = json.load(f)
else:
parser.error("No input specified. Use --input or pipe data to stdin.")
# Validate input
if not isinstance(events_data, list):
parser.error("Input must be a JSON array of events")
if len(events_data) < args.min_events:
parser.error(f"Minimum {args.min_events} events required")
# Reconstruct timeline
result = reconstructor.reconstruct_timeline(events_data)
# Format output
if args.format == "json":
output = format_json_output(result)
elif args.format == "markdown":
output = format_markdown_output(result)
else:
output = format_text_output(result)
# Write output
if args.output:
with open(args.output, 'w') as f:
f.write(output)
f.write('\n')
else:
print(output)
except FileNotFoundError as e:
print(f"Error: File not found - {e}", file=sys.stderr)
sys.exit(1)
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON - {e}", file=sys.stderr)
sys.exit(1)
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
sys.exit(1)
if __name__ == "__main__":
main()Phân loại, phân tầng, xác định đường leo thang và thu thập chứng cứ pháp y cho sự cố bảo mật theo SEV1-SEV4 và NIST SP 800-61.
---
name: "incident-response"
description: "Use when a security incident has been detected or declared and needs classification, triage, escalation path determination, and forensic evidence collection. Covers SEV1-SEV4 classification, false positive filtering, incident taxonomy, and NIST SP 800-61 lifecycle."
---
# Incident Response
Incident response skill for the full lifecycle from initial triage through forensic collection, severity declaration, and escalation routing. This is NOT threat hunting (see threat-detection) or post-incident compliance mapping (see governance/compliance-mapping) — this is about classifying, triaging, and managing declared security incidents.
---
## Table of Contents
- [Overview](#overview)
- [Incident Triage Tool](#incident-triage-tool)
- [Incident Classification](#incident-classification)
- [Severity Framework](#severity-framework)
- [False Positive Filtering](#false-positive-filtering)
- [Forensic Evidence Collection](#forensic-evidence-collection)
- [Escalation Paths](#escalation-paths)
- [Regulatory Notification Obligations](#regulatory-notification-obligations)
- [Workflows](#workflows)
- [Anti-Patterns](#anti-patterns)
- [Cross-References](#cross-references)
---
## Overview
### What This Skill Does
This skill provides the methodology and tooling for **incident triage and response** — classifying security events into typed incidents, scoring severity, filtering false positives, determining escalation paths, and initiating forensic evidence collection under chain-of-custody controls.
### Distinction from Other Security Skills
| Skill | Focus | Approach |
|-------|-------|----------|
| **incident-response** (this) | Active incidents | Reactive — classify, escalate, collect evidence |
| threat-detection | Pre-incident hunting | Proactive — find threats before alerts fire |
| cloud-security | Cloud posture assessment | Preventive — IAM, S3, network misconfiguration |
| red-team | Offensive simulation | Offensive — test detection and response capability |
### Prerequisites
A security event must be ingested before triage. Events can come from SIEM alerts, EDR detections, threat intel feeds, or user reports. The triage tool accepts JSON event payloads; see the input schema below.
---
## Incident Triage Tool
The `incident_triage.py` tool classifies events, checks false positives, scores severity, determines escalation, and performs forensic pre-analysis.
```bash
# Classify an event from JSON file
python3 scripts/incident_triage.py --input event.json --classify --json
# Classify with false positive filtering enabled
python3 scripts/incident_triage.py --input event.json --classify --false-positive-check --json
# Force a severity level for tabletop exercises
python3 scripts/incident_triage.py --input event.json --severity sev1 --json
# Read event from stdin
echo '{"event_type": "ransomware", "host": "prod-db-01", "raw_payload": {}}' | \
python3 scripts/incident_triage.py --classify --false-positive-check --json
```
### Input Event Schema
```json
{
"event_type": "ransomware",
"host": "prod-db-01",
"user": "svc_backup",
"source_ip": "10.1.2.3",
"timestamp": "2024-01-15T14:32:00Z",
"raw_payload": {}
}
```
### Exit Codes
| Code | Meaning | Required Response |
|------|---------|-------------------|
| 0 | SEV3/SEV4 or clean | Standard ticket-based handling |
| 1 | SEV2 — elevated | 1-hour bridge call, async coordination |
| 2 | SEV1 — critical | Immediate 15-minute war room, all-hands |
---
## Incident Classification
Security events are classified into 14 incident types. Classification drives default severity, MITRE technique mapping, and response SLA.
### Incident Taxonomy
| Incident Type | Default Severity | MITRE Technique | Response SLA |
|--------------|-----------------|-----------------|--------------|
| ransomware | SEV1 | T1486 | 15 minutes |
| data_exfiltration | SEV1 | T1048 | 15 minutes |
| apt_intrusion | SEV1 | T1566 | 15 minutes |
| supply_chain_compromise | SEV1 | T1195 | 15 minutes |
| domain_controller_breach | SEV1 | T1078.002 | 15 minutes |
| credential_compromise | SEV2 | T1110 | 1 hour |
| lateral_movement | SEV2 | T1021 | 1 hour |
| malware_infection | SEV2 | T1204 | 1 hour |
| insider_threat | SEV2 | T1078 | 1 hour |
| cloud_account_compromise | SEV2 | T1078.004 | 1 hour |
| unauthorized_access | SEV3 | T1190 | 4 hours |
| policy_violation | SEV3 | N/A | 4 hours |
| phishing_attempt | SEV4 | T1566.001 | 24 hours |
| security_alert | SEV4 | N/A | 24 hours |
### SEV Escalation Triggers
Any of the following automatically re-declare a higher severity:
| Trigger | New Severity |
|---------|-------------|
| Ransomware note found | SEV1 |
| Active exfiltration confirmed | SEV1 |
| CloudTrail or SIEM disabled | SEV1 |
| Domain controller access confirmed | SEV1 |
| Second system compromised | SEV1 |
| Exfiltration volume exceeds 1 GB | SEV2 minimum |
| C-suite account accessed | SEV2 minimum |
---
## Severity Framework
### SEV Level Matrix
| Level | Name | Criteria | Skills Invoked | Escalation Path |
|-------|------|----------|---------------|-----------------|
| SEV1 | Critical | Confirmed ransomware; active PII/PHI exfiltration (>10K records); domain controller breach; defense evasion (CloudTrail disabled); supply chain compromise | All skills (parallel) | SOC Lead → CISO → CEO → Board Chair |
| SEV2 | High | Confirmed unauthorized access to sensitive systems; credential compromise with elevated privileges; lateral movement confirmed; ransomware indicators without confirmed execution | triage + containment + forensics | SOC Lead → CISO |
| SEV3 | Medium | Suspected unauthorized access (unconfirmed); malware detected and contained; single account compromise (no priv escalation) | triage + containment | SOC Lead → Security Manager |
| SEV4 | Low | Security alert with no confirmed impact; informational indicator; policy violation with no data risk | triage only | L3 Analyst queue |
---
## False Positive Filtering
The triage tool applies five filters before escalating to prevent false positive inflation.
### False Positive Filter Types
| Filter | Description | Example Pattern |
|--------|-------------|----------------|
| CI/CD agent activity | Known build/deploy agents flagged as anomalies | jenkins, github-actions, circleci, gitlab-runner |
| Test environment tagging | Assets tagged as non-production | test-, staging-, dev-, sandbox- |
| Scheduled job patterns | Expected batch processes triggering alerts | cron, scheduled_task, batch_job, backup_ |
| Whitelisted identities | Explicitly approved service accounts | svc_monitoring, svc_backup, datadog-agent |
| Scanner activity | Known security scanners and vulnerability tools | nessus, qualys, rapid7, aws_inspector |
A confirmed false positive suppresses escalation and logs the suppression reason for audit purposes. Recurring false positives from the same source should be tuned out at the detection layer, not filtered repeatedly at triage.
---
## Forensic Evidence Collection
Evidence collection follows the DFRWS six-phase framework and the principle of volatile-first acquisition.
### DFRWS Six Phases
| Phase | Activity | Priority |
|-------|----------|----------|
| Identification | Identify what evidence exists and where | Immediate |
| Preservation | Prevent modification — write-block, snapshot, legal hold | Immediate |
| Collection | Acquire evidence in order of volatility | Immediate |
| Examination | Technical analysis of collected evidence | Within 2 hours |
| Analysis | Interpret findings in investigative context | Within 4 hours |
| Presentation | Produce findings report with chain of custody | Before incident closure |
### Volatile Evidence — Collect First
1. Live memory (RAM dump) — lost on reboot
2. Running processes and open network connections (`netstat`, `ps`)
3. Logged-in users and active sessions
4. System uptime and current time (for timeline anchoring)
5. Environment variables and loaded kernel modules
### Chain of Custody Requirements
Every evidence item must be recorded with:
- SHA-256 hash at acquisition time
- Acquisition timestamp in UTC with timezone offset
- Tool provenance (FTK Imager, Volatility, dd, AWS CloudTrail export)
- Investigator identity
- Transfer log (who had custody and when)
---
## Escalation Paths
### By Severity
| Severity | Immediate Contact | Bridge Call | External Notification |
|----------|------------------|-------------|----------------------|
| SEV1 | SOC Lead + CISO (15 min) | Immediate war room | Legal + PR standby; regulatory notification per deadline table |
| SEV2 | SOC Lead (30 min async) | 1-hour bridge | Legal notification if PII involved |
| SEV3 | Security Manager (4 hours) | Async only | None unless scope expands |
| SEV4 | L3 Analyst queue (24 hours) | None | None |
### By Incident Type
| Incident Type | Primary Escalation | Secondary |
|--------------|-------------------|-----------|
| Ransomware / APT | CISO + CEO | Board if data at risk |
| PII/PHI breach | Legal + CISO | Regulatory body (per deadline table) |
| Cloud account compromise | Cloud security team | CISO |
| Insider threat | HR + Legal + CISO | Law enforcement if criminal |
| Supply chain | CISO + Vendor management | Board |
---
## Regulatory Notification Obligations
The notification clock starts at incident declaration, not at investigation completion.
| Framework | Incident Type | Deadline | Penalty |
|-----------|--------------|----------|---------|
| GDPR (EU 2016/679) | Personal data breach | 72 hours after discovery | Up to 4% global revenue |
| PCI-DSS v4.0 | Cardholder data breach | 24 hours to acquirer | Card brand fines |
| HIPAA (45 CFR 164) | PHI breach (>500 individuals) | 60 days after discovery | Up to $1.9M per violation category |
| NY DFS 23 NYCRR 500 | Cybersecurity event | 72 hours to DFS | Regulatory sanctions |
| SEC Rule (17 CFR 229.106) | Material cybersecurity incident | 4 business days after materiality determination | SEC enforcement |
| CCPA / CPRA | Breach of sensitive PI | Without unreasonable delay | AG enforcement; private right of action |
| NIS2 (EU 2022/2555) | Significant incident (essential services) | 24-hour early warning; 72-hour notification | National authority sanctions |
**Operational rule:** If scope is unclear at declaration, assume the most restrictive applicable deadline and confirm scope within the first response window.
Full deadline reference: `references/regulatory-deadlines.md`
---
## Workflows
### Workflow 1: Quick Triage (15 Minutes)
For single alert requiring classification before escalation decision:
```bash
# 1. Classify the event with false positive filtering
python3 scripts/incident_triage.py --input alert.json \
--classify --false-positive-check --json
# 2. Review severity, escalation_path, and false_positive_flag in output
# 3. If severity = sev1 or sev2, page SOC Lead immediately
# 4. If false_positive_flag = true, document and close
```
**Decision**: Exit code 2 = SEV1 war room now. Exit code 1 = SEV2 bridge call within 30 minutes.
### Workflow 2: Full Incident Response (SEV1)
```
T+0 Detection arrives (SIEM alert, EDR, user report)
T+5 Classify with incident_triage.py --classify --false-positive-check
T+10 If SEV1: page CISO, open war room, start regulatory clock
T+15 Initiate forensic collection (volatile evidence first)
T+15 Containment assessment (parallel with forensics)
T+30 Human approval gate for any containment action
T+45 Execute approved containment
T+60 Assess containment effectiveness, brief Legal if PII/PHI scope
T+4h Final forensic evidence package, dwell time estimate
T+8h Eradication and recovery plan
T+72h Regulatory notification submission (if GDPR/NIS2 triggered)
```
```bash
# Full classification with forensic context
python3 scripts/incident_triage.py --input incident.json \
--classify --false-positive-check --severity sev1 --json > incident_triage_output.json
# Forensic pre-analysis
python3 scripts/incident_triage.py --input incident.json --json | \
jq '.forensic_findings, .chain_of_custody_steps'
```
### Workflow 3: Tabletop Exercise Simulation
Simulate incidents at specific severity levels without real events:
```bash
# Simulate SEV1 ransomware incident
echo '{"event_type": "ransomware", "host": "prod-db-01", "user": "svc_backup"}' | \
python3 scripts/incident_triage.py --classify --severity sev1 --json
# Simulate SEV2 credential compromise
echo '{"event_type": "credential_compromise", "user": "admin_user", "source_ip": "203.0.113.5"}' | \
python3 scripts/incident_triage.py --classify --false-positive-check --json
# Verify escalation paths for all 14 incident types
for type in ransomware data_exfiltration credential_compromise lateral_movement; do
echo "{\"event_type\": \"$type\"}" | python3 scripts/incident_triage.py --classify --json
done
```
---
## Anti-Patterns
1. **Starting the notification clock at investigation completion** — Regulatory clocks (GDPR 72 hours, PCI 24 hours) start at discovery, not investigation completion. Declaring late exposes the organization to maximum penalties even if the incident itself was minor.
2. **Containing before collecting volatile evidence** — Rebooting or isolating a system destroys RAM, running processes, and active connections. Forensic collection of volatile evidence must happen in parallel with containment, never after.
3. **Skipping false positive verification before escalation** — Escalating every alert to SEV1 degrades SOC credibility and causes alert fatigue. Always run false positive filters before paging the CISO.
4. **Undocumented incident command decisions** — Every decision made during a SEV1, including decisions made under uncertainty, must be logged in the evidence chain with timestamp and rationale. Undocumented decisions cannot be defended in regulatory investigations.
5. **Treating incident closure as investigation completion** — Incidents are closed when eradication and recovery are complete, not when the investigation is done. The forensic report and regulatory submissions may continue after operational closure.
6. **Single-source classification** — Classifying an incident from a single data source (one SIEM alert) without corroborating evidence frequently leads to misclassification. Collect at least two independent signals before declaring SEV1.
7. **Bypassing human approval gates for containment** — Automated containment actions (network isolation, credential revocation) taken without human approval can cause production outages, destroy evidence, and create liability. Human approval is non-negotiable for all mutating containment actions.
---
## Cross-References
| Skill | Relationship |
|-------|-------------|
| [threat-detection](../threat-detection/SKILL.md) | Confirmed hunting findings escalate to incident-response for triage and classification |
| [cloud-security](../cloud-security/SKILL.md) | Cloud posture findings (IAM compromise, S3 exposure) may trigger incident classification |
| [red-team](../red-team/SKILL.md) | Red team findings validate detection coverage; confirmed gaps become hunting hypotheses |
| [security-pen-testing](../security-pen-testing/SKILL.md) | Pen test vulnerabilities exploited in the wild escalate to incident-response for active incident handling |
FILE:references/regulatory-deadlines.md
# Regulatory Notification Deadlines
Reference table for incident notification deadlines under major regulatory frameworks. The notification clock starts at the moment an incident is declared, not at investigation completion.
**Operational rule:** If the scope of a breach is unclear at declaration time, assume the most restrictive applicable deadline and confirm scope within the first response window. Document the assumption and its resolution in the incident record.
---
## Deadline Summary Table
| Framework | Jurisdiction | Incident Type | Notification Deadline | Recipient | Penalty for Non-Compliance |
|-----------|-------------|--------------|----------------------|-----------|---------------------------|
| GDPR (EU 2016/679) | EU/EEA | Personal data breach | 72 hours after discovery | Supervisory Authority (DPA) | Up to 4% of global annual turnover or €20M |
| GDPR (EU 2016/679) | EU/EEA | Personal data breach affecting individual rights/freedoms | Without undue delay | Affected data subjects | Up to 4% of global annual turnover |
| PCI-DSS v4.0 | Global (card brands) | Cardholder data breach | 24 hours after confirmation | Acquiring bank and card brands | Fines per card brand schedule; potential card processing suspension |
| HIPAA (45 CFR §164.408) | United States | PHI breach (>500 individuals) | 60 calendar days after discovery | HHS Office for Civil Rights | $100–$50,000 per violation; up to $1.9M per violation category per year |
| HIPAA (45 CFR §164.406) | United States | PHI breach (>500 individuals in a state) | 60 days after discovery | Prominent media outlets in affected state | Same as above |
| HIPAA Small Breach | United States | PHI breach (<500 individuals) | Within 60 days of end of calendar year in which breach occurred | HHS (annual report) | Same as above |
| NY DFS 23 NYCRR 500.17 | New York State | Cybersecurity event affecting NY-regulated entity | 72 hours | NY DFS Superintendent | Regulatory sanctions, fines, license revocation |
| SEC Cybersecurity Rule (17 CFR §229.106) | United States (public companies) | Material cybersecurity incident | 4 business days after materiality determination | SEC Form 8-K filing (public disclosure) | SEC enforcement action; restatement risk |
| CCPA / CPRA | California, United States | Breach of sensitive personal information | Without unreasonable delay | CA Attorney General (if >500 CA residents affected) | Civil penalties up to $7,500 per intentional violation |
| NIS2 (EU 2022/2555) | EU/EEA (essential/important entities) | Significant incident | 24-hour early warning; 72-hour full notification | National CSIRT or competent authority | Up to €10M or 2% of global turnover |
| DORA (EU 2022/2554) | EU/EEA (financial sector) | Major ICT-related incident | Initial notification: 4 hours; intermediate: 72 hours; final: 1 month | Financial supervisory authority | National authority sanctions |
| SOX (for material incidents) | United States (public companies) | Financial system compromise creating material weakness | Immediate disclosure required | SEC, audit committee, auditors | Enforcement action; officer certification liability |
| Australia Privacy Act | Australia | Eligible data breach (serious harm likely) | 30 days after awareness | OAIC (Office of the Australian Information Commissioner) | Up to AUD 50M per serious contravention |
| PIPL (China) | China | Personal information breach | Immediately; notify individuals without delay | National Internet Information Office (CAC) | Up to ¥50M or 5% of prior year revenue |
---
## GDPR — Detailed Requirements
### Article 33 — Notification to Supervisory Authority
**When:** Any personal data breach where there is a risk to the rights and freedoms of individuals.
**Exception:** No notification required if the breach is unlikely to result in risk (e.g., the data was encrypted with a key that was not compromised, and the key cannot be recovered).
**What to include:**
1. Nature of the breach, including categories and approximate number of data subjects and records
2. Name and contact details of the Data Protection Officer
3. Likely consequences of the breach
4. Measures taken or proposed to address the breach, including mitigation
**Staggered notification:** If full information is not available within 72 hours, submit what is known and provide additional information in phases. Document why the information is being provided in phases.
### Article 34 — Notification to Data Subjects
**When:** When a breach is likely to result in high risk to the rights and freedoms of individuals.
**How:** In clear, plain language. Direct communication to the affected individuals.
**Exception:** Notification to individuals not required if:
- The personal data was protected by appropriate technical measures (e.g., encryption)
- The controller has taken subsequent measures that ensure high risk no longer materializes
- It would involve disproportionate effort (use public communication instead)
---
## PCI-DSS v4.0 — Detailed Requirements
### Requirement 12.10.5
Report compromises of cardholder data to the applicable payment brands and acquiring bank immediately upon detection of a suspected compromise. Do not wait for internal investigation to complete.
**Immediate actions required upon suspicion:**
1. Contact acquiring bank within 24 hours of suspicion (even if not yet confirmed)
2. Preserve all logs and evidence — do not modify or delete
3. Implement containment without destroying forensic evidence
4. Engage a PCI Forensic Investigator (PFI) from the approved list
**Card brand notification channels:**
- Visa: Visa Fraud Control
- Mastercard: Mastercard Fraud Control
- American Express: AmEx Security
- Discover: Discover Security
---
## HIPAA — Detailed Requirements
### 45 CFR §164.408 — Breach Notification to HHS
**Notification form:** HHS breach notification portal (https://www.hhs.gov/hipaa/for-professionals/breach-notification/)
**Content required:**
- Name of covered entity or business associate
- Nature of PHI involved (type of PHI, not specific records)
- Unauthorized persons who accessed or used the PHI
- Whether PHI was actually acquired or viewed
- Extent to which risk has been mitigated
### Breach Risk Assessment (45 CFR §164.402)
HIPAA provides a risk assessment safe harbor. A breach is presumed unless the covered entity can demonstrate (low probability PHI was compromised) based on:
1. Nature and extent of PHI involved
2. Who accessed the information
3. Whether PHI was actually acquired or viewed
4. Extent to which risk has been mitigated
Document this risk assessment in writing and retain for 6 years.
---
## Notification Clock Management
### Starting the Clock
Document the exact timestamp when the incident was declared in the incident record. This is the official start of all regulatory clocks.
### Parallel Tracking
Incidents often cross multiple frameworks simultaneously. Track all applicable clocks in parallel:
```
Incident declared: 2024-01-15T14:30:00Z
GDPR notification due: 2024-01-18T14:30:00Z (72 hours)
PCI notification due: 2024-01-16T14:30:00Z (24 hours)
HIPAA HHS notification: 2024-03-15T14:30:00Z (60 days)
NY DFS notification: 2024-01-18T14:30:00Z (72 hours)
```
### Notification Drafting
Prepare draft notifications in parallel with investigation. Do not wait until investigation is complete to begin drafting. All external regulatory communications must be reviewed by Legal and approved by CISO before transmission.
FILE:scripts/incident_triage.py
#!/usr/bin/env python3
"""
incident_triage.py — Incident Classification, Triage, and Escalation
Classifies security events into 14 incident types, applies false-positive
filters, scores severity (SEV1-SEV4), determines escalation path, and
performs forensic pre-analysis for confirmed incidents.
Usage:
echo '{"event_type": "ransomware", "raw_payload": {...}}' | python3 incident_triage.py
python3 incident_triage.py --input event.json --json
python3 incident_triage.py --classify --false-positive-check --input event.json --json
Exit codes:
0 SEV3/SEV4 or clean — standard handling
1 SEV2 — elevated response required
2 SEV1 — critical incident declared
"""
import argparse
import json
import sys
from datetime import datetime, timezone
from typing import Any, Dict, List, Optional, Tuple
# ---------------------------------------------------------------------------
# Constants — Forensic Pre-Analysis Base (reused from pre_analysis.py logic)
# ---------------------------------------------------------------------------
DWELL_CRITICAL = 720 # hours (30 days)
DWELL_HIGH = 168 # hours (7 days)
DWELL_MEDIUM = 24 # hours (1 day)
EVIDENCE_SOURCES = [
"siem_logs",
"edr_telemetry",
"network_pcap",
"dns_logs",
"proxy_logs",
"cloud_trail",
"authentication_logs",
"endpoint_filesystem",
"memory_dump",
"email_headers",
]
CHAIN_OF_CUSTODY_STEPS = [
"Identify and preserve volatile evidence (RAM, network connections)",
"Hash all collected artifacts (SHA-256) before analysis",
"Document collection timestamp and analyst identity",
"Transfer artifacts to isolated forensic workstation",
"Maintain write-blockers for disk images",
"Log every access to evidence with timestamps",
"Store originals in secure, access-controlled evidence vault",
"Maintain dual-custody chain for legal proceedings",
]
# ---------------------------------------------------------------------------
# Constants — Incident Taxonomy and Escalation
# ---------------------------------------------------------------------------
INCIDENT_TAXONOMY: Dict[str, Dict[str, Any]] = {
"ransomware": {
"default_severity": "sev1",
"mitre": "T1486",
"response_sla_minutes": 15,
},
"data_exfiltration": {
"default_severity": "sev1",
"mitre": "T1048",
"response_sla_minutes": 15,
},
"apt_intrusion": {
"default_severity": "sev1",
"mitre": "T1190",
"response_sla_minutes": 15,
},
"supply_chain_compromise": {
"default_severity": "sev1",
"mitre": "T1195",
"response_sla_minutes": 15,
},
"credential_compromise": {
"default_severity": "sev2",
"mitre": "T1078",
"response_sla_minutes": 60,
},
"lateral_movement": {
"default_severity": "sev2",
"mitre": "T1021",
"response_sla_minutes": 60,
},
"privilege_escalation": {
"default_severity": "sev2",
"mitre": "T1068",
"response_sla_minutes": 60,
},
"malware_detected": {
"default_severity": "sev2",
"mitre": "T1204",
"response_sla_minutes": 60,
},
"phishing": {
"default_severity": "sev3",
"mitre": "T1566",
"response_sla_minutes": 240,
},
"unauthorized_access": {
"default_severity": "sev3",
"mitre": "T1078",
"response_sla_minutes": 240,
},
"policy_violation": {
"default_severity": "sev4",
"mitre": "T1530",
"response_sla_minutes": 1440,
},
"vulnerability_discovered": {
"default_severity": "sev4",
"mitre": "T1190",
"response_sla_minutes": 1440,
},
"dos_attack": {
"default_severity": "sev3",
"mitre": "T1498",
"response_sla_minutes": 240,
},
"insider_threat": {
"default_severity": "sev2",
"mitre": "T1078.002",
"response_sla_minutes": 60,
},
}
FALSE_POSITIVE_INDICATORS = [
{
"name": "ci_cd_automation",
"description": "CI/CD pipeline service account activity",
"patterns": [
"jenkins", "github-actions", "gitlab-ci", "terraform",
"ansible", "circleci", "codepipeline",
],
},
{
"name": "test_environment",
"description": "Activity in test/dev/staging environment",
"patterns": [
"test", "dev", "staging", "sandbox", "qa", "nonprod", "non-prod",
],
},
{
"name": "scheduled_scanner",
"description": "Known security scanner or automated tool",
"patterns": [
"nessus", "qualys", "rapid7", "tenable", "crowdstrike",
"defender", "sentinel",
],
},
{
"name": "scheduled_batch_job",
"description": "Recurring batch process with expected behavior",
"patterns": [
"backup", "sync", "batch", "cron", "scheduled", "nightly", "weekly",
],
},
{
"name": "whitelisted_identity",
"description": "Identity in approved exception list",
"patterns": [
"svc-", "sa-", "system@", "automation@", "monitor@", "health-check",
],
},
]
ESCALATION_ROUTING: Dict[str, Dict[str, Any]] = {
"sev1": {
"escalate_to": "CISO + CEO + Board Chair (if data at risk)",
"bridge_call": True,
"war_room": True,
},
"sev2": {
"escalate_to": "SOC Lead + CISO",
"bridge_call": True,
"war_room": False,
},
"sev3": {
"escalate_to": "SOC Lead + Security Manager",
"bridge_call": False,
"war_room": False,
},
"sev4": {
"escalate_to": "L3 Analyst queue",
"bridge_call": False,
"war_room": False,
},
}
SEV_ESCALATION_TRIGGERS = [
{"indicator": "ransomware_note_found", "escalate_to": "sev1"},
{"indicator": "active_exfiltration_confirmed", "escalate_to": "sev1"},
{"indicator": "siem_disabled", "escalate_to": "sev1"},
{"indicator": "domain_controller_access", "escalate_to": "sev1"},
{"indicator": "second_system_compromised", "escalate_to": "sev1"},
]
# ---------------------------------------------------------------------------
# Forensic Pre-Analysis Functions (base pre_analysis.py logic)
# ---------------------------------------------------------------------------
def parse_forensic_fields(fact: dict) -> dict:
"""
Parse and normalise forensic-relevant fields from the raw event.
Returns a dict with keys: source_ip, destination_ip, user_account,
hostname, process_name, dwell_hours, iocs, raw_payload.
"""
raw = fact.get("raw_payload", {}) if isinstance(fact.get("raw_payload"), dict) else {}
def _pick(*keys: str, default: Any = None) -> Any:
"""Return first non-None value found across fact and raw_payload."""
for k in keys:
v = fact.get(k) or raw.get(k)
if v is not None:
return v
return default
source_ip = _pick("source_ip", "src_ip", "sourceIp", default="unknown")
destination_ip = _pick("destination_ip", "dst_ip", "dest_ip", "destinationIp", default="unknown")
user_account = _pick("user", "user_account", "username", "actor", "identity", default="unknown")
hostname = _pick("hostname", "host", "device", "computer_name", default="unknown")
process_name = _pick("process", "process_name", "executable", "image", default="unknown")
# Dwell time: accept hours directly or compute from timestamps
dwell_hours: float = 0.0
raw_dwell = _pick("dwell_hours", "dwell_time_hours", "dwell")
if raw_dwell is not None:
try:
dwell_hours = float(raw_dwell)
except (TypeError, ValueError):
dwell_hours = 0.0
else:
first_seen = _pick("first_seen", "first_observed", "initial_access_time")
last_seen = _pick("last_seen", "last_observed", "detection_time")
if first_seen and last_seen:
try:
fmt = "%Y-%m-%dT%H:%M:%SZ"
dt_first = datetime.strptime(str(first_seen), fmt)
dt_last = datetime.strptime(str(last_seen), fmt)
dwell_hours = max(0.0, (dt_last - dt_first).total_seconds() / 3600.0)
except (ValueError, TypeError):
dwell_hours = 0.0
iocs: List[str] = []
raw_iocs = _pick("iocs", "indicators", "indicators_of_compromise")
if isinstance(raw_iocs, list):
iocs = [str(i) for i in raw_iocs]
elif isinstance(raw_iocs, str):
iocs = [raw_iocs]
return {
"source_ip": source_ip,
"destination_ip": destination_ip,
"user_account": user_account,
"hostname": hostname,
"process_name": process_name,
"dwell_hours": dwell_hours,
"iocs": iocs,
"raw_payload": raw,
}
def assess_dwell_severity(dwell_hours: float) -> str:
"""
Map dwell time (hours) to a severity label.
Returns 'critical', 'high', 'medium', or 'low'.
"""
if dwell_hours >= DWELL_CRITICAL:
return "critical"
if dwell_hours >= DWELL_HIGH:
return "high"
if dwell_hours >= DWELL_MEDIUM:
return "medium"
return "low"
def build_ioc_summary(fields: dict) -> dict:
"""
Build a structured IOC summary from parsed forensic fields.
Returns a dict suitable for embedding in the triage output.
"""
iocs = fields.get("iocs", [])
dwell_hours = fields.get("dwell_hours", 0.0)
dwell_severity = assess_dwell_severity(dwell_hours)
# Classify IOCs by rough heuristic
ip_iocs = [i for i in iocs if _looks_like_ip(i)]
hash_iocs = [i for i in iocs if _looks_like_hash(i)]
domain_iocs = [i for i in iocs if not _looks_like_ip(i) and not _looks_like_hash(i)]
return {
"total_ioc_count": len(iocs),
"ip_indicators": ip_iocs,
"hash_indicators": hash_iocs,
"domain_url_indicators": domain_iocs,
"dwell_hours": round(dwell_hours, 2),
"dwell_severity": dwell_severity,
"evidence_sources_applicable": [
src for src in EVIDENCE_SOURCES
if _source_applicable(src, fields)
],
"chain_of_custody_steps": CHAIN_OF_CUSTODY_STEPS,
}
def _looks_like_ip(value: str) -> bool:
"""Heuristic: does the string look like an IPv4 address?"""
import re
return bool(re.match(r"^\d{1,3}(\.\d{1,3}){3}$", value.strip()))
def _looks_like_hash(value: str) -> bool:
"""Heuristic: does the string look like a hex hash (MD5/SHA1/SHA256)?"""
import re
return bool(re.match(r"^[0-9a-fA-F]{32,64}$", value.strip()))
def _source_applicable(source: str, fields: dict) -> bool:
"""Decide if an evidence source is relevant given parsed fields."""
mapping = {
"network_pcap": fields.get("source_ip") not in (None, "unknown"),
"edr_telemetry": fields.get("hostname") not in (None, "unknown"),
"authentication_logs": fields.get("user_account") not in (None, "unknown"),
"dns_logs": fields.get("destination_ip") not in (None, "unknown"),
"endpoint_filesystem": fields.get("process_name") not in (None, "unknown"),
"memory_dump": fields.get("process_name") not in (None, "unknown"),
}
return mapping.get(source, True)
# ---------------------------------------------------------------------------
# New Classification and Escalation Functions
# ---------------------------------------------------------------------------
def classify_incident(fact: dict) -> Tuple[str, float]:
"""
Classify incident type from event fields.
Performs keyword matching against INCIDENT_TAXONOMY keys and the
flattened string representation of raw_payload content.
Returns:
(incident_type, confidence) where confidence is 0.0–1.0.
Returns ("unknown", 0.0) when no match is found.
"""
# Build a single searchable string from the fact
searchable = _flatten_to_string(fact).lower()
scores: Dict[str, int] = {}
for incident_type in INCIDENT_TAXONOMY:
# The incident type slug itself is a keyword
slug_words = incident_type.replace("_", " ").split()
score = 0
for word in slug_words:
if word in searchable:
score += 2 # direct slug match carries more weight
# Additional keyword synonyms per type
synonyms = _get_synonyms(incident_type)
for syn in synonyms:
if syn in searchable:
score += 1
if score > 0:
scores[incident_type] = score
if not scores:
# Last resort: check explicit event_type field
event_type = str(fact.get("event_type", "")).lower().replace(" ", "_").replace("-", "_")
if event_type in INCIDENT_TAXONOMY:
return event_type, 0.6
return "unknown", 0.0
best_type = max(scores, key=lambda k: scores[k])
max_score = scores[best_type]
# Normalise confidence: cap at 1.0, scale by how much the best
# outscores alternatives
total_score = sum(scores.values()) or 1
raw_confidence = max_score / total_score
# Boost if event_type field matches
event_type = str(fact.get("event_type", "")).lower().replace(" ", "_").replace("-", "_")
if event_type == best_type:
raw_confidence = min(1.0, raw_confidence + 0.25)
confidence = round(min(1.0, raw_confidence + 0.1 * min(max_score, 5)), 2)
return best_type, confidence
def _flatten_to_string(obj: Any, depth: int = 0) -> str:
"""Recursively flatten any JSON-like object into a single string."""
if depth > 6:
return ""
if isinstance(obj, dict):
parts = []
for k, v in obj.items():
parts.append(str(k))
parts.append(_flatten_to_string(v, depth + 1))
return " ".join(parts)
if isinstance(obj, list):
return " ".join(_flatten_to_string(i, depth + 1) for i in obj)
return str(obj)
def _get_synonyms(incident_type: str) -> List[str]:
"""Return additional keyword synonyms for an incident type."""
synonyms_map: Dict[str, List[str]] = {
"ransomware": ["encrypt", "ransom", "locked", "decrypt", "wiper", "crypto"],
"data_exfiltration": ["exfil", "upload", "transfer", "leak", "dump", "steal", "exfiltrate"],
"apt_intrusion": ["apt", "nation-state", "targeted", "backdoor", "persistence", "c2", "c&c"],
"supply_chain_compromise": ["supply chain", "dependency", "package", "solarwinds", "xz", "npm"],
"credential_compromise": ["credential", "password", "brute force", "spray", "stuffing", "stolen"],
"lateral_movement": ["lateral", "pivot", "pass-the-hash", "wmi", "psexec", "rdp movement"],
"priv_escalation": ["privesc", "su_exec", "priv_change", "elevated_session", "priv_grant", "priv_abuse"],
"malware_detected": ["malware", "trojan", "virus", "worm", "keylogger", "spyware", "rat"],
"phishing": ["phish", "spear", "bec", "email", "lure", "credential harvest"],
"unauthorized_access": ["unauthorized", "unauthenticated", "brute", "login failed", "access denied"],
"policy_violation": ["policy", "dlp", "data loss", "violation", "compliance"],
"vulnerability_discovered": ["vulnerability", "cve", "exploit", "patch", "zero-day", "rce"],
"dos_attack": ["dos", "ddos", "flood", "amplification", "bandwidth", "exhaustion"],
"insider_threat": ["insider", "employee", "contractor", "abuse", "privilege misuse"],
}
return synonyms_map.get(incident_type, [])
def check_false_positives(fact: dict) -> List[str]:
"""
Check fact fields against FALSE_POSITIVE_INDICATORS pattern lists.
Returns a list of triggered false positive indicator names.
"""
searchable = _flatten_to_string(fact).lower()
triggered: List[str] = []
for indicator in FALSE_POSITIVE_INDICATORS:
for pattern in indicator["patterns"]:
if pattern.lower() in searchable:
triggered.append(indicator["name"])
break # one match per indicator is enough
return triggered
def get_escalation_path(incident_type: str, severity: str) -> dict:
"""
Return escalation routing for a given incident type and severity level.
Falls back to sev4 routing if severity is not recognised.
"""
sev_key = severity.lower()
routing = ESCALATION_ROUTING.get(sev_key, ESCALATION_ROUTING["sev4"]).copy()
# Augment with taxonomy SLA if available
taxonomy = INCIDENT_TAXONOMY.get(incident_type, {})
routing["incident_type"] = incident_type
routing["severity"] = sev_key
routing["response_sla_minutes"] = taxonomy.get("response_sla_minutes", 1440)
routing["mitre_technique"] = taxonomy.get("mitre", "N/A")
return routing
def check_sev_escalation_triggers(fact: dict) -> Optional[str]:
"""
Scan fact fields for any SEV escalation trigger indicators.
Returns the escalation target (e.g. 'sev1') if a trigger fires,
or None if no triggers are present.
"""
searchable = _flatten_to_string(fact).lower()
# Also inspect a flat list of explicit indicator flags
explicit_indicators: List[str] = []
if isinstance(fact.get("indicators"), list):
explicit_indicators = [str(i).lower() for i in fact["indicators"]]
if isinstance(fact.get("escalation_triggers"), list):
explicit_indicators += [str(i).lower() for i in fact["escalation_triggers"]]
for trigger in SEV_ESCALATION_TRIGGERS:
indicator_key = trigger["indicator"].replace("_", " ")
indicator_raw = trigger["indicator"].lower()
if (
indicator_key in searchable
or indicator_raw in searchable
or indicator_raw in explicit_indicators
):
return trigger["escalate_to"]
return None
# ---------------------------------------------------------------------------
# Severity Normalisation Helpers
# ---------------------------------------------------------------------------
_SEV_ORDER = {"sev1": 1, "sev2": 2, "sev3": 3, "sev4": 4}
def _sev_to_int(sev: str) -> int:
return _SEV_ORDER.get(sev.lower(), 4)
def _int_to_sev(n: int) -> str:
return {1: "sev1", 2: "sev2", 3: "sev3", 4: "sev4"}.get(n, "sev4")
def _escalate_sev(current: str, target: str) -> str:
"""Return the higher severity (lower SEV number)."""
return _int_to_sev(min(_sev_to_int(current), _sev_to_int(target)))
# ---------------------------------------------------------------------------
# Text Report
# ---------------------------------------------------------------------------
def _print_text_report(result: dict) -> None:
"""Print a human-readable triage report to stdout."""
sep = "=" * 70
print(sep)
print(" INCIDENT TRIAGE REPORT")
print(sep)
print(f" Timestamp : {result.get('timestamp_utc', 'N/A')}")
print(f" Incident Type : {result.get('incident_type', 'unknown').upper()}")
print(f" Severity : {result.get('severity', 'N/A').upper()}")
print(f" Confidence : {result.get('classification_confidence', 0.0):.0%}")
print(sep)
fp = result.get("false_positive_indicators", [])
if fp:
print(f"\n [!] FALSE POSITIVE FLAGS: {', '.join(fp)}")
print(" Review before escalating.")
esc_trigger = result.get("escalation_trigger_fired")
if esc_trigger:
print(f"\n [!] ESCALATION TRIGGER FIRED -> {esc_trigger.upper()}")
path = result.get("escalation_path", {})
print(f"\n Escalate To : {path.get('escalate_to', 'N/A')}")
print(f" Response SLA : {path.get('response_sla_minutes', 'N/A')} minutes")
print(f" Bridge Call : {'YES' if path.get('bridge_call') else 'no'}")
print(f" War Room : {'YES' if path.get('war_room') else 'no'}")
print(f" MITRE : {path.get('mitre_technique', 'N/A')}")
forensics = result.get("forensic_analysis", {})
if forensics:
print(f"\n Forensic Fields:")
print(f" Source IP : {forensics.get('source_ip', 'N/A')}")
print(f" User Account : {forensics.get('user_account', 'N/A')}")
print(f" Hostname : {forensics.get('hostname', 'N/A')}")
print(f" Process : {forensics.get('process_name', 'N/A')}")
print(f" Dwell (hrs) : {forensics.get('dwell_hours', 0.0)}")
print(f" Dwell Severity: {forensics.get('dwell_severity', 'N/A')}")
ioc_summary = result.get("ioc_summary", {})
if ioc_summary:
print(f"\n IOC Summary:")
print(f" Total IOCs : {ioc_summary.get('total_ioc_count', 0)}")
if ioc_summary.get("ip_indicators"):
print(f" IPs : {', '.join(ioc_summary['ip_indicators'])}")
if ioc_summary.get("hash_indicators"):
print(f" Hashes : {len(ioc_summary['hash_indicators'])} hash(es)")
print(f" Evidence Srcs : {', '.join(ioc_summary.get('evidence_sources_applicable', []))}")
print(f"\n Recommended Action: {result.get('recommended_action', 'N/A')}")
print(sep)
# ---------------------------------------------------------------------------
# Main Entry Point
# ---------------------------------------------------------------------------
def main() -> None:
parser = argparse.ArgumentParser(
description="Incident Classification, Triage, and Escalation",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
echo '{"event_type": "ransomware"}' | %(prog)s --json
%(prog)s --input event.json --classify --false-positive-check --json
%(prog)s --input event.json --severity sev1 --json
Exit codes:
0 SEV3/SEV4 or no confirmed incident
1 SEV2 — elevated response required
2 SEV1 — critical incident declared
""",
)
parser.add_argument(
"--input", "-i",
metavar="FILE",
help="JSON file path containing the security event (default: stdin)",
)
parser.add_argument(
"--json",
action="store_true",
help="Output results as JSON",
)
parser.add_argument(
"--classify",
action="store_true",
help="Run incident classification against INCIDENT_TAXONOMY",
)
parser.add_argument(
"--false-positive-check",
action="store_true",
dest="false_positive_check",
help="Run false positive filter checks",
)
parser.add_argument(
"--severity",
choices=["sev1", "sev2", "sev3", "sev4"],
help="Explicit severity override (skips taxonomy-derived severity)",
)
args = parser.parse_args()
# --- Load input ---
try:
if args.input:
with open(args.input, "r", encoding="utf-8") as fh:
raw_event = json.load(fh)
else:
raw_event = json.load(sys.stdin)
except json.JSONDecodeError as exc:
msg = {"error": f"Invalid JSON input: {exc}"}
if args.json:
print(json.dumps(msg, indent=2))
else:
print(f"Error: {msg['error']}", file=sys.stderr)
sys.exit(1)
except FileNotFoundError as exc:
msg = {"error": str(exc)}
if args.json:
print(json.dumps(msg, indent=2))
else:
print(f"Error: {msg['error']}", file=sys.stderr)
sys.exit(1)
# --- Forensic pre-analysis (base logic) ---
fields = parse_forensic_fields(raw_event)
ioc_summary = build_ioc_summary(fields)
forensic_analysis = {
"source_ip": fields["source_ip"],
"destination_ip": fields["destination_ip"],
"user_account": fields["user_account"],
"hostname": fields["hostname"],
"process_name": fields["process_name"],
"dwell_hours": fields["dwell_hours"],
"dwell_severity": assess_dwell_severity(fields["dwell_hours"]),
}
# --- Classification ---
incident_type = "unknown"
confidence = 0.0
if args.classify or not args.severity:
incident_type, confidence = classify_incident(raw_event)
# Override with explicit event_type if classify not run
if not args.classify:
et = str(raw_event.get("event_type", "")).lower().replace(" ", "_").replace("-", "_")
if et in INCIDENT_TAXONOMY:
incident_type = et
confidence = 0.75
# --- Determine base severity ---
if args.severity:
severity = args.severity.lower()
else:
taxonomy_entry = INCIDENT_TAXONOMY.get(incident_type, {})
severity = taxonomy_entry.get("default_severity", "sev4")
# Factor in dwell severity
dwell_sev_map = {"critical": "sev1", "high": "sev2", "medium": "sev3", "low": "sev4"}
dwell_derived = dwell_sev_map.get(forensic_analysis["dwell_severity"], "sev4")
severity = _escalate_sev(severity, dwell_derived)
# --- Escalation trigger check ---
escalation_trigger_fired: Optional[str] = None
trigger_result = check_sev_escalation_triggers(raw_event)
if trigger_result:
escalation_trigger_fired = trigger_result
severity = _escalate_sev(severity, trigger_result)
# --- False positive check ---
fp_indicators: List[str] = []
if args.false_positive_check:
fp_indicators = check_false_positives(raw_event)
# --- Escalation path ---
escalation_path = get_escalation_path(incident_type, severity)
# --- Recommended action ---
if fp_indicators:
recommended_action = (
f"Verify false positive flags before escalating: {', '.join(fp_indicators)}. "
"Confirm with asset owner and close or reclassify."
)
elif severity == "sev1":
recommended_action = (
"IMMEDIATE: Declare SEV1, open war room, page CISO and CEO. "
"Isolate affected systems, preserve evidence, activate IR playbook."
)
elif severity == "sev2":
recommended_action = (
"URGENT: Page SOC Lead and CISO. Open bridge call. "
"Contain impacted accounts/hosts and begin forensic collection."
)
elif severity == "sev3":
recommended_action = (
"Notify SOC Lead and Security Manager. "
"Investigate during business hours and document findings."
)
else:
recommended_action = (
"Queue for L3 Analyst review. "
"Document and track per standard operating procedure."
)
# --- Assemble output ---
result: Dict[str, Any] = {
"incident_type": incident_type,
"classification_confidence": confidence,
"severity": severity,
"false_positive_indicators": fp_indicators,
"escalation_trigger_fired": escalation_trigger_fired,
"escalation_path": escalation_path,
"forensic_analysis": forensic_analysis,
"ioc_summary": ioc_summary,
"recommended_action": recommended_action,
"taxonomy": INCIDENT_TAXONOMY.get(incident_type, {}),
"timestamp_utc": datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ"),
}
# --- Output ---
if args.json:
print(json.dumps(result, indent=2))
else:
_print_text_report(result)
# --- Exit code ---
if severity == "sev1":
sys.exit(2)
elif severity == "sev2":
sys.exit(1)
else:
sys.exit(0)
if __name__ == "__main__":
main()