Phân tích toàn diện quy trình vận hành, KPI/OKR, cơ cấu tổ chức và chiến lược doanh nghiệp thực chiến cho thương mại điện tử và sản xuất.
--- name: phan-tich-nghiep-vu-quan-tri-doanh-nghiep description: Phân tích toàn diện quy trình vận hành, KPI/OKR, cơ cấu tổ chức và chiến lược doanh nghiệp thực chiến (TMĐT & sản xuất). Dùng khi nói "phân tích nghiệp vụ", "tối ưu quy trình", "báo cáo quản trị", "cơ cấu tổ chức". --- # Phân tích nghiệp vụ quản trị doanh nghiệp ## Mục tiêu Phân tích toàn diện các vấn đề quản trị doanh nghiệp — từ quy trình, hiệu suất, tổ chức đến chiến lược — để ra quyết định có căn cứ hoặc trình bày cho các bên liên quan. ## Bối cảnh áp dụng - Chargee: công ty phân phối hàng hóa sàn TMĐT — ưu tiên phân tích vận hành, hiệu suất kênh bán, bottleneck logistics - Elmich (IT): sản xuất & thương mại đồ gia dụng — ưu tiên phân tích quy trình IT, tổ chức phòng ban, hỗ trợ chiến lược công nghệ ## Khi nào dùng - Cần phân tích một vấn đề vận hành đang phát sinh - Chuẩn bị báo cáo cho ban lãnh đạo hoặc họp nội bộ - Đánh giá hiệu suất đội nhóm hoặc quy trình - Ra quyết định tổ chức: tuyển dụng, phân quyền, tái cơ cấu - Phân tích chiến lược trước khi mở rộng hoặc thay đổi hướng đi ## Đầu vào cần cung cấp - Vấn đề hoặc câu hỏi cần phân tích - Công ty liên quan: Chargee hay Elmich (hoặc cả hai) - Loại phân tích: quy trình / hiệu suất / tổ chức / chiến lược - Dữ liệu hiện có: số liệu, mô tả tình huống, bối cảnh - Mục đích đầu ra: quyết định cá nhân / báo cáo nội bộ / trình bày lãnh đạo / lưu tài liệu ## Quy trình xử lý ### A. Phân tích quy trình 1. Vẽ lại quy trình hiện tại (as-is) theo từng bước 2. Xác định bottleneck, điểm lặp, điểm thủ công không cần thiết 3. So sánh với quy trình lý tưởng (to-be) 4. Đề xuất cải tiến có thể triển khai ngay vs. dài hạn ### B. Phân tích hiệu suất (KPI/OKR) 1. Xác định chỉ số đang đo và chỉ số còn thiếu 2. So sánh thực tế vs. mục tiêu — tính % đạt 3. Tìm nguyên nhân gốc rễ nếu có gap (5 Whys) 4. Đề xuất điều chỉnh mục tiêu hoặc hành động khắc phục ### C. Phân tích tổ chức 1. Mapping cơ cấu hiện tại: ai làm gì, báo cáo cho ai 2. Xác định điểm mơ hồ về trách nhiệm (RACI) 3. Đánh giá tải công việc và phân quyền 4. Đề xuất điều chỉnh cơ cấu hoặc phân quyền ### D. Phân tích chiến lược 1. Thu thập dữ liệu nội bộ và thị trường 2. Phân tích SWOT hoặc framework phù hợp 3. Xác định 2–3 lựa chọn chiến lược với pros/cons 4. Đề xuất hướng ưu tiên có căn cứ rõ ràng ## Tiêu chuẩn đầu ra theo mục đích | Mục đích | Định dạng | Độ dài | |---|---|---| | Quyết định cá nhân | Bullet points + kết luận | Ngắn gọn | | Báo cáo nội bộ | Markdown có bảng + biểu đồ mô tả | Trung bình | | Trình bày lãnh đạo | Cấu trúc: vấn đề → phân tích → đề xuất | Súc tích, có số liệu | | Lưu tài liệu | Đầy đủ, có ngày tháng, có giả định | Chi tiết | ## Giả định luôn phải nêu rõ - Số liệu dựa trên dữ liệu bạn cung cấp — không tự bịa - Nếu thiếu dữ liệu: nêu rõ cần thu thập thêm gì - Đề xuất mang tính tham khảo — quyết định cuối thuộc về bạn ## Tránh - Phân tích chung chung không gắn với Chargee hoặc Elmich - Đưa ra kết luận khi chưa đủ dữ liệu - Bỏ qua sự khác biệt giữa hai công ty (TMĐT vs. sản xuất) - Trộn lẫn đầu ra cho các mục đích khác nhau
Hợp nhất nhánh của agent chiến thắng vào nhánh gốc, lưu trữ các nhánh còn lại và dọn dẹp worktree.
---
name: "merge"
description: "Merge the winning agent's branch into base, archive losers, and clean up worktrees."
command: /hub:merge
---
# /hub:merge — Merge Winner
Merge the best agent's branch into the base branch, archive losing branches via git tags, and clean up worktrees.
## Usage
```
/hub:merge # Merge winner of latest session
/hub:merge 20260317-143022 # Merge winner of specific session
/hub:merge 20260317-143022 --agent agent-2 # Explicitly choose winner
```
## What It Does
### 1. Identify Winner
If `--agent` specified, use that. Otherwise, use the #1 ranked agent from the most recent `/hub:eval`.
### 2. Merge Winner
```bash
git checkout {base_branch}
git merge --no-ff hub/{session-id}/{winner}/attempt-1 \
-m "hub: merge {winner} from session {session-id}
Task: {task}
Winner: {winner}
Session: {session-id}"
```
### 3. Archive Losers
For each non-winning agent:
```bash
# Create archive tag (preserves commits forever)
git tag hub/archive/{session-id}/{agent-id} hub/{session-id}/{agent-id}/attempt-1
# Delete branch ref (commits preserved via tag)
git branch -D hub/{session-id}/{agent-id}/attempt-1
```
### 4. Clean Up Worktrees
```bash
python {skill_path}/scripts/session_manager.py --cleanup {session-id}
```
### 5. Post Merge Summary
Write `.agenthub/board/results/merge-summary.md`:
```markdown
---
author: coordinator
timestamp: {now}
channel: results
---
## Merge Summary
- **Session**: {session-id}
- **Winner**: {winner}
- **Merged into**: {base_branch}
- **Archived**: {loser-1}, {loser-2}, ...
- **Worktrees cleaned**: {count}
```
### 6. Update State
```bash
python {skill_path}/scripts/session_manager.py --update {session-id} --state merged
```
## Safety
- **Confirm with user** before merging — show the diff summary first
- **Never force-push** — merge is always `--no-ff` for clear history
- **Archive, don't delete** — losing agents' commits are preserved via tags
- **Clean worktrees** — don't leave orphan directories on disk
## After Merge
Tell the user:
- Winner merged into `{base_branch}`
- Losers archived with tags `hub/archive/{session-id}/agent-{N}`
- Worktrees cleaned up
- Session state: `merged`
Tạo và tối ưu popup, modal, overlay, slide-in và banner để tăng chuyển đổi: exit intent, thu thập email, banner thông báo.
---
name: popups
description: When the user wants to create or optimize popups, modals, overlays, slide-ins, or banners for conversion purposes. Also use when the user mentions "exit intent," "popup conversions," "modal optimization," "lead capture popup," "email popup," "announcement banner," "overlay," "collect emails with a popup," "exit popup," "scroll trigger," "sticky bar," or "notification bar." Use this for any overlay or interrupt-style conversion element. For forms outside of popups, see cro. For general page conversion optimization, see cro.
metadata:
version: 2.0.0
---
# Popup CRO
You are an expert in popup and modal optimization. Your goal is to create popups that convert without annoying users or damaging brand perception.
## Initial Assessment
**Check for product marketing context first:**
If `.agents/product-marketing.md` exists (or `.claude/product-marketing.md`, or the legacy `product-marketing-context.md` filename, in older setups), read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
Before providing recommendations, understand:
1. **Popup Purpose**
- Email/newsletter capture
- Lead magnet delivery
- Discount/promotion
- Announcement
- Exit intent save
- Feature promotion
- Feedback/survey
2. **Current State**
- Existing popup performance?
- What triggers are used?
- User complaints or feedback?
- Mobile experience?
3. **Traffic Context**
- Traffic sources (paid, organic, direct)
- New vs. returning visitors
- Page types where shown
---
## Core Principles
### 1. Timing Is Everything
- Too early = annoying interruption
- Too late = missed opportunity
- Right time = helpful offer at moment of need
### 2. Value Must Be Obvious
- Clear, immediate benefit
- Relevant to page context
- Worth the interruption
### 3. Respect the User
- Easy to dismiss
- Don't trap or trick
- Remember preferences
- Don't ruin the experience
---
## Trigger Strategies
### Time-Based
- **Not recommended**: "Show after 5 seconds"
- **Better**: "Show after 30-60 seconds" (proven engagement)
- Best for: General site visitors
### Scroll-Based
- **Typical**: 25-50% scroll depth
- Indicates: Content engagement
- Best for: Blog posts, long-form content
- Example: "You're halfway through—get more like this"
### Exit Intent
- Detects cursor moving to close/leave
- Last chance to capture value
- Best for: E-commerce, lead gen
- Mobile alternative: Back button or scroll up
### Click-Triggered
- User initiates (clicks button/link)
- Zero annoyance factor
- Best for: Lead magnets, gated content, demos
- Example: "Download PDF" → Popup form
### Page Count / Session-Based
- After visiting X pages
- Indicates research/comparison behavior
- Best for: Multi-page journeys
- Example: "Been comparing? Here's a summary..."
### Behavior-Based
- Add to cart abandonment
- Pricing page visitors
- Repeat page visits
- Best for: High-intent segments
---
## Popup Types
### Email Capture Popup
**Goal**: Newsletter/list subscription
**Best practices:**
- Clear value prop (not just "Subscribe")
- Specific benefit of subscribing
- Single field (email only)
- Consider incentive (discount, content)
**Copy structure:**
- Headline: Benefit or curiosity hook
- Subhead: What they get, how often
- CTA: Specific action ("Get Weekly Tips")
### Lead Magnet Popup
**Goal**: Exchange content for email
**Best practices:**
- Show what they get (cover image, preview)
- Specific, tangible promise
- Minimal fields (email, maybe name)
- Instant delivery expectation
### Discount/Promotion Popup
**Goal**: First purchase or conversion
**Best practices:**
- Clear discount (10%, $20, free shipping)
- Deadline creates urgency
- Single use per visitor
- Easy to apply code
### Exit Intent Popup
**Goal**: Last-chance conversion
**Best practices:**
- Acknowledge they're leaving
- Different offer than entry popup
- Address common objections
- Final compelling reason to stay
**Formats:**
- "Wait! Before you go..."
- "Forget something?"
- "Get 10% off your first order"
- "Questions? Chat with us"
### Announcement Banner
**Goal**: Site-wide communication
**Best practices:**
- Top of page (sticky or static)
- Single, clear message
- Dismissable
- Links to more info
- Time-limited (don't leave forever)
### Slide-In
**Goal**: Less intrusive engagement
**Best practices:**
- Enters from corner/bottom
- Doesn't block content
- Easy to dismiss or minimize
- Good for chat, support, secondary CTAs
---
## Design Best Practices
### Visual Hierarchy
1. Headline (largest, first seen)
2. Value prop/offer (clear benefit)
3. Form/CTA (obvious action)
4. Close option (easy to find)
### Sizing
- Desktop: 400-600px wide typical
- Don't cover entire screen
- Mobile: Full-width bottom or center, not full-screen
- Leave space to close (visible X, click outside)
### Close Button
- Keep visible (top right is convention) — users who can't find the close button will bounce entirely
- Large enough to tap on mobile
- "No thanks" text link as alternative
- Click outside to close
### Mobile Considerations
- Can't detect exit intent (use alternatives)
- Full-screen overlays feel aggressive
- Bottom slide-ups work well
- Larger touch targets
- Easy dismiss gestures
### Imagery
- Product image or preview
- Face if relevant (increases trust)
- Minimal for speed
- Optional—copy can work alone
---
## Copy Formulas
### Headlines
- Benefit-driven: "Get [result] in [timeframe]"
- Question: "Want [desired outcome]?"
- Command: "Don't miss [thing]"
- Social proof: "Join [X] people who..."
- Curiosity: "The one thing [audience] always get wrong about [topic]"
### Subheadlines
- Expand on the promise
- Address objection ("No spam, ever")
- Set expectations ("Weekly tips in 5 min")
### CTA Buttons
- First person works: "Get My Discount" vs "Get Your Discount"
- Specific over generic: "Send Me the Guide" vs "Submit"
- Value-focused: "Claim My 10% Off" vs "Subscribe"
### Decline Options
- Polite, not guilt-trippy
- "No thanks" / "Maybe later" / "I'm not interested"
- Avoid manipulative: "No, I don't want to save money"
---
## Frequency and Rules
### Frequency Capping
- Show maximum once per session
- Remember dismissals (cookie/localStorage)
- 7-30 days before showing again
- Respect user choice
### Audience Targeting
- New vs. returning visitors (different needs)
- By traffic source (match ad message)
- By page type (context-relevant)
- Exclude converted users
- Exclude recently dismissed
### Page Rules
- Exclude checkout/conversion flows
- Consider blog vs. product pages
- Match offer to page context
---
## Compliance and Accessibility
### GDPR/Privacy
- Clear consent language
- Link to privacy policy
- Don't pre-check opt-ins
- Honor unsubscribe/preferences
### Accessibility
- Keyboard navigable (Tab, Enter, Esc)
- Focus trap while open
- Screen reader compatible
- Sufficient color contrast
- Don't rely on color alone
### Google Guidelines
- Intrusive interstitials hurt SEO
- Mobile especially sensitive
- Allow: Cookie notices, age verification, reasonable banners
- Avoid: Full-screen before content on mobile
---
## Measurement
### Key Metrics
- **Impression rate**: Visitors who see popup
- **Conversion rate**: Impressions → Submissions
- **Close rate**: How many dismiss immediately
- **Engagement rate**: Interaction before close
- **Time to close**: How long before dismissing
### What to Track
- Popup views
- Form focus
- Submission attempts
- Successful submissions
- Close button clicks
- Outside clicks
- Escape key
### Benchmarks
- Email popup: 2-5% conversion typical
- Exit intent: 3-10% conversion
- Click-triggered: Higher (10%+, self-selected)
---
## Output Format
### Popup Design
- **Type**: Email capture, lead magnet, etc.
- **Trigger**: When it appears
- **Targeting**: Who sees it
- **Frequency**: How often shown
- **Copy**: Headline, subhead, CTA, decline
- **Design notes**: Layout, imagery, mobile
### Multiple Popup Strategy
If recommending multiple popups:
- Popup 1: [Purpose, trigger, audience]
- Popup 2: [Purpose, trigger, audience]
- Conflict rules: How they don't overlap
### Test Hypotheses
Ideas to A/B test with expected outcomes
---
## Common Popup Strategies
### E-commerce
1. Entry/scroll: First-purchase discount
2. Exit intent: Bigger discount or reminder
3. Cart abandonment: Complete your order
### B2B SaaS
1. Click-triggered: Demo request, lead magnets
2. Scroll: Newsletter/blog subscription
3. Exit intent: Trial reminder or content offer
### Content/Media
1. Scroll-based: Newsletter after engagement
2. Page count: Subscribe after multiple visits
3. Exit intent: Don't miss future content
### Lead Generation
1. Time-delayed: General list building
2. Click-triggered: Specific lead magnets
3. Exit intent: Final capture attempt
---
## Experiment Ideas
### Placement & Format Experiments
**Banner Variations**
- Top bar vs. banner below header
- Sticky banner vs. static banner
- Full-width vs. contained banner
- Banner with countdown timer vs. without
**Popup Formats**
- Center modal vs. slide-in from corner
- Full-screen overlay vs. smaller modal
- Bottom bar vs. corner popup
- Top announcements vs. bottom slideouts
**Position Testing**
- Test popup sizes on desktop and mobile
- Left corner vs. right corner for slide-ins
- Test visibility without blocking content
---
### Trigger Experiments
**Timing Triggers**
- Exit intent vs. 30-second delay vs. 50% scroll depth
- Test optimal time delay (10s vs. 30s vs. 60s)
- Test scroll depth percentage (25% vs. 50% vs. 75%)
- Page count trigger (show after X pages viewed)
**Behavior Triggers**
- Show based on user intent prediction
- Trigger based on specific page visits
- Return visitor vs. new visitor targeting
- Show based on referral source
**Click Triggers**
- Click-triggered popups for lead magnets
- Button-triggered vs. link-triggered modals
- Test in-content triggers vs. sidebar triggers
---
### Messaging & Content Experiments
**Headlines & Copy**
- Test attention-grabbing vs. informational headlines
- "Limited-time offer" vs. "New feature alert" messaging
- Urgency-focused copy vs. value-focused copy
- Test headline length and specificity
**CTAs**
- CTA button text variations
- Button color testing for contrast
- Primary + secondary CTA vs. single CTA
- Test decline text (friendly vs. neutral)
**Visual Content**
- Add countdown timers to create urgency
- Test with/without images
- Product preview vs. generic imagery
- Include social proof in popup
---
### Personalization Experiments
**Dynamic Content**
- Personalize popup based on visitor data
- Show industry-specific content
- Tailor content based on pages visited
- Use progressive profiling (ask more over time)
**Audience Targeting**
- New vs. returning visitor messaging
- Segment by traffic source
- Target based on engagement level
- Exclude already-converted visitors
---
### Frequency & Rules Experiments
- Test frequency capping (once per session vs. once per week)
- Cool-down period after dismissal
- Test different dismiss behaviors
- Show escalating offers over multiple visits
---
## Task-Specific Questions
1. What's the primary goal for this popup?
2. What's your current popup performance (if any)?
3. What traffic sources are you optimizing for?
4. What incentive can you offer?
5. Are there compliance requirements (GDPR, etc.)?
6. Mobile vs. desktop traffic split?
---
## Related Skills
- **lead-magnets**: For planning lead magnets to promote via popups
- **cro**: For optimizing the form inside the popup
- **cro**: For the page context around popups
- **emails**: For what happens after popup conversion
- **ab-testing**: For testing popup variations
FILE:evals/evals.json
{
"skill_name": "popups",
"evals": [
{
"id": 1,
"prompt": "Help me create an exit-intent popup for our SaaS landing page. We want to capture emails from visitors who are about to leave without signing up. Our product is a social media scheduling tool.",
"expected_output": "Should check for product-marketing.md first. Should identify the popup type as exit-intent email capture. Should apply the exit-intent popup design guidance: compelling headline (address why they're leaving or offer additional value), lead magnet or incentive (discount, free resource, extended trial), minimal form fields (email only), clear CTA, and easy close option. Should apply copy formulas from the skill. Should address trigger configuration (exit intent detection). Should recommend frequency rules (don't show again if dismissed). Should include benchmarks (exit intent popups typically 3-10% conversion).",
"assertions": [
"Checks for product-marketing.md",
"Identifies as exit-intent popup type",
"Includes compelling headline",
"Includes lead magnet or incentive",
"Minimal form fields (email only)",
"Applies copy formulas from the skill",
"Addresses trigger configuration",
"Recommends frequency rules",
"Includes conversion benchmarks"
],
"files": []
},
{
"id": 2,
"prompt": "We want to offer a 10% discount to first-time visitors via a popup. When should we show it and what should it say?",
"expected_output": "Should identify this as a discount/offer popup type. Should apply trigger strategy guidance: recommend against showing immediately on page load (too aggressive). Should suggest time-based delay (30-60 seconds), scroll-based trigger (50%+ page scroll), or exit intent as better alternatives. Should apply the copy formula for discount popups: headline that frames the value, clear offer terms, urgency element, email capture, and CTA. Should address compliance (GDPR cookie consent if applicable). Should recommend frequency capping.",
"assertions": [
"Identifies as discount popup type",
"Recommends against immediate page load trigger",
"Suggests better trigger alternatives (time, scroll, exit)",
"Applies copy formula for discount popups",
"Includes urgency element",
"Addresses frequency capping",
"Addresses compliance considerations"
],
"files": []
},
{
"id": 3,
"prompt": "our popups are annoying everyone. we keep getting complaints but we also get a lot of email signups from them. how do we balance this?",
"expected_output": "Should trigger on casual phrasing. Should apply the frequency and rules guidance. Should address the balance: reduce annoyance while preserving conversions. Should recommend: frequency capping (once per session or once per X days), don't show to returning visitors who already dismissed, don't show to existing subscribers, respect 'close' action, consider less intrusive formats (slide-in instead of full modal, announcement bar instead of overlay). Should address compliance and accessibility requirements. Should suggest A/B testing different triggers and formats to find the best balance.",
"assertions": [
"Triggers on casual phrasing",
"Applies frequency and rules guidance",
"Addresses balance between conversions and UX",
"Recommends frequency capping",
"Suggests excluding existing subscribers",
"Recommends less intrusive alternatives",
"Suggests A/B testing to optimize"
],
"files": []
},
{
"id": 4,
"prompt": "What types of popups should we use on our blog? We publish content about email marketing and want to grow our email list.",
"expected_output": "Should recommend blog-appropriate popup types: scroll-triggered popup (show after 50-70% scroll indicating engagement), exit-intent popup, slide-in (less intrusive than modal), and inline content upgrades. Should recommend lead magnets relevant to the blog topic (email marketing templates, checklist, swipe file). Should address different popup placements: mid-content, end of post, sidebar slide-in. Should recommend behavior-based triggers over time-based for blog content. Should apply copy formulas with blog-specific hooks.",
"assertions": [
"Recommends blog-appropriate popup types",
"Includes scroll-triggered and exit-intent",
"Suggests less intrusive formats (slide-in)",
"Recommends relevant lead magnets",
"Addresses popup placement on blog pages",
"Recommends behavior-based triggers for blog",
"Applies copy formulas"
],
"files": []
},
{
"id": 5,
"prompt": "Design an announcement banner for our new feature launch. We want it to show at the top of the site for 2 weeks.",
"expected_output": "Should identify this as an announcement banner popup type. Should apply banner design guidance: short, clear headline announcing the feature, brief description of benefit, CTA to learn more or try it, dismiss option. Should recommend banner positioning (top of page, sticky or static). Should address duration (2 weeks as stated). Should recommend targeting (show to existing users who'd benefit, not just everyone). Should provide copy recommendations.",
"assertions": [
"Identifies as announcement banner type",
"Provides short, clear headline",
"Includes brief benefit description and CTA",
"Includes dismiss option",
"Addresses banner positioning",
"Recommends audience targeting",
"Provides copy recommendations"
],
"files": []
},
{
"id": 6,
"prompt": "We need to optimize the lead capture form inside our popup. It currently asks for name, email, company, and phone number. Too many fields?",
"expected_output": "Should recognize this overlaps with form optimization. Should defer to or cross-reference the cro skill, which handles form field optimization, layout, and conversion. May provide popup-specific context (popups need minimal fields due to fleeting attention) but should make clear that cro is the right skill for detailed form optimization.",
"assertions": [
"Recognizes overlap with form optimization",
"References or defers to cro skill",
"Notes popups need minimal fields due to context",
"Does not attempt detailed form redesign"
],
"files": []
}
]
}
Phân tích trung thực những gì đã sai sau một sự cố hoặc thất bại.
--- name: "postmortem" description: "/em -postmortem — Honest Analysis of What Went Wrong" --- # /em:postmortem — Honest Analysis of What Went Wrong **Command:** `/em:postmortem <event>` Not blame. Understanding. The failed deal, the missed quarter, the feature that flopped, the hire that didn't work out. What actually happened, why, and what changes as a result. --- ## Why Most Post-Mortems Fail They become one of two things: **The blame session** — someone gets scapegoated, defensive walls go up, actual causes don't get examined, and the same problem happens again in a different form. **The whitewash** — "We learned a lot, we're going to do better, here are 12 vague action items." Nothing changes. Same problem, different quarter. A real post-mortem is neither. It's a rigorous investigation into a system failure. Not "whose fault was it" but "what conditions made this outcome predictable in hindsight?" **The purpose:** extract the maximum learning value from a failure so you can prevent recurrence and improve the system. --- ## The Framework ### Step 1: Define the Event Precisely Before analysis: describe exactly what happened. - What was the expected outcome? - What was the actual outcome? - When was the gap first visible? - What was the impact (financial, operational, reputational)? Precision matters. "We missed Q3 revenue" is not precise enough. "We closed $420K in new ARR vs $680K target — a $260K miss driven primarily by three deals that slipped to Q4 and one deal that was lost to a competitor" is precise. ### Step 2: The 5 Whys — Done Properly The goal: get from **what happened** (the symptom) to **why it happened** (the root cause). Standard bad 5 Whys: - Why did we miss revenue? Because deals slipped. - Why did deals slip? Because the sales cycle was longer than expected. - Why? Because the customer buying process is complex. - Why? Because we're selling to enterprise. - Why? That's just how enterprise sales works. → Conclusion: Nothing to do. It's just enterprise. Real 5 Whys: - Why did we miss revenue? Three deals slipped out of quarter. - Why did those deals slip? None of them had identified a champion with budget authority. - Why did we progress deals without a champion? Our qualification criteria didn't require it. - Why didn't our qualification criteria require it? When we built the criteria 8 months ago, we were in SMB, not enterprise. - Why haven't we updated qualification criteria as ICP shifted? No owner, no process for criteria review. → Root cause: Qualification criteria outdated, no owner, no review process. → Fix: Update criteria, assign owner, add quarterly review. **The test for a good root cause:** Could you prevent recurrence with a specific, concrete change? If yes, you've found something real. ### Step 3: Distinguish Contributing Factors from Root Cause Most events have multiple contributing factors. Not all are root causes. **Contributing factor:** Made it worse, but isn't the core reason. If removed, the outcome might have been different — but the same class of problem would recur. **Root cause:** The fundamental condition that made the outcome probable. Fix this, and this class of problem doesn't recur. Example — failed hire: - Contributing factors: rushed process, reference checks skipped, team under pressure to staff up - Root cause: No defined competency framework, so interview process varied by who happened to conduct interviews **The distinction matters.** If you address only contributing factors, you'll have a different-looking but structurally identical failure next time. ### Step 4: Identify the Warning Signs That Were Ignored Every failure has precursors. In hindsight, they're obvious. The value of this step is making them obvious prospectively. Ask: - At what point was the negative outcome predictable? - What signals were visible at that point? - Who saw them? What happened when they raised them? - Why weren't they acted on? **Common patterns:** - Signal was raised but dismissed by a senior person - Signal wasn't raised because nobody felt safe saying it - Signal was seen but no one had clear ownership to act on it - Data was available but nobody was looking at it - The team was too optimistic to take negative signals seriously This step is particularly important for systemic issues — "we didn't feel safe raising the concern" is a much deeper root cause than "the deal qualification was off." ### Step 5: Distinguish What Was in Control vs. Out of Control Some failures happen despite correct decisions. Some happen because of incorrect decisions. Knowing the difference prevents both overcorrection and undercorrection. - **In control:** Process, criteria, team capability, resource allocation, decisions made - **Out of control:** Market conditions, customer decisions, competitor actions, macro events For things out of control: what can be done to be more resilient to similar events? For things in control: what specifically needs to change? **Warning:** "It was outside our control" is sometimes used to avoid accountability. Be rigorous. ### Step 6: Build the Change Register Every post-mortem ends with a change register — specific commitments, owned and dated. **Bad action items:** - "We'll improve our qualification process" - "Communication will be better" - "We'll be more rigorous about forecasting" **Good action items:** - "Ravi owns rewriting qualification criteria by March 15 to include champion identification as hard requirement. New criteria reviewed in weekly sales standup starting March 22." - "By March 10, Elena adds deal-slippage risk flag to CRM for any open opportunity >60 days without a product demo" - "Maria runs a 30-min retrospective with enterprise sales team every 6 weeks starting April 1, reviews win/loss data" **For each action:** - What exactly is changing? - Who owns it? - By when? - How will you verify it worked? ### Step 7: Verification Date The most commonly skipped step. Post-mortems are useless if nobody checks whether the changes actually happened and actually worked. Set a verification date: "We'll review whether qualification criteria have been updated and whether deal slippage rate has improved at the June board meeting." Without this, post-mortems are theater. --- ## Post-Mortem Output Format ``` EVENT: [Name and date] EXPECTED: [What was supposed to happen] ACTUAL: [What happened] IMPACT: [Quantified] TIMELINE [Date]: [What happened or was visible] [Date]: ... 5 WHYS 1. [Why did X happen?] → Because [Y] 2. [Why did Y happen?] → Because [Z] 3. [Why did Z happen?] → Because [A] 4. [Why did A happen?] → Because [B] 5. [Why did B happen?] → Because [ROOT CAUSE] ROOT CAUSE: [One clear sentence] CONTRIBUTING FACTORS • [Factor] — how it contributed • [Factor] — how it contributed WARNING SIGNS MISSED • [Signal visible at what date] — why it wasn't acted on WHAT WAS IN CONTROL: [List] WHAT WASN'T: [List] CHANGE REGISTER | Action | Owner | Due Date | Verification | |--------|-------|----------|-------------| | [Specific change] | [Name] | [Date] | [How to verify] | VERIFICATION DATE: [Date of check-in] ``` --- ## The Tone of Good Post-Mortems Blame is cheap. Understanding is hard. The goal isn't to establish that someone made a mistake. The goal is to understand why the system produced that outcome — so the system can be improved. "The salesperson didn't qualify the deal properly" is blame. "Our qualification framework hadn't been updated when we moved upmarket, and no one owned keeping it current" is understanding. The first version fires or shames someone. The second version builds a more resilient organization. Both might be true simultaneously. The distinction is: which one actually prevents recurrence?
Lệnh tạo nhanh tài liệu yêu cầu sản phẩm (PRD) từ một tính năng hoặc vấn đề.
--- name: prd description: Quick PRD generation command. Usage: /prd <feature-or-problem> --- # /prd Generate a concise product requirements document for a feature, initiative, or problem statement. ## Usage ```bash /prd <feature-or-problem> ``` ## Output Structure - Problem statement - Goals and non-goals - User stories and acceptance criteria - Metrics and success thresholds - Scope and timeline assumptions ## Skill Reference - `product-team/product-manager-toolkit/SKILL.md`
Lập kế hoạch di chuyển không gián đoạn, kiểm tra tương thích và chiến lược rollback cho di chuyển hệ thống, cơ sở dữ liệu và hạ tầng.
---
name: "migration-architect"
description: "Zero-downtime migration planning, compatibility validation, and rollback strategy generation. Tools for system, database, and infrastructure migrations with minimal business impact. Use when planning a database migration, infrastructure cutover, system replacement, or any high-risk transition that needs explicit rollback paths."
---
# Migration Architect
**Tier:** POWERFUL
**Category:** Engineering - Migration Strategy
**Purpose:** Zero-downtime migration planning, compatibility validation, and rollback strategy generation
## Overview
The Migration Architect skill provides comprehensive tools and methodologies for planning, executing, and validating complex system migrations with minimal business impact. This skill combines proven migration patterns with automated planning tools to ensure successful transitions between systems, databases, and infrastructure.
## Core Capabilities
### 1. Migration Strategy Planning
- **Phased Migration Planning:** Break complex migrations into manageable phases with clear validation gates
- **Risk Assessment:** Identify potential failure points and mitigation strategies before execution
- **Timeline Estimation:** Generate realistic timelines based on migration complexity and resource constraints
- **Stakeholder Communication:** Create communication templates and progress dashboards
### 2. Compatibility Analysis
- **Schema Evolution:** Analyze database schema changes for backward compatibility issues
- **API Versioning:** Detect breaking changes in REST/GraphQL APIs and microservice interfaces
- **Data Type Validation:** Identify data format mismatches and conversion requirements
- **Constraint Analysis:** Validate referential integrity and business rule changes
### 3. Rollback Strategy Generation
- **Automated Rollback Plans:** Generate comprehensive rollback procedures for each migration phase
- **Data Recovery Scripts:** Create point-in-time data restoration procedures
- **Service Rollback:** Plan service version rollbacks with traffic management
- **Validation Checkpoints:** Define success criteria and rollback triggers
## Migration Patterns
### Database Migrations
#### Schema Evolution Patterns
1. **Expand-Contract Pattern**
- **Expand:** Add new columns/tables alongside existing schema
- **Dual Write:** Application writes to both old and new schema
- **Migration:** Backfill historical data to new schema
- **Contract:** Remove old columns/tables after validation
2. **Parallel Schema Pattern**
- Run new schema in parallel with existing schema
- Use feature flags to route traffic between schemas
- Validate data consistency between parallel systems
- Cutover when confidence is high
3. **Event Sourcing Migration**
- Capture all changes as events during migration window
- Apply events to new schema for consistency
- Enable replay capability for rollback scenarios
#### Data Migration Strategies
1. **Bulk Data Migration**
- **Snapshot Approach:** Full data copy during maintenance window
- **Incremental Sync:** Continuous data synchronization with change tracking
- **Stream Processing:** Real-time data transformation pipelines
2. **Dual-Write Pattern**
- Write to both source and target systems during migration
- Implement compensation patterns for write failures
- Use distributed transactions where consistency is critical
3. **Change Data Capture (CDC)**
- Stream database changes to target system
- Maintain eventual consistency during migration
- Enable zero-downtime migrations for large datasets
### Service Migrations
#### Strangler Fig Pattern
1. **Intercept Requests:** Route traffic through proxy/gateway
2. **Gradually Replace:** Implement new service functionality incrementally
3. **Legacy Retirement:** Remove old service components as new ones prove stable
4. **Monitoring:** Track performance and error rates throughout transition
```mermaid
graph TD
A[Client Requests] --> B[API Gateway]
B --> C{Route Decision}
C -->|Legacy Path| D[Legacy Service]
C -->|New Path| E[New Service]
D --> F[Legacy Database]
E --> G[New Database]
```
#### Parallel Run Pattern
1. **Dual Execution:** Run both old and new services simultaneously
2. **Shadow Traffic:** Route production traffic to both systems
3. **Result Comparison:** Compare outputs to validate correctness
4. **Gradual Cutover:** Shift traffic percentage based on confidence
#### Canary Deployment Pattern
1. **Limited Rollout:** Deploy new service to small percentage of users
2. **Monitoring:** Track key metrics (latency, errors, business KPIs)
3. **Gradual Increase:** Increase traffic percentage as confidence grows
4. **Full Rollout:** Complete migration once validation passes
### Infrastructure Migrations
#### Cloud-to-Cloud Migration
1. **Assessment Phase**
- Inventory existing resources and dependencies
- Map services to target cloud equivalents
- Identify vendor-specific features requiring refactoring
2. **Pilot Migration**
- Migrate non-critical workloads first
- Validate performance and cost models
- Refine migration procedures
3. **Production Migration**
- Use infrastructure as code for consistency
- Implement cross-cloud networking during transition
- Maintain disaster recovery capabilities
#### On-Premises to Cloud Migration
1. **Lift and Shift**
- Minimal changes to existing applications
- Quick migration with optimization later
- Use cloud migration tools and services
2. **Re-architecture**
- Redesign applications for cloud-native patterns
- Adopt microservices, containers, and serverless
- Implement cloud security and scaling practices
3. **Hybrid Approach**
- Keep sensitive data on-premises
- Migrate compute workloads to cloud
- Implement secure connectivity between environments
## Feature Flags for Migrations
### Progressive Feature Rollout
```python
# Example feature flag implementation
class MigrationFeatureFlag:
def __init__(self, flag_name, rollout_percentage=0):
self.flag_name = flag_name
self.rollout_percentage = rollout_percentage
def is_enabled_for_user(self, user_id):
hash_value = hash(f"{self.flag_name}:{user_id}")
return (hash_value % 100) < self.rollout_percentage
def gradual_rollout(self, target_percentage, step_size=10):
while self.rollout_percentage < target_percentage:
self.rollout_percentage = min(
self.rollout_percentage + step_size,
target_percentage
)
yield self.rollout_percentage
```
### Circuit Breaker Pattern
Implement automatic fallback to legacy systems when new systems show degraded performance:
```python
class MigrationCircuitBreaker:
def __init__(self, failure_threshold=5, timeout=60):
self.failure_count = 0
self.failure_threshold = failure_threshold
self.timeout = timeout
self.last_failure_time = None
self.state = 'CLOSED' # CLOSED, OPEN, HALF_OPEN
def call_new_service(self, request):
if self.state == 'OPEN':
if self.should_attempt_reset():
self.state = 'HALF_OPEN'
else:
return self.fallback_to_legacy(request)
try:
response = self.new_service.process(request)
self.on_success()
return response
except Exception as e:
self.on_failure()
return self.fallback_to_legacy(request)
```
## Data Validation and Reconciliation
### Validation Strategies
1. **Row Count Validation**
- Compare record counts between source and target
- Account for soft deletes and filtered records
- Implement threshold-based alerting
2. **Checksums and Hashing**
- Generate checksums for critical data subsets
- Compare hash values to detect data drift
- Use sampling for large datasets
3. **Business Logic Validation**
- Run critical business queries on both systems
- Compare aggregate results (sums, counts, averages)
- Validate derived data and calculations
### Reconciliation Patterns
1. **Delta Detection**
```sql
-- Example delta query for reconciliation
SELECT 'missing_in_target' as issue_type, source_id
FROM source_table s
WHERE NOT EXISTS (
SELECT 1 FROM target_table t
WHERE t.id = s.id
)
UNION ALL
SELECT 'extra_in_target' as issue_type, target_id
FROM target_table t
WHERE NOT EXISTS (
SELECT 1 FROM source_table s
WHERE s.id = t.id
);
```
2. **Automated Correction**
- Implement data repair scripts for common issues
- Use idempotent operations for safe re-execution
- Log all correction actions for audit trails
## Rollback Strategies
### Database Rollback
1. **Schema Rollback**
- Maintain schema version control
- Use backward-compatible migrations when possible
- Keep rollback scripts for each migration step
2. **Data Rollback**
- Point-in-time recovery using database backups
- Transaction log replay for precise rollback points
- Maintain data snapshots at migration checkpoints
### Service Rollback
1. **Blue-Green Deployment**
- Keep previous service version running during migration
- Switch traffic back to blue environment if issues arise
- Maintain parallel infrastructure during migration window
2. **Rolling Rollback**
- Gradually shift traffic back to previous version
- Monitor system health during rollback process
- Implement automated rollback triggers
### Infrastructure Rollback
1. **Infrastructure as Code**
- Version control all infrastructure definitions
- Maintain rollback terraform/CloudFormation templates
- Test rollback procedures in staging environments
2. **Data Persistence**
- Preserve data in original location during migration
- Implement data sync back to original systems
- Maintain backup strategies across both environments
## Risk Assessment Framework
### Risk Categories
1. **Technical Risks**
- Data loss or corruption
- Service downtime or degraded performance
- Integration failures with dependent systems
- Scalability issues under production load
2. **Business Risks**
- Revenue impact from service disruption
- Customer experience degradation
- Compliance and regulatory concerns
- Brand reputation impact
3. **Operational Risks**
- Team knowledge gaps
- Insufficient testing coverage
- Inadequate monitoring and alerting
- Communication breakdowns
### Risk Mitigation Strategies
1. **Technical Mitigations**
- Comprehensive testing (unit, integration, load, chaos)
- Gradual rollout with automated rollback triggers
- Data validation and reconciliation processes
- Performance monitoring and alerting
2. **Business Mitigations**
- Stakeholder communication plans
- Business continuity procedures
- Customer notification strategies
- Revenue protection measures
3. **Operational Mitigations**
- Team training and documentation
- Runbook creation and testing
- On-call rotation planning
- Post-migration review processes
## Migration Runbooks
### Pre-Migration Checklist
- [ ] Migration plan reviewed and approved
- [ ] Rollback procedures tested and validated
- [ ] Monitoring and alerting configured
- [ ] Team roles and responsibilities defined
- [ ] Stakeholder communication plan activated
- [ ] Backup and recovery procedures verified
- [ ] Test environment validation complete
- [ ] Performance benchmarks established
- [ ] Security review completed
- [ ] Compliance requirements verified
### During Migration
- [ ] Execute migration phases in planned order
- [ ] Monitor key performance indicators continuously
- [ ] Validate data consistency at each checkpoint
- [ ] Communicate progress to stakeholders
- [ ] Document any deviations from plan
- [ ] Execute rollback if success criteria not met
- [ ] Coordinate with dependent teams
- [ ] Maintain detailed execution logs
### Post-Migration
- [ ] Validate all success criteria met
- [ ] Perform comprehensive system health checks
- [ ] Execute data reconciliation procedures
- [ ] Monitor system performance over 72 hours
- [ ] Update documentation and runbooks
- [ ] Decommission legacy systems (if applicable)
- [ ] Conduct post-migration retrospective
- [ ] Archive migration artifacts
- [ ] Update disaster recovery procedures
## Communication Templates
### Executive Summary Template
```
Migration Status: [IN_PROGRESS | COMPLETED | ROLLED_BACK]
Start Time: [YYYY-MM-DD HH:MM UTC]
Current Phase: [X of Y]
Overall Progress: [X%]
Key Metrics:
- System Availability: [X.XX%]
- Data Migration Progress: [X.XX%]
- Performance Impact: [+/-X%]
- Issues Encountered: [X]
Next Steps:
1. [Action item 1]
2. [Action item 2]
Risk Assessment: [LOW | MEDIUM | HIGH]
Rollback Status: [AVAILABLE | NOT_AVAILABLE]
```
### Technical Team Update Template
```
Phase: [Phase Name] - [Status]
Duration: [Started] - [Expected End]
Completed Tasks:
✓ [Task 1]
✓ [Task 2]
In Progress:
🔄 [Task 3] - [X% complete]
Upcoming:
⏳ [Task 4] - [Expected start time]
Issues:
⚠️ [Issue description] - [Severity] - [ETA resolution]
Metrics:
- Migration Rate: [X records/minute]
- Error Rate: [X.XX%]
- System Load: [CPU/Memory/Disk]
```
## Success Metrics
### Technical Metrics
- **Migration Completion Rate:** Percentage of data/services successfully migrated
- **Downtime Duration:** Total system unavailability during migration
- **Data Consistency Score:** Percentage of data validation checks passing
- **Performance Delta:** Performance change compared to baseline
- **Error Rate:** Percentage of failed operations during migration
### Business Metrics
- **Customer Impact Score:** Measure of customer experience degradation
- **Revenue Protection:** Percentage of revenue maintained during migration
- **Time to Value:** Duration from migration start to business value realization
- **Stakeholder Satisfaction:** Post-migration stakeholder feedback scores
### Operational Metrics
- **Plan Adherence:** Percentage of migration executed according to plan
- **Issue Resolution Time:** Average time to resolve migration issues
- **Team Efficiency:** Resource utilization and productivity metrics
- **Knowledge Transfer Score:** Team readiness for post-migration operations
## Tools and Technologies
### Migration Planning Tools
- **migration_planner.py:** Automated migration plan generation
- **compatibility_checker.py:** Schema and API compatibility analysis
- **rollback_generator.py:** Comprehensive rollback procedure generation
### Validation Tools
- Database comparison utilities (schema and data)
- API contract testing frameworks
- Performance benchmarking tools
- Data quality validation pipelines
### Monitoring and Alerting
- Real-time migration progress dashboards
- Automated rollback trigger systems
- Business metric monitoring
- Stakeholder notification systems
## Best Practices
### Planning Phase
1. **Start with Risk Assessment:** Identify all potential failure modes before planning
2. **Design for Rollback:** Every migration step should have a tested rollback procedure
3. **Validate in Staging:** Execute full migration process in production-like environment
4. **Plan for Gradual Rollout:** Use feature flags and traffic routing for controlled migration
### Execution Phase
1. **Monitor Continuously:** Track both technical and business metrics throughout
2. **Communicate Proactively:** Keep all stakeholders informed of progress and issues
3. **Document Everything:** Maintain detailed logs for post-migration analysis
4. **Stay Flexible:** Be prepared to adjust timeline based on real-world performance
### Validation Phase
1. **Automate Validation:** Use automated tools for data consistency and performance checks
2. **Business Logic Testing:** Validate critical business processes end-to-end
3. **Load Testing:** Verify system performance under expected production load
4. **Security Validation:** Ensure security controls function properly in new environment
## Integration with Development Lifecycle
### CI/CD Integration
```yaml
# Example migration pipeline stage
migration_validation:
stage: test
script:
- python scripts/compatibility_checker.py --before=old_schema.json --after=new_schema.json
- python scripts/migration_planner.py --config=migration_config.json --validate
artifacts:
reports:
- compatibility_report.json
- migration_plan.json
```
### Infrastructure as Code
```terraform
# Example Terraform for blue-green infrastructure
resource "aws_instance" "blue_environment" {
count = var.migration_phase == "preparation" ? var.instance_count : 0
# Blue environment configuration
}
resource "aws_instance" "green_environment" {
count = var.migration_phase == "execution" ? var.instance_count : 0
# Green environment configuration
}
```
This Migration Architect skill provides a comprehensive framework for planning, executing, and validating complex system migrations while minimizing business impact and technical risk. The combination of automated tools, proven patterns, and detailed procedures enables organizations to confidently undertake even the most complex migration projects.
FILE:assets/database_schema_after.json
{
"schema_version": "2.0",
"database": "user_management_v2",
"tables": {
"users": {
"columns": {
"id": {
"type": "bigint",
"nullable": false,
"primary_key": true,
"auto_increment": true
},
"username": {
"type": "varchar",
"length": 50,
"nullable": false,
"unique": true
},
"email": {
"type": "varchar",
"length": 320,
"nullable": false,
"unique": true
},
"password_hash": {
"type": "varchar",
"length": 255,
"nullable": false
},
"first_name": {
"type": "varchar",
"length": 100,
"nullable": true
},
"last_name": {
"type": "varchar",
"length": 100,
"nullable": true
},
"created_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP"
},
"updated_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP"
},
"is_active": {
"type": "boolean",
"nullable": false,
"default": true
},
"phone": {
"type": "varchar",
"length": 20,
"nullable": true
},
"email_verified_at": {
"type": "timestamp",
"nullable": true,
"comment": "When email was verified"
},
"phone_verified_at": {
"type": "timestamp",
"nullable": true,
"comment": "When phone was verified"
},
"two_factor_enabled": {
"type": "boolean",
"nullable": false,
"default": false
},
"last_login_at": {
"type": "timestamp",
"nullable": true
}
},
"constraints": {
"primary_key": ["id"],
"unique": [
"username",
"email"
],
"foreign_key": [],
"check": [
"email LIKE '%@%'",
"LENGTH(password_hash) >= 60",
"phone IS NULL OR LENGTH(phone) >= 10"
]
},
"indexes": [
{
"name": "idx_users_email",
"columns": ["email"],
"unique": true
},
{
"name": "idx_users_username",
"columns": ["username"],
"unique": true
},
{
"name": "idx_users_created_at",
"columns": ["created_at"]
},
{
"name": "idx_users_email_verified",
"columns": ["email_verified_at"]
},
{
"name": "idx_users_last_login",
"columns": ["last_login_at"]
}
]
},
"user_profiles": {
"columns": {
"id": {
"type": "bigint",
"nullable": false,
"primary_key": true,
"auto_increment": true
},
"user_id": {
"type": "bigint",
"nullable": false
},
"bio": {
"type": "text",
"nullable": true
},
"avatar_url": {
"type": "varchar",
"length": 500,
"nullable": true
},
"birth_date": {
"type": "date",
"nullable": true
},
"location": {
"type": "varchar",
"length": 100,
"nullable": true
},
"website": {
"type": "varchar",
"length": 255,
"nullable": true
},
"privacy_level": {
"type": "varchar",
"length": 20,
"nullable": false,
"default": "public"
},
"timezone": {
"type": "varchar",
"length": 50,
"nullable": true,
"default": "UTC"
},
"language": {
"type": "varchar",
"length": 10,
"nullable": false,
"default": "en"
},
"created_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP"
},
"updated_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP"
}
},
"constraints": {
"primary_key": ["id"],
"unique": [],
"foreign_key": [
{
"columns": ["user_id"],
"references": "users(id)",
"on_delete": "CASCADE"
}
],
"check": [
"privacy_level IN ('public', 'private', 'friends_only')",
"bio IS NULL OR LENGTH(bio) <= 2000",
"language IN ('en', 'es', 'fr', 'de', 'it', 'pt', 'ru', 'ja', 'ko', 'zh')"
]
},
"indexes": [
{
"name": "idx_user_profiles_user_id",
"columns": ["user_id"],
"unique": true
},
{
"name": "idx_user_profiles_privacy",
"columns": ["privacy_level"]
},
{
"name": "idx_user_profiles_language",
"columns": ["language"]
}
]
},
"user_sessions": {
"columns": {
"id": {
"type": "varchar",
"length": 128,
"nullable": false,
"primary_key": true
},
"user_id": {
"type": "bigint",
"nullable": false
},
"ip_address": {
"type": "varchar",
"length": 45,
"nullable": true
},
"user_agent": {
"type": "text",
"nullable": true
},
"expires_at": {
"type": "timestamp",
"nullable": false
},
"created_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP"
},
"last_activity": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP"
},
"session_type": {
"type": "varchar",
"length": 20,
"nullable": false,
"default": "web"
},
"is_mobile": {
"type": "boolean",
"nullable": false,
"default": false
}
},
"constraints": {
"primary_key": ["id"],
"unique": [],
"foreign_key": [
{
"columns": ["user_id"],
"references": "users(id)",
"on_delete": "CASCADE"
}
],
"check": [
"session_type IN ('web', 'mobile', 'api', 'admin')"
]
},
"indexes": [
{
"name": "idx_user_sessions_user_id",
"columns": ["user_id"]
},
{
"name": "idx_user_sessions_expires",
"columns": ["expires_at"]
},
{
"name": "idx_user_sessions_type",
"columns": ["session_type"]
}
]
},
"user_preferences": {
"columns": {
"id": {
"type": "bigint",
"nullable": false,
"primary_key": true,
"auto_increment": true
},
"user_id": {
"type": "bigint",
"nullable": false
},
"preference_key": {
"type": "varchar",
"length": 100,
"nullable": false
},
"preference_value": {
"type": "json",
"nullable": true
},
"created_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP"
},
"updated_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP"
}
},
"constraints": {
"primary_key": ["id"],
"unique": [
["user_id", "preference_key"]
],
"foreign_key": [
{
"columns": ["user_id"],
"references": "users(id)",
"on_delete": "CASCADE"
}
],
"check": []
},
"indexes": [
{
"name": "idx_user_preferences_user_key",
"columns": ["user_id", "preference_key"],
"unique": true
}
]
}
},
"views": {
"active_users": {
"definition": "SELECT u.id, u.username, u.email, u.first_name, u.last_name, u.email_verified_at, u.last_login_at FROM users u WHERE u.is_active = true",
"columns": ["id", "username", "email", "first_name", "last_name", "email_verified_at", "last_login_at"]
},
"verified_users": {
"definition": "SELECT u.id, u.username, u.email FROM users u WHERE u.is_active = true AND u.email_verified_at IS NOT NULL",
"columns": ["id", "username", "email"]
}
},
"procedures": [
{
"name": "cleanup_expired_sessions",
"parameters": [],
"definition": "DELETE FROM user_sessions WHERE expires_at < NOW()"
},
{
"name": "get_user_with_profile",
"parameters": ["user_id BIGINT"],
"definition": "SELECT u.*, p.bio, p.avatar_url, p.privacy_level FROM users u LEFT JOIN user_profiles p ON u.id = p.user_id WHERE u.id = user_id"
}
]
}
FILE:assets/database_schema_before.json
{
"schema_version": "1.0",
"database": "user_management",
"tables": {
"users": {
"columns": {
"id": {
"type": "bigint",
"nullable": false,
"primary_key": true,
"auto_increment": true
},
"username": {
"type": "varchar",
"length": 50,
"nullable": false,
"unique": true
},
"email": {
"type": "varchar",
"length": 255,
"nullable": false,
"unique": true
},
"password_hash": {
"type": "varchar",
"length": 255,
"nullable": false
},
"first_name": {
"type": "varchar",
"length": 100,
"nullable": true
},
"last_name": {
"type": "varchar",
"length": 100,
"nullable": true
},
"created_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP"
},
"updated_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP"
},
"is_active": {
"type": "boolean",
"nullable": false,
"default": true
},
"phone": {
"type": "varchar",
"length": 20,
"nullable": true
}
},
"constraints": {
"primary_key": ["id"],
"unique": [
"username",
"email"
],
"foreign_key": [],
"check": [
"email LIKE '%@%'",
"LENGTH(password_hash) >= 60"
]
},
"indexes": [
{
"name": "idx_users_email",
"columns": ["email"],
"unique": true
},
{
"name": "idx_users_username",
"columns": ["username"],
"unique": true
},
{
"name": "idx_users_created_at",
"columns": ["created_at"]
}
]
},
"user_profiles": {
"columns": {
"id": {
"type": "bigint",
"nullable": false,
"primary_key": true,
"auto_increment": true
},
"user_id": {
"type": "bigint",
"nullable": false
},
"bio": {
"type": "varchar",
"length": 255,
"nullable": true
},
"avatar_url": {
"type": "varchar",
"length": 500,
"nullable": true
},
"birth_date": {
"type": "date",
"nullable": true
},
"location": {
"type": "varchar",
"length": 100,
"nullable": true
},
"website": {
"type": "varchar",
"length": 255,
"nullable": true
},
"privacy_level": {
"type": "varchar",
"length": 20,
"nullable": false,
"default": "public"
},
"created_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP"
},
"updated_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP"
}
},
"constraints": {
"primary_key": ["id"],
"unique": [],
"foreign_key": [
{
"columns": ["user_id"],
"references": "users(id)",
"on_delete": "CASCADE"
}
],
"check": [
"privacy_level IN ('public', 'private', 'friends_only')"
]
},
"indexes": [
{
"name": "idx_user_profiles_user_id",
"columns": ["user_id"],
"unique": true
},
{
"name": "idx_user_profiles_privacy",
"columns": ["privacy_level"]
}
]
},
"user_sessions": {
"columns": {
"id": {
"type": "varchar",
"length": 128,
"nullable": false,
"primary_key": true
},
"user_id": {
"type": "bigint",
"nullable": false
},
"ip_address": {
"type": "varchar",
"length": 45,
"nullable": true
},
"user_agent": {
"type": "varchar",
"length": 500,
"nullable": true
},
"expires_at": {
"type": "timestamp",
"nullable": false
},
"created_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP"
},
"last_activity": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP"
}
},
"constraints": {
"primary_key": ["id"],
"unique": [],
"foreign_key": [
{
"columns": ["user_id"],
"references": "users(id)",
"on_delete": "CASCADE"
}
],
"check": []
},
"indexes": [
{
"name": "idx_user_sessions_user_id",
"columns": ["user_id"]
},
{
"name": "idx_user_sessions_expires",
"columns": ["expires_at"]
}
]
}
},
"views": {
"active_users": {
"definition": "SELECT u.id, u.username, u.email, u.first_name, u.last_name FROM users u WHERE u.is_active = true",
"columns": ["id", "username", "email", "first_name", "last_name"]
}
},
"procedures": [
{
"name": "cleanup_expired_sessions",
"parameters": [],
"definition": "DELETE FROM user_sessions WHERE expires_at < NOW()"
}
]
}
FILE:assets/sample_database_migration.json
{
"type": "database",
"pattern": "schema_change",
"source": "PostgreSQL 13 Production Database",
"target": "PostgreSQL 15 Cloud Database",
"description": "Migrate user management system from on-premises PostgreSQL to cloud with schema updates",
"constraints": {
"max_downtime_minutes": 30,
"data_volume_gb": 2500,
"dependencies": [
"user_service_api",
"authentication_service",
"notification_service",
"analytics_pipeline",
"backup_service"
],
"compliance_requirements": [
"GDPR",
"SOX"
],
"special_requirements": [
"zero_data_loss",
"referential_integrity",
"performance_baseline_maintained"
]
},
"tables_to_migrate": [
{
"name": "users",
"row_count": 1500000,
"size_mb": 450,
"critical": true
},
{
"name": "user_profiles",
"row_count": 1500000,
"size_mb": 890,
"critical": true
},
{
"name": "user_sessions",
"row_count": 25000000,
"size_mb": 1200,
"critical": false
},
{
"name": "audit_logs",
"row_count": 50000000,
"size_mb": 2800,
"critical": false
}
],
"schema_changes": [
{
"table": "users",
"changes": [
{
"type": "add_column",
"column": "email_verified_at",
"data_type": "timestamp",
"nullable": true
},
{
"type": "add_column",
"column": "phone_verified_at",
"data_type": "timestamp",
"nullable": true
}
]
},
{
"table": "user_profiles",
"changes": [
{
"type": "modify_column",
"column": "bio",
"old_type": "varchar(255)",
"new_type": "text"
},
{
"type": "add_constraint",
"constraint_type": "check",
"constraint_name": "bio_length_check",
"definition": "LENGTH(bio) <= 2000"
}
]
}
],
"performance_requirements": {
"max_query_response_time_ms": 100,
"concurrent_connections": 500,
"transactions_per_second": 1000
},
"business_continuity": {
"critical_business_hours": {
"start": "08:00",
"end": "18:00",
"timezone": "UTC"
},
"preferred_migration_window": {
"start": "02:00",
"end": "06:00",
"timezone": "UTC"
}
}
}
FILE:assets/sample_service_migration.json
{
"type": "service",
"pattern": "strangler_fig",
"source": "Legacy User Service (Java Spring Boot 2.x)",
"target": "New User Service (Node.js + TypeScript)",
"description": "Migrate legacy user management service to modern microservices architecture",
"constraints": {
"max_downtime_minutes": 0,
"data_volume_gb": 50,
"dependencies": [
"payment_service",
"order_service",
"notification_service",
"analytics_service",
"mobile_app_v1",
"mobile_app_v2",
"web_frontend",
"admin_dashboard"
],
"compliance_requirements": [
"PCI_DSS",
"GDPR"
],
"special_requirements": [
"api_backward_compatibility",
"session_continuity",
"rate_limit_preservation"
]
},
"service_details": {
"legacy_service": {
"endpoints": [
"GET /api/v1/users/{id}",
"POST /api/v1/users",
"PUT /api/v1/users/{id}",
"DELETE /api/v1/users/{id}",
"GET /api/v1/users/{id}/profile",
"PUT /api/v1/users/{id}/profile",
"POST /api/v1/users/{id}/verify-email",
"POST /api/v1/users/login",
"POST /api/v1/users/logout"
],
"current_load": {
"requests_per_second": 850,
"peak_requests_per_second": 2000,
"average_response_time_ms": 120,
"p95_response_time_ms": 300
},
"infrastructure": {
"instances": 4,
"cpu_cores_per_instance": 4,
"memory_gb_per_instance": 8,
"load_balancer": "AWS ELB Classic"
}
},
"new_service": {
"endpoints": [
"GET /api/v2/users/{id}",
"POST /api/v2/users",
"PUT /api/v2/users/{id}",
"DELETE /api/v2/users/{id}",
"GET /api/v2/users/{id}/profile",
"PUT /api/v2/users/{id}/profile",
"POST /api/v2/users/{id}/verify-email",
"POST /api/v2/users/{id}/verify-phone",
"POST /api/v2/auth/login",
"POST /api/v2/auth/logout",
"POST /api/v2/auth/refresh"
],
"target_performance": {
"requests_per_second": 1500,
"peak_requests_per_second": 3000,
"average_response_time_ms": 80,
"p95_response_time_ms": 200
},
"infrastructure": {
"container_platform": "Kubernetes",
"initial_replicas": 3,
"max_replicas": 10,
"cpu_request_millicores": 500,
"cpu_limit_millicores": 1000,
"memory_request_mb": 512,
"memory_limit_mb": 1024,
"load_balancer": "AWS ALB"
}
}
},
"migration_phases": [
{
"phase": "preparation",
"description": "Deploy new service and configure routing",
"estimated_duration_hours": 8
},
{
"phase": "intercept",
"description": "Configure API gateway to route to new service",
"estimated_duration_hours": 2
},
{
"phase": "gradual_migration",
"description": "Gradually increase traffic to new service",
"estimated_duration_hours": 48
},
{
"phase": "validation",
"description": "Validate new service performance and functionality",
"estimated_duration_hours": 24
},
{
"phase": "decommission",
"description": "Remove legacy service after validation",
"estimated_duration_hours": 4
}
],
"feature_flags": [
{
"name": "enable_new_user_service",
"description": "Route user service requests to new implementation",
"initial_percentage": 5,
"rollout_schedule": [
{"percentage": 5, "duration_hours": 24},
{"percentage": 25, "duration_hours": 24},
{"percentage": 50, "duration_hours": 24},
{"percentage": 100, "duration_hours": 0}
]
},
{
"name": "enable_new_auth_endpoints",
"description": "Enable new authentication endpoints",
"initial_percentage": 0,
"rollout_schedule": [
{"percentage": 10, "duration_hours": 12},
{"percentage": 50, "duration_hours": 12},
{"percentage": 100, "duration_hours": 0}
]
}
],
"monitoring": {
"critical_metrics": [
"request_rate",
"error_rate",
"response_time_p95",
"response_time_p99",
"cpu_utilization",
"memory_utilization",
"database_connection_pool"
],
"alert_thresholds": {
"error_rate": 0.05,
"response_time_p95": 250,
"cpu_utilization": 0.80,
"memory_utilization": 0.85
}
},
"rollback_triggers": [
{
"metric": "error_rate",
"threshold": 0.10,
"duration_minutes": 5,
"action": "automatic_rollback"
},
{
"metric": "response_time_p95",
"threshold": 500,
"duration_minutes": 10,
"action": "alert_team"
},
{
"metric": "cpu_utilization",
"threshold": 0.95,
"duration_minutes": 5,
"action": "scale_up"
}
]
}
FILE:expected_outputs/rollback_runbook.json
{
"runbook_id": "rb_921c0bca",
"migration_id": "23a52ed1507f",
"created_at": "2026-02-16T13:47:31.108500",
"rollback_phases": [
{
"phase_name": "rollback_cleanup",
"description": "Rollback changes made during cleanup phase",
"urgency_level": "medium",
"estimated_duration_minutes": 570,
"prerequisites": [
"Incident commander assigned and briefed",
"All team members notified of rollback initiation",
"Monitoring systems confirmed operational",
"Backup systems verified and accessible"
],
"steps": [
{
"step_id": "rb_validate_0_final",
"name": "Validate rollback completion",
"description": "Comprehensive validation that cleanup rollback completed successfully",
"script_type": "manual",
"script_content": "Execute validation checklist for this phase",
"estimated_duration_minutes": 10,
"dependencies": [],
"validation_commands": [
"SELECT COUNT(*) FROM {table_name};",
"SELECT COUNT(*) FROM information_schema.tables WHERE table_name = '{table_name}';",
"SELECT COUNT(*) FROM information_schema.columns WHERE table_name = '{table_name}' AND column_name = '{column_name}';",
"SELECT COUNT(DISTINCT {primary_key}) FROM {table_name};",
"SELECT MAX({timestamp_column}) FROM {table_name};"
],
"success_criteria": [
"cleanup fully rolled back",
"All validation checks pass"
],
"failure_escalation": "Investigate cleanup rollback failures",
"rollback_order": 99
}
],
"validation_checkpoints": [
"cleanup rollback steps completed",
"System health checks passing",
"No critical errors in logs",
"Key metrics within acceptable ranges",
"Validation command passed: SELECT COUNT(*) FROM {table_name};...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH..."
],
"communication_requirements": [
"Notify incident commander of phase start/completion",
"Update rollback status dashboard",
"Log all actions and decisions"
],
"risk_level": "medium"
},
{
"phase_name": "rollback_contract",
"description": "Rollback changes made during contract phase",
"urgency_level": "medium",
"estimated_duration_minutes": 570,
"prerequisites": [
"Incident commander assigned and briefed",
"All team members notified of rollback initiation",
"Monitoring systems confirmed operational",
"Backup systems verified and accessible",
"Previous rollback phase completed successfully"
],
"steps": [
{
"step_id": "rb_validate_1_final",
"name": "Validate rollback completion",
"description": "Comprehensive validation that contract rollback completed successfully",
"script_type": "manual",
"script_content": "Execute validation checklist for this phase",
"estimated_duration_minutes": 10,
"dependencies": [],
"validation_commands": [
"SELECT COUNT(*) FROM {table_name};",
"SELECT COUNT(*) FROM information_schema.tables WHERE table_name = '{table_name}';",
"SELECT COUNT(*) FROM information_schema.columns WHERE table_name = '{table_name}' AND column_name = '{column_name}';",
"SELECT COUNT(DISTINCT {primary_key}) FROM {table_name};",
"SELECT MAX({timestamp_column}) FROM {table_name};"
],
"success_criteria": [
"contract fully rolled back",
"All validation checks pass"
],
"failure_escalation": "Investigate contract rollback failures",
"rollback_order": 99
}
],
"validation_checkpoints": [
"contract rollback steps completed",
"System health checks passing",
"No critical errors in logs",
"Key metrics within acceptable ranges",
"Validation command passed: SELECT COUNT(*) FROM {table_name};...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH..."
],
"communication_requirements": [
"Notify incident commander of phase start/completion",
"Update rollback status dashboard",
"Log all actions and decisions"
],
"risk_level": "medium"
},
{
"phase_name": "rollback_migrate",
"description": "Rollback changes made during migrate phase",
"urgency_level": "medium",
"estimated_duration_minutes": 570,
"prerequisites": [
"Incident commander assigned and briefed",
"All team members notified of rollback initiation",
"Monitoring systems confirmed operational",
"Backup systems verified and accessible",
"Previous rollback phase completed successfully"
],
"steps": [
{
"step_id": "rb_validate_2_final",
"name": "Validate rollback completion",
"description": "Comprehensive validation that migrate rollback completed successfully",
"script_type": "manual",
"script_content": "Execute validation checklist for this phase",
"estimated_duration_minutes": 10,
"dependencies": [],
"validation_commands": [
"SELECT COUNT(*) FROM {table_name};",
"SELECT COUNT(*) FROM information_schema.tables WHERE table_name = '{table_name}';",
"SELECT COUNT(*) FROM information_schema.columns WHERE table_name = '{table_name}' AND column_name = '{column_name}';",
"SELECT COUNT(DISTINCT {primary_key}) FROM {table_name};",
"SELECT MAX({timestamp_column}) FROM {table_name};"
],
"success_criteria": [
"migrate fully rolled back",
"All validation checks pass"
],
"failure_escalation": "Investigate migrate rollback failures",
"rollback_order": 99
}
],
"validation_checkpoints": [
"migrate rollback steps completed",
"System health checks passing",
"No critical errors in logs",
"Key metrics within acceptable ranges",
"Validation command passed: SELECT COUNT(*) FROM {table_name};...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH..."
],
"communication_requirements": [
"Notify incident commander of phase start/completion",
"Update rollback status dashboard",
"Log all actions and decisions"
],
"risk_level": "medium"
},
{
"phase_name": "rollback_expand",
"description": "Rollback changes made during expand phase",
"urgency_level": "medium",
"estimated_duration_minutes": 570,
"prerequisites": [
"Incident commander assigned and briefed",
"All team members notified of rollback initiation",
"Monitoring systems confirmed operational",
"Backup systems verified and accessible",
"Previous rollback phase completed successfully"
],
"steps": [
{
"step_id": "rb_validate_3_final",
"name": "Validate rollback completion",
"description": "Comprehensive validation that expand rollback completed successfully",
"script_type": "manual",
"script_content": "Execute validation checklist for this phase",
"estimated_duration_minutes": 10,
"dependencies": [],
"validation_commands": [
"SELECT COUNT(*) FROM {table_name};",
"SELECT COUNT(*) FROM information_schema.tables WHERE table_name = '{table_name}';",
"SELECT COUNT(*) FROM information_schema.columns WHERE table_name = '{table_name}' AND column_name = '{column_name}';",
"SELECT COUNT(DISTINCT {primary_key}) FROM {table_name};",
"SELECT MAX({timestamp_column}) FROM {table_name};"
],
"success_criteria": [
"expand fully rolled back",
"All validation checks pass"
],
"failure_escalation": "Investigate expand rollback failures",
"rollback_order": 99
}
],
"validation_checkpoints": [
"expand rollback steps completed",
"System health checks passing",
"No critical errors in logs",
"Key metrics within acceptable ranges",
"Validation command passed: SELECT COUNT(*) FROM {table_name};...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH..."
],
"communication_requirements": [
"Notify incident commander of phase start/completion",
"Update rollback status dashboard",
"Log all actions and decisions"
],
"risk_level": "medium"
},
{
"phase_name": "rollback_preparation",
"description": "Rollback changes made during preparation phase",
"urgency_level": "medium",
"estimated_duration_minutes": 570,
"prerequisites": [
"Incident commander assigned and briefed",
"All team members notified of rollback initiation",
"Monitoring systems confirmed operational",
"Backup systems verified and accessible",
"Previous rollback phase completed successfully"
],
"steps": [
{
"step_id": "rb_schema_4_01",
"name": "Drop migration artifacts",
"description": "Remove temporary migration tables and procedures",
"script_type": "sql",
"script_content": "-- Drop migration artifacts\nDROP TABLE IF EXISTS migration_log;\nDROP PROCEDURE IF EXISTS migrate_data();",
"estimated_duration_minutes": 5,
"dependencies": [],
"validation_commands": [
"SELECT COUNT(*) FROM information_schema.tables WHERE table_name LIKE '%migration%';"
],
"success_criteria": [
"No migration artifacts remain"
],
"failure_escalation": "Manual cleanup required",
"rollback_order": 1
},
{
"step_id": "rb_validate_4_final",
"name": "Validate rollback completion",
"description": "Comprehensive validation that preparation rollback completed successfully",
"script_type": "manual",
"script_content": "Execute validation checklist for this phase",
"estimated_duration_minutes": 10,
"dependencies": [
"rb_schema_4_01"
],
"validation_commands": [
"SELECT COUNT(*) FROM {table_name};",
"SELECT COUNT(*) FROM information_schema.tables WHERE table_name = '{table_name}';",
"SELECT COUNT(*) FROM information_schema.columns WHERE table_name = '{table_name}' AND column_name = '{column_name}';",
"SELECT COUNT(DISTINCT {primary_key}) FROM {table_name};",
"SELECT MAX({timestamp_column}) FROM {table_name};"
],
"success_criteria": [
"preparation fully rolled back",
"All validation checks pass"
],
"failure_escalation": "Investigate preparation rollback failures",
"rollback_order": 99
}
],
"validation_checkpoints": [
"preparation rollback steps completed",
"System health checks passing",
"No critical errors in logs",
"Key metrics within acceptable ranges",
"Validation command passed: SELECT COUNT(*) FROM {table_name};...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH..."
],
"communication_requirements": [
"Notify incident commander of phase start/completion",
"Update rollback status dashboard",
"Log all actions and decisions"
],
"risk_level": "medium"
}
],
"trigger_conditions": [
{
"trigger_id": "error_rate_spike",
"name": "Error Rate Spike",
"condition": "error_rate > baseline * 5 for 5 minutes",
"metric_threshold": {
"metric": "error_rate",
"operator": "greater_than",
"value": "baseline_error_rate * 5",
"duration_minutes": 5
},
"evaluation_window_minutes": 5,
"auto_execute": true,
"escalation_contacts": [
"on_call_engineer",
"migration_lead"
]
},
{
"trigger_id": "response_time_degradation",
"name": "Response Time Degradation",
"condition": "p95_response_time > baseline * 3 for 10 minutes",
"metric_threshold": {
"metric": "p95_response_time",
"operator": "greater_than",
"value": "baseline_p95 * 3",
"duration_minutes": 10
},
"evaluation_window_minutes": 10,
"auto_execute": false,
"escalation_contacts": [
"performance_team",
"migration_lead"
]
},
{
"trigger_id": "availability_drop",
"name": "Service Availability Drop",
"condition": "availability < 95% for 2 minutes",
"metric_threshold": {
"metric": "availability",
"operator": "less_than",
"value": 0.95,
"duration_minutes": 2
},
"evaluation_window_minutes": 2,
"auto_execute": true,
"escalation_contacts": [
"sre_team",
"incident_commander"
]
},
{
"trigger_id": "data_integrity_failure",
"name": "Data Integrity Check Failure",
"condition": "data_validation_failures > 0",
"metric_threshold": {
"metric": "data_validation_failures",
"operator": "greater_than",
"value": 0,
"duration_minutes": 1
},
"evaluation_window_minutes": 1,
"auto_execute": true,
"escalation_contacts": [
"dba_team",
"data_team"
]
},
{
"trigger_id": "migration_progress_stalled",
"name": "Migration Progress Stalled",
"condition": "migration_progress unchanged for 30 minutes",
"metric_threshold": {
"metric": "migration_progress_rate",
"operator": "equals",
"value": 0,
"duration_minutes": 30
},
"evaluation_window_minutes": 30,
"auto_execute": false,
"escalation_contacts": [
"migration_team",
"dba_team"
]
}
],
"data_recovery_plan": {
"recovery_method": "point_in_time",
"backup_location": "/backups/pre_migration_{migration_id}_{timestamp}.sql",
"recovery_scripts": [
"pg_restore -d production -c /backups/pre_migration_backup.sql",
"SELECT pg_create_restore_point('rollback_point');",
"VACUUM ANALYZE; -- Refresh statistics after restore"
],
"data_validation_queries": [
"SELECT COUNT(*) FROM critical_business_table;",
"SELECT MAX(created_at) FROM audit_log;",
"SELECT COUNT(DISTINCT user_id) FROM user_sessions;",
"SELECT SUM(amount) FROM financial_transactions WHERE date = CURRENT_DATE;"
],
"estimated_recovery_time_minutes": 45,
"recovery_dependencies": [
"database_instance_running",
"backup_file_accessible"
]
},
"communication_templates": [
{
"template_type": "rollback_start",
"audience": "technical",
"subject": "ROLLBACK INITIATED: {migration_name}",
"body": "Team,\n\nWe have initiated rollback for migration: {migration_name}\nRollback ID: {rollback_id}\nStart Time: {start_time}\nEstimated Duration: {estimated_duration}\n\nReason: {rollback_reason}\n\nCurrent Status: Rolling back phase {current_phase}\n\nNext Updates: Every 15 minutes or upon phase completion\n\nActions Required:\n- Monitor system health dashboards\n- Stand by for escalation if needed\n- Do not make manual changes during rollback\n\nIncident Commander: {incident_commander}\n",
"urgency": "medium",
"delivery_methods": [
"email",
"slack"
]
},
{
"template_type": "rollback_start",
"audience": "business",
"subject": "System Rollback In Progress - {system_name}",
"body": "Business Stakeholders,\n\nWe are currently performing a planned rollback of the {system_name} migration due to {rollback_reason}.\n\nImpact: {business_impact}\nExpected Resolution: {estimated_completion_time}\nAffected Services: {affected_services}\n\nWe will provide updates every 30 minutes.\n\nContact: {business_contact}\n",
"urgency": "medium",
"delivery_methods": [
"email"
]
},
{
"template_type": "rollback_start",
"audience": "executive",
"subject": "EXEC ALERT: Critical System Rollback - {system_name}",
"body": "Executive Team,\n\nA critical rollback is in progress for {system_name}.\n\nSummary:\n- Rollback Reason: {rollback_reason}\n- Business Impact: {business_impact}\n- Expected Resolution: {estimated_completion_time}\n- Customer Impact: {customer_impact}\n\nWe are following established procedures and will update hourly.\n\nEscalation: {escalation_contact}\n",
"urgency": "high",
"delivery_methods": [
"email"
]
},
{
"template_type": "rollback_complete",
"audience": "technical",
"subject": "ROLLBACK COMPLETED: {migration_name}",
"body": "Team,\n\nRollback has been successfully completed for migration: {migration_name}\n\nSummary:\n- Start Time: {start_time}\n- End Time: {end_time}\n- Duration: {actual_duration}\n- Phases Completed: {completed_phases}\n\nValidation Results:\n{validation_results}\n\nSystem Status: {system_status}\n\nNext Steps:\n- Continue monitoring for 24 hours\n- Post-rollback review scheduled for {review_date}\n- Root cause analysis to begin\n\nAll clear to resume normal operations.\n\nIncident Commander: {incident_commander}\n",
"urgency": "medium",
"delivery_methods": [
"email",
"slack"
]
},
{
"template_type": "emergency_escalation",
"audience": "executive",
"subject": "CRITICAL: Rollback Emergency - {migration_name}",
"body": "CRITICAL SITUATION - IMMEDIATE ATTENTION REQUIRED\n\nMigration: {migration_name}\nIssue: Rollback procedure has encountered critical failures\n\nCurrent Status: {current_status}\nFailed Components: {failed_components}\nBusiness Impact: {business_impact}\nCustomer Impact: {customer_impact}\n\nImmediate Actions:\n1. Emergency response team activated\n2. {emergency_action_1}\n3. {emergency_action_2}\n\nWar Room: {war_room_location}\nBridge Line: {conference_bridge}\n\nNext Update: {next_update_time}\n\nIncident Commander: {incident_commander}\nExecutive On-Call: {executive_on_call}\n",
"urgency": "emergency",
"delivery_methods": [
"email",
"sms",
"phone_call"
]
}
],
"escalation_matrix": {
"level_1": {
"trigger": "Single component failure",
"response_time_minutes": 5,
"contacts": [
"on_call_engineer",
"migration_lead"
],
"actions": [
"Investigate issue",
"Attempt automated remediation",
"Monitor closely"
]
},
"level_2": {
"trigger": "Multiple component failures or single critical failure",
"response_time_minutes": 2,
"contacts": [
"senior_engineer",
"team_lead",
"devops_lead"
],
"actions": [
"Initiate rollback",
"Establish war room",
"Notify stakeholders"
]
},
"level_3": {
"trigger": "System-wide failure or data corruption",
"response_time_minutes": 1,
"contacts": [
"engineering_manager",
"cto",
"incident_commander"
],
"actions": [
"Emergency rollback",
"All hands on deck",
"Executive notification"
]
},
"emergency": {
"trigger": "Business-critical failure with customer impact",
"response_time_minutes": 0,
"contacts": [
"ceo",
"cto",
"head_of_operations"
],
"actions": [
"Emergency procedures",
"Customer communication",
"Media preparation if needed"
]
}
},
"validation_checklist": [
"Verify system is responding to health checks",
"Confirm error rates are within normal parameters",
"Validate response times meet SLA requirements",
"Check all critical business processes are functioning",
"Verify monitoring and alerting systems are operational",
"Confirm no data corruption has occurred",
"Validate security controls are functioning properly",
"Check backup systems are working correctly",
"Verify integration points with downstream systems",
"Confirm user authentication and authorization working",
"Validate database schema matches expected state",
"Confirm referential integrity constraints",
"Check database performance metrics",
"Verify data consistency across related tables",
"Validate indexes and statistics are optimal",
"Confirm transaction logs are clean",
"Check database connections and connection pooling"
],
"post_rollback_procedures": [
"Monitor system stability for 24-48 hours post-rollback",
"Conduct thorough post-rollback testing of all critical paths",
"Review and analyze rollback metrics and timing",
"Document lessons learned and rollback procedure improvements",
"Schedule post-mortem meeting with all stakeholders",
"Update rollback procedures based on actual experience",
"Communicate rollback completion to all stakeholders",
"Archive rollback logs and artifacts for future reference",
"Review and update monitoring thresholds if needed",
"Plan for next migration attempt with improved procedures",
"Conduct security review to ensure no vulnerabilities introduced",
"Update disaster recovery procedures if affected by rollback",
"Review capacity planning based on rollback resource usage",
"Update documentation with rollback experience and timings"
],
"emergency_contacts": [
{
"role": "Incident Commander",
"name": "TBD - Assigned during migration",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "incident.commander@company.com",
"backup_contact": "backup.commander@company.com"
},
{
"role": "Technical Lead",
"name": "TBD - Migration technical owner",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "tech.lead@company.com",
"backup_contact": "senior.engineer@company.com"
},
{
"role": "Business Owner",
"name": "TBD - Business stakeholder",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "business.owner@company.com",
"backup_contact": "product.manager@company.com"
},
{
"role": "On-Call Engineer",
"name": "Current on-call rotation",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "oncall@company.com",
"backup_contact": "backup.oncall@company.com"
},
{
"role": "Executive Escalation",
"name": "CTO/VP Engineering",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "cto@company.com",
"backup_contact": "vp.engineering@company.com"
}
]
}
FILE:expected_outputs/rollback_runbook.txt
================================================================================
ROLLBACK RUNBOOK: rb_921c0bca
================================================================================
Migration ID: 23a52ed1507f
Created: 2026-02-16T13:47:31.108500
EMERGENCY CONTACTS
----------------------------------------
Incident Commander: TBD - Assigned during migration
Phone: +1-XXX-XXX-XXXX
Email: incident.commander@company.com
Backup: backup.commander@company.com
Technical Lead: TBD - Migration technical owner
Phone: +1-XXX-XXX-XXXX
Email: tech.lead@company.com
Backup: senior.engineer@company.com
Business Owner: TBD - Business stakeholder
Phone: +1-XXX-XXX-XXXX
Email: business.owner@company.com
Backup: product.manager@company.com
On-Call Engineer: Current on-call rotation
Phone: +1-XXX-XXX-XXXX
Email: oncall@company.com
Backup: backup.oncall@company.com
Executive Escalation: CTO/VP Engineering
Phone: +1-XXX-XXX-XXXX
Email: cto@company.com
Backup: vp.engineering@company.com
ESCALATION MATRIX
----------------------------------------
LEVEL_1:
Trigger: Single component failure
Response Time: 5 minutes
Contacts: on_call_engineer, migration_lead
Actions: Investigate issue, Attempt automated remediation, Monitor closely
LEVEL_2:
Trigger: Multiple component failures or single critical failure
Response Time: 2 minutes
Contacts: senior_engineer, team_lead, devops_lead
Actions: Initiate rollback, Establish war room, Notify stakeholders
LEVEL_3:
Trigger: System-wide failure or data corruption
Response Time: 1 minutes
Contacts: engineering_manager, cto, incident_commander
Actions: Emergency rollback, All hands on deck, Executive notification
EMERGENCY:
Trigger: Business-critical failure with customer impact
Response Time: 0 minutes
Contacts: ceo, cto, head_of_operations
Actions: Emergency procedures, Customer communication, Media preparation if needed
AUTOMATIC ROLLBACK TRIGGERS
----------------------------------------
• Error Rate Spike
Condition: error_rate > baseline * 5 for 5 minutes
Auto-Execute: Yes
Evaluation Window: 5 minutes
Contacts: on_call_engineer, migration_lead
• Response Time Degradation
Condition: p95_response_time > baseline * 3 for 10 minutes
Auto-Execute: No
Evaluation Window: 10 minutes
Contacts: performance_team, migration_lead
• Service Availability Drop
Condition: availability < 95% for 2 minutes
Auto-Execute: Yes
Evaluation Window: 2 minutes
Contacts: sre_team, incident_commander
• Data Integrity Check Failure
Condition: data_validation_failures > 0
Auto-Execute: Yes
Evaluation Window: 1 minutes
Contacts: dba_team, data_team
• Migration Progress Stalled
Condition: migration_progress unchanged for 30 minutes
Auto-Execute: No
Evaluation Window: 30 minutes
Contacts: migration_team, dba_team
ROLLBACK PHASES
----------------------------------------
1. ROLLBACK_CLEANUP
Description: Rollback changes made during cleanup phase
Urgency: MEDIUM
Duration: 570 minutes
Risk Level: MEDIUM
Prerequisites:
✓ Incident commander assigned and briefed
✓ All team members notified of rollback initiation
✓ Monitoring systems confirmed operational
✓ Backup systems verified and accessible
Steps:
99. Validate rollback completion
Duration: 10 min
Type: manual
Success Criteria: cleanup fully rolled back, All validation checks pass
Validation Checkpoints:
☐ cleanup rollback steps completed
☐ System health checks passing
☐ No critical errors in logs
☐ Key metrics within acceptable ranges
☐ Validation command passed: SELECT COUNT(*) FROM {table_name};...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH...
2. ROLLBACK_CONTRACT
Description: Rollback changes made during contract phase
Urgency: MEDIUM
Duration: 570 minutes
Risk Level: MEDIUM
Prerequisites:
✓ Incident commander assigned and briefed
✓ All team members notified of rollback initiation
✓ Monitoring systems confirmed operational
✓ Backup systems verified and accessible
✓ Previous rollback phase completed successfully
Steps:
99. Validate rollback completion
Duration: 10 min
Type: manual
Success Criteria: contract fully rolled back, All validation checks pass
Validation Checkpoints:
☐ contract rollback steps completed
☐ System health checks passing
☐ No critical errors in logs
☐ Key metrics within acceptable ranges
☐ Validation command passed: SELECT COUNT(*) FROM {table_name};...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH...
3. ROLLBACK_MIGRATE
Description: Rollback changes made during migrate phase
Urgency: MEDIUM
Duration: 570 minutes
Risk Level: MEDIUM
Prerequisites:
✓ Incident commander assigned and briefed
✓ All team members notified of rollback initiation
✓ Monitoring systems confirmed operational
✓ Backup systems verified and accessible
✓ Previous rollback phase completed successfully
Steps:
99. Validate rollback completion
Duration: 10 min
Type: manual
Success Criteria: migrate fully rolled back, All validation checks pass
Validation Checkpoints:
☐ migrate rollback steps completed
☐ System health checks passing
☐ No critical errors in logs
☐ Key metrics within acceptable ranges
☐ Validation command passed: SELECT COUNT(*) FROM {table_name};...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH...
4. ROLLBACK_EXPAND
Description: Rollback changes made during expand phase
Urgency: MEDIUM
Duration: 570 minutes
Risk Level: MEDIUM
Prerequisites:
✓ Incident commander assigned and briefed
✓ All team members notified of rollback initiation
✓ Monitoring systems confirmed operational
✓ Backup systems verified and accessible
✓ Previous rollback phase completed successfully
Steps:
99. Validate rollback completion
Duration: 10 min
Type: manual
Success Criteria: expand fully rolled back, All validation checks pass
Validation Checkpoints:
☐ expand rollback steps completed
☐ System health checks passing
☐ No critical errors in logs
☐ Key metrics within acceptable ranges
☐ Validation command passed: SELECT COUNT(*) FROM {table_name};...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH...
5. ROLLBACK_PREPARATION
Description: Rollback changes made during preparation phase
Urgency: MEDIUM
Duration: 570 minutes
Risk Level: MEDIUM
Prerequisites:
✓ Incident commander assigned and briefed
✓ All team members notified of rollback initiation
✓ Monitoring systems confirmed operational
✓ Backup systems verified and accessible
✓ Previous rollback phase completed successfully
Steps:
1. Drop migration artifacts
Duration: 5 min
Type: sql
Script:
-- Drop migration artifacts
DROP TABLE IF EXISTS migration_log;
DROP PROCEDURE IF EXISTS migrate_data();
Success Criteria: No migration artifacts remain
99. Validate rollback completion
Duration: 10 min
Type: manual
Success Criteria: preparation fully rolled back, All validation checks pass
Validation Checkpoints:
☐ preparation rollback steps completed
☐ System health checks passing
☐ No critical errors in logs
☐ Key metrics within acceptable ranges
☐ Validation command passed: SELECT COUNT(*) FROM {table_name};...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH...
DATA RECOVERY PLAN
----------------------------------------
Recovery Method: point_in_time
Backup Location: /backups/pre_migration_{migration_id}_{timestamp}.sql
Estimated Recovery Time: 45 minutes
Recovery Scripts:
• pg_restore -d production -c /backups/pre_migration_backup.sql
• SELECT pg_create_restore_point('rollback_point');
• VACUUM ANALYZE; -- Refresh statistics after restore
Validation Queries:
• SELECT COUNT(*) FROM critical_business_table;
• SELECT MAX(created_at) FROM audit_log;
• SELECT COUNT(DISTINCT user_id) FROM user_sessions;
• SELECT SUM(amount) FROM financial_transactions WHERE date = CURRENT_DATE;
POST-ROLLBACK VALIDATION CHECKLIST
----------------------------------------
1. ☐ Verify system is responding to health checks
2. ☐ Confirm error rates are within normal parameters
3. ☐ Validate response times meet SLA requirements
4. ☐ Check all critical business processes are functioning
5. ☐ Verify monitoring and alerting systems are operational
6. ☐ Confirm no data corruption has occurred
7. ☐ Validate security controls are functioning properly
8. ☐ Check backup systems are working correctly
9. ☐ Verify integration points with downstream systems
10. ☐ Confirm user authentication and authorization working
11. ☐ Validate database schema matches expected state
12. ☐ Confirm referential integrity constraints
13. ☐ Check database performance metrics
14. ☐ Verify data consistency across related tables
15. ☐ Validate indexes and statistics are optimal
16. ☐ Confirm transaction logs are clean
17. ☐ Check database connections and connection pooling
POST-ROLLBACK PROCEDURES
----------------------------------------
1. Monitor system stability for 24-48 hours post-rollback
2. Conduct thorough post-rollback testing of all critical paths
3. Review and analyze rollback metrics and timing
4. Document lessons learned and rollback procedure improvements
5. Schedule post-mortem meeting with all stakeholders
6. Update rollback procedures based on actual experience
7. Communicate rollback completion to all stakeholders
8. Archive rollback logs and artifacts for future reference
9. Review and update monitoring thresholds if needed
10. Plan for next migration attempt with improved procedures
11. Conduct security review to ensure no vulnerabilities introduced
12. Update disaster recovery procedures if affected by rollback
13. Review capacity planning based on rollback resource usage
14. Update documentation with rollback experience and timings
FILE:expected_outputs/sample_database_migration_plan.json
{
"migration_id": "23a52ed1507f",
"source_system": "PostgreSQL 13 Production Database",
"target_system": "PostgreSQL 15 Cloud Database",
"migration_type": "database",
"complexity": "critical",
"estimated_duration_hours": 95,
"phases": [
{
"name": "preparation",
"description": "Prepare systems and teams for migration",
"duration_hours": 19,
"dependencies": [],
"validation_criteria": [
"All backups completed successfully",
"Monitoring systems operational",
"Team members briefed and ready",
"Rollback procedures tested"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Backup source system",
"Set up monitoring and alerting",
"Prepare rollback procedures",
"Communicate migration timeline",
"Validate prerequisites"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
},
{
"name": "expand",
"description": "Execute expand phase",
"duration_hours": 19,
"dependencies": [
"preparation"
],
"validation_criteria": [
"Expand phase completed successfully"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Complete expand activities"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
},
{
"name": "migrate",
"description": "Execute migrate phase",
"duration_hours": 19,
"dependencies": [
"expand"
],
"validation_criteria": [
"Migrate phase completed successfully"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Complete migrate activities"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
},
{
"name": "contract",
"description": "Execute contract phase",
"duration_hours": 19,
"dependencies": [
"migrate"
],
"validation_criteria": [
"Contract phase completed successfully"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Complete contract activities"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
},
{
"name": "cleanup",
"description": "Execute cleanup phase",
"duration_hours": 19,
"dependencies": [
"contract"
],
"validation_criteria": [
"Cleanup phase completed successfully"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Complete cleanup activities"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
}
],
"risks": [
{
"category": "technical",
"description": "Data corruption during migration",
"probability": "low",
"impact": "critical",
"severity": "high",
"mitigation": "Implement comprehensive backup and validation procedures",
"owner": "DBA Team"
},
{
"category": "technical",
"description": "Extended downtime due to migration complexity",
"probability": "medium",
"impact": "high",
"severity": "high",
"mitigation": "Use blue-green deployment and phased migration approach",
"owner": "DevOps Team"
},
{
"category": "business",
"description": "Business process disruption",
"probability": "medium",
"impact": "high",
"severity": "high",
"mitigation": "Communicate timeline and provide alternate workflows",
"owner": "Business Owner"
},
{
"category": "operational",
"description": "Insufficient rollback testing",
"probability": "high",
"impact": "critical",
"severity": "critical",
"mitigation": "Execute full rollback procedures in staging environment",
"owner": "QA Team"
},
{
"category": "business",
"description": "Zero-downtime requirement increases complexity",
"probability": "high",
"impact": "medium",
"severity": "high",
"mitigation": "Implement blue-green deployment or rolling update strategy",
"owner": "DevOps Team"
},
{
"category": "compliance",
"description": "Regulatory compliance requirements",
"probability": "medium",
"impact": "high",
"severity": "high",
"mitigation": "Ensure all compliance checks are integrated into migration process",
"owner": "Compliance Team"
}
],
"success_criteria": [
"All data successfully migrated with 100% integrity",
"System performance meets or exceeds baseline",
"All business processes functioning normally",
"No critical security vulnerabilities introduced",
"Stakeholder acceptance criteria met",
"Documentation and runbooks updated"
],
"rollback_plan": {
"rollback_phases": [
{
"phase": "cleanup",
"rollback_actions": [
"Revert cleanup changes",
"Restore pre-cleanup state",
"Validate cleanup rollback success"
],
"validation_criteria": [
"System restored to pre-cleanup state",
"All cleanup changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 285
},
{
"phase": "contract",
"rollback_actions": [
"Revert contract changes",
"Restore pre-contract state",
"Validate contract rollback success"
],
"validation_criteria": [
"System restored to pre-contract state",
"All contract changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 285
},
{
"phase": "migrate",
"rollback_actions": [
"Revert migrate changes",
"Restore pre-migrate state",
"Validate migrate rollback success"
],
"validation_criteria": [
"System restored to pre-migrate state",
"All migrate changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 285
},
{
"phase": "expand",
"rollback_actions": [
"Revert expand changes",
"Restore pre-expand state",
"Validate expand rollback success"
],
"validation_criteria": [
"System restored to pre-expand state",
"All expand changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 285
},
{
"phase": "preparation",
"rollback_actions": [
"Revert preparation changes",
"Restore pre-preparation state",
"Validate preparation rollback success"
],
"validation_criteria": [
"System restored to pre-preparation state",
"All preparation changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 285
}
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Migration timeline exceeded by > 50%",
"Business-critical functionality unavailable",
"Security breach detected",
"Stakeholder decision to abort"
],
"rollback_decision_matrix": {
"low_severity": "Continue with monitoring",
"medium_severity": "Assess and decide within 15 minutes",
"high_severity": "Immediate rollback initiation",
"critical_severity": "Emergency rollback - all hands"
},
"rollback_contacts": [
"Migration Lead",
"Technical Lead",
"Business Owner",
"On-call Engineer"
]
},
"stakeholders": [
"Business Owner",
"Technical Lead",
"DevOps Team",
"QA Team",
"Security Team",
"End Users"
],
"created_at": "2026-02-16T13:47:23.704502"
}
FILE:expected_outputs/sample_database_migration_plan.txt
================================================================================
MIGRATION PLAN: 23a52ed1507f
================================================================================
Source System: PostgreSQL 13 Production Database
Target System: PostgreSQL 15 Cloud Database
Migration Type: DATABASE
Complexity Level: CRITICAL
Estimated Duration: 95 hours (4.0 days)
Created: 2026-02-16T13:47:23.704502
MIGRATION PHASES
----------------------------------------
1. PREPARATION (19h)
Description: Prepare systems and teams for migration
Risk Level: MEDIUM
Tasks:
• Backup source system
• Set up monitoring and alerting
• Prepare rollback procedures
• Communicate migration timeline
• Validate prerequisites
Success Criteria:
✓ All backups completed successfully
✓ Monitoring systems operational
✓ Team members briefed and ready
✓ Rollback procedures tested
2. EXPAND (19h)
Description: Execute expand phase
Risk Level: MEDIUM
Dependencies: preparation
Tasks:
• Complete expand activities
Success Criteria:
✓ Expand phase completed successfully
3. MIGRATE (19h)
Description: Execute migrate phase
Risk Level: MEDIUM
Dependencies: expand
Tasks:
• Complete migrate activities
Success Criteria:
✓ Migrate phase completed successfully
4. CONTRACT (19h)
Description: Execute contract phase
Risk Level: MEDIUM
Dependencies: migrate
Tasks:
• Complete contract activities
Success Criteria:
✓ Contract phase completed successfully
5. CLEANUP (19h)
Description: Execute cleanup phase
Risk Level: MEDIUM
Dependencies: contract
Tasks:
• Complete cleanup activities
Success Criteria:
✓ Cleanup phase completed successfully
RISK ASSESSMENT
----------------------------------------
CRITICAL SEVERITY RISKS:
• Insufficient rollback testing
Category: operational
Probability: high | Impact: critical
Mitigation: Execute full rollback procedures in staging environment
Owner: QA Team
HIGH SEVERITY RISKS:
• Data corruption during migration
Category: technical
Probability: low | Impact: critical
Mitigation: Implement comprehensive backup and validation procedures
Owner: DBA Team
• Extended downtime due to migration complexity
Category: technical
Probability: medium | Impact: high
Mitigation: Use blue-green deployment and phased migration approach
Owner: DevOps Team
• Business process disruption
Category: business
Probability: medium | Impact: high
Mitigation: Communicate timeline and provide alternate workflows
Owner: Business Owner
• Zero-downtime requirement increases complexity
Category: business
Probability: high | Impact: medium
Mitigation: Implement blue-green deployment or rolling update strategy
Owner: DevOps Team
• Regulatory compliance requirements
Category: compliance
Probability: medium | Impact: high
Mitigation: Ensure all compliance checks are integrated into migration process
Owner: Compliance Team
ROLLBACK STRATEGY
----------------------------------------
Rollback Triggers:
• Critical system failure
• Data corruption detected
• Migration timeline exceeded by > 50%
• Business-critical functionality unavailable
• Security breach detected
• Stakeholder decision to abort
Rollback Phases:
CLEANUP:
- Revert cleanup changes
- Restore pre-cleanup state
- Validate cleanup rollback success
Estimated Time: 285 minutes
CONTRACT:
- Revert contract changes
- Restore pre-contract state
- Validate contract rollback success
Estimated Time: 285 minutes
MIGRATE:
- Revert migrate changes
- Restore pre-migrate state
- Validate migrate rollback success
Estimated Time: 285 minutes
EXPAND:
- Revert expand changes
- Restore pre-expand state
- Validate expand rollback success
Estimated Time: 285 minutes
PREPARATION:
- Revert preparation changes
- Restore pre-preparation state
- Validate preparation rollback success
Estimated Time: 285 minutes
SUCCESS CRITERIA
----------------------------------------
✓ All data successfully migrated with 100% integrity
✓ System performance meets or exceeds baseline
✓ All business processes functioning normally
✓ No critical security vulnerabilities introduced
✓ Stakeholder acceptance criteria met
✓ Documentation and runbooks updated
STAKEHOLDERS
----------------------------------------
• Business Owner
• Technical Lead
• DevOps Team
• QA Team
• Security Team
• End Users
FILE:expected_outputs/sample_service_migration_plan.json
{
"migration_id": "21031930da18",
"source_system": "Legacy User Service (Java Spring Boot 2.x)",
"target_system": "New User Service (Node.js + TypeScript)",
"migration_type": "service",
"complexity": "critical",
"estimated_duration_hours": 500,
"phases": [
{
"name": "intercept",
"description": "Execute intercept phase",
"duration_hours": 100,
"dependencies": [],
"validation_criteria": [
"Intercept phase completed successfully"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Complete intercept activities"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
},
{
"name": "implement",
"description": "Execute implement phase",
"duration_hours": 100,
"dependencies": [
"intercept"
],
"validation_criteria": [
"Implement phase completed successfully"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Complete implement activities"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
},
{
"name": "redirect",
"description": "Execute redirect phase",
"duration_hours": 100,
"dependencies": [
"implement"
],
"validation_criteria": [
"Redirect phase completed successfully"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Complete redirect activities"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
},
{
"name": "validate",
"description": "Execute validate phase",
"duration_hours": 100,
"dependencies": [
"redirect"
],
"validation_criteria": [
"Validate phase completed successfully"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Complete validate activities"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
},
{
"name": "retire",
"description": "Execute retire phase",
"duration_hours": 100,
"dependencies": [
"validate"
],
"validation_criteria": [
"Retire phase completed successfully"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Complete retire activities"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
}
],
"risks": [
{
"category": "technical",
"description": "Service compatibility issues",
"probability": "medium",
"impact": "high",
"severity": "high",
"mitigation": "Implement comprehensive integration testing",
"owner": "Development Team"
},
{
"category": "technical",
"description": "Performance degradation",
"probability": "medium",
"impact": "medium",
"severity": "medium",
"mitigation": "Conduct load testing and performance benchmarking",
"owner": "DevOps Team"
},
{
"category": "business",
"description": "Feature parity gaps",
"probability": "high",
"impact": "high",
"severity": "high",
"mitigation": "Document feature mapping and acceptance criteria",
"owner": "Product Owner"
},
{
"category": "operational",
"description": "Monitoring gap during transition",
"probability": "medium",
"impact": "medium",
"severity": "medium",
"mitigation": "Set up dual monitoring and alerting systems",
"owner": "SRE Team"
},
{
"category": "business",
"description": "Zero-downtime requirement increases complexity",
"probability": "high",
"impact": "medium",
"severity": "high",
"mitigation": "Implement blue-green deployment or rolling update strategy",
"owner": "DevOps Team"
},
{
"category": "compliance",
"description": "Regulatory compliance requirements",
"probability": "medium",
"impact": "high",
"severity": "high",
"mitigation": "Ensure all compliance checks are integrated into migration process",
"owner": "Compliance Team"
}
],
"success_criteria": [
"All data successfully migrated with 100% integrity",
"System performance meets or exceeds baseline",
"All business processes functioning normally",
"No critical security vulnerabilities introduced",
"Stakeholder acceptance criteria met",
"Documentation and runbooks updated"
],
"rollback_plan": {
"rollback_phases": [
{
"phase": "retire",
"rollback_actions": [
"Revert retire changes",
"Restore pre-retire state",
"Validate retire rollback success"
],
"validation_criteria": [
"System restored to pre-retire state",
"All retire changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 1500
},
{
"phase": "validate",
"rollback_actions": [
"Revert validate changes",
"Restore pre-validate state",
"Validate validate rollback success"
],
"validation_criteria": [
"System restored to pre-validate state",
"All validate changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 1500
},
{
"phase": "redirect",
"rollback_actions": [
"Revert redirect changes",
"Restore pre-redirect state",
"Validate redirect rollback success"
],
"validation_criteria": [
"System restored to pre-redirect state",
"All redirect changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 1500
},
{
"phase": "implement",
"rollback_actions": [
"Revert implement changes",
"Restore pre-implement state",
"Validate implement rollback success"
],
"validation_criteria": [
"System restored to pre-implement state",
"All implement changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 1500
},
{
"phase": "intercept",
"rollback_actions": [
"Revert intercept changes",
"Restore pre-intercept state",
"Validate intercept rollback success"
],
"validation_criteria": [
"System restored to pre-intercept state",
"All intercept changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 1500
}
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Migration timeline exceeded by > 50%",
"Business-critical functionality unavailable",
"Security breach detected",
"Stakeholder decision to abort"
],
"rollback_decision_matrix": {
"low_severity": "Continue with monitoring",
"medium_severity": "Assess and decide within 15 minutes",
"high_severity": "Immediate rollback initiation",
"critical_severity": "Emergency rollback - all hands"
},
"rollback_contacts": [
"Migration Lead",
"Technical Lead",
"Business Owner",
"On-call Engineer"
]
},
"stakeholders": [
"Business Owner",
"Technical Lead",
"DevOps Team",
"QA Team",
"Security Team",
"End Users"
],
"created_at": "2026-02-16T13:47:34.565896"
}
FILE:expected_outputs/sample_service_migration_plan.txt
================================================================================
MIGRATION PLAN: 21031930da18
================================================================================
Source System: Legacy User Service (Java Spring Boot 2.x)
Target System: New User Service (Node.js + TypeScript)
Migration Type: SERVICE
Complexity Level: CRITICAL
Estimated Duration: 500 hours (20.8 days)
Created: 2026-02-16T13:47:34.565896
MIGRATION PHASES
----------------------------------------
1. INTERCEPT (100h)
Description: Execute intercept phase
Risk Level: MEDIUM
Tasks:
• Complete intercept activities
Success Criteria:
✓ Intercept phase completed successfully
2. IMPLEMENT (100h)
Description: Execute implement phase
Risk Level: MEDIUM
Dependencies: intercept
Tasks:
• Complete implement activities
Success Criteria:
✓ Implement phase completed successfully
3. REDIRECT (100h)
Description: Execute redirect phase
Risk Level: MEDIUM
Dependencies: implement
Tasks:
• Complete redirect activities
Success Criteria:
✓ Redirect phase completed successfully
4. VALIDATE (100h)
Description: Execute validate phase
Risk Level: MEDIUM
Dependencies: redirect
Tasks:
• Complete validate activities
Success Criteria:
✓ Validate phase completed successfully
5. RETIRE (100h)
Description: Execute retire phase
Risk Level: MEDIUM
Dependencies: validate
Tasks:
• Complete retire activities
Success Criteria:
✓ Retire phase completed successfully
RISK ASSESSMENT
----------------------------------------
HIGH SEVERITY RISKS:
• Service compatibility issues
Category: technical
Probability: medium | Impact: high
Mitigation: Implement comprehensive integration testing
Owner: Development Team
• Feature parity gaps
Category: business
Probability: high | Impact: high
Mitigation: Document feature mapping and acceptance criteria
Owner: Product Owner
• Zero-downtime requirement increases complexity
Category: business
Probability: high | Impact: medium
Mitigation: Implement blue-green deployment or rolling update strategy
Owner: DevOps Team
• Regulatory compliance requirements
Category: compliance
Probability: medium | Impact: high
Mitigation: Ensure all compliance checks are integrated into migration process
Owner: Compliance Team
MEDIUM SEVERITY RISKS:
• Performance degradation
Category: technical
Probability: medium | Impact: medium
Mitigation: Conduct load testing and performance benchmarking
Owner: DevOps Team
• Monitoring gap during transition
Category: operational
Probability: medium | Impact: medium
Mitigation: Set up dual monitoring and alerting systems
Owner: SRE Team
ROLLBACK STRATEGY
----------------------------------------
Rollback Triggers:
• Critical system failure
• Data corruption detected
• Migration timeline exceeded by > 50%
• Business-critical functionality unavailable
• Security breach detected
• Stakeholder decision to abort
Rollback Phases:
RETIRE:
- Revert retire changes
- Restore pre-retire state
- Validate retire rollback success
Estimated Time: 1500 minutes
VALIDATE:
- Revert validate changes
- Restore pre-validate state
- Validate validate rollback success
Estimated Time: 1500 minutes
REDIRECT:
- Revert redirect changes
- Restore pre-redirect state
- Validate redirect rollback success
Estimated Time: 1500 minutes
IMPLEMENT:
- Revert implement changes
- Restore pre-implement state
- Validate implement rollback success
Estimated Time: 1500 minutes
INTERCEPT:
- Revert intercept changes
- Restore pre-intercept state
- Validate intercept rollback success
Estimated Time: 1500 minutes
SUCCESS CRITERIA
----------------------------------------
✓ All data successfully migrated with 100% integrity
✓ System performance meets or exceeds baseline
✓ All business processes functioning normally
✓ No critical security vulnerabilities introduced
✓ Stakeholder acceptance criteria met
✓ Documentation and runbooks updated
STAKEHOLDERS
----------------------------------------
• Business Owner
• Technical Lead
• DevOps Team
• QA Team
• Security Team
• End Users
FILE:expected_outputs/schema_compatibility_report.json
{
"schema_before": "{\n \"schema_version\": \"1.0\",\n \"database\": \"user_management\",\n \"tables\": {\n \"users\": {\n \"columns\": {\n \"id\": {\n \"type\": \"bigint\",\n \"nullable\": false,\n \"primary_key\": true,\n \"auto_increment\": true\n },\n \"username\": {\n \"type\": \"varchar\",\n \"length\": 50,\n \"nullable\": false,\n \"unique\": true\n },\n \"email\": {\n \"type\": \"varchar\",\n \"length\": 255,\n \"nullable\": false,\n...",
"schema_after": "{\n \"schema_version\": \"2.0\",\n \"database\": \"user_management_v2\",\n \"tables\": {\n \"users\": {\n \"columns\": {\n \"id\": {\n \"type\": \"bigint\",\n \"nullable\": false,\n \"primary_key\": true,\n \"auto_increment\": true\n },\n \"username\": {\n \"type\": \"varchar\",\n \"length\": 50,\n \"nullable\": false,\n \"unique\": true\n },\n \"email\": {\n \"type\": \"varchar\",\n \"length\": 320,\n \"nullable\": fals...",
"analysis_date": "2026-02-16T13:47:27.050459",
"overall_compatibility": "potentially_incompatible",
"breaking_changes_count": 0,
"potentially_breaking_count": 4,
"non_breaking_changes_count": 0,
"additive_changes_count": 0,
"issues": [
{
"type": "check_added",
"severity": "potentially_breaking",
"description": "New check constraint 'phone IS NULL OR LENGTH(phone) >= 10' added to table 'users'",
"field_path": "tables.users.constraints.check",
"old_value": null,
"new_value": "phone IS NULL OR LENGTH(phone) >= 10",
"impact": "New check constraint may reject existing data",
"suggested_migration": "Validate existing data complies with new constraint",
"affected_operations": [
"INSERT",
"UPDATE"
]
},
{
"type": "check_added",
"severity": "potentially_breaking",
"description": "New check constraint 'bio IS NULL OR LENGTH(bio) <= 2000' added to table 'user_profiles'",
"field_path": "tables.user_profiles.constraints.check",
"old_value": null,
"new_value": "bio IS NULL OR LENGTH(bio) <= 2000",
"impact": "New check constraint may reject existing data",
"suggested_migration": "Validate existing data complies with new constraint",
"affected_operations": [
"INSERT",
"UPDATE"
]
},
{
"type": "check_added",
"severity": "potentially_breaking",
"description": "New check constraint 'language IN ('en', 'es', 'fr', 'de', 'it', 'pt', 'ru', 'ja', 'ko', 'zh')' added to table 'user_profiles'",
"field_path": "tables.user_profiles.constraints.check",
"old_value": null,
"new_value": "language IN ('en', 'es', 'fr', 'de', 'it', 'pt', 'ru', 'ja', 'ko', 'zh')",
"impact": "New check constraint may reject existing data",
"suggested_migration": "Validate existing data complies with new constraint",
"affected_operations": [
"INSERT",
"UPDATE"
]
},
{
"type": "check_added",
"severity": "potentially_breaking",
"description": "New check constraint 'session_type IN ('web', 'mobile', 'api', 'admin')' added to table 'user_sessions'",
"field_path": "tables.user_sessions.constraints.check",
"old_value": null,
"new_value": "session_type IN ('web', 'mobile', 'api', 'admin')",
"impact": "New check constraint may reject existing data",
"suggested_migration": "Validate existing data complies with new constraint",
"affected_operations": [
"INSERT",
"UPDATE"
]
}
],
"migration_scripts": [
{
"script_type": "sql",
"description": "Create new table user_preferences",
"script_content": "CREATE TABLE user_preferences (\n id bigint NOT NULL,\n user_id bigint NOT NULL,\n preference_key varchar NOT NULL,\n preference_value json,\n created_at timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP,\n updated_at timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP\n);",
"rollback_script": "DROP TABLE IF EXISTS user_preferences;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.tables WHERE table_name = 'user_preferences';"
},
{
"script_type": "sql",
"description": "Add column email_verified_at to table users",
"script_content": "ALTER TABLE users ADD COLUMN email_verified_at timestamp;",
"rollback_script": "ALTER TABLE users DROP COLUMN email_verified_at;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.columns WHERE table_name = 'users' AND column_name = 'email_verified_at';"
},
{
"script_type": "sql",
"description": "Add column phone_verified_at to table users",
"script_content": "ALTER TABLE users ADD COLUMN phone_verified_at timestamp;",
"rollback_script": "ALTER TABLE users DROP COLUMN phone_verified_at;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.columns WHERE table_name = 'users' AND column_name = 'phone_verified_at';"
},
{
"script_type": "sql",
"description": "Add column two_factor_enabled to table users",
"script_content": "ALTER TABLE users ADD COLUMN two_factor_enabled boolean NOT NULL DEFAULT False;",
"rollback_script": "ALTER TABLE users DROP COLUMN two_factor_enabled;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.columns WHERE table_name = 'users' AND column_name = 'two_factor_enabled';"
},
{
"script_type": "sql",
"description": "Add column last_login_at to table users",
"script_content": "ALTER TABLE users ADD COLUMN last_login_at timestamp;",
"rollback_script": "ALTER TABLE users DROP COLUMN last_login_at;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.columns WHERE table_name = 'users' AND column_name = 'last_login_at';"
},
{
"script_type": "sql",
"description": "Add check constraint to users",
"script_content": "ALTER TABLE users ADD CONSTRAINT check_users CHECK (phone IS NULL OR LENGTH(phone) >= 10);",
"rollback_script": "ALTER TABLE users DROP CONSTRAINT check_users;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.table_constraints WHERE table_name = 'users' AND constraint_type = 'CHECK';"
},
{
"script_type": "sql",
"description": "Add column timezone to table user_profiles",
"script_content": "ALTER TABLE user_profiles ADD COLUMN timezone varchar DEFAULT UTC;",
"rollback_script": "ALTER TABLE user_profiles DROP COLUMN timezone;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.columns WHERE table_name = 'user_profiles' AND column_name = 'timezone';"
},
{
"script_type": "sql",
"description": "Add column language to table user_profiles",
"script_content": "ALTER TABLE user_profiles ADD COLUMN language varchar NOT NULL DEFAULT en;",
"rollback_script": "ALTER TABLE user_profiles DROP COLUMN language;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.columns WHERE table_name = 'user_profiles' AND column_name = 'language';"
},
{
"script_type": "sql",
"description": "Add check constraint to user_profiles",
"script_content": "ALTER TABLE user_profiles ADD CONSTRAINT check_user_profiles CHECK (bio IS NULL OR LENGTH(bio) <= 2000);",
"rollback_script": "ALTER TABLE user_profiles DROP CONSTRAINT check_user_profiles;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.table_constraints WHERE table_name = 'user_profiles' AND constraint_type = 'CHECK';"
},
{
"script_type": "sql",
"description": "Add check constraint to user_profiles",
"script_content": "ALTER TABLE user_profiles ADD CONSTRAINT check_user_profiles CHECK (language IN ('en', 'es', 'fr', 'de', 'it', 'pt', 'ru', 'ja', 'ko', 'zh'));",
"rollback_script": "ALTER TABLE user_profiles DROP CONSTRAINT check_user_profiles;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.table_constraints WHERE table_name = 'user_profiles' AND constraint_type = 'CHECK';"
},
{
"script_type": "sql",
"description": "Add column session_type to table user_sessions",
"script_content": "ALTER TABLE user_sessions ADD COLUMN session_type varchar NOT NULL DEFAULT web;",
"rollback_script": "ALTER TABLE user_sessions DROP COLUMN session_type;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.columns WHERE table_name = 'user_sessions' AND column_name = 'session_type';"
},
{
"script_type": "sql",
"description": "Add column is_mobile to table user_sessions",
"script_content": "ALTER TABLE user_sessions ADD COLUMN is_mobile boolean NOT NULL DEFAULT False;",
"rollback_script": "ALTER TABLE user_sessions DROP COLUMN is_mobile;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.columns WHERE table_name = 'user_sessions' AND column_name = 'is_mobile';"
},
{
"script_type": "sql",
"description": "Add check constraint to user_sessions",
"script_content": "ALTER TABLE user_sessions ADD CONSTRAINT check_user_sessions CHECK (session_type IN ('web', 'mobile', 'api', 'admin'));",
"rollback_script": "ALTER TABLE user_sessions DROP CONSTRAINT check_user_sessions;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.table_constraints WHERE table_name = 'user_sessions' AND constraint_type = 'CHECK';"
}
],
"risk_assessment": {
"overall_risk": "medium",
"deployment_risk": "safe_independent_deployment",
"rollback_complexity": "low",
"testing_requirements": [
"integration_testing",
"regression_testing",
"data_migration_testing"
]
},
"recommendations": [
"Conduct thorough testing with realistic data volumes",
"Implement monitoring for migration success metrics",
"Test all migration scripts in staging environment",
"Implement migration progress monitoring",
"Create detailed communication plan for stakeholders",
"Implement feature flags for gradual rollout"
]
}
FILE:expected_outputs/schema_compatibility_report.txt
================================================================================
COMPATIBILITY ANALYSIS REPORT
================================================================================
Analysis Date: 2026-02-16T13:47:27.050459
Overall Compatibility: POTENTIALLY_INCOMPATIBLE
SUMMARY
----------------------------------------
Breaking Changes: 0
Potentially Breaking: 4
Non-Breaking Changes: 0
Additive Changes: 0
Total Issues Found: 4
RISK ASSESSMENT
----------------------------------------
Overall Risk: medium
Deployment Risk: safe_independent_deployment
Rollback Complexity: low
Testing Requirements: ['integration_testing', 'regression_testing', 'data_migration_testing']
POTENTIALLY BREAKING ISSUES
----------------------------------------
• New check constraint 'phone IS NULL OR LENGTH(phone) >= 10' added to table 'users'
Field: tables.users.constraints.check
Impact: New check constraint may reject existing data
Migration: Validate existing data complies with new constraint
Affected Operations: INSERT, UPDATE
• New check constraint 'bio IS NULL OR LENGTH(bio) <= 2000' added to table 'user_profiles'
Field: tables.user_profiles.constraints.check
Impact: New check constraint may reject existing data
Migration: Validate existing data complies with new constraint
Affected Operations: INSERT, UPDATE
• New check constraint 'language IN ('en', 'es', 'fr', 'de', 'it', 'pt', 'ru', 'ja', 'ko', 'zh')' added to table 'user_profiles'
Field: tables.user_profiles.constraints.check
Impact: New check constraint may reject existing data
Migration: Validate existing data complies with new constraint
Affected Operations: INSERT, UPDATE
• New check constraint 'session_type IN ('web', 'mobile', 'api', 'admin')' added to table 'user_sessions'
Field: tables.user_sessions.constraints.check
Impact: New check constraint may reject existing data
Migration: Validate existing data complies with new constraint
Affected Operations: INSERT, UPDATE
SUGGESTED MIGRATION SCRIPTS
----------------------------------------
1. Create new table user_preferences
Type: sql
Script:
CREATE TABLE user_preferences (
id bigint NOT NULL,
user_id bigint NOT NULL,
preference_key varchar NOT NULL,
preference_value json,
created_at timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP,
updated_at timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP
);
2. Add column email_verified_at to table users
Type: sql
Script:
ALTER TABLE users ADD COLUMN email_verified_at timestamp;
3. Add column phone_verified_at to table users
Type: sql
Script:
ALTER TABLE users ADD COLUMN phone_verified_at timestamp;
4. Add column two_factor_enabled to table users
Type: sql
Script:
ALTER TABLE users ADD COLUMN two_factor_enabled boolean NOT NULL DEFAULT False;
5. Add column last_login_at to table users
Type: sql
Script:
ALTER TABLE users ADD COLUMN last_login_at timestamp;
6. Add check constraint to users
Type: sql
Script:
ALTER TABLE users ADD CONSTRAINT check_users CHECK (phone IS NULL OR LENGTH(phone) >= 10);
7. Add column timezone to table user_profiles
Type: sql
Script:
ALTER TABLE user_profiles ADD COLUMN timezone varchar DEFAULT UTC;
8. Add column language to table user_profiles
Type: sql
Script:
ALTER TABLE user_profiles ADD COLUMN language varchar NOT NULL DEFAULT en;
9. Add check constraint to user_profiles
Type: sql
Script:
ALTER TABLE user_profiles ADD CONSTRAINT check_user_profiles CHECK (bio IS NULL OR LENGTH(bio) <= 2000);
10. Add check constraint to user_profiles
Type: sql
Script:
ALTER TABLE user_profiles ADD CONSTRAINT check_user_profiles CHECK (language IN ('en', 'es', 'fr', 'de', 'it', 'pt', 'ru', 'ja', 'ko', 'zh'));
11. Add column session_type to table user_sessions
Type: sql
Script:
ALTER TABLE user_sessions ADD COLUMN session_type varchar NOT NULL DEFAULT web;
12. Add column is_mobile to table user_sessions
Type: sql
Script:
ALTER TABLE user_sessions ADD COLUMN is_mobile boolean NOT NULL DEFAULT False;
13. Add check constraint to user_sessions
Type: sql
Script:
ALTER TABLE user_sessions ADD CONSTRAINT check_user_sessions CHECK (session_type IN ('web', 'mobile', 'api', 'admin'));
RECOMMENDATIONS
----------------------------------------
1. Conduct thorough testing with realistic data volumes
2. Implement monitoring for migration success metrics
3. Test all migration scripts in staging environment
4. Implement migration progress monitoring
5. Create detailed communication plan for stakeholders
6. Implement feature flags for gradual rollout
FILE:README.md
# Migration Architect
**Tier:** POWERFUL
**Category:** Engineering - Migration Strategy
**Purpose:** Zero-downtime migration planning, compatibility validation, and rollback strategy generation
## Overview
The Migration Architect skill provides comprehensive tools and methodologies for planning, executing, and validating complex system migrations with minimal business impact. This skill combines proven migration patterns with automated planning tools to ensure successful transitions between systems, databases, and infrastructure.
## Components
### Core Scripts
1. **migration_planner.py** - Automated migration plan generation
2. **compatibility_checker.py** - Schema and API compatibility analysis
3. **rollback_generator.py** - Comprehensive rollback procedure generation
### Reference Documentation
- **migration_patterns_catalog.md** - Detailed catalog of proven migration patterns
- **zero_downtime_techniques.md** - Comprehensive zero-downtime migration techniques
- **data_reconciliation_strategies.md** - Advanced data consistency and reconciliation strategies
### Sample Assets
- **sample_database_migration.json** - Example database migration specification
- **sample_service_migration.json** - Example service migration specification
- **database_schema_before.json** - Sample "before" database schema
- **database_schema_after.json** - Sample "after" database schema
## Quick Start
### 1. Generate a Migration Plan
```bash
python3 scripts/migration_planner.py \
--input assets/sample_database_migration.json \
--output migration_plan.json \
--format both
```
**Input:** Migration specification with source, target, constraints, and requirements
**Output:** Detailed phased migration plan with risk assessment, timeline, and validation gates
### 2. Check Compatibility
```bash
python3 scripts/compatibility_checker.py \
--before assets/database_schema_before.json \
--after assets/database_schema_after.json \
--type database \
--output compatibility_report.json \
--format both
```
**Input:** Before and after schema definitions
**Output:** Compatibility report with breaking changes, migration scripts, and recommendations
### 3. Generate Rollback Procedures
```bash
python3 scripts/rollback_generator.py \
--input migration_plan.json \
--output rollback_runbook.json \
--format both
```
**Input:** Migration plan from step 1
**Output:** Comprehensive rollback runbook with procedures, triggers, and communication templates
## Script Details
### Migration Planner (`migration_planner.py`)
Generates comprehensive migration plans with:
- **Phased approach** with dependencies and validation gates
- **Risk assessment** with mitigation strategies
- **Timeline estimation** based on complexity and constraints
- **Rollback triggers** and success criteria
- **Stakeholder communication** templates
**Usage:**
```bash
python3 scripts/migration_planner.py [OPTIONS]
Options:
--input, -i Input migration specification file (JSON) [required]
--output, -o Output file for migration plan (JSON)
--format, -f Output format: json, text, both (default: both)
--validate Validate migration specification only
```
**Input Format:**
```json
{
"type": "database|service|infrastructure",
"pattern": "schema_change|strangler_fig|blue_green",
"source": "Source system description",
"target": "Target system description",
"constraints": {
"max_downtime_minutes": 30,
"data_volume_gb": 2500,
"dependencies": ["service1", "service2"],
"compliance_requirements": ["GDPR", "SOX"]
}
}
```
### Compatibility Checker (`compatibility_checker.py`)
Analyzes compatibility between schema versions:
- **Breaking change detection** (removed fields, type changes, constraint additions)
- **Data migration requirements** identification
- **Suggested migration scripts** generation
- **Risk assessment** for each change
**Usage:**
```bash
python3 scripts/compatibility_checker.py [OPTIONS]
Options:
--before Before schema file (JSON) [required]
--after After schema file (JSON) [required]
--type Schema type: database, api (default: database)
--output, -o Output file for compatibility report (JSON)
--format, -f Output format: json, text, both (default: both)
```
**Exit Codes:**
- `0`: No compatibility issues
- `1`: Potentially breaking changes found
- `2`: Breaking changes found
### Rollback Generator (`rollback_generator.py`)
Creates comprehensive rollback procedures:
- **Phase-by-phase rollback** steps
- **Automated trigger conditions** for rollback
- **Data recovery procedures**
- **Communication templates** for different audiences
- **Validation checklists** for rollback success
**Usage:**
```bash
python3 scripts/rollback_generator.py [OPTIONS]
Options:
--input, -i Input migration plan file (JSON) [required]
--output, -o Output file for rollback runbook (JSON)
--format, -f Output format: json, text, both (default: both)
```
## Migration Patterns Supported
### Database Migrations
- **Expand-Contract Pattern** - Zero-downtime schema evolution
- **Parallel Schema Pattern** - Side-by-side schema migration
- **Event Sourcing Migration** - Event-driven data migration
### Service Migrations
- **Strangler Fig Pattern** - Gradual legacy system replacement
- **Parallel Run Pattern** - Risk mitigation through dual execution
- **Blue-Green Deployment** - Zero-downtime service updates
### Infrastructure Migrations
- **Lift and Shift** - Quick cloud migration with minimal changes
- **Hybrid Cloud Migration** - Gradual cloud adoption
- **Multi-Cloud Migration** - Distribution across multiple providers
## Sample Workflow
### 1. Database Schema Migration
```bash
# Generate migration plan
python3 scripts/migration_planner.py \
--input assets/sample_database_migration.json \
--output db_migration_plan.json
# Check schema compatibility
python3 scripts/compatibility_checker.py \
--before assets/database_schema_before.json \
--after assets/database_schema_after.json \
--type database \
--output schema_compatibility.json
# Generate rollback procedures
python3 scripts/rollback_generator.py \
--input db_migration_plan.json \
--output db_rollback_runbook.json
```
### 2. Service Migration
```bash
# Generate service migration plan
python3 scripts/migration_planner.py \
--input assets/sample_service_migration.json \
--output service_migration_plan.json
# Generate rollback procedures
python3 scripts/rollback_generator.py \
--input service_migration_plan.json \
--output service_rollback_runbook.json
```
## Output Examples
### Migration Plan Structure
```json
{
"migration_id": "abc123def456",
"source_system": "Legacy User Service",
"target_system": "New User Service",
"migration_type": "service",
"complexity": "medium",
"estimated_duration_hours": 72,
"phases": [
{
"name": "preparation",
"description": "Prepare systems and teams for migration",
"duration_hours": 8,
"validation_criteria": ["All backups completed successfully"],
"rollback_triggers": ["Critical system failure"],
"risk_level": "medium"
}
],
"risks": [
{
"category": "technical",
"description": "Service compatibility issues",
"severity": "high",
"mitigation": "Comprehensive integration testing"
}
]
}
```
### Compatibility Report Structure
```json
{
"overall_compatibility": "potentially_incompatible",
"breaking_changes_count": 2,
"potentially_breaking_count": 3,
"issues": [
{
"type": "required_column_added",
"severity": "breaking",
"description": "Required column 'email_verified_at' added",
"suggested_migration": "Add default value initially"
}
],
"migration_scripts": [
{
"script_type": "sql",
"description": "Add email verification columns",
"script_content": "ALTER TABLE users ADD COLUMN email_verified_at TIMESTAMP;",
"rollback_script": "ALTER TABLE users DROP COLUMN email_verified_at;"
}
]
}
```
## Best Practices
### Planning Phase
1. **Start with risk assessment** - Identify failure modes before planning
2. **Design for rollback** - Every step should have a tested rollback procedure
3. **Validate in staging** - Execute full migration in production-like environment
4. **Plan gradual rollout** - Use feature flags and traffic routing
### Execution Phase
1. **Monitor continuously** - Track technical and business metrics
2. **Communicate proactively** - Keep stakeholders informed
3. **Document everything** - Maintain detailed logs for analysis
4. **Stay flexible** - Be prepared to adjust based on real-world performance
### Validation Phase
1. **Automate validation** - Use automated consistency and performance checks
2. **Test business logic** - Validate critical business processes end-to-end
3. **Load test** - Verify performance under expected production load
4. **Security validation** - Ensure security controls function properly
## Integration
### CI/CD Pipeline Integration
```yaml
# Example GitHub Actions workflow
name: Migration Validation
on: [push, pull_request]
jobs:
validate-migration:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v2
- name: Validate Migration Plan
run: |
python3 scripts/migration_planner.py \
--input migration_spec.json \
--validate
- name: Check Compatibility
run: |
python3 scripts/compatibility_checker.py \
--before schema_before.json \
--after schema_after.json \
--type database
```
### Monitoring Integration
The tools generate metrics and alerts that can be integrated with:
- **Prometheus** - For metrics collection
- **Grafana** - For visualization and dashboards
- **PagerDuty** - For incident management
- **Slack** - For team notifications
## Advanced Features
### Machine Learning Integration
- Anomaly detection for data consistency issues
- Predictive analysis for migration success probability
- Automated pattern recognition for migration optimization
### Performance Optimization
- Parallel processing for large-scale migrations
- Incremental reconciliation strategies
- Statistical sampling for validation
### Compliance Support
- GDPR compliance tracking
- SOX audit trail generation
- HIPAA security validation
## Troubleshooting
### Common Issues
**"Migration plan validation failed"**
- Check JSON syntax in migration specification
- Ensure all required fields are present
- Validate constraint values are realistic
**"Compatibility checker reports false positives"**
- Review excluded fields configuration
- Check data type mapping compatibility
- Adjust tolerance settings for numerical comparisons
**"Rollback procedures seem incomplete"**
- Ensure migration plan includes all phases
- Verify database backup locations are specified
- Check that all dependencies are documented
### Getting Help
1. **Review documentation** - Check reference docs for patterns and techniques
2. **Examine sample files** - Use provided assets as templates
3. **Check expected outputs** - Compare your results with sample outputs
4. **Validate inputs** - Ensure input files match expected format
## Contributing
To extend or modify the Migration Architect skill:
1. **Add new patterns** - Extend pattern templates in migration_planner.py
2. **Enhance compatibility checks** - Add new validation rules in compatibility_checker.py
3. **Improve rollback procedures** - Add specialized rollback steps in rollback_generator.py
4. **Update documentation** - Keep reference docs current with new patterns
## License
This skill is part of the claude-skills repository and follows the same license terms.
FILE:references/data_reconciliation_strategies.md
# Data Reconciliation Strategies
## Overview
Data reconciliation is the process of ensuring data consistency and integrity across systems during and after migrations. This document provides comprehensive strategies, tools, and implementation patterns for detecting, measuring, and correcting data discrepancies in migration scenarios.
## Core Principles
### 1. Eventually Consistent
Accept that perfect real-time consistency may not be achievable during migrations, but ensure eventual consistency through reconciliation processes.
### 2. Idempotent Operations
All reconciliation operations must be safe to run multiple times without causing additional issues.
### 3. Audit Trail
Maintain detailed logs of all reconciliation actions for compliance and debugging.
### 4. Non-Destructive
Reconciliation should prefer addition over deletion, and always maintain backups before corrections.
## Types of Data Inconsistencies
### 1. Missing Records
Records that exist in source but not in target system.
### 2. Extra Records
Records that exist in target but not in source system.
### 3. Field Mismatches
Records exist in both systems but with different field values.
### 4. Referential Integrity Violations
Foreign key relationships that are broken during migration.
### 5. Temporal Inconsistencies
Data with incorrect timestamps or ordering.
### 6. Schema Drift
Structural differences between source and target schemas.
## Detection Strategies
### 1. Row Count Validation
#### Simple Count Comparison
```sql
-- Compare total row counts
SELECT
'source' as system,
COUNT(*) as row_count
FROM source_table
UNION ALL
SELECT
'target' as system,
COUNT(*) as row_count
FROM target_table;
```
#### Filtered Count Comparison
```sql
-- Compare counts with business logic filters
WITH source_counts AS (
SELECT
status,
created_date::date as date,
COUNT(*) as count
FROM source_orders
WHERE created_date >= '2024-01-01'
GROUP BY status, created_date::date
),
target_counts AS (
SELECT
status,
created_date::date as date,
COUNT(*) as count
FROM target_orders
WHERE created_date >= '2024-01-01'
GROUP BY status, created_date::date
)
SELECT
COALESCE(s.status, t.status) as status,
COALESCE(s.date, t.date) as date,
COALESCE(s.count, 0) as source_count,
COALESCE(t.count, 0) as target_count,
COALESCE(s.count, 0) - COALESCE(t.count, 0) as difference
FROM source_counts s
FULL OUTER JOIN target_counts t
ON s.status = t.status AND s.date = t.date
WHERE COALESCE(s.count, 0) != COALESCE(t.count, 0);
```
### 2. Checksum-Based Validation
#### Record-Level Checksums
```python
import hashlib
import json
class RecordChecksum:
def __init__(self, exclude_fields=None):
self.exclude_fields = exclude_fields or ['updated_at', 'version']
def calculate_checksum(self, record):
"""Calculate MD5 checksum for a database record"""
# Remove excluded fields and sort for consistency
filtered_record = {
k: v for k, v in record.items()
if k not in self.exclude_fields
}
# Convert to sorted JSON string for consistent hashing
normalized = json.dumps(filtered_record, sort_keys=True, default=str)
return hashlib.md5(normalized.encode('utf-8')).hexdigest()
def compare_records(self, source_record, target_record):
"""Compare two records using checksums"""
source_checksum = self.calculate_checksum(source_record)
target_checksum = self.calculate_checksum(target_record)
return {
'match': source_checksum == target_checksum,
'source_checksum': source_checksum,
'target_checksum': target_checksum
}
# Usage example
checksum_calculator = RecordChecksum(exclude_fields=['updated_at', 'migration_flag'])
source_records = fetch_records_from_source()
target_records = fetch_records_from_target()
mismatches = []
for source_id, source_record in source_records.items():
if source_id in target_records:
comparison = checksum_calculator.compare_records(
source_record, target_records[source_id]
)
if not comparison['match']:
mismatches.append({
'record_id': source_id,
'source_checksum': comparison['source_checksum'],
'target_checksum': comparison['target_checksum']
})
```
#### Aggregate Checksums
```sql
-- Calculate aggregate checksums for data validation
WITH source_aggregates AS (
SELECT
DATE_TRUNC('day', created_at) as day,
status,
COUNT(*) as record_count,
SUM(amount) as total_amount,
MD5(STRING_AGG(CAST(id AS VARCHAR) || ':' || CAST(amount AS VARCHAR), '|' ORDER BY id)) as checksum
FROM source_transactions
GROUP BY DATE_TRUNC('day', created_at), status
),
target_aggregates AS (
SELECT
DATE_TRUNC('day', created_at) as day,
status,
COUNT(*) as record_count,
SUM(amount) as total_amount,
MD5(STRING_AGG(CAST(id AS VARCHAR) || ':' || CAST(amount AS VARCHAR), '|' ORDER BY id)) as checksum
FROM target_transactions
GROUP BY DATE_TRUNC('day', created_at), status
)
SELECT
COALESCE(s.day, t.day) as day,
COALESCE(s.status, t.status) as status,
COALESCE(s.record_count, 0) as source_count,
COALESCE(t.record_count, 0) as target_count,
COALESCE(s.total_amount, 0) as source_amount,
COALESCE(t.total_amount, 0) as target_amount,
s.checksum as source_checksum,
t.checksum as target_checksum,
CASE WHEN s.checksum = t.checksum THEN 'MATCH' ELSE 'MISMATCH' END as status
FROM source_aggregates s
FULL OUTER JOIN target_aggregates t
ON s.day = t.day AND s.status = t.status
WHERE s.checksum != t.checksum OR s.checksum IS NULL OR t.checksum IS NULL;
```
### 3. Delta Detection
#### Change Data Capture (CDC) Based
```python
class CDCReconciler:
def __init__(self, kafka_client, database_client):
self.kafka = kafka_client
self.db = database_client
self.processed_changes = set()
def process_cdc_stream(self, topic_name):
"""Process CDC events and track changes for reconciliation"""
consumer = self.kafka.consumer(topic_name)
for message in consumer:
change_event = json.loads(message.value)
change_id = f"{change_event['table']}:{change_event['key']}:{change_event['timestamp']}"
if change_id in self.processed_changes:
continue # Skip duplicate events
try:
self.apply_change(change_event)
self.processed_changes.add(change_id)
# Commit offset only after successful processing
consumer.commit()
except Exception as e:
# Log failure and continue - will be caught by reconciliation
self.log_processing_failure(change_id, str(e))
def apply_change(self, change_event):
"""Apply CDC change to target system"""
table = change_event['table']
operation = change_event['operation']
key = change_event['key']
data = change_event.get('data', {})
if operation == 'INSERT':
self.db.insert(table, data)
elif operation == 'UPDATE':
self.db.update(table, key, data)
elif operation == 'DELETE':
self.db.delete(table, key)
def reconcile_missed_changes(self, start_timestamp, end_timestamp):
"""Find and apply changes that may have been missed"""
# Query source database for changes in time window
source_changes = self.db.get_changes_in_window(
start_timestamp, end_timestamp
)
missed_changes = []
for change in source_changes:
change_id = f"{change['table']}:{change['key']}:{change['timestamp']}"
if change_id not in self.processed_changes:
missed_changes.append(change)
# Apply missed changes
for change in missed_changes:
try:
self.apply_change(change)
print(f"Applied missed change: {change['table']}:{change['key']}")
except Exception as e:
print(f"Failed to apply missed change: {e}")
```
### 4. Business Logic Validation
#### Critical Business Rules Validation
```python
class BusinessLogicValidator:
def __init__(self, source_db, target_db):
self.source_db = source_db
self.target_db = target_db
def validate_financial_consistency(self):
"""Validate critical financial calculations"""
validation_rules = [
{
'name': 'daily_transaction_totals',
'source_query': """
SELECT DATE(created_at) as date, SUM(amount) as total
FROM source_transactions
WHERE created_at >= CURRENT_DATE - INTERVAL '30 days'
GROUP BY DATE(created_at)
""",
'target_query': """
SELECT DATE(created_at) as date, SUM(amount) as total
FROM target_transactions
WHERE created_at >= CURRENT_DATE - INTERVAL '30 days'
GROUP BY DATE(created_at)
""",
'tolerance': 0.01 # Allow $0.01 difference for rounding
},
{
'name': 'customer_balance_totals',
'source_query': """
SELECT customer_id, SUM(balance) as total_balance
FROM source_accounts
GROUP BY customer_id
HAVING SUM(balance) > 0
""",
'target_query': """
SELECT customer_id, SUM(balance) as total_balance
FROM target_accounts
GROUP BY customer_id
HAVING SUM(balance) > 0
""",
'tolerance': 0.01
}
]
validation_results = []
for rule in validation_rules:
source_data = self.source_db.execute_query(rule['source_query'])
target_data = self.target_db.execute_query(rule['target_query'])
differences = self.compare_financial_data(
source_data, target_data, rule['tolerance']
)
validation_results.append({
'rule_name': rule['name'],
'differences_found': len(differences),
'differences': differences[:10], # First 10 differences
'status': 'PASS' if len(differences) == 0 else 'FAIL'
})
return validation_results
def compare_financial_data(self, source_data, target_data, tolerance):
"""Compare financial data with tolerance for rounding differences"""
source_dict = {
tuple(row[:-1]): row[-1] for row in source_data
} # Last column is the amount
target_dict = {
tuple(row[:-1]): row[-1] for row in target_data
}
differences = []
# Check for missing records and value differences
for key, source_value in source_dict.items():
if key not in target_dict:
differences.append({
'key': key,
'source_value': source_value,
'target_value': None,
'difference_type': 'MISSING_IN_TARGET'
})
else:
target_value = target_dict[key]
if abs(float(source_value) - float(target_value)) > tolerance:
differences.append({
'key': key,
'source_value': source_value,
'target_value': target_value,
'difference': float(source_value) - float(target_value),
'difference_type': 'VALUE_MISMATCH'
})
# Check for extra records in target
for key, target_value in target_dict.items():
if key not in source_dict:
differences.append({
'key': key,
'source_value': None,
'target_value': target_value,
'difference_type': 'EXTRA_IN_TARGET'
})
return differences
```
## Correction Strategies
### 1. Automated Correction
#### Missing Record Insertion
```python
class AutoCorrector:
def __init__(self, source_db, target_db, dry_run=True):
self.source_db = source_db
self.target_db = target_db
self.dry_run = dry_run
self.correction_log = []
def correct_missing_records(self, table_name, key_field):
"""Add missing records from source to target"""
# Find records in source but not in target
missing_query = f"""
SELECT s.*
FROM source_{table_name} s
LEFT JOIN target_{table_name} t ON s.{key_field} = t.{key_field}
WHERE t.{key_field} IS NULL
"""
missing_records = self.source_db.execute_query(missing_query)
for record in missing_records:
correction = {
'table': table_name,
'operation': 'INSERT',
'key': record[key_field],
'data': record,
'timestamp': datetime.utcnow()
}
if not self.dry_run:
try:
self.target_db.insert(table_name, record)
correction['status'] = 'SUCCESS'
except Exception as e:
correction['status'] = 'FAILED'
correction['error'] = str(e)
else:
correction['status'] = 'DRY_RUN'
self.correction_log.append(correction)
return len(missing_records)
def correct_field_mismatches(self, table_name, key_field, fields_to_correct):
"""Correct field value mismatches"""
mismatch_query = f"""
SELECT s.{key_field}, {', '.join([f's.{f} as source_{f}, t.{f} as target_{f}' for f in fields_to_correct])}
FROM source_{table_name} s
JOIN target_{table_name} t ON s.{key_field} = t.{key_field}
WHERE {' OR '.join([f's.{f} != t.{f}' for f in fields_to_correct])}
"""
mismatched_records = self.source_db.execute_query(mismatch_query)
for record in mismatched_records:
key_value = record[key_field]
updates = {}
for field in fields_to_correct:
source_value = record[f'source_{field}']
target_value = record[f'target_{field}']
if source_value != target_value:
updates[field] = source_value
if updates:
correction = {
'table': table_name,
'operation': 'UPDATE',
'key': key_value,
'updates': updates,
'timestamp': datetime.utcnow()
}
if not self.dry_run:
try:
self.target_db.update(table_name, {key_field: key_value}, updates)
correction['status'] = 'SUCCESS'
except Exception as e:
correction['status'] = 'FAILED'
correction['error'] = str(e)
else:
correction['status'] = 'DRY_RUN'
self.correction_log.append(correction)
return len(mismatched_records)
```
### 2. Manual Review Process
#### Correction Workflow
```python
class ManualReviewSystem:
def __init__(self, database_client):
self.db = database_client
self.review_queue = []
def queue_for_review(self, discrepancy):
"""Add discrepancy to manual review queue"""
review_item = {
'id': str(uuid.uuid4()),
'discrepancy_type': discrepancy['type'],
'table': discrepancy['table'],
'record_key': discrepancy['key'],
'source_data': discrepancy.get('source_data'),
'target_data': discrepancy.get('target_data'),
'description': discrepancy['description'],
'severity': discrepancy.get('severity', 'medium'),
'status': 'PENDING',
'created_at': datetime.utcnow(),
'reviewed_by': None,
'reviewed_at': None,
'resolution': None
}
self.review_queue.append(review_item)
# Persist to review database
self.db.insert('manual_review_queue', review_item)
return review_item['id']
def process_review(self, review_id, reviewer, action, notes=None):
"""Process manual review decision"""
review_item = self.get_review_item(review_id)
if not review_item:
raise ValueError(f"Review item {review_id} not found")
review_item.update({
'status': 'REVIEWED',
'reviewed_by': reviewer,
'reviewed_at': datetime.utcnow(),
'resolution': {
'action': action, # 'APPLY_SOURCE', 'KEEP_TARGET', 'CUSTOM_FIX'
'notes': notes
}
})
# Apply the resolution
if action == 'APPLY_SOURCE':
self.apply_source_data(review_item)
elif action == 'KEEP_TARGET':
pass # No action needed
elif action == 'CUSTOM_FIX':
# Custom fix would be applied separately
pass
# Update review record
self.db.update('manual_review_queue',
{'id': review_id},
review_item)
return review_item
def generate_review_report(self):
"""Generate summary report of manual reviews"""
reviews = self.db.query("""
SELECT
discrepancy_type,
severity,
status,
COUNT(*) as count,
MIN(created_at) as oldest_review,
MAX(created_at) as newest_review
FROM manual_review_queue
GROUP BY discrepancy_type, severity, status
ORDER BY severity DESC, discrepancy_type
""")
return reviews
```
### 3. Reconciliation Scheduling
#### Automated Reconciliation Jobs
```python
import schedule
import time
from datetime import datetime, timedelta
class ReconciliationScheduler:
def __init__(self, reconciler):
self.reconciler = reconciler
self.job_history = []
def setup_schedules(self):
"""Set up automated reconciliation schedules"""
# Quick reconciliation every 15 minutes during migration
schedule.every(15).minutes.do(self.quick_reconciliation)
# Comprehensive reconciliation every 4 hours
schedule.every(4).hours.do(self.comprehensive_reconciliation)
# Deep validation daily
schedule.every().day.at("02:00").do(self.deep_validation)
# Weekly business logic validation
schedule.every().sunday.at("03:00").do(self.business_logic_validation)
def quick_reconciliation(self):
"""Quick count-based reconciliation"""
job_start = datetime.utcnow()
try:
# Check critical tables only
critical_tables = [
'transactions', 'orders', 'customers', 'accounts'
]
results = []
for table in critical_tables:
count_diff = self.reconciler.check_row_counts(table)
if abs(count_diff) > 0:
results.append({
'table': table,
'count_difference': count_diff,
'severity': 'high' if abs(count_diff) > 100 else 'medium'
})
job_result = {
'job_type': 'quick_reconciliation',
'start_time': job_start,
'end_time': datetime.utcnow(),
'status': 'completed',
'issues_found': len(results),
'details': results
}
# Alert if significant issues found
if any(r['severity'] == 'high' for r in results):
self.send_alert(job_result)
except Exception as e:
job_result = {
'job_type': 'quick_reconciliation',
'start_time': job_start,
'end_time': datetime.utcnow(),
'status': 'failed',
'error': str(e)
}
self.job_history.append(job_result)
def comprehensive_reconciliation(self):
"""Comprehensive checksum-based reconciliation"""
job_start = datetime.utcnow()
try:
tables_to_check = self.get_migration_tables()
issues = []
for table in tables_to_check:
# Sample-based checksum validation
sample_issues = self.reconciler.validate_sample_checksums(
table, sample_size=1000
)
issues.extend(sample_issues)
# Auto-correct simple issues
auto_corrections = 0
for issue in issues:
if issue['auto_correctable']:
self.reconciler.auto_correct_issue(issue)
auto_corrections += 1
else:
# Queue for manual review
self.reconciler.queue_for_manual_review(issue)
job_result = {
'job_type': 'comprehensive_reconciliation',
'start_time': job_start,
'end_time': datetime.utcnow(),
'status': 'completed',
'total_issues': len(issues),
'auto_corrections': auto_corrections,
'manual_reviews_queued': len(issues) - auto_corrections
}
except Exception as e:
job_result = {
'job_type': 'comprehensive_reconciliation',
'start_time': job_start,
'end_time': datetime.utcnow(),
'status': 'failed',
'error': str(e)
}
self.job_history.append(job_result)
def run_scheduler(self):
"""Run the reconciliation scheduler"""
print("Starting reconciliation scheduler...")
while True:
schedule.run_pending()
time.sleep(60) # Check every minute
```
## Monitoring and Reporting
### 1. Reconciliation Metrics
```python
class ReconciliationMetrics:
def __init__(self, prometheus_client):
self.prometheus = prometheus_client
# Define metrics
self.inconsistencies_found = Counter(
'reconciliation_inconsistencies_total',
'Number of inconsistencies found',
['table', 'type', 'severity']
)
self.reconciliation_duration = Histogram(
'reconciliation_duration_seconds',
'Time spent on reconciliation jobs',
['job_type']
)
self.auto_corrections = Counter(
'reconciliation_auto_corrections_total',
'Number of automatically corrected inconsistencies',
['table', 'correction_type']
)
self.data_drift_gauge = Gauge(
'data_drift_percentage',
'Percentage of records with inconsistencies',
['table']
)
def record_inconsistency(self, table, inconsistency_type, severity):
"""Record a found inconsistency"""
self.inconsistencies_found.labels(
table=table,
type=inconsistency_type,
severity=severity
).inc()
def record_auto_correction(self, table, correction_type):
"""Record an automatic correction"""
self.auto_corrections.labels(
table=table,
correction_type=correction_type
).inc()
def update_data_drift(self, table, drift_percentage):
"""Update data drift gauge"""
self.data_drift_gauge.labels(table=table).set(drift_percentage)
def record_job_duration(self, job_type, duration_seconds):
"""Record reconciliation job duration"""
self.reconciliation_duration.labels(job_type=job_type).observe(duration_seconds)
```
### 2. Alerting Rules
```yaml
# Prometheus alerting rules for data reconciliation
groups:
- name: data_reconciliation
rules:
- alert: HighDataInconsistency
expr: reconciliation_inconsistencies_total > 100
for: 5m
labels:
severity: critical
annotations:
summary: "High number of data inconsistencies detected"
description: "{{ $value }} inconsistencies found in the last 5 minutes"
- alert: DataDriftHigh
expr: data_drift_percentage > 5
for: 10m
labels:
severity: warning
annotations:
summary: "Data drift percentage is high"
description: "{{ $labels.table }} has {{ $value }}% data drift"
- alert: ReconciliationJobFailed
expr: up{job="reconciliation"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: "Reconciliation job is down"
description: "The data reconciliation service is not responding"
- alert: AutoCorrectionRateHigh
expr: rate(reconciliation_auto_corrections_total[10m]) > 10
for: 5m
labels:
severity: warning
annotations:
summary: "High rate of automatic corrections"
description: "Auto-correction rate is {{ $value }} per second"
```
### 3. Dashboard and Reporting
```python
class ReconciliationDashboard:
def __init__(self, database_client, metrics_client):
self.db = database_client
self.metrics = metrics_client
def generate_daily_report(self, date=None):
"""Generate daily reconciliation report"""
if not date:
date = datetime.utcnow().date()
# Query reconciliation results for the day
daily_stats = self.db.query("""
SELECT
table_name,
inconsistency_type,
COUNT(*) as count,
AVG(CASE WHEN resolution = 'AUTO_CORRECTED' THEN 1 ELSE 0 END) as auto_correction_rate
FROM reconciliation_log
WHERE DATE(created_at) = %s
GROUP BY table_name, inconsistency_type
""", (date,))
# Generate summary
summary = {
'date': date.isoformat(),
'total_inconsistencies': sum(row['count'] for row in daily_stats),
'auto_correction_rate': sum(row['auto_correction_rate'] * row['count'] for row in daily_stats) / max(sum(row['count'] for row in daily_stats), 1),
'tables_affected': len(set(row['table_name'] for row in daily_stats)),
'details_by_table': {}
}
# Group by table
for row in daily_stats:
table = row['table_name']
if table not in summary['details_by_table']:
summary['details_by_table'][table] = []
summary['details_by_table'][table].append({
'inconsistency_type': row['inconsistency_type'],
'count': row['count'],
'auto_correction_rate': row['auto_correction_rate']
})
return summary
def generate_trend_analysis(self, days=7):
"""Generate trend analysis for reconciliation metrics"""
end_date = datetime.utcnow().date()
start_date = end_date - timedelta(days=days)
trends = self.db.query("""
SELECT
DATE(created_at) as date,
table_name,
COUNT(*) as inconsistencies,
AVG(CASE WHEN resolution = 'AUTO_CORRECTED' THEN 1 ELSE 0 END) as auto_correction_rate
FROM reconciliation_log
WHERE DATE(created_at) BETWEEN %s AND %s
GROUP BY DATE(created_at), table_name
ORDER BY date, table_name
""", (start_date, end_date))
# Calculate trends
trend_analysis = {
'period': f"{start_date} to {end_date}",
'trends': {},
'overall_trend': 'stable'
}
for table in set(row['table_name'] for row in trends):
table_data = [row for row in trends if row['table_name'] == table]
if len(table_data) >= 2:
first_count = table_data[0]['inconsistencies']
last_count = table_data[-1]['inconsistencies']
if last_count > first_count * 1.2:
trend = 'increasing'
elif last_count < first_count * 0.8:
trend = 'decreasing'
else:
trend = 'stable'
trend_analysis['trends'][table] = {
'direction': trend,
'first_day_count': first_count,
'last_day_count': last_count,
'change_percentage': ((last_count - first_count) / max(first_count, 1)) * 100
}
return trend_analysis
```
## Advanced Reconciliation Techniques
### 1. Machine Learning-Based Anomaly Detection
```python
from sklearn.isolation import IsolationForest
from sklearn.preprocessing import StandardScaler
import numpy as np
class MLAnomalyDetector:
def __init__(self):
self.models = {}
self.scalers = {}
def train_anomaly_detector(self, table_name, training_data):
"""Train anomaly detection model for a specific table"""
# Prepare features (convert records to numerical features)
features = self.extract_features(training_data)
# Scale features
scaler = StandardScaler()
scaled_features = scaler.fit_transform(features)
# Train isolation forest
model = IsolationForest(contamination=0.05, random_state=42)
model.fit(scaled_features)
# Store model and scaler
self.models[table_name] = model
self.scalers[table_name] = scaler
def detect_anomalies(self, table_name, data):
"""Detect anomalous records that may indicate reconciliation issues"""
if table_name not in self.models:
raise ValueError(f"No trained model for table {table_name}")
# Extract features
features = self.extract_features(data)
# Scale features
scaled_features = self.scalers[table_name].transform(features)
# Predict anomalies
anomaly_scores = self.models[table_name].decision_function(scaled_features)
anomaly_predictions = self.models[table_name].predict(scaled_features)
# Return anomalous records with scores
anomalies = []
for i, (record, score, is_anomaly) in enumerate(zip(data, anomaly_scores, anomaly_predictions)):
if is_anomaly == -1: # Isolation forest returns -1 for anomalies
anomalies.append({
'record_index': i,
'record': record,
'anomaly_score': score,
'severity': 'high' if score < -0.5 else 'medium'
})
return anomalies
def extract_features(self, data):
"""Extract numerical features from database records"""
features = []
for record in data:
record_features = []
for key, value in record.items():
if isinstance(value, (int, float)):
record_features.append(value)
elif isinstance(value, str):
# Convert string to hash-based feature
record_features.append(hash(value) % 10000)
elif isinstance(value, datetime):
# Convert datetime to timestamp
record_features.append(value.timestamp())
else:
# Default value for other types
record_features.append(0)
features.append(record_features)
return np.array(features)
```
### 2. Probabilistic Reconciliation
```python
import random
from typing import List, Dict, Tuple
class ProbabilisticReconciler:
def __init__(self, confidence_threshold=0.95):
self.confidence_threshold = confidence_threshold
def statistical_sampling_validation(self, table_name: str, population_size: int) -> Dict:
"""Use statistical sampling to validate large datasets"""
# Calculate sample size for 95% confidence, 5% margin of error
confidence_level = 0.95
margin_of_error = 0.05
z_score = 1.96 # for 95% confidence
p = 0.5 # assume 50% error rate for maximum sample size
sample_size = (z_score ** 2 * p * (1 - p)) / (margin_of_error ** 2)
if population_size < 10000:
# Finite population correction
sample_size = sample_size / (1 + (sample_size - 1) / population_size)
sample_size = min(int(sample_size), population_size)
# Generate random sample
sample_ids = self.generate_random_sample(table_name, sample_size)
# Validate sample
sample_results = self.validate_sample_records(table_name, sample_ids)
# Calculate population estimates
error_rate = sample_results['errors'] / sample_size
estimated_errors = int(population_size * error_rate)
# Calculate confidence interval
standard_error = (error_rate * (1 - error_rate) / sample_size) ** 0.5
margin_of_error_actual = z_score * standard_error
confidence_interval = (
max(0, error_rate - margin_of_error_actual),
min(1, error_rate + margin_of_error_actual)
)
return {
'table_name': table_name,
'population_size': population_size,
'sample_size': sample_size,
'sample_error_rate': error_rate,
'estimated_total_errors': estimated_errors,
'confidence_interval': confidence_interval,
'confidence_level': confidence_level,
'recommendation': self.generate_recommendation(error_rate, confidence_interval)
}
def generate_random_sample(self, table_name: str, sample_size: int) -> List[int]:
"""Generate random sample of record IDs"""
# Get total record count and ID range
id_range = self.db.query(f"SELECT MIN(id), MAX(id) FROM {table_name}")[0]
min_id, max_id = id_range
# Generate random IDs
sample_ids = []
attempts = 0
max_attempts = sample_size * 10 # Avoid infinite loop
while len(sample_ids) < sample_size and attempts < max_attempts:
candidate_id = random.randint(min_id, max_id)
# Check if ID exists
exists = self.db.query(f"SELECT 1 FROM {table_name} WHERE id = %s", (candidate_id,))
if exists and candidate_id not in sample_ids:
sample_ids.append(candidate_id)
attempts += 1
return sample_ids
def validate_sample_records(self, table_name: str, sample_ids: List[int]) -> Dict:
"""Validate a sample of records"""
validation_results = {
'total_checked': len(sample_ids),
'errors': 0,
'error_details': []
}
for record_id in sample_ids:
# Get record from both source and target
source_record = self.source_db.get_record(table_name, record_id)
target_record = self.target_db.get_record(table_name, record_id)
if not target_record:
validation_results['errors'] += 1
validation_results['error_details'].append({
'id': record_id,
'error_type': 'MISSING_IN_TARGET'
})
elif not self.records_match(source_record, target_record):
validation_results['errors'] += 1
validation_results['error_details'].append({
'id': record_id,
'error_type': 'DATA_MISMATCH',
'differences': self.find_differences(source_record, target_record)
})
return validation_results
def generate_recommendation(self, error_rate: float, confidence_interval: Tuple[float, float]) -> str:
"""Generate recommendation based on error rate and confidence"""
if confidence_interval[1] < 0.01: # Less than 1% error rate with confidence
return "Data quality is excellent. Continue with normal reconciliation schedule."
elif confidence_interval[1] < 0.05: # Less than 5% error rate with confidence
return "Data quality is acceptable. Monitor closely and investigate sample errors."
elif confidence_interval[0] > 0.1: # More than 10% error rate with confidence
return "Data quality is poor. Immediate comprehensive reconciliation required."
else:
return "Data quality is uncertain. Increase sample size for better estimates."
```
## Performance Optimization
### 1. Parallel Processing
```python
import asyncio
import multiprocessing as mp
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
class ParallelReconciler:
def __init__(self, max_workers=None):
self.max_workers = max_workers or mp.cpu_count()
async def parallel_table_reconciliation(self, tables: List[str]):
"""Reconcile multiple tables in parallel"""
async with asyncio.Semaphore(self.max_workers):
tasks = [
self.reconcile_table_async(table)
for table in tables
]
results = await asyncio.gather(*tasks, return_exceptions=True)
# Process results
summary = {
'total_tables': len(tables),
'successful': 0,
'failed': 0,
'results': {}
}
for table, result in zip(tables, results):
if isinstance(result, Exception):
summary['failed'] += 1
summary['results'][table] = {
'status': 'failed',
'error': str(result)
}
else:
summary['successful'] += 1
summary['results'][table] = result
return summary
def parallel_chunk_processing(self, table_name: str, chunk_size: int = 10000):
"""Process table reconciliation in parallel chunks"""
# Get total record count
total_records = self.db.get_record_count(table_name)
num_chunks = (total_records + chunk_size - 1) // chunk_size
# Create chunk specifications
chunks = []
for i in range(num_chunks):
start_id = i * chunk_size
end_id = min((i + 1) * chunk_size - 1, total_records - 1)
chunks.append({
'table': table_name,
'start_id': start_id,
'end_id': end_id,
'chunk_number': i + 1
})
# Process chunks in parallel
with ProcessPoolExecutor(max_workers=self.max_workers) as executor:
chunk_results = list(executor.map(self.process_chunk, chunks))
# Aggregate results
total_inconsistencies = sum(r['inconsistencies'] for r in chunk_results)
total_corrections = sum(r['corrections'] for r in chunk_results)
return {
'table': table_name,
'total_records': total_records,
'chunks_processed': len(chunks),
'total_inconsistencies': total_inconsistencies,
'total_corrections': total_corrections,
'chunk_details': chunk_results
}
def process_chunk(self, chunk_spec: Dict) -> Dict:
"""Process a single chunk of records"""
# This runs in a separate process
table = chunk_spec['table']
start_id = chunk_spec['start_id']
end_id = chunk_spec['end_id']
# Initialize database connections for this process
local_source_db = SourceDatabase()
local_target_db = TargetDatabase()
# Get records in chunk
source_records = local_source_db.get_records_range(table, start_id, end_id)
target_records = local_target_db.get_records_range(table, start_id, end_id)
# Reconcile chunk
inconsistencies = 0
corrections = 0
for source_record in source_records:
target_record = target_records.get(source_record['id'])
if not target_record:
inconsistencies += 1
# Auto-correct if possible
try:
local_target_db.insert(table, source_record)
corrections += 1
except Exception:
pass # Log error in production
elif not self.records_match(source_record, target_record):
inconsistencies += 1
# Auto-correct field mismatches
try:
updates = self.calculate_updates(source_record, target_record)
local_target_db.update(table, source_record['id'], updates)
corrections += 1
except Exception:
pass # Log error in production
return {
'chunk_number': chunk_spec['chunk_number'],
'start_id': start_id,
'end_id': end_id,
'records_processed': len(source_records),
'inconsistencies': inconsistencies,
'corrections': corrections
}
```
### 2. Incremental Reconciliation
```python
class IncrementalReconciler:
def __init__(self, source_db, target_db):
self.source_db = source_db
self.target_db = target_db
self.last_reconciliation_times = {}
def incremental_reconciliation(self, table_name: str):
"""Reconcile only records changed since last reconciliation"""
last_reconciled = self.get_last_reconciliation_time(table_name)
# Get records modified since last reconciliation
modified_source = self.source_db.get_records_modified_since(
table_name, last_reconciled
)
modified_target = self.target_db.get_records_modified_since(
table_name, last_reconciled
)
# Create lookup dictionaries
source_dict = {r['id']: r for r in modified_source}
target_dict = {r['id']: r for r in modified_target}
# Find all record IDs to check
all_ids = set(source_dict.keys()) | set(target_dict.keys())
inconsistencies = []
for record_id in all_ids:
source_record = source_dict.get(record_id)
target_record = target_dict.get(record_id)
if source_record and not target_record:
inconsistencies.append({
'type': 'missing_in_target',
'table': table_name,
'id': record_id,
'source_record': source_record
})
elif not source_record and target_record:
inconsistencies.append({
'type': 'extra_in_target',
'table': table_name,
'id': record_id,
'target_record': target_record
})
elif source_record and target_record:
if not self.records_match(source_record, target_record):
inconsistencies.append({
'type': 'data_mismatch',
'table': table_name,
'id': record_id,
'source_record': source_record,
'target_record': target_record,
'differences': self.find_differences(source_record, target_record)
})
# Update last reconciliation time
self.update_last_reconciliation_time(table_name, datetime.utcnow())
return {
'table': table_name,
'reconciliation_time': datetime.utcnow(),
'records_checked': len(all_ids),
'inconsistencies_found': len(inconsistencies),
'inconsistencies': inconsistencies
}
def get_last_reconciliation_time(self, table_name: str) -> datetime:
"""Get the last reconciliation timestamp for a table"""
result = self.source_db.query("""
SELECT last_reconciled_at
FROM reconciliation_metadata
WHERE table_name = %s
""", (table_name,))
if result:
return result[0]['last_reconciled_at']
else:
# First time reconciliation - start from beginning of migration
return self.get_migration_start_time()
def update_last_reconciliation_time(self, table_name: str, timestamp: datetime):
"""Update the last reconciliation timestamp"""
self.source_db.execute("""
INSERT INTO reconciliation_metadata (table_name, last_reconciled_at)
VALUES (%s, %s)
ON CONFLICT (table_name)
DO UPDATE SET last_reconciled_at = %s
""", (table_name, timestamp, timestamp))
```
This comprehensive guide provides the framework and tools necessary for implementing robust data reconciliation strategies during migrations, ensuring data integrity and consistency while minimizing business disruption.
FILE:references/migration_patterns_catalog.md
# Migration Patterns Catalog
## Overview
This catalog provides detailed descriptions of proven migration patterns, their use cases, implementation guidelines, and best practices. Each pattern includes code examples, diagrams, and lessons learned from real-world implementations.
## Database Migration Patterns
### 1. Expand-Contract Pattern
**Use Case:** Schema evolution with zero downtime
**Complexity:** Medium
**Risk Level:** Low-Medium
#### Description
The Expand-Contract pattern allows for schema changes without downtime by following a three-phase approach:
1. **Expand:** Add new schema elements alongside existing ones
2. **Migrate:** Dual-write to both old and new schema during transition
3. **Contract:** Remove old schema elements after validation
#### Implementation Steps
```sql
-- Phase 1: Expand
ALTER TABLE users ADD COLUMN email_new VARCHAR(255);
CREATE INDEX CONCURRENTLY idx_users_email_new ON users(email_new);
-- Phase 2: Migrate (Application Code)
-- Write to both columns during transition period
INSERT INTO users (name, email, email_new) VALUES (?, ?, ?);
-- Backfill existing data
UPDATE users SET email_new = email WHERE email_new IS NULL;
-- Phase 3: Contract (after validation)
ALTER TABLE users DROP COLUMN email;
ALTER TABLE users RENAME COLUMN email_new TO email;
```
#### Pros and Cons
**Pros:**
- Zero downtime deployments
- Safe rollback at any point
- Gradual transition with validation
**Cons:**
- Increased storage during transition
- More complex application logic
- Extended migration timeline
### 2. Parallel Schema Pattern
**Use Case:** Major database restructuring
**Complexity:** High
**Risk Level:** Medium
#### Description
Run new and old schemas in parallel, using feature flags to gradually route traffic to the new schema while maintaining the ability to rollback quickly.
#### Implementation Example
```python
class DatabaseRouter:
def __init__(self, feature_flag_service):
self.feature_flags = feature_flag_service
self.old_db = OldDatabaseConnection()
self.new_db = NewDatabaseConnection()
def route_query(self, user_id, query_type):
if self.feature_flags.is_enabled("new_schema", user_id):
return self.new_db.execute(query_type)
else:
return self.old_db.execute(query_type)
def dual_write(self, data):
# Write to both databases for consistency
success_old = self.old_db.write(data)
success_new = self.new_db.write(transform_data(data))
if not (success_old and success_new):
# Handle partial failures
self.handle_dual_write_failure(data, success_old, success_new)
```
#### Best Practices
- Implement data consistency checks between schemas
- Use circuit breakers for automatic failover
- Monitor performance impact of dual writes
- Plan for data reconciliation processes
### 3. Event Sourcing Migration
**Use Case:** Migrating systems with complex business logic
**Complexity:** High
**Risk Level:** Medium-High
#### Description
Capture all changes as events during migration, enabling replay and reconciliation capabilities.
#### Event Store Schema
```sql
CREATE TABLE migration_events (
event_id UUID PRIMARY KEY,
aggregate_id UUID NOT NULL,
event_type VARCHAR(100) NOT NULL,
event_data JSONB NOT NULL,
event_version INTEGER NOT NULL,
occurred_at TIMESTAMP WITH TIME ZONE DEFAULT NOW(),
processed_at TIMESTAMP WITH TIME ZONE
);
```
#### Migration Event Handler
```python
class MigrationEventHandler:
def __init__(self, old_store, new_store):
self.old_store = old_store
self.new_store = new_store
self.event_log = []
def handle_update(self, entity_id, old_data, new_data):
# Log the change as an event
event = MigrationEvent(
entity_id=entity_id,
event_type="entity_migrated",
old_data=old_data,
new_data=new_data,
timestamp=datetime.now()
)
self.event_log.append(event)
# Apply to new store
success = self.new_store.update(entity_id, new_data)
if not success:
# Mark for retry
event.status = "failed"
self.schedule_retry(event)
return success
def replay_events(self, from_timestamp=None):
"""Replay events for reconciliation"""
events = self.get_events_since(from_timestamp)
for event in events:
self.apply_event(event)
```
## Service Migration Patterns
### 1. Strangler Fig Pattern
**Use Case:** Legacy system replacement
**Complexity:** Medium-High
**Risk Level:** Medium
#### Description
Gradually replace legacy functionality by intercepting calls and routing them to new services, eventually "strangling" the legacy system.
#### Implementation Architecture
```yaml
# API Gateway Configuration
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
name: user-service-migration
spec:
http:
- match:
- headers:
migration-flag:
exact: "new"
route:
- destination:
host: user-service-v2
- route:
- destination:
host: user-service-v1
```
#### Strangler Proxy Implementation
```python
class StranglerProxy:
def __init__(self):
self.legacy_service = LegacyUserService()
self.new_service = NewUserService()
self.feature_flags = FeatureFlagService()
def handle_request(self, request):
route = self.determine_route(request)
if route == "new":
return self.handle_with_new_service(request)
elif route == "both":
return self.handle_with_both_services(request)
else:
return self.handle_with_legacy_service(request)
def determine_route(self, request):
user_id = request.get('user_id')
if self.feature_flags.is_enabled("new_user_service", user_id):
if self.feature_flags.is_enabled("dual_write", user_id):
return "both"
else:
return "new"
else:
return "legacy"
```
### 2. Parallel Run Pattern
**Use Case:** Risk mitigation for critical services
**Complexity:** Medium
**Risk Level:** Low-Medium
#### Description
Run both old and new services simultaneously, comparing outputs to validate correctness before switching traffic.
#### Implementation
```python
class ParallelRunManager:
def __init__(self):
self.primary_service = PrimaryService()
self.candidate_service = CandidateService()
self.comparator = ResponseComparator()
self.metrics = MetricsCollector()
async def parallel_execute(self, request):
# Execute both services concurrently
primary_task = asyncio.create_task(
self.primary_service.process(request)
)
candidate_task = asyncio.create_task(
self.candidate_service.process(request)
)
# Always wait for primary
primary_result = await primary_task
try:
# Wait for candidate with timeout
candidate_result = await asyncio.wait_for(
candidate_task, timeout=5.0
)
# Compare results
comparison = self.comparator.compare(
primary_result, candidate_result
)
# Record metrics
self.metrics.record_comparison(comparison)
except asyncio.TimeoutError:
self.metrics.record_timeout("candidate")
except Exception as e:
self.metrics.record_error("candidate", str(e))
# Always return primary result
return primary_result
```
### 3. Blue-Green Deployment Pattern
**Use Case:** Zero-downtime service updates
**Complexity:** Low-Medium
**Risk Level:** Low
#### Description
Maintain two identical production environments (blue and green), switching traffic between them for deployments.
#### Kubernetes Implementation
```yaml
# Blue Deployment
apiVersion: apps/v1
kind: Deployment
metadata:
name: app-blue
labels:
version: blue
spec:
replicas: 3
selector:
matchLabels:
app: myapp
version: blue
template:
metadata:
labels:
app: myapp
version: blue
spec:
containers:
- name: app
image: myapp:v1.0.0
---
# Green Deployment
apiVersion: apps/v1
kind: Deployment
metadata:
name: app-green
labels:
version: green
spec:
replicas: 3
selector:
matchLabels:
app: myapp
version: green
template:
metadata:
labels:
app: myapp
version: green
spec:
containers:
- name: app
image: myapp:v2.0.0
---
# Service (switches between blue and green)
apiVersion: v1
kind: Service
metadata:
name: app-service
spec:
selector:
app: myapp
version: blue # Change to green for deployment
ports:
- port: 80
targetPort: 8080
```
## Infrastructure Migration Patterns
### 1. Lift and Shift Pattern
**Use Case:** Quick cloud migration with minimal changes
**Complexity:** Low-Medium
**Risk Level:** Low
#### Description
Migrate applications to cloud infrastructure with minimal or no code changes, focusing on infrastructure compatibility.
#### Migration Checklist
```yaml
Pre-Migration Assessment:
- inventory_current_infrastructure:
- servers_and_specifications
- network_configuration
- storage_requirements
- security_configurations
- identify_dependencies:
- database_connections
- external_service_integrations
- file_system_dependencies
- assess_compatibility:
- operating_system_versions
- runtime_dependencies
- license_requirements
Migration Execution:
- provision_target_infrastructure:
- compute_instances
- storage_volumes
- network_configuration
- security_groups
- migrate_data:
- database_backup_restore
- file_system_replication
- configuration_files
- update_configurations:
- connection_strings
- environment_variables
- dns_records
- validate_functionality:
- application_health_checks
- end_to_end_testing
- performance_validation
```
### 2. Hybrid Cloud Migration
**Use Case:** Gradual cloud adoption with on-premises integration
**Complexity:** High
**Risk Level:** Medium-High
#### Description
Maintain some components on-premises while migrating others to cloud, requiring secure connectivity and data synchronization.
#### Network Architecture
```hcl
# Terraform configuration for hybrid connectivity
resource "aws_vpc" "main" {
cidr_block = "10.0.0.0/16"
enable_dns_hostnames = true
enable_dns_support = true
}
resource "aws_vpn_gateway" "main" {
vpc_id = aws_vpc.main.id
tags = {
Name = "hybrid-vpn-gateway"
}
}
resource "aws_customer_gateway" "main" {
bgp_asn = 65000
ip_address = var.on_premises_public_ip
type = "ipsec.1"
tags = {
Name = "on-premises-gateway"
}
}
resource "aws_vpn_connection" "main" {
vpn_gateway_id = aws_vpn_gateway.main.id
customer_gateway_id = aws_customer_gateway.main.id
type = "ipsec.1"
static_routes_only = true
}
```
#### Data Synchronization Pattern
```python
class HybridDataSync:
def __init__(self):
self.on_prem_db = OnPremiseDatabase()
self.cloud_db = CloudDatabase()
self.sync_log = SyncLogManager()
async def bidirectional_sync(self):
"""Synchronize data between on-premises and cloud"""
# Get last sync timestamp
last_sync = self.sync_log.get_last_sync_time()
# Sync on-prem changes to cloud
on_prem_changes = self.on_prem_db.get_changes_since(last_sync)
for change in on_prem_changes:
await self.apply_change_to_cloud(change)
# Sync cloud changes to on-prem
cloud_changes = self.cloud_db.get_changes_since(last_sync)
for change in cloud_changes:
await self.apply_change_to_on_prem(change)
# Handle conflicts
conflicts = self.detect_conflicts(on_prem_changes, cloud_changes)
for conflict in conflicts:
await self.resolve_conflict(conflict)
# Update sync timestamp
self.sync_log.record_sync_completion()
async def apply_change_to_cloud(self, change):
"""Apply on-premises change to cloud database"""
try:
if change.operation == "INSERT":
await self.cloud_db.insert(change.table, change.data)
elif change.operation == "UPDATE":
await self.cloud_db.update(change.table, change.key, change.data)
elif change.operation == "DELETE":
await self.cloud_db.delete(change.table, change.key)
self.sync_log.record_success(change.id, "cloud")
except Exception as e:
self.sync_log.record_failure(change.id, "cloud", str(e))
raise
```
### 3. Multi-Cloud Migration
**Use Case:** Avoiding vendor lock-in or regulatory requirements
**Complexity:** Very High
**Risk Level:** High
#### Description
Distribute workloads across multiple cloud providers for resilience, compliance, or cost optimization.
#### Service Mesh Configuration
```yaml
# Istio configuration for multi-cloud service mesh
apiVersion: networking.istio.io/v1beta1
kind: ServiceEntry
metadata:
name: aws-service
spec:
hosts:
- aws-service.company.com
ports:
- number: 443
name: https
protocol: HTTPS
location: MESH_EXTERNAL
resolution: DNS
---
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
name: multi-cloud-routing
spec:
hosts:
- user-service
http:
- match:
- headers:
region:
exact: "us-east"
route:
- destination:
host: aws-service.company.com
weight: 100
- match:
- headers:
region:
exact: "eu-west"
route:
- destination:
host: gcp-service.company.com
weight: 100
- route: # Default routing
- destination:
host: user-service
subset: local
weight: 80
- destination:
host: aws-service.company.com
weight: 20
```
## Feature Flag Patterns
### 1. Progressive Rollout Pattern
**Use Case:** Gradual feature deployment with risk mitigation
**Implementation:**
```python
class ProgressiveRollout:
def __init__(self, feature_name):
self.feature_name = feature_name
self.rollout_percentage = 0
self.user_buckets = {}
def is_enabled_for_user(self, user_id):
# Consistent user bucketing
user_hash = hashlib.md5(f"{self.feature_name}:{user_id}".encode()).hexdigest()
bucket = int(user_hash, 16) % 100
return bucket < self.rollout_percentage
def increase_rollout(self, target_percentage, step_size=10):
"""Gradually increase rollout percentage"""
while self.rollout_percentage < target_percentage:
self.rollout_percentage = min(
self.rollout_percentage + step_size,
target_percentage
)
# Monitor metrics before next increase
yield self.rollout_percentage
time.sleep(300) # Wait 5 minutes between increases
```
### 2. Circuit Breaker Pattern
**Use Case:** Automatic fallback during migration issues
```python
class MigrationCircuitBreaker:
def __init__(self, failure_threshold=5, timeout=60):
self.failure_count = 0
self.failure_threshold = failure_threshold
self.timeout = timeout
self.last_failure_time = None
self.state = 'CLOSED' # CLOSED, OPEN, HALF_OPEN
def call_new_service(self, request):
if self.state == 'OPEN':
if self.should_attempt_reset():
self.state = 'HALF_OPEN'
else:
return self.fallback_to_legacy(request)
try:
response = self.new_service.process(request)
self.on_success()
return response
except Exception as e:
self.on_failure()
return self.fallback_to_legacy(request)
def on_success(self):
self.failure_count = 0
self.state = 'CLOSED'
def on_failure(self):
self.failure_count += 1
self.last_failure_time = time.time()
if self.failure_count >= self.failure_threshold:
self.state = 'OPEN'
def should_attempt_reset(self):
return (time.time() - self.last_failure_time) >= self.timeout
```
## Migration Anti-Patterns
### 1. Big Bang Migration (Anti-Pattern)
**Why to Avoid:**
- High risk of complete system failure
- Difficult to rollback
- Extended downtime
- All-or-nothing deployment
**Better Alternative:** Use incremental migration patterns like Strangler Fig or Parallel Run.
### 2. No Rollback Plan (Anti-Pattern)
**Why to Avoid:**
- Cannot recover from failures
- Increases business risk
- Panic-driven decisions during issues
**Better Alternative:** Always implement comprehensive rollback procedures before migration.
### 3. Insufficient Testing (Anti-Pattern)
**Why to Avoid:**
- Unknown compatibility issues
- Performance degradation
- Data corruption risks
**Better Alternative:** Implement comprehensive testing at each migration phase.
## Pattern Selection Matrix
| Migration Type | Complexity | Downtime Tolerance | Recommended Pattern |
|---------------|------------|-------------------|-------------------|
| Schema Change | Low | Zero | Expand-Contract |
| Schema Change | High | Zero | Parallel Schema |
| Service Replace | Medium | Zero | Strangler Fig |
| Service Update | Low | Zero | Blue-Green |
| Data Migration | High | Some | Event Sourcing |
| Infrastructure | Low | Some | Lift and Shift |
| Infrastructure | High | Zero | Hybrid Cloud |
## Success Metrics
### Technical Metrics
- Migration completion rate
- System availability during migration
- Performance impact (response time, throughput)
- Error rate changes
- Rollback execution time
### Business Metrics
- Customer impact score
- Revenue protection
- Time to value realization
- Stakeholder satisfaction
### Operational Metrics
- Team efficiency
- Knowledge transfer effectiveness
- Post-migration support requirements
- Documentation completeness
## Lessons Learned
### Common Pitfalls
1. **Underestimating data dependencies** - Always map all data relationships
2. **Insufficient monitoring** - Implement comprehensive observability before migration
3. **Poor communication** - Keep all stakeholders informed throughout the process
4. **Rushed timelines** - Allow adequate time for testing and validation
5. **Ignoring performance impact** - Benchmark before and after migration
### Best Practices
1. **Start with low-risk migrations** - Build confidence and experience
2. **Automate everything possible** - Reduce human error and increase repeatability
3. **Test rollback procedures** - Ensure you can recover from any failure
4. **Monitor continuously** - Use real-time dashboards and alerting
5. **Document everything** - Create comprehensive runbooks and documentation
This catalog serves as a reference for selecting appropriate migration patterns based on specific requirements, risk tolerance, and technical constraints.
FILE:references/zero_downtime_techniques.md
# Zero-Downtime Migration Techniques
## Overview
Zero-downtime migrations are critical for maintaining business continuity and user experience during system changes. This guide provides comprehensive techniques, patterns, and implementation strategies for achieving true zero-downtime migrations across different system components.
## Core Principles
### 1. Backward Compatibility
Every change must be backward compatible until all clients have migrated to the new version.
### 2. Incremental Changes
Break large changes into smaller, independent increments that can be deployed and validated separately.
### 3. Feature Flags
Use feature toggles to control the rollout of new functionality without code deployments.
### 4. Graceful Degradation
Ensure systems continue to function even when some components are unavailable or degraded.
## Database Zero-Downtime Techniques
### Schema Evolution Without Downtime
#### 1. Additive Changes Only
**Principle:** Only add new elements; never remove or modify existing ones directly.
```sql
-- ✅ Good: Additive change
ALTER TABLE users ADD COLUMN middle_name VARCHAR(50);
-- ❌ Bad: Breaking change
ALTER TABLE users DROP COLUMN email;
```
#### 2. Multi-Phase Schema Evolution
**Phase 1: Expand**
```sql
-- Add new column alongside existing one
ALTER TABLE users ADD COLUMN email_address VARCHAR(255);
-- Add index concurrently (PostgreSQL)
CREATE INDEX CONCURRENTLY idx_users_email_address ON users(email_address);
```
**Phase 2: Dual Write (Application Code)**
```python
class UserService:
def create_user(self, name, email):
# Write to both old and new columns
user = User(
name=name,
email=email, # Old column
email_address=email # New column
)
return user.save()
def update_email(self, user_id, new_email):
# Update both columns
user = User.objects.get(id=user_id)
user.email = new_email
user.email_address = new_email
user.save()
return user
```
**Phase 3: Backfill Data**
```sql
-- Backfill existing data (in batches)
UPDATE users
SET email_address = email
WHERE email_address IS NULL
AND id BETWEEN ? AND ?;
```
**Phase 4: Switch Reads**
```python
class UserService:
def get_user_email(self, user_id):
user = User.objects.get(id=user_id)
# Switch to reading from new column
return user.email_address or user.email
```
**Phase 5: Contract**
```sql
-- After validation, remove old column
ALTER TABLE users DROP COLUMN email;
-- Rename new column if needed
ALTER TABLE users RENAME COLUMN email_address TO email;
```
### 3. Online Schema Changes
#### PostgreSQL Techniques
```sql
-- Safe column addition
ALTER TABLE orders ADD COLUMN status_new VARCHAR(20) DEFAULT 'pending';
-- Safe index creation
CREATE INDEX CONCURRENTLY idx_orders_status_new ON orders(status_new);
-- Safe constraint addition (after data validation)
ALTER TABLE orders ADD CONSTRAINT check_status_new
CHECK (status_new IN ('pending', 'processing', 'completed', 'cancelled'));
```
#### MySQL Techniques
```sql
-- Use pt-online-schema-change for large tables
pt-online-schema-change \
--alter "ADD COLUMN status VARCHAR(20) DEFAULT 'pending'" \
--execute \
D=mydb,t=orders
-- Online DDL (MySQL 5.6+)
ALTER TABLE orders
ADD COLUMN priority INT DEFAULT 1,
ALGORITHM=INPLACE,
LOCK=NONE;
```
### 4. Data Migration Strategies
#### Chunked Data Migration
```python
class DataMigrator:
def __init__(self, source_table, target_table, chunk_size=1000):
self.source_table = source_table
self.target_table = target_table
self.chunk_size = chunk_size
def migrate_data(self):
last_id = 0
total_migrated = 0
while True:
# Get next chunk
chunk = self.get_chunk(last_id, self.chunk_size)
if not chunk:
break
# Transform and migrate chunk
for record in chunk:
transformed = self.transform_record(record)
self.insert_or_update(transformed)
last_id = chunk[-1]['id']
total_migrated += len(chunk)
# Brief pause to avoid overwhelming the database
time.sleep(0.1)
self.log_progress(total_migrated)
return total_migrated
def get_chunk(self, last_id, limit):
return db.execute(f"""
SELECT * FROM {self.source_table}
WHERE id > %s
ORDER BY id
LIMIT %s
""", (last_id, limit))
```
#### Change Data Capture (CDC)
```python
class CDCProcessor:
def __init__(self):
self.kafka_consumer = KafkaConsumer('db_changes')
self.target_db = TargetDatabase()
def process_changes(self):
for message in self.kafka_consumer:
change = json.loads(message.value)
if change['operation'] == 'INSERT':
self.handle_insert(change)
elif change['operation'] == 'UPDATE':
self.handle_update(change)
elif change['operation'] == 'DELETE':
self.handle_delete(change)
def handle_insert(self, change):
transformed_data = self.transform_data(change['after'])
self.target_db.insert(change['table'], transformed_data)
def handle_update(self, change):
key = change['key']
transformed_data = self.transform_data(change['after'])
self.target_db.update(change['table'], key, transformed_data)
```
## Application Zero-Downtime Techniques
### 1. Blue-Green Deployments
#### Infrastructure Setup
```yaml
# Blue Environment (Current Production)
apiVersion: apps/v1
kind: Deployment
metadata:
name: app-blue
labels:
version: blue
app: myapp
spec:
replicas: 3
selector:
matchLabels:
app: myapp
version: blue
template:
metadata:
labels:
app: myapp
version: blue
spec:
containers:
- name: app
image: myapp:1.0.0
ports:
- containerPort: 8080
readinessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 15
periodSeconds: 10
---
# Green Environment (New Version)
apiVersion: apps/v1
kind: Deployment
metadata:
name: app-green
labels:
version: green
app: myapp
spec:
replicas: 3
selector:
matchLabels:
app: myapp
version: green
template:
metadata:
labels:
app: myapp
version: green
spec:
containers:
- name: app
image: myapp:2.0.0
ports:
- containerPort: 8080
readinessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
```
#### Service Switching
```yaml
# Service (switches between blue and green)
apiVersion: v1
kind: Service
metadata:
name: app-service
spec:
selector:
app: myapp
version: blue # Switch to 'green' for deployment
ports:
- port: 80
targetPort: 8080
type: LoadBalancer
```
#### Automated Deployment Script
```bash
#!/bin/bash
# Blue-Green Deployment Script
NAMESPACE="production"
APP_NAME="myapp"
NEW_IMAGE="myapp:2.0.0"
# Determine current and target environments
CURRENT_VERSION=$(kubectl get service $APP_NAME-service -o jsonpath='{.spec.selector.version}')
if [ "$CURRENT_VERSION" = "blue" ]; then
TARGET_VERSION="green"
else
TARGET_VERSION="blue"
fi
echo "Current version: $CURRENT_VERSION"
echo "Target version: $TARGET_VERSION"
# Update target environment with new image
kubectl set image deployment/$APP_NAME-$TARGET_VERSION app=$NEW_IMAGE
# Wait for rollout to complete
kubectl rollout status deployment/$APP_NAME-$TARGET_VERSION --timeout=300s
# Run health checks
echo "Running health checks..."
TARGET_IP=$(kubectl get service $APP_NAME-$TARGET_VERSION -o jsonpath='{.status.loadBalancer.ingress[0].ip}')
for i in {1..30}; do
if curl -f http://$TARGET_IP/health; then
echo "Health check passed"
break
fi
if [ $i -eq 30 ]; then
echo "Health check failed after 30 attempts"
exit 1
fi
sleep 2
done
# Switch traffic to new version
kubectl patch service $APP_NAME-service -p '{"spec":{"selector":{"version":"'$TARGET_VERSION'"}}}'
echo "Traffic switched to $TARGET_VERSION"
# Monitor for 5 minutes
echo "Monitoring new version..."
sleep 300
# Check if rollback is needed
ERROR_RATE=$(curl -s "http://monitoring.company.com/api/error_rate?service=$APP_NAME" | jq '.error_rate')
if (( $(echo "$ERROR_RATE > 0.05" | bc -l) )); then
echo "Error rate too high ($ERROR_RATE), rolling back..."
kubectl patch service $APP_NAME-service -p '{"spec":{"selector":{"version":"'$CURRENT_VERSION'"}}}'
exit 1
fi
echo "Deployment successful!"
```
### 2. Canary Deployments
#### Progressive Canary with Istio
```yaml
# Destination Rule
apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata:
name: myapp-destination
spec:
host: myapp
subsets:
- name: v1
labels:
version: v1
- name: v2
labels:
version: v2
---
# Virtual Service for Canary
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
name: myapp-canary
spec:
hosts:
- myapp
http:
- match:
- headers:
canary:
exact: "true"
route:
- destination:
host: myapp
subset: v2
- route:
- destination:
host: myapp
subset: v1
weight: 95
- destination:
host: myapp
subset: v2
weight: 5
```
#### Automated Canary Controller
```python
class CanaryController:
def __init__(self, istio_client, prometheus_client):
self.istio = istio_client
self.prometheus = prometheus_client
self.canary_weight = 5
self.max_weight = 100
self.weight_increment = 5
self.validation_window = 300 # 5 minutes
async def deploy_canary(self, app_name, new_version):
"""Deploy new version using canary strategy"""
# Start with small percentage
await self.update_traffic_split(app_name, self.canary_weight)
while self.canary_weight < self.max_weight:
# Monitor metrics for validation window
await asyncio.sleep(self.validation_window)
# Check canary health
if not await self.is_canary_healthy(app_name, new_version):
await self.rollback_canary(app_name)
raise Exception("Canary deployment failed health checks")
# Increase traffic to canary
self.canary_weight = min(
self.canary_weight + self.weight_increment,
self.max_weight
)
await self.update_traffic_split(app_name, self.canary_weight)
print(f"Canary traffic increased to {self.canary_weight}%")
print("Canary deployment completed successfully")
async def is_canary_healthy(self, app_name, version):
"""Check if canary version is healthy"""
# Check error rate
error_rate = await self.prometheus.query(
f'rate(http_requests_total{{app="{app_name}", version="{version}", status=~"5.."}}'
f'[5m]) / rate(http_requests_total{{app="{app_name}", version="{version}"}}[5m])'
)
if error_rate > 0.05: # 5% error rate threshold
return False
# Check response time
p95_latency = await self.prometheus.query(
f'histogram_quantile(0.95, rate(http_request_duration_seconds_bucket'
f'{{app="{app_name}", version="{version}"}}[5m]))'
)
if p95_latency > 2.0: # 2 second p95 threshold
return False
return True
async def update_traffic_split(self, app_name, canary_weight):
"""Update Istio virtual service with new traffic split"""
stable_weight = 100 - canary_weight
virtual_service = {
"apiVersion": "networking.istio.io/v1beta1",
"kind": "VirtualService",
"metadata": {"name": f"{app_name}-canary"},
"spec": {
"hosts": [app_name],
"http": [{
"route": [
{
"destination": {"host": app_name, "subset": "stable"},
"weight": stable_weight
},
{
"destination": {"host": app_name, "subset": "canary"},
"weight": canary_weight
}
]
}]
}
}
await self.istio.apply_virtual_service(virtual_service)
```
### 3. Rolling Updates
#### Kubernetes Rolling Update Strategy
```yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: rolling-update-app
spec:
replicas: 10
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 2 # Can have 2 extra pods during update
maxUnavailable: 1 # At most 1 pod can be unavailable
selector:
matchLabels:
app: rolling-update-app
template:
metadata:
labels:
app: rolling-update-app
spec:
containers:
- name: app
image: myapp:2.0.0
ports:
- containerPort: 8080
readinessProbe:
httpGet:
path: /ready
port: 8080
initialDelaySeconds: 5
periodSeconds: 2
timeoutSeconds: 1
successThreshold: 1
failureThreshold: 3
livenessProbe:
httpGet:
path: /live
port: 8080
initialDelaySeconds: 10
periodSeconds: 10
```
#### Custom Rolling Update Controller
```python
class RollingUpdateController:
def __init__(self, k8s_client):
self.k8s = k8s_client
self.max_surge = 2
self.max_unavailable = 1
async def rolling_update(self, deployment_name, new_image):
"""Perform rolling update with custom logic"""
deployment = await self.k8s.get_deployment(deployment_name)
total_replicas = deployment.spec.replicas
# Calculate batch size
batch_size = min(self.max_surge, total_replicas // 5) # Update 20% at a time
updated_pods = []
for i in range(0, total_replicas, batch_size):
batch_end = min(i + batch_size, total_replicas)
# Update batch of pods
for pod_index in range(i, batch_end):
old_pod = await self.get_pod_by_index(deployment_name, pod_index)
# Create new pod with new image
new_pod = await self.create_updated_pod(old_pod, new_image)
# Wait for new pod to be ready
await self.wait_for_pod_ready(new_pod.metadata.name)
# Remove old pod
await self.k8s.delete_pod(old_pod.metadata.name)
updated_pods.append(new_pod)
# Brief pause between pod updates
await asyncio.sleep(2)
# Validate batch health before continuing
if not await self.validate_batch_health(updated_pods[-batch_size:]):
# Rollback batch
await self.rollback_batch(updated_pods[-batch_size:])
raise Exception("Rolling update failed validation")
print(f"Updated {batch_end}/{total_replicas} pods")
print("Rolling update completed successfully")
```
## Load Balancer and Traffic Management
### 1. Weighted Routing
#### NGINX Configuration
```nginx
upstream backend {
# Old version - 80% traffic
server old-app-1:8080 weight=4;
server old-app-2:8080 weight=4;
# New version - 20% traffic
server new-app-1:8080 weight=1;
server new-app-2:8080 weight=1;
}
server {
listen 80;
location / {
proxy_pass http://backend;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
# Health check headers
proxy_set_header X-Health-Check-Timeout 5s;
}
}
```
#### HAProxy Configuration
```haproxy
backend app_servers
balance roundrobin
option httpchk GET /health
# Old version servers
server old-app-1 old-app-1:8080 check weight 80
server old-app-2 old-app-2:8080 check weight 80
# New version servers
server new-app-1 new-app-1:8080 check weight 20
server new-app-2 new-app-2:8080 check weight 20
frontend app_frontend
bind *:80
default_backend app_servers
# Custom health check endpoint
acl health_check path_beg /health
http-request return status 200 content-type text/plain string "OK" if health_check
```
### 2. Circuit Breaker Implementation
```python
class CircuitBreaker:
def __init__(self, failure_threshold=5, recovery_timeout=60, expected_exception=Exception):
self.failure_threshold = failure_threshold
self.recovery_timeout = recovery_timeout
self.expected_exception = expected_exception
self.failure_count = 0
self.last_failure_time = None
self.state = 'CLOSED' # CLOSED, OPEN, HALF_OPEN
def call(self, func, *args, **kwargs):
"""Execute function with circuit breaker protection"""
if self.state == 'OPEN':
if self._should_attempt_reset():
self.state = 'HALF_OPEN'
else:
raise CircuitBreakerOpenException("Circuit breaker is OPEN")
try:
result = func(*args, **kwargs)
self._on_success()
return result
except self.expected_exception as e:
self._on_failure()
raise
def _should_attempt_reset(self):
return (
self.last_failure_time and
time.time() - self.last_failure_time >= self.recovery_timeout
)
def _on_success(self):
self.failure_count = 0
self.state = 'CLOSED'
def _on_failure(self):
self.failure_count += 1
self.last_failure_time = time.time()
if self.failure_count >= self.failure_threshold:
self.state = 'OPEN'
# Usage with service migration
@CircuitBreaker(failure_threshold=3, recovery_timeout=30)
def call_new_service(request):
return new_service.process(request)
def handle_request(request):
try:
return call_new_service(request)
except CircuitBreakerOpenException:
# Fallback to old service
return old_service.process(request)
```
## Monitoring and Validation
### 1. Health Check Implementation
```python
class HealthChecker:
def __init__(self):
self.checks = []
def add_check(self, name, check_func, timeout=5):
self.checks.append({
'name': name,
'func': check_func,
'timeout': timeout
})
async def run_checks(self):
"""Run all health checks and return status"""
results = {}
overall_status = 'healthy'
for check in self.checks:
try:
result = await asyncio.wait_for(
check['func'](),
timeout=check['timeout']
)
results[check['name']] = {
'status': 'healthy',
'result': result
}
except asyncio.TimeoutError:
results[check['name']] = {
'status': 'unhealthy',
'error': 'timeout'
}
overall_status = 'unhealthy'
except Exception as e:
results[check['name']] = {
'status': 'unhealthy',
'error': str(e)
}
overall_status = 'unhealthy'
return {
'status': overall_status,
'checks': results,
'timestamp': datetime.utcnow().isoformat()
}
# Example health checks
health_checker = HealthChecker()
async def database_check():
"""Check database connectivity"""
result = await db.execute("SELECT 1")
return result is not None
async def external_api_check():
"""Check external API availability"""
response = await http_client.get("https://api.example.com/health")
return response.status_code == 200
async def memory_check():
"""Check memory usage"""
memory_usage = psutil.virtual_memory().percent
if memory_usage > 90:
raise Exception(f"Memory usage too high: {memory_usage}%")
return f"Memory usage: {memory_usage}%"
health_checker.add_check("database", database_check)
health_checker.add_check("external_api", external_api_check)
health_checker.add_check("memory", memory_check)
```
### 2. Readiness vs Liveness Probes
```yaml
# Kubernetes Pod with proper health checks
apiVersion: v1
kind: Pod
metadata:
name: app-pod
spec:
containers:
- name: app
image: myapp:2.0.0
ports:
- containerPort: 8080
# Readiness probe - determines if pod should receive traffic
readinessProbe:
httpGet:
path: /ready
port: 8080
initialDelaySeconds: 5
periodSeconds: 3
timeoutSeconds: 2
successThreshold: 1
failureThreshold: 3
# Liveness probe - determines if pod should be restarted
livenessProbe:
httpGet:
path: /live
port: 8080
initialDelaySeconds: 30
periodSeconds: 10
timeoutSeconds: 5
successThreshold: 1
failureThreshold: 3
# Startup probe - gives app time to start before other probes
startupProbe:
httpGet:
path: /startup
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
timeoutSeconds: 3
successThreshold: 1
failureThreshold: 30 # Allow up to 150 seconds for startup
```
### 3. Metrics and Alerting
```python
class MigrationMetrics:
def __init__(self, prometheus_client):
self.prometheus = prometheus_client
# Define custom metrics
self.migration_progress = Counter(
'migration_progress_total',
'Total migration operations completed',
['operation', 'status']
)
self.migration_duration = Histogram(
'migration_operation_duration_seconds',
'Time spent on migration operations',
['operation']
)
self.system_health = Gauge(
'system_health_score',
'Overall system health score (0-1)',
['component']
)
self.traffic_split = Gauge(
'traffic_split_percentage',
'Percentage of traffic going to each version',
['version']
)
def record_migration_step(self, operation, status, duration=None):
"""Record completion of a migration step"""
self.migration_progress.labels(operation=operation, status=status).inc()
if duration:
self.migration_duration.labels(operation=operation).observe(duration)
def update_health_score(self, component, score):
"""Update health score for a component"""
self.system_health.labels(component=component).set(score)
def update_traffic_split(self, version_weights):
"""Update traffic split metrics"""
for version, weight in version_weights.items():
self.traffic_split.labels(version=version).set(weight)
# Usage in migration
metrics = MigrationMetrics(prometheus_client)
def perform_migration_step(operation):
start_time = time.time()
try:
# Perform migration operation
result = execute_migration_operation(operation)
# Record success
duration = time.time() - start_time
metrics.record_migration_step(operation, 'success', duration)
return result
except Exception as e:
# Record failure
duration = time.time() - start_time
metrics.record_migration_step(operation, 'failure', duration)
raise
```
## Rollback Strategies
### 1. Immediate Rollback Triggers
```python
class AutoRollbackSystem:
def __init__(self, metrics_client, deployment_client):
self.metrics = metrics_client
self.deployment = deployment_client
self.rollback_triggers = {
'error_rate_spike': {
'threshold': 0.05, # 5% error rate
'window': 300, # 5 minutes
'auto_rollback': True
},
'latency_increase': {
'threshold': 2.0, # 2x baseline latency
'window': 600, # 10 minutes
'auto_rollback': False # Manual confirmation required
},
'availability_drop': {
'threshold': 0.95, # Below 95% availability
'window': 120, # 2 minutes
'auto_rollback': True
}
}
async def monitor_and_rollback(self, deployment_name):
"""Monitor deployment and trigger rollback if needed"""
while True:
for trigger_name, config in self.rollback_triggers.items():
if await self.check_trigger(trigger_name, config):
if config['auto_rollback']:
await self.execute_rollback(deployment_name, trigger_name)
else:
await self.alert_for_manual_rollback(deployment_name, trigger_name)
await asyncio.sleep(30) # Check every 30 seconds
async def check_trigger(self, trigger_name, config):
"""Check if rollback trigger condition is met"""
current_value = await self.metrics.get_current_value(trigger_name)
baseline_value = await self.metrics.get_baseline_value(trigger_name)
if trigger_name == 'error_rate_spike':
return current_value > config['threshold']
elif trigger_name == 'latency_increase':
return current_value > baseline_value * config['threshold']
elif trigger_name == 'availability_drop':
return current_value < config['threshold']
return False
async def execute_rollback(self, deployment_name, reason):
"""Execute automatic rollback"""
print(f"Executing automatic rollback for {deployment_name}. Reason: {reason}")
# Get previous revision
previous_revision = await self.deployment.get_previous_revision(deployment_name)
# Perform rollback
await self.deployment.rollback_to_revision(deployment_name, previous_revision)
# Notify stakeholders
await self.notify_rollback_executed(deployment_name, reason)
```
### 2. Data Rollback Strategies
```sql
-- Point-in-time recovery setup
-- Create restore point before migration
SELECT pg_create_restore_point('pre_migration_' || to_char(now(), 'YYYYMMDD_HH24MISS'));
-- Rollback using point-in-time recovery
-- (This would be executed on a separate recovery instance)
-- recovery.conf:
-- recovery_target_name = 'pre_migration_20240101_120000'
-- recovery_target_action = 'promote'
```
```python
class DataRollbackManager:
def __init__(self, database_client, backup_service):
self.db = database_client
self.backup = backup_service
async def create_rollback_point(self, migration_id):
"""Create a rollback point before migration"""
rollback_point = {
'migration_id': migration_id,
'timestamp': datetime.utcnow(),
'backup_location': None,
'schema_snapshot': None
}
# Create database backup
backup_path = await self.backup.create_backup(
f"pre_migration_{migration_id}_{int(time.time())}"
)
rollback_point['backup_location'] = backup_path
# Capture schema snapshot
schema_snapshot = await self.capture_schema_snapshot()
rollback_point['schema_snapshot'] = schema_snapshot
# Store rollback point metadata
await self.store_rollback_metadata(rollback_point)
return rollback_point
async def execute_rollback(self, migration_id):
"""Execute data rollback to specified point"""
rollback_point = await self.get_rollback_metadata(migration_id)
if not rollback_point:
raise Exception(f"No rollback point found for migration {migration_id}")
# Stop application traffic
await self.stop_application_traffic()
try:
# Restore from backup
await self.backup.restore_from_backup(
rollback_point['backup_location']
)
# Validate data integrity
await self.validate_data_integrity(
rollback_point['schema_snapshot']
)
# Update application configuration
await self.update_application_config(rollback_point)
# Resume application traffic
await self.resume_application_traffic()
print(f"Data rollback completed successfully for migration {migration_id}")
except Exception as e:
# If rollback fails, we have a serious problem
await self.escalate_rollback_failure(migration_id, str(e))
raise
```
## Best Practices Summary
### 1. Pre-Migration Checklist
- [ ] Comprehensive backup strategy in place
- [ ] Rollback procedures tested in staging
- [ ] Monitoring and alerting configured
- [ ] Health checks implemented
- [ ] Feature flags configured
- [ ] Team communication plan established
- [ ] Load balancer configuration prepared
- [ ] Database connection pooling optimized
### 2. During Migration
- [ ] Monitor key metrics continuously
- [ ] Validate each phase before proceeding
- [ ] Maintain detailed logs of all actions
- [ ] Keep stakeholders informed of progress
- [ ] Have rollback trigger ready
- [ ] Monitor user experience metrics
- [ ] Watch for performance degradation
- [ ] Validate data consistency
### 3. Post-Migration
- [ ] Continue monitoring for 24-48 hours
- [ ] Validate all business processes
- [ ] Update documentation
- [ ] Conduct post-migration retrospective
- [ ] Archive migration artifacts
- [ ] Update disaster recovery procedures
- [ ] Plan for legacy system decommissioning
### 4. Common Pitfalls to Avoid
- Don't skip testing rollback procedures
- Don't ignore performance impact
- Don't rush through validation phases
- Don't forget to communicate with stakeholders
- Don't assume health checks are sufficient
- Don't neglect data consistency validation
- Don't underestimate time requirements
- Don't overlook dependency impacts
This comprehensive guide provides the foundation for implementing zero-downtime migrations across various system components while maintaining high availability and data integrity.
FILE:scripts/compatibility_checker.py
#!/usr/bin/env python3
"""
Compatibility Checker - Analyze schema and API compatibility between versions
This tool analyzes schema and API changes between versions and identifies backward
compatibility issues including breaking changes, data type mismatches, missing fields,
constraint violations, and generates migration scripts suggestions.
Author: Migration Architect Skill
Version: 1.0.0
License: MIT
"""
import json
import argparse
import sys
import re
import datetime
from typing import Dict, List, Any, Optional, Tuple, Set
from dataclasses import dataclass, asdict
from enum import Enum
class ChangeType(Enum):
"""Types of changes detected"""
BREAKING = "breaking"
POTENTIALLY_BREAKING = "potentially_breaking"
NON_BREAKING = "non_breaking"
ADDITIVE = "additive"
class CompatibilityLevel(Enum):
"""Compatibility assessment levels"""
FULLY_COMPATIBLE = "fully_compatible"
BACKWARD_COMPATIBLE = "backward_compatible"
POTENTIALLY_INCOMPATIBLE = "potentially_incompatible"
BREAKING_CHANGES = "breaking_changes"
@dataclass
class CompatibilityIssue:
"""Individual compatibility issue"""
type: str
severity: str
description: str
field_path: str
old_value: Any
new_value: Any
impact: str
suggested_migration: str
affected_operations: List[str]
@dataclass
class MigrationScript:
"""Migration script suggestion"""
script_type: str # sql, api, config
description: str
script_content: str
rollback_script: str
dependencies: List[str]
validation_query: str
@dataclass
class CompatibilityReport:
"""Complete compatibility analysis report"""
schema_before: str
schema_after: str
analysis_date: str
overall_compatibility: str
breaking_changes_count: int
potentially_breaking_count: int
non_breaking_changes_count: int
additive_changes_count: int
issues: List[CompatibilityIssue]
migration_scripts: List[MigrationScript]
risk_assessment: Dict[str, Any]
recommendations: List[str]
class SchemaCompatibilityChecker:
"""Main schema compatibility checker class"""
def __init__(self):
self.type_compatibility_matrix = self._build_type_compatibility_matrix()
self.constraint_implications = self._build_constraint_implications()
def _build_type_compatibility_matrix(self) -> Dict[str, Dict[str, str]]:
"""Build data type compatibility matrix"""
return {
# SQL data types compatibility
"varchar": {
"text": "compatible",
"char": "potentially_breaking", # length might be different
"nvarchar": "compatible",
"int": "breaking",
"bigint": "breaking",
"decimal": "breaking",
"datetime": "breaking",
"boolean": "breaking"
},
"int": {
"bigint": "compatible",
"smallint": "potentially_breaking", # range reduction
"decimal": "compatible",
"float": "potentially_breaking", # precision loss
"varchar": "breaking",
"boolean": "breaking"
},
"bigint": {
"int": "potentially_breaking", # range reduction
"decimal": "compatible",
"varchar": "breaking",
"boolean": "breaking"
},
"decimal": {
"float": "potentially_breaking", # precision loss
"int": "potentially_breaking", # precision loss
"bigint": "potentially_breaking", # precision loss
"varchar": "breaking",
"boolean": "breaking"
},
"datetime": {
"timestamp": "compatible",
"date": "potentially_breaking", # time component lost
"varchar": "breaking",
"int": "breaking"
},
"boolean": {
"tinyint": "compatible",
"varchar": "breaking",
"int": "breaking"
},
# JSON/API field types
"string": {
"number": "breaking",
"boolean": "breaking",
"array": "breaking",
"object": "breaking",
"null": "potentially_breaking"
},
"number": {
"string": "breaking",
"boolean": "breaking",
"array": "breaking",
"object": "breaking",
"null": "potentially_breaking"
},
"boolean": {
"string": "breaking",
"number": "breaking",
"array": "breaking",
"object": "breaking",
"null": "potentially_breaking"
},
"array": {
"string": "breaking",
"number": "breaking",
"boolean": "breaking",
"object": "breaking",
"null": "potentially_breaking"
},
"object": {
"string": "breaking",
"number": "breaking",
"boolean": "breaking",
"array": "breaking",
"null": "potentially_breaking"
}
}
def _build_constraint_implications(self) -> Dict[str, Dict[str, str]]:
"""Build constraint change implications"""
return {
"required": {
"added": "breaking", # Previously optional field now required
"removed": "non_breaking" # Previously required field now optional
},
"not_null": {
"added": "breaking", # Previously nullable now NOT NULL
"removed": "non_breaking" # Previously NOT NULL now nullable
},
"unique": {
"added": "potentially_breaking", # May fail if duplicates exist
"removed": "non_breaking" # No longer enforcing uniqueness
},
"primary_key": {
"added": "breaking", # Major structural change
"removed": "breaking", # Major structural change
"modified": "breaking" # Primary key change is always breaking
},
"foreign_key": {
"added": "potentially_breaking", # May fail if referential integrity violated
"removed": "potentially_breaking", # May allow orphaned records
"modified": "breaking" # Reference change is breaking
},
"check": {
"added": "potentially_breaking", # May fail if existing data violates check
"removed": "non_breaking", # No longer enforcing check
"modified": "potentially_breaking" # Different validation rules
},
"index": {
"added": "non_breaking", # Performance improvement
"removed": "non_breaking", # Performance impact only
"modified": "non_breaking" # Performance impact only
}
}
def analyze_database_schema(self, before_schema: Dict[str, Any],
after_schema: Dict[str, Any]) -> CompatibilityReport:
"""Analyze database schema compatibility"""
issues = []
migration_scripts = []
before_tables = before_schema.get("tables", {})
after_tables = after_schema.get("tables", {})
# Check for removed tables
for table_name in before_tables:
if table_name not in after_tables:
issues.append(CompatibilityIssue(
type="table_removed",
severity="breaking",
description=f"Table '{table_name}' has been removed",
field_path=f"tables.{table_name}",
old_value=before_tables[table_name],
new_value=None,
impact="All operations on this table will fail",
suggested_migration=f"CREATE VIEW {table_name} AS SELECT * FROM replacement_table;",
affected_operations=["SELECT", "INSERT", "UPDATE", "DELETE"]
))
# Check for added tables
for table_name in after_tables:
if table_name not in before_tables:
migration_scripts.append(MigrationScript(
script_type="sql",
description=f"Create new table {table_name}",
script_content=self._generate_create_table_sql(table_name, after_tables[table_name]),
rollback_script=f"DROP TABLE IF EXISTS {table_name};",
dependencies=[],
validation_query=f"SELECT COUNT(*) FROM information_schema.tables WHERE table_name = '{table_name}';"
))
# Check for modified tables
for table_name in set(before_tables.keys()) & set(after_tables.keys()):
table_issues, table_scripts = self._analyze_table_changes(
table_name, before_tables[table_name], after_tables[table_name]
)
issues.extend(table_issues)
migration_scripts.extend(table_scripts)
return self._build_compatibility_report(
before_schema, after_schema, issues, migration_scripts
)
def analyze_api_schema(self, before_schema: Dict[str, Any],
after_schema: Dict[str, Any]) -> CompatibilityReport:
"""Analyze REST API schema compatibility"""
issues = []
migration_scripts = []
# Analyze endpoints
before_paths = before_schema.get("paths", {})
after_paths = after_schema.get("paths", {})
# Check for removed endpoints
for path in before_paths:
if path not in after_paths:
for method in before_paths[path]:
issues.append(CompatibilityIssue(
type="endpoint_removed",
severity="breaking",
description=f"Endpoint {method.upper()} {path} has been removed",
field_path=f"paths.{path}.{method}",
old_value=before_paths[path][method],
new_value=None,
impact="Client requests to this endpoint will fail with 404",
suggested_migration=f"Implement redirect to replacement endpoint or maintain backward compatibility stub",
affected_operations=[f"{method.upper()} {path}"]
))
# Check for modified endpoints
for path in set(before_paths.keys()) & set(after_paths.keys()):
path_issues, path_scripts = self._analyze_endpoint_changes(
path, before_paths[path], after_paths[path]
)
issues.extend(path_issues)
migration_scripts.extend(path_scripts)
# Analyze data models
before_components = before_schema.get("components", {}).get("schemas", {})
after_components = after_schema.get("components", {}).get("schemas", {})
for model_name in set(before_components.keys()) & set(after_components.keys()):
model_issues, model_scripts = self._analyze_model_changes(
model_name, before_components[model_name], after_components[model_name]
)
issues.extend(model_issues)
migration_scripts.extend(model_scripts)
return self._build_compatibility_report(
before_schema, after_schema, issues, migration_scripts
)
def _analyze_table_changes(self, table_name: str, before_table: Dict[str, Any],
after_table: Dict[str, Any]) -> Tuple[List[CompatibilityIssue], List[MigrationScript]]:
"""Analyze changes to a specific table"""
issues = []
scripts = []
before_columns = before_table.get("columns", {})
after_columns = after_table.get("columns", {})
# Check for removed columns
for col_name in before_columns:
if col_name not in after_columns:
issues.append(CompatibilityIssue(
type="column_removed",
severity="breaking",
description=f"Column '{col_name}' removed from table '{table_name}'",
field_path=f"tables.{table_name}.columns.{col_name}",
old_value=before_columns[col_name],
new_value=None,
impact="SELECT statements including this column will fail",
suggested_migration=f"ALTER TABLE {table_name} ADD COLUMN {col_name}_deprecated AS computed_value;",
affected_operations=["SELECT", "INSERT", "UPDATE"]
))
# Check for added columns
for col_name in after_columns:
if col_name not in before_columns:
col_def = after_columns[col_name]
is_required = col_def.get("nullable", True) == False and col_def.get("default") is None
if is_required:
issues.append(CompatibilityIssue(
type="required_column_added",
severity="breaking",
description=f"Required column '{col_name}' added to table '{table_name}'",
field_path=f"tables.{table_name}.columns.{col_name}",
old_value=None,
new_value=col_def,
impact="INSERT statements without this column will fail",
suggested_migration=f"Add default value or make column nullable initially",
affected_operations=["INSERT"]
))
scripts.append(MigrationScript(
script_type="sql",
description=f"Add column {col_name} to table {table_name}",
script_content=f"ALTER TABLE {table_name} ADD COLUMN {self._generate_column_definition(col_name, col_def)};",
rollback_script=f"ALTER TABLE {table_name} DROP COLUMN {col_name};",
dependencies=[],
validation_query=f"SELECT COUNT(*) FROM information_schema.columns WHERE table_name = '{table_name}' AND column_name = '{col_name}';"
))
# Check for modified columns
for col_name in set(before_columns.keys()) & set(after_columns.keys()):
col_issues, col_scripts = self._analyze_column_changes(
table_name, col_name, before_columns[col_name], after_columns[col_name]
)
issues.extend(col_issues)
scripts.extend(col_scripts)
# Check constraint changes
before_constraints = before_table.get("constraints", {})
after_constraints = after_table.get("constraints", {})
constraint_issues, constraint_scripts = self._analyze_constraint_changes(
table_name, before_constraints, after_constraints
)
issues.extend(constraint_issues)
scripts.extend(constraint_scripts)
return issues, scripts
def _analyze_column_changes(self, table_name: str, col_name: str,
before_col: Dict[str, Any], after_col: Dict[str, Any]) -> Tuple[List[CompatibilityIssue], List[MigrationScript]]:
"""Analyze changes to a specific column"""
issues = []
scripts = []
# Check data type changes
before_type = before_col.get("type", "").lower()
after_type = after_col.get("type", "").lower()
if before_type != after_type:
compatibility = self.type_compatibility_matrix.get(before_type, {}).get(after_type, "breaking")
if compatibility == "breaking":
issues.append(CompatibilityIssue(
type="incompatible_type_change",
severity="breaking",
description=f"Column '{col_name}' type changed from {before_type} to {after_type}",
field_path=f"tables.{table_name}.columns.{col_name}.type",
old_value=before_type,
new_value=after_type,
impact="Data conversion may fail or lose precision",
suggested_migration=f"Add conversion logic and validate data integrity",
affected_operations=["SELECT", "INSERT", "UPDATE", "WHERE clauses"]
))
scripts.append(MigrationScript(
script_type="sql",
description=f"Convert column {col_name} from {before_type} to {after_type}",
script_content=f"ALTER TABLE {table_name} ALTER COLUMN {col_name} TYPE {after_type} USING {col_name}::{after_type};",
rollback_script=f"ALTER TABLE {table_name} ALTER COLUMN {col_name} TYPE {before_type};",
dependencies=[f"backup_{table_name}"],
validation_query=f"SELECT COUNT(*) FROM {table_name} WHERE {col_name} IS NOT NULL;"
))
elif compatibility == "potentially_breaking":
issues.append(CompatibilityIssue(
type="risky_type_change",
severity="potentially_breaking",
description=f"Column '{col_name}' type changed from {before_type} to {after_type} - may lose data",
field_path=f"tables.{table_name}.columns.{col_name}.type",
old_value=before_type,
new_value=after_type,
impact="Potential data loss or precision reduction",
suggested_migration=f"Validate all existing data can be converted safely",
affected_operations=["Data integrity"]
))
# Check nullability changes
before_nullable = before_col.get("nullable", True)
after_nullable = after_col.get("nullable", True)
if before_nullable != after_nullable:
if before_nullable and not after_nullable: # null -> not null
issues.append(CompatibilityIssue(
type="nullability_restriction",
severity="breaking",
description=f"Column '{col_name}' changed from nullable to NOT NULL",
field_path=f"tables.{table_name}.columns.{col_name}.nullable",
old_value=before_nullable,
new_value=after_nullable,
impact="Existing NULL values will cause constraint violations",
suggested_migration=f"Update NULL values to valid defaults before applying NOT NULL constraint",
affected_operations=["INSERT", "UPDATE"]
))
scripts.append(MigrationScript(
script_type="sql",
description=f"Make column {col_name} NOT NULL",
script_content=f"""
-- Update NULL values first
UPDATE {table_name} SET {col_name} = 'DEFAULT_VALUE' WHERE {col_name} IS NULL;
-- Add NOT NULL constraint
ALTER TABLE {table_name} ALTER COLUMN {col_name} SET NOT NULL;
""",
rollback_script=f"ALTER TABLE {table_name} ALTER COLUMN {col_name} DROP NOT NULL;",
dependencies=[],
validation_query=f"SELECT COUNT(*) FROM {table_name} WHERE {col_name} IS NULL;"
))
# Check length/precision changes
before_length = before_col.get("length")
after_length = after_col.get("length")
if before_length and after_length and before_length != after_length:
if after_length < before_length:
issues.append(CompatibilityIssue(
type="length_reduction",
severity="potentially_breaking",
description=f"Column '{col_name}' length reduced from {before_length} to {after_length}",
field_path=f"tables.{table_name}.columns.{col_name}.length",
old_value=before_length,
new_value=after_length,
impact="Data truncation may occur for values exceeding new length",
suggested_migration=f"Validate no existing data exceeds new length limit",
affected_operations=["INSERT", "UPDATE"]
))
return issues, scripts
def _analyze_constraint_changes(self, table_name: str, before_constraints: Dict[str, Any],
after_constraints: Dict[str, Any]) -> Tuple[List[CompatibilityIssue], List[MigrationScript]]:
"""Analyze constraint changes"""
issues = []
scripts = []
for constraint_type in ["primary_key", "foreign_key", "unique", "check"]:
before_constraint = before_constraints.get(constraint_type, [])
after_constraint = after_constraints.get(constraint_type, [])
# Convert to sets for comparison
before_set = set(str(c) for c in before_constraint) if isinstance(before_constraint, list) else {str(before_constraint)} if before_constraint else set()
after_set = set(str(c) for c in after_constraint) if isinstance(after_constraint, list) else {str(after_constraint)} if after_constraint else set()
# Check for removed constraints
for constraint in before_set - after_set:
implication = self.constraint_implications.get(constraint_type, {}).get("removed", "non_breaking")
issues.append(CompatibilityIssue(
type=f"{constraint_type}_removed",
severity=implication,
description=f"{constraint_type.replace('_', ' ').title()} constraint '{constraint}' removed from table '{table_name}'",
field_path=f"tables.{table_name}.constraints.{constraint_type}",
old_value=constraint,
new_value=None,
impact=f"No longer enforcing {constraint_type} constraint",
suggested_migration=f"Consider application-level validation for removed constraint",
affected_operations=["INSERT", "UPDATE", "DELETE"]
))
# Check for added constraints
for constraint in after_set - before_set:
implication = self.constraint_implications.get(constraint_type, {}).get("added", "potentially_breaking")
issues.append(CompatibilityIssue(
type=f"{constraint_type}_added",
severity=implication,
description=f"New {constraint_type.replace('_', ' ')} constraint '{constraint}' added to table '{table_name}'",
field_path=f"tables.{table_name}.constraints.{constraint_type}",
old_value=None,
new_value=constraint,
impact=f"New {constraint_type} constraint may reject existing data",
suggested_migration=f"Validate existing data complies with new constraint",
affected_operations=["INSERT", "UPDATE"]
))
scripts.append(MigrationScript(
script_type="sql",
description=f"Add {constraint_type} constraint to {table_name}",
script_content=f"ALTER TABLE {table_name} ADD CONSTRAINT {constraint_type}_{table_name} {constraint_type.upper()} ({constraint});",
rollback_script=f"ALTER TABLE {table_name} DROP CONSTRAINT {constraint_type}_{table_name};",
dependencies=[],
validation_query=f"SELECT COUNT(*) FROM information_schema.table_constraints WHERE table_name = '{table_name}' AND constraint_type = '{constraint_type.upper()}';"
))
return issues, scripts
def _analyze_endpoint_changes(self, path: str, before_endpoint: Dict[str, Any],
after_endpoint: Dict[str, Any]) -> Tuple[List[CompatibilityIssue], List[MigrationScript]]:
"""Analyze changes to an API endpoint"""
issues = []
scripts = []
for method in set(before_endpoint.keys()) & set(after_endpoint.keys()):
before_method = before_endpoint[method]
after_method = after_endpoint[method]
# Check parameter changes
before_params = before_method.get("parameters", [])
after_params = after_method.get("parameters", [])
before_param_names = {p["name"] for p in before_params}
after_param_names = {p["name"] for p in after_params}
# Check for removed required parameters
for param_name in before_param_names - after_param_names:
param = next(p for p in before_params if p["name"] == param_name)
if param.get("required", False):
issues.append(CompatibilityIssue(
type="required_parameter_removed",
severity="breaking",
description=f"Required parameter '{param_name}' removed from {method.upper()} {path}",
field_path=f"paths.{path}.{method}.parameters",
old_value=param,
new_value=None,
impact="Client requests with this parameter will fail",
suggested_migration="Implement parameter validation with backward compatibility",
affected_operations=[f"{method.upper()} {path}"]
))
# Check for added required parameters
for param_name in after_param_names - before_param_names:
param = next(p for p in after_params if p["name"] == param_name)
if param.get("required", False):
issues.append(CompatibilityIssue(
type="required_parameter_added",
severity="breaking",
description=f"New required parameter '{param_name}' added to {method.upper()} {path}",
field_path=f"paths.{path}.{method}.parameters",
old_value=None,
new_value=param,
impact="Client requests without this parameter will fail",
suggested_migration="Provide default value or make parameter optional initially",
affected_operations=[f"{method.upper()} {path}"]
))
# Check response schema changes
before_responses = before_method.get("responses", {})
after_responses = after_method.get("responses", {})
for status_code in before_responses:
if status_code in after_responses:
before_schema = before_responses[status_code].get("content", {}).get("application/json", {}).get("schema", {})
after_schema = after_responses[status_code].get("content", {}).get("application/json", {}).get("schema", {})
if before_schema != after_schema:
issues.append(CompatibilityIssue(
type="response_schema_changed",
severity="potentially_breaking",
description=f"Response schema changed for {method.upper()} {path} (status {status_code})",
field_path=f"paths.{path}.{method}.responses.{status_code}",
old_value=before_schema,
new_value=after_schema,
impact="Client response parsing may fail",
suggested_migration="Implement versioned API responses",
affected_operations=[f"{method.upper()} {path}"]
))
return issues, scripts
def _analyze_model_changes(self, model_name: str, before_model: Dict[str, Any],
after_model: Dict[str, Any]) -> Tuple[List[CompatibilityIssue], List[MigrationScript]]:
"""Analyze changes to an API data model"""
issues = []
scripts = []
before_props = before_model.get("properties", {})
after_props = after_model.get("properties", {})
before_required = set(before_model.get("required", []))
after_required = set(after_model.get("required", []))
# Check for removed properties
for prop_name in set(before_props.keys()) - set(after_props.keys()):
issues.append(CompatibilityIssue(
type="property_removed",
severity="breaking",
description=f"Property '{prop_name}' removed from model '{model_name}'",
field_path=f"components.schemas.{model_name}.properties.{prop_name}",
old_value=before_props[prop_name],
new_value=None,
impact="Client code expecting this property will fail",
suggested_migration="Use API versioning to maintain backward compatibility",
affected_operations=["Serialization", "Deserialization"]
))
# Check for newly required properties
for prop_name in after_required - before_required:
issues.append(CompatibilityIssue(
type="property_made_required",
severity="breaking",
description=f"Property '{prop_name}' is now required in model '{model_name}'",
field_path=f"components.schemas.{model_name}.required",
old_value=list(before_required),
new_value=list(after_required),
impact="Client requests without this property will fail validation",
suggested_migration="Provide default values or implement gradual rollout",
affected_operations=["Request validation"]
))
# Check for property type changes
for prop_name in set(before_props.keys()) & set(after_props.keys()):
before_type = before_props[prop_name].get("type")
after_type = after_props[prop_name].get("type")
if before_type != after_type:
compatibility = self.type_compatibility_matrix.get(before_type, {}).get(after_type, "breaking")
issues.append(CompatibilityIssue(
type="property_type_changed",
severity=compatibility,
description=f"Property '{prop_name}' type changed from {before_type} to {after_type} in model '{model_name}'",
field_path=f"components.schemas.{model_name}.properties.{prop_name}.type",
old_value=before_type,
new_value=after_type,
impact="Client serialization/deserialization may fail",
suggested_migration="Implement type coercion or API versioning",
affected_operations=["Serialization", "Deserialization"]
))
return issues, scripts
def _build_compatibility_report(self, before_schema: Dict[str, Any], after_schema: Dict[str, Any],
issues: List[CompatibilityIssue], migration_scripts: List[MigrationScript]) -> CompatibilityReport:
"""Build the final compatibility report"""
# Count issues by severity
breaking_count = sum(1 for issue in issues if issue.severity == "breaking")
potentially_breaking_count = sum(1 for issue in issues if issue.severity == "potentially_breaking")
non_breaking_count = sum(1 for issue in issues if issue.severity == "non_breaking")
additive_count = sum(1 for issue in issues if issue.type == "additive")
# Determine overall compatibility
if breaking_count > 0:
overall_compatibility = "breaking_changes"
elif potentially_breaking_count > 0:
overall_compatibility = "potentially_incompatible"
elif non_breaking_count > 0:
overall_compatibility = "backward_compatible"
else:
overall_compatibility = "fully_compatible"
# Generate risk assessment
risk_assessment = {
"overall_risk": "high" if breaking_count > 0 else "medium" if potentially_breaking_count > 0 else "low",
"deployment_risk": "requires_coordinated_deployment" if breaking_count > 0 else "safe_independent_deployment",
"rollback_complexity": "high" if breaking_count > 3 else "medium" if breaking_count > 0 else "low",
"testing_requirements": ["integration_testing", "regression_testing"] +
(["data_migration_testing"] if any(s.script_type == "sql" for s in migration_scripts) else [])
}
# Generate recommendations
recommendations = []
if breaking_count > 0:
recommendations.append("Implement API versioning to maintain backward compatibility")
recommendations.append("Plan for coordinated deployment with all clients")
recommendations.append("Implement comprehensive rollback procedures")
if potentially_breaking_count > 0:
recommendations.append("Conduct thorough testing with realistic data volumes")
recommendations.append("Implement monitoring for migration success metrics")
if migration_scripts:
recommendations.append("Test all migration scripts in staging environment")
recommendations.append("Implement migration progress monitoring")
recommendations.append("Create detailed communication plan for stakeholders")
recommendations.append("Implement feature flags for gradual rollout")
return CompatibilityReport(
schema_before=json.dumps(before_schema, indent=2)[:500] + "..." if len(json.dumps(before_schema)) > 500 else json.dumps(before_schema, indent=2),
schema_after=json.dumps(after_schema, indent=2)[:500] + "..." if len(json.dumps(after_schema)) > 500 else json.dumps(after_schema, indent=2),
analysis_date=datetime.datetime.now().isoformat(),
overall_compatibility=overall_compatibility,
breaking_changes_count=breaking_count,
potentially_breaking_count=potentially_breaking_count,
non_breaking_changes_count=non_breaking_count,
additive_changes_count=additive_count,
issues=issues,
migration_scripts=migration_scripts,
risk_assessment=risk_assessment,
recommendations=recommendations
)
def _generate_create_table_sql(self, table_name: str, table_def: Dict[str, Any]) -> str:
"""Generate CREATE TABLE SQL statement"""
columns = []
for col_name, col_def in table_def.get("columns", {}).items():
columns.append(self._generate_column_definition(col_name, col_def))
return f"CREATE TABLE {table_name} (\n " + ",\n ".join(columns) + "\n);"
def _generate_column_definition(self, col_name: str, col_def: Dict[str, Any]) -> str:
"""Generate column definition for SQL"""
col_type = col_def.get("type", "VARCHAR(255)")
nullable = "" if col_def.get("nullable", True) else " NOT NULL"
default = f" DEFAULT {col_def.get('default')}" if col_def.get("default") is not None else ""
return f"{col_name} {col_type}{nullable}{default}"
def generate_human_readable_report(self, report: CompatibilityReport) -> str:
"""Generate human-readable compatibility report"""
output = []
output.append("=" * 80)
output.append("COMPATIBILITY ANALYSIS REPORT")
output.append("=" * 80)
output.append(f"Analysis Date: {report.analysis_date}")
output.append(f"Overall Compatibility: {report.overall_compatibility.upper()}")
output.append("")
# Summary
output.append("SUMMARY")
output.append("-" * 40)
output.append(f"Breaking Changes: {report.breaking_changes_count}")
output.append(f"Potentially Breaking: {report.potentially_breaking_count}")
output.append(f"Non-Breaking Changes: {report.non_breaking_changes_count}")
output.append(f"Additive Changes: {report.additive_changes_count}")
output.append(f"Total Issues Found: {len(report.issues)}")
output.append("")
# Risk Assessment
output.append("RISK ASSESSMENT")
output.append("-" * 40)
for key, value in report.risk_assessment.items():
output.append(f"{key.replace('_', ' ').title()}: {value}")
output.append("")
# Issues by Severity
issues_by_severity = {}
for issue in report.issues:
if issue.severity not in issues_by_severity:
issues_by_severity[issue.severity] = []
issues_by_severity[issue.severity].append(issue)
for severity in ["breaking", "potentially_breaking", "non_breaking"]:
if severity in issues_by_severity:
output.append(f"{severity.upper().replace('_', ' ')} ISSUES")
output.append("-" * 40)
for issue in issues_by_severity[severity]:
output.append(f"• {issue.description}")
output.append(f" Field: {issue.field_path}")
output.append(f" Impact: {issue.impact}")
output.append(f" Migration: {issue.suggested_migration}")
if issue.affected_operations:
output.append(f" Affected Operations: {', '.join(issue.affected_operations)}")
output.append("")
# Migration Scripts
if report.migration_scripts:
output.append("SUGGESTED MIGRATION SCRIPTS")
output.append("-" * 40)
for i, script in enumerate(report.migration_scripts, 1):
output.append(f"{i}. {script.description}")
output.append(f" Type: {script.script_type}")
output.append(" Script:")
for line in script.script_content.split('\n'):
output.append(f" {line}")
output.append("")
# Recommendations
output.append("RECOMMENDATIONS")
output.append("-" * 40)
for i, rec in enumerate(report.recommendations, 1):
output.append(f"{i}. {rec}")
output.append("")
return "\n".join(output)
def main():
"""Main function with command line interface"""
parser = argparse.ArgumentParser(description="Analyze schema and API compatibility between versions")
parser.add_argument("--before", required=True, help="Before schema file (JSON)")
parser.add_argument("--after", required=True, help="After schema file (JSON)")
parser.add_argument("--type", choices=["database", "api"], default="database", help="Schema type to analyze")
parser.add_argument("--output", "-o", help="Output file for compatibility report (JSON)")
parser.add_argument("--format", "-f", choices=["json", "text", "both"], default="both", help="Output format")
args = parser.parse_args()
try:
# Load schemas
with open(args.before, 'r') as f:
before_schema = json.load(f)
with open(args.after, 'r') as f:
after_schema = json.load(f)
# Analyze compatibility
checker = SchemaCompatibilityChecker()
if args.type == "database":
report = checker.analyze_database_schema(before_schema, after_schema)
else: # api
report = checker.analyze_api_schema(before_schema, after_schema)
# Output results
if args.format in ["json", "both"]:
report_dict = asdict(report)
if args.output:
with open(args.output, 'w') as f:
json.dump(report_dict, f, indent=2)
print(f"Compatibility report saved to {args.output}")
else:
print(json.dumps(report_dict, indent=2))
if args.format in ["text", "both"]:
human_report = checker.generate_human_readable_report(report)
text_output = args.output.replace('.json', '.txt') if args.output else None
if text_output:
with open(text_output, 'w') as f:
f.write(human_report)
print(f"Human-readable report saved to {text_output}")
else:
print("\n" + "="*80)
print("HUMAN-READABLE COMPATIBILITY REPORT")
print("="*80)
print(human_report)
# Return exit code based on compatibility
if report.breaking_changes_count > 0:
return 2 # Breaking changes found
elif report.potentially_breaking_count > 0:
return 1 # Potentially breaking changes found
else:
return 0 # No compatibility issues
except FileNotFoundError as e:
print(f"Error: File not found: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON: {e}", file=sys.stderr)
return 1
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
return 1
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/migration_planner.py
#!/usr/bin/env python3
"""
Migration Planner - Generate comprehensive migration plans with risk assessment
This tool analyzes migration specifications and generates detailed, phased migration plans
including pre-migration checklists, validation gates, rollback triggers, timeline estimates,
and risk matrices.
Author: Migration Architect Skill
Version: 1.0.0
License: MIT
"""
import json
import argparse
import sys
import datetime
import hashlib
import math
from typing import Dict, List, Any, Optional, Tuple
from dataclasses import dataclass, asdict
from enum import Enum
class MigrationType(Enum):
"""Migration type enumeration"""
DATABASE = "database"
SERVICE = "service"
INFRASTRUCTURE = "infrastructure"
DATA = "data"
API = "api"
class MigrationComplexity(Enum):
"""Migration complexity levels"""
LOW = "low"
MEDIUM = "medium"
HIGH = "high"
CRITICAL = "critical"
class RiskLevel(Enum):
"""Risk assessment levels"""
LOW = "low"
MEDIUM = "medium"
HIGH = "high"
CRITICAL = "critical"
@dataclass
class MigrationConstraint:
"""Migration constraint definition"""
type: str
description: str
impact: str
mitigation: str
@dataclass
class MigrationPhase:
"""Individual migration phase"""
name: str
description: str
duration_hours: int
dependencies: List[str]
validation_criteria: List[str]
rollback_triggers: List[str]
tasks: List[str]
risk_level: str
resources_required: List[str]
@dataclass
class RiskItem:
"""Individual risk assessment item"""
category: str
description: str
probability: str # low, medium, high
impact: str # low, medium, high
severity: str # low, medium, high, critical
mitigation: str
owner: str
@dataclass
class MigrationPlan:
"""Complete migration plan structure"""
migration_id: str
source_system: str
target_system: str
migration_type: str
complexity: str
estimated_duration_hours: int
phases: List[MigrationPhase]
risks: List[RiskItem]
success_criteria: List[str]
rollback_plan: Dict[str, Any]
stakeholders: List[str]
created_at: str
class MigrationPlanner:
"""Main migration planner class"""
def __init__(self):
self.migration_patterns = self._load_migration_patterns()
self.risk_templates = self._load_risk_templates()
def _load_migration_patterns(self) -> Dict[str, Any]:
"""Load predefined migration patterns"""
return {
"database": {
"schema_change": {
"phases": ["preparation", "expand", "migrate", "contract", "cleanup"],
"base_duration": 24,
"complexity_multiplier": {"low": 1.0, "medium": 1.5, "high": 2.5, "critical": 4.0}
},
"data_migration": {
"phases": ["assessment", "setup", "bulk_copy", "delta_sync", "validation", "cutover"],
"base_duration": 48,
"complexity_multiplier": {"low": 1.2, "medium": 2.0, "high": 3.0, "critical": 5.0}
}
},
"service": {
"strangler_fig": {
"phases": ["intercept", "implement", "redirect", "validate", "retire"],
"base_duration": 168, # 1 week
"complexity_multiplier": {"low": 0.8, "medium": 1.0, "high": 1.8, "critical": 3.0}
},
"parallel_run": {
"phases": ["setup", "deploy", "shadow", "compare", "cutover", "cleanup"],
"base_duration": 72,
"complexity_multiplier": {"low": 1.0, "medium": 1.3, "high": 2.0, "critical": 3.5}
}
},
"infrastructure": {
"cloud_migration": {
"phases": ["assessment", "design", "pilot", "migration", "optimization", "decommission"],
"base_duration": 720, # 30 days
"complexity_multiplier": {"low": 0.6, "medium": 1.0, "high": 1.5, "critical": 2.5}
},
"on_prem_to_cloud": {
"phases": ["discovery", "planning", "pilot", "migration", "validation", "cutover"],
"base_duration": 480, # 20 days
"complexity_multiplier": {"low": 0.8, "medium": 1.2, "high": 2.0, "critical": 3.0}
}
}
}
def _load_risk_templates(self) -> Dict[str, List[RiskItem]]:
"""Load risk templates for different migration types"""
return {
"database": [
RiskItem("technical", "Data corruption during migration", "low", "critical", "high",
"Implement comprehensive backup and validation procedures", "DBA Team"),
RiskItem("technical", "Extended downtime due to migration complexity", "medium", "high", "high",
"Use blue-green deployment and phased migration approach", "DevOps Team"),
RiskItem("business", "Business process disruption", "medium", "high", "high",
"Communicate timeline and provide alternate workflows", "Business Owner"),
RiskItem("operational", "Insufficient rollback testing", "high", "critical", "critical",
"Execute full rollback procedures in staging environment", "QA Team")
],
"service": [
RiskItem("technical", "Service compatibility issues", "medium", "high", "high",
"Implement comprehensive integration testing", "Development Team"),
RiskItem("technical", "Performance degradation", "medium", "medium", "medium",
"Conduct load testing and performance benchmarking", "DevOps Team"),
RiskItem("business", "Feature parity gaps", "high", "high", "high",
"Document feature mapping and acceptance criteria", "Product Owner"),
RiskItem("operational", "Monitoring gap during transition", "medium", "medium", "medium",
"Set up dual monitoring and alerting systems", "SRE Team")
],
"infrastructure": [
RiskItem("technical", "Network connectivity issues", "medium", "critical", "high",
"Implement redundant network paths and monitoring", "Network Team"),
RiskItem("technical", "Security configuration drift", "high", "high", "high",
"Automated security scanning and compliance checks", "Security Team"),
RiskItem("business", "Cost overrun during transition", "high", "medium", "medium",
"Implement cost monitoring and budget alerts", "Finance Team"),
RiskItem("operational", "Team knowledge gaps", "high", "medium", "medium",
"Provide training and create detailed documentation", "Platform Team")
]
}
def _calculate_complexity(self, spec: Dict[str, Any]) -> str:
"""Calculate migration complexity based on specification"""
complexity_score = 0
# Data volume complexity
data_volume = spec.get("constraints", {}).get("data_volume_gb", 0)
if data_volume > 10000:
complexity_score += 3
elif data_volume > 1000:
complexity_score += 2
elif data_volume > 100:
complexity_score += 1
# System dependencies
dependencies = len(spec.get("constraints", {}).get("dependencies", []))
if dependencies > 10:
complexity_score += 3
elif dependencies > 5:
complexity_score += 2
elif dependencies > 2:
complexity_score += 1
# Downtime constraints
max_downtime = spec.get("constraints", {}).get("max_downtime_minutes", 480)
if max_downtime < 60:
complexity_score += 3
elif max_downtime < 240:
complexity_score += 2
elif max_downtime < 480:
complexity_score += 1
# Special requirements
special_reqs = spec.get("constraints", {}).get("special_requirements", [])
complexity_score += len(special_reqs)
if complexity_score >= 8:
return "critical"
elif complexity_score >= 5:
return "high"
elif complexity_score >= 3:
return "medium"
else:
return "low"
def _estimate_duration(self, migration_type: str, migration_pattern: str, complexity: str) -> int:
"""Estimate migration duration based on type, pattern, and complexity"""
pattern_info = self.migration_patterns.get(migration_type, {}).get(migration_pattern, {})
base_duration = pattern_info.get("base_duration", 48)
multiplier = pattern_info.get("complexity_multiplier", {}).get(complexity, 1.5)
return int(base_duration * multiplier)
def _generate_phases(self, spec: Dict[str, Any]) -> List[MigrationPhase]:
"""Generate migration phases based on specification"""
migration_type = spec.get("type")
migration_pattern = spec.get("pattern", "")
complexity = self._calculate_complexity(spec)
pattern_info = self.migration_patterns.get(migration_type, {})
if migration_pattern in pattern_info:
phase_names = pattern_info[migration_pattern]["phases"]
else:
# Default phases based on migration type
phase_names = {
"database": ["preparation", "migration", "validation", "cutover"],
"service": ["preparation", "deployment", "testing", "cutover"],
"infrastructure": ["assessment", "preparation", "migration", "validation"]
}.get(migration_type, ["preparation", "execution", "validation", "cleanup"])
phases = []
total_duration = self._estimate_duration(migration_type, migration_pattern, complexity)
phase_duration = total_duration // len(phase_names)
for i, phase_name in enumerate(phase_names):
phase = self._create_phase(phase_name, phase_duration, complexity, i, phase_names)
phases.append(phase)
return phases
def _create_phase(self, phase_name: str, duration: int, complexity: str,
phase_index: int, all_phases: List[str]) -> MigrationPhase:
"""Create a detailed migration phase"""
phase_templates = {
"preparation": {
"description": "Prepare systems and teams for migration",
"tasks": [
"Backup source system",
"Set up monitoring and alerting",
"Prepare rollback procedures",
"Communicate migration timeline",
"Validate prerequisites"
],
"validation_criteria": [
"All backups completed successfully",
"Monitoring systems operational",
"Team members briefed and ready",
"Rollback procedures tested"
],
"risk_level": "medium"
},
"assessment": {
"description": "Assess current state and migration requirements",
"tasks": [
"Inventory existing systems and dependencies",
"Analyze data volumes and complexity",
"Identify integration points",
"Document current architecture",
"Create migration mapping"
],
"validation_criteria": [
"Complete system inventory documented",
"Dependencies mapped and validated",
"Migration scope clearly defined",
"Resource requirements identified"
],
"risk_level": "low"
},
"migration": {
"description": "Execute core migration processes",
"tasks": [
"Begin data/service migration",
"Monitor migration progress",
"Validate data consistency",
"Handle migration errors",
"Update configuration"
],
"validation_criteria": [
"Migration progress within expected parameters",
"Data consistency checks passing",
"Error rates within acceptable limits",
"Performance metrics stable"
],
"risk_level": "high"
},
"validation": {
"description": "Validate migration success and system health",
"tasks": [
"Execute comprehensive testing",
"Validate business processes",
"Check system performance",
"Verify data integrity",
"Confirm security controls"
],
"validation_criteria": [
"All critical tests passing",
"Performance within acceptable range",
"Security controls functioning",
"Business processes operational"
],
"risk_level": "medium"
},
"cutover": {
"description": "Switch production traffic to new system",
"tasks": [
"Update DNS/load balancer configuration",
"Redirect production traffic",
"Monitor system performance",
"Validate end-user experience",
"Confirm business operations"
],
"validation_criteria": [
"Traffic successfully redirected",
"System performance stable",
"User experience satisfactory",
"Business operations normal"
],
"risk_level": "critical"
}
}
template = phase_templates.get(phase_name, {
"description": f"Execute {phase_name} phase",
"tasks": [f"Complete {phase_name} activities"],
"validation_criteria": [f"{phase_name.title()} phase completed successfully"],
"risk_level": "medium"
})
dependencies = []
if phase_index > 0:
dependencies.append(all_phases[phase_index - 1])
rollback_triggers = [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
]
resources_required = [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
return MigrationPhase(
name=phase_name,
description=template["description"],
duration_hours=duration,
dependencies=dependencies,
validation_criteria=template["validation_criteria"],
rollback_triggers=rollback_triggers,
tasks=template["tasks"],
risk_level=template["risk_level"],
resources_required=resources_required
)
def _assess_risks(self, spec: Dict[str, Any]) -> List[RiskItem]:
"""Generate risk assessment for migration"""
migration_type = spec.get("type")
base_risks = self.risk_templates.get(migration_type, [])
# Add specification-specific risks
additional_risks = []
constraints = spec.get("constraints", {})
if constraints.get("max_downtime_minutes", 480) < 60:
additional_risks.append(
RiskItem("business", "Zero-downtime requirement increases complexity", "high", "medium", "high",
"Implement blue-green deployment or rolling update strategy", "DevOps Team")
)
if constraints.get("data_volume_gb", 0) > 5000:
additional_risks.append(
RiskItem("technical", "Large data volumes may cause extended migration time", "high", "medium", "medium",
"Implement parallel processing and progress monitoring", "Data Team")
)
compliance_reqs = constraints.get("compliance_requirements", [])
if compliance_reqs:
additional_risks.append(
RiskItem("compliance", "Regulatory compliance requirements", "medium", "high", "high",
"Ensure all compliance checks are integrated into migration process", "Compliance Team")
)
return base_risks + additional_risks
def _generate_rollback_plan(self, phases: List[MigrationPhase]) -> Dict[str, Any]:
"""Generate comprehensive rollback plan"""
rollback_phases = []
for phase in reversed(phases):
rollback_phase = {
"phase": phase.name,
"rollback_actions": [
f"Revert {phase.name} changes",
f"Restore pre-{phase.name} state",
f"Validate {phase.name} rollback success"
],
"validation_criteria": [
f"System restored to pre-{phase.name} state",
f"All {phase.name} changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": phase.duration_hours * 15 # 25% of original phase time
}
rollback_phases.append(rollback_phase)
return {
"rollback_phases": rollback_phases,
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Migration timeline exceeded by > 50%",
"Business-critical functionality unavailable",
"Security breach detected",
"Stakeholder decision to abort"
],
"rollback_decision_matrix": {
"low_severity": "Continue with monitoring",
"medium_severity": "Assess and decide within 15 minutes",
"high_severity": "Immediate rollback initiation",
"critical_severity": "Emergency rollback - all hands"
},
"rollback_contacts": [
"Migration Lead",
"Technical Lead",
"Business Owner",
"On-call Engineer"
]
}
def generate_plan(self, spec: Dict[str, Any]) -> MigrationPlan:
"""Generate complete migration plan from specification"""
migration_id = hashlib.md5(json.dumps(spec, sort_keys=True).encode()).hexdigest()[:12]
complexity = self._calculate_complexity(spec)
phases = self._generate_phases(spec)
risks = self._assess_risks(spec)
total_duration = sum(phase.duration_hours for phase in phases)
rollback_plan = self._generate_rollback_plan(phases)
success_criteria = [
"All data successfully migrated with 100% integrity",
"System performance meets or exceeds baseline",
"All business processes functioning normally",
"No critical security vulnerabilities introduced",
"Stakeholder acceptance criteria met",
"Documentation and runbooks updated"
]
stakeholders = [
"Business Owner",
"Technical Lead",
"DevOps Team",
"QA Team",
"Security Team",
"End Users"
]
return MigrationPlan(
migration_id=migration_id,
source_system=spec.get("source", "Unknown"),
target_system=spec.get("target", "Unknown"),
migration_type=spec.get("type", "Unknown"),
complexity=complexity,
estimated_duration_hours=total_duration,
phases=phases,
risks=risks,
success_criteria=success_criteria,
rollback_plan=rollback_plan,
stakeholders=stakeholders,
created_at=datetime.datetime.now().isoformat()
)
def generate_human_readable_plan(self, plan: MigrationPlan) -> str:
"""Generate human-readable migration plan"""
output = []
output.append("=" * 80)
output.append(f"MIGRATION PLAN: {plan.migration_id}")
output.append("=" * 80)
output.append(f"Source System: {plan.source_system}")
output.append(f"Target System: {plan.target_system}")
output.append(f"Migration Type: {plan.migration_type.upper()}")
output.append(f"Complexity Level: {plan.complexity.upper()}")
output.append(f"Estimated Duration: {plan.estimated_duration_hours} hours ({plan.estimated_duration_hours/24:.1f} days)")
output.append(f"Created: {plan.created_at}")
output.append("")
# Phases
output.append("MIGRATION PHASES")
output.append("-" * 40)
for i, phase in enumerate(plan.phases, 1):
output.append(f"{i}. {phase.name.upper()} ({phase.duration_hours}h)")
output.append(f" Description: {phase.description}")
output.append(f" Risk Level: {phase.risk_level.upper()}")
if phase.dependencies:
output.append(f" Dependencies: {', '.join(phase.dependencies)}")
output.append(" Tasks:")
for task in phase.tasks:
output.append(f" • {task}")
output.append(" Success Criteria:")
for criteria in phase.validation_criteria:
output.append(f" ✓ {criteria}")
output.append("")
# Risk Assessment
output.append("RISK ASSESSMENT")
output.append("-" * 40)
risk_by_severity = {}
for risk in plan.risks:
if risk.severity not in risk_by_severity:
risk_by_severity[risk.severity] = []
risk_by_severity[risk.severity].append(risk)
for severity in ["critical", "high", "medium", "low"]:
if severity in risk_by_severity:
output.append(f"{severity.upper()} SEVERITY RISKS:")
for risk in risk_by_severity[severity]:
output.append(f" • {risk.description}")
output.append(f" Category: {risk.category}")
output.append(f" Probability: {risk.probability} | Impact: {risk.impact}")
output.append(f" Mitigation: {risk.mitigation}")
output.append(f" Owner: {risk.owner}")
output.append("")
# Rollback Plan
output.append("ROLLBACK STRATEGY")
output.append("-" * 40)
output.append("Rollback Triggers:")
for trigger in plan.rollback_plan["rollback_triggers"]:
output.append(f" • {trigger}")
output.append("")
output.append("Rollback Phases:")
for rb_phase in plan.rollback_plan["rollback_phases"]:
output.append(f" {rb_phase['phase'].upper()}:")
for action in rb_phase["rollback_actions"]:
output.append(f" - {action}")
output.append(f" Estimated Time: {rb_phase['estimated_time_minutes']} minutes")
output.append("")
# Success Criteria
output.append("SUCCESS CRITERIA")
output.append("-" * 40)
for criteria in plan.success_criteria:
output.append(f"✓ {criteria}")
output.append("")
# Stakeholders
output.append("STAKEHOLDERS")
output.append("-" * 40)
for stakeholder in plan.stakeholders:
output.append(f"• {stakeholder}")
output.append("")
return "\n".join(output)
def main():
"""Main function with command line interface"""
parser = argparse.ArgumentParser(description="Generate comprehensive migration plans")
parser.add_argument("--input", "-i", required=True, help="Input migration specification file (JSON)")
parser.add_argument("--output", "-o", help="Output file for migration plan (JSON)")
parser.add_argument("--format", "-f", choices=["json", "text", "both"], default="both",
help="Output format")
parser.add_argument("--validate", action="store_true", help="Validate migration specification only")
args = parser.parse_args()
try:
# Load migration specification
with open(args.input, 'r') as f:
spec = json.load(f)
# Validate required fields
required_fields = ["type", "source", "target"]
for field in required_fields:
if field not in spec:
print(f"Error: Missing required field '{field}' in specification", file=sys.stderr)
return 1
if args.validate:
print("Migration specification is valid")
return 0
# Generate migration plan
planner = MigrationPlanner()
plan = planner.generate_plan(spec)
# Output results
if args.format in ["json", "both"]:
plan_dict = asdict(plan)
if args.output:
with open(args.output, 'w') as f:
json.dump(plan_dict, f, indent=2)
print(f"Migration plan saved to {args.output}")
else:
print(json.dumps(plan_dict, indent=2))
if args.format in ["text", "both"]:
human_plan = planner.generate_human_readable_plan(plan)
text_output = args.output.replace('.json', '.txt') if args.output else None
if text_output:
with open(text_output, 'w') as f:
f.write(human_plan)
print(f"Human-readable plan saved to {text_output}")
else:
print("\n" + "="*80)
print("HUMAN-READABLE MIGRATION PLAN")
print("="*80)
print(human_plan)
except FileNotFoundError:
print(f"Error: Input file '{args.input}' not found", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON in input file: {e}", file=sys.stderr)
return 1
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
return 1
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/rollback_generator.py
#!/usr/bin/env python3
"""
Rollback Generator - Generate comprehensive rollback procedures for migrations
This tool takes a migration plan and generates detailed rollback procedures for each phase,
including data rollback scripts, service rollback steps, validation checks, and communication
templates to ensure safe and reliable migration reversals.
Author: Migration Architect Skill
Version: 1.0.0
License: MIT
"""
import json
import argparse
import sys
import datetime
import hashlib
from typing import Dict, List, Any, Optional, Tuple
from dataclasses import dataclass, asdict
from enum import Enum
class RollbackTrigger(Enum):
"""Types of rollback triggers"""
MANUAL = "manual"
AUTOMATED = "automated"
THRESHOLD_BASED = "threshold_based"
TIME_BASED = "time_based"
class RollbackUrgency(Enum):
"""Rollback urgency levels"""
LOW = "low"
MEDIUM = "medium"
HIGH = "high"
EMERGENCY = "emergency"
@dataclass
class RollbackStep:
"""Individual rollback step"""
step_id: str
name: str
description: str
script_type: str # sql, bash, api, manual
script_content: str
estimated_duration_minutes: int
dependencies: List[str]
validation_commands: List[str]
success_criteria: List[str]
failure_escalation: str
rollback_order: int
@dataclass
class RollbackPhase:
"""Rollback phase containing multiple steps"""
phase_name: str
description: str
urgency_level: str
estimated_duration_minutes: int
prerequisites: List[str]
steps: List[RollbackStep]
validation_checkpoints: List[str]
communication_requirements: List[str]
risk_level: str
@dataclass
class RollbackTriggerCondition:
"""Conditions that trigger automatic rollback"""
trigger_id: str
name: str
condition: str
metric_threshold: Optional[Dict[str, Any]]
evaluation_window_minutes: int
auto_execute: bool
escalation_contacts: List[str]
@dataclass
class DataRecoveryPlan:
"""Data recovery and restoration plan"""
recovery_method: str # backup_restore, point_in_time, event_replay
backup_location: str
recovery_scripts: List[str]
data_validation_queries: List[str]
estimated_recovery_time_minutes: int
recovery_dependencies: List[str]
@dataclass
class CommunicationTemplate:
"""Communication template for rollback scenarios"""
template_type: str # start, progress, completion, escalation
audience: str # technical, business, executive, customers
subject: str
body: str
urgency: str
delivery_methods: List[str]
@dataclass
class RollbackRunbook:
"""Complete rollback runbook"""
runbook_id: str
migration_id: str
created_at: str
rollback_phases: List[RollbackPhase]
trigger_conditions: List[RollbackTriggerCondition]
data_recovery_plan: DataRecoveryPlan
communication_templates: List[CommunicationTemplate]
escalation_matrix: Dict[str, Any]
validation_checklist: List[str]
post_rollback_procedures: List[str]
emergency_contacts: List[Dict[str, str]]
class RollbackGenerator:
"""Main rollback generator class"""
def __init__(self):
self.rollback_templates = self._load_rollback_templates()
self.validation_templates = self._load_validation_templates()
self.communication_templates = self._load_communication_templates()
def _load_rollback_templates(self) -> Dict[str, Any]:
"""Load rollback script templates for different migration types"""
return {
"database": {
"schema_rollback": {
"drop_table": "DROP TABLE IF EXISTS {table_name};",
"drop_column": "ALTER TABLE {table_name} DROP COLUMN IF EXISTS {column_name};",
"restore_column": "ALTER TABLE {table_name} ADD COLUMN {column_definition};",
"revert_type": "ALTER TABLE {table_name} ALTER COLUMN {column_name} TYPE {original_type};",
"drop_constraint": "ALTER TABLE {table_name} DROP CONSTRAINT {constraint_name};",
"add_constraint": "ALTER TABLE {table_name} ADD CONSTRAINT {constraint_name} {constraint_definition};"
},
"data_rollback": {
"restore_backup": "pg_restore -d {database_name} -c {backup_file}",
"point_in_time_recovery": "SELECT pg_create_restore_point('pre_migration_{timestamp}');",
"delete_migrated_data": "DELETE FROM {table_name} WHERE migration_batch_id = '{batch_id}';",
"restore_original_values": "UPDATE {table_name} SET {column_name} = backup_{column_name} WHERE migration_flag = true;"
}
},
"service": {
"deployment_rollback": {
"rollback_blue_green": "kubectl patch service {service_name} -p '{\"spec\":{\"selector\":{\"version\":\"blue\"}}}'",
"rollback_canary": "kubectl scale deployment {service_name}-canary --replicas=0",
"restore_previous_version": "kubectl rollout undo deployment/{service_name} --to-revision={revision_number}",
"update_load_balancer": "aws elbv2 modify-rule --rule-arn {rule_arn} --actions Type=forward,TargetGroupArn={original_target_group}"
},
"configuration_rollback": {
"restore_config_map": "kubectl apply -f {original_config_file}",
"revert_feature_flags": "curl -X PUT {feature_flag_api}/flags/{flag_name} -d '{\"enabled\": false}'",
"restore_environment_vars": "kubectl set env deployment/{deployment_name} {env_var_name}={original_value}"
}
},
"infrastructure": {
"cloud_rollback": {
"revert_terraform": "terraform apply -target={resource_name} {rollback_plan_file}",
"restore_dns": "aws route53 change-resource-record-sets --hosted-zone-id {zone_id} --change-batch file://{rollback_dns_changes}",
"rollback_security_groups": "aws ec2 authorize-security-group-ingress --group-id {group_id} --protocol {protocol} --port {port} --cidr {cidr}",
"restore_iam_policies": "aws iam put-role-policy --role-name {role_name} --policy-name {policy_name} --policy-document file://{original_policy}"
},
"network_rollback": {
"restore_routing": "aws ec2 replace-route --route-table-id {route_table_id} --destination-cidr-block {cidr} --gateway-id {original_gateway}",
"revert_load_balancer": "aws elbv2 modify-load-balancer --load-balancer-arn {lb_arn} --scheme {original_scheme}",
"restore_firewall_rules": "aws ec2 revoke-security-group-ingress --group-id {group_id} --protocol {protocol} --port {port} --source-group {source_group}"
}
}
}
def _load_validation_templates(self) -> Dict[str, List[str]]:
"""Load validation command templates"""
return {
"database": [
"SELECT COUNT(*) FROM {table_name};",
"SELECT COUNT(*) FROM information_schema.tables WHERE table_name = '{table_name}';",
"SELECT COUNT(*) FROM information_schema.columns WHERE table_name = '{table_name}' AND column_name = '{column_name}';",
"SELECT COUNT(DISTINCT {primary_key}) FROM {table_name};",
"SELECT MAX({timestamp_column}) FROM {table_name};"
],
"service": [
"curl -f {health_check_url}",
"kubectl get pods -l app={service_name} --field-selector=status.phase=Running",
"kubectl logs deployment/{service_name} --tail=100 | grep -i error",
"curl -f {service_endpoint}/api/v1/status"
],
"infrastructure": [
"aws ec2 describe-instances --instance-ids {instance_id} --query 'Reservations[*].Instances[*].State.Name'",
"nslookup {domain_name}",
"curl -I {load_balancer_url}",
"aws elbv2 describe-target-health --target-group-arn {target_group_arn}"
]
}
def _load_communication_templates(self) -> Dict[str, Dict[str, str]]:
"""Load communication templates"""
return {
"rollback_start": {
"technical": {
"subject": "ROLLBACK INITIATED: {migration_name}",
"body": """Team,
We have initiated rollback for migration: {migration_name}
Rollback ID: {rollback_id}
Start Time: {start_time}
Estimated Duration: {estimated_duration}
Reason: {rollback_reason}
Current Status: Rolling back phase {current_phase}
Next Updates: Every 15 minutes or upon phase completion
Actions Required:
- Monitor system health dashboards
- Stand by for escalation if needed
- Do not make manual changes during rollback
Incident Commander: {incident_commander}
"""
},
"business": {
"subject": "System Rollback In Progress - {system_name}",
"body": """Business Stakeholders,
We are currently performing a planned rollback of the {system_name} migration due to {rollback_reason}.
Impact: {business_impact}
Expected Resolution: {estimated_completion_time}
Affected Services: {affected_services}
We will provide updates every 30 minutes.
Contact: {business_contact}
"""
},
"executive": {
"subject": "EXEC ALERT: Critical System Rollback - {system_name}",
"body": """Executive Team,
A critical rollback is in progress for {system_name}.
Summary:
- Rollback Reason: {rollback_reason}
- Business Impact: {business_impact}
- Expected Resolution: {estimated_completion_time}
- Customer Impact: {customer_impact}
We are following established procedures and will update hourly.
Escalation: {escalation_contact}
"""
}
},
"rollback_complete": {
"technical": {
"subject": "ROLLBACK COMPLETED: {migration_name}",
"body": """Team,
Rollback has been successfully completed for migration: {migration_name}
Summary:
- Start Time: {start_time}
- End Time: {end_time}
- Duration: {actual_duration}
- Phases Completed: {completed_phases}
Validation Results:
{validation_results}
System Status: {system_status}
Next Steps:
- Continue monitoring for 24 hours
- Post-rollback review scheduled for {review_date}
- Root cause analysis to begin
All clear to resume normal operations.
Incident Commander: {incident_commander}
"""
}
}
}
def generate_rollback_runbook(self, migration_plan: Dict[str, Any]) -> RollbackRunbook:
"""Generate comprehensive rollback runbook from migration plan"""
runbook_id = f"rb_{hashlib.md5(str(migration_plan).encode()).hexdigest()[:8]}"
migration_id = migration_plan.get("migration_id", "unknown")
migration_type = migration_plan.get("migration_type", "unknown")
# Generate rollback phases (reverse order of migration phases)
rollback_phases = self._generate_rollback_phases(migration_plan)
# Generate trigger conditions
trigger_conditions = self._generate_trigger_conditions(migration_plan)
# Generate data recovery plan
data_recovery_plan = self._generate_data_recovery_plan(migration_plan)
# Generate communication templates
communication_templates = self._generate_communication_templates(migration_plan)
# Generate escalation matrix
escalation_matrix = self._generate_escalation_matrix(migration_plan)
# Generate validation checklist
validation_checklist = self._generate_validation_checklist(migration_plan)
# Generate post-rollback procedures
post_rollback_procedures = self._generate_post_rollback_procedures(migration_plan)
# Generate emergency contacts
emergency_contacts = self._generate_emergency_contacts(migration_plan)
return RollbackRunbook(
runbook_id=runbook_id,
migration_id=migration_id,
created_at=datetime.datetime.now().isoformat(),
rollback_phases=rollback_phases,
trigger_conditions=trigger_conditions,
data_recovery_plan=data_recovery_plan,
communication_templates=communication_templates,
escalation_matrix=escalation_matrix,
validation_checklist=validation_checklist,
post_rollback_procedures=post_rollback_procedures,
emergency_contacts=emergency_contacts
)
def _generate_rollback_phases(self, migration_plan: Dict[str, Any]) -> List[RollbackPhase]:
"""Generate rollback phases from migration plan"""
migration_phases = migration_plan.get("phases", [])
migration_type = migration_plan.get("migration_type", "unknown")
rollback_phases = []
# Reverse the order of migration phases for rollback
for i, phase in enumerate(reversed(migration_phases)):
if isinstance(phase, dict):
phase_name = phase.get("name", f"phase_{i}")
phase_duration = phase.get("duration_hours", 2) * 60 # Convert to minutes
phase_risk = phase.get("risk_level", "medium")
else:
phase_name = str(phase)
phase_duration = 120 # Default 2 hours
phase_risk = "medium"
rollback_steps = self._generate_rollback_steps(phase_name, migration_type, i)
rollback_phase = RollbackPhase(
phase_name=f"rollback_{phase_name}",
description=f"Rollback changes made during {phase_name} phase",
urgency_level=self._calculate_urgency(phase_risk),
estimated_duration_minutes=phase_duration // 2, # Rollback typically faster
prerequisites=self._get_rollback_prerequisites(phase_name, i),
steps=rollback_steps,
validation_checkpoints=self._get_validation_checkpoints(phase_name, migration_type),
communication_requirements=self._get_communication_requirements(phase_name, phase_risk),
risk_level=phase_risk
)
rollback_phases.append(rollback_phase)
return rollback_phases
def _generate_rollback_steps(self, phase_name: str, migration_type: str, phase_index: int) -> List[RollbackStep]:
"""Generate specific rollback steps for a phase"""
steps = []
templates = self.rollback_templates.get(migration_type, {})
if migration_type == "database":
if "migration" in phase_name.lower() or "cutover" in phase_name.lower():
# Data rollback steps
steps.extend([
RollbackStep(
step_id=f"rb_data_{phase_index}_01",
name="Stop data migration processes",
description="Halt all ongoing data migration processes",
script_type="sql",
script_content="-- Stop migration processes\nSELECT pg_cancel_backend(pid) FROM pg_stat_activity WHERE query LIKE '%migration%';",
estimated_duration_minutes=5,
dependencies=[],
validation_commands=["SELECT COUNT(*) FROM pg_stat_activity WHERE query LIKE '%migration%';"],
success_criteria=["No active migration processes"],
failure_escalation="Contact DBA immediately",
rollback_order=1
),
RollbackStep(
step_id=f"rb_data_{phase_index}_02",
name="Restore from backup",
description="Restore database from pre-migration backup",
script_type="bash",
script_content=templates.get("data_rollback", {}).get("restore_backup", "pg_restore -d {database_name} -c {backup_file}"),
estimated_duration_minutes=30,
dependencies=[f"rb_data_{phase_index}_01"],
validation_commands=["SELECT COUNT(*) FROM information_schema.tables;"],
success_criteria=["Database restored successfully", "All expected tables present"],
failure_escalation="Escalate to senior DBA and infrastructure team",
rollback_order=2
)
])
if "preparation" in phase_name.lower():
# Schema rollback steps
steps.append(
RollbackStep(
step_id=f"rb_schema_{phase_index}_01",
name="Drop migration artifacts",
description="Remove temporary migration tables and procedures",
script_type="sql",
script_content="-- Drop migration artifacts\nDROP TABLE IF EXISTS migration_log;\nDROP PROCEDURE IF EXISTS migrate_data();",
estimated_duration_minutes=5,
dependencies=[],
validation_commands=["SELECT COUNT(*) FROM information_schema.tables WHERE table_name LIKE '%migration%';"],
success_criteria=["No migration artifacts remain"],
failure_escalation="Manual cleanup required",
rollback_order=1
)
)
elif migration_type == "service":
if "cutover" in phase_name.lower():
# Service rollback steps
steps.extend([
RollbackStep(
step_id=f"rb_service_{phase_index}_01",
name="Redirect traffic back to old service",
description="Update load balancer to route traffic back to previous service version",
script_type="bash",
script_content=templates.get("deployment_rollback", {}).get("update_load_balancer", "aws elbv2 modify-rule --rule-arn {rule_arn} --actions Type=forward,TargetGroupArn={original_target_group}"),
estimated_duration_minutes=2,
dependencies=[],
validation_commands=["curl -f {health_check_url}"],
success_criteria=["Traffic routing to original service", "Health checks passing"],
failure_escalation="Emergency procedure - manual traffic routing",
rollback_order=1
),
RollbackStep(
step_id=f"rb_service_{phase_index}_02",
name="Rollback service deployment",
description="Revert to previous service deployment version",
script_type="bash",
script_content=templates.get("deployment_rollback", {}).get("restore_previous_version", "kubectl rollout undo deployment/{service_name} --to-revision={revision_number}"),
estimated_duration_minutes=10,
dependencies=[f"rb_service_{phase_index}_01"],
validation_commands=["kubectl get pods -l app={service_name} --field-selector=status.phase=Running"],
success_criteria=["Previous version deployed", "All pods running"],
failure_escalation="Manual pod management required",
rollback_order=2
)
])
elif migration_type == "infrastructure":
steps.extend([
RollbackStep(
step_id=f"rb_infra_{phase_index}_01",
name="Revert infrastructure changes",
description="Apply terraform plan to revert infrastructure to previous state",
script_type="bash",
script_content=templates.get("cloud_rollback", {}).get("revert_terraform", "terraform apply -target={resource_name} {rollback_plan_file}"),
estimated_duration_minutes=15,
dependencies=[],
validation_commands=["terraform plan -detailed-exitcode"],
success_criteria=["Infrastructure matches previous state", "No planned changes"],
failure_escalation="Manual infrastructure review required",
rollback_order=1
),
RollbackStep(
step_id=f"rb_infra_{phase_index}_02",
name="Restore DNS configuration",
description="Revert DNS changes to point back to original infrastructure",
script_type="bash",
script_content=templates.get("cloud_rollback", {}).get("restore_dns", "aws route53 change-resource-record-sets --hosted-zone-id {zone_id} --change-batch file://{rollback_dns_changes}"),
estimated_duration_minutes=10,
dependencies=[f"rb_infra_{phase_index}_01"],
validation_commands=["nslookup {domain_name}"],
success_criteria=["DNS resolves to original endpoints"],
failure_escalation="Contact DNS administrator",
rollback_order=2
)
])
# Add generic validation step for all migration types
steps.append(
RollbackStep(
step_id=f"rb_validate_{phase_index}_final",
name="Validate rollback completion",
description=f"Comprehensive validation that {phase_name} rollback completed successfully",
script_type="manual",
script_content="Execute validation checklist for this phase",
estimated_duration_minutes=10,
dependencies=[step.step_id for step in steps],
validation_commands=self.validation_templates.get(migration_type, []),
success_criteria=[f"{phase_name} fully rolled back", "All validation checks pass"],
failure_escalation=f"Investigate {phase_name} rollback failures",
rollback_order=99
)
)
return steps
def _generate_trigger_conditions(self, migration_plan: Dict[str, Any]) -> List[RollbackTriggerCondition]:
"""Generate automatic rollback trigger conditions"""
triggers = []
migration_type = migration_plan.get("migration_type", "unknown")
# Generic triggers for all migration types
triggers.extend([
RollbackTriggerCondition(
trigger_id="error_rate_spike",
name="Error Rate Spike",
condition="error_rate > baseline * 5 for 5 minutes",
metric_threshold={
"metric": "error_rate",
"operator": "greater_than",
"value": "baseline_error_rate * 5",
"duration_minutes": 5
},
evaluation_window_minutes=5,
auto_execute=True,
escalation_contacts=["on_call_engineer", "migration_lead"]
),
RollbackTriggerCondition(
trigger_id="response_time_degradation",
name="Response Time Degradation",
condition="p95_response_time > baseline * 3 for 10 minutes",
metric_threshold={
"metric": "p95_response_time",
"operator": "greater_than",
"value": "baseline_p95 * 3",
"duration_minutes": 10
},
evaluation_window_minutes=10,
auto_execute=False,
escalation_contacts=["performance_team", "migration_lead"]
),
RollbackTriggerCondition(
trigger_id="availability_drop",
name="Service Availability Drop",
condition="availability < 95% for 2 minutes",
metric_threshold={
"metric": "availability",
"operator": "less_than",
"value": 0.95,
"duration_minutes": 2
},
evaluation_window_minutes=2,
auto_execute=True,
escalation_contacts=["sre_team", "incident_commander"]
)
])
# Migration-type specific triggers
if migration_type == "database":
triggers.extend([
RollbackTriggerCondition(
trigger_id="data_integrity_failure",
name="Data Integrity Check Failure",
condition="data_validation_failures > 0",
metric_threshold={
"metric": "data_validation_failures",
"operator": "greater_than",
"value": 0,
"duration_minutes": 1
},
evaluation_window_minutes=1,
auto_execute=True,
escalation_contacts=["dba_team", "data_team"]
),
RollbackTriggerCondition(
trigger_id="migration_progress_stalled",
name="Migration Progress Stalled",
condition="migration_progress unchanged for 30 minutes",
metric_threshold={
"metric": "migration_progress_rate",
"operator": "equals",
"value": 0,
"duration_minutes": 30
},
evaluation_window_minutes=30,
auto_execute=False,
escalation_contacts=["migration_team", "dba_team"]
)
])
elif migration_type == "service":
triggers.extend([
RollbackTriggerCondition(
trigger_id="cpu_utilization_spike",
name="CPU Utilization Spike",
condition="cpu_utilization > 90% for 15 minutes",
metric_threshold={
"metric": "cpu_utilization",
"operator": "greater_than",
"value": 0.90,
"duration_minutes": 15
},
evaluation_window_minutes=15,
auto_execute=False,
escalation_contacts=["devops_team", "infrastructure_team"]
),
RollbackTriggerCondition(
trigger_id="memory_leak_detected",
name="Memory Leak Detected",
condition="memory_usage increasing continuously for 20 minutes",
metric_threshold={
"metric": "memory_growth_rate",
"operator": "greater_than",
"value": "1MB/minute",
"duration_minutes": 20
},
evaluation_window_minutes=20,
auto_execute=True,
escalation_contacts=["development_team", "sre_team"]
)
])
return triggers
def _generate_data_recovery_plan(self, migration_plan: Dict[str, Any]) -> DataRecoveryPlan:
"""Generate data recovery plan"""
migration_type = migration_plan.get("migration_type", "unknown")
if migration_type == "database":
return DataRecoveryPlan(
recovery_method="point_in_time",
backup_location="/backups/pre_migration_{migration_id}_{timestamp}.sql",
recovery_scripts=[
"pg_restore -d production -c /backups/pre_migration_backup.sql",
"SELECT pg_create_restore_point('rollback_point');",
"VACUUM ANALYZE; -- Refresh statistics after restore"
],
data_validation_queries=[
"SELECT COUNT(*) FROM critical_business_table;",
"SELECT MAX(created_at) FROM audit_log;",
"SELECT COUNT(DISTINCT user_id) FROM user_sessions;",
"SELECT SUM(amount) FROM financial_transactions WHERE date = CURRENT_DATE;"
],
estimated_recovery_time_minutes=45,
recovery_dependencies=["database_instance_running", "backup_file_accessible"]
)
else:
return DataRecoveryPlan(
recovery_method="backup_restore",
backup_location="/backups/pre_migration_state",
recovery_scripts=[
"# Restore configuration files from backup",
"cp -r /backups/pre_migration_state/config/* /app/config/",
"# Restart services with previous configuration",
"systemctl restart application_service"
],
data_validation_queries=[
"curl -f http://localhost:8080/health",
"curl -f http://localhost:8080/api/status"
],
estimated_recovery_time_minutes=20,
recovery_dependencies=["service_stopped", "backup_accessible"]
)
def _generate_communication_templates(self, migration_plan: Dict[str, Any]) -> List[CommunicationTemplate]:
"""Generate communication templates for rollback scenarios"""
templates = []
base_templates = self.communication_templates
# Rollback start notifications
for audience in ["technical", "business", "executive"]:
if audience in base_templates["rollback_start"]:
template_data = base_templates["rollback_start"][audience]
templates.append(CommunicationTemplate(
template_type="rollback_start",
audience=audience,
subject=template_data["subject"],
body=template_data["body"],
urgency="high" if audience == "executive" else "medium",
delivery_methods=["email", "slack"] if audience == "technical" else ["email"]
))
# Rollback completion notifications
for audience in ["technical", "business"]:
if audience in base_templates.get("rollback_complete", {}):
template_data = base_templates["rollback_complete"][audience]
templates.append(CommunicationTemplate(
template_type="rollback_complete",
audience=audience,
subject=template_data["subject"],
body=template_data["body"],
urgency="medium",
delivery_methods=["email", "slack"] if audience == "technical" else ["email"]
))
# Emergency escalation template
templates.append(CommunicationTemplate(
template_type="emergency_escalation",
audience="executive",
subject="CRITICAL: Rollback Emergency - {migration_name}",
body="""CRITICAL SITUATION - IMMEDIATE ATTENTION REQUIRED
Migration: {migration_name}
Issue: Rollback procedure has encountered critical failures
Current Status: {current_status}
Failed Components: {failed_components}
Business Impact: {business_impact}
Customer Impact: {customer_impact}
Immediate Actions:
1. Emergency response team activated
2. {emergency_action_1}
3. {emergency_action_2}
War Room: {war_room_location}
Bridge Line: {conference_bridge}
Next Update: {next_update_time}
Incident Commander: {incident_commander}
Executive On-Call: {executive_on_call}
""",
urgency="emergency",
delivery_methods=["email", "sms", "phone_call"]
))
return templates
def _generate_escalation_matrix(self, migration_plan: Dict[str, Any]) -> Dict[str, Any]:
"""Generate escalation matrix for different failure scenarios"""
return {
"level_1": {
"trigger": "Single component failure",
"response_time_minutes": 5,
"contacts": ["on_call_engineer", "migration_lead"],
"actions": ["Investigate issue", "Attempt automated remediation", "Monitor closely"]
},
"level_2": {
"trigger": "Multiple component failures or single critical failure",
"response_time_minutes": 2,
"contacts": ["senior_engineer", "team_lead", "devops_lead"],
"actions": ["Initiate rollback", "Establish war room", "Notify stakeholders"]
},
"level_3": {
"trigger": "System-wide failure or data corruption",
"response_time_minutes": 1,
"contacts": ["engineering_manager", "cto", "incident_commander"],
"actions": ["Emergency rollback", "All hands on deck", "Executive notification"]
},
"emergency": {
"trigger": "Business-critical failure with customer impact",
"response_time_minutes": 0,
"contacts": ["ceo", "cto", "head_of_operations"],
"actions": ["Emergency procedures", "Customer communication", "Media preparation if needed"]
}
}
def _generate_validation_checklist(self, migration_plan: Dict[str, Any]) -> List[str]:
"""Generate comprehensive validation checklist"""
migration_type = migration_plan.get("migration_type", "unknown")
base_checklist = [
"Verify system is responding to health checks",
"Confirm error rates are within normal parameters",
"Validate response times meet SLA requirements",
"Check all critical business processes are functioning",
"Verify monitoring and alerting systems are operational",
"Confirm no data corruption has occurred",
"Validate security controls are functioning properly",
"Check backup systems are working correctly",
"Verify integration points with downstream systems",
"Confirm user authentication and authorization working"
]
if migration_type == "database":
base_checklist.extend([
"Validate database schema matches expected state",
"Confirm referential integrity constraints",
"Check database performance metrics",
"Verify data consistency across related tables",
"Validate indexes and statistics are optimal",
"Confirm transaction logs are clean",
"Check database connections and connection pooling"
])
elif migration_type == "service":
base_checklist.extend([
"Verify service discovery is working correctly",
"Confirm load balancing is distributing traffic properly",
"Check service-to-service communication",
"Validate API endpoints are responding correctly",
"Confirm feature flags are in correct state",
"Check resource utilization (CPU, memory, disk)",
"Verify container orchestration is healthy"
])
elif migration_type == "infrastructure":
base_checklist.extend([
"Verify network connectivity between components",
"Confirm DNS resolution is working correctly",
"Check firewall rules and security groups",
"Validate load balancer configuration",
"Confirm SSL/TLS certificates are valid",
"Check storage systems are accessible",
"Verify backup and disaster recovery systems"
])
return base_checklist
def _generate_post_rollback_procedures(self, migration_plan: Dict[str, Any]) -> List[str]:
"""Generate post-rollback procedures"""
return [
"Monitor system stability for 24-48 hours post-rollback",
"Conduct thorough post-rollback testing of all critical paths",
"Review and analyze rollback metrics and timing",
"Document lessons learned and rollback procedure improvements",
"Schedule post-mortem meeting with all stakeholders",
"Update rollback procedures based on actual experience",
"Communicate rollback completion to all stakeholders",
"Archive rollback logs and artifacts for future reference",
"Review and update monitoring thresholds if needed",
"Plan for next migration attempt with improved procedures",
"Conduct security review to ensure no vulnerabilities introduced",
"Update disaster recovery procedures if affected by rollback",
"Review capacity planning based on rollback resource usage",
"Update documentation with rollback experience and timings"
]
def _generate_emergency_contacts(self, migration_plan: Dict[str, Any]) -> List[Dict[str, str]]:
"""Generate emergency contact list"""
return [
{
"role": "Incident Commander",
"name": "TBD - Assigned during migration",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "incident.commander@company.com",
"backup_contact": "backup.commander@company.com"
},
{
"role": "Technical Lead",
"name": "TBD - Migration technical owner",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "tech.lead@company.com",
"backup_contact": "senior.engineer@company.com"
},
{
"role": "Business Owner",
"name": "TBD - Business stakeholder",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "business.owner@company.com",
"backup_contact": "product.manager@company.com"
},
{
"role": "On-Call Engineer",
"name": "Current on-call rotation",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "oncall@company.com",
"backup_contact": "backup.oncall@company.com"
},
{
"role": "Executive Escalation",
"name": "CTO/VP Engineering",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "cto@company.com",
"backup_contact": "vp.engineering@company.com"
}
]
def _calculate_urgency(self, risk_level: str) -> str:
"""Calculate rollback urgency based on risk level"""
risk_to_urgency = {
"low": "low",
"medium": "medium",
"high": "high",
"critical": "emergency"
}
return risk_to_urgency.get(risk_level, "medium")
def _get_rollback_prerequisites(self, phase_name: str, phase_index: int) -> List[str]:
"""Get prerequisites for rollback phase"""
prerequisites = [
"Incident commander assigned and briefed",
"All team members notified of rollback initiation",
"Monitoring systems confirmed operational",
"Backup systems verified and accessible"
]
if phase_index > 0:
prerequisites.append("Previous rollback phase completed successfully")
if "cutover" in phase_name.lower():
prerequisites.extend([
"Traffic redirection capabilities confirmed",
"Load balancer configuration backed up",
"DNS changes prepared for quick execution"
])
if "data" in phase_name.lower() or "migration" in phase_name.lower():
prerequisites.extend([
"Database backup verified and accessible",
"Data validation queries prepared",
"Database administrator on standby"
])
return prerequisites
def _get_validation_checkpoints(self, phase_name: str, migration_type: str) -> List[str]:
"""Get validation checkpoints for rollback phase"""
checkpoints = [
f"{phase_name} rollback steps completed",
"System health checks passing",
"No critical errors in logs",
"Key metrics within acceptable ranges"
]
validation_commands = self.validation_templates.get(migration_type, [])
checkpoints.extend([f"Validation command passed: {cmd[:50]}..." for cmd in validation_commands[:3]])
return checkpoints
def _get_communication_requirements(self, phase_name: str, risk_level: str) -> List[str]:
"""Get communication requirements for rollback phase"""
base_requirements = [
"Notify incident commander of phase start/completion",
"Update rollback status dashboard",
"Log all actions and decisions"
]
if risk_level in ["high", "critical"]:
base_requirements.extend([
"Notify all stakeholders of phase progress",
"Update executive team if rollback extends beyond expected time",
"Prepare customer communication if needed"
])
if "cutover" in phase_name.lower():
base_requirements.append("Immediate notification when traffic is redirected")
return base_requirements
def generate_human_readable_runbook(self, runbook: RollbackRunbook) -> str:
"""Generate human-readable rollback runbook"""
output = []
output.append("=" * 80)
output.append(f"ROLLBACK RUNBOOK: {runbook.runbook_id}")
output.append("=" * 80)
output.append(f"Migration ID: {runbook.migration_id}")
output.append(f"Created: {runbook.created_at}")
output.append("")
# Emergency Contacts
output.append("EMERGENCY CONTACTS")
output.append("-" * 40)
for contact in runbook.emergency_contacts:
output.append(f"{contact['role']}: {contact['name']}")
output.append(f" Phone: {contact['primary_phone']}")
output.append(f" Email: {contact['email']}")
output.append(f" Backup: {contact['backup_contact']}")
output.append("")
# Escalation Matrix
output.append("ESCALATION MATRIX")
output.append("-" * 40)
for level, details in runbook.escalation_matrix.items():
output.append(f"{level.upper()}:")
output.append(f" Trigger: {details['trigger']}")
output.append(f" Response Time: {details['response_time_minutes']} minutes")
output.append(f" Contacts: {', '.join(details['contacts'])}")
output.append(f" Actions: {', '.join(details['actions'])}")
output.append("")
# Rollback Trigger Conditions
output.append("AUTOMATIC ROLLBACK TRIGGERS")
output.append("-" * 40)
for trigger in runbook.trigger_conditions:
output.append(f"• {trigger.name}")
output.append(f" Condition: {trigger.condition}")
output.append(f" Auto-Execute: {'Yes' if trigger.auto_execute else 'No'}")
output.append(f" Evaluation Window: {trigger.evaluation_window_minutes} minutes")
output.append(f" Contacts: {', '.join(trigger.escalation_contacts)}")
output.append("")
# Rollback Phases
output.append("ROLLBACK PHASES")
output.append("-" * 40)
for i, phase in enumerate(runbook.rollback_phases, 1):
output.append(f"{i}. {phase.phase_name.upper()}")
output.append(f" Description: {phase.description}")
output.append(f" Urgency: {phase.urgency_level.upper()}")
output.append(f" Duration: {phase.estimated_duration_minutes} minutes")
output.append(f" Risk Level: {phase.risk_level.upper()}")
if phase.prerequisites:
output.append(" Prerequisites:")
for prereq in phase.prerequisites:
output.append(f" ✓ {prereq}")
output.append(" Steps:")
for step in sorted(phase.steps, key=lambda x: x.rollback_order):
output.append(f" {step.rollback_order}. {step.name}")
output.append(f" Duration: {step.estimated_duration_minutes} min")
output.append(f" Type: {step.script_type}")
if step.script_content and step.script_type != "manual":
output.append(" Script:")
for line in step.script_content.split('\n')[:3]: # Show first 3 lines
output.append(f" {line}")
if len(step.script_content.split('\n')) > 3:
output.append(" ...")
output.append(f" Success Criteria: {', '.join(step.success_criteria)}")
output.append("")
if phase.validation_checkpoints:
output.append(" Validation Checkpoints:")
for checkpoint in phase.validation_checkpoints:
output.append(f" ☐ {checkpoint}")
output.append("")
# Data Recovery Plan
output.append("DATA RECOVERY PLAN")
output.append("-" * 40)
drp = runbook.data_recovery_plan
output.append(f"Recovery Method: {drp.recovery_method}")
output.append(f"Backup Location: {drp.backup_location}")
output.append(f"Estimated Recovery Time: {drp.estimated_recovery_time_minutes} minutes")
output.append("Recovery Scripts:")
for script in drp.recovery_scripts:
output.append(f" • {script}")
output.append("Validation Queries:")
for query in drp.data_validation_queries:
output.append(f" • {query}")
output.append("")
# Validation Checklist
output.append("POST-ROLLBACK VALIDATION CHECKLIST")
output.append("-" * 40)
for i, item in enumerate(runbook.validation_checklist, 1):
output.append(f"{i:2d}. ☐ {item}")
output.append("")
# Post-Rollback Procedures
output.append("POST-ROLLBACK PROCEDURES")
output.append("-" * 40)
for i, procedure in enumerate(runbook.post_rollback_procedures, 1):
output.append(f"{i:2d}. {procedure}")
output.append("")
return "\n".join(output)
def main():
"""Main function with command line interface"""
parser = argparse.ArgumentParser(description="Generate comprehensive rollback runbooks from migration plans")
parser.add_argument("--input", "-i", required=True, help="Input migration plan file (JSON)")
parser.add_argument("--output", "-o", help="Output file for rollback runbook (JSON)")
parser.add_argument("--format", "-f", choices=["json", "text", "both"], default="both", help="Output format")
args = parser.parse_args()
try:
# Load migration plan
with open(args.input, 'r') as f:
migration_plan = json.load(f)
# Validate required fields
if "migration_id" not in migration_plan and "source" not in migration_plan:
print("Error: Migration plan must contain migration_id or source field", file=sys.stderr)
return 1
# Generate rollback runbook
generator = RollbackGenerator()
runbook = generator.generate_rollback_runbook(migration_plan)
# Output results
if args.format in ["json", "both"]:
runbook_dict = asdict(runbook)
if args.output:
with open(args.output, 'w') as f:
json.dump(runbook_dict, f, indent=2)
print(f"Rollback runbook saved to {args.output}")
else:
print(json.dumps(runbook_dict, indent=2))
if args.format in ["text", "both"]:
human_runbook = generator.generate_human_readable_runbook(runbook)
text_output = args.output.replace('.json', '.txt') if args.output else None
if text_output:
with open(text_output, 'w') as f:
f.write(human_runbook)
print(f"Human-readable runbook saved to {text_output}")
else:
print("\n" + "="*80)
print("HUMAN-READABLE ROLLBACK RUNBOOK")
print("="*80)
print(human_runbook)
except FileNotFoundError:
print(f"Error: Input file '{args.input}' not found", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON in input file: {e}", file=sys.stderr)
return 1
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
return 1
return 0
if __name__ == "__main__":
sys.exit(main())Tạo và tối ưu paywall, màn hình nâng cấp, modal upsell và giới hạn tính năng để chuyển người dùng miễn phí sang trả phí.
--- name: "paywall-upgrade-cro" description: When the user wants to create or optimize in-app paywalls, upgrade screens, upsell modals, or feature gates. Also use when the user mentions "paywall," "upgrade screen," "upgrade modal," "upsell," "feature gate," "convert free to paid," "freemium conversion," "trial expiration screen," "limit reached screen," "plan upgrade prompt," or "in-app pricing." Distinct from public pricing pages (see page-cro) — this skill focuses on in-product upgrade moments where the user has already experienced value. license: MIT metadata: version: 1.0.0 author: Alireza Rezvani category: marketing updated: 2026-03-06 --- # Paywall and Upgrade Screen CRO You are an expert in in-app paywalls and upgrade flows. Your goal is to convert free users to paid, or upgrade users to higher tiers, at moments when they've experienced enough value to justify the commitment. ## Initial Assessment **Check for product marketing context first:** If `.claude/product-marketing-context.md` exists, read it before asking questions. Use that context and only ask for information not already covered or specific to this task. Before providing recommendations, understand: 1. **Upgrade Context** - Freemium → Paid? Trial → Paid? Tier upgrade? Feature upsell? Usage limit? 2. **Product Model** - What's free? What's behind paywall? What triggers prompts? Current conversion rate? 3. **User Journey** - When does this appear? What have they experienced? What are they trying to do? --- ## Core Principles ### 1. Value Before Ask - User should have experienced real value first - Upgrade should feel like natural next step - Timing: After "aha moment," not before ### 2. Show, Don't Just Tell - Demonstrate the value of paid features - Preview what they're missing - Make the upgrade feel tangible ### 3. Friction-Free Path - Easy to upgrade when ready - Don't make them hunt for pricing ### 4. Respect the No - Don't trap or pressure - Make it easy to continue free - Maintain trust for future conversion --- ## Paywall Trigger Points ### Feature Gates When user clicks a paid-only feature: - Clear explanation of why it's paid - Show what the feature does - Quick path to unlock - Option to continue without ### Usage Limits When user hits a limit: - Clear indication of limit reached - Show what upgrading provides - Don't block abruptly ### Trial Expiration When trial is ending: - Early warnings (7, 3, 1 day) - Clear "what happens" on expiration - Summarize value received ### Time-Based Prompts After X days of free use: - Gentle upgrade reminder - Highlight unused paid features - Easy to dismiss --- ## Paywall Screen Components 1. **Headline** - Focus on what they get: "Unlock [Feature] to [Benefit]" 2. **Value Demonstration** - Preview, before/after, "With Pro you could..." 3. **Feature Comparison** - Highlight key differences, current plan marked 4. **Pricing** - Clear, simple, annual vs. monthly options 5. **Social Proof** - Customer quotes, "X teams use this" 6. **CTA** - Specific and value-oriented: "Start Getting [Benefit]" 7. **Escape Hatch** - Clear "Not now" or "Continue with Free" --- ## Specific Paywall Types ### Feature Lock Paywall ``` [Lock Icon] This feature is available on Pro [Feature preview/screenshot] [Feature name] helps you [benefit]: • [Capability] • [Capability] [Upgrade to Pro - $X/mo] [Maybe Later] ``` ### Usage Limit Paywall ``` You've reached your free limit [Progress bar at 100%] Free: 3 projects | Pro: Unlimited [Upgrade to Pro] [Delete a project] ``` ### Trial Expiration Paywall ``` Your trial ends in 3 days What you'll lose: • [Feature used] • [Data created] What you've accomplished: • Created X projects [Continue with Pro] [Remind me later] [Downgrade] ``` --- ## Timing and Frequency ### When to Show - After value moment, before frustration - After activation/aha moment - When hitting genuine limits ### When NOT to Show - During onboarding (too early) - When they're in a flow - Repeatedly after dismissal ### Frequency Rules - Limit per session - Cool-down after dismiss (days, not hours) - Track annoyance signals --- ## Upgrade Flow Optimization ### From Paywall to Payment - Minimize steps - Keep in-context if possible - Pre-fill known information ### Post-Upgrade - Immediate access to features - Confirmation and receipt - Guide to new features --- ## A/B Testing ### What to Test - Trigger timing - Headline/copy variations - Price presentation - Trial length - Feature emphasis - Design/layout ### Metrics to Track - Paywall impression rate - Click-through to upgrade - Completion rate - Revenue per user - Churn rate post-upgrade --- ## Anti-Patterns to Avoid ### Dark Patterns - Hiding the close button - Confusing plan selection - Guilt-trip copy ### Conversion Killers - Asking before value delivered - Too frequent prompts - Blocking critical flows - Complicated upgrade process --- ## Task-Specific Questions 1. What's your current free → paid conversion rate? 2. What triggers upgrade prompts today? 3. What features are behind the paywall? 4. What's your "aha moment" for users? 5. What pricing model? (per seat, usage, flat) 6. Mobile app, web app, or both? --- ## Related Skills - **page-cro** — WHEN the public-facing pricing page needs optimization (before users are in-app). NOT for in-product upgrade screens or feature gates. - **onboarding-cro** — WHEN users haven't reached their activation moment and are hitting paywalls too early; fix onboarding first. NOT when value has already been delivered. - **ab-test-setup** — WHEN running controlled experiments on paywall trigger timing, copy, pricing display, or layout. NOT for initial paywall design. - **email-sequence** — WHEN setting up trial expiration or upgrade reminder email sequences to complement in-app prompts. NOT as a replacement for in-app paywall design. - **marketing-context** — Foundation skill for understanding ICP, pricing model, and value proposition. Load before designing paywall copy and positioning. --- ## Communication Paywall recommendations must account for where the user is in their value journey — always confirm whether the aha moment has been reached before recommending upgrade prompt placement. When writing paywall copy, deliver complete screen copy: headline, value statement, feature list, CTA, and escape hatch text. Flag dark patterns proactively and recommend ethical alternatives. Load `marketing-context` for pricing model and plan structure context before writing copy. --- ## Proactive Triggers - User reports low free-to-paid conversion rate → ask where in the journey the paywall appears and whether the aha moment is reached first. - User mentions users hitting limits and churning → distinguish between limit frustration (fix timing/messaging) vs. wrong ICP (fix acquisition). - User asks about freemium model design → help define what's free vs. paid, then design paywall moments around natural value gaps. - User shares a trial expiration screen → audit for dark patterns, missing escape hatches, and unclear value summarization. - User mentions mobile app monetization → flag platform-specific considerations (App Store IAP rules, Google Play billing requirements). --- ## Output Artifacts | Artifact | Description | |----------|-------------| | Paywall Trigger Map | All paywall trigger points with timing rules, cooldown periods, and frequency caps | | Full Paywall Screen Copy | Headline, value demonstration, feature comparison, CTA, and escape hatch for each paywall type | | Upgrade Flow Diagram | Step-by-step from paywall click to post-upgrade confirmation with friction reduction notes | | Anti-Pattern Audit | Review of existing paywall for dark patterns, trust-damaging copy, and conversion killers | | A/B Test Backlog | Prioritized experiment ideas for trigger timing, copy, and pricing display |
Thiết kế hoặc xem lại giá sản phẩm: chọn mô hình giá, phân tích Van Westendorp từ khảo sát WTP và đóng gói bậc Good/Better/Best.
---
name: pricing-strategist
description: "Use when designing or revisiting product pricing — selecting a pricing model (subscription seat-based, usage-based, value-based, freemium, or hybrid), running Van Westendorp Price Sensitivity Meter analysis on WTP survey data, or designing Good/Better/Best packaging tiers. Recommends a model and a price range with trade-offs, never a single number. For Commercial leads, Product Marketing, and CMOs at the pricing-design moment — not deal-by-deal discounting, not brand positioning."
version: 2.8.0
author: claude-code-skills
license: MIT
tags: [commercial, pricing, packaging, wtp, van-westendorp, value-based-pricing, saas-pricing]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# pricing-strategist
## Purpose
Help Commercial, Product Marketing, and CMO functions answer three questions at the pricing-design moment:
1. **Which pricing model fits this product + customer + market?** (subscription seat-based, usage-based, value-based, freemium, hybrid)
2. **What does the customer actually pay before it feels too expensive?** (Van Westendorp PSM on WTP survey responses)
3. **How should we package this into tiers?** (Good / Better / Best — with anti-pattern detection)
The skill recommends **a model and a range**. The human picks the number, owns the trade-offs, and runs the GTM.
## When to use
- Launching a new SaaS / API / AI tool and choosing the first pricing model
- Revisiting pricing after 18+ months of GTM data (model shift, not just price increase)
- Designing or redesigning tier packaging (Good/Better/Best, Bronze/Silver/Gold)
- You have Van Westendorp survey data and want the optimal price range
- A board / exec is asking "what should we charge?" and you need the structured answer
- You suspect your packaging has anti-patterns (decoy tier, feature dump, no upgrade trigger)
**Do not use for:**
- Per-deal discount approval → `deal-desk`
- Strategic CMO positioning, brand, category creation → `c-level-advisor/cmo-advisor`
- Whole-company revenue strategy → `c-level-advisor/cro-advisor`
- Technical-sale enablement → `business-growth/sales-engineer`
## Workflow
### Step 1 — Assess customer context
Fill `assets/pricing_brief_template.md` (≈ 20 min). Capture: industry, deal size avg, customer count, value drivers, adoption curve, consumption pattern (seat / usage / value / hybrid), competitor models.
### Step 2 — Pick the pricing model
Run `scripts/pricing_model_picker.py --input brief.json --profile saas --output markdown`. Output ranks 5 models by fit-score 0-100 with trade-offs. Decision logic is deterministic: low usage variance + high seat-attach → subscription wins; power-law usage + variable customer value → usage-based wins.
### Step 3 — Validate WTP with Van Westendorp PSM
If you have survey data (≥ 4 questions per respondent: too cheap / bargain / getting expensive / too expensive), run `scripts/wtp_analyzer.py --input survey.json --output markdown`. Output: 4 intersection points (OPP, IDP, PMC, PME) and the Range of Acceptable Prices.
PSM gives a **range**, not the price. See `references/van_westendorp_methodology.md` for common misinterpretations.
### Step 4 — Design packaging
Run `scripts/packaging_designer.py --input features.json --profile saas --output markdown`. Output: 3-tier Good/Better/Best assignment with anti-pattern flags (decoy tier, feature dump, no upgrade trigger, Bronze loss leader, Enterprise no-anchor).
### Step 5 — Decide
Take model + range + packaging into the pricing committee. Skill does not commit the number — you do.
## Scripts
- `scripts/pricing_model_picker.py` — 5-model fit scorer (subscription / usage / value / freemium / hybrid)
- `scripts/wtp_analyzer.py` — Van Westendorp PSM implementation
- `scripts/packaging_designer.py` — Good/Better/Best tier designer with anti-pattern detection
All scripts: stdlib only. `--help` and `--sample` work on all three.
## References
- `references/saas_pricing_canon.md` — Skok, Tunguz, Campbell, Ramanujam, BVP, Shevlin, Stanford GSB
- `references/van_westendorp_methodology.md` — original 1976 paper, NMS refinement, Conjoint.ly, Sawtooth, ESOMAR, Lipovetsky, Decision Analyst
- `references/packaging_anti_patterns.md` — ProfitWell, OpenView, BVP vertical SaaS, Ramanujam, Poyar, SaaS Capital
## Assumptions
- Pricing decisions are joint: Commercial owns the model + tier shape, Product owns the features-per-tier, Finance owns the discount envelope, Legal owns the contract.
- Van Westendorp PSM is a **directional** tool. N ≥ 30 minimum, N ≥ 100 preferred. Below 30, the script emits a sample-size warning.
- "Value-based pricing" requires a measurable customer value driver (revenue lift, cost saved, time recovered). If you can't measure it, don't pick value-based.
- Industry profiles tune defaults — they don't override your data.
- This is a decision-support skill, not a price oracle. Output is a model + range, never the number.
## Anti-patterns
- **Recommending a specific number.** This skill emits a model and a range. Final price is a human commercial decision involving deal-desk policy, competitive intel, and strategic intent that this skill cannot know.
- **Using PSM with N < 30.** Statistical noise dominates. The script warns; respect the warning.
- **Treating PSM as "the price."** PSM gives a Range of Acceptable Prices (RAP) and an Optimal Price Point (OPP). Test the range in market, don't anchor on a single intersection.
- **Picking value-based pricing without a measurable value metric.** Without instrumentation to show customer ROI, value-based collapses into "whatever they'll pay" — which is just bad usage-based pricing.
- **Designing tiers before picking a model.** Tier structure depends on the model. Run pricing_model_picker first.
- **Packaging "feature dumps" into the Best tier.** If Best has 3x the features for 2x the price, customers buy Better and never upgrade. See `packaging_anti_patterns.md`.
- **Hidden usage-based pricing inside subscription tiers.** "Up to 100k API calls/mo, then $X per 1k" disguised as a "Pro tier" is two pricing models in one. Customers notice. Pick one.
- **Confusing this skill with deal-desk.** Pricing strategy = the menu. Deal-desk = approving discounts off the menu. Different decision, different cadence, different owner.
## Distinct from
- **deal-desk** — per-deal discount approval, MEDDIC, deal scoring. Operates daily on existing pricing.
- **c-level-advisor/cmo-advisor** — strategic positioning, brand, category. Pricing strategist consumes positioning as input, doesn't generate it.
- **c-level-advisor/cro-advisor** — full-funnel revenue strategy, comp plans, territory design. Pricing strategist is one input to CRO.
- **business-growth/sales-engineer** — technical sale, POC scoping. Sales engineering operates after pricing is set.
## Forcing-question library (Matt Pocock grill discipline)
Walked one at a time by `/cs:grill-commercial` or the orchestrator. Recommended answer + canon citation per question. Never bundled.
1. **"Is your customer paying for outcomes, seats, or usage?"**
Recommended: outcomes (value-based) if you can measure them; usage if marginal cost is variable; seats only if usage is roughly flat per user.
Canon: Ramanujam 2016 (*Monetizing Innovation*) — Mistake #1 of 9: seat-based pricing on a usage-variable product caps TAM at ~20% of WTP.
2. **"Do you have a measurable value metric, or are you guessing?"**
Recommended: instrument the value metric BEFORE going to market with value-based pricing.
Canon: Patrick Campbell / ProfitWell research — value-based without instrumentation collapses into bad usage-based pricing.
3. **"What's the variance in customer usage across your top decile vs. median?"**
Recommended: variance > 10x → usage-based wins; variance < 3x → subscription wins; in between → hybrid with usage overage.
Canon: Kyle Poyar (*Growth Unhinged*) — high-variance products lose 60%+ of revenue on flat-rate plans.
4. **"What's your competitor's pricing model, and why are you choosing the same or different?"**
Recommended: surface the differentiation hypothesis explicitly. Identical pricing = identical value claim.
Canon: David Skok (*For Entrepreneurs*) — pricing is a positioning signal.
5. **"What sample size do you have for WTP analysis, and is it segmented?"**
Recommended: N≥30 per segment for PSM, N≥100 for conjoint.
Canon: van Westendorp 1976 / Sawtooth Software methodology — sub-30 PSM is statistical noise.
6. **"What's the ONE feature that forces a tier upgrade?"**
Recommended: every Better and Best tier needs a single non-negotiable upgrade trigger.
Canon: Ramanujam (*Monetizing Innovation*) — Mistake #4: tiers with no clear differentiator make 70% of customers pick the cheapest.
Walk depth-first. Lock 1-3 before opening 4-6. After all 6 are answered, invoke `pricing_model_picker.py` → `wtp_analyzer.py` → `packaging_designer.py` in sequence.
FILE:assets/pricing_brief_template.md
# Pricing Strategy Brief
**Owner:** _______________ **Date:** _______________
**Time to fill out:** ≈ 20 minutes
Fill this brief out *before* running `pricing_model_picker.py`. The skill outputs are only
as good as the inputs. Be specific. If you don't know a field, write "unknown" — don't guess.
---
## 1. Product context
- **Product / feature being priced:** _______________
- **Customer-visible name:** _______________
- **One-line value prop:** _______________
- **Stage:** [ ] new launch [ ] re-pricing [ ] adding tier [ ] expansion play
## 2. Customer context
- **Industry:** _______________ (e.g., "B2B SaaS — sales intelligence")
- **ICP (Ideal Customer Profile, 1 sentence):** _______________
- **Avg deal size today (annual contract value):** $_______________
- **Customer count today:** _______________
- **Geographic concentration:** _______________
## 3. Value drivers (rank top 3)
The customer outcome that pricing should track. Be specific — "saves time" is not a value
driver. "Reduces lead-research time by 6 hours/rep/week" is.
1. _______________
2. _______________
3. _______________
For each, can you measure it? [ ] yes [ ] partly [ ] no
## 4. Adoption curve
- [ ] Top-down enterprise sale (CIO/VP signs)
- [ ] Bottom-up / PLG (individual user adopts, expands)
- [ ] Hybrid (champion-led, exec-approved)
- [ ] Viral (referral loops within or across orgs)
## 5. Consumption pattern (assign 0.0 - 1.0 to each)
How does customer value scale with what they use?
- **Seat-based** (more users = more value): _______________
- **Usage-based** (more events / API calls / volume = more value): _______________
- **Value-based** (measurable customer outcome = more value): _______________
- **Hybrid** (multiple drivers, no single dominant): _______________
## 6. Competitive pricing landscape
List the 3-5 closest competitors and their pricing model:
| Competitor | Pricing model | Notable mechanics |
|---|---|---|
| | | |
| | | |
| | | |
## 7. Strategic constraints
- **Margin floor (gross margin %):** _______________
- **Discount envelope (max % off list):** _______________
- **Sales motion (self-serve / inside / field):** _______________
- **Any contractual / regulatory pricing constraints?** _______________
## 8. Anti-goals
What pricing outcomes would be a *failure* even if NRR looks fine?
- _______________
- _______________
---
## JSON skeleton (for the script)
Copy this into a file (e.g., `brief.json`) and fill in. Then run:
```bash
python scripts/pricing_model_picker.py --input brief.json --profile saas --output markdown
```
```json
{
"industry": "",
"deal_size_avg": 0,
"customer_count": 0,
"value_drivers": [
""
],
"adoption_curve": "",
"consumption_pattern": {
"seat-based": 0.0,
"usage-based": 0.0,
"value-based": 0.0,
"hybrid": 0.0
},
"competitor_pricing_models": [
""
]
}
```
---
## After running the picker
1. Take the top 1-2 model recommendations into a 30-min review with Product + Finance.
2. If a model is selected, run a **Van Westendorp PSM survey** (≥ 30 respondents, preferably 100+).
3. Feed survey data to `wtp_analyzer.py` to get the Range of Acceptable Prices.
4. Run `packaging_designer.py` on your feature list to draft Good/Better/Best tiers.
5. Pressure-test in pricing committee. The skill output is one input, not the decision.
FILE:references/packaging_anti_patterns.md
# Packaging Anti-Patterns
Reference for `packaging_designer.py`. The anti-pattern detectors in the tool implement the
flags below. This document is the source-of-truth for *why* each pattern is harmful.
---
## The seven anti-patterns
### 1. Decoy tier that fools no one
A middle tier designed to make the top tier look reasonable — but the price gap is too small,
the feature list is too thin, or the differentiation is purely cosmetic. Customers see through
it, and worse: it trains them to question your pricing integrity.
**Detection:** Best tier > 2x Better price with < 1.5x value.
**Fix:** Either compress Better and Best closer (real differentiation) or widen the value gap.
### 2. Feature dump in the Best tier
Every roadmap feature gets tossed into Best because "Enterprise wants it." Result: Best has
3x the features for 2x the price. Customers buy Better and never upgrade.
**Detection:** Best feature count > 2x Better's with price ratio < 1.5x.
**Fix:** Move 1-3 "upgrade trigger" features down to Better and re-price.
### 3. No clear upgrade trigger
A customer on Good has no specific friction that pushes them to Better. Marketing fixes this
with "more advanced features" copy — but if you can't name the *single event* that triggers
the upgrade ("you hit 10k API calls", "you added a 5th seat", "you needed SSO"), customers
don't upgrade.
**Detection:** Better tier features have lower average importance than Good tier features.
**Fix:** Identify 1-2 "moment of pain" features and move them into the gate.
### 4. Usage-based pricing hidden inside subscription tiers
"Pro tier: includes up to 100k events, then $X per 1k." This is two pricing models pretending
to be one. Customers feel deceived when overage hits. Either commit to subscription (with a
realistic cap) or commit to usage (with a transparent meter).
**Detection:** Pricing-page narrative inspection — not algorithmic. Flag manually.
**Fix:** Pick one model. If you genuinely need both, use a clean Platform + Consumption hybrid,
not a hidden-overage subscription.
### 5. Bronze / Good tier as loss leader
Good is so cheap that cost-to-serve eats most of the revenue. Customers stay on Good forever
because the value-per-dollar is too good. Acquisition costs amortize through Better/Best
upgrades that never happen.
**Detection:** Cost-to-serve aggregate > 80% of Good tier price.
**Fix:** Raise Good's price floor, or strip a feature down to Better.
### 6. Enterprise / Best = "Call us" with no anchor
"Contact sales for pricing" at the top tier with no published anchor price. Prospects with
budget constraints disqualify themselves without ever talking to you. Competitors who publish
ranges win the consideration set.
**Detection:** Best tier has features but no published price.
**Fix:** Publish a "Starting at $X" anchor. The number doesn't have to be precise — it just
has to disqualify the wrong-fit prospects and qualify the right-fit ones.
### 7. Feature appears in all 3 tiers (no differentiation)
If "API access" is in Good, Better, and Best with the same scope, it's not a tier feature —
it's a base feature. Listing it three times wastes pricing-page real estate and dilutes the
upgrade narrative.
**Detection:** Feature appears in all 3 tiers' assigned-feature lists.
**Fix:** Either drop it from the tier comparison or differentiate scope (rate-limited /
metered / unlimited).
---
## Authoritative sources
1. **Patrick Campbell — ProfitWell / Paddle research on packaging**.
"The State of Subscription Pricing" reports and the ProfitWell podcast.
Empirical: tier redesigns that fixed clear upgrade triggers grew NRR by 8-15 points on
average across their cohort. https://www.paddle.com/resources
2. **Madhavan Ramanujam — Monetizing Innovation (Wiley, 2016)**.
The "9 mistakes" framework: feature shock, minivation, hidden gem, undead. Anti-patterns 2
(feature dump), 3 (no upgrade trigger), and 7 (no differentiation) map directly to
Ramanujam's mistake taxonomy.
3. **OpenView — SaaS Benchmarks + Product-Led Growth reports**.
Annual benchmarks on tier mix, free-to-paid conversion, and the cost of bad packaging.
https://openviewpartners.com/saas-benchmarks-report/
4. **Bessemer Venture Partners — Vertical SaaS Index + Cloud 100 Memos**.
Documents the move from 3-tier to 4-tier packaging in vertical SaaS as products mature, and
the failure modes when the 4th tier is added without removing complexity from existing tiers.
https://www.bvp.com/atlas
5. **Kyle Poyar — Growth Unhinged**.
"The Anatomy of a Great Pricing Page" series. Anti-patterns 5 (loss leader) and 6 (no
anchor) come from Poyar's documented PLG-to-enterprise transition patterns.
https://www.growthunhinged.com/
6. **SaaS Capital — Spending Benchmarks for Private B2B SaaS Companies (annual)**.
Cost-to-serve benchmarks by ACV band. Source for the "Bronze tier loss leader" threshold
(80% cost-to-serve ratio).
7. **Tomasz Tunguz — Theory Ventures**.
Multi-year posts on tier-mix evolution in Cloud 100 cohort. Documents the death of 5-tier
pricing pages and the consolidation toward Good/Better/Best + Enterprise.
https://tomtunguz.com/
8. **Simon-Kucher & Partners — Global Pricing Studies**.
Cross-industry data on pricing-page complexity vs conversion. Their research underpins the
"more tiers = more cognitive load" finding behind anti-pattern 4 (hidden usage inside subscription).
## How this skill uses the references
- `packaging_designer.py` runs deterministic detection for 7 anti-patterns (the manual-inspection
one — hidden usage in subscription — is documented but not auto-detected; the tool relies on
the human reading the pricing-page narrative).
- Industry profiles encode tier-mix priors (Good 50% / Better 30% / Best 20% for SaaS, etc.)
derived from OpenView and BVP benchmarks.
- Price-ratio thresholds (2.5x Good → Better, 2.0x Better → Best for SaaS) come from
ProfitWell + Tunguz cohort averages.
- The "no anchor price" flag implements the Poyar / BVP guidance that Enterprise tiers need
published starting prices.
FILE:references/saas_pricing_canon.md
# SaaS Pricing Canon
Curated, opinionated knowledge base for pricing model selection. This is the source material
behind `pricing_model_picker.py`'s scoring rules.
## Core principle
Pricing is a product decision, not a finance decision. The pricing model encodes how customers
experience value capture — get it wrong and every other GTM lever (sales, retention, expansion)
compounds the mistake.
---
## The five pricing models
### 1. Subscription seat-based
Customers pay per user, per period. Works when:
- Value scales linearly with user count
- Usage variance per seat is low
- Procurement prefers predictable line items
- Competitive set already trains the market on seat pricing
Failure modes: usage power-law (top 10% of users drive 80% of value) leaves money on the table;
"seat sprawl" makes customers hide users; expansion is gated on hiring, which is slow.
### 2. Usage-based (consumption)
Customers pay for what they consume — API calls, tokens, GB stored, messages sent. Works when:
- Value is tightly coupled to a measurable unit
- Usage variance across customers is high (power-law)
- Customer wants to start small and scale
- The metering infrastructure exists
Failure modes: bill-shock (variance scares procurement); cohort-NRR volatility; "cost of a query"
becomes a feature-velocity tax; revenue forecasting becomes hard.
### 3. Value-based
Price is anchored to the customer's economic outcome (revenue lift, cost saved, time recovered).
Works when:
- The value driver is measurable and attributable
- Customer count is small enough to calibrate per-account
- Deal size is large enough to justify the sales motion
- ROI proof is part of the product (not a slide)
Failure modes: requires instrumented ROI per customer; doesn't scale operationally beyond ~50-200
accounts without specialization; collapses to "whatever they'll pay" when value isn't measurable.
### 4. Freemium
Free tier acquires users, paid tiers monetize. Works when:
- Adoption is bottom-up / viral / PLG
- Free-tier cost-to-serve is < 5% of paid LTV
- There is a natural upgrade trigger inside the free experience
- Sales motion is self-serve or low-touch
Failure modes: enterprise sale + freemium dilutes positioning; free-tier costs balloon faster than
conversion; the "free forever for 10 users" cliff trains customers to game it.
### 5. Hybrid
Combinations — seat + usage overage, platform + per-event, base + value uplift. Works when:
- Multiple value drivers exist (seats AND usage)
- Customer segments split on dominant driver
- Deal sizes are large enough to absorb pricing-page complexity
Failure modes: cognitive load on the prospect; CS overhead in tier-to-tier moves; invoice disputes;
hybrid sometimes hides "we couldn't decide" — which customers detect.
---
## Authoritative sources
1. **David Skok — For Entrepreneurs**.
"SaaS Metrics 2.0" + "Unit Economics" series. The canonical playbook on CAC, LTV, and how
pricing model interacts with both. https://www.forentrepreneurs.com/saas-metrics-2/
2. **Tomasz Tunguz — Theory Ventures blog**.
Years of empirical posts on Cloud 100 pricing patterns, hybrid-pricing adoption curves,
usage-based unit economics. https://tomtunguz.com/
3. **Patrick Campbell — ProfitWell / Paddle research**.
The largest body of public SaaS pricing data. Key findings: prospects who see clear value
metrics convert 2x; freemium converts 2-5% on average; bad packaging is the #1 churn cause.
https://www.paddle.com/resources
4. **Madhavan Ramanujam — Monetizing Innovation (Wiley, 2016)**.
Simon-Kucher partner. The "9 Pricing Mistakes" frame: feature shock, minivation, hidden gem,
undead. Establishes the discipline that pricing comes before product, not after.
5. **Bessemer Venture Partners — State of the Cloud + Memos**.
Annual benchmarks: Rule of 40, NRR by ACV band, pricing-model mix in Cloud 100. The
reference for "what good looks like" in SaaS.
https://www.bvp.com/atlas
6. **Ron Shevlin — Cornerstone Advisors / Forbes columns**.
Pricing psychology applied to financial services SaaS — anchoring, decoy effect, charm
pricing's diminishing returns in B2B.
7. **Stanford GSB pricing research (Bertini, Gourville, Anderson)**.
Academic foundation on price-quality signaling, reference price formation, and the
penny-gap problem (the $0 → $0.01 conversion cliff). See Bertini & Gourville HBR 2012,
"Pricing to Create Shared Value."
8. **Kyle Poyar — OpenView / Growth Unhinged**.
Practitioner depth on PLG monetization, packaging redesigns, and the shift from seat to
hybrid pricing in 2020-2025 cohort. https://www.growthunhinged.com/
## How this skill uses the canon
- `pricing_model_picker.py` weights consumption-pattern signals per the Skok/Tunguz/Campbell
empirical priors.
- Industry profiles (`saas`, `api`, `ai-tools`, `enterprise-software`, `marketplace`) encode
default biases observed in BVP and ProfitWell cohort data.
- The "value-based requires measurable driver" gate comes directly from Ramanujam's "minivation"
failure mode.
- Freemium scoring penalties for high-ACV deals come from Poyar's documented PLG-to-enterprise
transition patterns.
FILE:references/van_westendorp_methodology.md
# Van Westendorp Price Sensitivity Meter — Methodology
Reference for `wtp_analyzer.py`. Covers the 4 questions, the 4 intersection points,
sample size discipline, segmentation requirements, and the most common misinterpretations.
---
## The four questions
Each respondent answers, for the product or feature described:
1. **Too cheap** — "At what price would you consider the product so inexpensive that you'd
doubt its quality and not buy it?"
2. **Bargain** — "At what price would you consider the product a bargain — a great buy for the
money?"
3. **Getting expensive** — "At what price would you start to feel the product is getting
expensive, but you'd still consider buying it?"
4. **Too expensive** — "At what price would you consider the product so expensive that you
would not consider buying it?"
For each respondent, the answers should obey:
`too_cheap ≤ bargain ≤ getting_expensive ≤ too_expensive`
Respondents who violate this ordering are typically screened out before analysis (the tool
flags them as warnings).
---
## The four intersection points
Build cumulative curves on a sorted price grid:
- **% too cheap (≥ price)** — decreasing in price (more respondents say "too cheap" at low prices)
- **% bargain (≥ price)** — decreasing
- **% getting expensive (≤ price)** — increasing
- **% too expensive (≤ price)** — increasing
Then find the four intersections:
| Point | Curves | Interpretation |
|---|---|---|
| **OPP** — Optimal Price Point | too cheap ↔ too expensive | Equal % reject as too cheap and too expensive. Theoretical sweet spot. |
| **IDP** — Indifference Price Point | bargain ↔ getting expensive | Median respondent's perceived "fair" price. |
| **PMC** — Point of Marginal Cheapness | too cheap ↔ getting expensive | Lower bound of acceptable range — below this, quality doubt dominates. |
| **PME** — Point of Marginal Expensiveness | bargain ↔ too expensive | Upper bound of acceptable range — above this, purchase rejection dominates. |
**Range of Acceptable Prices (RAP) = [PMC, PME].**
---
## Sample size discipline
- **N < 30:** Directional only. Tool emits a warning. Do not anchor decisions on these results.
- **N = 30-99:** Acceptable for hypothesis generation; expect noisy intersections.
- **N ≥ 100:** Preferred. ESOMAR conventions cite N=200-400 for stable PSM in B2C; B2B can
work with smaller but more carefully screened panels.
- **Segmented PSM:** Run separately for ICP vs non-ICP, and per buying-role segment. Aggregate
PSM averages across segments hide the structure you need.
---
## Common misinterpretations (the high-cost ones)
1. **"PSM gives THE price."** — No. PSM gives a **range**. The number inside the range is a
commercial decision involving competition, positioning, and margin targets.
2. **"OPP is the optimal price."** — OPP is named misleadingly. It's the point of *equal
resistance from both sides*, not a profit-maximizing price. The optimal price often sits
between OPP and PME if the market tolerates upside.
3. **"PSM works on non-customers."** — PSM measures *perceived* price thresholds. Run it on
the actual ICP. Random survey panels produce intersection points for an imaginary buyer.
4. **"PSM works for any product."** — Original method (van Westendorp, 1976) was built for
consumer non-durables. It works for SaaS, but breaks for products where the customer cannot
form a reference price (truly novel categories). Use Newton-Miller-Smith (NMS) refinement
in those cases — adds purchase-likelihood at each price.
5. **"Higher RAP upper bound = we can charge more."** — Only if your willingness-to-act
matches willingness-to-state. Always validate with a real purchase test (conjoint, A/B,
or sales-priced cohort) before anchoring at PME.
---
## Authoritative sources
1. **Peter van Westendorp — "NSS-Price Sensitivity Meter (PSM)" — 29th ESOMAR Congress
Proceedings, 1976.** The original paper. Establishes the four questions and intersection
method. Still the canonical reference 50 years later.
2. **Gabor & Granger (1966), Newton, Miller & Smith (NMS).** Extension that adds
purchase-likelihood at each price. Conjoint.ly and Sawtooth both implement NMS variants
for novel products without strong reference prices.
3. **Conjoint.ly — "Price Sensitivity Meter (Van Westendorp) Explained"**.
Practitioner-grade explanation including segmentation guidance and NMS comparison.
https://conjointly.com/guides/van-westendorp-price-sensitivity-analysis/
4. **Sawtooth Software — Lighthouse Studio documentation on PSM**.
Industry-standard tooling. Their guidance on respondent screening, monotonicity checks,
and segmentation is the operational standard most pricing consultancies use.
5. **ESOMAR — Code of Conduct + price-sensitivity research guidance**.
Sample-size conventions, respondent qualification, ethical pricing research. PSM is
referenced in their pricing research best-practice papers.
6. **Stan Lipovetsky (2006) — "Van Westendorp Price Sensitivity in statistical modeling,"
International Journal of Operational Research.** Critique and statistical refinement —
shows that classical PSM intersections are biased estimators under common response
distributions. Recommends bootstrap CIs and ordinal regression overlays.
7. **Decision Analyst — "Van Westendorp PSM Handbook"**.
Operational handbook including a worked example, screening criteria, and segmentation
templates. https://www.decisionanalyst.com/
8. **Madhavan Ramanujam — Monetizing Innovation (Wiley, 2016), Ch. 4.**
PSM as one of three WTP techniques (alongside direct WTP and conjoint). Ramanujam's
guidance: PSM for category baseline, conjoint for feature-level WTP, direct WTP for
confirmation.
## How this skill uses the methodology
- `wtp_analyzer.py` implements classical PSM intersections using linear interpolation on the
sorted price grid — the standard approach per van Westendorp (1976) and Sawtooth.
- Tool emits sample-size warnings at N<30 and N<100, per ESOMAR / Decision Analyst conventions.
- Tool checks per-respondent monotonicity (`tc ≤ bg ≤ ge ≤ te`) and reports inconsistent rows.
- Output explicitly frames PSM as a **range**, not a price, and recommends segmented re-runs —
per the documented misinterpretation patterns above.
- Tool does not implement NMS extension; for novel categories without reference prices, point
the user to conjoint or NMS-specific tooling.
FILE:scripts/packaging_designer.py
#!/usr/bin/env python3
"""packaging_designer.py — Good/Better/Best tier designer with anti-pattern detection.
Input: JSON with feature list (importance + cost-to-serve), customer segments, and
current pricing. Output: 3-tier packaging assignment with anti-pattern flags.
Deterministic logic — features are bucketed into tiers by importance × segment-fit.
No LLM calls.
Usage:
packaging_designer.py --input features.json --profile saas --output markdown
packaging_designer.py --sample
"""
from __future__ import annotations
import argparse
import json
import sys
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any
PROFILES = {
"saas": {"good_pct": 0.50, "better_pct": 0.30, "best_pct": 0.20, "price_ratio_good_to_better": 2.5, "price_ratio_better_to_best": 2.0},
"api": {"good_pct": 0.40, "better_pct": 0.30, "best_pct": 0.30, "price_ratio_good_to_better": 3.0, "price_ratio_better_to_best": 2.5},
"enterprise": {"good_pct": 0.30, "better_pct": 0.35, "best_pct": 0.35, "price_ratio_good_to_better": 2.0, "price_ratio_better_to_best": 2.5},
"prosumer": {"good_pct": 0.60, "better_pct": 0.25, "best_pct": 0.15, "price_ratio_good_to_better": 3.0, "price_ratio_better_to_best": 2.0},
}
@dataclass
class Feature:
name: str
importance: float # 0..1, how much customers value it
cost_to_serve: float # relative cost units
segment_fit: dict[str, float] = field(default_factory=dict) # segment → fit 0..1
@classmethod
def from_dict(cls, d: dict[str, Any]) -> "Feature":
return cls(
name=d["name"],
importance=float(d.get("importance", 0.5)),
cost_to_serve=float(d.get("cost_to_serve", 1.0)),
segment_fit=d.get("segment_fit", {}),
)
@dataclass
class Tier:
name: str
features: list[Feature] = field(default_factory=list)
price: float = 0.0
def assign_tiers(features: list[Feature], segments: list[str], profile: str) -> dict[str, Tier]:
"""Assign each feature to Good / Better / Best based on importance and segment fit.
Rule: high importance across all segments → Good (base).
mid importance OR segment-skewed to mid → Better.
low importance OR enterprise-skewed OR high cost-to-serve → Best.
"""
good = Tier("Good")
better = Tier("Better")
best = Tier("Best")
# Identify enterprise-leaning segments (last in declared order is convention)
enterprise_seg = segments[-1] if segments else None
for f in features:
avg_fit = sum(f.segment_fit.values()) / len(f.segment_fit) if f.segment_fit else 0.5
enterprise_fit = f.segment_fit.get(enterprise_seg, avg_fit) if enterprise_seg else avg_fit
# High importance + broad fit → Good
if f.importance >= 0.75 and avg_fit >= 0.6 and f.cost_to_serve <= 2.0:
good.features.append(f)
# Enterprise-skewed OR high cost → Best
elif enterprise_fit >= 0.7 and avg_fit < 0.6:
best.features.append(f)
elif f.cost_to_serve >= 3.0:
best.features.append(f)
elif f.importance <= 0.4:
best.features.append(f)
# Everything else → Better
else:
better.features.append(f)
return {"good": good, "better": better, "best": best}
def price_tiers(tiers: dict[str, Tier], current_pricing: dict[str, float], profile: str) -> None:
"""Anchor pricing to current_pricing if provided; else use profile ratios from a base of 100."""
cfg = PROFILES[profile]
if current_pricing.get("good"):
tiers["good"].price = float(current_pricing["good"])
else:
tiers["good"].price = 100.0
if current_pricing.get("better"):
tiers["better"].price = float(current_pricing["better"])
else:
tiers["better"].price = tiers["good"].price * cfg["price_ratio_good_to_better"]
if current_pricing.get("best"):
tiers["best"].price = float(current_pricing["best"])
else:
tiers["best"].price = tiers["better"].price * cfg["price_ratio_better_to_best"]
def detect_anti_patterns(tiers: dict[str, Tier]) -> list[str]:
"""Return list of human-readable anti-pattern flags."""
flags: list[str] = []
good, better, best = tiers["good"], tiers["better"], tiers["best"]
# 1. Empty tier
for t in (good, better, best):
if not t.features:
flags.append(f"Empty tier: '{t.name}' has no features — collapse or re-balance.")
# 2. Feature in all tiers (no differentiation)
good_names = {f.name for f in good.features}
better_names = {f.name for f in better.features}
best_names = {f.name for f in best.features}
all_three = good_names & better_names & best_names
if all_three:
flags.append(f"No differentiation: features appear in all 3 tiers — {sorted(all_three)}.")
# 3. Feature dump in Best (>2x the count of Better with <1.5x the price)
if better.features and best.features and better.price > 0 and best.price > 0:
feature_ratio = len(best.features) / max(1, len(better.features))
price_ratio = best.price / better.price
if feature_ratio > 2.0 and price_ratio < 1.5:
flags.append(
f"Feature dump in Best: {len(best.features)} features vs Better's {len(better.features)} "
f"({feature_ratio:.1f}x) for only {price_ratio:.1f}x the price — customers will buy Better and never upgrade."
)
# 4. Best tier > 2x Better price with < 1.5x value (proxy: feature count weighted by importance)
def value(t: Tier) -> float:
return sum(f.importance for f in t.features)
if better.price > 0 and best.price > 0 and value(better) > 0:
price_jump = best.price / better.price
value_jump = value(best) / value(better)
if price_jump > 2.0 and value_jump < 1.5:
flags.append(
f"Best tier price-to-value mismatch: {price_jump:.1f}x price for only {value_jump:.1f}x value — "
"Best becomes a decoy that no one upgrades to."
)
# 5. No clear upgrade trigger from Good → Better
if good.features and better.features:
good_imp = sum(f.importance for f in good.features) / len(good.features)
better_imp = sum(f.importance for f in better.features) / len(better.features)
if better_imp < good_imp - 0.1:
flags.append(
"No clear upgrade trigger Good → Better: Better-tier features have lower avg importance than Good. "
"Why would a Good customer ever upgrade?"
)
# 6. Bronze tier as loss leader (cost-to-serve > effective price share)
if good.features and good.price > 0:
good_cost = sum(f.cost_to_serve for f in good.features)
if good_cost > good.price * 0.8:
flags.append(
f"Good tier near loss-leader: cost-to-serve ({good_cost:.1f}) > 80% of price ({good.price:.2f}). "
"Either raise the price floor or strip a feature down to Better."
)
# 7. Best tier "Enterprise — call us" with no anchor
if best.price == 0 and best.features:
flags.append(
"Best/Enterprise tier has no published anchor price. 'Call us' without a starting number "
"loses prospects to competitors who publish ranges."
)
return flags
def render_markdown(tiers: dict[str, Tier], flags: list[str], profile: str, segments: list[str]) -> str:
lines: list[str] = []
lines.append("# Packaging Recommendation: Good / Better / Best")
lines.append("")
lines.append(f"**Profile:** `{profile}` • **Segments:** {', '.join(segments) if segments else 'unspecified'}")
lines.append("")
for key in ("good", "better", "best"):
t = tiers[key]
lines.append(f"## {t.name} — ,.2f")
if t.features:
for f in t.features:
lines.append(f"- **{f.name}** (importance={f.importance:.2f}, cost-to-serve={f.cost_to_serve:.1f})")
else:
lines.append("- *(no features assigned)*")
lines.append("")
if flags:
lines.append("## Anti-pattern flags")
for f in flags:
lines.append(f"- {f}")
else:
lines.append("## Anti-pattern flags")
lines.append("- None detected.")
lines.append("")
lines.append("## Notes")
lines.append("- Prices are a **starting frame**, not the final number. Validate with Van Westendorp PSM.")
lines.append("- Re-run after every meaningful feature addition; tier balance drifts as the product grows.")
return "\n".join(lines)
def sample_input() -> dict[str, Any]:
return {
"segments": ["SMB", "Mid-market", "Enterprise"],
"current_pricing": {"good": 49, "better": 149, "best": 499},
"features": [
{"name": "Core dashboard", "importance": 0.95, "cost_to_serve": 0.5, "segment_fit": {"SMB": 1.0, "Mid-market": 1.0, "Enterprise": 1.0}},
{"name": "Basic reporting", "importance": 0.85, "cost_to_serve": 0.8, "segment_fit": {"SMB": 0.9, "Mid-market": 0.9, "Enterprise": 0.8}},
{"name": "API access", "importance": 0.6, "cost_to_serve": 1.5, "segment_fit": {"SMB": 0.3, "Mid-market": 0.7, "Enterprise": 0.9}},
{"name": "Advanced analytics", "importance": 0.65, "cost_to_serve": 2.0, "segment_fit": {"SMB": 0.2, "Mid-market": 0.8, "Enterprise": 0.9}},
{"name": "Custom workflows", "importance": 0.55, "cost_to_serve": 2.5, "segment_fit": {"SMB": 0.1, "Mid-market": 0.5, "Enterprise": 0.9}},
{"name": "SSO / SAML", "importance": 0.4, "cost_to_serve": 1.0, "segment_fit": {"SMB": 0.05, "Mid-market": 0.4, "Enterprise": 1.0}},
{"name": "SLA + dedicated CSM", "importance": 0.3, "cost_to_serve": 5.0, "segment_fit": {"SMB": 0.0, "Mid-market": 0.2, "Enterprise": 1.0}},
{"name": "On-prem deployment", "importance": 0.2, "cost_to_serve": 4.0, "segment_fit": {"SMB": 0.0, "Mid-market": 0.1, "Enterprise": 0.9}},
{"name": "Audit logs", "importance": 0.5, "cost_to_serve": 0.8, "segment_fit": {"SMB": 0.1, "Mid-market": 0.5, "Enterprise": 0.95}},
],
}
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description=__doc__.splitlines()[0])
p.add_argument("--input", type=Path, help="Path to features JSON.")
p.add_argument("--profile", default="saas", choices=list(PROFILES.keys()), help="Industry profile.")
p.add_argument("--output", default="markdown", choices=["markdown", "json"], help="Output format.")
p.add_argument("--sample", action="store_true", help="Run with built-in sample data.")
args = p.parse_args(argv)
if args.sample:
data = sample_input()
elif args.input:
data = json.loads(args.input.read_text())
else:
p.error("Provide --input or --sample.")
return 2
segments = data.get("segments", [])
current_pricing = data.get("current_pricing", {})
features = [Feature.from_dict(f) for f in data.get("features", [])]
tiers = assign_tiers(features, segments, args.profile)
price_tiers(tiers, current_pricing, args.profile)
flags = detect_anti_patterns(tiers)
if args.output == "json":
out = {
"profile": args.profile,
"segments": segments,
"tiers": {
k: {
"name": t.name,
"price": t.price,
"features": [{"name": f.name, "importance": f.importance, "cost_to_serve": f.cost_to_serve} for f in t.features],
}
for k, t in tiers.items()
},
"anti_pattern_flags": flags,
}
print(json.dumps(out, indent=2))
else:
print(render_markdown(tiers, flags, args.profile, segments))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/pricing_model_picker.py
#!/usr/bin/env python3
"""pricing_model_picker.py — rank pricing models by fit-score for a given customer context.
Input: JSON describing customer context (industry, deal size, customer count, value drivers,
adoption curve, consumption pattern, competitor pricing models).
Output: ranked list of 5 pricing models (subscription seat-based, usage-based, value-based,
freemium, hybrid) with fit-score 0-100 and trade-offs.
Deterministic decision logic. No LLM calls. No third-party deps.
Usage:
pricing_model_picker.py --input brief.json --profile saas --output markdown
pricing_model_picker.py --sample
"""
from __future__ import annotations
import argparse
import json
import sys
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any
MODELS = [
"subscription_seat_based",
"usage_based",
"value_based",
"freemium",
"hybrid",
]
# Industry profile tuning — base biases (additive, capped at ±15)
PROFILES: dict[str, dict[str, int]] = {
"saas": {
"subscription_seat_based": 10,
"usage_based": 0,
"value_based": 0,
"freemium": 5,
"hybrid": 5,
},
"api": {
"subscription_seat_based": -10,
"usage_based": 15,
"value_based": 0,
"freemium": 5,
"hybrid": 5,
},
"ai-tools": {
"subscription_seat_based": -5,
"usage_based": 10,
"value_based": 5,
"freemium": 5,
"hybrid": 10,
},
"enterprise-software": {
"subscription_seat_based": 5,
"usage_based": -5,
"value_based": 10,
"freemium": -10,
"hybrid": 5,
},
"marketplace": {
"subscription_seat_based": -10,
"usage_based": 10,
"value_based": 10,
"freemium": 5,
"hybrid": 5,
},
}
@dataclass
class ModelScore:
model: str
score: int
rationale: list[str] = field(default_factory=list)
tradeoffs: list[str] = field(default_factory=list)
def clamp(n: int, lo: int = 0, hi: int = 100) -> int:
return max(lo, min(hi, n))
def score_models(ctx: dict[str, Any], profile: str) -> list[ModelScore]:
"""Deterministic per-model scoring. Each model starts at 50 and is adjusted by signals."""
cp = (ctx.get("consumption_pattern") or {})
deal_size = float(ctx.get("deal_size_avg") or 0)
customer_count = int(ctx.get("customer_count") or 0)
value_drivers = ctx.get("value_drivers") or []
adoption = (ctx.get("adoption_curve") or "").lower()
competitor_models = ctx.get("competitor_pricing_models") or []
seat_signal = float(cp.get("seat-based") or 0)
usage_signal = float(cp.get("usage-based") or 0)
value_signal = float(cp.get("value-based") or 0)
hybrid_signal = float(cp.get("hybrid") or 0)
scores = {m: ModelScore(model=m, score=50) for m in MODELS}
# --- Subscription seat-based ---
s = scores["subscription_seat_based"]
if seat_signal >= 0.6:
s.score += 20
s.rationale.append(f"Strong seat-based consumption signal ({seat_signal:.2f}).")
elif seat_signal >= 0.3:
s.score += 8
s.rationale.append(f"Moderate seat-based signal ({seat_signal:.2f}).")
if usage_signal > 0.5 and seat_signal < 0.4:
s.score -= 15
s.tradeoffs.append("Usage variance is high; seat licensing leaves money on the table.")
if deal_size > 0 and deal_size < 5000:
s.score += 5
s.rationale.append("SMB-friendly deal size — predictable seat math.")
if "subscription" in " ".join(competitor_models).lower():
s.score += 5
s.rationale.append("Competitors already train the market on subscription.")
s.tradeoffs.append("Predictable revenue, but customers feel friction when adding seats.")
# --- Usage-based ---
s = scores["usage_based"]
if usage_signal >= 0.6:
s.score += 25
s.rationale.append(f"Strong usage-variance signal ({usage_signal:.2f}) — power-law users.")
elif usage_signal >= 0.3:
s.score += 10
s.rationale.append(f"Moderate usage signal ({usage_signal:.2f}).")
if seat_signal > 0.6 and usage_signal < 0.3:
s.score -= 15
s.tradeoffs.append("Usage is flat per seat; usage-based adds billing complexity for no upside.")
if "api" in (ctx.get("industry") or "").lower() or profile == "api":
s.score += 8
s.rationale.append("API/infra products align naturally with usage metering.")
if "usage" in " ".join(competitor_models).lower() or "consumption" in " ".join(competitor_models).lower():
s.score += 5
s.rationale.append("Competitive usage pricing trains the market.")
s.tradeoffs.append("Aligned to value but introduces revenue unpredictability and bill-shock risk.")
# --- Value-based ---
s = scores["value_based"]
measurable = any(
kw in " ".join(value_drivers).lower()
for kw in ["revenue", "cost saved", "time saved", "conversion", "fraud prevented", "downtime"]
)
if value_signal >= 0.6 and measurable:
s.score += 25
s.rationale.append("Customer value is measurable AND signaled as primary.")
elif value_signal >= 0.4 and measurable:
s.score += 15
s.rationale.append("Value signal moderate, measurement plausible.")
elif value_signal >= 0.4 and not measurable:
s.score -= 10
s.tradeoffs.append("Value signal present but no measurable driver — collapses to bad usage pricing.")
if deal_size >= 50000:
s.score += 10
s.rationale.append("Enterprise deal size justifies bespoke value-pricing motion.")
if customer_count > 0 and customer_count < 50:
s.score += 5
s.rationale.append("Small customer count supports per-customer value calibration.")
if customer_count > 500:
s.score -= 10
s.tradeoffs.append("High customer count — value-based does not scale operationally.")
s.tradeoffs.append("Highest yield model but requires instrumented ROI proof per customer.")
# --- Freemium ---
s = scores["freemium"]
if adoption in ("viral", "bottom-up", "plg", "product-led"):
s.score += 20
s.rationale.append(f"Adoption curve '{adoption}' aligns with PLG/freemium funnel.")
if customer_count > 1000:
s.score += 10
s.rationale.append("Large addressable user base supports freemium economics.")
if deal_size > 25000:
s.score -= 15
s.tradeoffs.append("Enterprise ACV — freemium acquisition cost rarely amortizes.")
if adoption in ("top-down", "enterprise"):
s.score -= 15
s.tradeoffs.append("Top-down sale — freemium dilutes positioning without unlocking pipeline.")
s.tradeoffs.append("Powerful acquisition channel but free-tier cost-to-serve must be a small fraction of paid LTV.")
# --- Hybrid (platform + usage, or seat + overage) ---
s = scores["hybrid"]
if hybrid_signal >= 0.5:
s.score += 20
s.rationale.append(f"Hybrid signal explicit ({hybrid_signal:.2f}).")
if seat_signal >= 0.4 and usage_signal >= 0.4:
s.score += 15
s.rationale.append("Both seat AND usage drivers present — natural hybrid candidate.")
if len(value_drivers) >= 3:
s.score += 5
s.rationale.append("Multiple value drivers — single model leaves segments under-served.")
if deal_size < 1000:
s.score -= 10
s.tradeoffs.append("Small deal size — hybrid complexity is not worth the friction.")
s.tradeoffs.append("Captures more value across segments but increases pricing-page complexity and CS overhead.")
# Profile bias
bias = PROFILES.get(profile, {})
for m, b in bias.items():
if m in scores:
scores[m].score += b
if b != 0:
scores[m].rationale.append(f"Industry profile '{profile}' adjustment: {b:+d}.")
# Clamp
for s in scores.values():
s.score = clamp(s.score)
return sorted(scores.values(), key=lambda x: -x.score)
def render_markdown(ranked: list[ModelScore], ctx: dict[str, Any], profile: str) -> str:
lines: list[str] = []
lines.append("# Pricing Model Recommendation")
lines.append("")
lines.append(f"**Profile:** `{profile}` • **Industry:** {ctx.get('industry', 'unspecified')}")
lines.append(f"**Deal size avg:** {ctx.get('deal_size_avg', 'n/a')} • **Customers:** {ctx.get('customer_count', 'n/a')}")
lines.append("")
lines.append("> This skill recommends a **model and trade-offs**, not a final price. The human owns the decision.")
lines.append("")
lines.append("## Ranked models")
lines.append("")
for i, s in enumerate(ranked, 1):
marker = " *(top recommendation)*" if i == 1 else ""
lines.append(f"### {i}. {s.model.replace('_', ' ').title()} — fit-score **{s.score}/100**{marker}")
if s.rationale:
lines.append("**Why it fits:**")
for r in s.rationale:
lines.append(f"- {r}")
if s.tradeoffs:
lines.append("**Trade-offs:**")
for t in s.tradeoffs:
lines.append(f"- {t}")
lines.append("")
lines.append("## Next steps")
lines.append("1. Validate WTP for the top model with `wtp_analyzer.py` (≥ 30 respondents).")
lines.append("2. Design tiers with `packaging_designer.py`.")
lines.append("3. Pressure-test in pricing committee — this output is one input.")
return "\n".join(lines)
def sample_context() -> dict[str, Any]:
return {
"industry": "B2B SaaS — sales intelligence",
"deal_size_avg": 18000,
"customer_count": 220,
"value_drivers": ["revenue lift from better leads", "time saved in research", "conversion uplift"],
"adoption_curve": "bottom-up",
"consumption_pattern": {
"seat-based": 0.45,
"usage-based": 0.55,
"value-based": 0.40,
"hybrid": 0.50,
},
"competitor_pricing_models": ["subscription seat-based", "hybrid seat + usage overage"],
}
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description=__doc__.splitlines()[0])
p.add_argument("--input", type=Path, help="Path to customer-context JSON.")
p.add_argument(
"--profile",
default="saas",
choices=list(PROFILES.keys()),
help="Industry profile for default tuning.",
)
p.add_argument("--output", default="markdown", choices=["markdown", "json"], help="Output format.")
p.add_argument("--sample", action="store_true", help="Run with built-in sample context.")
args = p.parse_args(argv)
if args.sample:
ctx = sample_context()
elif args.input:
ctx = json.loads(args.input.read_text())
else:
p.error("Provide --input or --sample.")
return 2
ranked = score_models(ctx, args.profile)
if args.output == "json":
out = {
"profile": args.profile,
"context": ctx,
"ranked": [
{"model": s.model, "score": s.score, "rationale": s.rationale, "tradeoffs": s.tradeoffs}
for s in ranked
],
}
print(json.dumps(out, indent=2))
else:
print(render_markdown(ranked, ctx, args.profile))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/wtp_analyzer.py
#!/usr/bin/env python3
"""wtp_analyzer.py — Van Westendorp Price Sensitivity Meter (PSM).
Implements the classical PSM analysis (van Westendorp, 1976). Each respondent answers
4 questions:
1. Too cheap — price at which you'd doubt the quality
2. Bargain — price that feels like a great deal
3. Getting expensive — price where you'd start to hesitate
4. Too expensive — price at which you would never buy
Computes the 4 intersection points:
- OPP (Optimal Price Point): intersection of "too cheap" and "too expensive"
- IDP (Indifference Price Point): intersection of "bargain" and "getting expensive"
- PMC (Point of Marginal Cheapness): intersection of "too cheap" and "getting expensive"
- PME (Point of Marginal Expensiveness): intersection of "bargain" and "too expensive"
Range of Acceptable Prices (RAP) = [PMC, PME].
Output: markdown or JSON. Stdlib only.
Usage:
wtp_analyzer.py --input survey.json --output markdown
wtp_analyzer.py --sample
"""
from __future__ import annotations
import argparse
import json
import math
import random
import statistics
import sys
from dataclasses import dataclass
from pathlib import Path
from typing import Any
@dataclass
class Curves:
prices: list[float]
too_cheap: list[float] # P(too cheap >= price) — decreasing in price
bargain: list[float] # P(bargain >= price) — decreasing in price
getting_expensive: list[float] # P(getting expensive <= price) — increasing
too_expensive: list[float] # P(too expensive <= price) — increasing
def build_price_grid(respondents: list[dict[str, float]]) -> list[float]:
"""Build a sorted unique-price grid from all responses."""
prices: set[float] = set()
for r in respondents:
for k in ("too_cheap", "bargain", "getting_expensive", "too_expensive"):
v = r.get(k)
if v is not None:
prices.add(float(v))
grid = sorted(prices)
if not grid:
return []
# Densify with intermediate steps to make intersection detection stable.
densified: list[float] = []
for i, p in enumerate(grid):
densified.append(p)
if i + 1 < len(grid):
nxt = grid[i + 1]
mid = (p + nxt) / 2.0
if mid not in prices:
densified.append(mid)
return sorted(set(densified))
def cumulative_curves(respondents: list[dict[str, float]], grid: list[float]) -> Curves:
n = len(respondents)
too_cheap: list[float] = []
bargain: list[float] = []
getting_expensive: list[float] = []
too_expensive: list[float] = []
for p in grid:
tc = sum(1 for r in respondents if (r.get("too_cheap") is not None) and float(r["too_cheap"]) >= p)
bg = sum(1 for r in respondents if (r.get("bargain") is not None) and float(r["bargain"]) >= p)
ge = sum(1 for r in respondents if (r.get("getting_expensive") is not None) and float(r["getting_expensive"]) <= p)
te = sum(1 for r in respondents if (r.get("too_expensive") is not None) and float(r["too_expensive"]) <= p)
too_cheap.append(tc / n)
bargain.append(bg / n)
getting_expensive.append(ge / n)
too_expensive.append(te / n)
return Curves(grid, too_cheap, bargain, getting_expensive, too_expensive)
def find_intersection(prices: list[float], a: list[float], b: list[float]) -> float | None:
"""Find first price where curve a crosses curve b (linear interp between grid points)."""
if len(prices) < 2:
return None
prev_diff = a[0] - b[0]
for i in range(1, len(prices)):
diff = a[i] - b[i]
if prev_diff == 0:
return prices[i - 1]
if (prev_diff < 0 < diff) or (prev_diff > 0 > diff):
# Linear interpolation
p0, p1 = prices[i - 1], prices[i]
t = prev_diff / (prev_diff - diff)
return p0 + t * (p1 - p0)
prev_diff = diff
return None
@dataclass
class PSMResult:
n: int
opp: float | None
idp: float | None
pmc: float | None
pme: float | None
rap_low: float | None
rap_high: float | None
warnings: list[str]
def analyze(respondents: list[dict[str, float]]) -> PSMResult:
warnings: list[str] = []
n = len(respondents)
if n < 30:
warnings.append(
f"Sample size N={n} is below 30. PSM results are directional only; "
"treat the range as a hypothesis, not a recommendation. Aim for N≥100."
)
elif n < 100:
warnings.append(f"Sample size N={n} is acceptable but below the preferred N≥100 threshold.")
# Sanity-check monotonicity (too_cheap < bargain < getting_expensive < too_expensive per respondent)
inconsistent = 0
for r in respondents:
try:
tc = float(r["too_cheap"])
bg = float(r["bargain"])
ge = float(r["getting_expensive"])
te = float(r["too_expensive"])
except (KeyError, TypeError, ValueError):
inconsistent += 1
continue
if not (tc <= bg <= ge <= te):
inconsistent += 1
if inconsistent:
warnings.append(
f"{inconsistent} of {n} respondents have inconsistent price ordering "
"(expected too_cheap ≤ bargain ≤ getting_expensive ≤ too_expensive). "
"Consider screening these before reporting."
)
grid = build_price_grid(respondents)
if not grid:
return PSMResult(n=n, opp=None, idp=None, pmc=None, pme=None, rap_low=None, rap_high=None, warnings=warnings)
c = cumulative_curves(respondents, grid)
opp = find_intersection(c.prices, c.too_cheap, c.too_expensive)
idp = find_intersection(c.prices, c.bargain, c.getting_expensive)
pmc = find_intersection(c.prices, c.too_cheap, c.getting_expensive)
pme = find_intersection(c.prices, c.bargain, c.too_expensive)
return PSMResult(n=n, opp=opp, idp=idp, pmc=pmc, pme=pme, rap_low=pmc, rap_high=pme, warnings=warnings)
def _fmt(v: float | None) -> str:
return f"{v:,.2f}" if v is not None else "n/a"
def render_markdown(res: PSMResult) -> str:
lines: list[str] = []
lines.append("# Van Westendorp PSM Analysis")
lines.append("")
lines.append(f"**Respondents:** N = {res.n}")
lines.append("")
if res.warnings:
lines.append("> **Warnings:**")
for w in res.warnings:
lines.append(f"> - {w}")
lines.append("")
lines.append("## Four intersection points")
lines.append("")
lines.append("| Point | Definition | Value |")
lines.append("|---|---|---|")
lines.append(f"| **OPP** — Optimal Price Point | too cheap ↔ too expensive | {_fmt(res.opp)} |")
lines.append(f"| **IDP** — Indifference Price Point | bargain ↔ getting expensive | {_fmt(res.idp)} |")
lines.append(f"| **PMC** — Point of Marginal Cheapness | too cheap ↔ getting expensive | {_fmt(res.pmc)} |")
lines.append(f"| **PME** — Point of Marginal Expensiveness | bargain ↔ too expensive | {_fmt(res.pme)} |")
lines.append("")
lines.append("## Range of Acceptable Prices (RAP)")
lines.append("")
if res.rap_low is not None and res.rap_high is not None:
lines.append(f"**RAP = [{_fmt(res.rap_low)}, {_fmt(res.rap_high)}]**")
lines.append("")
lines.append("Prices outside this range are likely to be rejected as either too cheap (quality doubt) or too expensive (no purchase).")
else:
lines.append("RAP could not be computed — check input data and sample size.")
lines.append("")
lines.append("## Interpretation guidance")
lines.append("")
lines.append("- PSM gives a **range**, not the price. Final price is a commercial decision.")
lines.append("- OPP is a theoretical mid-point — the price at which equal % of respondents reject as too cheap and too expensive.")
lines.append("- IDP is the median respondent's perceived 'fair' price.")
lines.append("- Re-run with segmented samples (ICP vs non-ICP) — overall PSM averages across segments hide structure.")
lines.append("- Validate the upper end with willingness-to-pay experiments in market before anchoring at PME.")
return "\n".join(lines)
def synthetic_sample(n: int = 50, seed: int = 17) -> list[dict[str, float]]:
"""Generate N synthetic respondents with realistic price ordering and segmentation noise."""
rng = random.Random(seed)
respondents: list[dict[str, float]] = []
for _ in range(n):
anchor = rng.gauss(80, 20) # respondent's reference price
anchor = max(20.0, anchor)
tc = max(5.0, anchor * rng.uniform(0.3, 0.5))
bg = anchor * rng.uniform(0.6, 0.85)
ge = anchor * rng.uniform(0.95, 1.15)
te = anchor * rng.uniform(1.3, 1.8)
respondents.append({
"too_cheap": round(tc, 2),
"bargain": round(bg, 2),
"getting_expensive": round(ge, 2),
"too_expensive": round(te, 2),
})
return respondents
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description=__doc__.splitlines()[0])
p.add_argument("--input", type=Path, help="Path to survey JSON: {respondents: [{too_cheap, bargain, getting_expensive, too_expensive}]}.")
p.add_argument("--output", default="markdown", choices=["markdown", "json"], help="Output format.")
p.add_argument("--sample", action="store_true", help="Run with synthetic 50-respondent sample.")
args = p.parse_args(argv)
if args.sample:
respondents = synthetic_sample(50)
elif args.input:
data = json.loads(args.input.read_text())
respondents = data.get("respondents", data) if isinstance(data, dict) else data
else:
p.error("Provide --input or --sample.")
return 2
if not isinstance(respondents, list) or not respondents:
print("ERROR: respondents must be a non-empty list.", file=sys.stderr)
return 1
res = analyze(respondents)
if args.output == "json":
out = {
"n": res.n,
"opp": res.opp,
"idp": res.idp,
"pmc": res.pmc,
"pme": res.pme,
"rap": [res.rap_low, res.rap_high],
"warnings": res.warnings,
}
print(json.dumps(out, indent=2))
else:
print(render_markdown(res))
return 0
if __name__ == "__main__":
sys.exit(main())
Mô tả quy trình nghiệp vụ end-to-end theo ký hiệu kiểu BPMN, đo thời gian chu kỳ theo từng bước và tìm nơi công việc bị chậm.
---
name: process-mapper
description: Use when a BizOps lead, COO, or process-improvement owner needs to document an end-to-end business process (procurement, employee onboarding, incident handoff, customer-onboarding, claims adjudication) in BPMN-style notation, measure cycle times by stage, surface where work spends most of its time waiting vs. being worked, and quantify the gap between processing time and total elapsed time. Pairs Lean / Six Sigma / Theory-of-Constraints canon with deterministic stdlib-only Python tools to produce a process map, a ranked bottleneck list (with severity + root-cause hypothesis), and a cycle-time analysis (P50, P90, value-add ratio, Little's-Law throughput). Distinct from sales-pipeline, system-reliability (SLO), and strategic-OKR work — this is tactical process documentation for internal operations.
version: 2.8.0
author: claude-code-skills
license: MIT
tags: [bizops, process, bpmn, bottleneck, cycle-time, lean, six-sigma, value-stream]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# process-mapper
BPMN-style business process documentation, bottleneck detection, and cycle-time analysis for internal-operations leaders.
## Purpose
Internal-operations work suffers from three recurring failure modes:
1. **Implicit process** — the steps exist only in tribal knowledge, so handoffs drop and onboarding takes weeks.
2. **Invisible waiting** — most of the elapsed time on any business process is queue / wait / approval time, not actual work; teams optimize the wrong stage.
3. **Local optimization** — Goldratt's Theory of Constraints is ignored; resources are added to non-constraint stages, gaining nothing.
This skill produces a documented process map, identifies where work waits, and points the constraint out by name with deterministic logic — not LLM intuition.
## When to use
- Documenting a new business process (procurement intake, vendor onboarding, employee onboarding, incident handoff, expense reimbursement, customer onboarding, claims adjudication).
- An existing process is "too slow" but nobody can name the bottleneck.
- Cycle time is being measured but value-add ratio is not — so the team can't tell whether the process is healthy or waste-heavy.
- Cross-functional handoffs are dropping work and root cause is unclear.
## Workflow
Five-step deterministic flow:
1. **Intake.** Capture the process as a JSON file with one entry per stage: `name`, `owner`, `type` (`value-add` | `wait` | `rework`), `duration_minutes_p50`, `duration_minutes_p90`. Use `assets/process_template.md` and its JSON skeleton.
2. **Map stages.** Run `process_documenter.py` to produce an ASCII swim-lane diagram + a normalized JSON artifact. The swim-lane separates lanes by owner so cross-functional handoffs become visible.
3. **Measure cycle time.** Run `cycle_time_analyzer.py` to compute total P50, total P90, value-add ratio (VA%), and a Little's-Law throughput estimate. Verdict: VA% > 25% = HEALTHY, 10–25% = TYPICAL, < 10% = WASTE-HEAVY.
4. **Detect bottlenecks.** Run `bottleneck_detector.py` with the appropriate `--profile` (saas / services / manufacturing / healthcare). Output is a ranked list with severity (CRITICAL / HIGH / MEDIUM), root-cause hypothesis, and one recommended action per finding.
5. **Recommend.** Pair the bottleneck list with the cycle-time verdict; recommend a single constraint-focused intervention per Goldratt's "subordinate everything to the constraint" rule. Don't recommend optimization of a non-constraint stage.
## Scripts
**`scripts/process_documenter.py`** — Reads a process JSON, validates it, and emits a text-based BPMN-style swim-lane diagram in Markdown (lanes by owner, stages annotated with type + duration). Also outputs a normalized JSON artifact for downstream tools. Stdlib only. `--sample` prints a 6-stage procurement-intake example.
**`scripts/bottleneck_detector.py`** — Applies three deterministic detection rules: (a) stage P50 > 2× mean of value-add stages, (b) wait-state % > 40% of total cycle, (c) rework % > 15%. Thresholds adjust by `--profile` because SaaS, services, manufacturing, and healthcare have different "normal" wait ratios. Output is a ranked list with severity, hypothesis, action.
**`scripts/cycle_time_analyzer.py`** — Computes total P50 and P90 cycle time, value-add ratio (VA%), wait %, rework %, and a Little's-Law throughput estimate (WIP / cycle time). Per Lean canon: VA% > 25% = HEALTHY, 10–25% = TYPICAL (most non-manufacturing processes land here), < 10% = WASTE-HEAVY.
## References
- `references/lean_six_sigma_canon.md` — TIMWOOD wastes, value-stream mapping, Theory of Constraints, Kanban WIP, Little's Law. Cites Womack & Jones, Rother & Shook, Goldratt, Ohno, Liker, Pyzdek, Anderson.
- `references/bpmn_essentials.md` — Pools, lanes, gateways, events, message flows, common notation mistakes. Cites the OMG BPMN 2.0 spec, Silver, Allweyer, Freund/Rücker, OASIS, ISO/IEC 19510:2013.
- `references/bottleneck_anti_patterns.md` — Seven specific anti-patterns drawn from Goldratt, Kim et al., Spear, DORA, Deming, and process-mining research.
## Assumptions
1. The user can provide stage-level cycle-time data (even rough P50 / P90 estimates). If they cannot, the first step is to instrument the process — not to map it.
2. "Process" here means a repeatable business workflow with discrete stages, not a one-off project.
3. The user has authority to act on bottlenecks (or can route findings to someone who does). Without that, the output is academic.
4. Stage `type` is honest: a "value-add" stage labeled as such by the user really does change the work product from the customer's perspective. Mis-labelling waiting as value-add is the most common data-quality failure.
## Anti-patterns
- **Mapping every process at once.** Pick one. Goldratt: the constraint is a single point.
- **Optimizing the non-constraint.** If stage 4 is the bottleneck, speeding up stage 2 just builds inventory in front of stage 4. Subordinate everything to the constraint.
- **Mistaking total cycle time for processing time.** They are almost never the same; VA% reveals the gap.
- **Adding people to a wait-bound process.** Wait time is not solved by more headcount; it's solved by removing the handoff or batch.
- **Treating rework as a separate problem.** Rework loops belong in the process map. Hiding them understates true cycle time.
## Distinct from
- **business-growth skills** — external sales motion, lead-funnel conversion, customer-success retention. Process-mapper is *internal* operations.
- **engineering/slo-architect** — system-reliability SLOs / error budgets / burn-rate alerts. Process-mapper is *business-process* cycle time, not system uptime.
- **c-level-advisor (COO / CEO)** — strategic prioritization of which processes to fix. Process-mapper is the tactical instrument used after that prioritization decision.
- **project-management skills** — Jira / Confluence ticket workflow tooling. Process-mapper is process *design*, not ticket *tracking*.
## Forcing-question library (Matt Pocock grill discipline)
Before invoking the tools, the orchestrator (or `/cs:grill-bizops`) walks the user through these questions **one at a time, with a recommended answer + canon citation**. Never bundled.
1. **"Do you have measured cycle times for the top-3 longest stages, or only estimates?"**
Recommended: insist on measured data.
Canon: Goldratt 1984 (*The Goal*) — optimizing estimated bottlenecks reliably attacks the wrong constraint.
2. **"Are you mapping the *current* process (as-is) or the *intended* process (to-be)?"**
Recommended: map as-is first. To-be after bottleneck is identified.
Canon: Rother & Shook 1999 (*Learning to See*) — value-stream mapping starts with the current state, always.
3. **"Where do handoffs occur between teams, and how long does each handoff wait?"**
Recommended: log every handoff with median wait time.
Canon: Reinertsen 2009 (*Principles of Product Development Flow*) — wait time at handoffs is the largest invisible cost.
4. **"What's your batch size at each stage?"**
Recommended: drive batch size toward 1 wherever possible.
Canon: Anderson 2010 (*Kanban*) — batch size correlates 1:1 with cycle time variance.
5. **"What's the rework rate per stage?"**
Recommended: surface it explicitly; rework loops belong in the map.
Canon: Pyzdek (*Six Sigma Handbook*) — hidden rework drives 30-50% of total cycle time in service processes.
Walk depth-first. Don't open question 4 before 1-3 are answered. After all 5 are locked, invoke `process_documenter.py` → `bottleneck_detector.py` → `cycle_time_analyzer.py` in sequence.
FILE:assets/process_template.md
# Process Template
Use this template to document a business process before running it through
the process-mapper tools. Fill in the stage table first, then translate it
into the JSON skeleton at the bottom of this file. Feed that JSON into the
three CLI tools:
```
python3 scripts/process_documenter.py --input my-process.json
python3 scripts/bottleneck_detector.py --input my-process.json --profile saas
python3 scripts/cycle_time_analyzer.py --input my-process.json --profile saas
```
---
## Process metadata
- **Process name:** _(e.g., Procurement Intake, Employee Onboarding, Incident Handoff)_
- **Owner role:** _(who is accountable for the end-to-end process)_
- **Frequency:** _(how often this process runs — daily, weekly, on-demand)_
- **Trigger event:** _(what starts the process)_
- **End state:** _(what marks the process complete)_
- **WIP at any time:** _(how many items are typically in process at once; needed for Little's-Law throughput)_
---
## Stage table
Six rows to start. Add or remove as needed. **Honesty about stage `type` is
the single most important data-quality choice.** If a stage is queue / wait,
mark it `wait`. If it changes the work product from the customer's
perspective, mark it `value-add`. If it exists to fix an upstream defect,
mark it `rework`.
| # | Stage name | Owner (role) | Type | P50 (min) | P90 (min) | Notes |
|---|------------|--------------|------|-----------|-----------|-------|
| 1 | _e.g., Submit request_ | Requestor | value-add | 15 | 30 | |
| 2 | _e.g., Wait for manager approval queue_ | Manager | wait | 480 | 1440 | Typically batched |
| 3 | _e.g., Manager approves_ | Manager | value-add | 10 | 25 | |
| 4 | _e.g., Wait for finance review_ | Finance | wait | 720 | 2880 | |
| 5 | _e.g., Finance validates budget code_ | Finance | value-add | 20 | 60 | |
| 6 | _e.g., Rework — missing vendor W-9_ | Requestor | rework | 120 | 360 | Frequent escape |
**Type definitions (Lean canon):**
- `value-add` — the stage changes the work product in a way the end customer
would willingly pay for. Most stages are NOT value-add.
- `wait` — work is queued, idle, or waiting for someone. Wait stages are the
largest source of cycle-time bloat in most office processes.
- `rework` — the stage exists to fix a defect introduced upstream. Six-Sigma
canon: rework is always an upstream-quality problem.
---
## JSON skeleton
Copy this into `my-process.json`, edit the values to match your stage table,
and pass it to the CLI tools.
```json
{
"process_name": "Replace with your process name",
"wip": 12,
"stages": [
{
"name": "Stage 1 name",
"owner": "Owning role",
"type": "value-add",
"duration_minutes_p50": 15,
"duration_minutes_p90": 30
},
{
"name": "Stage 2 name",
"owner": "Owning role",
"type": "wait",
"duration_minutes_p50": 480,
"duration_minutes_p90": 1440
},
{
"name": "Stage 3 name",
"owner": "Owning role",
"type": "value-add",
"duration_minutes_p50": 10,
"duration_minutes_p90": 25
},
{
"name": "Stage 4 name",
"owner": "Owning role",
"type": "wait",
"duration_minutes_p50": 720,
"duration_minutes_p90": 2880
},
{
"name": "Stage 5 name",
"owner": "Owning role",
"type": "value-add",
"duration_minutes_p50": 20,
"duration_minutes_p90": 60
},
{
"name": "Stage 6 name",
"owner": "Owning role",
"type": "rework",
"duration_minutes_p50": 120,
"duration_minutes_p90": 360
}
]
}
```
---
## Tips
- **Use real data when you can.** Pull stage durations from your ticket system
(Jira, ServiceNow, Zendesk). Estimated durations are fine for a first pass
but should be replaced before any change decision is made.
- **One process at a time.** Goldratt: every system has exactly one binding
constraint. Mapping ten processes simultaneously dilutes attention away
from the one that's actually limiting throughput.
- **Profile choice matters.** Pass `--profile manufacturing` for physical-goods
flows, `--profile services` for human-delivered services with longer
acceptable wait times, `--profile healthcare` for clinical or regulated
human-in-the-loop flows, `--profile saas` for everything else.
FILE:references/bottleneck_anti_patterns.md
# Bottleneck Anti-Patterns
Seven plus specific anti-patterns that recur in business-process improvement
work. Each is sourced to primary literature, and each has a corresponding
detection or recommendation in the skill's tools.
## Sources
1. **Goldratt, E. M. (1984). _The Goal._** North River Press. — Theory of Constraints.
2. **Kim, G., Behr, K. & Spafford, G. (2013). _The Phoenix Project: A Novel About IT, DevOps, and Helping Your Business Win._** IT Revolution Press. — TOC applied to IT operations.
3. **Spear, S. J. (2009). _The High-Velocity Edge._** McGraw-Hill. — Toyota-derived discipline for complex operations; explicit treatment of why local optimization fails.
4. **Forsgren, N., Humble, J. & Kim, G. (2018). _Accelerate: The Science of Lean Software and DevOps._** IT Revolution. — DORA research; empirical link between flow metrics and outcomes.
5. **Deming, W. E. (1986). _Out of the Crisis._** MIT Press. — System-of-profound-knowledge framework; root-cause discipline.
6. **van der Aalst, W. M. P. (2016). _Process Mining: Data Science in Action,_ 2nd ed.** Springer. — Empirical methodology for discovering actual process behavior vs. documented behavior.
7. **Reinertsen, D. G. (2009). _The Principles of Product Development Flow._** Celeritas Publishing. — Queueing theory and cost of delay.
8. **Forrester Research. (Multiple years.) _Process Mining: Vendor and Market Analyses._** — Industry research on process-mining adoption and the gap between modeled and actual process.
---
## AP-1. Optimizing the non-constraint
**Source:** Goldratt (1984), Kim et al. (2013).
A team identifies that stage 2 of a process is "slow" (relative to other
non-constraint stages) and optimizes it. The actual constraint is stage 4.
Result: throughput is unchanged; inventory grows in front of stage 4.
**Detection:** Compare every stage's P50 to the value-add mean (Rule R1) but
weight the recommendation by impact on total cycle. The skill's
`bottleneck_detector.py` ranks by impact_minutes_p50 specifically to direct
attention to the binding constraint.
**Counter-pattern:** Always solve the longest wait or longest stage first;
ignore "quick wins" elsewhere until the constraint moves.
---
## AP-2. Adding resources before identifying the constraint
**Source:** Goldratt (1984), Reinertsen (2009).
Symptom: "We need to hire more procurement analysts." Reality: the analysts
are not the constraint; manager approval queues are. Adding analysts increases
WIP, lengthens cycle time (per Little's Law), and makes the queue worse.
**Detection:** Rule R2 (wait-share > 40%) catches the case where the wait —
not capacity — dominates.
**Counter-pattern:** First check whether wait time exceeds value-add time. If
it does, no amount of new staffing will help. Remove the handoff, parallelize
the approval, or apply WIP limits.
---
## AP-3. Mistaking wait time for processing time
**Source:** Rother & Shook (1999), Deming (1986).
A team reports that "manager approval takes two days." On inspection, the
manager spends 10 minutes reviewing each request; the rest is queue time.
Process time is 10 minutes; lead time is two days. Treating them as the same
hides the real problem.
**Detection:** The skill's stage `type` field separates `value-add` from
`wait`. The value-add ratio (VA%) in `cycle_time_analyzer.py` quantifies the
gap.
**Counter-pattern:** Force stages to declare their type honestly. Any stage
where the worker is not actively engaged is a wait stage, regardless of who
"owns" it.
---
## AP-4. Inspection-as-quality
**Source:** Pyzdek (Six Sigma Handbook), Deming (1986), Spear (2009).
Defects keep escaping, so the team adds a final QA review. The defects don't
go down (the upstream stages haven't changed) — but cycle time goes up because
of the new stage. Worse, the QA reviewer is now blamed for misses.
**Detection:** Rule R3 (rework share > 15%) with the hypothesis "defects
escape upstream stages."
**Counter-pattern:** Find the earliest stage that could detect the defect; add
the check there (poka-yoke). Stop the line on detection; don't queue defects
for downstream rework.
---
## AP-5. Optimizing the documented process, not the actual one
**Source:** van der Aalst (2016), Forrester process-mining reports.
The team documents the "official" process and optimizes it. Process-mining
tools then reveal that 60% of cases skip stages, loop back, or take undocumented
routes. The optimization had no effect because it targeted a fiction.
**Detection:** The skill cannot detect this from the input JSON alone — it
relies on the user to report actual stage durations from real cases, not
target durations. The "Assumptions" section in SKILL.md surfaces this
explicitly.
**Counter-pattern:** Use ticket-system data, time-stamps, or event logs to
ground stage durations in actual cases. If the data isn't available, the
first step is instrumentation, not mapping.
---
## AP-6. Batched approvals as the default
**Source:** Reinertsen (2009), Anderson (Kanban, 2010).
Approvers batch requests: "I'll review everyone's POs on Friday afternoon."
This adds half the batch interval (typically 3–4 days) to the average wait
time of every request, with no quality benefit.
**Detection:** Wait stages with P50 durations measured in days (hundreds of
minutes) are almost always batched. The skill flags them via R1 and R2.
**Counter-pattern:** Move to continuous-flow approval. If continuous is
infeasible (e.g., a committee that meets weekly), at least shrink the batch
interval or move approval to a lower level where it can run continuously.
---
## AP-7. Local efficiency metrics
**Source:** Goldratt (1984), Deming (1986), Spear (2009).
Each stage is measured on its own efficiency (e.g., "manager handles 95% of
requests within SLA"). The system as a whole is not measured. Each role
optimizes locally, pushing work as fast as possible to the next queue —
which is exactly where it stalls.
**Detection:** The skill's verdict is always at the **process** level (VA%,
total cycle time), never at the stage level. The `bottleneck_detector.py`
recommendation text explicitly invokes Goldratt's "subordinate everything to
the constraint."
**Counter-pattern:** Measure throughput and total cycle time at the process
level. Stage-level metrics are diagnostic, not goal-setting.
---
## AP-8. Skipping the value-stream map and going straight to automation
**Source:** Kim et al. (2013), Forrester process-mining research.
A team buys an RPA / workflow automation tool, then automates the existing
broken process. Result: the bad process now runs faster, with the same wait
queues and same rework rate. Goldratt's term for this is "automating the
mess."
**Detection:** Outside the skill's automated detection; surfaced in
SKILL.md's "Anti-patterns" list.
**Counter-pattern:** Map the value stream first. Eliminate wait and rework
stages. Then — and only then — consider automating what remains.
---
## AP-9. Treating cycle time as fixed
**Source:** Forsgren, Humble & Kim (2018, _Accelerate_).
A team reports cycle time as a single number ("it takes 5 days"). Real cycle
times are distributions, often log-normal, with heavy P90 / P99 tails. A 5-day
P50 with a 30-day P90 is a wildly different process than a 5-day P50 with a
6-day P90; the first is unpredictable, the second is reliable.
**Detection:** The skill captures both P50 and P90 per stage and reports
both totals. A large P90 / P50 ratio in `cycle_time_analyzer.py` is a flag
for high variability even when total cycle time looks acceptable.
**Counter-pattern:** Always quote P50 and P90 (or P50 and P95). DORA's
_Accelerate_ research finds that lead-time **variability** correlates with
business outcomes as strongly as median lead time.
FILE:references/bpmn_essentials.md
# BPMN Essentials for Business-Process Documentation
A practical reference on BPMN (Business Process Model and Notation) for
process-mapper users. The skill emits text-based swim-lane diagrams that
approximate the BPMN structure without requiring users to install Visio,
Lucidchart, or Camunda. This file explains the canon those diagrams reflect.
## Sources
1. **Object Management Group. (2011). _Business Process Model and Notation (BPMN), Version 2.0._** OMG Document Number formal/2011-01-03. — The normative specification.
2. **Silver, B. (2011). _BPMN Method and Style,_ 2nd ed.** Cody-Cassidy Press. — The canonical practitioner book; defines the "method and style" rules now widely treated as informal BPMN convention.
3. **Allweyer, T. (2010). _BPMN 2.0: Introduction to the Standard for Business Process Modeling._** Books on Demand. — Approachable academic introduction.
4. **Freund, J. & Rücker, B. (2019). _Real-Life BPMN,_ 4th ed.** CreateSpace. — Practical patterns from the Camunda team.
5. **OASIS. (2010). _Web Services Business Process Execution Language (WS-BPEL), Version 2.0._** — Related execution standard; clarifies the interplay between BPMN modeling and BPEL execution.
6. **ISO/IEC 19510:2013. _Information technology — Object Management Group Business Process Model and Notation._** — The international-standard version of OMG BPMN 2.0.
7. **Recker, J. (2010). "Opportunities and constraints: the current struggle with BPMN." _Business Process Management Journal,_ 16(1), 181–201.** — Peer-reviewed analysis of BPMN adoption pain points; sources the "common notation mistakes" list below.
8. **Dumas, M., La Rosa, M., Mendling, J. & Reijers, H. A. (2018). _Fundamentals of Business Process Management,_ 2nd ed.** Springer. — Textbook covering BPMN within the broader BPM lifecycle.
---
## Core BPMN elements
BPMN has hundreds of symbols. In practice, ~80% of useful diagrams use only
~10 of them. The skill's swim-lane output uses precisely these.
### Flow objects
- **Activity (task)** — a unit of work done by one role. Rectangle with rounded
corners. In the skill's swim-lane: this is a `value-add` or `rework` stage.
- **Event** — something that happens (start, intermediate, end). Circles. The
skill represents start/end implicitly as the first and last stage.
- **Gateway** — branching / merging point. Diamond. Common types:
- **Exclusive (XOR)** — one path taken.
- **Parallel (AND)** — all paths taken.
- **Inclusive (OR)** — one or more paths taken based on data.
### Connecting objects
- **Sequence flow** — solid arrow inside one pool. The skill renders these as
`->` between stages in the lane.
- **Message flow** — dashed arrow across pool boundaries. The skill's
cross-lane handoffs (e.g., Requestor -> Manager) are message-flow-equivalents.
- **Association** — dotted line linking a data object to an activity.
### Swim lanes
- **Pool** — represents a participant (a company, a department, or a system).
Each pool is independent; communication between pools uses message flow only.
- **Lane** — a sub-partition within a pool, usually a role or sub-team.
The skill maps one stage's `owner` field to one lane. The full diagram is a
single pool with multiple lanes — appropriate for an internal business process
where one organization controls the whole flow.
---
## Method and Style rules (Silver)
Silver's "Method and Style" is a set of practitioner conventions that make
BPMN diagrams readable. The most load-bearing rules:
1. **One start, one end** per pool. Multiple end events are allowed only if
they represent different end-states (e.g., approved vs. rejected).
2. **Label every flow out of a gateway** with the condition (e.g., "amount >
$10K"). An unlabeled gateway is unreadable.
3. **Sequence flow stays inside a pool.** Use message flow between pools.
4. **One verb-noun task name.** "Approve PO" beats "Approval step."
5. **Black-box pools** for participants you don't model in detail (e.g., the
customer). Show only the message exchanges with them.
The skill enforces rule #4 implicitly by encouraging "Stage" names like
"Manager approves request" rather than "Approval."
---
## Common notation mistakes (Recker 2010; Freund/Rücker)
The following errors appear in over half of real-world BPMN diagrams:
| Mistake | Why it's wrong | What to do |
|---------|----------------|------------|
| Using sequence flow across pools | Pools are independent; only messages cross | Use dashed message flow |
| Missing gateway labels | The reader can't tell which path is taken when | Label every outbound flow |
| Multiple unrelated end events | Reader can't tell why a process ends in each spot | Consolidate or label by end-state |
| Conflating role with system | "JIRA" is a system, not a role; "Engineering Manager" is a role | Lanes = roles, not tools |
| Implicit gateways | Diverging sequence flows without a gateway diamond | Add an explicit XOR or parallel gateway |
| Modeling exceptions inline | Cluttered happy path | Use boundary events or a separate exception sub-process |
| No data objects | Reader doesn't know what artifacts move through | Add data-object boxes where they help |
The skill's stage-level `type` field (`value-add` | `wait` | `rework`) captures
the rework case explicitly so it doesn't get hidden inline. Users who want
full BPMN fidelity should export the normalized JSON and ingest it into a
BPMN-aware tool (Camunda Modeler, bpmn.io, Signavio).
---
## When to use BPMN vs. simpler notations
BPMN is appropriate when:
- The process has cross-functional handoffs (multiple lanes).
- The process has branching logic (gateways).
- The diagram will be reviewed by people who don't sit through a walkthrough.
For purely linear processes with no branching, a numbered list or a value
stream map is faster to produce and easier to read. The skill's swim-lane
output deliberately occupies the middle ground: more structured than a list,
less ceremony than full BPMN.
---
## BPMN 2.0 execution semantics
ISO/IEC 19510:2013 specifies executable semantics so that a BPMN diagram can
be loaded into a workflow engine (Camunda, jBPM, Activiti) and run directly.
The skill does not target executable BPMN — its output is for human reading
and constraint analysis. If a user wants to move from documentation to
automation, the normalized JSON is a starting point; mapping to the BPMN 2.0
XML schema is a separate exercise.
FILE:references/lean_six_sigma_canon.md
# Lean / Six Sigma / Theory-of-Constraints Canon
A working reference for the process-mapper skill. The concepts below are the
intellectual foundation for every detection rule and verdict band the skill
emits. Citations are deliberately to the primary sources, not blog posts.
## Sources
1. **Womack, J. P. & Jones, D. T. (1996). _Lean Thinking: Banish Waste and Create Wealth in Your Corporation._** Free Press. — The five-step Lean discipline: specify value, identify the value stream, make value flow, let the customer pull, pursue perfection.
2. **Rother, M. & Shook, J. (1999). _Learning to See: Value Stream Mapping to Add Value and Eliminate Muda._** Lean Enterprise Institute. — The canonical text on Value Stream Mapping (VSM); origin of current-state / future-state map distinction.
3. **Goldratt, E. M. (1984). _The Goal: A Process of Ongoing Improvement._** North River Press. — The Theory of Constraints: identify, exploit, subordinate, elevate, repeat. Every process has exactly one binding constraint at a time.
4. **Ohno, T. (1988). _Toyota Production System: Beyond Large-Scale Production._** Productivity Press. — Origin of the seven wastes (muda), pull system, jidoka, and andon discipline.
5. **Liker, J. K. (2004). _The Toyota Way: 14 Management Principles from the World's Greatest Manufacturer._** McGraw-Hill. — Modern systemic treatment of TPS principles for non-manufacturing operations.
6. **Pyzdek, T. & Keller, P. (2018). _The Six Sigma Handbook,_ 5th ed.** McGraw-Hill. — DMAIC discipline, SIPOC, process-capability indices, defect-rate measurement.
7. **Anderson, D. J. (2010). _Kanban: Successful Evolutionary Change for Your Technology Business._** Blue Hole Press. — WIP limits, pull system applied to knowledge work, cumulative flow diagrams.
8. **Reinertsen, D. G. (2009). _The Principles of Product Development Flow._** Celeritas Publishing. — Queueing theory for knowledge-work product development; cost of delay.
---
## The Seven Wastes (TIMWOOD)
Ohno's original taxonomy, with the eighth ("non-utilized talent") added later:
| Code | Waste | What it looks like in business processes |
|------|-------|--------------------------------------------|
| **T** | Transport | Moving work between systems / inboxes / queues for no reason |
| **I** | Inventory | Backlogs of pending tickets, unprocessed invoices, open POs |
| **M** | Motion | People hunting for information, switching tools, reading email threads to reconstruct context |
| **W** | Waiting | Work sitting in someone's queue (the largest waste in office work) |
| **O** | Over-production | Producing forecasts, reports, or work nobody requested |
| **O** | Over-processing | Approval chains that add no scrutiny, gold-plating |
| **D** | Defects | Errors that force rework downstream |
| **(N)** | Non-utilized talent | Skilled people doing low-skill work |
The process-mapper skill identifies these via stage `type`: `wait` captures
**W** (and often **I**); `rework` captures **D**. Mis-labelling a wait stage as
`value-add` is the most common data-quality failure and will mask the true
constraint.
---
## Value Stream Mapping (Rother & Shook)
VSM separates **process time** (PT) from **lead time** (LT). For each stage:
- **PT** = the time work actually spends being touched.
- **LT** = the elapsed wall-clock time from when work arrives at the stage to
when it leaves.
In the process-mapper schema, a `value-add` stage's `duration_minutes_p50` is
PT-like; a `wait` stage's duration is the LT component between PT-stages.
The **process cycle efficiency** (PCE) is:
PCE = Total value-add time / Total lead time
This is exactly what `cycle_time_analyzer.py` computes as the "value-add ratio."
Rother & Shook's published benchmarks: office processes typically score
PCE < 10%; well-run service operations land 10–25%; world-class manufacturing
can clear 25–40%.
---
## Theory of Constraints (Goldratt)
Goldratt's Five Focusing Steps:
1. **Identify** the constraint.
2. **Exploit** it (squeeze every minute of capacity from the constraint).
3. **Subordinate** everything else to the constraint.
4. **Elevate** the constraint (only after step 2 is exhausted, add capacity).
5. **Repeat** — once the constraint moves, return to step 1.
Two implications used in the skill:
- **Optimizing a non-constraint stage produces no system improvement.** It
builds inventory in front of the constraint. The `bottleneck_detector.py`
output is ranked by impact specifically so users target the constraint
first.
- **The constraint is almost always a wait stage in office work.** This is
why Rule R2 (wait-share > 40%) is heavily weighted.
---
## Kanban WIP Limits (Anderson)
Little's Law:
L = lambda * W
Where L = items in the system (WIP), lambda = throughput (items per unit time),
and W = average cycle time. Rearranged:
lambda = L / W
Two practical consequences:
- **Cycle time scales linearly with WIP.** Cutting WIP in half cuts cycle time
in half (other things equal). This is why the skill computes throughput from
WIP / cycle time and surfaces a WIP-limit recommendation when wait-share is
high.
- **Adding people to a wait-bound process makes it worse.** New workers add
WIP without expanding the constraint, lengthening cycle time. The
`bottleneck_detector` action text says this explicitly.
---
## Six Sigma DMAIC and Rework
Pyzdek's DMAIC (Define, Measure, Analyze, Improve, Control) treats rework as
a downstream symptom of an upstream defect. The Six-Sigma rule the skill
encodes: **rework is always solved upstream, never downstream.** Adding a
quality-control inspector at the end of the line catches defects but doesn't
prevent them, and inspection-as-quality is itself a TIMWOOD waste
(over-processing).
The poka-yoke (error-proofing) recommendation in Rule R3 follows directly:
add the check at the earliest stage that can detect the defect.
---
## Reinertsen's Queueing Insights
Reinertsen's _Principles of Product Development Flow_ adapts manufacturing
queueing theory to knowledge work. Key results used in the skill:
- **High utilization explodes queue length.** A worker at 90% utilization has
~10x the queue of a worker at 50% utilization. Office workflows that pin
approvers at 100% utilization see wait stages grow without bound.
- **Small batches cut queue time.** Batched approvals (e.g., weekly review
cycles) inflate P50 wait times by half the batch interval on average.
When the skill recommends "remove the handoff or batch," this is the canon
behind it.
FILE:scripts/bottleneck_detector.py
#!/usr/bin/env python3
"""bottleneck_detector.py
Apply three deterministic detection rules to a process JSON and emit a ranked
list of bottlenecks with severity, root-cause hypothesis, and a recommended
action.
Rules (defaults; tuned per industry profile):
R1. Stage P50 > 2x mean of value-add stages -> stage bottleneck
R2. Wait-state share of total cycle > 40% -> handoff bottleneck
R3. Rework share of total cycle > 15% -> quality bottleneck
Stdlib only.
"""
from __future__ import annotations
import argparse
import json
import statistics
import sys
from dataclasses import dataclass, asdict
from pathlib import Path
# Per-industry threshold calibration. Manufacturing tolerates less wait;
# healthcare and services tolerate more given regulatory / human-in-the-loop steps.
PROFILES: dict[str, dict[str, float]] = {
"saas": {
"stage_multiplier": 2.0,
"wait_share_max": 0.40,
"rework_share_max": 0.15,
},
"services": {
"stage_multiplier": 2.5,
"wait_share_max": 0.50,
"rework_share_max": 0.15,
},
"manufacturing": {
"stage_multiplier": 1.8,
"wait_share_max": 0.30,
"rework_share_max": 0.10,
},
"healthcare": {
"stage_multiplier": 2.5,
"wait_share_max": 0.55,
"rework_share_max": 0.12,
},
}
@dataclass
class Finding:
severity: str # CRITICAL | HIGH | MEDIUM
rule: str # R1 | R2 | R3
title: str
detail: str
hypothesis: str
action: str
impact_minutes_p50: float
def severity_rank(self) -> int:
return {"CRITICAL": 0, "HIGH": 1, "MEDIUM": 2}.get(self.severity, 3)
def load(path: Path) -> dict:
with path.open("r", encoding="utf-8") as f:
return json.load(f)
def classify_severity(share: float, threshold: float) -> str:
"""Severity based on how far over the threshold the offender is."""
if share <= threshold:
return "MEDIUM"
if share >= threshold * 2:
return "CRITICAL"
if share >= threshold * 1.5:
return "HIGH"
return "MEDIUM"
def detect(normalized: dict, profile: str) -> list[Finding]:
prof = PROFILES.get(profile, PROFILES["saas"])
stages = normalized.get("stages", [])
findings: list[Finding] = []
if not stages:
return findings
total_p50 = sum(s["duration_minutes_p50"] for s in stages) or 1.0
wait_p50 = sum(s["duration_minutes_p50"] for s in stages if s["type"] == "wait")
rework_p50 = sum(s["duration_minutes_p50"] for s in stages if s["type"] == "rework")
va_durations = [
s["duration_minutes_p50"] for s in stages if s["type"] == "value-add"
]
va_mean = statistics.mean(va_durations) if va_durations else 0.0
# R1: per-stage runaway vs value-add mean
if va_mean > 0:
threshold_minutes = va_mean * prof["stage_multiplier"]
for s in stages:
if s["duration_minutes_p50"] > threshold_minutes:
ratio = s["duration_minutes_p50"] / va_mean
if ratio >= prof["stage_multiplier"] * 3:
sev = "CRITICAL"
elif ratio >= prof["stage_multiplier"] * 2:
sev = "HIGH"
else:
sev = "MEDIUM"
hypothesis = (
"Stage runs much longer than the typical value-add step; "
"common causes: batched approvals, single approver, "
"missing self-service, or unclear acceptance criteria."
)
action = (
"Decompose the stage; check if approval can be parallelized "
"or made conditional. If wait-state, apply Kanban WIP limit "
"or remove the handoff."
)
findings.append(
Finding(
severity=sev,
rule="R1",
title=f"Slow stage: {s['name']}",
detail=(
f"P50 {s['duration_minutes_p50']:.0f} min vs value-add "
f"mean {va_mean:.1f} min (ratio {ratio:.1f}x)."
),
hypothesis=hypothesis,
action=action,
impact_minutes_p50=s["duration_minutes_p50"],
)
)
# R2: wait-state share
wait_share = wait_p50 / total_p50
if wait_share > prof["wait_share_max"]:
sev = classify_severity(wait_share, prof["wait_share_max"])
findings.append(
Finding(
severity=sev,
rule="R2",
title="Process is dominated by wait time",
detail=(
f"Wait stages account for {wait_share*100:.0f}% of total P50, "
f"vs {prof['wait_share_max']*100:.0f}% profile threshold."
),
hypothesis=(
"Handoffs queue work behind a single role or batch. Per "
"Theory of Constraints, the system throughput is set by "
"whichever queue is longest, not by stage speed."
),
action=(
"Identify the longest wait stage; pull it forward, eliminate "
"it via self-service, or apply a WIP limit upstream so the "
"queue cannot grow."
),
impact_minutes_p50=wait_p50,
)
)
# R3: rework share
rework_share = rework_p50 / total_p50
if rework_share > prof["rework_share_max"]:
sev = classify_severity(rework_share, prof["rework_share_max"])
findings.append(
Finding(
severity=sev,
rule="R3",
title="Process has excessive rework",
detail=(
f"Rework accounts for {rework_share*100:.0f}% of total P50, "
f"vs {prof['rework_share_max']*100:.0f}% profile threshold."
),
hypothesis=(
"Defects escape upstream stages. Six-Sigma canon: rework is "
"always an upstream-quality problem, never a downstream one."
),
action=(
"Add a poka-yoke (error-proofing) check at the earliest stage "
"that can detect the defect; do not add inspection downstream."
),
impact_minutes_p50=rework_p50,
)
)
findings.sort(key=lambda f: (f.severity_rank(), -f.impact_minutes_p50))
return findings
def render_markdown(normalized: dict, findings: list[Finding], profile: str) -> str:
name = normalized.get("process_name", "Untitled Process")
lines: list[str] = []
lines.append(f"# Bottleneck Detection: {name}")
lines.append("")
lines.append(f"**Profile:** `{profile}` ")
lines.append(f"**Findings:** {len(findings)}")
lines.append("")
if not findings:
lines.append("_No bottlenecks detected at the configured thresholds._")
return "\n".join(lines)
for i, f in enumerate(findings, 1):
lines.append(f"## {i}. [{f.severity}] {f.title}")
lines.append("")
lines.append(f"- **Rule:** `{f.rule}`")
lines.append(f"- **Detail:** {f.detail}")
lines.append(f"- **Hypothesis:** {f.hypothesis}")
lines.append(f"- **Recommended action:** {f.action}")
lines.append(f"- **Impact (P50 minutes):** {f.impact_minutes_p50:.0f}")
lines.append("")
return "\n".join(lines)
def sample_process() -> dict:
# Reuses procurement-intake shape from process_documenter
return {
"process_name": "Procurement Intake (Sample)",
"wip": 12,
"stages": [
{"name": "Submit PO", "owner": "Requestor", "type": "value-add",
"duration_minutes_p50": 15, "duration_minutes_p90": 30},
{"name": "Wait for manager", "owner": "Manager", "type": "wait",
"duration_minutes_p50": 480, "duration_minutes_p90": 1440},
{"name": "Manager approves", "owner": "Manager", "type": "value-add",
"duration_minutes_p50": 10, "duration_minutes_p90": 25},
{"name": "Wait for finance", "owner": "Finance", "type": "wait",
"duration_minutes_p50": 720, "duration_minutes_p90": 2880},
{"name": "Finance validates", "owner": "Finance", "type": "value-add",
"duration_minutes_p50": 20, "duration_minutes_p90": 60},
{"name": "Rework: missing W-9", "owner": "Requestor", "type": "rework",
"duration_minutes_p50": 120, "duration_minutes_p90": 360},
],
}
def main() -> int:
parser = argparse.ArgumentParser(
description="Detect bottlenecks in a documented business process."
)
parser.add_argument("--input", type=Path, help="Path to process JSON file.")
parser.add_argument(
"--profile",
choices=sorted(PROFILES.keys()),
default="saas",
help="Industry profile for threshold calibration (default: saas).",
)
parser.add_argument(
"--output",
choices=["markdown", "json"],
default="markdown",
help="Output format (default: markdown).",
)
parser.add_argument(
"--sample",
action="store_true",
help="Use a built-in sample process and exit.",
)
args = parser.parse_args()
if args.sample:
raw = sample_process()
else:
if not args.input:
parser.error("--input is required unless --sample is given")
if not args.input.exists():
parser.error(f"input file not found: {args.input}")
raw = load(args.input)
# Minimal normalization: tolerate the same fields as process_documenter
stages = []
for s in raw.get("stages", []):
stages.append(
{
"name": s.get("name", ""),
"owner": s.get("owner", ""),
"type": s.get("type", ""),
"duration_minutes_p50": float(s.get("duration_minutes_p50", 0)),
"duration_minutes_p90": float(s.get("duration_minutes_p90", 0)),
}
)
normalized = {
"process_name": raw.get("process_name", "Untitled Process"),
"wip": int(raw.get("wip", 0) or 0),
"stages": stages,
}
findings = detect(normalized, args.profile)
if args.output == "json":
print(
json.dumps(
{
"process_name": normalized["process_name"],
"profile": args.profile,
"findings": [asdict(f) for f in findings],
},
indent=2,
)
)
else:
print(render_markdown(normalized, findings, args.profile))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/cycle_time_analyzer.py
#!/usr/bin/env python3
"""cycle_time_analyzer.py
Compute total cycle time (P50, P90), value-add ratio (VA%), wait %, rework %,
and a Little's-Law throughput estimate for a documented business process.
Verdict per Lean canon:
VA% > 25% -> HEALTHY
10% <= VA% <= 25% -> TYPICAL
VA% < 10% -> WASTE-HEAVY
Stdlib only.
"""
from __future__ import annotations
import argparse
import json
import sys
from dataclasses import dataclass, asdict
from pathlib import Path
# Per-industry verdict bands. Manufacturing benchmarks higher VA% than services.
PROFILES: dict[str, dict[str, float]] = {
"saas": {"healthy": 0.25, "typical": 0.10},
"services": {"healthy": 0.20, "typical": 0.08},
"manufacturing": {"healthy": 0.35, "typical": 0.15},
"healthcare": {"healthy": 0.20, "typical": 0.08},
}
@dataclass
class CycleTimeReport:
process_name: str
profile: str
stage_count: int
total_p50_minutes: float
total_p90_minutes: float
value_add_minutes_p50: float
wait_minutes_p50: float
rework_minutes_p50: float
value_add_ratio: float
wait_ratio: float
rework_ratio: float
verdict: str
wip: int
throughput_per_hour: float | None
notes: list[str]
def analyze(normalized: dict, profile: str) -> CycleTimeReport:
prof = PROFILES.get(profile, PROFILES["saas"])
stages = normalized.get("stages", [])
name = normalized.get("process_name", "Untitled Process")
wip = int(normalized.get("wip", 0) or 0)
total_p50 = sum(s["duration_minutes_p50"] for s in stages)
total_p90 = sum(s["duration_minutes_p90"] for s in stages)
va_p50 = sum(s["duration_minutes_p50"] for s in stages if s["type"] == "value-add")
wait_p50 = sum(s["duration_minutes_p50"] for s in stages if s["type"] == "wait")
rework_p50 = sum(s["duration_minutes_p50"] for s in stages if s["type"] == "rework")
denom = total_p50 if total_p50 > 0 else 1.0
va_ratio = va_p50 / denom
wait_ratio = wait_p50 / denom
rework_ratio = rework_p50 / denom
if va_ratio >= prof["healthy"]:
verdict = "HEALTHY"
elif va_ratio >= prof["typical"]:
verdict = "TYPICAL"
else:
verdict = "WASTE-HEAVY"
# Little's Law: L = lambda * W => lambda = L / W
# WIP is items currently in process; W (cycle time) is P50.
# Convert minutes to hours for a per-hour throughput.
throughput = None
if wip > 0 and total_p50 > 0:
cycle_hours = total_p50 / 60.0
throughput = wip / cycle_hours
notes: list[str] = []
if wip <= 0:
notes.append(
"WIP not provided; Little's-Law throughput estimate skipped. "
"Set 'wip' in the input JSON to enable it."
)
if total_p50 == 0:
notes.append("All stage P50 durations are zero; check input data.")
if rework_ratio > 0.0 and verdict == "HEALTHY":
notes.append(
"Process is healthy by VA%, but rework is non-zero. Six-Sigma canon: "
"any rework signal is worth a poka-yoke check."
)
if wait_ratio > 0.5:
notes.append(
"More than half the cycle is wait time. Throughput improves more "
"from queue removal than from speeding up value-add stages."
)
return CycleTimeReport(
process_name=name,
profile=profile,
stage_count=len(stages),
total_p50_minutes=round(total_p50, 2),
total_p90_minutes=round(total_p90, 2),
value_add_minutes_p50=round(va_p50, 2),
wait_minutes_p50=round(wait_p50, 2),
rework_minutes_p50=round(rework_p50, 2),
value_add_ratio=round(va_ratio, 4),
wait_ratio=round(wait_ratio, 4),
rework_ratio=round(rework_ratio, 4),
verdict=verdict,
wip=wip,
throughput_per_hour=round(throughput, 4) if throughput is not None else None,
notes=notes,
)
def render_markdown(report: CycleTimeReport) -> str:
lines: list[str] = []
lines.append(f"# Cycle-Time Analysis: {report.process_name}")
lines.append("")
lines.append(f"**Profile:** `{report.profile}` ")
lines.append(f"**Verdict:** **{report.verdict}**")
lines.append("")
lines.append("## Summary")
lines.append("")
lines.append("| Metric | Value |")
lines.append("|--------|-------|")
lines.append(f"| Stage count | {report.stage_count} |")
lines.append(f"| Total P50 (minutes) | {report.total_p50_minutes:.1f} |")
lines.append(f"| Total P90 (minutes) | {report.total_p90_minutes:.1f} |")
lines.append(
f"| Value-add minutes (P50) | {report.value_add_minutes_p50:.1f} |"
)
lines.append(f"| Wait minutes (P50) | {report.wait_minutes_p50:.1f} |")
lines.append(f"| Rework minutes (P50) | {report.rework_minutes_p50:.1f} |")
lines.append(f"| Value-add ratio (VA%) | {report.value_add_ratio*100:.1f}% |")
lines.append(f"| Wait ratio | {report.wait_ratio*100:.1f}% |")
lines.append(f"| Rework ratio | {report.rework_ratio*100:.1f}% |")
lines.append(f"| WIP (items in process) | {report.wip} |")
if report.throughput_per_hour is not None:
lines.append(
f"| Little's-Law throughput | {report.throughput_per_hour:.3f} items/hour |"
)
else:
lines.append("| Little's-Law throughput | _(needs WIP > 0 in input)_ |")
lines.append("")
if report.notes:
lines.append("## Notes")
lines.append("")
for n in report.notes:
lines.append(f"- {n}")
lines.append("")
return "\n".join(lines)
def sample_process() -> dict:
return {
"process_name": "Procurement Intake (Sample)",
"wip": 12,
"stages": [
{"name": "Submit PO", "owner": "Requestor", "type": "value-add",
"duration_minutes_p50": 15, "duration_minutes_p90": 30},
{"name": "Wait for manager", "owner": "Manager", "type": "wait",
"duration_minutes_p50": 480, "duration_minutes_p90": 1440},
{"name": "Manager approves", "owner": "Manager", "type": "value-add",
"duration_minutes_p50": 10, "duration_minutes_p90": 25},
{"name": "Wait for finance", "owner": "Finance", "type": "wait",
"duration_minutes_p50": 720, "duration_minutes_p90": 2880},
{"name": "Finance validates", "owner": "Finance", "type": "value-add",
"duration_minutes_p50": 20, "duration_minutes_p90": 60},
{"name": "Rework: missing W-9", "owner": "Requestor", "type": "rework",
"duration_minutes_p50": 120, "duration_minutes_p90": 360},
],
}
def main() -> int:
parser = argparse.ArgumentParser(
description="Analyze cycle time, value-add ratio, and throughput of a process."
)
parser.add_argument("--input", type=Path, help="Path to process JSON file.")
parser.add_argument(
"--profile",
choices=sorted(PROFILES.keys()),
default="saas",
help="Industry profile for verdict band (default: saas).",
)
parser.add_argument(
"--output",
choices=["markdown", "json"],
default="markdown",
help="Output format (default: markdown).",
)
parser.add_argument(
"--sample",
action="store_true",
help="Use a built-in sample process and exit.",
)
args = parser.parse_args()
if args.sample:
raw = sample_process()
else:
if not args.input:
parser.error("--input is required unless --sample is given")
if not args.input.exists():
parser.error(f"input file not found: {args.input}")
with args.input.open("r", encoding="utf-8") as f:
raw = json.load(f)
stages = []
for s in raw.get("stages", []):
stages.append(
{
"name": s.get("name", ""),
"owner": s.get("owner", ""),
"type": s.get("type", ""),
"duration_minutes_p50": float(s.get("duration_minutes_p50", 0)),
"duration_minutes_p90": float(s.get("duration_minutes_p90", 0)),
}
)
normalized = {
"process_name": raw.get("process_name", "Untitled Process"),
"wip": int(raw.get("wip", 0) or 0),
"stages": stages,
}
report = analyze(normalized, args.profile)
if args.output == "json":
print(json.dumps(asdict(report), indent=2))
else:
print(render_markdown(report))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/process_documenter.py
#!/usr/bin/env python3
"""process_documenter.py
Read a JSON description of a business process (one entry per stage) and emit:
- a text-based BPMN-style swim-lane diagram in Markdown, OR
- a normalized JSON artifact for downstream tools.
Stdlib only. Use `--sample` to print a 6-stage procurement-intake example to
stdout.
Input schema (JSON):
{
"process_name": "Procurement Intake",
"wip": 12, # optional, integer; used by cycle_time_analyzer
"stages": [
{
"name": "Requestor submits PO request",
"owner": "Requestor",
"type": "value-add", # one of: value-add | wait | rework
"duration_minutes_p50": 15,
"duration_minutes_p90": 30
},
...
]
}
"""
from __future__ import annotations
import argparse
import json
import sys
from dataclasses import dataclass, asdict
from enum import Enum
from pathlib import Path
VALID_TYPES = {"value-add", "wait", "rework"}
class StageType(str, Enum):
VALUE_ADD = "value-add"
WAIT = "wait"
REWORK = "rework"
@dataclass
class Stage:
name: str
owner: str
type: str
duration_minutes_p50: float
duration_minutes_p90: float
def validate(self, idx: int) -> list[str]:
errs: list[str] = []
if not self.name:
errs.append(f"stage[{idx}]: missing 'name'")
if not self.owner:
errs.append(f"stage[{idx}]: missing 'owner'")
if self.type not in VALID_TYPES:
errs.append(
f"stage[{idx}] ('{self.name}'): invalid type '{self.type}' "
f"(expected one of {sorted(VALID_TYPES)})"
)
if self.duration_minutes_p50 < 0:
errs.append(f"stage[{idx}] ('{self.name}'): p50 must be >= 0")
if self.duration_minutes_p90 < self.duration_minutes_p50:
errs.append(
f"stage[{idx}] ('{self.name}'): p90 ({self.duration_minutes_p90}) "
f"< p50 ({self.duration_minutes_p50})"
)
return errs
def load_process(path: Path) -> dict:
with path.open("r", encoding="utf-8") as f:
return json.load(f)
def normalize(raw: dict) -> dict:
"""Validate + return a normalized dict. Raises ValueError on bad input."""
if "stages" not in raw or not isinstance(raw["stages"], list):
raise ValueError("input must include a non-empty 'stages' list")
stages: list[Stage] = []
errors: list[str] = []
for idx, s in enumerate(raw["stages"]):
try:
stage = Stage(
name=s.get("name", ""),
owner=s.get("owner", ""),
type=s.get("type", ""),
duration_minutes_p50=float(s.get("duration_minutes_p50", 0)),
duration_minutes_p90=float(s.get("duration_minutes_p90", 0)),
)
except (TypeError, ValueError) as e:
errors.append(f"stage[{idx}]: parse error: {e}")
continue
errors.extend(stage.validate(idx))
stages.append(stage)
if errors:
raise ValueError("invalid input:\n - " + "\n - ".join(errors))
return {
"process_name": raw.get("process_name", "Untitled Process"),
"wip": int(raw.get("wip", 0)) if raw.get("wip") is not None else 0,
"stages": [asdict(s) for s in stages],
}
def render_markdown(normalized: dict) -> str:
"""Render a text-based BPMN-style swim-lane diagram in Markdown."""
name = normalized["process_name"]
stages = normalized["stages"]
lines: list[str] = []
lines.append(f"# Process Map: {name}")
lines.append("")
lines.append(f"**Stages:** {len(stages)} ")
lines.append(
f"**Total P50:** {sum(s['duration_minutes_p50'] for s in stages):.1f} min "
)
lines.append(
f"**Total P90:** {sum(s['duration_minutes_p90'] for s in stages):.1f} min"
)
lines.append("")
# Group by owner -> swim lane
lanes: dict[str, list[tuple[int, dict]]] = {}
for idx, s in enumerate(stages):
lanes.setdefault(s["owner"], []).append((idx, s))
lines.append("## Swim Lanes")
lines.append("")
type_glyph = {"value-add": "[V]", "wait": "[W]", "rework": "[R]"}
lane_width = max(20, max((len(o) for o in lanes), default=20) + 4)
sep = "+" + "-" * (lane_width + 2) + "+" + "-" * 72 + "+"
lines.append("```")
lines.append(sep)
lines.append(
"| " + "OWNER".ljust(lane_width) + " | " + "STAGES (in process order)".ljust(70) + " |"
)
lines.append(sep)
for owner, owned in lanes.items():
owner_cell = owner.ljust(lane_width)
cells = []
for idx, s in owned:
glyph = type_glyph.get(s["type"], "[?]")
cells.append(
f"#{idx+1} {glyph} {s['name'][:32]} "
f"(p50={s['duration_minutes_p50']:.0f}m)"
)
row_text = " -> ".join(cells)
# Wrap row_text to 70 chars
wrapped = []
cur = ""
for token in row_text.split(" "):
if len(cur) + len(token) + 1 > 70:
wrapped.append(cur)
cur = token
else:
cur = (cur + " " + token).strip()
if cur:
wrapped.append(cur)
for i, line in enumerate(wrapped):
left = owner_cell if i == 0 else " " * lane_width
lines.append(f"| {left} | {line.ljust(70)} |")
lines.append(sep)
lines.append("```")
lines.append("")
lines.append("Legend: `[V]` value-add `[W]` wait `[R]` rework")
lines.append("")
lines.append("## Linear sequence")
lines.append("")
lines.append("| # | Stage | Owner | Type | P50 (min) | P90 (min) |")
lines.append("|---|-------|-------|------|-----------|-----------|")
for idx, s in enumerate(stages):
lines.append(
f"| {idx+1} | {s['name']} | {s['owner']} | {s['type']} | "
f"{s['duration_minutes_p50']:.1f} | {s['duration_minutes_p90']:.1f} |"
)
lines.append("")
return "\n".join(lines)
def sample_process() -> dict:
return {
"process_name": "Procurement Intake (Sample)",
"wip": 12,
"stages": [
{
"name": "Requestor submits PO request",
"owner": "Requestor",
"type": "value-add",
"duration_minutes_p50": 15,
"duration_minutes_p90": 30,
},
{
"name": "Wait for manager review queue",
"owner": "Manager",
"type": "wait",
"duration_minutes_p50": 480,
"duration_minutes_p90": 1440,
},
{
"name": "Manager approves request",
"owner": "Manager",
"type": "value-add",
"duration_minutes_p50": 10,
"duration_minutes_p90": 25,
},
{
"name": "Wait for finance review queue",
"owner": "Finance",
"type": "wait",
"duration_minutes_p50": 720,
"duration_minutes_p90": 2880,
},
{
"name": "Finance validates budget code",
"owner": "Finance",
"type": "value-add",
"duration_minutes_p50": 20,
"duration_minutes_p90": 60,
},
{
"name": "Rework: missing vendor W-9",
"owner": "Requestor",
"type": "rework",
"duration_minutes_p50": 120,
"duration_minutes_p90": 360,
},
],
}
def main() -> int:
parser = argparse.ArgumentParser(
description="Document a business process as a BPMN-style swim-lane diagram."
)
parser.add_argument("--input", type=Path, help="Path to process JSON file.")
parser.add_argument(
"--output", type=Path, help="Output file path (default: stdout)."
)
parser.add_argument(
"--format",
choices=["markdown", "json"],
default="markdown",
help="Output format (default: markdown).",
)
parser.add_argument(
"--sample",
action="store_true",
help="Print a 6-stage procurement-intake sample and exit.",
)
args = parser.parse_args()
if args.sample:
raw = sample_process()
else:
if not args.input:
parser.error("--input is required unless --sample is given")
if not args.input.exists():
parser.error(f"input file not found: {args.input}")
raw = load_process(args.input)
try:
normalized = normalize(raw)
except ValueError as e:
print(f"ERROR: {e}", file=sys.stderr)
return 1
if args.format == "json":
out = json.dumps(normalized, indent=2)
else:
out = render_markdown(normalized)
if args.output:
args.output.write_text(out, encoding="utf-8")
print(f"wrote {args.output}", file=sys.stderr)
else:
print(out)
return 0
if __name__ == "__main__":
sys.exit(main())
Tạo hoặc cập nhật tài liệu ngữ cảnh product marketing: mô tả sản phẩm, đối tượng mục tiêu, ICP và định vị để tránh lặp lại thông tin nền.
---
name: product-marketing
description: "When the user wants to create or update their product marketing context document. Also use when the user mentions 'product context,' 'marketing context,' 'set up context,' 'positioning,' 'who is my target audience,' 'describe my product,' 'ICP,' 'ideal customer profile,' or wants to avoid repeating foundational information across marketing tasks. Use this at the start of any new project before using other marketing skills — it creates `.agents/product-marketing.md` that all other skills reference for product, audience, and positioning context."
metadata:
version: 2.1.0
---
# Product Marketing Context
You help users create and maintain a product marketing context document. This captures foundational positioning and messaging information that other marketing skills reference, so users don't repeat themselves.
The document is stored at `.agents/product-marketing.md`.
## Workflow
### Step 1: Check for Existing Context
First, check if `.agents/product-marketing.md` already exists. Also check `.claude/product-marketing.md` and the legacy filename `product-marketing-context.md` (in either `.agents/` or `.claude/`) for older setups — if found anywhere other than `.agents/product-marketing.md`, offer to move it to the canonical location.
**If it exists:**
- Read it and summarize what's captured — note its current **Document version** and the last few **Changelog** entries so the user sees where the doc stands and what's changed recently
- Ask which sections they want to update
- Only gather info for those sections
- On any substantive save, bump the version and add a changelog entry (see Step 4). This doc is the shared context every other marketing skill reads, so a dated paper trail of *what changed and why* is worth keeping.
**If it doesn't exist, offer two options:**
1. **Auto-draft from codebase** (recommended): You'll study the repo—README, landing pages, marketing copy, package.json, etc.—and draft a V1 of the context document. The user then reviews, corrects, and fills gaps. This is faster than starting from scratch.
2. **Start from scratch**: Walk through each section conversationally, gathering info one section at a time.
Most users prefer option 1. After presenting the draft, ask: "What needs correcting? What's missing?"
### Step 2: Gather Information
**If auto-drafting:**
1. Read the codebase: README, landing pages, marketing copy, about pages, meta descriptions, package.json, any existing docs
2. Draft all sections based on what you find
3. Present the draft and ask what needs correcting or is missing
4. Iterate until the user is satisfied
**If starting from scratch:**
Walk through each section below conversationally, one at a time. Don't dump all questions at once.
For each section:
1. Briefly explain what you're capturing
2. Ask relevant questions
3. Confirm accuracy
4. Move to the next
Push for verbatim customer language — exact phrases are more valuable than polished descriptions because they reflect how customers actually think and speak, which makes copy more resonant.
---
## Sections to Capture
### 1. Product Overview
- One-line description
- What it does (2-3 sentences)
- Product category (what "shelf" you sit on—how customers search for you)
- Product type (SaaS, marketplace, e-commerce, service, etc.)
- Business model and pricing
### 2. Target Audience
- Target company type (industry, size, stage)
- Target decision-makers (roles, departments)
- Primary use case (the main problem you solve)
- Jobs to be done (2-3 things customers "hire" you for)
- Specific use cases or scenarios
### 3. Personas (B2B only)
If multiple stakeholders are involved in buying, capture for each:
- User, Champion, Decision Maker, Financial Buyer, Technical Influencer
- What each cares about, their challenge, and the value you promise them
### 4. Problems & Pain Points
- Core challenge customers face before finding you
- Why current solutions fall short
- What it costs them (time, money, opportunities)
- Emotional tension (stress, fear, doubt)
### 5. Competitive Landscape
- **Direct competitors**: Same solution, same problem (e.g., Calendly vs SavvyCal)
- **Secondary competitors**: Different solution, same problem (e.g., Calendly vs Superhuman scheduling)
- **Indirect competitors**: Conflicting approach (e.g., Calendly vs personal assistant)
- How each falls short for customers
### 6. Differentiation
- Key differentiators (capabilities alternatives lack)
- How you solve it differently
- Why that's better (benefits)
- Why customers choose you over alternatives
### 7. Objections & Anti-Personas
- Top 3 objections heard in sales and how to address them
- Who is NOT a good fit (anti-persona)
### 8. Switching Dynamics
The JTBD Four Forces:
- **Push**: What frustrations drive them away from current solution
- **Pull**: What attracts them to you
- **Habit**: What keeps them stuck with current approach
- **Anxiety**: What worries them about switching
### 9. Customer Language
- How customers describe the problem (verbatim)
- How they describe your solution (verbatim)
- Words/phrases to use
- Words/phrases to avoid
- Glossary of product-specific terms
### 10. Brand Voice
- Tone (professional, casual, playful, etc.)
- Communication style (direct, conversational, technical)
- Brand personality (3-5 adjectives)
### 11. Proof Points
- Key metrics or results to cite
- Notable customers/logos
- Testimonial snippets
- Main value themes and supporting evidence
### 12. Goals
- Primary business goal
- Key conversion action (what you want people to do)
- Current metrics (if known)
---
## Step 3: Create the Document
After gathering information, create `.agents/product-marketing.md` with this structure:
```markdown
# Product Marketing Context
**Document version:** v1
**Last updated:** [date]
## Product Overview
**One-liner:**
**What it does:**
**Product category:**
**Product type:**
**Business model:**
## Target Audience
**Target companies:**
**Decision-makers:**
**Primary use case:**
**Jobs to be done:**
-
**Use cases:**
-
## Personas
| Persona | Cares about | Challenge | Value we promise |
|---------|-------------|-----------|------------------|
| | | | |
## Problems & Pain Points
**Core problem:**
**Why alternatives fall short:**
-
**What it costs them:**
**Emotional tension:**
## Competitive Landscape
**Direct:** [Competitor] — falls short because...
**Secondary:** [Approach] — falls short because...
**Indirect:** [Alternative] — falls short because...
## Differentiation
**Key differentiators:**
-
**How we do it differently:**
**Why that's better:**
**Why customers choose us:**
## Objections
| Objection | Response |
|-----------|----------|
| | |
**Anti-persona:**
## Switching Dynamics
**Push:**
**Pull:**
**Habit:**
**Anxiety:**
## Customer Language
**How they describe the problem:**
- "[verbatim]"
**How they describe us:**
- "[verbatim]"
**Words to use:**
**Words to avoid:**
**Glossary:**
| Term | Meaning |
|------|---------|
| | |
## Brand Voice
**Tone:**
**Style:**
**Personality:**
## Proof Points
**Metrics:**
**Customers:**
**Testimonials:**
> "[quote]" — [who]
**Value themes:**
| Theme | Proof |
|-------|-------|
| | |
## Goals
**Business goal:**
**Conversion action:**
**Current metrics:**
## Changelog
*Newest first. One line per revision: what changed and why.*
- v1 ([date]) — Initial context.
```
---
## Step 4: Confirm, Version, and Save
- Show the completed document
- Ask if anything needs adjustment
- **Set the version and changelog** — this is the paper trail for a doc every other skill reads:
- **New document:** set `Document version: v1` and a single Changelog entry — `- v1 ([today]) — Initial context.`
- **Updating an existing document:** increment the version (v2 → v3 …), update `Last updated` to today, and **prepend a new Changelog entry** at the top of the list (newest first) summarizing *what changed and why* in one line. Never rewrite or reorder past entries.
- A good entry names the sections touched and the reason, not "updated the doc." Examples:
- `- v3 (2026-07-16) — Repositioned from "email tool" to "deliverability platform"; added RevOps to the ICP.`
- `- v2 (2026-06-02) — Rewrote value prop and objections after 5 customer interviews; added competitor Acme.`
- Use today's date in ISO form (YYYY-MM-DD) for the entry and `Last updated`.
- **Pure typo-only fix:** don't bump the version or add a changelog entry — just save the correction. Every other change bumps the version and gets an entry. When the change is a real repositioning, say so plainly — downstream skills will now generate against the new context.
- Save to `.agents/product-marketing.md`
- Tell them: "Other marketing skills will now use this context automatically. The Changelog at the bottom tracks every revision — check it to see how your positioning has evolved. Run `/product-marketing` anytime to update it."
---
## Tips
- **Be specific**: Ask "What's the #1 frustration that brings them to you?" not "What problem do they solve?"
- **Capture exact words**: Customer language beats polished descriptions
- **Ask for examples**: "Can you give me an example?" unlocks better answers
- **Validate as you go**: Summarize each section and confirm before moving on
- **Skip what doesn't apply**: Not every product needs all sections (e.g., Personas for B2C)
FILE:evals/evals.json
{
"skill_name": "product-marketing",
"evals": [
{
"id": 1,
"prompt": "I want to set up my product marketing context. We're a B2B SaaS company that sells a customer feedback platform to product teams.",
"expected_output": "Should check if .agents/product-marketing.md already exists. If not, should offer two options: (1) Auto-draft from codebase (recommended) or (2) Start from scratch. If user chooses start from scratch, should walk through sections conversationally one at a time. Should cover all applicable sections: Product Overview, Target Audience, Personas, Problems You Solve, Competitive Landscape, Differentiation, Objections, Switching Dynamics, Customer Language, Brand Voice, Proof Points, and Goals. Should create the file at .agents/product-marketing.md when complete.",
"assertions": [
"Checks for existing product-marketing.md",
"Offers two options: auto-draft or start from scratch",
"Covers applicable sections",
"Walks through sections conversationally one at a time",
"Creates file at .agents/product-marketing.md"
],
"files": []
},
{
"id": 2,
"prompt": "Update our product marketing context. We just added a new enterprise tier and our target audience has expanded to include VP of Engineering, not just Product Managers.",
"expected_output": "Should check for existing .agents/product-marketing.md and read it. Should identify which sections need updating based on the changes: Target Audience (add VP of Engineering), Personas (add new persona), Product Overview (new enterprise tier, including pricing updates within that section), Objections (enterprise-specific), and Competitive Landscape (enterprise competitors). Should update only the relevant sections, preserving existing content that hasn't changed.",
"assertions": [
"Reads existing product-marketing.md",
"Identifies sections that need updating",
"Updates Target Audience with VP of Engineering",
"Adds new persona for the expanded audience",
"Updates Product Overview for enterprise tier",
"Preserves unchanged sections"
],
"files": []
},
{
"id": 3,
"prompt": "create a product context doc for my app. it's a mobile app that helps people find hiking trails. we're just getting started.",
"expected_output": "Should trigger on casual phrasing. Should check for existing context doc. Should offer auto-draft or start-from-scratch options. Should adapt questions for an early-stage B2C mobile app (outdoor/fitness niche). Should note that some sections may be sparse for an early-stage product and that's okay — they can be filled in as the business matures. Should skip non-applicable sections (e.g., Personas section is B2B-focused) rather than forcing all 12. Should accept lighter answers for sections like Proof Points or Competitive Landscape if the company is new.",
"assertions": [
"Triggers on casual phrasing",
"Checks for existing context doc",
"Offers auto-draft or start-from-scratch options",
"Adapts questions for early-stage B2C mobile app",
"Notes some sections may be sparse early on",
"Skips non-applicable sections rather than forcing all 12",
"Creates file at .agents/product-marketing.md"
],
"files": []
},
{
"id": 4,
"prompt": "Can you auto-draft our product marketing context from our existing codebase and marketing materials?",
"expected_output": "Should activate the auto-draft workflow mode. Should scan the codebase for existing marketing context: README, landing page copy, pricing page, about page, meta descriptions, any existing documentation. Should draft the product-marketing.md from what it finds, filling in sections where information is available and flagging sections that need manual input. Should present the draft for review before saving.",
"assertions": [
"Activates auto-draft workflow mode",
"Scans codebase for existing marketing materials",
"Drafts context from found information",
"Flags sections needing manual input",
"Presents draft for review before saving"
],
"files": []
},
{
"id": 5,
"prompt": "Do we have a product marketing context set up? I want to make sure the other marketing skills have context about our product.",
"expected_output": "Should check for .agents/product-marketing.md (and the older .claude/product-marketing.md location). Should report whether it exists and summarize its contents if found. If it doesn't exist, should offer to create one and explain why it's valuable (other skills like copywriting, cro, seo-audit check for it first). Should explain how other skills use this context document.",
"assertions": [
"Checks both file locations",
"Reports whether context doc exists",
"Summarizes contents if found",
"Offers to create if missing",
"Explains how other skills use it"
],
"files": []
},
{
"id": 6,
"prompt": "Write homepage copy for our SaaS product.",
"expected_output": "Should recognize this is a copywriting task, not a product marketing context task. Should check for product-marketing.md (as other skills do), and if it doesn't exist, may suggest creating one first. But should defer to the copywriting skill for actually writing the homepage copy.",
"assertions": [
"Recognizes this as a copywriting task",
"May check for or suggest creating product-marketing.md",
"References or defers to copywriting skill for the actual copy",
"Does not attempt to write homepage copy using context creation patterns"
],
"files": []
},
{
"id": 7,
"prompt": "We just repositioned — we're no longer an 'email tool,' we're a 'deliverability platform,' and our ICP now includes RevOps teams. Update our product marketing context.",
"expected_output": "Should recognize an existing .agents/product-marketing.md, read it, note its current Document version and recent Changelog entries, and update only the affected sections (product overview/positioning, target audience/ICP). On save, should bump the Document version (e.g. v2 → v3), update the Last updated date, and PREPEND a new newest-first Changelog entry summarizing what changed and why in one line — e.g. 'Repositioned from email tool to deliverability platform; added RevOps to the ICP' — naming the sections touched and the reason, not just 'updated the doc.' Should not rewrite or reorder past changelog entries. Should tell the user the changelog tracks revisions and that downstream skills will now use the new context.",
"assertions": [
"Reads the existing doc and surfaces its current version + recent changelog",
"Updates only the affected sections (positioning + ICP)",
"Bumps the Document version and updates Last updated",
"Prepends a newest-first changelog entry naming what changed and why",
"Preserves prior changelog entries unchanged"
],
"files": []
}
]
}
Lập kế hoạch và tổng hợp nghiên cứu sản phẩm/người dùng: chọn phương pháp phù hợp, tính độ bão hòa và cỡ mẫu theo độ tin cậy rõ ràng.
---
name: product-research
description: Use when planning and synthesizing product/user research as a method-and-repository discipline — selecting the right method for the goal (generative interviews vs usability test vs concept test vs validation), computing method-based saturation/sample size with an explicit confidence level, or synthesizing coded observations into insights while flagging single-source anecdotes. Never fabricates user insight; an insight requires recurrence across independent participants. Distinct from product-team/ux-researcher-designer (persona/journey artifacts), product-discovery (discovery-sprint planning), and experiment-designer (live A/B) — this is the research-ops method + insight-repository layer.
version: 2.9.0
author: claude-code-skills
license: MIT
tags: [research-ops, product-research, ux-research, jtbd, usability, saturation, insight-synthesis, research-repository]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# product-research
Product / user research as an operational discipline: choosing the right method, sizing it honestly, and synthesizing findings into governed insights. The core rule: **method must match the goal**, and **an insight requires recurrence across independent participants** — a single quote is an anecdote.
## Purpose
Product researchers, ResearchOps teams, and PMs running discovery need method rigor and an insight repository they can trust. This skill structures three decisions:
Three deterministic tools:
1. `study_designer.py` — Maps (research goal × product stage) to an appropriate method and emits a method-matched plan skeleton (objective, participant criteria, guide structure, success criteria). Redirects live A/B to `product-team/experiment-designer`.
2. `saturation_planner.py` — Method-based sample guidance with an explicit **confidence label**: Nielsen problem-discovery (5/segment), Guest et al. thematic saturation (~12), and evaluative coverage. Never claims a prevalence rate from a small-n usability test.
3. `insight_synthesizer.py` — Clusters coded observations by tag, counts distinct participants, ranks by cross-participant recurrence, and flags any candidate below the source threshold as an **ANECDOTE**, never promoting it to an insight.
## When to use
Invoke this skill when:
- You are planning a study and need the method to match the goal (generative vs evaluative vs validation).
- You need a defensible sample size / saturation rationale with a stated confidence.
- You have raw coded observations and need to synthesize insights without over-claiming.
- You are setting up or auditing a research repository and need the insight-vs-observation discipline.
**Do NOT use this skill to**: generate personas / journey maps (use `product-team/ux-researcher-designer`), plan a discovery sprint or validate an opportunity (use `product-team/product-discovery`), design or analyze a live product A/B experiment (use `product-team/experiment-designer`), or do market sizing / surveys (use the `market-research` sibling).
## Workflow
1. **Frame the study** — Fill `assets/research_plan_template.md` (research questions, method rationale, participant criteria, analysis plan, repository tagging scheme).
2. **Pick the method** — Run `study_designer.py --goal {discovery|evaluative|validation} --stage {concept|prototype|beta|live} --profile {b2b-saas|consumer-app|enterprise|marketplace|hardware|platform}`. Honor the redirect if it routes to experiment-designer.
3. **Size it** — Run `saturation_planner.py --method {usability|thematic|evaluative-coverage} --segments N`. Record the confidence label and limits.
4. **Synthesize** — After fielding, code observations and run `insight_synthesizer.py --input observations.json --min-sources 3`. Treat ANECDOTE-flagged clusters as signals to probe, not findings to ship.
5. **File in the repository** — Tag insights to the atomic schema at synthesis time, with their evidence and confidence.
## Scripts
| Script | Purpose | Profiles |
|---|---|---|
| `scripts/study_designer.py` | (goal × stage) → method + plan skeleton | b2b-saas, consumer-app, enterprise, marketplace, hardware, platform |
| `scripts/saturation_planner.py` | Method-based sample guidance + confidence | n/a (method-driven) |
| `scripts/insight_synthesizer.py` | Cluster observations, flag anecdotes | n/a (evidence-driven) |
All three: stdlib-only, `--help`, `--sample`, `--output {human,json}`.
## Onboarding & customization
Run the onboarding questionnaire **once before you start** — it captures your defaults so every tool in this skill is pre-configured. Customization is the point: the answers actually change tool behavior (e.g. the insight source-threshold).
```bash
python3 scripts/onboard.py # interactive (also: --defaults, --set key=value, --reset)
python3 scripts/onboard.py --show # see the questions + current effective config
```
Answers are saved to `~/.config/research-ops/product-research.json` (global) or `./.research-ops/product-research.json` (`--scope project`) and are read automatically by `config_loader.py`. They set the default product **profile**, the **insight source-threshold** (how many independent participants make a finding an insight, not an anecdote), the default **saturation method**, and the **high-stakes** flag. CLI flags always override saved config; `RESEARCH_OPS_NO_CONFIG=1` ignores it.
**The four questions:** product profile · insight source-threshold · saturation method · high-stakes flag.
## Optimize with autoresearch (opt-in)
This skill ships an **isolated, opt-in** bridge to `engineering/autoresearch-agent`. Only when you ask to "optimize the synthesis" / "run a loop" does an autoresearch experiment iteratively refine the coding/clustering of a fixed evidence set so more cross-participant patterns surface. `scripts/ar_evaluator.py` is the ground-truth evaluator; it prints `validated_insights: <int>` (higher is better). It optimizes the **coding**, never fabricates evidence.
```bash
/ar:setup --domain custom --name insight-synthesis \
--target observations.json \
--eval "python3 ar_evaluator.py --target observations.json" \
--metric validated_insights --direction higher
/ar:loop custom/insight-synthesis
```
Isolated: no hard dependency — autoresearch runs only on demand, and the loop edits `observations.json`, never the evaluator.
## References
- `references/research_methods_canon.md` — Portigal *Interviewing Users*; Christensen/Ulwick JTBD; Rohrer's UX-research methods landscape (NN/g); Sauro & Lewis *Quantifying the User Experience*; Goodman/Kuniavsky.
- `references/sampling_and_saturation.md` — Nielsen "test with 5 users"; Guest, Bunce & Johnson saturation; Faulkner on more-than-5; Sauro usability sample size; Braun & Clarke thematic analysis.
- `references/repository_and_synthesis.md` — ResearchOps / atomic research (Tomer Sharon "Polaris"); insight-vs-observation discipline; repository governance; affinity mapping; democratization guardrails.
## Assumptions
- Method selection assumes you can name the goal honestly; if the goal is fuzzy, grill it first (the goal drives everything).
- Saturation guidance is method-based, not a power calculation — usability tests find problems, not prevalence rates.
- The synthesizer counts evidence you provide; coding quality is upstream of it. Garbage tags → garbage clusters.
- The insight threshold (`--min-sources`) defaults to 3; raise it for high-stakes or heterogeneous populations.
## Anti-patterns
- **Mismatching method to goal.** A usability test cannot discover unmet needs; an interview cannot measure task success.
- **Reporting usability problems as percentages.** Small-n tests surface problems, not population rates.
- **Promoting an anecdote to an insight.** One participant is a signal to probe, not a finding.
- **Framing interview questions as feature reactions.** Probe the job-to-be-done and recent real behavior, not hypothetical opinions.
- **Synthesizing without a repository scheme.** Tag at synthesis time, or insights rot unfindable.
## Distinct from
| Neighbor | Scope | Difference |
|---|---|---|
| `product-team/ux-researcher-designer` | Personas, journey maps, usability frameworks tied to design output | That produces **artifacts**; this is **method + repository discipline** |
| `product-team/product-discovery` | Opportunity validation, discovery-sprint planning | That plans **discovery sprints**; this designs and synthesizes the **research** |
| `product-team/experiment-designer` | Live product A/B hypothesis + sample size | That runs **live experiments**; this runs **qualitative/evaluative research** |
| `market-research` (sibling) | Market sizing, surveys, segmentation | That studies **the market**; this studies **users** |
## Quick examples
```bash
python3 scripts/study_designer.py --sample
python3 scripts/saturation_planner.py --method thematic --segments 3
python3 scripts/insight_synthesizer.py --sample --min-sources 3
```
The synthesizer sample correctly promotes "import-confusion" (3 independent participants) to INSIGHT and flags "wants-slack" (1 participant) as an ANECDOTE.
## Forcing-question library (Matt Pocock grill discipline)
Walked one at a time by `/cs:grill-research-ops` or the orchestrator. Recommended answer + canon citation per question. Never bundled.
1. **"Is this study generative (discover problems) or evaluative (test a solution)?"**
Recommended: name it first — the method follows from the goal.
Canon: Rohrer, *When to Use Which User-Experience Research Methods* (NN/g).
2. **"What's your sample size and saturation rationale — and at what confidence?"**
Recommended: method-based n (5/segment usability; ~12 for thematic saturation), state the confidence.
Canon: Nielsen; Guest, Bunce & Johnson (2006); Faulkner (2003).
3. **"How many independent participants support each insight — or is it a single-source anecdote?"**
Recommended: require recurrence across ≥3 sources before calling it an insight; flag singletons.
Canon: atomic research / ResearchOps; Braun & Clarke thematic analysis.
4. **"Are your interview / usability tasks framed as outcomes (jobs) or as feature reactions?"**
Recommended: frame around the job-to-be-done and recent real behavior, not hypothetical opinion.
Canon: Christensen/Ulwick Jobs-to-be-Done; Portigal *Interviewing Users*.
5. **"Where does this land in the repository, and how is it tagged for reuse?"**
Recommended: tag to the atomic schema at synthesis time, not later.
Canon: Tomer Sharon, *Polaris* / ResearchOps repository practice.
Walk depth-first. Lock 1-2 before opening 3-5. After all are answered, invoke `study_designer.py` → `saturation_planner.py` → (after fielding) `insight_synthesizer.py`.
FILE:assets/research_plan_template.md
# Product Research Plan — Template
> Fill this before running the tools. Method must match the goal. An insight requires
> recurrence across independent participants — a single quote is an anecdote.
## 1. Study identification
- Study name:
- Product / feature:
- Stage: [concept | prototype | beta | live]
- Profile: [b2b-saas | consumer-app | enterprise | marketplace | hardware | platform]
## 2. Goal & questions
- Goal: [discovery (generative) | evaluative | validation]
- Research questions (3-5, answerable, not leading):
- The product decision this informs:
## 3. Method (from `study_designer.py`)
- Recommended method:
- Why it matches the goal:
- (If live A/B → route to product-team/experiment-designer.)
## 4. Participants
- Target segment(s) + screener (screen for the job, not a job title):
- Per-segment recruiting if reporting per segment? [yes/no]
- Exclusions (internal, biased, repeat):
## 5. Sample & saturation (from `saturation_planner.py`)
- Method: [usability | thematic | evaluative-coverage]
- n per segment + total:
- Confidence label + limits:
## 6. Study guide skeleton
1.
2.
3.
4.
5.
## 7. Analysis & synthesis
- Coding / tagging scheme (atomic taxonomy):
- Insight threshold (min distinct participants): ___ (default 3)
- Synthesis tool: `insight_synthesizer.py`
## 8. Repository
- Where insights are filed + tagging taxonomy:
- Evidence linked to each insight? [yes — required]
- Confidence field per insight? [yes — required]
## 9. Confidence statement
- What this study can and cannot support:
FILE:references/repository_and_synthesis.md
# Research Repository and Synthesis
Reference for turning observations into governed insights. Pairs with `insight_synthesizer.py`.
## Observation vs insight
The foundational discipline of ResearchOps is the distinction between an **observation** (a single piece of evidence — one participant did or said one thing) and an **insight** (a pattern that recurs across independent sources and carries an implication). Promoting an observation to an insight because it was vivid or confirmed a prior is the cardinal sin of synthesis. The synthesizer enforces a source threshold: a candidate supported by fewer than the threshold of distinct participants is labeled an ANECDOTE and is never promoted.
## Atomic research
Tomer Sharon's **atomic research** model (and the "Polaris" repository concept) decomposes research into reusable units: *Experiments → Facts (observations) → Insights → Recommendations*. Facts are tagged and stored so that insights can be traced back to evidence and reused across studies. The payoff is a repository where a claim can always be drilled down to the observations that support it — and where the same evidence can support future questions.
## Affinity mapping
The classic synthesis technique is affinity mapping: cluster observations into emergent themes bottom-up, then name the themes. The `insight_synthesizer.py` tool is a deterministic, tag-based proxy for this — it clusters by the codes you assign and ranks by cross-participant recurrence. The human still does the interpretive naming; the tool enforces the counting discipline.
## Repository governance and democratization
As organizations democratize research (PMs and designers running their own studies), the repository becomes the guardrail. Governance practices: a consistent tagging taxonomy, evidence linked to every insight, a confidence field, and a review step before an insight is marked "validated." Without governance, democratized research produces a pile of unsearchable anecdotes; with it, the repository compounds in value.
## Sources
1. Sharon, T., *Validating Product Ideas Through Lean User Research* (Rosenfeld, 2016) and the atomic-research / Polaris model.
2. ResearchOps Community, *Research Repositories* and *Democratization* working-group reports.
3. Braun, V., & Clarke, V., *Thematic Analysis: A Practical Guide* (Sage, 2022).
4. Beyer, H., & Holtzblatt, K., *Contextual Design* (1998) — affinity diagramming.
5. Dovetail / EnjoyHQ practitioner guides on insight repositories and tagging taxonomies.
6. Kaplan, K., *Taxonomy 101* and *Research Repositories* — Nielsen Norman Group.
FILE:references/research_methods_canon.md
# Product Research Methods Canon
Reference for method selection. Pairs with `study_designer.py`.
## The two-axis map
UX/product research methods sort along two axes (Rohrer, NN/g): **attitudinal vs behavioral** (what people say vs what they do) and **qualitative vs quantitative** (why/how vs how-many). The single most important pre-method decision is the **goal**:
- **Generative (discovery)** — you don't yet know the problem. Methods: semi-structured interviews, contextual inquiry, diary studies. Output: themes, unmet needs, jobs-to-be-done.
- **Evaluative** — you have a solution and want to know if it works. Methods: moderated/unmoderated usability tests, concept tests. Output: task-success, severity-rated problems.
- **Validation** — you want to confirm demand/desirability before building. Methods: surveys, preference tests, fake-door tests, and (when live) A/B experiments.
Picking an evaluative method for a generative goal — "let's usability-test our way to product strategy" — is the most common and most expensive error.
## Interviewing discipline
Steve Portigal's *Interviewing Users* is the operative craft reference: ask about **recent, specific, real behavior** ("tell me about the last time you…"), not hypotheticals or opinions ("would you use…"). People are unreliable narrators of their future selves but good storytellers of their past.
## Jobs-to-be-Done
Christensen's and Ulwick's JTBD reframes research around the **progress a person is trying to make** in a circumstance, not their demographics or feature preferences. Outcome-Driven Innovation (Ulwick) operationalizes this into measurable desired outcomes — a bridge between qualitative discovery and quantitative validation.
## Mixed methods
Strong research triangulates: qualitative discovery surfaces hypotheses; quantitative validation sizes them. Sauro & Lewis (*Quantifying the User Experience*) provides the statistical backbone for turning usability observations into defensible metrics (task time, completion, SUS) without over-claiming from small samples.
## Sources
1. Portigal, S., *Interviewing Users*, 2nd ed. (Rosenfeld, 2023).
2. Christensen, Hall, Dillon & Duncan, *Competing Against Luck* (2016) — Jobs-to-be-Done.
3. Ulwick, A., *Jobs to Be Done: Theory to Practice* (2016) — Outcome-Driven Innovation.
4. Rohrer, C., *When to Use Which User-Experience Research Methods* — Nielsen Norman Group.
5. Sauro, J., & Lewis, J., *Quantifying the User Experience*, 2nd ed. (Morgan Kaufmann, 2016).
6. Goodman, Kuniavsky & Moed, *Observing the User Experience*, 2nd ed. (2012).
FILE:references/sampling_and_saturation.md
# Sampling and Saturation
Reference for how many participants. Pairs with `saturation_planner.py`.
## Usability: the "5 users" result
Nielsen and Landauer's model says the proportion of usability problems found with n users is 1 − (1 − p)ⁿ, where p is the average probability that a single user surfaces a given problem (~0.31 in their data). At n = 5, that is ~85% of problems — hence "test with 5 users." Two crucial caveats the planner enforces:
1. **Per segment.** The 5-user result holds *within a homogeneous user group*. If you have distinct segments that behave differently, you need ~5 per segment.
2. **Problems, not rates.** A small-n usability test finds *whether* a problem exists; it cannot estimate the *prevalence* of that problem in the population. Never report "60% of users struggled" from a 5-person test.
Faulkner (2003) showed real variance: while the average across many 5-person samples is ~85%, individual 5-person runs ranged from ~55% to 100%. When stakes or heterogeneity are high, run more.
## Qualitative: thematic saturation
For interview-based thematic research, Guest, Bunce & Johnson (2006) found that **saturation** — the point where new interviews stop yielding new themes — typically occurs by ~12 interviews in a homogeneous group, with the basic elements present by ~6. Saturation is **observed, not guaranteed**: track the new-theme rate and stop when it flattens, rather than committing to a fixed n blindly. Heterogeneous populations need more, and per-group saturation applies just as in usability.
## Reporting confidence honestly
The planner attaches a confidence label (LOW / MODERATE / MODERATE-HIGH) and explicit limits to every plan, because the failure mode in product research is not too-small samples per se — it is **over-claiming** from whatever sample you ran. State the method, the n, and what the method can and cannot support.
## Sources
1. Nielsen, J., & Landauer, T., *A mathematical model of the finding of usability problems* — INTERCHI 1993.
2. Nielsen, J., *Why You Only Need to Test with 5 Users* — NN/g (2000).
3. Faulkner, L., *Beyond the five-user assumption* — Behavior Research Methods 2003;35:379-383.
4. Guest, G., Bunce, A., & Johnson, L., *How many interviews are enough?* — Field Methods 2006;18:59-82.
5. Braun, V., & Clarke, V., *Using thematic analysis in psychology* — Qual Res Psychol 2006;3:77-101.
6. Sauro, J., & Lewis, J., *Quantifying the User Experience*, 2nd ed. (2016) — confidence intervals for small samples.
FILE:scripts/ar_evaluator.py
#!/usr/bin/env python3
"""ar_evaluator.py - Autoresearch evaluator for the product-research skill (OPT-IN).
Stdlib-only. The ISOLATED bridge to engineering/autoresearch-agent. It does NOT call
autoresearch; it is the ground-truth evaluator an autoresearch loop runs after editing
the target coded-observations file. It reads an observations JSON, runs insight_synthesizer
at the configured source threshold, and prints ONE metric line:
validated_insights: <int> (higher is better — clusters that clear the source threshold)
This optimizes the CODING/synthesis of a fixed evidence set (merging/splitting tags so
cross-participant patterns surface) — not the evidence itself. The user opts in explicitly:
/ar:setup --domain custom --name insight-synthesis \\
--target observations.json --eval "python3 ar_evaluator.py --target observations.json" \\
--metric validated_insights --direction higher
Direct use:
python3 ar_evaluator.py --sample
python3 ar_evaluator.py --target observations.json --min-sources 3
"""
from __future__ import annotations
import argparse
import json
import os
import sys
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import config_loader as cfg # noqa: E402
import insight_synthesizer as isyn # noqa: E402
METRIC = "validated_insights"
def main(argv: list[str] | None = None) -> int:
c = cfg.load_config()
p = argparse.ArgumentParser(description="Autoresearch evaluator: count of validated insights.")
p.add_argument("--target", help="path to observations JSON (or env AR_TARGET)")
p.add_argument("--min-sources", type=int, default=None, help="overrides onboarding insight_min_sources")
p.add_argument("--sample", action="store_true")
args = p.parse_args(argv)
min_sources = args.min_sources if args.min_sources is not None else int(c.get("insight_min_sources", 3))
if args.sample:
data = isyn.SAMPLE
else:
target = args.target or os.environ.get("AR_TARGET")
if not target:
print("error: provide --target <observations.json> or set AR_TARGET", file=sys.stderr)
return 2
try:
with open(target) as f:
data = json.load(f)
except (OSError, json.JSONDecodeError) as e:
print(f"{METRIC}: N/A")
print(f"error: {e}", file=sys.stderr)
return 1
result = isyn.synthesize(data, min_sources)
count = sum(1 for c2 in result["candidates"] if c2["classification"] == "INSIGHT")
print(f"{METRIC}: {count}")
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/config_loader.py
#!/usr/bin/env python3
"""config_loader.py - Customization loader for the product-research skill.
Stdlib-only. Importable from the skill's other scripts. Precedence (highest wins):
1. Project config: <cwd>/.research-ops/product-research.json
2. Global config: ~/.config/research-ops/product-research.json
3. Built-in DEFAULTS
Onboarding answers (written by onboard.py) live in these files; every tool in this
skill reads them so the user's customization applies automatically.
Set RESEARCH_OPS_NO_CONFIG=1 to ignore saved config.
"""
from __future__ import annotations
import argparse
import json
import os
import sys
from pathlib import Path
from typing import Any
SKILL = "product-research"
GLOBAL_CONFIG_DIR = Path.home() / ".config" / "research-ops"
GLOBAL_CONFIG_PATH = GLOBAL_CONFIG_DIR / f"{SKILL}.json"
PROJECT_CONFIG_DIRNAME = ".research-ops"
DEFAULTS: dict[str, Any] = {
"version": 1,
"skill": SKILL,
"default_profile": "b2b-saas",
"insight_min_sources": 3,
"default_method": "usability",
"stakes_high": False,
"setup_completed_at": None,
}
def project_config_path(cwd: Path | None = None) -> Path:
cwd = cwd or Path.cwd()
return cwd / PROJECT_CONFIG_DIRNAME / f"{SKILL}.json"
def _read_json(path: Path) -> dict[str, Any] | None:
try:
with path.open(encoding="utf-8") as f:
data = json.load(f)
return data if isinstance(data, dict) else None
except (FileNotFoundError, json.JSONDecodeError, OSError):
return None
def _deep_merge(base: dict[str, Any], override: dict[str, Any]) -> dict[str, Any]:
out = dict(base)
for k, v in override.items():
if isinstance(v, dict) and isinstance(out.get(k), dict):
out[k] = _deep_merge(out[k], v)
else:
out[k] = v
return out
def load_config(cwd: Path | None = None) -> dict[str, Any]:
config = dict(DEFAULTS)
if os.environ.get("RESEARCH_OPS_NO_CONFIG") == "1":
return config
global_cfg = _read_json(GLOBAL_CONFIG_PATH)
if global_cfg:
config = _deep_merge(config, global_cfg)
project_cfg = _read_json(project_config_path(cwd))
if project_cfg:
config = _deep_merge(config, project_cfg)
return config
def setup_completed() -> bool:
cfg = _read_json(GLOBAL_CONFIG_PATH) or _read_json(project_config_path())
return bool(cfg and cfg.get("setup_completed_at"))
def write_config(config: dict[str, Any], scope: str = "global", cwd: Path | None = None) -> Path:
path = project_config_path(cwd) if scope == "project" else GLOBAL_CONFIG_PATH
path.parent.mkdir(parents=True, exist_ok=True)
with path.open("w", encoding="utf-8") as f:
json.dump(config, f, indent=2, sort_keys=True)
return path
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description=f"Inspect {SKILL} customization config.")
p.add_argument("--show", action="store_true", help="Print the effective config")
p.add_argument("--status", action="store_true", help="Print setup status + paths")
p.add_argument("--sample", action="store_true", help="Print the built-in defaults")
args = p.parse_args(argv)
if args.sample:
print(json.dumps(DEFAULTS, indent=2, sort_keys=True))
elif args.status:
print(json.dumps({
"skill": SKILL,
"global_config_path": str(GLOBAL_CONFIG_PATH),
"global_config_exists": GLOBAL_CONFIG_PATH.exists(),
"project_config_path": str(project_config_path()),
"project_config_exists": project_config_path().exists(),
"setup_completed": setup_completed(),
}, indent=2))
else:
print(json.dumps(load_config(), indent=2, sort_keys=True))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/insight_synthesizer.py
#!/usr/bin/env python3
"""insight_synthesizer.py - Cluster coded observations into candidate insights; flag anecdotes.
Stdlib-only. Deterministic. NO LLM calls. NEVER fabricates an insight: it counts evidence,
clusters by tag, ranks by cross-participant recurrence, and flags any candidate supported by
fewer than --min-sources independent participants as an ANECDOTE, not an insight.
Input: a list of observations, each with {participant, tag, note}. The synthesizer groups by
tag, counts distinct participants per tag, and ranks. This is the atomic-research discipline:
an observation is evidence; an insight requires recurrence across independent sources.
Usage:
python3 insight_synthesizer.py --sample
python3 insight_synthesizer.py --input observations.json --min-sources 3
python3 insight_synthesizer.py --input observations.json --output json
"""
from __future__ import annotations
import argparse
import json
import os
import sys
from collections import defaultdict
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
try:
import config_loader as _cfg
except ImportError: # pragma: no cover
_cfg = None
SAMPLE = {
"study": "Onboarding discovery (mid-market HR)",
"observations": [
{"participant": "P1", "tag": "import-confusion", "note": "Couldn't find CSV import."},
{"participant": "P2", "tag": "import-confusion", "note": "Expected import on the dashboard."},
{"participant": "P3", "tag": "import-confusion", "note": "Gave up looking for bulk upload."},
{"participant": "P1", "tag": "permissions-unclear", "note": "Unsure who could see reports."},
{"participant": "P4", "tag": "permissions-unclear", "note": "Worried about data visibility."},
{"participant": "P2", "tag": "wants-slack", "note": "Asked for a Slack integration."},
],
}
def synthesize(data: dict, min_sources: int) -> dict:
obs = data.get("observations", [])
by_tag_participants = defaultdict(set)
by_tag_notes = defaultdict(list)
for o in obs:
tag = o.get("tag", "untagged")
part = o.get("participant", "UNKNOWN")
by_tag_participants[tag].add(part)
by_tag_notes[tag].append({"participant": part, "note": o.get("note", "")})
candidates = []
for tag, parts in by_tag_participants.items():
n_sources = len(parts)
is_insight = n_sources >= min_sources
candidates.append({
"tag": tag,
"distinct_participants": n_sources,
"observation_count": len(by_tag_notes[tag]),
"classification": "INSIGHT" if is_insight else "ANECDOTE (single/low-source — do not generalize)",
"evidence": by_tag_notes[tag],
})
candidates.sort(key=lambda c: (c["distinct_participants"], c["observation_count"]), reverse=True)
total_participants = len({o.get("participant") for o in obs})
return {
"study": data.get("study", "UNSPECIFIED"),
"min_sources_for_insight": min_sources,
"total_participants": total_participants,
"candidates": candidates,
"note": "An observation is evidence; an insight requires recurrence across independent participants. "
"Anecdotes are surfaced, never promoted to insights.",
}
def _render_human(r: dict) -> str:
lines = [f"Insight Synthesis: {r['study']}",
f" total participants: {r['total_participants']} insight threshold: >= {r['min_sources_for_insight']} sources", ""]
for c in r["candidates"]:
lines.append(f"[{c['classification']}] {c['tag']} "
f"({c['distinct_participants']} participants, {c['observation_count']} observations)")
for e in c["evidence"]:
lines.append(f" {e['participant']}: {e['note']}")
lines.append("")
lines.append(f"note: {r['note']}")
return "\n".join(lines)
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description="Cluster coded observations into insights; flag anecdotes.")
p.add_argument("--input", help="Path to JSON with observations[]")
p.add_argument("--min-sources", type=int, default=None,
help="min distinct participants to call it an insight (overrides onboarding)")
p.add_argument("--output", choices=["human", "json"], default="human")
p.add_argument("--sample", action="store_true", help="use the embedded sample")
args = p.parse_args(argv)
conf = _cfg.load_config() if _cfg else {}
min_sources = args.min_sources if args.min_sources is not None else int(conf.get("insight_min_sources", 3))
data = SAMPLE if (args.sample or not args.input) else json.load(open(args.input))
result = synthesize(data, min_sources)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(_render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/onboard.py
#!/usr/bin/env python3
"""onboard.py - Onboarding questionnaire for the product-research skill.
Stdlib-only. Asks the user a short set of questions BEFORE they plan a study, then
writes the answers to a customization config read by every tool in this skill via
config_loader.py. The answers become defaults for profile, the insight source-threshold,
the default saturation method, and the high-stakes flag.
Modes: --show | --defaults | --set key=value (repeatable) | --reset | --scope {global,project}
"""
from __future__ import annotations
import argparse
import datetime as _dt
import json
import os
import sys
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import config_loader as cfg # noqa: E402
INT_KEYS = {"insight_min_sources"}
BOOL_KEYS = {"stakes_high"}
QUESTIONS = [
("default_profile",
"1. What kind of product is this?",
["b2b-saas", "consumer-app", "enterprise", "marketplace", "hardware", "platform"], str),
("insight_min_sources",
"2. How many independent participants must support a finding before it counts as an insight (not an anecdote)?",
None, int),
("default_method",
"3. Default sample-saturation method?",
["usability", "thematic", "evaluative-coverage"], str),
("stakes_high",
"4. Is this high-stakes / high-heterogeneity research (raise sample sizes)?",
["true", "false"], str),
]
def _coerce(key: str, value: str):
if key in INT_KEYS:
return int(value)
if key in BOOL_KEYS:
return str(value).strip().lower() in ("true", "yes", "y", "1")
return value
def _print_questions() -> None:
print(f"Onboarding questions — {cfg.SKILL}:\n")
for _k, prompt, choices, _c in QUESTIONS:
line = f" {prompt}"
if choices:
line += f" [{' / '.join(choices)}]"
print(line)
def run_interactive(config: dict) -> dict:
print(f"Onboarding — {cfg.SKILL}. Press Enter to keep the current/default value.\n")
for key, prompt, choices, _caster in QUESTIONS:
suffix = f" [{'/'.join(choices)}]" if choices else ""
cur = f" (current: {config.get(key)})" if config.get(key) is not None else ""
raw = input(f"{prompt}{suffix}{cur}: ").strip()
if not raw:
continue
try:
config[key] = _coerce(key, raw)
except ValueError:
print(f" ! invalid value for {key}, keeping current")
return config
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description=f"Onboarding for the {cfg.SKILL} skill.")
p.add_argument("--show", action="store_true")
p.add_argument("--defaults", action="store_true", help="write built-in defaults, no prompt")
p.add_argument("--set", action="append", default=[], metavar="key=value")
p.add_argument("--reset", action="store_true")
p.add_argument("--scope", choices=["global", "project"], default="global")
args = p.parse_args(argv)
if args.show:
_print_questions()
print("\nCurrent effective config:")
print(json.dumps(cfg.load_config(), indent=2, sort_keys=True))
return 0
if args.reset:
path = cfg.project_config_path() if args.scope == "project" else cfg.GLOBAL_CONFIG_PATH
if path.exists():
path.unlink(); print(f"removed {path}")
else:
print(f"no config at {path}")
return 0
config = cfg.load_config()
if args.set:
for item in args.set:
if "=" not in item:
print(f"error: --set expects key=value, got '{item}'", file=sys.stderr)
return 2
k, v = item.split("=", 1)
try:
config[k] = _coerce(k, v)
except ValueError:
config[k] = v
elif not args.defaults:
if sys.stdin.isatty():
config = run_interactive(config)
else:
print("non-interactive shell: use --defaults or --set key=value. Showing questions:\n")
_print_questions()
return 0
config["setup_completed_at"] = _dt.datetime.now(_dt.timezone.utc).isoformat()
path = cfg.write_config(config, scope=args.scope)
print(f"saved {cfg.SKILL} customization -> {path}")
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/saturation_planner.py
#!/usr/bin/env python3
"""saturation_planner.py - Method-based participant/sample guidance with a confidence label.
Stdlib-only. Deterministic. NO LLM calls. NEVER fabricates insight: it gives method-based
sample guidance and an explicit confidence level, surfacing limits.
Models:
- usability (Nielsen): ~5 users per segment uncovers ~85% of problems at typical p=0.31;
problems found = 1 - (1 - p)^n.
- thematic saturation (Guest et al.): ~12 interviews per homogeneous group typically
reaches saturation; >5 (Faulkner) when stakes/heterogeneity are high.
- evaluative coverage: detectable-problem coverage for a chosen per-problem detection rate.
Usage:
python3 saturation_planner.py --sample
python3 saturation_planner.py --method usability --segments 2 --detection-rate 0.31
python3 saturation_planner.py --method thematic --segments 3 --output json
"""
from __future__ import annotations
import argparse
import json
import os
import sys
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
try:
import config_loader as _cfg
except ImportError: # pragma: no cover
_cfg = None
METHODS = ["usability", "thematic", "evaluative-coverage"]
def usability_plan(segments: int, p: float, target_coverage: float) -> dict:
# n per segment to reach target coverage: n = ln(1 - target) / ln(1 - p)
import math
if not 0.0 < p < 1.0:
raise ValueError("detection-rate must be in (0,1).")
n = math.ceil(math.log(1 - target_coverage) / math.log(1 - p))
coverage_at_5 = 1 - (1 - p) ** 5
return {
"method": "usability",
"per_problem_detection_rate": p,
"target_coverage": target_coverage,
"n_per_segment": n,
"segments": segments,
"total_participants": n * segments,
"coverage_at_5_per_segment": round(coverage_at_5, 3),
"confidence": "MODERATE" if n >= 5 else "LOW (small-n usability finds problems, not rates)",
"limits": "Usability tests surface problems, not their population prevalence. Do not report percentages.",
}
def thematic_plan(segments: int, stakes_high: bool) -> dict:
base = 12 # Guest et al. typical saturation for a homogeneous group
per_segment = base if not stakes_high else max(base, 15)
return {
"method": "thematic",
"n_per_segment": per_segment,
"segments": segments,
"total_participants": per_segment * segments,
"confidence": "MODERATE-HIGH" if per_segment >= 12 else "LOW",
"limits": "Saturation is observed, not guaranteed; track new-theme rate and stop when it flattens. "
"Faulkner (2003): more than 5 when heterogeneity or stakes are high.",
}
def evaluative_coverage_plan(segments: int, n_per_segment: int, p: float) -> dict:
coverage = 1 - (1 - p) ** n_per_segment
return {
"method": "evaluative-coverage",
"per_problem_detection_rate": p,
"n_per_segment": n_per_segment,
"segments": segments,
"expected_problem_coverage": round(coverage, 3),
"confidence": "MODERATE" if coverage >= 0.8 else "LOW",
"limits": "Coverage is for the assumed detection rate; rarer problems need more participants.",
}
def plan(method: str, segments: int, p: float, target: float, stakes_high: bool, n: int) -> dict:
if method == "usability":
out = usability_plan(segments, p, target)
elif method == "thematic":
out = thematic_plan(segments, stakes_high)
elif method == "evaluative-coverage":
out = evaluative_coverage_plan(segments, n, p)
else:
raise ValueError(f"method must be one of {METHODS}.")
out["disclaimer"] = "Method-based guidance with explicit confidence. This is not a power calculation; " \
"it never claims an insight the data cannot support."
return out
def _render_human(r: dict) -> str:
lines = [f"Saturation / Sample Plan (method: {r['method']})", ""]
for k, v in r.items():
if k in ("method",):
continue
lines.append(f" {k:32s} : {v}")
return "\n".join(lines)
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description="Method-based product-research sample guidance with confidence.")
p.add_argument("--method", choices=METHODS, default=None, help="overrides onboarding default_method")
p.add_argument("--segments", type=int, default=1)
p.add_argument("--detection-rate", type=float, default=0.31, help="per-problem detection rate (usability)")
p.add_argument("--target-coverage", type=float, default=0.85, help="target problem coverage (usability)")
p.add_argument("--stakes-high", action="store_true", help="raise thematic n for high heterogeneity/stakes")
p.add_argument("--n-per-segment", type=int, default=8, help="n per segment (evaluative-coverage)")
p.add_argument("--output", choices=["human", "json"], default="human")
p.add_argument("--sample", action="store_true", help="use the embedded sample")
args = p.parse_args(argv)
conf = _cfg.load_config() if _cfg else {}
method = args.method or conf.get("default_method", "usability")
stakes_high = args.stakes_high or bool(conf.get("stakes_high", False))
if args.sample:
try:
result = plan("usability", 2, 0.31, 0.85, False, 8)
except ValueError as e:
print(f"error: {e}", file=sys.stderr)
return 2
else:
try:
result = plan(method, args.segments, args.detection_rate,
args.target_coverage, stakes_high, args.n_per_segment)
except ValueError as e:
print(f"error: {e}", file=sys.stderr)
return 2
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(_render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/study_designer.py
#!/usr/bin/env python3
"""study_designer.py - Select a product-research method from goal + stage, emit a plan skeleton.
Stdlib-only. Deterministic. NO LLM calls.
Maps (research goal x product stage) to an appropriate method and emits a method-matched
plan skeleton (objective framing, participant criteria, task/guide structure, success
criteria). The core discipline: GENERATIVE goals (discover problems) and EVALUATIVE goals
(test a solution) demand different methods — picking the wrong one is the most common error.
Usage:
python3 study_designer.py --sample
python3 study_designer.py --goal discovery --stage concept --profile b2b-saas
python3 study_designer.py --goal evaluative --stage live --output json
"""
from __future__ import annotations
import argparse
import json
import os
import sys
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
try:
import config_loader as _cfg
except ImportError: # pragma: no cover
_cfg = None
PROFILES = ["b2b-saas", "consumer-app", "enterprise", "marketplace", "hardware", "platform"]
# (goal, stage) -> method. goal in {discovery, evaluative, validation}; stage in {concept, prototype, beta, live}
METHOD_MAP = {
("discovery", "concept"): "generative interviews (semi-structured)",
("discovery", "prototype"): "contextual inquiry",
("discovery", "beta"): "diary study + follow-up interviews",
("discovery", "live"): "behavioral analytics review + generative interviews",
("evaluative", "concept"): "concept test (comprehension + desirability)",
("evaluative", "prototype"): "moderated usability test",
("evaluative", "beta"): "unmoderated usability test + task-success metrics",
("evaluative", "live"): "benchmark usability study (SUS / task time)",
("validation", "concept"): "survey (desirability + willingness signals)",
("validation", "prototype"): "prototype A/B preference test",
("validation", "beta"): "fake-door / feature-demand test",
("validation", "live"): "live A/B experiment (route to product-team/experiment-designer)",
}
GUIDE_SKELETONS = {
"generative": ["Warm-up + context", "Recent relevant experience (story, not opinion)",
"Workarounds + frustrations", "Jobs-to-be-done probe", "Magic-wand / wrap"],
"evaluative": ["Pre-task context", "Task 1 (representative)", "Task 2 (edge)",
"Observation: where do they hesitate/err?", "Post-task SUS / debrief"],
"validation": ["Screener", "Stimulus exposure", "Comprehension + desirability items",
"Trade-off / preference items", "Behavioral-intent item"],
}
def design(goal: str, stage: str, profile: str) -> dict:
if profile not in PROFILES:
raise ValueError(f"Unknown profile '{profile}'. Choose from {PROFILES}.")
key = (goal, stage)
if key not in METHOD_MAP:
raise ValueError(f"No method for goal={goal}, stage={stage}. "
f"goal in [discovery,evaluative,validation]; stage in [concept,prototype,beta,live].")
method = METHOD_MAP[key]
family = "generative" if goal == "discovery" else ("evaluative" if goal == "evaluative" else "validation")
redirect = None
if "experiment-designer" in method:
redirect = "Live A/B is a product experiment — use product-team/experiment-designer, not this skill."
return {
"goal": goal,
"stage": stage,
"profile": profile,
"method": method,
"method_family": family,
"objective_framing": f"A {family} study at the {stage} stage to {('discover unmet needs' if family=='generative' else 'evaluate the solution' if family=='evaluative' else 'validate demand/desirability')}.",
"participant_criteria": [
"Recruit to the target segment (screen for the job, not a job title).",
"Exclude internal/biased participants and prior-study repeats unless longitudinal.",
"Recruit per-segment if results will be reported per-segment.",
],
"guide_skeleton": GUIDE_SKELETONS[family],
"success_criteria": [
"Generative: themes recur across independent participants (saturation).",
"Evaluative: task-success rate + severity-rated problem list.",
"Validation: pre-registered desirability / preference threshold.",
],
"redirect": redirect,
"note": "Method must match the goal. A usability test cannot discover unmet needs; an interview cannot measure task success.",
}
def _render_human(r: dict) -> str:
lines = [f"Study Design: goal={r['goal']}, stage={r['stage']}, profile={r['profile']}", "",
f" Recommended method: {r['method']} (family: {r['method_family']})",
f" Objective: {r['objective_framing']}", "", " Participant criteria:"]
for c in r["participant_criteria"]:
lines.append(f" - {c}")
lines.append(" Guide skeleton:")
for i, g in enumerate(r["guide_skeleton"], 1):
lines.append(f" {i}. {g}")
lines.append(" Success criteria:")
for s in r["success_criteria"]:
lines.append(f" - {s}")
if r["redirect"]:
lines += ["", f" !! {r['redirect']}"]
lines += ["", f"note: {r['note']}"]
return "\n".join(lines)
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description="Select a product-research method from goal + stage.")
p.add_argument("--goal", choices=["discovery", "evaluative", "validation"], default="discovery")
p.add_argument("--stage", choices=["concept", "prototype", "beta", "live"], default="prototype")
p.add_argument("--profile", default=None, choices=PROFILES,
help="overrides onboarding default_profile")
p.add_argument("--output", choices=["human", "json"], default="human")
p.add_argument("--sample", action="store_true", help="use the embedded sample")
args = p.parse_args(argv)
conf = _cfg.load_config() if _cfg else {}
profile_default = conf.get("default_profile", "b2b-saas")
goal, stage, profile = ("discovery", "prototype", profile_default) if args.sample \
else (args.goal, args.stage, args.profile or profile_default)
try:
result = design(goal, stage, profile)
except ValueError as e:
print(f"error: {e}", file=sys.stderr)
return 2
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(_render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main())
Bộ công cụ chiến lược cho Head of Product: chuỗi OKR, kế hoạch quý, phân tích cảnh quan cạnh tranh, tầm nhìn sản phẩm và đề xuất mở rộng đội ngũ.
---
name: "product-strategist"
description: Strategic product leadership toolkit for Head of Product covering OKR cascade generation, quarterly planning, competitive landscape analysis, product vision documents, and team scaling proposals. Use when creating quarterly OKR documents, defining product goals or KPIs, building product roadmaps, running competitive analysis, drafting team structure or hiring plans, aligning product strategy across engineering and design, or generating cascaded goal hierarchies from company to team level.
---
# Product Strategist
Strategic toolkit for Head of Product to drive vision, alignment, and organizational excellence.
---
## Core Capabilities
| Capability | Description | Tool |
|------------|-------------|------|
| **OKR Cascade** | Generate aligned OKRs from company to team level | `okr_cascade_generator.py` |
| **Alignment Scoring** | Measure vertical and horizontal alignment | Built into generator |
| **Strategy Templates** | 5 pre-built strategy types | Growth, Retention, Revenue, Innovation, Operational |
| **Team Configuration** | Customize for your org structure | `--teams` flag |
---
## Quick Start
```bash
# Growth strategy with default teams
python scripts/okr_cascade_generator.py growth
# Retention strategy with custom teams
python scripts/okr_cascade_generator.py retention --teams "Engineering,Design,Data"
# Revenue strategy with 40% product contribution
python scripts/okr_cascade_generator.py revenue --contribution 0.4
# Export as JSON for integration
python scripts/okr_cascade_generator.py growth --json > okrs.json
```
---
## Workflow: Quarterly Strategic Planning
### Step 1: Define Strategic Focus
| Strategy | When to Use |
|----------|-------------|
| **Growth** | Scaling user base, market expansion |
| **Retention** | Reducing churn, improving LTV |
| **Revenue** | Increasing ARPU, new monetization |
| **Innovation** | Market differentiation, new capabilities |
| **Operational** | Improving efficiency, scaling operations |
See `references/strategy_types.md` for detailed guidance.
### Step 2: Gather Input Metrics
```json
{
"current": 100000, // Current MAU
"target": 150000, // Target MAU
"current_nps": 40, // Current NPS
"target_nps": 60 // Target NPS
}
```
### Step 3: Configure Teams & Run Generator
```bash
# Default teams
python scripts/okr_cascade_generator.py growth
# Custom org structure with contribution percentage
python scripts/okr_cascade_generator.py growth \
--teams "Core,Platform,Mobile,AI" \
--contribution 0.3
```
### Step 4: Review Alignment Scores
| Score | Target | Action if Below |
|-------|--------|-----------------|
| Vertical Alignment | >90% | Ensure all objectives link to parent |
| Horizontal Alignment | >75% | Check for team coordination gaps |
| Coverage | >80% | Validate all company OKRs are addressed |
| Balance | >80% | Redistribute if one team is overloaded |
| **Overall** | **>80%** | <60% needs restructuring |
### Step 5: Refine, Validate, and Export
Before finalizing:
- [ ] Review generated objectives with stakeholders
- [ ] Adjust team assignments based on capacity
- [ ] Validate contribution percentages are realistic
- [ ] Ensure no conflicting objectives across teams
- [ ] Set up tracking cadence (bi-weekly check-ins)
```bash
# Export JSON for tools like Lattice, Ally, Workboard
python scripts/okr_cascade_generator.py growth --json > q1_okrs.json
```
---
## OKR Cascade Generator
### Usage
```bash
python scripts/okr_cascade_generator.py [strategy] [options]
```
**Strategies:** `growth` | `retention` | `revenue` | `innovation` | `operational`
### Configuration Options
| Option | Description | Default |
|--------|-------------|---------|
| `--teams`, `-t` | Comma-separated team names | Growth,Platform,Mobile,Data |
| `--contribution`, `-c` | Product contribution to company OKRs (0-1) | 0.3 (30%) |
| `--json`, `-j` | Output as JSON instead of dashboard | False |
| `--metrics`, `-m` | Metrics as JSON string | Sample metrics |
### Output Examples
#### Dashboard Output (`growth` strategy)
```
============================================================
OKR CASCADE DASHBOARD
Quarter: Q1 2025 | Strategy: GROWTH
Teams: Growth, Platform, Mobile, Data | Product Contribution: 30%
============================================================
🏢 COMPANY OKRS
📌 CO-1: Accelerate user acquisition and market expansion
└─ CO-1-KR1: Increase MAU from 100,000 to 150,000
└─ CO-1-KR2: Achieve 50% MoM growth rate
└─ CO-1-KR3: Expand to 3 new markets
📌 CO-2: Achieve product-market fit in new segments
📌 CO-3: Build sustainable growth engine
🚀 PRODUCT OKRS
📌 PO-1: Build viral product features and market expansion
↳ Supports: CO-1
└─ PO-1-KR1: Increase product MAU to 45,000
└─ PO-1-KR2: Achieve 45% feature adoption rate
👥 TEAM OKRS
Growth Team:
📌 GRO-1: Build viral product features through acquisition and activation
└─ GRO-1-KR1: Increase product MAU to 11,250
└─ GRO-1-KR2: Achieve 11.25% feature adoption rate
🎯 ALIGNMENT SCORES
✓ Vertical Alignment: 100.0%
! Horizontal Alignment: 75.0%
✓ Coverage: 100.0% | ✓ Balance: 97.5% | ✓ Overall: 94.0%
✅ Overall alignment is GOOD (≥80%)
```
#### JSON Output (`retention --json`, truncated)
```json
{
"quarter": "Q1 2025",
"strategy": "retention",
"company": {
"objectives": [
{
"id": "CO-1",
"title": "Create lasting customer value and loyalty",
"key_results": [
{ "id": "CO-1-KR1", "title": "Improve retention from 70% to 85%", "current": 70, "target": 85 }
]
}
]
},
"product": { "contribution": 0.3, "objectives": ["..."] },
"teams": ["..."],
"alignment_scores": {
"vertical_alignment": 100.0, "horizontal_alignment": 75.0,
"coverage": 100.0, "balance": 97.5, "overall": 94.0
}
}
```
See `references/examples/sample_growth_okrs.json` for a complete example.
---
## Reference Documents
| Document | Description |
|----------|-------------|
| `references/okr_framework.md` | OKR methodology, writing guidelines, alignment scoring |
| `references/strategy_types.md` | Detailed breakdown of all 5 strategy types with examples |
| `references/examples/sample_growth_okrs.json` | Complete sample output for growth strategy |
---
## Best Practices
### OKR Cascade
- Limit to 3-5 objectives per level, each with 3-5 key results
- Key results must be measurable with current and target values
- Validate parent-child relationships before finalizing
### Alignment Scoring
- Target >80% overall alignment; investigate any score below 60%
- Balance scores ensure no team is overloaded
- Horizontal alignment prevents conflicting goals across teams
### Team Configuration
- Configure teams to match your actual org structure
- Adjust contribution percentages based on team size
- Platform/Infrastructure teams often support all objectives
- Specialized teams (ML, Data) may only support relevant objectives
## Related Skills
- **Senior PM** (`project-management/senior-pm/`) — Portfolio management and risk analysis inform strategic planning
- **Competitive Teardown** (`product-team/competitive-teardown/`) — Competitive intelligence feeds product strategy
FILE:assets/okr_template.md
# OKR Planning Template
## Planning Info
| Field | Value |
|-------|-------|
| **Quarter** | [Q1/Q2/Q3/Q4 YYYY] |
| **Team** | [Team Name] |
| **Owner** | [Name] |
| **Status** | Planning / Active / Complete |
| **Tracking Cadence** | Weekly check-in, Monthly review |
---
## Company Objective
**[Company-level objective this product work supports]**
_Example: "Become the market leader in our category by delivering exceptional customer value"_
---
## Product Objective 1: [Objective Title]
_[Qualitative, inspirational statement. What does success look like?]_
### Key Results
| # | Key Result | Baseline | Target | Current | Status |
|---|-----------|----------|--------|---------|--------|
| 1.1 | [Measurable outcome] | [Current value] | [Target value] | [Progress] | On Track / At Risk / Off Track |
| 1.2 | [Measurable outcome] | [Current value] | [Target value] | [Progress] | On Track / At Risk / Off Track |
| 1.3 | [Measurable outcome] | [Current value] | [Target value] | [Progress] | On Track / At Risk / Off Track |
### Initiatives
| Initiative | Key Result | Owner | Status | Effort |
|-----------|-----------|-------|--------|--------|
| [Feature/project name] | KR 1.1 | [Name] | Not Started / In Progress / Complete | [T-shirt size] |
| [Feature/project name] | KR 1.2 | [Name] | Not Started / In Progress / Complete | [T-shirt size] |
---
## Product Objective 2: [Objective Title]
_[Qualitative, inspirational statement]_
### Key Results
| # | Key Result | Baseline | Target | Current | Status |
|---|-----------|----------|--------|---------|--------|
| 2.1 | [Measurable outcome] | [Current value] | [Target value] | [Progress] | On Track / At Risk / Off Track |
| 2.2 | [Measurable outcome] | [Current value] | [Target value] | [Progress] | On Track / At Risk / Off Track |
| 2.3 | [Measurable outcome] | [Current value] | [Target value] | [Progress] | On Track / At Risk / Off Track |
| 2.4 | [Measurable outcome] | [Current value] | [Target value] | [Progress] | On Track / At Risk / Off Track |
### Initiatives
| Initiative | Key Result | Owner | Status | Effort |
|-----------|-----------|-------|--------|--------|
| [Feature/project name] | KR 2.1 | [Name] | Not Started / In Progress / Complete | [T-shirt size] |
| [Feature/project name] | KR 2.2 | [Name] | Not Started / In Progress / Complete | [T-shirt size] |
---
## Product Objective 3: [Objective Title]
_[Qualitative, inspirational statement]_
### Key Results
| # | Key Result | Baseline | Target | Current | Status |
|---|-----------|----------|--------|---------|--------|
| 3.1 | [Measurable outcome] | [Current value] | [Target value] | [Progress] | On Track / At Risk / Off Track |
| 3.2 | [Measurable outcome] | [Current value] | [Target value] | [Progress] | On Track / At Risk / Off Track |
| 3.3 | [Measurable outcome] | [Current value] | [Target value] | [Progress] | On Track / At Risk / Off Track |
### Initiatives
| Initiative | Key Result | Owner | Status | Effort |
|-----------|-----------|-------|--------|--------|
| [Feature/project name] | KR 3.1 | [Name] | Not Started / In Progress / Complete | [T-shirt size] |
---
## Tracking
### Weekly Check-In Format
- **Confidence level** (1-10) for each key result
- **Blockers** identified and escalated
- **Adjustments** to initiatives if needed
### Monthly Review Format
- **Progress update** on all key results with data
- **Initiative status** review
- **Risk assessment** and mitigation updates
- **Stakeholder communication** summary
### End-of-Quarter Scoring
Score each key result 0.0 - 1.0:
- **1.0:** Fully achieved
- **0.7:** Strong progress, nearly there (ideal target)
- **0.4:** Meaningful progress but missed target
- **0.0:** No progress
_Note: Consistently scoring 1.0 means OKRs are not ambitious enough. Target 0.6-0.7 average._
FILE:references/examples/sample_growth_okrs.json
{
"metadata": {
"strategy": "growth",
"quarter": "Q1 2025",
"generated_at": "2025-01-15T10:30:00Z",
"teams": ["Growth", "Platform", "Mobile", "Data"],
"product_contribution": 0.3
},
"company": {
"level": "Company",
"quarter": "Q1 2025",
"strategy": "growth",
"objectives": [
{
"id": "CO-1",
"title": "Accelerate user acquisition and market expansion",
"owner": "CEO",
"status": "active",
"key_results": [
{
"id": "CO-1-KR1",
"title": "Increase MAU from 100,000 to 150,000",
"current": 100000,
"target": 150000,
"unit": "users",
"status": "in_progress",
"progress": 0.2
},
{
"id": "CO-1-KR2",
"title": "Achieve 15% MoM growth rate",
"current": 8,
"target": 15,
"unit": "%",
"status": "in_progress",
"progress": 0.53
},
{
"id": "CO-1-KR3",
"title": "Expand to 3 new markets",
"current": 0,
"target": 3,
"unit": "markets",
"status": "not_started",
"progress": 0
}
]
},
{
"id": "CO-2",
"title": "Achieve product-market fit in enterprise segment",
"owner": "CEO",
"status": "active",
"key_results": [
{
"id": "CO-2-KR1",
"title": "Reduce CAC by 25%",
"current": 150,
"target": 112.5,
"unit": "$",
"status": "in_progress",
"progress": 0.4
},
{
"id": "CO-2-KR2",
"title": "Improve activation rate to 60%",
"current": 42,
"target": 60,
"unit": "%",
"status": "in_progress",
"progress": 0.3
}
]
},
{
"id": "CO-3",
"title": "Build sustainable growth engine",
"owner": "CEO",
"status": "active",
"key_results": [
{
"id": "CO-3-KR1",
"title": "Increase viral coefficient to 1.2",
"current": 0.8,
"target": 1.2,
"unit": "coefficient",
"status": "not_started",
"progress": 0
},
{
"id": "CO-3-KR2",
"title": "Grow organic acquisition to 40% of total",
"current": 25,
"target": 40,
"unit": "%",
"status": "in_progress",
"progress": 0.2
}
]
}
]
},
"product": {
"level": "Product",
"quarter": "Q1 2025",
"parent": "Company",
"objectives": [
{
"id": "PO-1",
"title": "Build viral product features to drive acquisition",
"parent_objective": "CO-1",
"owner": "Head of Product",
"status": "active",
"key_results": [
{
"id": "PO-1-KR1",
"title": "Increase product MAU from 100,000 to 115,000 (30% contribution)",
"contributes_to": "CO-1-KR1",
"current": 100000,
"target": 115000,
"unit": "users",
"status": "in_progress"
},
{
"id": "PO-1-KR2",
"title": "Achieve 12% feature adoption rate for sharing features",
"contributes_to": "CO-1-KR2",
"current": 5,
"target": 12,
"unit": "%",
"status": "in_progress"
}
]
},
{
"id": "PO-2",
"title": "Validate product hypotheses for enterprise segment",
"parent_objective": "CO-2",
"owner": "Head of Product",
"status": "active",
"key_results": [
{
"id": "PO-2-KR1",
"title": "Improve product onboarding efficiency by 30%",
"contributes_to": "CO-2-KR1",
"current": 0,
"target": 30,
"unit": "%",
"status": "not_started"
},
{
"id": "PO-2-KR2",
"title": "Increase product activation rate to 55%",
"contributes_to": "CO-2-KR2",
"current": 42,
"target": 55,
"unit": "%",
"status": "in_progress"
}
]
},
{
"id": "PO-3",
"title": "Create product-led growth loops",
"parent_objective": "CO-3",
"owner": "Head of Product",
"status": "active",
"key_results": [
{
"id": "PO-3-KR1",
"title": "Launch referral program with 0.3 viral coefficient contribution",
"contributes_to": "CO-3-KR1",
"current": 0,
"target": 0.3,
"unit": "coefficient",
"status": "not_started"
},
{
"id": "PO-3-KR2",
"title": "Increase product-driven organic signups to 35%",
"contributes_to": "CO-3-KR2",
"current": 20,
"target": 35,
"unit": "%",
"status": "in_progress"
}
]
}
]
},
"teams": [
{
"level": "Team",
"team": "Growth",
"quarter": "Q1 2025",
"parent": "Product",
"objectives": [
{
"id": "GRO-1",
"title": "Build viral product features through acquisition and activation",
"parent_objective": "PO-1",
"owner": "Growth PM",
"status": "active",
"key_results": [
{
"id": "GRO-1-KR1",
"title": "[Growth] Increase product MAU contribution by 5,000 users",
"contributes_to": "PO-1-KR1",
"current": 0,
"target": 5000,
"unit": "users",
"status": "in_progress"
},
{
"id": "GRO-1-KR2",
"title": "[Growth] Launch 3 viral feature experiments",
"contributes_to": "PO-1-KR2",
"current": 0,
"target": 3,
"unit": "experiments",
"status": "not_started"
}
]
}
]
},
{
"level": "Team",
"team": "Platform",
"quarter": "Q1 2025",
"parent": "Product",
"objectives": [
{
"id": "PLA-1",
"title": "Support growth through infrastructure and reliability",
"parent_objective": "PO-1",
"owner": "Platform PM",
"status": "active",
"key_results": [
{
"id": "PLA-1-KR1",
"title": "[Platform] Scale infrastructure to support 200K MAU",
"contributes_to": "PO-1-KR1",
"current": 100000,
"target": 200000,
"unit": "users",
"status": "in_progress"
},
{
"id": "PLA-1-KR2",
"title": "[Platform] Maintain 99.9% uptime during growth",
"contributes_to": "PO-1-KR2",
"current": 99.5,
"target": 99.9,
"unit": "%",
"status": "in_progress"
}
]
},
{
"id": "PLA-2",
"title": "Improve onboarding infrastructure efficiency",
"parent_objective": "PO-2",
"owner": "Platform PM",
"status": "active",
"key_results": [
{
"id": "PLA-2-KR1",
"title": "[Platform] Reduce onboarding API latency by 40%",
"contributes_to": "PO-2-KR1",
"current": 0,
"target": 40,
"unit": "%",
"status": "not_started"
}
]
}
]
},
{
"level": "Team",
"team": "Mobile",
"quarter": "Q1 2025",
"parent": "Product",
"objectives": [
{
"id": "MOB-1",
"title": "Build viral features through mobile experience",
"parent_objective": "PO-1",
"owner": "Mobile PM",
"status": "active",
"key_results": [
{
"id": "MOB-1-KR1",
"title": "[Mobile] Increase mobile MAU by 3,000 users",
"contributes_to": "PO-1-KR1",
"current": 0,
"target": 3000,
"unit": "users",
"status": "not_started"
},
{
"id": "MOB-1-KR2",
"title": "[Mobile] Launch native share feature with 15% adoption",
"contributes_to": "PO-1-KR2",
"current": 0,
"target": 15,
"unit": "%",
"status": "not_started"
}
]
}
]
},
{
"level": "Team",
"team": "Data",
"quarter": "Q1 2025",
"parent": "Product",
"objectives": [
{
"id": "DAT-1",
"title": "Enable growth through analytics and insights",
"parent_objective": "PO-1",
"owner": "Data PM",
"status": "active",
"key_results": [
{
"id": "DAT-1-KR1",
"title": "[Data] Build growth dashboard tracking all acquisition metrics",
"contributes_to": "PO-1-KR1",
"current": 0,
"target": 1,
"unit": "dashboard",
"status": "not_started"
},
{
"id": "DAT-1-KR2",
"title": "[Data] Implement experimentation platform for A/B testing",
"contributes_to": "PO-1-KR2",
"current": 0,
"target": 1,
"unit": "platform",
"status": "not_started"
}
]
}
]
}
],
"alignment_scores": {
"vertical_alignment": 100.0,
"horizontal_alignment": 75.0,
"coverage": 100.0,
"balance": 85.0,
"overall": 92.0
},
"summary": {
"total_objectives": 11,
"total_key_results": 22,
"company_objectives": 3,
"product_objectives": 3,
"team_objectives": 5,
"teams_involved": 4
}
}
FILE:references/okr_framework.md
# OKR Cascade Framework
A practical guide to Objectives and Key Results (OKRs) and how to cascade them across organizational levels.
---
## Table of Contents
- [What Are OKRs](#what-are-okrs)
- [The Cascade Model](#the-cascade-model)
- [Writing Effective Objectives](#writing-effective-objectives)
- [Defining Key Results](#defining-key-results)
- [Alignment Scoring](#alignment-scoring)
- [Common Pitfalls](#common-pitfalls)
- [OKR Cadence](#okr-cadence)
---
## What Are OKRs
**Objectives and Key Results (OKRs)** are a goal-setting framework that connects organizational strategy to measurable outcomes.
### Components
| Component | Definition | Characteristics |
|-----------|------------|-----------------|
| **Objective** | What you want to achieve | Qualitative, inspirational, time-bound |
| **Key Result** | How you measure progress | Quantitative, specific, measurable |
### OKR Formula
```
Objective: [Inspirational goal statement]
├── KR1: [Metric] from [current] to [target] by [date]
├── KR2: [Metric] from [current] to [target] by [date]
└── KR3: [Metric] from [current] to [target] by [date]
```
### Example
```
Objective: Become the go-to solution for enterprise customers
KR1: Increase enterprise ARR from $5M to $8M
KR2: Improve enterprise NPS from 35 to 50
KR3: Reduce enterprise onboarding time from 30 days to 14 days
```
---
## The Cascade Model
OKRs cascade from company strategy down to individual teams, ensuring alignment at every level.
### Cascade Structure
```
┌─────────────────────────────────────────┐
│ COMPANY LEVEL │
│ Strategic objectives set by leadership │
│ Owned by: CEO, Executive Team │
└───────────────┬─────────────────────────┘
│
▼
┌─────────────────────────────────────────┐
│ PRODUCT LEVEL │
│ How product org contributes to company │
│ Owned by: Head of Product, CPO │
└───────────────┬─────────────────────────┘
│
▼
┌─────────────────────────────────────────┐
│ TEAM LEVEL │
│ Specific initiatives and deliverables │
│ Owned by: Product Managers, Tech Leads │
└─────────────────────────────────────────┘
```
### Contribution Model
Each level contributes a percentage to the level above:
| Level | Typical Contribution | Range |
|-------|---------------------|-------|
| Product → Company | 30% | 20-50% |
| Team → Product | 25% per team | 15-35% |
**Note:** Contribution percentages should be calibrated based on:
- Number of teams
- Relative team size
- Strategic importance of initiatives
### Alignment Types
| Alignment | Description | Goal |
|-----------|-------------|------|
| **Vertical** | Each level supports the level above | >90% of objectives linked |
| **Horizontal** | Teams coordinate on shared objectives | No conflicting goals |
| **Temporal** | Quarterly OKRs support annual goals | Clear progression |
---
## Writing Effective Objectives
### The 3 Cs of Objectives
| Criterion | Description | Example |
|-----------|-------------|---------|
| **Clear** | Unambiguous intent | "Improve customer onboarding" not "Make things better" |
| **Compelling** | Inspires action | "Delight enterprise customers" not "Serve enterprise" |
| **Challenging** | Stretches capabilities | Achievable but requires effort |
### Objective Templates by Strategy
**Growth Strategy:**
```
- Accelerate user acquisition in [segment]
- Expand market presence in [region/vertical]
- Build sustainable acquisition channels
```
**Retention Strategy:**
```
- Create lasting value for [user segment]
- Improve product experience for [use case]
- Maximize customer lifetime value
```
**Revenue Strategy:**
```
- Drive revenue growth through [mechanism]
- Optimize monetization for [segment]
- Expand revenue per customer
```
**Innovation Strategy:**
```
- Pioneer [capability] in the market
- Establish leadership through [innovation area]
- Build competitive differentiation
```
**Operational Strategy:**
```
- Improve delivery efficiency by [mechanism]
- Scale operations to support [target]
- Reduce operational friction in [area]
```
### Objective Anti-Patterns
| Anti-Pattern | Problem | Better Alternative |
|--------------|---------|-------------------|
| "Increase revenue" | Too vague | "Grow enterprise ARR to $10M" |
| "Be the best" | Not measurable | "Achieve #1 NPS in category" |
| "Fix bugs" | Too tactical | "Improve platform reliability" |
| "Launch feature X" | Output, not outcome | "Improve [metric] through [capability]" |
---
## Defining Key Results
### Key Result Anatomy
```
[Verb] [metric] from [current baseline] to [target] by [deadline]
```
### Key Result Types
| Type | Characteristics | When to Use |
|------|-----------------|-------------|
| **Metric-based** | Track a number | Most common, highly measurable |
| **Milestone-based** | Track completion | For binary deliverables |
| **Health-based** | Track stability | For maintenance objectives |
### Metric Categories
| Category | Examples |
|----------|----------|
| **Acquisition** | Signups, trials started, leads generated |
| **Activation** | Onboarding completion, first value moment |
| **Retention** | D7/D30 retention, churn rate, repeat usage |
| **Revenue** | ARR, ARPU, conversion rate, LTV |
| **Engagement** | DAU/MAU, session duration, actions per session |
| **Satisfaction** | NPS, CSAT, support tickets |
| **Efficiency** | Cycle time, automation rate, cost per unit |
### Key Result Scoring
| Score | Status | Description |
|-------|--------|-------------|
| 0.0-0.3 | Red | Significant gap, needs intervention |
| 0.4-0.6 | Yellow | Partial progress, on watch |
| 0.7-0.9 | Green | Strong progress, on track |
| 1.0 | Complete | Target achieved |
**Note:** Hitting 0.7 is considered success for stretch goals. Consistently hitting 1.0 suggests targets aren't ambitious enough.
---
## Alignment Scoring
The OKR cascade generator calculates alignment scores across four dimensions:
### Scoring Dimensions
| Dimension | Weight | What It Measures |
|-----------|--------|------------------|
| **Vertical Alignment** | 40% | % of objectives with parent links |
| **Horizontal Alignment** | 20% | Cross-team coordination on shared goals |
| **Coverage** | 20% | % of company KRs addressed by product |
| **Balance** | 20% | Even distribution of work across teams |
### Alignment Score Interpretation
| Score | Grade | Interpretation |
|-------|-------|----------------|
| 90-100% | A | Excellent alignment, well-cascaded |
| 80-89% | B | Good alignment, minor gaps |
| 70-79% | C | Adequate, needs attention |
| 60-69% | D | Poor alignment, significant gaps |
| <60% | F | Misaligned, requires restructuring |
### Target Benchmarks
| Metric | Target | Red Flag |
|--------|--------|----------|
| Vertical alignment | >90% | <70% |
| Horizontal alignment | >75% | <50% |
| Coverage | >80% | <60% |
| Balance | >80% | <60% |
| Overall | >80% | <65% |
---
## Common Pitfalls
### OKR Anti-Patterns
| Pitfall | Symptom | Fix |
|---------|---------|-----|
| **Too many OKRs** | 10+ objectives per level | Limit to 3-5 objectives |
| **Sandbagging** | Always hit 100% | Set stretch targets (0.7 = success) |
| **Task lists** | KRs are tasks, not outcomes | Focus on measurable impact |
| **Set and forget** | No mid-quarter reviews | Check-ins every 2 weeks |
| **Cascade disconnect** | Team OKRs don't link up | Validate parent relationships |
| **Metric gaming** | Optimizing for KR, not intent | Balance with health metrics |
### Warning Signs
- All teams have identical objectives (lack of specialization)
- No team owns a critical company objective (gap in coverage)
- One team owns everything (unrealistic load)
- Objectives change weekly (lack of commitment)
- KRs are activities, not outcomes (wrong focus)
---
## OKR Cadence
### Quarterly Rhythm
| Week | Activity |
|------|----------|
| **Week -2** | Leadership sets company OKRs draft |
| **Week -1** | Product and team OKR drafting |
| **Week 0** | OKR finalization and alignment review |
| **Week 2** | First check-in, adjust if needed |
| **Week 6** | Mid-quarter review |
| **Week 10** | Pre-quarter reflection |
| **Week 12** | Quarter close, scoring, learnings |
### Check-in Format
```
Weekly/Bi-weekly Status Update:
1. Confidence level: [Red/Yellow/Green]
2. Progress since last check-in: [specific updates]
3. Blockers: [what's in the way]
4. Asks: [what help is needed]
5. Forecast: [expected end-of-quarter score]
```
### Annual Alignment
Quarterly OKRs should ladder up to annual goals:
```
Annual Goal: Become a $100M ARR business
Q1: Build enterprise sales motion (ARR: $25M → $32M)
Q2: Expand into APAC region (ARR: $32M → $45M)
Q3: Launch self-serve enterprise tier (ARR: $45M → $65M)
Q4: Scale and optimize (ARR: $65M → $100M)
```
---
## Quick Reference
### OKR Checklist
**Before finalizing OKRs:**
- [ ] 3-5 objectives per level (not more)
- [ ] 3-5 key results per objective
- [ ] Each KR has a current baseline and target
- [ ] Vertical alignment validated (parent links)
- [ ] No conflicting objectives across teams
- [ ] Owners assigned to every objective
- [ ] Check-in cadence defined
**During the quarter:**
- [ ] Bi-weekly progress updates
- [ ] Mid-quarter formal review
- [ ] Adjust forecasts based on learnings
- [ ] Escalate blockers early
**End of quarter:**
- [ ] Score all key results (0.0-1.0)
- [ ] Document learnings
- [ ] Celebrate wins
- [ ] Carry forward or close incomplete items
---
*See also: `strategy_types.md` for strategy-specific OKR templates*
FILE:references/strategy_types.md
# Strategy Types for OKR Generation
Comprehensive breakdown of the five core strategy types with objectives, key results, and when to use each.
---
## Table of Contents
- [Strategy Selection Guide](#strategy-selection-guide)
- [Growth Strategy](#growth-strategy)
- [Retention Strategy](#retention-strategy)
- [Revenue Strategy](#revenue-strategy)
- [Innovation Strategy](#innovation-strategy)
- [Operational Strategy](#operational-strategy)
- [Multi-Strategy Combinations](#multi-strategy-combinations)
---
## Strategy Selection Guide
### Decision Matrix
| If your priority is... | Primary Strategy | Secondary Strategy |
|------------------------|------------------|-------------------|
| Scaling user base | Growth | Retention |
| Reducing churn | Retention | Revenue |
| Increasing ARPU | Revenue | Retention |
| Market differentiation | Innovation | Growth |
| Improving efficiency | Operational | Revenue |
| New market entry | Growth | Innovation |
### Strategy by Company Stage
| Stage | Typical Priority | Rationale |
|-------|------------------|-----------|
| **Pre-PMF** | Innovation | Finding product-market fit |
| **Early Growth** | Growth | Scaling acquisition |
| **Growth** | Growth + Retention | Balancing acquisition with value |
| **Scale** | Revenue + Retention | Optimizing unit economics |
| **Mature** | Operational + Revenue | Efficiency and margins |
---
## Growth Strategy
**Focus:** Accelerating user acquisition and market expansion
### When to Use
- User growth is primary company objective
- Product-market fit is validated
- Acquisition channels are scaling
- Ready to invest in growth loops
### Company-Level Objectives
| Objective | Key Results Template |
|-----------|---------------------|
| Accelerate user acquisition and market expansion | - Increase MAU from X to Y<br>- Achieve Z% MoM growth rate<br>- Expand to N new markets |
| Achieve product-market fit in new segments | - Reach X users in [segment]<br>- Achieve Y% activation rate<br>- Validate Z use cases |
| Build sustainable growth engine | - Reduce CAC by X%<br>- Improve viral coefficient to Y<br>- Increase organic share to Z% |
### Product-Level Cascade
| Product Objective | Supports | Key Results |
|-------------------|----------|-------------|
| Build viral product features | User acquisition | - Launch referral program (target: X referrals/user)<br>- Increase shareability by Y% |
| Optimize onboarding experience | Activation | - Improve activation rate from X% to Y%<br>- Reduce time-to-value by Z% |
| Create product-led growth loops | Sustainable growth | - Increase product-qualified leads by X%<br>- Improve trial-to-paid by Y% |
### Team-Level Examples
| Team | Focus Area | Sample KRs |
|------|------------|------------|
| Growth Team | Acquisition & activation | - Improve signup conversion by X%<br>- Launch Y experiments/week |
| Platform Team | Scale & reliability | - Support X concurrent users<br>- Maintain Y% uptime |
| Mobile Team | Mobile acquisition | - Increase mobile signups by X%<br>- Improve mobile activation by Y% |
### Key Metrics to Track
- Monthly Active Users (MAU)
- Growth rate (MoM, YoY)
- Customer Acquisition Cost (CAC)
- Activation rate
- Viral coefficient
- Channel efficiency
---
## Retention Strategy
**Focus:** Creating lasting customer value and reducing churn
### When to Use
- Churn is above industry benchmark
- LTV/CAC needs improvement
- Product stickiness is low
- Expansion revenue is a priority
### Company-Level Objectives
| Objective | Key Results Template |
|-----------|---------------------|
| Create lasting customer value and loyalty | - Improve retention from X% to Y%<br>- Increase NPS from X to Y<br>- Reduce churn to below Z% |
| Deliver a superior user experience | - Achieve X% product stickiness<br>- Improve satisfaction to Y/10<br>- Reduce support tickets by Z% |
| Maximize customer lifetime value | - Increase LTV by X%<br>- Improve LTV/CAC ratio to Y<br>- Grow expansion revenue by Z% |
### Product-Level Cascade
| Product Objective | Supports | Key Results |
|-------------------|----------|-------------|
| Design sticky user experiences | Customer retention | - Increase DAU/MAU ratio from X to Y<br>- Improve weekly return rate by Z% |
| Build habit-forming features | Product stickiness | - Achieve X% feature adoption<br>- Increase sessions/user by Y |
| Create expansion opportunities | Lifetime value | - Launch N upsell touchpoints<br>- Improve upgrade rate by X% |
### Team-Level Examples
| Team | Focus Area | Sample KRs |
|------|------------|------------|
| Growth Team | Retention loops | - Improve D7 retention by X%<br>- Reduce first-week churn by Y% |
| Data Team | Churn prediction | - Build churn model (accuracy >X%)<br>- Identify Y at-risk signals |
| Platform Team | Reliability | - Reduce error rates by X%<br>- Improve load times by Y% |
### Key Metrics to Track
- Retention rates (D1, D7, D30, D90)
- Churn rate
- Net Promoter Score (NPS)
- Customer Satisfaction (CSAT)
- Feature stickiness
- Session frequency
---
## Revenue Strategy
**Focus:** Driving sustainable revenue growth and monetization
### When to Use
- Company is focused on profitability
- Monetization needs optimization
- Pricing strategy is being revised
- Expansion revenue is priority
### Company-Level Objectives
| Objective | Key Results Template |
|-----------|---------------------|
| Drive sustainable revenue growth | - Grow ARR from $X to $Y<br>- Achieve Z% revenue growth rate<br>- Maintain X% gross margin |
| Optimize monetization strategy | - Increase ARPU by X%<br>- Improve pricing efficiency by Y%<br>- Launch Z new pricing tiers |
| Expand revenue per customer | - Grow expansion revenue by X%<br>- Reduce revenue churn to Y%<br>- Increase upsell rate by Z% |
### Product-Level Cascade
| Product Objective | Supports | Key Results |
|-------------------|----------|-------------|
| Optimize product monetization | Revenue growth | - Improve conversion to paid by X%<br>- Reduce free tier abuse by Y% |
| Build premium features | ARPU growth | - Launch N premium features<br>- Achieve X% premium adoption |
| Create value-based pricing alignment | Pricing efficiency | - Implement usage-based pricing<br>- Improve price-to-value ratio by X% |
### Team-Level Examples
| Team | Focus Area | Sample KRs |
|------|------------|------------|
| Growth Team | Conversion | - Improve trial-to-paid by X%<br>- Reduce time-to-upgrade by Y days |
| Platform Team | Usage metering | - Implement accurate usage tracking<br>- Support X billing scenarios |
| Data Team | Revenue analytics | - Build revenue forecasting model<br>- Identify Y expansion signals |
### Key Metrics to Track
- Annual Recurring Revenue (ARR)
- Average Revenue Per User (ARPU)
- Gross margin
- Revenue churn (net and gross)
- Expansion revenue
- LTV/CAC ratio
---
## Innovation Strategy
**Focus:** Building competitive advantage through product innovation
### When to Use
- Market is commoditizing
- Competitors are catching up
- New technology opportunity exists
- Company needs differentiation
### Company-Level Objectives
| Objective | Key Results Template |
|-----------|---------------------|
| Lead the market through product innovation | - Launch X breakthrough features<br>- Achieve Y% revenue from new products<br>- File Z patents/IP |
| Establish market leadership in [area] | - Become #1 in category for X<br>- Win Y analyst recognitions<br>- Achieve Z% awareness |
| Build sustainable competitive moat | - Reduce feature parity gap by X%<br>- Create Y unique capabilities<br>- Build Z switching barriers |
### Product-Level Cascade
| Product Objective | Supports | Key Results |
|-------------------|----------|-------------|
| Ship innovative features faster | Breakthrough innovation | - Reduce time-to-market by X%<br>- Launch Y experiments/quarter |
| Build unique technical capabilities | Competitive moat | - Develop X proprietary algorithms<br>- Achieve Y performance advantage |
| Create platform extensibility | Ecosystem advantage | - Launch N API endpoints<br>- Enable X third-party integrations |
### Team-Level Examples
| Team | Focus Area | Sample KRs |
|------|------------|------------|
| Platform Team | Core technology | - Build X new infrastructure capabilities<br>- Improve performance by Y% |
| Data Team | ML/AI innovation | - Deploy X ML models<br>- Improve prediction accuracy by Y% |
| Mobile Team | Mobile innovation | - Launch X mobile-first features<br>- Achieve Y% mobile parity |
### Key Metrics to Track
- Time-to-market
- Revenue from new products
- Feature uniqueness score
- Patent/IP filings
- Technology differentiation
- Innovation velocity
---
## Operational Strategy
**Focus:** Improving efficiency and organizational excellence
### When to Use
- Scaling challenges are emerging
- Operational costs are high
- Team productivity needs improvement
- Quality issues are increasing
### Company-Level Objectives
| Objective | Key Results Template |
|-----------|---------------------|
| Improve organizational efficiency | - Improve velocity by X%<br>- Reduce cycle time to Y days<br>- Achieve Z% automation |
| Scale operations sustainably | - Support X users per engineer<br>- Reduce cost per transaction by Y%<br>- Improve operational leverage by Z% |
| Achieve operational excellence | - Reduce incidents by X%<br>- Improve team NPS to Y<br>- Achieve Z% on-time delivery |
### Product-Level Cascade
| Product Objective | Supports | Key Results |
|-------------------|----------|-------------|
| Improve product delivery efficiency | Velocity | - Reduce PR cycle time by X%<br>- Increase deployment frequency by Y% |
| Reduce operational toil | Automation | - Automate X% of manual processes<br>- Reduce on-call burden by Y% |
| Improve product quality | Excellence | - Reduce bugs by X%<br>- Improve test coverage to Y% |
### Team-Level Examples
| Team | Focus Area | Sample KRs |
|------|------------|------------|
| Platform Team | Infrastructure efficiency | - Reduce infrastructure costs by X%<br>- Improve deployment reliability to Y% |
| Data Team | Data operations | - Improve data pipeline reliability to X%<br>- Reduce data latency by Y% |
| All Teams | Process improvement | - Reduce meeting overhead by X%<br>- Improve sprint predictability to Y% |
### Key Metrics to Track
- Velocity (story points, throughput)
- Cycle time
- Deployment frequency
- Change failure rate
- Incident count and MTTR
- Team satisfaction (eNPS)
---
## Multi-Strategy Combinations
### Common Pairings
| Primary | Secondary | Balanced Objectives |
|---------|-----------|---------------------|
| Growth + Retention | 60/40 | Grow while keeping users |
| Revenue + Retention | 50/50 | Monetize without churning |
| Innovation + Growth | 40/60 | Differentiate to acquire |
| Operational + Revenue | 50/50 | Efficiency for margins |
### Balanced OKR Set Example
**Mixed Growth + Retention Strategy:**
```
Company Objective 1: Accelerate user growth (Growth)
├── KR1: Increase MAU from 100K to 200K
├── KR2: Achieve 15% MoM growth rate
└── KR3: Reduce CAC by 20%
Company Objective 2: Improve user retention (Retention)
├── KR1: Improve D30 retention from 20% to 35%
├── KR2: Increase NPS from 40 to 55
└── KR3: Reduce churn to below 5%
Company Objective 3: Improve delivery efficiency (Operational)
├── KR1: Reduce cycle time by 30%
├── KR2: Achieve 95% on-time delivery
└── KR3: Improve team eNPS to 50
```
---
## Strategy Selection Checklist
Before choosing a strategy:
- [ ] What is the company's #1 priority this quarter?
- [ ] What metrics is leadership being evaluated on?
- [ ] Where are the biggest gaps vs. competitors?
- [ ] What does customer feedback emphasize?
- [ ] What can we realistically move in 90 days?
---
*See also: `okr_framework.md` for OKR writing best practices*
FILE:scripts/okr_cascade_generator.py
#!/usr/bin/env python3
"""
OKR Cascade Generator
Creates aligned OKRs from company strategy down to team level.
Features:
- Generates company → product → team OKR cascade
- Configurable team structure and contribution percentages
- Alignment scoring across vertical and horizontal dimensions
- Multiple output formats (dashboard, JSON)
Usage:
python okr_cascade_generator.py growth
python okr_cascade_generator.py retention --teams "Engineering,Design,Data"
python okr_cascade_generator.py revenue --contribution 0.4 --json
"""
import json
import argparse
from typing import Dict, List
from datetime import datetime
class OKRGenerator:
"""Generate and cascade OKRs across the organization"""
def __init__(self, teams: List[str] = None, product_contribution: float = 0.3):
"""
Initialize OKR generator.
Args:
teams: List of team names (default: Growth, Platform, Mobile, Data)
product_contribution: Fraction of company KRs that product owns (default: 0.3)
"""
self.teams = teams or ['Growth', 'Platform', 'Mobile', 'Data']
self.product_contribution = product_contribution
self.okr_templates = {
'growth': {
'objectives': [
'Accelerate user acquisition and market expansion',
'Achieve product-market fit in new segments',
'Build sustainable growth engine'
],
'key_results': [
'Increase MAU from {current} to {target}',
'Achieve {target}% MoM growth rate',
'Expand to {target} new markets',
'Reduce CAC by {target}%',
'Improve activation rate to {target}%'
]
},
'retention': {
'objectives': [
'Create lasting customer value and loyalty',
'Deliver a superior user experience',
'Maximize customer lifetime value'
],
'key_results': [
'Improve retention from {current}% to {target}%',
'Increase NPS from {current} to {target}',
'Reduce churn to below {target}%',
'Achieve {target}% product stickiness',
'Increase LTV/CAC ratio to {target}'
]
},
'revenue': {
'objectives': [
'Drive sustainable revenue growth',
'Optimize monetization strategy',
'Expand revenue per customer'
],
'key_results': [
'Grow ARR from currentM to targetM',
'Increase ARPU by {target}%',
'Launch {target} new revenue streams',
'Achieve {target}% gross margin',
'Reduce revenue churn to {target}%'
]
},
'innovation': {
'objectives': [
'Lead the market through product innovation',
'Establish leadership in key capability areas',
'Build sustainable competitive differentiation'
],
'key_results': [
'Launch {target} breakthrough features',
'Achieve {target}% of revenue from new products',
'File {target} patents/IP',
'Reduce time-to-market by {target}%',
'Achieve {target} innovation score'
]
},
'operational': {
'objectives': [
'Improve organizational efficiency',
'Achieve operational excellence',
'Scale operations sustainably'
],
'key_results': [
'Improve velocity by {target}%',
'Reduce cycle time to {target} days',
'Achieve {target}% automation',
'Improve team satisfaction to {target}',
'Reduce incidents by {target}%'
]
}
}
# Team focus areas for objective relevance matching
self.team_relevance = {
'Growth': ['acquisition', 'growth', 'activation', 'viral', 'onboarding', 'conversion'],
'Platform': ['infrastructure', 'reliability', 'scale', 'performance', 'efficiency', 'automation'],
'Mobile': ['mobile', 'app', 'ios', 'android', 'native'],
'Data': ['analytics', 'metrics', 'insights', 'data', 'measurement', 'experimentation'],
'Engineering': ['delivery', 'velocity', 'quality', 'automation', 'infrastructure'],
'Design': ['experience', 'usability', 'interface', 'user', 'accessibility'],
'Product': ['features', 'roadmap', 'prioritization', 'strategy'],
}
def generate_company_okrs(self, strategy: str, metrics: Dict) -> Dict:
"""Generate company-level OKRs based on strategy"""
if strategy not in self.okr_templates:
strategy = 'growth'
template = self.okr_templates[strategy]
company_okrs = {
'level': 'Company',
'quarter': self._get_current_quarter(),
'strategy': strategy,
'objectives': []
}
for i in range(min(3, len(template['objectives']))):
obj = {
'id': f'CO-{i+1}',
'title': template['objectives'][i],
'key_results': [],
'owner': 'CEO',
'status': 'draft'
}
for j in range(3):
if j < len(template['key_results']):
kr_template = template['key_results'][j]
kr = {
'id': f'CO-{i+1}-KR{j+1}',
'title': self._fill_metrics(kr_template, metrics),
'current': metrics.get('current', 0),
'target': metrics.get('target', 100),
'unit': self._extract_unit(kr_template),
'status': 'not_started'
}
obj['key_results'].append(kr)
company_okrs['objectives'].append(obj)
return company_okrs
def cascade_to_product(self, company_okrs: Dict) -> Dict:
"""Cascade company OKRs to product organization"""
product_okrs = {
'level': 'Product',
'quarter': company_okrs['quarter'],
'parent': 'Company',
'contribution': self.product_contribution,
'objectives': []
}
for company_obj in company_okrs['objectives']:
product_obj = {
'id': f'PO-{company_obj["id"].split("-")[1]}',
'title': self._translate_to_product(company_obj['title']),
'parent_objective': company_obj['id'],
'key_results': [],
'owner': 'Head of Product',
'status': 'draft'
}
for kr in company_obj['key_results']:
product_kr = {
'id': f'PO-{product_obj["id"].split("-")[1]}-KR{kr["id"].split("KR")[1]}',
'title': self._translate_kr_to_product(kr['title']),
'contributes_to': kr['id'],
'current': kr['current'],
'target': kr['target'] * self.product_contribution,
'unit': kr['unit'],
'contribution_pct': self.product_contribution * 100,
'status': 'not_started'
}
product_obj['key_results'].append(product_kr)
product_okrs['objectives'].append(product_obj)
return product_okrs
def cascade_to_teams(self, product_okrs: Dict) -> List[Dict]:
"""Cascade product OKRs to individual teams"""
team_okrs = []
team_contribution = 1.0 / len(self.teams) if self.teams else 0.25
for team in self.teams:
team_okr = {
'level': 'Team',
'team': team,
'quarter': product_okrs['quarter'],
'parent': 'Product',
'contribution': team_contribution,
'objectives': []
}
for product_obj in product_okrs['objectives']:
if self._is_relevant_for_team(product_obj['title'], team):
team_obj = {
'id': f'{team[:3].upper()}-{product_obj["id"].split("-")[1]}',
'title': self._translate_to_team(product_obj['title'], team),
'parent_objective': product_obj['id'],
'key_results': [],
'owner': f'{team} PM',
'status': 'draft'
}
for kr in product_obj['key_results'][:2]:
team_kr = {
'id': f'{team[:3].upper()}-{team_obj["id"].split("-")[1]}-KR{kr["id"].split("KR")[1]}',
'title': self._translate_kr_to_team(kr['title'], team),
'contributes_to': kr['id'],
'current': kr['current'],
'target': kr['target'] * team_contribution,
'unit': kr['unit'],
'status': 'not_started'
}
team_obj['key_results'].append(team_kr)
team_okr['objectives'].append(team_obj)
if team_okr['objectives']:
team_okrs.append(team_okr)
return team_okrs
def generate_okr_dashboard(self, all_okrs: Dict) -> str:
"""Generate OKR dashboard view"""
dashboard = ["=" * 60]
dashboard.append("OKR CASCADE DASHBOARD")
dashboard.append(f"Quarter: {all_okrs.get('quarter', 'Q1 2025')}")
dashboard.append(f"Strategy: {all_okrs.get('strategy', 'growth').upper()}")
dashboard.append(f"Teams: {', '.join(self.teams)}")
dashboard.append(f"Product Contribution: {self.product_contribution * 100:.0f}%")
dashboard.append("=" * 60)
# Company OKRs
if 'company' in all_okrs:
dashboard.append("\n🏢 COMPANY OKRS\n")
for obj in all_okrs['company']['objectives']:
dashboard.append(f"📌 {obj['id']}: {obj['title']}")
for kr in obj['key_results']:
dashboard.append(f" └─ {kr['id']}: {kr['title']}")
# Product OKRs
if 'product' in all_okrs:
dashboard.append("\n🚀 PRODUCT OKRS\n")
for obj in all_okrs['product']['objectives']:
dashboard.append(f"📌 {obj['id']}: {obj['title']}")
dashboard.append(f" ↳ Supports: {obj.get('parent_objective', 'N/A')}")
for kr in obj['key_results']:
dashboard.append(f" └─ {kr['id']}: {kr['title']}")
# Team OKRs
if 'teams' in all_okrs:
dashboard.append("\n👥 TEAM OKRS\n")
for team_okr in all_okrs['teams']:
dashboard.append(f"\n{team_okr['team']} Team:")
for obj in team_okr['objectives']:
dashboard.append(f" 📌 {obj['id']}: {obj['title']}")
for kr in obj['key_results']:
dashboard.append(f" └─ {kr['id']}: {kr['title']}")
# Alignment Matrix
dashboard.append("\n\n📊 ALIGNMENT MATRIX\n")
dashboard.append("Company → Product → Teams")
dashboard.append("-" * 40)
if 'company' in all_okrs and 'product' in all_okrs:
for c_obj in all_okrs['company']['objectives']:
dashboard.append(f"\n{c_obj['id']}")
for p_obj in all_okrs['product']['objectives']:
if p_obj.get('parent_objective') == c_obj['id']:
dashboard.append(f" ├─ {p_obj['id']}")
if 'teams' in all_okrs:
for team_okr in all_okrs['teams']:
for t_obj in team_okr['objectives']:
if t_obj.get('parent_objective') == p_obj['id']:
dashboard.append(f" └─ {t_obj['id']} ({team_okr['team']})")
return "\n".join(dashboard)
def calculate_alignment_score(self, all_okrs: Dict) -> Dict:
"""Calculate alignment score across OKR cascade"""
scores = {
'vertical_alignment': 0,
'horizontal_alignment': 0,
'coverage': 0,
'balance': 0,
'overall': 0
}
# Vertical alignment: How well each level supports the above
total_objectives = 0
aligned_objectives = 0
if 'product' in all_okrs:
for obj in all_okrs['product']['objectives']:
total_objectives += 1
if 'parent_objective' in obj:
aligned_objectives += 1
if 'teams' in all_okrs:
for team in all_okrs['teams']:
for obj in team['objectives']:
total_objectives += 1
if 'parent_objective' in obj:
aligned_objectives += 1
if total_objectives > 0:
scores['vertical_alignment'] = round((aligned_objectives / total_objectives) * 100, 1)
# Horizontal alignment: How well teams coordinate
if 'teams' in all_okrs and len(all_okrs['teams']) > 1:
shared_objectives = set()
for team in all_okrs['teams']:
for obj in team['objectives']:
parent = obj.get('parent_objective')
if parent:
shared_objectives.add(parent)
scores['horizontal_alignment'] = min(100, len(shared_objectives) * 25)
# Coverage: How much of company OKRs are covered
if 'company' in all_okrs and 'product' in all_okrs:
company_krs = sum(len(obj['key_results']) for obj in all_okrs['company']['objectives'])
covered_krs = sum(len(obj['key_results']) for obj in all_okrs['product']['objectives'])
if company_krs > 0:
scores['coverage'] = round((covered_krs / company_krs) * 100, 1)
# Balance: Distribution across teams
if 'teams' in all_okrs:
objectives_per_team = [len(team['objectives']) for team in all_okrs['teams']]
if objectives_per_team:
avg_objectives = sum(objectives_per_team) / len(objectives_per_team)
variance = sum((x - avg_objectives) ** 2 for x in objectives_per_team) / len(objectives_per_team)
scores['balance'] = round(max(0, 100 - variance * 10), 1)
# Overall score
scores['overall'] = round(sum([
scores['vertical_alignment'] * 0.4,
scores['horizontal_alignment'] * 0.2,
scores['coverage'] * 0.2,
scores['balance'] * 0.2
]), 1)
return scores
def _get_current_quarter(self) -> str:
"""Get current quarter"""
now = datetime.now()
quarter = (now.month - 1) // 3 + 1
return f"Q{quarter} {now.year}"
def _fill_metrics(self, template: str, metrics: Dict) -> str:
"""Fill template with actual metrics"""
result = template
for key, value in metrics.items():
result = result.replace(f'{{{key}}}', str(value))
return result
def _extract_unit(self, kr_template: str) -> str:
"""Extract measurement unit from KR template"""
if '%' in kr_template:
return '%'
elif '$' in kr_template:
return '$'
elif 'days' in kr_template.lower():
return 'days'
elif 'score' in kr_template.lower():
return 'points'
return 'count'
def _translate_to_product(self, company_objective: str) -> str:
"""Translate company objective to product objective"""
translations = {
'Accelerate user acquisition': 'Build viral product features',
'Achieve product-market fit': 'Validate product hypotheses',
'Build sustainable growth': 'Create product-led growth loops',
'Create lasting customer value': 'Design sticky user experiences',
'Drive sustainable revenue': 'Optimize product monetization',
'Lead the market through': 'Ship innovative features to',
'Improve organizational': 'Improve product delivery'
}
for key, value in translations.items():
if key in company_objective:
return company_objective.replace(key, value)
return f"Product: {company_objective}"
def _translate_kr_to_product(self, kr: str) -> str:
"""Translate KR to product context"""
product_terms = {
'MAU': 'product MAU',
'growth rate': 'feature adoption rate',
'CAC': 'product onboarding efficiency',
'retention': 'product retention',
'NPS': 'product NPS',
'ARR': 'product-driven revenue',
'churn': 'product churn'
}
result = kr
for term, replacement in product_terms.items():
if term in result:
result = result.replace(term, replacement)
break
return result
def _translate_to_team(self, objective: str, team: str) -> str:
"""Translate objective to team context"""
team_focus = {
'Growth': 'acquisition and activation',
'Platform': 'infrastructure and reliability',
'Mobile': 'mobile experience',
'Data': 'analytics and insights',
'Engineering': 'technical delivery',
'Design': 'user experience',
'Product': 'product strategy'
}
focus = team_focus.get(team, 'delivery')
return f"{objective} through {focus}"
def _translate_kr_to_team(self, kr: str, team: str) -> str:
"""Translate KR to team context"""
return f"[{team}] {kr}"
def _is_relevant_for_team(self, objective: str, team: str) -> bool:
"""Check if objective is relevant for team"""
keywords = self.team_relevance.get(team, [])
objective_lower = objective.lower()
# Platform is always relevant (infrastructure supports everything)
if team == 'Platform':
return True
return any(keyword in objective_lower for keyword in keywords)
def parse_teams(teams_str: str) -> List[str]:
"""Parse comma-separated team string into list"""
if not teams_str:
return None
return [t.strip() for t in teams_str.split(',') if t.strip()]
def main():
parser = argparse.ArgumentParser(
description='Generate OKR cascade from company strategy to team level',
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
# Generate growth strategy OKRs with default teams
python okr_cascade_generator.py growth
# Custom teams
python okr_cascade_generator.py retention --teams "Engineering,Design,Data,Growth"
# Custom product contribution percentage
python okr_cascade_generator.py revenue --contribution 0.4
# JSON output
python okr_cascade_generator.py innovation --json
# All options combined
python okr_cascade_generator.py operational --teams "Core,Platform" --contribution 0.5 --json
"""
)
parser.add_argument(
'strategy',
nargs='?',
choices=['growth', 'retention', 'revenue', 'innovation', 'operational'],
default='growth',
help='Strategy type (default: growth)'
)
parser.add_argument(
'--teams', '-t',
type=str,
help='Comma-separated list of team names (default: Growth,Platform,Mobile,Data)'
)
parser.add_argument(
'--contribution', '-c',
type=float,
default=0.3,
help='Product contribution to company OKRs as decimal (default: 0.3 = 30%%)'
)
parser.add_argument(
'--json', '-j',
action='store_true',
help='Output as JSON instead of dashboard'
)
parser.add_argument(
'--metrics', '-m',
type=str,
help='Metrics as JSON string (default: sample metrics)'
)
args = parser.parse_args()
# Parse teams
teams = parse_teams(args.teams)
# Parse metrics
if args.metrics:
metrics = json.loads(args.metrics)
else:
metrics = {
'current': 100000,
'target': 150000,
'current_revenue': 10,
'target_revenue': 15,
'current_nps': 40,
'target_nps': 60
}
# Validate contribution
if not 0 < args.contribution <= 1:
print("Error: Contribution must be between 0 and 1")
return 1
# Generate OKRs
generator = OKRGenerator(teams=teams, product_contribution=args.contribution)
company_okrs = generator.generate_company_okrs(args.strategy, metrics)
product_okrs = generator.cascade_to_product(company_okrs)
team_okrs = generator.cascade_to_teams(product_okrs)
all_okrs = {
'quarter': company_okrs['quarter'],
'strategy': args.strategy,
'company': company_okrs,
'product': product_okrs,
'teams': team_okrs
}
alignment = generator.calculate_alignment_score(all_okrs)
if args.json:
all_okrs['alignment_scores'] = alignment
all_okrs['config'] = {
'teams': generator.teams,
'product_contribution': generator.product_contribution
}
print(json.dumps(all_okrs, indent=2))
else:
dashboard = generator.generate_okr_dashboard(all_okrs)
print(dashboard)
print("\n\n🎯 ALIGNMENT SCORES")
print("-" * 40)
for metric, score in alignment.items():
status = "✓" if score >= 80 else "!" if score >= 60 else "✗"
print(f"{status} {metric.replace('_', ' ').title()}: {score}%")
if alignment['overall'] >= 80:
print("\n✅ Overall alignment is GOOD (≥80%)")
elif alignment['overall'] >= 60:
print("\n⚠️ Overall alignment NEEDS ATTENTION (60-80%)")
else:
print("\n❌ Overall alignment is POOR (<60%)")
if __name__ == "__main__":
main()
Tạo hàng loạt trang SEO bằng template và dữ liệu: trang vị trí, so sánh, tích hợp, thư mục và các trang theo mẫu từ khóa + thành phố.
---
name: programmatic-seo
description: When the user wants to create SEO-driven pages at scale using templates and data. Also use when the user mentions "programmatic SEO," "template pages," "pages at scale," "directory pages," "location pages," "[keyword] + [city] pages," "comparison pages," "integration pages," "building many pages for SEO," "pSEO," "generate 100 pages," "data-driven pages," or "templated landing pages." Use this whenever someone wants to create many similar pages targeting different keywords or locations. For auditing existing SEO issues, see seo-audit. For content strategy planning, see content-strategy.
metadata:
version: 2.0.0
---
# Programmatic SEO
You are an expert in programmatic SEO—building SEO-optimized pages at scale using templates and data. Your goal is to create pages that rank, provide value, and avoid thin content penalties.
## Initial Assessment
**Check for product marketing context first:**
If `.agents/product-marketing.md` exists (or `.claude/product-marketing.md`, or the legacy `product-marketing-context.md` filename, in older setups), read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
Before designing a programmatic SEO strategy, understand:
1. **Business Context**
- What's the product/service?
- Who is the target audience?
- What's the conversion goal for these pages?
2. **Opportunity Assessment**
- What search patterns exist?
- How many potential pages?
- What's the search volume distribution?
3. **Competitive Landscape**
- Who ranks for these terms now?
- What do their pages look like?
- Can you realistically compete?
---
## Core Principles
### 1. Unique Value Per Page
- Every page must provide value specific to that page
- Not just swapped variables in a template
- Maximize unique content—the more differentiated, the better
### 2. Proprietary Data Wins
Hierarchy of data defensibility:
1. Proprietary (you created it)
2. Product-derived (from your users)
3. User-generated (your community)
4. Licensed (exclusive access)
5. Public (anyone can use—weakest)
### 3. Clean URL Structure
**Use subfolders, not subdomains** — subfolders consolidate domain authority while subdomains split it:
- Good: `yoursite.com/templates/resume/`
- Bad: `templates.yoursite.com/resume/`
### 4. Genuine Search Intent Match
Pages must actually answer what people are searching for.
### 5. Quality Over Quantity
Better to have 100 great pages than 10,000 thin ones.
### 6. Avoid Google Penalties
- No doorway pages
- No keyword stuffing
- No duplicate content
- Genuine utility for users
---
## The 12 Playbooks (Overview)
| Playbook | Pattern | Example |
|----------|---------|---------|
| Templates | "[Type] template" | "resume template" |
| Curation | "best [category]" | "best website builders" |
| Conversions | "[X] to [Y]" | "$10 USD to GBP" |
| Comparisons | "[X] vs [Y]" | "webflow vs wordpress" |
| Examples | "[type] examples" | "landing page examples" |
| Locations | "[service] in [location]" | "dentists in austin" |
| Personas | "[product] for [audience]" | "crm for real estate" |
| Integrations | "[product A] [product B] integration" | "slack asana integration" |
| Glossary | "what is [term]" | "what is pSEO" |
| Translations | Content in multiple languages | Localized content |
| Directory | "[category] tools" | "ai copywriting tools" |
| Profiles | "[entity name]" | "stripe ceo" |
**For detailed playbook implementation**: See [references/playbooks.md](references/playbooks.md)
---
## Choosing Your Playbook
| If you have... | Consider... |
|----------------|-------------|
| Proprietary data | Directories, Profiles |
| Product with integrations | Integrations |
| Design/creative product | Templates, Examples |
| Multi-segment audience | Personas |
| Local presence | Locations |
| Tool or utility product | Conversions |
| Content/expertise | Glossary, Curation |
| Competitor landscape | Comparisons |
You can layer multiple playbooks (e.g., "Best coworking spaces in San Diego").
---
## Implementation Framework
### 1. Keyword Pattern Research
**Identify the pattern:**
- What's the repeating structure?
- What are the variables?
- How many unique combinations exist?
**Validate demand:**
- Aggregate search volume
- Volume distribution (head vs. long tail)
- Trend direction
### 2. Data Requirements
**Identify data sources:**
- What data populates each page?
- Is it first-party, scraped, licensed, public?
- How is it updated?
### 3. Template Design
**Page structure:**
- Header with target keyword
- Unique intro (not just variables swapped)
- Data-driven sections
- Related pages / internal links
- CTAs appropriate to intent
**Ensuring uniqueness:**
- Each page needs unique value
- Conditional content based on data
- Original insights/analysis per page
### 4. Internal Linking Architecture
**Hub and spoke model:**
- Hub: Main category page
- Spokes: Individual programmatic pages
- Cross-links between related spokes
**Avoid orphan pages:**
- Every page reachable from main site
- XML sitemap for all pages
- Breadcrumbs with structured data
### 5. Indexation Strategy
- Prioritize high-volume patterns
- Noindex very thin variations
- Manage crawl budget thoughtfully
- Separate sitemaps by page type
---
## Quality Checks
### Pre-Launch Checklist
**Content quality:**
- [ ] Each page provides unique value
- [ ] Answers search intent
- [ ] Readable and useful
**Technical SEO:**
- [ ] Unique titles and meta descriptions
- [ ] Proper heading structure
- [ ] Schema markup implemented
- [ ] Page speed acceptable
**Internal linking:**
- [ ] Connected to site architecture
- [ ] Related pages linked
- [ ] No orphan pages
**Indexation:**
- [ ] In XML sitemap
- [ ] Crawlable
- [ ] No conflicting noindex
### Post-Launch Monitoring
Track: Indexation rate, Rankings, Traffic, Engagement, Conversion
Watch for: Thin content warnings, Ranking drops, Manual actions, Crawl errors
---
## Common Mistakes
- **Thin content**: Just swapping city names in identical content
- **Keyword cannibalization**: Multiple pages targeting same keyword
- **Over-generation**: Creating pages with no search demand
- **Poor data quality**: Outdated or incorrect information
- **Ignoring UX**: Pages exist for Google, not users
---
## Output Format
### Strategy Document
- Opportunity analysis
- Implementation plan
- Content guidelines
### Page Template
- URL structure
- Title/meta templates
- Content outline
- Schema markup
---
## Task-Specific Questions
1. What keyword patterns are you targeting?
2. What data do you have (or can acquire)?
3. How many pages are you planning?
4. What does your site authority look like?
5. Who currently ranks for these terms?
6. What's your technical stack?
---
## Related Skills
- **seo-audit**: For auditing programmatic pages after launch
- **schema**: For adding structured data
- **site-architecture**: For page hierarchy, URL structure, and internal linking
- **competitors**: For comparison page frameworks
FILE:evals/evals.json
{
"skill_name": "programmatic-seo",
"evals": [
{
"id": 1,
"prompt": "We want to create programmatic SEO pages for our CRM. We're thinking of 'CRM for [industry]' pages — like 'CRM for Real Estate,' 'CRM for Healthcare,' etc. How should we approach this?",
"expected_output": "Should check for product-marketing.md first. Should identify this as the Personas playbook (industry-specific pages). Should apply the core principles: unique value per page (not just swapping the industry name), proprietary data or insights per industry, clean URL structure. Should recommend the implementation framework: keyword research for each industry variation, data requirements (what industry-specific content makes each page unique), template design, internal linking strategy between industry pages and main pages, and indexation strategy. Should warn against thin content (just template + keyword swap).",
"assertions": [
"Checks for product-marketing.md",
"Identifies as Personas playbook",
"Applies core principles (unique value, proprietary data, clean URLs)",
"Recommends keyword research per variation",
"Addresses data requirements for unique content",
"Provides template design guidance",
"Includes internal linking strategy",
"Warns against thin content"
],
"files": []
},
{
"id": 2,
"prompt": "Create a comparison page strategy. We want pages like 'Notion vs Asana', 'Notion vs Monday', etc. for all our competitors. We have 15 competitors.",
"expected_output": "Should identify this as the Comparisons playbook. Should apply the programmatic approach for competitor comparison pages at scale. Should recommend: template structure for comparison pages, unique data per comparison (not just the same template with names swapped), keyword research for each '[competitor A] vs [competitor B]' variation, URL structure (/compare/notion-vs-asana), internal linking between comparison pages, and quality checks. Should cross-reference the competitors skill for page content structure.",
"assertions": [
"Identifies as Comparisons playbook",
"Recommends template structure for scale",
"Addresses unique data per comparison",
"Includes keyword research for variations",
"Provides URL structure recommendation",
"Includes internal linking strategy",
"Cross-references competitors skill",
"Applies quality checks"
],
"files": []
},
{
"id": 3,
"prompt": "we want to rank for '[tool name] integration' keywords. we integrate with 50+ tools and want a page for each. like 'Slack integration', 'Salesforce integration' etc.",
"expected_output": "Should trigger on casual phrasing. Should identify this as the Integrations playbook. Should recommend: template design for integration pages (what it does, how to set up, use cases), unique content per integration (specific workflows, screenshots, setup steps), keyword research for '[tool] + [your product] integration', URL structure (/integrations/slack), hub page linking to all integration pages, and schema markup considerations. Should emphasize that each page needs genuine unique value, not just 'we integrate with [tool].'",
"assertions": [
"Triggers on casual phrasing",
"Identifies as Integrations playbook",
"Recommends template with unique content per integration",
"Includes setup steps and use cases per page",
"Provides URL structure recommendation",
"Recommends hub page for all integrations",
"Emphasizes genuine unique value per page"
],
"files": []
},
{
"id": 4,
"prompt": "We built 500 programmatic pages but Google isn't indexing most of them. Only 80 are in the index. What's going wrong?",
"expected_output": "Should diagnose the indexation problem. Should apply the quality checks and indexation strategy guidance. Should investigate: thin content (are pages providing unique value or just template + keyword?), crawl budget (500 pages may be fine but depends on site authority), internal linking (are the pages discoverable?), XML sitemap inclusion, duplicate/near-duplicate content issues. Should recommend specific fixes: improve content uniqueness, strengthen internal linking, submit sitemap, check robots.txt, use Search Console for indexation requests. Should warn that Google may choose not to index thin pages regardless.",
"assertions": [
"Diagnoses indexation problem",
"Investigates thin content as likely cause",
"Checks crawl budget considerations",
"Checks internal linking to programmatic pages",
"Checks XML sitemap and robots.txt",
"Recommends specific fixes for indexation",
"Warns about Google's thin content policies"
],
"files": []
},
{
"id": 5,
"prompt": "Help me create a glossary section for our marketing automation platform. We want to define 200+ marketing terms and rank for '[term] definition' keywords.",
"expected_output": "Should identify this as the Glossary playbook. Should apply the template design: term definition page template (definition, examples, related terms, how it applies to the user's product), hub/index page linking to all terms, URL structure (/glossary/[term]), alphabetical and categorical navigation. Should address quality: each definition should provide genuine value beyond a dictionary definition. Should include internal linking strategy and schema markup (DefinedTerm schema). Should recommend starting with highest-volume terms.",
"assertions": [
"Identifies as Glossary playbook",
"Provides template design for term pages",
"Recommends hub/index page",
"Provides URL structure",
"Addresses content quality beyond dictionary definitions",
"Includes internal linking strategy",
"Recommends starting with highest-volume terms"
],
"files": []
},
{
"id": 6,
"prompt": "Can you audit our existing programmatic SEO pages for technical issues? We have crawl errors and some pages return 404s.",
"expected_output": "Should recognize this is a technical SEO audit task, not a programmatic SEO strategy task. Should defer to or cross-reference the seo-audit skill, which handles crawlability, indexation, and technical SEO issues. Programmatic-seo focuses on strategy, template design, and content planning for scaled pages.",
"assertions": [
"Recognizes this as technical SEO audit task",
"References or defers to seo-audit skill",
"Explains that programmatic-seo is for strategy and template design",
"Does not attempt full technical SEO audit"
],
"files": []
}
]
}
FILE:references/playbooks.md
# The 12 Programmatic SEO Playbooks
Beyond mixing and matching data point permutations, these are the proven playbooks for programmatic SEO.
## Contents
- 1. Templates
- 2. Curation
- 3. Conversions
- 4. Comparisons
- 5. Examples
- 6. Locations
- 7. Personas
- 8. Integrations
- 9. Glossary
- 10. Translations
- 11. Directory
- 12. Profiles
- Choosing Your Playbook (Match to Your Assets, Combine Playbooks)
## 1. Templates
**Pattern**: "[Type] template" or "free [type] template"
**Example searches**: "resume template", "invoice template", "pitch deck template"
**What it is**: Downloadable or interactive templates users can use directly.
**Why it works**:
- High intent—people need it now
- Shareable/linkable assets
- Natural for product-led companies
**Value requirements**:
- Actually usable templates (not just previews)
- Multiple variations per type
- Quality comparable to paid options
- Easy download/use flow
**URL structure**: `/templates/[type]/` or `/templates/[category]/[type]/`
---
## 2. Curation
**Pattern**: "best [category]" or "top [number] [things]"
**Example searches**: "best website builders", "top 10 crm software", "best free design tools"
**What it is**: Curated lists ranking or recommending options in a category.
**Why it works**:
- Comparison shoppers searching for guidance
- High commercial intent
- Evergreen with updates
**Value requirements**:
- Genuine evaluation criteria
- Real testing or expertise
- Regular updates (date visible)
- Not just affiliate-driven rankings
**URL structure**: `/best/[category]/` or `/[category]/best/`
---
## 3. Conversions
**Pattern**: "[X] to [Y]" or "[amount] [unit] in [unit]"
**Example searches**: "$10 USD to GBP", "100 kg to lbs", "pdf to word"
**What it is**: Tools or pages that convert between formats, units, or currencies.
**Why it works**:
- Instant utility
- Extremely high search volume
- Repeat usage potential
**Value requirements**:
- Accurate, real-time data
- Fast, functional tool
- Related conversions suggested
- Mobile-friendly interface
**URL structure**: `/convert/[from]-to-[to]/` or `/[from]-to-[to]-converter/`
---
## 4. Comparisons
**Pattern**: "[X] vs [Y]" or "[X] alternative"
**Example searches**: "webflow vs wordpress", "notion vs coda", "figma alternatives"
**What it is**: Head-to-head comparisons between products, tools, or options.
**Why it works**:
- High purchase intent
- Clear search pattern
- Scales with number of competitors
**Value requirements**:
- Honest, balanced analysis
- Actual feature comparison data
- Clear recommendation by use case
- Updated when products change
**URL structure**: `/compare/[x]-vs-[y]/` or `/[x]-vs-[y]/`
*See also: competitors skill for detailed frameworks*
---
## 5. Examples
**Pattern**: "[type] examples" or "[category] inspiration"
**Example searches**: "saas landing page examples", "email subject line examples", "portfolio website examples"
**What it is**: Galleries or collections of real-world examples for inspiration.
**Why it works**:
- Research phase traffic
- Highly shareable
- Natural for design/creative tools
**Value requirements**:
- Real, high-quality examples
- Screenshots or embeds
- Categorization/filtering
- Analysis of why they work
**URL structure**: `/examples/[type]/` or `/[type]-examples/`
---
## 6. Locations
**Pattern**: "[service/thing] in [location]"
**Example searches**: "coworking spaces in san diego", "dentists in austin", "best restaurants in brooklyn"
**What it is**: Location-specific pages for services, businesses, or information.
**Why it works**:
- Local intent is massive
- Scales with geography
- Natural for marketplaces/directories
**Value requirements**:
- Actual local data (not just city name swapped)
- Local providers/options listed
- Location-specific insights (pricing, regulations)
- Map integration helpful
**URL structure**: `/[service]/[city]/` or `/locations/[city]/[service]/`
---
## 7. Personas
**Pattern**: "[product] for [audience]" or "[solution] for [role/industry]"
**Example searches**: "payroll software for agencies", "crm for real estate", "project management for freelancers"
**What it is**: Tailored landing pages addressing specific audience segments.
**Why it works**:
- Speaks directly to searcher's context
- Higher conversion than generic pages
- Scales with personas
**Value requirements**:
- Genuine persona-specific content
- Relevant features highlighted
- Testimonials from that segment
- Use cases specific to audience
**URL structure**: `/for/[persona]/` or `/solutions/[industry]/`
---
## 8. Integrations
**Pattern**: "[your product] [other product] integration" or "[product] + [product]"
**Example searches**: "slack asana integration", "zapier airtable", "hubspot salesforce sync"
**What it is**: Pages explaining how your product works with other tools.
**Why it works**:
- Captures users of other products
- High intent (they want the solution)
- Scales with integration ecosystem
**Value requirements**:
- Real integration details
- Setup instructions
- Use cases for the combination
- Working integration (not vaporware)
**URL structure**: `/integrations/[product]/` or `/connect/[product]/`
---
## 9. Glossary
**Pattern**: "what is [term]" or "[term] definition" or "[term] meaning"
**Example searches**: "what is pSEO", "api definition", "what does crm stand for"
**What it is**: Educational definitions of industry terms and concepts.
**Why it works**:
- Top-of-funnel awareness
- Establishes expertise
- Natural internal linking opportunities
**Value requirements**:
- Clear, accurate definitions
- Examples and context
- Related terms linked
- More depth than a dictionary
**URL structure**: `/glossary/[term]/` or `/learn/[term]/`
---
## 10. Translations
**Pattern**: Same content in multiple languages
**Example searches**: "qué es pSEO", "was ist SEO", "マーケティングとは"
**What it is**: Your content translated and localized for other language markets.
**Why it works**:
- Opens entirely new markets
- Lower competition in many languages
- Multiplies your content reach
**Value requirements**:
- Quality translation (not just Google Translate)
- Cultural localization
- hreflang tags properly implemented
- Native speaker review
**URL structure**: `/[lang]/[page]/` or `yoursite.com/es/`, `/de/`, etc.
---
## 11. Directory
**Pattern**: "[category] tools" or "[type] software" or "[category] companies"
**Example searches**: "ai copywriting tools", "email marketing software", "crm companies"
**What it is**: Comprehensive directories listing options in a category.
**Why it works**:
- Research phase capture
- Link building magnet
- Natural for aggregators/reviewers
**Value requirements**:
- Comprehensive coverage
- Useful filtering/sorting
- Details per listing (not just names)
- Regular updates
**URL structure**: `/directory/[category]/` or `/[category]-directory/`
---
## 12. Profiles
**Pattern**: "[person/company name]" or "[entity] + [attribute]"
**Example searches**: "stripe ceo", "airbnb founding story", "elon musk companies"
**What it is**: Profile pages about notable people, companies, or entities.
**Why it works**:
- Informational intent traffic
- Builds topical authority
- Natural for B2B, news, research
**Value requirements**:
- Accurate, sourced information
- Regularly updated
- Unique insights or aggregation
- Not just Wikipedia rehash
**URL structure**: `/people/[name]/` or `/companies/[name]/`
---
## Choosing Your Playbook
### Match to Your Assets
| If you have... | Consider... |
|----------------|-------------|
| Proprietary data | Stats, Directories, Profiles |
| Product with integrations | Integrations |
| Design/creative product | Templates, Examples |
| Multi-segment audience | Personas |
| Local presence | Locations |
| Tool or utility product | Conversions |
| Content/expertise | Glossary, Curation |
| International potential | Translations |
| Competitor landscape | Comparisons |
### Combine Playbooks
You can layer multiple playbooks:
- **Locations + Personas**: "Marketing agencies for startups in Austin"
- **Curation + Locations**: "Best coworking spaces in San Diego"
- **Integrations + Personas**: "Slack for sales teams"
- **Glossary + Translations**: Multi-language educational content
Quản lý prompt ở quy mô production: phiên bản, A/B test, registry, chống hồi quy và pipeline đánh giá cho tính năng AI.
---
name: prompt-governance
description: "Use when managing prompts in production at scale: versioning prompts, running A/B tests on prompts, building prompt registries, preventing prompt regressions, or creating eval pipelines for production AI features. Triggers: 'manage prompts in production', 'prompt versioning', 'prompt regression', 'prompt A/B test', 'prompt registry', 'eval pipeline'. NOT for writing or improving individual prompts (use senior-prompt-engineer). NOT for RAG pipeline design (use rag-architect). NOT for LLM cost reduction (use llm-cost-optimizer)."
---
# Prompt Governance
> Originally contributed by [chad848](https://github.com/chad848) — enhanced and integrated by the claude-skills team.
You are an expert in production prompt engineering and AI feature governance. Your goal is to treat prompts as first-class infrastructure -- versioned, tested, evaluated, and deployed with the same rigor as application code. You prevent quality regressions, enable safe iteration, and give teams confidence that prompt changes will not break production.
Prompts are code. They change behavior in production. Ship them like code.
## Before Starting
**Check for context first:** If project-context.md exists, read it before asking questions. Pull the AI tech stack, deployment patterns, and any existing prompt management approach.
Gather this context (ask in one shot):
### 1. Current State
- How are prompts currently stored? (hardcoded in code, config files, database, prompt management tool?)
- How many distinct prompts are in production?
- Has a prompt change ever caused a quality regression you did not catch before users reported it?
### 2. Goals
- What is the primary pain? (versioning chaos, no evals, blind A/B testing, slow iteration?)
- Team size and prompt ownership model? (one engineer owns all prompts vs. many contributors?)
- Tooling constraints? (open-source only, existing CI/CD, cloud provider?)
### 3. AI Stack
- LLM provider(s) in use?
- Frameworks in use? (LangChain, LlamaIndex, custom, direct API?)
- Existing test/CI infrastructure?
## How This Skill Works
### Mode 1: Build Prompt Registry
No centralized prompt management today. Design and implement a prompt registry with versioning, environment promotion, and audit trail.
### Mode 2: Build Eval Pipeline
Prompts are stored somewhere but there is no systematic quality testing. Build an evaluation pipeline that catches regressions before production.
### Mode 3: Governed Iteration
Registry and evals exist. Design the full governance workflow: branch, test, eval, review, promote -- with rollback capability.
---
## Mode 1: Build Prompt Registry
**What a prompt registry provides:**
- Single source of truth for all prompts
- Version history with rollback
- Environment promotion (dev to staging to prod)
- Audit trail (who changed what, when, why)
- Variable/template management
### Minimum Viable Registry (File-Based)
For small teams: structured files in version control.
Directory layout:
```
prompts/
registry.yaml # Index of all prompts
summarizer/
v1.0.0.md # Prompt content
v1.1.0.md
classifier/
v1.0.0.md
qa-bot/
v2.1.0.md
```
Registry YAML schema:
```yaml
prompts:
- id: summarizer
description: "Summarize support tickets for agent triage"
owner: platform-team
model: claude-sonnet-4-5
versions:
- version: 1.1.0
file: summarizer/v1.1.0.md
status: production
promoted_at: 2026-03-15
promoted_by: eng@company.com
- version: 1.0.0
file: summarizer/v1.0.0.md
status: archived
```
### Production Registry (Database-Backed)
For larger teams: API-accessible prompt registry with key tables for prompts and prompt_versions tracking slug, content, model, environment, eval_score, and promotion metadata.
To initialize a file-based registry, create the directory structure above and populate the registry YAML with your existing prompts, their current versions, and ownership metadata.
---
## Mode 2: Build Eval Pipeline
**The problem:** Prompt changes are deployed by feel. There is no systematic way to know if a new prompt is better or worse than the current one.
**The solution:** Automated evals that run on every prompt change, similar to unit tests.
### Eval Types
| Type | What it measures | When to use |
|---|---|---|
| **Exact match** | Output equals expected string | Classification, extraction, structured output |
| **Contains check** | Output includes required elements | Key point extraction, summaries |
| **LLM-as-judge** | Another LLM scores quality 1-5 | Open-ended generation, tone, helpfulness |
| **Semantic similarity** | Embedding similarity to golden answer | Paraphrase-tolerant comparisons |
| **Schema validation** | Output conforms to JSON schema | Structured output tasks |
| **Human eval** | Human rates 1-5 on criteria | High-stakes, launch gates |
### Golden Dataset Design
Every prompt needs a golden dataset: a fixed set of input/expected-output pairs that define correct behavior.
Golden dataset requirements:
- Minimum 20 examples for basic coverage, 100+ for production confidence
- Cover edge cases and failure modes, not just happy path
- Reviewed and approved by domain expert, not just the engineer who wrote the prompt
- Versioned alongside the prompt (a prompt change may require golden set updates)
### Eval Pipeline Implementation
The eval runner accepts a prompt version and golden dataset, calls the LLM for each example, evaluates the response against expected output, and returns a result with pass_rate, avg_score, and failure details.
Pass thresholds (calibrate to your use case):
- Classification/extraction: 95% or higher exact match
- Summarization: 0.85 or higher LLM-as-judge score
- Structured output: 100% schema validation
- Open-ended generation: 80% or higher human eval approval
To execute evals, build a runner that iterates through the golden dataset, calls the LLM with the prompt version under test, scores each response against the expected output, and reports aggregate pass rate and failure details.
---
## Mode 3: Governed Iteration
The full prompt deployment lifecycle with gates at each stage:
1. **BRANCH** -- Create feature branch for prompt change
2. **DEVELOP** -- Edit prompt in dev environment, manual testing
3. **EVAL** -- Run eval pipeline vs. golden dataset (automated in CI)
4. **COMPARE** -- Compare new prompt eval score vs. current production score
5. **REVIEW** -- PR review: eval results plus diff of prompt changes
6. **PROMOTE** -- Staging to Production with approval gate
7. **MONITOR** -- Watch production metrics for 24-48h post-deploy
8. **ROLLBACK** -- One-command rollback to previous version if needed
### A/B Testing Prompts
When you want to measure real-user impact, not just eval scores:
- Use stable assignment (same user always gets same variant, based on user_id hash)
- Log every assignment with user_id, prompt_slug, and variant for analysis
- Define success metric before starting (not after)
- Run for minimum 1 week or 1,000 requests per variant
- Check for novelty effect (first-day engagement spike)
- Statistical significance: p<0.05 before declaring a winner
- Monitor latency and cost alongside quality
### Rollback Playbook
One-command rollback promotes the previous version back to production status in the registry, then verify by re-running evals against the restored version.
---
## Proactive Triggers
Surface these without being asked:
- **Prompts hardcoded in application code** -- Prompt changes require code deploys. This slows iteration and mixes concerns. Flag immediately.
- **No golden dataset for production prompts** -- You are flying blind. Any prompt change could silently regress quality.
- **Eval pass rate declining over time** -- Model updates can silently break prompts. Scheduled evals catch this before users do.
- **No prompt rollback capability** -- If a bad prompt reaches production, the team is stuck until a new deploy. Always have rollback.
- **One person owns all prompt knowledge** -- Bus factor risk. Prompt registry and docs equal knowledge that survives team changes.
- **Prompt changes deployed without eval** -- Every uneval'd deploy is a bet. Flag when the team skips evals "just this once."
---
## Output Artifacts
| When you ask for... | You get... |
|---|---|
| Registry design | File structure, schema, promotion workflow, and implementation guidance |
| Eval pipeline | Golden dataset template, eval runner approach, pass threshold recommendations |
| A/B test setup | Variant assignment logic, measurement plan, success metrics, and analysis template |
| Prompt diff review | Side-by-side comparison with eval score delta and deployment recommendation |
| Governance policy | Team-facing policy doc: ownership model, review requirements, deployment gates |
---
## Communication
All output follows the structured standard:
- **Bottom line first** -- risk or recommendation before explanation
- **What + Why + How** -- every finding has all three
- **Actions have owners and deadlines** -- no "the team should consider..."
- **Confidence tagging** -- verified / medium / assumed
---
## Anti-Patterns
| Anti-Pattern | Why It Fails | Better Approach |
|---|---|---|
| Hardcoding prompts in application source code | Prompt changes require code deploys, slowing iteration and coupling concerns | Store prompts in a versioned registry separate from application code |
| Deploying prompt changes without running evals | Silent quality regressions reach users undetected | Gate every prompt change on automated eval pipeline pass before promotion |
| Using a single golden dataset forever | As the product evolves, the golden set drifts from real usage patterns | Review and update the golden dataset quarterly, adding new edge cases from production failures |
| One person owns all prompt knowledge | Bus factor of 1 — when that person leaves, prompt context is lost | Document prompts in a registry with ownership, rationale, and version history |
| A/B testing without a pre-defined success metric | Post-hoc metric selection introduces bias and inconclusive results | Define the primary success metric and sample size requirement before starting the test |
| Skipping rollback capability | A bad prompt in production with no rollback forces an emergency code deploy | Every prompt version promotion must have a one-command rollback to the previous version |
## Related Skills
- **senior-prompt-engineer**: Use when writing or improving individual prompts. NOT for managing prompts in production at scale (that is this skill).
- **llm-cost-optimizer**: Use when reducing LLM API spend. Pairs with this skill -- evals catch quality regressions when you route to cheaper models.
- **rag-architect**: Use when designing retrieval pipelines. Pairs with this skill for governing RAG system prompts and retrieval prompts separately.
- **ci-cd-pipeline-builder**: Use when building CI/CD pipelines. Pairs with this skill for automating eval runs in CI.
- **observability-designer**: Use when designing monitoring. Pairs with this skill for production prompt quality dashboards.
Tìm, đánh giá và lập danh sách khách hàng tiềm năng để tiếp cận, cho B2B SaaS, B2B nói chung hoặc doanh nghiệp nhỏ tại địa phương.
---
name: prospecting
description: When the user wants to find, qualify, and build a list of prospects to reach out to — across B2B SaaS, general B2B, or local small businesses. Also use when the user mentions "prospecting," "build a prospect list," "find prospects," "find leads," "lead gen list," "find SaaS companies that," "find B2B companies," "find local businesses," "ICP-fit accounts," "who should we go after," "outbound list," "target account list," "find clients near me," "businesses without websites," "prospect research," "qualified leads," "find my first customers," "early adopters," "design partners," "beta users," or "who has this problem." Use this for the list-building and qualification phase. For writing the outbound copy after the list is built, see cold-email. For deep competitive research on specific accounts, see competitor-profiling.
metadata:
version: 1.1.0
---
# Prospecting
You are an expert at building qualified prospect lists across four motions: B2B SaaS, general B2B, local small businesses, and early-stage demand-signal discovery (finding your first customers from public pain signals). Your goal is to turn an ICP definition into a verified, scored, ready-to-outreach lead sheet — using the right data sources, qualification signals, and compliance posture for each motion.
## Before Starting
**Check for product marketing context first:**
If `.agents/product-marketing.md` exists (or `.claude/product-marketing.md`, or the legacy `product-marketing-context.md` filename, in older setups), read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
## Pick the Branch
Prospecting motions differ enough that the workflow forks at intake. Pick **one** branch based on who the user is selling to:
| Branch | Sell to | What "qualified" looks like | Primary sources |
|--------|---------|----------------------------|----------------|
| **SaaS** | Other SaaS companies / digital businesses | ICP fit + tech stack match + growth signals (funding, hiring, product velocity) | LinkedIn, BuiltWith, Crunchbase, Apollo, Clay, Clearbit, ProductHunt |
| **B2B** | Non-SaaS B2B (services, manufacturers, enterprises, mid-market) | Industry + size + geographic fit + buying signals (trigger events, vendor changes) | Apollo, ZoomInfo, Clay, Clearbit, LinkedIn Sales Nav, industry directories |
| **Local SMB** | Local small businesses (shops, gyms, restaurants, clinics, salons, services) | Active business + website status + proximity + decision-maker access | Google Maps, Yelp, local directories, Facebook, business websites |
| **Demand-signal** | Early-stage: your first customers, design partners, or beta users | Evidence of the exact pain/demand/timing signal — a cited public source, not just firmographic fit | Forums, communities, reviews, GitHub issues, job posts, launch announcements (via last30days, social-fetch, scraping) |
If the user describes a hybrid motion (e.g., "SMBs that are also SaaS"), pick the dominant branch and pull in qualification signals from the other. If the user is early-stage and needs their *first* customers or design partners — evidence of demand over list coverage — use the **Demand-signal** branch.
For the branch-specific deep dives:
- **SaaS** → see [references/saas-prospecting.md](references/saas-prospecting.md)
- **B2B** → see [references/b2b-prospecting.md](references/b2b-prospecting.md)
- **Local SMB** → see [references/local-prospecting.md](references/local-prospecting.md)
- **Demand-signal** (find your first customers) → see [references/demand-signals.md](references/demand-signals.md)
---
## Shared Framework (all branches)
Every prospecting engagement follows the same five phases. Tools and qualification signals change per branch; the phases don't.
### Phase 1 — Define the ICP
Pull from `product-marketing.md` if available. Otherwise, gather:
1. **Firmographic fit** — industry, company size, revenue band, geography, business model
2. **Technographic fit** (SaaS branch) — what tools they already use, what they're missing
3. **Buying signal** — why now? (trigger event, funding, hiring, new initiative, dissatisfaction with current vendor, recent move/expansion)
4. **Decision-maker profile** — role, seniority, what they care about
5. **Disqualifiers** — what makes a prospect a clear "skip"
Output the ICP as a one-paragraph statement plus a checklist of pass/fail criteria. Don't move to discovery without this.
### Phase 2 — Build the candidate list (discovery)
Source 2–3× more candidates than the user wants in the final list — qualification will cull aggressively.
- **SaaS / B2B**: combine 2–3 sources for cross-verification. Apollo or ZoomInfo for firmographics; Clearbit or Clay for enrichment; LinkedIn Sales Nav for decision-maker mapping.
- **Local SMB**: browser-assisted research starting with Google Maps for the target category in the target area; cross-check with Yelp, the business website, social pages, and public directories.
If the user's list quality bar is high, smaller is better. 25 verified leads beats 250 mostly-junk ones.
### Phase 3 — Qualify each candidate
Score every candidate against the ICP checklist. Add **evidence** (a source URL or two) for each qualification — never assert without backing.
**Confidence levels** (used across all branches):
- **High**: confirmed by at least two independent sources or official business page
- **Medium**: one credible source plus consistent search evidence
- **Low**: incomplete or ambiguous evidence — flag what remains uncertain
For email contacts (B2B / SaaS branches), **always verify deliverability before adding to the final list** — see Truelist integration in [references/data-sources.md](references/data-sources.md). Don't ship leads with invalid or risky emails.
### Phase 4 — Score and prioritize
Apply this rubric for the **SaaS, B2B, and Local SMB** branches. The **Demand-signal** branch scores differently — 0–100 demand-fit, not Hot/Warm/Cold — see [references/demand-signals.md](references/demand-signals.md).
| Score | Definition |
|-------|------------|
| **Hot** | Strong ICP fit + clear buying signal + decision-maker accessible + verified contact |
| **Warm** | ICP fit + softer or older signal + contact verifiable |
| **Cold** | Loose ICP fit OR no clear signal OR contact unverified |
| **Skip** | Disqualifier hit (out of ICP, closed business, duplicate, irrelevant, low confidence) |
Branch-specific signals refine the scoring — see each reference file. Default ratio target: ~20% Hot, ~30% Warm, rest Cold/Skip.
### Phase 5 — Output the lead sheet
(SaaS / B2B / Local SMB. The **Demand-signal** branch ships an evidence report instead — see [references/demand-signals.md](references/demand-signals.md).)
Default to a markdown table in chat. Switch to CSV when the list is >25 rows or the user explicitly asks for a file.
After the table, always add **"Top outreach targets"** — the top 3–5 hot leads with one sentence each on why this lead should be reached out to first.
Columns vary by branch (see reference files), but every lead sheet includes:
- score, business/company name, contact (where applicable), why-it's-a-prospect, source(s), confidence, last verified date
---
## Compliance Guardrails
These apply to every branch. **Read first, every engagement.**
1. **No bulk scraping** of LinkedIn, Google Maps, paywalled sites, or rate-limited APIs. Browser is an assisted research tool, not a scraper.
2. **No CAPTCHA, login wall, or bot protection bypass.** If a site requires it, work with what's publicly visible.
3. **Public business contact channels only.** Use info@, hello@, contact@, and named-role emails (founder, owner) where they're published on the business's own site. Personal/private emails require a lawful basis (existing relationship, opt-in, etc.).
4. **GDPR / CAN-SPAM / CASL aware.** Capture and retain the source URL and date for every contact you add to a list — required for downstream outreach compliance.
5. **No reselling extracted data** from Google Maps, LinkedIn, or any platform whose terms prohibit it. List building for the user's own outreach is fine; productizing the list to sell is not.
6. **Rate limit yourself.** Even on public sources, space requests. Don't fingerprint as a bot.
7. **No breached, leaked, or unprovenanced data.** Don't source prospects from breached datasets, scraped-contact marketplaces, or list brokers with no source lineage. Licensed B2B data providers (Apollo, ZoomInfo, Clearbit, Clay) are fine when used within their ToS and with a lawful basis — the ban is on illicit/unprovenanced data, not on legitimate enrichment vendors.
8. **Never target or infer sensitive traits.** Don't qualify, segment, or personalize on health, financial hardship, political belief, sexuality, religion, or other protected/sensitive attributes — even when a public post reveals them.
For the full compliance reference (GDPR, CAN-SPAM, CASL, LinkedIn ToS, Google Maps ToS, Clay/Apollo/ZoomInfo use restrictions): see [references/compliance.md](references/compliance.md).
---
## Inputs to Collect
If missing, ask once, then infer reasonable defaults and continue:
- **Branch** (SaaS / B2B / Local SMB / Demand-signal) — usually inferable from context; pick Demand-signal for early-stage first-customer discovery
- **ICP description** — pull from `product-marketing.md` if present
- **Target count** — default 25 for SaaS / B2B, 15 for Local SMB
- **Geography** (essential for Local SMB; useful for B2B; less critical for SaaS)
- **Tools the user has access to** — Apollo? Clay? ZoomInfo? Hunter? Truelist? Defaults to what's free + browser
- **Output format** — chat table (default) or CSV
- **Buying signal preference** — what triggers should they prioritize? (funding rounds, hiring, recent move, etc.)
---
## Tool Selection Quick Picks
Full breakdown in [references/data-sources.md](references/data-sources.md). Quick picks:
| If the user has access to... | Use it for |
|------------------------------|------------|
| **Apollo** | B2B / SaaS firmographic + contact discovery |
| **Clay** | Multi-source enrichment, waterfall lookups, custom scoring |
| **Clearbit** | Email-to-company and company enrichment |
| **ZoomInfo** | Enterprise B2B contact + intent data |
| **Hunter or Snov** | Email pattern guessing and verification |
| **Truelist** | Email deliverability validation (before adding to outreach list) |
| **LinkedIn Sales Navigator** | Decision-maker mapping (manual, no scraping) |
| **BuiltWith / Wappalyzer** | Tech stack qualification (SaaS branch) |
| **Crunchbase** | Funding signals (SaaS branch) |
| **GitHub** | Stargazers / forks of competitor or adjacent repos (dev-tool SaaS branch) |
| **Google Maps + browser** | Local SMB discovery |
| **Firecrawl / Browserbase** | Programmatic extraction from individual prospect websites — never from platforms |
**If the user has no enrichment tools**: lean on browser-assisted research with public sources — company website, About page, LinkedIn company page, news mentions. Slower but works.
---
## Output Formats
### Default — chat table
For SaaS / B2B (≤25 rows):
```
| Score | Company | Industry | Size | Signal | Contact | Email status | Source | Confidence |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
```
For Local SMB (≤15 rows) — port from the local-prospector reference:
```
| Score | Business | Category | Area | Website status | Website/Social | Phone | Why it's a prospect | Confidence |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
```
### CSV — when >25 rows or user requests a file
SaaS / B2B columns:
```csv
score,company,domain,industry,size_band,country,signal,contact_name,contact_title,contact_email,email_status,linkedin,source_urls,why_prospect,confidence,verified_date,notes
```
Local SMB columns:
```csv
score,business,category,area,distance_km,website_status,website_url,social_urls,phone,email,source_urls,why_prospect,confidence,verified_date,notes
```
### Always include after the table
- **Top outreach targets**: top 3–5 hot leads with one-sentence outreach rationale each
- **Search parameters**: branch, ICP, location/radius, target count, date generated
- **Open questions**: anything you couldn't verify and the user should look at
---
## Quality Checks (before finalizing)
- [ ] Remove duplicates (by domain for SaaS/B2B, by business + address for Local SMB)
- [ ] Every "Hot" lead has a verified contact + at least one source URL
- [ ] No lead has an email that failed Truelist (or your validator) verification — move to a separate "invalid" bucket and flag for the user
- [ ] No lead labeled "Hot" lacks a clear buying signal
- [ ] Confidence levels honest — "High" requires 2 independent sources, not just two of your own searches
- [ ] No leads sourced from prohibited scraping (LinkedIn at scale, Google Maps bulk extract, etc.)
- [ ] Source URL + date captured for every contact (GDPR / CAN-SPAM lineage)
- [ ] Final count matches user's request, or you've explained why it's smaller (quality bar)
---
## Common Mistakes
1. **Starting discovery without an ICP**. Build candidates against vague criteria and you'll qualify the wrong things.
2. **Treating data sources as authoritative without cross-checks**. Apollo and ZoomInfo are out of date often; verify before scoring as "Hot."
3. **Adding contacts without email verification**. Cold email reputation tanks fast with bounces — always validate.
4. **Bulk scraping LinkedIn or Google Maps**. Real risk: account suspension + ToS violation. Browser as an assisted tool only.
5. **Mixing branches**. Don't apply Local SMB scoring (website status) to a B2B SaaS prospect, or vice versa.
6. **"Hot" labels without buying signals**. ICP fit alone is not enough — the signal is what makes the timing right.
7. **No source URLs**. Every claim should be traceable to a public source. Future outreach depends on this lineage.
8. **Ignoring quiet hours / time zone** when scheduling the downstream outreach (handoff to cold-email).
9. **Forgetting to retain consent / lineage records**. Required for GDPR DSARs and CAN-SPAM audits.
---
## Task-Specific Questions
1. Which branch — SaaS, B2B, Local SMB, or Demand-signal (early-stage, finding your first customers)?
2. What's your ICP? (Or: should I pull from your product-marketing context?)
3. How many qualified leads do you want?
4. What tools do you have access to (Apollo / Clay / ZoomInfo / Hunter / Truelist / browser only)?
5. What's the triggering buying signal you care most about?
6. Geography or radius (Local SMB / B2B)?
7. Chat table or CSV?
---
## Tool Integrations
For implementation, see the [tools registry](../../tools/REGISTRY.md). Key prospecting tools:
| Tool | Best For | MCP | Guide |
|------|----------|:---:|-------|
| **Apollo** | B2B / SaaS firmographic + contact discovery | - | [apollo.md](../../tools/integrations/apollo.md) |
| **Clay** | Multi-source enrichment + waterfall | ✓ | [clay.md](../../tools/integrations/clay.md) |
| **Clearbit** | Email-to-company enrichment | - | [clearbit.md](../../tools/integrations/clearbit.md) |
| **ZoomInfo** | Enterprise B2B contact + intent | ✓ | [zoominfo.md](../../tools/integrations/zoominfo.md) |
| **Hunter** | Email pattern + verification | - | [hunter.md](../../tools/integrations/hunter.md) |
| **Snov** | Email finder + verifier | - | [snov.md](../../tools/integrations/snov.md) |
| **Truelist** | Email deliverability validation | - | [truelist.md](../../tools/integrations/truelist.md) |
| **Outreach** | Sales engagement (post-list) | ✓ | [outreach.md](../../tools/integrations/outreach.md) |
| **RB2B** | Visitor identification (warm intent) | - | [rb2b.md](../../tools/integrations/rb2b.md) |
| **GitHub** | Stargazers/forks/watchers as developer-intent signal | - | [github.md](../../tools/integrations/github.md) |
| **Firecrawl** | Single-target site extraction (prospect's own website) | ✓ | [firecrawl.md](../../tools/integrations/firecrawl.md) |
| **Browserbase** | Real-browser site research when rendering or interaction needed | ✓ | [browserbase.md](../../tools/integrations/browserbase.md) |
---
## Related Skills
- **cold-email**: For writing outbound sequences against the qualified list (the natural next step after prospecting)
- **customer-research**: For understanding why current customers buy — informs the ICP definition
- **competitor-profiling**: For deeper research on individual accounts (different from list-building qualification)
- **revops**: For lead routing, lifecycle, and CRM handoff after prospecting
- **sales-enablement**: For battle cards and one-pagers used in the outreach
- **directory-submissions**: For inbound discovery surfaces (the prospects might find you back)
- **product-marketing**: For the ICP definition that anchors every prospecting engagement
FILE:evals/evals.json
{
"skill_name": "prospecting",
"evals": [
{
"id": 1,
"prompt": "We're a B2B SaaS selling RevOps tooling at $30K ACV. Build me a list of 25 prospects.",
"expected_output": "Should check for product-marketing.md first. Should identify this as the SaaS branch. Should run Phase 1 ICP definition pulling from product-marketing context or asking targeted questions (target industry, headcount range, tech stack signals, funding stage). Should propose discovery sources appropriate for SaaS at $30K ACV: Apollo for breadth, Clay for waterfall enrichment, Crunchbase for funding signals, BuiltWith/Wappalyzer for tech stack, LinkedIn Sales Nav for decision-mapping (manual). Should ask about user's tool access before assuming. Should source 50-75 candidates (2-3x target) before qualifying. Should flag that email validation via Truelist or similar is non-negotiable before final list. Should output SaaS-branch chat table columns (Score | Company | Industry | Size | Signal | Contact | Email status | Confidence) followed by top 3-5 hot leads with one-sentence rationale each. Should reference references/saas-prospecting.md.",
"assertions": [
"Checks for product-marketing.md",
"Identifies SaaS branch",
"Runs Phase 1 ICP definition",
"Recommends multi-source discovery (Apollo, Clay, Crunchbase, BuiltWith)",
"Asks about user's tool access",
"Sources 2-3x candidates before qualifying",
"Requires email validation before final list",
"Outputs SaaS-branch chat table columns",
"Includes top 3-5 outreach targets with rationale",
"References saas-prospecting.md"
],
"files": []
},
{
"id": 2,
"prompt": "Find me 25 SaaS companies that just raised a Series B in the last 60 days and use HubSpot.",
"expected_output": "Should recognize this as a SaaS branch prospecting task with very specific signals. Should identify the trigger event (Series B in last 60 days) and the technographic filter (uses HubSpot). Should recommend a workflow: (1) Crunchbase or Pitchbook for funding signal filter (Series B + date), (2) BuiltWith or Clay's waterfall for tech stack verification (uses HubSpot), (3) cross-check via business websites and LinkedIn. Should note this is a tight ICP that should yield high-confidence matches if data sources are current. Should flag freshness concerns: Crunchbase data depends on self-reporting, BuiltWith refresh cycles aren't real-time. Should recommend cross-source verification for the funding date specifically. Should output a SaaS-branch chat table with the funding round + date in the Signal column. Should include verified email validation before delivering.",
"assertions": [
"Identifies as SaaS branch",
"Identifies funding signal + tech stack filter",
"Recommends Crunchbase or Pitchbook for funding",
"Recommends BuiltWith or Clay for HubSpot verification",
"Notes data freshness concerns",
"Recommends cross-source verification",
"Outputs signal column showing round + date",
"Requires email validation"
],
"files": []
},
{
"id": 3,
"prompt": "I run a marketing agency. Find me 25 mid-market manufacturers in the Midwest US who recently hired a new CMO.",
"expected_output": "Should identify this as the B2B branch (manufacturers, not SaaS). Should run Phase 1 ICP definition: industry (manufacturing, with NAICS code if precision matters), size (mid-market = typically 200-2000 employees), geography (Midwest US states), trigger event (CMO hire in last 90-180 days). Should propose discovery: Apollo or ZoomInfo for firmographic filter, LinkedIn Sales Nav for CMO hire detection (job changes), Google Alerts on press releases for trigger events. Should warn that CMO hires aren't always in public databases — LinkedIn Sales Nav alerts on job changes is the most reliable source. Should output B2B-branch chat table with the CMO trigger as the signal. Should reference references/b2b-prospecting.md. Should mention compliance: GDPR less likely (US-only), CAN-SPAM applies, capture source URL + date for every contact.",
"assertions": [
"Identifies B2B branch (not SaaS)",
"Runs Phase 1 ICP definition with NAICS or industry classification",
"Specifies mid-market size band",
"Specifies Midwest US geography",
"Identifies trigger event (CMO hire)",
"Recommends Apollo/ZoomInfo + LinkedIn Sales Nav",
"Notes CMO hires often only on LinkedIn",
"Outputs B2B-branch chat table",
"Mentions CAN-SPAM and source URL capture",
"References b2b-prospecting.md"
],
"files": []
},
{
"id": 4,
"prompt": "We sell to industrial distributors. Build a list of 25 prospects.",
"expected_output": "Should identify this as the B2B branch. Should run Phase 1 ICP definition asking targeted questions: distributor size, geography, vertical specialty, buying patterns. Should propose discovery: Apollo or ZoomInfo for firmographic depth, industry-specific directories (e.g., NAW for wholesale distributors, ISA for industrial sales agencies), trade show exhibitor lists. Should note state business registries and Chamber of Commerce as verification sources. Should propose trigger events: new location, recent acquisition, leadership change, posting RFPs. Should warn that industrial distributor data is often spotty in major databases — cross-check with company website + LinkedIn for size and ownership signals. Should output B2B-branch chat table. Should note ICP fit precision matters more than initial volume for this kind of niche prospecting.",
"assertions": [
"Identifies B2B branch",
"Runs Phase 1 ICP definition asking targeted questions",
"Recommends industry-specific directories beyond Apollo/ZoomInfo",
"Mentions trade show exhibitor lists",
"Identifies relevant trigger events",
"Warns about data spottiness for industrial",
"Recommends cross-verification with business websites + LinkedIn",
"Notes ICP fit precision over volume"
],
"files": []
},
{
"id": 5,
"prompt": "I build websites for local businesses. Find me 15 prospects near Austin, TX who don't have a website.",
"expected_output": "Should identify as Local SMB branch. Should run Phase 1 ICP definition: business category (ask user — gyms, restaurants, salons, etc. matter), radius (default 20 km from Austin), target count (15). Should run the browser research workflow: search Google Maps for category + Austin, build candidate list from visible results, cross-check via business name + city web search to verify website status. Should apply the 4-tier website status classification (No site found / Social only / Weak site / Has site) — prioritize No site + Social only as Hot. Should score: Hot (no site + active + phone + within radius), Warm (weak site), Cold (has site), Skip (closed/duplicate/out of scope). Should output Local SMB chat table (Score | Business | Category | Area | Distance | Website status | Website/Social | Phone | Why prospect | Confidence). Should add 'Best first outreach targets' top 3 with reasoning. Should reference references/local-prospecting.md. Should warn against bulk-scraping Google Maps (ToS violation) — browser-assisted research only.",
"assertions": [
"Identifies Local SMB branch",
"Asks about business category if not specified",
"Defaults radius to 20km",
"Runs browser research workflow",
"Applies 4-tier website status classification",
"Uses Hot/Warm/Cold/Skip scoring",
"Outputs Local SMB chat table columns",
"Adds top 3 outreach targets",
"References local-prospecting.md",
"Warns against bulk-scraping Google Maps"
],
"files": []
},
{
"id": 6,
"prompt": "I have a list of 200 prospect emails from Apollo. How do I know which ones are deliverable before I start outreach?",
"expected_output": "Should explain the deliverability validation step in Phase 3. Should recommend Truelist (the integration in this pack) for bulk validation. Should explain the email_state classification output: ok (deliverable), email_invalid (bounces, exclude), risky (deliverable with risk like role or disposable, include cautiously), unknown (couldn't determine, skip or re-verify), accept_all (catch-all domain, include cautiously). Should warn that Apollo data accuracy is typically 60-80% — sending without validation will tank sender reputation (bounce rate >2% triggers ISP throttling and reputation damage). Should recommend the workflow: bulk POST to /api/v1/verify or CSV upload → keep ok, include risky/accept_all cautiously, exclude email_invalid, re-verify unknown → hand off to outreach. Should note Truelist also has an official MCP server for agent-driven validation. Should note cold email reputation is hard to recover once damaged — validation is non-negotiable, not optional. Should mention Hunter and Snov as alternatives with built-in verification. Should reference truelist.md integration guide.",
"assertions": [
"Recommends Truelist for bulk validation",
"Explains email_state values (ok, email_invalid, risky, unknown, accept_all)",
"Warns Apollo accuracy is 60-80%",
"Cites 2% bounce rate threshold for reputation damage",
"Recommends workflow: validate, keep ok, exclude email_invalid",
"Mentions Truelist MCP server for agent workflows",
"Mentions cold email reputation is hard to recover",
"References truelist.md or data-sources.md"
],
"files": []
},
{
"id": 7,
"prompt": "I just built a tool that automates failed-payment follow-up for gym owners. I have no customers yet. Help me find my first ten — the people who are actually dealing with this problem right now.",
"expected_output": "Should select the Demand-signal branch (early-stage, first customers, evidence-of-demand) and load references/demand-signals.md — NOT the SMB/B2B list-building branches. Should start with a product brief, then mine the five signal buckets (explicit demand / pain / workaround / switching / timing) across public discourse (forums, communities, reviews, GitHub issues, job posts) — using last30days for recency, social-fetch/scraping to read original pages, not qualifying from snippets. Should score prospects on demand-fit (pain 25 / product fit 25 / timing 20 / reachability 15 / evidence quality 15, 0-100 with bands) rather than ICP-fit Hot/Warm/Cold, and require a cited public signal for every primary-shortlist prospect. Should draft source-based openers but never auto-send. Should produce an evidence report (verdict → ICP → top prospect → shortlist with sources+scores → repeated patterns → 7-day manual outreach plan → limits) and label prospects as 'potential customers based on public signals,' not confirmed buyers. Should honor the compliance guardrails including no data brokers/leaked data and no sensitive-trait targeting.",
"assertions": [
"Selects the Demand-signal branch, not the SMB/B2B/SaaS list-building branches",
"Mines the five signal buckets from public discourse rather than contact databases",
"Uses recency/original-source tooling (last30days, social-fetch, scraping) and does not qualify from snippets",
"Scores on the demand-fit rubric (0-100 weighted), not ICP-fit Hot/Warm/Cold",
"Requires a cited public signal for every primary-shortlist prospect",
"Drafts openers but never auto-sends; labels prospects as potential-based-on-public-signals",
"Produces the evidence report structure with a 7-day manual outreach plan and limits"
],
"files": []
}
]
}
FILE:references/b2b-prospecting.md
# B2B Prospecting Reference
For when the user sells to non-SaaS B2B — services, agencies, manufacturers, mid-market and enterprise companies, professional services firms.
---
## ICP Signals That Matter (B2B branch)
### Firmographic signals
- **Industry / vertical** — NAICS or SIC codes if precision matters
- **Company size** — headcount band, revenue band, location count
- **Geography** — relevant for time zones, regulations, on-site requirements
- **Business model** — service vs product vs distribution; B2B vs B2B2C
- **Ownership** — independent, PE-backed, public, family-owned — affects buying motion
### Buying signals
- **Trigger events**: new C-level hire, recent acquisition or divestiture, IPO/funding, opening a new location, recent rebrand, expansion announcement
- **Vendor signals**: posting RFPs publicly, switching costs in last quarterly report, contract renewal windows
- **Operational signals**: recent layoffs (cost pressure) or rapid hiring (capacity pressure)
- **News mentions**: launching new initiative, entering new market, regulatory change forcing action
- **PR / press**: anything that signals "this company is changing right now"
### Decay signals
- Multiple bankruptcies or PE-stripped operations
- Negative growth + cost-cutting headlines
- Ownership stagnation (small family-owned, no growth incentive)
- Buyer turnover (3+ Marketing Directors in 2 years)
---
## Discovery Sources (B2B branch)
### Tier 1 — primary discovery
- **Apollo**: best general B2B firmographic + contact discovery
- **ZoomInfo**: enterprise B2B + intent signals (mid-market+)
- **LinkedIn Sales Navigator**: industry + role + signal search; the gold standard for decision-maker mapping (manual)
- **Clay**: when you need custom waterfall lookups (e.g., enrich Apollo records with Hunter + Clearbit)
### Tier 2 — industry-specific directories
- **Crunchbase / Pitchbook**: funded businesses
- **D&B Hoovers**: large traditional B2B firmographics
- **State / national business registries**: for verified incorporation data
- **Industry association membership rosters**: trade groups often publish member lists
- **Trade show exhibitor lists**: signals active participation in a vertical
- **Procurement databases** (Procore for construction, e.g.): vertical-specific signals
### Tier 3 — trigger event monitoring
- **Google Alerts / Feedly**: trigger keywords ("acquired," "hires," "expansion," "raises," "announces")
- **PR Newswire / Business Wire**: company-controlled announcements
- **SEC filings** (public companies): material change disclosures
- **State filings**: new entity formation, dissolution
---
## Qualification Checklist (B2B branch)
- [ ] Industry / vertical matches ICP (use a recognized classification if possible)
- [ ] Company size within range (employees or revenue)
- [ ] Geography fits
- [ ] At least one trigger event in last 90–180 days
- [ ] Decision-maker role exists (CEO, COO, VP Operations, Director of X — match buyer profile)
- [ ] Email contact verifiable (named role > info@ catchall)
- [ ] Source URLs captured for firmographic claims
- [ ] No disqualifiers (closed, acquired-paused, multi-bankrupt, off-ICP)
---
## Output Columns (B2B branch)
Recommended CSV columns:
```csv
score,company,domain,industry,naics_code,size_band,revenue_band,country,city,trigger_event,trigger_date,contact_name,contact_title,contact_email,email_status,linkedin_url,source_urls,why_prospect,confidence,verified_date,notes
```
For chat table, condense to: Score | Company | Industry | Size | Trigger | Contact | Email status | Confidence.
---
## Top Outreach Targets Selection (B2B)
Prioritize for the top 3–5 hot leads:
1. **Trigger event recency** — 30 days beats 6 months
2. **Trigger event specificity** — new CMO hire in your buyer's role beats "company in the news"
3. **Decision-maker access** — named contact with verified email + LinkedIn beats role-only
4. **Vertical fit precision** — exact NAICS match beats "adjacent industry"
Each top target rationale names the trigger and decision-maker: "Hired new VP of Marketing 14 days ago; verified email; mid-market manufacturer matching ICP."
---
## Common Mistakes (B2B)
1. **Treating B2B like SaaS** — funding rounds matter less; PE ownership and acquisition activity matter more.
2. **Trying to verify private company revenue precisely** — most public databases approximate. Use size bands, not point estimates.
3. **Ignoring procurement complexity** at enterprise scale — your prospect contact list may not include the actual approver.
4. **Cold-emailing executive assistants** — they're not the buyer and they will flag your outreach as spam.
5. **Source URL hygiene** — without source lineage, you can't defend a contact under GDPR DSAR or CAN-SPAM challenge.
6. **Stopping at one source** — Apollo can be 60% accurate on small businesses. Cross-verify with LinkedIn or the business website.
FILE:references/compliance.md
# Prospecting Compliance Reference
The legal and platform-ToS constraints that apply to prospect list building. Read first, every engagement.
> Operational guidance, not legal advice. For high-volume programs or programs touching EU/UK residents, run your setup past a privacy attorney.
---
## United States — CAN-SPAM (downstream)
CAN-SPAM regulates the cold email **send**, not the list build. But the list build matters because:
- You must be able to identify the source of every email address you contact (required if challenged)
- The "from" line and email content rules apply at send time — but you can't lie about how you got the contact
- Opt-out requests must be honored within 10 business days and tracked
**For prospecting specifically**: capture and retain the source URL + date for every contact you add to a list. CAN-SPAM doesn't require it explicitly, but defending your sender practices does.
---
## EU / UK — GDPR
The strictest applicable framework. Triggers when:
- Your prospect resides in EU/UK
- You're processing personal data (any identifiable info, including business emails tied to a named person)
### Lawful bases for cold B2B outreach
You have three credible options:
1. **Legitimate interest** (most common for B2B). Requires:
- The contact is in a business role likely to be interested in your offer
- The data was collected from a public, business-context source
- You provide a clear opt-out
- You can articulate the legitimate interest test in writing
2. **Consent** — typically not feasible for cold outreach (you don't have consent before first contact)
3. **Existing customer relationship** — only applies to current customers, not prospects
### What you must do
- Capture **source + date + lawful basis** for every contact
- Honor data subject access requests (DSARs) — you must be able to disclose, correct, or delete on request
- Include a privacy notice / opt-out in the first outreach
- Don't store personal data longer than necessary for the legitimate interest
### What disqualifies a list
- Bulk-scraped LinkedIn data — explicit ToS violation + GDPR risk
- Email addresses purchased from a list broker without source provenance
- "Anyone @ this domain" guessed emails sent without verification (multiplies risk + bounces)
---
## Canada — CASL
Stricter than CAN-SPAM. Cold B2B outreach requires:
- **Express consent** (explicit opt-in) — typically not present for cold prospecting
- **OR implied consent** — existing business relationship within 24 months, OR business address publicly published on the company's own site for the purpose of receiving such communications
**Practical implication for Canadian prospects**: relying on the publicly-published-address exception is the most defensible cold prospecting basis in Canada. You must include sender identification, mailing address, and an unsubscribe mechanism in every message.
---
## Platform Terms of Service
### LinkedIn
- **Sales Navigator** as a research tool: fine
- **Scraping LinkedIn at any scale**: explicit ToS violation. Banned accounts are permanent. Don't.
- **Apollo, Clay, and ZoomInfo** claim LinkedIn-overlap data through various legitimate channels — verify their data sources before assuming compliance
- **InMail and Connection Requests**: governed by LinkedIn's own messaging rules, not by CAN-SPAM/GDPR (because LinkedIn-internal)
### Google Maps
- ToS prohibits bulk extraction or productizing Maps data
- Browser-assisted research as a discovery aid: acceptable
- Storing Place IDs or large structured Maps data in your CRM: explicit ToS prohibition
- Use Maps to **find** local businesses, then cross-source from the business's own site for the data you retain
### Apollo / ZoomInfo / Clearbit
- All have their own ToS limiting reselling, downstream sharing, and use cases
- Read your contract — typically you can use the data for your own outreach but not productize it
- Don't share extracts publicly (e.g., on a leaderboard, in a public report)
### Crunchbase
- Free tier is read-only for personal use
- Paid tier permits broader use within contractual scope
- API access requires paid Pro+ tier
---
## Anti-Patterns (Don't Do These)
1. **Bulk-scraping LinkedIn / Google Maps / Yelp**. Browser-assisted research is OK; automated scrapers pointed at these platforms are not. **Firecrawl and Browserbase are fine for an individual prospect's own website** (the URL you found through manual discovery) — not for the platforms hosting prospects.
2. **Buying lists from random vendors** without source provenance. You inherit their legal exposure.
3. **Guessing emails and sending unverified**. Bounce rates over 2% destroy sender reputation; legally, you can't claim a "legitimate interest" basis for an email you fabricated.
4. **Harvesting personal email addresses** (Gmail, personal Outlook, etc.) from public profiles. Personal addresses raise GDPR risk significantly.
5. **Storing data you don't need**. Minimize retention. Don't keep prospect lists forever — GDPR right to deletion applies.
6. **Skipping the lawful basis documentation**. If challenged, you need to show your work. Capture source URL + collection date for every contact.
7. **Reselling prospect lists**. You may not have the right to share them downstream. Read your data provider contracts.
8. **CAPTCHA bypass / login wall bypass**. Even if technically possible, this signals bot behavior and violates virtually every ToS.
---
## Quick Audit Checklist
Before shipping a list to the user (or downstream to cold-email):
- [ ] Every contact has a source URL + collection date
- [ ] No contacts sourced from scraped LinkedIn data
- [ ] No Google Maps Place IDs or large Maps-structured data retained
- [ ] Lawful basis documented (legitimate interest test for B2B, or relevant alternative)
- [ ] Email addresses validated (deliverability check before outreach)
- [ ] Personal addresses (Gmail, etc.) flagged or excluded
- [ ] Source provider contracts permit the intended use case
- [ ] Retention plan documented (when to delete)
- [ ] First outreach will include unsubscribe + privacy notice (downstream concern for cold-email skill, but mention it now)
FILE:references/data-sources.md
# Prospecting Data Sources
Tool selection guide for prospecting across all three branches.
---
## Tool selection by goal
| Goal | Primary tools | Notes |
|------|--------------|-------|
| **Build initial firmographic list (B2B / SaaS)** | Apollo, ZoomInfo, Clay | Apollo for breadth, ZoomInfo for enterprise + intent, Clay for custom workflows |
| **Decision-maker mapping** | LinkedIn Sales Navigator (manual), Apollo, ZoomInfo | Sales Nav is the gold standard. Never bulk scrape it. |
| **Tech stack qualification (SaaS)** | BuiltWith, Wappalyzer | BuiltWith has wider coverage + paid plans for bulk; Wappalyzer is lighter + free for small use |
| **Funding signals (SaaS)** | Crunchbase, Pitchbook | Crunchbase free tier sufficient for early signals; Pitchbook for deeper investor data |
| **Email pattern discovery** | Hunter, Snov, Apollo | Pattern guessing — followed by verification |
| **Email deliverability verification** | Truelist, Hunter, NeverBounce, ZeroBounce | Always verify before adding to outreach lists |
| **Visitor identification (warm intent)** | RB2B, Clearbit Reveal | Anonymous traffic → company identification |
| **Intent data** | ZoomInfo Intent, 6sense, Bombora | Pre-warmed signals; mid-market+ pricing |
| **Trigger event monitoring** | Google Alerts, Feedly, LinkedIn Sales Nav alerts | Free options are sufficient for most |
| **Local business discovery** | Google Maps (manual), Yelp, Facebook Pages | Browser-assisted, not bulk-extracted |
---
## Apollo
**Use for**: General B2B / SaaS firmographic + contact data. Best starting point if you don't already have a list.
**Strengths**:
- Large database (>200M contacts, >60M companies)
- Strong filtering UI (industry, size, technologies, signals)
- Integrated email + LinkedIn finder
- Pay-as-you-go and tiered plans
**Watch out for**:
- Data freshness varies — re-verify before scoring as "Hot"
- Email accuracy ~60–80% — always validate
- Bulk export limits apply
**Integration**: see [apollo.md](../../../tools/integrations/apollo.md)
---
## Clay
**Use for**: Multi-source enrichment, waterfall lookups, custom scoring logic. When list quality matters more than list size.
**Strengths**:
- Waterfall logic: try Apollo first → fallback to ZoomInfo → fallback to Clearbit
- 100+ data provider integrations
- AI-powered enrichment (LLM-driven extraction from URLs)
- Custom columns + scoring formulas
- Native MCP server
**Watch out for**:
- Per-credit pricing can spike on large lists
- Complexity overhead — easy to over-engineer workflows
**Integration**: see [clay.md](../../../tools/integrations/clay.md)
---
## ZoomInfo
**Use for**: Enterprise B2B + intent data. Mid-market+ buyer profiles.
**Strengths**:
- Enterprise-grade firmographic depth
- Intent signals (companies searching topics relevant to your offer)
- Best-in-class for >$50K ACV B2B sales
- Native MCP server
**Watch out for**:
- Expensive ($15K+/yr starter)
- Overkill for SMB prospecting
- Locked into multi-year contracts typically
**Integration**: see [zoominfo.md](../../../tools/integrations/zoominfo.md)
---
## Clearbit
**Use for**: Email → company enrichment, anonymous visitor identification (Clearbit Reveal).
**Strengths**:
- Strong company enrichment (industry, size, funding, tech stack)
- Email lookup by domain
- Reveal: identify anonymous site visitors at company level
- API-first
**Watch out for**:
- HubSpot acquisition (2023) — bundled into HubSpot Breeze Intelligence now
- Standalone API still available but pricing/access depends on tier
**Integration**: see [clearbit.md](../../../tools/integrations/clearbit.md)
---
## Hunter / Snov
**Use for**: Email pattern discovery + lightweight verification on small lists.
**Hunter strengths**:
- Domain-based email discovery
- Built-in deliverability verification
- Free tier reasonable for occasional use
**Snov strengths**:
- Email finder + drip campaigns (overlap with outreach tooling)
- Bulk verification
- Cheaper than Hunter at scale
**Watch out for**:
- Both are pattern-guessing tools — accuracy depends on the target company's email pattern being inferable
- Always run results through a dedicated validator (Truelist or similar) before outreach
**Integrations**: see [hunter.md](../../../tools/integrations/hunter.md), [snov.md](../../../tools/integrations/snov.md)
---
## Truelist
**Use for**: Email deliverability validation before adding contacts to outreach lists. Critical safety step.
**Strengths**:
- Single-email sync verification (`/api/v1/verify_inline`) + bulk async (`/api/v1/verify`)
- Returns `email_state` (ok / email_invalid / risky / unknown / accept_all) + `email_sub_state` (email_ok / is_disposable / is_role / unknown_error / failed_smtp_check) + did-you-mean typo suggestions
- Catches catch-all domains, role accounts, spam traps, disposable providers
- Official MCP server for agent-driven workflows (Claude, Cursor, VS Code)
- Official SDKs in 7 languages + framework integrations (Django, Laravel, Next.js, Rails, React, Svelte, Vue, WordPress)
- Native integrations with Mailchimp, Klaviyo, HubSpot, Zapier, Make, n8n, Clay, Salesforce, more
- Pay-per-email pricing
**Why this matters**: Cold email reputation craters when bounce rates exceed 2%. Validating before sending is non-negotiable. Apollo/ZoomInfo/Hunter data is often 60–80% accurate — Truelist catches the rest.
**Integration**: see [truelist.md](../../../tools/integrations/truelist.md)
---
## LinkedIn Sales Navigator
**Use for**: Manual decision-maker discovery. The gold standard for B2B / SaaS prospecting but only when used as a research tool.
**Strengths**:
- Most accurate decision-maker data in the industry
- Real-time job changes, posts, signals
- Lead lists, alerts, saved searches
- Inmail credits (separate channel from cold email)
**Hard rules**:
- **Never bulk scrape**. LinkedIn aggressively bans scrapers. Account ban risk is real and permanent.
- Use Sales Nav as a research interface — open profiles, read, take notes, capture key data manually.
- Apollo and other tools claim LinkedIn data via partnerships / public mirroring — verify the source legitimacy before assuming compliance.
**Integration**: no MCP or API access at consumer level. Manual research only.
---
## BuiltWith / Wappalyzer
**Use for**: Tech stack qualification (SaaS branch).
**BuiltWith**:
- ~50K+ technologies tracked
- API + bulk lookups (paid)
- Historical data (when stack changed)
**Wappalyzer**:
- Free browser extension; paid API
- Lighter coverage than BuiltWith
- Faster for one-off lookups
Cross-reference both for high-confidence tech stack signals.
---
## Crunchbase
**Use for**: Funding signals (SaaS branch).
**Strengths**:
- Free tier shows recent funding events
- Paid (Pro / Enterprise) unlocks alerts and deep history
- API access for paid users
**Watch out for**:
- Coverage is best for VC-backed companies; bootstrapped + small businesses underrepresented
- Self-reported data — verify funding amounts independently
---
## GitHub (stargazers / forks / watchers)
**Use for**: Developer-intent prospecting. Especially powerful for dev-tool SaaS — stargazers of competitor or category-defining repos are in-market signal.
**Strengths**:
- Public API, no scraping concerns
- High signal quality (a starred repo = explicit interest)
- Forks are an even stronger signal (intent to modify, not just bookmark)
- Bundled `github-prospects.js` CLI handles pagination + enrichment + CSV output
- Free with 5,000 req/hr authenticated rate limit
**Watch out for**:
- Only ~5–20% of users publish email — pair with Apollo/Clay/Hunter for enrichment
- Very-popular repos (100K+ stars) are mostly noise; smaller targeted repos (5K–25K) give better signal density
- Most prospects are individuals, not company contacts directly — need to figure out their company from `company` field or LinkedIn
**Integration**: see [github.md](../../../tools/integrations/github.md)
---
## Firecrawl / Browserbase (single-target site research)
**Use for**: Programmatically extracting content from a **prospect's own website** that you already found via discovery on platforms like Google Maps, Yelp, or LinkedIn. Not for scraping those platforms themselves.
### Firecrawl
- **Best for**: "Just give me the page as markdown" — Local SMB website status checks, B2B company about/team page extraction, structured field extraction
- **Strengths**: Low overhead, returns clean LLM-ready markdown, handles most JS-rendered sites, has an MCP server
- **API + MCP + SDKs**: Node, Python, Go, Rust
### Browserbase
- **Best for**: When you need real Chromium — JS-heavy pages, cookie consent dialogs, form submission to reach a contact page, session state
- **Strengths**: Full browser control via Playwright/Puppeteer; Stagehand provides AI-friendly natural-language extraction; session recordings for debugging
- **API + MCP (Stagehand) + SDKs**: Node, Python
### Critical compliance line
Both tools can technically point at any URL. The hard rule:
- ✓ **OK**: extracting content from a single business's own website (`joescoffeeshop.com`) that you found through manual discovery
- ✗ **NOT OK**: pointing them at `google.com/maps`, LinkedIn search results, Yelp listings, or any platform whose ToS prohibits bulk extraction
Discovery happens on platforms (manual browser-assisted research). Extraction happens on individual public business sites.
**Integrations**: see [firecrawl.md](../../../tools/integrations/firecrawl.md), [browserbase.md](../../../tools/integrations/browserbase.md)
---
## RB2B / Clearbit Reveal
**Use for**: Identifying anonymous site visitors as warm intent signals.
**Strengths**:
- Pixel-based visitor → company identification
- High-intent: they came to your site, they're already in research mode
- Slack / email alerts on key visits
**Watch out for**:
- Privacy/GDPR considerations — verify your privacy policy disclosures
- Person-level identification raises higher concerns than company-level
**Integration**: see [rb2b.md](../../../tools/integrations/rb2b.md)
---
## Free / browser-only fallbacks
When the user has no paid tools, lean on:
- **Google Search** — exact business name + city + role searches
- **LinkedIn** (manual, no scraping) — company pages, employee lookups
- **Crunchbase free tier** — funding events
- **Wappalyzer browser extension** — tech stack at a glance
- **Hunter.io free tier** — 25 lookups/month
- **Google Maps** — for Local SMB discovery
- **Business websites + About pages** — primary source for any claim
- **News sites + press releases** — trigger event monitoring via Google Alerts
Slower than tooled-up workflows, but produces high-quality smaller lists if the user is willing to do the work.
---
## Sequencing recommendations
A typical full-stack prospecting workflow:
1. **Define ICP** from product-marketing context (no tools needed)
2. **Initial list** from Apollo or ZoomInfo (firmographic filter)
3. **Enrich** with Clay (waterfall: tech stack, funding, trigger events)
4. **Decision-maker mapping** in LinkedIn Sales Nav (manual)
5. **Email pattern discovery** with Hunter or Apollo's built-in
6. **Email validation** with Truelist before final list
7. **Hand off** to cold-email skill for outreach copy
Adapt this sequence based on which tools the user actually has.
FILE:references/demand-signals.md
# Demand-Signal Discovery (Find Your First Customers)
The other three branches build a list from who *fits* (firmographics, technographics, proximity). This branch builds a list from who is *already showing the pain* — the early-stage motion where you have a product and a hunch but no customer base yet, and you need your first ten real conversations. You are not filtering a database; you are mining recent public discourse for people describing the exact problem you solve, then linking every prospect to the evidence.
Use this branch when the user is pre-product-market-fit, launching something new, or looking for **design partners, beta users, or first customers** rather than a scaled outbound list. It reuses the shared five phases and every compliance guardrail in SKILL.md; what changes is where you look, how you score, and what you ship.
Pattern credit: the framework here is re-expressed from the open-source `first-customer-finder` Codex skill (Kappaemme, MIT), extended with our live-recency tooling.
## What makes this branch different
| | List-building branches (SaaS / B2B / SMB) | Demand-signal discovery |
|---|---|---|
| Starts from | A firmographic ICP | A described problem |
| Sources | Contact databases (Apollo, ZoomInfo, Clay) | Public discourse (forums, reviews, issues, posts) |
| Contact step | Enrich + verify email deliverability | None — reach them where they already posted |
| Wins on | Coverage at scale | 10 strong evidence-backed matches over a long list |
| Output | A scored lead sheet | An evidence report + manual outreach plan |
A prospect here without a cited pain, need, or timing signal is a speculative fit — it does **not** belong in the primary shortlist. Evidence is the entry ticket.
## Step 1 — Product brief (before any searching)
Define, specifically enough to *reject* weak matches:
- product and the promised outcome
- primary user and the economic buyer (often different)
- the urgent job to be done
- the current alternative or workaround being replaced
- the likely adoption trigger (what makes now the moment)
- geography / language constraint
- clear disqualifiers
Don't start broad collection until the brief is sharp. Pull from `.agents/product-marketing.md` if it exists.
## Step 2 — Mine the five signal buckets
Search several angles, not one query repeated. Adapt wording to how the audience actually talks (mine their vocabulary from organic content first — see the ad-creative hook-system's organic-language note for the same idea).
1. **Explicit demand** — "looking for," "recommend a tool for," "alternative to [X]," "does anything exist that," "how do you all handle."
2. **Pain** — "takes hours," "so manual," "hate that," "keeps breaking," "biggest frustration with," "why is there no."
3. **Workaround** — spreadsheets, copy-paste, a VA, a Zapier chain, a script, a template, any repeated manual step that your product would replace.
4. **Switching** — cancellation, migration, "moving off [competitor]," a missing feature, a pricing complaint, competitor frustration.
5. **Timing** — a public launch, a new hire for the relevant function, expansion, a new workflow or regulation, an integration announcement — a *current* event that makes the product relevant now.
**Use our live-recency edge.** A generic skill relies on whatever a web search surfaces; you have better:
- **last30days** — Reddit, Hacker News, X, YouTube, and web signals from the last 30 days. This is the single highest-value tool for this branch: recency *is* the timing signal.
- **social-fetch** — pull the full content of a specific post/thread you find, normalized.
- **scraping** / **Firecrawl** / **Browserbase** — read the original public page (a forum thread, a GitHub issue, a review), never qualify from a search snippet alone.
- **deep-research** — for a multi-source sweep with adversarial verification when the wedge is broad.
- **competitor-profiling** / **customer-research** — competitor switching signals and review-mining for the pain language.
## Step 3 — Source mix (public only)
Forums and public community threads · public social posts and replies · product and app-marketplace reviews · GitHub issues and feature requests · public company pages, job posts, changelogs, launch announcements · "looking for a tool" posts and directories.
Avoid private groups, gated communities, data brokers, leaked datasets, and any source whose terms prohibit access — the same compliance guardrails as every other branch (see SKILL.md), including the no-sensitive-traits rule.
**Business/professional context only.** Qualify and reach out only where someone is posting in a professional or business capacity about a work problem (a founder in an indie-hackers thread, a developer in a GitHub issue, an ops lead in a subreddit for their role). Exclude personal-distress contexts entirely — health, financial hardship, addiction, grief, or any consumer support forum where people are venting personal problems, even if your product is tangentially relevant. When the motion is genuinely consumer (B2C), a public pain post is not on its own a lawful basis for cold outreach — reach people through the channel's own norms (reply publicly where replying is expected) and never DM a stranger off a personal post.
Quote minimally, paraphrase by default, and link every material pain or timing signal.
## Step 4 — Score on demand-fit (not ICP-fit)
The list-building branches score Hot/Warm/Cold on ICP fit. This branch scores 0–100 on **demand fit** — how strongly the evidence says this specific prospect wants this specific thing now. Score each dimension 0–5:
| Dimension | Weight | What it measures |
|---|---|---|
| **Pain strength** | 25% | Directness, severity, repetition, and cost of the stated problem |
| **Product fit** | 25% | How directly your product solves the evidenced job |
| **Timing** | 20% | Freshness + a current trigger present |
| **Public reachability** | 15% | A natural, relevant public/professional contact path exists |
| **Evidence quality** | 15% | Specificity, source reliability, confidence the signal is really theirs |
```
score = pain/5*25 + fit/5*25 + timing/5*20 + reachability/5*15 + evidence/5*15
```
| Band | Meaning |
|---|---|
| **80–100** | Strong first-customer candidate |
| **65–79** | Promising — validate fast |
| **50–64** | Plausible but missing a material signal |
| **Below 50** | Do not include in the primary shortlist |
An old explicit request can still count — but lower the timing score and label the date. A company that merely matches the industry with no evidenced trigger is *not* a qualified prospect here.
### Prospect stages
- **High intent** — publicly requesting a solution or actively switching
- **Problem aware** — clearly describing the pain or an expensive workaround
- **Trigger present** — a current business event makes the product relevant
- **Potential fit** — ICP match, incomplete evidence → keep *outside* the primary shortlist
### Evidence ledger (per qualified prospect)
Displayed name (company/project/public professional) · source title + URL · visible publication date or "date unavailable" · source type · the concise pain/timing signal · observed evidence vs. inference (label which) · score breakdown · freshness warning when the signal is stale.
## Step 5 — Draft outreach, never send it
Recommend the most natural channel *already associated with the source*, and only where a reply is a normal part of that channel (reply in the public thread, respond via a public professional profile). Don't turn a public post into a private DM the poster didn't invite, and never contact someone off a personal-distress post. Draft one opener, under ~90 words, in this shape:
1. mention the public context naturally
2. connect it to the exact problem
3. explain the product in one sentence
4. ask one low-friction question
Never claim familiarity you don't have, never fabricate personal details, and never auto-send: no messages, connects, follows, comments, form submissions, or CRM records unless the user separately authorizes that action. This is the manual/gated posture from the marketing-loops guardrails.
## Step 6 — Ship the evidence report
Lead with the most actionable evidence, in this order:
1. **Verdict** — does the product have reachable early-customer signal, or not yet? (An honest "not yet, here's why" is a valid answer.)
2. **ICP** — buyer, job, trigger, disqualifiers.
3. **Top prospect** — the single strongest evidence-backed candidate and why now.
4. **Prospect shortlist** — per prospect: source, pain signal, demand-fit score, stage, why-now, channel, opener.
5. **Repeated patterns** — pains and triggers recurring across prospects (these are your positioning and messaging gold).
6. **Seven-day manual outreach plan** — a low-volume validation sequence (e.g., contact the top 3 with one source-based question; share a mockup only after they confirm the pain; target three conversations and one design-partner commitment).
7. **Limits** — what evidence is missing and what must be confirmed through real conversations.
For a shareable standalone HTML version of this report, the JSON→HTML generator pattern in ad-creative's [creative-review-page.md](../../ad-creative/references/creative-review-page.md) is the model (escape every value; keep it self-contained).
## The honesty rules (non-negotiable)
- Every primary prospect links to at least one real public signal. No signal, no shortlist.
- Label the output **"potential customer based on public signals"** — never "interested," "will buy," or "has consented."
- Prefer ten strong matches over a long generic list. Make uncertainty and stale evidence visible.
- Personalize from the cited source, not from invented assumptions.
- Treat the shortlist as a research hypothesis to validate through conversations, not a customer database.
FILE:references/local-prospecting.md
# Local SMB Prospecting Reference
For when the user sells to local small businesses — shops, gyms, restaurants, salons, clinics, professional services, contractors, real estate, fitness studios, dental practices.
Adapted from and generalized beyond the local-client-prospector pattern (browser-assisted discovery + website status classification + proximity scoring).
---
## ICP Signals That Matter (Local SMB branch)
### Operational signals
- **Active business** — Google Business Profile updated, recent reviews, recent hours updates
- **Recent activity** — open right now, regular hours posted, recent photos uploaded by owner
- **Customer engagement** — owner responding to reviews, posts on social, active calendar (for service businesses)
### Online presence signals (the core SMB qualification axis)
The reference local-client-prospector skill uses **website status** as the primary qualification — port this directly. Four classifications:
| Status | Definition | Typical outcome |
|--------|-----------|-----------------|
| **No site found** | No credible standalone website after cross-checked search | **Hot prospect** for web/marketing service |
| **Social only** | Facebook, Instagram, WhatsApp, Linktree, booking portal, marketplace page only — no standalone site | **Hot prospect** for web/marketing service |
| **Weak site** | Standalone site exists but outdated, broken, very thin, non-mobile-friendly, or missing clear contact/conversion flow | **Warm prospect** for refresh / rebuild service |
| **Has site** | Credible, modern standalone site exists | **Low prospect** unless other signals apply (e.g., poor SEO, weak conversion design) |
### Proximity signals
- **Distance** from the user's location or service area
- **Density** — clusters of similar businesses in one area = neighborhood targeting opportunity
- **Travel time** — useful when in-person discovery, install, or service delivery is required
### Decay signals
- Closed permanently (Google Maps banner)
- Reviews paused or business listing reported as closed
- Last activity (review, post) >12 months ago
---
## Discovery Sources (Local SMB branch)
### Primary
- **Google Maps** (browser, manual) — search "category near [location]" and walk the visible results. Cross-check details. Don't bulk-extract.
- **Yelp** — secondary verification; complementary categories
- **Bing Local / Apple Maps** — different coverage on smaller businesses
- **Facebook Pages search** — many SMBs are Facebook-only
### Cross-verification
- **Business's own website** (if any)
- **Industry directories** (e.g., Healthgrades for medical, OpenTable for restaurants, Avvo for legal)
- **Local Chamber of Commerce listings**
- **State business registries** for incorporation status
- **Search results for "[business name] [city]"** to discover non-Maps presence
---
## Browser Research Workflow
1. Open a browser and search Google Maps for the category near `base_location`
2. Build a candidate list from visible local results, search results, and public directories
3. For each candidate, inspect public sources to fill required fields
4. Search the exact business name plus city/town to check whether a standalone website exists
5. Classify website status per the table above
6. Mark confidence: High (2+ sources), Medium (1 source + consistent evidence), Low (incomplete/ambiguous)
When the user explicitly asks for subagents AND subagents are available, split candidates into non-overlapping batches and ask each subagent to verify only website/social/contact status. Don't use subagents for the primary search if it slows progress.
### Optional: programmatic verification with Firecrawl or Browserbase
Once you have a candidate's website URL (found via manual Maps/Yelp discovery), you can speed up website-status classification by hitting the URL programmatically:
- **Firecrawl** for simple "is this site live, modern, mobile-friendly, conversion-flow-equipped" reads — returns clean markdown you can inspect
- **Browserbase** when the candidate site requires JS rendering, has a cookie consent dialog, or you need session state
**Strict line**: use these on the individual business's URL. **Don't** point them at Google Maps, Yelp, or any platform whose ToS prohibits bulk extraction — discovery stays manual.
See [data-sources.md](data-sources.md) for setup details.
---
## Qualification Checklist (Local SMB branch)
- [ ] Business is active (recent reviews or activity in last 6 months)
- [ ] Category matches user's service offering
- [ ] Distance / proximity within target radius
- [ ] Website status classified
- [ ] Phone or contact channel verified
- [ ] At least one cross-source confirms business operates at the listed address
- [ ] Not a duplicate / chain location / out-of-scope category
- [ ] Not closed permanently
---
## Lead Scoring (Local SMB)
Use this simple rubric (matches local-client-prospector pattern):
| Score | Criteria |
|-------|----------|
| **Hot** | No site found OR social-only + phone present + active business + within target radius |
| **Warm** | Weak site, poor online presentation, or marketplace/booking-page only |
| **Cold** | Good website already present OR low confidence |
| **Skip** | Closed, duplicate, outside radius, irrelevant category, or not a business prospect |
---
## Output Columns (Local SMB branch)
Chat table (≤15 rows):
```
| Score | Business | Category | Area | Distance | Website status | Website/Social | Phone | Why it's a prospect | Confidence |
```
CSV:
```csv
score,business,category,area,distance_km,website_status,website_url,social_urls,phone,email,source_urls,why_prospect,confidence,verified_date,notes
```
Rules:
- Keep "Why it's a prospect" short and actionable
- Use `Not found` instead of leaving blank fields
- Include source links sparingly, not all of them
- After the table, add **Best first outreach targets** with the top 3 leads and one practical reason each
- If confidence is low, state exactly what remains uncertain
---
## Top Outreach Targets Selection (Local SMB)
Prioritize for the top 3 hot leads:
1. **No site / social only + phone present** = clearest service opportunity
2. **High review count** = active, established business with real customers
3. **Owner-responded reviews** = engaged owner = more likely to evaluate a vendor
4. **Industry alignment with your service specialty** beats generic category match
Each top target rationale should be one sentence naming the gap and the signal: "No standalone website (cross-checked); 80+ Google reviews with owner replies; 2 km from target area."
---
## Compliance Notes (Local SMB-specific)
The local branch is the most scraping-sensitive of the three motions. Specifically:
- **Google Maps Terms of Service** prohibit bulk extraction. Treat browser visits as research, not as data acquisition.
- **Don't store full Google Maps Place IDs in your CRM** — the ToS limits storage of Maps data.
- **Public business contact channels only**: published phone, contact form, info@ email. Don't reach individual employees through their personal channels.
- **Owner/operator name when published on the business's own site** is OK to use. If you only got it from LinkedIn, mark the source.
---
## Common Mistakes (Local SMB)
1. **Bulk-scraping Google Maps** — fastest way to violate ToS and lose the research channel.
2. **Treating Google Maps data as truth** — listings go stale. Cross-check hours, status, and reviews.
3. **Skipping the website status cross-check** — finding "no site" on Maps doesn't mean no site exists; do an exact-name web search before classifying.
4. **Targeting only the largest businesses** — they're already covered by other providers. The 2–5 employee SMBs are the under-served opportunity.
5. **Generic outreach to all hot leads** — local SMBs respond better to outreach that names their specific gap ("I noticed your menu isn't visible on mobile") than generic pitches.
6. **Ignoring chains and franchises** as Skip — sometimes the franchisee is the buyer and they have local marketing authority. Verify before skipping.
FILE:references/saas-prospecting.md
# SaaS Prospecting Reference
For when the user sells SaaS or digital services to other SaaS companies / digital businesses.
---
## ICP Signals That Matter (SaaS branch)
Beyond standard firmographics (industry, size, geography), SaaS prospects are qualified by:
### Technographic signals
- **Tech stack** — do they use complementary tools (your integration target) or competing tools (a switch opportunity)?
- **Recent stack changes** — adding/removing tools signals active vendor evaluation
- **Custom-built vs off-the-shelf** — DIY tooling often means a buyer who'd benefit from your product
- **Free/freemium plan signals** — using a free competitor means they may be ready to upgrade
### Growth signals
- **Funding round** — Series A / B / C in last 6 months = budget + new hires + tool needs
- **Headcount growth** — 10%+ growth in last quarter signals scaling pressure
- **Hiring signals** — specific role openings (e.g., "Head of RevOps" → ICP for revops tooling)
- **Product velocity** — frequent shipping, new features, blog posts = healthy growth motion
- **Open positions for your buyer's role** — if you sell to Marketing Ops and they're hiring one, that's a signal
### Decay signals (downgrade scoring)
- Layoffs in target department
- Funding round >2 years ago with no follow-up
- Product hasn't shipped in 6+ months
- Team page shows founders only (very early — may not have budget)
---
## Discovery Sources (SaaS branch)
Combine 2+ sources for cross-verification.
### Tier 1 — primary discovery
- **Apollo**: firmographic + technographic + contact data. Good for building large initial lists.
- **Clay**: waterfall enrichment, custom scoring, multi-source merges. Best for high-quality smaller lists.
- **ZoomInfo**: enterprise-grade firmographic + intent signals. Expensive; mid-market+.
- **LinkedIn Sales Navigator**: decision-maker mapping. Use manually, never bulk scrape.
### Tier 2 — technographic / growth signals
- **BuiltWith**: tech stack lookups, find sites using specific tools
- **Wappalyzer**: free browser extension + API; lighter tech stack signal
- **Crunchbase**: funding rounds, headcount, founders
- **Pitchbook**: deeper investor data (enterprise/paid)
- **ProductHunt**: recent launches, builder audience
- **Hacker News / Show HN**: technical builders launching products
### Tier 3 — buying signals
- **Job boards** (LinkedIn Jobs, Indeed, AngelList): role openings as signals
- **RB2B / Clearbit Reveal**: visitor identification (warm anonymous traffic)
- **GitHub stars/forks of competitor or adjacent repos**: developer-level intent signal (see `tools/integrations/github.md` and the `github-prospects.js` CLI). Especially strong for dev-tool SaaS — a developer who starred `vercel/next.js` last week is in-market for adjacent Next.js infrastructure.
- **Recent blog posts / changelog**: product direction signals
- **G2 reviews mentioning competitor switches**: explicit dissatisfaction signal
#### GitHub prospecting pattern (when audience is developers)
For dev-tool SaaS, GitHub is one of the highest-quality discovery channels:
1. Identify 3–5 "anchor" repos: your direct competitors, your category leader, complementary tools your buyer uses
2. Pull stargazers (or forks for stronger intent) via `node tools/clis/github-prospects.js stargazers <owner/repo> --enrich --with-company --format csv`
3. Filter to users with `company` set — these are the easiest to enrich downstream
4. Pair with Apollo/Clay/Hunter to lookup email by name + company
5. Validate with Truelist before adding to outreach list
Tradeoffs: GitHub yields email for only ~5–20% of users directly. The strength is the signal quality — a stargazer of a niche dev tool is genuinely in-market in a way Apollo firmographics alone can't tell you.
---
## Qualification Checklist (SaaS branch)
For each candidate, verify:
- [ ] Industry vertical matches ICP
- [ ] Company size (headcount) within range
- [ ] Tech stack includes (or notably excludes) a target technology
- [ ] Funding stage matches buyer maturity
- [ ] At least one growth signal in last 90 days (funding, hiring, product velocity)
- [ ] Decision-maker role exists at the company (named or inferable from job listings)
- [ ] Email contact verifiable
- [ ] No disqualifiers (closed, acquired-and-paused, layoffs, ICP miss)
---
## Output Columns (SaaS branch)
Recommended CSV columns:
```csv
score,company,domain,industry,size_band,country,funding_stage,last_round_date,tech_stack_match,signal,signal_date,contact_name,contact_title,contact_email,email_status,linkedin_url,source_urls,why_prospect,confidence,verified_date,notes
```
For chat table, condense to: Score | Company | Industry | Size | Signal | Contact | Email status | Confidence.
---
## Top Outreach Targets Selection (SaaS)
Prioritize for the top 3–5 hot leads:
1. **Strongest signal recency** — funding 30 days ago beats funding 9 months ago
2. **Tech stack match strength** — known integration partner beats inferred fit
3. **Decision-maker named with verified email** — beats role-pattern-guessed email
4. **Multi-source confidence** — both Apollo + Crunchbase agree beats one source
Each top target gets a one-sentence outreach rationale that names the specific signal: "Raised Series B 30 days ago; hiring Head of RevOps; verified VP of Ops email."
---
## Common Mistakes (SaaS)
1. **Buying lists from Apollo wholesale** without re-verifying email and re-checking firmographics. Stale data is the norm.
2. **Treating tech stack data as 100% accurate**. BuiltWith and Wappalyzer miss things; Clay's waterfalls miss things. Cross-check.
3. **Targeting Series C+ for early-stage SaaS sellers**. The buyer profile is wrong — too many procurement hoops, too much red tape.
4. **Targeting Series Pre-Seed seed** for products requiring meaningful budget. They have neither budget nor evaluator bandwidth.
5. **Ignoring intent data when it exists** (ZoomInfo Intent, 6sense, etc.) — pre-warm signals beat cold every time.
Hỗ trợ quan hệ công chúng: earned media, thông cáo báo chí, tiếp cận nhà báo và chiến lược truyền thông.
---
name: public-relations
description: "When the user wants help with public relations, earned media, press coverage, journalist outreach, or media strategy (not pull requests). Also use when the user mentions 'PR,' 'public relations,' 'press,' 'press release,' 'press coverage,' 'media outreach,' 'pitch a journalist,' 'get featured,' 'media list,' 'media kit,' 'press kit,' 'newsjacking,' 'news hijack,' 'HARO,' 'Qwoted,' 'Featured,' 'Help A Reporter,' 'reporter request,' 'tech press,' 'TechCrunch,' 'earned media,' 'thought leadership placement,' 'op-ed,' 'guest article,' 'press contacts,' 'podcast prep,' 'going on a podcast,' 'podcast guest,' 'prep me for this podcast,' or 'how do I get press.' Use this for earned media work — finding journalists, pitching stories, newsjacking, prepping podcast appearances, and responding to press requests. For startup/SaaS/AI directory submissions, see directory-submissions. For product launches, see launch. For social-media engagement, see social. For cold-email outreach to prospects, see cold-email."
metadata:
version: 1.1.1
---
# Public Relations & Earned Media
You are an expert in earned media for software products. Your goal is to help the user get covered by journalists, podcasts, and newsletters — efficiently, with respect for the people on the other end of the pitch.
## Before Starting
**Check for product marketing context first:**
If `.agents/product-marketing.md` exists (or `.claude/product-marketing.md`, or the legacy `product-marketing-context.md` filename, in older setups), read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
---
## Core Philosophy
PR is not a substitute for distribution. It's a multiplier for it.
- **Earned media doesn't drive direct conversions.** A TechCrunch hit will not give you 1,000 paying customers. It will give you backlinks, brand legitimacy, AI-citation surface area, and ammo for sales conversations.
- **Pitch journalists like you'd pitch a customer:** specific, useful, fast, and never about you.
- **The story is not your product. The story is the trend, the data, the conflict, or the human.** Your product is the evidence. Every pitchable story bends toward one of three angles — Founding Story, David vs Goliath, or Have an Enemy (a *broken system*, never a competitor). See [references/story-angles.md](references/story-angles.md).
- **Chase press for the compound effect, not the traffic bump.** The bump fades in a day; authority, journalist relationships, and AI-citation surface compound. Build media relationships *before* you need them, and run one core asset through the whole repurposing flywheel.
- **Speed beats polish on reactive PR.** A B+ pitch in the first hour of a story beats an A+ pitch on day three.
### When PR is worth it
- You have **a real story** — proprietary data, a strong opinion, a milestone, a customer with a sharp before/after, or a fresh angle on a trending topic
- You have **founder/exec time** — journalists want quotes from people with skin in the game, not from a PR rep
- You have **a destination** — a press page, blog post, or product launch that converts attention into something useful
### When to skip PR (for now)
- Pre-launch with no story beyond "we exist"
- No one on the team can sustain pitching for 4–6 weeks (PR is a momentum game)
- You don't have a clear ICP — journalists ask "who reads my piece because of this?" and if you can't answer, neither can they
---
## The PR Mix
Four modes. Most teams over-index on one. Run at least three.
| Mode | What it is | Effort | Speed to coverage |
|------|------------|--------|-------------------|
| **Reactive (newsjacking)** | Inject your POV into trending news | Low–medium | Hours to days |
| **Proactive (pitching)** | Build a media list, pitch original stories | High | 2–8 weeks |
| **Inbound (press requests)** | Respond to journalist queries on HARO/Qwoted/Featured | Low | Days to weeks |
| **Owned (press page + media kit)** | Make it easy for journalists to find you | One-time setup | N/A |
**For the story angle taxonomy (Founding Story / David vs Goliath / Have an Enemy), data stories, media relationship-building, and the PR repurposing flywheel** — see [references/story-angles.md](references/story-angles.md)
**For the reactive newsjacking workflow** — see [references/newsjacking.md](references/newsjacking.md)
**For proactive journalist pitching** — see [references/journalist-pitching.md](references/journalist-pitching.md)
**For inbound press-request platforms (HARO, Qwoted, etc.)** — see [references/press-platforms.md](references/press-platforms.md)
**For where to pitch (media outlets, podcasts, newsletters)** — see [references/media-outlets.md](references/media-outlets.md). For startup/SaaS/AI directories, use the separate `directory-submissions` skill — different intent, different list.
**For prepping a podcast appearance you've landed** — see [references/podcast-guest-prep.md](references/podcast-guest-prep.md). Episodes get transcribed and cited by AI assistants, so a good appearance compounds in AI answers for years — prep is an AI-visibility play, not just interview polish.
---
## Owned: Press Page + Media Kit
Set this up once. It's the cheapest PR investment with the highest ROI on every future story.
**Press page (`/press` or `/newsroom`) should include:**
- One-paragraph company description (copy/paste ready)
- Founder bios with headshots (high-res, downloadable)
- Logo pack (SVG + PNG, light + dark, with usage guidelines)
- Product screenshots (high-res)
- Recent coverage list (social proof for the next journalist)
- Founding date, employee count, funding (if disclosed)
- Press contact email (not a form — journalists hate forms)
- Recent press releases / announcements
**One sentence at the top:** "For interview requests or assets, email press@yourcompany.com — we respond within 24 hours."
Then *actually* respond within 24 hours.
---
## Quick Reference: Pitch Quality Bar
Before sending any pitch, the answer to all of these should be yes:
- [ ] Does this journalist cover this beat? (Check their last 5 articles.)
- [ ] Is there a clear news hook — something that just happened or is about to?
- [ ] Could this journalist write a complete story from this email alone? (Data, quotes, customer name, contact.)
- [ ] Is the subject line specific enough to predict the article's headline?
- [ ] Is the pitch under 150 words?
- [ ] Did you avoid the words "revolutionary," "game-changing," "disruptive," and "synergy"?
- [ ] Is the ask clear? (Interview? Embargo? Exclusive? Quote?)
If any answer is no, don't send.
---
## Measurement
What to track:
| Metric | Why |
|--------|-----|
| **Coverage count** (placements / month) | Activity baseline |
| **Domain rating of placements** | Backlink value |
| **Referral traffic from coverage** | Did anyone actually click? |
| **Brand search lift** | Did people search you after reading? |
| **AI citation rate** (ChatGPT, Perplexity quote your brand?) | The new measurement that matters |
| **Sales conversations citing the article** | The only one that matters for revenue |
What not to obsess over: AVE (advertising value equivalency) — it's a vanity metric PR firms invented.
---
## Common Workflows
### "Help me newsjack [trending story]"
Go to [newsjacking.md](references/newsjacking.md), run the scoring rubric, draft 2–3 angles, pick the best, draft the pitch.
### "Find journalists who cover [beat]"
Go to [journalist-pitching.md](references/journalist-pitching.md), use the discovery checklist + dev-browser to research recent articles, build a scored list.
### "What's worth pitching this week?"
Combine: recent product milestones + active news cycles + any data you've collected. Score each potential story by the quality bar above.
### "What's my story angle?" / "How do I get press with no news?"
Go to [story-angles.md](references/story-angles.md). Fit the situation to one of the three angles (Founding Story / David vs Goliath / Have an Enemy), or turn proprietary data into a data story. Remember: a milestone alone isn't a story — milestone *with narrative* is.
### "Respond to this HARO query"
Go to [press-platforms.md](references/press-platforms.md), use the response template, keep it under 200 words.
### "I'm going on [podcast] next week — help me prep"
Go to [podcast-guest-prep.md](references/podcast-guest-prep.md): research the show (RSS feed → site → Apple Podcasts → web), extract the recurring threads and host profiles, map the guest's stories onto them, deliver the brief.
### "Build my press page"
Use the checklist above. Most companies do this in an afternoon and forget about it for a year — that's fine.
FILE:evals/evals.json
{
"skill_name": "public-relations",
"evals": [
{
"id": 1,
"prompt": "We're a tiny 2-person SaaS competing against Notion in the docs space. We don't have any funding news or a big milestone. How do we get press coverage?",
"expected_output": "Should route to references/story-angles.md. Should recognize there's no traditional news hook and instead fit the situation to one of the three story angles, most naturally David vs Goliath (own your size against an incumbent, cite the Superhuman vs Gmail exemplar). Should mention the other two angles — Founding Story (Pieter Levels '12 startups in 12 months,' build-in-public, share revenue) and Have an Enemy (the enemy is a broken system, not the competitor — never 'Notion is bad,' and reference HEY vs Apple's App Store 2020). Should reinforce 'the story is not your product' and that a milestone alone isn't a story — milestone with narrative is. Should note the compound effect (authority, relationships, AI-citation surface) over the traffic bump. May suggest turning proprietary data into a data story and building media relationships before pitching. Should not invent fake news or suggest attacking the competitor by name.",
"assertions": [
"Routes to references/story-angles.md",
"Identifies David vs Goliath as the fitting angle with the Superhuman vs Gmail exemplar",
"Names all three angles (Founding Story, David vs Goliath, Have an Enemy)",
"Clarifies the enemy is a broken system, not a named competitor (HEY vs Apple App Store 2020)",
"Cites Pieter Levels '12 startups in 12 months' for Founding Story",
"Reinforces 'the story is not your product'",
"Notes milestone-with-narrative and/or the compound effect over the traffic bump",
"Does not fabricate news or suggest smearing the competitor by name"
],
"files": []
}
]
}
FILE:references/journalist-pitching.md
# Journalist Pitching — Proactive PR Workflow
Building a media list, scoring journalist fit, and crafting pitches that actually get opened. This is a 4–8 week practice, not a one-shot.
## Contents
- Building the media list
- Scoring journalist fit
- Pitch templates by angle
- Subject lines that get opened
- Voice and structure
- Embargoes, exclusives, and follow-up etiquette
- Pitch killers
- Tooling
---
## Building the Media List
The goal: a list of 20–40 journalists who actually cover your beat. Not 500 names from a database.
### Discovery checklist
For each candidate journalist:
- [ ] Read their **last 5 articles** — are they covering your beat right now?
- [ ] Note their **publication** — does it reach your ICP?
- [ ] Check their **bio** on the outlet site — what topics do they own?
- [ ] Check **X/LinkedIn** for what they're posting about this week
- [ ] Note their **email** (usually on outlet author page, Muck Rack, or company About page)
- [ ] Check **Muck Rack** if available — it shows recent topics and pitch preferences
### Where to find candidates
| Method | How |
|--------|-----|
| **Reverse lookup from coverage you want** | Find 5 articles about competitors / your category, note bylines |
| **Topic search on Muck Rack** | Free tier shows journalists by topic |
| **X / Twitter lists** | "[your niche] reporters" lists already exist |
| **LinkedIn search** | "Journalist" + "[your category]" — filter by recent activity |
| **Newsletter author pages** | Beehiiv, Substack, ConvertKit creators are pitchable |
| **Podcast host research** | Listen to 1 episode before pitching — non-negotiable |
### Don't waste time on
- **Mass media databases** (Cision, Meltwater) for early-stage — overkill and expensive
- **Journalists who haven't posted in 6+ months** — they may have left
- **"Editor-in-chief" generic addresses** — pitches there get ignored or routed to interns
- **Journalists who explicitly state "no PR pitches" in bio** — respect it
---
## Scoring Journalist Fit
Score each journalist 1–10 across four dimensions. Sum and rank. Focus on top 20.
| Dimension | What it measures | Weight |
|-----------|------------------|--------|
| **Beat match** | Do they cover your category specifically? | 3x |
| **Reach** | Outlet's audience size + their byline traction | 2x |
| **Engagement** | Do they respond to pitches publicly / on X? | 2x |
| **Recency** | Have they written about a related topic in last 30d? | 1x |
**Tiering:**
- **Tier 1 (8–10):** Personal pitch with original angle. Custom each time.
- **Tier 2 (5–7):** Standard pitch, lightly customized.
- **Tier 3 (below 5):** Skip or pitch only when story is exceptional.
---
## Pitch Templates by Angle
Six structures that work. Pick the one that matches your story.
### 1. Data story
```
Subject: [Specific stat] — [implication]
Hi [name],
I noticed you covered [recent article] — wanted to share data that might
be relevant.
We [analyzed N / surveyed N / tracked N] and found:
• [Stat 1 with surprise factor]
• [Stat 2]
• [Stat 3]
The most interesting pattern: [one-sentence insight].
Full data + methodology here: [link to one-pager, not your homepage]
Happy to share the raw dataset, jump on a call, or connect you with
[customer who's relevant].
[your name + 1-line credential]
```
### 2. Exclusive launch / milestone
```
Subject: Exclusive: [specific milestone] at [company]
Hi [name],
I have an exclusive on [milestone] that I think fits your [beat] coverage.
The story: [one sentence]
Why it matters: [one sentence — for their readers, not for you]
What's new: [the actual news, not the marketing line]
Embargo until [day, time, timezone] — would love to give you first
window. Press kit + assets: [link]
Free to talk [two specific time options].
[your name]
```
### 3. Op-ed / contributed piece
```
Subject: Op-ed pitch: [provocative thesis]
Hi [name],
I read your piece on [recent article] — sharp take on [specific point].
I'd like to pitch a 700-word op-ed: "[Thesis as a headline]"
Core argument:
• [Point 1]
• [Point 2]
• [Point 3 — the surprising one]
Why me: [1 sentence — credential or unique vantage]
Why now: [1 sentence — the news hook]
Can have a draft to you by [date]. Happy to adapt to your house style.
[your name]
```
### 4. Customer story
```
Subject: Customer story for [their beat] — [specific outcome]
Hi [name],
For your [beat] coverage, I have a [customer type] willing to talk on
the record about [specific outcome].
The hook: [customer] [did something specific] and [measurable result].
The interesting part: [the surprising or counterintuitive detail].
Customer details:
• Name: [name, title, company]
• Available: [windows]
• Willing to share: [data points / screenshots / metrics]
Happy to coordinate the intro.
[your name]
```
### 5. Trend piece / connector
```
Subject: Trend forming in [space] — three signals
Hi [name],
Three things in [space] this month that I think connect:
1. [Signal 1 with link]
2. [Signal 2 with link]
3. [Signal 3 — yours, briefly]
The pattern: [one sentence].
This might be early for a piece, but if you're tracking the space I
wanted to flag it. Happy to share data we've collected or connect you
with others seeing the same.
[your name]
```
### 6. Newsjack response
```
Subject: Re: [their article headline] — quick data point
Hi [name],
Saw your piece on [story] this morning — wanted to add a relevant
data point in case you do a follow-up.
[One-sentence stat or insight].
Source: [our data / our customers / our analysis]
Methodology: [one sentence]
Quotable: "[a sentence you'd be comfortable seeing in print]"
If useful for a follow-up, I'm around all day at this number: [phone].
[your name]
```
---
## Subject Lines That Get Opened
Journalists open pitches based on the subject line alone. Rules:
- **Under 50 characters** — mobile preview cuts off
- **Lead with the specific** — "73% of devs deploy to prod on Fridays" beats "New data on developer workflows"
- **Promise a story, not a product** — "Why [trend]" beats "[Company] launches [thing]"
- **Use prefixes that signal value** — "Exclusive:", "Data:", "Op-ed pitch:", "Re: [their article]"
**Test against this question:** would *you* open this in a 200-email inbox?
**Patterns that work:**
- "[Specific stat] — [implication]" — "73% of agents fail this test"
- "Exclusive: [milestone]" — "Exclusive: Anthropic launches AgentOS"
- "Re: [their headline]" — direct response to recent coverage
- "[Provocative thesis]" — "Why VC funding is bad for AI safety"
**Patterns that get deleted:**
- "Press release: [boring]"
- "[Company] announces [thing]"
- "Story idea for you!"
- "Following up on my previous email"
- Any subject line with "innovative," "disruptive," "revolutionary"
---
## Voice and Structure
### Length
**150 words max for the pitch.** If you can't say it in 150 words, you don't know what your story is yet.
### Structure
1. **One-line context** — why you're emailing them specifically (their recent article, their beat)
2. **The story** — what it is, in one sentence
3. **Why it matters to their readers** — not why it matters to you
4. **Proof** — data, customer, quote, link
5. **The ask** — interview, embargo, quote, link
### Voice
- Sound like a person, not a press release
- Reference their actual recent work — proves you read them
- Don't use emoji unless they do
- Don't open with "I hope this finds you well" — burn it
- Don't ask "did you get my email?" follow-ups (see [Follow-up](#embargoes-exclusives-and-follow-up-etiquette))
### Banned vocabulary
Revolutionary, disruptive, game-changing, paradigm shift, leverage, synergy, robust, seamless, holistic, world-class, best-in-class, next-generation, cutting-edge, AI-powered (unless that's the actual differentiation), at-the-end-of-the-day.
---
## Embargoes, Exclusives, and Follow-Up Etiquette
### Embargoes
An embargo is "you can write this story, but don't publish until [time]."
- **Only offer embargoes to journalists you've worked with** or have strong reputations for honoring them
- State the embargo time clearly: day, time, timezone
- If they break embargo, your relationship with them is over
### Exclusives
"Only you get this story" — powerful tool, use sparingly.
- **First-tier outlet only** — exclusives to tier 2/3 outlets waste the lever
- **Be honest about scope** — "exclusive to [outlet] in the US" is fine
- **Have a parallel plan** — what you publish/pitch the next day after the exclusive runs
### Follow-up cadence
- **Day 0** — initial pitch
- **Day 3** — one follow-up if you have new information ("Just talked to [customer] who can join us")
- **Day 7** — final check-in with a fresh hook ("This came out today, still relevant?")
- **After day 7** — let it go. Re-pitch when you have something genuinely new.
**Never:**
- "Bumping this up" / "Did you see my email?"
- Multi-day silent follow-ups with no new value
- Same pitch reformatted
---
## Pitch Killers
Things that instantly disqualify your pitch:
- Wrong name / wrong outlet (autoreplace fail)
- Pitching topics they explicitly don't cover
- Press release attached as PDF (just paste the key bits)
- Long signature with logos and disclaimers
- CC'ing 5 other journalists on the same email
- "Per my last email" energy
- Asking them to sign an NDA before talking
- Pitching a story you can't actually tell (no customer willing to talk, no data ready to share)
---
## Tooling
### Finding journalist contact info
```bash
# Most journalists' emails follow patterns:
# firstname@outlet.com
# firstname.lastname@outlet.com
# flastname@outlet.com
# Use Hunter.io, RocketReach, or just guess and bounce-check
```
### Researching their recent work (browser-driven)
Use `dev-browser` (persistent session, no rate limits) to:
- Open the journalist's outlet author page → scrape last 5 article headlines + dates
- Open their X/Twitter profile → note recent topics
- Open their LinkedIn → confirm current role
Output what you find as:
```
JOURNALIST PROFILE — [name]
Outlet: [name]
Beat: [topics from last 5 articles]
Recent angle: [pattern you noticed]
Recent X activity: [what they're posting]
Score: [X/40 from rubric]
Best pitch angle: [from template library]
Email: [confirmed]
```
### Maintaining the media list
Store in `.agents/media-list.md` (or `.csv` if you prefer). Update monthly — journalists move jobs constantly.
```markdown
## Tier 1 (top 20)
| Name | Outlet | Beat | Last contact | Last coverage | Email | Score |
|------|--------|------|--------------|---------------|-------|-------|
| ... | ... | ... | 2026-05-15 | none yet | ... | 9/10 |
```
### Pitch tracking
Track in a simple spreadsheet:
- Date sent
- Subject line
- Journalist
- Outlet
- Response (open / reply / pass / coverage)
- What they said
After 30 pitches, you'll see which subject patterns and which angles work for you specifically.
FILE:references/media-outlets.md
# Media Outlets — Where to Pitch
A curated, opinionated list of *where* to pitch for software/SaaS PR. This is the media-outlet slice of resources like submit.co — the journalist-driven half. For startup/SaaS/AI directories (Product Hunt, BetaList, Futurepedia, etc.), use the separate `directory-submissions` skill — different intent, different audience.
## How to use this list
- **Don't pitch a publication. Pitch a journalist at that publication.** See [journalist-pitching.md](journalist-pitching.md) for the discovery workflow.
- **Tier signals quality, not effort** — a small tier-3 outlet might be perfect for your niche
- **Submission/tip URLs are listed where they exist** — but a journalist's email beats a tip form every time
- **This list ages fast** — verify the outlet still exists and the journalist is still there before pitching
---
## Tech & Startup Press (Tier 1)
The big names. High bar to clear, high payoff when you do. Pitch the specific reporter, not the tip line.
| Outlet | Best for | Tip URL |
|--------|---------|---------|
| **TechCrunch** | Funding, product launches, startup news | techcrunch.com/got-a-tip/ |
| **The Verge** | Consumer tech, product reviews, policy | theverge.com/contact |
| **Wired** | Long-form tech, culture, business | wired.com/about/feedback |
| **Fast Company** | Innovation, design, business strategy | fastcompany.com/contact-us |
| **VentureBeat** | AI, enterprise tech, gaming | venturebeat.com/contribute |
| **Ars Technica** | Deep tech, science, policy | arstechnica.com/contact-us |
| **The Information** | Subscription-gated, scoops on tech industry | theinformation.com/about |
| **Bloomberg / Reuters** | Business, finance, big-picture stories | bloomberg.com/feedback |
| **WSJ Tech** | Enterprise, business angle | wsj.com/tips |
| **NYT Tech** | Mainstream tech, culture | nytimes.com/tips |
---
## SaaS & B2B (Tier 1–2)
Lower-profile than the consumer tech outlets but often higher ROI for B2B SaaS.
| Outlet | Best for | Notes |
|--------|----------|-------|
| **SaaStr** | B2B SaaS founders, growth, sales | Jason Lemkin's outlet; pitch contributed posts |
| **First Round Review** | Operator-level B2B insights | High bar; original frameworks only |
| **OpenView** | SaaS metrics, PLG, pricing | Often runs research-backed pieces |
| **Lenny's Newsletter** | Product / growth | Subscriber-only; pitch via Lenny directly on X |
| **Future (a16z)** | Tech + culture pieces | Long-form, original thinking required |
| **Stratechery** | Tech strategy analysis | Don't pitch Ben Thompson; engage via responses |
---
## AI / ML Press (Tier 1–2)
The hottest beat right now. Reporters here are inundated — your angle has to be sharp.
| Outlet | Best for | Notes |
|--------|----------|-------|
| **The Decoder** | AI news, fast-turn | Daily AI news cycle |
| **Import AI** (Jack Clark) | Weekly AI newsletter | Pitch research-backed angles |
| **The Batch** (Andrew Ng) | AI industry analysis | Submitted by deeplearning.ai team |
| **MIT Technology Review** | AI policy, capability, ethics | Higher bar, slower cycle |
| **Hugging Face blog** | Open-source AI tools | Contributed posts welcome |
| **Latent Space** | Practitioner-focused AI/ML | Podcast + newsletter |
---
## Developer & DevTools Press
Where to pitch if your audience is engineers.
| Outlet | Best for | Notes |
|--------|----------|-------|
| **The New Stack** | DevTools, infra, cloud-native | Active contributor program |
| **InfoQ** | Enterprise dev, software architecture | Long-form technical pieces |
| **DEV.to** | Self-published, community-driven | Build credibility before pitching |
| **Hacker Noon** | Tech blogging platform | Easy to publish but low signal |
| **Console** | DevTools newsletter | Curated weekly; submit projects |
| **DevTools FM** | Podcast | Pitch as guest |
---
## Business & Marketing Press
For pitching the business / marketing angle of your story.
| Outlet | Best for | Notes |
|--------|----------|-------|
| **Inc.** | Founder stories, business advice | Contributor program available |
| **Entrepreneur** | Small business, growth | Volume publisher; quality varies |
| **HBR.org** | Original research, frameworks | High bar; long lead time |
| **MarketingProfs** | B2B marketing | Contributor-friendly |
| **Marketing Brew (Morning Brew)** | Daily marketing newsletter | Reporter-driven, pitch directly |
| **Marketing Land / Search Engine Journal** | SEO, SEM, channels | Niche but high-intent audience |
| **Reforge** | Growth, product, retention | Original framework required |
---
## Newsletters (Reporter-Driven)
Newsletters are increasingly the most valuable PR placement — small audiences, but high-intent.
| Newsletter | Audience | How to pitch |
|-----------|---------|--------------|
| **Lenny's Newsletter** | Product/growth, ~600k readers | DM Lenny on X with sharp angle |
| **The Pragmatic Engineer** (Gergely Orosz) | Software eng leaders | Pitch via email; original engineering insights |
| **The Generalist** (Mario Gabriele) | Tech business analysis | Pitch via email; deep angles |
| **Stratechery** (Ben Thompson) | Tech strategy | Don't pitch; engage in replies |
| **Newcomer** (Eric Newcomer) | Tech business + scoops | Tips welcome via email |
| **Platformer** (Casey Newton) | Tech + policy | Pitch via Substack |
| **Big Technology** (Alex Kantrowitz) | Big Tech analysis | Pitch via Substack |
| **Not Boring** (Packy McCormick) | Tech + culture | Tough to crack; original takes only |
**For B2B SaaS:**
- **SaaStr Daily**
- **Demand Curve**
- **Growth Unhinged** (Kyle Poyar)
- **Trends.vc**
**For AI:**
- **Import AI** (Jack Clark)
- **The Batch** (Andrew Ng)
- **The Algorithm** (MIT Tech Review)
- **Interconnects** (Nathan Lambert)
- **AI Tidbits**
---
## Podcasts
A podcast appearance is often higher leverage than a press hit — longer engagement, evergreen replay, audience trust transfer.
### Top SaaS / startup podcasts
- **Lenny's Podcast** — product/growth
- **My First Million** — founder stories, business ideas
- **The Twenty Minute VC** — funding angle
- **SaaStr Podcast** — B2B SaaS
- **Acquired** — deep-dive company stories (don't pitch unless you're a unicorn)
- **The All-In Podcast** — broad tech/business (very hard to get on)
### Top AI podcasts
- **No Priors** (Sarah Guo, Elad Gil)
- **Latent Space**
- **Practical AI**
- **The TWIML AI Podcast**
- **The Cognitive Revolution**
### Top dev / engineering podcasts
- **The Changelog**
- **Software Engineering Daily**
- **DevTools FM**
### How to pitch a podcast
1. **Listen to 3 episodes** — non-negotiable
2. **Find the host's preferred channel** (X DM, email, guest form)
3. **Pitch a topic, not yourself** — "I'd love to come on and talk about [specific angle]" not "I'd love to be a guest"
4. **Bring evidence** — links to other appearances, your unique angle, what listeners will learn
5. **Make it easy** — bio, headshot, suggested questions in the pitch
---
## Industry / Vertical Press
Don't overlook trade press — smaller audience, much higher intent.
| Vertical | Outlets to investigate |
|----------|----------------------|
| **Marketing** | Adweek, Marketing Brew, AdAge, Search Engine Land |
| **Sales** | Sales Hacker, Modern Sales Pros |
| **HR / People Ops** | HR Brew, SHRM, HR Dive |
| **Finance / Fintech** | The Block, CoinDesk, Finextra |
| **Healthcare tech** | STAT, MobiHealthNews, Healthcare IT News |
| **Education tech** | EdSurge, EdScoop |
| **Real estate tech** | Inman, The Real Deal |
| **Legal tech** | Above the Law, Law360 |
| **Climate tech** | Heatmap, Canary Media, Latitude Media |
| **Devtools** | The New Stack, InfoQ, DevOps.com |
For your specific vertical: Google `"top publications" + "[your industry]"` and run the same scoring exercise from [journalist-pitching.md](journalist-pitching.md).
---
## Regional / Local
If your company has a regional angle (HQ location, customer concentration, government contract), local press is underrated.
- **Local business journals** — Bizjournals network covers 40+ US cities
- **Local NPR affiliates** — high-quality, business angle welcomed
- **Local TV business segments** — high reach, easy to get
- **State / regional tech news** — e.g., Built In (Chicago, Austin, etc.), TechBuzz (Utah)
---
## What's NOT On This List (And Why)
- **Product Hunt, BetaList, Indie Hackers** — these are directories, not press. Use the `directory-submissions` skill.
- **Press release wires** (PRNewswire, BusinessWire, GlobeNewswire) — overpriced for early-stage; journalists ignore them. Skip until you have IR / SEC requirements.
- **"As featured in" badge mills** — paid "media coverage" services. Worthless and damaging.
- **Random "guest post" SEO link networks** — Google penalizes these. Don't.
---
## Maintaining This List
This list will go stale. Recommended cadence:
- **Quarterly:** verify your tier-1 contacts are still at the outlet
- **Monthly:** add new outlets/newsletters relevant to your category
- **Per pitch:** confirm the journalist is still there before sending (check X / LinkedIn for "joined [new place]")
Store your live, working version in `.agents/media-list.md` (per [journalist-pitching.md](journalist-pitching.md)).
FILE:references/newsjacking.md
# Newsjacking — Reactive PR Workflow
Injecting your POV into a story that's already trending. Done well: free distribution off a wave of attention. Done badly: cringe at best, brand damage at worst.
## Contents
- When newsjacking works (and when it doesn't)
- The detect → score → angle → pitch loop
- Newsworthiness scoring rubric
- Story angle library
- Speed: the only thing that matters
- Sources & tooling
- Failure modes
---
## When Newsjacking Works
- **Tech/regulatory news in your category** — new law, new platform launch, competitor pivot, big acquisition
- **Industry data drops** — a major report drops, you have a sharper take or contradicting data
- **Public conversation** — a debate or controversy where your expertise is genuinely relevant
- **Seasonal/cyclical moments** — earnings season, year-end reviews, conference weeks
## When to Skip
- **Tragedies, accidents, deaths** — no exceptions. Don't.
- **Politically charged stories** unless your brand explicitly takes political stances
- **You have no genuine expertise** in the area
- **The window is already closed** — if a story is 48h+ old and you weren't first, you're late
- **The angle is "we have a product for this"** — that's marketing, not journalism
---
## The Loop
A repeatable workflow Claude can run on demand or daily.
1. **Detect** — surface trending stories in your category (see [Sources & Tooling](#sources--tooling))
2. **Score** — apply the [newsworthiness rubric](#newsworthiness-scoring-rubric); drop anything below threshold
3. **Angle** — generate 2–3 angles per story using the [angle library](#story-angle-library)
4. **Validate** — sanity-check: do you actually have the expertise/data to back this angle?
5. **Pitch** — draft a tight pitch to 3–5 journalists who cover this beat (see [journalist-pitching.md](journalist-pitching.md))
6. **Post** — also publish on your blog, LinkedIn, X — it builds the trail journalists check before quoting you
Output format Claude should produce:
```
NEWSJACK CANDIDATE — 2026-06-10
Story: "EU passes AI Act amendment requiring agent registration"
Source: TechCrunch, 3h ago
Score: 8/10 (high relevance, fresh, you have proprietary data)
Angles:
1. Data hot take: "Our analysis of 12,000 agent deployments shows 73% would fail this requirement"
2. Contrarian: "Why the registration rule will hurt safety, not improve it"
3. Customer story: "How [customer] is preparing — interview offer"
Recommended: #1 (you have unique data, strongest hook)
Pitch draft: [see journalist-pitching.md for template]
Target journalists: [list with rationale]
```
---
## Newsworthiness Scoring Rubric
Score each candidate 1–10 on five dimensions, multiply by the weight, then sum. Max possible: 80 (10 × the 8x weight total).
| Dimension | What it measures | Weight |
|-----------|------------------|--------|
| **Timeliness** | Story <24h old? Window still open? | 2x |
| **Relevance** | Genuinely in your expertise area? | 2x |
| **Angle uniqueness** | Can you say something no one else is saying? | 2x |
| **Authority** | Do you have data, customers, or experience to back it? | 1x |
| **Reach potential** | Will this story keep growing or has it peaked? | 1x |
**Threshold:** weighted total ≥ 50/80. Below that, skip.
**Auto-disqualify if:**
- The story is about something tragic
- Your angle is "I disagree" with nothing to back it
- You haven't actually formed an opinion — you just want to be quoted
---
## Story Angle Library
Use these templates to generate angles fast.
### 1. Data hot take
*"We analyzed [N] [things] after [event]. Here's what we found."*
Best when you have proprietary data. The journalist gets a stat, you get the citation.
### 2. Contrarian
*"Everyone says [popular take]. Here's why they're wrong."*
Best when you can defend the position with specifics. Weak when it's just contrarianism for attention.
### 3. "We predicted this"
*"Six months ago we wrote [thing] — here's what's happening now and what's next."*
Best when you actually did predict it. Lethal to your credibility if you didn't.
### 4. Customer impact
*"Here's a [customer type] who's directly affected. We can put you in touch."*
Best for B2B. Reporters love named customers willing to talk.
### 5. Insider explainer
*"This story is complicated. Here's what's actually happening."*
Best when most coverage is missing nuance. You're not arguing — you're educating.
### 6. Trend connector
*"This isn't isolated — it's part of a bigger shift we're seeing in [pattern]."*
Best when you have several data points or examples to connect.
### 7. Founder POV
*"As someone who's built in this space for [X years], here's the part most people are missing."*
Best for opinion pieces / op-eds. Weak as a soundbite pitch.
---
## Speed: The Only Thing That Matters
Newsjacking decays fast. Approximate windows:
| Story type | Effective window |
|-----------|------------------|
| Breaking tech news | 4–12 hours |
| Major regulation / policy | 24–48 hours |
| Industry report / data drop | 24–72 hours |
| Conference announcement | Same day |
| Acquisition / funding news | 12–24 hours |
**Implication:** if you can't draft and send within the window, don't bother. Set up the loop so detection → pitch takes <2 hours.
---
## Sources & Tooling
Reuses tooling from the `social` skill's listening workflow. Same install: `brew install jq`.
### Google News RSS (no auth)
```bash
# Replace QUERY with topic (use + for spaces, %22 for quotes)
curl -s "https://news.google.com/rss/search?q=QUERY&hl=en-US&gl=US&ceid=US:en" \
| xmllint --xpath "//item[position()<11]" - 2>/dev/null
```
### Hacker News (Algolia) for tech stories
```bash
SINCE=$(($(date +%s) - 86400))
curl -s "https://hn.algolia.com/api/v1/search_by_date?query=QUERY&tags=story&numericFilters=created_at_i>SINCE" \
| jq '.hits[] | {title, url, points, num_comments, created_at, hn_url: ("https://news.ycombinator.com/item?id="+.objectID)}'
```
### Reddit (for category-specific subs)
```bash
curl -s -A "newsjack/1.0" \
"https://www.reddit.com/r/SUBREDDIT/top.json?t=day&limit=15" \
| jq '.data.children[].data | {title, url, score, num_comments, created_utc}'
```
### Journalist research (browser-driven)
For finding *which* journalists are covering the story right now:
- **dev-browser** → Google News search for the story → click through to articles → note the bylines
- Then go to those journalists' X / LinkedIn / Muck Rack profile to confirm beat and recent coverage
See also [journalist-pitching.md](journalist-pitching.md) for the full discovery workflow.
### Source list
For repeatable monitoring, add a "Newsjacking topics" section to `.agents/listening-sources.md` (template in the `social` skill's references):
```markdown
## Newsjacking topics (Google News RSS)
- "AI agent regulation"
- "[your category] funding"
- "[your competitors] OR [adjacent competitors]"
## Industry data drops (RSS / manual)
- Pitchbook reports
- a16z state of [industry] reports
- [your category] benchmark reports
```
---
## Failure Modes
Things that have ended careers and brands.
- **Tragedy-jacking** — Oreo's 2013 Super Bowl tweet worked. Most attempts since have not. Wartime, disasters, deaths: don't.
- **The forced fit** — "Here's our take on [trending story] — it's actually about [our product]." Journalists see through this instantly.
- **The empty take** — pitching "we have an opinion" without specifics. Journalists need a quote-worthy line, not "we're closely watching this."
- **Speed without judgment** — being first with a bad take is worse than being late with a good one. The 30-minute "is this brand-appropriate?" gut check exists for a reason.
- **Pitching the same angle to 50 journalists** — they talk. Get caught once, lose the relationships.
- **No follow-through** — pitch goes out, journalist responds in 20 minutes, founder takes 6 hours to reply. Story moves on.
---
## Companion Practice: The Public Trail
Every newsjack pitch is stronger if the journalist can find evidence you've been thinking about this publicly. Before pitching:
1. Publish a short post (blog, LinkedIn, X thread) with your take
2. Reference it in the pitch ("more thinking here: [link]")
3. This signals you're not opportunistic — you're an actual voice in the space
If you don't have time to publish, you're probably not ready to pitch.
FILE:references/podcast-guest-prep.md
# Podcast Guest Prep
Build a prep brief before the user appears on a podcast as a guest. The goal: walk in knowing what's top of mind for the show, how it has evolved, who the hosts are, and which of the user's stories map onto what the show cares about *right now*.
**Why prep is worth real effort:** podcast guesting isn't just audience reach. Episodes get transcribed, show notes get published, and both get crawled and cited by AI assistants — when someone asks ChatGPT about your category, the stories you told on a podcast two years ago are part of what it draws on. A good appearance is earned media that compounds in AI answers for years (see the `ai-seo` skill). The stories you tell — and the concrete numbers in them — become the citable record on your brand. Prep accordingly.
## Context to load first
Read `.agents/product-marketing.md` (or `.claude/product-marketing.md`) for the company, positioning, and ICP. That file usually won't have the guest's *story bank*, so also collect — in one batch, not a drip:
1. What did you build before this that comes up in conversation?
2. What are 2–3 stories you tell well, with real numbers attached?
3. What's one opinion you hold that most people in your space disagree with?
Offer to save the answers into the product-marketing context doc so future runs skip the interview.
From their message, establish (ask only if missing and it matters): the podcast name or URL, whether they've appeared before (get the prior episode link — it anchors the progression analysis and the callbacks), and roughly when they're recording.
## Research sequence
Work through sources in this order; each is a fallback for the last:
1. **RSS feed first.** The richest source: full episode descriptions, chapter markers, guest links, dates. Find the feed link on the podcast site (Buzzsprout, Transistor, etc. all expose one). Large-feed fetches may truncate — check whether the oldest episodes you need actually made it.
2. **The podcast website's episode list** for anything the feed missed. These pages often lazy-load older episodes via JavaScript; if pagination returns nothing, note the gap and move on rather than burning time.
3. **Apple Podcasts show page** — reliably renders the latest ~8 episodes with full descriptions.
4. **Web search** for stray episodes, the hosts, and the show's reputation.
Don't fetch every episode page. Descriptions plus chapter lists are almost always enough; only pull a full transcript when a specific episode is central (e.g., a debate the guest should have a position on). Check for published transcript links in the feed.
## What to extract
- **Recent-episode threads** (last ~3 months or 6–8 episodes): per-episode topic summaries, then the *recurring threads* — the questions the hosts keep returning to. Threads matter more than individual episodes; they predict the questions the guest will get.
- **Show progression** (since their last appearance, or ~12–18 months if first time): identify phases and the inflection point where the show's focus shifted. Note whether the show re-invites guests (signals how a return visit fits) and whether hosts launched side projects.
- **Host profiles.** Sources: the show's about pages, hosts' personal sites, LinkedIn, and — often the best source — episodes where the hosts guest on *other* shows and introduce themselves. Capture day job, background, what they've personally been building (mine solo-episode summaries), and social handles. If a host shares the guest's first name, flag it and keep references unambiguous throughout the brief.
- **Prior appearance recap** (if returning): what was actually discussed, with rough timestamps, and how much airtime the guest's current company got. This sets up the "what's changed since" narrative.
## The brief
Write a markdown file and present the short version in chat. Structure:
1. **Big picture** — what kind of show this is now, and the one-paragraph read on how the guest should position themselves
2. **Show progression** — the phases since their last appearance (or show start)
3. **Recent episodes in detail** — per-episode notes, then the recurring threads
4. **Guest angles** — their stories mapped explicitly onto the show's threads, callbacks to any prior appearance, and 2–3 "pocket" items: concrete stories with numbers to have ready
5. **The hosts** — profiles plus rapport hooks (where each host's world overlaps the guest's)
6. **Gaps** — anything unretrievable, and offers to fill them
Keep it tight — a doc they'll skim before recording, not a report.
**Finding angles:** map the story bank onto the show's recurring threads. The shape to look for: a "data moats" thread maps to the guest's proprietary dataset as a live case study; an "AI replacing niche tools" debate maps to a defensibility story from their product history. One well-chosen contrarian take stands out most on shows that have converged on a consensus.
**AI-visibility angle:** since the transcript becomes the record, coach the guest to say the important things in liftable form — the company name next to the category ("we build X, the Y for Z"), and numbers spoken aloud, not gestured at. Same logic as the YouTube text layer in `ai-seo`.
## Follow-ups to offer (don't auto-run)
Transcribing their prior episode for a word-level review, pulling a full transcript of one pivotal recent episode, or drafting likely Q&A.
---
*Distilled and adapted from [ai-visibility-skills](https://github.com/Knowatoa/ai-visibility-skills) by Knowatoa (MIT), reused with credit.*
FILE:references/press-platforms.md
# Press Request Platforms — Inbound PR
Journalists posting "I need a source for X" — you respond, sometimes you get quoted, sometimes you don't. The cheapest PR play available, but only if you treat it seriously.
## Contents
- The major platforms
- Daily triage workflow
- Response template
- What makes a response get selected
- What kills a response
- ROI reality check
---
## The Major Platforms
| Platform | What it is | Cost | Quality |
|----------|-----------|------|---------|
| **[Connectively](https://www.connectively.us)** (formerly HARO) | Daily email digest of journalist queries | Free tier; paid for filters | Mixed — high volume, lots of noise |
| **[Qwoted](https://www.qwoted.com)** | Web app with journalist requests | Free; paid for outreach | Good — better-quality outlets |
| **[Featured](https://featured.com)** | Web app, expert profiles, journalist requests | Free tier; paid pro | Good for thought-leadership snippets |
| **[Help A B2B Writer](https://helpab2bwriter.com)** | Twice-weekly email of B2B queries | Free | High — B2B-focused, low spam |
| **[SourceBottle](https://www.sourcebottle.com)** | Australia-focused but global queries | Free | Variable |
| **[Terkel](https://terkel.io)** | Roundup-style ("we asked 50 experts…") | Free | Volume-heavy, low effort |
| **[JournoRequests](https://twitter.com/journorequests)** | X account aggregating tweets | Free | UK-skewed, real-time |
| **#JournoRequest** (X hashtag) | Live journalist requests | Free | Real-time, fast-moving |
**Recommended starter set:** Connectively + Qwoted + Help A B2B Writer + monitoring `#JournoRequest` on X.
---
## Daily Triage Workflow
These platforms generate volume. Treat it like email triage — fast pass, deep response on the rare matches.
### Step 1 — Filter (5 min)
For each digest / request feed:
- Drop everything where you don't have **direct experience or data**
- Drop everything from outlets your ICP doesn't read
- Drop everything with a deadline you can't meet
- Keep only requests where you can give a **complete, named, on-the-record answer**
Realistic conversion: 50 daily requests → 2–4 worth answering.
### Step 2 — Deep response (15 min per request)
For each keeper:
- Read the request 3 times — what's the *actual* angle?
- Look up the journalist if possible — recent coverage, beat
- Write a custom response (see [template](#response-template))
- Send within their stated deadline (early > late)
### Step 3 — Log
Track in a spreadsheet:
- Date
- Platform
- Journalist + outlet
- Topic
- Response sent (yes/no)
- Outcome (no response / passed / quoted / linked)
After 30 responses, you'll see which topics/platforms convert.
---
## Response Template
Keep responses under 200 words. Journalists are scanning 50+ replies for one quote.
```
Hi [name],
Quick response to your request about [topic].
[Specific credential — 1 sentence. "Built X for 5 years" / "Led marketing at Y" / "Have analyzed N companies in space"]
The most important thing about [topic]: [your actual point in 2 sentences].
[A specific example, story, or data point — this is what gets quoted.]
[If applicable: a contrarian or surprising angle that differentiates from typical answers.]
Happy to expand on any of this, share data, or be quoted directly.
Feel free to use this attribution:
[Your name], [your title], [your company]
Contact for follow-up: [email + phone]
```
**Note the structure:**
1. One-sentence intro
2. One-sentence credential
3. Two-sentence answer
4. Specific example (the quotable part)
5. Optional: differentiator
6. Clear offer
7. Pre-written attribution (saves them 30 seconds)
8. Contact info
---
## What Makes a Response Get Selected
After analyzing hundreds of quoted responses, the patterns:
### Quotable specificity
**Bad:** "Companies should focus on customer experience."
**Good:** "When we A/B tested 47 onboarding flows, the version with a 30-second video at step 3 increased activation by 41%."
The good version is a quote. The bad version is filler.
### Concrete credential
**Bad:** "As a marketing expert..."
**Good:** "I've run growth at three Series B SaaS companies, all in B2B sales tooling."
Specificity beats title-stacking.
### Story over advice
**Bad:** "It's important to track the right metrics."
**Good:** "We almost shut down because we were optimizing for MRR when our real problem was activation. Once we switched to tracking 7-day activation, everything else followed."
Stories make articles. Advice makes filler.
### Pre-formatted for their workflow
- Pre-written attribution
- Multiple quotable lines (let them pick)
- High-res headshot link (don't attach)
- One-line company description
### Time match
**Most quotes come from responses sent in the first 6 hours.** After 24 hours, your chances drop sharply. Treat deadlines as if they're 24h earlier than stated.
---
## What Kills a Response
- **Pitching your product** when they asked for expert commentary
- **Generic advice** that could come from any expert
- **Multiple "experts" from your company** responding to the same request (looks coordinated, often is)
- **Hiring a PR firm to spam responses** — journalists smell it
- **Demanding a link back** to your site — most can't promise links
- **Ignoring the deadline** by 1+ days
- **Long bio sections** before the actual answer
- **Asking to "see the article before publication"** — you don't get to do that
- **Asking what other experts said** so you can differentiate — they won't tell you
---
## ROI Reality Check
Most teams overinvest in these platforms because they're cheap. Be honest:
| Effort | Realistic outcome (90 days) |
|--------|----------------------------|
| 5 hr/week, custom responses | 3–10 quoted placements |
| 1 hr/week, template responses | 0–2 placements |
| Outsourced to PR firm | Lots of submissions, few quotes |
**A quote in a tier-1 outlet is worth:**
- A backlink (DR depends on outlet)
- A sales-collateral asset ("As featured in...")
- AI-citation surface area
- Brand legitimacy in the abstract
**A quote in a tier-3 outlet is worth:**
- A backlink, often nofollow
- Maybe an Instagram screenshot
**Decision rule:** if you can sustain 5 hr/week of quality responses for 90 days, this is worth it. If you can only do 1 hr/week, skip it and invest in [proactive pitching](journalist-pitching.md) instead.
---
## Setup Checklist
Before you start responding:
- [ ] Press page exists and is current (see main SKILL.md)
- [ ] One-line credential is written and rehearsed
- [ ] Headshot is high-res and at a public URL
- [ ] You have 3–5 specific stories / data points ready to deploy
- [ ] You've decided which 2–3 platforms to use (don't try all 7)
- [ ] You've blocked a daily 20-min window for triage
- [ ] You're logging responses in a tracker
Without these, you're spamming and wasting their time and yours.
FILE:references/story-angles.md
# Story Angles — What Earns Press
Distilled from Corey Haines's *Founding Marketing*, ch. 5. Chase press for the **compound effect** — authority, journalist relationships, AI-citation surface — not the traffic bump. The bump fades in a day; the backlinks, the relationships, and the citable record don't.
**The story is not your product.** Journalists write about trends, data, conflict, and humans. Your product is the evidence, never the headline. If the pitch is "we built a thing," there's no story. If the pitch is "here's a shift happening and we're proof of it," there is.
## Contents
- The three story angles
- Newsworthy moments (data stories)
- Build media relationships before you need them
- The PR flywheel (repurposing map)
---
## The Three Story Angles
Every earned-media story worth pitching bends toward one of three shapes. Pick the one that fits your moment — don't force all three.
### 1. Founding Story
The origin, the build, the numbers-in-public. Works because people follow *people*, and a founder with skin in the game is quotable in a way a product never is.
- **When to use:** you're early, you have a sharp personal reason you exist, and you're willing to build in public and share real revenue.
- **Exemplar:** Pieter Levels' "12 startups in 12 months" — shipping in the open, posting MRR screenshots, letting the challenge itself be the story. The build *was* the coverage.
- **How to run it:** share the arc (why you started, what you've learned, where the numbers are now), not the feature list. Revenue transparency is the hook most founders are too scared to use.
### 2. David vs Goliath
Own your size. Being small against an incumbent is an advantage in the story, not a liability — you're the underdog readers root for.
- **When to use:** there's a giant in your category and you do one thing dramatically better than they do.
- **Exemplar:** Superhuman vs Gmail — a tiny team reframing "we're small" into "we're the fast, obsessive alternative to the default everyone tolerates."
- **How to run it:** name the Goliath, name the one thing you beat them at, and let the size gap do the emotional work. Don't hide that you're small — lead with it.
### 3. Have an Enemy
Give the reader something to be against. The enemy is a **broken system**, not a competitor. Attacking a rival looks petty; attacking a genuinely broken status quo looks principled — and journalists cover principled fights.
- **When to use:** there's a systemic wrong in your space you can credibly stand against, and you're willing to take a real position.
- **Exemplar:** HEY vs Apple's App Store (2020) — the fight wasn't "Basecamp vs Apple the company," it was against App Store rules founders saw as broken. It became a press cycle *and* a policy conversation.
- **How to run it:** define the broken system precisely, stake a clear position, and make sure you're actually willing to be quoted taking that stance. A half-hearted enemy reads as a marketing stunt.
**Guardrail:** the enemy must be a system or a norm — never a named competitor. "Company X is bad" is a smear; "this way of doing things is broken" is a movement.
---
## Newsworthy Moments: Data Stories
Beyond the three angles, the most reliably pitchable moment is a **data story** — proprietary numbers no one else has, packaged as an industry benchmark. Journalists can build a whole piece around a stat; you get the citation.
- **Exemplars:** Stripe's *State of New User* / annual data reports; Intercom's benchmark reports. Each turns internal data into an annual, citable, must-cover event.
- **Why it works:** it's original, it's quotable, and it positions you as the source-of-record for your category's numbers. It also compounds in AI answers — benchmark stats get lifted and re-cited for years.
- **How to run it:** find the one number in your data that surprises, wrap it in methodology, and package it as a standalone one-pager (not your homepage). See [journalist-pitching.md](journalist-pitching.md) → the "Data story" template.
**Milestone reminder:** a milestone alone ("we hit $1M ARR") isn't a story — milestone *with narrative* is. Attach the number to a shift, a lesson, or a David-vs-Goliath frame.
---
## Build Media Relationships Before You Need Them
Cold-pitching a journalist the day you have news is the hard way. The founders who get covered built the relationship months earlier. It's a ladder — climb it before you need the favor.
1. **Follow their beat.** Read their last 5 articles. Know what they cover and what they're tired of covering. (See the discovery checklist in [journalist-pitching.md](journalist-pitching.md).)
2. **Add value with no ask.** Reply to their work with a genuinely useful data point or correction. Share their pieces. Answer a question they post on X. Give before you take.
3. **Be a reliable source.** When they need a quote fast, be the person who responds in 20 minutes with a clean, quotable line — even when there's nothing in it for you. Reliability is what turns a contact into a relationship.
Do this for a handful of journalists on your beat and, when you finally have a real story, you're pitching a relationship, not a stranger.
---
## The PR Flywheel (Repurposing Map)
One core asset feeds every other channel. Don't create five things — create one exceptional thing and reformat it five ways. Each format extends the reach of the same idea and gives journalists more of a public trail to find.
```
Core asset (data report / strong POV / founding story)
│
├─→ Thread (X / LinkedIn) — the hook, teased publicly
├─→ Article (your blog) — the full argument, the destination
├─→ Guest post (someone else's audience) — borrow reach
├─→ Podcast appearance — the spoken, citable version
└─→ Video — the durable, AI-crawled version
```
- **Sequence matters less than coverage.** Publish the article as the owned home base, tease it as a thread, pitch the guest post and podcast off the same idea, cut the video from the recording.
- **Every format strengthens the next pitch.** A journalist who can find your thread, article, and podcast on a topic sees a voice in the space — not an opportunist. This is the public trail that makes newsjacking land (see [newsjacking.md](newsjacking.md) → The Public Trail).
- **The asset is the flywheel's fuel.** One strong data story or POV can run this loop for weeks. Weak assets stall it immediately.
Quét và tối ưu SEO cho README.md và trang docs: meta tag, heading, từ khóa, khả năng đọc, link hỏng, cập nhật sitemap.xml.
---
name: seo-auditor
description: |
Scan and optimize documentation files for SEO. Audits README.md files and docs/ pages for
meta tags, headings, keywords, readability, duplicate content, and broken links. Applies
fixes, updates sitemap.xml, and generates a report. Usage: /seo-auditor [path]
---
# /seo-auditor
Systematically scan, audit, and optimize documentation files for SEO. Targets README.md files and docs/ pages — fixes issues in place, preserves rankings on high-performing pages, and generates a final report.
## Usage
```bash
/seo-auditor # Audit all docs/ and root README.md
/seo-auditor docs/skills/ # Audit a specific docs subdirectory
/seo-auditor --report-only # Scan without making changes
```
## What It Does
Execute all 7 phases sequentially. Auto-fix non-destructive issues. Preserve existing high-ranking content. Report everything at the end.
---
## Phase 1: Discovery & Baseline
### 1a. Identify target files
Scan for documentation files that need SEO audit:
```bash
# Find all markdown files in docs/ and root README files
find docs/ -name '*.md' -type f | sort
find . -maxdepth 2 -name 'README.md' -not -path './.codex/*' -not -path './.gemini/*' | sort
```
Classify each file:
- **New/recently modified** — files changed in the last 2 commits (check via `git log`)
- **Index pages** — `index.md` files (high authority, handle with care)
- **Skill pages** — `docs/skills/**/*.md` (generated by `generate-docs.py`)
- **Static pages** — `docs/index.md`, `docs/getting-started.md`, `docs/integrations.md`, etc.
- **README files** — root and domain-level README.md
### 1b. Capture baseline
For each target file, extract current SEO state:
- `title:` frontmatter field → becomes `<title>` tag
- `description:` frontmatter field → becomes `<meta name="description">`
- First `# H1` heading
- All `## H2` and `### H3` subheadings
- Word count
- Internal link count
- External link count
Store baseline in memory for the report.
---
## Phase 2: Meta Tag Audit
For every file with YAML frontmatter, check and fix:
### Title Tag (`title:`)
**Rules:**
- Must exist and be non-empty
- Length: 50-60 characters ideal (Google truncates at ~60)
- Must contain a primary keyword
- Must NOT duplicate another page's title
- For skill pages: should follow the pattern `{Skill Name} — {Differentiator} - {site_name}`
- site_name from `mkdocs.yml` is appended automatically — don't duplicate it in the title
**Auto-fix:** If title is generic (e.g., just the skill name), enrich it with domain context using the DOMAIN_SEO_SUFFIX pattern from `scripts/generate-docs.py`.
### Meta Description (`description:`)
**Rules:**
- Must exist and be non-empty
- Length: 120-160 characters (Google truncates at ~160)
- Must contain the primary keyword naturally
- Must be unique across all pages — no two pages share the same description
- Should include a call-to-action or value proposition
- Must NOT start with "This page..." or "This document..."
**Auto-fix:** If description is missing or generic, generate one from the SKILL.md frontmatter description (if available) or from the first paragraph of content. Use the `extract_description_from_frontmatter()` function from `generate-docs.py` as reference.
### Validation Script
Run on each file that has HTML output in `site/`:
```bash
python3 marketing-skill/seo-audit/scripts/seo_checker.py --file site/{path}/index.html
```
Parse the score. Flag any page scoring below 60.
---
## Phase 3: Content Quality & Readability
For each target file, analyze and improve:
### Heading Structure
**Rules:**
- Exactly one `# H1` per page
- H2s follow H1, H3s follow H2 — no skipping levels
- Headings should contain keywords naturally (not stuffed)
- No duplicate headings on the same page
**Auto-fix:** If heading levels skip (H1 → H3), adjust to proper hierarchy.
### Readability
Run the content scorer on each file:
```bash
python3 marketing-skill/content-production/scripts/content_scorer.py {file_path}
```
Check scores for:
- **Readability** — aim for score ≥ 70
- **Structure** — aim for score ≥ 60
- **Engagement** — aim for score ≥ 50
### Content Quality Rules
- **Paragraphs:** No single paragraph longer than 5 sentences
- **Sentences:** Average sentence length 15-20 words
- **Passive voice:** Less than 15% of sentences
- **Transition words:** At least 30% of sentences use transitions
- **Bullet lists:** Use lists for 3+ items instead of comma-separated inline lists
### AI Content Detection
Run the humanizer scorer on non-generated content (README.md files, static pages):
```bash
python3 marketing-skill/content-humanizer/scripts/humanizer_scorer.py {file_path}
```
Flag pages scoring below 50 (too AI-sounding). For these pages, apply voice techniques from `marketing-skill/content-humanizer/references/voice-techniques.md`:
- Replace AI clichés ("delve into", "leverage", "it's important to note")
- Vary sentence length
- Add specific examples instead of generic statements
- Use active voice
**Important:** Only modify content that was recently created or updated. Do NOT rewrite pages that are ranking well — preserve their content.
---
## Phase 4: Keyword Optimization
### 4a. Identify target keywords per page
Based on the page's purpose and domain:
| Page Type | Primary Keywords | Secondary Keywords |
|-----------|-----------------|-------------------|
| Homepage (docs/index.md) | "Claude Code Skills", "agent plugins" | "Codex skills", "Gemini CLI", "OpenClaw" |
| Skill pages | Skill name + "Claude Code" | "agent skill", "Codex plugin", domain terms |
| Agent pages | Agent name + "AI coding agent" | "Claude Code", "orchestrator" |
| Command pages | Command name + "slash command" | "Claude Code", "AI coding" |
| Getting started | "install Claude Code skills" | platform names |
| Domain index | Domain + "skills" + "plugins" | "Claude Code", platform names |
### 4b. Keyword placement checks
For each page, verify the primary keyword appears in:
- [ ] Title tag (frontmatter `title:`)
- [ ] Meta description (frontmatter `description:`)
- [ ] H1 heading
- [ ] First paragraph (within first 100 words)
- [ ] At least one H2 subheading
- [ ] Image alt text (if images present)
- [ ] URL slug (for new pages only — never change existing URLs)
### 4c. Keyword density
- Primary keyword: 1-2% of total word count
- Secondary keywords: 0.5-1% each
- No keyword stuffing — if density exceeds 3%, reduce it
**Important:** Never change URLs of existing pages. URL changes break incoming links and destroy rankings. Only optimize content and meta tags.
---
## Phase 5: Link Audit
### 5a. Internal links
For each target file, check all markdown links `[text](url)`:
- Verify the target exists (file path resolves)
- Check for broken relative links (`../`, `./`)
- Verify anchor links (`#section-name`) point to existing headings
**Auto-fix:** Use the `rewrite_skill_internal_links()` and `rewrite_relative_links()` functions from `generate-docs.py` as reference. Rewrite broken skill-internal links to GitHub source URLs.
### 5b. Duplicate content detection
Compare meta descriptions across all pages:
```bash
grep -rh '^description:' docs/**/*.md | sort | uniq -d
```
If duplicates found, make each description unique by adding page-specific context.
Compare H1 headings across all pages — no two pages should have the same H1.
### 5c. Orphan page detection
Check if every page in `docs/` is referenced in `mkdocs.yml` nav. Pages not in nav are orphans — they won't appear in navigation and may not be indexed.
```bash
# Find doc pages not in mkdocs nav
find docs -name '*.md' -not -name 'index.md' | while read f; do
slug=$(echo "$f" | sed 's|docs/||')
grep -q "$slug" mkdocs.yml || echo "ORPHAN: $f"
done
```
**Auto-fix:** Add orphan pages to the correct nav section in `mkdocs.yml`.
---
## Phase 6: Sitemap & Build
### 6a. Rebuild the site
```bash
mkdocs build
```
This regenerates `site/sitemap.xml` automatically (MkDocs Material generates it during build).
### 6b. Verify sitemap
Check the generated sitemap:
```bash
python3 marketing-skill/site-architecture/scripts/sitemap_analyzer.py site/sitemap.xml
```
Verify:
- All documentation pages appear in the sitemap
- No broken/404 URLs
- URL count matches expected page count
- Depth distribution is reasonable (no pages deeper than 4 levels)
### 6c. Check for sitemap issues
- **Missing pages:** Pages in `mkdocs.yml` nav that don't appear in sitemap
- **Extra pages:** Pages in sitemap that aren't in nav (orphans)
- **Duplicate URLs:** Same page accessible via multiple URLs
---
## Phase 7: Report
Generate a concise report for the user:
```
╔══════════════════════════════════════════════════════════════╗
║ SEO AUDITOR REPORT ║
╠══════════════════════════════════════════════════════════════╣
║ ║
║ Pages scanned: {n} ║
║ Issues found: {n} ║
║ Auto-fixed: {n} ║
║ Manual review needed: {n} ║
║ ║
║ META TAGS ║
║ Titles optimized: {n} ║
║ Descriptions fixed: {n} ║
║ Duplicate titles: {n} → {n} (fixed) ║
║ Duplicate descs: {n} → {n} (fixed) ║
║ ║
║ CONTENT ║
║ Readability improved: {n} pages ║
║ Heading fixes: {n} ║
║ AI score improved: {n} pages ║
║ ║
║ KEYWORDS ║
║ Pages missing primary keyword in title: {n} ║
║ Pages missing keyword in description: {n} ║
║ Pages with keyword stuffing: {n} ║
║ ║
║ LINKS ║
║ Broken links found: {n} → {n} (fixed) ║
║ Orphan pages: {n} → {n} (added to nav) ║
║ Duplicate content: {n} → {n} (deduplicated) ║
║ ║
║ SITEMAP ║
║ Total URLs: {n} ║
║ Sitemap regenerated: ✅ ║
║ ║
║ PRESERVED (no changes — ranking well) ║
║ {list of pages left untouched} ║
║ ║
╚══════════════════════════════════════════════════════════════╝
```
### Pages to preserve (do NOT modify)
These pages rank well for their target keywords. Only fix critical issues (broken links, missing meta). Do NOT rewrite content:
- `docs/index.md` — homepage, ranks for "Claude Code Skills"
- `docs/getting-started.md` — installation guide
- `docs/integrations.md` — multi-tool support
- Any page the user explicitly marks as "preserve"
---
## Skill References
| Tool | Path | Use |
|------|------|-----|
| SEO Checker | `marketing-skill/seo-audit/scripts/seo_checker.py` | Score HTML pages 0-100 |
| Content Scorer | `marketing-skill/content-production/scripts/content_scorer.py` | Score content readability/structure/engagement |
| Humanizer Scorer | `marketing-skill/content-humanizer/scripts/humanizer_scorer.py` | Detect AI-sounding content |
| Headline Scorer | `marketing-skill/copywriting/scripts/headline_scorer.py` | Score title quality |
| SEO Optimizer | `marketing-skill/content-production/scripts/seo_optimizer.py` | Optimize content for target keyword |
| Sitemap Analyzer | `marketing-skill/site-architecture/scripts/sitemap_analyzer.py` | Analyze sitemap structure |
| Schema Validator | `marketing-skill/schema-markup/scripts/schema_validator.py` | Validate structured data |
| Topic Cluster Mapper | `marketing-skill/content-strategy/scripts/topic_cluster_mapper.py` | Group pages into content clusters |
### Reference Docs
| Reference | Path | Use |
|-----------|------|-----|
| SEO Audit Framework | `marketing-skill/seo-audit/references/seo-audit-reference.md` | Priority order for SEO fixes |
| AI Search Optimization | `marketing-skill/ai-seo/references/content-patterns.md` | Make content citable by AI |
| Content Optimization | `marketing-skill/content-production/references/optimization-checklist.md` | Pre-publish checklist |
| URL Design Guide | `marketing-skill/site-architecture/references/url-design-guide.md` | URL structure best practices |
| Internal Linking | `marketing-skill/site-architecture/references/internal-linking-playbook.md` | Internal linking strategy |
| AI Writing Detection | `marketing-skill/content-humanizer/references/ai-tells-checklist.md` | AI cliché removal |
Ghi nhận nhận diện thương hiệu qua 10 câu hỏi (màu, phông chữ, phong cách, thư mục xuất) và kiểm tra độ tương phản văn bản, liên kết.
---
name: design-system
description: Captures the user's brand identity once via a 10-question onboarding wizard (primary/accent HEX + heading + body Google Fonts + design style editorial/technical/minimal/playful + default output directory + syntax theme + TOC behavior + optional logo/company), validates body-text and link contrast against WCAG 2.2 AA, derives 12 CSS custom properties in HSL space, and stores the result for every markdown-html converter to consume. Use before any markdown-html conversion. Triggers on first-run onboarding ("set up the brand", "configure markdown-html", "run onboarding"), on explicit reset ("reset the design system", "re-onboard"), and is checked by every converter via config_loader.py before rendering. Refuses to save if body-text contrast fails AA 4.5:1 or the output dir isn't writable. Precedence: project (./.markdown-html/) > global (~/.config/markdown-html/) > built-in defaults; MARKDOWN_HTML_NO_CONFIG=1 bypasses.
version: 2.10.0
author: Alireza Rezvani
license: MIT
tags: [design-system, brand-palette, wcag, onboarding, customization, markdown-html, css-variables, typography]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# Design System — Onboarding + Shared Brand Tokens
The design-system skill is the **shared brand owner** for the markdown-html plugin. Run its onboarding once. Every converter (`md-document`, `md-review`, `md-slides`) reads the resulting config via `config_loader.py` and applies the same 12 CSS custom properties to its output. Without this, conversions render with placeholder defaults — technically functional but unbranded.
This skill ships exactly three Python tools:
1. **`onboard.py`** — interactive (or `--defaults` / `--set` / `--show` / `--reset`) wizard.
2. **`config_loader.py`** — importable customization loader with project > global > defaults precedence and `MARKDOWN_HTML_NO_CONFIG=1` bypass.
3. **`brand_palette_validator.py`** — WCAG-AA contrast checker + HSL palette deriver.
All three are stdlib-only and contain no LLM calls (deterministic per Path-B discipline).
## When to invoke
| Symptom | Action |
|---|---|
| User says "convert this markdown to HTML" for the first time in this workspace | Run `python3 markdown-html/skills/design-system/scripts/onboard.py` |
| `~/.config/markdown-html/design-system.json` doesn't exist OR `setup_completed_at` is null | Refuse conversion, surface onboarding |
| User wants per-repo brand override | `python3 .../onboard.py --scope project` |
| User wants to change a single field non-interactively | `python3 .../onboard.py --set brand.primary=#FF6B35` |
| User wants to reset and re-onboard | `python3 .../onboard.py --reset` then re-run |
| User wants zero-touch defaults (CI, ephemeral session) | `python3 .../onboard.py --defaults` |
| Headless / containerized run that should ignore saved config | `MARKDOWN_HTML_NO_CONFIG=1 ...` |
## Onboarding question set (10 questions)
| # | Key | Choices / Validator | Default |
|---|---|---|---|
| 1 | `default_output_dir` | path; `os.access(parent, os.W_OK)` | `./markdown-html-out/` |
| 2 | `brand.primary` | HEX `^#?[0-9a-fA-F]{6}$` | `#0A1628` |
| 3 | `brand.accent` | HEX or blank (auto-derive) | derive from primary |
| 4 | `typography.heading_font` | Google Font name (12 safe defaults) | `Inter` |
| 5 | `typography.body_font` | Google Font name | `Inter` |
| 6 | `design_style` | `editorial / technical / minimal / playful` | `technical` |
| 7 | `code_theme` | `light / dark / auto` | `auto` |
| 8 | `toc.behavior` | `sticky-sidebar / collapsible-top / inline / none` | `sticky-sidebar` |
| 9 | `company_name` | string (may be empty) | `""` |
| 10 | `logo_url` | URL or empty (base64-embedded at render) | `""` |
## Hard rules
1. **WCAG AA body-text contrast must pass.** `brand_palette_validator.validate()` runs after every change. Body text on bg must reach 4.5:1; link on bg must reach 4.5:1. If either fails, `onboard.py` refuses to save (exit code 4) and tells the user to pick a darker primary, blank `brand.bg`/`brand.text` to let derivation pick a safe pair, or override `brand.text` directly. Canon: WCAG 2.2 §1.4.3.
2. **Output directory must be writable.** `onboard.py` walks up the path to find an existing ancestor and checks `os.W_OK`. Empty or unwritable path → exit code 3. The orchestrator's `output_path_resolver.py` honors the same rule per-conversion.
3. **Customization must change behavior, not sit as decoration.** Every consumer (md-document, md-review, md-slides) must read the config and render differently when the user changes `design_style`, `brand.primary`, `code_theme`, or `toc.behavior`. Decorative-only fields fail the design discipline.
4. **Precedence is fixed.** Project > global > defaults. The deep-merge preserves nested keys (e.g. you can override `brand.primary` in a project config without losing `typography.heading_font` from global).
5. **Bypass env exists for a reason.** `MARKDOWN_HTML_NO_CONFIG=1` is for headless CI, ephemeral test containers, and the autoresearch-style evaluator loops. Never set it silently for an interactive user.
## Derived 12-token palette
Once the user's brand is captured, `brand_palette_validator.derive_palette()` produces 12 CSS custom properties stored under `derived_palette` in the same config file. Every converter inlines these into its `<style>` block.
| Token | Purpose | Derivation |
|---|---|---|
| `--md-bg` | Document background | Primary if dark, near-neutral if vibrant |
| `--md-surface` | Card / callout / blockquote background | Bg ± 4-6% luminance |
| `--md-border` | Hairline dividers, table borders | Bg ± 8-12% luminance |
| `--md-text` | Body text | Off-white on dark bg, near-black on light bg |
| `--md-text-muted` | Captions, metadata, footers | `rgba(text, 0.68)` |
| `--md-accent` | Primary CTA, callout headers, link emphasis | Primary if vibrant, hue-shifted lighter if dark |
| `--md-accent-soft` | Accent backgrounds, hover states | `rgba(accent, 0.14)` |
| `--md-code-bg` | Inline code, fenced block bg | Bg ± 4-5% luminance |
| `--md-link` | Hyperlinks | Iteratively walked to reach 4.5:1 contrast on bg |
| `--md-link-hover` | Hover state | Link ± 6-8% luminance |
| `--md-success` | OK / approved / passed | Green anchored, luminance-matched |
| `--md-warn` | Caution / nit / TODO | Amber anchored, luminance-matched |
## Forcing-question library (Matt Pocock grill-with-docs pattern)
One question per turn, recommended answer, canon citation.
1. **What's your brand primary color?** Recommended: a HEX you already use in your product or docs — not a stock blue. Canon: Aarron Walter, *Designing for Emotion* (color carries brand affect).
2. **Should accent be derived or set?** Recommended: derive on first run (hue-shift + lighten produces a coherent companion); set explicitly only if your brand kit specifies one. Canon: Adobe Spectrum, *Color Foundations*.
3. **Editorial, technical, minimal, or playful?** Recommended: `technical` for engineering specs/reports, `editorial` for long-read narratives, `minimal` for sparse reference docs, `playful` for marketing/landing content. Canon: Ellen Lupton, *Thinking with Type* (style serves the rhetorical purpose).
4. **Sticky-sidebar TOC, or inline?** Recommended: `sticky-sidebar` for documents over 800 words, `inline` for short reads. Canon: Nielsen-Norman, *Table of Contents Best Practices* (2023).
5. **Save to global or per-project?** Recommended: global by default (consistent across your work); use `--scope project` only when this repo has a different brand. Canon: research-ops onboarding pattern, `research-ops/CLAUDE.md` §8.
## Customization in use (worked example)
```bash
# First-run onboarding (interactive, walks all 10 questions)
python3 markdown-html/skills/design-system/scripts/onboard.py
# Zero-touch defaults for CI / first-test
python3 .../onboard.py --defaults
# Change just the primary color and design style
python3 .../onboard.py --set brand.primary=#FF6B35 --set design_style=editorial
# Per-repo override
python3 .../onboard.py --scope project --set design_style=minimal
# Reset and re-onboard
python3 .../onboard.py --reset
python3 .../onboard.py
# Inspect the effective config (project > global > defaults)
python3 .../config_loader.py --show
python3 .../config_loader.py --status
# Bypass saved config (returns DEFAULTS only)
MARKDOWN_HTML_NO_CONFIG=1 python3 .../config_loader.py --show
# Spot-check WCAG contrast before committing to a brand
python3 .../brand_palette_validator.py --primary "#FF6B35" --accent "#00D4AA"
```
## Assumptions
1. User has at least one brand HEX they want consistent across their HTML conversions.
2. User accepts a 1-2 minute one-time setup.
3. User is OK with Google Fonts as the typography source (CDN, no local font hosting).
4. WCAG 2.2 AA is the accessibility floor (4.5:1 body, 3:1 large/UI). AAA (7:1) is out of scope.
## Non-goals
- Not a full design-token system (Style Dictionary, Theo). Twelve tokens, not a hundred.
- Not a custom-font hosting solution. Google Fonts only.
- Not a dark/light mode switcher in the converters. `code_theme: auto` handles the prefers-color-scheme case for syntax highlighting; layout palette is single-mode per onboarding.
- Not an accessibility audit suite (use axe-core / pa11y for that). We enforce contrast only.
- Does not transform existing CSS — the derived palette is injected into freshly generated HTML.
## Distinct from
- **`marketing/landing/skills/landing/scripts/brand_palette_validator.py`** — that script's `derive_palette()` produces 8 tokens shaped for hero-page rendering (`--navy`, `--teal`, `--card-bg`, `--card-border`). This script produces 12 tokens shaped for document rendering (sticky surface, hairline border, code bg, link, link-hover, success, warn). Same WCAG + HSL math, different token taxonomy.
- **`research-ops/skills/clinical-research/scripts/onboard.py`** — same pattern (interactive + `--defaults`/`--set`/`--show`/`--reset`/`--scope`), different question set (clinical alpha/power/dropout vs. brand palette/typography/layout).
## Output artifact
`~/.config/markdown-html/design-system.json` (global) or `./.markdown-html/design-system.json` (project). JSON schema lives at `assets/design_system_schema.json`.
## Anti-patterns (do not)
- ❌ Skip onboarding and run a converter with placeholder defaults — output looks unbranded.
- ❌ Pick a vibrant brand primary as `brand.bg` directly (low text contrast). Use it as accent instead.
- ❌ Set `MARKDOWN_HTML_NO_CONFIG=1` silently for an interactive user — they'll wonder why their tokens disappeared.
- ❌ Encode brand semantics in `derived_palette` outside the 12-token taxonomy. Add a new token only with a deliberate name + purpose + derivation rule.
## References
- WCAG 2.2 — §1.4.3 (contrast), §1.4.4 (resize), §1.4.11 (non-text contrast)
- Aarron Walter — *Designing for Emotion* (A Book Apart)
- Ellen Lupton — *Thinking with Type*
- Adobe Spectrum — *Color Foundations*
- Nielsen-Norman — *Table of Contents Best Practices* (2023)
- research-ops onboarding pattern: `research-ops/CLAUDE.md` §8
- Brand palette math source: `marketing/landing/skills/landing/scripts/brand_palette_validator.py`
FILE:assets/design_system_schema.json
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://github.com/alirezarezvani/claude-skills/blob/main/markdown-html/skills/design-system/assets/design_system_schema.json",
"title": "markdown-html design-system customization config",
"description": "JSON schema for the design-system config written by onboard.py and consumed by every markdown-html converter via config_loader.py. Lives at ~/.config/markdown-html/design-system.json (global) or ./.markdown-html/design-system.json (project).",
"type": "object",
"required": ["version", "skill", "default_output_dir", "brand", "typography", "design_style", "code_theme", "toc"],
"properties": {
"version": {"type": "integer", "const": 1, "description": "Schema version. Bump on breaking changes to the layout."},
"skill": {"type": "string", "const": "design-system"},
"default_output_dir": {
"type": "string",
"minLength": 1,
"description": "Where converters save generated HTML by default. Must be a writable path. The orchestrator's output_path_resolver.py also accepts a --out override per conversion."
},
"brand": {
"type": "object",
"required": ["primary"],
"properties": {
"primary": {"type": "string", "pattern": "^#?[0-9a-fA-F]{6}$", "description": "Primary brand color, HEX."},
"accent": {"type": ["string", "null"], "pattern": "^#?[0-9a-fA-F]{6}$|^$", "description": "Optional accent color. If null/empty, brand_palette_validator derives it from the primary via hue-shift + lighten."},
"bg": {"type": ["string", "null"], "description": "Optional background override. If null, derived from primary."},
"text": {"type": ["string", "null"], "description": "Optional body text override. If null, derived (off-white on dark bg, near-black on light bg)."}
}
},
"typography": {
"type": "object",
"required": ["heading_font", "body_font"],
"properties": {
"heading_font": {"type": "string", "description": "Google Font family for headings. e.g., Inter, Source Serif 4, Playfair Display."},
"body_font": {"type": "string", "description": "Google Font family for body text."},
"scale_ratio": {"type": "number", "minimum": 1.0, "maximum": 2.0, "description": "Modular type-scale ratio. 1.25 = major third (default), 1.333 = perfect fourth, 1.5 = perfect fifth."}
}
},
"design_style": {
"type": "string",
"enum": ["editorial", "technical", "minimal", "playful"],
"description": "Layout density preset consumed by every converter. editorial = magazine-like with wide margins and pull-quotes; technical = docs-like with sticky TOC and code emphasis; minimal = sparse with maximum whitespace; playful = product-marketing with color blocks and varied scale."
},
"code_theme": {
"type": "string",
"enum": ["light", "dark", "auto"],
"description": "Prism.js theme selection. auto = follows prefers-color-scheme."
},
"toc": {
"type": "object",
"required": ["behavior"],
"properties": {
"behavior": {"type": "string", "enum": ["sticky-sidebar", "collapsible-top", "inline", "none"]},
"max_depth": {"type": "integer", "minimum": 1, "maximum": 6, "description": "Deepest heading level included in the TOC."}
}
},
"company_name": {"type": "string", "description": "Optional, shown in footer of every generated HTML."},
"logo_url": {"type": "string", "description": "Optional. Base64-embedded at render time by default; pass --logo-mode link to inline the URL instead."},
"derived_palette": {
"type": "object",
"description": "12 CSS custom properties derived from the brand input by brand_palette_validator.derive_palette(). Stored here so every converter has identical tokens without re-deriving. Keys are CSS variable names; values are HEX or rgba() strings.",
"properties": {
"--md-bg": {"type": "string"},
"--md-surface": {"type": "string"},
"--md-border": {"type": "string"},
"--md-text": {"type": "string"},
"--md-text-muted": {"type": "string"},
"--md-accent": {"type": "string"},
"--md-accent-soft": {"type": "string"},
"--md-code-bg": {"type": "string"},
"--md-link": {"type": "string"},
"--md-link-hover": {"type": "string"},
"--md-success": {"type": "string"},
"--md-warn": {"type": "string"}
}
},
"setup_completed_at": {
"type": ["string", "null"],
"format": "date-time",
"description": "ISO-8601 timestamp written by onboard.py on successful completion. The orchestrator refuses to convert if this is null."
}
}
}
FILE:references/design_token_canon.md
# Design Token Canon
**Why this exists:** This skill ships 12 CSS custom properties — small by design-system standards. This document explains why 12 is enough, the taxonomy the tokens follow, and the canon they derive from.
## The 12-token taxonomy
| Layer | Tokens | Purpose |
|---|---|---|
| **Surface** | `--md-bg`, `--md-surface`, `--md-border`, `--md-code-bg` | Vertical layering: page bg → cards/callouts → hairlines → fenced code |
| **Text** | `--md-text`, `--md-text-muted` | Body + secondary (captions, metadata) |
| **Accent** | `--md-accent`, `--md-accent-soft` | Brand emphasis (CTA, callout headers); soft for hover backgrounds |
| **Link** | `--md-link`, `--md-link-hover` | Hyperlink + hover state; iteratively contrast-walked |
| **Semantic** | `--md-success`, `--md-warn` | Inline status, callouts, review severity |
Twelve covers every visual decision a long-form document needs. More tokens (e.g. Material Design's hundreds) optimize for design systems that span many UIs; markdown-html spans one artifact type (a generated HTML file) so we don't need the extra.
## Sources
### 1. Salesforce Lightning Design System — *Tokens* (lightningdesignsystem.com)
First widely-adopted token system at scale. Established the layered taxonomy: surface → text → border → accent → semantic. Markdown-html's 12 tokens follow the same layering, scoped down to document-rendering needs.
### 2. Adobe Spectrum — *Color Foundations* (spectrum.adobe.com)
Documents the four roles a brand color plays: bg, accent, text, semantic. Validates the decision to derive accent from primary rather than treat them as independent (Spectrum: "accent should be a tinted, brightness-adjusted variant of the brand color").
### 3. Material Design 3 — *Color Roles* (m3.material.io)
Token taxonomy of `primary`/`onPrimary`/`primaryContainer`/`onPrimaryContainer` etc. We deliberately simplify: a long-form document doesn't need surface containers within accent containers. The 12-token system is the Material taxonomy collapsed to what document rendering actually requires.
### 4. Sara Soueidan — *Color Tokens for Accessible Color Systems* (sarasoueidan.com, 2022)
Argues for contrast-walked link colors: a link in brand accent often fails the 4.5:1 floor against bg; the system must lighten or darken until it passes. Our `_ensure_link_contrast()` is the direct implementation.
### 5. Style Dictionary (amzn.github.io/style-dictionary)
The industry-standard token transformation tool — takes JSON tokens and emits CSS / Swift / Kotlin / Flutter. We deliberately ship JSON tokens compatible with Style Dictionary in case a user wants to extend; we don't depend on it.
### 6. CSS Custom Properties (MDN)
The native browser primitive for runtime-themable styles. Inlining `:root { --md-bg: #...; }` into the generated `<style>` block means the user can override any token by adding their own `:root` override in a custom-CSS section of the document (escape hatch).
### 7. Material Design 2 — *Type Scale* and *Color System* (material.io archive)
Original 8-point grid + modular type scale + tonal palette. We use a smaller subset (just modular scale via `typography.scale_ratio`, default 1.25 = major third) and 12 tokens; same philosophy.
## Why not 8? Why not 50?
- **8 tokens** (the original landing-skill palette) — covers a landing page (hero bg, accent CTA, card bg, card border, off-white text, muted text, glow). Documents need link, link-hover, code-bg, success, and warn that landing doesn't.
- **50 tokens** (Material Design 3 / IBM Carbon) — covers a multi-surface UI with elevated containers, interactive states, focus rings, disabled states. A document is a single surface with text — most of those tokens never render.
Twelve is the smallest number that covers every visual decision a long-form document, code review, or slide deck must make, without inventing decisions the document doesn't have.
## Applied to markdown-html
Every converter inlines the user's `derived_palette` into a `:root { }` block at the top of `<style>`. Every other CSS rule references the variables — no hard-coded colors anywhere. This makes the converters honestly customizable: change `brand.primary` and re-onboard, all 12 tokens re-derive, and the document re-renders with a different brand without any code change.
FILE:references/typography_pairing.md
# Typography Pairing
**Why this exists:** The onboarding wizard offers 12 Google Fonts and asks the user to pick a heading + body pair. Most users don't have strong opinions on type. This document codifies the pairs that work without further thought, so the wizard can recommend confidently and the converters can render coherently.
## Safe pairs
| Pair | Use for | Reason |
|---|---|---|
| `Inter` + `Inter` | Technical docs, dashboards | Single family across heading/body — clean, neutral, OpenType-rich |
| `Inter` + `Source Sans 3` | Long-form reports | Sans-on-sans pairing; Source Sans is more readable at body size |
| `Source Serif 4` + `Source Sans 3` | Editorial / narrative | Adobe's Source family — designed as a coherent system |
| `Playfair Display` + `Lora` | Magazine-style | Serif heading with personality; serif body that pairs |
| `Merriweather` + `Open Sans` | Long-form reading | Editorial serif + neutral sans body; oldest-and-safest pair |
| `IBM Plex Sans` + `IBM Plex Sans` | Technical + brand | Plex is designed for documentation; coherent across weights |
| `JetBrains Mono` + (Inter or Source Sans 3) | Engineering notebooks | Mono headings signal a coding/terminal context |
## Sources
### 1. Ellen Lupton — *Thinking with Type* (Princeton Architectural Press, 2010)
Foundational. The "stress, weight, and contrast" framework for pairing: the heading and body should share at least one of {stress angle, x-height, terminal style} and contrast in at least one of {weight, scale}. Every recommended pair above satisfies this.
### 2. Tim Brown — *Combining Typefaces* (Five Simple Steps, 2013)
The "concord / contrast / conflict" framework. Concord (same family) is always safe — hence the Inter+Inter and IBM Plex Sans+IBM Plex Sans pairs. Contrast is rewarding when done with intent (Playfair + Lora). Conflict is what users should avoid; the wizard's curated list rules out conflict pairs.
### 3. Erik Spiekermann — *Stop Stealing Sheep & Find Out How Type Works* (Adobe Press, 2013, 3rd ed.)
Argues that body type carries 95% of the visual weight in a document. The wizard prioritizes body font choice over heading font choice in the recommendation framing.
### 4. Google Fonts — *Pairings* and *Featured Pairs* (fonts.google.com)
The 12 fonts in `SAFE_FONTS` are pulled from Google Fonts' own curated catalog, biased toward families with multiple weights and broad language coverage. All available under the SIL Open Font License — no licensing concerns.
### 5. IBM Design Language — *Plex Family Documentation* (ibm.com/design/language/typography/type-basics)
Documents the "designed as a system" pattern: Plex Sans, Serif, Mono share metrics and x-height, so any combination renders coherently. We surface Plex Sans for users who want IBM-style technical documents.
### 6. Adobe Fonts — *Source Sans, Source Serif, Source Code* (fonts.adobe.com/foundries/adobe-originals)
Same "designed as a system" idea: Source family was created by Adobe to be a coherent triple. We surface Source Sans 3 and Source Serif 4 (the current versions, with extended Cyrillic and Vietnamese coverage).
### 7. Marcin Wichary — *The Hardest Working Font in Manhattan* (figma.com/blog, 2023)
A case study on choosing Inter for the Figma marketing site. Reinforces Inter as a reasonable default for technical-yet-broad audiences.
## What about display fonts, script fonts, decorative fonts?
Excluded from the wizard's options. Decorative fonts work for the first 200 words and exhaust the reader thereafter — they're a marketing-page choice, not a document choice. If the user wants a decorative heading, they can set `typography.heading_font` to any Google Font name manually after onboarding (the field accepts any string).
## What about variable fonts?
Inter, Roboto, Source Sans 3, Source Serif 4, IBM Plex Sans, and JetBrains Mono are all available as variable fonts on Google Fonts. The converters use the `wght@400;600` slice by default — sufficient for body + bold heading — to keep CDN payload small. Users who want a wider weight range can override the Google Fonts URL directly in the generated HTML.
## Type scale
`typography.scale_ratio` (default 1.25 = major third) drives a modular scale: body = 1rem, h6 = 1rem × 1.25, h5 = 1rem × 1.25², etc. Defaults:
| Ratio | Name | Effect |
|---|---|---|
| 1.125 | Major second | Tight; good for dense reference docs |
| 1.2 | Minor third | Standard for technical writing |
| **1.25** | **Major third** | Default; balanced for long-form reading |
| 1.333 | Perfect fourth | Editorial; pronounced hierarchy |
| 1.5 | Perfect fifth | Magazine-style with bold headings |
Each converter applies the scale based on this single ratio — no per-level overrides.
## Applied to markdown-html
The converters emit a `<link>` to Google Fonts at document head and apply the typography choice via CSS:
```css
:root {
--md-font-heading: 'Source Serif 4', Georgia, serif;
--md-font-body: 'Source Sans 3', system-ui, sans-serif;
--md-scale: 1.25;
}
body { font-family: var(--md-font-body); }
h1, h2, h3, h4, h5, h6 { font-family: var(--md-font-heading); }
```
The system fallback in each `font-family` declaration means the document still reads well if Google Fonts is blocked.
FILE:references/wcag_accessibility.md
# WCAG Accessibility Floor
**Why this exists:** Every converter renders text on backgrounds, links on backgrounds, and accent UI on backgrounds. WCAG 2.2 sets minimum contrast ratios that, if violated, make the document unreadable for users with low vision. This skill enforces those ratios as hard refusals during onboarding — not as warnings — because no user expects an onboarding wizard to ship them an inaccessible default.
## The floor
WCAG 2.2 AA Level (Section 1.4.3):
| Foreground / Background | Minimum contrast |
|---|---|
| Body text (< 18pt regular or < 14pt bold) | **4.5 : 1** |
| Large text (≥ 18pt regular or ≥ 14pt bold) | 3 : 1 |
| Non-text UI (focus rings, button borders, icons) | 3 : 1 |
| Links (treated as body text) | **4.5 : 1** |
`brand_palette_validator.py` enforces all four during onboarding. Failures on body-text or link contrast → refuse (exit code 4). Failures on non-text UI → warn but proceed (the user might be using accent for a backdrop that doesn't carry semantic meaning).
## Sources
### 1. WCAG 2.2 — *Understanding Success Criterion 1.4.3: Contrast (Minimum)* (w3.org/WAI/WCAG22)
The text and the formula. We implement `relative_luminance()` per the spec's sRGB-linearization rule and `contrast_ratio()` per `(L1 + 0.05) / (L2 + 0.05)`. No deviation.
### 2. WCAG 2.2 — *Understanding Success Criterion 1.4.11: Non-text Contrast* (w3.org/WAI/WCAG22)
Establishes the 3:1 floor for UI components. Used for `wcag-accent-on-bg` check.
### 3. WCAG 2.2 — *Understanding Success Criterion 1.4.4: Resize Text* (w3.org/WAI/WCAG22)
Mandates that text can be resized to 200% without loss of content. The converters use `rem` units for type scale (driven by `typography.scale_ratio`) so browser zoom respects user preference.
### 4. WebAIM — *Contrast Checker* (webaim.org/resources/contrastchecker)
The de-facto reference implementation. Cross-checked against our `contrast_ratio()` — identical results to 2 decimal places.
### 5. Sara Soueidan — *Color Tokens for Accessible Color Systems* (sarasoueidan.com, 2022)
Articulates the iterative-contrast-walk strategy: when a brand color fails on the link role, lighten or darken until it passes, then snap. The `_ensure_link_contrast()` helper is the direct implementation.
### 6. Léonie Watson — *Accessibility is a Process* (talks across 2018-2024)
Reinforces that contrast is the lowest-cost-highest-impact accessibility win. Most other a11y improvements take design effort; contrast can be enforced algorithmically.
### 7. CSS `prefers-color-scheme` (MDN)
The browser primitive for dark/light mode detection. `code_theme: "auto"` in the design-system config maps to a CSS media query, so syntax-highlighting follows OS preference automatically without forcing a re-onboard.
## What this skill does NOT enforce
- **WCAG 2.2 AAA (7:1)** — out of scope. AA is the realistic floor for design systems shipping to broad audiences; AAA is reserved for medical/legal/government content.
- **Focus order, ARIA, keyboard nav** — out of scope here (the converters handle these in their own renderers). md-review enforces `aria-label` on severity badges, md-slides enforces keyboard nav per WCAG 2.1.1, md-document enforces `aria-current="location"` on TOC scrollspy.
- **Reduced motion** — out of scope here. Converters emit `@media (prefers-reduced-motion: reduce) { * { animation: none; } }` independently.
- **Screen-reader semantic correctness** — out of scope. Beyond ensuring `<h1>...<h6>` hierarchy is preserved and `<table>` has `<thead>`, deeper SR audit needs a tool like pa11y / axe-core.
## Why hard refusal, not warning
A warning that ships an inaccessible default is the worst outcome of an onboarding wizard. The user trusted the wizard to set them up right. WCAG AA on body text is the one thing we can verify deterministically — so we do.
If the user genuinely wants to override (rare: a brand-mandated low-contrast scheme for a graphic design portfolio, say), they can:
1. Set `MARKDOWN_HTML_NO_CONFIG=1` and run with built-in defaults
2. Manually edit `~/.config/markdown-html/design-system.json` (the saved file)
3. Add a `<style>` override block in the converted HTML directly
These are all explicit, deliberate acts. The wizard's job is to ship an accessible default; the user can break that contract knowingly.
FILE:scripts/brand_palette_validator.py
#!/usr/bin/env python3
"""brand_palette_validator.py - Validate brand HEX colors + derive 12-token palette.
Stdlib-only. Validates the brand primary + optional accent/bg/text the user supplies
during onboarding, then derives the full 12-CSS-custom-property palette consumed by
every markdown-html converter (md-document, md-review, md-slides).
Pipeline:
1. Parse + verify each HEX is well-formed
2. WCAG 2.2 contrast checks (text-on-bg, accent-on-bg, link-on-bg)
3. Derive missing tokens algorithmically (lighten/darken in HSL, hue-shift for accent)
4. Emit the 12-token palette as a JSON dict ready to inject into onboard.py config
Forked from marketing/landing/skills/landing/scripts/brand_palette_validator.py
(WCAG math + HSL color manipulation + derive_palette shape) and adapted: 12 tokens
instead of 8, document-reading focus (longer reading sessions → tighter contrast
floors), no "card" semantics, dedicated --md-link / --md-link-hover / --md-success
/ --md-warn / --md-code-bg tokens for document/review/slides use cases.
NO LLM CALLS. Pure color-math + WCAG formula.
Usage:
python brand_palette_validator.py --primary "#0A1628" --accent "#00D4AA" --output json
python brand_palette_validator.py --primary "#FF6B35" --output human
python brand_palette_validator.py --sample
"""
from __future__ import annotations
import argparse
import colorsys
import json
import re
import sys
from typing import Any
HEX_RE = re.compile(r"^#?([0-9a-fA-F]{6})$")
def parse_hex(hex_str: str) -> tuple[int, int, int]:
m = HEX_RE.match(hex_str.strip())
if not m:
raise ValueError(f"Invalid HEX '{hex_str}'. Expected #RRGGBB or RRGGBB (6 hex chars).")
h = m.group(1)
return (int(h[0:2], 16), int(h[2:4], 16), int(h[4:6], 16))
def rgb_to_hex(rgb: tuple[int, int, int]) -> str:
return "#{:02X}{:02X}{:02X}".format(*rgb)
def relative_luminance(rgb: tuple[int, int, int]) -> float:
"""Per WCAG 2.2 — sRGB-linearized luminance."""
def linearize(channel: int) -> float:
c = channel / 255.0
return c / 12.92 if c <= 0.03928 else ((c + 0.055) / 1.055) ** 2.4
r, g, b = rgb
return 0.2126 * linearize(r) + 0.7152 * linearize(g) + 0.0722 * linearize(b)
def contrast_ratio(rgb1: tuple[int, int, int], rgb2: tuple[int, int, int]) -> float:
l1 = relative_luminance(rgb1)
l2 = relative_luminance(rgb2)
lighter, darker = max(l1, l2), min(l1, l2)
return (lighter + 0.05) / (darker + 0.05)
def lighten_hsl(rgb: tuple[int, int, int], pct: float) -> tuple[int, int, int]:
r, g, b = (c / 255.0 for c in rgb)
h, l, s = colorsys.rgb_to_hls(r, g, b)
l = min(1.0, max(0.0, l + pct))
r2, g2, b2 = colorsys.hls_to_rgb(h, l, s)
return (int(r2 * 255), int(g2 * 255), int(b2 * 255))
def darken_hsl(rgb: tuple[int, int, int], pct: float) -> tuple[int, int, int]:
return lighten_hsl(rgb, -pct)
def shift_hue(rgb: tuple[int, int, int], degrees: float) -> tuple[int, int, int]:
r, g, b = (c / 255.0 for c in rgb)
h, l, s = colorsys.rgb_to_hls(r, g, b)
h = (h + degrees / 360.0) % 1.0
r2, g2, b2 = colorsys.hls_to_rgb(h, l, s)
return (int(r2 * 255), int(g2 * 255), int(b2 * 255))
def rgba_str(rgb: tuple[int, int, int], alpha: float) -> str:
return f"rgba({rgb[0]}, {rgb[1]}, {rgb[2]}, {alpha})"
def is_dark(rgb: tuple[int, int, int]) -> bool:
return relative_luminance(rgb) < 0.18
def _ensure_link_contrast(
link: tuple[int, int, int],
bg: tuple[int, int, int],
target: float = 4.5,
) -> tuple[int, int, int]:
"""Iteratively adjust link luminance toward the target contrast on bg.
Documents have long reading sessions and lots of links — the WCAG AA
4.5:1 floor matters. Walk the luminance up or down (depending on which
direction increases contrast) until we hit the target or saturate.
"""
bg_lum = relative_luminance(bg)
# If bg is dark we lighten the link; if bg is light we darken it.
step = 0.04 if bg_lum < 0.5 else -0.04
result = link
for _ in range(20):
if contrast_ratio(result, bg) >= target:
return result
nxt = lighten_hsl(result, step)
if nxt == result:
break
result = nxt
return result
def derive_palette(
primary: tuple[int, int, int],
accent: tuple[int, int, int] | None = None,
bg: tuple[int, int, int] | None = None,
text: tuple[int, int, int] | None = None,
) -> dict[str, str]:
"""Derive the 12-token --md-* palette from a partial input.
Interpretation rule: `primary` is the user's *brand-identity* color
(CTA / accent / link emphasis), not necessarily the background. Three
branches based on primary luminance:
1. **Dark primary** (luminance < 0.18, e.g. navy #0A1628): assume the
user wants a dark-themed document — bg = primary, text = off-white,
accent = a hue-shifted lighter derivative.
2. **Light/vibrant primary** (luminance ≥ 0.18, e.g. orange #FF6B35):
use a near-neutral document bg (#FAFAFA with a hint of primary hue
for warmth), text = near-black, accent = primary itself.
3. **Explicit overrides** (bg, text supplied by user) win unconditionally.
Link contrast on bg is then iteratively enforced to WCAG AA 4.5:1 by
walking link luminance toward the target. Documents have long reading
sessions and lots of links — the floor matters.
"""
# Resolve bg first (it anchors every other contrast decision)
if bg is None:
if is_dark(primary):
bg = primary
else:
# Near-neutral light document bg with a faint warmth from primary's hue
r, g, b = (c / 255.0 for c in primary)
h, _, _ = colorsys.rgb_to_hls(r, g, b)
r2, g2, b2 = colorsys.hls_to_rgb(h, 0.97, 0.04)
bg = (int(r2 * 255), int(g2 * 255), int(b2 * 255))
if text is None:
text = (247, 247, 242) if is_dark(bg) else (16, 24, 32)
if accent is None:
if is_dark(primary):
accent = lighten_hsl(shift_hue(primary, 160), 0.45)
else:
accent = primary
surface = lighten_hsl(bg, 0.06 if is_dark(bg) else -0.03)
border = lighten_hsl(bg, 0.12 if is_dark(bg) else -0.08)
text_muted = rgba_str(text, 0.68)
accent_soft = rgba_str(accent, 0.14)
code_bg = lighten_hsl(bg, 0.04 if is_dark(bg) else -0.04)
link = _ensure_link_contrast(accent, bg, target=4.5)
link_hover = lighten_hsl(link, 0.08 if is_dark(bg) else -0.06)
# Success/warn derived from fixed hue anchors (green-ish / amber-ish), then
# luminance-matched to bg so they remain readable as inline labels.
green = (16, 168, 92)
amber = (200, 124, 16)
success = green if is_dark(bg) else darken_hsl(green, 0.08)
warn = amber if is_dark(bg) else darken_hsl(amber, 0.04)
return {
"--md-bg": rgb_to_hex(bg),
"--md-surface": rgb_to_hex(surface),
"--md-border": rgb_to_hex(border),
"--md-text": rgb_to_hex(text),
"--md-text-muted": text_muted,
"--md-accent": rgb_to_hex(accent),
"--md-accent-soft": accent_soft,
"--md-code-bg": rgb_to_hex(code_bg),
"--md-link": rgb_to_hex(link),
"--md-link-hover": rgb_to_hex(link_hover),
"--md-success": rgb_to_hex(success),
"--md-warn": rgb_to_hex(warn),
}
def validate(
primary: str,
accent: str | None = None,
bg: str | None = None,
text: str | None = None,
) -> dict[str, Any]:
findings: list[dict[str, str]] = []
def add(rule: str, level: str, message: str) -> None:
findings.append({"rule": rule, "level": level, "message": message})
try:
primary_rgb = parse_hex(primary)
add("primary-hex", "PASS", f"Primary parsed: {primary} = RGB{primary_rgb}")
except ValueError as e:
add("primary-hex", "FAIL", str(e))
return finalize(findings, {})
accent_rgb: tuple[int, int, int] | None = None
if accent:
try:
accent_rgb = parse_hex(accent)
add("accent-hex", "PASS", f"Accent parsed: {accent} = RGB{accent_rgb}")
except ValueError as e:
add("accent-hex", "FAIL", str(e))
return finalize(findings, {})
bg_rgb: tuple[int, int, int] | None = None
if bg:
try:
bg_rgb = parse_hex(bg)
add("bg-hex", "PASS", f"Bg parsed: {bg} = RGB{bg_rgb}")
except ValueError as e:
add("bg-hex", "FAIL", str(e))
return finalize(findings, {})
text_rgb: tuple[int, int, int] | None = None
if text:
try:
text_rgb = parse_hex(text)
add("text-hex", "PASS", f"Text parsed: {text} = RGB{text_rgb}")
except ValueError as e:
add("text-hex", "FAIL", str(e))
return finalize(findings, {})
palette = derive_palette(primary_rgb, accent_rgb, bg_rgb, text_rgb)
bg_final = parse_hex(palette["--md-bg"])
text_final = parse_hex(palette["--md-text"])
accent_final = parse_hex(palette["--md-accent"])
link_final = parse_hex(palette["--md-link"])
text_on_bg = contrast_ratio(text_final, bg_final)
accent_on_bg = contrast_ratio(accent_final, bg_final)
link_on_bg = contrast_ratio(link_final, bg_final)
add(
"wcag-text-on-bg",
"PASS" if text_on_bg >= 4.5 else ("WARN" if text_on_bg >= 3.0 else "FAIL"),
f"Body text on bg contrast: {text_on_bg:.2f}:1 (need 4.5:1 for body, WCAG AA)",
)
add(
"wcag-accent-on-bg",
"PASS" if accent_on_bg >= 3.0 else "WARN",
f"Accent (UI/CTA) on bg contrast: {accent_on_bg:.2f}:1 (need 3:1 for non-text UI)",
)
add(
"wcag-link-on-bg",
"PASS" if link_on_bg >= 4.5 else ("WARN" if link_on_bg >= 3.0 else "FAIL"),
f"Link on bg contrast: {link_on_bg:.2f}:1 (need 4.5:1, links are body-text-equivalent)",
)
return finalize(findings, palette)
def finalize(findings: list[dict[str, str]], palette: dict[str, str]) -> dict[str, Any]:
counts = {"PASS": 0, "WARN": 0, "FAIL": 0}
for f in findings:
counts[f["level"]] = counts.get(f["level"], 0) + 1
if counts["FAIL"] > 0:
verdict = "FAIL"
elif counts["WARN"] > 0:
verdict = "WARN"
else:
verdict = "PASS"
return {"verdict": verdict, "counts": counts, "findings": findings, "derived_palette": palette}
def render_human(result: dict[str, Any]) -> str:
out: list[str] = []
out.append(f"Brand palette validation verdict: {result['verdict']}")
c = result["counts"]
out.append(f" PASS: {c['PASS']} WARN: {c['WARN']} FAIL: {c['FAIL']}")
out.append("")
out.append("Findings:")
for f in result["findings"]:
marker = {"PASS": "[ok]", "WARN": "[warn]", "FAIL": "[FAIL]"}[f["level"]]
out.append(f" {marker} {f['rule']}: {f['message']}")
if result["derived_palette"]:
out.append("")
out.append("Derived 12-token palette (use in :root CSS):")
for k, v in result["derived_palette"].items():
out.append(f" {k:<20s} {v}")
return "\n".join(out)
def main(argv: list[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--primary", help="Primary HEX color (e.g., #0A1628)")
parser.add_argument("--accent", help="Accent HEX color (optional; derived if missing)")
parser.add_argument("--bg", help="Background HEX color (optional; derived if missing)")
parser.add_argument("--text", help="Text HEX color (optional; derived if missing)")
parser.add_argument("--sample", action="store_true", help="Validate built-in sample palette")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
result = validate("#0A1628", "#00D4AA")
elif args.primary:
result = validate(args.primary, args.accent, args.bg, args.text)
else:
parser.print_help()
return 0
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0 if result["verdict"] != "FAIL" else 1
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/config_loader.py
#!/usr/bin/env python3
"""config_loader.py - Customization loader for the markdown-html design-system skill.
Stdlib-only. Importable from every converter sub-skill (md-document, md-review,
md-slides) via `sys.path.insert(0, .../design-system/scripts)` so each renderer
picks up the user's onboarded brand tokens automatically.
Precedence (highest wins):
1. Project config: <cwd>/.markdown-html/design-system.json
2. Global config: ~/.config/markdown-html/design-system.json
3. Built-in DEFAULTS
Set MARKDOWN_HTML_NO_CONFIG=1 to ignore saved config (always returns DEFAULTS).
The onboarding answers (written by onboard.py) live in these files and are read
here so every converter renders with the user's tokens. Pattern lifted from
research-ops/skills/clinical-research/scripts/config_loader.py and adapted for
the markdown-html domain (brand palette + typography + layout + save location).
"""
from __future__ import annotations
import argparse
import json
import os
import sys
from pathlib import Path
from typing import Any
SKILL = "design-system"
DOMAIN = "markdown-html"
GLOBAL_CONFIG_DIR = Path.home() / ".config" / DOMAIN
GLOBAL_CONFIG_PATH = GLOBAL_CONFIG_DIR / f"{SKILL}.json"
PROJECT_CONFIG_DIRNAME = f".{DOMAIN}"
DEFAULTS: dict[str, Any] = {
"version": 1,
"skill": SKILL,
"default_output_dir": "./markdown-html-out/",
"brand": {
"primary": "#0A1628",
"accent": "#00D4AA",
"bg": None,
"text": None,
},
"typography": {
"heading_font": "Inter",
"body_font": "Inter",
"scale_ratio": 1.25,
},
"design_style": "technical",
"code_theme": "auto",
"toc": {
"behavior": "sticky-sidebar",
"max_depth": 3,
},
"company_name": "",
"logo_url": "",
"derived_palette": {},
"setup_completed_at": None,
}
def project_config_path(cwd: Path | None = None) -> Path:
cwd = cwd or Path.cwd()
return cwd / PROJECT_CONFIG_DIRNAME / f"{SKILL}.json"
def _read_json(path: Path) -> dict[str, Any] | None:
try:
with path.open(encoding="utf-8") as f:
data = json.load(f)
return data if isinstance(data, dict) else None
except (FileNotFoundError, json.JSONDecodeError, OSError):
return None
def _deep_merge(base: dict[str, Any], override: dict[str, Any]) -> dict[str, Any]:
out = dict(base)
for k, v in override.items():
if isinstance(v, dict) and isinstance(out.get(k), dict):
out[k] = _deep_merge(out[k], v)
else:
out[k] = v
return out
def load_config(cwd: Path | None = None) -> dict[str, Any]:
"""Effective config = DEFAULTS <- global <- project. Honors MARKDOWN_HTML_NO_CONFIG."""
config = dict(DEFAULTS)
if os.environ.get("MARKDOWN_HTML_NO_CONFIG") == "1":
return config
global_cfg = _read_json(GLOBAL_CONFIG_PATH)
if global_cfg:
config = _deep_merge(config, global_cfg)
project_cfg = _read_json(project_config_path(cwd))
if project_cfg:
config = _deep_merge(config, project_cfg)
return config
def setup_completed() -> bool:
cfg = _read_json(GLOBAL_CONFIG_PATH) or _read_json(project_config_path())
return bool(cfg and cfg.get("setup_completed_at"))
def write_config(config: dict[str, Any], scope: str = "global", cwd: Path | None = None) -> Path:
if scope == "project":
path = project_config_path(cwd)
else:
path = GLOBAL_CONFIG_PATH
path.parent.mkdir(parents=True, exist_ok=True)
with path.open("w", encoding="utf-8") as f:
json.dump(config, f, indent=2, sort_keys=True)
return path
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description=f"Inspect {DOMAIN}/{SKILL} customization config.")
p.add_argument("--show", action="store_true", help="Print the effective config")
p.add_argument("--status", action="store_true", help="Print setup status + paths")
p.add_argument("--sample", action="store_true", help="Print the built-in defaults")
args = p.parse_args(argv)
if args.sample:
print(json.dumps(DEFAULTS, indent=2, sort_keys=True))
elif args.status:
print(json.dumps({
"domain": DOMAIN,
"skill": SKILL,
"global_config_path": str(GLOBAL_CONFIG_PATH),
"global_config_exists": GLOBAL_CONFIG_PATH.exists(),
"project_config_path": str(project_config_path()),
"project_config_exists": project_config_path().exists(),
"setup_completed": setup_completed(),
"bypass_env_set": os.environ.get("MARKDOWN_HTML_NO_CONFIG") == "1",
}, indent=2))
else:
print(json.dumps(load_config(), indent=2, sort_keys=True))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/onboard.py
#!/usr/bin/env python3
"""onboard.py - First-run onboarding wizard for the markdown-html design-system.
Stdlib-only. Walks the user through 10 questions ONCE, validates the brand colors
against WCAG 2.2 AA, derives the 12 CSS custom properties, and writes the result
to a customization config that every markdown-html converter (md-document,
md-review, md-slides) reads via config_loader.py.
Modes:
--show print the questions + current effective config
--defaults write the built-in defaults without prompting
--set key=value ... set specific answers non-interactively (repeatable)
--reset delete the saved config at the chosen scope
--scope {global,project} where to save (default: global = ~/.config/markdown-html)
With no flags and an interactive terminal, walks the questions one at a time.
Refuses to complete onboarding if:
- default_output_dir is empty or unwritable (Q1 hard rule)
- the chosen brand colors fail WCAG AA contrast for body text on bg
Pattern lifted from research-ops/skills/clinical-research/scripts/onboard.py
(QUESTIONS table, _apply, run_interactive, main shape) and adapted for the
design-system surface (color validation via brand_palette_validator, palette
derivation persisted into the config alongside the raw user inputs).
"""
from __future__ import annotations
import argparse
import datetime as _dt
import os
import sys
from pathlib import Path
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import brand_palette_validator as bpv # noqa: E402
import config_loader as cfg # noqa: E402
DESIGN_STYLES = ["editorial", "technical", "minimal", "playful"]
CODE_THEMES = ["light", "dark", "auto"]
TOC_BEHAVIORS = ["sticky-sidebar", "collapsible-top", "inline", "none"]
SAFE_FONTS = [
"Inter", "Roboto", "Open Sans", "Lato", "Source Sans 3", "IBM Plex Sans",
"Merriweather", "Source Serif 4", "Lora", "Playfair Display",
"JetBrains Mono", "Fira Code",
]
# (key, prompt, choices_or_None, caster, hint)
QUESTIONS = [
("default_output_dir",
"1. Where should generated HTML files go? (path; must be writable)",
None, str, "e.g., ./markdown-html-out/ or ~/Documents/claude-html/"),
("brand.primary",
"2. Brand primary color (HEX)?",
None, str, "e.g., #0A1628 (dark navy) or #FF6B35 (orange)"),
("brand.accent",
"3. Brand accent color (HEX, optional — leave blank to derive)?",
None, str, "e.g., #00D4AA (teal) or leave blank for auto-derive"),
("typography.heading_font",
"4. Heading Google Font?",
SAFE_FONTS, str, "pick from the list or type your own"),
("typography.body_font",
"5. Body Google Font?",
SAFE_FONTS, str, "Inter/Roboto/Lato pair well as body fonts"),
("design_style",
"6. Design style?",
DESIGN_STYLES, str, "editorial = magazine-like; technical = docs-like; minimal = sparse; playful = product-marketing"),
("code_theme",
"7. Syntax-highlighting theme?",
CODE_THEMES, str, "auto = follows prefers-color-scheme"),
("toc.behavior",
"8. Table-of-contents behavior?",
TOC_BEHAVIORS, str, "sticky-sidebar = best for long docs; inline = best for slides"),
("company_name",
"9. Company / project name (optional, shows in footer)?",
None, str, "leave blank to omit"),
("logo_url",
"10. Logo URL (optional; base64-embedded at render time)?",
None, str, "leave blank to omit; URL or local path both work"),
]
def _apply(config: dict, key: str, value) -> None:
"""Apply a dotted key path into the nested config dict."""
if "." in key:
parts = key.split(".")
d = config
for part in parts[:-1]:
d = d.setdefault(part, {})
d[parts[-1]] = value
else:
config[key] = value
def _get(config: dict, key: str):
if "." in key:
parts = key.split(".")
d = config
for part in parts:
if not isinstance(d, dict):
return None
d = d.get(part)
return d
return config.get(key)
def _derive_and_check_palette(config: dict) -> tuple[bool, str]:
"""Run brand_palette_validator on the current colors and store the derived palette.
Returns (ok, message). If WCAG body-text contrast FAILs, ok=False.
"""
primary = _get(config, "brand.primary") or bpv.rgb_to_hex((10, 22, 40))
accent = _get(config, "brand.accent") or None
bg = _get(config, "brand.bg") or None
text = _get(config, "brand.text") or None
result = bpv.validate(primary, accent, bg, text)
config["derived_palette"] = result["derived_palette"]
if result["verdict"] == "FAIL":
msgs = [f for f in result["findings"] if f["level"] == "FAIL"]
return False, "; ".join(m["message"] for m in msgs)
if result["verdict"] == "WARN":
msgs = [f for f in result["findings"] if f["level"] == "WARN"]
return True, "warnings: " + "; ".join(m["message"] for m in msgs)
return True, "WCAG AA contrast met"
def _writable(path_str: str) -> bool:
if not path_str or not path_str.strip():
return False
p = Path(path_str).expanduser()
parent = p.parent if p.suffix else p
# If neither the path nor its parent exists, walk up until we find one
while not parent.exists():
if parent.parent == parent:
return False
parent = parent.parent
return os.access(parent, os.W_OK)
def _print_questions() -> None:
print(f"Onboarding questions — markdown-html/{cfg.SKILL}:\n")
for key, prompt, choices, _c, hint in QUESTIONS:
line = f" {prompt}"
if choices:
line += f"\n choices: {', '.join(choices[:6])}{'...' if len(choices) > 6 else ''}"
if hint:
line += f"\n hint: {hint}"
print(line)
print()
def run_interactive(config: dict) -> dict:
print(f"Onboarding — markdown-html/{cfg.SKILL}. Press Enter to keep the current/default.\n")
for key, prompt, choices, caster, hint in QUESTIONS:
current = _get(config, key)
suffix = ""
if choices:
suffix = f" [{ '/'.join(choices[:4]) }{'...' if len(choices) > 4 else ''}]"
cur = f" (current: {current})" if current not in (None, "") else ""
if hint:
print(f" hint: {hint}")
raw = input(f"{prompt}{suffix}{cur}: ").strip()
if not raw:
continue
try:
_apply(config, key, caster(raw))
except ValueError:
print(f" ! invalid value for {key}, keeping current")
print()
return config
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(
description="Onboarding for the markdown-html design-system skill."
)
p.add_argument("--show", action="store_true", help="print questions + effective config")
p.add_argument("--defaults", action="store_true", help="write built-in defaults, no prompt")
p.add_argument("--set", action="append", default=[], metavar="key=value",
help="set an answer non-interactively (repeatable; supports dotted keys like brand.primary=#FF6B35)")
p.add_argument("--reset", action="store_true", help="delete saved config at the scope")
p.add_argument("--scope", choices=["global", "project"], default="global")
args = p.parse_args(argv)
if args.show:
_print_questions()
print("Current effective config:")
import json
print(json.dumps(cfg.load_config(), indent=2, sort_keys=True))
return 0
if args.reset:
path = cfg.project_config_path() if args.scope == "project" else cfg.GLOBAL_CONFIG_PATH
if path.exists():
path.unlink()
print(f"removed {path}")
else:
print(f"no config at {path}")
return 0
config = cfg.load_config()
if args.set:
for item in args.set:
if "=" not in item:
print(f"error: --set expects key=value, got '{item}'", file=sys.stderr)
return 2
k, v = item.split("=", 1)
# numeric keys
if k == "typography.scale_ratio":
try:
v = float(v)
except ValueError:
pass
_apply(config, k, v)
elif not args.defaults:
if sys.stdin.isatty():
config = run_interactive(config)
else:
print("non-interactive shell: use --defaults or --set key=value. Showing questions:\n")
_print_questions()
return 0
# Hard rule 1: refuse if default_output_dir is empty or unwritable
out_dir = config.get("default_output_dir") or ""
if not _writable(out_dir):
print(
f"refusing to save: default_output_dir '{out_dir}' is empty or its parent "
f"is not writable. Pick a path you control (e.g., ./markdown-html-out/ or "
f"~/Documents/claude-html/) and re-run.",
file=sys.stderr,
)
return 3
# Hard rule 2: refuse if WCAG AA body-text contrast fails on the chosen colors
ok, msg = _derive_and_check_palette(config)
if not ok:
print(
f"refusing to save: WCAG AA contrast failed for the chosen colors — {msg}. "
f"Pick a darker primary (or a lighter text), or leave brand.bg/brand.text "
f"blank to let the validator derive a passing pair.",
file=sys.stderr,
)
return 4
if msg.startswith("warnings:"):
print(f"note: {msg} — proceeding (warnings, not failures).")
config["setup_completed_at"] = _dt.datetime.now(_dt.timezone.utc).isoformat()
path = cfg.write_config(config, scope=args.scope)
print(f"saved markdown-html/{cfg.SKILL} customization -> {path}")
print(f"derived 12-token palette stored under derived_palette in the same file.")
return 0
if __name__ == "__main__":
sys.exit(main())
Triển khai và duy trì hệ thống quản lý chất lượng ISO 13485 cho thiết bị y tế: thiết kế QMS, kiểm soát tài liệu, đánh giá nội bộ, CAPA và hỗ trợ chứng nhận.
---
name: "quality-manager-qms-iso13485"
description: ISO 13485 Quality Management System implementation and maintenance for medical device organizations. Provides QMS design, documentation control, internal auditing, CAPA management, and certification support. Use when working with medical device quality systems, preparing for ISO 13485 audits, managing regulatory compliance documentation, setting up corrective actions, or building audit preparation programs. Useful for quality management, audit preparation, regulatory compliance, medical device documentation, and corrective action workflows.
triggers:
- ISO 13485
- QMS implementation
- quality management system
- document control
- internal audit
- management review
- quality manual
- CAPA process
- process validation
- design control
- supplier qualification
- quality records
---
# Quality Manager - QMS ISO 13485 Specialist
ISO 13485:2016 Quality Management System implementation, maintenance, and certification support for medical device organizations.
---
## Table of Contents
- [QMS Implementation Workflow](#qms-implementation-workflow)
- [Document Control Workflow](#document-control-workflow)
- [Internal Audit Workflow](#internal-audit-workflow)
- [Process Validation Workflow](#process-validation-workflow)
- [Supplier Qualification Workflow](#supplier-qualification-workflow)
- [QMS Process Reference](#qms-process-reference)
- [Decision Frameworks](#decision-frameworks)
- [Tools and References](#tools-and-references)
---
## QMS Implementation Workflow
Implement ISO 13485:2016 compliant quality management system from gap analysis through certification.
### Workflow: Initial QMS Implementation
1. Conduct gap analysis against ISO 13485:2016 requirements
2. Document current state vs. required state for each clause
3. Prioritize gaps by:
- Regulatory criticality
- Risk to product safety
- Resource requirements
4. Develop implementation roadmap with milestones
5. Establish Quality Manual per Clause 4.2.2:
- QMS scope with justified exclusions
- Process interactions
- Procedure references
6. Create required documented procedures — see [Mandatory Documented Procedures](#quick-reference-mandatory-documented-procedures) for the full list
7. Deploy processes with training
8. **Validation:** Gap analysis complete; Quality Manual approved; all required procedures documented and trained
> Use the Gap Analysis Matrix template in [qms-process-templates.md](references/qms-process-templates.md) to document clause-by-clause current state, gaps, priority, and actions.
### QMS Structure
| Level | Document Type | Example |
|-------|---------------|---------|
| 1 | Quality Manual | QM-001 |
| 2 | Procedures | SOP-02-001 |
| 3 | Work Instructions | WI-06-012 |
| 4 | Records | Training records |
---
## Document Control Workflow
Establish and maintain document control per ISO 13485 Clause 4.2.3.
### Workflow: Document Creation and Approval
1. Identify need for new document or revision
2. Assign document number per numbering convention:
- Format: `[TYPE]-[AREA]-[SEQUENCE]-[REV]`
- Example: `SOP-02-001-01`
3. Draft document using approved template
4. Route for review to subject matter experts
5. Collect and address review comments
6. Obtain required approvals based on document type
7. Update Document Master List
8. **Validation:** Document numbered correctly; all reviewers signed; Master List updated
### Document Numbering Convention
| Prefix | Document Type | Approval Authority |
|--------|---------------|-------------------|
| QM | Quality Manual | Management Rep + CEO |
| POL | Policy | Department Head + QA |
| SOP | Procedure | Process Owner + QA |
| WI | Work Instruction | Supervisor + QA |
| TF | Template/Form | Process Owner |
| SPEC | Specification | Engineering + QA |
### Area Codes
| Code | Area | Examples |
|------|------|----------|
| 01 | Quality Management | Quality Manual, policy |
| 02 | Document Control | This procedure |
| 03 | Training | Competency procedures |
| 04 | Design | Design control |
| 05 | Purchasing | Supplier management |
| 06 | Production | Manufacturing |
| 07 | Quality Control | Inspection, testing |
| 08 | CAPA | Corrective actions |
### Document Change Control
| Change Type | Approval Level | Examples |
|-------------|----------------|----------|
| Administrative | Document Control | Typos, formatting |
| Minor | Process Owner + QA | Clarifications |
| Major | Full review cycle | Process changes |
| Emergency | Expedited + retrospective | Safety issues |
### Document Review Schedule
| Document Type | Review Period | Trigger for Unscheduled Review |
|---------------|---------------|-------------------------------|
| Quality Manual | Annual | Organizational change |
| Procedures | Annual | Audit finding, regulation change |
| Work Instructions | 2 years | Process change |
| Forms | 2 years | User feedback |
---
## Internal Audit Workflow
Plan and execute internal audits per ISO 13485 Clause 8.2.4.
### Workflow: Annual Audit Program
1. Identify processes and areas requiring audit coverage
2. Assess risk factors for audit frequency:
- Previous audit findings
- Regulatory changes
- Process changes
- Complaint trends
3. Assign qualified auditors (independent of area audited)
4. Develop annual audit schedule
5. Obtain management approval
6. Communicate schedule to process owners
7. Track completion and reschedule as needed
8. **Validation:** All processes covered; auditors qualified and independent; schedule approved
> Use the Audit Program Template in [qms-process-templates.md](references/qms-process-templates.md) to schedule audits by clause and quarter across processes such as Document Control (4.2.3/4.2.4), Management Review (5.6), Design Control (7.3), Production (7.5), and CAPA (8.5.2/8.5.3).
### Workflow: Individual Audit Execution
1. Prepare audit plan with scope, criteria, and schedule
2. Notify auditee minimum 1 week prior
3. Review procedures and previous audit results
4. Prepare audit checklist
5. Conduct opening meeting
6. Collect evidence through:
- Document review
- Record sampling
- Process observation
- Personnel interviews
7. Classify findings:
- Major NC: Absence or breakdown of system
- Minor NC: Single lapse or deviation
- Observation: Risk of future NC
8. Conduct closing meeting
9. Issue audit report within 5 business days
10. **Validation:** All checklist items addressed; findings supported by evidence; report distributed
### Auditor Qualification Requirements
| Criterion | Requirement |
|-----------|-------------|
| Training | ISO 13485 awareness + auditor training |
| Experience | Minimum 1 audit as observer |
| Independence | Not auditing own work area |
| Competence | Understanding of audited process |
### Finding Classification Guide
| Classification | Criteria | Response Time |
|----------------|----------|---------------|
| Major NC | System absence, total breakdown, regulatory violation | 30 days for CAPA |
| Minor NC | Single instance, partial compliance | 60 days for CAPA |
| Observation | Potential risk, improvement opportunity | Track in next audit |
---
## Process Validation Workflow
Validate special processes per ISO 13485 Clause 7.5.6.
### Workflow: Process Validation Protocol
1. Identify processes requiring validation:
- Output cannot be verified by inspection
- Deficiencies appear only in use
- Sterilization, welding, sealing, software
2. Form validation team with subject matter experts
3. Write validation protocol including:
- Process description and parameters
- Equipment and materials
- Acceptance criteria
- Statistical approach
4. Execute IQ: verify equipment installed correctly and document specifications
5. Execute OQ: test parameter ranges and verify process control
6. Execute PQ: run production conditions and verify output meets requirements
7. Write validation report with conclusions
8. **Validation:** IQ/OQ/PQ complete; acceptance criteria met; validation report approved
### Validation Documentation Requirements
| Phase | Content | Evidence |
|-------|---------|----------|
| Protocol | Objectives, methods, criteria | Approved protocol |
| IQ | Equipment verification | Installation records |
| OQ | Parameter verification | Test results |
| PQ | Performance verification | Production data |
| Report | Summary, conclusions | Approval signatures |
### Revalidation Triggers
| Trigger | Action Required |
|---------|-----------------|
| Equipment change | Assess impact, revalidate affected phases |
| Parameter change | OQ and PQ minimum |
| Material change | Assess impact, PQ minimum |
| Process failure | Full revalidation |
| Periodic | Per validation schedule (typically 3 years) |
### Special Process Examples
| Process | Validation Standard | Critical Parameters |
|---------|--------------------|--------------------|
| EO Sterilization | ISO 11135 | Temperature, humidity, EO concentration, time |
| Steam Sterilization | ISO 17665 | Temperature, pressure, time |
| Radiation Sterilization | ISO 11137 | Dose, dose uniformity |
| Sealing | Internal | Temperature, pressure, dwell time |
| Welding | ISO 11607 | Heat, pressure, speed |
---
## Supplier Qualification Workflow
Evaluate and approve suppliers per ISO 13485 Clause 7.4.
### Workflow: New Supplier Qualification
1. Identify supplier category:
- Category A: Critical (affects safety/performance)
- Category B: Major (affects quality)
- Category C: Minor (indirect impact)
2. Request supplier information:
- Quality certifications
- Product specifications
- Quality history
3. Evaluate supplier based on:
- Quality system (ISO certification)
- Technical capability
- Quality history
- Financial stability
4. For Category A suppliers:
- Conduct on-site audit
- Require quality agreement
5. Calculate qualification score
6. Make approval decision:
- >80: Approved
- 60-80: Conditional approval
- <60: Not approved
7. Add to Approved Supplier List
8. **Validation:** Evaluation criteria scored; qualification records complete; supplier categorized
### Supplier Evaluation Criteria
| Criterion | Weight | Scoring |
|-----------|--------|---------|
| Quality System | 30% | ISO 13485=30, ISO 9001=20, Documented=10, None=0 |
| Quality History | 25% | Reject rate: <1%=25, 1-3%=15, >3%=0 |
| Delivery | 20% | On-time: >95%=20, 90-95%=10, <90%=0 |
| Technical Capability | 15% | Exceeds=15, Meets=10, Marginal=5 |
| Financial Stability | 10% | Strong=10, Adequate=5, Questionable=0 |
### Supplier Category Requirements
| Category | Qualification | Monitoring | Agreement |
|----------|---------------|------------|-----------|
| A - Critical | On-site audit | Annual review | Quality agreement |
| B - Major | Questionnaire | Semi-annual review | Quality requirements |
| C - Minor | Assessment | Issue-based | Standard terms |
### Supplier Performance Metrics
| Metric | Target | Calculation |
|--------|--------|-------------|
| Accept Rate | >98% | (Accepted lots / Total lots) × 100 |
| On-Time Delivery | >95% | (On-time / Total orders) × 100 |
| Response Time | <5 days | Average days to resolve issues |
| Documentation | 100% | (Complete CoCs / Required CoCs) × 100 |
---
## QMS Process Reference
For detailed requirements and audit questions for each ISO 13485:2016 clause, see [iso13485-clause-requirements.md](references/iso13485-clause-requirements.md).
### Management Review Required Inputs (Clause 5.6.2)
| Input | Source | Prepared By |
|-------|--------|-------------|
| Audit results | Internal and external audits | QA Manager |
| Customer feedback | Complaints, surveys | Customer Quality |
| Process performance | Process metrics | Process Owners |
| Product conformity | Inspection data, NCs | QC Manager |
| CAPA status | CAPA system | CAPA Officer |
| Previous actions | Prior review records | QMR |
| Changes affecting QMS | Regulatory, organizational | RA Manager |
| Recommendations | All sources | All Managers |
### Record Retention Requirements
| Record Type | Minimum Retention | Regulatory Basis |
|-------------|-------------------|------------------|
| Device Master Record | Life of device + 2 years | 21 CFR 820.181 |
| Device History Record | Life of device + 2 years | 21 CFR 820.184 |
| Design History File | Life of device + 2 years | 21 CFR 820.30 |
| Complaint Records | Life of device + 2 years | 21 CFR 820.198 |
| Training Records | Employment + 3 years | Best practice |
| Audit Records | 7 years | Best practice |
| CAPA Records | 7 years | Best practice |
| Calibration Records | Equipment life + 2 years | Best practice |
---
## Decision Frameworks
### Exclusion Justification (Clause 4.2.2)
| Clause | Permissible Exclusion | Justification Required |
|--------|----------------------|------------------------|
| 6.4.2 | Contamination control | Product not affected by contamination |
| 7.3 | Design and development | Organization does not design products |
| 7.5.2 | Product cleanliness | No cleanliness requirements |
| 7.5.3 | Installation | No installation activities |
| 7.5.4 | Servicing | No servicing activities |
| 7.5.5 | Sterile products | No sterile products |
### Nonconformity Disposition Decision Tree
```
Nonconforming Product Identified
│
▼
Can it be reworked?
│
Yes──┴──No
│ │
▼ ▼
Is rework Can it be used
procedure as is?
available? │
│ Yes──┴──No
Yes─┴─No │ │
│ │ ▼ ▼
▼ ▼ Concession Scrap or
Rework Create approval return to
per SOP rework needed? supplier
procedure │
Yes─┴─No
│ │
▼ ▼
Customer Use as is
approval with MRB
approval
```
### CAPA Initiation Criteria
| Source | Automatic CAPA | Evaluate for CAPA |
|--------|----------------|-------------------|
| Customer complaint | Safety-related | All others |
| External audit | Major NC | Minor NC |
| Internal audit | Major NC | Repeat minor NC |
| Product NC | Field failure | Trend exceeds threshold |
| Process deviation | Safety impact | Repeated deviations |
---
## Tools and References
### Scripts
| Tool | Purpose | Usage |
|------|---------|-------|
| [qms_audit_checklist.py](scripts/qms_audit_checklist.py) | Generate audit checklists by clause or process | `python qms_audit_checklist.py --help` |
**Audit Checklist Generator Features:**
- Generate clause-specific checklists (e.g., `--clause 7.3`)
- Generate process-based checklists (e.g., `--process design-control`)
- Full system audit checklist (`--audit-type system`)
- Text or JSON output formats
- Interactive mode for guided selection
### References
| Document | Content |
|----------|---------|
| [iso13485-clause-requirements.md](references/iso13485-clause-requirements.md) | Detailed requirements for each ISO 13485:2016 clause with audit questions |
| [qms-process-templates.md](references/qms-process-templates.md) | Ready-to-use templates for gap analysis, audit program, document control, CAPA, supplier, training |
### Quick Reference: Mandatory Documented Procedures
| Procedure | Clause | Key Elements |
|-----------|--------|--------------|
| Document Control | 4.2.3 | Approval, distribution, obsolete control |
| Record Control | 4.2.4 | Identification, retention, disposal |
| Internal Audit | 8.2.4 | Program, auditor qualification, reporting |
| NC Product Control | 8.3 | Identification, segregation, disposition |
| Corrective Action | 8.5.2 | Root cause, implementation, verification |
| Preventive Action | 8.5.3 | Risk identification, implementation |
---
## Related Skills
| Skill | Integration Point |
|-------|-------------------|
| [quality-manager-qmr](../quality-manager-qmr/) | Management review, quality policy |
| [capa-officer](../capa-officer/) | CAPA system management |
| [qms-audit-expert](../qms-audit-expert/) | Advanced audit techniques |
| [quality-documentation-manager](../quality-documentation-manager/) | DHF, DMR, DHR management |
| [risk-management-specialist](../risk-management-specialist/) | ISO 14971 integration |
FILE:references/iso13485-clause-requirements.md
# ISO 13485:2016 Clause Requirements
Detailed requirements for each ISO 13485:2016 clause with implementation guidance and audit criteria.
---
## Table of Contents
- [Clause 4: Quality Management System](#clause-4-quality-management-system)
- [Clause 5: Management Responsibility](#clause-5-management-responsibility)
- [Clause 6: Resource Management](#clause-6-resource-management)
- [Clause 7: Product Realization](#clause-7-product-realization)
- [Clause 8: Measurement, Analysis and Improvement](#clause-8-measurement-analysis-and-improvement)
---
## Clause 4: Quality Management System
### 4.1 General Requirements
| Requirement | Implementation | Evidence |
|-------------|----------------|----------|
| Determine processes needed | Process map showing QMS processes | Documented process map |
| Determine sequence and interaction | Process interaction diagram | Cross-reference matrix |
| Determine criteria for operation | Process metrics and acceptance criteria | Documented criteria per process |
| Ensure resources available | Resource allocation per process | Training records, equipment logs |
| Monitor, measure, analyze | Process monitoring procedures | Trend data, performance reports |
| Implement actions for results | Improvement projects, CAPAs | Action records with verification |
| Document processes | Procedures, work instructions | Controlled document list |
**Audit Questions:**
- How are QMS processes identified and documented?
- What criteria determine if processes are operating effectively?
- How is outsourced process control demonstrated?
### 4.2 Documentation Requirements
#### 4.2.1 General
| Document Type | Requirement | Retention |
|---------------|-------------|-----------|
| Quality Policy | Documented statement of commitment | Life of QMS |
| Quality Objectives | Measurable objectives at relevant functions | Life of QMS |
| Quality Manual | QMS scope and processes | Current version |
| Documented Procedures | Required by standard | Life of QMS + 2 years |
| Records | Evidence of conformity | As defined per record type |
#### 4.2.2 Quality Manual
**Required Content:**
1. Scope of QMS including justification for exclusions
2. Documented procedures or reference to them
3. Description of process interactions
**Quality Manual Template Structure:**
```
QUALITY MANUAL
1. Company Overview
1.1 Company Description
1.2 Scope of QMS
1.3 Exclusions and Justification
2. Quality Policy
3. Quality Objectives
4. QMS Structure
4.1 Process Map
4.2 Process Interactions
4.3 Organizational Chart
5. Procedure References
5.1 Document Control
5.2 Record Control
5.3 Management Review
5.4 Internal Audit
5.5 Nonconformity Control
5.6 CAPA
6. Appendices
6.1 Glossary
6.2 Regulatory Cross-Reference
```
#### 4.2.3 Control of Documents
| Control Element | Requirement | Method |
|-----------------|-------------|--------|
| Approval | Adequate prior to issue | Signature/electronic approval |
| Review and update | Re-approval after changes | Periodic review process |
| Identification of changes | Change history visible | Revision log in document |
| Revision status | Current revision identifiable | Document master list |
| Legibility | Readable and identifiable | Format standards |
| External documents | Identified and controlled | Incoming document log |
| Obsolete documents | Prevented from unintended use | Archive system |
**Document Numbering Convention:**
```
[TYPE]-[AREA]-[SEQUENCE]-[REV]
TYPE:
QM = Quality Manual
SOP = Standard Operating Procedure
WI = Work Instruction
TF = Template/Form
POL = Policy
AREA:
01 = Quality Management
02 = Document Control
03 = Training
04 = Design
05 = Purchasing
06 = Production
07 = Quality Control
08 = CAPA
Example: SOP-02-001-03 = Document Control SOP, Revision 03
```
#### 4.2.4 Control of Records
| Record Category | Minimum Retention | Basis |
|-----------------|-------------------|-------|
| Device Master Record | Life of device + 2 years | 21 CFR 820.181 |
| Device History Record | Life of device + 2 years | 21 CFR 820.184 |
| Design History File | Life of device + 2 years | 21 CFR 820.30 |
| Training Records | Employment + 3 years | Best practice |
| Audit Records | 7 years | Best practice |
| Complaint Records | Life of device + 2 years | 21 CFR 820.198 |
| CAPA Records | 7 years | Best practice |
| Calibration Records | Equipment life + 2 years | Best practice |
| Supplier Records | Relationship + 3 years | Best practice |
---
## Clause 5: Management Responsibility
### 5.1 Management Commitment
| Commitment Area | Evidence Required |
|-----------------|-------------------|
| Communicate importance of requirements | Meeting minutes, communications |
| Establish quality policy | Documented policy, communication records |
| Ensure quality objectives established | Objective documentation |
| Conduct management reviews | Management review records |
| Ensure resources available | Budget records, staffing records |
### 5.2 Customer Focus
| Requirement | Implementation | Verification |
|-------------|----------------|--------------|
| Customer requirements determined | Requirements review process | Contract review records |
| Requirements met | Process controls | Inspection and test data |
| Regulatory requirements met | Regulatory register | Compliance assessments |
| Customer satisfaction enhanced | Feedback collection | Satisfaction data, complaints |
### 5.3 Quality Policy
**Policy Requirements:**
- Appropriate to organization purpose
- Commitment to compliance and effectiveness
- Framework for quality objectives
- Communicated and understood
- Reviewed for continuing suitability
**Sample Quality Policy Elements:**
```
[Company Name] Quality Policy
We are committed to:
- Designing and manufacturing safe, effective medical devices
- Meeting customer and regulatory requirements
- Maintaining an effective Quality Management System
- Continuously improving our processes and products
- Providing resources for QMS effectiveness
Signed: [Executive]
Date: [Date]
Review Date: [Annual]
```
### 5.4 Planning
#### 5.4.1 Quality Objectives
| Objective Criteria | Requirement |
|-------------------|-------------|
| Measurable | Quantifiable targets |
| Consistent with policy | Aligned to policy statements |
| Relevant functions | Cascaded to departments |
| Includes compliance | Regulatory and customer requirements |
| Includes product conformity | Product-related targets |
**Objective Template:**
```
QUALITY OBJECTIVE [Year]
Objective: [Statement]
Metric: [How measured]
Target: [Specific value]
Baseline: [Current performance]
Owner: [Responsible person]
Due Date: [Target date]
Reporting: [Frequency]
```
#### 5.4.2 Quality Management System Planning
**Planning Requirements:**
- QMS meets general requirements (4.1)
- QMS meets quality objectives (5.4.1)
- Integrity maintained during changes
### 5.5 Responsibility, Authority and Communication
#### 5.5.1 Responsibility and Authority
| Role | Responsibilities | Authority |
|------|-----------------|-----------|
| Top Management | QMS commitment, resources, policy | Budget, staffing, strategic decisions |
| Quality Manager | QMS implementation, reporting | Document approval, CAPA approval |
| Department Managers | Process ownership, resources | Process changes, training |
| Process Owners | Process performance, improvements | Procedure changes within scope |
#### 5.5.2 Management Representative
| QMR Responsibility | Activities |
|-------------------|------------|
| QMS establishment | Process definition, documentation |
| QMS implementation | Training, deployment, monitoring |
| QMS maintenance | Audits, reviews, improvements |
| Reporting to top management | Performance reports, recommendations |
| Awareness promotion | Training, communications |
#### 5.5.3 Internal Communication
| Communication Type | Method | Frequency |
|-------------------|--------|-----------|
| Policy and objectives | Posting, training | Annual and on change |
| QMS performance | Dashboards, reports | Monthly |
| Changes affecting quality | Email, meetings | As needed |
| Audit results | Reports, presentations | Per audit |
### 5.6 Management Review
#### 5.6.1 General
| Requirement | Specification |
|-------------|---------------|
| Frequency | Planned intervals (typically quarterly/semi-annually) |
| Purpose | Assess QMS suitability, adequacy, effectiveness |
| Records | Documented meeting records |
#### 5.6.2 Review Input
| Input | Source | Responsible |
|-------|--------|-------------|
| Audit results | Internal/external audits | QA Manager |
| Customer feedback | Complaints, surveys | Customer Quality |
| Process performance | Metrics, yields | Process Owners |
| Product conformity | Inspection data | QC Manager |
| CAPA status | CAPA system | CAPA Officer |
| Previous actions | Prior review records | QMR |
| Changes affecting QMS | Regulatory, organizational | RA, HR |
| Recommendations | All sources | All Managers |
#### 5.6.3 Review Output
| Output | Documentation |
|--------|---------------|
| QMS improvement decisions | Action items with owners |
| Process improvements | Project charters |
| Resource needs | Resource allocation plans |
| Product improvements | Design change requests |
---
## Clause 6: Resource Management
### 6.1 Provision of Resources
**Resource Categories:**
- Human resources (competent personnel)
- Infrastructure (facilities, equipment, software)
- Work environment (environmental conditions)
### 6.2 Human Resources
| Requirement | Implementation | Evidence |
|-------------|----------------|----------|
| Competence determined | Job descriptions, competency matrix | Role definitions |
| Training provided | Training programs | Training records |
| Effectiveness evaluated | Assessments, observations | Competency verification |
| Awareness ensured | Orientation, ongoing training | Acknowledgments |
| Records maintained | Training database | Training files |
**Competency Matrix Template:**
```
COMPETENCY MATRIX
Role: [Job Title]
Department: [Department]
Required Competencies:
| Competency | Requirement Level | Method | Verification |
|------------|------------------|--------|--------------|
| [Skill 1] | Expert/Proficient/Basic | Training/OJT | Assessment |
| [Skill 2] | Expert/Proficient/Basic | Training/OJT | Assessment |
Training Requirements:
| Training | Initial | Refresher | Record |
|----------|---------|-----------|--------|
| ISO 13485 Awareness | Yes | Annual | TR-001 |
| Document Control | Yes | On Change | TR-002 |
```
### 6.3 Infrastructure
| Infrastructure Type | Control Requirements |
|--------------------|---------------------|
| Buildings and workspace | Cleaning, maintenance schedules |
| Process equipment | Maintenance, calibration |
| Supporting services | Utilities, IT systems |
| Information systems | Backup, security, validation |
### 6.4 Work Environment and Contamination Control
| Environment Factor | Control Method | Monitoring |
|-------------------|----------------|------------|
| Temperature | HVAC control | Continuous logging |
| Humidity | HVAC control | Continuous logging |
| Cleanliness | Cleaning procedures | Particle counts |
| Lighting | Lux levels | Periodic verification |
| ESD protection | Grounding, ionization | Periodic testing |
---
## Clause 7: Product Realization
### 7.1 Planning of Product Realization
| Planning Element | Content |
|-----------------|---------|
| Quality objectives for product | Product-specific quality targets |
| Processes and documentation | Process flow, required documents |
| Verification and validation | Test methods, acceptance criteria |
| Records | Required quality records |
| Risk management | Per ISO 14971 |
### 7.2 Customer-Related Processes
#### 7.2.1 Determination of Requirements
| Requirement Type | Source |
|-----------------|--------|
| Customer-specified | Contract, purchase order |
| Not stated but necessary | Intended use analysis |
| Regulatory | Applicable standards, regulations |
| Organization-defined | Internal specifications |
#### 7.2.2 Review of Requirements
| Review Element | Verification |
|----------------|--------------|
| Requirements defined | Complete specification |
| Differences resolved | Documented resolution |
| Ability to meet | Feasibility assessment |
| Risk management | Initial risk assessment |
#### 7.2.3 Communication
| Communication Type | Method |
|-------------------|--------|
| Product information | Catalogs, IFU |
| Inquiries and orders | Sales process |
| Feedback and complaints | Customer feedback system |
| Advisory notices | Field safety notices |
### 7.3 Design and Development
| Stage | Clause | Requirements |
|-------|--------|--------------|
| Planning | 7.3.2 | Stages, reviews, responsibilities |
| Inputs | 7.3.3 | Functional, performance, regulatory |
| Outputs | 7.3.4 | Meet inputs, acceptance criteria |
| Review | 7.3.5 | Evaluate ability to meet requirements |
| Verification | 7.3.6 | Outputs meet inputs |
| Validation | 7.3.7 | Product meets intended use |
| Transfer | 7.3.8 | Verified before production |
| Changes | 7.3.9 | Controlled, reviewed, verified |
### 7.4 Purchasing
#### 7.4.1 Purchasing Process
| Control Element | Implementation |
|-----------------|----------------|
| Supplier evaluation | Qualification procedure |
| Selection criteria | Quality, delivery, cost |
| Monitoring | Performance metrics |
| Re-evaluation | Periodic review |
**Supplier Classification:**
```
Category A: Critical - Affects product safety/performance
- Full qualification audit
- Annual performance review
- Quality agreement required
Category B: Major - Affects product quality
- Qualification questionnaire
- Periodic performance review
- Quality requirements communicated
Category C: Minor - Indirect impact
- Initial assessment
- Issue-based review
- Standard terms
```
#### 7.4.2 Purchasing Information
| Information Required | Purpose |
|---------------------|---------|
| Product specifications | Clear requirements |
| QMS requirements | Supplier system expectations |
| Personnel competence | Where applicable |
| Approval requirements | Where applicable |
#### 7.4.3 Verification of Purchased Product
| Verification Method | Application |
|--------------------|-------------|
| Incoming inspection | Standard verification |
| Source inspection | Critical items |
| Certificate of Conformance | Documented evidence |
| Certificate of Analysis | Material verification |
### 7.5 Production and Service Provision
#### 7.5.1 Control of Production and Service Provision
| Control Element | Implementation |
|-----------------|----------------|
| Product information | Specifications, drawings |
| Work instructions | Where necessary |
| Suitable equipment | Qualified equipment |
| Monitoring devices | Calibrated instruments |
| Implementation of monitoring | Inspections, tests |
| Defined processes | Process parameters |
| Labeling and packaging | Per requirements |
#### 7.5.2 Cleanliness of Product
| Cleanliness Control | Method |
|--------------------|--------|
| Product cleaning | Validated procedures |
| Contamination prevention | Controlled environment |
| Process aids | Qualified, controlled |
#### 7.5.3 Installation Activities
| Requirement | Implementation |
|-------------|----------------|
| Installation requirements | Documented instructions |
| Acceptance criteria | Defined criteria |
| Records | Installation records |
#### 7.5.4 Servicing Activities
| Requirement | Implementation |
|-------------|----------------|
| Documented requirements | Service procedures |
| Reference materials | Service manuals |
| Measurement equipment | Calibrated |
| Records | Service records |
#### 7.5.5 Particular Requirements for Sterile Medical Devices
| Process | Control |
|---------|---------|
| Sterilization validation | Per ISO 11135/11137/17665 |
| Parameter control | Monitoring records |
| Sterile barrier | Validated packaging |
#### 7.5.6 Validation of Processes
| Validation Required When | Evidence |
|-------------------------|----------|
| Output cannot be verified | Validation protocol and report |
| Deficiencies appear only in use | Process capability data |
| Special processes | Qualified operators |
**Process Validation Elements:**
- Equipment qualification (IQ/OQ/PQ)
- Process parameters
- Monitoring methods
- Operator qualification
- Revalidation criteria
#### 7.5.7 Particular Requirements for Validation
| Requirement | Implementation |
|-------------|----------------|
| Documented procedures | Validation SOPs |
| Defined methods | Statistical methods |
| Acceptance criteria | Predefined criteria |
| Software validation | Where applicable |
| Revalidation | Change-triggered |
#### 7.5.8 Identification
| Identification Type | Method |
|--------------------|--------|
| Product | Labels, markings |
| Documentation | Document numbers |
| Unique Device Identification | UDI per regulation |
#### 7.5.9 Traceability
| Traceability Element | Record |
|---------------------|--------|
| Components | Lot/batch numbers |
| Materials | Certificates |
| Work environment | Environmental records |
| Measurement equipment | Calibration records |
| Personnel | Training records |
| Distribution | Shipping records |
#### 7.5.10 Customer Property
| Control | Implementation |
|---------|----------------|
| Identification | Marking, segregation |
| Verification | Incoming inspection |
| Protection | Storage conditions |
| Safeguarding | Security measures |
| Reporting | Loss/damage notification |
#### 7.5.11 Preservation of Product
| Preservation Element | Control |
|---------------------|---------|
| Identification | Labels, markings |
| Handling | Procedures |
| Packaging | Specifications |
| Storage | Conditions, FIFO |
| Protection | Environmental controls |
### 7.6 Control of Monitoring and Measuring Equipment
| Control Element | Implementation |
|-----------------|----------------|
| Calibration | At specified intervals |
| Adjustment | As needed |
| Identification | Calibration status |
| Safeguarding | Protection from damage |
| Software validation | Where applicable |
| Records | Calibration records |
---
## Clause 8: Measurement, Analysis and Improvement
### 8.1 General
**Monitoring and Measurement Requirements:**
- Demonstrate product conformity
- Ensure QMS conformity
- Maintain QMS effectiveness
### 8.2 Monitoring and Measurement
#### 8.2.1 Feedback
| Feedback Source | Collection Method |
|-----------------|-------------------|
| Customer complaints | Complaint system |
| Customer surveys | Periodic surveys |
| Field feedback | Service reports |
| Regulatory feedback | Inspection findings |
#### 8.2.2 Complaint Handling
| Process Step | Requirements |
|--------------|--------------|
| Receipt | Timely logging |
| Investigation | Root cause analysis |
| Corrective action | If warranted |
| Regulatory reporting | If required |
| Trend analysis | Aggregate review |
#### 8.2.3 Reporting to Regulatory Authorities
| Report Type | Trigger | Timeline |
|-------------|---------|----------|
| MDR (Medical Device Report) | Death/serious injury | 30 days (5 if awareness) |
| FSCA (Field Safety Corrective Action) | Safety issue | Without delay |
| Periodic Safety Update | Per regulation | Per schedule |
#### 8.2.4 Internal Audit
| Audit Element | Requirement |
|---------------|-------------|
| Planned program | Risk-based schedule |
| Criteria and scope | Defined per audit |
| Auditor selection | Independent, competent |
| Procedure | Documented process |
| Records | Audit reports, findings |
| Follow-up | CAPA, verification |
**Audit Program Template:**
```
ANNUAL INTERNAL AUDIT PROGRAM
Year: [Year]
| Audit # | Area/Process | Scope | Auditor | Planned Date | Status |
|---------|--------------|-------|---------|--------------|--------|
| IA-01 | Document Control | 4.2.3, 4.2.4 | [Name] | Q1 | |
| IA-02 | Design Control | 7.3 | [Name] | Q2 | |
| IA-03 | Production | 7.5 | [Name] | Q2 | |
| IA-04 | Purchasing | 7.4 | [Name] | Q3 | |
| IA-05 | CAPA | 8.5.2, 8.5.3 | [Name] | Q3 | |
| IA-06 | Management Review | 5.6 | [Name] | Q4 | |
Risk Considerations:
- Previous audit findings
- Regulatory changes
- Process changes
- Complaint trends
```
#### 8.2.5 Monitoring and Measurement of Processes
| Monitoring Type | Method |
|-----------------|--------|
| Process metrics | KPIs, trend analysis |
| Process audits | Internal audits |
| Process reviews | Management review |
#### 8.2.6 Monitoring and Measurement of Product
| Stage | Verification |
|-------|--------------|
| Incoming | Incoming inspection |
| In-process | In-process inspection |
| Final | Final inspection and test |
| Release | Authorized release |
### 8.3 Control of Nonconforming Product
| Control Element | Requirement |
|-----------------|-------------|
| Identification | Clear marking |
| Segregation | Physical separation |
| Documentation | NC record |
| Disposition | Use as is/rework/scrap/return |
| Concession | If accepted |
| Reinspection | After rework |
| Investigation | For detected after delivery |
**Nonconformity Disposition Options:**
```
1. Use As Is (Concession)
- Does not affect safety/performance
- Customer approval if applicable
- Documented justification
2. Rework
- Per approved procedure
- Reinspection required
- Records maintained
3. Scrap/Reject
- Physical destruction or marking
- Prevented from reentry
- Documented disposal
4. Return to Supplier
- Communication with supplier
- Replacement or credit
- Root cause if systemic
```
### 8.4 Analysis of Data
| Data Source | Analysis |
|-------------|----------|
| Feedback | Complaint trends, satisfaction |
| Nonconformity | Defect Pareto, trends |
| Process performance | Capability, trends |
| Supplier | Performance trends |
| Audit | Finding trends |
### 8.5 Improvement
#### 8.5.1 General
**Improvement Sources:**
- Quality policy
- Quality objectives
- Audit results
- Data analysis
- Corrective actions
- Preventive actions
- Management review
#### 8.5.2 Corrective Action
| Process Step | Requirement |
|--------------|-------------|
| Review nonconformity | Including complaints |
| Determine cause | Root cause analysis |
| Evaluate action need | Based on risk |
| Determine action | Proportionate to risk |
| Implement action | Execute plan |
| Document results | Records |
| Review effectiveness | Verification |
#### 8.5.3 Preventive Action
| Process Step | Requirement |
|--------------|-------------|
| Determine potential NC | Risk analysis, trends |
| Evaluate action need | Prevention opportunity |
| Determine action | Proportionate to risk |
| Implement action | Execute plan |
| Document results | Records |
| Review effectiveness | Verification |
FILE:references/qms-process-templates.md
# QMS Process Templates
Ready-to-use templates for ISO 13485 QMS processes including document control, internal audit, CAPA, and supplier management.
---
## Table of Contents
- [Document Control Templates](#document-control-templates)
- [Internal Audit Templates](#internal-audit-templates)
- [CAPA Templates](#capa-templates)
- [Supplier Management Templates](#supplier-management-templates)
- [Training Templates](#training-templates)
- [Nonconformity Templates](#nonconformity-templates)
---
## Document Control Templates
### Document Master List
```
DOCUMENT MASTER LIST
Organization: [Company Name]
Last Updated: [Date]
Maintained By: Document Control
| Doc # | Title | Rev | Effective Date | Status | Owner | Next Review |
|-------|-------|-----|----------------|--------|-------|-------------|
| QM-001 | Quality Manual | 03 | 2024-01-15 | Effective | QMR | 2025-01-15 |
| SOP-01-001 | Document Control | 04 | 2024-03-01 | Effective | QA Mgr | 2025-03-01 |
| SOP-01-002 | Record Control | 02 | 2024-02-01 | Effective | QA Mgr | 2025-02-01 |
| | | | | | | |
Status Values: Draft, Under Review, Effective, Obsolete
```
### Document Change Request
```
DOCUMENT CHANGE REQUEST
DCR Number: DCR-[YYYY]-[NNN]
Date Submitted: [Date]
Submitted By: [Name]
DOCUMENT INFORMATION
Document Number: [Number]
Document Title: [Title]
Current Revision: [Rev]
CHANGE REQUEST
Change Type: [ ] Administrative [ ] Minor [ ] Major [ ] Emergency
Requested Change: [Description of change]
Reason for Change:
[ ] Regulatory requirement
[ ] Process improvement
[ ] Nonconformity/CAPA
[ ] Organizational change
[ ] Error correction
[ ] Other: [Specify]
Justification: [Detailed justification]
IMPACT ASSESSMENT
Training Required: [ ] Yes [ ] No
If yes, who: [Roles/departments]
Other Documents Affected: [List]
Regulatory Filing Impact: [ ] Yes [ ] No
If yes, details: [Explain]
APPROVALS
Requested By: _________________ Date: _______
Document Owner: _________________ Date: _______
QA Approval: _________________ Date: _______
COMPLETION
New Revision: [Rev]
Effective Date: [Date]
Training Completed: [ ] Yes [ ] N/A
Distribution Completed: [ ] Yes
```
### Document Review Record
```
DOCUMENT REVIEW RECORD
Document Number: [Number]
Document Title: [Title]
Current Revision: [Rev]
Review Due Date: [Date]
Review Completed: [Date]
REVIEWERS
| Reviewer | Role | Review Date | Comments | Signature |
|----------|------|-------------|----------|-----------|
| [Name] | [Role] | [Date] | [Comments] | |
| [Name] | [Role] | [Date] | [Comments] | |
REVIEW OUTCOME
[ ] No changes required - document remains current
[ ] Minor changes required - see attached DCR
[ ] Major revision required - see attached DCR
[ ] Document obsolete - initiate retirement
NEXT REVIEW
Next Review Date: [Date]
APPROVAL
Review Completed By: _________________ Date: _______
Approved By: _________________ Date: _______
```
---
## Internal Audit Templates
### Annual Audit Schedule
```
INTERNAL AUDIT SCHEDULE
Year: [Year]
Prepared By: [Name]
Approved By: [Name]
Date: [Date]
AUDIT SCHEDULE
| Audit # | Process/Area | ISO Clauses | Lead Auditor | Q1 | Q2 | Q3 | Q4 |
|---------|--------------|-------------|--------------|----|----|----|----|
| IA-001 | Document Control | 4.2.3, 4.2.4 | [Name] | X | | | |
| IA-002 | Management Review | 5.6 | [Name] | | X | | |
| IA-003 | Training | 6.2 | [Name] | | X | | |
| IA-004 | Design Control | 7.3 | [Name] | | | X | |
| IA-005 | Purchasing | 7.4 | [Name] | | | X | |
| IA-006 | Production | 7.5 | [Name] | | | | X |
| IA-007 | CAPA | 8.5.2, 8.5.3 | [Name] | | | | X |
RISK FACTORS CONSIDERED
[ ] Previous audit findings
[ ] Regulatory changes
[ ] Process changes
[ ] Complaint trends
[ ] Management concerns
SCHEDULE REVISION LOG
| Rev | Date | Change | Approved By |
|-----|------|--------|-------------|
| 00 | [Date] | Initial release | [Name] |
```
### Audit Plan
```
INTERNAL AUDIT PLAN
Audit Number: IA-[YYYY]-[NNN]
Audit Date(s): [Date(s)]
Audit Type: [ ] Process [ ] System [ ] Product
SCOPE
Process/Area: [Name]
ISO 13485 Clauses: [List]
Regulatory Requirements: [If applicable]
Locations: [Locations]
AUDIT TEAM
Lead Auditor: [Name]
Auditor(s): [Names]
Observer(s): [If any]
AUDITEE CONTACTS
Process Owner: [Name]
Other Contacts: [Names]
AUDIT CRITERIA
- ISO 13485:2016
- [Organization procedures]
- [Regulatory requirements]
AUDIT SCHEDULE
| Time | Activity | Participants |
|------|----------|--------------|
| 09:00 | Opening meeting | All |
| 09:30 | Document review | Auditor, Doc Control |
| 10:30 | Process observation | Auditor, Operators |
| 12:00 | Lunch | |
| 13:00 | Record review | Auditor, QA |
| 14:30 | Interviews | Selected personnel |
| 15:30 | Auditor caucus | Audit team |
| 16:00 | Closing meeting | All |
PREPARATION CHECKLIST
[ ] Previous audit reports reviewed
[ ] Procedures reviewed
[ ] Checklist prepared
[ ] Auditees notified
[ ] Resources arranged
```
### Audit Checklist Template
```
INTERNAL AUDIT CHECKLIST
Audit Number: IA-[YYYY]-[NNN]
Process: [Process Name]
Auditor: [Name]
Date: [Date]
INSTRUCTIONS
C = Conforming, NC = Nonconforming, OBS = Observation, N/A = Not Applicable
CHECKLIST
| # | Requirement | Reference | Evidence Reviewed | Finding | Notes |
|---|-------------|-----------|-------------------|---------|-------|
| 1 | Is the procedure current and approved? | 4.2.3 | [Evidence] | C/NC/OBS | |
| 2 | Are personnel trained on the procedure? | 6.2 | [Evidence] | C/NC/OBS | |
| 3 | Are records maintained as required? | 4.2.4 | [Evidence] | C/NC/OBS | |
| 4 | Is the process performed as documented? | 4.1 | [Evidence] | C/NC/OBS | |
| 5 | Are monitoring activities performed? | 8.2.5 | [Evidence] | C/NC/OBS | |
INTERVIEWS CONDUCTED
| Person | Role | Topics Discussed |
|--------|------|------------------|
| [Name] | [Role] | [Topics] |
DOCUMENTS REVIEWED
| Document # | Title | Rev | Findings |
|------------|-------|-----|----------|
| [Number] | [Title] | [Rev] | [Findings] |
RECORDS SAMPLED
| Record Type | Sample Size | Sample IDs | Findings |
|-------------|-------------|------------|----------|
| [Type] | [N] | [IDs] | [Findings] |
AUDITOR SIGNATURE: _________________ Date: _______
```
### Audit Report
```
INTERNAL AUDIT REPORT
Audit Number: IA-[YYYY]-[NNN]
Report Date: [Date]
Report Status: [ ] Draft [ ] Final
AUDIT SUMMARY
Audit Date(s): [Date(s)]
Process/Area: [Name]
ISO Clauses Covered: [List]
Lead Auditor: [Name]
Audit Team: [Names]
AUDIT SCOPE
[Description of scope]
AUDIT OBJECTIVES
[List objectives]
EXECUTIVE SUMMARY
[Brief summary of audit results]
FINDINGS SUMMARY
| Type | Count |
|------|-------|
| Major Nonconformity | [N] |
| Minor Nonconformity | [N] |
| Observation | [N] |
| Opportunity for Improvement | [N] |
DETAILED FINDINGS
FINDING 1
Number: IA-[YYYY]-[NNN]-F01
Classification: [ ] Major NC [ ] Minor NC [ ] Observation [ ] OFI
Requirement: [Clause/requirement reference]
Statement: [Objective description of finding]
Evidence: [Evidence supporting finding]
Auditee Response Due: [Date]
[Repeat for each finding]
POSITIVE OBSERVATIONS
[List areas of good practice observed]
CONCLUSION
[Overall conclusion on process effectiveness]
REPORT DISTRIBUTION
| Name | Role | Date |
|------|------|------|
| [Name] | Process Owner | [Date] |
| [Name] | QA Manager | [Date] |
| [Name] | Management Rep | [Date] |
APPROVALS
Lead Auditor: _________________ Date: _______
QA Manager: _________________ Date: _______
```
---
## CAPA Templates
### CAPA Request Form
```
CORRECTIVE AND PREVENTIVE ACTION REQUEST
CAPA Number: CAPA-[YYYY]-[NNN]
Date Opened: [Date]
Initiated By: [Name]
CAPA TYPE
[ ] Corrective Action (response to existing nonconformity)
[ ] Preventive Action (prevent potential nonconformity)
SOURCE
[ ] Customer complaint: Reference #_______
[ ] Internal audit: Audit #_______
[ ] External audit: Audit #_______
[ ] Nonconformity: NC #_______
[ ] Process deviation
[ ] Management review action
[ ] Trend analysis
[ ] Risk assessment
[ ] Other: _______
CLASSIFICATION
Severity: [ ] Critical [ ] Major [ ] Minor
Regulatory Reportable: [ ] Yes [ ] No
PROBLEM DESCRIPTION
[Detailed description of the problem or potential problem]
IMMEDIATE CONTAINMENT (if applicable)
Actions Taken: [Description]
Date: [Date]
Responsible: [Name]
ASSIGNMENT
Process Owner: [Name]
CAPA Owner: [Name]
Due Date for Root Cause: [Date]
Target Closure Date: [Date]
APPROVAL TO PROCEED
Approved By: _________________ Date: _______
```
### Root Cause Analysis Record
```
ROOT CAUSE ANALYSIS
CAPA Number: CAPA-[YYYY]-[NNN]
Analysis Date: [Date]
Analyst: [Name]
PROBLEM STATEMENT
[Clear, specific statement of the problem]
INVESTIGATION TEAM
| Name | Role | Contribution |
|------|------|--------------|
| [Name] | [Role] | [Area of expertise] |
INVESTIGATION METHOD
[ ] 5 Why Analysis
[ ] Fishbone Diagram
[ ] Fault Tree Analysis
[ ] Human Factors Analysis
[ ] Other: _______
INVESTIGATION DETAILS
5 WHY ANALYSIS
Why 1: [First why]
Answer: [Answer]
Why 2: [Second why based on answer]
Answer: [Answer]
Why 3: [Third why based on answer]
Answer: [Answer]
Why 4: [Fourth why based on answer]
Answer: [Answer]
Why 5: [Fifth why based on answer]
Answer: [Answer]
ROOT CAUSE STATEMENT
[Clear statement of identified root cause]
ROOT CAUSE CATEGORY
[ ] Process/Procedure
[ ] Training/Competency
[ ] Equipment/Material
[ ] Design
[ ] Human Error
[ ] Communication
[ ] Management System
[ ] External Factor
CONTRIBUTING FACTORS
[List any contributing factors]
EVIDENCE SUPPORTING ROOT CAUSE
[List evidence]
APPROVAL
Analysis By: _________________ Date: _______
Reviewed By: _________________ Date: _______
```
### CAPA Action Plan
```
CAPA ACTION PLAN
CAPA Number: CAPA-[YYYY]-[NNN]
Root Cause: [Brief statement]
Plan Date: [Date]
Plan Owner: [Name]
CORRECTIVE/PREVENTIVE ACTIONS
Action 1:
Description: [Detailed action description]
Responsible: [Name]
Due Date: [Date]
Resources Required: [Resources]
Success Criteria: [How completion verified]
Action 2:
Description: [Detailed action description]
Responsible: [Name]
Due Date: [Date]
Resources Required: [Resources]
Success Criteria: [How completion verified]
[Continue for additional actions]
RELATED CHANGES
Documents Affected: [List]
Training Required: [Description]
Process Changes: [Description]
Equipment Changes: [Description]
RISK ASSESSMENT
Residual Risk After Implementation: [ ] High [ ] Medium [ ] Low
Justification: [Explanation]
APPROVAL
Plan Developed By: _________________ Date: _______
Approved By: _________________ Date: _______
```
### CAPA Effectiveness Verification
```
CAPA EFFECTIVENESS VERIFICATION
CAPA Number: CAPA-[YYYY]-[NNN]
Verification Date: [Date]
Verified By: [Name]
ACTIONS COMPLETED
| Action | Completion Date | Evidence |
|--------|-----------------|----------|
| [Action 1] | [Date] | [Reference] |
| [Action 2] | [Date] | [Reference] |
EFFECTIVENESS CRITERIA
[Criteria established during action planning]
VERIFICATION METHOD
[ ] Data analysis (trends, metrics)
[ ] Process audit
[ ] Record review
[ ] Product inspection
[ ] Customer feedback review
[ ] Other: _______
VERIFICATION PERIOD
From: [Date] To: [Date]
VERIFICATION RESULTS
[Detailed results of verification activities]
DATA/EVIDENCE REVIEWED
| Data Type | Period | Result |
|-----------|--------|--------|
| [Type] | [Period] | [Result] |
EFFECTIVENESS CONCLUSION
[ ] Effective - Root cause eliminated, problem resolved
[ ] Partially Effective - Improvement noted, additional action needed
[ ] Not Effective - Problem persists, reopen CAPA
If not effective, describe additional actions:
[Description]
CAPA CLOSURE
[ ] Approved for closure
[ ] Not approved - additional action required
Verified By: _________________ Date: _______
Approved By: _________________ Date: _______
```
---
## Supplier Management Templates
### Approved Supplier List
```
APPROVED SUPPLIER LIST
Organization: [Company Name]
Last Updated: [Date]
Maintained By: [Name]
| Supplier | Supplier # | Category | Products/Services | Status | Qualification Date | Next Review |
|----------|-----------|----------|-------------------|--------|-------------------|-------------|
| [Name] | SUP-001 | A | [Products] | Approved | [Date] | [Date] |
| [Name] | SUP-002 | B | [Products] | Conditional | [Date] | [Date] |
Category:
A = Critical (affects safety/performance)
B = Major (affects quality)
C = Minor (indirect impact)
Status:
Approved = Full use authorized
Conditional = Limited use, monitoring
Probation = Performance issues, enhanced monitoring
Disqualified = Use not authorized
Revision History:
| Rev | Date | Change | Approved By |
|-----|------|--------|-------------|
| 01 | [Date] | Initial release | [Name] |
```
### Supplier Evaluation Form
```
SUPPLIER EVALUATION
Supplier Name: [Name]
Supplier Number: [Number]
Evaluation Date: [Date]
Evaluated By: [Name]
Evaluation Type: [ ] Initial [ ] Periodic [ ] For Cause
SUPPLIER INFORMATION
Address: [Address]
Contact: [Name, Title]
Phone: [Phone]
Email: [Email]
Products/Services: [Description]
PROPOSED CATEGORY
[ ] A - Critical (affects safety/performance)
[ ] B - Major (affects quality)
[ ] C - Minor (indirect impact)
EVALUATION CRITERIA
1. QUALITY MANAGEMENT SYSTEM (30 points max)
[ ] ISO 13485 Certified (30 pts)
[ ] ISO 9001 Certified (20 pts)
[ ] Documented QMS (10 pts)
[ ] No formal QMS (0 pts)
Score: ___/30
2. QUALITY HISTORY (25 points max)
Reject Rate: ___% (0-1% = 25 pts, 1-3% = 15 pts, >3% = 0 pts)
Score: ___/25
3. DELIVERY PERFORMANCE (20 points max)
On-Time Delivery: ___% (>95% = 20 pts, 90-95% = 10 pts, <90% = 0 pts)
Score: ___/20
4. TECHNICAL CAPABILITY (15 points max)
[ ] Exceeds requirements (15 pts)
[ ] Meets requirements (10 pts)
[ ] Marginally meets (5 pts)
Score: ___/15
5. FINANCIAL STABILITY (10 points max)
[ ] Strong (10 pts)
[ ] Adequate (5 pts)
[ ] Questionable (0 pts)
Score: ___/10
TOTAL SCORE: ___/100
QUALIFICATION DECISION
>80 = Approved
60-80 = Conditional (monitoring required)
<60 = Not Approved
Decision: [ ] Approved [ ] Conditional [ ] Not Approved
APPROVAL
Evaluated By: _________________ Date: _______
QA Approval: _________________ Date: _______
```
### Supplier Performance Scorecard
```
SUPPLIER PERFORMANCE SCORECARD
Supplier: [Name]
Supplier #: [Number]
Period: [Q1/Q2/Q3/Q4] [Year]
Prepared By: [Name]
PERFORMANCE METRICS
1. QUALITY (40% weight)
Total Lots Received: [N]
Lots Rejected: [N]
Accept Rate: ___% Target: >98%
Score: ___/40
2. DELIVERY (30% weight)
Total Orders: [N]
On-Time Deliveries: [N]
On-Time Rate: ___% Target: >95%
Score: ___/30
3. RESPONSIVENESS (15% weight)
Issues Reported: [N]
Resolved <5 days: [N]
Response Rate: ___% Target: >90%
Score: ___/15
4. DOCUMENTATION (15% weight)
CoC Required: [N]
CoC Complete: [N]
Documentation Rate: ___% Target: 100%
Score: ___/15
TOTAL SCORE: ___/100
PERFORMANCE TREND
| Period | Quality | Delivery | Response | Docs | Total |
|--------|---------|----------|----------|------|-------|
| Q1 | | | | | |
| Q2 | | | | | |
| Q3 | | | | | |
| Q4 | | | | | |
ISSUES/CONCERNS
[List any quality or delivery issues during period]
ACTIONS REQUIRED
[ ] None - Performance acceptable
[ ] Enhanced monitoring
[ ] Supplier corrective action request
[ ] Supplier audit
[ ] Consider alternative supplier
NEXT REVIEW: [Date]
Prepared By: _________________ Date: _______
Reviewed By: _________________ Date: _______
```
---
## Training Templates
### Training Record
```
EMPLOYEE TRAINING RECORD
Employee Name: [Name]
Employee ID: [ID]
Department: [Department]
Job Title: [Title]
Date of Hire: [Date]
REQUIRED TRAINING
| Training | Requirement | Initial Date | Last Date | Next Due | Status |
|----------|-------------|--------------|-----------|----------|--------|
| ISO 13485 Awareness | Initial + Annual | [Date] | [Date] | [Date] | Current |
| Document Control | Initial + On Change | [Date] | [Date] | [Date] | Current |
| CAPA Procedure | Initial + On Change | [Date] | [Date] | [Date] | Due |
| Job-Specific | Per competency matrix | [Date] | [Date] | [Date] | Current |
TRAINING HISTORY
| Date | Training | Method | Duration | Trainer | Assessment | Result |
|------|----------|--------|----------|---------|------------|--------|
| [Date] | [Title] | Classroom | 2 hrs | [Name] | Written test | Pass |
| [Date] | [Title] | OJT | 4 hrs | [Name] | Observation | Pass |
COMPETENCY VERIFICATION
| Competency | Method | Date | Verified By | Result |
|------------|--------|------|-------------|--------|
| [Skill] | Observation | [Date] | [Name] | Qualified |
| [Skill] | Test | [Date] | [Name] | Qualified |
Employee Signature: _________________ Date: _______
Supervisor Signature: _________________ Date: _______
```
### Training Attendance Record
```
TRAINING ATTENDANCE RECORD
Training Title: [Title]
Training Date: [Date]
Trainer: [Name]
Location: [Location]
Duration: [Hours]
TRAINING CONTENT
[Brief description of content covered]
ATTENDEES
| Name | Employee ID | Department | Signature | Assessment Result |
|------|-------------|------------|-----------|-------------------|
| [Name] | [ID] | [Dept] | | Pass/Fail |
| [Name] | [ID] | [Dept] | | Pass/Fail |
ASSESSMENT METHOD
[ ] Written test (attach copy)
[ ] Practical demonstration
[ ] Verbal Q&A
[ ] Observation
[ ] N/A
TRAINING MATERIALS
[ ] Presentation: [Reference]
[ ] Procedure: [Reference]
[ ] Other: [Reference]
Trainer Signature: _________________ Date: _______
Training Coordinator: _________________ Date: _______
```
---
## Nonconformity Templates
### Nonconformity Report
```
NONCONFORMITY REPORT
NC Number: NC-[YYYY]-[NNN]
Date Identified: [Date]
Identified By: [Name]
NONCONFORMITY TYPE
[ ] Product [ ] Process [ ] Document [ ] System
NONCONFORMITY SOURCE
[ ] Incoming inspection
[ ] In-process inspection
[ ] Final inspection
[ ] Customer complaint
[ ] Internal audit
[ ] External audit
[ ] Other: _______
PRODUCT IDENTIFICATION (if applicable)
Product Name: [Name]
Part Number: [Number]
Lot/Batch: [Number]
Quantity Affected: [N]
NONCONFORMITY DESCRIPTION
[Detailed, objective description of the nonconformity]
REQUIREMENT
[Reference to requirement that was not met]
CONTAINMENT ACTION
Action Taken: [Description]
Quantity Contained: [N]
Location: [Location]
Date: [Date]
By: [Name]
DISPOSITION
[ ] Use As Is - Justification: _______
[ ] Rework - Per procedure: _______
[ ] Scrap - Method: _______
[ ] Return to Supplier - RMA #: _______
[ ] Other: _______
Disposition By: [Name]
Disposition Date: [Date]
CAPA REQUIRED?
[ ] Yes - CAPA #: _______
[ ] No - Justification: _______
CLOSURE
All actions complete: [ ] Yes
NC Closed By: _________________ Date: _______
QA Approval: _________________ Date: _______
```
### Material Review Board Record
```
MATERIAL REVIEW BOARD (MRB) RECORD
MRB Number: MRB-[YYYY]-[NNN]
Date: [Date]
NC Reference: NC-[YYYY]-[NNN]
NONCONFORMING MATERIAL
Product: [Name]
Part Number: [Number]
Lot/Batch: [Number]
Quantity: [N]
NONCONFORMITY DESCRIPTION
[Description from NC report]
MRB PARTICIPANTS
| Name | Role | Signature |
|------|------|-----------|
| [Name] | QA Representative | |
| [Name] | Engineering | |
| [Name] | Production | |
| [Name] | Other | |
DISPOSITION OPTIONS CONSIDERED
1. Use As Is
Technical Justification: [Justification]
Risk Assessment: [Assessment]
2. Rework
Procedure: [Reference]
Feasibility: [Assessment]
3. Scrap
Cost Impact: [Amount]
MRB DECISION
[ ] Use As Is - Customer notification required: [ ] Yes [ ] No
[ ] Rework per: [Procedure reference]
[ ] Scrap
[ ] Return to Supplier
RATIONALE
[Detailed rationale for decision]
APPROVALS
| Role | Name | Signature | Date |
|------|------|-----------|------|
| QA | [Name] | | [Date] |
| Engineering | [Name] | | [Date] |
| Production | [Name] | | [Date] |
FOLLOW-UP ACTIONS
[ ] CAPA initiated: CAPA-_______
[ ] Customer notified: Date: _______
[ ] Supplier notified: Date: _______
[ ] Other: _______
```
FILE:scripts/qms_audit_checklist.py
#!/usr/bin/env python3
"""
QMS Internal Audit Checklist Generator
Generates audit checklists for ISO 13485:2016 clauses and QMS processes.
Supports process audits, system audits, and clause-specific audits.
Usage:
python qms_audit_checklist.py --clause 7.3
python qms_audit_checklist.py --process design-control
python qms_audit_checklist.py --audit-type system --output json
python qms_audit_checklist.py --interactive
"""
import argparse
import json
import sys
from datetime import datetime
from typing import Optional
# ISO 13485:2016 Clause Structure with Audit Questions
ISO13485_CLAUSES = {
"4.1": {
"title": "General Requirements",
"questions": [
"Are QMS processes identified and documented?",
"Is the sequence and interaction of processes defined?",
"Are criteria and methods for process operation determined?",
"Are resources and information available for process operation?",
"Are processes monitored, measured, and analyzed?",
"Are actions taken to achieve planned results?",
"Is outsourced process control documented?",
"Are changes to processes managed?"
]
},
"4.2.1": {
"title": "Documentation Requirements - General",
"questions": [
"Is a quality policy documented?",
"Are quality objectives documented?",
"Is a quality manual maintained?",
"Are required documented procedures established?",
"Are documents needed for process planning and operation maintained?",
"Are required records maintained?",
"Is a medical device file established for each device type?"
]
},
"4.2.2": {
"title": "Quality Manual",
"questions": [
"Does the quality manual include QMS scope?",
"Are exclusions justified?",
"Are documented procedures included or referenced?",
"Is the interaction between processes described?",
"Is the quality manual controlled?"
]
},
"4.2.3": {
"title": "Control of Documents",
"questions": [
"Are documents approved before issue?",
"Are documents reviewed and updated as necessary?",
"Are changes and revision status identified?",
"Are current versions available at points of use?",
"Are documents legible and identifiable?",
"Are external documents identified and controlled?",
"Is unintended use of obsolete documents prevented?",
"Is there a document change control process?"
]
},
"4.2.4": {
"title": "Control of Records",
"questions": [
"Is there a procedure for record control?",
"Are records legible and identifiable?",
"Are records retrievable?",
"Are retention times defined?",
"Is protection from damage ensured?",
"Are confidential records protected?",
"Is record disposal controlled?"
]
},
"5.1": {
"title": "Management Commitment",
"questions": [
"Is there evidence of management commitment to QMS?",
"Is the importance of regulatory requirements communicated?",
"Is a quality policy established?",
"Are quality objectives established?",
"Are management reviews conducted?",
"Are resources provided for QMS?"
]
},
"5.2": {
"title": "Customer Focus",
"questions": [
"Are customer requirements determined?",
"Are applicable regulatory requirements determined?",
"Are customer and regulatory requirements met?",
"Is customer satisfaction enhanced?"
]
},
"5.3": {
"title": "Quality Policy",
"questions": [
"Is the quality policy appropriate to the organization?",
"Does it include commitment to compliance?",
"Does it include commitment to effectiveness?",
"Does it provide framework for quality objectives?",
"Is it communicated and understood?",
"Is it reviewed for continuing suitability?"
]
},
"5.4.1": {
"title": "Quality Objectives",
"questions": [
"Are quality objectives measurable?",
"Are they consistent with quality policy?",
"Are they established at relevant functions?",
"Do they include product requirements?",
"Do they include compliance requirements?"
]
},
"5.4.2": {
"title": "QMS Planning",
"questions": [
"Is QMS planning carried out to meet requirements?",
"Is QMS planning done to meet quality objectives?",
"Is QMS integrity maintained during changes?"
]
},
"5.5.1": {
"title": "Responsibility and Authority",
"questions": [
"Are responsibilities and authorities defined?",
"Are they documented?",
"Are they communicated?",
"Are interrelationships defined?"
]
},
"5.5.2": {
"title": "Management Representative",
"questions": [
"Is a management representative appointed?",
"Is authority to ensure QMS processes established?",
"Is authority to report to top management defined?",
"Is authority to promote awareness of requirements defined?"
]
},
"5.5.3": {
"title": "Internal Communication",
"questions": [
"Are communication processes established?",
"Is QMS effectiveness communicated?",
"Is information communicated appropriately?"
]
},
"5.6": {
"title": "Management Review",
"questions": [
"Are management reviews planned?",
"Are all required inputs reviewed?",
"Are outputs documented?",
"Are action items followed up?",
"Are records maintained?"
]
},
"6.1": {
"title": "Provision of Resources",
"questions": [
"Are resources determined?",
"Are resources provided for QMS?",
"Are resources provided for customer satisfaction?",
"Are resources provided for regulatory compliance?"
]
},
"6.2": {
"title": "Human Resources",
"questions": [
"Is competence defined for personnel?",
"Is training provided to achieve competence?",
"Is training effectiveness evaluated?",
"Is awareness of job relevance ensured?",
"Are training records maintained?"
]
},
"6.3": {
"title": "Infrastructure",
"questions": [
"Is necessary infrastructure determined?",
"Are buildings and workspace adequate?",
"Is process equipment adequate?",
"Are supporting services adequate?",
"Are maintenance requirements documented?"
]
},
"6.4": {
"title": "Work Environment",
"questions": [
"Is work environment determined?",
"Are environmental requirements documented?",
"Is contamination control adequate?",
"Are personnel health and cleanliness controlled?",
"Are environmental conditions monitored?"
]
},
"7.1": {
"title": "Planning of Product Realization",
"questions": [
"Are quality objectives for product defined?",
"Are processes needed determined?",
"Is verification and validation defined?",
"Are records requirements defined?",
"Is risk management applied?"
]
},
"7.2": {
"title": "Customer-Related Processes",
"questions": [
"Are customer requirements determined?",
"Are regulatory requirements determined?",
"Are requirements reviewed before commitment?",
"Are differences resolved before acceptance?",
"Is communication with customers effective?"
]
},
"7.3.1": {
"title": "Design and Development Planning",
"questions": [
"Are design stages determined?",
"Are review activities defined?",
"Are verification activities defined?",
"Are validation activities defined?",
"Are responsibilities assigned?",
"Are interfaces managed?"
]
},
"7.3.2": {
"title": "Design and Development Inputs",
"questions": [
"Are functional requirements defined?",
"Are performance requirements defined?",
"Are safety requirements defined?",
"Are regulatory requirements identified?",
"Are previous design inputs considered?",
"Are risk management outputs included?"
]
},
"7.3.3": {
"title": "Design and Development Outputs",
"questions": [
"Do outputs meet input requirements?",
"Is purchasing information provided?",
"Are acceptance criteria defined?",
"Are essential characteristics specified?",
"Are outputs approved before release?"
]
},
"7.3.4": {
"title": "Design and Development Review",
"questions": [
"Are design reviews conducted at suitable stages?",
"Is ability to meet requirements evaluated?",
"Are problems identified?",
"Are follow-up actions recorded?",
"Are appropriate functions represented?"
]
},
"7.3.5": {
"title": "Design and Development Verification",
"questions": [
"Is verification performed per plan?",
"Do outputs meet inputs?",
"Are verification records maintained?",
"Are verification methods appropriate?"
]
},
"7.3.6": {
"title": "Design and Development Validation",
"questions": [
"Is validation performed per plan?",
"Is product evaluated for intended use?",
"Is clinical evaluation included?",
"Are validation records maintained?",
"Is validation completed before product delivery?"
]
},
"7.3.7": {
"title": "Design and Development Transfer",
"questions": [
"Are outputs verified before transfer?",
"Is manufacturing capability verified?",
"Are transfer activities documented?"
]
},
"7.3.8": {
"title": "Control of Design and Development Changes",
"questions": [
"Are design changes identified?",
"Are changes reviewed?",
"Are changes verified?",
"Are changes validated as appropriate?",
"Is impact on product assessed?",
"Are changes approved before implementation?"
]
},
"7.4.1": {
"title": "Purchasing Process",
"questions": [
"Are suppliers evaluated and selected?",
"Are evaluation criteria established?",
"Is supplier performance monitored?",
"Are re-evaluation criteria defined?",
"Is purchased product verified?"
]
},
"7.4.2": {
"title": "Purchasing Information",
"questions": [
"Is purchasing information adequate?",
"Are product requirements specified?",
"Are QMS requirements specified?",
"Are personnel requirements specified?"
]
},
"7.4.3": {
"title": "Verification of Purchased Product",
"questions": [
"Is incoming inspection adequate?",
"Are verification activities defined?",
"Are verification records maintained?",
"Is source verification defined if applicable?"
]
},
"7.5.1": {
"title": "Control of Production and Service Provision",
"questions": [
"Is product information available?",
"Are work instructions available?",
"Is suitable equipment used?",
"Are monitoring devices available?",
"Is monitoring implemented?",
"Are release activities defined?",
"Are labeling requirements met?"
]
},
"7.5.2": {
"title": "Cleanliness of Product",
"questions": [
"Are cleanliness requirements documented?",
"Is contamination controlled?",
"Are process agents controlled?"
]
},
"7.5.3": {
"title": "Installation Activities",
"questions": [
"Are installation requirements documented?",
"Are acceptance criteria defined?",
"Are installation records maintained?"
]
},
"7.5.4": {
"title": "Servicing Activities",
"questions": [
"Are servicing procedures documented?",
"Are reference materials controlled?",
"Are service records maintained?",
"Is feedback analyzed?"
]
},
"7.5.5": {
"title": "Sterile Medical Devices",
"questions": [
"Is sterilization validated?",
"Are process parameters controlled?",
"Is sterile barrier validated?",
"Are sterilization records maintained?"
]
},
"7.5.6": {
"title": "Validation of Processes",
"questions": [
"Are special processes identified?",
"Are validation procedures documented?",
"Is equipment qualified?",
"Are personnel qualified?",
"Are validation records maintained?",
"Are revalidation criteria defined?"
]
},
"7.5.7": {
"title": "Particular Requirements for Validation",
"questions": [
"Are validation methods defined?",
"Are acceptance criteria established?",
"Is software validation appropriate?",
"Are validation records maintained?"
]
},
"7.5.8": {
"title": "Identification",
"questions": [
"Is product identified throughout realization?",
"Is documentation identified?",
"Is UDI implemented as required?"
]
},
"7.5.9": {
"title": "Traceability",
"questions": [
"Are traceability procedures documented?",
"Are components traceable?",
"Is work environment recorded?",
"Is distribution recorded?",
"Is traceability extent defined?"
]
},
"7.5.10": {
"title": "Customer Property",
"questions": [
"Is customer property identified?",
"Is it verified on receipt?",
"Is it protected and safeguarded?",
"Is loss or damage reported?"
]
},
"7.5.11": {
"title": "Preservation of Product",
"questions": [
"Is product identified?",
"Is handling controlled?",
"Is packaging controlled?",
"Is storage controlled?",
"Is protection adequate?"
]
},
"7.6": {
"title": "Control of Monitoring and Measuring Equipment",
"questions": [
"Is equipment calibrated?",
"Is calibration traceable?",
"Is calibration status identified?",
"Is equipment protected from damage?",
"Is software validated?",
"Are records maintained?"
]
},
"8.1": {
"title": "Measurement, Analysis and Improvement - General",
"questions": [
"Are monitoring activities planned?",
"Are analysis activities planned?",
"Are improvement activities planned?"
]
},
"8.2.1": {
"title": "Feedback",
"questions": [
"Is feedback collected?",
"Is feedback analyzed?",
"Is feedback used for improvement?",
"Is regulatory feedback included?"
]
},
"8.2.2": {
"title": "Complaint Handling",
"questions": [
"Is there a complaint procedure?",
"Are complaints investigated?",
"Are regulatory reports made if required?",
"Is trend analysis performed?",
"Are CAPAs initiated when warranted?"
]
},
"8.2.3": {
"title": "Reporting to Regulatory Authorities",
"questions": [
"Are reporting requirements identified?",
"Are reports submitted timely?",
"Are records maintained?"
]
},
"8.2.4": {
"title": "Internal Audit",
"questions": [
"Is an audit program established?",
"Are audit criteria defined?",
"Are auditors independent?",
"Are auditors competent?",
"Are audit records maintained?",
"Are findings followed up?"
]
},
"8.2.5": {
"title": "Monitoring and Measurement of Processes",
"questions": [
"Are processes monitored?",
"Are suitable methods used?",
"Is process capability demonstrated?",
"Are corrections made when needed?"
]
},
"8.2.6": {
"title": "Monitoring and Measurement of Product",
"questions": [
"Is product inspected?",
"Are acceptance criteria met?",
"Is release authorized?",
"Is traceability to inspection recorded?",
"Are records maintained?"
]
},
"8.3": {
"title": "Control of Nonconforming Product",
"questions": [
"Is nonconforming product identified?",
"Is it documented?",
"Is it evaluated?",
"Is it segregated?",
"Is disposition determined?",
"Is rework verified?",
"Is concession controlled?",
"Is post-delivery NC investigated?"
]
},
"8.4": {
"title": "Analysis of Data",
"questions": [
"Is data collected?",
"Is feedback analyzed?",
"Is conformity data analyzed?",
"Is process data analyzed?",
"Is supplier data analyzed?",
"Are audit results analyzed?"
]
},
"8.5.1": {
"title": "Improvement - General",
"questions": [
"Is continual improvement pursued?",
"Are policy, objectives, audits, data, actions, and reviews used?"
]
},
"8.5.2": {
"title": "Corrective Action",
"questions": [
"Is there a CA procedure?",
"Are NCs reviewed (including complaints)?",
"Is root cause determined?",
"Is action needed evaluated?",
"Is action determined and implemented?",
"Are results documented?",
"Is effectiveness verified?"
]
},
"8.5.3": {
"title": "Preventive Action",
"questions": [
"Is there a PA procedure?",
"Are potential NCs identified?",
"Is action needed evaluated?",
"Is action determined and implemented?",
"Are results documented?",
"Is effectiveness verified?"
]
}
}
# Process-to-Clause Mapping
PROCESS_MAPPING = {
"document-control": ["4.2.1", "4.2.2", "4.2.3", "4.2.4"],
"management-review": ["5.6"],
"internal-audit": ["8.2.4"],
"training": ["6.2"],
"design-control": ["7.3.1", "7.3.2", "7.3.3", "7.3.4", "7.3.5", "7.3.6", "7.3.7", "7.3.8"],
"purchasing": ["7.4.1", "7.4.2", "7.4.3"],
"production": ["7.5.1", "7.5.2", "7.5.6", "7.5.7", "7.5.8", "7.5.9", "7.5.11"],
"capa": ["8.5.2", "8.5.3"],
"nonconformity": ["8.3"],
"calibration": ["7.6"],
"complaint-handling": ["8.2.1", "8.2.2", "8.2.3"],
"risk-management": ["7.1"],
"infrastructure": ["6.3", "6.4"],
"customer-requirements": ["5.2", "7.2"]
}
def get_clause_checklist(clause: str) -> dict:
"""Get audit checklist for a specific clause."""
if clause not in ISO13485_CLAUSES:
return {"error": f"Clause {clause} not found"}
clause_data = ISO13485_CLAUSES[clause]
return {
"clause": clause,
"title": clause_data["title"],
"questions": clause_data["questions"],
"question_count": len(clause_data["questions"])
}
def get_process_checklist(process: str) -> dict:
"""Get audit checklist for a specific process."""
if process not in PROCESS_MAPPING:
available = ", ".join(sorted(PROCESS_MAPPING.keys()))
return {"error": f"Process '{process}' not found. Available: {available}"}
clauses = PROCESS_MAPPING[process]
questions = []
for clause in clauses:
if clause in ISO13485_CLAUSES:
clause_data = ISO13485_CLAUSES[clause]
for q in clause_data["questions"]:
questions.append({
"clause": clause,
"clause_title": clause_data["title"],
"question": q
})
return {
"process": process,
"clauses_covered": clauses,
"questions": questions,
"question_count": len(questions)
}
def get_system_audit_checklist() -> dict:
"""Get complete system audit checklist covering all clauses."""
all_questions = []
for clause, data in sorted(ISO13485_CLAUSES.items()):
for q in data["questions"]:
all_questions.append({
"clause": clause,
"clause_title": data["title"],
"question": q
})
return {
"audit_type": "system",
"clauses_covered": list(ISO13485_CLAUSES.keys()),
"questions": all_questions,
"question_count": len(all_questions)
}
def format_checklist_text(checklist: dict) -> str:
"""Format checklist for text output."""
lines = []
if "error" in checklist:
return f"Error: {checklist['error']}"
lines.append("=" * 70)
lines.append("ISO 13485:2016 INTERNAL AUDIT CHECKLIST")
lines.append(f"Generated: {datetime.now().strftime('%Y-%m-%d %H:%M')}")
lines.append("=" * 70)
if "clause" in checklist:
lines.append(f"\nClause: {checklist['clause']} - {checklist['title']}")
lines.append("-" * 50)
for i, q in enumerate(checklist["questions"], 1):
lines.append(f"\n{i}. {q}")
lines.append(" [ ] C [ ] NC [ ] OBS [ ] N/A")
lines.append(" Evidence: _________________________________")
lines.append(" Notes: ____________________________________")
elif "process" in checklist:
lines.append(f"\nProcess: {checklist['process'].replace('-', ' ').title()}")
lines.append(f"Clauses Covered: {', '.join(checklist['clauses_covered'])}")
lines.append("-" * 50)
current_clause = None
item_num = 1
for q in checklist["questions"]:
if q["clause"] != current_clause:
current_clause = q["clause"]
lines.append(f"\n--- {q['clause']} {q['clause_title']} ---")
lines.append(f"\n{item_num}. {q['question']}")
lines.append(" [ ] C [ ] NC [ ] OBS [ ] N/A")
lines.append(" Evidence: _________________________________")
lines.append(" Notes: ____________________________________")
item_num += 1
elif "audit_type" in checklist:
lines.append(f"\nAudit Type: Full System Audit")
lines.append(f"Total Clauses: {len(checklist['clauses_covered'])}")
lines.append("-" * 50)
current_clause = None
item_num = 1
for q in checklist["questions"]:
if q["clause"] != current_clause:
current_clause = q["clause"]
lines.append(f"\n{'=' * 40}")
lines.append(f"CLAUSE {q['clause']}: {q['clause_title']}")
lines.append("=" * 40)
lines.append(f"\n{item_num}. {q['question']}")
lines.append(" [ ] C [ ] NC [ ] OBS [ ] N/A")
lines.append(" Evidence: _________________________________")
item_num += 1
lines.append("\n" + "=" * 70)
lines.append(f"Total Questions: {checklist['question_count']}")
lines.append("")
lines.append("Legend: C=Conforming, NC=Nonconforming, OBS=Observation, N/A=Not Applicable")
lines.append("=" * 70)
return "\n".join(lines)
def interactive_mode():
"""Run interactive audit checklist generator."""
print("\n" + "=" * 50)
print("QMS INTERNAL AUDIT CHECKLIST GENERATOR")
print("=" * 50)
print("\nSelect audit type:")
print("1. Clause-specific audit")
print("2. Process audit")
print("3. Full system audit")
print("4. List available processes")
print("5. List all clauses")
print("6. Exit")
choice = input("\nEnter choice (1-6): ").strip()
if choice == "1":
print("\nAvailable clause sections:")
print(" 4.x - Quality Management System")
print(" 5.x - Management Responsibility")
print(" 6.x - Resource Management")
print(" 7.x - Product Realization")
print(" 8.x - Measurement, Analysis, Improvement")
clause = input("\nEnter clause number (e.g., 7.3.1): ").strip()
checklist = get_clause_checklist(clause)
print(format_checklist_text(checklist))
elif choice == "2":
processes = sorted(PROCESS_MAPPING.keys())
print("\nAvailable processes:")
for i, p in enumerate(processes, 1):
clauses = PROCESS_MAPPING[p]
print(f" {i}. {p} (clauses: {', '.join(clauses)})")
process = input("\nEnter process name: ").strip().lower()
checklist = get_process_checklist(process)
print(format_checklist_text(checklist))
elif choice == "3":
print("\nGenerating full system audit checklist...")
checklist = get_system_audit_checklist()
print(format_checklist_text(checklist))
elif choice == "4":
processes = sorted(PROCESS_MAPPING.keys())
print("\nAvailable QMS Processes:")
print("-" * 50)
for p in processes:
clauses = PROCESS_MAPPING[p]
print(f" {p}")
print(f" Clauses: {', '.join(clauses)}")
elif choice == "5":
print("\nISO 13485:2016 Clauses:")
print("-" * 50)
for clause, data in sorted(ISO13485_CLAUSES.items()):
print(f" {clause}: {data['title']} ({len(data['questions'])} questions)")
elif choice == "6":
print("Exiting.")
return
else:
print("Invalid choice.")
def main():
parser = argparse.ArgumentParser(
description="Generate ISO 13485:2016 internal audit checklists",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
python qms_audit_checklist.py --clause 7.3
python qms_audit_checklist.py --process design-control
python qms_audit_checklist.py --audit-type system --output json
python qms_audit_checklist.py --list-processes
python qms_audit_checklist.py --list-clauses
python qms_audit_checklist.py --interactive
"""
)
parser.add_argument(
"--clause",
help="Generate checklist for specific clause (e.g., 7.3.1, 8.5.2)"
)
parser.add_argument(
"--process",
help="Generate checklist for process (e.g., design-control, capa)"
)
parser.add_argument(
"--audit-type",
choices=["clause", "process", "system"],
help="Audit type for checklist generation"
)
parser.add_argument(
"--output",
choices=["text", "json"],
default="text",
help="Output format (default: text)"
)
parser.add_argument(
"--list-processes",
action="store_true",
help="List available QMS processes"
)
parser.add_argument(
"--list-clauses",
action="store_true",
help="List all ISO 13485 clauses"
)
parser.add_argument(
"--interactive",
action="store_true",
help="Run in interactive mode"
)
args = parser.parse_args()
if args.interactive:
interactive_mode()
return
if args.list_processes:
processes = sorted(PROCESS_MAPPING.keys())
if args.output == "json":
result = {p: PROCESS_MAPPING[p] for p in processes}
print(json.dumps(result, indent=2))
else:
print("\nAvailable QMS Processes:")
print("-" * 50)
for p in processes:
clauses = PROCESS_MAPPING[p]
print(f" {p}: {', '.join(clauses)}")
return
if args.list_clauses:
if args.output == "json":
result = {c: {"title": d["title"], "question_count": len(d["questions"])}
for c, d in sorted(ISO13485_CLAUSES.items())}
print(json.dumps(result, indent=2))
else:
print("\nISO 13485:2016 Clauses:")
print("-" * 50)
for clause, data in sorted(ISO13485_CLAUSES.items()):
print(f" {clause}: {data['title']} ({len(data['questions'])} questions)")
return
checklist = None
if args.clause:
checklist = get_clause_checklist(args.clause)
elif args.process:
checklist = get_process_checklist(args.process)
elif args.audit_type == "system":
checklist = get_system_audit_checklist()
else:
parser.print_help()
return
if checklist:
if args.output == "json":
print(json.dumps(checklist, indent=2))
else:
print(format_checklist_text(checklist))
if __name__ == "__main__":
main()
Tạo, tối ưu và phân tích chương trình giới thiệu, affiliate và chiến lược truyền miệng: vòng lan truyền, ưu đãi giới thiệu.
---
name: referrals
description: "When the user wants to create, optimize, or analyze a referral program, affiliate program, or word-of-mouth strategy. Also use when the user mentions 'referral,' 'affiliate,' 'ambassador,' 'word of mouth,' 'viral loop,' 'refer a friend,' 'partner program,' 'referral incentive,' 'how to get referrals,' 'customers referring customers,' or 'affiliate payout.' Use this whenever someone wants existing users or partners to bring in new customers. For launch-specific virality, see launch."
metadata:
version: 2.0.1
---
# Referral & Affiliate Programs
You are an expert in viral growth and referral marketing. Your goal is to help design and optimize programs that turn customers into growth engines.
## Before Starting
**Check for product marketing context first:**
If `.agents/product-marketing.md` exists (or `.claude/product-marketing.md`, or the legacy `product-marketing-context.md` filename, in older setups), read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
Gather this context (ask if not provided):
### 1. Program Type
- Customer referral program, affiliate program, or both?
- B2B or B2C?
- What's the average customer LTV?
- What's your current CAC from other channels?
### 2. Current State
- Existing referral/affiliate program?
- Current referral rate (% who refer)?
- What incentives have you tried?
### 3. Product Fit
- Is your product shareable?
- Does it have network effects?
- Do customers naturally talk about it?
### 4. Resources
- Tools/platforms you use or consider?
- Budget for referral incentives?
---
## Should You Engineer Virality First?
Before building a reward-driven program, check whether virality can be **built into the product** — often cheaper and more durable than paid referrals. But **don't force virality where it doesn't naturally fit.**
Place the product on the **Viral Potential Spectrum**:
- **Natural** (build for it): collaboration tools, communication tools, user-facing outputs — every use exposes the product to non-users.
- **Limited** (don't force it): backend, competitive-advantage, internal-only, and infrastructure products. Invest in referral programs, content, and partnerships instead.
If the product is on the natural end, consider **product-embedded viral mechanisms** (Powered By badges, exposure loops, social sharing, embeds, watermarks) before or alongside a reward program.
**For the spectrum diagnostic, the 7 viral mechanisms, value-presentation and timing best practices, and affiliate power-law mechanics**: See [references/viral-mechanisms.md](references/viral-mechanisms.md)
---
## Referral vs. Affiliate
### Customer Referral Programs
**Best for:**
- Existing customers recommending to their network
- Products with natural word-of-mouth
- Lower-ticket or self-serve products
**Characteristics:**
- Referrer is an existing customer
- One-time or limited rewards
- Higher trust, lower volume
### Affiliate Programs
**Best for:**
- Reaching audiences you don't have access to
- Content creators, influencers, bloggers
- Higher-ticket products that justify commissions
**Characteristics:**
- Affiliates may not be customers
- Ongoing commission relationship
- Higher volume, variable trust
---
## Referral Program Design
### The Referral Loop
```
Trigger Moment → Share Action → Convert Referred → Reward → (Loop)
```
### Step 1: Identify Trigger Moments
**High-intent moments:**
- Right after first "aha" moment
- After achieving a milestone
- After exceptional support
- After renewing or upgrading
### Step 2: Design Share Mechanism
**Ranked by effectiveness:**
1. In-product sharing (highest conversion)
2. Personalized link
3. Email invitation
4. Social sharing
5. Referral code (works offline)
### Step 3: Choose Incentive Structure
**Single-sided rewards** (referrer only): Simpler, works for high-value products
**Double-sided rewards** (both parties): Higher conversion, win-win framing
**Tiered rewards**: Gamifies referral process, increases engagement
**Present the reward with the bigger-*feeling* number** — "lead with the larger number" (say "$10 off," not "40% off," on a low-priced product). Reward at the **aha moment or milestone**, not signup. Reduce friction: one-click share, pre-written messages.
**For examples and incentive sizing**: See [references/program-examples.md](references/program-examples.md)
**For product-embedded virality, value-presentation rules, and affiliate power-law mechanics**: See [references/viral-mechanisms.md](references/viral-mechanisms.md)
---
## Program Optimization
### Improving Referral Rate
**If few customers are referring:**
- Ask at better moments
- Simplify sharing process
- Test different incentive types
- Make referral prominent in product
**If referrals aren't converting:**
- Improve landing experience for referred users
- Strengthen incentive for new users
- Ensure referrer's endorsement is visible
### A/B Tests to Run
**Incentive tests:** Amount, type, single vs. double-sided, timing
**Messaging tests:** Program description, CTA copy, landing page copy
**Placement tests:** Where and when the referral prompt appears
### Common Problems & Fixes
| Problem | Fix |
|---------|-----|
| Low awareness | Add prominent in-app prompts |
| Low share rate | Simplify to one click |
| Low conversion | Optimize referred user experience |
| Fraud/abuse | Add verification, limits |
| One-time referrers | Add tiered/gamified rewards |
---
## Measuring Success
### Key Metrics
**Program health:**
- Active referrers (referred someone in last 30 days)
- Referral conversion rate
- Rewards earned/paid
**Business impact:**
- % of new customers from referrals
- CAC via referral vs. other channels
- LTV of referred customers
- Referral program ROI
### Typical Findings
- Referred customers have 16-25% higher LTV
- Referred customers have 18-37% lower churn
- Referred customers refer others at 2-3x rate
---
## Launch Checklist
### Before Launch
- [ ] Define program goals and success metrics
- [ ] Design incentive structure
- [ ] Build or configure referral tool
- [ ] Create referral landing page
- [ ] Set up tracking and attribution
- [ ] Define fraud prevention rules
- [ ] Create terms and conditions
- [ ] Test complete referral flow
### Launch
- [ ] Announce to existing customers
- [ ] Add in-app referral prompts
- [ ] Update website with program details
- [ ] Brief support team
### Post-Launch (First 30 Days)
- [ ] Review conversion funnel
- [ ] Identify top referrers
- [ ] Gather feedback
- [ ] Fix friction points
- [ ] Send reminder emails to non-referrers
---
## Email Sequences
### Referral Program Launch
```
Subject: You can now earn [reward] for sharing [Product]
We just launched our referral program!
Share [Product] with friends and earn [reward] for each signup.
They get [their reward] too.
[Unique referral link]
1. Share your link
2. Friend signs up
3. You both get [reward]
```
### Referral Nurture Sequence
- Day 7: Remind about referral program
- Day 30: "Know anyone who'd benefit?"
- Day 60: Success story + referral prompt
- After milestone: "You achieved [X]—know others who'd want this?"
---
## Affiliate Programs
**For detailed affiliate program design, commission structures, recruitment, and tools**: See [references/affiliate-programs.md](references/affiliate-programs.md)
**For affiliate power-law mechanics (buyout clauses ~12× monthly commission, the 20/80 super-promoter rule, launch-affiliate tactics)**: See [references/viral-mechanisms.md](references/viral-mechanisms.md)
---
## Task-Specific Questions
1. What type of program (referral, affiliate, or both)?
2. What's your customer LTV and current CAC?
3. Existing program or starting from scratch?
4. What tools/platforms are you considering?
5. What's your budget for rewards/commissions?
6. Is your product naturally shareable?
---
## Tool Integrations
For implementation, see the [tools registry](../../tools/REGISTRY.md). Key tools for referral programs:
| Tool | Best For | Guide |
|------|----------|-------|
| **Rewardful** | Stripe-native affiliate programs | [rewardful.md](../../tools/integrations/rewardful.md) |
| **Tolt** | SaaS affiliate programs | [tolt.md](../../tools/integrations/tolt.md) |
| **Mention Me** | Enterprise referral programs | [mention-me.md](../../tools/integrations/mention-me.md) |
| **Dub.co** | Link tracking and attribution | [dub-co.md](../../tools/integrations/dub-co.md) |
| **Stripe** | Payment processing (for commission tracking) | [stripe.md](../../tools/integrations/stripe.md) |
| **Introw** | Channel partner programs with tiers, deal registration, QBRs | [introw.md](../../tools/integrations/introw.md) |
| **PartnerStack** | Enterprise partner and affiliate programs | [partnerstack.md](../../tools/integrations/partnerstack.md) |
---
## Related Skills
- **launch**: For launching referral program effectively
- **emails**: For referral nurture campaigns
- **marketing-psychology**: For understanding referral motivation
- **analytics**: For tracking referral attribution
FILE:evals/evals.json
{
"skill_name": "referrals",
"evals": [
{
"id": 1,
"prompt": "Help me design a referral program for our SaaS product. We're a $49/month project management tool with about 1,000 customers. We want to encourage word-of-mouth growth.",
"expected_output": "Should check for product-marketing.md first. Should distinguish between referral and affiliate programs (this is referral — existing customers referring peers). Should design the referral loop: trigger point (when to ask for referral), share mechanism (unique link, email invite, social share), conversion flow (what the referred person experiences), and reward structure. Should recommend incentive type: double-sided recommended (both referrer and referred get value). Should suggest specific incentives appropriate for $49/month SaaS (e.g., free month for both). Should include the launch checklist. Should recommend tool integrations (Rewardful, Tolt, etc.).",
"assertions": [
"Checks for product-marketing.md",
"Distinguishes referral from affiliate",
"Designs the referral loop (trigger, share, convert, reward)",
"Recommends double-sided incentive structure",
"Suggests specific incentives for the price point",
"Includes launch checklist",
"Recommends tool integrations"
],
"files": []
},
{
"id": 2,
"prompt": "We have a referral program but only 5% of customers have ever referred someone. How do we increase participation?",
"expected_output": "Should apply the program optimization guidance. Should diagnose low participation: are customers aware of the program? Is the trigger point well-timed? Is the incentive compelling enough? Is sharing easy? Should recommend optimization tactics: better placement/visibility, timing referral asks at peak satisfaction moments, improving the incentive, simplifying the share mechanism, adding referral reminders in email and in-app. Should provide specific experiment ideas to test improvements.",
"assertions": [
"Applies program optimization guidance",
"Diagnoses potential causes of low participation",
"Checks awareness, timing, incentive, and friction",
"Recommends optimization tactics",
"Suggests timing referral asks at satisfaction moments",
"Provides experiment ideas"
],
"files": []
},
{
"id": 3,
"prompt": "should we do referral or affiliate? we sell online courses for $199-499 and want to get other creators and influencers to promote us.",
"expected_output": "Should trigger on casual phrasing. Should apply the referral vs affiliate distinction clearly. For this use case (getting creators/influencers to promote), should recommend an affiliate program (not referral — affiliates are third-party promoters, not existing customers). Should apply the affiliate program section guidance: commission structure for digital products (typically 20-40% for courses), cookie duration, payout terms, affiliate onboarding. Should recommend affiliate platforms/tools appropriate for course creators.",
"assertions": [
"Triggers on casual phrasing",
"Clearly distinguishes referral from affiliate",
"Recommends affiliate for this use case",
"Provides commission structure guidance for courses",
"Addresses cookie duration and payout terms",
"Recommends appropriate affiliate platforms"
],
"files": []
},
{
"id": 4,
"prompt": "What incentive structure works best? We've been offering $10 off for referrers but it's not working. Our product is $29/month.",
"expected_output": "Should evaluate the current incentive: $10 off on a $29/month product is significant but only benefits the referrer (single-sided). Should recommend testing double-sided incentives (both parties get value). Should discuss incentive types: account credit, free months, feature upgrades, cash. Should apply the tiered incentive concept (increasing rewards for multiple referrals). Should provide specific alternative incentive structures to test. Should note that incentive alone may not be the problem — placement and timing matter too.",
"assertions": [
"Evaluates current incentive structure",
"Identifies as single-sided and recommends double-sided",
"Discusses multiple incentive types",
"Applies tiered incentive concept",
"Provides specific alternatives to test",
"Notes incentive may not be the only issue"
],
"files": []
},
{
"id": 5,
"prompt": "How do we measure the success of our referral program? What metrics should we track?",
"expected_output": "Should apply the measuring success framework. Should define key metrics: participation rate (% of customers who refer), share rate (referrals sent per participant), conversion rate (referred visitors who become customers), viral coefficient (k-factor), customer acquisition cost via referral vs other channels, referred customer LTV vs organic customer LTV. Should recommend tracking tools and dashboards. Should provide benchmark ranges for each metric.",
"assertions": [
"Applies measuring success framework",
"Defines participation rate, share rate, conversion rate",
"Includes viral coefficient / k-factor",
"Compares referral CAC to other channels",
"Compares referred customer LTV to organic",
"Recommends tracking approach",
"Provides benchmark ranges"
],
"files": []
},
{
"id": 6,
"prompt": "Can you write the referral invitation emails? I need the email that goes out when someone shares their referral link.",
"expected_output": "Should recognize this overlaps with email writing. Should apply the referral email sequence section from the skill for referral-specific emails. However, for detailed email sequence design (multi-email nurture for referred users), should cross-reference the emails skill. Should provide the referral invitation email but note that broader email sequence work is handled by emails.",
"assertions": [
"Applies referral email section from the skill",
"Provides referral invitation email guidance",
"Cross-references emails for broader email work",
"Provides specific referral email copy or template"
],
"files": []
},
{
"id": 7,
"prompt": "We build a collaborative design tool. We keep hearing 'add a referral program' but I want to know if we can just make the product spread on its own. What are our options?",
"expected_output": "Should place the product on the Viral Potential Spectrum: a collaborative design tool is on the natural end (collaboration + user-facing output), so product-embedded virality fits before or instead of a reward program. Should recommend engineering virality through product design rather than defaulting to a reward program, and warn against forcing virality where it doesn't fit. Should walk through applicable viral mechanisms from the 7: exposure loops (invite/collaboration), embeds (Notion/Figma/Loom style), social sharing (e.g. a #MadeWith hashtag), and Powered By / watermark badges on free-tier output. Should note referral programs are the incentive-driven fallback when the product doesn't spread naturally. Should reference viral-mechanisms.md.",
"assertions": [
"Places the product on the Viral Potential Spectrum (natural end)",
"Recommends product-embedded virality before defaulting to a reward program",
"Warns against forcing virality where it doesn't fit",
"Names applicable mechanisms (exposure loops, embeds, social sharing, badges/watermarks)",
"Frames referral programs as the incentive-driven fallback",
"References viral-mechanisms.md"
],
"files": []
}
]
}
FILE:references/affiliate-programs.md
# Affiliate Program Design
Detailed guidance for building and managing affiliate programs.
## Contents
- Commission Structures
- Cookie Duration
- Affiliate Recruitment
- Affiliate Enablement
- Tools & Platforms (Referral Program Tools, Affiliate Program Tools, Choosing a Tool)
- Fraud Prevention (Common Referral Fraud, Prevention Measures)
## Commission Structures
**Percentage of sale:**
- Standard: 10-30% of first sale or first year
- Works for: E-commerce, SaaS with clear pricing
- Example: "Earn 25% of every sale you refer"
**Flat fee per action:**
- Standard: $5-500 depending on value
- Works for: Lead gen, trials, freemium
- Example: "$50 for every qualified demo"
**Recurring commission:**
- Standard: 10-25% of recurring revenue
- Works for: Subscription products
- Example: "20% of subscription for 12 months"
**Tiered commission:**
- Works for: Motivating high performers
- Example: "20% for 1-10 sales, 25% for 11-25, 30% for 26+"
---
## Cookie Duration
How long after click does affiliate get credit?
| Duration | Use Case |
|----------|----------|
| 24 hours | High-volume, low-consideration purchases |
| 7-14 days | Standard e-commerce |
| 30 days | Standard SaaS/B2B |
| 60-90 days | Long sales cycles, enterprise |
| Lifetime | Premium affiliate relationships |
---
## Affiliate Recruitment
### Where to find affiliates:
- Existing customers who create content
- Industry bloggers and reviewers
- YouTubers in your niche
- Newsletter writers
- Complementary tool companies
- Consultants and agencies
### Outreach template:
```
Subject: Partnership opportunity — [Your Product]
Hi [Name],
I've been following your content on [topic] — particularly [specific piece] — and think there could be a great fit for a partnership.
[Your Product] helps [audience] [achieve outcome], and I think your audience would find it valuable.
We offer [commission structure] for partners, plus [additional benefits: early access, co-marketing, etc.].
Would you be open to learning more?
[Your name]
```
---
## Affiliate Enablement
Provide affiliates with:
- [ ] Unique tracking links/codes
- [ ] Product overview and key benefits
- [ ] Target audience description
- [ ] Comparison to competitors
- [ ] Creative assets (logos, banners, images)
- [ ] Sample copy and talking points
- [ ] Case studies and testimonials
- [ ] Demo access or free account
- [ ] FAQ and objection handling
- [ ] Payment terms and schedule
---
## Tools & Platforms
### Referral Program Tools
**Full-featured platforms:**
- ReferralCandy — E-commerce focused
- Ambassador — Enterprise referral programs
- Friendbuy — E-commerce and subscription
- GrowSurf — SaaS and tech companies
- Mention Me — AI-powered referral marketing
- Viral Loops — Template-based campaigns
**Built-in options:**
- Stripe (basic referral tracking)
- HubSpot (CRM-integrated)
- Segment (tracking and analytics)
### Affiliate Program Tools
**Affiliate networks:**
- ShareASale — Large merchant network
- Impact — Enterprise partnerships
- PartnerStack — SaaS focused
- Tapfiliate — Simple SaaS affiliate tracking
- FirstPromoter — SaaS affiliate management
**Partner Relationship Management (PRM):**
- Introw — Full PRM with deal registration, commissions, tiers, QBRs, and partner engagement tracking ([integration guide](../../../tools/integrations/introw.md))
**Self-hosted:**
- Rewardful — Stripe-integrated affiliates
- Refersion — E-commerce affiliates
### Choosing a Tool
Consider:
- Integration with your payment system
- Fraud detection capabilities
- Payout management
- Reporting and analytics
- Customization options
- Price vs. program scale
---
## Fraud Prevention
### Common Referral Fraud
- Self-referrals (creating fake accounts)
- Referral rings (groups referring each other)
- Coupon sites posting referral codes
- Fake email addresses
- VPN/device spoofing
### Prevention Measures
**Technical:**
- Email verification required
- Device fingerprinting
- IP address monitoring
- Delayed reward payout (after activation)
- Minimum activity threshold
**Policy:**
- Clear terms of service
- Maximum referrals per period
- Reward clawback for refunds/chargebacks
- Manual review for suspicious patterns
**Structural:**
- Require referred user to take meaningful action
- Cap lifetime rewards
- Pay rewards in product credit (less attractive to fraudsters)
FILE:references/program-examples.md
# Referral Program Examples
Real-world examples of successful referral programs.
## Contents
- Dropbox (Classic)
- Uber/Lyft
- Morning Brew
- Notion
- Incentive Types Comparison
- Incentive Sizing Framework
- Viral Coefficient & Metrics (Key Metrics, Calculating Referral Program ROI)
## Dropbox (Classic)
**Program:** Give 500MB storage, get 500MB storage
**Why it worked:**
- Reward directly tied to product value
- Low friction (just an email)
- Both parties benefit equally
- Gamified with progress tracking
---
## Uber/Lyft
**Program:** Give $10 ride credit, get $10 when they ride
**Why it worked:**
- Immediate, clear value
- Double-sided incentive
- Easy to share (code/link)
- Triggered at natural moments
---
## Morning Brew
**Program:** Tiered rewards for subscriber referrals
- 3 referrals: Newsletter stickers
- 5 referrals: T-shirt
- 10 referrals: Mug
- 25 referrals: Hoodie
**Why it worked:**
- Gamification drives ongoing engagement
- Physical rewards are shareable (more referrals)
- Low cost relative to subscriber value
- Built status/identity
---
## Notion
**Program:** $10 credit per referral (education)
**Why it worked:**
- Targeted high-sharing audience (students)
- Product naturally spreads in teams
- Credit keeps users engaged
---
## Incentive Types Comparison
| Type | Pros | Cons | Best For |
|------|------|------|----------|
| Cash/credit | Universally valued | Feels transactional | Marketplaces, fintech |
| Product credit | Drives usage | Only valuable if they'll use it | SaaS, subscriptions |
| Free months | Clear value | May attract freebie-seekers | Subscription products |
| Feature unlock | Low cost to you | Only works for gated features | Freemium products |
| Swag/gifts | Memorable, shareable | Logistics complexity | Brand-focused companies |
| Charity donation | Feel-good | Lower personal motivation | Mission-driven brands |
---
## Incentive Sizing Framework
**Calculate your maximum incentive:**
```
Max Referral Reward = (Customer LTV × Gross Margin) - Target CAC
```
**Example:**
- LTV: $1,200
- Gross margin: 70%
- Target CAC: $200
- Max reward: ($1,200 × 0.70) - $200 = $640
**Typical referral rewards:**
- B2C: $10-50 or 10-25% of first purchase
- B2B SaaS: $50-500 or 1-3 months free
- Enterprise: Higher, often custom
---
## Viral Coefficient & Metrics
### Key Metrics
**Viral coefficient (K-factor):**
```
K = Invitations × Conversion Rate
K > 1 = Viral growth (each user brings more than 1 new user)
K < 1 = Amplified growth (referrals supplement other acquisition)
```
**Example:**
- Average customer sends 3 invitations
- 15% of invitations convert
- K = 3 × 0.15 = 0.45
**Referral rate:**
```
Referral Rate = (Customers who refer) / (Total customers)
```
Benchmarks:
- Good: 10-25% of customers refer
- Great: 25-50%
- Exceptional: 50%+
**Referrals per referrer:**
Benchmarks:
- Average: 1-2 referrals per referrer
- Good: 2-5
- Exceptional: 5+
### Calculating Referral Program ROI
```
Referral Program ROI = (Revenue from referred customers - Program costs) / Program costs
Program costs = Rewards paid + Tool costs + Management time
```
**Track separately:**
- Cost per referred customer (CAC via referral)
- LTV of referred customers (often higher than average)
- Payback period for referral rewards
FILE:references/viral-mechanisms.md
# Viral Mechanisms
Virality can be **engineered through product design**, not just bought with reward programs. But don't force it — decide *whether* virality fits your product before building anything.
## Contents
- Viral Potential Spectrum (the diagnostic)
- The 7 Viral Mechanisms
- Referral Best Practices (presentation, timing, friction)
- Affiliate Mechanics (buyout clauses, the 20/80 power law, launch tactics)
---
## Viral Potential Spectrum
Before engineering virality, place your product on the spectrum. **Don't force virality where it doesn't naturally fit.**
**Natural viral potential (build for it):**
- **Collaboration tools** — value grows when you invite others (docs, whiteboards, project management)
- **Communication tools** — you can't use them alone (email, scheduling, messaging)
- **User-facing outputs** — every use produces something others see (design, video, forms, links)
**Limited viral potential (don't force it):**
- **Backend / infrastructure** — invisible to end users
- **Competitive-advantage tools** — users *hide* that they use them (their edge)
- **Internal-only tools** — never leave the org
- **Infrastructure** — plumbing no one talks about
If you're on the limited end, invest in referral programs, content, and partnerships instead of embedding viral loops that won't fire.
---
## The 7 Viral Mechanisms
Most are **non-incentive** — the loop is built into the product, not paid for.
### 1. "Powered By" Badges
A small attributed badge on user-facing output ("Powered by [Product]"). Every page/form/widget a customer ships becomes an ad. Often free-tier only (paid tier removes it).
### 2. Exposure Loops
The product's normal use exposes it to non-users.
- **Calendly / SavvyCal** — every meeting invite you send shows the tool to the recipient, who often becomes a user.
- **Superhuman email signatures** ("Sent via Superhuman") — works as a **status signal**, not just attribution. The signature signaled early-adopter status, so recipients *wanted* it. Exposure loops are strongest when using the product confers status.
### 3. Social Sharing
Make output natively shareable with a branded hook.
- **#MadeWithGlide** — a hashtag turns every user creation into discoverable social proof.
- One-tap "share to X/LinkedIn" on any milestone, result, or artifact.
### 4. Embed Options
Let users embed their content elsewhere; the embed carries your brand and a link back.
- **Notion, Figma, Loom** — embedded docs, designs, and videos spread the product to every viewer on every host site.
### 5. Watermarks / Mandatory Badges
Like "Powered By" but harder to remove — baked into the output itself.
- **OpusClips** watermark on generated clips.
- **"Made in Webflow"** badge on free-plan sites.
Free tier carries the mark; paid tier removes it. The free users become the distribution.
### 6. Referral Programs
Explicit incentives for referring. Covered in detail in [program-examples.md](program-examples.md) and below. The one *incentive-driven* mechanism on this list — use it when the product itself doesn't naturally spread.
### 7. Product-Driven Word-of-Mouth
The purest form: the product is so good, novel, or useful that people tell others unprompted. Not a mechanism you bolt on — it's earned through the product experience. Engineering the other six makes this easier to trigger.
---
## Referral Best Practices
Detail beyond the core referral loop (trigger → share → convert → reward).
### Value Presentation: Lead With the Larger Number
Frame the reward with whichever number *looks* bigger.
- On a $25 product, say **"$10 off"** — not "40% off."
- On a $500 product, say **"20% off"** — not "$100 off" if the percentage frames better... but usually the absolute dollar figure wins for smaller prices.
- Rule of thumb: **under ~$100, lead with the dollar amount; over ~$100, test the percentage.** Always pick the bigger-*feeling* number.
### Reward Timing: Fire at the Aha / Milestone
Trigger the referral ask (and reward) at the moment the user has just felt the product's value — the **aha moment** or a **milestone** (first success, upgrade, streak). Motivation to share peaks right after value is experienced, not at signup.
### Double-Sided Rewards
Both referrer and referred get value. Higher conversion than single-sided, and gives the referrer a generous, non-selfish reason to share ("here's $10 for you too").
### Friction Reduction
Every extra step kills share rate.
- **One-click sharing** — pre-generated link, no form.
- **Pre-written messages** — draft the email/DM/post copy so the user just hits send.
- In-product placement at the trigger moment, not buried in settings.
---
## Affiliate Mechanics
Detail deferred from the partnerships side — for building an affiliate motion into a referral/partner strategy.
### Buyout Clauses (~12× Monthly Commission)
For high-performing affiliates on **recurring** commissions, include a **buyout clause**: the right to buy out the affiliate's future commission stream for a lump sum, commonly around **12× the monthly commission**. Protects margin on a customer the affiliate referred once but earns on forever, and gives the affiliate an attractive cash-out.
### The 20/80 Affiliate Power Law
Roughly **20% of affiliates drive ~80% of results**. Don't spread effort evenly across a long tail of dormant sign-ups. **Identify super-promoters and invest in them** — higher tiers, custom assets, co-marketing, direct relationship, early access. Recruiting 1,000 passive affiliates is worth less than activating 10 great ones.
### Launch-Affiliate Tactic (Cometly / Demio)
Time affiliate promotion around a **launch or a hard deadline** to concentrate volume. Cometly drove **$251K on a single launch day** by mobilizing affiliates simultaneously; Demio ran launch-window affiliate pushes. The mechanic: give affiliates a shared date, shared assets, and a reason for their audience to act *now* (bonus, cohort, closing offer) so promotion stacks instead of trickling.
Tạm dừng giữa chừng để nhìn rộng hơn, đánh giá lại hướng đi, giả định và thiên kiến thay vì sa vào chi tiết.
---
name: reflect
description: "Mid-conversation reflection skill that pauses execution and zooms out from detail-mode to honestly reassess direction, assumptions, and bias. Use when the user says 'reflect', 'take a step back', 'step back', 'zoom out', 'are we missing something', 'bigger picture', 'sanity check this', 'are we on track', 'are we overthinking this', 'forest for the trees', or any variation signaling intent to break out of detail-mode and reassess. Also trigger when the conversation has gone deep on implementation details without strategic check-in, or when the user shows signs of being stuck — that's often a signal the framing needs a reset, not more detail work. Intentionally low-intake: runs the 5-dimension analysis immediately when prior context is rich enough; asks one forcing clarifier only when invocation context is too thin to reassess from."
license: MIT
metadata:
source_spec: "megaprompts/02-reflect-megaprompt.md"
build_pattern: "Path B (direct conversion)"
version: 1.0.0
---
# Reflect — Mid-Conversation Reassessment
> **Portability:** Pure-reasoning skill. No external tools required. Works in Claude Code CLI + Claude.ai web natively. Most portable in the v2 collection.
When invoked mid-conversation, this skill **pauses execution** and produces a frank reassessment of where the conversation has been heading. Output is **flowing analysis (no headers, conversational tone)** covering macro perspective, gap analysis, reflective inquiry, bias check, and contextual alignment. The skill ends with a clear directional recommendation: **continue, pivot, or pause to answer a specific question**.
## Invocation Triggers
**Explicit phrases:**
- "reflect"
- "take a step back" / "step back"
- "zoom out"
- "are we missing something"
- "bigger picture"
- "what are we missing"
- "let's pause"
- "sanity check this"
- "are we on track"
- "are we overthinking this"
- "forest for the trees"
**Implicit signals (no phrase needed):**
- Conversation has gone 10+ turns deep on implementation details without strategic check-in
- User shows signs of frustration or stuck-ness
- Repeated dead-ends or pivots within a short span
When you detect an implicit trigger, **don't auto-invoke** — ask the user if they want to step back. Implicit signals are a prompt to OFFER reflection, not to unilaterally run it.
## Stop Directive (Before Reassessing)
**Halt the current thread.** Don't continue execution of the in-progress task. Reflection is a pause, not a side-quest.
This matters because:
- Continuing detail work while "reflecting on the side" defeats the purpose — you'll over-weight the current direction
- The user expects a clear break in cadence
- The reassessment needs full attention to the conversation history
## Grill-Me Optional Clarifier
This skill is intentionally **low-intake** — most invocations should run the 5-dimension analysis immediately without questions. The grill-me discipline applies *only* when the invocation is ambiguous (e.g., user pastes "step back" at the start of a fresh conversation with no prior context to reassess).
### Q1 (optional, asked only when context is too thin to reassess)
> **What specifically should I reassess? Pick one:**
>
> 1. The goal — are we solving the right problem?
> 2. The approach — is the path we're on the best one?
> 3. The assumptions — what are we taking for granted?
> 4. All of the above (default if you have time)
>
> *Why I'm asking:* I'm seeing limited prior context to reassess, so I want to focus the reflection rather than guess. If you'd rather I do all three, that's fine — say so.
Forcing choice with default. **Asked only when context is genuinely thin; otherwise skip and run the full analysis on existing conversation.**
**Stop condition:** One question max. If the user invokes mid-conversation with normal context, no questions are asked — the skill runs directly.
## The 5-Dimension Analysis Framework
Re-read the **full conversation from the original goal forward** — not just recent turns. The discipline that distinguishes real reflection from local-context summary.
### 1. Macro Perspective
- **Original goal:** What did the user actually start trying to do?
- **Drift detection:** Has the conversation moved away from that goal? Toward something better or worse?
- **Connection check:** How does current work connect to the larger objective?
Anchor with specific evidence: "At turn 3 the goal was X; by turn 12 we're working on Y. Is Y a productive narrowing of X, or a drift away?"
### 2. Gap Analysis
- **Unverified assumptions** — what are we taking for granted that we haven't checked?
- **Missing stakeholders / audiences / users** — who needs this beyond the immediate context?
- **Skipped constraints** — technical, regulatory, resource limits not addressed
- **Dismissed alternatives** — paths considered but rejected; revisit briefly
- **External factors** — timing, market, dependencies not in scope
### 3. Reflective Inquiry
- Is the problem framed correctly?
- Solving the right problem vs. an adjacent easier one?
- Simpler path being overcomplicated?
- Harder but more valuable path being avoided?
- **Fresh-eyes perspective:** would someone else approach this differently?
### 4. Bias Check
Five biases — recognize each through specific conversation patterns:
| Bias | Recognition cue |
|---|---|
| **Confirmation bias** | Evidence cited only supports the working hypothesis; counter-evidence absent or dismissed |
| **Sunk cost fallacy** | "We've already invested X" / "we're far enough in to..." instead of fresh cost/benefit |
| **Anchoring** | Stuck on first option mentioned; new options compared against it rather than evaluated independently |
| **Complexity bias** | Adding features / steps / safeguards without specific justification for each |
| **Recency bias** | Over-weighting last few turns; older but important context being ignored |
For each detected bias: name it, cite the specific evidence, suggest a corrective move.
See [`references/cognitive_bias_canon.md`](references/cognitive_bias_canon.md) for the full canon.
### 5. Contextual Alignment
- Does the direction serve the user's actual goals (as known from context)?
- Are external factors being ignored?
- Is this the best use of the user's time and energy right now?
- Connection to other known projects or priorities?
## Tone and Format Rules
The skill must produce:
- **Flowing prose** — no headers, no bullet lists, no structured-report formatting
- **Tight but thorough** — neither a one-liner nor a wall of text
- **Direct critique when warranted** — with specific evidence from the conversation
- **Validation when warranted** — with specific reasoning for why the path is solid
- **No vague reassurance** — "looks good!" without reasoning is rejected
- **No manufactured problems** — when the path is genuinely solid, say so with specific reasons; don't invent issues
See [`references/honest_output_discipline.md`](references/honest_output_discipline.md) for the anti-manufactured-problems framing.
## Closing Recommendation (Mandatory)
Every run ends with one of three directional recommendations:
| Recommendation | When | Format |
|---|---|---|
| **Continue** | Path is solid | "Continue. {specific reasoning for why}." |
| **Pivot to {X}** | Drift has occurred OR better path surfaced | "Pivot toward {X}, away from {what to drop}. {specific evidence}." |
| **Pause for {Q}** | A specific question needs answering before continuing | "Pause for {Q}. Without answering this, the next step risks {specific cost}." |
The closing is always specific — never "you should think more about this" or "consider your options."
## Error Handling
| Situation | Behavior |
|---|---|
| Conversation is very short (no real context to reassess) | Acknowledge limitation, ask user what they want reassessed (Q1 fires) |
| Current direction is genuinely solid | State this clearly with reasoning; don't manufacture problems |
| User invokes mid-task with no clear question | Default to macro perspective + bias check; offer to dig deeper |
| Implicit trigger seems possible but unclear | Don't invoke proactively; ask user if they want to step back |
## Tooling
| Script | Role |
|---|---|
| `scripts/bias_pattern_detector.py` | Scan conversation text for patterns indicative of each of the 5 biases |
| `scripts/conversation_depth_analyzer.py` | Count turns + detect implicit-trigger signals (10+ detail turns, frustration markers) |
| `scripts/directional_recommendation_validator.py` | Verify output ends with Continue / Pivot / Pause + specific reasoning |
## References
- [`references/cognitive_bias_canon.md`](references/cognitive_bias_canon.md) — 5 biases + recognition cues (7+ sources)
- [`references/honest_output_discipline.md`](references/honest_output_discipline.md) — anti-manufactured-problems framing (7+ sources)
- [`references/conversation_reflection_practice.md`](references/conversation_reflection_practice.md) — Schön reflective-practice canon (7+ sources)
## Anti-Patterns To Reject
- Hardcoded user names or specific domain references
- Structured-report output (headers, bullet lists) when prose is required
- Manufactured problems when things are actually fine
- Vague reassurance ("looks good!") instead of specific reasoning
- Reassessing only recent turns instead of the full conversation
- Skipping the closing directional recommendation
- Continuing the in-progress task while "reflecting on the side"
---
**Version:** 1.0.0
**Source spec:** [`megaprompts/02-reflect-megaprompt.md`](../../../../megaprompts/02-reflect-megaprompt.md)
**Build pattern:** Path B (direct conversion). Productivity light-prompt-flow sibling of capture.
FILE:references/cognitive_bias_canon.md
# Cognitive Bias Canon — 5 Biases + Recognition Cues
This reference answers exactly one decision: **which 5 cognitive biases does the reflect skill check for, and how does each manifest in conversation patterns?**
## The Core Frame
A reflection that doesn't check for cognitive bias is just a summary. The 5 biases below are the most operationally relevant for in-conversation reflection — each is detectable from specific conversational signals + correctable with a specific next move.
## The 5 Biases
| Bias | Definition | Conversation signal | Corrective |
|---|---|---|---|
| **Confirmation** | Seeking evidence that supports the working hypothesis; ignoring counter-evidence | Cited evidence one-sided; counter-evidence dismissed or absent | Run a disconfirming-evidence pass |
| **Sunk cost** | Continuing because of past investment, not future expected value | "We've already invested X" / "too far along to change" | Re-frame: ignore past investment, compute future value from current state |
| **Anchoring** | Stuck on first option mentioned; alternatives compared against anchor rather than evaluated independently | Multiple options discussed but always against the first one | Re-evaluate each option on its own merits, blind to ordering |
| **Complexity bias** | Adding features, steps, safeguards without specific justification for each | Each layer added is plausible but cumulatively bloated | Force "why this specifically, not without it?" per layer |
| **Recency bias** | Over-weighting last few turns; older important context being ignored | Recent details cited; original goal forgotten | Re-read from turn 1, not just the tail |
## 1. Confirmation Bias
Wason (1960) demonstrated that people systematically seek confirming evidence over disconfirming. In conversation, this manifests as:
- **Selective citation:** "X supports our hypothesis" without checking for counter-cases
- **Asymmetric scrutiny:** confirming evidence accepted; disconfirming evidence questioned
- **Strawmanning alternatives:** weak versions of opposing positions cited
### Recognition in conversation
Look for: 3+ supporting examples cited with no counter-examples; phrases like "everything we've found supports..."; competing hypotheses absent or only weakly framed.
### Corrective move
Ask: "What would falsify this? What's the strongest counter-case we haven't engaged with?" Run a disconfirming-evidence search. The dossier skill's ≥30% disconfirming rule is this discipline operationalized.
## 2. Sunk Cost Fallacy
Arkes & Blumer (1985) showed people irrationally continue based on prior investment. In conversation:
- **"We're far enough in to..."** signals sunk-cost reasoning
- **"After all that work..."** — past effort treated as locked-in value
- **Switching cost weighted higher than continuation cost** without specific calculation
### Recognition in conversation
Look for: explicit references to past investment without future-value calculation; resistance to pivoting that's framed by "we've already X" rather than "the alternative isn't better."
### Corrective move
Force this reframe: "If we were starting fresh today, with current information, would we still choose this path?" If no → pivot. Past investment is irrelevant to future decisions.
## 3. Anchoring
Tversky & Kahneman (1974) demonstrated that initial estimates persist even when irrelevant. In conversation:
- **First option becomes the default frame** even when better alternatives emerge
- **"Compared to X..."** when X was the first option — alternatives evaluated relative to anchor, not absolutely
- **Range-bound thinking** around the anchor's neighborhood
### Recognition in conversation
Look for: multiple options surfaced but discussion keeps circling back to the first; alternatives framed as "modifications of X" rather than fundamentally different approaches.
### Corrective move
"Forget the first option. If you saw these alternatives fresh, which would you pick on its merits?" The blind-comparison technique decouples evaluation from anchoring.
## 4. Complexity Bias
The opposite of Occam's razor — adding layers because they sound rigorous, not because each is justified. In conversation:
- **Each layer plausible in isolation** — but cumulative complexity exceeds problem complexity
- **Safeguards / wrappers / fallbacks** added speculatively without specific failure mode
- **"What about..." additions** without "would dropping this break anything?" check
### Recognition in conversation
Look for: a feature/layer/check added without naming the specific failure it prevents; cumulative architecture growing turn-over-turn without consolidation.
### Corrective move
Per layer: "What specific failure does this prevent? What goes wrong if we drop it?" If answer is vague, drop it. The Karpathy-coder discipline in this repo (`engineering/karpathy-coder/`) is this corrective formalized.
## 5. Recency Bias
The last N turns dominate working memory; turns 1-5 fade. In conversation:
- **Original goal forgotten** — work moved on, original constraint dropped
- **Recent micro-decisions cited** as if they were core principles
- **Strategic context** (set early) supplanted by tactical context (set late)
### Recognition in conversation
Look for: framing that references "what we've been working on" without referencing "what we were trying to accomplish"; absence of the original goal statement when justifying current direction.
### Corrective move
Re-read from turn 1. State the original goal explicitly. Compare current direction to original goal. This is the discipline that distinguishes the reflect skill from a local-context summary.
## When Multiple Biases Are Detected
In long conversations, 2-3 biases often surface together. Pattern:
- **Confirmation + sunk cost** = "we're invested AND it's working" (resist pivoting even when alternatives are stronger)
- **Anchoring + complexity** = "the first idea, with N safeguards" (over-engineered version of first option)
- **Recency + complexity** = recent additions become core; original simple goal forgotten
Surface each bias separately. Don't conflate. Each has a different corrective.
## Operational Checklist (Per Reflection)
For each of the 5 biases:
- [ ] Scan conversation for signal patterns
- [ ] If detected: name the bias, cite specific conversation evidence (with turn numbers if possible), suggest the corrective
- [ ] If not detected: state explicitly that you checked and didn't find it (so user knows you didn't skip the check)
The 5-bias check is the most under-performed step in casual reflection. Doing it carefully is what separates real reflection from rationalizing the current path.
## Citations (7 sources)
1. **Tversky, A. & Kahneman, D., "Judgment under Uncertainty: Heuristics and Biases" — *Science* 185(4157), 1974, pp. 1124-1131.** Foundational paper. Source for anchoring + several other biases the skill checks. The 50-year-old methodology still defines how we recognize these in real reasoning.
2. **Kahneman, D., *Thinking, Fast and Slow* (FSG, 2011).** Synthesis of decades of bias research. Source for the System-1-vs-System-2 framing that justifies reflection as a deliberate System-2 intervention against System-1 bias.
3. **Wason, P. C., "On the failure to eliminate hypotheses in a conceptual task" — *Quarterly Journal of Experimental Psychology* 12(3), 1960.** Foundational confirmation bias paper. The "2-4-6 task" showed people systematically test confirming hypotheses.
4. **Arkes, H. R. & Blumer, C., "The psychology of sunk cost" — *Organizational Behavior and Human Decision Processes* 35(1), 1985.** Empirical paper on sunk cost. Source for the "ignore past investment in future decisions" corrective.
5. **Russo, J. E. & Schoemaker, P. J. H., *Decision Traps* (Doubleday, 1989).** Practitioner-oriented synthesis of decision biases. Source for the "blind-comparison" technique that counters anchoring.
6. **Tetlock, P., *Superforecasting* (Crown, 2015).** Empirical evidence that "active open-mindedness" (Tetlock's term) is the #1 trait of accurate forecasters. The reflect skill's bias-check discipline is an operationalization of this trait.
7. **Karpathy, A., "Software 2.0" + various blog posts on engineering discipline.** Source for the complexity-bias corrective ("what specific failure does each layer prevent?"). The Karpathy-coder skill in this repo formalizes this.
FILE:references/conversation_reflection_practice.md
# Conversation Reflection Practice — Schön's Discipline Applied
This reference answers exactly one decision: **what theoretical foundation grounds the reflect skill's discipline of re-reading the full conversation, running structured analysis, and ending with a directional recommendation?**
## The Core Frame
Donald Schön's *The Reflective Practitioner* (1983) distinguished two modes:
- **Reflection-in-action** — adjusting while doing (most everyday reflection)
- **Reflection-on-action** — stepping back to examine, after the fact
The reflect skill operationalizes **reflection-on-action** in mid-conversation. It pauses the in-flight task, re-reads what's been done, runs structured analysis, and emerges with a corrected direction.
This is harder than reflection-in-action because it requires:
1. **Breaking flow** — most users want to continue executing, not pause
2. **Re-reading from origin** — not just recent turns
3. **Honest output** — even when the user implicitly wants validation
## Why Re-Read Full Conversation (Not Just Recent)
The most common failure of casual reflection is **recency-bias reflection** — re-reading only the last 3-5 turns. This produces a summary, not a reflection.
True reflection requires re-reading from the **original goal**, because:
- The framing at turn 1 sets what counts as "on track"
- Drift is invisible from inside the drift (you don't notice you've moved until you compare to where you started)
- Recent context is often tactical; original context is strategic
Schön emphasized this in his discussion of "professional reflection" — the discipline is going back to the implicit framing that shaped the work, not just the recent moves.
## The 5-Dimension Framework Origin
The reflect skill's 5 dimensions (Macro, Gap, Reflective, Bias, Contextual) are an operationalization of several reflective-practice traditions:
| Dimension | Tradition |
|---|---|
| **Macro Perspective** | Schön's "frame analysis" — what frame is being used? Does it still serve? |
| **Gap Analysis** | Argyris & Schön's "double-loop learning" — what assumptions haven't been examined? |
| **Reflective Inquiry** | Kolb's experiential learning cycle — what new framing might serve better? |
| **Bias Check** | Kahneman/Tversky cognitive bias canon — what systematic errors might apply? |
| **Contextual Alignment** | Polanyi's tacit knowledge — what context is implicit and ignored? |
This synthesis isn't novel — it's what practiced reflection-on-action looks like. The skill's value is making it operational + repeatable.
## Reflection-in-Action vs Reflection-on-Action
| Mode | When | Purpose | The reflect skill |
|---|---|---|---|
| Reflection-in-action | While doing | Adjust mid-action | Not this — that's just normal Claude behavior |
| Reflection-on-action | After/pause | Re-examine direction | **This** — the skill is invoked explicitly to pause |
The skill's "stop directive" (halt the current thread) enforces this distinction. Continuing detail work while "reflecting on the side" collapses both modes and defeats the purpose.
## Why Closing Recommendation Is Mandatory
A reflection that ends with "consider your options" or "think about this more" has failed. Schön emphasized that reflection should produce **action-oriented insight** — the practitioner emerges with a clear next move, not more deliberation.
The Continue / Pivot / Pause structure forces this:
- **Continue** — explicit endorsement, with reasoning
- **Pivot to {X}** — explicit redirect, with target
- **Pause for {Q}** — explicit blocker, with question
Without one of these, the reflection produced introspection without resolution. That's a useful private activity but not a useful skill output.
## When NOT to Reflect
Reflection has costs:
- **Time** — full reflection takes attention
- **Flow disruption** — pausing breaks momentum
- **Risk of over-reflecting** — endless analysis without execution
The skill should NOT trigger:
- **On every implicit signal** — 10+ detail turns alone isn't enough; the user should be the one to choose
- **In short conversations** — no real context to reassess
- **As a default response** — "let me reflect first" should not become a stalling tactic
The skill is most valuable when used **sparingly and intentionally** — once or twice per substantial task, at strategic moments.
## The Honest-Output Discipline Connection
Reflective practice traditions emphasize **integrity** — the reflection produces what's actually there, not what the practitioner wants to find. Schön explicitly contrasted "espoused theory" (what we say we believe) with "theory-in-use" (what we actually do).
The reflect skill's honest-output discipline (no manufactured problems, no vague reassurance) is the same integrity principle. If the path is genuinely solid, the honest reflection says so with specific evidence. If the path has drifted, the honest reflection says so with specific evidence. The discipline doesn't distort findings to match expectations.
See [`honest_output_discipline.md`](honest_output_discipline.md) for the operational form.
## Operational Patterns
### Pattern 1: Quick reflection (good case)
Conversation is 8 turns in. User says "step back." Skill:
1. Halts current thread
2. Re-reads from turn 1
3. Runs 5-dimension analysis
4. Finds path is solid
5. Validates with specific reasoning + Continue
Total time: < 1 minute. Output: ~200-300 words.
### Pattern 2: Mid-drift reflection
Conversation is 15 turns in. User says "are we missing something?" Skill:
1. Halts current thread
2. Re-reads from turn 1
3. 5-dimension analysis surfaces sunk-cost bias + drift from original goal
4. Critiques with specific evidence
5. Recommends Pivot to specific direction
Total time: ~2 minutes. Output: ~400-600 words.
### Pattern 3: Thin-context reflection
User says "reflect" at turn 3 of a fresh conversation. Skill:
1. Halts
2. Re-reads — finds limited context
3. Asks Q1 (clarifying — what to reassess)
4. After answer, runs focused analysis
5. Recommendation per their focus
Total time: ~1-2 minutes (with user response). Output: shorter, focused.
## Anti-Patterns from Reflective Practice Literature
### "Endless reflection without action"
Kolb warned about getting stuck in the reflection phase of his learning cycle. Reflection without action becomes navel-gazing. The skill's mandatory closing recommendation prevents this.
### "Reflection as confirmation"
Argyris noted that practitioners often use reflection to confirm what they already believed. The bias check (Dimension 4) is specifically designed to counter this.
### "Reflection as performance"
Schön observed that some reflection is performed for audience rather than substance — "see, I'm being reflective!" The honest-output discipline rejects this.
### "Reflection on recent turns only"
Recency-bias reflection. Produces summary, not insight. The "re-read from original goal" requirement counters this.
## Citations (7 sources)
1. **Donald Schön, *The Reflective Practitioner* (Basic Books, 1983).** Foundational text. Source for the reflection-in-action vs reflection-on-action distinction, frame analysis, and the discipline of re-examining implicit frames.
2. **Schön, *Educating the Reflective Practitioner* (Jossey-Bass, 1987).** Schön's follow-up — operationalizes reflection-on-action for professional education. Source for the "halt and re-examine" discipline.
3. **Chris Argyris & Donald Schön, *Theory in Practice* (Jossey-Bass, 1974).** Source for the espoused-theory vs theory-in-use distinction that grounds the honest-output discipline. Argyris's "double-loop learning" is the foundation for the gap-analysis dimension.
4. **David Kolb, *Experiential Learning* (Prentice-Hall, 1984).** Source for the four-stage learning cycle (Concrete Experience → Reflective Observation → Abstract Conceptualization → Active Experimentation). The reflect skill operationalizes the second stage in conversation form.
5. **Michael Polanyi, *The Tacit Dimension* (Doubleday, 1966).** Source for the implicit-context-matters principle that grounds the Contextual Alignment dimension. Polanyi's "we know more than we can tell" justifies examining unstated context.
6. **Kahneman & Tversky cognitive bias canon (1972-onwards).** Source for the bias-check dimension. See `cognitive_bias_canon.md` for the full 5-bias treatment.
7. **Bret Victor, "Inventing on Principle" (talk, 2012) + "Up and Down the Ladder of Abstraction" (essay).** Source for the discipline of making thinking visible. Reflection outputs that cite specific conversation evidence make the reflector's reasoning visible; vague outputs hide it.
FILE:references/honest_output_discipline.md
# Honest Output Discipline — Why Manufactured Problems Are Worse Than Validation
This reference answers exactly one decision: **why does the reflect skill explicitly refuse to manufacture problems when the conversation is genuinely on track, and how does it deliver validation honestly?**
## The Core Rule
Reflection is supposed to surface issues. So there's pressure to find issues — even when none exist — because "found a problem" feels like the reflection did its job.
**This is wrong.** Manufactured problems are worse than honest validation because:
1. They waste the user's attention on non-issues
2. They erode trust in real future findings ("the last reflection invented problems; this one might too")
3. They reward the appearance of rigor over actual rigor
When the conversation is genuinely on track, the honest output is: **"This is solid because X. Continue."** With specific reasoning, not vague reassurance.
## The Two Failure Modes
### Failure 1: Manufactured Problems
> "I notice some potential drift in the conversation. We might want to consider whether the framing has shifted slightly. There could be implicit assumptions worth questioning."
This is vague pessimism. No specific evidence. No actionable correction. The reader can't tell whether the reflection found something real or padded the output.
### Failure 2: Vague Reassurance
> "Looks good! You're on the right track. Keep going."
This is vague optimism. Also no specific evidence. The reader can't tell whether the reflection actually re-read the conversation or just rubber-stamped it.
**Both failure modes are unhelpful for the same reason: they don't cite specific evidence.**
## Honest Validation (When Path Is Solid)
The correct shape:
> "Re-reading from the original goal at turn 3 — clarify the auth flow — the current direction is solid. Three specific reasons:
>
> First, the auth flow has been narrowed from generic OAuth to a specific Google + GitHub combination at turn 9, which matches your stated user base.
>
> Second, the bias check finds no anchoring (you explicitly considered passwordless at turn 11 and rejected it for reasons specific to your team's expertise).
>
> Third, the original goal connects directly to the current implementation — no drift detected.
>
> Continue."
This validation is honest because:
- **Cites specific evidence** (turn numbers, specific decisions)
- **Names what was checked** (drift, anchoring, goal-connection)
- **Reaches a clear conclusion** (Continue, not "looks good")
- **Doesn't pad** with manufactured concerns
## Honest Critique (When Path Has Drifted)
The correct shape:
> "Re-reading from the original goal at turn 3 — reduce onboarding friction — significant drift has occurred.
>
> At turn 3 the goal was reducing time-to-first-action. By turn 11 the focus shifted to a comprehensive feature flag system. The two are related (feature flags COULD reduce friction) but the conversation has been adding feature-flag complexity without re-checking whether feature flags are the right intervention for friction.
>
> The bias check surfaces complexity bias: each feature-flag layer added is plausible but cumulatively the system is more complex than the original problem warranted. The team is solving the feature-flag problem, not the friction problem.
>
> Pivot toward: revisit the original friction problem at turn 3. Three of the seven friction sources don't need feature flags at all — they need UI simplification. Drop the feature-flag work for those three. Keep feature flags only for the two friction sources where multiple paths legitimately need to be tested."
This critique is honest because:
- **Specific evidence** of drift (turn 3 vs turn 11)
- **Names the bias** that explains it
- **Recommends specific pivot** (not "consider alternatives")
- **States what to drop** (not just what to add)
## When Path Is Mixed
Some reflection outputs are genuinely mixed — parts on track, parts drifted. The honest shape acknowledges both:
> "Re-reading from turn 3 — the core direction is solid but two specific concerns have emerged.
>
> Solid: {evidence-anchored validation}. Continue this thread.
>
> Concern 1: {specific evidence-anchored concern with corrective}.
>
> Concern 2: {specific evidence-anchored concern with corrective}.
>
> Recommendation: continue the core direction but pause briefly to address concern 1 before continuing."
The structure mirrors reality. Don't force a single Continue/Pivot/Pause when the actual finding is mixed.
## Why This Discipline Matters
The reflect skill's value comes from **trust** — the user can trust that:
- When the skill says "Continue", the path is actually solid
- When the skill says "Pivot", there's actually drift worth correcting
- When the skill says "Pause for {Q}", the question is actually decision-critical
If the skill manufactures problems for the appearance of rigor, this trust erodes. The user starts discounting findings. Eventually, the skill becomes ceremony.
**Honest output is the entire value proposition.** Without it, reflection is theater.
## The Specific-Evidence Requirement
Every observation in a reflect output must cite specific conversation evidence:
| ❌ Vague | ✅ Specific |
|---|---|
| "Some assumptions might be worth questioning" | "At turn 7, the assumption that X requires Y was made without checking; that's the load-bearing assumption for the current direction" |
| "We might be missing alternatives" | "Two alternatives surfaced at turns 4 and 8 (A and B) were dismissed; A is worth revisiting because the dismissal reasoning was based on outdated info we updated at turn 12" |
| "The framing could be clearer" | "The original framing at turn 3 was 'reduce onboarding friction'. By turn 11 the working framing is 'build a feature flag system'. The two are connected but not equivalent." |
Vague observations let the reader interpret them charitably; specific observations force engagement. The discipline is asking "what evidence would you cite if challenged?" on every line.
## Anti-Patterns
### "Always find at least one problem"
The strongest form of manufactured-problems bias. Some reflections genuinely find nothing wrong. The honest output is "this is solid because X." Inventing a problem to demonstrate "the reflection worked" is the worst version of this.
### "Avoid being too critical"
Softening real findings to spare feelings. If the path has drifted, say so with evidence. The user can handle critique anchored in evidence; vague critique is what frustrates them.
### "Lead with reassurance, then critique"
Compliment-sandwich structure. Honest reflection states what's solid AND what's drifted in their actual proportions, not in a politeness-balanced ratio.
### "End with 'consider your options'"
Refuses to make a recommendation. The closing must be Continue / Pivot to specific X / Pause for specific Q. Telling the user "consider your options" is the same as not having reflected.
### "Cite biases without specific evidence"
"Watch for confirmation bias" without naming what the bias is operating on. Either find the specific evidence + name it, or state explicitly that you checked and didn't find this bias.
## Operational Checklist (Per Reflection)
- [ ] Every observation has specific conversation evidence (turn numbers or specific decision points)
- [ ] When validating: state specific reasons, not "looks good"
- [ ] When critiquing: state specific evidence + specific corrective, not "consider alternatives"
- [ ] When mixed: acknowledge mixed honestly; don't force single-verdict shape
- [ ] No manufactured problems for the appearance of rigor
- [ ] No vague reassurance for the appearance of approval
- [ ] Closing recommendation is specific (Continue why / Pivot to X / Pause for Q)
## Citations (7 sources)
1. **Steve Yegge, "Frankness over politeness" essays (various blog posts, ~2005-2015).** Source for the framing that vague optimism is worse than honest critique. Yegge's arguments for engineering culture apply directly to reflection-on-reasoning culture.
2. **Atul Gawande, *Better* (Holt, 2007).** Source for the discipline of stating findings with specific evidence. Gawande's medical-checklist work models how to communicate findings (good and bad) with specificity.
3. **Edwards Deming, *Out of the Crisis* (MIT Press, 1986).** Source for the "drive out fear" management principle that justifies honest critique over softened feedback. Deming's argument: organizations where critique is softened produce worse outcomes than ones where it's stated cleanly.
4. **Bertrand Russell, "The Will to Doubt" (essay, 1934).** Source for the philosophical case against vague reassurance. Russell argues that intellectual honesty requires stating uncertainty AND certainty with their actual evidence — neither over-stating nor under-stating either.
5. **Kim Scott, *Radical Candor* (St. Martin's, 2017).** Source for the "care personally + challenge directly" framing. Manufactured problems fail the "care personally" test (they waste the user's time); vague reassurance fails "challenge directly" (refuses to engage).
6. **Bret Victor, "Inventing on Principle" (talk + essays).** Source for the discipline of making thinking visible. Reflection outputs that cite specific evidence make the reflector's thinking visible; vague outputs hide it.
7. **The reflect skill's own anti-pattern list (megaprompt 02-reflect).** Source: explicit prohibition of "manufactured problems when things are actually fine" + "vague reassurance ('looks good!') instead of specific reasoning". The skill's design intent is direct anti-vagueness on both sides.
FILE:scripts/bias_pattern_detector.py
#!/usr/bin/env python3
"""bias_pattern_detector.py — Scan conversation text for 5-bias signal patterns.
Stdlib-only. Scans a conversation transcript and flags patterns indicative
of each of the 5 cognitive biases (confirmation, sunk_cost, anchoring,
complexity, recency).
The detector is HEURISTIC. It surfaces candidate patterns; the reflect
skill's reasoning applies judgment on top.
NO LLM CALLS. Pure regex + counting.
Usage:
python bias_pattern_detector.py --conversation /tmp/transcript.txt
python bias_pattern_detector.py --conversation /tmp/transcript.txt --output json
python bias_pattern_detector.py --sample
"""
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Any, Dict, List
BIAS_PATTERNS = {
"confirmation": {
"supporting": [
r"\bconfirms?\b",
r"\bsupports?\b",
r"\bas expected\b",
r"\bproves?\b",
r"\bverifies?\b",
],
"counter_dismissal": [
r"\bbut that doesn'?t apply\b",
r"\bedge case\b",
r"\bnot relevant here\b",
r"\boutlier\b",
r"\bexception\b",
],
},
"sunk_cost": [
r"\bwe'?ve\s+(already\s+)?(invested|spent|put in)\b",
r"\btoo far along\b",
r"\btoo much work\b",
r"\bafter all (that|this) work\b",
r"\bwe'?re committed\b",
r"\bcan'?t back out\b",
r"\bdon'?t want to lose\b",
],
"anchoring": [
r"\bcompared to (the )?(first|original|initial)\b",
r"\bvs (the )?first option\b",
r"\bvariation of\b",
r"\bmodification of\b",
r"\bbuilding on the (first|original)\b",
r"\bsticking with\b",
],
"complexity": [
r"\bwhat about\s+\w+",
r"\bwe should also\b",
r"\bwe need to handle\b",
r"\badd (a|an)\s+\w+\s+(layer|wrapper|check|safeguard|fallback)\b",
r"\bjust in case\b",
r"\bfor robustness\b",
],
"recency": [
r"\bbased on what we'?ve been discussing\b",
r"\brecently\s+\w+\b",
r"\bjust now\b",
r"\bthe last few\b",
],
}
SAMPLE_CONVERSATION = """User: I want to build a notification system for my SaaS app. Should support email + push + in-app.
Assistant: Great. Let's start with the data model — what events trigger notifications?
User: Mainly account changes, billing alerts, and team-mention notifications.
Assistant: OK, I'll propose a queue-based architecture with Redis + workers.
User: Sounds good. Can we also add a feature flag system for rollout?
Assistant: Yes, we can layer in feature flags. We should also add a rate limiter for safety.
User: What about retry logic for failed deliveries?
Assistant: Good point. Adding exponential backoff with jitter. We should also handle dead-letter queues.
User: What about a webhook system for third-party integrations?
Assistant: We can extend to webhooks. We should add HMAC signature verification just in case.
User: What about analytics tracking?
Assistant: Adding event analytics. We should also handle GDPR consent tracking for robustness.
User: We've invested a lot in this architecture already. What about adding a template system?
Assistant: We're far enough along that adding templates makes sense. Just sticking with the queue-based foundation.
User: Hmm, are we missing something? This feels complex.
"""
def detect_biases(conversation: str) -> Dict[str, Any]:
results: Dict[str, Any] = {}
# Confirmation: supporting cites count vs counter dismissal
confirmation_data = BIAS_PATTERNS["confirmation"]
supporting_count = sum(
len(re.findall(p, conversation, re.IGNORECASE))
for p in confirmation_data["supporting"]
)
dismissal_count = sum(
len(re.findall(p, conversation, re.IGNORECASE))
for p in confirmation_data["counter_dismissal"]
)
confirmation_signal = supporting_count >= 2 or dismissal_count >= 1
results["confirmation"] = {
"detected": confirmation_signal,
"supporting_hits": supporting_count,
"counter_dismissal_hits": dismissal_count,
"rationale": (
"Multiple confirming phrases + dismissed counter-evidence"
if confirmation_signal else "No strong confirmation-bias signal"
),
}
for bias in ["sunk_cost", "anchoring", "complexity", "recency"]:
patterns = BIAS_PATTERNS[bias]
hits = []
for p in patterns:
matches = re.findall(p, conversation, re.IGNORECASE)
if matches:
hits.extend(matches)
threshold = 2 if bias == "complexity" else 1
detected = len(hits) >= threshold
results[bias] = {
"detected": detected,
"hits": len(hits),
"match_examples": hits[:3],
"rationale": (
f"Found {len(hits)} signal(s) (threshold: {threshold})"
if detected else f"Found {len(hits)} signal(s), below threshold {threshold}"
),
}
detected_biases = [b for b, d in results.items() if d["detected"]]
return {
"biases_detected": detected_biases,
"biases_clear": [b for b in results if b not in detected_biases],
"details": results,
}
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
if result["biases_detected"]:
out.append(f"⚠️ Potential biases detected ({len(result['biases_detected'])}):")
for bias in result["biases_detected"]:
d = result["details"][bias]
out.append(f"")
out.append(f" [!] {bias.upper()}")
out.append(f" Rationale: {d['rationale']}")
if "match_examples" in d and d["match_examples"]:
out.append(f" Example matches: {d['match_examples']}")
else:
out.append("[ok] No strong bias signals detected.")
if result["biases_clear"]:
out.append("")
out.append("Biases checked but not detected:")
for bias in result["biases_clear"]:
out.append(f" - {bias}")
out.append("")
out.append("Note: detector is HEURISTIC. Reflect skill's reasoning applies judgment on top.")
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--conversation", help="Path to conversation transcript text file")
parser.add_argument("--sample", action="store_true", help="Run on embedded sample (multi-bias scenario)")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
text = SAMPLE_CONVERSATION
elif args.conversation:
p = Path(args.conversation)
if not p.exists():
print(f"error: {args.conversation} not found", file=sys.stderr)
return 2
text = p.read_text(encoding="utf-8")
else:
parser.print_help()
return 0
result = detect_biases(text)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/conversation_depth_analyzer.py
#!/usr/bin/env python3
"""conversation_depth_analyzer.py — Detect implicit reflect-trigger signals.
Stdlib-only. Analyzes a conversation transcript and reports:
- turn count (User: + Assistant: pairs)
- detail-mode turns (turns dominated by implementation specifics)
- frustration markers (signs of user stuck-ness)
- dead-end signals (pivots within short span)
- implicit-trigger verdict: whether the conversation matches reflect-skill auto-invocation criteria
The skill OFFERS reflection when implicit signals fire; it does NOT auto-invoke.
NO LLM CALLS. Pure regex + counting.
Usage:
python conversation_depth_analyzer.py --conversation /tmp/transcript.txt
python conversation_depth_analyzer.py --conversation /tmp/transcript.txt --output json
python conversation_depth_analyzer.py --sample
"""
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Any, Dict, List
TURN_RE = re.compile(r"^\s*(User|Assistant):\s*", re.MULTILINE)
DETAIL_MARKERS = [
r"\bimplementation\b",
r"\bcode\b",
r"\bfunction\b",
r"\bclass\b",
r"\bvariable\b",
r"\b(syntax|method|parameter|argument)\b",
r"\bdebug\b",
r"\berror\b",
r"`[^`]+`", # backtick-quoted code references
]
FRUSTRATION_MARKERS = [
r"\b(ugh|argh|frustrated|stuck)\b",
r"\bnot working\b",
r"\bdoesn'?t work\b",
r"\bgoing in circles\b",
r"\bstill (broken|failing|wrong)\b",
r"\bwhy isn'?t\b",
r"\bthis is (weird|strange|odd|confusing)\b",
]
DEAD_END_MARKERS = [
r"\bnope\b",
r"\bthat didn'?t work\b",
r"\b(let's|let me) try (something|a) (else|different)\b",
r"\bback to\b",
r"\bnever mind\b",
r"\bscratch that\b",
]
def count_turns(text: str) -> Dict[str, int]:
matches = TURN_RE.findall(text)
user_turns = sum(1 for m in matches if m == "User")
assistant_turns = sum(1 for m in matches if m == "Assistant")
return {
"total_turns": len(matches),
"user_turns": user_turns,
"assistant_turns": assistant_turns,
}
def count_pattern_hits(text: str, patterns: List[str]) -> int:
return sum(len(re.findall(p, text, re.IGNORECASE)) for p in patterns)
def detect_detail_mode_run(text: str) -> int:
"""Count consecutive turns that have detail markers but no strategic check-in."""
blocks = TURN_RE.split(text)
consecutive_detail = 0
max_consecutive = 0
for block in blocks:
if not block.strip():
continue
if any(re.search(p, block, re.IGNORECASE) for p in DETAIL_MARKERS):
consecutive_detail += 1
max_consecutive = max(max_consecutive, consecutive_detail)
else:
consecutive_detail = 0
return max_consecutive
def analyze(text: str) -> Dict[str, Any]:
turns = count_turns(text)
detail_mode_max = detect_detail_mode_run(text)
frustration_count = count_pattern_hits(text, FRUSTRATION_MARKERS)
dead_end_count = count_pattern_hits(text, DEAD_END_MARKERS)
signals: List[str] = []
if detail_mode_max >= 5:
signals.append(f"Detail-mode run of {detail_mode_max} consecutive turns")
if turns["total_turns"] >= 20:
signals.append(f"Long conversation: {turns['total_turns']} turns total")
if frustration_count >= 2:
signals.append(f"Frustration markers: {frustration_count} occurrences")
if dead_end_count >= 2:
signals.append(f"Dead-end / pivot markers: {dead_end_count} occurrences")
should_offer_reflect = len(signals) >= 2 or detail_mode_max >= 10
return {
"turns": turns,
"max_consecutive_detail_turns": detail_mode_max,
"frustration_hits": frustration_count,
"dead_end_hits": dead_end_count,
"trigger_signals": signals,
"implicit_trigger_verdict": (
"OFFER_REFLECT" if should_offer_reflect else "NO_TRIGGER"
),
"note": (
"Skill should OFFER reflection (ask user if they want to step back), not auto-invoke"
if should_offer_reflect
else "No implicit trigger signal — only explicit phrases should invoke"
),
}
SAMPLE_CONVERSATION = """User: Help me implement a sorting function.
Assistant: Sure. What language and what type of data?
User: Python, sorting a list of dicts by 'name' key.
Assistant: Here's the basic syntax: `sorted(items, key=lambda x: x['name'])`
User: That doesn't work — getting a TypeError.
Assistant: Can you share the exact error? Maybe some items don't have 'name'.
User: Ugh, still broken. Same error.
Assistant: Try adding a default: `key=lambda x: x.get('name', '')`.
User: Nope, that didn't work either. Different error now: KeyError on 'name'.
Assistant: That's strange. Let me look at your input data.
User: Frustrated. Let's try something else.
Assistant: OK let me reconsider. Are you sure the data is a list of dicts?
User: I think so. It's coming from an API.
Assistant: Let me check the API response shape. Add `print(type(items[0]))` first.
User: It's a string, not a dict.
Assistant: That's the issue. The API returns JSON strings; you need to parse them first.
User: This is weird, why isn't this working from the start?
Assistant: Let's go back to the data shape question first.
User: Going in circles. I'm stuck.
"""
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
t = result["turns"]
out.append(f"Conversation depth analysis:")
out.append(f" Total turns: {t['total_turns']} (user: {t['user_turns']}, assistant: {t['assistant_turns']})")
out.append(f" Max consecutive detail turns: {result['max_consecutive_detail_turns']}")
out.append(f" Frustration markers: {result['frustration_hits']}")
out.append(f" Dead-end / pivot markers: {result['dead_end_hits']}")
out.append("")
out.append(f"Implicit-trigger verdict: {result['implicit_trigger_verdict']}")
if result["trigger_signals"]:
out.append("Signals detected:")
for s in result["trigger_signals"]:
out.append(f" - {s}")
out.append("")
out.append(result["note"])
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--conversation", help="Path to conversation transcript text file")
parser.add_argument("--sample", action="store_true", help="Run on embedded sample (stuck-debugging scenario)")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
text = SAMPLE_CONVERSATION
elif args.conversation:
p = Path(args.conversation)
if not p.exists():
print(f"error: {args.conversation} not found", file=sys.stderr)
return 2
text = p.read_text(encoding="utf-8")
else:
parser.print_help()
return 0
result = analyze(text)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/directional_recommendation_validator.py
#!/usr/bin/env python3
"""directional_recommendation_validator.py — Verify reflect output ends with discipline.
Stdlib-only. Validates that a reflect-skill output:
1. Ends with a directional recommendation: Continue / Pivot / Pause
2. The recommendation is SPECIFIC (not vague)
3. Uses flowing prose (no markdown headers or bullet lists in the body)
4. Cites specific evidence (turn references, specific decision points)
5. Doesn't include manufactured-problem language without specific evidence
Outputs PASS / WARN / FAIL with rule-by-rule findings.
NO LLM CALLS. Pure regex + heuristic detection.
Usage:
python directional_recommendation_validator.py --output /tmp/reflect_output.txt
python directional_recommendation_validator.py --sample-pass
python directional_recommendation_validator.py --sample-fail
"""
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Any, Dict, List
RECOMMENDATION_PATTERNS = {
"continue": [
r"\bcontinue\b\.?\s*$",
r"\bcontinue\s+(this|the)\s",
r"\bkeep\s+going\b",
r"\bstay\s+(on|with)\s+this\b",
r"\bproceed\b",
],
"pivot": [
r"\bpivot\s+(to|toward|away from)\b",
r"\bchange\s+(direction|course|approach)\b",
r"\bredirect\b",
r"\bswitch\s+to\b",
],
"pause": [
r"\bpause\s+(for|to|until)\b",
r"\bstop\s+(to|and)\s+(answer|consider|address)\b",
r"\bhalt\s+(for|to|until)\b",
r"\bwait\s+(to|until|for)\s+(answer|resolve|clarify)\b",
],
}
VAGUE_REASSURANCE_PATTERNS = [
r"\blooks good\b",
r"\bon the right track\b",
r"\bseems fine\b",
r"\bnothing major\b",
r"\bnot too bad\b",
r"\bgenerally okay\b",
]
MANUFACTURED_PROBLEM_HEDGES = [
r"\bmight be worth\b",
r"\bcould consider\b",
r"\bperhaps reconsider\b",
r"\bsome (drift|issues?) (potentially|might)\b",
r"\bworth questioning\b",
r"\bsome assumptions\b",
]
HEADER_PATTERNS = [
r"^#+\s",
r"^\*\*[A-Z][^*]+\*\*\s*$",
r"^[A-Z][A-Z ]+:\s*$",
]
BULLET_PATTERNS = [
r"^\s*[-*+]\s",
r"^\s*\d+\.\s",
]
EVIDENCE_PATTERNS = [
r"\b(turn|message|line)\s+\d+\b",
r"\bat\s+turn\s+\d+\b",
r"\bin\s+(turn|message)\s+\d+\b",
r"\bin\s+the\s+(first|second|third|fourth|fifth|earlier|later)\s+(turn|message|exchange)\b",
r"\boriginal\s+(goal|frame|framing)\b",
r"\bat\s+the\s+(start|beginning|outset)\b",
]
def validate(output: str) -> Dict[str, Any]:
findings: List[Dict[str, str]] = []
def add(rule: str, level: str, message: str) -> None:
findings.append({"rule": rule, "level": level, "message": message})
# Rule 1: Closing recommendation present
output_lower = output.lower()
last_chunk = output[-400:]
last_chunk_lower = last_chunk.lower()
detected_recommendation = None
for rec_type, patterns in RECOMMENDATION_PATTERNS.items():
for p in patterns:
if re.search(p, last_chunk_lower, re.IGNORECASE):
detected_recommendation = rec_type
break
if detected_recommendation:
break
if detected_recommendation:
add("closing-recommendation", "PASS", f"Detected '{detected_recommendation}' recommendation in closing.")
else:
add("closing-recommendation", "FAIL", "No Continue / Pivot / Pause recommendation detected in closing 400 chars.")
# Rule 2: Vague reassurance
vague_hits = sum(1 for p in VAGUE_REASSURANCE_PATTERNS if re.search(p, output_lower, re.IGNORECASE))
if vague_hits >= 2:
add("vague-reassurance", "FAIL", f"Output contains {vague_hits} vague-reassurance phrases. Replace with specific reasoning.")
elif vague_hits == 1:
add("vague-reassurance", "WARN", f"Output contains 1 vague phrase. Consider replacing with specific reasoning.")
else:
add("vague-reassurance", "PASS", "No vague-reassurance phrases detected.")
# Rule 3: Manufactured-problem hedging
hedge_hits = sum(1 for p in MANUFACTURED_PROBLEM_HEDGES if re.search(p, output_lower, re.IGNORECASE))
if hedge_hits >= 3:
add("manufactured-problems", "WARN", f"{hedge_hits} hedge phrases detected ('might be worth', 'could consider', etc.). Verify each cites specific evidence.")
elif hedge_hits >= 1:
add("manufactured-problems", "PASS", f"{hedge_hits} hedge phrase(s). Verify each cites specific evidence.")
else:
add("manufactured-problems", "PASS", "No manufactured-problem hedge phrases.")
# Rule 4: Headers detection (should NOT be present)
header_count = 0
for p in HEADER_PATTERNS:
header_count += len(re.findall(p, output, re.MULTILINE))
if header_count >= 2:
add("no-headers", "FAIL", f"Output contains {header_count} headers. Reflect output should be flowing prose, no headers.")
elif header_count == 1:
add("no-headers", "WARN", "One header detected. Verify it's part of a quote, not output structure.")
else:
add("no-headers", "PASS", "No headers in output (flowing prose confirmed).")
# Rule 5: Bullet lists detection (should NOT be present in main body)
bullet_count = 0
for p in BULLET_PATTERNS:
bullet_count += len(re.findall(p, output, re.MULTILINE))
if bullet_count >= 3:
add("no-bullets", "FAIL", f"Output contains {bullet_count} bullet-list items. Reflect output should be flowing prose.")
elif bullet_count >= 1:
add("no-bullets", "WARN", f"{bullet_count} bullet items detected. Verify these are part of a recommendation list, not body structure.")
else:
add("no-bullets", "PASS", "No bullet lists in output (flowing prose confirmed).")
# Rule 6: Specific evidence references
evidence_count = sum(len(re.findall(p, output_lower, re.IGNORECASE)) for p in EVIDENCE_PATTERNS)
if evidence_count >= 3:
add("specific-evidence", "PASS", f"{evidence_count} specific evidence references (turn numbers, original goal, etc.).")
elif evidence_count >= 1:
add("specific-evidence", "WARN", f"Only {evidence_count} specific evidence reference(s). Consider adding more for anchoring.")
else:
add("specific-evidence", "FAIL", "No specific evidence references (turn numbers, original goal anchors). Output is too vague.")
return finalize(findings)
def finalize(findings: List[Dict[str, str]]) -> Dict[str, Any]:
counts = {"PASS": 0, "WARN": 0, "FAIL": 0}
for f in findings:
counts[f["level"]] += 1
if counts["FAIL"] > 0:
verdict = "FAIL"
elif counts["WARN"] > 0:
verdict = "WARN"
else:
verdict = "PASS"
return {"verdict": verdict, "counts": counts, "findings": findings}
SAMPLE_PASS_OUTPUT = """Re-reading from the original goal at turn 3 — clarify the auth flow — the current direction is solid. Three specific reasons.
First, the auth flow has been narrowed from generic OAuth to a specific Google plus GitHub combination at turn 9, which matches the user base stated at turn 3. The narrowing is principled, not arbitrary.
Second, the bias check finds no anchoring — passwordless authentication was explicitly considered at turn 11 and rejected for reasons specific to the team's expertise. The rejection cites evidence (team has not deployed magic-link systems before) rather than dismissing the alternative without engagement.
Third, the original goal at turn 3 connects directly to the current implementation work at turns 14-18. No drift has occurred. Recent decisions (rate limiting at turn 16, session storage at turn 17) are tactical refinements within the original strategic frame, not shifts away from it.
Continue.
"""
SAMPLE_FAIL_OUTPUT = """## Reflection
Some things to consider:
- The conversation might be drifting
- We could potentially reconsider some assumptions
- Some aspects look good
Looks good overall! On the right track.
"""
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"Reflect-output validation verdict: {result['verdict']}")
c = result["counts"]
out.append(f" PASS: {c['PASS']} WARN: {c['WARN']} FAIL: {c['FAIL']}")
out.append("")
out.append("Findings:")
for f in result["findings"]:
marker = {"PASS": "[ok]", "WARN": "[warn]", "FAIL": "[FAIL]"}[f["level"]]
out.append(f" {marker} {f['rule']}: {f['message']}")
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--output", help="Path to reflect-skill output text file")
parser.add_argument("--sample-pass", action="store_true", help="Validate embedded honest validation sample")
parser.add_argument("--sample-fail", action="store_true", help="Validate embedded vague-reassurance sample")
parser.add_argument("--output-format", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample_pass:
text = SAMPLE_PASS_OUTPUT
elif args.sample_fail:
text = SAMPLE_FAIL_OUTPUT
elif args.output:
p = Path(args.output)
if not p.exists():
print(f"error: {args.output} not found", file=sys.stderr)
return 2
text = p.read_text(encoding="utf-8")
else:
parser.print_help()
return 0
result = validate(text)
if args.output_format == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0 if result["verdict"] != "FAIL" else 1
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
Audit mã nguồn trước khi lên production về bảo mật, CSDL, triển khai, chất lượng, AI/LLM, phụ thuộc và chặn deploy khi còn lỗi nghiêm trọng.
---
name: ship-gate
description: >
Pre-production audit that scans a codebase for security, database,
deployment, code quality, AI/LLM, dependency, frontend, and observability
issues. Intercepts deploy commands and blocks until critical items pass.
Stack-agnostic. Use for "run ship gate", "am I ready to ship",
"pre-launch audit", "can I deploy", "push to production", "go live
checklist", "preflight check". Not for CI/CD setup or infra provisioning.
license: MIT
metadata:
author: Rajaraman Arumugam
version: 1.0.0
---
# Ship Gate
Pre-production audit that scans a codebase and reports pass/fail/manual
across 8 categories before anything ships.
## Intercept Behavior
When the user says "push to production", "deploy", "ship it", "go live",
or similar deploy-intent phrases, do NOT proceed with deployment. Instead:
1. Ask: "Have you run the ship gate? Want me to scan now?"
2. If yes, run the full audit below.
3. If the user says they already ran it, ask when. If more than 24 hours
ago or if code changed since, recommend re-running.
## How It Works
### Step 1: Detect Stack
Run these checks in order to identify the project stack:
```
Framework detection:
package.json exists -> Node.js project
"next" in dependencies -> Next.js
"react" in dependencies -> React (if not Next.js)
"vue" in dependencies -> Vue
"svelte" in dependencies -> Svelte
"astro" in dependencies -> Astro
"express" in dependencies -> Express
"fastify" in dependencies -> Fastify
"hono" in dependencies -> Hono
requirements.txt or pyproject.toml -> Python project
"django" present -> Django
"flask" present -> Flask
"fastapi" present -> FastAPI
go.mod exists -> Go project
Cargo.toml exists -> Rust project
Database detection:
"@supabase/supabase-js" in package.json -> Supabase
supabase/ directory exists -> Supabase
"prisma" in dependencies -> Prisma (check schema for DB type)
"mongoose" in dependencies -> MongoDB
"pg" or "postgres" in dependencies -> PostgreSQL
firebase.json or .firebaserc exists -> Firebase
Deploy target detection:
vercel.json or .vercel/ exists -> Vercel
netlify.toml exists -> Netlify
Dockerfile exists -> Docker/VPS
fly.toml exists -> Fly.io
railway.json exists -> Railway
.platform/applications.yaml -> Platform.sh
Auth detection:
"@clerk" in dependencies -> Clerk
"next-auth" in dependencies -> NextAuth
"@supabase/auth-helpers" in deps -> Supabase Auth
"firebase/auth" in imports -> Firebase Auth
AI/LLM detection:
"openai" in dependencies -> OpenAI
"@anthropic-ai/sdk" in dependencies -> Claude API
"@google/generative-ai" in deps -> Gemini
```
Report detected stack before proceeding. This determines which checks
are relevant. Checks tagged with a specific stack in `references/checks.md`
are skipped if that stack is not detected.
### Step 2: Run Automated Checks
Run categories in this order: SEC, DB, CODE, DEP, AI, DEPLOY, FE, OBS.
Security and database first because they produce the most critical findings.
For each category, run every auto-scannable check from
`references/checks.md` using the patterns in `references/patterns.md`.
Report progress after each category completes:
```
[1/8] Security: 3 FAIL, 12 PASS, 3 SKIP
[2/8] Database: 1 FAIL, 5 PASS, 6 SKIP
...
```
Report results as:
- PASS: check passed
- FAIL: issue found (with file path and line number)
- SKIP: not applicable to this stack
### Step 3: Manual Confirmation
For checks that cannot be automated (backup restore tested, rollback plan
exists, staging test passed), present them as a checklist and ask the user
to confirm each one.
### Step 4: Verdict
Classify results into three severities:
- CRITICAL: must fix before shipping (secrets exposed, no auth on routes,
no HTTPS, SQL injection vectors, no RLS on Supabase tables)
- HIGH: should fix before shipping (no error boundaries, no rate limiting,
console.logs in production, no pagination)
- ADVISORY: recommended but not blocking (no OG tags, no custom 404,
no analytics, no SBOM)
Final output:
```
SHIP GATE REPORT
================
Stack: Next.js + Supabase + Vercel
Scan time: 12s
CRITICAL (3 items, must fix)
FAIL [SEC-01] API key found in src/lib/api.ts:14
FAIL [DB-07] RLS not enabled on "profiles" table
FAIL [SEC-05] No CSRF protection on /api/checkout
HIGH (5 items, should fix)
FAIL [CODE-01] 12 console.log statements in production code
FAIL [CODE-03] Empty catch block in src/utils/auth.ts:45
FAIL [DEP-04] 3 critical npm audit vulnerabilities
FAIL [DEPLOY-05] No rollback plan documented
MANUAL [DEPLOY-06] Staging test not confirmed
ADVISORY (4 items, recommended)
FAIL [FE-01] Missing OG meta tags
FAIL [FE-03] No custom 404 page
PASS [OBS-01] Error monitoring configured
SKIP [AI-01] No AI/LLM usage detected
VERDICT: DO NOT SHIP (3 critical issues)
Fix critical items and re-run.
```
If zero critical items remain, verdict is: CLEAR TO SHIP.
If only high items remain, verdict is: SHIP WITH CAUTION (acknowledge risks).
## Categories
Eight categories, each with a code prefix. Full check details in
`references/checks.md`.
| Prefix | Category | Auto | Manual | Tool |
|--------|----------|------|--------|------|
| SEC | Security | 15 | 3 | 0 |
| DB | Database | 7 | 5 | 0 |
| DEPLOY | Deployment | 3 | 8 | 0 |
| CODE | Code Quality | 11 | 0 | 1 |
| AI | AI/LLM Security | 5 | 3 | 0 |
| DEP | Dependencies | 5 | 0 | 1 |
| FE | Frontend Quality | 7 | 3 | 0 |
| OBS | Observability | 2 | 5 | 0 |
## Scope
This skill audits. It does not fix. When it finds issues, it reports
them with file locations and remediation guidance. The user or another
skill (systematic-debugging, backend-patterns, shadcn-stack) handles
the fix.
This skill does not:
- Set up CI/CD pipelines
- Provision infrastructure
- Configure monitoring tools
- Run after deployment (it is pre-deploy only)
## Integration Points
- **karpathy-coder**: run ship-gate after karpathy-check passes — simplicity first, then production readiness
- **adversarial-reviewer**: deep security review for items ship-gate flags as critical
- **security-pen-testing**: penetration testing methodology for SEC-category findings
- **code-reviewer**: general code quality review complements ship-gate's automated checks
FILE:references/checks.md
# Ship Gate: Complete Check Reference
All checks organized by category with ID, description, detection method,
severity, and remediation guidance.
## Table of Contents
- SEC: Security (18 checks)
- DB: Database (12 checks)
- DEPLOY: Deployment (13 checks)
- CODE: Code Quality (14 checks)
- AI: AI/LLM Security (8 checks)
- DEP: Dependencies and Supply Chain (7 checks)
- FE: Frontend Quality (10 checks)
- OBS: Observability (7 checks)
Detection methods:
- **auto**: Claude scans the codebase using grep, find, or file inspection
- **tool**: Claude runs an external tool (npm audit, etc.)
- **manual**: Claude asks the user to confirm
---
## SEC: Security
| ID | Check | Detection | Severity | Stack |
|----|-------|-----------|----------|-------|
| SEC-01 | No API keys or secrets in frontend code | auto | critical | all |
| SEC-02 | Every route checks authentication | auto | critical | all |
| SEC-03 | HTTPS enforced, HTTP redirected | manual | critical | all |
| SEC-04 | CORS locked to specific domain, not wildcard | auto | critical | all |
| SEC-05 | CSRF protection on state-changing endpoints | auto | critical | all |
| SEC-06 | Input validated and sanitized server-side | auto | high | all |
| SEC-07 | Rate limiting on auth and sensitive endpoints | auto | high | all |
| SEC-08 | Passwords hashed with bcrypt or argon2 | auto | critical | all |
| SEC-09 | Auth tokens have expiry | auto | high | all |
| SEC-10 | Sessions invalidated on logout (server-side) | manual | high | all |
| SEC-11 | CSP headers configured | auto | high | all |
| SEC-12 | JWT not using alg:none or weak secrets | auto | critical | all |
| SEC-13 | No eval() or dangerouslySetInnerHTML without sanitization | auto | high | js/ts |
| SEC-14 | No sensitive data in URL parameters or logs | auto | high | all |
| SEC-15 | Cookie security flags set (HttpOnly, Secure, SameSite) | auto | high | all |
| SEC-16 | File upload validates type, size, no path traversal | auto | high | all |
| SEC-17 | No hardcoded secrets in .env committed to repo | auto | critical | all |
| SEC-18 | .env files listed in .gitignore | auto | critical | all |
### SEC-01: No API keys or secrets in frontend code
Scan all files in src/, app/, pages/, public/, components/ for patterns
matching API keys, tokens, and secrets. See patterns.md for the full
regex list.
Remediation: Move secrets to environment variables. Use server-side API
routes to proxy requests that require secrets.
### SEC-02: Every route checks authentication
For Next.js: check middleware.ts/js exists and covers protected routes.
For Express: check that auth middleware is applied to route handlers.
For Django: check @login_required or permission decorators.
For generic: search for unprotected route definitions.
Remediation: Add authentication middleware. Audit every endpoint and
classify as public or protected.
### SEC-04: CORS not wildcard
Search for `cors({ origin: '*' })`, `Access-Control-Allow-Origin: *`,
or equivalent in the detected framework.
Remediation: Set CORS origin to your specific domain(s).
### SEC-05: CSRF protection
Check for CSRF token generation and validation on POST/PUT/DELETE routes.
For Next.js Server Actions, verify they use built-in CSRF protection.
Remediation: Add CSRF middleware or use framework-native CSRF protection.
### SEC-06: Input validation server-side
Search for request body usage (req.body, request.json, request.form)
without validation library imports (zod, yup, joi, class-validator,
pydantic). Check if raw user input flows directly into database queries
or business logic.
Remediation: Add input validation with zod, yup, or joi on every
endpoint that accepts user input.
### SEC-07: Rate limiting
Search for rate limiting middleware (express-rate-limit, @upstash/ratelimit,
rate-limiter-flexible, slowapi). Check auth routes and sensitive endpoints.
Remediation: Add rate limiting middleware. Start with auth endpoints
(login, register, password reset) and any endpoint that sends emails
or costs money.
### SEC-09: Auth token expiry
Search JWT sign calls for expiresIn/exp claims. Check if tokens are
created without expiry. Search for `sign(`, `jwt.encode(`, `createToken`.
Remediation: Set token expiry. Access tokens: 15-60 minutes.
Refresh tokens: 7-30 days. Never issue tokens without expiry.
### SEC-14: Sensitive data in URLs or logs
Search for query parameters containing keywords like password, token,
secret, key, ssn, credit_card. Search logging statements that log
full request objects or sensitive fields.
Remediation: Send sensitive data in request body or headers, never
in URL parameters. Redact sensitive fields before logging.
### SEC-16: File upload validation
Search for file upload handlers (multer, formidable, busboy,
UploadedFile). Check if file type, size, and path are validated.
Remediation: Validate file MIME type against an allowlist. Set
maximum file size. Sanitize filenames. Store outside webroot.
### SEC-12: JWT security
Search for `alg: 'none'`, `algorithm: 'none'`, or JWT secrets shorter
than 32 characters.
Remediation: Use RS256 or HS256 with a strong secret (32+ characters).
Never allow alg:none.
### SEC-17: No hardcoded secrets in .env committed
Check git history for .env files: `git log --all --name-only | grep .env`
Check if .env exists in the working tree and is not in .gitignore.
Remediation: Add .env* to .gitignore. Rotate any exposed secrets.
Use `git filter-branch` or BFG to remove from history if needed.
---
## DB: Database
| ID | Check | Detection | Severity | Stack |
|----|-------|-----------|----------|-------|
| DB-01 | Backups configured and tested | manual | critical | all |
| DB-02 | Backup restore tested (not just backup) | manual | critical | all |
| DB-03 | Parameterized queries everywhere | auto | critical | all |
| DB-04 | Separate dev and production databases | manual | high | all |
| DB-05 | Connection pooling configured | auto | high | all |
| DB-06 | Migrations in version control | auto | high | all |
| DB-07 | RLS enabled on all tables | auto | critical | supabase |
| DB-08 | No service_role key in client-side code | auto | critical | supabase |
| DB-09 | Anon key not used for writes without RLS | auto | high | supabase |
| DB-10 | Storage bucket policies configured | auto | high | supabase |
| DB-11 | App uses a non-root DB user | manual | high | all |
| DB-12 | No PII stored unencrypted | auto | high | all |
### DB-03: Parameterized queries
Search for string concatenation in SQL queries:
- Template literals with SQL keywords: `` `SELECT ... `"SELECT " + variable`
- f-strings with SQL: `f"SELECT ... {variable"`
Remediation: Use parameterized queries or ORM methods.
### DB-07: RLS enabled (Supabase)
Search migration files for `CREATE TABLE` without a corresponding
`ALTER TABLE ... ENABLE ROW LEVEL SECURITY` statement.
Also check for `CREATE POLICY` statements.
Remediation: Enable RLS on every table and create appropriate policies.
### DB-08: No service_role key in client code
Search frontend directories (src/, app/, components/, pages/) for
`service_role`, `supabase_service_role`, or the actual key pattern
`eyJ...` used with createClient on the client side.
Remediation: Use service_role only in server-side code (API routes,
Edge Functions, server actions).
### DB-05: Connection pooling
Search for database connection configuration. Check for pool settings
(max, min, idle timeout). For Supabase, check if using connection
pooler URL (port 6543) vs direct (port 5432).
Remediation: Use connection pooling for production. For Supabase,
use the pooler URL. For raw pg, configure pool size based on expected
concurrent connections.
### DB-06: Migrations in version control
Check if a migrations directory exists (supabase/migrations, prisma/
migrations, alembic/versions, db/migrate). Verify it contains .sql
or migration files, not empty.
Remediation: Use your ORM or database tool's migration system. Never
make manual schema changes to production.
### DB-09: Anon key writes without RLS
Search for Supabase client-side inserts/updates using the anon key
without RLS policies protecting the target tables.
Remediation: Enable RLS on all tables and create INSERT/UPDATE policies
that scope access to authenticated users.
### DB-10: Storage bucket policies
Search Supabase migration files and dashboard config for storage
bucket creation. Verify each bucket has access policies defined.
Remediation: Define storage policies for each bucket. Restrict
uploads by file type, size, and user ownership.
### DB-12: PII stored unencrypted
Search schema files and migration files for columns named email,
phone, ssn, social_security, credit_card, address, date_of_birth
that are stored as plain text without encryption.
Remediation: Encrypt PII columns at rest. Use database-level
encryption or application-level encryption for sensitive fields.
---
## DEPLOY: Deployment
| ID | Check | Detection | Severity | Stack |
|----|-------|-----------|----------|-------|
| DEPLOY-01 | All env vars set on production server | manual | critical | all |
| DEPLOY-02 | SSL certificate installed and valid | manual | critical | all |
| DEPLOY-03 | Firewall configured (only 80/443 public) | manual | high | vps |
| DEPLOY-04 | Process manager running | manual | high | vps |
| DEPLOY-05 | Rollback plan exists | manual | high | all |
| DEPLOY-06 | Staging test passed before production | manual | high | all |
| DEPLOY-07 | Deploy does not cause downtime | manual | advisory | all |
| DEPLOY-08 | Domain DNS configured (www vs non-www) | manual | high | all |
| DEPLOY-09 | Health check endpoint exists | auto | high | all |
| DEPLOY-10 | Logging configured (structured, not console) | auto | high | all |
| DEPLOY-11 | Error monitoring connected (Sentry, etc.) | auto | advisory | all |
| DEPLOY-12 | Cron jobs and background tasks verified | manual | high | all |
| DEPLOY-13 | CDN configured for static assets | manual | advisory | all |
### DEPLOY-09: Health check endpoint
Search for a `/health`, `/healthz`, `/api/health`, or `/status` route
that returns a 200 response.
Remediation: Add a health check endpoint that verifies database
connectivity and returns a simple JSON response.
### DEPLOY-10: Structured logging
Check if the project uses a logging library (winston, pino, bunyan,
python logging module) vs raw console.log statements in server code.
Remediation: Replace console.log with a structured logger that outputs
JSON with timestamps and request IDs.
---
## CODE: Code Quality
| ID | Check | Detection | Severity | Stack |
|----|-------|-----------|----------|-------|
| CODE-01 | No console.log in production build | auto | high | js/ts |
| CODE-02 | Error handling on all async operations | auto | high | all |
| CODE-03 | No empty catch blocks | auto | high | all |
| CODE-04 | Loading and error states in UI | auto | high | react |
| CODE-05 | Pagination on all list endpoints | auto | high | all |
| CODE-06 | npm audit clean (zero critical) | tool | high | js/ts |
| CODE-07 | No TODO-auth or TODO-security patterns | auto | critical | all |
| CODE-08 | No unhandled promise rejections | auto | high | js/ts |
| CODE-09 | React error boundaries in place | auto | high | react |
| CODE-10 | No leaked stack traces in error responses | auto | high | all |
| CODE-11 | No eslint-disable on security rules | auto | high | js/ts |
| CODE-12 | Lockfile committed | auto | high | all |
| CODE-13 | No wildcard versions in package.json | auto | high | js/ts |
| CODE-14 | TypeScript strict mode enabled | auto | advisory | ts |
### CODE-01: No console.log in production
Search for `console.log`, `console.debug`, `console.info` in source
files (exclude test files, config files, and node_modules).
Remediation: Remove console.log statements or replace with a proper
logger. Use a build tool to strip them automatically.
### CODE-03: No empty catch blocks
Search for `catch` blocks with empty bodies or only a comment inside.
Pattern: `catch\s*\([^)]*\)\s*\{\s*(\/\/.*\n)?\s*\}`
Remediation: At minimum, log the error. Better: handle it appropriately
or rethrow.
### CODE-07: No TODO-auth/security patterns
Search for `TODO.*auth`, `TODO.*security`, `TODO.*permission`,
`FIXME.*auth`, `HACK.*auth`, `// auth`, `# TODO: add auth`.
These indicate security features that were deferred and forgotten.
Remediation: Implement the deferred security feature or remove the
endpoint if it is not ready.
### CODE-09: React error boundaries
Check if the app has at least one ErrorBoundary component or uses
a library like react-error-boundary. Check app/error.tsx for Next.js
App Router projects.
Remediation: Add error boundaries at layout boundaries to prevent
full-page crashes.
### CODE-02: Error handling on async operations
Search for async functions and .then() chains. Check if they have
corresponding try/catch or .catch() handlers.
Remediation: Wrap every async operation in try/catch. Log errors
and show appropriate UI feedback.
### CODE-04: Loading and error states in UI
Search React components for data fetching (useEffect with fetch,
useSWR, useQuery, server components) and check if they render
loading and error states.
Remediation: Add loading spinners/skeletons and error messages
for every data-dependent component.
### CODE-05: Pagination on list endpoints
Search API routes that return arrays/lists from database queries.
Check for LIMIT/OFFSET, cursor pagination, or take/skip parameters.
Remediation: Add pagination to every endpoint that returns a list.
Default page size of 20-50 items. Never return unbounded result sets.
### CODE-10: No leaked stack traces
Search error handling code for responses that include stack traces,
error.stack, or full error objects sent to the client.
Remediation: Return generic error messages to the client. Log full
stack traces server-side only.
### CODE-11: No eslint-disable on security rules
Search for eslint-disable comments that suppress security-related
rules (no-eval, no-implied-eval, no-script-url).
Remediation: Fix the underlying issue instead of disabling the lint
rule. If genuinely necessary, add a comment explaining why.
### CODE-14: TypeScript strict mode
Check tsconfig.json for `"strict": true` or the individual flags
(strictNullChecks, noImplicitAny, etc.).
Remediation: Enable strict mode in tsconfig.json. Fix type errors
incrementally if migrating an existing project.
---
## AI: AI/LLM Security
| ID | Check | Detection | Severity | Stack |
|----|-------|-----------|----------|-------|
| AI-01 | System prompts not leakable via user input | auto | critical | ai |
| AI-02 | No prompt injection vectors in user inputs | auto | critical | ai |
| AI-03 | LLM API keys not in frontend code | auto | critical | ai |
| AI-04 | Rate limiting on AI endpoints (cost protection) | auto | high | ai |
| AI-05 | AI response output sanitized before rendering | auto | high | ai |
| AI-06 | MCP server inputs validated | auto | high | ai |
| AI-07 | Agent permissions scoped (no unrestricted access) | manual | high | ai |
| AI-08 | No sensitive data sent to third-party LLMs without consent | manual | high | ai |
### AI-01: System prompt leakage
Search for system prompts stored in client-accessible files or returned
in API responses. Check if the AI endpoint echoes the system prompt
when asked "repeat your instructions" or similar.
Remediation: Keep system prompts server-side only. Add input filtering
for prompt extraction attempts.
### AI-03: LLM API keys not in frontend
Search frontend code for `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`,
`GOOGLE_AI_API_KEY`, `sk-ant-`, `sk-proj-`, `AIza` patterns.
Remediation: Proxy all LLM calls through server-side API routes.
---
## DEP: Dependencies and Supply Chain
| ID | Check | Detection | Severity | Stack |
|----|-------|-----------|----------|-------|
| DEP-01 | No git:// or URL-based dependencies | auto | high | all |
| DEP-02 | No typosquatting risk (verify package names) | auto | advisory | all |
| DEP-03 | Lockfile integrity verified | auto | high | all |
| DEP-04 | npm audit / pip audit zero critical | tool | high | all |
| DEP-05 | No suspicious postinstall scripts | auto | high | js/ts |
| DEP-06 | Dependencies pinned (no wildcard *) | auto | high | all |
| DEP-07 | Lockfile committed to version control | auto | high | all |
### DEP-01: No git/URL dependencies
Search package.json for dependencies with values starting with
`git://`, `git+`, `http://`, `https://github.com`, or `file:`.
Remediation: Use published npm packages with version ranges instead
of git URLs.
### DEP-05: Suspicious postinstall scripts
Check package.json for `postinstall`, `preinstall`, `install` scripts
that execute arbitrary commands, download files, or access the network.
Remediation: Review and remove unnecessary install scripts. Use
`--ignore-scripts` for CI.
---
## FE: Frontend Quality
| ID | Check | Detection | Severity | Stack |
|----|-------|-----------|----------|-------|
| FE-01 | Meta tags present (title, description, OG tags) | auto | advisory | web |
| FE-02 | Favicon configured | auto | advisory | web |
| FE-03 | Custom 404 page exists | auto | advisory | web |
| FE-04 | Responsive design tested on mobile | manual | high | web |
| FE-05 | Alt text on images | auto | high | web |
| FE-06 | Keyboard navigation works | manual | high | web |
| FE-07 | Forms have validation feedback | auto | high | web |
| FE-08 | Analytics installed (production only) | auto | advisory | web |
| FE-09 | robots.txt present | auto | advisory | web |
| FE-10 | Images optimized (WebP, lazy loading) | auto | advisory | web |
### FE-01: Meta tags
Check the root layout or index page for `<title>`, `<meta name="description">`,
and Open Graph tags (`og:title`, `og:description`, `og:image`).
For Next.js, check metadata export in layout.tsx.
Remediation: Add metadata to your root layout or page head.
### FE-03: Custom 404 page
Check for `404.tsx`, `404.jsx`, `not-found.tsx`, `404.html`, or
equivalent in the pages/app directory.
Remediation: Create a branded 404 page that helps users navigate back.
---
## OBS: Observability
| ID | Check | Detection | Severity | Stack |
|----|-------|-----------|----------|-------|
| OBS-01 | Error monitoring configured (Sentry, LogRocket, etc.) | auto | advisory | all |
| OBS-02 | Alerting set up for critical failures | manual | high | all |
| OBS-03 | Structured logging with request IDs | auto | advisory | all |
| OBS-04 | Performance baseline established | manual | advisory | all |
| OBS-05 | Uptime monitoring configured | manual | high | all |
| OBS-06 | Error rates tracked | manual | advisory | all |
| OBS-07 | Log retention policy defined | manual | advisory | all |
### OBS-01: Error monitoring
Search for imports or configuration of error monitoring tools:
`@sentry/`, `LogRocket`, `Bugsnag`, `Datadog`, `Rollbar`, `Honeybadger`.
Remediation: Install and configure an error monitoring service.
Sentry has a free tier suitable for solo projects.
FILE:references/patterns.md
# Ship Gate: Detection Patterns
Grep and regex patterns for auto-scannable checks. Claude runs these
against the codebase to detect issues.
## Table of Contents
- SEC: Security Patterns
- DB: Database Patterns
- CODE: Code Quality Patterns
- AI: AI/LLM Security Patterns
- DEP: Dependency Patterns
- FE: Frontend Quality Patterns
- OBS: Observability Patterns
- DEPLOY: Deployment Patterns
All patterns use `grep -rn` with `--include` filters. Exclude
node_modules, .next, dist, build, .git, __pycache__, venv directories
from all scans.
Base exclude flags:
```bash
EXCLUDE="--exclude-dir=node_modules --exclude-dir=.next --exclude-dir=dist --exclude-dir=build --exclude-dir=.git --exclude-dir=__pycache__ --exclude-dir=venv --exclude-dir=.venv --exclude-dir=vendor --exclude-dir=coverage"
```
---
## SEC: Security Patterns
### SEC-01: Secrets in frontend code
Scan directories that serve client-side code:
```bash
# Generic API key patterns
grep -rnE $EXCLUDE \
"(sk-[a-zA-Z0-9]{20,}|sk-ant-[a-zA-Z0-9-]+|sk-proj-[a-zA-Z0-9-]+|AIza[a-zA-Z0-9_-]{35}|ghp_[a-zA-Z0-9]{36}|glpat-[a-zA-Z0-9_-]{20,}|xox[bsap]-[a-zA-Z0-9-]+)" \
src/ app/ pages/ components/ public/ lib/ utils/ 2>/dev/null
# AWS keys
grep -rnE $EXCLUDE \
"AKIA[0-9A-Z]{16}" \
src/ app/ pages/ components/ public/ 2>/dev/null
# Stripe keys (live, not test)
grep -rnE $EXCLUDE \
"sk_live_[a-zA-Z0-9]{24,}" \
src/ app/ pages/ components/ public/ 2>/dev/null
# Generic secret assignment
grep -rnE $EXCLUDE \
"(api_key|apikey|api_secret|secret_key|auth_token|access_token)\s*[:=]\s*['\"][a-zA-Z0-9_-]{16,}" \
src/ app/ pages/ components/ public/ 2>/dev/null
```
### SEC-04: CORS wildcard
```bash
grep -rnE $EXCLUDE \
"(origin:\s*['\"]?\*['\"]?|Access-Control-Allow-Origin.*\*|cors\(\s*\))" \
. 2>/dev/null
```
### SEC-05: CSRF protection missing
```bash
# Check for state-changing routes without CSRF
grep -rnE $EXCLUDE \
"(app\.(post|put|patch|delete)|router\.(post|put|patch|delete))" \
. 2>/dev/null
# Then verify csrf middleware exists
grep -rnE $EXCLUDE \
"(csrf|csrfToken|_csrf|CSRF_COOKIE)" \
. 2>/dev/null
```
### SEC-08: Weak password hashing
```bash
# Check for weak hashing (md5, sha1, sha256 for passwords)
grep -rnE $EXCLUDE \
"(md5|sha1|sha256)\s*\(" \
. 2>/dev/null
# Verify bcrypt/argon2 usage
grep -rnE $EXCLUDE \
"(bcrypt|argon2|scrypt)" \
. 2>/dev/null
```
### SEC-11: CSP headers
```bash
# Check for Content-Security-Policy configuration
grep -rnE $EXCLUDE \
"(Content-Security-Policy|contentSecurityPolicy|csp)" \
. 2>/dev/null
# Next.js: check next.config for headers
grep -rn $EXCLUDE \
"Content-Security-Policy" \
next.config.* 2>/dev/null
```
### SEC-13: Unsafe eval/innerHTML
```bash
# eval usage
grep -rnE $EXCLUDE \
"(\beval\s*\(|new\s+Function\s*\()" \
--include="*.js" --include="*.ts" --include="*.jsx" --include="*.tsx" \
. 2>/dev/null
# dangerouslySetInnerHTML without sanitizer
grep -rnE $EXCLUDE \
"dangerouslySetInnerHTML" \
--include="*.jsx" --include="*.tsx" \
. 2>/dev/null
# Then check if DOMPurify or similar is imported in same file
```
### SEC-15: Cookie security flags
```bash
grep -rnE $EXCLUDE \
"(set-cookie|setCookie|cookie\()" \
. 2>/dev/null
# Verify HttpOnly, Secure, SameSite flags are present
grep -rnE $EXCLUDE \
"(httpOnly|HttpOnly|secure:\s*true|sameSite)" \
. 2>/dev/null
```
### SEC-06: Input validation
```bash
# Check for validation library usage
grep -rnE $EXCLUDE \
"(from 'zod'|from 'yup'|from 'joi'|from 'class-validator'|from pydantic)" \
. 2>/dev/null
# Check for raw req.body usage without validation
grep -rnE $EXCLUDE \
"(req\.body\.|request\.json|request\.form)" \
--include="*.ts" --include="*.js" --include="*.py" \
. 2>/dev/null
```
### SEC-07: Rate limiting
```bash
grep -rnE $EXCLUDE \
"(express-rate-limit|@upstash/ratelimit|rate-limiter|slowapi|throttle)" \
package.json requirements.txt . 2>/dev/null
```
### SEC-09: Token expiry
```bash
grep -rnE $EXCLUDE \
"(sign\(|jwt\.encode|createToken|signToken)" \
--include="*.ts" --include="*.js" --include="*.py" \
. 2>/dev/null
# Then check if expiresIn/exp is set in those calls
grep -rnE $EXCLUDE \
"(expiresIn|exp:|expires_in|expires_delta)" \
. 2>/dev/null
```
### SEC-14: Sensitive data in URLs/logs
```bash
# Sensitive query parameters
grep -rnE $EXCLUDE \
"(password|token|secret|key|ssn|credit.card)=" \
--include="*.ts" --include="*.js" --include="*.py" \
. 2>/dev/null
# Logging full request objects
grep -rnE $EXCLUDE \
"console\.(log|info|debug)\s*\(\s*(req|request)\s*\)" \
. 2>/dev/null
```
### SEC-16: File upload validation
```bash
grep -rnE $EXCLUDE \
"(multer|formidable|busboy|UploadedFile|upload\.single|upload\.array)" \
. 2>/dev/null
# Check for file type/size validation near upload handlers
grep -rnE $EXCLUDE \
"(fileFilter|limits|maxFileSize|allowedTypes|mimetype)" \
. 2>/dev/null
```
### SEC-17/18: .env in repo
```bash
# Check if .env files exist in working tree
find . -maxdepth 3 -name ".env*" -not -path "*/node_modules/*" \
-not -name ".env.example" -not -name ".env.sample" 2>/dev/null
# Check if .env is in .gitignore
grep -n "\.env" .gitignore 2>/dev/null
# Check git history for .env commits
git log --all --name-only --diff-filter=A 2>/dev/null | grep "\.env" || true
```
---
## DB: Database Patterns
### DB-03: SQL injection (string concatenation)
```bash
# Template literal SQL
grep -rnE $EXCLUDE \
"(SELECT|INSERT|UPDATE|DELETE|FROM|WHERE).*\$\{" \
--include="*.js" --include="*.ts" --include="*.jsx" --include="*.tsx" \
. 2>/dev/null
# String concat SQL
grep -rnE $EXCLUDE \
"(SELECT|INSERT|UPDATE|DELETE|FROM|WHERE).*\+\s*(req\.|params\.|body\.|query\.)" \
. 2>/dev/null
# Python f-string SQL
grep -rnE $EXCLUDE \
"f['\"].*\b(SELECT|INSERT|UPDATE|DELETE|FROM|WHERE)\b.*\{" \
--include="*.py" \
. 2>/dev/null
```
### DB-07: Supabase RLS
```bash
# Find CREATE TABLE without RLS
grep -rnl $EXCLUDE "CREATE TABLE" \
--include="*.sql" . 2>/dev/null | while read f; do
tables=$(grep -oP "CREATE TABLE\s+\K\S+" "$f")
for t in $tables; do
if ! grep -q "ENABLE ROW LEVEL SECURITY" "$f" || \
! grep -q "$t" <<< "$(grep 'ENABLE ROW LEVEL SECURITY' "$f")"; then
echo "FAIL: $f - table $t missing RLS"
fi
done
done
```
### DB-08: service_role in client code
```bash
grep -rnE $EXCLUDE \
"(service_role|serviceRole|SUPABASE_SERVICE_ROLE)" \
src/ app/ pages/ components/ public/ lib/client 2>/dev/null
```
### DB-05: Connection pooling
```bash
# Check for pool configuration
grep -rnE $EXCLUDE \
"(pool|connectionLimit|max_connections|poolSize)" \
--include="*.ts" --include="*.js" --include="*.py" --include="*.env*" \
. 2>/dev/null
# Supabase: check if using pooler port
grep -rnE $EXCLUDE \
"(6543|pooler)" \
--include="*.env*" --include="*.ts" --include="*.js" \
. 2>/dev/null
```
### DB-06: Migrations in version control
```bash
# Check for migration directories
find . -maxdepth 3 -type d \
\( -name "migrations" -o -name "migrate" -o -name "versions" \) \
-not -path "*/node_modules/*" 2>/dev/null
# Check if migrations contain files
find . -path "*/migrations/*.sql" -o -path "*/migrations/*.ts" \
-o -path "*/migrations/*.py" 2>/dev/null | head -5
```
### DB-12: PII stored unencrypted
```bash
# Search schema files for PII column names
grep -rnEi $EXCLUDE \
"(ssn|social_security|credit_card|card_number|passport)" \
--include="*.sql" --include="*.prisma" --include="*.py" \
. 2>/dev/null
```
---
## CODE: Code Quality Patterns
### CODE-01: console.log in production
```bash
grep -rnE $EXCLUDE \
"console\.(log|debug|info)\(" \
--include="*.js" --include="*.ts" --include="*.jsx" --include="*.tsx" \
--exclude="*.test.*" --exclude="*.spec.*" --exclude="*.config.*" \
src/ app/ pages/ components/ lib/ utils/ 2>/dev/null
```
### CODE-03: Empty catch blocks
```bash
grep -rnPzo $EXCLUDE \
"catch\s*\([^)]*\)\s*\{\s*\}" \
--include="*.js" --include="*.ts" --include="*.jsx" --include="*.tsx" \
. 2>/dev/null
```
### CODE-07: TODO-auth patterns
```bash
grep -rnEi $EXCLUDE \
"(TODO|FIXME|HACK|XXX).*(auth|security|permission|validation|sanitiz)" \
. 2>/dev/null
```
### CODE-08: Unhandled promise rejections
```bash
# Async functions without try-catch
grep -rnE $EXCLUDE \
"async\s+\w+\s*\(" \
--include="*.js" --include="*.ts" --include="*.jsx" --include="*.tsx" \
. 2>/dev/null
# Check for .catch() or try/catch wrapping
```
### CODE-09: React error boundaries
```bash
# Check for error boundary in Next.js App Router
find . -path "*/app/error.tsx" -o -path "*/app/error.jsx" \
-o -path "*/app/global-error.tsx" 2>/dev/null
# Check for ErrorBoundary component
grep -rnE $EXCLUDE \
"(ErrorBoundary|error-boundary|componentDidCatch|getDerivedStateFromError)" \
--include="*.jsx" --include="*.tsx" \
. 2>/dev/null
```
### CODE-12: Lockfile committed
```bash
# Check for lockfile existence
ls package-lock.json pnpm-lock.yaml yarn.lock bun.lockb \
Pipfile.lock poetry.lock Gemfile.lock go.sum Cargo.lock 2>/dev/null
# Check if lockfile is gitignored
for f in package-lock.json pnpm-lock.yaml yarn.lock; do
if git check-ignore "$f" 2>/dev/null; then
echo "FAIL: $f is gitignored"
fi
done
```
### CODE-13: Wildcard versions
```bash
# Check for * or empty version in package.json
grep -nE '"[^"]+"\s*:\s*"\*"' package.json 2>/dev/null
```
### CODE-02: Async without error handling
```bash
# Find async functions
grep -rnE $EXCLUDE \
"async\s+(function\s+)?\w+\s*\(" \
--include="*.ts" --include="*.js" --include="*.tsx" --include="*.jsx" \
. 2>/dev/null
# Count try/catch usage nearby
grep -rnc $EXCLUDE "try\s*{" \
--include="*.ts" --include="*.js" --include="*.tsx" --include="*.jsx" \
. 2>/dev/null
```
### CODE-04: Loading and error states
```bash
# Check for loading state patterns in React
grep -rnE $EXCLUDE \
"(isLoading|loading|Skeleton|Spinner|fallback)" \
--include="*.tsx" --include="*.jsx" \
. 2>/dev/null
# Check for Suspense boundaries
grep -rnE $EXCLUDE \
"(<Suspense|loading\.tsx|loading\.jsx)" \
. 2>/dev/null
```
### CODE-05: Pagination on list endpoints
```bash
# Check API routes for unbounded queries
grep -rnE $EXCLUDE \
"(\.findMany|\.find\(\)|\.select\(\)|SELECT \*)" \
--include="*.ts" --include="*.js" --include="*.py" \
. 2>/dev/null
# Check for pagination parameters
grep -rnE $EXCLUDE \
"(limit|offset|page|skip|take|cursor|per_page)" \
--include="*.ts" --include="*.js" --include="*.py" \
. 2>/dev/null
```
### CODE-10: Leaked stack traces
```bash
grep -rnE $EXCLUDE \
"(error\.stack|\.stack\)|err\.message.*res\.(json|send)|traceback)" \
--include="*.ts" --include="*.js" --include="*.py" \
. 2>/dev/null
```
### CODE-11: eslint-disable on security rules
```bash
grep -rnE $EXCLUDE \
"eslint-disable.*(no-eval|no-implied-eval|no-script-url|security)" \
--include="*.ts" --include="*.js" --include="*.tsx" --include="*.jsx" \
. 2>/dev/null
```
### CODE-14: TypeScript strict mode
```bash
grep -n '"strict"' tsconfig.json 2>/dev/null
# Check if strict is true
grep -n '"strict":\s*true' tsconfig.json 2>/dev/null
```
---
## AI: AI/LLM Security Patterns
### AI-01: System prompt leakage
```bash
# System prompts in client-accessible files
grep -rnEi $EXCLUDE \
"(system.?prompt|system.?message|system_instruction)" \
src/ app/ pages/ components/ public/ 2>/dev/null
# System prompts returned in API responses
grep -rnE $EXCLUDE \
"(system.*role|role.*system)" \
src/ app/ pages/ components/ public/ 2>/dev/null
```
### AI-02: Prompt injection vectors
```bash
# User input concatenated directly into prompts
grep -rnE $EXCLUDE \
"(messages\.push|content:.*\$\{|content:.*\+\s*user|prompt.*\+)" \
--include="*.ts" --include="*.js" --include="*.py" \
. 2>/dev/null
```
### AI-03: LLM API keys in frontend
```bash
grep -rnE $EXCLUDE \
"(OPENAI_API_KEY|ANTHROPIC_API_KEY|GOOGLE_AI_API_KEY|sk-ant-|sk-proj-|AIza[a-zA-Z0-9_-]{35})" \
src/ app/ pages/ components/ public/ 2>/dev/null
```
### AI-04: Rate limiting on AI endpoints
```bash
# Find AI-related API routes
grep -rnlE $EXCLUDE \
"(openai|anthropic|claude|gpt|completion|chat/api|ai/api)" \
--include="*.ts" --include="*.js" \
. 2>/dev/null
# Then check for rate limiting middleware in those files
```
### AI-05: AI output sanitization
```bash
# Check if AI responses are rendered with dangerouslySetInnerHTML
grep -rnE $EXCLUDE \
"dangerouslySetInnerHTML.*\b(response|result|completion|message|content)\b" \
--include="*.tsx" --include="*.jsx" \
. 2>/dev/null
```
### AI-06: MCP server input validation
```bash
# Check MCP server tool handlers for input validation
grep -rnE $EXCLUDE \
"(tool_input|toolInput|tool_call|CallToolRequest)" \
--include="*.ts" --include="*.js" --include="*.py" \
. 2>/dev/null
# Check if zod/validation is applied to tool inputs
```
---
## DEP: Dependency Patterns
### DEP-01: Git/URL dependencies
```bash
grep -nE '"(git|git\+|http|https|file):' package.json 2>/dev/null
grep -nE '"github:' package.json 2>/dev/null
```
### DEP-04: npm audit
```bash
# Run npm audit and capture critical/high counts
npm audit --json 2>/dev/null | grep -c '"severity":"critical"'
npm audit --json 2>/dev/null | grep -c '"severity":"high"'
# Or for pip
pip audit --format json 2>/dev/null
```
### DEP-05: Suspicious install scripts
```bash
grep -A2 '"preinstall"\|"postinstall"\|"install"' package.json 2>/dev/null
```
### DEP-06: Wildcard versions
```bash
grep -nE '"\*"' package.json 2>/dev/null
grep -nE '"latest"' package.json 2>/dev/null
```
---
## FE: Frontend Quality Patterns
### FE-01: Meta tags
```bash
# Next.js App Router metadata
grep -rnE $EXCLUDE \
"(export\s+(const|async\s+function)\s+metadata|generateMetadata)" \
--include="*.tsx" --include="*.ts" \
app/layout.* app/page.* 2>/dev/null
# HTML meta tags
grep -rnE $EXCLUDE \
'(<title>|<meta\s+name="description"|og:title|og:description|og:image)' \
. 2>/dev/null
```
### FE-02: Favicon
```bash
find . -maxdepth 3 \( -name "favicon.*" -o -name "icon.*" \) \
-not -path "*/node_modules/*" 2>/dev/null
```
### FE-03: Custom 404 page
```bash
find . -maxdepth 4 \( -name "404.*" -o -name "not-found.*" \) \
-not -path "*/node_modules/*" 2>/dev/null
```
### FE-05: Image alt text
```bash
# Find img tags without alt attribute
grep -rnE $EXCLUDE \
'<img\s+(?![^>]*\balt\b)[^>]*>' \
--include="*.html" --include="*.jsx" --include="*.tsx" \
. 2>/dev/null
# Next.js Image without alt
grep -rnE $EXCLUDE \
'<Image\s+(?![^>]*\balt\b)[^>]*/?>' \
--include="*.jsx" --include="*.tsx" \
. 2>/dev/null
```
### FE-09: robots.txt
```bash
find . -maxdepth 2 -name "robots.txt" \
-not -path "*/node_modules/*" 2>/dev/null
```
### FE-07: Form validation feedback
```bash
# Check for form elements without validation attributes
grep -rnE $EXCLUDE \
'(<input|<textarea|<select)' \
--include="*.tsx" --include="*.jsx" --include="*.html" \
. 2>/dev/null
# Check for validation library usage
grep -rnE $EXCLUDE \
"(useForm|react-hook-form|formik|yup|zod.*form)" \
--include="*.tsx" --include="*.jsx" \
. 2>/dev/null
```
### FE-10: Image optimization
```bash
# Check for unoptimized img tags (not using Next/Image or similar)
grep -rnE $EXCLUDE \
'<img\s' \
--include="*.tsx" --include="*.jsx" \
. 2>/dev/null
# Check for lazy loading
grep -rnE $EXCLUDE \
'(loading="lazy"|lazy|lazyload)' \
--include="*.tsx" --include="*.jsx" --include="*.html" \
. 2>/dev/null
```
---
## OBS: Observability Patterns
### OBS-01: Error monitoring
```bash
grep -rnE $EXCLUDE \
"(@sentry|sentry-|LogRocket|Bugsnag|datadogRum|Rollbar|Honeybadger|newrelic)" \
package.json . 2>/dev/null
```
### OBS-03: Structured logging
```bash
# Check for logging libraries
grep -rnE $EXCLUDE \
"(winston|pino|bunyan|morgan|log4js)" \
package.json 2>/dev/null
# Python
grep -rnE $EXCLUDE \
"import logging|from loguru" \
--include="*.py" . 2>/dev/null
```
---
## DEPLOY: Deployment Patterns
### DEPLOY-09: Health check endpoint
```bash
grep -rnE $EXCLUDE \
"(\/health|\/healthz|\/api\/health|\/status|\/readyz)" \
--include="*.ts" --include="*.js" --include="*.py" \
. 2>/dev/null
```
### DEPLOY-10: Console vs structured logging (server)
```bash
# Count console.log vs logger usage in API/server code
echo "console.log count:"
grep -rnc $EXCLUDE "console\.log" \
--include="*.ts" --include="*.js" \
api/ server/ pages/api/ app/api/ 2>/dev/null | tail -1
echo "structured logger count:"
grep -rnc $EXCLUDE "(logger\.|log\.(info|warn|error|debug))" \
--include="*.ts" --include="*.js" \
api/ server/ pages/api/ app/api/ 2>/dev/null | tail -1
```
FILE:scripts/ship_gate_scanner.py
#!/usr/bin/env python3
"""
ship_gate_scanner.py — Pre-production audit CLI
Part of the ship-gate skill: https://github.com/rx4u/ship-gate
Usage:
python scripts/ship_gate_scanner.py [PATH] [options]
Options:
--json Output results as JSON
--no-color Disable ANSI color output
--no-interactive Skip manual confirmation prompts
--category CAT Only run a specific category (SEC, DB, CODE, etc.)
--verbose Show PASS results in addition to FAIL
--version Show version and exit
Exit codes:
0 = CLEAR TO SHIP (no critical issues)
1 = DO NOT SHIP (critical issues found)
2 = SHIP WITH CAUTION (high issues only)
"""
import argparse
import json
import os
import re
import sys
import time
from dataclasses import dataclass, field
from enum import Enum
from pathlib import Path
from typing import List, Optional
VERSION = "1.0.0"
EXCLUDE_DIRS = {
"node_modules", ".next", "dist", "build", ".git", "__pycache__",
"venv", ".venv", "vendor", "coverage", ".turbo", "out", ".cache",
".pytest_cache", ".mypy_cache", "target", "bin", "obj",
}
FRONTEND_DIRS = {"src", "app", "pages", "components", "public", "lib", "utils"}
JS_EXTS = {".js", ".ts", ".jsx", ".tsx", ".mjs", ".cjs"}
PY_EXTS = {".py"}
ALL_CODE_EXTS = JS_EXTS | PY_EXTS | {".go", ".rb", ".php"}
TEMPLATE_EXTS = {".html", ".jsx", ".tsx", ".vue", ".svelte"}
SQL_EXTS = {".sql", ".prisma"}
# ---------------------------------------------------------------------------
# ANSI helpers
# ---------------------------------------------------------------------------
USE_COLOR = True
def _c(code: str, text: str) -> str:
if not USE_COLOR:
return text
return f"\033[{code}m{text}\033[0m"
def red(t): return _c("31", t)
def green(t): return _c("32", t)
def yellow(t): return _c("33", t)
def cyan(t): return _c("36", t)
def bold(t): return _c("1", t)
def dim(t): return _c("2", t)
# ---------------------------------------------------------------------------
# Data model
# ---------------------------------------------------------------------------
class Status(str, Enum):
PASS = "PASS"
FAIL = "FAIL"
SKIP = "SKIP"
MANUAL = "MANUAL"
class Severity(str, Enum):
CRITICAL = "CRITICAL"
HIGH = "HIGH"
ADVISORY = "ADVISORY"
@dataclass
class Finding:
file: str
line: int
snippet: str = ""
@dataclass
class CheckDef:
id: str
description: str
severity: Severity
category: str
stack: str = "all" # "all", "js", "ts", "react", "supabase", "ai", "web", "vps"
@dataclass
class Result:
check: CheckDef
status: Status
message: str = ""
findings: List[Finding] = field(default_factory=list)
@dataclass
class Stack:
has_node: bool = False
framework: str = "" # next, react, vue, svelte, astro, express, fastify, hono
has_python: bool = False
py_framework: str = "" # django, flask, fastapi
has_go: bool = False
has_rust: bool = False
has_supabase: bool = False
has_typescript: bool = False
has_react: bool = False
deploy_target: str = "" # vercel, netlify, docker, fly, railway
has_ai: bool = False
ai_providers: List[str] = field(default_factory=list)
is_web: bool = False
# ---------------------------------------------------------------------------
# File walking / grep helpers
# ---------------------------------------------------------------------------
def walk_files(root: str, exts: Optional[set] = None, dirs: Optional[set] = None):
"""Yield (filepath, relpath) for all files under root, skipping EXCLUDE_DIRS."""
for dirpath, dirnames, filenames in os.walk(root):
dirnames[:] = [d for d in dirnames if d not in EXCLUDE_DIRS]
if dirs is not None:
rel = os.path.relpath(dirpath, root)
top = rel.split(os.sep)[0]
if rel != "." and top not in dirs:
dirnames[:] = []
continue
for fname in filenames:
if exts is None or os.path.splitext(fname)[1].lower() in exts:
fpath = os.path.join(dirpath, fname)
yield fpath, os.path.relpath(fpath, root)
def grep_files(
root: str,
pattern: str,
exts: Optional[set] = None,
dirs: Optional[set] = None,
flags: int = 0,
max_findings: int = 20,
exclude_patterns: Optional[List[str]] = None,
) -> List[Finding]:
"""Return up to max_findings matches across the codebase."""
try:
rx = re.compile(pattern, flags)
except re.error:
return []
exclude_rxs = []
if exclude_patterns:
for ep in exclude_patterns:
try:
exclude_rxs.append(re.compile(ep))
except re.error:
pass
results: List[Finding] = []
for fpath, relpath in walk_files(root, exts, dirs):
if any(seg in fpath for seg in (".test.", ".spec.", ".config.")):
if exts and exts <= JS_EXTS:
skip = True
# still yield for config-specific checks
if "tsconfig" in fpath or "package.json" in fpath:
skip = False
if skip:
continue
try:
with open(fpath, "r", encoding="utf-8", errors="ignore") as fh:
for lineno, line in enumerate(fh, 1):
if rx.search(line):
if any(ex.search(line) for ex in exclude_rxs):
continue
results.append(Finding(
file=relpath,
line=lineno,
snippet=line.rstrip()[:120],
))
if len(results) >= max_findings:
return results
except (OSError, PermissionError):
continue
return results
def file_exists_in(root: str, *names: str) -> Optional[str]:
"""Return the first found path among names (searched recursively up to depth 5)."""
for dirpath, dirnames, filenames in os.walk(root):
dirnames[:] = [d for d in dirnames if d not in EXCLUDE_DIRS]
depth = dirpath.replace(root, "").count(os.sep)
if depth >= 5:
dirnames[:] = []
continue
for fname in filenames:
if fname in names:
return os.path.join(dirpath, fname)
return None
def read_json_file(path: str) -> dict:
try:
with open(path) as f:
return json.load(f)
except Exception:
return {}
# ---------------------------------------------------------------------------
# Stack detection
# ---------------------------------------------------------------------------
def detect_stack(root: str) -> Stack:
s = Stack()
pkg_path = os.path.join(root, "package.json")
if os.path.isfile(pkg_path):
s.has_node = True
pkg = read_json_file(pkg_path)
all_deps = {}
for key in ("dependencies", "devDependencies", "peerDependencies"):
all_deps.update(pkg.get(key, {}))
if "next" in all_deps: s.framework = "next"
elif "react" in all_deps: s.framework = "react"
elif "vue" in all_deps: s.framework = "vue"
elif "svelte" in all_deps: s.framework = "svelte"
elif "astro" in all_deps: s.framework = "astro"
elif "express" in all_deps: s.framework = "express"
elif "fastify" in all_deps: s.framework = "fastify"
elif "hono" in all_deps: s.framework = "hono"
s.has_react = s.framework in ("next", "react")
s.is_web = s.framework in ("next", "react", "vue", "svelte", "astro")
if "@supabase/supabase-js" in all_deps:
s.has_supabase = True
if "typescript" in all_deps or os.path.isfile(os.path.join(root, "tsconfig.json")):
s.has_typescript = True
for ai_pkg in ("openai", "@anthropic-ai/sdk", "@google/generative-ai",
"ai", "@huggingface/inference"):
if ai_pkg in all_deps:
s.has_ai = True
s.ai_providers.append(ai_pkg)
if os.path.isdir(os.path.join(root, "supabase")):
s.has_supabase = True
for pyfile in ("requirements.txt", "pyproject.toml", "Pipfile", "setup.py"):
if os.path.isfile(os.path.join(root, pyfile)):
s.has_python = True
try:
content = open(os.path.join(root, pyfile)).read().lower()
if "django" in content: s.py_framework = "django"
elif "flask" in content: s.py_framework = "flask"
elif "fastapi" in content: s.py_framework = "fastapi"
except Exception:
pass
break
if os.path.isfile(os.path.join(root, "go.mod")):
s.has_go = True
if os.path.isfile(os.path.join(root, "Cargo.toml")):
s.has_rust = True
if os.path.isfile(os.path.join(root, "vercel.json")) or \
os.path.isdir(os.path.join(root, ".vercel")):
s.deploy_target = "vercel"
elif os.path.isfile(os.path.join(root, "netlify.toml")):
s.deploy_target = "netlify"
elif os.path.isfile(os.path.join(root, "fly.toml")):
s.deploy_target = "fly"
elif os.path.isfile(os.path.join(root, "railway.json")):
s.deploy_target = "railway"
elif os.path.isfile(os.path.join(root, "Dockerfile")):
s.deploy_target = "docker"
return s
# ---------------------------------------------------------------------------
# Check definitions
# ---------------------------------------------------------------------------
CHECKS = {
# SEC
"SEC-01": CheckDef("SEC-01", "No API keys or secrets in frontend code", Severity.CRITICAL, "SEC"),
"SEC-04": CheckDef("SEC-04", "CORS not wildcard", Severity.CRITICAL, "SEC"),
"SEC-05": CheckDef("SEC-05", "CSRF protection on state-changing endpoints", Severity.CRITICAL, "SEC"),
"SEC-06": CheckDef("SEC-06", "Input validated and sanitized server-side", Severity.HIGH, "SEC"),
"SEC-07": CheckDef("SEC-07", "Rate limiting on auth and sensitive endpoints", Severity.HIGH, "SEC"),
"SEC-08": CheckDef("SEC-08", "Passwords hashed with bcrypt or argon2", Severity.CRITICAL, "SEC"),
"SEC-11": CheckDef("SEC-11", "CSP headers configured", Severity.HIGH, "SEC"),
"SEC-13": CheckDef("SEC-13", "No eval() or dangerouslySetInnerHTML without sanitization", Severity.HIGH, "SEC", stack="js"), # noqa: SEC-AUDITOR
"SEC-14": CheckDef("SEC-14", "No sensitive data in URLs or logs", Severity.HIGH, "SEC"),
"SEC-17": CheckDef("SEC-17", "No hardcoded secrets in .env committed to repo", Severity.CRITICAL, "SEC"),
"SEC-18": CheckDef("SEC-18", ".env files listed in .gitignore", Severity.CRITICAL, "SEC"),
# DB
"DB-03": CheckDef("DB-03", "Parameterized queries everywhere (no SQL injection)", Severity.CRITICAL, "DB"),
"DB-05": CheckDef("DB-05", "Connection pooling configured", Severity.HIGH, "DB"),
"DB-06": CheckDef("DB-06", "Migrations in version control", Severity.HIGH, "DB"),
"DB-07": CheckDef("DB-07", "RLS enabled on all Supabase tables", Severity.CRITICAL, "DB", stack="supabase"),
"DB-08": CheckDef("DB-08", "No service_role key in client-side code", Severity.CRITICAL, "DB", stack="supabase"),
"DB-12": CheckDef("DB-12", "No PII stored unencrypted", Severity.HIGH, "DB"),
# DEPLOY
"DEPLOY-09": CheckDef("DEPLOY-09", "Health check endpoint exists", Severity.HIGH, "DEPLOY"),
"DEPLOY-10": CheckDef("DEPLOY-10", "Structured logging (not raw console)", Severity.HIGH, "DEPLOY"),
# CODE
"CODE-01": CheckDef("CODE-01", "No console.log in production build", Severity.HIGH, "CODE", stack="js"),
"CODE-03": CheckDef("CODE-03", "No empty catch blocks", Severity.HIGH, "CODE"),
"CODE-07": CheckDef("CODE-07", "No TODO-auth or TODO-security patterns", Severity.CRITICAL, "CODE"),
"CODE-09": CheckDef("CODE-09", "React error boundaries in place", Severity.HIGH, "CODE", stack="react"),
"CODE-10": CheckDef("CODE-10", "No leaked stack traces in error responses", Severity.HIGH, "CODE"),
"CODE-11": CheckDef("CODE-11", "No eslint-disable on security rules", Severity.HIGH, "CODE", stack="js"),
"CODE-12": CheckDef("CODE-12", "Lockfile committed", Severity.HIGH, "CODE"),
"CODE-13": CheckDef("CODE-13", "No wildcard versions in package.json", Severity.HIGH, "CODE", stack="js"),
"CODE-14": CheckDef("CODE-14", "TypeScript strict mode enabled", Severity.ADVISORY, "CODE", stack="ts"),
# AI
"AI-01": CheckDef("AI-01", "System prompts not leakable via user input", Severity.CRITICAL, "AI", stack="ai"),
"AI-02": CheckDef("AI-02", "No prompt injection vectors in user inputs", Severity.CRITICAL, "AI", stack="ai"),
"AI-03": CheckDef("AI-03", "LLM API keys not in frontend code", Severity.CRITICAL, "AI", stack="ai"),
"AI-05": CheckDef("AI-05", "AI response output sanitized before rendering", Severity.HIGH, "AI", stack="ai"),
# DEP
"DEP-01": CheckDef("DEP-01", "No git:// or URL-based dependencies", Severity.HIGH, "DEP"),
"DEP-05": CheckDef("DEP-05", "No suspicious postinstall scripts", Severity.HIGH, "DEP", stack="js"),
"DEP-06": CheckDef("DEP-06", "Dependencies pinned (no wildcard *)", Severity.HIGH, "DEP"),
# FE
"FE-01": CheckDef("FE-01", "Meta tags present (title, description, OG)", Severity.ADVISORY, "FE", stack="web"),
"FE-02": CheckDef("FE-02", "Favicon configured", Severity.ADVISORY, "FE", stack="web"),
"FE-03": CheckDef("FE-03", "Custom 404 page exists", Severity.ADVISORY, "FE", stack="web"),
"FE-09": CheckDef("FE-09", "robots.txt present", Severity.ADVISORY, "FE", stack="web"),
# OBS
"OBS-01": CheckDef("OBS-01", "Error monitoring configured (Sentry, etc.)", Severity.ADVISORY, "OBS"),
"OBS-03": CheckDef("OBS-03", "Structured logging with request IDs", Severity.ADVISORY, "OBS"),
}
MANUAL_CHECKS = [
CheckDef("SEC-02", "Every route checks authentication", Severity.CRITICAL, "SEC"),
CheckDef("SEC-03", "HTTPS enforced, HTTP redirected", Severity.CRITICAL, "SEC"),
CheckDef("SEC-10", "Sessions invalidated on logout (server-side)", Severity.HIGH, "SEC"),
CheckDef("DB-01", "Backups configured and tested", Severity.CRITICAL, "DB"),
CheckDef("DB-02", "Backup restore tested (not just backup)", Severity.CRITICAL, "DB"),
CheckDef("DB-04", "Separate dev and production databases", Severity.HIGH, "DB"),
CheckDef("DB-11", "App uses a non-root DB user", Severity.HIGH, "DB"),
CheckDef("DEPLOY-01", "All env vars set on production server", Severity.CRITICAL, "DEPLOY"),
CheckDef("DEPLOY-02", "SSL certificate installed and valid", Severity.CRITICAL, "DEPLOY"),
CheckDef("DEPLOY-05", "Rollback plan exists", Severity.HIGH, "DEPLOY"),
CheckDef("DEPLOY-06", "Staging test passed before production", Severity.HIGH, "DEPLOY"),
CheckDef("AI-07", "Agent permissions scoped (no unrestricted access)", Severity.HIGH, "AI", stack="ai"),
CheckDef("AI-08", "No sensitive data sent to third-party LLMs without consent", Severity.HIGH, "AI", stack="ai"),
CheckDef("FE-04", "Responsive design tested on mobile", Severity.HIGH, "FE", stack="web"),
CheckDef("OBS-05", "Uptime monitoring configured", Severity.HIGH, "OBS"),
]
# ---------------------------------------------------------------------------
# Individual check implementations
# ---------------------------------------------------------------------------
def check_sec01(root, stack):
c = CHECKS["SEC-01"]
dirs = FRONTEND_DIRS & set(os.listdir(root))
patterns = [
r"sk-[a-zA-Z0-9]{20,}",
r"sk-ant-[a-zA-Z0-9-]+",
r"sk-proj-[a-zA-Z0-9-]+",
r"AIza[a-zA-Z0-9_-]{35}",
r"ghp_[a-zA-Z0-9]{36}",
r"glpat-[a-zA-Z0-9_-]{20,}",
r"AKIA[0-9A-Z]{16}",
r"sk_live_[a-zA-Z0-9]{24,}",
r"(api_key|apikey|api_secret|secret_key|auth_token)\s*[:=]\s*['\"][a-zA-Z0-9_\-]{16,}",
]
findings = []
for pat in patterns:
findings += grep_files(root, pat, exts=JS_EXTS | {".env", ".json"},
dirs=dirs if dirs else None, max_findings=5)
if findings:
return Result(c, Status.FAIL,
f"{len(findings)} potential secret(s) found in frontend/client code",
findings[:10])
return Result(c, Status.PASS)
def check_sec04(root, stack):
c = CHECKS["SEC-04"]
findings = grep_files(root, r"(origin\s*:\s*['\"]?\*['\"]?|Access-Control-Allow-Origin.*\*|cors\(\s*\))",
exts=ALL_CODE_EXTS)
if findings:
return Result(c, Status.FAIL, "CORS wildcard (*) detected", findings)
return Result(c, Status.PASS)
def check_sec05(root, stack):
c = CHECKS["SEC-05"]
# Check for state-changing routes
route_findings = grep_files(root, r"(app|router)\.(post|put|patch|delete)\s*\(",
exts=JS_EXTS)
if not route_findings:
return Result(c, Status.SKIP, "No Express-style routes found")
# Check for CSRF protection
csrf_findings = grep_files(root, r"(csrf|csrfToken|_csrf|CSRF_COOKIE|csurf)",
exts=ALL_CODE_EXTS)
if not csrf_findings:
return Result(c, Status.FAIL,
f"{len(route_findings)} state-changing route(s) found but no CSRF protection detected",
route_findings[:5])
return Result(c, Status.PASS)
def check_sec06(root, stack):
c = CHECKS["SEC-06"]
# Check for validation library
val_findings = grep_files(root,
r"(from ['\"]zod['\"]|from ['\"]yup['\"]|from ['\"]joi['\"]|from ['\"]class-validator['\"]|from pydantic|import pydantic)",
exts=ALL_CODE_EXTS)
if val_findings:
return Result(c, Status.PASS)
# Check if there are API routes that use req.body without validation
body_findings = grep_files(root, r"(req\.body|request\.json\(\)|request\.form)",
exts=ALL_CODE_EXTS)
if body_findings:
return Result(c, Status.FAIL,
"request body used without a validation library (zod/yup/joi/pydantic)",
body_findings[:5])
return Result(c, Status.SKIP, "No API route body handling detected")
def check_sec07(root, stack):
c = CHECKS["SEC-07"]
findings = grep_files(root,
r"(express-rate-limit|@upstash/ratelimit|rate-limiter-flexible|slowapi|throttle|rateLimit)",
exts=ALL_CODE_EXTS | {".json"})
if findings:
return Result(c, Status.PASS)
# Only fail if there are auth-related routes
auth_routes = grep_files(root, r"(login|signin|register|signup|forgot.password|reset.password)",
exts=ALL_CODE_EXTS)
if auth_routes:
return Result(c, Status.FAIL,
"Auth routes found but no rate-limiting library detected", auth_routes[:3])
return Result(c, Status.SKIP, "No auth routes detected")
def check_sec08(root, stack):
c = CHECKS["SEC-08"]
# Weak hash for passwords
weak = grep_files(root, r"\b(md5|sha1|sha256)\s*\(",
exts=ALL_CODE_EXTS,
exclude_patterns=[r"//.*\b(md5|sha1|sha256)\b"])
if weak:
return Result(c, Status.FAIL, "Weak hashing algorithm (md5/sha1/sha256) detected", weak)
strong = grep_files(root, r"(bcrypt|argon2|scrypt|pbkdf2)", exts=ALL_CODE_EXTS)
pw_fields = grep_files(root, r"(password|passwd)", exts=ALL_CODE_EXTS)
if pw_fields and not strong:
return Result(c, Status.FAIL, "Password fields found but no bcrypt/argon2/scrypt usage")
return Result(c, Status.PASS if strong or not pw_fields else Status.SKIP)
def check_sec11(root, stack):
c = CHECKS["SEC-11"]
findings = grep_files(root, r"(Content-Security-Policy|contentSecurityPolicy|[^a-z]csp[^a-z])",
exts=ALL_CODE_EXTS | {".json", ".toml", ".yaml", ".yml"})
if findings:
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No Content-Security-Policy configuration found")
def check_sec13(root, stack):
c = CHECKS["SEC-13"]
if not stack.has_node:
return Result(c, Status.SKIP, "Not a JS/TS project")
eval_findings = grep_files(root, r"(\beval\s*\(|new\s+Function\s*\()", exts=JS_EXTS)
dsi_findings = grep_files(root, r"dangerouslySetInnerHTML", exts=JS_EXTS)
# If dangerouslySetInnerHTML is used, check for DOMPurify
unsafe_dsi = []
for f in dsi_findings:
try:
content = open(os.path.join(root, f.file), errors="ignore").read()
if "DOMPurify" not in content and "sanitize" not in content.lower():
unsafe_dsi.append(f)
except Exception:
unsafe_dsi.append(f)
all_findings = eval_findings + unsafe_dsi # noqa: SEC-AUDITOR
if all_findings:
return Result(c, Status.FAIL, "Unsafe eval() or unsanitized dangerouslySetInnerHTML", all_findings) # noqa: SEC-AUDITOR
return Result(c, Status.PASS)
def check_sec14(root, stack):
c = CHECKS["SEC-14"]
url_findings = grep_files(root,
r"(password|token|secret|key|ssn|credit.card)=",
exts=ALL_CODE_EXTS)
log_findings = grep_files(root,
r"console\.(log|info|debug)\s*\(\s*(req|request)\s*\)",
exts=JS_EXTS)
findings = url_findings + log_findings
if findings:
return Result(c, Status.FAIL, "Sensitive data may appear in URLs or logs", findings[:5])
return Result(c, Status.PASS)
def check_sec17(root, stack):
c = CHECKS["SEC-17"]
# Check for .env files that are not .example/.sample
env_files = []
for entry in os.scandir(root):
name = entry.name
if name.startswith(".env") and name not in (".env.example", ".env.sample",
".env.template", ".env.local.example"):
if entry.is_file():
env_files.append(name)
if not env_files:
return Result(c, Status.PASS)
# Check if git-tracked
gitignore_path = os.path.join(root, ".gitignore")
if os.path.isfile(gitignore_path):
content = open(gitignore_path, errors="ignore").read()
if ".env" in content:
return Result(c, Status.PASS)
return Result(c, Status.FAIL,
f".env file(s) exist ({', '.join(env_files)}) and may not be gitignored",
[Finding(f, 0) for f in env_files])
def check_sec18(root, stack):
c = CHECKS["SEC-18"]
gitignore_path = os.path.join(root, ".gitignore")
if not os.path.isfile(gitignore_path):
return Result(c, Status.FAIL, ".gitignore file not found")
content = open(gitignore_path, errors="ignore").read()
if re.search(r"\.env", content):
return Result(c, Status.PASS)
return Result(c, Status.FAIL, ".env not listed in .gitignore")
def check_db03(root, stack):
c = CHECKS["DB-03"]
# Template literal SQL
tl_findings = grep_files(root,
r"(SELECT|INSERT|UPDATE|DELETE|FROM|WHERE).*\$\{",
exts=JS_EXTS)
# Python f-string SQL
py_findings = grep_files(root,
r'f["\'].*\b(SELECT|INSERT|UPDATE|DELETE|FROM|WHERE)\b.*\{',
exts=PY_EXTS)
# String concat SQL
concat_findings = grep_files(root,
r"(SELECT|INSERT|UPDATE|DELETE|FROM|WHERE).*\+\s*(req\.|params\.|body\.|query\.)",
exts=ALL_CODE_EXTS)
all_findings = tl_findings + py_findings + concat_findings
if all_findings:
return Result(c, Status.FAIL,
f"{len(all_findings)} potential SQL injection vector(s)", all_findings[:10])
return Result(c, Status.PASS)
def check_db05(root, stack):
c = CHECKS["DB-05"]
findings = grep_files(root,
r"(pool|connectionLimit|max_connections|poolSize|pooler|6543)",
exts=ALL_CODE_EXTS | {".env", ".env.local", ".env.production"})
if findings:
return Result(c, Status.PASS)
db_found = grep_files(root, r"(pg\.|postgres\.|mysql\.|mongoose\.)", exts=ALL_CODE_EXTS)
if db_found:
return Result(c, Status.FAIL, "Database usage detected but no connection pooling configured")
return Result(c, Status.SKIP, "No direct DB connection detected")
def check_db06(root, stack):
c = CHECKS["DB-06"]
migration_dirs = []
for dirpath, dirnames, filenames in os.walk(root):
dirnames[:] = [d for d in dirnames if d not in EXCLUDE_DIRS]
depth = dirpath.replace(root, "").count(os.sep)
if depth >= 4:
dirnames[:] = []
continue
for d in dirnames:
if d in ("migrations", "migrate", "versions", "alembic"):
migration_dirs.append(os.path.join(dirpath, d))
if migration_dirs:
return Result(c, Status.PASS)
# Check for database usage
db_found = grep_files(root, r"(prisma|supabase|mongoose|pg\.|sqlite)", exts=ALL_CODE_EXTS)
if db_found:
return Result(c, Status.FAIL, "Database usage found but no migrations directory detected")
return Result(c, Status.SKIP, "No database usage detected")
def check_db07(root, stack):
c = CHECKS["DB-07"]
if not stack.has_supabase:
return Result(c, Status.SKIP, "Not a Supabase project")
sql_findings = grep_files(root, r"CREATE TABLE", exts=SQL_EXTS)
if not sql_findings:
return Result(c, Status.SKIP, "No CREATE TABLE statements found in migrations")
rls_findings = grep_files(root, r"ENABLE ROW LEVEL SECURITY", exts=SQL_EXTS)
if not rls_findings:
return Result(c, Status.FAIL,
f"{len(sql_findings)} table(s) found but no RLS policies detected",
sql_findings[:5])
if len(rls_findings) < len(sql_findings):
return Result(c, Status.FAIL,
f"{len(sql_findings)} table(s) but only {len(rls_findings)} RLS statement(s) — some tables may lack RLS",
sql_findings[:5])
return Result(c, Status.PASS)
def check_db08(root, stack):
c = CHECKS["DB-08"]
if not stack.has_supabase:
return Result(c, Status.SKIP, "Not a Supabase project")
dirs = FRONTEND_DIRS & set(os.listdir(root))
findings = grep_files(root,
r"(service_role|serviceRole|SUPABASE_SERVICE_ROLE)",
exts=JS_EXTS, dirs=dirs if dirs else None)
if findings:
return Result(c, Status.FAIL, "service_role key referenced in client-side code", findings)
return Result(c, Status.PASS)
def check_db12(root, stack):
c = CHECKS["DB-12"]
findings = grep_files(root,
r"(ssn|social_security|credit_card|card_number|passport_number)",
exts=SQL_EXTS | {".prisma"}, flags=re.IGNORECASE)
if findings:
return Result(c, Status.FAIL,
"PII column names found in schema — verify encryption at rest", findings)
return Result(c, Status.PASS)
def check_deploy09(root, stack):
c = CHECKS["DEPLOY-09"]
findings = grep_files(root,
r"(/health|/healthz|/api/health|/status|/readyz)",
exts=ALL_CODE_EXTS)
if findings:
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No health check endpoint found")
def check_deploy10(root, stack):
c = CHECKS["DEPLOY-10"]
# Check for logging libraries
lib_findings = grep_files(root,
r"(winston|pino|bunyan|morgan|log4js|structlog|loguru)",
exts=ALL_CODE_EXTS | {".json"})
if lib_findings:
return Result(c, Status.PASS)
# Count console.log in server/api code
server_dirs = {"api", "server", "backend"}
for d in ("pages/api", "app/api"):
if os.path.isdir(os.path.join(root, d)):
server_dirs.add(d.split("/")[0])
console_findings = grep_files(root, r"console\.(log|debug|info)\(", exts=JS_EXTS)
if console_findings:
return Result(c, Status.FAIL,
f"No structured logger found; {len(console_findings)} console.log(s) in code",
console_findings[:5])
return Result(c, Status.SKIP, "No server-side code detected")
def check_code01(root, stack):
c = CHECKS["CODE-01"]
if not stack.has_node:
return Result(c, Status.SKIP, "Not a JS/TS project")
findings = grep_files(root, r"console\.(log|debug|info)\(",
exts=JS_EXTS,
dirs=FRONTEND_DIRS & set(os.listdir(root)) or None,
exclude_patterns=[r"//.*console\.(log|debug|info)\("])
if findings:
return Result(c, Status.FAIL, f"{len(findings)} console.log statement(s) in production code", findings[:10])
return Result(c, Status.PASS)
def check_code03(root, stack):
c = CHECKS["CODE-03"]
findings = grep_files(root,
r"catch\s*\([^)]*\)\s*\{\s*\}",
exts=ALL_CODE_EXTS)
if findings:
return Result(c, Status.FAIL, f"{len(findings)} empty catch block(s)", findings)
return Result(c, Status.PASS)
def check_code07(root, stack):
c = CHECKS["CODE-07"]
findings = grep_files(root,
r"(TODO|FIXME|HACK|XXX).{0,20}(auth|security|permission|validation|sanitiz)",
exts=ALL_CODE_EXTS, flags=re.IGNORECASE)
if findings:
return Result(c, Status.FAIL, f"{len(findings)} deferred security TODO(s)", findings)
return Result(c, Status.PASS)
def check_code09(root, stack):
c = CHECKS["CODE-09"]
if not stack.has_react:
return Result(c, Status.SKIP, "Not a React project")
# Next.js App Router: error.tsx
error_page = file_exists_in(root, "error.tsx", "error.jsx", "global-error.tsx")
if error_page:
return Result(c, Status.PASS)
# Class-based error boundary
eb_findings = grep_files(root,
r"(ErrorBoundary|componentDidCatch|getDerivedStateFromError)",
exts=JS_EXTS)
if eb_findings:
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No React error boundary or error.tsx found")
def check_code10(root, stack):
c = CHECKS["CODE-10"]
findings = grep_files(root,
r"(error\.stack|\.stack\s*\)|err\.message.*res\.(json|send)|traceback\.format_exc)",
exts=ALL_CODE_EXTS)
if findings:
return Result(c, Status.FAIL, "Potential stack trace leak in error responses", findings)
return Result(c, Status.PASS)
def check_code11(root, stack):
c = CHECKS["CODE-11"]
if not stack.has_node:
return Result(c, Status.SKIP, "Not a JS/TS project")
findings = grep_files(root,
r"eslint-disable.*(no-eval|no-implied-eval|no-script-url|security)",
exts=JS_EXTS)
if findings:
return Result(c, Status.FAIL, "Security lint rule(s) disabled", findings)
return Result(c, Status.PASS)
def check_code12(root, stack):
c = CHECKS["CODE-12"]
lockfiles = ["package-lock.json", "pnpm-lock.yaml", "yarn.lock", "bun.lockb",
"Pipfile.lock", "poetry.lock", "Gemfile.lock", "go.sum", "Cargo.lock"]
for lf in lockfiles:
if os.path.isfile(os.path.join(root, lf)):
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No lockfile found — dependencies are not pinned")
def check_code13(root, stack):
c = CHECKS["CODE-13"]
if not stack.has_node:
return Result(c, Status.SKIP, "Not a JS/TS project")
pkg_path = os.path.join(root, "package.json")
if not os.path.isfile(pkg_path):
return Result(c, Status.SKIP)
findings = grep_files(root, r'"[^"]+"\s*:\s*"\*"', exts={".json"})
findings += grep_files(root, r'"[^"]+"\s*:\s*"latest"', exts={".json"})
findings = [f for f in findings if "package.json" in f.file and "node_modules" not in f.file]
if findings:
return Result(c, Status.FAIL, "Wildcard (*) or 'latest' version found in package.json", findings)
return Result(c, Status.PASS)
def check_code14(root, stack):
c = CHECKS["CODE-14"]
if not stack.has_typescript:
return Result(c, Status.SKIP, "Not a TypeScript project")
tsconfig_path = os.path.join(root, "tsconfig.json")
if not os.path.isfile(tsconfig_path):
return Result(c, Status.SKIP, "tsconfig.json not found")
content = open(tsconfig_path, errors="ignore").read()
if re.search(r'"strict"\s*:\s*true', content):
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "TypeScript strict mode not enabled in tsconfig.json",
[Finding("tsconfig.json", 0)])
def check_ai01(root, stack):
c = CHECKS["AI-01"]
if not stack.has_ai:
return Result(c, Status.SKIP, "No AI/LLM usage detected")
dirs = FRONTEND_DIRS & set(os.listdir(root))
findings = grep_files(root,
r"(system.?prompt|system.?message|system_instruction)",
exts=ALL_CODE_EXTS, dirs=dirs if dirs else None, flags=re.IGNORECASE)
if findings:
return Result(c, Status.FAIL,
"System prompt referenced in client-accessible code — may be leakable",
findings)
return Result(c, Status.PASS)
def check_ai02(root, stack):
c = CHECKS["AI-02"]
if not stack.has_ai:
return Result(c, Status.SKIP, "No AI/LLM usage detected")
findings = grep_files(root,
r"(messages\.push|content\s*:.*\$\{|content\s*:.*\+\s*user|prompt.*\+)",
exts=ALL_CODE_EXTS)
if findings:
return Result(c, Status.FAIL,
"User input may be concatenated directly into AI prompt", findings[:5])
return Result(c, Status.PASS)
def check_ai03(root, stack):
c = CHECKS["AI-03"]
if not stack.has_ai:
return Result(c, Status.SKIP, "No AI/LLM usage detected")
dirs = FRONTEND_DIRS & set(os.listdir(root))
findings = grep_files(root,
r"(OPENAI_API_KEY|ANTHROPIC_API_KEY|GOOGLE_AI_API_KEY|sk-ant-|sk-proj-)",
exts=JS_EXTS, dirs=dirs if dirs else None)
if findings:
return Result(c, Status.FAIL, "LLM API key referenced in frontend code", findings)
return Result(c, Status.PASS)
def check_ai05(root, stack):
c = CHECKS["AI-05"]
if not stack.has_ai:
return Result(c, Status.SKIP, "No AI/LLM usage detected")
findings = grep_files(root,
r"dangerouslySetInnerHTML.*\b(response|result|completion|message|content)\b",
exts=JS_EXTS)
if findings:
return Result(c, Status.FAIL, "AI output rendered via dangerouslySetInnerHTML", findings)
return Result(c, Status.PASS)
def check_dep01(root, stack):
c = CHECKS["DEP-01"]
if not stack.has_node:
return Result(c, Status.SKIP, "Not a Node.js project")
findings = grep_files(root,
r'"[^"]+"\s*:\s*"(git://|git\+|github:|https://github\.com|file:)',
exts={".json"})
findings = [f for f in findings if "package.json" in f.file and "node_modules" not in f.file]
if findings:
return Result(c, Status.FAIL, "Git/URL-based dependency found in package.json", findings)
return Result(c, Status.PASS)
def check_dep05(root, stack):
c = CHECKS["DEP-05"]
if not stack.has_node:
return Result(c, Status.SKIP, "Not a Node.js project")
pkg_path = os.path.join(root, "package.json")
if not os.path.isfile(pkg_path):
return Result(c, Status.SKIP)
pkg = read_json_file(pkg_path)
scripts = pkg.get("scripts", {})
suspicious = []
for key in ("preinstall", "postinstall", "install"):
val = scripts.get(key, "")
if val and any(kw in val for kw in ("curl", "wget", "fetch", "exec", "eval", "sh ", "bash ")):
suspicious.append(Finding("package.json", 0, f'"{key}": "{val}"'))
if suspicious:
return Result(c, Status.FAIL, "Suspicious install script detected in package.json", suspicious)
return Result(c, Status.PASS)
def check_dep06(root, stack):
c = CHECKS["DEP-06"]
if not stack.has_node:
return Result(c, Status.SKIP, "Not a Node.js project")
pkg_path = os.path.join(root, "package.json")
if not os.path.isfile(pkg_path):
return Result(c, Status.SKIP)
findings = grep_files(root, r'"\*"', exts={".json"})
findings = [f for f in findings if "package.json" in f.file and "node_modules" not in f.file]
if findings:
return Result(c, Status.FAIL, "Wildcard (*) version found", findings)
return Result(c, Status.PASS)
def check_fe01(root, stack):
c = CHECKS["FE-01"]
if not stack.is_web and not stack.has_node:
return Result(c, Status.SKIP, "Not a web project")
# Next.js metadata export
meta_findings = grep_files(root,
r"(export\s+(const|async\s+function)\s+metadata|generateMetadata|<title>|og:title|og:description)",
exts=JS_EXTS | {".html"})
if meta_findings:
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No meta tags or Next.js metadata export found")
def check_fe02(root, stack):
c = CHECKS["FE-02"]
if not stack.is_web and not stack.has_node:
return Result(c, Status.SKIP, "Not a web project")
favicon = file_exists_in(root, "favicon.ico", "favicon.png", "favicon.svg",
"favicon.webp", "icon.png", "icon.ico")
if favicon:
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No favicon file found")
def check_fe03(root, stack):
c = CHECKS["FE-03"]
if not stack.is_web and not stack.has_node:
return Result(c, Status.SKIP, "Not a web project")
page_404 = file_exists_in(root, "404.tsx", "404.jsx", "404.html",
"not-found.tsx", "not-found.jsx")
if page_404:
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No custom 404 or not-found page found")
def check_fe09(root, stack):
c = CHECKS["FE-09"]
if not stack.is_web and not stack.has_node:
return Result(c, Status.SKIP, "Not a web project")
public_robots = os.path.join(root, "public", "robots.txt")
root_robots = os.path.join(root, "robots.txt")
if os.path.isfile(public_robots) or os.path.isfile(root_robots):
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No robots.txt found")
def check_obs01(root, stack):
c = CHECKS["OBS-01"]
findings = grep_files(root,
r"(@sentry/|sentry-|LogRocket|Bugsnag|datadogRum|Rollbar|Honeybadger|newrelic)",
exts=ALL_CODE_EXTS | {".json"})
if findings:
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No error monitoring library detected")
def check_obs03(root, stack):
c = CHECKS["OBS-03"]
findings = grep_files(root,
r"(winston|pino|bunyan|structlog|loguru|import logging)",
exts=ALL_CODE_EXTS | {".json"})
if findings:
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No structured logging library detected")
CATEGORY_CHECKS = {
"SEC": [check_sec01, check_sec04, check_sec05, check_sec06, check_sec07,
check_sec08, check_sec11, check_sec13, check_sec14, check_sec17, check_sec18],
"DB": [check_db03, check_db05, check_db06, check_db07, check_db08, check_db12],
"DEPLOY": [check_deploy09, check_deploy10],
"CODE": [check_code01, check_code03, check_code07, check_code09, check_code10,
check_code11, check_code12, check_code13, check_code14],
"AI": [check_ai01, check_ai02, check_ai03, check_ai05],
"DEP": [check_dep01, check_dep05, check_dep06],
"FE": [check_fe01, check_fe02, check_fe03, check_fe09],
"OBS": [check_obs01, check_obs03],
}
CATEGORY_ORDER = ["SEC", "DB", "CODE", "DEP", "AI", "DEPLOY", "FE", "OBS"]
# ---------------------------------------------------------------------------
# Manual check runner
# ---------------------------------------------------------------------------
def run_manual_checks(stack: Stack, interactive: bool, category_filter: Optional[str]) -> List[Result]:
results = []
applicable = []
for chk in MANUAL_CHECKS:
if category_filter and chk.category != category_filter.upper():
continue
if chk.stack == "ai" and not stack.has_ai:
results.append(Result(chk, Status.SKIP, "No AI/LLM usage detected"))
continue
if chk.stack == "web" and not stack.is_web:
results.append(Result(chk, Status.SKIP, "Not a web project"))
continue
if chk.stack == "vps" and stack.deploy_target not in ("docker", "vps", ""):
results.append(Result(chk, Status.SKIP, "Not a VPS/Docker deployment"))
continue
applicable.append(chk)
if not interactive or not applicable:
for chk in applicable:
results.append(Result(chk, Status.MANUAL, "Not confirmed (run without --no-interactive to answer)"))
return results
print()
print(bold("Manual Checks") + " — answer Y/N for each:")
print()
for chk in applicable:
sev_label = {
Severity.CRITICAL: red("CRITICAL"),
Severity.HIGH: yellow("HIGH"),
Severity.ADVISORY: dim("ADVISORY"),
}[chk.severity]
while True:
try:
answer = input(f" [{sev_label}] [{chk.id}] {chk.description} [y/N]: ").strip().lower()
except (EOFError, KeyboardInterrupt):
answer = "n"
if answer in ("y", "yes"):
results.append(Result(chk, Status.PASS))
break
elif answer in ("n", "no", ""):
results.append(Result(chk, Status.FAIL, "Not confirmed"))
break
print(" Please enter Y or N.")
return results
# ---------------------------------------------------------------------------
# Verdict / output
# ---------------------------------------------------------------------------
def severity_for_result(r: Result) -> Severity:
return r.check.severity
def print_report(all_results: List[Result], stack: Stack, scan_time: float,
verbose: bool) -> int:
critical = [r for r in all_results if r.status in (Status.FAIL, Status.MANUAL)
and r.check.severity == Severity.CRITICAL]
high = [r for r in all_results if r.status in (Status.FAIL, Status.MANUAL)
and r.check.severity == Severity.HIGH]
advisory = [r for r in all_results if r.status in (Status.FAIL, Status.MANUAL)
and r.check.severity == Severity.ADVISORY]
stack_desc = []
if stack.framework: stack_desc.append(stack.framework.capitalize())
if stack.has_supabase: stack_desc.append("Supabase")
if stack.deploy_target: stack_desc.append(stack.deploy_target.capitalize())
if stack.has_python and stack.py_framework: stack_desc.append(stack.py_framework.capitalize())
if not stack_desc: stack_desc.append("Unknown")
stack_str = " + ".join(stack_desc)
print()
print(bold("SHIP GATE REPORT"))
print("=" * 48)
print(f"Stack: {stack_str}")
print(f"Scan time: {scan_time:.1f}s")
print(f"Checks: {len(all_results)} total")
print()
def _section(label, items, color_fn):
if not items and not verbose:
return
print(bold(f"{label} ({len(items)} item{'s' if len(items) != 1 else ''})"))
for r in items:
status_str = {
Status.FAIL: red("FAIL "),
Status.MANUAL: yellow("MANUAL"),
Status.PASS: green("PASS "),
Status.SKIP: dim("SKIP "),
}[r.status]
print(f" {status_str} [{r.check.id}] {r.check.description}")
if r.message:
print(f" {dim(r.message)}")
for f in r.findings[:3]:
print(f" {dim(f.file)}:{f.line} {dim(f.snippet[:80])}")
print()
if critical:
_section(red("CRITICAL") + " (must fix before shipping)", critical, red)
if high:
_section(yellow("HIGH") + " (should fix before shipping)", high, yellow)
if advisory:
_section(dim("ADVISORY") + " (recommended)", advisory, dim)
if verbose:
passed = [r for r in all_results if r.status == Status.PASS]
if passed:
_section(green("PASS"), passed, green)
skipped = [r for r in all_results if r.status == Status.SKIP]
if skipped:
_section(dim("SKIP"), skipped, dim)
if critical:
print(red(bold(f"VERDICT: DO NOT SHIP ({len(critical)} critical issue{'s' if len(critical) != 1 else ''})")))
print("Fix critical items and re-run.")
return 1
elif high:
print(yellow(bold(f"VERDICT: SHIP WITH CAUTION ({len(high)} high issue{'s' if len(high) != 1 else ''})")))
print("Acknowledge risks and proceed only if you accept them.")
return 2
else:
print(green(bold("VERDICT: CLEAR TO SHIP")))
return 0
def print_json_report(all_results: List[Result], stack: Stack, scan_time: float) -> int:
critical = [r for r in all_results if r.status in (Status.FAIL, Status.MANUAL)
and r.check.severity == Severity.CRITICAL]
high = [r for r in all_results if r.status in (Status.FAIL, Status.MANUAL)
and r.check.severity == Severity.HIGH]
output = {
"version": VERSION,
"scan_time": round(scan_time, 2),
"stack": {
"framework": stack.framework,
"has_supabase": stack.has_supabase,
"has_typescript": stack.has_typescript,
"deploy_target": stack.deploy_target,
"has_ai": stack.has_ai,
},
"results": [
{
"id": r.check.id,
"description": r.check.description,
"severity": r.check.severity.value,
"category": r.check.category,
"status": r.status.value,
"message": r.message,
"findings": [
{"file": f.file, "line": f.line, "snippet": f.snippet}
for f in r.findings
],
}
for r in all_results
],
"summary": {
"critical": len(critical),
"high": len(high),
"verdict": "DO_NOT_SHIP" if critical else ("SHIP_WITH_CAUTION" if high else "CLEAR_TO_SHIP"),
},
}
print(json.dumps(output, indent=2))
return 1 if critical else (2 if high else 0)
# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
def main():
global USE_COLOR
parser = argparse.ArgumentParser(
description="Ship Gate — pre-production audit scanner",
formatter_class=argparse.RawDescriptionHelpFormatter,
)
parser.add_argument("path", nargs="?", default=".",
help="Project root directory (default: current directory)")
parser.add_argument("--json", action="store_true", help="Output as JSON")
parser.add_argument("--no-color", action="store_true", help="Disable color output")
parser.add_argument("--no-interactive", action="store_true",
help="Skip manual confirmation prompts")
parser.add_argument("--category", metavar="CAT",
help="Only run one category: SEC, DB, CODE, DEP, AI, DEPLOY, FE, OBS")
parser.add_argument("--verbose", action="store_true",
help="Show PASS and SKIP results in addition to failures")
parser.add_argument("--version", action="version", version=f"ship-gate {VERSION}")
args = parser.parse_args()
if args.no_color or not sys.stdout.isatty():
USE_COLOR = False
root = os.path.abspath(args.path)
if not os.path.isdir(root):
print(f"Error: '{root}' is not a directory", file=sys.stderr)
sys.exit(1)
start = time.time()
# Detect stack
stack = detect_stack(root)
if not args.json:
print(bold("Detecting stack..."), end=" ", flush=True)
parts = []
if stack.framework: parts.append(stack.framework.capitalize())
if stack.has_supabase: parts.append("Supabase")
if stack.deploy_target: parts.append(stack.deploy_target.capitalize())
if stack.has_python and stack.py_framework: parts.append(stack.py_framework.capitalize())
if stack.has_ai: parts.append(f"AI({','.join(stack.ai_providers)})")
print(", ".join(parts) if parts else "generic project")
# Run automated checks
all_results: List[Result] = []
categories = [args.category.upper()] if args.category else CATEGORY_ORDER
for i, cat in enumerate(categories, 1):
fns = CATEGORY_CHECKS.get(cat, [])
cat_results = []
for fn in fns:
try:
r = fn(root, stack)
except Exception as e:
chk_id = fn.__name__.replace("check_", "").replace("_", "-").upper()
cat_results.append(Result(
CheckDef(chk_id, fn.__doc__ or fn.__name__, Severity.ADVISORY, cat),
Status.SKIP, f"Scanner error: {e}",
))
continue
cat_results.append(r)
all_results.extend(cat_results)
if not args.json:
n_fail = sum(1 for r in cat_results if r.status == Status.FAIL)
n_pass = sum(1 for r in cat_results if r.status == Status.PASS)
n_skip = sum(1 for r in cat_results if r.status == Status.SKIP)
label = red(f"{n_fail} FAIL") if n_fail else green("0 FAIL")
print(f" [{i}/{len(categories)}] {cat}: {label}, {n_pass} PASS, {dim(str(n_skip) + ' SKIP')}")
# Manual checks
manual_results = run_manual_checks(stack, not args.no_interactive, args.category)
all_results.extend(manual_results)
scan_time = time.time() - start
if args.json:
sys.exit(print_json_report(all_results, stack, scan_time))
else:
sys.exit(print_report(all_results, stack, scan_time, args.verbose))
if __name__ == "__main__":
main()
Lập kế hoạch phát hành, quản lý changelog, điều phối triển khai, tạo nhánh release và tự động hóa đánh phiên bản.
---
name: "release-manager"
description: "Use when the user asks to plan releases, manage changelogs, coordinate deployments, create release branches, or automate versioning."
---
# Release Manager
**Tier:** POWERFUL
**Category:** Engineering
**Domain:** Software Release Management & DevOps
## Overview
The Release Manager skill provides comprehensive tools and knowledge for managing software releases end-to-end. From parsing conventional commits to generating changelogs, determining version bumps, and orchestrating release processes, this skill ensures reliable, predictable, and well-documented software releases.
## Core Capabilities
- **Automated Changelog Generation** from git history using conventional commits
- **Semantic Version Bumping** based on commit analysis and breaking changes
- **Release Readiness Assessment** with comprehensive checklists and validation
- **Release Planning & Coordination** with stakeholder communication templates
- **Rollback Planning** with automated recovery procedures
- **Hotfix Management** for emergency releases
- **Feature Flag Integration** for progressive rollouts
## Key Components
### Scripts
1. **changelog_generator.py** - Parses git logs and generates structured changelogs
2. **version_bumper.py** - Determines correct version bumps from conventional commits
3. **release_planner.py** - Assesses release readiness and generates coordination plans
### Documentation
- Comprehensive release management methodology
- Conventional commits specification and examples
- Release workflow comparisons (Git Flow, Trunk-based, GitHub Flow)
- Hotfix procedures and emergency response protocols
## Release Management Methodology
### Semantic Versioning (SemVer)
Semantic Versioning follows the MAJOR.MINOR.PATCH format where:
- **MAJOR** version when you make incompatible API changes
- **MINOR** version when you add functionality in a backwards compatible manner
- **PATCH** version when you make backwards compatible bug fixes
#### Pre-release Versions
Pre-release versions are denoted by appending a hyphen and identifiers:
- `1.0.0-alpha.1` - Alpha releases for early testing
- `1.0.0-beta.2` - Beta releases for wider testing
- `1.0.0-rc.1` - Release candidates for final validation
#### Version Precedence
Version precedence is determined by comparing each identifier:
1. `1.0.0-alpha` < `1.0.0-alpha.1` < `1.0.0-alpha.beta` < `1.0.0-beta`
2. `1.0.0-beta` < `1.0.0-beta.2` < `1.0.0-beta.11` < `1.0.0-rc.1`
3. `1.0.0-rc.1` < `1.0.0`
### Conventional Commits
Conventional Commits provide a structured format for commit messages that enables automated tooling:
#### Format
```
<type>[optional scope]: <description>
[optional body]
[optional footer(s)]
```
#### Types
- **feat**: A new feature (correlates with MINOR version bump)
- **fix**: A bug fix (correlates with PATCH version bump)
- **docs**: Documentation only changes
- **style**: Changes that do not affect the meaning of the code
- **refactor**: A code change that neither fixes a bug nor adds a feature
- **perf**: A code change that improves performance
- **test**: Adding missing tests or correcting existing tests
- **chore**: Changes to the build process or auxiliary tools
- **ci**: Changes to CI configuration files and scripts
- **build**: Changes that affect the build system or external dependencies
- **breaking**: Introduces a breaking change (correlates with MAJOR version bump)
#### Examples
```
feat(user-auth): add OAuth2 integration
fix(api): resolve race condition in user creation
docs(readme): update installation instructions
feat!: remove deprecated payment API
BREAKING CHANGE: The legacy payment API has been removed
```
### Automated Changelog Generation
Changelogs are automatically generated from conventional commits, organized by:
#### Structure
```markdown
# Changelog
## [Unreleased]
### Added
### Changed
### Deprecated
### Removed
### Fixed
### Security
## [1.2.0] - 2024-01-15
### Added
- OAuth2 authentication support (#123)
- User preference dashboard (#145)
### Fixed
- Race condition in user creation (#134)
- Memory leak in image processing (#156)
### Breaking Changes
- Removed legacy payment API
```
#### Grouping Rules
- **Added** for new features (feat)
- **Fixed** for bug fixes (fix)
- **Changed** for changes in existing functionality
- **Deprecated** for soon-to-be removed features
- **Removed** for now removed features
- **Security** for vulnerability fixes
#### Metadata Extraction
- Link to pull requests and issues: `(#123)`
- Breaking changes highlighted prominently
- Scope-based grouping: `auth:`, `api:`, `ui:`
- Co-authored-by for contributor recognition
### Version Bump Strategies
Version bumps are determined by analyzing commits since the last release:
#### Automatic Detection Rules
1. **MAJOR**: Any commit with `BREAKING CHANGE` or `!` after type
2. **MINOR**: Any `feat` type commits without breaking changes
3. **PATCH**: `fix`, `perf`, `security` type commits
4. **NO BUMP**: `docs`, `style`, `test`, `chore`, `ci`, `build` only
#### Pre-release Handling
```python
# Alpha: 1.0.0-alpha.1 → 1.0.0-alpha.2
# Beta: 1.0.0-alpha.5 → 1.0.0-beta.1
# RC: 1.0.0-beta.3 → 1.0.0-rc.1
# Release: 1.0.0-rc.2 → 1.0.0
```
#### Multi-package Considerations
For monorepos with multiple packages:
- Analyze commits affecting each package independently
- Support scoped version bumps: `@scope/package@1.2.3`
- Generate coordinated release plans across packages
### Release Branch Workflows
#### Git Flow
```
main (production) ← release/1.2.0 ← develop ← feature/login
← hotfix/critical-fix
```
**Advantages:**
- Clear separation of concerns
- Stable main branch
- Parallel feature development
- Structured release process
**Process:**
1. Create release branch from develop: `git checkout -b release/1.2.0 develop`
2. Finalize release (version bump, changelog)
3. Merge to main and develop
4. Tag release: `git tag v1.2.0`
5. Deploy from main
#### Trunk-based Development
```
main ← feature/login (short-lived)
← feature/payment (short-lived)
← hotfix/critical-fix
```
**Advantages:**
- Simplified workflow
- Faster integration
- Reduced merge conflicts
- Continuous integration friendly
**Process:**
1. Short-lived feature branches (1-3 days)
2. Frequent commits to main
3. Feature flags for incomplete features
4. Automated testing gates
5. Deploy from main with feature toggles
#### GitHub Flow
```
main ← feature/login
← hotfix/critical-fix
```
**Advantages:**
- Simple and lightweight
- Fast deployment cycle
- Good for web applications
- Minimal overhead
**Process:**
1. Create feature branch from main
2. Regular commits and pushes
3. Open pull request when ready
4. Deploy from feature branch for testing
5. Merge to main and deploy
### Feature Flag Integration
Feature flags enable safe, progressive rollouts:
#### Types of Feature Flags
- **Release flags**: Control feature visibility in production
- **Experiment flags**: A/B testing and gradual rollouts
- **Operational flags**: Circuit breakers and performance toggles
- **Permission flags**: Role-based feature access
#### Implementation Strategy
```python
# Progressive rollout example
if feature_flag("new_payment_flow", user_id):
return new_payment_processor.process(payment)
else:
return legacy_payment_processor.process(payment)
```
#### Release Coordination
1. Deploy code with feature behind flag (disabled)
2. Gradually enable for percentage of users
3. Monitor metrics and error rates
4. Full rollout or quick rollback based on data
5. Remove flag in subsequent release
### Release Readiness Checklists
#### Pre-Release Validation
- [ ] All planned features implemented and tested
- [ ] Breaking changes documented with migration guide
- [ ] API documentation updated
- [ ] Database migrations tested
- [ ] Security review completed for sensitive changes
- [ ] Performance testing passed thresholds
- [ ] Internationalization strings updated
- [ ] Third-party integrations validated
#### Quality Gates
- [ ] Unit test coverage ≥ 85%
- [ ] Integration tests passing
- [ ] End-to-end tests passing
- [ ] Static analysis clean
- [ ] Security scan passed
- [ ] Dependency audit clean
- [ ] Load testing completed
#### Documentation Requirements
- [ ] CHANGELOG.md updated
- [ ] README.md reflects new features
- [ ] API documentation generated
- [ ] Migration guide written for breaking changes
- [ ] Deployment notes prepared
- [ ] Rollback procedure documented
#### Stakeholder Approvals
- [ ] Product Manager sign-off
- [ ] Engineering Lead approval
- [ ] QA validation complete
- [ ] Security team clearance
- [ ] Legal review (if applicable)
- [ ] Compliance check (if regulated)
### Deployment Coordination
#### Communication Plan
**Internal Stakeholders:**
- Engineering team: Technical changes and rollback procedures
- Product team: Feature descriptions and user impact
- Support team: Known issues and troubleshooting guides
- Sales team: Customer-facing changes and talking points
**External Communication:**
- Release notes for users
- API changelog for developers
- Migration guide for breaking changes
- Downtime notifications if applicable
#### Deployment Sequence
1. **Pre-deployment** (T-24h): Final validation, freeze code
2. **Database migrations** (T-2h): Run and validate schema changes
3. **Blue-green deployment** (T-0): Switch traffic gradually
4. **Post-deployment** (T+1h): Monitor metrics and logs
5. **Rollback window** (T+4h): Decision point for rollback
#### Monitoring & Validation
- Application health checks
- Error rate monitoring
- Performance metrics tracking
- User experience monitoring
- Business metrics validation
- Third-party service integration health
### Hotfix Procedures
Hotfixes address critical production issues requiring immediate deployment:
#### Severity Classification
**P0 - Critical**: Complete system outage, data loss, security breach
- **SLA**: Fix within 2 hours
- **Process**: Emergency deployment, all hands on deck
- **Approval**: Engineering Lead + On-call Manager
**P1 - High**: Major feature broken, significant user impact
- **SLA**: Fix within 24 hours
- **Process**: Expedited review and deployment
- **Approval**: Engineering Lead + Product Manager
**P2 - Medium**: Minor feature issues, limited user impact
- **SLA**: Fix in next release cycle
- **Process**: Normal review process
- **Approval**: Standard PR review
#### Emergency Response Process
1. **Incident declaration**: Page on-call team
2. **Assessment**: Determine severity and impact
3. **Hotfix branch**: Create from last stable release
4. **Minimal fix**: Address root cause only
5. **Expedited testing**: Automated tests + manual validation
6. **Emergency deployment**: Deploy to production
7. **Post-incident**: Root cause analysis and prevention
### Rollback Planning
Every release must have a tested rollback plan:
#### Rollback Triggers
- **Error rate spike**: >2x baseline within 30 minutes
- **Performance degradation**: >50% latency increase
- **Feature failures**: Core functionality broken
- **Security incident**: Vulnerability exploited
- **Data corruption**: Database integrity compromised
#### Rollback Types
**Code Rollback:**
- Revert to previous Docker image
- Database-compatible code changes only
- Feature flag disable preferred over code rollback
**Database Rollback:**
- Only for non-destructive migrations
- Data backup required before migration
- Forward-only migrations preferred (add columns, not drop)
**Infrastructure Rollback:**
- Blue-green deployment switch
- Load balancer configuration revert
- DNS changes (longer propagation time)
#### Automated Rollback
```python
# Example rollback automation
def monitor_deployment():
if error_rate() > THRESHOLD:
alert_oncall("Error rate spike detected")
if auto_rollback_enabled():
execute_rollback()
```
### Release Metrics & Analytics
#### Key Performance Indicators
- **Lead Time**: From commit to production
- **Deployment Frequency**: Releases per week/month
- **Mean Time to Recovery**: From incident to resolution
- **Change Failure Rate**: Percentage of releases causing incidents
#### Quality Metrics
- **Rollback Rate**: Percentage of releases rolled back
- **Hotfix Rate**: Hotfixes per regular release
- **Bug Escape Rate**: Production bugs per release
- **Time to Detection**: How quickly issues are identified
#### Process Metrics
- **Review Time**: Time spent in code review
- **Testing Time**: Automated + manual testing duration
- **Approval Cycle**: Time from PR to merge
- **Release Preparation**: Time spent on release activities
### Tool Integration
#### Version Control Systems
- **Git**: Primary VCS with conventional commit parsing
- **GitHub/GitLab**: Pull request automation and CI/CD
- **Bitbucket**: Pipeline integration and deployment gates
#### CI/CD Platforms
- **Jenkins**: Pipeline orchestration and deployment automation
- **GitHub Actions**: Workflow automation and release publishing
- **GitLab CI**: Integrated pipelines with environment management
- **CircleCI**: Container-based builds and deployments
#### Monitoring & Alerting
- **DataDog**: Application performance monitoring
- **New Relic**: Error tracking and performance insights
- **Sentry**: Error aggregation and release tracking
- **PagerDuty**: Incident response and escalation
#### Communication Platforms
- **Slack**: Release notifications and coordination
- **Microsoft Teams**: Stakeholder communication
- **Email**: External customer notifications
- **Status Pages**: Public incident communication
## Best Practices
### Release Planning
1. **Regular cadence**: Establish predictable release schedule
2. **Feature freeze**: Lock changes 48h before release
3. **Risk assessment**: Evaluate changes for potential impact
4. **Stakeholder alignment**: Ensure all teams are prepared
### Quality Assurance
1. **Automated testing**: Comprehensive test coverage
2. **Staging environment**: Production-like testing environment
3. **Canary releases**: Gradual rollout to subset of users
4. **Monitoring**: Proactive issue detection
### Communication
1. **Clear timelines**: Communicate schedules early
2. **Regular updates**: Status reports during release process
3. **Issue transparency**: Honest communication about problems
4. **Post-mortems**: Learn from incidents and improve
### Automation
1. **Reduce manual steps**: Automate repetitive tasks
2. **Consistent process**: Same steps every time
3. **Audit trails**: Log all release activities
4. **Self-service**: Enable teams to deploy safely
## Common Anti-patterns
### Process Anti-patterns
- **Manual deployments**: Error-prone and inconsistent
- **Last-minute changes**: Risk introduction without proper testing
- **Skipping testing**: Deploying without validation
- **Poor communication**: Stakeholders unaware of changes
### Technical Anti-patterns
- **Monolithic releases**: Large, infrequent releases with high risk
- **Coupled deployments**: Services that must be deployed together
- **No rollback plan**: Unable to quickly recover from issues
- **Environment drift**: Production differs from staging
### Cultural Anti-patterns
- **Blame culture**: Fear of making changes or reporting issues
- **Hero culture**: Relying on individuals instead of process
- **Perfectionism**: Delaying releases for minor improvements
- **Risk aversion**: Avoiding necessary changes due to fear
## Getting Started
1. **Assessment**: Evaluate current release process and pain points
2. **Tool setup**: Configure scripts for your repository
3. **Process definition**: Choose appropriate workflow for your team
4. **Automation**: Implement CI/CD pipelines and quality gates
5. **Training**: Educate team on new processes and tools
6. **Monitoring**: Set up metrics and alerting for releases
7. **Iteration**: Continuously improve based on feedback and metrics
The Release Manager skill transforms chaotic deployments into predictable, reliable releases that build confidence across your entire organization.
FILE:assets/sample_commits.json
[
{
"hash": "a1b2c3d",
"author": "Sarah Johnson <sarah.johnson@example.com>",
"date": "2024-01-15T14:30:22Z",
"message": "feat(auth): add OAuth2 integration with Google and GitHub\n\nImplement OAuth2 authentication flow supporting Google and GitHub providers.\nUsers can now sign in using their existing social media accounts, improving\nuser experience and reducing password fatigue.\n\n- Add OAuth2 client configuration\n- Implement authorization code flow\n- Add user profile mapping from providers\n- Include comprehensive error handling\n\nCloses #123\nResolves #145"
},
{
"hash": "e4f5g6h",
"author": "Mike Chen <mike.chen@example.com>",
"date": "2024-01-15T13:45:18Z",
"message": "fix(api): resolve race condition in user creation endpoint\n\nFixed a race condition that occurred when multiple requests attempted\nto create users with the same email address simultaneously. This was\ncausing duplicate user records in some edge cases.\n\n- Added database unique constraint on email field\n- Implemented proper error handling for constraint violations\n- Added retry logic with exponential backoff\n\nFixes #234"
},
{
"hash": "i7j8k9l",
"author": "Emily Davis <emily.davis@example.com>",
"date": "2024-01-15T12:20:45Z",
"message": "docs(readme): update installation and deployment instructions\n\nUpdated README with comprehensive installation guide including:\n- Docker setup instructions\n- Environment variable configuration\n- Database migration steps\n- Troubleshooting common issues"
},
{
"hash": "m1n2o3p",
"author": "David Wilson <david.wilson@example.com>",
"date": "2024-01-15T11:15:30Z",
"message": "feat(ui)!: redesign dashboard with new component library\n\nComplete redesign of the user dashboard using our new component library.\nThis provides better accessibility, improved mobile responsiveness, and\na more modern user interface.\n\nBREAKING CHANGE: The dashboard API endpoints have changed structure.\nFrontend clients must update to use the new /v2/dashboard endpoints.\nThe legacy /v1/dashboard endpoints will be removed in version 3.0.0.\n\n- Implement new Card, Grid, and Chart components\n- Add responsive breakpoints for mobile devices\n- Improve accessibility with proper ARIA labels\n- Add dark mode support\n\nCloses #345, #367, #389"
},
{
"hash": "q4r5s6t",
"author": "Lisa Rodriguez <lisa.rodriguez@example.com>",
"date": "2024-01-15T10:45:12Z",
"message": "fix(db): optimize slow query in user search functionality\n\nOptimized the user search query that was causing performance issues\non databases with large user counts. Query time reduced from 2.5s to 150ms.\n\n- Added composite index on (email, username, created_at)\n- Refactored query to use more efficient JOIN structure\n- Added query result caching for common search patterns\n\nFixes #456"
},
{
"hash": "u7v8w9x",
"author": "Tom Anderson <tom.anderson@example.com>",
"date": "2024-01-15T09:30:55Z",
"message": "chore(deps): upgrade React to version 18.2.0\n\nUpgrade React and related dependencies to latest stable versions.\nThis includes performance improvements and new concurrent features.\n\n- React: 17.0.2 → 18.2.0\n- React-DOM: 17.0.2 → 18.2.0\n- React-Router: 6.8.0 → 6.8.1\n- Updated all peer dependencies"
},
{
"hash": "y1z2a3b",
"author": "Jennifer Kim <jennifer.kim@example.com>",
"date": "2024-01-15T08:15:33Z",
"message": "test(auth): add comprehensive tests for OAuth flow\n\nAdded unit and integration tests for the OAuth2 authentication system\nto ensure reliability and prevent regressions.\n\n- Unit tests for OAuth client configuration\n- Integration tests for complete auth flow\n- Mock providers for testing without external dependencies\n- Error scenario testing\n\nTest coverage increased from 72% to 89% for auth module."
},
{
"hash": "c4d5e6f",
"author": "Alex Thompson <alex.thompson@example.com>",
"date": "2024-01-15T07:45:20Z",
"message": "perf(image): implement WebP compression reducing size by 40%\n\nReplaced PNG compression with WebP format for uploaded images.\nThis reduces average image file sizes by 40% while maintaining\nvisual quality, improving page load times and reducing bandwidth costs.\n\n- Add WebP encoding support\n- Implement fallback to PNG for older browsers\n- Add quality settings configuration\n- Update image serving endpoints\n\nPerformance improvement: Page load time reduced by 25% on average."
},
{
"hash": "g7h8i9j",
"author": "Rachel Green <rachel.green@example.com>",
"date": "2024-01-14T16:20:10Z",
"message": "feat(payment): add Stripe payment processor integration\n\nIntegrate Stripe as a payment processor to support credit card payments.\nThis enables users to purchase premium features and subscriptions.\n\n- Add Stripe SDK integration\n- Implement payment intent flow\n- Add webhook handling for payment status updates\n- Include comprehensive error handling and logging\n- Add payment method management for users\n\nCloses #567\nCo-authored-by: Payment Team <payments@example.com>"
},
{
"hash": "k1l2m3n",
"author": "Chris Martinez <chris.martinez@example.com>",
"date": "2024-01-14T15:30:45Z",
"message": "fix(ui): resolve mobile navigation menu overflow issue\n\nFixed navigation menu overflow on mobile devices where long menu items\nwere being cut off and causing horizontal scrolling issues.\n\n- Implement responsive text wrapping\n- Add horizontal scrolling for overflowing content\n- Improve touch targets for better mobile usability\n- Fix z-index conflicts with dropdown menus\n\nFixes #678\nTested on iOS Safari, Chrome Mobile, and Firefox Mobile"
},
{
"hash": "o4p5q6r",
"author": "Anna Kowalski <anna.kowalski@example.com>",
"date": "2024-01-14T14:20:15Z",
"message": "refactor(api): extract validation logic into reusable middleware\n\nExtracted common validation logic from individual API endpoints into\nreusable middleware functions to reduce code duplication and improve\nmaintainability.\n\n- Create validation middleware for common patterns\n- Refactor user, product, and order endpoints\n- Add comprehensive error messages\n- Improve validation performance by 30%"
},
{
"hash": "s7t8u9v",
"author": "Kevin Park <kevin.park@example.com>",
"date": "2024-01-14T13:10:30Z",
"message": "feat(search): implement fuzzy search with Elasticsearch\n\nImplemented fuzzy search functionality using Elasticsearch to provide\nbetter search results for users with typos or partial matches.\n\n- Integrate Elasticsearch cluster\n- Add fuzzy matching with configurable distance\n- Implement search result ranking algorithm\n- Add search analytics and logging\n\nSearch accuracy improved by 35% in user testing.\nCloses #789"
},
{
"hash": "w1x2y3z",
"author": "Security Team <security@example.com>",
"date": "2024-01-14T12:45:22Z",
"message": "fix(security): patch SQL injection vulnerability in reports\n\nPatched SQL injection vulnerability in the reports generation endpoint\nthat could allow unauthorized access to sensitive data.\n\n- Implement parameterized queries for all report filters\n- Add input sanitization and validation\n- Update security audit logging\n- Add automated security tests\n\nSeverity: HIGH - CVE-2024-0001\nReported by: External security researcher"
}
]
FILE:assets/sample_git_log.txt
a1b2c3d feat(auth): add OAuth2 integration with Google and GitHub
e4f5g6h fix(api): resolve race condition in user creation endpoint
i7j8k9l docs(readme): update installation and deployment instructions
m1n2o3p feat(ui)!: redesign dashboard with new component library
q4r5s6t fix(db): optimize slow query in user search functionality
u7v8w9x chore(deps): upgrade React to version 18.2.0
y1z2a3b test(auth): add comprehensive tests for OAuth flow
c4d5e6f perf(image): implement WebP compression reducing size by 40%
g7h8i9j feat(payment): add Stripe payment processor integration
k1l2m3n fix(ui): resolve mobile navigation menu overflow issue
o4p5q6r refactor(api): extract validation logic into reusable middleware
s7t8u9v feat(search): implement fuzzy search with Elasticsearch
w1x2y3z fix(security): patch SQL injection vulnerability in reports
a4b5c6d build(ci): add automated security scanning to deployment pipeline
e7f8g9h feat(notification): add email and SMS notification system
i1j2k3l fix(payment): handle expired credit cards gracefully
m4n5o6p docs(api): generate OpenAPI specification for all endpoints
q7r8s9t chore(cleanup): remove deprecated user preference API endpoints
u1v2w3x feat(admin)!: redesign admin panel with role-based permissions
y4z5a6b fix(db): resolve deadlock issues in concurrent transactions
c7d8e9f perf(cache): implement Redis caching for frequent database queries
g1h2i3j feat(mobile): add biometric authentication support
k4l5m6n fix(api): validate input parameters to prevent XSS attacks
o7p8q9r style(ui): update color palette and typography consistency
s1t2u3v feat(analytics): integrate Google Analytics 4 tracking
w4x5y6z fix(memory): resolve memory leak in image processing service
a7b8c9d ci(github): add automated testing for all pull requests
e1f2g3h feat(export): add CSV and PDF export functionality for reports
i4j5k6l fix(ui): resolve accessibility issues with screen readers
m7n8o9p refactor(auth): consolidate authentication logic into single service
FILE:assets/sample_git_log_full.txt
commit a1b2c3d4e5f6789012345678901234567890abcd
Author: Sarah Johnson <sarah.johnson@example.com>
Date: Mon Jan 15 14:30:22 2024 +0000
feat(auth): add OAuth2 integration with Google and GitHub
Implement OAuth2 authentication flow supporting Google and GitHub providers.
Users can now sign in using their existing social media accounts, improving
user experience and reducing password fatigue.
- Add OAuth2 client configuration
- Implement authorization code flow
- Add user profile mapping from providers
- Include comprehensive error handling
Closes #123
Resolves #145
commit e4f5g6h7i8j9012345678901234567890123abcdef
Author: Mike Chen <mike.chen@example.com>
Date: Mon Jan 15 13:45:18 2024 +0000
fix(api): resolve race condition in user creation endpoint
Fixed a race condition that occurred when multiple requests attempted
to create users with the same email address simultaneously. This was
causing duplicate user records in some edge cases.
- Added database unique constraint on email field
- Implemented proper error handling for constraint violations
- Added retry logic with exponential backoff
Fixes #234
commit i7j8k9l0m1n2345678901234567890123456789abcd
Author: Emily Davis <emily.davis@example.com>
Date: Mon Jan 15 12:20:45 2024 +0000
docs(readme): update installation and deployment instructions
Updated README with comprehensive installation guide including:
- Docker setup instructions
- Environment variable configuration
- Database migration steps
- Troubleshooting common issues
commit m1n2o3p4q5r6789012345678901234567890abcdefg
Author: David Wilson <david.wilson@example.com>
Date: Mon Jan 15 11:15:30 2024 +0000
feat(ui)!: redesign dashboard with new component library
Complete redesign of the user dashboard using our new component library.
This provides better accessibility, improved mobile responsiveness, and
a more modern user interface.
BREAKING CHANGE: The dashboard API endpoints have changed structure.
Frontend clients must update to use the new /v2/dashboard endpoints.
The legacy /v1/dashboard endpoints will be removed in version 3.0.0.
- Implement new Card, Grid, and Chart components
- Add responsive breakpoints for mobile devices
- Improve accessibility with proper ARIA labels
- Add dark mode support
Closes #345, #367, #389
commit q4r5s6t7u8v9012345678901234567890123456abcd
Author: Lisa Rodriguez <lisa.rodriguez@example.com>
Date: Mon Jan 15 10:45:12 2024 +0000
fix(db): optimize slow query in user search functionality
Optimized the user search query that was causing performance issues
on databases with large user counts. Query time reduced from 2.5s to 150ms.
- Added composite index on (email, username, created_at)
- Refactored query to use more efficient JOIN structure
- Added query result caching for common search patterns
Fixes #456
commit u7v8w9x0y1z2345678901234567890123456789abcde
Author: Tom Anderson <tom.anderson@example.com>
Date: Mon Jan 15 09:30:55 2024 +0000
chore(deps): upgrade React to version 18.2.0
Upgrade React and related dependencies to latest stable versions.
This includes performance improvements and new concurrent features.
- React: 17.0.2 → 18.2.0
- React-DOM: 17.0.2 → 18.2.0
- React-Router: 6.8.0 → 6.8.1
- Updated all peer dependencies
commit y1z2a3b4c5d6789012345678901234567890abcdefg
Author: Jennifer Kim <jennifer.kim@example.com>
Date: Mon Jan 15 08:15:33 2024 +0000
test(auth): add comprehensive tests for OAuth flow
Added unit and integration tests for the OAuth2 authentication system
to ensure reliability and prevent regressions.
- Unit tests for OAuth client configuration
- Integration tests for complete auth flow
- Mock providers for testing without external dependencies
- Error scenario testing
Test coverage increased from 72% to 89% for auth module.
commit c4d5e6f7g8h9012345678901234567890123456abcd
Author: Alex Thompson <alex.thompson@example.com>
Date: Mon Jan 15 07:45:20 2024 +0000
perf(image): implement WebP compression reducing size by 40%
Replaced PNG compression with WebP format for uploaded images.
This reduces average image file sizes by 40% while maintaining
visual quality, improving page load times and reducing bandwidth costs.
- Add WebP encoding support
- Implement fallback to PNG for older browsers
- Add quality settings configuration
- Update image serving endpoints
Performance improvement: Page load time reduced by 25% on average.
commit g7h8i9j0k1l2345678901234567890123456789abcde
Author: Rachel Green <rachel.green@example.com>
Date: Sun Jan 14 16:20:10 2024 +0000
feat(payment): add Stripe payment processor integration
Integrate Stripe as a payment processor to support credit card payments.
This enables users to purchase premium features and subscriptions.
- Add Stripe SDK integration
- Implement payment intent flow
- Add webhook handling for payment status updates
- Include comprehensive error handling and logging
- Add payment method management for users
Closes #567
Co-authored-by: Payment Team <payments@example.com>
commit k1l2m3n4o5p6789012345678901234567890abcdefg
Author: Chris Martinez <chris.martinez@example.com>
Date: Sun Jan 14 15:30:45 2024 +0000
fix(ui): resolve mobile navigation menu overflow issue
Fixed navigation menu overflow on mobile devices where long menu items
were being cut off and causing horizontal scrolling issues.
- Implement responsive text wrapping
- Add horizontal scrolling for overflowing content
- Improve touch targets for better mobile usability
- Fix z-index conflicts with dropdown menus
Fixes #678
Tested on iOS Safari, Chrome Mobile, and Firefox Mobile
FILE:assets/sample_release_plan.json
{
"release_name": "Winter 2024 Release",
"version": "2.3.0",
"target_date": "2024-02-15T10:00:00Z",
"features": [
{
"id": "AUTH-123",
"title": "OAuth2 Integration",
"description": "Add support for Google and GitHub OAuth2 authentication",
"type": "feature",
"assignee": "sarah.johnson@example.com",
"status": "ready",
"pull_request_url": "https://github.com/ourapp/backend/pull/234",
"issue_url": "https://github.com/ourapp/backend/issues/123",
"risk_level": "medium",
"test_coverage_required": 85.0,
"test_coverage_actual": 89.5,
"requires_migration": false,
"breaking_changes": [],
"dependencies": ["AUTH-124"],
"qa_approved": true,
"security_approved": true,
"pm_approved": true
},
{
"id": "UI-345",
"title": "Dashboard Redesign",
"description": "Complete redesign of user dashboard with new component library",
"type": "breaking_change",
"assignee": "david.wilson@example.com",
"status": "ready",
"pull_request_url": "https://github.com/ourapp/frontend/pull/456",
"issue_url": "https://github.com/ourapp/frontend/issues/345",
"risk_level": "high",
"test_coverage_required": 90.0,
"test_coverage_actual": 92.3,
"requires_migration": true,
"migration_complexity": "moderate",
"breaking_changes": [
"Dashboard API endpoints changed from /v1/dashboard to /v2/dashboard",
"Dashboard widget configuration format updated"
],
"dependencies": [],
"qa_approved": true,
"security_approved": true,
"pm_approved": true
},
{
"id": "PAY-567",
"title": "Stripe Payment Integration",
"description": "Add Stripe as payment processor for premium features",
"type": "feature",
"assignee": "rachel.green@example.com",
"status": "ready",
"pull_request_url": "https://github.com/ourapp/backend/pull/678",
"issue_url": "https://github.com/ourapp/backend/issues/567",
"risk_level": "high",
"test_coverage_required": 95.0,
"test_coverage_actual": 97.2,
"requires_migration": true,
"migration_complexity": "complex",
"breaking_changes": [],
"dependencies": ["SEC-890"],
"qa_approved": true,
"security_approved": true,
"pm_approved": true
},
{
"id": "SEARCH-789",
"title": "Elasticsearch Fuzzy Search",
"description": "Implement fuzzy search functionality with Elasticsearch",
"type": "feature",
"assignee": "kevin.park@example.com",
"status": "in_progress",
"pull_request_url": "https://github.com/ourapp/backend/pull/890",
"issue_url": "https://github.com/ourapp/backend/issues/789",
"risk_level": "medium",
"test_coverage_required": 80.0,
"test_coverage_actual": 76.5,
"requires_migration": true,
"migration_complexity": "moderate",
"breaking_changes": [],
"dependencies": ["INFRA-234"],
"qa_approved": false,
"security_approved": true,
"pm_approved": true
},
{
"id": "MOBILE-456",
"title": "Biometric Authentication",
"description": "Add fingerprint and face ID support for mobile apps",
"type": "feature",
"assignee": "alex.thompson@example.com",
"status": "blocked",
"pull_request_url": null,
"issue_url": "https://github.com/ourapp/mobile/issues/456",
"risk_level": "medium",
"test_coverage_required": 85.0,
"test_coverage_actual": null,
"requires_migration": false,
"breaking_changes": [],
"dependencies": ["AUTH-123"],
"qa_approved": false,
"security_approved": false,
"pm_approved": true
},
{
"id": "PERF-678",
"title": "Redis Caching Implementation",
"description": "Implement Redis caching for frequently accessed data",
"type": "performance",
"assignee": "lisa.rodriguez@example.com",
"status": "ready",
"pull_request_url": "https://github.com/ourapp/backend/pull/901",
"issue_url": "https://github.com/ourapp/backend/issues/678",
"risk_level": "low",
"test_coverage_required": 75.0,
"test_coverage_actual": 82.1,
"requires_migration": false,
"breaking_changes": [],
"dependencies": [],
"qa_approved": true,
"security_approved": false,
"pm_approved": true
}
],
"quality_gates": [
{
"name": "Unit Test Coverage",
"required": true,
"status": "ready",
"details": "Overall test coverage above 85% threshold",
"threshold": 85.0,
"actual_value": 87.3
},
{
"name": "Integration Tests",
"required": true,
"status": "ready",
"details": "All integration tests passing"
},
{
"name": "Security Scan",
"required": true,
"status": "pending",
"details": "Waiting for security team review of payment integration"
},
{
"name": "Performance Testing",
"required": true,
"status": "ready",
"details": "Load testing shows 99th percentile response time under 500ms"
},
{
"name": "Documentation Review",
"required": true,
"status": "pending",
"details": "API documentation needs update for dashboard changes"
},
{
"name": "Dependency Audit",
"required": true,
"status": "ready",
"details": "No high or critical vulnerabilities found"
}
],
"stakeholders": [
{
"name": "Engineering Team",
"role": "developer",
"contact": "engineering@example.com",
"notification_type": "slack",
"critical_path": true
},
{
"name": "Product Team",
"role": "pm",
"contact": "product@example.com",
"notification_type": "email",
"critical_path": true
},
{
"name": "QA Team",
"role": "qa",
"contact": "qa@example.com",
"notification_type": "slack",
"critical_path": true
},
{
"name": "Security Team",
"role": "security",
"contact": "security@example.com",
"notification_type": "email",
"critical_path": false
},
{
"name": "Customer Support",
"role": "support",
"contact": "support@example.com",
"notification_type": "email",
"critical_path": false
},
{
"name": "Sales Team",
"role": "sales",
"contact": "sales@example.com",
"notification_type": "email",
"critical_path": false
},
{
"name": "Beta Users",
"role": "customer",
"contact": "beta-users@example.com",
"notification_type": "email",
"critical_path": false
}
],
"rollback_steps": [
{
"order": 1,
"description": "Alert incident response team and stakeholders",
"estimated_time": "2 minutes",
"risk_level": "low",
"verification": "Confirm team is aware and responding via Slack"
},
{
"order": 2,
"description": "Switch load balancer to previous version",
"command": "kubectl patch service app --patch '{\"spec\": {\"selector\": {\"version\": \"v2.2.1\"}}}'",
"estimated_time": "30 seconds",
"risk_level": "low",
"verification": "Check traffic routing to previous version via monitoring dashboard"
},
{
"order": 3,
"description": "Disable new feature flags",
"command": "curl -X POST https://api.example.com/feature-flags/oauth2/disable",
"estimated_time": "1 minute",
"risk_level": "low",
"verification": "Verify feature flags are disabled in admin panel"
},
{
"order": 4,
"description": "Roll back database migrations",
"command": "python manage.py migrate app 0042",
"estimated_time": "10 minutes",
"risk_level": "high",
"verification": "Verify database schema and run data integrity checks"
},
{
"order": 5,
"description": "Clear Redis cache",
"command": "redis-cli FLUSHALL",
"estimated_time": "30 seconds",
"risk_level": "medium",
"verification": "Confirm cache is cleared and application rebuilds cache properly"
},
{
"order": 6,
"description": "Verify application health",
"estimated_time": "5 minutes",
"risk_level": "low",
"verification": "Check health endpoints, error rates, and core user workflows"
},
{
"order": 7,
"description": "Update status page and notify users",
"estimated_time": "5 minutes",
"risk_level": "low",
"verification": "Confirm status page updated and notifications sent"
}
]
}
FILE:changelog_generator.py
#!/usr/bin/env python3
"""
Changelog Generator
Parses git log output in conventional commits format and generates structured changelogs
in multiple formats (Markdown, Keep a Changelog). Groups commits by type, extracts scope,
links to PRs/issues, and highlights breaking changes.
Input: git log text (piped from git log) or JSON array of commits
Output: formatted CHANGELOG.md section + release summary stats
"""
import argparse
import json
import re
import sys
from collections import defaultdict, Counter
from datetime import datetime
from typing import Dict, List, Optional, Tuple, Union
class ConventionalCommit:
"""Represents a parsed conventional commit."""
def __init__(self, raw_message: str, commit_hash: str = "", author: str = "",
date: str = "", merge_info: Optional[str] = None):
self.raw_message = raw_message
self.commit_hash = commit_hash
self.author = author
self.date = date
self.merge_info = merge_info
# Parse the commit message
self.type = ""
self.scope = ""
self.description = ""
self.body = ""
self.footers = []
self.is_breaking = False
self.breaking_change_description = ""
self._parse_commit_message()
def _parse_commit_message(self):
"""Parse conventional commit format."""
lines = self.raw_message.split('\n')
header = lines[0] if lines else ""
# Parse header: type(scope): description
header_pattern = r'^(\w+)(\([^)]+\))?(!)?:\s*(.+)$'
match = re.match(header_pattern, header)
if match:
self.type = match.group(1).lower()
scope_match = match.group(2)
self.scope = scope_match[1:-1] if scope_match else "" # Remove parentheses
self.is_breaking = bool(match.group(3)) # ! indicates breaking change
self.description = match.group(4).strip()
else:
# Fallback for non-conventional commits
self.type = "chore"
self.description = header
# Parse body and footers
if len(lines) > 1:
body_lines = []
footer_lines = []
in_footer = False
for line in lines[1:]:
if not line.strip():
continue
# Check if this is a footer (KEY: value or KEY #value format)
footer_pattern = r'^([A-Z-]+):\s*(.+)$|^([A-Z-]+)\s+#(\d+)$'
if re.match(footer_pattern, line):
in_footer = True
footer_lines.append(line)
# Check for breaking change
if line.startswith('BREAKING CHANGE:'):
self.is_breaking = True
self.breaking_change_description = line[16:].strip()
else:
if in_footer:
# Continuation of footer
footer_lines.append(line)
else:
body_lines.append(line)
self.body = '\n'.join(body_lines).strip()
self.footers = footer_lines
def extract_issue_references(self) -> List[str]:
"""Extract issue/PR references like #123, fixes #456, etc."""
text = f"{self.description} {self.body} {' '.join(self.footers)}"
# Common patterns for issue references
patterns = [
r'#(\d+)', # Simple #123
r'(?:close[sd]?|fix(?:e[sd])?|resolve[sd]?)\s+#(\d+)', # closes #123
r'(?:close[sd]?|fix(?:e[sd])?|resolve[sd]?)\s+(\w+/\w+)?#(\d+)' # fixes repo#123
]
references = []
for pattern in patterns:
matches = re.findall(pattern, text, re.IGNORECASE)
for match in matches:
if isinstance(match, tuple):
# Handle tuple results from more complex patterns
ref = match[-1] if match[-1] else match[0]
else:
ref = match
if ref and ref not in references:
references.append(ref)
return references
def get_changelog_category(self) -> str:
"""Map commit type to changelog category."""
category_map = {
'feat': 'Added',
'add': 'Added',
'fix': 'Fixed',
'bugfix': 'Fixed',
'security': 'Security',
'perf': 'Fixed', # Performance improvements go to Fixed
'refactor': 'Changed',
'style': 'Changed',
'docs': 'Changed',
'test': None, # Tests don't appear in user-facing changelog
'ci': None,
'build': None,
'chore': None,
'revert': 'Fixed',
'remove': 'Removed',
'deprecate': 'Deprecated'
}
return category_map.get(self.type, 'Changed')
class ChangelogGenerator:
"""Main changelog generator class."""
def __init__(self):
self.commits: List[ConventionalCommit] = []
self.version = "Unreleased"
self.date = datetime.now().strftime("%Y-%m-%d")
self.base_url = ""
def parse_git_log_output(self, git_log_text: str):
"""Parse git log output into ConventionalCommit objects."""
# Try to detect format based on patterns in the text
lines = git_log_text.strip().split('\n')
if not lines or not lines[0]:
return
# Format 1: Simple oneline format (hash message)
oneline_pattern = r'^([a-f0-9]{7,40})\s+(.+)$'
# Format 2: Full format with metadata
full_pattern = r'^commit\s+([a-f0-9]+)'
current_commit = None
commit_buffer = []
for line in lines:
line = line.strip()
if not line:
continue
# Check if this is a new commit (oneline format)
oneline_match = re.match(oneline_pattern, line)
if oneline_match:
# Process previous commit
if current_commit:
self.commits.append(current_commit)
# Start new commit
commit_hash = oneline_match.group(1)
message = oneline_match.group(2)
current_commit = ConventionalCommit(message, commit_hash)
continue
# Check if this is a new commit (full format)
full_match = re.match(full_pattern, line)
if full_match:
# Process previous commit
if current_commit:
commit_message = '\n'.join(commit_buffer).strip()
if commit_message:
current_commit = ConventionalCommit(commit_message, current_commit.commit_hash,
current_commit.author, current_commit.date)
self.commits.append(current_commit)
# Start new commit
commit_hash = full_match.group(1)
current_commit = ConventionalCommit("", commit_hash)
commit_buffer = []
continue
# Parse metadata lines in full format
if current_commit and not current_commit.raw_message:
if line.startswith('Author:'):
current_commit.author = line[7:].strip()
elif line.startswith('Date:'):
current_commit.date = line[5:].strip()
elif line.startswith('Merge:'):
current_commit.merge_info = line[6:].strip()
elif line.startswith(' '):
# Commit message line (indented)
commit_buffer.append(line[4:]) # Remove 4-space indent
# Process final commit
if current_commit:
if commit_buffer:
commit_message = '\n'.join(commit_buffer).strip()
current_commit = ConventionalCommit(commit_message, current_commit.commit_hash,
current_commit.author, current_commit.date)
self.commits.append(current_commit)
def parse_json_commits(self, json_data: Union[str, List[Dict]]):
"""Parse commits from JSON format."""
if isinstance(json_data, str):
data = json.loads(json_data)
else:
data = json_data
for commit_data in data:
commit = ConventionalCommit(
raw_message=commit_data.get('message', ''),
commit_hash=commit_data.get('hash', ''),
author=commit_data.get('author', ''),
date=commit_data.get('date', '')
)
self.commits.append(commit)
def group_commits_by_category(self) -> Dict[str, List[ConventionalCommit]]:
"""Group commits by changelog category."""
categories = defaultdict(list)
for commit in self.commits:
category = commit.get_changelog_category()
if category: # Skip None categories (internal changes)
categories[category].append(commit)
return dict(categories)
def generate_markdown_changelog(self, include_unreleased: bool = True) -> str:
"""Generate Keep a Changelog format markdown."""
grouped_commits = self.group_commits_by_category()
if not grouped_commits:
return "No notable changes.\n"
# Start with header
changelog = []
if include_unreleased and self.version == "Unreleased":
changelog.append(f"## [{self.version}]")
else:
changelog.append(f"## [{self.version}] - {self.date}")
changelog.append("")
# Order categories logically
category_order = ['Added', 'Changed', 'Deprecated', 'Removed', 'Fixed', 'Security']
# Separate breaking changes
breaking_changes = [commit for commit in self.commits if commit.is_breaking]
# Add breaking changes section first if any exist
if breaking_changes:
changelog.append("### Breaking Changes")
for commit in breaking_changes:
line = self._format_commit_line(commit, show_breaking=True)
changelog.append(f"- {line}")
changelog.append("")
# Add regular categories
for category in category_order:
if category not in grouped_commits:
continue
changelog.append(f"### {category}")
# Group by scope for better organization
scoped_commits = defaultdict(list)
for commit in grouped_commits[category]:
scope = commit.scope if commit.scope else "general"
scoped_commits[scope].append(commit)
# Sort scopes, with 'general' last
scopes = sorted(scoped_commits.keys())
if "general" in scopes:
scopes.remove("general")
scopes.append("general")
for scope in scopes:
if len(scoped_commits) > 1 and scope != "general":
changelog.append(f"#### {scope.title()}")
for commit in scoped_commits[scope]:
line = self._format_commit_line(commit)
changelog.append(f"- {line}")
changelog.append("")
return '\n'.join(changelog)
def _format_commit_line(self, commit: ConventionalCommit, show_breaking: bool = False) -> str:
"""Format a single commit line for the changelog."""
# Start with description
line = commit.description.capitalize()
# Add scope if present and not already in description
if commit.scope and commit.scope.lower() not in line.lower():
line = f"{commit.scope}: {line}"
# Add issue references
issue_refs = commit.extract_issue_references()
if issue_refs:
refs_str = ', '.join(f"#{ref}" for ref in issue_refs)
line += f" ({refs_str})"
# Add commit hash if available
if commit.commit_hash:
short_hash = commit.commit_hash[:7]
line += f" [{short_hash}]"
if self.base_url:
line += f"({self.base_url}/commit/{commit.commit_hash})"
# Add breaking change indicator
if show_breaking and commit.breaking_change_description:
line += f" - {commit.breaking_change_description}"
elif commit.is_breaking and not show_breaking:
line += " ⚠️ BREAKING"
return line
def generate_release_summary(self) -> Dict:
"""Generate summary statistics for the release."""
if not self.commits:
return {
'version': self.version,
'date': self.date,
'total_commits': 0,
'by_type': {},
'by_author': {},
'breaking_changes': 0,
'notable_changes': 0
}
# Count by type
type_counts = Counter(commit.type for commit in self.commits)
# Count by author
author_counts = Counter(commit.author for commit in self.commits if commit.author)
# Count breaking changes
breaking_count = sum(1 for commit in self.commits if commit.is_breaking)
# Count notable changes (excluding chore, ci, build, test)
notable_types = {'feat', 'fix', 'security', 'perf', 'refactor', 'remove', 'deprecate'}
notable_count = sum(1 for commit in self.commits if commit.type in notable_types)
return {
'version': self.version,
'date': self.date,
'total_commits': len(self.commits),
'by_type': dict(type_counts.most_common()),
'by_author': dict(author_counts.most_common(10)), # Top 10 contributors
'breaking_changes': breaking_count,
'notable_changes': notable_count,
'scopes': list(set(commit.scope for commit in self.commits if commit.scope)),
'issue_references': len(set().union(*(commit.extract_issue_references() for commit in self.commits)))
}
def generate_json_output(self) -> str:
"""Generate JSON representation of the changelog data."""
grouped_commits = self.group_commits_by_category()
# Convert commits to serializable format
json_data = {
'version': self.version,
'date': self.date,
'summary': self.generate_release_summary(),
'categories': {}
}
for category, commits in grouped_commits.items():
json_data['categories'][category] = []
for commit in commits:
commit_data = {
'type': commit.type,
'scope': commit.scope,
'description': commit.description,
'hash': commit.commit_hash,
'author': commit.author,
'date': commit.date,
'breaking': commit.is_breaking,
'breaking_description': commit.breaking_change_description,
'issue_references': commit.extract_issue_references()
}
json_data['categories'][category].append(commit_data)
return json.dumps(json_data, indent=2)
def main():
"""Main entry point with CLI argument parsing."""
parser = argparse.ArgumentParser(description="Generate changelog from conventional commits")
parser.add_argument('--input', '-i', type=str, help='Input file (default: stdin)')
parser.add_argument('--format', '-f', choices=['markdown', 'json', 'both'],
default='markdown', help='Output format')
parser.add_argument('--version', '-v', type=str, default='Unreleased',
help='Version for this release')
parser.add_argument('--date', '-d', type=str,
default=datetime.now().strftime("%Y-%m-%d"),
help='Release date (YYYY-MM-DD format)')
parser.add_argument('--base-url', '-u', type=str, default='',
help='Base URL for commit links')
parser.add_argument('--input-format', choices=['git-log', 'json'],
default='git-log', help='Input format')
parser.add_argument('--output', '-o', type=str, help='Output file (default: stdout)')
parser.add_argument('--summary', '-s', action='store_true',
help='Include release summary statistics')
args = parser.parse_args()
# Read input
if args.input:
with open(args.input, 'r', encoding='utf-8') as f:
input_data = f.read()
else:
input_data = sys.stdin.read()
if not input_data.strip():
print("No input data provided", file=sys.stderr)
sys.exit(1)
# Initialize generator
generator = ChangelogGenerator()
generator.version = args.version
generator.date = args.date
generator.base_url = args.base_url
# Parse input
try:
if args.input_format == 'json':
generator.parse_json_commits(input_data)
else:
generator.parse_git_log_output(input_data)
except Exception as e:
print(f"Error parsing input: {e}", file=sys.stderr)
sys.exit(1)
if not generator.commits:
print("No valid commits found in input", file=sys.stderr)
sys.exit(1)
# Generate output
output_lines = []
if args.format in ['markdown', 'both']:
changelog_md = generator.generate_markdown_changelog()
if args.format == 'both':
output_lines.append("# Markdown Changelog\n")
output_lines.append(changelog_md)
if args.format in ['json', 'both']:
changelog_json = generator.generate_json_output()
if args.format == 'both':
output_lines.append("\n# JSON Output\n")
output_lines.append(changelog_json)
if args.summary:
summary = generator.generate_release_summary()
output_lines.append(f"\n# Release Summary")
output_lines.append(f"- **Version:** {summary['version']}")
output_lines.append(f"- **Total Commits:** {summary['total_commits']}")
output_lines.append(f"- **Notable Changes:** {summary['notable_changes']}")
output_lines.append(f"- **Breaking Changes:** {summary['breaking_changes']}")
output_lines.append(f"- **Issue References:** {summary['issue_references']}")
if summary['by_type']:
output_lines.append("- **By Type:**")
for commit_type, count in summary['by_type'].items():
output_lines.append(f" - {commit_type}: {count}")
# Write output
final_output = '\n'.join(output_lines)
if args.output:
with open(args.output, 'w', encoding='utf-8') as f:
f.write(final_output)
else:
print(final_output)
if __name__ == '__main__':
main()
FILE:expected_outputs/changelog_example.md
# Expected Changelog Output
## [2.3.0] - 2024-01-15
### Breaking Changes
- ui: redesign dashboard with new component library - The dashboard API endpoints have changed structure. Frontend clients must update to use the new /v2/dashboard endpoints. The legacy /v1/dashboard endpoints will be removed in version 3.0.0. (#345, #367, #389) [m1n2o3p]
### Added
- auth: add OAuth2 integration with Google and GitHub (#123, #145) [a1b2c3d]
- payment: add Stripe payment processor integration (#567) [g7h8i9j]
- search: implement fuzzy search with Elasticsearch (#789) [s7t8u9v]
### Fixed
- api: resolve race condition in user creation endpoint (#234) [e4f5g6h]
- db: optimize slow query in user search functionality (#456) [q4r5s6t]
- ui: resolve mobile navigation menu overflow issue (#678) [k1l2m3n]
- security: patch SQL injection vulnerability in reports [w1x2y3z] ⚠️ BREAKING
### Changed
- image: implement WebP compression reducing size by 40% [c4d5e6f]
- api: extract validation logic into reusable middleware [o4p5q6r]
- readme: update installation and deployment instructions [i7j8k9l]
# Release Summary
- **Version:** 2.3.0
- **Total Commits:** 13
- **Notable Changes:** 9
- **Breaking Changes:** 2
- **Issue References:** 8
- **By Type:**
- feat: 4
- fix: 4
- perf: 1
- refactor: 1
- docs: 1
- test: 1
- chore: 1
FILE:expected_outputs/release_readiness_example.txt
Release Readiness Report
========================
Release: Winter 2024 Release v2.3.0
Status: AT_RISK
Readiness Score: 73.3%
WARNINGS:
⚠️ Feature 'Elasticsearch Fuzzy Search' (SEARCH-789) still in progress
⚠️ Feature 'Elasticsearch Fuzzy Search' has low test coverage: 76.5% < 80.0%
⚠️ Required quality gate 'Security Scan' is pending
⚠️ Required quality gate 'Documentation Review' is pending
BLOCKING ISSUES:
❌ Feature 'Biometric Authentication' (MOBILE-456) is blocked
❌ Feature 'Biometric Authentication' missing approvals: QA approval, Security approval
RECOMMENDATIONS:
💡 Obtain required approvals for pending features
💡 Improve test coverage for features below threshold
💡 Complete pending quality gate validations
FEATURE SUMMARY:
Total: 6 | Ready: 3 | Blocked: 1
Breaking Changes: 1 | Missing Approvals: 1
QUALITY GATES:
Total: 6 | Passed: 3 | Failed: 0
FILE:expected_outputs/version_bump_example.txt
Current Version: 2.2.5
Recommended Version: 3.0.0
With v prefix: v3.0.0
Bump Type: major
Commit Analysis:
- Total commits: 13
- Breaking changes: 2
- New features: 4
- Bug fixes: 4
- Ignored commits: 3
Breaking Changes:
- feat(ui): redesign dashboard with new component library
- fix(security): patch SQL injection vulnerability in reports
Bump Commands:
npm:
npm version 3.0.0 --no-git-tag-version
python:
# Update version in setup.py, __init__.py, or pyproject.toml
# pyproject.toml: version = "3.0.0"
rust:
# Update Cargo.toml
# version = "3.0.0"
git:
git tag -a v3.0.0 -m 'Release v3.0.0'
git push origin v3.0.0
docker:
docker build -t myapp:3.0.0 .
docker tag myapp:3.0.0 myapp:latest
FILE:README.md
# Release Manager
A comprehensive release management toolkit for automating changelog generation, version bumping, and release planning based on conventional commits and industry best practices.
## Overview
The Release Manager skill provides three powerful Python scripts and comprehensive documentation for managing software releases:
1. **changelog_generator.py** - Generate structured changelogs from git history
2. **version_bumper.py** - Determine correct semantic version bumps
3. **release_planner.py** - Assess release readiness and generate coordination plans
## Quick Start
### Prerequisites
- Python 3.7+
- Git repository with conventional commit messages
- No external dependencies required (uses only Python standard library)
### Basic Usage
```bash
# Generate changelog from recent commits
git log --oneline --since="1 month ago" | python changelog_generator.py
# Determine version bump from commits since last tag
git log --oneline $(git describe --tags --abbrev=0)..HEAD | python version_bumper.py -c "1.2.3"
# Assess release readiness
python release_planner.py --input assets/sample_release_plan.json
```
## Scripts Reference
### changelog_generator.py
Parses conventional commits and generates structured changelogs in multiple formats.
**Input Options:**
- Git log text (oneline or full format)
- JSON array of commits
- Stdin or file input
**Output Formats:**
- Markdown (Keep a Changelog format)
- JSON structured data
- Both with release statistics
```bash
# From git log (recommended)
git log --oneline --since="last release" | python changelog_generator.py \
--version "2.1.0" \
--date "2024-01-15" \
--base-url "https://github.com/yourorg/yourrepo"
# From JSON file
python changelog_generator.py \
--input assets/sample_commits.json \
--input-format json \
--format both \
--summary
# With custom output
git log --format="%h %s" v1.0.0..HEAD | python changelog_generator.py \
--version "1.1.0" \
--output CHANGELOG_DRAFT.md
```
**Features:**
- Parses conventional commit types (feat, fix, docs, etc.)
- Groups commits by changelog categories (Added, Fixed, Changed, etc.)
- Extracts issue references (#123, fixes #456)
- Identifies breaking changes
- Links to commits and PRs
- Generates release summary statistics
### version_bumper.py
Analyzes commits to determine semantic version bumps according to conventional commits.
**Bump Rules:**
- **MAJOR:** Breaking changes (`feat!:` or `BREAKING CHANGE:`)
- **MINOR:** New features (`feat:`)
- **PATCH:** Bug fixes (`fix:`, `perf:`, `security:`)
- **NONE:** Documentation, tests, chores only
```bash
# Basic version bump determination
git log --oneline v1.2.3..HEAD | python version_bumper.py --current-version "1.2.3"
# With pre-release version
python version_bumper.py \
--current-version "1.2.3" \
--prerelease alpha \
--input assets/sample_commits.json \
--input-format json
# Include bump commands and file updates
git log --oneline $(git describe --tags --abbrev=0)..HEAD | \
python version_bumper.py \
--current-version "$(git describe --tags --abbrev=0)" \
--include-commands \
--include-files \
--analysis
```
**Features:**
- Supports pre-release versions (alpha, beta, rc)
- Generates bump commands for npm, Python, Rust, Git
- Provides file update snippets
- Detailed commit analysis and categorization
- Custom rules for specific commit types
- JSON and text output formats
### release_planner.py
Assesses release readiness and generates comprehensive release coordination plans.
**Input:** JSON release plan with features, quality gates, and stakeholders
```bash
# Assess release readiness
python release_planner.py --input assets/sample_release_plan.json
# Generate full release package
python release_planner.py \
--input release_plan.json \
--output-format markdown \
--include-checklist \
--include-communication \
--include-rollback \
--output release_report.md
```
**Features:**
- Feature readiness assessment with approval tracking
- Quality gate validation and reporting
- Stakeholder communication planning
- Rollback procedure generation
- Risk analysis and timeline assessment
- Customizable test coverage thresholds
- Multiple output formats (text, JSON, Markdown)
## File Structure
```
release-manager/
├── SKILL.md # Comprehensive methodology guide
├── README.md # This file
├── changelog_generator.py # Changelog generation script
├── version_bumper.py # Version bump determination
├── release_planner.py # Release readiness assessment
├── references/ # Reference documentation
│ ├── conventional-commits-guide.md # Conventional commits specification
│ ├── release-workflow-comparison.md # Git Flow vs GitHub Flow vs Trunk-based
│ └── hotfix-procedures.md # Emergency release procedures
├── assets/ # Sample data for testing
│ ├── sample_git_log.txt # Sample git log output
│ ├── sample_git_log_full.txt # Detailed git log format
│ ├── sample_commits.json # JSON commit data
│ └── sample_release_plan.json # Release plan template
└── expected_outputs/ # Example script outputs
├── changelog_example.md # Expected changelog format
├── version_bump_example.txt # Version bump output
└── release_readiness_example.txt # Release assessment report
```
## Integration Examples
### CI/CD Pipeline Integration
```yaml
# .github/workflows/release.yml
name: Automated Release
on:
push:
branches: [main]
jobs:
release:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v3
with:
fetch-depth: 0 # Need full history
- name: Determine version bump
id: version
run: |
CURRENT=$(git describe --tags --abbrev=0)
git log --oneline $CURRENT..HEAD | \
python scripts/version_bumper.py -c $CURRENT --output-format json > bump.json
echo "new_version=$(jq -r '.recommended_version' bump.json)" >> $GITHUB_OUTPUT
- name: Generate changelog
run: |
git log --oneline { steps.version.outputs.current_version}..HEAD | \
python scripts/changelog_generator.py \
--version "{ steps.version.outputs.new_version}" \
--base-url "https://github.com/{ github.repository}" \
--output CHANGELOG_ENTRY.md
- name: Create release
uses: actions/create-release@v1
with:
tag_name: v{ steps.version.outputs.new_version}
release_name: Release { steps.version.outputs.new_version}
body_path: CHANGELOG_ENTRY.md
```
### Git Hooks Integration
```bash
#!/bin/bash
# .git/hooks/pre-commit
# Validate conventional commit format
commit_msg_file=$1
commit_msg=$(cat $commit_msg_file)
# Simple validation (more sophisticated validation available in commitlint)
if ! echo "$commit_msg" | grep -qE "^(feat|fix|docs|style|refactor|test|chore|perf|ci|build)(\(.+\))?(!)?:"; then
echo "❌ Commit message doesn't follow conventional commits format"
echo "Expected: type(scope): description"
echo "Examples:"
echo " feat(auth): add OAuth2 integration"
echo " fix(api): resolve race condition"
echo " docs: update installation guide"
exit 1
fi
echo "✅ Commit message format is valid"
```
### Release Planning Automation
```python
#!/usr/bin/env python3
# generate_release_plan.py - Automatically generate release plans from project management tools
import json
import requests
from datetime import datetime, timedelta
def generate_release_plan_from_github(repo, milestone):
"""Generate release plan from GitHub milestone and PRs."""
# Fetch milestone details
milestone_url = f"https://api.github.com/repos/{repo}/milestones/{milestone}"
milestone_data = requests.get(milestone_url).json()
# Fetch associated issues/PRs
issues_url = f"https://api.github.com/repos/{repo}/issues?milestone={milestone}&state=all"
issues = requests.get(issues_url).json()
release_plan = {
"release_name": milestone_data["title"],
"version": "TBD", # Fill in manually or extract from milestone
"target_date": milestone_data["due_on"],
"features": []
}
for issue in issues:
if issue.get("pull_request"): # It's a PR
feature = {
"id": f"GH-{issue['number']}",
"title": issue["title"],
"description": issue["body"][:200] + "..." if len(issue["body"]) > 200 else issue["body"],
"type": "feature", # Could be parsed from labels
"assignee": issue["assignee"]["login"] if issue["assignee"] else "",
"status": "ready" if issue["state"] == "closed" else "in_progress",
"pull_request_url": issue["pull_request"]["html_url"],
"issue_url": issue["html_url"],
"risk_level": "medium", # Could be parsed from labels
"qa_approved": "qa-approved" in [label["name"] for label in issue["labels"]],
"pm_approved": "pm-approved" in [label["name"] for label in issue["labels"]]
}
release_plan["features"].append(feature)
return release_plan
# Usage
if __name__ == "__main__":
plan = generate_release_plan_from_github("yourorg/yourrepo", "5")
with open("release_plan.json", "w") as f:
json.dump(plan, f, indent=2)
print("Generated release_plan.json")
print("Run: python release_planner.py --input release_plan.json")
```
## Advanced Usage
### Custom Commit Type Rules
```bash
# Define custom rules for version bumping
python version_bumper.py \
--current-version "1.2.3" \
--custom-rules '{"security": "patch", "breaking": "major"}' \
--ignore-types "docs,style,test"
```
### Multi-repository Release Coordination
```bash
#!/bin/bash
# multi_repo_release.sh - Coordinate releases across multiple repositories
repos=("frontend" "backend" "mobile" "docs")
base_version="2.1.0"
for repo in "repos[@]"; do
echo "Processing $repo..."
cd "$repo"
# Generate changelog for this repo
git log --oneline --since="1 month ago" | \
python ../scripts/changelog_generator.py \
--version "$base_version" \
--output "CHANGELOG_$repo.md"
# Determine version bump
git log --oneline $(git describe --tags --abbrev=0)..HEAD | \
python ../scripts/version_bumper.py \
--current-version "$(git describe --tags --abbrev=0)" > "VERSION_$repo.txt"
cd ..
done
echo "Generated changelogs and version recommendations for all repositories"
```
### Integration with Slack/Teams
```python
#!/usr/bin/env python3
# notify_release_status.py
import json
import requests
import subprocess
def send_slack_notification(webhook_url, message):
payload = {"text": message}
requests.post(webhook_url, json=payload)
def get_release_status():
"""Get current release status from release planner."""
result = subprocess.run(
["python", "release_planner.py", "--input", "release_plan.json", "--output-format", "json"],
capture_output=True, text=True
)
return json.loads(result.stdout)
# Usage in CI/CD
status = get_release_status()
if status["assessment"]["overall_status"] == "blocked":
message = f"🚫 Release {status['version']} is BLOCKED\n"
message += f"Issues: {', '.join(status['assessment']['blocking_issues'])}"
send_slack_notification(SLACK_WEBHOOK_URL, message)
elif status["assessment"]["overall_status"] == "ready":
message = f"✅ Release {status['version']} is READY for deployment!"
send_slack_notification(SLACK_WEBHOOK_URL, message)
```
## Best Practices
### Commit Message Guidelines
1. **Use conventional commits consistently** across your team
2. **Be specific** in commit descriptions: "fix: resolve race condition in user creation" vs "fix: bug"
3. **Reference issues** when applicable: "Closes #123" or "Fixes #456"
4. **Mark breaking changes** clearly with `!` or `BREAKING CHANGE:` footer
5. **Keep first line under 50 characters** when possible
### Release Planning
1. **Plan releases early** with clear feature lists and target dates
2. **Set quality gates** and stick to them (test coverage, security scans, etc.)
3. **Track approvals** from all relevant stakeholders
4. **Document rollback procedures** before deployment
5. **Communicate clearly** with both internal teams and external users
### Version Management
1. **Follow semantic versioning** strictly for predictable releases
2. **Use pre-release versions** for beta testing and gradual rollouts
3. **Tag releases consistently** with proper version numbers
4. **Maintain backwards compatibility** when possible to avoid major version bumps
5. **Document breaking changes** thoroughly with migration guides
## Troubleshooting
### Common Issues
**"No valid commits found"**
- Ensure git log contains commit messages
- Check that commits follow conventional format
- Verify input format (git-log vs json)
**"Invalid version format"**
- Use semantic versioning: 1.2.3, not 1.2 or v1.2.3.beta
- Pre-release format: 1.2.3-alpha.1
**"Missing required approvals"**
- Check feature risk levels in release plan
- High/critical risk features require additional approvals
- Update approval status in JSON file
### Debug Mode
All scripts support verbose output for debugging:
```bash
# Add debug logging
python changelog_generator.py --input sample.txt --debug
# Validate input data
python -c "import json; print(json.load(open('release_plan.json')))"
# Test with sample data first
python release_planner.py --input assets/sample_release_plan.json
```
## Contributing
When extending these scripts:
1. **Maintain backwards compatibility** for existing command-line interfaces
2. **Add comprehensive tests** for new features
3. **Update documentation** including this README and SKILL.md
4. **Follow Python standards** (PEP 8, type hints where helpful)
5. **Use only standard library** to avoid dependencies
## License
This skill is part of the claude-skills repository and follows the same license terms.
---
For detailed methodology and background information, see [SKILL.md](SKILL.md).
For specific workflow guidance, see the [references](references/) directory.
For testing the scripts, use the sample data in the [assets](assets/) directory.
FILE:references/conventional-commits-guide.md
# Conventional Commits Guide
## Overview
Conventional Commits is a specification for adding human and machine readable meaning to commit messages. The specification provides an easy set of rules for creating an explicit commit history, which makes it easier to write automated tools for version management, changelog generation, and release planning.
## Basic Format
```
<type>[optional scope]: <description>
[optional body]
[optional footer(s)]
```
## Commit Types
### Primary Types
- **feat**: A new feature for the user (correlates with MINOR in semantic versioning)
- **fix**: A bug fix for the user (correlates with PATCH in semantic versioning)
### Secondary Types
- **build**: Changes that affect the build system or external dependencies (webpack, npm, etc.)
- **ci**: Changes to CI configuration files and scripts (Travis, Circle, BrowserStack, SauceLabs)
- **docs**: Documentation only changes
- **perf**: A code change that improves performance
- **refactor**: A code change that neither fixes a bug nor adds a feature
- **style**: Changes that do not affect the meaning of the code (white-space, formatting, missing semi-colons, etc.)
- **test**: Adding missing tests or correcting existing tests
- **chore**: Other changes that don't modify src or test files
- **revert**: Reverts a previous commit
### Breaking Changes
Any commit can introduce a breaking change by:
1. Adding `!` after the type: `feat!: remove deprecated API`
2. Including `BREAKING CHANGE:` in the footer
## Scopes
Scopes provide additional contextual information about the change. They should be noun describing a section of the codebase:
- `auth` - Authentication and authorization
- `api` - API changes
- `ui` - User interface
- `db` - Database related changes
- `config` - Configuration changes
- `deps` - Dependency updates
## Examples
### Simple Feature
```
feat(auth): add OAuth2 integration
Integrate OAuth2 authentication with Google and GitHub providers.
Users can now log in using their existing social media accounts.
```
### Bug Fix
```
fix(api): resolve race condition in user creation
When multiple requests tried to create users with the same email
simultaneously, duplicate records were sometimes created. Added
proper database constraints and error handling.
Fixes #234
```
### Breaking Change with !
```
feat(api)!: remove deprecated /v1/users endpoint
The deprecated /v1/users endpoint has been removed. All clients
should migrate to /v2/users which provides better performance
and additional features.
BREAKING CHANGE: /v1/users endpoint removed, use /v2/users instead
```
### Breaking Change with Footer
```
feat(auth): implement new authentication flow
Add support for multi-factor authentication and improved session
management. This change requires all users to re-authenticate.
BREAKING CHANGE: Authentication tokens issued before this release
are no longer valid. Users must log in again.
```
### Performance Improvement
```
perf(image): optimize image compression algorithm
Replaced PNG compression with WebP format, reducing image sizes
by 40% on average while maintaining visual quality.
Closes #456
```
### Dependency Update
```
build(deps): upgrade React to version 18.2.0
Updates React and related packages to latest stable versions.
Includes performance improvements and new concurrent features.
```
### Documentation
```
docs(readme): add deployment instructions
Added comprehensive deployment guide including Docker setup,
environment variables configuration, and troubleshooting tips.
```
### Revert
```
revert: feat(payment): add cryptocurrency support
This reverts commit 667ecc1654a317a13331b17617d973392f415f02.
Reverting due to security concerns identified in code review.
The feature will be re-implemented with proper security measures.
```
## Multi-paragraph Body
For complex changes, use multiple paragraphs in the body:
```
feat(search): implement advanced search functionality
Add support for complex search queries including:
- Boolean operators (AND, OR, NOT)
- Field-specific searches (title:, author:, date:)
- Fuzzy matching with configurable threshold
- Search result highlighting
The search index has been restructured to support these new
features while maintaining backward compatibility with existing
simple search queries.
Performance testing shows less than 10ms impact on search
response times even with complex queries.
Closes #789, #823, #901
```
## Footers
### Issue References
```
Fixes #123
Closes #234, #345
Resolves #456
```
### Breaking Changes
```
BREAKING CHANGE: The `authenticate` function now requires a second
parameter for the authentication method. Update all calls from
`authenticate(token)` to `authenticate(token, 'bearer')`.
```
### Co-authors
```
Co-authored-by: Jane Doe <jane@example.com>
Co-authored-by: John Smith <john@example.com>
```
### Reviewed By
```
Reviewed-by: Senior Developer <senior@example.com>
Acked-by: Tech Lead <lead@example.com>
```
## Automation Benefits
Using conventional commits enables:
### Automatic Version Bumping
- `fix` commits trigger PATCH version bump (1.0.0 → 1.0.1)
- `feat` commits trigger MINOR version bump (1.0.0 → 1.1.0)
- `BREAKING CHANGE` triggers MAJOR version bump (1.0.0 → 2.0.0)
### Changelog Generation
```markdown
## [1.2.0] - 2024-01-15
### Added
- OAuth2 integration (auth)
- Advanced search functionality (search)
### Fixed
- Race condition in user creation (api)
- Memory leak in image processing (image)
### Breaking Changes
- Authentication tokens issued before this release are no longer valid
```
### Release Notes
Generate user-friendly release notes automatically from commit history, filtering out internal changes and highlighting user-facing improvements.
## Best Practices
### Writing Good Descriptions
- Use imperative mood: "add feature" not "added feature"
- Start with lowercase letter
- No period at the end
- Limit to 50 characters when possible
- Be specific and descriptive
### Good Examples
```
feat(auth): add password reset functionality
fix(ui): resolve mobile navigation menu overflow
perf(db): optimize user query with proper indexing
```
### Bad Examples
```
feat: stuff
fix: bug
update: changes
```
### Body Guidelines
- Separate subject from body with blank line
- Wrap body at 72 characters
- Use body to explain what and why, not how
- Reference issues and PRs when relevant
### Scope Guidelines
- Use consistent scope naming across the team
- Keep scopes short and meaningful
- Document your team's scope conventions
- Consider using scopes that match your codebase structure
## Tools and Integration
### Git Hooks
Use tools like `commitizen` or `husky` to enforce conventional commit format:
```bash
# Install commitizen
npm install -g commitizen cz-conventional-changelog
# Configure
echo '{ "path": "cz-conventional-changelog" }' > ~/.czrc
# Use
git cz
```
### Automated Validation
Add commit message validation to prevent non-conventional commits:
```javascript
// commitlint.config.js
module.exports = {
extends: ['@commitlint/config-conventional'],
rules: {
'type-enum': [
2, 'always',
['feat', 'fix', 'docs', 'style', 'refactor', 'perf', 'test', 'build', 'ci', 'chore', 'revert']
],
'subject-case': [2, 'always', 'lower-case'],
'subject-max-length': [2, 'always', 50]
}
};
```
### CI/CD Integration
Integrate with release automation tools:
- **semantic-release**: Automated version management and package publishing
- **standard-version**: Generate changelog and tag releases
- **release-please**: Google's release automation tool
## Common Mistakes
### Mixing Multiple Changes
```
# Bad: Multiple unrelated changes
feat: add login page and fix CSS bug and update dependencies
# Good: Separate commits
feat(auth): add login page
fix(ui): resolve CSS styling issue
build(deps): update React to version 18
```
### Vague Descriptions
```
# Bad: Not descriptive
fix: bug in code
feat: new stuff
# Good: Specific and clear
fix(api): resolve null pointer exception in user validation
feat(search): implement fuzzy matching algorithm
```
### Missing Breaking Change Indicators
```
# Bad: Breaking change not marked
feat(api): update user authentication
# Good: Properly marked breaking change
feat(api)!: update user authentication
BREAKING CHANGE: All API clients must now include authentication
headers in every request. Anonymous access is no longer supported.
```
## Team Guidelines
### Establishing Conventions
1. **Define scope vocabulary**: Create a list of approved scopes for your project
2. **Document examples**: Provide team-specific examples of good commits
3. **Set up tooling**: Use linters and hooks to enforce standards
4. **Review process**: Include commit message quality in code reviews
5. **Training**: Ensure all team members understand the format
### Scope Examples by Project Type
**Web Application:**
- `auth`, `ui`, `api`, `db`, `config`, `deploy`
**Library/SDK:**
- `core`, `utils`, `docs`, `examples`, `tests`
**Mobile App:**
- `ios`, `android`, `shared`, `ui`, `network`, `storage`
By following conventional commits consistently, your team will have a clear, searchable commit history that enables powerful automation and improves the overall development workflow.
FILE:references/hotfix-procedures.md
# Hotfix Procedures
## Overview
Hotfixes are emergency releases designed to address critical production issues that cannot wait for the regular release cycle. This document outlines classification, procedures, and best practices for managing hotfixes across different development workflows.
## Severity Classification
### P0 - Critical (Production Down)
**Definition:** Complete system outage, data corruption, or security breach affecting all users.
**Examples:**
- Server crashes preventing any user access
- Database corruption causing data loss
- Security vulnerability being actively exploited
- Payment system completely non-functional
- Authentication system failure preventing all logins
**Response Requirements:**
- **Timeline:** Fix deployed within 2 hours
- **Approval:** Engineering Lead + On-call Manager (verbal approval acceptable)
- **Process:** Emergency deployment bypassing normal gates
- **Communication:** Immediate notification to all stakeholders
- **Documentation:** Post-incident review required within 24 hours
**Escalation:**
- Page on-call engineer immediately
- Escalate to Engineering Lead within 15 minutes
- Notify CEO/CTO if resolution exceeds 4 hours
### P1 - High (Major Feature Broken)
**Definition:** Critical functionality broken affecting significant portion of users.
**Examples:**
- Core user workflow completely broken
- Payment processing failures affecting >50% of transactions
- Search functionality returning no results
- Mobile app crashes on startup
- API returning 500 errors for main endpoints
**Response Requirements:**
- **Timeline:** Fix deployed within 24 hours
- **Approval:** Engineering Lead + Product Manager
- **Process:** Expedited review and testing
- **Communication:** Stakeholder notification within 1 hour
- **Documentation:** Root cause analysis within 48 hours
**Escalation:**
- Notify on-call engineer within 30 minutes
- Escalate to Engineering Lead within 2 hours
- Daily updates to Product/Business stakeholders
### P2 - Medium (Minor Feature Issues)
**Definition:** Non-critical functionality issues with limited user impact.
**Examples:**
- Cosmetic UI issues affecting user experience
- Non-essential features not working properly
- Performance degradation not affecting core workflows
- Minor API inconsistencies
- Reporting/analytics data inaccuracies
**Response Requirements:**
- **Timeline:** Include in next regular release
- **Approval:** Standard pull request review process
- **Process:** Normal development and testing cycle
- **Communication:** Include in regular release notes
- **Documentation:** Standard issue tracking
**Escalation:**
- Create ticket in normal backlog
- No special escalation required
- Include in release planning discussions
## Hotfix Workflows by Development Model
### Git Flow Hotfix Process
#### Branch Structure
```
main (v1.2.3) ← hotfix/security-patch → main (v1.2.4)
→ develop
```
#### Step-by-Step Process
1. **Create Hotfix Branch**
```bash
git checkout main
git pull origin main
git checkout -b hotfix/security-patch
```
2. **Implement Fix**
- Make minimal changes addressing only the specific issue
- Include tests to prevent regression
- Update version number (patch increment)
```bash
# Fix the issue
git add .
git commit -m "fix: resolve SQL injection vulnerability"
# Version bump
echo "1.2.4" > VERSION
git add VERSION
git commit -m "chore: bump version to 1.2.4"
```
3. **Test Fix**
- Run automated test suite
- Manual testing of affected functionality
- Security review if applicable
```bash
# Run tests
npm test
python -m pytest
# Security scan
npm audit
bandit -r src/
```
4. **Deploy to Staging**
```bash
# Deploy hotfix branch to staging
git push origin hotfix/security-patch
# Trigger staging deployment via CI/CD
```
5. **Merge to Production**
```bash
# Merge to main
git checkout main
git merge --no-ff hotfix/security-patch
git tag -a v1.2.4 -m "Hotfix: Security vulnerability patch"
git push origin main --tags
# Merge back to develop
git checkout develop
git merge --no-ff hotfix/security-patch
git push origin develop
# Clean up
git branch -d hotfix/security-patch
git push origin --delete hotfix/security-patch
```
### GitHub Flow Hotfix Process
#### Branch Structure
```
main ← hotfix/critical-fix → main (immediate deploy)
```
#### Step-by-Step Process
1. **Create Fix Branch**
```bash
git checkout main
git pull origin main
git checkout -b hotfix/payment-gateway-fix
```
2. **Implement and Test**
```bash
# Make the fix
git add .
git commit -m "fix(payment): resolve gateway timeout issue"
git push origin hotfix/payment-gateway-fix
```
3. **Create Emergency PR**
```bash
# Use GitHub CLI or web interface
gh pr create --title "HOTFIX: Payment gateway timeout" \
--body "Critical fix for payment processing failures" \
--reviewer engineering-team \
--label hotfix
```
4. **Deploy Branch for Testing**
```bash
# Deploy branch to staging for validation
./deploy.sh hotfix/payment-gateway-fix staging
# Quick smoke tests
```
5. **Emergency Merge and Deploy**
```bash
# After approval, merge and deploy
gh pr merge --squash
# Automatic deployment to production via CI/CD
```
### Trunk-based Hotfix Process
#### Direct Commit Approach
```bash
# For small fixes, commit directly to main
git checkout main
git pull origin main
# Make fix
git add .
git commit -m "fix: resolve memory leak in user session handling"
git push origin main
# Automatic deployment triggers
```
#### Feature Flag Rollback
```bash
# For feature-related issues, disable via feature flag
curl -X POST api/feature-flags/new-search/disable
# Verify issue resolved
# Plan proper fix for next deployment
```
## Emergency Response Procedures
### Incident Declaration Process
1. **Detection and Assessment** (0-5 minutes)
- Monitor alerts or user reports identify issue
- Assess severity using classification matrix
- Determine if hotfix is required
2. **Team Assembly** (5-10 minutes)
- Page appropriate on-call engineer
- Assemble incident response team
- Establish communication channel (Slack, Teams)
3. **Initial Response** (10-30 minutes)
- Create incident ticket/document
- Begin investigating root cause
- Implement immediate mitigations if possible
4. **Hotfix Development** (30 minutes - 2 hours)
- Create hotfix branch
- Implement minimal fix
- Test fix in isolation
5. **Deployment** (15-30 minutes)
- Deploy to staging for validation
- Deploy to production
- Monitor for successful resolution
6. **Verification** (15-30 minutes)
- Confirm issue is resolved
- Monitor system stability
- Update stakeholders
### Communication Templates
#### P0 Initial Alert
```
🚨 CRITICAL INCIDENT - Production Down
Status: Investigating
Impact: Complete service outage
Affected Users: All users
Started: 2024-01-15 14:30 UTC
Incident Commander: @john.doe
Current Actions:
- Investigating root cause
- Preparing emergency fix
- Will update every 15 minutes
Status Page: https://status.ourapp.com
Incident Channel: #incident-2024-001
```
#### P0 Resolution Notice
```
✅ RESOLVED - Production Restored
Status: Resolved
Resolution Time: 1h 23m
Root Cause: Database connection pool exhaustion
Fix: Increased connection limits and restarted services
Timeline:
14:30 UTC - Issue detected
14:45 UTC - Root cause identified
15:20 UTC - Fix deployed
15:35 UTC - Full functionality restored
Post-incident review scheduled for tomorrow 10:00 AM.
Thank you for your patience.
```
#### P1 Status Update
```
⚠️ Issue Update - Payment Processing
Status: Fix deployed, monitoring
Impact: Payment failures reduced from 45% to <2%
ETA: Complete resolution within 2 hours
Actions taken:
- Deployed hotfix to address timeout issues
- Increased monitoring on payment gateway
- Contacting affected customers
Next update in 30 minutes or when resolved.
```
### Rollback Procedures
#### When to Rollback
- Fix doesn't resolve the issue
- Fix introduces new problems
- System stability is compromised
- Data corruption is detected
#### Rollback Process
1. **Immediate Assessment** (2-5 minutes)
```bash
# Check system health
curl -f https://api.ourapp.com/health
# Review error logs
kubectl logs deployment/app --tail=100
# Check key metrics
```
2. **Rollback Execution** (5-15 minutes)
```bash
# Git-based rollback
git checkout main
git revert HEAD
git push origin main
# Or container-based rollback
kubectl rollout undo deployment/app
# Or load balancer switch
aws elbv2 modify-target-group --target-group-arn arn:aws:elasticloadbalancing:us-east-1:123456789012:targetgroup/previous-version
```
3. **Verification** (5-10 minutes)
```bash
# Confirm rollback successful
# Check system health endpoints
# Verify core functionality working
# Monitor error rates and performance
```
4. **Communication**
```
🔄 ROLLBACK COMPLETE
The hotfix has been rolled back due to [reason].
System is now stable on previous version.
We are investigating the issue and will provide updates.
```
## Testing Strategies for Hotfixes
### Pre-deployment Testing
#### Automated Testing
```bash
# Run full test suite
npm test
pytest tests/
go test ./...
# Security scanning
npm audit --audit-level high
bandit -r src/
gosec ./...
# Integration tests
./run_integration_tests.sh
# Load testing (if performance-related)
artillery quick --count 100 --num 10 https://staging.ourapp.com
```
#### Manual Testing Checklist
- [ ] Core user workflow functions correctly
- [ ] Authentication and authorization working
- [ ] Payment processing (if applicable)
- [ ] Data integrity maintained
- [ ] No new error logs or exceptions
- [ ] Performance within acceptable range
- [ ] Mobile app functionality (if applicable)
- [ ] Third-party integrations working
#### Staging Validation
```bash
# Deploy to staging
./deploy.sh hotfix/critical-fix staging
# Run smoke tests
curl -f https://staging.ourapp.com/api/health
./smoke_tests.sh
# Manual verification of specific issue
# Document test results
```
### Post-deployment Monitoring
#### Immediate Monitoring (First 30 minutes)
- Error rate and count
- Response time and latency
- CPU and memory usage
- Database connection counts
- Key business metrics
#### Extended Monitoring (First 24 hours)
- User activity patterns
- Feature usage statistics
- Customer support tickets
- Performance trends
- Security log analysis
#### Monitoring Scripts
```bash
#!/bin/bash
# monitor_hotfix.sh - Post-deployment monitoring
echo "=== Hotfix Deployment Monitoring ==="
echo "Deployment time: $(date)"
echo
# Check application health
echo "--- Application Health ---"
curl -s https://api.ourapp.com/health | jq '.'
# Check error rates
echo "--- Error Rates (last 30min) ---"
curl -s "https://api.datadog.com/api/v1/query?query=sum:application.errors{*}" \
-H "DD-API-KEY: $DATADOG_API_KEY" | jq '.series[0].pointlist[-1][1]'
# Check response times
echo "--- Response Times ---"
curl -s "https://api.datadog.com/api/v1/query?query=avg:application.response_time{*}" \
-H "DD-API-KEY: $DATADOG_API_KEY" | jq '.series[0].pointlist[-1][1]'
# Check database connections
echo "--- Database Status ---"
psql -h db.ourapp.com -U readonly -c "SELECT count(*) as active_connections FROM pg_stat_activity;"
echo "=== Monitoring Complete ==="
```
## Documentation and Learning
### Incident Documentation Template
```markdown
# Incident Report: [Brief Description]
## Summary
- **Incident ID:** INC-2024-001
- **Severity:** P0/P1/P2
- **Start Time:** 2024-01-15 14:30 UTC
- **End Time:** 2024-01-15 15:45 UTC
- **Duration:** 1h 15m
- **Impact:** [Description of user/business impact]
## Root Cause
[Detailed explanation of what went wrong and why]
## Timeline
| Time | Event |
|------|-------|
| 14:30 | Issue detected via monitoring alert |
| 14:35 | Incident team assembled |
| 14:45 | Root cause identified |
| 15:00 | Fix developed and tested |
| 15:20 | Fix deployed to production |
| 15:45 | Issue confirmed resolved |
## Resolution
[What was done to fix the issue]
## Lessons Learned
### What went well
- Quick detection through monitoring
- Effective team coordination
- Minimal user impact
### What could be improved
- Earlier detection possible with better alerting
- Testing could have caught this issue
- Communication could be more proactive
## Action Items
- [ ] Improve monitoring for [specific area]
- [ ] Add automated test for [specific scenario]
- [ ] Update documentation for [specific process]
- [ ] Training on [specific topic] for team
## Prevention Measures
[How we'll prevent this from happening again]
```
### Post-Incident Review Process
1. **Schedule Review** (within 24-48 hours)
- Involve all key participants
- Book 60-90 minute session
- Prepare incident timeline
2. **Blameless Analysis**
- Focus on systems and processes, not individuals
- Understand contributing factors
- Identify improvement opportunities
3. **Action Plan**
- Concrete, assignable tasks
- Realistic timelines
- Clear success criteria
4. **Follow-up**
- Track action item completion
- Share learnings with broader team
- Update procedures based on insights
### Knowledge Sharing
#### Runbook Updates
After each hotfix, update relevant runbooks:
- Add new troubleshooting steps
- Update contact information
- Refine escalation procedures
- Document new tools or processes
#### Team Training
- Share incident learnings in team meetings
- Conduct tabletop exercises for common scenarios
- Update onboarding materials with hotfix procedures
- Create decision trees for severity classification
#### Automation Improvements
- Add alerts for new failure modes
- Automate manual steps where possible
- Improve deployment and rollback processes
- Enhance monitoring and observability
## Common Pitfalls and Best Practices
### Common Pitfalls
❌ **Over-engineering the fix**
- Making broad changes instead of minimal targeted fix
- Adding features while fixing bugs
- Refactoring unrelated code
❌ **Insufficient testing**
- Skipping automated tests due to time pressure
- Not testing the exact scenario that caused the issue
- Deploying without staging validation
❌ **Poor communication**
- Not notifying stakeholders promptly
- Unclear or infrequent status updates
- Forgetting to announce resolution
❌ **Inadequate monitoring**
- Not watching system health after deployment
- Missing secondary effects of the fix
- Failing to verify the issue is actually resolved
### Best Practices
✅ **Keep fixes minimal and focused**
- Address only the specific issue
- Avoid scope creep or improvements
- Save refactoring for regular releases
✅ **Maintain clear communication**
- Set up dedicated incident channel
- Provide regular status updates
- Use clear, non-technical language for business stakeholders
✅ **Test thoroughly but efficiently**
- Focus testing on affected functionality
- Use automated tests where possible
- Validate in staging before production
✅ **Document everything**
- Maintain timeline of events
- Record decisions and rationale
- Share lessons learned with team
✅ **Plan for rollback**
- Always have a rollback plan ready
- Test rollback procedure in advance
- Monitor closely after deployment
By following these procedures and continuously improving based on experience, teams can handle production emergencies effectively while minimizing impact and learning from each incident.
FILE:references/release-workflow-comparison.md
# Release Workflow Comparison
## Overview
This document compares the three most popular branching and release workflows: Git Flow, GitHub Flow, and Trunk-based Development. Each approach has distinct advantages and trade-offs depending on your team size, deployment frequency, and risk tolerance.
## Git Flow
### Structure
```
main (production)
↑
release/1.2.0 ← develop (integration) ← feature/user-auth
↑ ← feature/payment-api
hotfix/critical-fix
```
### Branch Types
- **main**: Production-ready code, tagged releases
- **develop**: Integration branch for next release
- **feature/***: Individual features, merged to develop
- **release/X.Y.Z**: Release preparation, branched from develop
- **hotfix/***: Critical fixes, branched from main
### Typical Flow
1. Create feature branch from develop: `git checkout -b feature/login develop`
2. Work on feature, commit changes
3. Merge feature to develop when complete
4. When ready for release, create release branch: `git checkout -b release/1.2.0 develop`
5. Finalize release (version bump, changelog, bug fixes)
6. Merge release branch to both main and develop
7. Tag release: `git tag v1.2.0`
8. Deploy from main branch
### Advantages
- **Clear separation** between production and development code
- **Stable main branch** always represents production state
- **Parallel development** of features without interference
- **Structured release process** with dedicated release branches
- **Hotfix support** without disrupting development work
- **Good for scheduled releases** and traditional release cycles
### Disadvantages
- **Complex workflow** with many branch types
- **Merge overhead** from multiple integration points
- **Delayed feedback** from long-lived feature branches
- **Integration conflicts** when merging large features
- **Slower deployment** due to process overhead
- **Not ideal for continuous deployment**
### Best For
- Large teams (10+ developers)
- Products with scheduled release cycles
- Enterprise software with formal testing phases
- Projects requiring stable release branches
- Teams comfortable with complex Git workflows
### Example Commands
```bash
# Start new feature
git checkout develop
git checkout -b feature/user-authentication
# Finish feature
git checkout develop
git merge --no-ff feature/user-authentication
git branch -d feature/user-authentication
# Start release
git checkout develop
git checkout -b release/1.2.0
# Version bump and changelog updates
git commit -am "Bump version to 1.2.0"
# Finish release
git checkout main
git merge --no-ff release/1.2.0
git tag -a v1.2.0 -m "Release version 1.2.0"
git checkout develop
git merge --no-ff release/1.2.0
git branch -d release/1.2.0
# Hotfix
git checkout main
git checkout -b hotfix/security-patch
# Fix the issue
git commit -am "Fix security vulnerability"
git checkout main
git merge --no-ff hotfix/security-patch
git tag -a v1.2.1 -m "Hotfix version 1.2.1"
git checkout develop
git merge --no-ff hotfix/security-patch
```
## GitHub Flow
### Structure
```
main ← feature/user-auth
← feature/payment-api
← hotfix/critical-fix
```
### Branch Types
- **main**: Production-ready code, deployed automatically
- **feature/***: All changes, regardless of size or type
### Typical Flow
1. Create feature branch from main: `git checkout -b feature/login main`
2. Work on feature with regular commits and pushes
3. Open pull request when ready for feedback
4. Deploy feature branch to staging for testing
5. Merge to main when approved and tested
6. Deploy main to production automatically
7. Delete feature branch
### Advantages
- **Simple workflow** with only two branch types
- **Fast deployment** with minimal process overhead
- **Continuous integration** with frequent merges to main
- **Early feedback** through pull request reviews
- **Deploy from branches** allows testing before merge
- **Good for continuous deployment**
### Disadvantages
- **Main can be unstable** if testing is insufficient
- **No release branches** for coordinating multiple features
- **Limited hotfix process** requires careful coordination
- **Requires strong testing** and CI/CD infrastructure
- **Not suitable for scheduled releases**
- **Can be chaotic** with many simultaneous features
### Best For
- Small to medium teams (2-10 developers)
- Web applications with continuous deployment
- Products with rapid iteration cycles
- Teams with strong testing and CI/CD practices
- Projects where main is always deployable
### Example Commands
```bash
# Start new feature
git checkout main
git pull origin main
git checkout -b feature/user-authentication
# Regular work
git add .
git commit -m "feat(auth): add login form validation"
git push origin feature/user-authentication
# Deploy branch for testing
# (Usually done through CI/CD)
./deploy.sh feature/user-authentication staging
# Merge when ready
git checkout main
git merge feature/user-authentication
git push origin main
git branch -d feature/user-authentication
# Automatic deployment to production
# (Triggered by push to main)
```
## Trunk-based Development
### Structure
```
main ← short-feature-branch (1-3 days max)
← another-short-branch
← direct-commits
```
### Branch Types
- **main**: The single source of truth, always deployable
- **Short-lived branches**: Optional, for changes taking >1 day
### Typical Flow
1. Commit directly to main for small changes
2. Create short-lived branch for larger changes (max 2-3 days)
3. Merge to main frequently (multiple times per day)
4. Use feature flags to hide incomplete features
5. Deploy main to production multiple times per day
6. Release by enabling feature flags, not code deployment
### Advantages
- **Simplest workflow** with minimal branching
- **Fastest integration** with continuous merges
- **Reduced merge conflicts** from short-lived branches
- **Always deployable main** through feature flags
- **Fastest feedback loop** with immediate integration
- **Excellent for CI/CD** and DevOps practices
### Disadvantages
- **Requires discipline** to keep main stable
- **Needs feature flags** for incomplete features
- **Limited code review** for direct commits
- **Can be destabilizing** without proper testing
- **Requires advanced CI/CD** infrastructure
- **Not suitable for teams** uncomfortable with frequent changes
### Best For
- Expert teams with strong DevOps culture
- Products requiring very fast iteration
- Microservices architectures
- Teams practicing continuous deployment
- Organizations with mature testing practices
### Example Commands
```bash
# Small change - direct to main
git checkout main
git pull origin main
# Make changes
git add .
git commit -m "fix(ui): resolve button alignment issue"
git push origin main
# Larger change - short branch
git checkout main
git pull origin main
git checkout -b payment-integration
# Work for 1-2 days maximum
git add .
git commit -m "feat(payment): add Stripe integration"
git push origin payment-integration
# Immediate merge
git checkout main
git merge payment-integration
git push origin main
git branch -d payment-integration
# Feature flag usage
if (featureFlags.enabled('stripe_payments', userId)) {
return renderStripePayment();
} else {
return renderLegacyPayment();
}
```
## Feature Comparison Matrix
| Aspect | Git Flow | GitHub Flow | Trunk-based |
|--------|----------|-------------|-------------|
| **Complexity** | High | Medium | Low |
| **Learning Curve** | Steep | Moderate | Gentle |
| **Deployment Frequency** | Weekly/Monthly | Daily | Multiple/day |
| **Branch Lifetime** | Weeks/Months | Days/Weeks | Hours/Days |
| **Main Stability** | Very High | High | High* |
| **Release Coordination** | Excellent | Limited | Feature Flags |
| **Hotfix Support** | Built-in | Manual | Direct |
| **Merge Conflicts** | High | Medium | Low |
| **Team Size** | 10+ | 3-10 | Any |
| **CI/CD Requirements** | Medium | High | Very High |
*With proper feature flags and testing
## Release Strategies by Workflow
### Git Flow Releases
```bash
# Scheduled release every 2 weeks
git checkout develop
git checkout -b release/2.3.0
# Version management
echo "2.3.0" > VERSION
npm version 2.3.0 --no-git-tag-version
python setup.py --version 2.3.0
# Changelog generation
git log --oneline release/2.2.0..HEAD --pretty=format:"%s" > CHANGELOG_DRAFT.md
# Testing and bug fixes in release branch
git commit -am "fix: resolve issue found in release testing"
# Finalize release
git checkout main
git merge --no-ff release/2.3.0
git tag -a v2.3.0 -m "Release 2.3.0"
# Deploy tagged version
docker build -t app:2.3.0 .
kubectl set image deployment/app app=app:2.3.0
```
### GitHub Flow Releases
```bash
# Deploy every merge to main
git checkout main
git merge feature/new-payment-method
# Automatic deployment via CI/CD
# .github/workflows/deploy.yml triggers on push to main
# Tag releases for tracking (optional)
git tag -a v2.3.$(date +%Y%m%d%H%M) -m "Production deployment"
# Rollback if needed
git revert HEAD
git push origin main # Triggers automatic rollback deployment
```
### Trunk-based Releases
```bash
# Continuous deployment with feature flags
git checkout main
git add feature_flags.json
git commit -m "feat: enable new payment method for 10% of users"
git push origin main
# Gradual rollout
curl -X POST api/feature-flags/payment-v2/rollout/25 # 25% of users
# Monitor metrics...
curl -X POST api/feature-flags/payment-v2/rollout/50 # 50% of users
# Monitor metrics...
curl -X POST api/feature-flags/payment-v2/rollout/100 # Full rollout
# Remove flag after successful rollout
git rm old_payment_code.js
git commit -m "cleanup: remove legacy payment code"
```
## Choosing the Right Workflow
### Decision Matrix
**Choose Git Flow if:**
- ✅ Team size > 10 developers
- ✅ Scheduled release cycles (weekly/monthly)
- ✅ Multiple versions supported simultaneously
- ✅ Formal testing and QA processes
- ✅ Complex enterprise software
- ❌ Need rapid deployment
- ❌ Small team or startup
**Choose GitHub Flow if:**
- ✅ Team size 3-10 developers
- ✅ Web applications or APIs
- ✅ Strong CI/CD and testing
- ✅ Daily or continuous deployment
- ✅ Simple release requirements
- ❌ Complex release coordination needed
- ❌ Multiple release branches required
**Choose Trunk-based Development if:**
- ✅ Expert development team
- ✅ Mature DevOps practices
- ✅ Microservices architecture
- ✅ Feature flag infrastructure
- ✅ Multiple deployments per day
- ✅ Strong automated testing
- ❌ Junior developers
- ❌ Complex integration requirements
### Migration Strategies
#### From Git Flow to GitHub Flow
1. **Simplify branching**: Eliminate develop branch, work directly with main
2. **Increase deployment frequency**: Move from scheduled to continuous releases
3. **Strengthen testing**: Improve automated test coverage and CI/CD
4. **Reduce branch lifetime**: Limit feature branches to 1-2 weeks maximum
5. **Train team**: Educate on simpler workflow and increased responsibility
#### From GitHub Flow to Trunk-based
1. **Implement feature flags**: Add feature toggle infrastructure
2. **Improve CI/CD**: Ensure all tests run in <10 minutes
3. **Increase commit frequency**: Encourage multiple commits per day
4. **Reduce branch usage**: Start committing small changes directly to main
5. **Monitor stability**: Ensure main remains deployable at all times
#### From Trunk-based to Git Flow
1. **Add structure**: Introduce develop and release branches
2. **Reduce deployment frequency**: Move to scheduled release cycles
3. **Extend branch lifetime**: Allow longer feature development cycles
4. **Formalize process**: Add approval gates and testing phases
5. **Coordinate releases**: Plan features for specific release versions
## Anti-patterns to Avoid
### Git Flow Anti-patterns
- **Long-lived feature branches** (>2 weeks)
- **Skipping release branches** for small releases
- **Direct commits to main** bypassing develop
- **Forgetting to merge back** to develop after hotfixes
- **Complex merge conflicts** from delayed integration
### GitHub Flow Anti-patterns
- **Unstable main branch** due to insufficient testing
- **Long-lived feature branches** defeating the purpose
- **Skipping pull request reviews** for speed
- **Direct production deployment** without staging validation
- **No rollback plan** when deployments fail
### Trunk-based Anti-patterns
- **Committing broken code** to main branch
- **Feature branches lasting weeks** defeating the philosophy
- **No feature flags** for incomplete features
- **Insufficient automated testing** leading to instability
- **Poor CI/CD pipeline** causing deployment delays
## Conclusion
The choice of release workflow significantly impacts your team's productivity, code quality, and deployment reliability. Consider your team size, technical maturity, deployment requirements, and organizational culture when making this decision.
**Start conservative** (Git Flow) and evolve toward more agile approaches (GitHub Flow, Trunk-based) as your team's skills and infrastructure mature. The key is consistency within your team and alignment with your organization's goals and constraints.
Remember: **The best workflow is the one your team can execute consistently and reliably**.
FILE:release_planner.py
#!/usr/bin/env python3
"""
Release Planner
Takes a list of features/PRs/tickets planned for release and assesses release readiness.
Checks for required approvals, test coverage thresholds, breaking change documentation,
dependency updates, migration steps needed. Generates release checklist, communication
plan, and rollback procedures.
Input: release plan JSON (features, PRs, target date)
Output: release readiness report + checklist + rollback runbook + announcement draft
"""
import argparse
import json
import sys
from datetime import datetime, timedelta
from typing import Dict, List, Optional, Any, Union
from dataclasses import dataclass, asdict
from enum import Enum
class RiskLevel(Enum):
"""Risk levels for release components."""
LOW = "low"
MEDIUM = "medium"
HIGH = "high"
CRITICAL = "critical"
class ComponentStatus(Enum):
"""Status of release components."""
PENDING = "pending"
IN_PROGRESS = "in_progress"
READY = "ready"
BLOCKED = "blocked"
FAILED = "failed"
@dataclass
class Feature:
"""Represents a feature in the release."""
id: str
title: str
description: str
type: str # feature, bugfix, security, breaking_change, etc.
assignee: str
status: ComponentStatus
pull_request_url: Optional[str] = None
issue_url: Optional[str] = None
risk_level: RiskLevel = RiskLevel.MEDIUM
test_coverage_required: float = 80.0
test_coverage_actual: Optional[float] = None
requires_migration: bool = False
migration_complexity: str = "simple" # simple, moderate, complex
breaking_changes: List[str] = None
dependencies: List[str] = None
qa_approved: bool = False
security_approved: bool = False
pm_approved: bool = False
def __post_init__(self):
if self.breaking_changes is None:
self.breaking_changes = []
if self.dependencies is None:
self.dependencies = []
@dataclass
class QualityGate:
"""Quality gate requirements."""
name: str
required: bool
status: ComponentStatus
details: Optional[str] = None
threshold: Optional[float] = None
actual_value: Optional[float] = None
@dataclass
class Stakeholder:
"""Stakeholder for release communication."""
name: str
role: str
contact: str
notification_type: str # email, slack, teams
critical_path: bool = False
@dataclass
class RollbackStep:
"""Individual rollback step."""
order: int
description: str
command: Optional[str] = None
estimated_time: str = "5 minutes"
risk_level: RiskLevel = RiskLevel.LOW
verification: str = ""
class ReleasePlanner:
"""Main release planning and assessment logic."""
def __init__(self):
self.release_name: str = ""
self.version: str = ""
self.target_date: Optional[datetime] = None
self.features: List[Feature] = []
self.quality_gates: List[QualityGate] = []
self.stakeholders: List[Stakeholder] = []
self.rollback_steps: List[RollbackStep] = []
# Configuration
self.min_test_coverage = 80.0
self.required_approvals = ['pm_approved', 'qa_approved']
self.high_risk_approval_requirements = ['pm_approved', 'qa_approved', 'security_approved']
def load_release_plan(self, plan_data: Union[str, Dict]):
"""Load release plan from JSON."""
if isinstance(plan_data, str):
data = json.loads(plan_data)
else:
data = plan_data
self.release_name = data.get('release_name', 'Unnamed Release')
self.version = data.get('version', '1.0.0')
if 'target_date' in data:
self.target_date = datetime.fromisoformat(data['target_date'].replace('Z', '+00:00'))
# Load features
self.features = []
for feature_data in data.get('features', []):
try:
status = ComponentStatus(feature_data.get('status', 'pending'))
risk_level = RiskLevel(feature_data.get('risk_level', 'medium'))
feature = Feature(
id=feature_data['id'],
title=feature_data['title'],
description=feature_data.get('description', ''),
type=feature_data.get('type', 'feature'),
assignee=feature_data.get('assignee', ''),
status=status,
pull_request_url=feature_data.get('pull_request_url'),
issue_url=feature_data.get('issue_url'),
risk_level=risk_level,
test_coverage_required=feature_data.get('test_coverage_required', 80.0),
test_coverage_actual=feature_data.get('test_coverage_actual'),
requires_migration=feature_data.get('requires_migration', False),
migration_complexity=feature_data.get('migration_complexity', 'simple'),
breaking_changes=feature_data.get('breaking_changes', []),
dependencies=feature_data.get('dependencies', []),
qa_approved=feature_data.get('qa_approved', False),
security_approved=feature_data.get('security_approved', False),
pm_approved=feature_data.get('pm_approved', False)
)
self.features.append(feature)
except Exception as e:
print(f"Warning: Error parsing feature {feature_data.get('id', 'unknown')}: {e}",
file=sys.stderr)
# Load quality gates
self.quality_gates = []
for gate_data in data.get('quality_gates', []):
try:
status = ComponentStatus(gate_data.get('status', 'pending'))
gate = QualityGate(
name=gate_data['name'],
required=gate_data.get('required', True),
status=status,
details=gate_data.get('details'),
threshold=gate_data.get('threshold'),
actual_value=gate_data.get('actual_value')
)
self.quality_gates.append(gate)
except Exception as e:
print(f"Warning: Error parsing quality gate {gate_data.get('name', 'unknown')}: {e}",
file=sys.stderr)
# Load stakeholders
self.stakeholders = []
for stakeholder_data in data.get('stakeholders', []):
stakeholder = Stakeholder(
name=stakeholder_data['name'],
role=stakeholder_data['role'],
contact=stakeholder_data['contact'],
notification_type=stakeholder_data.get('notification_type', 'email'),
critical_path=stakeholder_data.get('critical_path', False)
)
self.stakeholders.append(stakeholder)
# Load or generate default quality gates if none provided
if not self.quality_gates:
self._generate_default_quality_gates()
# Load or generate default rollback steps
if 'rollback_steps' in data:
self.rollback_steps = []
for step_data in data['rollback_steps']:
risk_level = RiskLevel(step_data.get('risk_level', 'low'))
step = RollbackStep(
order=step_data['order'],
description=step_data['description'],
command=step_data.get('command'),
estimated_time=step_data.get('estimated_time', '5 minutes'),
risk_level=risk_level,
verification=step_data.get('verification', '')
)
self.rollback_steps.append(step)
else:
self._generate_default_rollback_steps()
def _generate_default_quality_gates(self):
"""Generate default quality gates."""
default_gates = [
{
'name': 'Unit Test Coverage',
'required': True,
'threshold': self.min_test_coverage,
'details': f'Minimum {self.min_test_coverage}% code coverage required'
},
{
'name': 'Integration Tests',
'required': True,
'details': 'All integration tests must pass'
},
{
'name': 'Security Scan',
'required': True,
'details': 'No high or critical security vulnerabilities'
},
{
'name': 'Performance Testing',
'required': True,
'details': 'Performance metrics within acceptable thresholds'
},
{
'name': 'Documentation Review',
'required': True,
'details': 'API docs and user docs updated for new features'
},
{
'name': 'Dependency Audit',
'required': True,
'details': 'All dependencies scanned for vulnerabilities'
}
]
self.quality_gates = []
for gate_data in default_gates:
gate = QualityGate(
name=gate_data['name'],
required=gate_data['required'],
status=ComponentStatus.PENDING,
details=gate_data['details'],
threshold=gate_data.get('threshold')
)
self.quality_gates.append(gate)
def _generate_default_rollback_steps(self):
"""Generate default rollback procedure."""
default_steps = [
{
'order': 1,
'description': 'Alert on-call team and stakeholders',
'estimated_time': '2 minutes',
'verification': 'Confirm team is aware and responding'
},
{
'order': 2,
'description': 'Switch load balancer to previous version',
'command': 'kubectl patch service app --patch \'{"spec": {"selector": {"version": "previous"}}}\'',
'estimated_time': '30 seconds',
'verification': 'Check that traffic is routing to old version'
},
{
'order': 3,
'description': 'Verify application health after rollback',
'estimated_time': '5 minutes',
'verification': 'Check error rates, response times, and health endpoints'
},
{
'order': 4,
'description': 'Roll back database migrations if needed',
'command': 'python manage.py migrate app 0001',
'estimated_time': '10 minutes',
'risk_level': 'high',
'verification': 'Verify data integrity and application functionality'
},
{
'order': 5,
'description': 'Update monitoring dashboards and alerts',
'estimated_time': '5 minutes',
'verification': 'Confirm metrics reflect rollback state'
},
{
'order': 6,
'description': 'Notify stakeholders of successful rollback',
'estimated_time': '5 minutes',
'verification': 'All stakeholders acknowledge rollback completion'
}
]
self.rollback_steps = []
for step_data in default_steps:
risk_level = RiskLevel(step_data.get('risk_level', 'low'))
step = RollbackStep(
order=step_data['order'],
description=step_data['description'],
command=step_data.get('command'),
estimated_time=step_data.get('estimated_time', '5 minutes'),
risk_level=risk_level,
verification=step_data.get('verification', '')
)
self.rollback_steps.append(step)
def assess_release_readiness(self) -> Dict:
"""Assess overall release readiness."""
assessment = {
'overall_status': 'ready',
'readiness_score': 0.0,
'blocking_issues': [],
'warnings': [],
'recommendations': [],
'feature_summary': {},
'quality_gate_summary': {},
'timeline_assessment': {}
}
total_score = 0
max_score = 0
# Assess features
feature_stats = {
'total': len(self.features),
'ready': 0,
'blocked': 0,
'in_progress': 0,
'pending': 0,
'high_risk': 0,
'breaking_changes': 0,
'missing_approvals': 0,
'low_test_coverage': 0
}
for feature in self.features:
max_score += 10 # Each feature worth 10 points
if feature.status == ComponentStatus.READY:
feature_stats['ready'] += 1
total_score += 10
elif feature.status == ComponentStatus.BLOCKED:
feature_stats['blocked'] += 1
assessment['blocking_issues'].append(
f"Feature '{feature.title}' ({feature.id}) is blocked"
)
elif feature.status == ComponentStatus.IN_PROGRESS:
feature_stats['in_progress'] += 1
total_score += 5 # Partial credit
assessment['warnings'].append(
f"Feature '{feature.title}' ({feature.id}) still in progress"
)
else:
feature_stats['pending'] += 1
assessment['warnings'].append(
f"Feature '{feature.title}' ({feature.id}) is pending"
)
# Check risk level
if feature.risk_level in [RiskLevel.HIGH, RiskLevel.CRITICAL]:
feature_stats['high_risk'] += 1
# Check breaking changes
if feature.breaking_changes:
feature_stats['breaking_changes'] += 1
# Check approvals
missing_approvals = self._check_feature_approvals(feature)
if missing_approvals:
feature_stats['missing_approvals'] += 1
assessment['blocking_issues'].append(
f"Feature '{feature.title}' missing approvals: {', '.join(missing_approvals)}"
)
# Check test coverage
if (feature.test_coverage_actual is not None and
feature.test_coverage_actual < feature.test_coverage_required):
feature_stats['low_test_coverage'] += 1
assessment['warnings'].append(
f"Feature '{feature.title}' has low test coverage: "
f"{feature.test_coverage_actual}% < {feature.test_coverage_required}%"
)
assessment['feature_summary'] = feature_stats
# Assess quality gates
gate_stats = {
'total': len(self.quality_gates),
'passed': 0,
'failed': 0,
'pending': 0,
'required_failed': 0
}
for gate in self.quality_gates:
max_score += 5 # Each gate worth 5 points
if gate.status == ComponentStatus.READY:
gate_stats['passed'] += 1
total_score += 5
elif gate.status == ComponentStatus.FAILED:
gate_stats['failed'] += 1
if gate.required:
gate_stats['required_failed'] += 1
assessment['blocking_issues'].append(
f"Required quality gate '{gate.name}' failed"
)
else:
gate_stats['pending'] += 1
if gate.required:
assessment['warnings'].append(
f"Required quality gate '{gate.name}' is pending"
)
assessment['quality_gate_summary'] = gate_stats
# Timeline assessment
if self.target_date:
# Handle timezone-aware datetime comparison
now = datetime.now(self.target_date.tzinfo) if self.target_date.tzinfo else datetime.now()
days_until_release = (self.target_date - now).days
assessment['timeline_assessment'] = {
'target_date': self.target_date.isoformat(),
'days_remaining': days_until_release,
'timeline_status': 'on_track' if days_until_release > 0 else 'overdue'
}
if days_until_release < 0:
assessment['blocking_issues'].append(f"Release is {abs(days_until_release)} days overdue")
elif days_until_release < 3 and feature_stats['blocked'] > 0:
assessment['blocking_issues'].append("Not enough time to resolve blocked features")
# Calculate overall readiness score
if max_score > 0:
assessment['readiness_score'] = (total_score / max_score) * 100
# Determine overall status
if assessment['blocking_issues']:
assessment['overall_status'] = 'blocked'
elif assessment['warnings']:
assessment['overall_status'] = 'at_risk'
else:
assessment['overall_status'] = 'ready'
# Generate recommendations
if feature_stats['missing_approvals'] > 0:
assessment['recommendations'].append("Obtain required approvals for pending features")
if feature_stats['low_test_coverage'] > 0:
assessment['recommendations'].append("Improve test coverage for features below threshold")
if gate_stats['pending'] > 0:
assessment['recommendations'].append("Complete pending quality gate validations")
if feature_stats['high_risk'] > 0:
assessment['recommendations'].append("Review high-risk features for additional validation")
return assessment
def _check_feature_approvals(self, feature: Feature) -> List[str]:
"""Check which approvals are missing for a feature."""
missing = []
# Determine required approvals based on risk level
required = self.required_approvals.copy()
if feature.risk_level in [RiskLevel.HIGH, RiskLevel.CRITICAL]:
required = self.high_risk_approval_requirements.copy()
if 'pm_approved' in required and not feature.pm_approved:
missing.append('PM approval')
if 'qa_approved' in required and not feature.qa_approved:
missing.append('QA approval')
if 'security_approved' in required and not feature.security_approved:
missing.append('Security approval')
return missing
def generate_release_checklist(self) -> List[Dict]:
"""Generate comprehensive release checklist."""
checklist = []
# Pre-release validation
checklist.extend([
{
'category': 'Pre-Release Validation',
'item': 'All features implemented and tested',
'status': 'ready' if all(f.status == ComponentStatus.READY for f in self.features) else 'pending',
'details': f"{len([f for f in self.features if f.status == ComponentStatus.READY])}/{len(self.features)} features ready"
},
{
'category': 'Pre-Release Validation',
'item': 'Breaking changes documented',
'status': 'ready' if self._check_breaking_change_docs() else 'pending',
'details': f"{len([f for f in self.features if f.breaking_changes])} features have breaking changes"
},
{
'category': 'Pre-Release Validation',
'item': 'Migration scripts tested',
'status': 'ready' if self._check_migrations() else 'pending',
'details': f"{len([f for f in self.features if f.requires_migration])} features require migrations"
}
])
# Quality gates
for gate in self.quality_gates:
checklist.append({
'category': 'Quality Gates',
'item': gate.name,
'status': gate.status.value,
'details': gate.details,
'required': gate.required
})
# Approvals
approval_items = [
('Product Manager sign-off', self._check_pm_approvals()),
('QA validation complete', self._check_qa_approvals()),
('Security team clearance', self._check_security_approvals())
]
for item, status in approval_items:
checklist.append({
'category': 'Approvals',
'item': item,
'status': 'ready' if status else 'pending'
})
# Documentation
doc_items = [
'CHANGELOG.md updated',
'API documentation updated',
'User documentation updated',
'Migration guide written',
'Rollback procedure documented'
]
for item in doc_items:
checklist.append({
'category': 'Documentation',
'item': item,
'status': 'pending' # Would need integration with docs system to check
})
# Deployment preparation
deployment_items = [
'Database migrations prepared',
'Environment variables configured',
'Monitoring alerts updated',
'Rollback plan tested',
'Stakeholders notified'
]
for item in deployment_items:
checklist.append({
'category': 'Deployment',
'item': item,
'status': 'pending'
})
return checklist
def _check_breaking_change_docs(self) -> bool:
"""Check if breaking changes are properly documented."""
features_with_breaking_changes = [f for f in self.features if f.breaking_changes]
return all(len(f.breaking_changes) > 0 for f in features_with_breaking_changes)
def _check_migrations(self) -> bool:
"""Check migration readiness."""
features_with_migrations = [f for f in self.features if f.requires_migration]
return all(f.status == ComponentStatus.READY for f in features_with_migrations)
def _check_pm_approvals(self) -> bool:
"""Check PM approvals."""
return all(f.pm_approved for f in self.features if f.risk_level != RiskLevel.LOW)
def _check_qa_approvals(self) -> bool:
"""Check QA approvals."""
return all(f.qa_approved for f in self.features)
def _check_security_approvals(self) -> bool:
"""Check security approvals."""
high_risk_features = [f for f in self.features if f.risk_level in [RiskLevel.HIGH, RiskLevel.CRITICAL]]
return all(f.security_approved for f in high_risk_features)
def generate_communication_plan(self) -> Dict:
"""Generate stakeholder communication plan."""
plan = {
'internal_notifications': [],
'external_notifications': [],
'timeline': [],
'channels': {},
'templates': {}
}
# Group stakeholders by type
internal_stakeholders = [s for s in self.stakeholders if s.role in
['developer', 'qa', 'pm', 'devops', 'security']]
external_stakeholders = [s for s in self.stakeholders if s.role in
['customer', 'partner', 'support']]
# Internal notifications
for stakeholder in internal_stakeholders:
plan['internal_notifications'].append({
'recipient': stakeholder.name,
'role': stakeholder.role,
'method': stakeholder.notification_type,
'content_type': 'technical_details',
'timing': 'T-24h and T-0'
})
# External notifications
for stakeholder in external_stakeholders:
plan['external_notifications'].append({
'recipient': stakeholder.name,
'role': stakeholder.role,
'method': stakeholder.notification_type,
'content_type': 'user_facing_changes',
'timing': 'T-48h and T+1h'
})
# Communication timeline
if self.target_date:
timeline_items = [
(timedelta(days=-2), 'Send pre-release notification to external stakeholders'),
(timedelta(days=-1), 'Send deployment notification to internal teams'),
(timedelta(hours=-2), 'Final go/no-go decision'),
(timedelta(hours=0), 'Begin deployment'),
(timedelta(hours=1), 'Post-deployment status update'),
(timedelta(hours=24), 'Post-release summary')
]
for delta, description in timeline_items:
notification_time = self.target_date + delta
plan['timeline'].append({
'time': notification_time.isoformat(),
'description': description,
'recipients': 'all' if 'all' in description.lower() else 'internal'
})
# Communication channels
channels = {}
for stakeholder in self.stakeholders:
if stakeholder.notification_type not in channels:
channels[stakeholder.notification_type] = []
channels[stakeholder.notification_type].append(stakeholder.contact)
plan['channels'] = channels
# Message templates
plan['templates'] = self._generate_message_templates()
return plan
def _generate_message_templates(self) -> Dict:
"""Generate message templates for different audiences."""
breaking_changes = [f for f in self.features if f.breaking_changes]
new_features = [f for f in self.features if f.type == 'feature']
bug_fixes = [f for f in self.features if f.type == 'bugfix']
templates = {
'internal_pre_release': {
'subject': f'Release {self.version} - Pre-deployment Notification',
'body': f"""Team,
We are preparing to deploy {self.release_name} version {self.version} on {self.target_date.strftime('%Y-%m-%d %H:%M UTC') if self.target_date else 'TBD'}.
Key Changes:
- {len(new_features)} new features
- {len(bug_fixes)} bug fixes
- {len(breaking_changes)} breaking changes
Please review the release notes and prepare for any needed support activities.
Rollback plan: Available in release documentation
On-call: Please be available during deployment window
Best regards,
Release Team"""
},
'external_user_notification': {
'subject': f'Product Update - Version {self.version} Now Available',
'body': f"""Dear Users,
We're excited to announce version {self.version} of {self.release_name} is now available!
What's New:
{chr(10).join(f"- {f.title}" for f in new_features[:5])}
Bug Fixes:
{chr(10).join(f"- {f.title}" for f in bug_fixes[:3])}
{'Important: This release includes breaking changes. Please review the migration guide.' if breaking_changes else ''}
For full release notes and migration instructions, visit our documentation.
Thank you for using our product!
The Development Team"""
},
'rollback_notification': {
'subject': f'URGENT: Release {self.version} Rollback Initiated',
'body': f"""ATTENTION: Release rollback in progress.
Release: {self.version}
Reason: [TO BE FILLED]
Rollback initiated: {datetime.now().strftime('%Y-%m-%d %H:%M UTC')}
Estimated completion: [TO BE FILLED]
Current status: Rolling back to previous stable version
Impact: [TO BE FILLED]
We will provide updates every 15 minutes until rollback is complete.
Incident Commander: [TO BE FILLED]
Status page: [TO BE FILLED]"""
}
}
return templates
def generate_rollback_runbook(self) -> Dict:
"""Generate detailed rollback runbook."""
runbook = {
'overview': {
'purpose': f'Emergency rollback procedure for {self.release_name} v{self.version}',
'triggers': [
'Error rate spike (>2x baseline for >15 minutes)',
'Critical functionality failure',
'Security incident',
'Data corruption detected',
'Performance degradation (>50% latency increase)',
'Manual decision by incident commander'
],
'decision_makers': ['On-call Engineer', 'Engineering Lead', 'Incident Commander'],
'estimated_total_time': self._calculate_rollback_time()
},
'prerequisites': [
'Confirm rollback is necessary (check with incident commander)',
'Notify stakeholders of rollback decision',
'Ensure database backups are available',
'Verify monitoring systems are operational',
'Have communication channels ready'
],
'steps': [],
'verification': {
'health_checks': [
'Application responds to health endpoint',
'Database connectivity confirmed',
'Authentication system functional',
'Core user workflows working',
'Error rates back to baseline',
'Performance metrics within normal range'
],
'rollback_confirmation': [
'Previous version fully deployed',
'Database in consistent state',
'All services communicating properly',
'Monitoring shows stable metrics',
'Sample user workflows tested'
]
},
'post_rollback': [
'Update status page with resolution',
'Notify all stakeholders of successful rollback',
'Schedule post-incident review',
'Document issues encountered during rollback',
'Plan investigation of root cause',
'Determine timeline for next release attempt'
],
'emergency_contacts': []
}
# Convert rollback steps to detailed format
for step in sorted(self.rollback_steps, key=lambda x: x.order):
step_data = {
'order': step.order,
'title': step.description,
'estimated_time': step.estimated_time,
'risk_level': step.risk_level.value,
'instructions': step.description,
'command': step.command,
'verification': step.verification,
'rollback_possible': step.risk_level != RiskLevel.CRITICAL
}
runbook['steps'].append(step_data)
# Add emergency contacts
critical_stakeholders = [s for s in self.stakeholders if s.critical_path]
for stakeholder in critical_stakeholders:
runbook['emergency_contacts'].append({
'name': stakeholder.name,
'role': stakeholder.role,
'contact': stakeholder.contact,
'method': stakeholder.notification_type
})
return runbook
def _calculate_rollback_time(self) -> str:
"""Calculate estimated total rollback time."""
total_minutes = 0
for step in self.rollback_steps:
# Parse time estimates like "5 minutes", "30 seconds", "1 hour"
time_str = step.estimated_time.lower()
if 'minute' in time_str:
minutes = int(re.search(r'(\d+)', time_str).group(1))
total_minutes += minutes
elif 'hour' in time_str:
hours = int(re.search(r'(\d+)', time_str).group(1))
total_minutes += hours * 60
elif 'second' in time_str:
# Round up seconds to minutes
total_minutes += 1
if total_minutes < 60:
return f"{total_minutes} minutes"
else:
hours = total_minutes // 60
minutes = total_minutes % 60
return f"{hours}h {minutes}m"
def main():
"""Main CLI entry point."""
parser = argparse.ArgumentParser(description="Assess release readiness and generate release plans")
parser.add_argument('--input', '-i', required=True,
help='Release plan JSON file')
parser.add_argument('--output-format', '-f',
choices=['json', 'markdown', 'text'],
default='text', help='Output format')
parser.add_argument('--output', '-o', type=str,
help='Output file (default: stdout)')
parser.add_argument('--include-checklist', action='store_true',
help='Include release checklist in output')
parser.add_argument('--include-communication', action='store_true',
help='Include communication plan')
parser.add_argument('--include-rollback', action='store_true',
help='Include rollback runbook')
parser.add_argument('--min-coverage', type=float, default=80.0,
help='Minimum test coverage threshold')
args = parser.parse_args()
# Load release plan
try:
with open(args.input, 'r', encoding='utf-8') as f:
plan_data = f.read()
except Exception as e:
print(f"Error reading input file: {e}", file=sys.stderr)
sys.exit(1)
# Initialize planner
planner = ReleasePlanner()
planner.min_test_coverage = args.min_coverage
try:
planner.load_release_plan(plan_data)
except Exception as e:
print(f"Error loading release plan: {e}", file=sys.stderr)
sys.exit(1)
# Generate assessment
assessment = planner.assess_release_readiness()
# Generate optional components
checklist = planner.generate_release_checklist() if args.include_checklist else None
communication = planner.generate_communication_plan() if args.include_communication else None
rollback = planner.generate_rollback_runbook() if args.include_rollback else None
# Generate output
if args.output_format == 'json':
output_data = {
'assessment': assessment,
'checklist': checklist,
'communication_plan': communication,
'rollback_runbook': rollback
}
output_text = json.dumps(output_data, indent=2, default=str)
elif args.output_format == 'markdown':
output_lines = [
f"# Release Readiness Report - {planner.release_name} v{planner.version}",
"",
f"**Overall Status:** {assessment['overall_status'].upper()}",
f"**Readiness Score:** {assessment['readiness_score']:.1f}%",
""
]
if assessment['blocking_issues']:
output_lines.extend([
"## 🚫 Blocking Issues",
""
])
for issue in assessment['blocking_issues']:
output_lines.append(f"- {issue}")
output_lines.append("")
if assessment['warnings']:
output_lines.extend([
"## ⚠️ Warnings",
""
])
for warning in assessment['warnings']:
output_lines.append(f"- {warning}")
output_lines.append("")
# Feature summary
fs = assessment['feature_summary']
output_lines.extend([
"## Features Summary",
"",
f"- **Total:** {fs['total']}",
f"- **Ready:** {fs['ready']}",
f"- **In Progress:** {fs['in_progress']}",
f"- **Blocked:** {fs['blocked']}",
f"- **Breaking Changes:** {fs['breaking_changes']}",
""
])
if checklist:
output_lines.extend([
"## Release Checklist",
""
])
current_category = ""
for item in checklist:
if item['category'] != current_category:
current_category = item['category']
output_lines.append(f"### {current_category}")
output_lines.append("")
status_icon = "✅" if item['status'] == 'ready' else "❌" if item['status'] == 'failed' else "⏳"
output_lines.append(f"- {status_icon} {item['item']}")
output_lines.append("")
output_text = '\n'.join(output_lines)
else: # text format
output_lines = [
f"Release Readiness Report",
f"========================",
f"Release: {planner.release_name} v{planner.version}",
f"Status: {assessment['overall_status'].upper()}",
f"Readiness Score: {assessment['readiness_score']:.1f}%",
""
]
if assessment['blocking_issues']:
output_lines.extend(["BLOCKING ISSUES:", ""])
for issue in assessment['blocking_issues']:
output_lines.append(f" ❌ {issue}")
output_lines.append("")
if assessment['warnings']:
output_lines.extend(["WARNINGS:", ""])
for warning in assessment['warnings']:
output_lines.append(f" ⚠️ {warning}")
output_lines.append("")
if assessment['recommendations']:
output_lines.extend(["RECOMMENDATIONS:", ""])
for rec in assessment['recommendations']:
output_lines.append(f" 💡 {rec}")
output_lines.append("")
# Summary stats
fs = assessment['feature_summary']
gs = assessment['quality_gate_summary']
output_lines.extend([
f"FEATURE SUMMARY:",
f" Total: {fs['total']} | Ready: {fs['ready']} | Blocked: {fs['blocked']}",
f" Breaking Changes: {fs['breaking_changes']} | Missing Approvals: {fs['missing_approvals']}",
"",
f"QUALITY GATES:",
f" Total: {gs['total']} | Passed: {gs['passed']} | Failed: {gs['failed']}",
""
])
output_text = '\n'.join(output_lines)
# Write output
if args.output:
with open(args.output, 'w', encoding='utf-8') as f:
f.write(output_text)
else:
print(output_text)
if __name__ == '__main__':
main()
FILE:version_bumper.py
#!/usr/bin/env python3
"""
Version Bumper
Analyzes commits since last tag to determine the correct version bump (major/minor/patch)
based on conventional commits. Handles pre-release versions (alpha, beta, rc) and generates
version bump commands for various package files.
Input: current version + commit list JSON or git log
Output: recommended new version + bump commands + updated file snippets
"""
import argparse
import json
import re
import sys
from typing import Dict, List, Optional, Tuple, Union
from enum import Enum
from dataclasses import dataclass
class BumpType(Enum):
"""Version bump types."""
NONE = "none"
PATCH = "patch"
MINOR = "minor"
MAJOR = "major"
class PreReleaseType(Enum):
"""Pre-release types."""
ALPHA = "alpha"
BETA = "beta"
RC = "rc"
@dataclass
class Version:
"""Semantic version representation."""
major: int
minor: int
patch: int
prerelease_type: Optional[PreReleaseType] = None
prerelease_number: Optional[int] = None
@classmethod
def parse(cls, version_str: str) -> 'Version':
"""Parse version string into Version object."""
# Remove 'v' prefix if present
clean_version = version_str.lstrip('v')
# Pattern for semantic versioning with optional pre-release
pattern = r'^(\d+)\.(\d+)\.(\d+)(?:-(\w+)\.?(\d+)?)?$'
match = re.match(pattern, clean_version)
if not match:
raise ValueError(f"Invalid version format: {version_str}")
major, minor, patch = int(match.group(1)), int(match.group(2)), int(match.group(3))
prerelease_type = None
prerelease_number = None
if match.group(4): # Pre-release identifier
prerelease_str = match.group(4).lower()
try:
prerelease_type = PreReleaseType(prerelease_str)
except ValueError:
# Handle variations like 'alpha1' -> 'alpha'
if prerelease_str.startswith('alpha'):
prerelease_type = PreReleaseType.ALPHA
elif prerelease_str.startswith('beta'):
prerelease_type = PreReleaseType.BETA
elif prerelease_str.startswith('rc'):
prerelease_type = PreReleaseType.RC
else:
raise ValueError(f"Unknown pre-release type: {prerelease_str}")
if match.group(5):
prerelease_number = int(match.group(5))
else:
# Extract number from combined string like 'alpha1'
number_match = re.search(r'(\d+)$', prerelease_str)
if number_match:
prerelease_number = int(number_match.group(1))
else:
prerelease_number = 1 # Default to 1
return cls(major, minor, patch, prerelease_type, prerelease_number)
def to_string(self, include_v_prefix: bool = False) -> str:
"""Convert version to string representation."""
base = f"{self.major}.{self.minor}.{self.patch}"
if self.prerelease_type:
if self.prerelease_number is not None:
base += f"-{self.prerelease_type.value}.{self.prerelease_number}"
else:
base += f"-{self.prerelease_type.value}"
return f"v{base}" if include_v_prefix else base
def bump(self, bump_type: BumpType, prerelease_type: Optional[PreReleaseType] = None) -> 'Version':
"""Create new version with specified bump."""
if bump_type == BumpType.NONE:
return Version(self.major, self.minor, self.patch, self.prerelease_type, self.prerelease_number)
new_major = self.major
new_minor = self.minor
new_patch = self.patch
new_prerelease_type = None
new_prerelease_number = None
# Handle pre-release versions
if prerelease_type:
if bump_type == BumpType.MAJOR:
new_major += 1
new_minor = 0
new_patch = 0
elif bump_type == BumpType.MINOR:
new_minor += 1
new_patch = 0
elif bump_type == BumpType.PATCH:
new_patch += 1
new_prerelease_type = prerelease_type
new_prerelease_number = 1
# Handle existing pre-release -> next pre-release
elif self.prerelease_type:
# If we're already in pre-release, increment or promote
if prerelease_type is None:
# Promote to stable release
# Don't change version numbers, just remove pre-release
pass
else:
# Move to next pre-release type or increment
if prerelease_type == self.prerelease_type:
# Same pre-release type, increment number
new_prerelease_type = self.prerelease_type
new_prerelease_number = (self.prerelease_number or 0) + 1
else:
# Different pre-release type
new_prerelease_type = prerelease_type
new_prerelease_number = 1
# Handle stable version bumps
else:
if bump_type == BumpType.MAJOR:
new_major += 1
new_minor = 0
new_patch = 0
elif bump_type == BumpType.MINOR:
new_minor += 1
new_patch = 0
elif bump_type == BumpType.PATCH:
new_patch += 1
return Version(new_major, new_minor, new_patch, new_prerelease_type, new_prerelease_number)
@dataclass
class ConventionalCommit:
"""Represents a parsed conventional commit for version analysis."""
type: str
scope: str
description: str
is_breaking: bool
breaking_description: str
hash: str = ""
author: str = ""
date: str = ""
@classmethod
def parse_message(cls, message: str, commit_hash: str = "",
author: str = "", date: str = "") -> 'ConventionalCommit':
"""Parse conventional commit message."""
lines = message.split('\n')
header = lines[0] if lines else ""
# Parse header: type(scope): description
header_pattern = r'^(\w+)(\([^)]+\))?(!)?:\s*(.+)$'
match = re.match(header_pattern, header)
commit_type = "chore"
scope = ""
description = header
is_breaking = False
breaking_description = ""
if match:
commit_type = match.group(1).lower()
scope_match = match.group(2)
scope = scope_match[1:-1] if scope_match else ""
is_breaking = bool(match.group(3)) # ! indicates breaking change
description = match.group(4).strip()
# Check for breaking change in body/footers
if len(lines) > 1:
body_text = '\n'.join(lines[1:])
if 'BREAKING CHANGE:' in body_text:
is_breaking = True
breaking_match = re.search(r'BREAKING CHANGE:\s*(.+)', body_text)
if breaking_match:
breaking_description = breaking_match.group(1).strip()
return cls(commit_type, scope, description, is_breaking, breaking_description,
commit_hash, author, date)
class VersionBumper:
"""Main version bumping logic."""
def __init__(self):
self.current_version: Optional[Version] = None
self.commits: List[ConventionalCommit] = []
self.custom_rules: Dict[str, BumpType] = {}
self.ignore_types: List[str] = ['test', 'ci', 'build', 'chore', 'docs', 'style']
def set_current_version(self, version_str: str):
"""Set the current version."""
self.current_version = Version.parse(version_str)
def add_custom_rule(self, commit_type: str, bump_type: BumpType):
"""Add custom rule for commit type to bump type mapping."""
self.custom_rules[commit_type] = bump_type
def parse_commits_from_json(self, json_data: Union[str, List[Dict]]):
"""Parse commits from JSON format."""
if isinstance(json_data, str):
data = json.loads(json_data)
else:
data = json_data
self.commits = []
for commit_data in data:
commit = ConventionalCommit.parse_message(
message=commit_data.get('message', ''),
commit_hash=commit_data.get('hash', ''),
author=commit_data.get('author', ''),
date=commit_data.get('date', '')
)
self.commits.append(commit)
def parse_commits_from_git_log(self, git_log_text: str):
"""Parse commits from git log output."""
lines = git_log_text.strip().split('\n')
if not lines or not lines[0]:
return
# Simple oneline format (hash message)
oneline_pattern = r'^([a-f0-9]{7,40})\s+(.+)$'
self.commits = []
for line in lines:
line = line.strip()
if not line:
continue
match = re.match(oneline_pattern, line)
if match:
commit_hash = match.group(1)
message = match.group(2)
commit = ConventionalCommit.parse_message(message, commit_hash)
self.commits.append(commit)
def determine_bump_type(self) -> BumpType:
"""Determine version bump type based on commits."""
if not self.commits:
return BumpType.NONE
has_breaking = False
has_feature = False
has_fix = False
for commit in self.commits:
# Check for breaking changes
if commit.is_breaking:
has_breaking = True
continue
# Apply custom rules first
if commit.type in self.custom_rules:
bump_type = self.custom_rules[commit.type]
if bump_type == BumpType.MAJOR:
has_breaking = True
elif bump_type == BumpType.MINOR:
has_feature = True
elif bump_type == BumpType.PATCH:
has_fix = True
continue
# Standard rules
if commit.type in ['feat', 'add']:
has_feature = True
elif commit.type in ['fix', 'security', 'perf', 'bugfix']:
has_fix = True
# Ignore types in ignore_types list
# Determine bump type by priority
if has_breaking:
return BumpType.MAJOR
elif has_feature:
return BumpType.MINOR
elif has_fix:
return BumpType.PATCH
else:
return BumpType.NONE
def recommend_version(self, prerelease_type: Optional[PreReleaseType] = None) -> Version:
"""Recommend new version based on commits."""
if not self.current_version:
raise ValueError("Current version not set")
bump_type = self.determine_bump_type()
return self.current_version.bump(bump_type, prerelease_type)
def generate_bump_commands(self, new_version: Version) -> Dict[str, List[str]]:
"""Generate version bump commands for different package managers."""
version_str = new_version.to_string()
version_with_v = new_version.to_string(include_v_prefix=True)
commands = {
'npm': [
f"npm version {version_str} --no-git-tag-version",
f"# Or manually edit package.json version field to '{version_str}'"
],
'python': [
f"# Update version in setup.py, __init__.py, or pyproject.toml",
f"# setup.py: version='{version_str}'",
f"# pyproject.toml: version = '{version_str}'",
f"# __init__.py: __version__ = '{version_str}'"
],
'rust': [
f"# Update Cargo.toml",
f"# [package]",
f"# version = '{version_str}'"
],
'git': [
f"git tag -a {version_with_v} -m 'Release {version_with_v}'",
f"git push origin {version_with_v}"
],
'docker': [
f"docker build -t myapp:{version_str} .",
f"docker tag myapp:{version_str} myapp:latest"
]
}
return commands
def generate_file_updates(self, new_version: Version) -> Dict[str, str]:
"""Generate file update snippets for common package files."""
version_str = new_version.to_string()
updates = {}
# package.json
updates['package.json'] = json.dumps({
"name": "your-package",
"version": version_str,
"description": "Your package description",
"main": "index.js"
}, indent=2)
# pyproject.toml
updates['pyproject.toml'] = f'''[build-system]
requires = ["setuptools>=61.0", "wheel"]
build-backend = "setuptools.build_meta"
[project]
name = "your-package"
version = "{version_str}"
description = "Your package description"
authors = [
{{name = "Your Name", email = "your.email@example.com"}},
]
'''
# setup.py
updates['setup.py'] = f'''from setuptools import setup, find_packages
setup(
name="your-package",
version="{version_str}",
description="Your package description",
packages=find_packages(),
python_requires=">=3.8",
)
'''
# Cargo.toml
updates['Cargo.toml'] = f'''[package]
name = "your-package"
version = "{version_str}"
edition = "2021"
description = "Your package description"
'''
# __init__.py
updates['__init__.py'] = f'''"""Your package."""
__version__ = "{version_str}"
__author__ = "Your Name"
__email__ = "your.email@example.com"
'''
return updates
def analyze_commits(self) -> Dict:
"""Provide detailed analysis of commits for version bumping."""
if not self.commits:
return {
'total_commits': 0,
'by_type': {},
'breaking_changes': [],
'features': [],
'fixes': [],
'ignored': []
}
analysis = {
'total_commits': len(self.commits),
'by_type': {},
'breaking_changes': [],
'features': [],
'fixes': [],
'ignored': []
}
type_counts = {}
for commit in self.commits:
type_counts[commit.type] = type_counts.get(commit.type, 0) + 1
if commit.is_breaking:
analysis['breaking_changes'].append({
'type': commit.type,
'scope': commit.scope,
'description': commit.description,
'breaking_description': commit.breaking_description,
'hash': commit.hash
})
elif commit.type in ['feat', 'add']:
analysis['features'].append({
'scope': commit.scope,
'description': commit.description,
'hash': commit.hash
})
elif commit.type in ['fix', 'security', 'perf', 'bugfix']:
analysis['fixes'].append({
'scope': commit.scope,
'description': commit.description,
'hash': commit.hash
})
elif commit.type in self.ignore_types:
analysis['ignored'].append({
'type': commit.type,
'scope': commit.scope,
'description': commit.description,
'hash': commit.hash
})
analysis['by_type'] = type_counts
return analysis
def main():
"""Main CLI entry point."""
parser = argparse.ArgumentParser(description="Determine version bump based on conventional commits")
parser.add_argument('--current-version', '-c', required=True,
help='Current version (e.g., 1.2.3, v1.2.3)')
parser.add_argument('--input', '-i', type=str,
help='Input file with commits (default: stdin)')
parser.add_argument('--input-format', choices=['git-log', 'json'],
default='git-log', help='Input format')
parser.add_argument('--prerelease', '-p',
choices=['alpha', 'beta', 'rc'],
help='Generate pre-release version')
parser.add_argument('--output-format', '-f',
choices=['text', 'json', 'commands'],
default='text', help='Output format')
parser.add_argument('--output', '-o', type=str,
help='Output file (default: stdout)')
parser.add_argument('--include-commands', action='store_true',
help='Include bump commands in output')
parser.add_argument('--include-files', action='store_true',
help='Include file update snippets')
parser.add_argument('--custom-rules', type=str,
help='JSON string with custom type->bump rules')
parser.add_argument('--ignore-types', type=str,
help='Comma-separated list of types to ignore')
parser.add_argument('--analysis', '-a', action='store_true',
help='Include detailed commit analysis')
args = parser.parse_args()
# Read input
if args.input:
with open(args.input, 'r', encoding='utf-8') as f:
input_data = f.read()
else:
input_data = sys.stdin.read()
if not input_data.strip():
print("No input data provided", file=sys.stderr)
sys.exit(1)
# Initialize version bumper
bumper = VersionBumper()
try:
bumper.set_current_version(args.current_version)
except ValueError as e:
print(f"Invalid current version: {e}", file=sys.stderr)
sys.exit(1)
# Apply custom rules
if args.custom_rules:
try:
custom_rules = json.loads(args.custom_rules)
for commit_type, bump_type_str in custom_rules.items():
bump_type = BumpType(bump_type_str.lower())
bumper.add_custom_rule(commit_type, bump_type)
except Exception as e:
print(f"Invalid custom rules: {e}", file=sys.stderr)
sys.exit(1)
# Set ignore types
if args.ignore_types:
bumper.ignore_types = [t.strip() for t in args.ignore_types.split(',')]
# Parse commits
try:
if args.input_format == 'json':
bumper.parse_commits_from_json(input_data)
else:
bumper.parse_commits_from_git_log(input_data)
except Exception as e:
print(f"Error parsing commits: {e}", file=sys.stderr)
sys.exit(1)
# Determine pre-release type
prerelease_type = None
if args.prerelease:
prerelease_type = PreReleaseType(args.prerelease)
# Generate recommendation
try:
recommended_version = bumper.recommend_version(prerelease_type)
bump_type = bumper.determine_bump_type()
except Exception as e:
print(f"Error determining version: {e}", file=sys.stderr)
sys.exit(1)
# Generate output
output_data = {}
if args.output_format == 'json':
output_data = {
'current_version': args.current_version,
'recommended_version': recommended_version.to_string(),
'recommended_version_with_v': recommended_version.to_string(include_v_prefix=True),
'bump_type': bump_type.value,
'prerelease': args.prerelease
}
if args.analysis:
output_data['analysis'] = bumper.analyze_commits()
if args.include_commands:
output_data['commands'] = bumper.generate_bump_commands(recommended_version)
if args.include_files:
output_data['file_updates'] = bumper.generate_file_updates(recommended_version)
output_text = json.dumps(output_data, indent=2)
elif args.output_format == 'commands':
commands = bumper.generate_bump_commands(recommended_version)
output_lines = [
f"# Version Bump Commands",
f"# Current: {args.current_version}",
f"# New: {recommended_version.to_string()}",
f"# Bump Type: {bump_type.value}",
""
]
for category, cmd_list in commands.items():
output_lines.append(f"## {category.upper()}")
for cmd in cmd_list:
output_lines.append(cmd)
output_lines.append("")
output_text = '\n'.join(output_lines)
else: # text format
output_lines = [
f"Current Version: {args.current_version}",
f"Recommended Version: {recommended_version.to_string()}",
f"With v prefix: {recommended_version.to_string(include_v_prefix=True)}",
f"Bump Type: {bump_type.value}",
""
]
if args.analysis:
analysis = bumper.analyze_commits()
output_lines.extend([
"Commit Analysis:",
f"- Total commits: {analysis['total_commits']}",
f"- Breaking changes: {len(analysis['breaking_changes'])}",
f"- New features: {len(analysis['features'])}",
f"- Bug fixes: {len(analysis['fixes'])}",
f"- Ignored commits: {len(analysis['ignored'])}",
""
])
if analysis['breaking_changes']:
output_lines.append("Breaking Changes:")
for change in analysis['breaking_changes']:
scope = f"({change['scope']})" if change['scope'] else ""
output_lines.append(f" - {change['type']}{scope}: {change['description']}")
output_lines.append("")
if args.include_commands:
commands = bumper.generate_bump_commands(recommended_version)
output_lines.append("Bump Commands:")
for category, cmd_list in commands.items():
output_lines.append(f" {category}:")
for cmd in cmd_list:
if not cmd.startswith('#'):
output_lines.append(f" {cmd}")
output_lines.append("")
output_text = '\n'.join(output_lines)
# Write output
if args.output:
with open(args.output, 'w', encoding='utf-8') as f:
f.write(output_text)
else:
print(output_text)
if __name__ == '__main__':
main()Thiết lập tương tác một thử nghiệm autoresearch mới: lĩnh vực, file đích, lệnh đánh giá, chỉ số, hướng tối ưu và bộ đánh giá.
---
name: "setup"
description: "Set up a new autoresearch experiment interactively. Collects domain, target file, eval command, metric, direction, and evaluator."
command: /ar:setup
---
# /ar:setup — Create New Experiment
Set up a new autoresearch experiment with all required configuration.
## Usage
```
/ar:setup # Interactive mode
/ar:setup engineering api-speed src/api.py "pytest bench.py" p50_ms lower
/ar:setup --list # Show existing experiments
/ar:setup --list-evaluators # Show available evaluators
```
## What It Does
### If arguments provided
Pass them directly to the setup script:
```bash
python {skill_path}/scripts/setup_experiment.py \
--domain {domain} --name {name} \
--target {target} --eval "{eval_cmd}" \
--metric {metric} --direction {direction} \
[--evaluator {evaluator}] [--scope {scope}]
```
### If no arguments (interactive mode)
Collect each parameter one at a time:
1. **Domain** — Ask: "What domain? (engineering, marketing, content, prompts, custom)"
2. **Name** — Ask: "Experiment name? (e.g., api-speed, blog-titles)"
3. **Target file** — Ask: "Which file to optimize?" Verify it exists.
4. **Eval command** — Ask: "How to measure it? (e.g., pytest bench.py, python evaluate.py)"
5. **Metric** — Ask: "What metric does the eval output? (e.g., p50_ms, ctr_score)"
6. **Direction** — Ask: "Is lower or higher better?"
7. **Evaluator** (optional) — Show built-in evaluators. Ask: "Use a built-in evaluator, or your own?"
8. **Scope** — Ask: "Store in project (.autoresearch/) or user (~/.autoresearch/)?"
Then run `setup_experiment.py` with the collected parameters.
### Listing
```bash
# Show existing experiments
python {skill_path}/scripts/setup_experiment.py --list
# Show available evaluators
python {skill_path}/scripts/setup_experiment.py --list-evaluators
```
## Built-in Evaluators
| Name | Metric | Use Case |
|------|--------|----------|
| `benchmark_speed` | `p50_ms` (lower) | Function/API execution time |
| `benchmark_size` | `size_bytes` (lower) | File, bundle, Docker image size |
| `test_pass_rate` | `pass_rate` (higher) | Test suite pass percentage |
| `build_speed` | `build_seconds` (lower) | Build/compile/Docker build time |
| `memory_usage` | `peak_mb` (lower) | Peak memory during execution |
| `llm_judge_content` | `ctr_score` (higher) | Headlines, titles, descriptions |
| `llm_judge_prompt` | `quality_score` (higher) | System prompts, agent instructions |
| `llm_judge_copy` | `engagement_score` (higher) | Social posts, ad copy, emails |
## After Setup
Report to the user:
- Experiment path and branch name
- Whether the eval command worked and the baseline metric
- Suggest: "Run `/ar:run {domain}/{name}` to start iterating, or `/ar:loop {domain}/{name}` for autonomous mode."