@admin
Tạo, lập kế hoạch và tối ưu lead magnet để thu thập email và khách hàng tiềm năng: nội dung gated, ebook, cheat sheet, checklist, template tải về.
---
name: lead-magnets
description: When the user wants to create, plan, or optimize a lead magnet for email capture or lead generation. Also use when the user mentions "lead magnet," "gated content," "content upgrade," "downloadable," "ebook," "cheat sheet," "checklist," "template download," "opt-in," "freebie," "PDF download," "resource library," "content offer," "email capture content," "Notion template," "spreadsheet template," or "what should I give away for emails." Use this for planning what to create and how to distribute it. For interactive tools as lead magnets, see free-tools. For writing the actual content, see copywriting. For the email sequence after capture, see emails.
metadata:
version: 2.0.0
---
# Lead Magnets
You are an expert in lead magnet strategy. Your goal is to help plan lead magnets that capture emails, generate qualified leads, and naturally lead to product adoption.
## Before Planning
**Check for product marketing context first:**
If `.agents/product-marketing.md` exists (or `.claude/product-marketing.md`, or the legacy `product-marketing-context.md` filename, in older setups), read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
Gather this context (ask if not provided):
### 1. Business Context
- What does the company do?
- Who is the ideal customer?
- What problems does your product solve?
### 2. Current Lead Generation
- How do you currently capture leads?
- What lead magnets or offers do you have?
- What's your current conversion rate on email capture?
### 3. Content Assets
- What existing content could be repurposed? (blog posts, guides, data)
- What expertise can you package?
- What templates or tools do you use internally?
### 4. Goals
- Primary goal: email list growth, lead quality, product education?
- Target audience stage: awareness, consideration, or decision?
- Timeline and resource constraints?
---
## Lead Magnet Principles
### 1. Solve a Specific Problem
- Address one clear pain point, not a broad topic
- "How to write cold emails that get replies" > "Marketing guide"
### 2. Match the Buyer Stage
- Awareness leads need education
- Consideration leads need comparison and evaluation
- Decision leads need implementation help
### 3. High Perceived Value, Low Time Investment
- Should look like it's worth paying for
- Consumable in under 30 minutes (ideally under 10)
- Immediate, actionable takeaway
### 4. Natural Path to Product
- Solves a problem your product also solves
- Creates awareness of a gap your product fills
- Demonstrates your expertise in the space
### 5. Easy to Consume
- One clear format (don't mix ebook + video + spreadsheet)
- Works on mobile
- No special software required
---
## Lead Magnet Types
| Type | Best For | Effort | Time to Create |
|------|----------|--------|----------------|
| Checklist | Quick wins, process steps | Low | 1-2 hours |
| Cheat sheet | Reference material, shortcuts | Low | 2-4 hours |
| Template (doc/spreadsheet/Notion) | Repeatable processes, workflows | Low-Med | 2-8 hours |
| Swipe file | Inspiration, examples | Medium | 4-8 hours |
| Ebook/guide | Deep education, authority | High | 1-3 weeks |
| Mini-course (email) | Education + nurture | Medium | 1-2 weeks |
| Mini-course (video) | Education + personality | High | 2-4 weeks |
| Quiz/assessment | Segmentation, engagement | Medium | 1-2 weeks |
| Webinar | Authority, live engagement | Medium | 1 week prep |
| Resource library | Ongoing value, return visits | High | Ongoing |
| Free trial/community access | Product experience | Varies | Varies |
**For detailed creation guidance per format**: See [references/format-guide.md](references/format-guide.md)
---
## Matching Lead Magnets to Buyer Stage
### Awareness Stage
Goal: Educate on the problem. Attract people who don't know you yet.
| Format | Example |
|--------|---------|
| Checklist | "10-Point Website Audit Checklist" |
| Cheat sheet | "SEO Cheat Sheet for Beginners" |
| Ebook/guide | "The Complete Guide to Email Marketing" |
| Quiz | "What Type of Marketer Are You?" |
### Consideration Stage
Goal: Help evaluate solutions. Build trust and demonstrate expertise.
| Format | Example |
|--------|---------|
| Comparison template | "CRM Comparison Spreadsheet" |
| Assessment | "Marketing Maturity Assessment" |
| Case study collection | "5 Companies That 3x'd Their Pipeline" |
| Webinar | "How to Choose the Right Analytics Tool" |
### Decision Stage
Goal: Help implement. Remove friction to purchase.
| Format | Example |
|--------|---------|
| Template | "Ready-to-Use Sales Email Templates" |
| Free trial | "14-Day Free Trial" |
| Implementation guide | "Migration Checklist: Switch in 30 Minutes" |
| ROI calculator | "Calculate Your Savings" (→ see **free-tools**) |
---
## Gating Strategy
### Gating Options
| Approach | When to Use | Trade-off |
|----------|-------------|-----------|
| **Full gate** | High-value content, bottom-funnel | Max capture, lower reach |
| **Partial gate** | Preview + full version | Balance of reach and capture |
| **Ungated + optional** | Top-funnel education | Max reach, lower capture |
| **Content upgrade** | Blog post + bonus | Contextual, high-intent |
### What to Ask For
- **Email only** — highest conversion, lowest friction
- **Email + name** — enables personalization, slight friction increase
- **Email + company/role** — better lead qualification, more friction
- **Multi-field** — only for high-value offers (webinars, demos)
Rule of thumb: Ask for the minimum needed. Every extra field reduces conversion by 5-10%.
### How to Frame the Exchange
- Make the value obvious: "Get the full 25-page guide free"
- Show a preview: table of contents, first page, sample results
- Add social proof: "Downloaded by 5,000+ marketers"
- Reduce risk: "No spam. Unsubscribe anytime."
**For form optimization**: See **cro** skill
**For popup implementation**: See **popups** skill
---
## Landing Page & Delivery
### Landing Page Structure
1. **Headline** — Clear benefit: what they'll get and why it matters
2. **Preview/mockup** — Visual of the lead magnet (cover, screenshot, sample page)
3. **What's inside** — 3-5 bullet points of key takeaways
4. **Social proof** — Download count, testimonials, logos
5. **Form** — Minimal fields, clear CTA button
6. **FAQ** — Address hesitations (Is it really free? What format?)
**For landing page optimization**: See **cro** skill
### Delivery Methods
| Method | Pros | Cons |
|--------|------|------|
| **Instant download** | Immediate gratification | No email verification |
| **Email delivery** | Verifies email, starts relationship | Slight delay |
| **Thank you page + email** | Best of both—instant access + email copy | Slightly more complex |
| **Drip delivery** | Builds habit, multiple touchpoints | Only for courses/series |
### Thank You Page Optimization
Don't waste the thank you page. After they've converted:
- Confirm delivery ("Check your inbox")
- Offer a next step (book a demo, start trial, join community)
- Share on social (pre-written tweet/post)
- Recommend related content
---
## Promotion & Distribution
### Blog CTAs & Content Upgrades
- Add relevant CTAs within blog posts (inline, end-of-post)
- Create post-specific content upgrades (bonus checklist for a how-to post)
- Content upgrades convert 2-5x better than generic sidebar CTAs
### Exit-Intent & Popups
- Trigger on exit intent or scroll depth
- Match the popup offer to the page content
- **See popups** for implementation
### Social Media
- Share snippets and teasers from the lead magnet
- Create carousel posts from key points
- Use the lead magnet as the CTA in your bio/profile
- **See social** for social strategy
### Paid Promotion
- Facebook/Instagram lead ads for top-funnel lead magnets
- Google Ads for high-intent lead magnets (templates, tools)
- LinkedIn for B2B lead magnets
- Retarget blog visitors with lead magnet ads
- **See ads** for campaign strategy
### Partner Co-Promotion
- Cross-promote with complementary brands
- Guest webinars with partner audiences
- Include in partner newsletters
- Bundle in resource collections
---
## Measuring Success
### Key Metrics
| Metric | What It Tells You | Benchmark |
|--------|-------------------|-----------|
| **Landing page conversion rate** | Offer attractiveness | 20-40% (warm traffic), 5-15% (cold) |
| **Cost per lead** | Acquisition efficiency | Varies by channel and industry |
| **Lead-to-customer rate** | Lead quality | 1-5% (B2B), varies widely |
| **Email engagement** | Content relevance | 30-50% open, 2-5% click |
| **Time to conversion** | Nurture effectiveness | Track by lead magnet source |
**For detailed benchmarks by format and industry**: See [references/benchmarks.md](references/benchmarks.md)
### A/B Testing Ideas
- **Headline**: Benefit-focused vs. curiosity-driven
- **Format**: Checklist vs. guide on same topic
- **Gate level**: Full gate vs. partial preview
- **Form fields**: Email-only vs. email + name
- **CTA copy**: "Download Free Guide" vs. "Get Your Copy"
- **Delivery**: Instant download vs. email delivery
### Lead Quality Signals
Good lead magnet attracted quality leads if:
- Higher-than-average email engagement
- Leads progress to trial/demo at expected rates
- Low unsubscribe rate after delivery
- Leads match ICP demographics
---
## Output Format
When creating a lead magnet strategy, provide:
### 1. Lead Magnet Recommendation
- Format and topic
- Target buyer stage
- Why this format for this audience
- Estimated creation effort
### 2. Content Outline
- Key sections/components
- Length and scope
- What makes it unique or valuable
### 3. Gating & Capture Plan
- What to gate and how
- Form fields
- Landing page structure
### 4. Distribution Plan
- Promotion channels
- Content upgrade opportunities
- Paid amplification (if applicable)
### 5. Measurement Plan
- KPIs and targets
- What to A/B test first
---
## Task-Specific Questions
1. What existing content or expertise could you turn into a lead magnet?
2. Where does your audience spend time online?
3. What's the most common question prospects ask before buying?
4. Do you have an email nurture sequence set up for new leads?
5. What's your budget for design and promotion?
---
## Related Skills
- **free-tools**: For interactive tools as lead magnets (calculators, graders, quizzes)
- **copywriting**: For writing the lead magnet content itself
- **emails**: For nurture sequences after lead capture
- **cro**: For optimizing lead magnet landing pages
- **popups**: For popup-based lead capture
- **cro**: For optimizing capture forms
- **content-strategy**: For content planning and topic selection
- **analytics**: For measuring lead magnet performance
- **ads**: For paid promotion of lead magnets
- **social**: For social media promotion
FILE:evals/evals.json
{
"skill_name": "lead-magnets",
"evals": [
{
"id": 1,
"prompt": "We're a B2B SaaS selling project management software to marketing agencies. What lead magnet should we create?",
"expected_output": "Should check for product-marketing.md first. Should ask about current lead gen, existing content assets, and primary goal (list growth, lead quality, product education). Should apply Lead Magnet Principles: solve a specific problem (not 'agency marketing'), match buyer stage, high perceived value + low time investment, natural path to product. Should recommend a specific format suited to a busy agency audience — likely a template (Notion/spreadsheet) or checklist over an ebook. Examples: 'Agency Project Profitability Calculator' (decision stage, naturally leads to project management), 'Client Onboarding Checklist for Agencies' (consideration), 'The Agency Capacity Planning Template' (decision stage). Should justify the choice by matching buyer stage and effort/value ratio. Should outline content, gating, landing page, distribution, and measurement plan.",
"assertions": [
"Checks for product-marketing.md",
"Asks about buyer stage and goal",
"Applies the 5 principles",
"Recommends specific format with rationale",
"Examples match the audience and product",
"Outlines all 5 output sections (recommendation, content, gating, distribution, measurement)"
],
"files": []
},
{
"id": 2,
"prompt": "We have a 50-page ebook we spent 3 months writing. Conversion on the landing page is only 4%. Should we keep iterating?",
"expected_output": "Should diagnose this as a likely mismatch on Lead Magnet Principles, especially #3 (high perceived value, low time investment — consumable in under 30 minutes, ideally under 10). Should warn 50 pages may signal too much effort to consume — flag this as a possible cause. Should recommend A/B testing the format (chunking the ebook into a 5-part email mini-course, releasing as a checklist + ebook combo, or breaking into shorter topic-specific guides). Should review landing page structure: headline, preview/mockup, what's inside, social proof, form fields, FAQ. Should suggest testing partial gate (preview first 5 pages) vs full gate. Should ask about traffic source — 4% on cold traffic might be acceptable while 4% on warm traffic is low. Should reference cro skill for landing page optimization and ab-testing for test design.",
"assertions": [
"Diagnoses likely cause as length/effort mismatch",
"Recommends format A/B test",
"Suggests breaking into shorter formats",
"Reviews landing page structure",
"Asks about traffic source (cold vs warm)",
"Cross-references cro or ab-testing skill"
],
"files": []
},
{
"id": 3,
"prompt": "Our lead form asks for name, email, company, role, company size, and phone. We're not getting enough signups. Could the form be the problem?",
"expected_output": "Should immediately flag form length as a likely culprit. Should cite the rule of thumb: every extra field reduces conversion 5-10%. Should recommend reducing to the minimum needed: ideally email only (highest conversion), or email + name if personalization matters. Should explain when multi-field is justified (only for high-value offers like webinars or demos). Should ask what information is actually used in follow-up — fields that aren't used should be removed. Should suggest progressive profiling: capture email now, ask for more fields later via enrichment or follow-up forms. Should reference cro skill for form optimization specifically.",
"assertions": [
"Flags form length as likely culprit",
"Cites 5-10% per field rule",
"Recommends reducing to email or email + name",
"Asks what fields are actually used",
"Suggests progressive profiling",
"Cross-references cro skill"
],
"files": []
},
{
"id": 4,
"prompt": "What's the difference between a lead magnet and a free tool? Should I build one or the other?",
"expected_output": "Should explain the distinction: lead magnets are static content offers (ebooks, checklists, templates) while free tools are interactive (calculators, graders, quizzes). Should explain when to build which. Lead magnets: faster to ship (hours-days), works well for awareness/consideration education, lower ongoing maintenance, lead quality varies. Free tools: longer build time (weeks-months), higher engagement and shareability, naturally segment leads by tool usage, can rank for SEO ('X calculator', 'Y grader'), higher lead quality typically. Should recommend lead magnet first if speed matters, free tool if you can invest the build time and have repeatable user inputs that produce a meaningful output. Should defer to free-tools skill for tool strategy specifically.",
"assertions": [
"Distinguishes static content from interactive tool",
"Compares effort to build",
"Compares SEO and shareability characteristics",
"Recommends based on speed vs investment trade-off",
"Defers to free-tools skill"
],
"files": []
},
{
"id": 5,
"prompt": "We have a top-performing blog post on email subject lines. Can we use it as a lead magnet?",
"expected_output": "Should recommend creating a content upgrade specific to the post rather than gating the post itself (post-specific content upgrades convert 2-5x better than generic sidebar CTAs). Should suggest specific upgrade ideas: '50 Email Subject Line Templates' (template format, decision stage), 'Subject Line Cheat Sheet PDF' (cheat sheet format, awareness/consideration), 'Subject Line Swipe File' (collection of high-performing examples with annotations). Should explain content upgrades convert better because they match what the reader is already engaged with — relevance + intent are higher than generic offers. Should recommend keeping the blog post ungated (preserve SEO) and offering the upgrade as an inline or end-of-post CTA. Should reference cro for placement and copywriting for the upgrade itself.",
"assertions": [
"Recommends content upgrade over gating the post",
"Cites 2-5x improvement vs generic CTAs",
"Suggests specific upgrade formats with rationale",
"Keeps blog post ungated to preserve SEO",
"Explains why upgrades convert better"
],
"files": []
},
{
"id": 6,
"prompt": "Our checklist gets a lot of downloads but very few of them ever sign up for a trial. Is the lead magnet broken?",
"expected_output": "Should diagnose this as a lead quality / buyer stage mismatch problem. Should ask whether the checklist is awareness-stage content drawing people who aren't ready to buy. Should check Lead Quality Signals: higher-than-average email engagement, leads progress to trial/demo at expected rates, low unsubscribe rate, leads match ICP demographics. Should review the principle: lead magnets should create a natural path to product. If a checklist for total beginners attracts beginners, that's working as designed but they won't convert quickly — they need nurture. Should recommend reviewing the nurture sequence (cross-reference emails skill) and checking whether the offer matches the right buyer stage for the goal. May suggest creating a consideration- or decision-stage lead magnet (template, ROI calculator, comparison spreadsheet) that pulls higher-intent leads. Should track time to conversion by lead magnet source.",
"assertions": [
"Diagnoses as lead quality / buyer stage mismatch",
"Asks about ICP fit of leads",
"References Lead Quality Signals",
"Cross-references emails skill for nurture",
"Suggests a decision-stage lead magnet alternative",
"Mentions tracking time to conversion by source"
],
"files": []
}
]
}
FILE:references/benchmarks.md
# Lead Magnet Benchmarks
Reference data for planning and evaluating lead magnet performance.
---
## Conversion Rate Benchmarks
### By Format Type
| Format | Landing Page Conversion | Notes |
|--------|------------------------|-------|
| Checklist | 30-50% | High because low commitment |
| Cheat sheet | 25-40% | Quick reference appeal |
| Template | 25-45% | Immediate utility drives conversion |
| Ebook/guide | 20-35% | Higher commitment, lower rate |
| Quiz | 30-50% | Engagement drives completion |
| Webinar | 20-40% (registration) | 30-50% attendance rate of registrants |
| Mini-course | 15-30% | Higher commitment, higher quality leads |
| Free trial | 5-15% | High intent but high friction |
### By Traffic Source
| Source | Expected Conversion | Why |
|--------|-------------------|-----|
| Blog content upgrade | 3-8% of post readers | Contextually relevant |
| Dedicated landing page (organic) | 20-40% | High intent |
| Dedicated landing page (paid) | 10-25% | Cold traffic |
| Exit-intent popup | 2-5% of visitors | Interruption-based |
| Sidebar/banner CTA | 0.5-2% | Low engagement |
| Social media link | 10-20% | Warm but browsing |
### By Industry (Landing Page)
| Industry | Average Conversion |
|----------|-------------------|
| SaaS/Tech | 15-25% |
| Marketing/Agency | 20-35% |
| Finance | 10-20% |
| E-commerce | 10-20% |
| Education | 20-35% |
| Health/Wellness | 15-25% |
---
## Lead Quality Indicators
### Signals of High-Quality Leads
- Open first 3 emails at 40%+ rate
- Click through to content or product pages
- Return to site within 30 days
- Match ICP demographics (role, company size, industry)
- Progress to trial, demo, or purchase within 90 days
### Signals of Low-Quality Leads
- Unsubscribe within first 3 emails
- Never open beyond delivery email
- Use disposable email addresses
- Don't match target customer profile
- Downloaded for the content, no product interest
### Quality vs. Quantity by Format
| Format | Lead Volume | Lead Quality | Net Value |
|--------|-------------|-------------|-----------|
| Generic ebook | High | Low-Medium | Medium |
| Specific template | Medium | High | High |
| Industry report | Medium | Medium-High | High |
| Quiz/assessment | High | Medium (segmentable) | High |
| Webinar | Low-Medium | High | High |
| Checklist | High | Low-Medium | Medium |
| Free trial | Low | Very High | Very High |
---
## Cost Benchmarks
### Cost Per Lead by Channel
| Channel | Typical CPL | Notes |
|---------|-------------|-------|
| Organic search | $0-5 | Lowest, but slow to build |
| Blog content upgrade | $0-2 | Nearly free if you have traffic |
| Facebook/Instagram Ads | $3-15 | B2C lower, B2B higher |
| Google Ads | $10-50 | High intent, higher cost |
| LinkedIn Ads | $25-75 | B2B, expensive but qualified |
| Partner co-promotion | $0-5 | Depends on relationship |
### Creation Cost by Format
| Format | DIY Cost | With Designer/Freelancer |
|--------|----------|-------------------------|
| Checklist | Free | $100-300 |
| Cheat sheet | Free | $200-500 |
| Template | Free | $100-500 |
| Ebook (10-25 pages) | Free | $500-2,000 |
| Quiz | $0-100/mo (tool) | $500-2,000 |
| Webinar | Free (Zoom) | $500-1,500 (production) |
| Mini-course (email) | Free | $500-1,500 (copywriting) |
| Video course | $0-200 (gear) | $2,000-5,000 |
---
## Timeline Expectations
### Time to Create
| Format | Solo Creator | With Team |
|--------|-------------|-----------|
| Checklist | 1-2 hours | Same day |
| Cheat sheet | 2-4 hours | Same day |
| Template | 2-8 hours | 1-2 days |
| Swipe file | 4-8 hours | 1-2 days |
| Ebook | 1-3 weeks | 1-2 weeks |
| Quiz | 1-2 weeks | 1 week |
| Webinar prep | 1 week | 3-5 days |
| Mini-course | 1-2 weeks | 1 week |
### Time to See Results
| Phase | Timeline |
|-------|----------|
| First leads | Immediately with existing traffic or paid |
| Organic traffic growth | 2-6 months (SEO) |
| Meaningful lead volume | 1-3 months |
| Measurable impact on pipeline | 3-6 months |
| Full ROI assessment | 6-12 months |
**Note**: These benchmarks are general guidelines. Your actual results depend on audience, niche, traffic volume, and offer quality. Start measuring from day one and build your own benchmarks.
FILE:references/format-guide.md
# Lead Magnet Format Guide
Detailed creation guidance for each lead magnet format.
## Contents
- Ebooks & Guides
- Checklists
- Cheat Sheets
- Templates & Spreadsheets
- Swipe Files
- Mini-Courses
- Quizzes & Assessments
- Webinars & Workshops
---
## Ebooks & Guides
**Best for**: Building authority, deep education, awareness-stage leads
**Structure**:
1. Title page with professional design
2. Table of contents
3. Introduction — frame the problem, set expectations
4. 3-7 chapters — one key concept per chapter
5. Summary — recap key takeaways
6. CTA — next step toward your product
**Guidelines**:
- Ideal length: 10-25 pages (shorter is fine if valuable)
- Include visuals: charts, diagrams, screenshots
- Use callout boxes for key stats or quotes
- End each chapter with a quick takeaway
- Don't pad — density beats length
**Tools**: Canva, Google Docs → PDF, Notion export, Designrr, Beacon.by
---
## Checklists
**Best for**: Process-oriented tasks, quick wins, implementation help
**Structure**:
- Title: "[Number]-Point [Topic] Checklist"
- Numbered or checkbox items
- Group into logical sections if 10+ items
- Brief explanation per item (1-2 sentences)
**Guidelines**:
- Keep to 1-2 pages
- Use actionable language ("Verify X", "Set up Y", "Remove Z")
- Order by workflow sequence or priority
- Make it printable — clean layout, generous spacing
- Include a "done" checkbox for each item
**What works**: Step-by-step processes, audit criteria, launch checklists, setup guides
---
## Cheat Sheets
**Best for**: Reference material, shortcuts, quick-lookup information
**Structure**:
- One page (two pages max)
- Organized by category or workflow
- Dense but scannable
- Visual hierarchy with headers and grouping
**Guidelines**:
- Optimize for quick reference, not reading
- Use tables, grids, or columns
- Include formulas, shortcuts, or code snippets
- Design for printing or saving as desktop reference
- Bold the most important items
**What works**: Keyboard shortcuts, formula references, terminology glossaries, decision matrices
---
## Templates & Spreadsheets
**Best for**: Repeatable processes, planning, tracking
### Spreadsheet Templates (Google Sheets / Excel)
- Include a "How to Use" tab with instructions
- Pre-fill with example data
- Use data validation for dropdown fields
- Add conditional formatting for visual cues
- Lock formula cells, leave input cells editable
- Include a "Make a Copy" link (Google Sheets)
### Notion Templates
- Provide a duplicate link
- Include a getting-started guide
- Pre-populate with example content
- Use Notion's database features (views, filters, relations)
- Keep it simple — don't over-engineer
### Document Templates
- Provide in multiple formats (Google Doc, Word, PDF)
- Include placeholder text with [BRACKETS] for customization
- Add inline instructions in a different color
- Make it immediately usable with minimal editing
**Key principle**: Templates should be usable within 5 minutes of downloading.
---
## Swipe Files
**Best for**: Inspiration, examples, learning from others
**Structure**:
- Curated collection of 15-50 examples
- Organized by category, type, or use case
- Each example includes:
- The example itself (screenshot, text, link)
- Why it works (2-3 bullet annotations)
- How to adapt it (1-2 sentences)
**Guidelines**:
- Quality over quantity — curate ruthlessly
- Add your analysis, don't just collect
- Organize for browsing (categories, tags)
- Update periodically with fresh examples
- Credit original sources
**What works**: Email subject lines, landing pages, ad copy, CTAs, onboarding flows, pricing pages
---
## Mini-Courses
### Email-Based Mini-Courses
- 3-5 emails delivered over 5-7 days
- One lesson per email, one concept per lesson
- Each email: teach → example → exercise
- Progressive difficulty (build on previous lessons)
- Final email: summary + CTA for product or next step
### Video-Based Mini-Courses
- 3-5 videos, 5-15 minutes each
- Host on unlisted YouTube, Loom, or course platform
- Deliver links via email drip
- Include worksheets or exercises per lesson
- More personal — builds stronger connection
**Cadence**: Every 1-2 days. Don't stretch too thin or compress too tight.
**Key principle**: Each lesson should deliver standalone value. If someone only watches lesson 2, they should still learn something useful.
---
## Quizzes & Assessments
**Best for**: Engagement, segmentation, personalized results
**Question Design**:
- 5-10 questions (sweet spot: 7)
- Multiple choice only — no open-ended
- Questions should feel insightful, not obvious
- Progress indicator ("Question 3 of 7")
**Result Segmentation**:
- 3-5 result categories
- Each result: name, description, personalized recommendations
- Tailor follow-up emails by result type
- Share-worthy result format ("I got: Growth Stage Marketer!")
**Implementation**: Gate results behind email capture. The quiz itself is ungated — the personalized results require an email.
**For building interactive quizzes**: See **free-tools** skill for technical implementation guidance.
---
## Webinars & Workshops
### Live Webinars
- 30-45 minutes teaching + 15 minutes Q&A
- Structure: Hook → Teach (3 key points) → Demo/example → CTA
- Promote 1-2 weeks in advance
- Send 3 reminder emails (confirmation, day before, 1 hour before)
- Record for replay (extends value)
### Evergreen Webinars
- Pre-recorded, available on demand
- Same structure as live but tighter editing
- Always-on lead generation
- Gate with email registration
- Automated follow-up sequence
**Follow-up**: Send replay link + summary + CTA within 24 hours. Continue with nurture sequence.
**Key principle**: Teach something genuinely useful. A webinar that's just a sales pitch will damage trust.
Hỗ trợ quan hệ công chúng: earned media, thông cáo báo chí, tiếp cận nhà báo và chiến lược truyền thông.
---
name: public-relations
description: "When the user wants help with public relations, earned media, press coverage, journalist outreach, or media strategy (not pull requests). Also use when the user mentions 'PR,' 'public relations,' 'press,' 'press release,' 'press coverage,' 'media outreach,' 'pitch a journalist,' 'get featured,' 'media list,' 'media kit,' 'press kit,' 'newsjacking,' 'news hijack,' 'HARO,' 'Qwoted,' 'Featured,' 'Help A Reporter,' 'reporter request,' 'tech press,' 'TechCrunch,' 'earned media,' 'thought leadership placement,' 'op-ed,' 'guest article,' 'press contacts,' 'podcast prep,' 'going on a podcast,' 'podcast guest,' 'prep me for this podcast,' or 'how do I get press.' Use this for earned media work — finding journalists, pitching stories, newsjacking, prepping podcast appearances, and responding to press requests. For startup/SaaS/AI directory submissions, see directory-submissions. For product launches, see launch. For social-media engagement, see social. For cold-email outreach to prospects, see cold-email."
metadata:
version: 1.1.1
---
# Public Relations & Earned Media
You are an expert in earned media for software products. Your goal is to help the user get covered by journalists, podcasts, and newsletters — efficiently, with respect for the people on the other end of the pitch.
## Before Starting
**Check for product marketing context first:**
If `.agents/product-marketing.md` exists (or `.claude/product-marketing.md`, or the legacy `product-marketing-context.md` filename, in older setups), read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
---
## Core Philosophy
PR is not a substitute for distribution. It's a multiplier for it.
- **Earned media doesn't drive direct conversions.** A TechCrunch hit will not give you 1,000 paying customers. It will give you backlinks, brand legitimacy, AI-citation surface area, and ammo for sales conversations.
- **Pitch journalists like you'd pitch a customer:** specific, useful, fast, and never about you.
- **The story is not your product. The story is the trend, the data, the conflict, or the human.** Your product is the evidence. Every pitchable story bends toward one of three angles — Founding Story, David vs Goliath, or Have an Enemy (a *broken system*, never a competitor). See [references/story-angles.md](references/story-angles.md).
- **Chase press for the compound effect, not the traffic bump.** The bump fades in a day; authority, journalist relationships, and AI-citation surface compound. Build media relationships *before* you need them, and run one core asset through the whole repurposing flywheel.
- **Speed beats polish on reactive PR.** A B+ pitch in the first hour of a story beats an A+ pitch on day three.
### When PR is worth it
- You have **a real story** — proprietary data, a strong opinion, a milestone, a customer with a sharp before/after, or a fresh angle on a trending topic
- You have **founder/exec time** — journalists want quotes from people with skin in the game, not from a PR rep
- You have **a destination** — a press page, blog post, or product launch that converts attention into something useful
### When to skip PR (for now)
- Pre-launch with no story beyond "we exist"
- No one on the team can sustain pitching for 4–6 weeks (PR is a momentum game)
- You don't have a clear ICP — journalists ask "who reads my piece because of this?" and if you can't answer, neither can they
---
## The PR Mix
Four modes. Most teams over-index on one. Run at least three.
| Mode | What it is | Effort | Speed to coverage |
|------|------------|--------|-------------------|
| **Reactive (newsjacking)** | Inject your POV into trending news | Low–medium | Hours to days |
| **Proactive (pitching)** | Build a media list, pitch original stories | High | 2–8 weeks |
| **Inbound (press requests)** | Respond to journalist queries on HARO/Qwoted/Featured | Low | Days to weeks |
| **Owned (press page + media kit)** | Make it easy for journalists to find you | One-time setup | N/A |
**For the story angle taxonomy (Founding Story / David vs Goliath / Have an Enemy), data stories, media relationship-building, and the PR repurposing flywheel** — see [references/story-angles.md](references/story-angles.md)
**For the reactive newsjacking workflow** — see [references/newsjacking.md](references/newsjacking.md)
**For proactive journalist pitching** — see [references/journalist-pitching.md](references/journalist-pitching.md)
**For inbound press-request platforms (HARO, Qwoted, etc.)** — see [references/press-platforms.md](references/press-platforms.md)
**For where to pitch (media outlets, podcasts, newsletters)** — see [references/media-outlets.md](references/media-outlets.md). For startup/SaaS/AI directories, use the separate `directory-submissions` skill — different intent, different list.
**For prepping a podcast appearance you've landed** — see [references/podcast-guest-prep.md](references/podcast-guest-prep.md). Episodes get transcribed and cited by AI assistants, so a good appearance compounds in AI answers for years — prep is an AI-visibility play, not just interview polish.
---
## Owned: Press Page + Media Kit
Set this up once. It's the cheapest PR investment with the highest ROI on every future story.
**Press page (`/press` or `/newsroom`) should include:**
- One-paragraph company description (copy/paste ready)
- Founder bios with headshots (high-res, downloadable)
- Logo pack (SVG + PNG, light + dark, with usage guidelines)
- Product screenshots (high-res)
- Recent coverage list (social proof for the next journalist)
- Founding date, employee count, funding (if disclosed)
- Press contact email (not a form — journalists hate forms)
- Recent press releases / announcements
**One sentence at the top:** "For interview requests or assets, email press@yourcompany.com — we respond within 24 hours."
Then *actually* respond within 24 hours.
---
## Quick Reference: Pitch Quality Bar
Before sending any pitch, the answer to all of these should be yes:
- [ ] Does this journalist cover this beat? (Check their last 5 articles.)
- [ ] Is there a clear news hook — something that just happened or is about to?
- [ ] Could this journalist write a complete story from this email alone? (Data, quotes, customer name, contact.)
- [ ] Is the subject line specific enough to predict the article's headline?
- [ ] Is the pitch under 150 words?
- [ ] Did you avoid the words "revolutionary," "game-changing," "disruptive," and "synergy"?
- [ ] Is the ask clear? (Interview? Embargo? Exclusive? Quote?)
If any answer is no, don't send.
---
## Measurement
What to track:
| Metric | Why |
|--------|-----|
| **Coverage count** (placements / month) | Activity baseline |
| **Domain rating of placements** | Backlink value |
| **Referral traffic from coverage** | Did anyone actually click? |
| **Brand search lift** | Did people search you after reading? |
| **AI citation rate** (ChatGPT, Perplexity quote your brand?) | The new measurement that matters |
| **Sales conversations citing the article** | The only one that matters for revenue |
What not to obsess over: AVE (advertising value equivalency) — it's a vanity metric PR firms invented.
---
## Common Workflows
### "Help me newsjack [trending story]"
Go to [newsjacking.md](references/newsjacking.md), run the scoring rubric, draft 2–3 angles, pick the best, draft the pitch.
### "Find journalists who cover [beat]"
Go to [journalist-pitching.md](references/journalist-pitching.md), use the discovery checklist + dev-browser to research recent articles, build a scored list.
### "What's worth pitching this week?"
Combine: recent product milestones + active news cycles + any data you've collected. Score each potential story by the quality bar above.
### "What's my story angle?" / "How do I get press with no news?"
Go to [story-angles.md](references/story-angles.md). Fit the situation to one of the three angles (Founding Story / David vs Goliath / Have an Enemy), or turn proprietary data into a data story. Remember: a milestone alone isn't a story — milestone *with narrative* is.
### "Respond to this HARO query"
Go to [press-platforms.md](references/press-platforms.md), use the response template, keep it under 200 words.
### "I'm going on [podcast] next week — help me prep"
Go to [podcast-guest-prep.md](references/podcast-guest-prep.md): research the show (RSS feed → site → Apple Podcasts → web), extract the recurring threads and host profiles, map the guest's stories onto them, deliver the brief.
### "Build my press page"
Use the checklist above. Most companies do this in an afternoon and forget about it for a year — that's fine.
FILE:evals/evals.json
{
"skill_name": "public-relations",
"evals": [
{
"id": 1,
"prompt": "We're a tiny 2-person SaaS competing against Notion in the docs space. We don't have any funding news or a big milestone. How do we get press coverage?",
"expected_output": "Should route to references/story-angles.md. Should recognize there's no traditional news hook and instead fit the situation to one of the three story angles, most naturally David vs Goliath (own your size against an incumbent, cite the Superhuman vs Gmail exemplar). Should mention the other two angles — Founding Story (Pieter Levels '12 startups in 12 months,' build-in-public, share revenue) and Have an Enemy (the enemy is a broken system, not the competitor — never 'Notion is bad,' and reference HEY vs Apple's App Store 2020). Should reinforce 'the story is not your product' and that a milestone alone isn't a story — milestone with narrative is. Should note the compound effect (authority, relationships, AI-citation surface) over the traffic bump. May suggest turning proprietary data into a data story and building media relationships before pitching. Should not invent fake news or suggest attacking the competitor by name.",
"assertions": [
"Routes to references/story-angles.md",
"Identifies David vs Goliath as the fitting angle with the Superhuman vs Gmail exemplar",
"Names all three angles (Founding Story, David vs Goliath, Have an Enemy)",
"Clarifies the enemy is a broken system, not a named competitor (HEY vs Apple App Store 2020)",
"Cites Pieter Levels '12 startups in 12 months' for Founding Story",
"Reinforces 'the story is not your product'",
"Notes milestone-with-narrative and/or the compound effect over the traffic bump",
"Does not fabricate news or suggest smearing the competitor by name"
],
"files": []
}
]
}
FILE:references/journalist-pitching.md
# Journalist Pitching — Proactive PR Workflow
Building a media list, scoring journalist fit, and crafting pitches that actually get opened. This is a 4–8 week practice, not a one-shot.
## Contents
- Building the media list
- Scoring journalist fit
- Pitch templates by angle
- Subject lines that get opened
- Voice and structure
- Embargoes, exclusives, and follow-up etiquette
- Pitch killers
- Tooling
---
## Building the Media List
The goal: a list of 20–40 journalists who actually cover your beat. Not 500 names from a database.
### Discovery checklist
For each candidate journalist:
- [ ] Read their **last 5 articles** — are they covering your beat right now?
- [ ] Note their **publication** — does it reach your ICP?
- [ ] Check their **bio** on the outlet site — what topics do they own?
- [ ] Check **X/LinkedIn** for what they're posting about this week
- [ ] Note their **email** (usually on outlet author page, Muck Rack, or company About page)
- [ ] Check **Muck Rack** if available — it shows recent topics and pitch preferences
### Where to find candidates
| Method | How |
|--------|-----|
| **Reverse lookup from coverage you want** | Find 5 articles about competitors / your category, note bylines |
| **Topic search on Muck Rack** | Free tier shows journalists by topic |
| **X / Twitter lists** | "[your niche] reporters" lists already exist |
| **LinkedIn search** | "Journalist" + "[your category]" — filter by recent activity |
| **Newsletter author pages** | Beehiiv, Substack, ConvertKit creators are pitchable |
| **Podcast host research** | Listen to 1 episode before pitching — non-negotiable |
### Don't waste time on
- **Mass media databases** (Cision, Meltwater) for early-stage — overkill and expensive
- **Journalists who haven't posted in 6+ months** — they may have left
- **"Editor-in-chief" generic addresses** — pitches there get ignored or routed to interns
- **Journalists who explicitly state "no PR pitches" in bio** — respect it
---
## Scoring Journalist Fit
Score each journalist 1–10 across four dimensions. Sum and rank. Focus on top 20.
| Dimension | What it measures | Weight |
|-----------|------------------|--------|
| **Beat match** | Do they cover your category specifically? | 3x |
| **Reach** | Outlet's audience size + their byline traction | 2x |
| **Engagement** | Do they respond to pitches publicly / on X? | 2x |
| **Recency** | Have they written about a related topic in last 30d? | 1x |
**Tiering:**
- **Tier 1 (8–10):** Personal pitch with original angle. Custom each time.
- **Tier 2 (5–7):** Standard pitch, lightly customized.
- **Tier 3 (below 5):** Skip or pitch only when story is exceptional.
---
## Pitch Templates by Angle
Six structures that work. Pick the one that matches your story.
### 1. Data story
```
Subject: [Specific stat] — [implication]
Hi [name],
I noticed you covered [recent article] — wanted to share data that might
be relevant.
We [analyzed N / surveyed N / tracked N] and found:
• [Stat 1 with surprise factor]
• [Stat 2]
• [Stat 3]
The most interesting pattern: [one-sentence insight].
Full data + methodology here: [link to one-pager, not your homepage]
Happy to share the raw dataset, jump on a call, or connect you with
[customer who's relevant].
[your name + 1-line credential]
```
### 2. Exclusive launch / milestone
```
Subject: Exclusive: [specific milestone] at [company]
Hi [name],
I have an exclusive on [milestone] that I think fits your [beat] coverage.
The story: [one sentence]
Why it matters: [one sentence — for their readers, not for you]
What's new: [the actual news, not the marketing line]
Embargo until [day, time, timezone] — would love to give you first
window. Press kit + assets: [link]
Free to talk [two specific time options].
[your name]
```
### 3. Op-ed / contributed piece
```
Subject: Op-ed pitch: [provocative thesis]
Hi [name],
I read your piece on [recent article] — sharp take on [specific point].
I'd like to pitch a 700-word op-ed: "[Thesis as a headline]"
Core argument:
• [Point 1]
• [Point 2]
• [Point 3 — the surprising one]
Why me: [1 sentence — credential or unique vantage]
Why now: [1 sentence — the news hook]
Can have a draft to you by [date]. Happy to adapt to your house style.
[your name]
```
### 4. Customer story
```
Subject: Customer story for [their beat] — [specific outcome]
Hi [name],
For your [beat] coverage, I have a [customer type] willing to talk on
the record about [specific outcome].
The hook: [customer] [did something specific] and [measurable result].
The interesting part: [the surprising or counterintuitive detail].
Customer details:
• Name: [name, title, company]
• Available: [windows]
• Willing to share: [data points / screenshots / metrics]
Happy to coordinate the intro.
[your name]
```
### 5. Trend piece / connector
```
Subject: Trend forming in [space] — three signals
Hi [name],
Three things in [space] this month that I think connect:
1. [Signal 1 with link]
2. [Signal 2 with link]
3. [Signal 3 — yours, briefly]
The pattern: [one sentence].
This might be early for a piece, but if you're tracking the space I
wanted to flag it. Happy to share data we've collected or connect you
with others seeing the same.
[your name]
```
### 6. Newsjack response
```
Subject: Re: [their article headline] — quick data point
Hi [name],
Saw your piece on [story] this morning — wanted to add a relevant
data point in case you do a follow-up.
[One-sentence stat or insight].
Source: [our data / our customers / our analysis]
Methodology: [one sentence]
Quotable: "[a sentence you'd be comfortable seeing in print]"
If useful for a follow-up, I'm around all day at this number: [phone].
[your name]
```
---
## Subject Lines That Get Opened
Journalists open pitches based on the subject line alone. Rules:
- **Under 50 characters** — mobile preview cuts off
- **Lead with the specific** — "73% of devs deploy to prod on Fridays" beats "New data on developer workflows"
- **Promise a story, not a product** — "Why [trend]" beats "[Company] launches [thing]"
- **Use prefixes that signal value** — "Exclusive:", "Data:", "Op-ed pitch:", "Re: [their article]"
**Test against this question:** would *you* open this in a 200-email inbox?
**Patterns that work:**
- "[Specific stat] — [implication]" — "73% of agents fail this test"
- "Exclusive: [milestone]" — "Exclusive: Anthropic launches AgentOS"
- "Re: [their headline]" — direct response to recent coverage
- "[Provocative thesis]" — "Why VC funding is bad for AI safety"
**Patterns that get deleted:**
- "Press release: [boring]"
- "[Company] announces [thing]"
- "Story idea for you!"
- "Following up on my previous email"
- Any subject line with "innovative," "disruptive," "revolutionary"
---
## Voice and Structure
### Length
**150 words max for the pitch.** If you can't say it in 150 words, you don't know what your story is yet.
### Structure
1. **One-line context** — why you're emailing them specifically (their recent article, their beat)
2. **The story** — what it is, in one sentence
3. **Why it matters to their readers** — not why it matters to you
4. **Proof** — data, customer, quote, link
5. **The ask** — interview, embargo, quote, link
### Voice
- Sound like a person, not a press release
- Reference their actual recent work — proves you read them
- Don't use emoji unless they do
- Don't open with "I hope this finds you well" — burn it
- Don't ask "did you get my email?" follow-ups (see [Follow-up](#embargoes-exclusives-and-follow-up-etiquette))
### Banned vocabulary
Revolutionary, disruptive, game-changing, paradigm shift, leverage, synergy, robust, seamless, holistic, world-class, best-in-class, next-generation, cutting-edge, AI-powered (unless that's the actual differentiation), at-the-end-of-the-day.
---
## Embargoes, Exclusives, and Follow-Up Etiquette
### Embargoes
An embargo is "you can write this story, but don't publish until [time]."
- **Only offer embargoes to journalists you've worked with** or have strong reputations for honoring them
- State the embargo time clearly: day, time, timezone
- If they break embargo, your relationship with them is over
### Exclusives
"Only you get this story" — powerful tool, use sparingly.
- **First-tier outlet only** — exclusives to tier 2/3 outlets waste the lever
- **Be honest about scope** — "exclusive to [outlet] in the US" is fine
- **Have a parallel plan** — what you publish/pitch the next day after the exclusive runs
### Follow-up cadence
- **Day 0** — initial pitch
- **Day 3** — one follow-up if you have new information ("Just talked to [customer] who can join us")
- **Day 7** — final check-in with a fresh hook ("This came out today, still relevant?")
- **After day 7** — let it go. Re-pitch when you have something genuinely new.
**Never:**
- "Bumping this up" / "Did you see my email?"
- Multi-day silent follow-ups with no new value
- Same pitch reformatted
---
## Pitch Killers
Things that instantly disqualify your pitch:
- Wrong name / wrong outlet (autoreplace fail)
- Pitching topics they explicitly don't cover
- Press release attached as PDF (just paste the key bits)
- Long signature with logos and disclaimers
- CC'ing 5 other journalists on the same email
- "Per my last email" energy
- Asking them to sign an NDA before talking
- Pitching a story you can't actually tell (no customer willing to talk, no data ready to share)
---
## Tooling
### Finding journalist contact info
```bash
# Most journalists' emails follow patterns:
# firstname@outlet.com
# firstname.lastname@outlet.com
# flastname@outlet.com
# Use Hunter.io, RocketReach, or just guess and bounce-check
```
### Researching their recent work (browser-driven)
Use `dev-browser` (persistent session, no rate limits) to:
- Open the journalist's outlet author page → scrape last 5 article headlines + dates
- Open their X/Twitter profile → note recent topics
- Open their LinkedIn → confirm current role
Output what you find as:
```
JOURNALIST PROFILE — [name]
Outlet: [name]
Beat: [topics from last 5 articles]
Recent angle: [pattern you noticed]
Recent X activity: [what they're posting]
Score: [X/40 from rubric]
Best pitch angle: [from template library]
Email: [confirmed]
```
### Maintaining the media list
Store in `.agents/media-list.md` (or `.csv` if you prefer). Update monthly — journalists move jobs constantly.
```markdown
## Tier 1 (top 20)
| Name | Outlet | Beat | Last contact | Last coverage | Email | Score |
|------|--------|------|--------------|---------------|-------|-------|
| ... | ... | ... | 2026-05-15 | none yet | ... | 9/10 |
```
### Pitch tracking
Track in a simple spreadsheet:
- Date sent
- Subject line
- Journalist
- Outlet
- Response (open / reply / pass / coverage)
- What they said
After 30 pitches, you'll see which subject patterns and which angles work for you specifically.
FILE:references/media-outlets.md
# Media Outlets — Where to Pitch
A curated, opinionated list of *where* to pitch for software/SaaS PR. This is the media-outlet slice of resources like submit.co — the journalist-driven half. For startup/SaaS/AI directories (Product Hunt, BetaList, Futurepedia, etc.), use the separate `directory-submissions` skill — different intent, different audience.
## How to use this list
- **Don't pitch a publication. Pitch a journalist at that publication.** See [journalist-pitching.md](journalist-pitching.md) for the discovery workflow.
- **Tier signals quality, not effort** — a small tier-3 outlet might be perfect for your niche
- **Submission/tip URLs are listed where they exist** — but a journalist's email beats a tip form every time
- **This list ages fast** — verify the outlet still exists and the journalist is still there before pitching
---
## Tech & Startup Press (Tier 1)
The big names. High bar to clear, high payoff when you do. Pitch the specific reporter, not the tip line.
| Outlet | Best for | Tip URL |
|--------|---------|---------|
| **TechCrunch** | Funding, product launches, startup news | techcrunch.com/got-a-tip/ |
| **The Verge** | Consumer tech, product reviews, policy | theverge.com/contact |
| **Wired** | Long-form tech, culture, business | wired.com/about/feedback |
| **Fast Company** | Innovation, design, business strategy | fastcompany.com/contact-us |
| **VentureBeat** | AI, enterprise tech, gaming | venturebeat.com/contribute |
| **Ars Technica** | Deep tech, science, policy | arstechnica.com/contact-us |
| **The Information** | Subscription-gated, scoops on tech industry | theinformation.com/about |
| **Bloomberg / Reuters** | Business, finance, big-picture stories | bloomberg.com/feedback |
| **WSJ Tech** | Enterprise, business angle | wsj.com/tips |
| **NYT Tech** | Mainstream tech, culture | nytimes.com/tips |
---
## SaaS & B2B (Tier 1–2)
Lower-profile than the consumer tech outlets but often higher ROI for B2B SaaS.
| Outlet | Best for | Notes |
|--------|----------|-------|
| **SaaStr** | B2B SaaS founders, growth, sales | Jason Lemkin's outlet; pitch contributed posts |
| **First Round Review** | Operator-level B2B insights | High bar; original frameworks only |
| **OpenView** | SaaS metrics, PLG, pricing | Often runs research-backed pieces |
| **Lenny's Newsletter** | Product / growth | Subscriber-only; pitch via Lenny directly on X |
| **Future (a16z)** | Tech + culture pieces | Long-form, original thinking required |
| **Stratechery** | Tech strategy analysis | Don't pitch Ben Thompson; engage via responses |
---
## AI / ML Press (Tier 1–2)
The hottest beat right now. Reporters here are inundated — your angle has to be sharp.
| Outlet | Best for | Notes |
|--------|----------|-------|
| **The Decoder** | AI news, fast-turn | Daily AI news cycle |
| **Import AI** (Jack Clark) | Weekly AI newsletter | Pitch research-backed angles |
| **The Batch** (Andrew Ng) | AI industry analysis | Submitted by deeplearning.ai team |
| **MIT Technology Review** | AI policy, capability, ethics | Higher bar, slower cycle |
| **Hugging Face blog** | Open-source AI tools | Contributed posts welcome |
| **Latent Space** | Practitioner-focused AI/ML | Podcast + newsletter |
---
## Developer & DevTools Press
Where to pitch if your audience is engineers.
| Outlet | Best for | Notes |
|--------|----------|-------|
| **The New Stack** | DevTools, infra, cloud-native | Active contributor program |
| **InfoQ** | Enterprise dev, software architecture | Long-form technical pieces |
| **DEV.to** | Self-published, community-driven | Build credibility before pitching |
| **Hacker Noon** | Tech blogging platform | Easy to publish but low signal |
| **Console** | DevTools newsletter | Curated weekly; submit projects |
| **DevTools FM** | Podcast | Pitch as guest |
---
## Business & Marketing Press
For pitching the business / marketing angle of your story.
| Outlet | Best for | Notes |
|--------|----------|-------|
| **Inc.** | Founder stories, business advice | Contributor program available |
| **Entrepreneur** | Small business, growth | Volume publisher; quality varies |
| **HBR.org** | Original research, frameworks | High bar; long lead time |
| **MarketingProfs** | B2B marketing | Contributor-friendly |
| **Marketing Brew (Morning Brew)** | Daily marketing newsletter | Reporter-driven, pitch directly |
| **Marketing Land / Search Engine Journal** | SEO, SEM, channels | Niche but high-intent audience |
| **Reforge** | Growth, product, retention | Original framework required |
---
## Newsletters (Reporter-Driven)
Newsletters are increasingly the most valuable PR placement — small audiences, but high-intent.
| Newsletter | Audience | How to pitch |
|-----------|---------|--------------|
| **Lenny's Newsletter** | Product/growth, ~600k readers | DM Lenny on X with sharp angle |
| **The Pragmatic Engineer** (Gergely Orosz) | Software eng leaders | Pitch via email; original engineering insights |
| **The Generalist** (Mario Gabriele) | Tech business analysis | Pitch via email; deep angles |
| **Stratechery** (Ben Thompson) | Tech strategy | Don't pitch; engage in replies |
| **Newcomer** (Eric Newcomer) | Tech business + scoops | Tips welcome via email |
| **Platformer** (Casey Newton) | Tech + policy | Pitch via Substack |
| **Big Technology** (Alex Kantrowitz) | Big Tech analysis | Pitch via Substack |
| **Not Boring** (Packy McCormick) | Tech + culture | Tough to crack; original takes only |
**For B2B SaaS:**
- **SaaStr Daily**
- **Demand Curve**
- **Growth Unhinged** (Kyle Poyar)
- **Trends.vc**
**For AI:**
- **Import AI** (Jack Clark)
- **The Batch** (Andrew Ng)
- **The Algorithm** (MIT Tech Review)
- **Interconnects** (Nathan Lambert)
- **AI Tidbits**
---
## Podcasts
A podcast appearance is often higher leverage than a press hit — longer engagement, evergreen replay, audience trust transfer.
### Top SaaS / startup podcasts
- **Lenny's Podcast** — product/growth
- **My First Million** — founder stories, business ideas
- **The Twenty Minute VC** — funding angle
- **SaaStr Podcast** — B2B SaaS
- **Acquired** — deep-dive company stories (don't pitch unless you're a unicorn)
- **The All-In Podcast** — broad tech/business (very hard to get on)
### Top AI podcasts
- **No Priors** (Sarah Guo, Elad Gil)
- **Latent Space**
- **Practical AI**
- **The TWIML AI Podcast**
- **The Cognitive Revolution**
### Top dev / engineering podcasts
- **The Changelog**
- **Software Engineering Daily**
- **DevTools FM**
### How to pitch a podcast
1. **Listen to 3 episodes** — non-negotiable
2. **Find the host's preferred channel** (X DM, email, guest form)
3. **Pitch a topic, not yourself** — "I'd love to come on and talk about [specific angle]" not "I'd love to be a guest"
4. **Bring evidence** — links to other appearances, your unique angle, what listeners will learn
5. **Make it easy** — bio, headshot, suggested questions in the pitch
---
## Industry / Vertical Press
Don't overlook trade press — smaller audience, much higher intent.
| Vertical | Outlets to investigate |
|----------|----------------------|
| **Marketing** | Adweek, Marketing Brew, AdAge, Search Engine Land |
| **Sales** | Sales Hacker, Modern Sales Pros |
| **HR / People Ops** | HR Brew, SHRM, HR Dive |
| **Finance / Fintech** | The Block, CoinDesk, Finextra |
| **Healthcare tech** | STAT, MobiHealthNews, Healthcare IT News |
| **Education tech** | EdSurge, EdScoop |
| **Real estate tech** | Inman, The Real Deal |
| **Legal tech** | Above the Law, Law360 |
| **Climate tech** | Heatmap, Canary Media, Latitude Media |
| **Devtools** | The New Stack, InfoQ, DevOps.com |
For your specific vertical: Google `"top publications" + "[your industry]"` and run the same scoring exercise from [journalist-pitching.md](journalist-pitching.md).
---
## Regional / Local
If your company has a regional angle (HQ location, customer concentration, government contract), local press is underrated.
- **Local business journals** — Bizjournals network covers 40+ US cities
- **Local NPR affiliates** — high-quality, business angle welcomed
- **Local TV business segments** — high reach, easy to get
- **State / regional tech news** — e.g., Built In (Chicago, Austin, etc.), TechBuzz (Utah)
---
## What's NOT On This List (And Why)
- **Product Hunt, BetaList, Indie Hackers** — these are directories, not press. Use the `directory-submissions` skill.
- **Press release wires** (PRNewswire, BusinessWire, GlobeNewswire) — overpriced for early-stage; journalists ignore them. Skip until you have IR / SEC requirements.
- **"As featured in" badge mills** — paid "media coverage" services. Worthless and damaging.
- **Random "guest post" SEO link networks** — Google penalizes these. Don't.
---
## Maintaining This List
This list will go stale. Recommended cadence:
- **Quarterly:** verify your tier-1 contacts are still at the outlet
- **Monthly:** add new outlets/newsletters relevant to your category
- **Per pitch:** confirm the journalist is still there before sending (check X / LinkedIn for "joined [new place]")
Store your live, working version in `.agents/media-list.md` (per [journalist-pitching.md](journalist-pitching.md)).
FILE:references/newsjacking.md
# Newsjacking — Reactive PR Workflow
Injecting your POV into a story that's already trending. Done well: free distribution off a wave of attention. Done badly: cringe at best, brand damage at worst.
## Contents
- When newsjacking works (and when it doesn't)
- The detect → score → angle → pitch loop
- Newsworthiness scoring rubric
- Story angle library
- Speed: the only thing that matters
- Sources & tooling
- Failure modes
---
## When Newsjacking Works
- **Tech/regulatory news in your category** — new law, new platform launch, competitor pivot, big acquisition
- **Industry data drops** — a major report drops, you have a sharper take or contradicting data
- **Public conversation** — a debate or controversy where your expertise is genuinely relevant
- **Seasonal/cyclical moments** — earnings season, year-end reviews, conference weeks
## When to Skip
- **Tragedies, accidents, deaths** — no exceptions. Don't.
- **Politically charged stories** unless your brand explicitly takes political stances
- **You have no genuine expertise** in the area
- **The window is already closed** — if a story is 48h+ old and you weren't first, you're late
- **The angle is "we have a product for this"** — that's marketing, not journalism
---
## The Loop
A repeatable workflow Claude can run on demand or daily.
1. **Detect** — surface trending stories in your category (see [Sources & Tooling](#sources--tooling))
2. **Score** — apply the [newsworthiness rubric](#newsworthiness-scoring-rubric); drop anything below threshold
3. **Angle** — generate 2–3 angles per story using the [angle library](#story-angle-library)
4. **Validate** — sanity-check: do you actually have the expertise/data to back this angle?
5. **Pitch** — draft a tight pitch to 3–5 journalists who cover this beat (see [journalist-pitching.md](journalist-pitching.md))
6. **Post** — also publish on your blog, LinkedIn, X — it builds the trail journalists check before quoting you
Output format Claude should produce:
```
NEWSJACK CANDIDATE — 2026-06-10
Story: "EU passes AI Act amendment requiring agent registration"
Source: TechCrunch, 3h ago
Score: 8/10 (high relevance, fresh, you have proprietary data)
Angles:
1. Data hot take: "Our analysis of 12,000 agent deployments shows 73% would fail this requirement"
2. Contrarian: "Why the registration rule will hurt safety, not improve it"
3. Customer story: "How [customer] is preparing — interview offer"
Recommended: #1 (you have unique data, strongest hook)
Pitch draft: [see journalist-pitching.md for template]
Target journalists: [list with rationale]
```
---
## Newsworthiness Scoring Rubric
Score each candidate 1–10 on five dimensions, multiply by the weight, then sum. Max possible: 80 (10 × the 8x weight total).
| Dimension | What it measures | Weight |
|-----------|------------------|--------|
| **Timeliness** | Story <24h old? Window still open? | 2x |
| **Relevance** | Genuinely in your expertise area? | 2x |
| **Angle uniqueness** | Can you say something no one else is saying? | 2x |
| **Authority** | Do you have data, customers, or experience to back it? | 1x |
| **Reach potential** | Will this story keep growing or has it peaked? | 1x |
**Threshold:** weighted total ≥ 50/80. Below that, skip.
**Auto-disqualify if:**
- The story is about something tragic
- Your angle is "I disagree" with nothing to back it
- You haven't actually formed an opinion — you just want to be quoted
---
## Story Angle Library
Use these templates to generate angles fast.
### 1. Data hot take
*"We analyzed [N] [things] after [event]. Here's what we found."*
Best when you have proprietary data. The journalist gets a stat, you get the citation.
### 2. Contrarian
*"Everyone says [popular take]. Here's why they're wrong."*
Best when you can defend the position with specifics. Weak when it's just contrarianism for attention.
### 3. "We predicted this"
*"Six months ago we wrote [thing] — here's what's happening now and what's next."*
Best when you actually did predict it. Lethal to your credibility if you didn't.
### 4. Customer impact
*"Here's a [customer type] who's directly affected. We can put you in touch."*
Best for B2B. Reporters love named customers willing to talk.
### 5. Insider explainer
*"This story is complicated. Here's what's actually happening."*
Best when most coverage is missing nuance. You're not arguing — you're educating.
### 6. Trend connector
*"This isn't isolated — it's part of a bigger shift we're seeing in [pattern]."*
Best when you have several data points or examples to connect.
### 7. Founder POV
*"As someone who's built in this space for [X years], here's the part most people are missing."*
Best for opinion pieces / op-eds. Weak as a soundbite pitch.
---
## Speed: The Only Thing That Matters
Newsjacking decays fast. Approximate windows:
| Story type | Effective window |
|-----------|------------------|
| Breaking tech news | 4–12 hours |
| Major regulation / policy | 24–48 hours |
| Industry report / data drop | 24–72 hours |
| Conference announcement | Same day |
| Acquisition / funding news | 12–24 hours |
**Implication:** if you can't draft and send within the window, don't bother. Set up the loop so detection → pitch takes <2 hours.
---
## Sources & Tooling
Reuses tooling from the `social` skill's listening workflow. Same install: `brew install jq`.
### Google News RSS (no auth)
```bash
# Replace QUERY with topic (use + for spaces, %22 for quotes)
curl -s "https://news.google.com/rss/search?q=QUERY&hl=en-US&gl=US&ceid=US:en" \
| xmllint --xpath "//item[position()<11]" - 2>/dev/null
```
### Hacker News (Algolia) for tech stories
```bash
SINCE=$(($(date +%s) - 86400))
curl -s "https://hn.algolia.com/api/v1/search_by_date?query=QUERY&tags=story&numericFilters=created_at_i>SINCE" \
| jq '.hits[] | {title, url, points, num_comments, created_at, hn_url: ("https://news.ycombinator.com/item?id="+.objectID)}'
```
### Reddit (for category-specific subs)
```bash
curl -s -A "newsjack/1.0" \
"https://www.reddit.com/r/SUBREDDIT/top.json?t=day&limit=15" \
| jq '.data.children[].data | {title, url, score, num_comments, created_utc}'
```
### Journalist research (browser-driven)
For finding *which* journalists are covering the story right now:
- **dev-browser** → Google News search for the story → click through to articles → note the bylines
- Then go to those journalists' X / LinkedIn / Muck Rack profile to confirm beat and recent coverage
See also [journalist-pitching.md](journalist-pitching.md) for the full discovery workflow.
### Source list
For repeatable monitoring, add a "Newsjacking topics" section to `.agents/listening-sources.md` (template in the `social` skill's references):
```markdown
## Newsjacking topics (Google News RSS)
- "AI agent regulation"
- "[your category] funding"
- "[your competitors] OR [adjacent competitors]"
## Industry data drops (RSS / manual)
- Pitchbook reports
- a16z state of [industry] reports
- [your category] benchmark reports
```
---
## Failure Modes
Things that have ended careers and brands.
- **Tragedy-jacking** — Oreo's 2013 Super Bowl tweet worked. Most attempts since have not. Wartime, disasters, deaths: don't.
- **The forced fit** — "Here's our take on [trending story] — it's actually about [our product]." Journalists see through this instantly.
- **The empty take** — pitching "we have an opinion" without specifics. Journalists need a quote-worthy line, not "we're closely watching this."
- **Speed without judgment** — being first with a bad take is worse than being late with a good one. The 30-minute "is this brand-appropriate?" gut check exists for a reason.
- **Pitching the same angle to 50 journalists** — they talk. Get caught once, lose the relationships.
- **No follow-through** — pitch goes out, journalist responds in 20 minutes, founder takes 6 hours to reply. Story moves on.
---
## Companion Practice: The Public Trail
Every newsjack pitch is stronger if the journalist can find evidence you've been thinking about this publicly. Before pitching:
1. Publish a short post (blog, LinkedIn, X thread) with your take
2. Reference it in the pitch ("more thinking here: [link]")
3. This signals you're not opportunistic — you're an actual voice in the space
If you don't have time to publish, you're probably not ready to pitch.
FILE:references/podcast-guest-prep.md
# Podcast Guest Prep
Build a prep brief before the user appears on a podcast as a guest. The goal: walk in knowing what's top of mind for the show, how it has evolved, who the hosts are, and which of the user's stories map onto what the show cares about *right now*.
**Why prep is worth real effort:** podcast guesting isn't just audience reach. Episodes get transcribed, show notes get published, and both get crawled and cited by AI assistants — when someone asks ChatGPT about your category, the stories you told on a podcast two years ago are part of what it draws on. A good appearance is earned media that compounds in AI answers for years (see the `ai-seo` skill). The stories you tell — and the concrete numbers in them — become the citable record on your brand. Prep accordingly.
## Context to load first
Read `.agents/product-marketing.md` (or `.claude/product-marketing.md`) for the company, positioning, and ICP. That file usually won't have the guest's *story bank*, so also collect — in one batch, not a drip:
1. What did you build before this that comes up in conversation?
2. What are 2–3 stories you tell well, with real numbers attached?
3. What's one opinion you hold that most people in your space disagree with?
Offer to save the answers into the product-marketing context doc so future runs skip the interview.
From their message, establish (ask only if missing and it matters): the podcast name or URL, whether they've appeared before (get the prior episode link — it anchors the progression analysis and the callbacks), and roughly when they're recording.
## Research sequence
Work through sources in this order; each is a fallback for the last:
1. **RSS feed first.** The richest source: full episode descriptions, chapter markers, guest links, dates. Find the feed link on the podcast site (Buzzsprout, Transistor, etc. all expose one). Large-feed fetches may truncate — check whether the oldest episodes you need actually made it.
2. **The podcast website's episode list** for anything the feed missed. These pages often lazy-load older episodes via JavaScript; if pagination returns nothing, note the gap and move on rather than burning time.
3. **Apple Podcasts show page** — reliably renders the latest ~8 episodes with full descriptions.
4. **Web search** for stray episodes, the hosts, and the show's reputation.
Don't fetch every episode page. Descriptions plus chapter lists are almost always enough; only pull a full transcript when a specific episode is central (e.g., a debate the guest should have a position on). Check for published transcript links in the feed.
## What to extract
- **Recent-episode threads** (last ~3 months or 6–8 episodes): per-episode topic summaries, then the *recurring threads* — the questions the hosts keep returning to. Threads matter more than individual episodes; they predict the questions the guest will get.
- **Show progression** (since their last appearance, or ~12–18 months if first time): identify phases and the inflection point where the show's focus shifted. Note whether the show re-invites guests (signals how a return visit fits) and whether hosts launched side projects.
- **Host profiles.** Sources: the show's about pages, hosts' personal sites, LinkedIn, and — often the best source — episodes where the hosts guest on *other* shows and introduce themselves. Capture day job, background, what they've personally been building (mine solo-episode summaries), and social handles. If a host shares the guest's first name, flag it and keep references unambiguous throughout the brief.
- **Prior appearance recap** (if returning): what was actually discussed, with rough timestamps, and how much airtime the guest's current company got. This sets up the "what's changed since" narrative.
## The brief
Write a markdown file and present the short version in chat. Structure:
1. **Big picture** — what kind of show this is now, and the one-paragraph read on how the guest should position themselves
2. **Show progression** — the phases since their last appearance (or show start)
3. **Recent episodes in detail** — per-episode notes, then the recurring threads
4. **Guest angles** — their stories mapped explicitly onto the show's threads, callbacks to any prior appearance, and 2–3 "pocket" items: concrete stories with numbers to have ready
5. **The hosts** — profiles plus rapport hooks (where each host's world overlaps the guest's)
6. **Gaps** — anything unretrievable, and offers to fill them
Keep it tight — a doc they'll skim before recording, not a report.
**Finding angles:** map the story bank onto the show's recurring threads. The shape to look for: a "data moats" thread maps to the guest's proprietary dataset as a live case study; an "AI replacing niche tools" debate maps to a defensibility story from their product history. One well-chosen contrarian take stands out most on shows that have converged on a consensus.
**AI-visibility angle:** since the transcript becomes the record, coach the guest to say the important things in liftable form — the company name next to the category ("we build X, the Y for Z"), and numbers spoken aloud, not gestured at. Same logic as the YouTube text layer in `ai-seo`.
## Follow-ups to offer (don't auto-run)
Transcribing their prior episode for a word-level review, pulling a full transcript of one pivotal recent episode, or drafting likely Q&A.
---
*Distilled and adapted from [ai-visibility-skills](https://github.com/Knowatoa/ai-visibility-skills) by Knowatoa (MIT), reused with credit.*
FILE:references/press-platforms.md
# Press Request Platforms — Inbound PR
Journalists posting "I need a source for X" — you respond, sometimes you get quoted, sometimes you don't. The cheapest PR play available, but only if you treat it seriously.
## Contents
- The major platforms
- Daily triage workflow
- Response template
- What makes a response get selected
- What kills a response
- ROI reality check
---
## The Major Platforms
| Platform | What it is | Cost | Quality |
|----------|-----------|------|---------|
| **[Connectively](https://www.connectively.us)** (formerly HARO) | Daily email digest of journalist queries | Free tier; paid for filters | Mixed — high volume, lots of noise |
| **[Qwoted](https://www.qwoted.com)** | Web app with journalist requests | Free; paid for outreach | Good — better-quality outlets |
| **[Featured](https://featured.com)** | Web app, expert profiles, journalist requests | Free tier; paid pro | Good for thought-leadership snippets |
| **[Help A B2B Writer](https://helpab2bwriter.com)** | Twice-weekly email of B2B queries | Free | High — B2B-focused, low spam |
| **[SourceBottle](https://www.sourcebottle.com)** | Australia-focused but global queries | Free | Variable |
| **[Terkel](https://terkel.io)** | Roundup-style ("we asked 50 experts…") | Free | Volume-heavy, low effort |
| **[JournoRequests](https://twitter.com/journorequests)** | X account aggregating tweets | Free | UK-skewed, real-time |
| **#JournoRequest** (X hashtag) | Live journalist requests | Free | Real-time, fast-moving |
**Recommended starter set:** Connectively + Qwoted + Help A B2B Writer + monitoring `#JournoRequest` on X.
---
## Daily Triage Workflow
These platforms generate volume. Treat it like email triage — fast pass, deep response on the rare matches.
### Step 1 — Filter (5 min)
For each digest / request feed:
- Drop everything where you don't have **direct experience or data**
- Drop everything from outlets your ICP doesn't read
- Drop everything with a deadline you can't meet
- Keep only requests where you can give a **complete, named, on-the-record answer**
Realistic conversion: 50 daily requests → 2–4 worth answering.
### Step 2 — Deep response (15 min per request)
For each keeper:
- Read the request 3 times — what's the *actual* angle?
- Look up the journalist if possible — recent coverage, beat
- Write a custom response (see [template](#response-template))
- Send within their stated deadline (early > late)
### Step 3 — Log
Track in a spreadsheet:
- Date
- Platform
- Journalist + outlet
- Topic
- Response sent (yes/no)
- Outcome (no response / passed / quoted / linked)
After 30 responses, you'll see which topics/platforms convert.
---
## Response Template
Keep responses under 200 words. Journalists are scanning 50+ replies for one quote.
```
Hi [name],
Quick response to your request about [topic].
[Specific credential — 1 sentence. "Built X for 5 years" / "Led marketing at Y" / "Have analyzed N companies in space"]
The most important thing about [topic]: [your actual point in 2 sentences].
[A specific example, story, or data point — this is what gets quoted.]
[If applicable: a contrarian or surprising angle that differentiates from typical answers.]
Happy to expand on any of this, share data, or be quoted directly.
Feel free to use this attribution:
[Your name], [your title], [your company]
Contact for follow-up: [email + phone]
```
**Note the structure:**
1. One-sentence intro
2. One-sentence credential
3. Two-sentence answer
4. Specific example (the quotable part)
5. Optional: differentiator
6. Clear offer
7. Pre-written attribution (saves them 30 seconds)
8. Contact info
---
## What Makes a Response Get Selected
After analyzing hundreds of quoted responses, the patterns:
### Quotable specificity
**Bad:** "Companies should focus on customer experience."
**Good:** "When we A/B tested 47 onboarding flows, the version with a 30-second video at step 3 increased activation by 41%."
The good version is a quote. The bad version is filler.
### Concrete credential
**Bad:** "As a marketing expert..."
**Good:** "I've run growth at three Series B SaaS companies, all in B2B sales tooling."
Specificity beats title-stacking.
### Story over advice
**Bad:** "It's important to track the right metrics."
**Good:** "We almost shut down because we were optimizing for MRR when our real problem was activation. Once we switched to tracking 7-day activation, everything else followed."
Stories make articles. Advice makes filler.
### Pre-formatted for their workflow
- Pre-written attribution
- Multiple quotable lines (let them pick)
- High-res headshot link (don't attach)
- One-line company description
### Time match
**Most quotes come from responses sent in the first 6 hours.** After 24 hours, your chances drop sharply. Treat deadlines as if they're 24h earlier than stated.
---
## What Kills a Response
- **Pitching your product** when they asked for expert commentary
- **Generic advice** that could come from any expert
- **Multiple "experts" from your company** responding to the same request (looks coordinated, often is)
- **Hiring a PR firm to spam responses** — journalists smell it
- **Demanding a link back** to your site — most can't promise links
- **Ignoring the deadline** by 1+ days
- **Long bio sections** before the actual answer
- **Asking to "see the article before publication"** — you don't get to do that
- **Asking what other experts said** so you can differentiate — they won't tell you
---
## ROI Reality Check
Most teams overinvest in these platforms because they're cheap. Be honest:
| Effort | Realistic outcome (90 days) |
|--------|----------------------------|
| 5 hr/week, custom responses | 3–10 quoted placements |
| 1 hr/week, template responses | 0–2 placements |
| Outsourced to PR firm | Lots of submissions, few quotes |
**A quote in a tier-1 outlet is worth:**
- A backlink (DR depends on outlet)
- A sales-collateral asset ("As featured in...")
- AI-citation surface area
- Brand legitimacy in the abstract
**A quote in a tier-3 outlet is worth:**
- A backlink, often nofollow
- Maybe an Instagram screenshot
**Decision rule:** if you can sustain 5 hr/week of quality responses for 90 days, this is worth it. If you can only do 1 hr/week, skip it and invest in [proactive pitching](journalist-pitching.md) instead.
---
## Setup Checklist
Before you start responding:
- [ ] Press page exists and is current (see main SKILL.md)
- [ ] One-line credential is written and rehearsed
- [ ] Headshot is high-res and at a public URL
- [ ] You have 3–5 specific stories / data points ready to deploy
- [ ] You've decided which 2–3 platforms to use (don't try all 7)
- [ ] You've blocked a daily 20-min window for triage
- [ ] You're logging responses in a tracker
Without these, you're spamming and wasting their time and yours.
FILE:references/story-angles.md
# Story Angles — What Earns Press
Distilled from Corey Haines's *Founding Marketing*, ch. 5. Chase press for the **compound effect** — authority, journalist relationships, AI-citation surface — not the traffic bump. The bump fades in a day; the backlinks, the relationships, and the citable record don't.
**The story is not your product.** Journalists write about trends, data, conflict, and humans. Your product is the evidence, never the headline. If the pitch is "we built a thing," there's no story. If the pitch is "here's a shift happening and we're proof of it," there is.
## Contents
- The three story angles
- Newsworthy moments (data stories)
- Build media relationships before you need them
- The PR flywheel (repurposing map)
---
## The Three Story Angles
Every earned-media story worth pitching bends toward one of three shapes. Pick the one that fits your moment — don't force all three.
### 1. Founding Story
The origin, the build, the numbers-in-public. Works because people follow *people*, and a founder with skin in the game is quotable in a way a product never is.
- **When to use:** you're early, you have a sharp personal reason you exist, and you're willing to build in public and share real revenue.
- **Exemplar:** Pieter Levels' "12 startups in 12 months" — shipping in the open, posting MRR screenshots, letting the challenge itself be the story. The build *was* the coverage.
- **How to run it:** share the arc (why you started, what you've learned, where the numbers are now), not the feature list. Revenue transparency is the hook most founders are too scared to use.
### 2. David vs Goliath
Own your size. Being small against an incumbent is an advantage in the story, not a liability — you're the underdog readers root for.
- **When to use:** there's a giant in your category and you do one thing dramatically better than they do.
- **Exemplar:** Superhuman vs Gmail — a tiny team reframing "we're small" into "we're the fast, obsessive alternative to the default everyone tolerates."
- **How to run it:** name the Goliath, name the one thing you beat them at, and let the size gap do the emotional work. Don't hide that you're small — lead with it.
### 3. Have an Enemy
Give the reader something to be against. The enemy is a **broken system**, not a competitor. Attacking a rival looks petty; attacking a genuinely broken status quo looks principled — and journalists cover principled fights.
- **When to use:** there's a systemic wrong in your space you can credibly stand against, and you're willing to take a real position.
- **Exemplar:** HEY vs Apple's App Store (2020) — the fight wasn't "Basecamp vs Apple the company," it was against App Store rules founders saw as broken. It became a press cycle *and* a policy conversation.
- **How to run it:** define the broken system precisely, stake a clear position, and make sure you're actually willing to be quoted taking that stance. A half-hearted enemy reads as a marketing stunt.
**Guardrail:** the enemy must be a system or a norm — never a named competitor. "Company X is bad" is a smear; "this way of doing things is broken" is a movement.
---
## Newsworthy Moments: Data Stories
Beyond the three angles, the most reliably pitchable moment is a **data story** — proprietary numbers no one else has, packaged as an industry benchmark. Journalists can build a whole piece around a stat; you get the citation.
- **Exemplars:** Stripe's *State of New User* / annual data reports; Intercom's benchmark reports. Each turns internal data into an annual, citable, must-cover event.
- **Why it works:** it's original, it's quotable, and it positions you as the source-of-record for your category's numbers. It also compounds in AI answers — benchmark stats get lifted and re-cited for years.
- **How to run it:** find the one number in your data that surprises, wrap it in methodology, and package it as a standalone one-pager (not your homepage). See [journalist-pitching.md](journalist-pitching.md) → the "Data story" template.
**Milestone reminder:** a milestone alone ("we hit $1M ARR") isn't a story — milestone *with narrative* is. Attach the number to a shift, a lesson, or a David-vs-Goliath frame.
---
## Build Media Relationships Before You Need Them
Cold-pitching a journalist the day you have news is the hard way. The founders who get covered built the relationship months earlier. It's a ladder — climb it before you need the favor.
1. **Follow their beat.** Read their last 5 articles. Know what they cover and what they're tired of covering. (See the discovery checklist in [journalist-pitching.md](journalist-pitching.md).)
2. **Add value with no ask.** Reply to their work with a genuinely useful data point or correction. Share their pieces. Answer a question they post on X. Give before you take.
3. **Be a reliable source.** When they need a quote fast, be the person who responds in 20 minutes with a clean, quotable line — even when there's nothing in it for you. Reliability is what turns a contact into a relationship.
Do this for a handful of journalists on your beat and, when you finally have a real story, you're pitching a relationship, not a stranger.
---
## The PR Flywheel (Repurposing Map)
One core asset feeds every other channel. Don't create five things — create one exceptional thing and reformat it five ways. Each format extends the reach of the same idea and gives journalists more of a public trail to find.
```
Core asset (data report / strong POV / founding story)
│
├─→ Thread (X / LinkedIn) — the hook, teased publicly
├─→ Article (your blog) — the full argument, the destination
├─→ Guest post (someone else's audience) — borrow reach
├─→ Podcast appearance — the spoken, citable version
└─→ Video — the durable, AI-crawled version
```
- **Sequence matters less than coverage.** Publish the article as the owned home base, tease it as a thread, pitch the guest post and podcast off the same idea, cut the video from the recording.
- **Every format strengthens the next pitch.** A journalist who can find your thread, article, and podcast on a topic sees a voice in the space — not an opportunist. This is the public trail that makes newsjacking land (see [newsjacking.md](newsjacking.md) → The Public Trail).
- **The asset is the flywheel's fuel.** One strong data story or POV can run this loop for weeks. Weak assets stall it immediately.
Lãnh đạo marketing: định vị thương hiệu, mô hình tăng trưởng, phân bổ ngân sách marketing và thiết kế tổ chức.
---
name: "cmo-advisor"
description: "Marketing leadership for scaling companies. Brand positioning, growth model design, marketing budget allocation, and marketing org design. Use when designing brand strategy, selecting growth models (PLG vs sales-led vs community-led), allocating marketing budgets, building marketing teams, or when user mentions CMO, brand strategy, growth model, CAC, LTV, channel mix, or marketing ROI."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: c-level
domain: cmo-leadership
updated: 2026-03-05
python-tools: marketing_budget_modeler.py, growth_model_simulator.py
frameworks: brand-positioning, growth-frameworks, marketing-org
---
# CMO Advisor
Strategic marketing leadership — brand positioning, growth model design, budget allocation, and org design. Not campaign execution or content creation; those have their own skills. This is the engine.
## Keywords
CMO, chief marketing officer, brand strategy, brand positioning, growth model, product-led growth, PLG, sales-led growth, community-led growth, marketing budget, CAC, customer acquisition cost, LTV, lifetime value, channel mix, marketing ROI, pipeline contribution, marketing org, category design, competitive positioning, growth loops, payback period, MQL, pipeline coverage
## Quick Start
```bash
# Model budget allocation across channels, project MQL output by scenario
python scripts/marketing_budget_modeler.py
# Project MRR growth by model, show impact of channel mix shifts
python scripts/growth_model_simulator.py
```
**Reference docs (load when needed):**
- `references/brand_positioning.md` — category design, messaging architecture, battlecards, rebrand framework
- `references/growth_frameworks.md` — PLG/SLG/CLG playbooks, growth loops, switching models
- `references/marketing_org.md` — team structure by stage, hiring sequence, agency vs. in-house
---
## The Four CMO Questions
Every CMO must own answers to these — no one else in the C-suite can:
1. **Who are we for?** — ICP, positioning, category
2. **Why do they choose us?** — Differentiation, messaging, brand
3. **How do they find us?** — Growth model, channel mix, demand gen
4. **Is it working?** — CAC, LTV:CAC, pipeline contribution, payback period
---
## Core Responsibilities (Brief)
**Brand & Positioning** — Define category, build messaging architecture, maintain competitive differentiation. Details → `references/brand_positioning.md`
**Growth Model** — Choose and operate the right acquisition engine: PLG, sales-led, community-led, or hybrid. The growth model determines team structure, budget, and what "working" means. Details → `references/growth_frameworks.md`
**Marketing Budget** — Allocate from revenue target backward: new customers needed → conversion rates by stage → MQLs needed → spend by channel based on CAC. Run `marketing_budget_modeler.py` for scenarios.
**Marketing Org** — Structure follows growth model. Hire in sequence: generalist first, then specialist in the working channel, then PMM, then marketing ops. Details → `references/marketing_org.md`
**Channel Mix** — Audit quarterly: MQLs, cost, CAC, payback, trend. Scale what's improving. Cut what's worsening. Don't optimize a channel that isn't in the strategy.
**Board Reporting** — Pipeline contribution, CAC by channel, payback period, LTV:CAC. Not impressions. Not MQLs in isolation.
---
## Key Diagnostic Questions
Ask these before making any strategic recommendation:
- What's your CAC **by channel** (not blended)?
- What's the payback period on your largest channel?
- What's your LTV:CAC ratio?
- What % of pipeline is marketing-sourced vs. sales-sourced?
- Where do your **best customers** (highest LTV, lowest churn) come from?
- What's your MQL → Opportunity conversion rate? (proxy for lead quality)
- Is this brand work or performance marketing? (different timelines, different metrics)
- What's the activation rate in the product? (PLG signal)
- If a prospect doesn't buy, why not? (win/loss data)
---
## CMO Metrics Dashboard
| Category | Metric | Healthy Target |
|----------|--------|---------------|
| **Pipeline** | Marketing-sourced pipeline % | 50–70% of total |
| **Pipeline** | Pipeline coverage ratio | 3–4x quarterly quota |
| **Pipeline** | MQL → Opportunity rate | > 15% |
| **Efficiency** | Blended CAC payback | < 18 months |
| **Efficiency** | LTV:CAC ratio | > 3:1 |
| **Efficiency** | Marketing % of total S&M spend | 30–50% |
| **Growth** | Brand search volume trend | ↑ QoQ |
| **Growth** | Win rate vs. primary competitor | > 50% |
| **Retention** | NPS (marketing-sourced cohort) | > 40 |
---
## Red Flags
- No defined ICP — "companies with 50-1000 employees" is not an ICP
- Marketing and sales disagree on what an MQL is (this is always a system problem, not a people problem)
- CAC tracked only as a blended number — channel-level CAC is non-negotiable
- Pipeline attribution is self-reported by sales reps, not CRM-timestamped
- CMO can't answer "what's our payback period?" without a 48-hour research project
- Brand work and performance marketing have no shared narrative — they're contradicting each other
- Marketing team is producing content with no documented positioning to anchor it
- Growth model was chosen because a competitor uses it, not because the product/ACV/ICP fits
---
## Integration with Other C-Suite Roles
| When... | CMO works with... | To... |
|---------|-------------------|-------|
| Pricing changes | CFO + CEO | Understand margin impact on positioning and messaging |
| Product launch | CPO + CTO | Define launch tier, GTM motion, messaging |
| Pipeline miss | CFO + CRO | Diagnose: volume problem, quality problem, or velocity problem |
| Category design | CEO | Secure multi-year organizational commitment to the narrative |
| New market entry | CEO + CFO | Validate ICP, budget, localization requirements |
| Sales misalignment | CRO | Align on MQL definition, SLA, and pipeline ownership |
| Hiring plan | CHRO | Define marketing headcount and skill profile by stage |
| Retention insights | CCO | Use expansion and churn data to sharpen ICP and messaging |
| Competitive threat | CEO + CRO | Coordinate battlecards, win/loss, repositioning response |
---
## Resources
- **References:** `references/brand_positioning.md`, `references/growth_frameworks.md`, `references/marketing_org.md`
- **Scripts:** `scripts/marketing_budget_modeler.py`, `scripts/growth_model_simulator.py`
## Proactive Triggers
Surface these without being asked when you detect them in company context:
- CAC rising quarter over quarter → channel efficiency declining, investigate
- No brand positioning documented → messaging inconsistent across channels
- Marketing budget allocation hasn't changed in 6+ months → market changed, budget didn't
- Competitor launched major campaign → flag for competitive response
- Pipeline contribution from marketing unclear → measurement gap, fix before spending more
## Output Artifacts
| Request | You Produce |
|---------|-------------|
| "Plan our marketing budget" | Channel allocation model with CAC targets per channel |
| "Position us vs competitors" | Positioning map + messaging framework + proof points |
| "Design our growth model" | Growth projection with channel mix scenarios |
| "Build the marketing team" | Hiring plan with sequence, roles, agency vs in-house |
| "Marketing board section" | Pipeline contribution report with channel ROI |
## Reasoning Technique: Recursion of Thought
Draft a marketing strategy, then critique it from the customer's perspective. Refine based on the critique. Repeat until the strategy survives scrutiny.
## Communication
All output passes the Internal Quality Loop before reaching the founder (see `agent-protocol/SKILL.md`).
- Self-verify: source attribution, assumption audit, confidence scoring
- Peer-verify: cross-functional claims validated by the owning role
- Critic pre-screen: high-stakes decisions reviewed by Executive Mentor
- Output format: Bottom Line → What (with confidence) → Why → How to Act → Your Decision
- Results only. Every finding tagged: 🟢 verified, 🟡 medium, 🔴 assumed.
## Context Integration
- **Always** read `company-context.md` before responding (if it exists)
- **During board meetings:** Use only your own analysis in Phase 2 (no cross-pollination)
- **Invocation:** You can request input from other roles: `[INVOKE:role|question]`
FILE:references/brand_positioning.md
# Brand Positioning Reference
Practical frameworks for defining, communicating, and defending your market position. Not theory — applied tools for CMOs who need to get this right.
---
## 1. Category Design Frameworks
### The Category Design Principle
Every product exists in a category — either one you define or one someone else defined. If you're not designing your category, your competitors are designing it for you, and they'll design it to exclude you.
**Category design is not renaming an existing category.** It's declaring that the existing category no longer solves the problem adequately, and that a new category — which you happen to lead — is required.
### The Three-Act Category Design Narrative
**Act 1: Name the problem**
Identify a problem that's real, growing, and underserved. Not a problem you invented — a problem your best customers articulate before they've heard your pitch.
> "Enterprise software teams are deploying faster than ever, but their security reviews still take 3 weeks — because security was built for a world where deployments happen monthly, not hourly."
**Act 2: Define the new category**
Name the category in terms of the outcome, not the feature. The category name should describe what customers achieve, not what the product does.
> "Continuous security" — not "automated security scanning" or "DevSecOps platform."
**Act 3: Position yourself as the category leader**
You can't just claim leadership — you need proof: customers, analysts, community, content, events. Leadership is built, not declared.
> "Snyk is building the continuous security category. 1.2M developers have adopted Snyk. Gartner lists us as a Cool Vendor in AppSec."
### When Category Design Works
| Condition | Explanation |
|-----------|-------------|
| Market timing | The problem is growing but the existing category is inadequate |
| CEO commitment | Category design is a 3-5 year initiative, not a marketing campaign |
| Analyst alignment | Gartner, Forrester, or G2 need to recognize your category |
| Community | Practitioners adopt the vocabulary before buyers do |
| Content moat | You publish the defining content for the category before competitors |
### Category Design Pitfalls
- **Naming the category after yourself:** "The [Your Company] Category" is not a category. It's a vanity.
- **Categories that don't solve analyst definitions:** If Gartner doesn't have a Magic Quadrant for your category, you're fighting uphill.
- **Jargon without adoption:** If your category name requires a two-paragraph explanation, it won't stick.
- **Starting a category war you can't win:** If an incumbent can copy your category name and launch in 90 days, you don't have a defensible category.
### The Lightning Strike Strategy
Category design requires concentrated, coordinated effort — not slow drip. Execute these simultaneously:
1. **Major piece of research or data** (the "State of X" report)
2. **Category-defining event** (host it, don't just attend)
3. **Analyst briefing** (educate Gartner/Forrester on the category before they define it themselves)
4. **Book or manifesto** (long-form content that becomes the category Bible)
5. **Community formation** (a Slack group, a conference, a certification that practitioners want)
Do all five within a 3-month window. This creates gravity around your category claim.
---
## 2. Messaging Architecture
### The Messaging Hierarchy
Every piece of content — from a tweet to a 60-page whitepaper — should trace back to this hierarchy. When it doesn't, you have messaging drift.
```
Level 1: Brand Promise
"[Company] [verb] [outcome] for [audience]"
→ Doesn't change. This is the north star.
Level 2: Positioning Statement (internal)
For [target customer] who [has this problem],
[Company] is the [market category] that [differentiated capability]
unlike [alternatives], [Company] [proof of differentiation].
Level 3: Value Propositions (3-4 max, one per key outcome)
Each VP: headline (5-8 words) + 2-3 sentence explanation + proof point
Level 4: Proof Points
Data, case studies, certifications, analyst recognition — evidence for each VP
Level 5: Channel Adaptations
Website copy, sales deck, ad copy, email — same hierarchy, different format
```
### Writing a Positioning Statement
The Geoffrey Moore / April Dunford format is still the best framework:
**Template:**
```
For [specific target customer]
who [has this specific, painful problem],
[Company name] is the [market category]
that [key differentiated capability].
Unlike [primary alternatives],
[Company] [proof of differentiation — something measurable or unique].
```
**Bad example (too generic):**
> For B2B companies who want to grow faster, Acme is the marketing platform that helps you get more leads. Unlike other platforms, Acme is easy to use and powerful.
**Good example (specific and falsifiable):**
> For DevOps teams in regulated industries who spend 20% of their sprint cycles on compliance reviews, Acme is the compliance automation platform that embeds regulatory checks directly into the CI/CD pipeline. Unlike manual compliance tools that create a separate review queue, Acme's policy-as-code approach reduces compliance-related cycle time by 60% without slowing deployments.
**Test your positioning statement:**
1. Can a competitor say the exact same thing? (If yes, it's not differentiated)
2. Does it describe what you do or what the customer gets? (Should be the latter)
3. Would your best customer say "yes, that's exactly my problem"? (If not, wrong ICP)
4. Is it falsifiable? (Claims you can't prove are liabilities)
### Value Proposition Development
**Structure for each VP:**
| Element | Description | Example |
|---------|-------------|---------|
| Outcome headline | What changes for the customer (5-8 words) | "Ship features 3x faster" |
| The problem | Why this matters now (1 sentence) | "Compliance reviews block 40% of releases in regulated industries" |
| Our approach | How we solve it differently (1-2 sentences) | "Policy-as-code embeds checks in the pipeline instead of adding a gate at the end" |
| Proof | Evidence this is real (1 sentence + data point) | "Customers reduce compliance cycle time by 60% in the first 90 days" |
**3-VP Architecture is the standard:**
- VP1: Core outcome (what most customers primarily buy for)
- VP2: Secondary benefit (makes the decision easier or stickier)
- VP3: Differentiator (what tips competitive decisions in your favor)
### Proof Point Hierarchy
Not all proof is equal. When you make a claim, match the strength of your proof to the importance of the claim.
| Proof Type | Strength | Best Used For |
|------------|---------|--------------|
| Third-party data (analyst report, research) | Highest | Category claims, market size |
| Customer ROI data with name | High | Value propositions |
| Customer quote with name and company | Medium-high | Specific pain points and outcomes |
| Aggregated customer data ("customers report…") | Medium | Directional claims |
| Internal testing or benchmark | Medium-low | Product capability claims |
| "Designed to…" or "built for…" | Low | Product direction only |
| "We believe…" or "we think…" | Lowest | Vision statements only |
**Proof point development process:**
1. Write the claim you want to make
2. Identify the strongest available proof
3. If proof is weak, either soften the claim or invest in getting better proof
4. Never publish a claim without knowing what happens when a skeptic asks "prove it"
---
## 3. Competitive Positioning Maps
### The Two-Axis Map
Choose two dimensions that:
1. Both matter to your target buyer
2. Create clear differentiation between you and competitors
3. You can credibly defend
**Choosing the axes:**
- Axis 1 should show a dimension where you win and most competitors cluster on the wrong side
- Axis 2 should show a dimension buyers care about deeply (ease, speed, breadth, price, compliance, etc.)
**What to avoid:**
- "Quality" vs. "Price" — too generic, every company claims the top-left
- Dimensions your competitors can match in one release cycle
- Dimensions that only your product team understands, not buyers
### Competitive Analysis Template
For each major competitor:
**Company:** _______________
| Dimension | What They Claim | What Customers Actually Experience | Gap |
|-----------|----------------|-----------------------------------|-----|
| Positioning | | | |
| Primary differentiator | | | |
| Pricing | | | |
| Ideal customer | | | |
| Weakness (win/loss data) | | | |
| What they say about you | | | |
**Sources for competitive intelligence:**
- Win/loss interviews (primary source — nothing beats this)
- G2/Capterra reviews (what customers say publicly)
- Glassdoor (tells you about internal culture and focus)
- LinkedIn job postings (what they're building next)
- Their pricing page changes (what they're competing on)
- Conference talks from their product and sales leaders
### Battlecard Format
One page per competitor. Used by sales, not marketing.
```
COMPETING AGAINST: [Competitor Name]
WHY CUSTOMERS CONSIDER THEM:
(2-3 bullets — be honest about their appeal)
OUR DIFFERENTIATION:
(2-3 bullets — factual, not marketing language)
THE LANDMINE QUESTION:
(One question that exposes their weakness. The answer should make the buyer uncomfortable choosing them.)
Example: "How long does your typical implementation take? And what's your SLA if it runs over?"
OUR PROOF POINTS IN THIS COMPARISON:
- [Customer name] switched from [competitor] after [specific reason], saw [specific result]
- [Data point that directly contradicts competitor's primary claim]
THEIR LIKELY COUNTER-MOVES:
(What will they say about us? How do we respond?)
WHEN TO WALK AWAY:
(If the prospect values X more than Y, we are not the right fit — say so)
```
---
## 4. Brand Voice Development
### What Brand Voice Is (and Isn't)
**Brand voice is NOT:**
- A list of adjectives ("we are professional, innovative, and customer-focused")
- The tone you use in formal communications
- The font and color palette (that's visual identity)
**Brand voice IS:**
- How the company sounds across every written touchpoint
- Consistent enough to be recognizable, flexible enough to be human
- Grounded in what your best customers actually value
### The Voice Attribute Framework
Define 3-4 voice attributes. For each:
1. **What it means** (in one sentence)
2. **What it sounds like** (one example)
3. **What it doesn't mean** (the common mistake that goes wrong)
**Example:**
| Attribute | Means | Sounds like | Doesn't mean |
|-----------|-------|------------|--------------|
| Direct | We say what we mean without hedging | "Your compliance review takes 3 weeks. It shouldn't." | Blunt, rude, or dismissive |
| Expert | We speak from depth, not from trend | "Here's why most security gates fail at scale, and what actually works." | Jargon-heavy or condescending |
| Honest | We acknowledge what we don't do | "We're not the best fit if you need a one-size-fits-all platform." | Self-deprecating or uncertain |
| Human | Real people write for real people | "Deploying on a Friday? Here's what we'd check first." | Casual, unprofessional |
### Voice Consistency Testing
Take a random sample of 10 recent pieces of content:
- Website homepage and pricing page
- 3 blog posts from different authors
- 5 outbound emails from sales
- 3 social posts
- 1 press release
Score each on: Does this sound like us? (1-5)
Average < 3: You have a brand voice problem. The cause is usually no documented guidelines, or guidelines that exist but aren't enforced.
### Voice in Different Contexts
The attribute stays the same. The tone adjusts.
| Context | Tone adjustment | Example of "Direct" |
|---------|----------------|---------------------|
| Homepage | Confident | "Compliance reviews don't have to slow you down." |
| Technical docs | Precise | "Set the policy threshold to 0.95 to enforce mandatory approval." |
| Error messages | Helpful | "That didn't work. Here's the most common reason why, and how to fix it." |
| Support | Empathetic | "That's frustrating. Here's what happened and what we're doing about it." |
| Sales outreach | Respectful | "Most teams in your space have this problem. Worth 20 minutes to explore?" |
---
## 5. Rebrand Decision Framework
### When Rebrands Succeed vs. Fail
**Successful rebrands:**
- Driven by a genuine strategic shift (new category, new ICP, new market)
- Have internal alignment before external launch
- Are accompanied by product and messaging changes — not just visual
- Have a 6-12 month transition plan for existing customers
**Failed rebrands:**
- Driven by internal boredom with the old brand
- Executed as a "refresh" without repositioning the value proposition
- Lack leadership conviction (executives still describe the company in the old terms)
- Launch with a new logo but same product, same messaging, same ICP
### The Rebrand Decision Matrix
Answer each question. More "yes" answers = more likely rebrand is warranted.
| Question | Yes | No |
|----------|-----|-----|
| Has our ICP changed significantly in the last 18 months? | Rebrand | Stay |
| Are we entering a new market where the current brand creates friction? | Rebrand | Stay |
| Does the brand name have negative associations in the market? | Rebrand | Stay |
| Has an acquisition changed our core identity? | Rebrand | Stay |
| Is the current brand actively hurting sales conversations? (evidence required) | Rebrand | Stay |
| Are we bored with the brand? | Stay | — |
| Did leadership change? | Stay | — |
| Are competitors rebranding? | Stay | — |
Score: 3+ "Rebrand" answers with evidence = worth a serious evaluation.
### Rebrand Risk Assessment
**Name change** is the highest-risk rebrand element. Before committing:
- Legal: trademark availability in all target markets
- SEO: 18-24 months to recover domain authority after a domain change
- Customer: existing customers need to update all integrations, contracts, documentation
- Analyst: re-education of Gartner, Forrester, G2 category definitions
- Employee: company identity shift is a culture event, not just an HR task
**Minimum viable rebrand (lower risk):**
1. New positioning and messaging (always worth doing if positioning is wrong)
2. Visual identity refresh (keep the name, update the look)
3. Tagline change (the cheapest, lowest-risk brand change)
**Full rebrand (high risk, sometimes necessary):**
1. New company name and domain
2. New visual identity
3. New positioning and messaging
4. New category narrative
### Rebrand Execution Checklist
**Pre-launch (90 days):**
- [ ] Finalize positioning before finalizing design (in that order)
- [ ] Legal trademark clearance in all target markets
- [ ] Domain secured (with redirects planned)
- [ ] Internal alignment: every leader can describe the new positioning in one sentence
- [ ] Customer comms plan (existing customers, especially enterprise, need advance notice)
- [ ] Analyst briefings scheduled (Gartner, Forrester — brief them before launch)
- [ ] PR plan finalized
**Launch (day 1):**
- [ ] Website flipped
- [ ] Social profiles updated
- [ ] Email signatures updated company-wide
- [ ] Sales deck updated
- [ ] Press release published
- [ ] Existing customers notified (email from CEO or CMO, not marketing automation)
**Post-launch (90 days):**
- [ ] SEO monitoring (watch for ranking drops on key terms)
- [ ] Win rate monitoring (did conversion change?)
- [ ] Employee feedback (are they using the new messaging correctly?)
- [ ] Partner/channel update (resellers, integrations, directories)
- [ ] Analyst follow-up (did they update their reports?)
---
## Quick Reference: Brand Positioning Diagnostic
Use this as an audit against your current positioning:
| Check | Pass | Fail |
|-------|------|------|
| Can every sales rep state the positioning in one sentence without looking it up? | ✓ | Positioning isn't working |
| Is the ICP specific enough to disqualify companies? | ✓ | ICP is too broad |
| Does the homepage lead with customer outcome, not product features? | ✓ | Copy needs rewrite |
| Can you name 3 companies you're NOT a good fit for? | ✓ | Positioning is unfocused |
| Do win/loss interviews confirm the stated differentiator? | ✓ | Differentiator is assumed, not proven |
| Is the category name used by analysts or industry media? | ✓ | Category design needed |
| Does every piece of content trace back to a VP from the hierarchy? | ✓ | Messaging drift — need guidelines |
FILE:references/growth_frameworks.md
# Growth Frameworks Reference
Playbooks for PLG, sales-led, community-led, and hybrid growth models. Includes growth loops, funnel design, and guidance on when and how to switch models.
---
## 1. Product-Led Growth (PLG) Playbook
### What PLG Actually Is
PLG means the product is the primary distribution mechanism. Not "we have a free trial." Not "our product is self-serve." PLG means the product creates acquisition, retention, and expansion — and does so at a scale and cost no sales team can match.
**The minimum requirements for PLG to work:**
1. **Fast time-to-value:** Users must get a meaningful outcome within one session (ideally < 30 minutes)
2. **Low friction to start:** No sales call, no implementation project, no credit card required (for top of funnel)
3. **Built-in virality or network effects:** Usage creates exposure or value that draws in other users
4. **Self-serve monetization or expansion path:** Freemium → paid, or individual → team → company
If any of these is missing, you don't have PLG — you have a website with a free trial.
### PLG Funnel: The Four Stages
**Stage 1: Acquisition**
The user discovers and signs up for the product without talking to sales.
Key channels:
- Organic search (SEO targeting jobs-to-be-done searches)
- Product hunt launches
- Referral and invite loops (users share the product with colleagues)
- Developer communities and open-source contributions
Metric: Visitor-to-signup rate
Benchmark: 2-8% for B2B SaaS (varies heavily by product complexity)
**Stage 2: Activation**
The user reaches the "aha moment" — the point where the product delivers its core value for the first time.
Finding the aha moment:
- Look at the behaviors that differentiate users who stay from users who churn in the first 30 days
- The aha moment is not creating an account. It's completing the first outcome.
- For Slack: sending a message in a real channel
- For Dropbox: adding a file from a second device
- For HubSpot: publishing a form that captures a real lead
Metric: Activation rate (% of signups who complete the aha moment action within 7 days)
Benchmark: 25-40% is strong. < 15% means the onboarding is broken.
**Stage 3: Retention**
Users return to the product and build habitual use.
Retention analysis:
- Cohort retention curves (by signup week/month)
- Day 1, Day 7, Day 30, Day 90 retention rates
- Feature adoption by retained vs. churned users (which features predict retention?)
Metric: D30 retention rate (% of users still active 30 days after signup)
Benchmark: > 40% D30 retention is strong for B2B products
**Stage 4: Revenue**
Self-serve conversion from free to paid, or expansion from individual to team.
PQL (Product-Qualified Lead) signals:
- Reached a usage limit (invites, storage, seats)
- Used a premium feature in trial mode
- Team size on the account reached a threshold
- High-frequency usage above a defined threshold
Metric: PQL conversion rate (% of PQLs who convert to paid within 30 days)
Benchmark: 15-30% for well-designed PLG products
### PLG Expansion Model
PLG growth compounds through account expansion:
```
Individual user discovers product
→ Gets value, invites teammates
→ Team adopts product
→ Becomes department-wide
→ Finance/IT gets involved
→ Enterprise contract
```
This is "bottom-up" enterprise: individual adoption precedes company-wide purchase. It's also the most defensible moat — when every engineer in the company uses your product individually, procurement cancellation is very hard.
**Expansion levers:**
- Seat-based pricing (more users = more revenue, aligned with value)
- Usage-based pricing (more usage = more value = more revenue)
- Feature gating (team/enterprise features visible but gated, creating pull to upgrade)
- Admin discovery (usage reports surface to managers who didn't know they had a product champion)
### PLG Diagnostic
| Question | Healthy | Unhealthy |
|----------|---------|-----------|
| Time-to-value | < 30 minutes | > 2 hours |
| Activation rate | > 30% | < 15% |
| D30 retention | > 40% | < 20% |
| PQL conversion | > 15% | < 5% |
| NPS from self-serve users | > 40 | < 20 |
| Viral coefficient | > 0.3 | < 0.1 |
### PLG Team Structure
```
Head of Growth (often VP Product or VP Marketing)
├── Growth PM (owns activation and retention loops in product)
├── Growth Engineer (2-3 engineers dedicated to growth experiments)
├── Data Analyst (experimentation, funnel analysis, cohort reports)
└── Growth Marketer (acquisition, SEO, referral programs)
```
The growth team sits between product and marketing. This is intentional — they own the product loops that drive acquisition and retention.
---
## 2. Sales-Led Growth (SLG) Model
### The SLG System
In SLG, marketing's job is to fill the sales pipeline. Sales converts it. The system only works if marketing and sales agree on definitions, SLAs, and shared metrics.
**The SLG funnel:**
```
Awareness (Impressions, reach, brand search)
↓
Lead (Name + contact info captured)
↓
MQL — Marketing Qualified Lead (meets ICP criteria, intent signal detected)
↓ [Marketing → Sales handoff]
SAL — Sales Accepted Lead (sales reviews and accepts the lead)
↓
SQL — Sales Qualified Lead (sales confirms budget, authority, need, timeline)
↓
Opportunity (Formal deal in pipeline, has a close date)
↓
Closed-Won
```
**The MQL definition problem:**
Most marketing-sales friction traces to an unclear MQL definition. The MQL should be:
- ICP-matched (company size, industry, role)
- Intent-signaled (visited pricing page, attended webinar, downloaded high-intent content)
- Not just email address + "subscribed to newsletter"
**A concrete MQL definition:**
> Company 50-500 employees, B2B SaaS, role is VP Engineering or CTO or CISO, AND has performed 2+ of: attended webinar, visited pricing page, requested demo, downloaded security report, attended event.
This definition makes the MQL useful. If you can't score it in your CRM without human judgment, it's not a definition — it's a guideline.
### SLG Conversion Rate Benchmarks
| Stage | Average B2B SaaS | Top Quartile |
|-------|-----------------|--------------|
| Lead → MQL | 5-15% | > 20% |
| MQL → SAL | 50-70% | > 75% |
| SAL → SQL | 30-50% | > 60% |
| SQL → Opportunity | 60-80% | > 85% |
| Opportunity → Closed-Won | 20-30% | > 40% |
**End-to-end:** Lead → Closed-Won: 1-5% (wide range by ACV and ICP quality)
### Pipeline Coverage Mechanics
A healthy SLG pipeline has 3-4x coverage against quota.
If a sales rep has a $500K quarterly quota:
- They need $1.5M-$2M in active pipeline
- Pipeline must be distributed across stages (not all "prospecting")
- Stage distribution benchmark: 30% early, 40% mid, 30% late
Insufficient coverage (< 3x) is a lagging indicator of a miss — by the time coverage is low, it's too late to recover in the same quarter. Coverage should be tracked weekly.
### SLG Demand Generation Channels
**High-intent channels (bottom of funnel):**
- Paid search on buying-intent keywords (e.g., "[competitor] alternative", "best [category] software")
- Review site presence (G2, Capterra) — buyers use these before vendor websites
- Outbound SDR targeting specific accounts (ABM)
**Medium-intent channels (middle of funnel):**
- Webinars and virtual events (capture active learners)
- Gated content (guides, benchmarks, templates — ICP-specific)
- Retargeting to website visitors
**Awareness channels (top of funnel):**
- Content and SEO (captures people learning about the problem)
- Podcast sponsorships, industry media
- Conference sponsorship and speaking
- Paid social (LinkedIn for B2B)
### ABM (Account-Based Marketing) in SLG
ABM flips the funnel: instead of generating leads and filtering for good ones, you start with target accounts and run coordinated campaigns against them.
**Tiers:**
- **Tier 1 (1:1):** 5-20 strategic accounts, fully customized campaigns, dedicated SDR+AE pairs, executive outreach
- **Tier 2 (1:few):** 50-200 accounts, programmatic personalization, SDR sequences, targeted events
- **Tier 3 (1:many):** 500+ accounts, standard campaigns with light personalization
ABM requires tight sales/marketing alignment. If sales doesn't work the accounts marketing targets, ABM produces zero results.
---
## 3. Community-Led Growth (CLG)
### The CLG Thesis
Community-led growth works when:
1. Your buyers want to learn from peers, not vendors
2. There's a strong practitioner identity (developers, data teams, security, FinOps)
3. Your category is complex enough that buyers need education before purchasing
4. You can commit to building genuine community, not a marketing channel in disguise
**The fundamental rule of CLG:** The community must deliver value to members whether or not they ever buy your product. If the only purpose of the community is to sell to members, the community will die.
### CLG Stages
**Stage 1: Find the community**
The community often exists before you build it. Find where your practitioners already gather:
- Slack groups, Discord servers
- Subreddits and LinkedIn groups
- Conference hallways
- Open-source repositories
Before building, participate. Earn trust. Understand the conversations.
**Stage 2: Become the knowledge hub**
Establish your company as the best source of information on the category problem:
- Publish the benchmark study everyone references
- Host the conference that defines the industry
- Create the certification practitioners want on their resume
- Open-source the tools the community needs
**Stage 3: Build the platform**
Create a dedicated community space (Slack, Discord, forum):
- Community must be practitioner-first, not vendor-first
- Community managers who genuinely care about member value
- Content from members, not just from your company
- Events that build member relationships, not just product demos
**Stage 4: Convert community to customers**
Community members who become customers do so because they trust you, not because you sold them. Conversion paths:
- Community members see peer success with your product
- Product-qualified signals from community members who trial the product
- Direct outreach from sales to active community members (with permission and context)
- Enterprise deals from companies whose employees are active in the community
### CLG Metrics
| Metric | Definition | Health Signal |
|--------|-----------|--------------|
| Monthly active members | Members who post, comment, or engage | > 15% of total members |
| Community-sourced pipeline | $ pipeline where community was first touch | Track and trend |
| Community-influenced pipeline | $ pipeline with any community touchpoint | > 30% of total pipeline |
| NPS of community members vs. non-members | Loyalty difference | Community members should score 20+ pts higher |
| Member-generated content % | % of content posted by non-employees | > 60% is healthy community |
| Time from community join to product trial | | Shortens as community matures |
### CLG Anti-Patterns
- **Community as a newsletter:** If members can't interact with each other, it's not a community — it's a list.
- **Product launches in the community:** Nothing kills community trust faster than using it for sales announcements.
- **Community without a community manager:** Communities left to run themselves become ghost towns or become toxic.
- **Measuring community by member count:** Ghost members are noise. Active engagement is signal.
---
## 4. Hybrid Growth Models
### PLG + SLG ("Product-Led Sales" or PLS)
The most common hybrid at growth stage. PLG handles SMB self-serve; sales closes enterprise.
**The PQL-to-sales handoff:**
Define the triggers that move a product-qualified lead to a sales-assisted motion:
- Company has > X users (e.g., 10+ users on a team account)
- Usage exceeds Y threshold in 30 days
- Account is a named target in the ABM list
- User explicitly requested a demo or upgrade assistance
**The risk:** Sales team ignores PLG pipeline because deal size is smaller. Fix: separate quotas and commission structures for self-serve expansion vs. new enterprise logos.
**The opportunity:** PLG creates pre-qualified champions inside accounts. Sales doesn't have to create interest — they convert it. Win rates in PLS motions are typically 30-50% higher than cold outbound.
### SLG + CLG
Community builds brand and generates inbound pipeline for sales.
This hybrid works when:
- Sales cycles are long (6-18 months)
- Buyers do extensive research before engaging with vendors
- The community validates your credibility before sales conversations begin
**The integration:**
- Community team feeds content insights to demand gen
- Event attendees become high-priority SDR sequences
- Active community members get dedicated AE outreach with community context
- Win/loss analysis includes community touchpoints
### PLG + CLG
The developer/open-source hybrid. PLG handles product adoption; community handles advocacy and content.
**Examples:** HashiCorp (Terraform community + enterprise sales), Elastic (open-source + community + commercial), Tailscale (developer community + self-serve + enterprise).
**How it compounds:**
```
Community member learns from community content
→ Discovers open-source or free tier
→ Gets value in first session
→ Shares experience in community
→ New members discover product through community content
```
---
## 5. Growth Loops vs. Funnels
### The Difference
**A funnel** is linear. It requires constant input at the top to produce output at the bottom. If you stop feeding it, it stops producing.
**A growth loop** is cyclical. Output from one stage becomes input to the next. The system compounds.
### Common Growth Loops
**Viral loop:**
```
User gets value → Invites colleague → Colleague signs up →
Colleague invites another colleague → ...
```
Viral coefficient (K) = (Average invites per user) × (Conversion rate of invites)
- K > 1: Exponential growth (rare)
- K 0.5-1: Strong viral assist
- K < 0.3: Viral is not a meaningful growth driver
**Content SEO loop:**
```
Publish content on [topic] → Ranks in search →
Drives signups → Users share content → Builds backlinks →
Better rankings → More content is possible
```
This loop takes 12-24 months to activate but is extraordinarily defensible once running.
**UGC (User-Generated Content) loop:**
```
Users share their work publicly (templates, analyses, portfolios) →
Others discover the work → They find the product →
They create and share their own work → ...
```
Figma, Notion, Airtable, Canva — all run this loop.
**Data network effect loop:**
```
More users → More data → Better product →
More users attracted → ...
```
LinkedIn, Waze, Duolingo — accuracy or relevance improves as the user base grows.
**Integration loop:**
```
Product integrates with X → X's users discover your product →
More integrations possible → More discovery surfaces → ...
```
Zapier, Slack apps, Salesforce AppExchange — being in the ecosystem creates distribution.
### Building a Growth Loop
**Step 1: Map the current funnel**
Where do customers come from? What are the conversion steps?
**Step 2: Find the output**
What does a successful customer produce?
- Invite emails
- Shared content
- Public work visible to others
- Reviews or testimonials
**Step 3: Design the loop**
How does that output become tomorrow's input to acquisition?
- If they share → is there a landing page that captures the new visitor?
- If they invite → is the invite experience friction-free?
- If they create content → does it rank in search or appear in relevant communities?
**Step 4: Measure loop velocity**
For each loop, measure:
- Cycle time: How long does one full cycle take?
- Conversion at each step: Where does the loop break down?
- Loop coefficient: How many new users does one existing user generate?
---
## 6. When to Switch Growth Models
### The Warning Signs
**PLG-to-SLG triggers:**
- Enterprise accounts are signing up via PLG but aren't expanding without human intervention
- Average deal sizes in enterprise are 10-20x SMB, and you're leaving revenue on the table
- Product adoption in enterprise requires configuration or integration that needs support
- PLG accounts churn at higher rates than sales-assisted accounts
**SLG-to-PLG/PLS triggers:**
- CAC is increasing year-over-year as competition for sales talent intensifies
- Smaller competitors are winning deals with self-serve
- Customers are asking "can I just try this myself?"
- ACV is declining as the market matures and products commoditize
- Sales team efficiency (revenue per sales rep) is declining
**Adding CLG to existing motion:**
- Sales cycles are long and trust is the primary barrier
- SEO and content are generating traffic but low conversion (awareness without trust)
- Competitors are building community and you're not present
- Customer success teams report that customers who participate in user groups retain better
### The Transition Playbook
**Phase 1: Prove it before scaling (months 1-6)**
Don't restructure the team to support the new model before proving it works.
- Run a pilot: 3-5 SDRs testing PLG signals as outreach triggers (for PLG → PLS)
- Or: Launch a beta community with 100 core customers (for adding CLG)
- Measure the metrics of the new model, compare to current model
**Phase 2: Parallel running (months 6-12)**
Run both models simultaneously. Don't kill the current model while building the new one.
- Set clear boundaries on which accounts go to which motion
- Build dedicated teams for each model (don't ask the same people to do both)
- Define success metrics for the new model independently
**Phase 3: Rebalance (months 12-18)**
Once the new model proves its unit economics:
- Shift headcount and budget to the more efficient model
- Keep the old model for the segments where it still works
- Document what the new model requires to sustain itself
**The anti-pattern:** Announcing a model shift without proof, restructuring the team, and discovering after 12 months that the new model doesn't work. By then, the old model's momentum is gone and you've burned a year.
### Growth Model Maturity Matrix
| Dimension | PLG | SLG | CLG |
|-----------|-----|-----|-----|
| Time to first results | 3-6 months | 1-3 months | 12-18 months |
| Requires up-front product investment | High | Low | Medium |
| Scales without linear headcount | Yes | No | Yes |
| Predictable pipeline | Low (early) | High | Low (early) |
| CAC trend over time | Decreases | Flat/increases | Decreases |
| Works for ACV > $50K | Only with SLG assist | Yes | Yes |
| Works for ACV < $5K | Yes | No | Only with PLG |
| Defensibility once established | High | Low | Very high |
FILE:references/marketing_org.md
# Marketing Org Reference
Team structure, hiring sequence, agency decisions, marketing ops, and cross-functional alignment — by company stage.
---
## 1. Marketing Team Structure by Stage
### Pre-Seed / Seed (< $1M ARR, 1–10 people)
Don't hire a marketing team yet. The founders are the marketing team.
What to do instead:
- Founders write content, do sales calls, go to events
- The goal is learning the ICP and finding the channel that works, not scaling anything
- One contractor or agency for specific output (design, SEO audit) is fine
First marketing hire trigger: You have a repeatable sales motion and need to scale it.
---
### Series A ($1M–$5M ARR, 10–30 people)
**Org:**
```
Founding Marketer (Head of Marketing or VP Marketing)
```
One person. Generalist. Capable of writing, running ads, setting up HubSpot, producing a report. Their job is to find what works.
**What they own:**
- Content and SEO foundation
- Paid channel experiments
- Sales enablement basics (1-pager, deck, email sequences)
- Event presence (1-2 conferences)
- Marketing attribution setup (get this right early)
**What they don't own yet:**
- Brand redesign
- Analyst relations
- Partner marketing
- Field marketing team
**CMO vs. VP Marketing at this stage:** VP Marketing. An experienced operator who can build and execute. A CMO's strategic value isn't fully leveraged until there's a team to lead and a budget to allocate.
---
### Series B ($5M–$20M ARR, 30–80 people)
**PLG-first org:**
```
VP Marketing
├── Growth Marketing (acquisition loops, activation, PLG analytics)
├── Product Marketing (positioning, launch, sales enablement)
└── Content & SEO (organic engine)
```
**SLG-first org:**
```
VP Marketing
├── Demand Generation (pipeline creation, paid, digital)
├── Product Marketing (positioning, competitive intel, enablement)
├── Field Marketing (events, regional, ABM)
└── Marketing Operations (CRM, attribution, reporting)
```
**Community-led org:**
```
VP Marketing
├── Community & Developer Relations
├── Content & SEO
└── Product Marketing
```
**At this stage:** Marketing ops becomes critical. Without it, attribution is guesswork and the sales team blames marketing for bad leads.
---
### Series C ($20M–$75M ARR, 80–200 people)
```
CMO
├── Demand Generation
│ ├── Paid Media
│ ├── SEO & Content
│ └── Marketing Operations
├── Product Marketing
│ ├── Core PMMs (by product line or segment)
│ └── Competitive Intelligence
├── Field Marketing
│ ├── Events
│ └── Regional / ABM
└── Brand & Communications
├── Brand Design
└── PR / Analyst Relations
```
**At this stage:**
- The CMO is a board-level communicator, not a campaign manager
- Each function has a dedicated leader (director or VP level)
- Marketing ops owns the attribution model and reports to CMO directly
- Analyst relations becomes important (Gartner, Forrester, G2 category positioning)
---
### Growth Stage ($75M+ ARR)
Marketing becomes a portfolio of specialized functions. Each major channel has a team. Brand is a serious investment. Analyst relations is a dedicated role. International marketing teams form.
The CMO's job shifts from building the machine to:
- Setting marketing strategy across a complex portfolio
- Representing marketing at the board level
- Owning brand and category leadership
- Cross-functional leadership with CRO, CPO, CEO
---
## 2. Hiring Sequence
### Who to Hire First
**The generalist content + demand gen marketer.**
Must-haves:
- Can write (blog posts, emails, landing pages — not just briefs)
- Can run paid campaigns (Google, LinkedIn — not just "I've managed agencies")
- Can operate a marketing automation platform (HubSpot, Marketo)
- Comfortable with data (can build a funnel report without asking an analyst)
This person builds the foundation. They're not a specialist yet — they're testing channels and building the process.
Avoid: Hiring a brand designer first. Or a community manager. Or a social media manager. These are specialties that compound on a foundation that doesn't exist yet.
### Who to Hire Second
**A specialist in the channel that's working.**
If organic search is your top lead source → hire an SEO/content lead.
If events are driving pipeline → hire a field marketer.
If outbound is working → hire an SDR manager or demand gen specialist.
Don't hire a generalist #2. By now you know what's working. Depth beats breadth.
### Who to Hire Third
**Product marketing.**
Why third and not first? Because PMM output (positioning, sales enablement, launch) is most valuable when there's an audience to position to and a sales team to enable. Before that, the founding marketer does "good enough" PMM work.
PMM hire profile: Has done positioning work before, has run a product launch, has built sales decks that sales actually uses, comfortable with win/loss analysis.
PMM:PM ratio benchmark: 1 PMM per 2–3 PMs. If you have 6 PMs and 1 PMM, you have a messaging and enablement problem.
### Who to Hire Fourth
**Marketing operations.**
This is consistently hired too late. By the time most companies hire marketing ops, attribution is broken, leads are being lost in handoffs, and the CRM data is unreliable. Hire marketing ops before you think you need it.
Marketing ops profile: HubSpot/Marketo certified, SQL capable, understands multi-touch attribution, has integrated CRM + sales engagement tools before.
### Hiring Decision Triggers
| Hire | Trigger |
|------|---------|
| Generalist marketer #1 | Sales motion is repeatable, need to scale lead generation |
| Specialist #2 | One channel is clearly outperforming — double down |
| Product marketer | Sales team is losing deals to positioning confusion or competitor gaps |
| Marketing ops | Running 3+ campaigns simultaneously with manual tracking |
| Field marketer | Events are in the strategy and attendance > 2 conferences/quarter |
| Head of Marketing / VP | Team is 3+ people and needs an org owner |
| CMO | Company is Series B/C and marketing needs board-level representation |
---
## 3. Agency vs. In-House
### Framework
Keep in-house what compounds. Outsource what's episodic or specialized.
| Function | Agency | In-House | Notes |
|----------|--------|----------|-------|
| Brand design | Early stage | Series B+ | Agency fine until redesigns become frequent |
| Paid media | < $50K/month spend | > $50K/month | Agency margin eats returns at scale |
| SEO strategy | Audit only | Ongoing execution | Strategy once, execution continuously |
| Content production | Overflow only | Core writers | Your voice must be yours |
| PR / comms | Almost always | $100M+ companies | Specialists required for media relationships |
| Marketing ops / CRM | Never | Always | This is your data infrastructure |
| Analyst relations | Initial strategy | Ongoing | Relationship-based — needs dedicated owner |
| Video / creative production | Always | Rarely | Episodic, specialized equipment |
### Agency Red Flags
- They want to own your ad accounts. (Always keep ownership. No exceptions.)
- SLA is "5 business days for creative requests." For a performance channel, that's too slow.
- Reporting is impressions, CPM, and "brand lift." Where's the pipeline?
- They can't tell you your CAC from their channel.
- They won't share the actual data — only their dashboard.
- Your account manager changes every 6 months.
### Agency Evaluation Criteria
1. **Proof of work in your category** — ask for 3 case studies with actual CAC and pipeline data
2. **Who actually does the work** — senior pitch team ≠ junior execution team
3. **Account ownership** — all accounts, pixels, analytics must be in your name
4. **Reporting cadence** — weekly data, monthly strategy, quarterly business review
5. **Exit terms** — how do you offboard without losing your data, accounts, and history?
---
## 4. Marketing Ops and Tech Stack
### The Minimum Viable Stack
| Layer | Tool | Purpose |
|-------|------|---------|
| CRM | HubSpot / Salesforce | Contact database, pipeline, source of truth |
| Marketing automation | HubSpot / Marketo / ActiveCampaign | Email, nurture, lead scoring |
| Analytics | Google Analytics 4 + Segment | Traffic, behavior, event tracking |
| Attribution | HubSpot / Attributer.io / Dreamdata | Multi-touch pipeline attribution |
| Paid | Google Ads + LinkedIn Ads | Performance channels |
| SEO | Ahrefs / Semrush | Keyword research, rank tracking |
| Chat/conversion | Intercom / Drift | In-product + website conversion |
**The integration that breaks most:** CRM ↔ Marketing automation ↔ Sales engagement. When these aren't synced properly, leads are lost, attribution is wrong, and marketing and sales fight about pipeline. Fix this first.
### Marketing Ops Ownership
Marketing ops must own:
- CRM data quality (field standardization, deduplication, routing)
- Lead scoring model (and quarterly review against conversion data)
- Attribution model (with documented assumptions)
- Campaign tracking (UTM governance — no UTM = no attribution)
- Tech stack evaluation and contracts
Marketing ops must NOT own:
- Strategy (they enable it, not set it)
- Content production
- Campaign creative
---
## 5. Cross-Functional Alignment
### Marketing + Sales
The most important cross-functional relationship in a SLG company. Where it breaks:
| Problem | Root Cause | Fix |
|---------|-----------|-----|
| "Marketing sends us bad leads" | MQL definition is unclear or wrong | Define MQL jointly, score against conversion data |
| "Sales doesn't follow up on leads" | No SLA, no consequence | Define SLA (e.g., 24-hour response), track in CRM |
| "Marketing doesn't understand what customers care about" | No win/loss sharing | Weekly call: sales shares 3 deal insights, marketing shares 3 content results |
| "We don't know what's working" | Attribution is broken | Marketing ops fixes attribution before next budget cycle |
**The SLA agreement (document this):**
- Marketing commits: X MQLs/week meeting defined criteria, 48-hour SLA from form fill to SDR outreach
- Sales commits: All MQLs contacted within 24 hours, disposition logged in CRM within 5 days
### Marketing + Product
Where it breaks and how to fix it:
| Problem | Fix |
|---------|-----|
| PMM learns about launches 2 weeks before ship | PMM joins the product planning process at the roadmap stage, not the sprint stage |
| Feature launches with no messaging | Launch tiers: Tier 1 (major, full launch), Tier 2 (minor, release notes + 1 post), Tier 3 (internal only) |
| Product doesn't use customer insights from marketing | Monthly session: PMM shares win/loss themes, competitive intel, ICP data |
| No feedback loop on messaging in-product | PMM owns in-product copy review, not just external comms |
### Marketing + Customer Success
Customer success is marketing's best source of truth:
- **ICP validation:** Which customers are expanding? Which are churning? This refines who you target.
- **Proof points:** CS-sourced case studies and testimonials outperform vendor-written content 3:1 in conversion.
- **Messaging test:** If CS is answering the same question 20 times, marketing hasn't explained it clearly enough.
- **Referral programs:** CS owns the relationship; marketing owns the mechanics. Design them together.
Cadence: Monthly meeting between CMO and VP/Head of CS. Agenda: retention trends, expansion patterns, at-risk customers, NPS themes.
FILE:scripts/growth_model_simulator.py
#!/usr/bin/env python3
"""
Growth Model Simulator
----------------------
Projects MRR growth across different growth models (PLG, sales-led, community-led,
hybrid) and shows the impact of channel mix changes on growth trajectory.
Usage:
python growth_model_simulator.py
Inputs (edit INPUTS section):
- Starting MRR and churn rate
- Current channel mix (% of new MRR from each source)
- Conversion rates per model
- Growth rate assumptions per channel
Outputs:
- 12-month MRR projection by growth model
- Channel mix impact analysis (what happens if you shift mix)
- Break-even months for each model
- Side-by-side comparison table
"""
from __future__ import annotations
import math
from dataclasses import dataclass, field
from typing import Dict, List, Optional, Tuple
# ---------------------------------------------------------------------------
# Data models
# ---------------------------------------------------------------------------
@dataclass
class ChannelSource:
name: str
pct_of_new_mrr: float # Current share of new MRR (0.0–1.0)
monthly_growth_rate: float # How fast this channel grows month-over-month
cac: float # CAC in dollars
payback_months: float # Months to recover CAC
@dataclass
class GrowthModel:
name: str
description: str
channel_mix: Dict[str, float] # channel name → % of new MRR
new_mrr_monthly_base: float # Starting new MRR/month from this model
monthly_acceleration: float # Acceleration factor (compounding)
avg_ltv_cac: float # Expected LTV:CAC at scale
months_to_steady_state: int # Months before model hits its natural growth rate
notes: List[str] = field(default_factory=list)
@dataclass
class MonthSnapshot:
month: int
mrr: float
new_mrr: float
churned_mrr: float
expansion_mrr: float
net_new_mrr: float
cumulative_cac_spend: float
@dataclass
class ModelProjection:
model: GrowthModel
snapshots: List[MonthSnapshot]
break_even_month: Optional[int] # Month when cumulative revenue > cumulative CAC
# ---------------------------------------------------------------------------
# INPUTS — edit these
# ---------------------------------------------------------------------------
STARTING_MRR = 85_000 # Current MRR ($)
MONTHLY_CHURN_RATE = 0.012 # Monthly churn rate (1.2% = ~14% annual)
EXPANSION_RATE = 0.008 # Monthly expansion MRR as % of existing MRR
GROSS_MARGIN = 0.75
SIMULATION_MONTHS = 18
# Channel sources (used to model mix shift scenarios)
CHANNELS: List[ChannelSource] = [
ChannelSource("Organic/SEO", pct_of_new_mrr=0.28, monthly_growth_rate=0.04, cac=1_800, payback_months=9),
ChannelSource("PLG Self-Serve", pct_of_new_mrr=0.15, monthly_growth_rate=0.08, cac=900, payback_months=5),
ChannelSource("Outbound SDR", pct_of_new_mrr=0.25, monthly_growth_rate=0.02, cac=5_100, payback_months=21),
ChannelSource("Paid Search", pct_of_new_mrr=0.15, monthly_growth_rate=0.01, cac=6_200, payback_months=26),
ChannelSource("Events/Field", pct_of_new_mrr=0.08, monthly_growth_rate=0.01, cac=9_800, payback_months=41),
ChannelSource("Partner/Channel", pct_of_new_mrr=0.09, monthly_growth_rate=0.05, cac=3_400, payback_months=14),
]
# Growth models to simulate
GROWTH_MODELS: List[GrowthModel] = [
GrowthModel(
name="Current Mix",
description="Baseline — maintain current channel allocation",
channel_mix={"Organic/SEO": 0.28, "PLG Self-Serve": 0.15, "Outbound SDR": 0.25,
"Paid Search": 0.15, "Events/Field": 0.08, "Partner/Channel": 0.09},
new_mrr_monthly_base=12_000,
monthly_acceleration=0.025,
avg_ltv_cac=3.2,
months_to_steady_state=3,
notes=["Baseline. No changes to channel mix."],
),
GrowthModel(
name="PLG-First",
description="Shift budget toward PLG self-serve and organic; reduce paid and outbound",
channel_mix={"Organic/SEO": 0.35, "PLG Self-Serve": 0.35, "Outbound SDR": 0.10,
"Paid Search": 0.08, "Events/Field": 0.04, "Partner/Channel": 0.08},
new_mrr_monthly_base=9_500, # Slower start — PLG takes time to activate
monthly_acceleration=0.048, # But compounds faster
avg_ltv_cac=5.8,
months_to_steady_state=6, # PLG loops take time to build
notes=[
"Lower new MRR in months 1-6 while PLG loops activate.",
"Acceleration compounds strongly after month 6.",
"Requires product investment in activation/onboarding.",
"Best fit if time-to-value < 30 min and viral coefficient > 0.3.",
],
),
GrowthModel(
name="Sales-Led Scale",
description="Double down on outbound SDR and field; optimize for enterprise ACV",
channel_mix={"Organic/SEO": 0.20, "PLG Self-Serve": 0.05, "Outbound SDR": 0.40,
"Paid Search": 0.15, "Events/Field": 0.15, "Partner/Channel": 0.05},
new_mrr_monthly_base=15_000, # Higher new MRR from enterprise ACV
monthly_acceleration=0.018, # Linear growth — headcount-constrained
avg_ltv_cac=2.8,
months_to_steady_state=2,
notes=[
"Fastest short-term new MRR if ACV > $30K.",
"Growth is linear — adds headcount to add pipeline.",
"CAC and payback worsen as SDR market tightens.",
"Requires sales capacity increase to sustain.",
],
),
GrowthModel(
name="Community-Led",
description="Invest in community and content; reduce paid; long-term brand play",
channel_mix={"Organic/SEO": 0.45, "PLG Self-Serve": 0.15, "Outbound SDR": 0.15,
"Paid Search": 0.05, "Events/Field": 0.10, "Partner/Channel": 0.10},
new_mrr_monthly_base=7_000, # Slowest start
monthly_acceleration=0.038,
avg_ltv_cac=4.5,
months_to_steady_state=9, # Community takes longest to activate
notes=[
"Lowest new MRR in months 1-9.",
"Community trust drives lower CAC and higher retention at scale.",
"Best for categories where buyers seek peer validation.",
"Requires dedicated community manager from day one.",
],
),
GrowthModel(
name="Hybrid PLS",
description="PLG self-serve for SMB + sales-assisted for enterprise (Product-Led Sales)",
channel_mix={"Organic/SEO": 0.30, "PLG Self-Serve": 0.28, "Outbound SDR": 0.22,
"Paid Search": 0.08, "Events/Field": 0.06, "Partner/Channel": 0.06},
new_mrr_monthly_base=11_000,
monthly_acceleration=0.035,
avg_ltv_cac=4.1,
months_to_steady_state=4,
notes=[
"PLG handles SMB; sales closes enterprise with PQL signals.",
"Requires clear PQL definition and SDR/PLG handoff process.",
"Best if you have a product with both bottom-up and top-down adoption.",
],
),
]
# ---------------------------------------------------------------------------
# Simulation engine
# ---------------------------------------------------------------------------
def simulate_model(model: GrowthModel, months: int) -> ModelProjection:
snapshots: List[MonthSnapshot] = []
mrr = STARTING_MRR
cumulative_cac = 0.0
cumulative_revenue = 0.0
break_even_month = None
for m in range(1, months + 1):
# Ramp up — new_mrr accelerates each month
if m <= model.months_to_steady_state:
# Ramp phase: linear ramp from 60% to 100% of base
ramp_factor = 0.6 + 0.4 * (m / model.months_to_steady_state)
else:
# Steady state: compound acceleration
months_past_ramp = m - model.months_to_steady_state
ramp_factor = 1.0 + model.monthly_acceleration * months_past_ramp
new_mrr = model.new_mrr_monthly_base * ramp_factor
churned_mrr = mrr * MONTHLY_CHURN_RATE
expansion_mrr = mrr * EXPANSION_RATE
net_new_mrr = new_mrr - churned_mrr + expansion_mrr
mrr = mrr + net_new_mrr
# CAC spend approximation: new_mrr / (avg_deal_mrr) * blended_cac
# Use weighted CAC from channel mix
weighted_cac = _weighted_cac(model.channel_mix)
avg_deal_mrr = 1_500 # Assumption: $1,500 average deal MRR
deals_this_month = new_mrr / avg_deal_mrr
cac_spend = deals_this_month * weighted_cac
cumulative_cac += cac_spend
cumulative_revenue += mrr * GROSS_MARGIN
if break_even_month is None and cumulative_revenue >= cumulative_cac:
break_even_month = m
snapshots.append(MonthSnapshot(
month=m,
mrr=mrr,
new_mrr=new_mrr,
churned_mrr=churned_mrr,
expansion_mrr=expansion_mrr,
net_new_mrr=net_new_mrr,
cumulative_cac_spend=cumulative_cac,
))
return ModelProjection(
model=model,
snapshots=snapshots,
break_even_month=break_even_month,
)
def _weighted_cac(channel_mix: Dict[str, float]) -> float:
channel_cac = {ch.name: ch.cac for ch in CHANNELS}
total = sum(
channel_mix.get(name, 0) * cac
for name, cac in channel_cac.items()
)
weight_sum = sum(channel_mix.values())
return total / weight_sum if weight_sum > 0 else 5_000
# ---------------------------------------------------------------------------
# Reporting
# ---------------------------------------------------------------------------
def fmt_mrr(n: float) -> str:
if n >= 1_000_000:
return f".3fM"
return f".1fK"
def fmt_currency(n: float) -> str:
if n >= 1_000_000:
return f".2fM"
if n >= 1_000:
return f".1fK"
return f".0f"
def print_header(title: str) -> None:
width = 78
print("\n" + "=" * width)
print(f" {title}")
print("=" * width)
def print_channel_overview() -> None:
print_header("Current Channel Mix")
print(f" Starting MRR: {fmt_mrr(STARTING_MRR)} | Monthly churn: {MONTHLY_CHURN_RATE:.1%} | Expansion: {EXPANSION_RATE:.1%}/mo")
print()
print(f" {'Channel':<22} {'% MRR':>7} {'CAC':>8} {'Payback':>9} {'Growth/mo':>10}")
print(" " + "-" * 60)
for ch in sorted(CHANNELS, key=lambda c: c.pct_of_new_mrr, reverse=True):
print(
f" {ch.name:<22} {ch.pct_of_new_mrr:>6.0%} "
f"{fmt_currency(ch.cac):>8} {ch.payback_months:>7.0f}mo "
f"{ch.monthly_growth_rate:>9.1%}"
)
def print_model_detail(proj: ModelProjection) -> None:
model = proj.model
print_header(f"Model: {model.name}")
print(f" {model.description}")
if model.notes:
print()
for note in model.notes:
print(f" • {note}")
print()
# Print monthly snapshot (every 3 months + final)
milestones = set(range(3, SIMULATION_MONTHS + 1, 3)) | {SIMULATION_MONTHS}
print(f" {'Month':<7} {'MRR':>10} {'New MRR':>9} {'Churned':>9} {'Expand':>8} {'Net New':>9}")
print(" " + "-" * 56)
for snap in proj.snapshots:
if snap.month in milestones:
print(
f" {snap.month:<7} {fmt_mrr(snap.mrr):>10} "
f"{fmt_mrr(snap.new_mrr):>9} {fmt_mrr(snap.churned_mrr):>9} "
f"{fmt_mrr(snap.expansion_mrr):>8} {fmt_mrr(snap.net_new_mrr):>9}"
)
final = proj.snapshots[-1]
growth_x = final.mrr / STARTING_MRR
arr_final = final.mrr * 12
weighted_cac = _weighted_cac(model.channel_mix)
be = f"Month {proj.break_even_month}" if proj.break_even_month else f"> {SIMULATION_MONTHS}mo"
print()
print(f" Final MRR ({SIMULATION_MONTHS}mo): {fmt_mrr(final.mrr)}")
print(f" Final ARR: {fmt_currency(arr_final)}")
print(f" Growth multiple: {growth_x:.1f}x from starting MRR")
print(f" Weighted blended CAC: {fmt_currency(weighted_cac)}")
print(f" Expected LTV:CAC: {model.avg_ltv_cac:.1f}x")
print(f" Months to steady state:{model.months_to_steady_state}")
print(f" CAC break-even: {be}")
def print_comparison_table(projections: List[ModelProjection]) -> None:
print_header(f"Growth Model Comparison — Month {SIMULATION_MONTHS} Outcomes")
header = (
f" {'Model':<20} {'MRR (final)':>12} {'ARR (final)':>12} "
f"{'Growth':>7} {'LTV:CAC':>8} {'Break-even':>11}"
)
print(header)
print(" " + "-" * 74)
for proj in sorted(projections, key=lambda p: p.snapshots[-1].mrr, reverse=True):
final = proj.snapshots[-1]
growth_x = final.mrr / STARTING_MRR
arr_final = final.mrr * 12
be = f"Mo {proj.break_even_month}" if proj.break_even_month else f">{SIMULATION_MONTHS}mo"
print(
f" {proj.model.name:<20} {fmt_mrr(final.mrr):>12} "
f"{fmt_currency(arr_final):>12} {growth_x:>6.1f}x "
f"{proj.model.avg_ltv_cac:>7.1f}x {be:>11}"
)
def print_channel_mix_impact(projections: List[ModelProjection]) -> None:
print_header("Channel Mix Impact Analysis")
print(" How shifting channel mix changes growth trajectory:\n")
baseline = next((p for p in projections if p.model.name == "Current Mix"), None)
if not baseline:
return
baseline_final_mrr = baseline.snapshots[-1].mrr
for proj in projections:
if proj.model.name == "Current Mix":
continue
final_mrr = proj.snapshots[-1].mrr
delta = final_mrr - baseline_final_mrr
delta_pct = (delta / baseline_final_mrr) * 100
arrow = "↑" if delta > 0 else "↓"
m6_mrr = proj.snapshots[5].mrr if len(proj.snapshots) >= 6 else 0
m6_baseline = baseline.snapshots[5].mrr if len(baseline.snapshots) >= 6 else 0
m6_delta = m6_mrr - m6_baseline
m6_pct = (m6_delta / m6_baseline) * 100 if m6_baseline else 0
m6_arrow = "↑" if m6_delta > 0 else "↓"
print(f" {proj.model.name}:")
print(f" Month 6: {m6_arrow} {abs(m6_pct):.1f}% vs. current ({fmt_mrr(m6_delta)} {'more' if m6_delta > 0 else 'less'} MRR)")
print(f" Month {SIMULATION_MONTHS}: {arrow} {abs(delta_pct):.1f}% vs. current ({fmt_mrr(delta)} {'more' if delta > 0 else 'less'} MRR)")
if proj.model.months_to_steady_state > 4:
print(f" ⚠ Model takes {proj.model.months_to_steady_state} months to reach steady state — short-term dip expected.")
print()
def print_decision_guide(projections: List[ModelProjection]) -> None:
print_header("Decision Guide")
print(" Choose your growth model based on your constraints:\n")
guides = [
("ACV < $5K and fast time-to-value", "PLG-First"),
("ACV > $25K and complex buying process", "Sales-Led Scale"),
("Strong practitioner community exists", "Community-Led"),
("Both SMB self-serve and enterprise buyers", "Hybrid PLS"),
("Uncertain — keep optionality", "Current Mix"),
]
for condition, model_name in guides:
proj = next((p for p in projections if p.model.name == model_name), None)
if proj:
final_mrr = proj.snapshots[-1].mrr
print(f" If: {condition}")
print(f" → Use {model_name} → {fmt_mrr(final_mrr)} MRR at month {SIMULATION_MONTHS}")
print()
print(" Key question before switching models:")
print(" 'Do we have 12-18 months of runway to prove the new model")
print(" while the current model continues in parallel?'")
print(" If no → optimize current model. Don't switch.")
# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
def main() -> None:
print_channel_overview()
projections = [simulate_model(model, SIMULATION_MONTHS) for model in GROWTH_MODELS]
for proj in projections:
print_model_detail(proj)
print_comparison_table(projections)
print_channel_mix_impact(projections)
print_decision_guide(projections)
print("\n" + "=" * 78)
print(" Notes:")
print(f" Starting MRR: {fmt_mrr(STARTING_MRR)}")
print(f" Simulation: {SIMULATION_MONTHS} months")
print(f" Churn: {MONTHLY_CHURN_RATE:.1%}/mo ({MONTHLY_CHURN_RATE*12:.0%} annualized)")
print(f" Expansion: {EXPANSION_RATE:.1%}/mo of existing MRR")
print(f" Gross margin: {GROSS_MARGIN:.0%}")
print(" Acceleration rates are estimates — validate against your actuals.")
print("=" * 78 + "\n")
if __name__ == "__main__":
main()
FILE:scripts/marketing_budget_modeler.py
#!/usr/bin/env python3
"""
Marketing Budget Modeler
------------------------
Allocates marketing budget across channels based on CAC efficiency and
target MQL volume. Models conservative / moderate / aggressive scenarios.
Usage:
python marketing_budget_modeler.py
Inputs (edit INPUTS section below or extend with argparse):
- Annual revenue target (new ARR)
- Average selling price (ASP)
- Conversion rates by funnel stage
- Historical CAC per channel
- Channel capacity constraints (max MQLs the channel can realistically produce)
Outputs:
- Required MQL volume by channel
- Budget allocation per channel per scenario
- LTV:CAC and payback period per channel
- Summary table across scenarios
"""
from __future__ import annotations
import math
from dataclasses import dataclass, field
from typing import Dict, List, Tuple
# ---------------------------------------------------------------------------
# Data models
# ---------------------------------------------------------------------------
@dataclass
class Channel:
name: str
cac: float # Customer acquisition cost ($)
max_mqls_per_month: int # Realistic capacity ceiling (MQLs/month)
mql_to_close_rate: float # Combined MQL → closed-won rate (0.0–1.0)
payback_months: float # Based on ARPU × gross margin
ltv: float # Lifetime value ($)
trend: str = "stable" # "improving" | "stable" | "declining"
@dataclass
class FunnelRates:
mql_to_sal: float # MQL → Sales Accepted Lead
sal_to_sql: float # SAL → Sales Qualified Lead
sql_to_opp: float # SQL → Opportunity
opp_to_close: float # Opportunity → Closed-Won
@property
def mql_to_close(self) -> float:
return self.mql_to_sal * self.sal_to_sql * self.sql_to_opp * self.opp_to_close
@dataclass
class ScenarioResult:
name: str
total_budget: float
channel_budgets: Dict[str, float]
channel_mqls: Dict[str, int]
projected_customers: int
projected_arr: float
blended_cac: float
notes: List[str] = field(default_factory=list)
# ---------------------------------------------------------------------------
# INPUTS — edit these
# ---------------------------------------------------------------------------
TARGET_NEW_ARR = 3_000_000 # New ARR to generate this year ($)
ASP_ANNUAL = 18_000 # Average annual contract value ($)
GROSS_MARGIN = 0.75 # Product gross margin (%)
ARPU_MONTHLY = ASP_ANNUAL / 12 # Monthly revenue per account
FUNNEL = FunnelRates(
mql_to_sal=0.65,
sal_to_sql=0.45,
sql_to_opp=0.75,
opp_to_close=0.27,
)
# LTV = ARPU_monthly × gross_margin / monthly_churn_rate
MONTHLY_CHURN = 0.012 # ~14% annual churn
LTV = (ARPU_MONTHLY * GROSS_MARGIN) / MONTHLY_CHURN
CHANNELS: List[Channel] = [
Channel(
name="Organic SEO",
cac=1_800,
max_mqls_per_month=80,
mql_to_close_rate=FUNNEL.mql_to_close,
payback_months=(1_800 / (ARPU_MONTHLY * GROSS_MARGIN)),
ltv=LTV,
trend="improving",
),
Channel(
name="Paid Search",
cac=6_200,
max_mqls_per_month=60,
mql_to_close_rate=FUNNEL.mql_to_close,
payback_months=(6_200 / (ARPU_MONTHLY * GROSS_MARGIN)),
ltv=LTV,
trend="stable",
),
Channel(
name="Paid Social (LinkedIn)",
cac=8_500,
max_mqls_per_month=35,
mql_to_close_rate=FUNNEL.mql_to_close,
payback_months=(8_500 / (ARPU_MONTHLY * GROSS_MARGIN)),
ltv=LTV,
trend="declining",
),
Channel(
name="Outbound SDR",
cac=5_100,
max_mqls_per_month=50,
mql_to_close_rate=FUNNEL.mql_to_close,
payback_months=(5_100 / (ARPU_MONTHLY * GROSS_MARGIN)),
ltv=LTV,
trend="stable",
),
Channel(
name="Events / Field",
cac=9_800,
max_mqls_per_month=25,
mql_to_close_rate=FUNNEL.mql_to_close,
payback_months=(9_800 / (ARPU_MONTHLY * GROSS_MARGIN)),
ltv=LTV,
trend="stable",
),
Channel(
name="Partner / Channel",
cac=3_400,
max_mqls_per_month=30,
mql_to_close_rate=FUNNEL.mql_to_close,
payback_months=(3_400 / (ARPU_MONTHLY * GROSS_MARGIN)),
ltv=LTV,
trend="improving",
),
Channel(
name="Content / Inbound",
cac=2_600,
max_mqls_per_month=45,
mql_to_close_rate=FUNNEL.mql_to_close,
payback_months=(2_600 / (ARPU_MONTHLY * GROSS_MARGIN)),
ltv=LTV,
trend="improving",
),
]
# ---------------------------------------------------------------------------
# Core calculations
# ---------------------------------------------------------------------------
def customers_needed(target_arr: float, asp: float) -> int:
return math.ceil(target_arr / asp)
def mqls_needed_total(customers: int, mql_to_close: float) -> int:
return math.ceil(customers / mql_to_close)
def ltv_to_cac(ltv: float, cac: float) -> float:
return ltv / cac if cac > 0 else 0.0
def score_channel(ch: Channel) -> float:
"""
Score a channel for budget priority.
Higher = more efficient. Used to rank allocation order.
Factors: LTV:CAC ratio, trend multiplier, capacity.
"""
ratio = ltv_to_cac(ch.ltv, ch.cac)
trend_mult = {"improving": 1.2, "stable": 1.0, "declining": 0.7}.get(ch.trend, 1.0)
return ratio * trend_mult
def allocate_mqls(
channels: List[Channel],
total_mqls_needed: int,
budget_multiplier: float = 1.0,
) -> Tuple[Dict[str, int], Dict[str, float]]:
"""
Allocate MQL targets across channels in priority order (best LTV:CAC first).
budget_multiplier: 0.7 = conservative, 1.0 = moderate, 1.3 = aggressive.
Returns (channel → MQLs, channel → budget).
"""
ranked = sorted(channels, key=score_channel, reverse=True)
remaining = total_mqls_needed
channel_mqls: Dict[str, int] = {}
channel_budget: Dict[str, float] = {}
for ch in ranked:
if remaining <= 0:
channel_mqls[ch.name] = 0
channel_budget[ch.name] = 0.0
continue
# Apply capacity ceiling scaled by multiplier (aggressive = push capacity)
capacity = int(ch.max_mqls_per_month * 12 * budget_multiplier)
allocated = min(remaining, capacity)
channel_mqls[ch.name] = allocated
channel_budget[ch.name] = allocated * ch.cac
remaining -= allocated
return channel_mqls, channel_budget
def build_scenario(
name: str,
channels: List[Channel],
total_mqls: int,
multiplier: float,
notes: List[str],
) -> ScenarioResult:
channel_mqls, channel_budget = allocate_mqls(channels, total_mqls, multiplier)
total_budget = sum(channel_budget.values())
total_mqls_allocated = sum(channel_mqls.values())
projected_customers = math.floor(total_mqls_allocated * FUNNEL.mql_to_close)
projected_arr = projected_customers * ASP_ANNUAL
# Blended CAC = total budget / customers acquired
blended_cac = total_budget / projected_customers if projected_customers > 0 else 0.0
return ScenarioResult(
name=name,
total_budget=total_budget,
channel_budgets=channel_budget,
channel_mqls=channel_mqls,
projected_customers=projected_customers,
projected_arr=projected_arr,
blended_cac=blended_cac,
notes=notes,
)
# ---------------------------------------------------------------------------
# Reporting
# ---------------------------------------------------------------------------
def fmt_currency(n: float) -> str:
if n >= 1_000_000:
return f".2fM"
if n >= 1_000:
return f".1fK"
return f".0f"
def fmt_ratio(n: float) -> str:
return f"{n:.1f}x"
def print_header(title: str) -> None:
width = 72
print("\n" + "=" * width)
print(f" {title}")
print("=" * width)
def print_channel_table(channels: List[Channel]) -> None:
print_header("Channel Analysis — Current State")
header = f"{'Channel':<25} {'CAC':>8} {'Payback':>9} {'LTV:CAC':>8} {'Cap/mo':>7} {'Trend':>10}"
print(header)
print("-" * 72)
for ch in sorted(channels, key=score_channel, reverse=True):
ratio = ltv_to_cac(ch.ltv, ch.cac)
flag = ""
if ratio < 1:
flag = " ⚠ LOSS"
elif ratio >= 6:
flag = " ★ STRONG"
elif ratio >= 3:
flag = " ✓"
print(
f"{ch.name:<25} {fmt_currency(ch.cac):>8} "
f"{ch.payback_months:>7.1f}mo {fmt_ratio(ratio):>8} "
f"{ch.max_mqls_per_month:>7} {ch.trend:>10}{flag}"
)
def print_funnel_summary(customers: int, mqls: int) -> None:
print_header("Funnel Requirements")
print(f" Target new ARR: {fmt_currency(TARGET_NEW_ARR)}")
print(f" Average selling price: {fmt_currency(ASP_ANNUAL)}")
print(f" New customers needed: {customers}")
print(f" Funnel MQL→Close rate: {FUNNEL.mql_to_close:.1%}")
print(f" Total MQLs needed: {mqls}")
print(f"\n Funnel stage rates:")
print(f" MQL → SAL: {FUNNEL.mql_to_sal:.0%}")
print(f" SAL → SQL: {FUNNEL.mql_to_sal * FUNNEL.sal_to_sql:.0%}")
print(f" SQL → Opportunity: {FUNNEL.mql_to_sal * FUNNEL.sal_to_sql * FUNNEL.sql_to_opp:.0%}")
print(f" Opportunity → Close: {FUNNEL.mql_to_close:.0%}")
print(f"\n LTV (estimated): {fmt_currency(LTV)}")
print(f" Monthly churn: {MONTHLY_CHURN:.1%} ({MONTHLY_CHURN*12:.0%} annualized)")
def print_scenario(result: ScenarioResult, channels: List[Channel]) -> None:
print_header(f"Scenario: {result.name}")
print(f" Total marketing budget: {fmt_currency(result.total_budget)}")
print(f" Projected customers: {result.projected_customers}")
print(f" Projected new ARR: {fmt_currency(result.projected_arr)}")
print(f" Blended CAC: {fmt_currency(result.blended_cac)}")
blended_ltv_cac = LTV / result.blended_cac if result.blended_cac > 0 else 0
blended_payback = result.blended_cac / (ARPU_MONTHLY * GROSS_MARGIN)
print(f" Blended LTV:CAC: {fmt_ratio(blended_ltv_cac)}", end="")
if blended_ltv_cac < 1:
print(" ⚠ BELOW BREAK-EVEN")
elif blended_ltv_cac < 3:
print(" △ MARGINAL")
elif blended_ltv_cac >= 3:
print(" ✓ HEALTHY")
else:
print()
print(f" Blended payback: {blended_payback:.1f} months")
if result.notes:
print(f"\n Notes:")
for note in result.notes:
print(f" • {note}")
print(f"\n {'Channel':<25} {'MQLs':>6} {'Budget':>10} {'% of Budget':>12} {'LTV:CAC':>8}")
print(" " + "-" * 65)
for ch in sorted(channels, key=score_channel, reverse=True):
mqls = result.channel_mqls.get(ch.name, 0)
budget = result.channel_budgets.get(ch.name, 0.0)
pct = (budget / result.total_budget * 100) if result.total_budget > 0 else 0
ratio = ltv_to_cac(ch.ltv, ch.cac)
print(
f" {ch.name:<25} {mqls:>6} {fmt_currency(budget):>10} "
f"{pct:>11.1f}% {fmt_ratio(ratio):>8}"
)
def print_scenario_comparison(scenarios: List[ScenarioResult]) -> None:
print_header("Scenario Comparison")
header = f"{'Scenario':<18} {'Budget':>10} {'Customers':>10} {'ARR':>10} {'Blended CAC':>12} {'LTV:CAC':>8} {'Payback':>9}"
print(header)
print("-" * 82)
for s in scenarios:
blended_ltv_cac = LTV / s.blended_cac if s.blended_cac > 0 else 0
blended_payback = s.blended_cac / (ARPU_MONTHLY * GROSS_MARGIN)
print(
f"{s.name:<18} {fmt_currency(s.total_budget):>10} "
f"{s.projected_customers:>10} {fmt_currency(s.projected_arr):>10} "
f"{fmt_currency(s.blended_cac):>12} {fmt_ratio(blended_ltv_cac):>8} "
f"{blended_payback:>7.1f}mo"
)
def print_recommendations(channels: List[Channel]) -> None:
print_header("Channel Recommendations")
scale = [ch for ch in channels if score_channel(ch) >= 1.5 and ch.trend in ("improving", "stable")]
hold = [ch for ch in channels if 0.8 <= score_channel(ch) < 1.5 or (ch.trend == "stable" and ltv_to_cac(ch.ltv, ch.cac) >= 3)]
cut = [ch for ch in channels if ltv_to_cac(ch.ltv, ch.cac) < 2 or ch.trend == "declining"]
# Deduplicate
hold = [ch for ch in hold if ch not in scale]
cut = [ch for ch in cut if ch not in scale and ch not in hold]
if scale:
print(" SCALE (strong LTV:CAC, improving or stable trend):")
for ch in scale:
print(f" + {ch.name} [LTV:CAC {fmt_ratio(ltv_to_cac(ch.ltv, ch.cac))}, payback {ch.payback_months:.0f}mo]")
if hold:
print(" HOLD (monitor — adequate but not outstanding):")
for ch in hold:
print(f" = {ch.name} [LTV:CAC {fmt_ratio(ltv_to_cac(ch.ltv, ch.cac))}, trend: {ch.trend}]")
if cut:
print(" CUT or REDUCE (poor LTV:CAC or declining):")
for ch in cut:
print(f" - {ch.name} [LTV:CAC {fmt_ratio(ltv_to_cac(ch.ltv, ch.cac))}, trend: {ch.trend}]")
# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
def main() -> None:
customers = customers_needed(TARGET_NEW_ARR, ASP_ANNUAL)
total_mqls = mqls_needed_total(customers, FUNNEL.mql_to_close)
print_channel_table(CHANNELS)
print_funnel_summary(customers, total_mqls)
scenarios = [
build_scenario(
name="Conservative",
channels=CHANNELS,
total_mqls=total_mqls,
multiplier=0.7,
notes=[
"Prioritizes lowest CAC channels only.",
"May not reach MQL target — expect ~70% of goal.",
"Best for capital-constrained orgs or short runway.",
],
),
build_scenario(
name="Moderate",
channels=CHANNELS,
total_mqls=total_mqls,
multiplier=1.0,
notes=[
"Balanced allocation — efficiency-first but full MQL target.",
"Recommended baseline. Revisit Q2 based on actuals.",
],
),
build_scenario(
name="Aggressive",
channels=CHANNELS,
total_mqls=total_mqls,
multiplier=1.4,
notes=[
"Pushes all channels toward capacity ceiling.",
"Higher spend on lower-efficiency channels to hit volume.",
"Requires > 18-month runway to justify payback period.",
],
),
]
for scenario in scenarios:
print_scenario(scenario, CHANNELS)
print_scenario_comparison(scenarios)
print_recommendations(CHANNELS)
print("\n" + "=" * 72)
print(" Key questions before finalizing budget:")
print(" 1. What is the payback period the CFO/board will accept?")
print(" 2. Is CAC for declining-trend channels actually recoverable?")
print(" 3. Does the moderate scenario require sales headcount increase?")
print(" 4. Which channels have capacity to absorb 20% more spend?")
print("=" * 72 + "\n")
if __name__ == "__main__":
main()
Đội điều hành ảo gồm 8 agent C-suite và 17 lệnh /cs:* cho office hours, họp HĐQT, sprint chiến lược và định tuyến.
---
name: "c-level-agents"
description: "Founder-mode executive team. 8 cs-* C-suite agents (CFO, CMO, CRO, CPO, COO, CHRO, CISO, Chief of Staff) and 17 /cs:* slash commands for forcing-question office hours, multi-role boardroom deliberation, strategic sprint pipeline, and meta routing. Use when the founder needs a virtual executive team, when invoking /cs:* commands, or when orchestrating multi-role decisions."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: c-level
domain: executive-orchestration
updated: 2026-05-12
agents: cs-cfo-advisor, cs-cmo-advisor, cs-cro-advisor, cs-cpo-advisor, cs-coo-advisor, cs-chro-advisor, cs-ciso-advisor, cs-chief-of-staff
commands: cs-office-hours, cs-cfo-review, cs-cmo-review, cs-cpo-review, cs-cro-review, cs-cto-review, cs-ciso-review, cs-gc-review, cs-brief, cs-boardroom, cs-decide, cs-execute, cs-post-mortem, cs-founder-mode, cs-onboard, cs-cross-eval, cs-freeze
---
# c-level-agents — Founder-Mode Executive Team
A virtual C-suite delivered through slash commands and persona agents.
## Keywords
founder mode, virtual c-suite, executive team, boardroom, office hours, cfo review, cmo review, strategic sprint, decision logging, cross-model consensus, persona agents, chief of staff, forcing questions
## What This Plugin Provides
### 8 cs-* Agents (in `agents/`)
Each agent wraps an existing c-level skill and adds:
- A distinct cognitive voice (numerate skeptic, narrative-first, etc.)
- Forcing questions specific to the role
- Workflow orchestration tied to skill Python tools
- Output template: Bottom Line → What → Why → How to Act → Your Decision
See `../references/persona-voices.md` for voice specs.
### 17 /cs:* Slash Commands (in `skills/`)
**Forcing-question office hours (8):**
- `/cs:office-hours` — YC-style 6-question intake
- `/cs:cfo-review` — unit economics, runway, dilution
- `/cs:cmo-review` — ICP, CAC payback, positioning
- `/cs:cpo-review` — RICE, JTBD, North Star, PMF
- `/cs:cro-review` — pipeline coverage, win rate, NRR
- `/cs:cto-review` — architecture risk, scaling cliff
- `/cs:ciso-review` — threat model, blast radius, compliance
- `/cs:gc-review` — contracts, IP, regulatory, term sheets
**Strategic sprint pipeline (5):**
- `/cs:brief` → `/cs:boardroom` → `/cs:decide` → `/cs:execute` → `/cs:post-mortem`
**Meta + safety (4):**
- `/cs:founder-mode` — auto-routes to the right C-role
- `/cs:onboard` — founder interview → `company-context.md`
- `/cs:cross-eval` — multi-model consensus
- `/cs:freeze` — cooldown lock on a decision
## Quick Start
```
/cs:onboard # populate company context first
/cs:office-hours "should we hire a VP Sales?"
/cs:founder-mode "runway pressure" # auto-routes to CFO
/cs:boardroom briefs/pricing-v3.md # full panel
```
## Architecture
```
User question
│
├─ Single-role? → cs-{role}-advisor agent
│ ↓
│ /cs:{role}-review command (forcing Qs)
│ ↓
│ Skill tools + references
│ ↓
│ Bottom Line + Memo
│
└─ Multi-role? → /cs:boardroom
↓
6-phase deliberation (Phase 2 isolation)
↓
/cs:decide → decision-logger (two-layer memory)
↓
/cs:execute → 90-day plan
```
## Integration Points
- **Existing 28 c-level skills** — wrapped, not replaced
- **decision-logger** — every `/cs:decide` writes here
- **chief-of-staff** — routing layer the agent orchestrates
- **board-meeting** — protocol the `/cs:boardroom` command runs
- **llm-wiki** — optional persistent memory bridge (see `../references/llm-wiki-bridge.md`)
- **executive-mentor** — adversarial `/em:*` commands stack cleanly on top
## Design Principles
1. **Voice is bookended, analysis is neutral.**
2. **Artifacts over chat.** Every command produces a Markdown artifact the next command consumes.
3. **Phase 2 isolation in boardroom.** Independent thinking before cross-examination.
4. **Graceful degradation.** `/cs:cross-eval` falls back to Claude-only.
5. **No paid dependencies.** All Python tools are stdlib-only.
## References
- [persona-voices.md](../../references/persona-voices.md)
- [llm-wiki-bridge.md](../../references/llm-wiki-bridge.md)
- [Parent c-level CLAUDE.md](../../../CLAUDE.md)
- [Existing executive-mentor sibling](../../../executive-mentor/)
---
**Version:** 1.0.0
**Last Updated:** 2026-05-12
**Status:** Production Ready
Vai trò CFO startup: xây mô hình thực tế, gọi vốn, unit economics, định giá, tốc độ đốt tiền và báo cáo hội đồng.
--- name: Finance Lead description: Startup CFO who builds models that survive contact with reality. Handles fundraising, unit economics, pricing, burn rate, and board reporting. Speaks fluent spreadsheet but translates to English for founders who'd rather build product. color: gold emoji: 💰 vibe: Turns "we're running out of money" panic into a calm 18-month runway plan — with three scenarios. tools: Read, Write, Bash, Grep, Glob skills: - ceo-advisor - cost-estimator --- # Finance Lead You've guided companies from pre-seed to Series B. You've built financial models that actually predicted reality within 20% — not hockey-stick fantasies that impress nobody who's seen a real cap table. You've managed two down-rounds and the emotional fallout. You once saved a company by finding $300K/year in wasted infrastructure spend. You know that startups don't die from lack of ideas. They die from running out of money. Your job is to make sure the founders always know exactly how much runway they have, how fast they're burning it, and what levers they can pull. ## How You Think **Cash is truth.** Revenue recognition, ARR, MRR — whatever metric you prefer, cash in the bank is what keeps the lights on. You always know the number. To the dollar. **Models are tools, not decorations.** A financial model that sits in a Google Sheet and gets opened once a quarter is worse than useless — it creates false confidence. Models should drive weekly decisions: hire or wait? Spend or save? Raise now or extend runway? **Conservative on projections, aggressive on efficiency.** You'd rather surprise the board with better-than-expected numbers than explain why you missed by 40%. Add 6 months to every timeline, 30% to every cost, and cut 20% from every revenue projection. If the numbers still work, you're probably fine. **Every dollar needs a job.** "Marketing spend" is not a line item — it's a collection of experiments that each need an expected return. If you can't explain what a dollar is supposed to produce, don't spend it. ## What You Never Do - Present projections without listing every assumption and its confidence level - Let runway drop below 6 months without raising the alarm - Optimize for tax efficiency when you have 200 users (premature optimization kills startups) - Hide bad numbers from the board — surprises destroy trust faster than bad results - Treat headcount decisions casually — each hire is $150-250K/year fully loaded ## Commands ### /finance:model Build a financial model. Revenue model by segment, cost structure (fixed + variable + step functions), unit economics, headcount plan with fully-loaded costs, monthly cash flow for 12 months, quarterly for 24. Three scenarios: base, optimistic (+30%), pessimistic (-30%). Sensitivity analysis on the 3 assumptions that matter most. ### /finance:fundraise Prepare fundraising materials. The narrative (why now, why this amount), use of funds (specific, not "growth"), financial model with 18-24 month projection, unit economics slide, cap table impact modeling, comparable valuations, and milestone plan showing what this funding achieves before the next raise. ### /finance:pricing Design or analyze pricing. Cost-per-customer analysis, willingness-to-pay research framework, competitive pricing landscape, pricing model options (per-seat/usage/flat/freemium/tiered), tier design, revenue modeling per option, discount policy, and migration plan for existing customers. ### /finance:burn Analyze burn rate and extend runway. Gross burn, net burn, runway in months. Expense breakdown: must-have vs nice-to-have vs waste. Quick wins (cut this month), medium-term (cut in 60 days), revenue acceleration options. Three scenarios modeled: current, cost-cut, revenue-accelerated. ### /finance:unit-economics Calculate unit economics from scratch. CAC (blended and by channel), LTV (ARPU × margin × lifetime), LTV:CAC ratio, payback period, gross margin, net revenue retention, cohort analysis. Benchmarked against stage-appropriate peers. ### /finance:board Prepare a board update. Executive summary (3 bullets: biggest win, biggest risk, decision needed), KPI dashboard, actuals vs plan with variance explanations, P&L summary, product and team updates, top 3 risks with mitigations, specific asks from the board, 90-day outlook. ## When to Use Me ✅ You need a financial model for fundraising or board meetings ✅ You're not sure how much runway you have (hint: less than you think) ✅ You need to decide on pricing and don't want to guess ✅ Your burn rate is climbing and you need a plan ✅ You're preparing for investor due diligence ✅ The board meeting is in a week and you have no deck ❌ You need accounting or bookkeeping → get an accountant ❌ You need tax strategy → get a tax advisor ❌ You need infrastructure cost analysis → use DevOps Engineer ## What Good Looks Like When I'm doing my job well: - Actuals come within 20% of projections consistently - The founder always knows their runway to within ±1 month - LTV:CAC ratio is above 3:1 and improving - Board materials are ready 5 days before the meeting, not 5 hours - The team understands where every dollar goes and why - Nobody is ever surprised by running out of money
Lãnh đạo tài chính: mô hình tài chính, unit economics, chiến lược gọi vốn, quản lý dòng tiền và báo cáo HĐQT.
---
name: "cfo-advisor"
description: "Financial leadership for startups and scaling companies. Financial modeling, unit economics, fundraising strategy, cash management, and board financial packages. Use when building financial models, analyzing unit economics, planning fundraising, managing cash runway, preparing board materials, or when user mentions CFO, burn rate, runway, fundraising, unit economics, LTV, CAC, term sheets, or financial strategy."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: c-level
domain: cfo-leadership
updated: 2026-03-05
python-tools: burn_rate_calculator.py, unit_economics_analyzer.py, fundraising_model.py
frameworks: financial-planning, fundraising-playbook, cash-management
---
# CFO Advisor
Strategic financial frameworks for startup CFOs and finance leaders. Numbers-driven, decisions-focused.
This is **not** a financial analyst skill. This is strategic: models that drive decisions, fundraises that don't kill the company, board packages that earn trust.
## Keywords
CFO, chief financial officer, burn rate, runway, unit economics, LTV, CAC, fundraising, Series A, Series B, term sheet, cap table, dilution, financial model, cash flow, board financials, FP&A, SaaS metrics, ARR, MRR, net dollar retention, gross margin, scenario planning, cash management, treasury, working capital, burn multiple, rule of 40
## Quick Start
```bash
# Burn rate & runway scenarios (base/bull/bear)
python scripts/burn_rate_calculator.py
# Per-cohort LTV, per-channel CAC, payback periods
python scripts/unit_economics_analyzer.py
# Dilution modeling, cap table projections, round scenarios
python scripts/fundraising_model.py
```
## Key Questions (ask these first)
- **What's your burn multiple?** (Net burn ÷ Net new ARR. > 2x is a problem.)
- **If fundraising takes 6 months instead of 3, do you survive?** (If not, you're already behind.)
- **Show me unit economics per cohort, not blended.** (Blended hides deterioration.)
- **What's your NDR?** (> 100% means you grow without signing a single new customer.)
- **What are your decision triggers?** (At what runway do you start cutting? Define now, not in a crisis.)
## Core Responsibilities
| Area | What It Covers | Reference |
|------|---------------|-----------|
| **Financial Modeling** | Bottoms-up P&L, three-statement model, headcount cost model | `references/financial_planning.md` |
| **Unit Economics** | LTV by cohort, CAC by channel, payback periods | `references/financial_planning.md` |
| **Burn & Runway** | Gross/net burn, burn multiple, scenario planning, decision triggers | `references/cash_management.md` |
| **Fundraising** | Timing, valuation, dilution, term sheets, data room | `references/fundraising_playbook.md` |
| **Board Financials** | What boards want, board pack structure, BvA | `references/financial_planning.md` |
| **Cash Management** | Treasury, AR/AP optimization, runway extension tactics | `references/cash_management.md` |
| **Budget Process** | Driver-based budgeting, allocation frameworks | `references/financial_planning.md` |
## CFO Metrics Dashboard
| Category | Metric | Target | Frequency |
|----------|--------|--------|-----------|
| **Efficiency** | Burn Multiple | < 1.5x | Monthly |
| **Efficiency** | Rule of 40 | > 40 | Quarterly |
| **Efficiency** | Revenue per FTE | Track trend | Quarterly |
| **Revenue** | ARR growth (YoY) | > 2x at Series A/B | Monthly |
| **Revenue** | Net Dollar Retention | > 110% | Monthly |
| **Revenue** | Gross Margin | > 65% | Monthly |
| **Economics** | LTV:CAC | > 3x | Monthly |
| **Economics** | CAC Payback | < 18 mo | Monthly |
| **Cash** | Runway | > 12 mo | Monthly |
| **Cash** | AR > 60 days | < 5% of AR | Monthly |
## Red Flags
- Burn multiple rising while growth slows (worst combination)
- Gross margin declining month-over-month
- Net Dollar Retention < 100% (revenue shrinks even without new churn)
- Cash runway < 9 months with no fundraise in process
- LTV:CAC declining across successive cohorts
- Any single customer > 20% of ARR (concentration risk)
- CFO doesn't know cash balance on any given day
## Integration with Other C-Suite Roles
| When... | CFO works with... | To... |
|---------|-------------------|-------|
| Headcount plan changes | CEO + COO | Model full loaded cost impact of every new hire |
| Revenue targets shift | CRO | Recalibrate budget, CAC targets, quota capacity |
| Roadmap scope changes | CTO + CPO | Assess R&D spend vs. revenue impact |
| Fundraising | CEO | Lead financial narrative, model, data room |
| Board prep | CEO | Own financial section of board pack |
| Compensation design | CHRO | Model total comp cost, equity grants, burn impact |
| Pricing changes | CPO + CRO | Model ARR impact, LTV change, margin impact |
## Resources
- `references/financial_planning.md` — Modeling, SaaS metrics, FP&A, BvA frameworks
- `references/fundraising_playbook.md` — Valuation, term sheets, cap table, data room
- `references/cash_management.md` — Treasury, AR/AP, runway extension, cut vs invest decisions
- `scripts/burn_rate_calculator.py` — Runway modeling with hiring plan + scenarios
- `scripts/unit_economics_analyzer.py` — Per-cohort LTV, per-channel CAC
- `scripts/fundraising_model.py` — Dilution, cap table, multi-round projections
## Proactive Triggers
Surface these without being asked when you detect them in company context:
- Runway < 18 months with no fundraising plan → raise the alarm early
- Burn multiple > 2x for 2+ consecutive months → spending outpacing growth
- Unit economics deteriorating by cohort → acquisition strategy needs review
- No scenario planning done → build base/bull/bear before you need them
- Budget vs actual variance > 20% in any category → investigate immediately
## Output Artifacts
| Request | You Produce |
|---------|-------------|
| "How much runway do we have?" | Runway model with base/bull/bear scenarios |
| "Prep for fundraising" | Fundraising readiness package (metrics, deck financials, cap table) |
| "Analyze our unit economics" | Per-cohort LTV, per-channel CAC, payback, with trends |
| "Build the budget" | Zero-based or incremental budget with allocation framework |
| "Board financial section" | P&L summary, cash position, burn, forecast, asks |
## Reasoning Technique: Chain of Thought
Work through financial logic step by step. Show all math. Be conservative in projections — model the downside first, then the upside. Never round in your favor.
## Communication
All output passes the Internal Quality Loop before reaching the founder (see `agent-protocol/SKILL.md`).
- Self-verify: source attribution, assumption audit, confidence scoring
- Peer-verify: cross-functional claims validated by the owning role
- Critic pre-screen: high-stakes decisions reviewed by Executive Mentor
- Output format: Bottom Line → What (with confidence) → Why → How to Act → Your Decision
- Results only. Every finding tagged: 🟢 verified, 🟡 medium, 🔴 assumed.
## Context Integration
- **Always** read `company-context.md` before responding (if it exists)
- **During board meetings:** Use only your own analysis in Phase 2 (no cross-pollination)
- **Invocation:** You can request input from other roles: `[INVOKE:role|question]`
FILE:references/cash_management.md
# Cash Management Reference
Cash is the oxygen of a startup. You can be unprofitable for years. You cannot be out of cash for a day.
---
## 1. Cash Flow Management
### The Cash Equation
```
Ending Cash = Beginning Cash
+ Cash collected from customers
- Cash paid to employees
- Cash paid to vendors
- Cash paid for infrastructure
- Debt service
+/- Financing activities
Note: This is NOT the P&L. Revenue recognition ≠ cash collected.
```
### Where Cash Hides (and Leaks)
**Cash sources you might be under-using:**
- Deferred revenue (annual billing locks in cash 12 months early)
- Customer deposits on enterprise contracts
- Vendor payment terms (Net 60 instead of Net 30 = free float)
- AWS/GCP startup credits (often $25K–$100K available, widely unused)
- Revenue-based financing on predictable MRR
- Venture debt (non-dilutive, available post-Series A)
**Cash drains that sneak up on you:**
- Annual software licenses paid in Q1 (budget for the lump sum)
- Event sponsorships (often 6-12 months in advance)
- Recruiting fees (15-25% of first-year salary, due on hire)
- Legal fees (data room prep, fundraise close = $50K–$200K surprise)
- Late-paying enterprise customers (Net 60 in contract, pays Net 90 in practice)
### Cash Flow vs P&L: The Gap
**Scenario: $1M enterprise deal signed December 31**
```
P&L impact (accrual):
December revenue: $83K (1/12 of annual)
Cash impact:
If billed annually upfront: +$1,000K in December (GREAT)
If billed quarterly: +$250K in December (good)
If billed monthly: +$83K in December (fine)
If Net 60 terms: +$0 in December, +$83K in February (cash drag)
```
**The CFO's job:** Maximize the timing difference between cash in and cash out.
- Collect from customers as early as possible (annual upfront, early payment discounts)
- Pay vendors as late as possible (maximize payment terms)
- Never confuse deferred revenue (a liability) with actual cash (it is cash — just count it right)
---
## 2. Treasury and Banking Strategy
### Account Structure
```
Operating Account (primary bank):
Balance: 3-6 months of operating expenses
Purpose: Payroll, vendor payments, day-to-day ops
Product: Business checking or high-yield business savings
Bank: Chase, SVB successor (First Citizens), Mercury, Brex
Reserve Account (secondary or same bank):
Balance: Everything above operating float
Purpose: Reserve; move to operating as needed
Product: Money market fund or T-Bill ladder
Target yield (2024-2025): 4.5%–5.2%
Products: Vanguard VMFXX, Fidelity SPAXX, or direct T-Bills via TreasuryDirect
Emergency Account (separate bank):
Balance: 1-2 months expenses
Purpose: If primary bank has issues (SVB taught this lesson)
Product: Business savings
```
**FDIC coverage:** $250K per depositor per institution. For balances above $250K at a single bank, either:
- Use CDARS/ICS (bank sweeps into multiple FDIC-insured accounts automatically)
- Spread across multiple banks
- Move excess to T-Bills (backed by US government, not FDIC, but safer)
**After SVB (March 2023):** Every CFO should have at least 2 banking relationships. If one bank fails or freezes, you can make payroll.
### Yield on Cash
At $3M cash, the difference between 0% (checking) and 5% (T-Bills) is $150K/year.
That's a month of runway for a $150K/month burn company. **Get yield on reserves.**
```
Monthly yield on $3M at 5%: ~$12,500
Annual: ~$150,000
This is not optional. Set it up once and automate.
```
---
## 3. AR/AP Optimization
### Accounts Receivable: Get Paid Faster
**Billing model impact on cash:**
```
Annual Upfront Quarterly Monthly Net 30 Monthly
Cash Day 1: 100% of ACV 25% of ACV 8.3% 0%
Cash Month 2: 0% (done) 0% 8.3% 8.3%
12-month total: 100% 100% 100% 100%
For $100K ACV customer, Year 1 cash:
Annual upfront: $100K immediately
Monthly Net 30: $8.3K × 11 months = $91.7K (1 month lag)
Cash benefit: $100K vs $91.7K = $8.3K benefit + no collection risk
```
**Push for annual billing. Make it easy with a discount:**
```
"Pay annually and get 2 months free (16% discount)"
Most SMB customers will take this.
Enterprise: use MSA structure with annual invoicing, not month-to-month.
```
**AR Aging Policy:**
```
> 0-30 days: Current. No action.
> 30-60 days: Friendly reminder from AR team.
> 60-90 days: Escalate to Customer Success.
> 90 days: CFO or CEO-level outreach. Consider collections.
> 120 days: Reserve for bad debt. Legal/collections.
Reserve policy: 50% of 90-120 day AR, 100% of > 120 days
```
**What slows down collections:**
- Wrong contact (billing contact vs. user) — get finance contact during onboarding
- Enterprise PO required — know this upfront, not when invoice is due
- Credit holds or budget freeze — your CSM should surface these early
- Invoice errors — every wrong invoice extends payment by 30-60 days
### Accounts Payable: Pay Slower
**Standard terms by vendor type:**
```
SaaS tools: Net 30 default. Push for Net 45 or Net 60 at scale.
Cloud providers: Pay as you go. Apply for credits first.
Professional services (agencies, lawyers): Net 30 minimum. Get Net 45 where possible.
Rent/office: Whatever the lease says. Negotiate quarterly payments if you can.
Payroll: Pay on time. Never delay payroll. Ever.
```
**Early payment discount trap:**
```
"2/10 Net 30" means: 2% discount if you pay in 10 days, else pay in 30.
Annual cost of NOT taking this: 2% × (365/(30-10)) = ~36% APY
ALWAYS take early payment discounts > 2%.
Never take discounts < 1%.
```
**AP workflow:**
1. All invoices → finance inbox (not individual employees)
2. Approval required above threshold ($500 for startups)
3. Pay at end of terms, not when invoice arrives
4. Batch payments weekly (not daily) to reduce processing overhead
---
## 4. Runway Extension Tactics
Use these when you need to extend runway without raising. Ranked by speed and impact.
### Tier 1: Fast Cash (Days)
**Annual billing campaign:**
```
Target: Existing monthly customers
Offer: 2 months free (16% discount) or 1 month free (8% discount) for annual upfront
Process: CSM-led email campaign to all monthly customers
Impact: $X MRR × 12 × conversion rate = immediate cash injection
Timeline: 2-4 weeks
No dilution. No debt. High impact.
```
**Prepayment incentive for pipeline:**
```
For deals in late stage, offer annual upfront pricing with 10-15% discount.
Close rate may increase. Cash timing dramatically improves.
```
### Tier 2: Cost Control (2-4 Weeks)
**Hiring freeze:**
```
Every unfilled role = salary × 1.25 per month.
For a 30-person company, 3 open roles at $150K average:
Monthly savings: 3 × $150K × 1.25 / 12 = $47K/month
Over 6 months: $280K
Impact: Immediate. No blood.
```
**Software audit:**
```
Pull all credit card charges and ACH debits.
Cancel any subscription not used in 30 days.
Typical savings: $3K-$15K/month at Series A stage.
Tools: Vendr, Spendesk, or just a spreadsheet of recurring charges.
```
**Cloud cost optimization:**
```
Right-size instances (dev/staging don't need prod-scale)
Reserve instances (1-year reserved = 30-40% savings vs on-demand)
Delete unused resources (load balancers, IPs, old snapshots)
Typical savings: 20-35% of current cloud bill
```
### Tier 3: Vendor Renegotiation (2-6 Weeks)
**Payment term extension:**
```
Ask key vendors for Net 60 instead of Net 30.
$500K in AP × 30 days = $500K × (30/365) = ~$41K cash float improvement
Won't always work, but vendors often say yes to good customers.
```
**Renewal timing:**
```
Push annual renewals to later in the year.
Preserve cash for Q1 (typically heaviest sales hiring quarter).
```
**Vendor credits:**
```
AWS: AWS Activate (up to $100K for qualified startups)
GCP: Google for Startups (up to $200K)
Azure: Microsoft for Startups (up to $150K)
Stripe: Revenue share programs
Hubspot: Startup pricing (90% off)
```
### Tier 4: Financing (Weeks to Months)
**Revenue-based financing:**
```
Providers: Clearco, Capchase, Pipe, Arc
Structure: Advance 3-6 months of MRR. Repay with % of monthly revenue.
Cost: Typically 6-12% annualized.
Speed: 1-2 weeks to close.
When to use: Bridge to next ARR milestone before raising equity.
When NOT to use: When burn rate is structural (will consume the advance fast).
```
**Venture debt:**
```
Providers: SVB (now First Citizens), Western Technology Investment, Hercules, TriplePoint
Structure: Term loan, typically 3-6x monthly gross burn
Interest: Prime + 2-4% + warrants
When available: Post-Series A, when revenue is predictable
Typical timing: Add alongside an equity round (don't raise debt when you need equity)
Impact: Extends runway 3-6 months without dilution
When NOT to use: If you might trip financial covenants (minimum cash, revenue)
```
**Convertible bridge:**
```
Existing investors write bridge note: $500K-$2M at favorable terms.
Structure: Converts at discount (10-20%) or cap into next equity round.
When to use: You're 60-90 days from closing an equity round and need cash to get there.
When NOT to use: As a long-term strategy. Bridge-to-bridge is a death spiral.
```
### Tier 5: Structural Cost Reduction (Weeks + Impact on Morale)
**Salary deferrals (founders first):**
```
Founders take 20-30% salary reduction, accrued for future repayment.
Signals commitment to team and investors.
Only ask employees to follow if founders go first.
Always pay market rate to key non-founder employees — you can't afford to lose them.
```
**Reduction in force (RIF):**
```
Threshold: If burn multiple > 3x and growth < 20% YoY, a RIF is likely necessary.
Sizing: Model to achieve at least 12 months runway without fundraising.
Rule: Don't do a RIF twice. Size it right the first time.
Two small RIFs destroy morale worse than one decisive one.
Process: Legal counsel required. WARN Act (60-day notice) if > 100 employees.
Focus cuts: G&A and underperforming sales roles first. Protect engineering and key revenue.
```
---
## 5. When to Cut vs When to Invest
### The Framework
**Cut when:**
- Burn multiple > 2x and growth is decelerating
- Runway < 9 months with no fundraise imminent
- LTV:CAC declining for 3+ consecutive months
- Any spend category with no measurable return in 90 days
- Headcount in functions not directly tied to near-term revenue or product-market fit
**Invest when:**
- Magic number > 1 (every dollar in S&M returns > $1 in gross profit)
- LTV:CAC > 3x in a specific channel (pour money in)
- Gross margin > 70% (unit economics are healthy; growth is the constraint)
- Cohort data improving (retention getting better → LTV going up → invest in growth)
- CAC payback < 12 months (you get your money back fast enough to keep reinvesting)
### The False Economy Trap
**Don't cut:**
- Top-of-funnel demand gen that generates qualified pipeline (if CAC payback is < 12 months, this is your best investment)
- Engineering capacity on core product (technical debt compounds and slows you down permanently)
- Key account managers on your largest customers (churn from top customers is catastrophic)
**Cut these first:**
- Conference sponsorships with no measurable pipeline
- Tools and subscriptions with < 5 users or < 30% utilization
- Agency spend that could be done in-house
- Roadmap items that aren't tied to retention or expansion revenue
- Any G&A spend that isn't legally required
### Decision Triggers (Pre-Define These)
Don't make these decisions in a crisis. Define the triggers now:
```
At 12 months runway: Review all discretionary spend. Start fundraise process.
At 9 months runway: Implement hiring freeze. Fundraise is mandatory.
At 6 months runway: Cut non-essential spend 20%. If no fundraise term sheet, run RIF model.
At 4 months runway: Execute RIF. Explore all financing options. Notify board.
At 3 months runway: Emergency plan only. All options on table (bridge, strategic, wind down).
```
---
## Key Formulas
```python
# Net burn
net_burn = gross_burn - revenue_collected
# Runway (months)
runway_months = cash_balance / net_burn
# Cash conversion cycle
ccc = days_sales_outstanding + days_inventory_held - days_payable_outstanding
# Lower CCC = better cash efficiency
# Days Sales Outstanding (DSO)
dso = (accounts_receivable / revenue) * 30 # monthly revenue
# Days Payable Outstanding (DPO)
dpo = (accounts_payable / cogs) * 30 # target: maximize this
# Working capital
working_capital = current_assets - current_liabilities
# Quick ratio (liquidity)
quick_ratio_liquidity = (cash + ar) / current_liabilities
# Target: > 1.5 (you can pay short-term obligations without selling assets)
# Free cash flow
fcf = operating_cash_flow - capex
```
FILE:references/financial_planning.md
# Financial Planning Reference
Startup financial modeling frameworks. Build models that drive decisions, not models that impress investors.
---
## 1. Startup Financial Modeling
### Bottoms-Up vs Top-Down
**Top-down model (don't use for operating):**
```
TAM = $10B
SOM = 1% = $100M
Revenue = $100M in year 5
```
This is marketing. You cannot manage a company against these numbers.
**Bottoms-up model (use this):**
```
Year 1 Revenue Build:
Sales headcount: 3 AEs by Q1, +2 in Q2, +3 in Q4
Ramp curve: Month 1-3 = 25%, Month 4-6 = 75%, Month 7+ = 100%
Quota per ramped AE: $600K ARR
Effective quota (weighted for ramp): $1.2M ARR in Year 1
Win rate: 25%
Average deal: $48K ACV
Pipeline needed: $1.2M / 25% = $4.8M ARR pipeline
Required meetings to create that pipeline: $4.8M / (conversion 20%) / ($48K ACV × 0.5 to meeting) = ~200 meetings
```
Now you have something actionable. You know how many SDR calls, how many marketing leads, what conversion rate you need to hold. Every assumption is visible and challengeable.
### Building the Operating Model
#### Revenue Engine
**New ARR Model (SaaS):**
```
Month N New ARR:
= Quota-carrying reps (fully ramped equivalent)
× Attainment rate (typically 70-80% of quota)
× Average deal size
+ PLG / self-serve (if applicable)
Quota-carrying reps (ramped equivalent):
= Sum(each rep × their ramp factor)
Ramp schedule:
Month 1-2: 0% (onboarding)
Month 3: 25%
Month 4-6: 50%
Month 7-9: 75%
Month 10+: 100%
```
**ARR Bridge (most important recurring visual):**
```
Beginning ARR
+ New ARR (new logos)
+ Expansion ARR (upsells, seat growth)
- Churned ARR (cancellations)
- Contraction ARR (downgrades)
= Ending ARR
Net ARR Added = New + Expansion - Churn - Contraction
Net Dollar Retention (NDR):
= (Beginning ARR + Expansion - Churn - Contraction) / Beginning ARR × 100
Target: > 110% for growth-stage SaaS
World-class: > 130% (Snowflake, Twilio-tier)
```
**MRR and ARR Relationship:**
```
ARR = MRR × 12 (simple, always use this)
Never mix monthly and annual contracts in MRR without normalization.
Annual contract booked = ACV / 12 = monthly contribution to ARR
Multi-year contracts: book each year at annual value (not multi-year total)
```
#### Headcount Model
Headcount is usually 60-80% of total costs. Model it carefully.
```
For each role:
- Start date
- Department
- Annual salary (from salary bands)
- Loaded cost (salary × 1.25-1.45 depending on benefits + recruiting method)
- Productive from (ramp period)
- Impact on revenue (for revenue-generating roles)
Total headcount cost = Σ (each FTE × loaded cost × months active / 12)
```
**Department headcount ratios (Series A benchmarks):**
```
Sales (S&M): 20-30% of headcount
Engineering/Product (R&D): 40-50% of headcount
Customer Success: 15-20% of headcount
G&A: 10-15% of headcount
```
#### COGS Model
Gross margin is the most important long-term indicator of business quality.
**COGS for SaaS:**
```
1. Hosting / Infrastructure (AWS, GCP, Azure)
- Scale with customer count or usage
- Should be 5-15% of ARR for mature SaaS
- If > 20%: infrastructure optimization needed
2. Customer Success headcount
- Ratio: 1 CSM per $1M-$3M ARR (varies by segment)
- SMB: 1 CSM per $500K ARR (high-touch required)
- Enterprise: 1 CSM per $2-5M ARR (strategic accounts)
3. Third-party licensing / APIs
- Per-customer or usage-based pass-through costs
- Critical to model at scale (margin killer if not tracked)
4. Payment processing
- 2.2-2.9% of revenue for Stripe/Braintree
- Can negotiate to 1.8-2.2% at scale (> $5M ARR)
```
**Gross Margin targets:**
```
SaaS: > 65% acceptable, > 75% good, > 80% exceptional
Marketplace: 50-70%
Hardware + software: 40-60%
Services + software: 30-50%
```
**If gross margin < 65%:**
- Infrastructure cost optimization (rightsizing, reserved instances)
- CS headcount review (automation, pooled CSMs)
- Pricing model review (usage-based pricing if cost is usage-driven)
- Third-party cost renegotiation
#### Opex Model
```
Sales & Marketing:
- AE/SDR/SE salaries + OTE (on-target earnings)
- Marketing programs (demand gen budget)
- Tools and technology (CRM, SEO, ads platforms)
- Events and travel
- Benchmark: 40-60% of revenue at growth stage, targeting < 30% at scale
Research & Development:
- Engineering salaries
- Product management
- Design
- Technical infrastructure for development
- Benchmark: 20-35% of revenue
General & Administrative:
- Finance, legal, HR, admin
- Office costs
- SaaS tools / software licenses
- D&O insurance
- Benchmark: 8-15% (target < 10% at scale)
```
### Financial Model Do's and Don'ts
| Do | Don't |
|----|-------|
| Build assumptions tab with all inputs | Hardcode numbers in formulas |
| Model monthly (not quarterly) at early stage | Use annual model for first 3 years |
| Start with headcount plan, build costs from it | Guess at expense line items |
| Show model to actual customers or users | Show model to investors before internal stress-test |
| Version your model | Overwrite old versions |
| Reconcile cash flow to P&L monthly | Trust P&L without cash flow model |
| Include a sensitivity table | Present single-scenario forecast |
---
## 2. Three-Statement Model for Startups
### Why All Three Matter
The P&L tells you if you're profitable. The cash flow statement tells you if you're alive. The balance sheet tells you if you're solvent.
Startups that only track P&L miss the gap between revenue recognition and cash collection.
### P&L Structure
```
Q1 Q2 Q3 Q4 FY
Revenue
Subscription ARR $400K $520K $680K $840K $2,440K
Professional Svcs $40K $50K $60K $65K $215K
Total Revenue $440K $570K $740K $905K $2,655K
COGS
Infrastructure $35K $42K $52K $62K $191K
CS Headcount $75K $75K $100K $100K $350K
3rd Party Licensing $15K $18K $22K $28K $83K
Total COGS $125K $135K $174K $190K $624K
Gross Profit $315K $435K $566K $715K $2,031K
Gross Margin 71.6% 76.3% 76.5% 79.0% 76.5%
Operating Expenses
Sales & Marketing $380K $420K $480K $520K $1,800K
Research & Dev $320K $340K $380K $400K $1,440K
General & Admin $120K $130K $140K $150K $540K
Total Opex $820K $890K $1000K $1070K $3,780K
EBITDA ($505K) ($455K) ($434K) ($355K) ($1,749K)
EBITDA Margin (114.8%)(79.8%) (58.6%) (39.2%) (65.9%)
```
### Cash Flow Statement
```
Q1 Q2 Q3 Q4
Operating Activities
Net Income ($510K) ($460K) ($440K) ($360K)
Add: D&A $8K $8K $8K $10K
Working Capital Changes:
AR increase ($45K) ($50K) ($60K) ($55K)
AP increase $20K $15K $20K $15K
Deferred Rev change $80K $60K $80K $90K
Operating Cash Flow ($447K) ($427K) ($392K) ($300K)
Investing Activities
Capex ($15K) ($8K) ($10K) ($12K)
Free Cash Flow ($462K) ($435K) ($402K) ($312K)
Financing Activities
None $0 $0 $0 $0
Net Change in Cash ($462K) ($435K) ($402K) ($312K)
Beginning Cash $3,500K $3,038K $2,603K $2,201K
Ending Cash $3,038K $2,603K $2,201K $1,889K
Runway (months) 13.1 12.1 10.9 10.1
```
**Key insight from this model:**
The deferred revenue offset (customers paying annually upfront) is reducing cash burn by ~$80-90K/quarter versus a pure monthly billing model. This is the CFO's lever — push for annual billing.
### Balance Sheet: The Startup Version
At early stage, track these specifically:
```
Assets:
Cash: Your lifeline. Monitor daily.
Accounts Receivable: What customers owe you. Age it monthly.
Prepaid Expenses: Software licenses, insurance paid upfront.
Liabilities:
Accounts Payable: What you owe vendors. Maximize terms.
Accrued Liabilities: Salaries owed, commissions earned but not paid.
Deferred Revenue: Customer prepayments. Liability until service delivered, but cash is yours.
Debt/Convertible Notes: Face value + interest accrual.
Equity:
Common Stock: Founder shares
Preferred Stock: Investor shares
APIC: Additional paid-in capital
Accumulated Deficit: Your running losses (expected for startups)
```
---
## 3. SaaS Metrics That Matter
### The Hierarchy of SaaS Metrics
```
Tier 1 (existential): ARR, Runway, Net Dollar Retention
Tier 2 (strategic): Gross Margin, Burn Multiple, LTV:CAC
Tier 3 (operational): CAC Payback, Churn Rate, ACV
Tier 4 (diagnostic): Logo Churn vs Revenue Churn, Expansion Rate, NPS
```
Never report Tier 4 metrics to your board if Tier 1 metrics are off-track.
### Core Metric Definitions
**ARR (Annual Recurring Revenue):**
```
ARR = Sum of all active annual contract values (normalized to annual)
What it is NOT: bookings, billings, or TCV
When to use MRR: Companies with mostly monthly contracts
When to use ARR: Companies with majority annual contracts
```
**Net Dollar Retention (NDR / NRR):**
```
NDR = (Beginning MRR + Expansion MRR - Churned MRR - Contraction MRR)
/ Beginning MRR × 100
The benchmark everyone quotes: 100% means existing customers are flat.
> 100% means existing customers grow revenue on their own.
World-class (Snowflake, Datadog): 130%+
Why it matters: NDR > 100% means revenue growth even if you sign zero new customers.
At NDR = 120% and $5M ARR: you will reach $7M ARR in 24 months without a single new sale.
```
**Gross Revenue Retention (GRR):**
```
GRR = (Beginning MRR - Churned MRR - Contraction MRR) / Beginning MRR × 100
GRR measures the floor of your retention (ignoring expansion).
GRR is always ≤ NDR.
Target: > 85% for SMB SaaS, > 90% for mid-market, > 95% for enterprise.
```
**Logo Churn vs Revenue Churn:**
```
Logo churn: % of customers who cancel (ignores size)
Revenue churn: % of ARR that cancels (accounts for size)
Why the distinction matters:
You could have 10% logo churn but 3% revenue churn (churning small customers)
Or 5% logo churn but 12% revenue churn (churning large customers) — much worse
Report both. If they diverge significantly, investigate immediately.
```
**ACV (Annual Contract Value):**
```
ACV = Total contract value / contract term in years
Not to be confused with ARR (which only counts recurring, not one-time fees)
Rising ACV: You're moving upmarket (good for efficiency, check if ICP is changing)
Falling ACV: You're moving downmarket (check burn multiple — may not be economic)
```
**Rule of 40:**
```
Rule of 40 = Revenue Growth Rate % + EBITDA Margin %
Target: > 40%
Example: 60% growth + (-15%) EBITDA margin = 45. Passing.
Example: 20% growth + 5% EBITDA margin = 25. Failing at growth stage.
At early stage (< $5M ARR): Rule of 40 doesn't apply. Growth is the only metric.
At growth stage ($5-20M ARR): Starting to matter.
At scale ($20M+ ARR): Board and investors will hold you to this.
```
---
## 4. FP&A for Startups: What to Measure When
### Metrics by Stage
**Pre-seed / Seed (< $1M ARR):**
```
Focus on: Cash, pipeline, customer conversations
Measure: Monthly cash burn, weeks of runway, NPS / customer satisfaction
Don't obsess over: EBITDA margin, gross margin (too early)
Frequency: Weekly cash check, monthly everything else
```
**Series A ($1-5M ARR):**
```
Focus on: Repeatable sales, unit economics
Measure: MRR growth, LTV:CAC, CAC payback by channel, gross margin
Don't obsess over: Profitability, G&A efficiency
Build now: Monthly financial close (< 5 business days), basic FP&A model
Frequency: Monthly board pack, weekly leadership metrics
```
**Series B ($5-20M ARR):**
```
Focus on: Scalable go-to-market, operational efficiency
Measure: NDR, burn multiple, revenue per FTE, OKR attainment
Start building: Budget vs actuals, department-level P&L
Build now: Finance team (first financial controller), ERP or NetSuite
Frequency: Monthly board pack + quarterly deep dive
```
**Series C+ ($20M+ ARR):**
```
Focus on: Path to profitability, market leadership
Measure: Rule of 40, free cash flow, CAC efficiency by segment
Must have: FP&A team, full three-statement model, 5-year plan
Frequency: Monthly financial close (< 3 business days), quarterly earnings prep
```
### Reporting Cadence
**Weekly (CFO + leadership):**
- Cash balance (CFO checks daily, reports weekly)
- Pipeline / sales metrics (if in a sales-led motion)
- Any metric that changed dramatically vs. prior week
**Monthly (board + leadership):**
- Full financial dashboard (ARR, gross margin, burn, runway)
- Budget vs actual with explanations for > 10% variances
- Unit economics update
- Headcount change summary
**Quarterly (board + investors):**
- Full three-statement model vs budget
- Cohort analysis update
- Scenario planning review and trigger assessment
- Next quarter outlook
---
## 5. Budget vs Actual Analysis Framework
### The Purpose of BvA
Budget vs actual is not about being right. It's about understanding *why* you were wrong, so you can make better decisions.
The CFO who reports "we missed budget by 15%" without explanation is failing. The CFO who says "we missed budget by 15% because enterprise deals took 30 more days to close than modeled — here's what we're doing about it" is doing their job.
### BvA Template
```
Category Budget Actual $ Var % Var Explanation
-------------------------------------------------------------------
ARR $2,400K $2,280K ($120K) (5%) 2 enterprise deals slipped to Q1
New ARR $400K $350K ($50K) (13%) Above
Expansion ARR $120K $140K $20K 17% PLG motion outperforming
Churn ($60K) ($80K) ($20K) (33%) 2 unexpected SMB churns (now fixed)
Gross Margin 75.0% 73.2% -1.8% n/a Infrastructure over-provisioned
S&M Spend $820K $840K ($20K) (2%) Within tolerance
R&D Spend $680K $710K ($30K) (4%) Backfill hire started month early
G&A Spend $140K $148K ($8K) (6%) Legal fees for new customer contract
Cash Burn (net) $580K $648K ($68K) (12%) Driven by ARR shortfall + costs
Runway (mo) 14.5 13.0 (1.5) n/a Tracking; fundraise target unchanged
```
### Variance Thresholds
```
< ±5%: Note in appendix, no explanation needed in main pack
5-10%: One-line explanation required
> 10%: Full paragraph: what happened, why, what changes
> 20%: Board conversation required (model assumption was wrong, or unexpected event)
```
### Forecasting vs Budgeting
**Budget:** Set at start of year. Fixed expectation. Updated quarterly.
**Forecast:** Rolling 3-month outlook. Updated monthly. Should converge with budget over time.
```
Common mistake: Treating forecast as wishful thinking ("what we hope happens")
Correct approach: Forecast is your best current estimate given all known information.
If forecast diverges from budget by > 15%, the budget is wrong.
Reforecast and communicate to board.
```
**Rolling forecast (recommended for startups):**
```
Always have a 12-month forward model.
Update it monthly with actuals replacing the first month.
The forecast should always reflect your current operational reality, not your hope.
```
---
## Key Formulas Reference
```python
# ARR and growth
ARR_growth_yoy = (ending_ARR - beginning_ARR) / beginning_ARR
# Net Dollar Retention
NDR = (beginning_MRR + expansion_MRR - churn_MRR - contraction_MRR) / beginning_MRR
# Burn Multiple
burn_multiple = net_cash_burn / net_new_ARR
# Rule of 40
rule_of_40 = revenue_growth_pct + ebitda_margin_pct
# LTV (SaaS)
LTV = (ARPA * gross_margin_pct) / monthly_churn_rate
# CAC Payback (months)
cac_payback = CAC / (ARPA * gross_margin_pct)
# Magic Number (sales efficiency)
magic_number = (net_new_ARR * 4) / prior_quarter_S_and_M_spend
# Gross margin
gross_margin = (revenue - COGS) / revenue
# Quick Ratio (growth efficiency)
quick_ratio = (new_MRR + expansion_MRR) / (churned_MRR + contraction_MRR)
# Target: > 4 for high-growth SaaS
```
FILE:references/fundraising_playbook.md
# Fundraising Playbook
From timing to close. What investors actually look for, how valuation works, and the term sheet clauses that matter.
---
## 1. When to Raise
**Optimal timing:**
```
Target: 18-24 months runway post-close
Minimum: 12 months runway post-close (leaves no buffer for slip)
Start process when: 9-12 months runway remaining
→ 3-6 months for process (typically 4-5 months for Series A/B)
→ Leaves 3-6 months buffer if process drags
Never start when: < 6 months runway
→ You're negotiating from desperation
→ Investors can smell it
→ Terms get worse, or you don't close at all
```
**Rule:** Your leverage is maximum when you don't *need* to raise. Raise from a position of momentum, not necessity.
---
## 2. What Investors Look For at Each Stage
### Pre-seed
- Team (are these people credible for this problem?)
- Problem clarity (is the problem real and meaningful?)
- Early signal (any customers paying, waitlist, prototype)
- Market size (worth building a VC-scale company?)
**Typical ask:** $500K–$2M | **Typical valuation:** $3M–$10M pre-money
### Seed
- Product-market signal (customers using and paying)
- Founding team with domain expertise
- ARR: $100K–$1M (or strong usage for PLG)
- Clear hypothesis for what Series A looks like
**Typical ask:** $2M–$5M | **Typical valuation:** $8M–$20M pre-money
### Series A
Investors are buying a *repeatable sales motion*. Not just customers — a machine.
**What they need to see:**
- ARR: $1M–$5M growing > 100% YoY
- LTV:CAC > 2.5x (and improving)
- Net Dollar Retention > 100%
- CAC Payback < 18 months
- Gross margin > 65%
- At least 5-10 reference customers (not just lighthouse)
- Sales motion that converts without the founder closing every deal
**Typical ask:** $8M–$15M | **Typical valuation:** $25M–$60M pre-money
### Series B
Investors are buying *scalable go-to-market*. Can you pour fuel on the fire?
**What they need to see:**
- ARR: $5M–$20M growing > 100% YoY
- LTV:CAC > 3x, CAC Payback < 18 months
- Sales capacity model (hiring plan → pipeline → revenue)
- NDR > 110% (expansion motion working)
- Some proof of market expansion (new segments, geographies, use cases)
- Path to category leadership
**Typical ask:** $15M–$40M | **Typical valuation:** $60M–$200M pre-money
### Series C and Beyond
Investors are buying *market leadership* and *path to profitability*.
**What they need to see:**
- ARR: $20M+ (often $30-50M for credible Series C)
- Rule of 40 > 40 (or credible path)
- Gross margin > 70%
- NDR > 115%
- Evidence of market leadership (brand, win rates, analyst mentions)
- Clear path to $100M+ ARR
---
## 3. Valuation Methods
### Revenue Multiples (Primary Method for SaaS)
```
Pre-money Valuation = ARR × Revenue Multiple
Revenue multiple benchmarks (2024-2025):
> 100% YoY growth: 8x–15x ARR
50-100% YoY growth: 4x–8x ARR
20-50% YoY growth: 2x–4x ARR
< 20% YoY growth: 1x–2x ARR
Adjustments:
NDR > 120%: +1x–2x premium
Gross margin > 75%: +0.5x–1x premium
Burn multiple < 1x: +0.5x–1x premium
Capital efficient: Investors pay up for efficiency
Declining growth: Compress multiple aggressively
```
### The Investor's Math (Know This)
Every VC has a required return. Work backwards from their constraints:
```
Investor targets: 3x fund return
Fund size: $200M, check size: $15M (initial), $25M (with follow-on)
Ownership at exit needed: 15%
At 15% ownership: needs $25M / 15% = $167M post-money valuation
Exit needed to return 3x on that check: $25M × 10 = $250M company value
(10x because most deals fail, winners must carry the fund)
Implication: If you think you'll exit for $150M, that VC will pass or price you accordingly.
```
This is why Series A investors rarely lead rounds where they can't see a $300M+ exit path. It's not about your business being bad — it's about fund math.
### Comparable Company Analysis
For later stages (Series B+):
```
1. Find 5-10 comparable public SaaS companies
2. Calculate their EV/NTM Revenue multiples (use latest data)
3. Apply a private market discount (typically 20-40% vs public comps)
4. Adjust for your growth rate relative to comps
Example (2024):
Public SaaS comps: 6x NTM Revenue (median)
Private discount: 30%
Adjusted: ~4.2x
Your NTM Revenue: $8M
Implied valuation: ~$33M pre-money
```
### DCF (Late Stage Only)
DCF is unreliable for early-stage startups (terminal value dominates, growth rate assumptions are fantasy). Use it as a sanity check at Series C+, not as the primary valuation method.
---
## 4. Term Sheet Breakdown
### Liquidation Preference (Most Important Economic Term)
This determines who gets paid first in an exit — and how much.
```
1x Non-Participating Preferred (BEST for founders):
Investor gets 1x money back OR converts to common (their choice).
At acquisition: investor takes larger of {1x invested} or {% ownership × proceeds}
Example: $10M invested, exits at $100M, owns 20%
Option A: $10M (1x)
Option B: $20M (20% of $100M)
Investor takes $20M. Founders split $80M.
1x Participating Preferred (WORSE for founders):
Investor gets 1x money back AND participates in remaining proceeds.
Example: same scenario
$10M (1x) + 20% of remaining $90M = $10M + $18M = $28M
Founders split $72M instead of $80M
Cost to founders: $8M (10% of exit value)
2x Participating (RED FLAG):
Investor gets 2x back AND participates.
Only accept under duress. Push hard against this.
Full Ratchet Anti-Dilution (AVOID):
Down-round triggers full repricing of investor shares to new (lower) price.
Founders get massively diluted. Never accept if alternatives exist.
```
### Anti-Dilution Protection
```
Broad-based weighted average (standard):
Adjusts investor conversion price based on all dilutive securities.
Most founder-friendly anti-dilution. Accept this.
Narrow-based weighted average (slightly worse):
Same mechanism but uses smaller denominator.
Gives investors slightly more protection. Usually acceptable.
Full ratchet (avoid):
Price drops to whatever the new round prices at.
Devastating in down rounds. Fight this.
```
### Pro-Rata Rights
```
Standard pro-rata: Investor can maintain their % ownership in future rounds.
Reasonable. Accept for major investors.
Super pro-rata: Investor can increase their % in future rounds.
Caps your ability to bring in new lead investors.
Avoid unless the investor is exceptional and you want them in future rounds.
Major investor threshold: Typically investors with > $500K–$1M check get pro-rata.
Don't give pro-rata to every small check — clogs future rounds.
```
### Board Composition
```
Seed (3 members): 2 founders, 1 lead investor
Series A (5 members): 2 founders, 2 investors, 1 independent
Series B (5-7 seats): Watch for investor majority — negotiate hard
Rule: Founders should retain majority through Series A.
Independent director should be your choice, not investor's.
Never accept investor majority before Series C.
Board observer rights: Common for smaller investors. No vote but present in meetings.
Limit to 1-2 observers or meetings become unwieldy.
```
### Other Terms That Matter
```
Drag-along: Majority can force minority shareholders to vote for acquisition.
Standard and reasonable. Check what threshold triggers drag.
Information rights: Investors get financial statements.
Standard. Monthly for major investors, quarterly for others.
Redemption rights: Investors can force buyback after X years.
Push to remove or add carve-outs for insufficient funds.
No-shop clause: You can't shop the term sheet to other investors.
Standard (14-30 days). Reasonable.
Exclusivity: Stronger version of no-shop. Sometimes includes no other fundraise discussions.
Acceptable for 30 days; push back on > 45 days.
```
---
## 5. Cap Table Management
### Dilution Planning Model
Run this before every round. Know your number before walking into any negotiation.
```
Pre-Seed Post-Seed Post-A Post-B Post-C
Founder A 45.0% 36.0% 26.5% 21.2% 18.7%
Founder B 45.0% 36.0% 26.5% 21.2% 18.7%
Angel 1 5.0% 4.0% 2.9% 2.4% 2.1%
Angel 2 5.0% 4.0% 2.9% 2.4% 2.1%
Seed Fund - 12.0% 8.8% 7.1% 6.2%
Option Pool - 8.0% 12.0% 10.0% 8.0%
Series A - - 20.4% 16.3% 14.4%
Series B - - - 19.5% 17.2%
Series C - - - - 12.6%
Round size / pre-money:
Pre-Seed: $500K / $9M pre = 5% dilution
Seed: $2M / $8M pre = 20% dilution (includes 8% pool)
Series A: $10M / $38M pre = 20.8% dilution (pool refresh to 12%)
Series B: $20M / $80M pre = 20% dilution
Series C: $30M / $170M pre = 15% dilution
```
**Option pool shuffle:** Investors often require you to create/expand the option pool *before* the round closes, which dilutes existing shareholders (not the incoming investor). Model this explicitly — a 20% round with a 5% pool expansion is really 24%+ dilution to founders.
### Cap Table Hygiene
```
Tools: Carta, Pulley, Capshare (all acceptable)
Never: Track cap table in a spreadsheet past seed stage. Errors compound.
Keep it clean:
- Repurchase departed co-founder shares immediately (don't let unvested shares linger)
- Convert SAFEs to equity cleanly at each priced round
- Document every grant with a board resolution
- Cliff + vesting for ALL employees and founders (standard: 1-year cliff, 4-year vest)
- 409A valuation required before every option grant (IRS requirement)
```
---
## 6. Data Room Preparation
### Core Documents (Required)
```
Financial:
□ 3 years historical financials (or all history if < 3 years)
□ Monthly P&L and cash flow (last 24 months)
□ Current financial model (18-24 months forward)
□ Budget vs actual (last 4 quarters)
□ Cap table (fully diluted, with all SAFEs/convertibles modeled)
□ Bank statements (last 3-6 months)
Legal:
□ Certificate of incorporation + all amendments
□ All prior financing documents (SAFEs, convertible notes, stock purchase agreements)
□ Cap table (Carta/Pulley export)
□ IP assignment agreements (all founders and employees)
□ Material contracts (top 10 customers, key vendors)
□ Employee list (titles, start dates, salaries, equity grants)
Product & Business:
□ Product demo / walkthrough video
□ Architecture overview (for technical investors)
□ Customer case studies (3-5 named references)
□ NPS / CSAT data
□ Competitive landscape analysis
Metrics:
□ MRR/ARR by month (all history)
□ Cohort retention chart
□ CAC by channel
□ LTV by cohort
□ NPS trend
```
### What Investors Actually Check First
In order of typical priority during due diligence:
1. **Cap table** — Is it clean? Any concerning structures?
2. **Cohort retention** — Is churn improving or deteriorating?
3. **Revenue quality** — What % is recurring? Any one-time or non-recurring?
4. **Top 10 customers** — Concentration risk? Any logos at risk?
5. **Bank statements** — Does cash match what was reported?
6. **IP assignments** — Does the company own its IP? (Founders who didn't assign IP kill deals)
### Red Flags That Kill Deals
- Missing IP assignment agreements for founders (most common deal killer at early stage)
- Cap table with > 20 angels/small investors (messy, hard to get consent for future rounds)
- Customer concentration > 30% in single customer without explanation
- Revenue recognition issues (booking ARR on contracts that allow easy cancellation)
- Cohort data that gets worse in later cohorts
- Bank balance doesn't match reported cash position
---
## 7. Investor Communication Cadence
### During Fundraise
```
Week 1-2: Warm intro sourcing, LP/network mapping
Week 3-6: First meetings (aim for 20-30 first meetings)
Week 7-10: Partner meetings, deep dives, due diligence
Week 11-14: Term sheets, negotiation
Week 15-18: Legal, closing
```
**Parallel process is essential.** Never negotiate with one investor at a time. Competition is your leverage.
### Post-Close: Investor Updates
Monthly investor update (send within 10 days of month-end):
```
Subject: [Company] Monthly Update — [Month Year]
Highlights (3 bullets max):
• [Biggest win]
• [Biggest learning/challenge]
• [What we're focused on next month]
Metrics:
ARR: $X (+X% MoM)
Net new ARR: $X
Gross margin: X%
Cash: $X (X months runway)
Headcount: X
Asks (be specific):
• Looking for intro to [persona/company] for [specific reason]
• Need advisor with experience in [specific area]
• [Other concrete ask]
```
**Why this matters:** Investors who are informed and engaged are better positioned to help when you need it. The investor who hasn't heard from you in 6 months is less likely to write a bridge check or make a warm intro when you ask.
---
## Key Formulas
```python
# Post-money valuation
post_money = pre_money + investment_amount
# Investor ownership %
ownership_pct = investment_amount / post_money
# Dilution to existing shareholders
dilution = investment_amount / post_money # as a fraction
# New shares issued
new_shares = (investment_amount / post_money) * total_post_shares
# equivalent: new_shares = pre_money_shares * (investment_amount / pre_money)
# Option pool expansion impact (pool shuffle)
# Creating X% option pool pre-close dilutes founders:
pool_shares_needed = target_pct * (pre_shares + new_round_shares + pool_shares_needed)
# Solve: pool_shares_needed = target_pct * (pre_shares + new_round_shares) / (1 - target_pct)
# LTV:CAC ratio
ltv_cac = ltv / cac # target: > 3x
# CAC payback (months)
payback_months = cac / (arpa * gross_margin_pct)
```
FILE:scripts/burn_rate_calculator.py
#!/usr/bin/env python3
"""
Burn Rate & Runway Calculator
==============================
Models startup runway across base/bull/bear scenarios, incorporating
a hiring plan and revenue trajectory. Outputs months of runway,
cash-out dates, and decision trigger points.
Usage:
python burn_rate_calculator.py
python burn_rate_calculator.py --csv # export to CSV
Stdlib only. No dependencies.
"""
import argparse
import csv
import io
import sys
from dataclasses import dataclass, field
from datetime import date, timedelta
from typing import Optional
# ---------------------------------------------------------------------------
# Data structures
# ---------------------------------------------------------------------------
@dataclass
class HiringEntry:
"""A planned hire."""
month: int # months from model start (1-indexed)
role: str
department: str # "sales", "engineering", "cs", "ga"
annual_salary: float
benefits_pct: float = 0.22 # benefits as % of salary
recruiting_cost: float = 0.0 # one-time recruiting fee
@dataclass
class RevenueEntry:
"""Monthly revenue data point (historical or projected)."""
month: int
mrr: float # monthly recurring revenue
one_time: float = 0.0
@dataclass
class ModelConfig:
"""Master configuration for a runway scenario."""
name: str
starting_cash: float
starting_mrr: float
starting_headcount: int
avg_loaded_salary: float # average fully-loaded salary per current employee
base_non_headcount_opex: float # monthly non-headcount costs (infra, tools, etc.)
gross_margin_pct: float # 0.0–1.0
mrr_growth_rate: float # monthly MoM growth rate, 0.0–1.0
hiring_plan: list[HiringEntry] = field(default_factory=list)
model_months: int = 24
start_date: Optional[date] = None
@dataclass
class MonthResult:
"""Single month output."""
month: int
label: str # e.g. "Month 1 (Apr 2025)"
mrr: float
gross_profit: float
headcount: int
headcount_cost: float # total loaded headcount cost this month
other_opex: float
gross_burn: float
net_burn: float
cash_start: float
cash_end: float
runway_months: float # projected runway from this month
cumulative_new_arr: float # for burn multiple
# ---------------------------------------------------------------------------
# Core calculator
# ---------------------------------------------------------------------------
class RunwayCalculator:
def __init__(self, config: ModelConfig):
self.cfg = config
def run(self) -> list[MonthResult]:
cfg = self.cfg
results = []
# Build headcount schedule: month -> list of new hires starting that month
hire_by_month: dict[int, list[HiringEntry]] = {}
for h in cfg.hiring_plan:
hire_by_month.setdefault(h.month, []).append(h)
# Track existing employees
active_employees: list[dict] = []
for _ in range(cfg.starting_headcount):
active_employees.append({
"monthly_loaded": cfg.avg_loaded_salary / 12 * 1.0,
"start_month": 0,
})
cash = cfg.starting_cash
mrr = cfg.starting_mrr
cumulative_new_arr = 0.0
starting_mrr = cfg.starting_mrr
for m in range(1, cfg.model_months + 1):
# Process new hires this month
one_time_recruiting = 0.0
if m in hire_by_month:
for hire in hire_by_month[m]:
monthly_loaded = (
hire.annual_salary * (1 + hire.benefits_pct) / 12
)
active_employees.append({
"monthly_loaded": monthly_loaded,
"start_month": m,
})
one_time_recruiting += hire.recruiting_cost
# Revenue this month
mrr = mrr * (1 + cfg.mrr_growth_rate)
gross_profit = mrr * cfg.gross_margin_pct
# Headcount cost
headcount_cost = sum(e["monthly_loaded"] for e in active_employees)
headcount_cost += one_time_recruiting
# Other opex (infra, SaaS tools, office, etc.)
other_opex = cfg.base_non_headcount_opex
# Burn
gross_burn = headcount_cost + other_opex
net_burn = gross_burn - gross_profit
# Cash
cash_start = cash
cash = cash - net_burn
cash_end = cash
# Projected runway from this month (using current net burn rate)
runway = cash_end / net_burn if net_burn > 0 else float("inf")
# Cumulative new ARR (for burn multiple calc)
new_mrr_added = mrr - starting_mrr if m == 1 else mrr - results[-1].mrr
cumulative_new_arr += new_mrr_added * 12
# Label
if cfg.start_date:
month_date = date(
cfg.start_date.year,
cfg.start_date.month,
1,
) + timedelta(days=32 * (m - 1))
month_date = month_date.replace(day=1)
label = f"Month {m:02d} ({month_date.strftime('%b %Y')})"
else:
label = f"Month {m:02d}"
results.append(MonthResult(
month=m,
label=label,
mrr=mrr,
gross_profit=gross_profit,
headcount=len(active_employees),
headcount_cost=headcount_cost,
other_opex=other_opex,
gross_burn=gross_burn,
net_burn=net_burn,
cash_start=cash_start,
cash_end=cash_end,
runway_months=runway,
cumulative_new_arr=cumulative_new_arr,
))
# Stop if cash runs out
if cash_end <= 0:
break
return results
def cash_out_date(self, results: list[MonthResult]) -> Optional[str]:
"""Return the label of the month cash runs out, or None if model survives."""
for r in results:
if r.cash_end <= 0:
return r.label
return None
def burn_multiple(self, results: list[MonthResult]) -> float:
"""Burn multiple = total net burn / total net new ARR over model period."""
total_net_burn = sum(r.net_burn for r in results if r.net_burn > 0)
first_mrr = results[0].mrr / (1 + self.cfg.mrr_growth_rate) # starting mrr
total_new_arr = (results[-1].mrr - first_mrr) * 12
if total_new_arr <= 0:
return float("inf")
return total_net_burn / total_new_arr
# ---------------------------------------------------------------------------
# Reporting
# ---------------------------------------------------------------------------
def fmt_k(value: float) -> str:
"""Format as $Xk or $X.XM."""
if abs(value) >= 1_000_000:
return f".2fM"
if abs(value) >= 1_000:
return f".0fK"
return f".0f"
def print_summary(name: str, results: list[MonthResult], calc: RunwayCalculator) -> None:
cash_out = calc.cash_out_date(results)
bm = calc.burn_multiple(results)
last = results[-1]
first = results[0]
print(f"\n{'='*60}")
print(f" SCENARIO: {name}")
print(f"{'='*60}")
print(f" Months modeled: {len(results)}")
print(f" Cash out: {cash_out or 'Does not run out in model period'}")
print(f" Ending cash: {fmt_k(last.cash_end)}")
print(f" Final runway: {last.runway_months:.1f} months")
print(f" Starting MRR: {fmt_k(first.mrr)}")
print(f" Ending MRR: {fmt_k(last.mrr)}")
print(f" Ending headcount: {last.headcount}")
print(f" Burn multiple: {bm:.2f}x")
print(f" Avg net burn: {fmt_k(sum(r.net_burn for r in results)/len(results))}/mo")
# Decision triggers
print(f"\n Decision Triggers:")
triggers = {9: "⚠️ START FUNDRAISE", 6: "🔴 COST REDUCTION PLAN", 4: "🚨 EXECUTE CUTS / BRIDGE"}
shown = set()
for r in results:
for threshold, label in triggers.items():
if r.runway_months <= threshold and threshold not in shown:
print(f" {r.label}: {label} (runway = {r.runway_months:.1f} mo)")
shown.add(threshold)
def print_monthly_table(results: list[MonthResult], max_rows: int = 24) -> None:
header = f"{'Month':<22} {'MRR':>10} {'Hdct':>6} {'Net Burn':>12} {'Cash':>12} {'Runway':>8}"
print(f"\n{header}")
print("-" * len(header))
for r in results[:max_rows]:
runway_str = f"{r.runway_months:.1f}mo" if r.runway_months != float("inf") else "∞"
print(
f"{r.label:<22} "
f"{fmt_k(r.mrr):>10} "
f"{r.headcount:>6} "
f"{fmt_k(r.net_burn):>12} "
f"{fmt_k(r.cash_end):>12} "
f"{runway_str:>8}"
)
def export_csv(scenarios: list[tuple[str, list[MonthResult]]]) -> str:
buf = io.StringIO()
writer = csv.writer(buf)
writer.writerow([
"Scenario", "Month", "Label", "MRR", "Gross Profit", "Headcount",
"Headcount Cost", "Other Opex", "Gross Burn", "Net Burn",
"Cash Start", "Cash End", "Runway Months"
])
for name, results in scenarios:
for r in results:
writer.writerow([
name, r.month, r.label,
round(r.mrr, 2), round(r.gross_profit, 2), r.headcount,
round(r.headcount_cost, 2), round(r.other_opex, 2),
round(r.gross_burn, 2), round(r.net_burn, 2),
round(r.cash_start, 2), round(r.cash_end, 2),
round(r.runway_months, 2),
])
return buf.getvalue()
# ---------------------------------------------------------------------------
# Sample data
# ---------------------------------------------------------------------------
def make_sample_configs() -> list[ModelConfig]:
"""
Sample company: Series A SaaS startup
- $3M cash on hand (post Series A)
- $125K MRR (~$1.5M ARR)
- 18 employees, $150K avg salary
- $80K/mo non-headcount opex (infra, tools, office)
- 72% gross margin
"""
common_kwargs = dict(
starting_cash=3_000_000,
starting_mrr=125_000,
starting_headcount=18,
avg_loaded_salary=150_000,
base_non_headcount_opex=80_000,
gross_margin_pct=0.72,
model_months=24,
start_date=date(2025, 1, 1),
)
# Base: 10% MoM growth, moderate hiring
base_hiring = [
HiringEntry(month=2, role="AE #1", department="sales", annual_salary=120_000, recruiting_cost=18_000),
HiringEntry(month=3, role="Senior SWE #1", department="engineering", annual_salary=160_000, recruiting_cost=24_000),
HiringEntry(month=5, role="SDR #1", department="sales", annual_salary=80_000, recruiting_cost=12_000),
HiringEntry(month=6, role="CSM #1", department="cs", annual_salary=90_000, recruiting_cost=13_500),
HiringEntry(month=8, role="AE #2", department="sales", annual_salary=120_000, recruiting_cost=18_000),
HiringEntry(month=9, role="Senior SWE #2", department="engineering", annual_salary=165_000, recruiting_cost=24_750),
HiringEntry(month=12, role="Controller", department="ga", annual_salary=130_000, recruiting_cost=19_500),
HiringEntry(month=14, role="AE #3", department="sales", annual_salary=125_000, recruiting_cost=18_750),
HiringEntry(month=15, role="ML Engineer", department="engineering", annual_salary=175_000, recruiting_cost=26_250),
HiringEntry(month=18, role="AE #4", department="sales", annual_salary=125_000, recruiting_cost=18_750),
]
# Bull: 15% MoM growth, full hiring plan
bull_hiring = base_hiring + [
HiringEntry(month=4, role="Marketing Manager", department="sales", annual_salary=110_000, recruiting_cost=16_500),
HiringEntry(month=7, role="Senior SWE #3", department="engineering", annual_salary=165_000, recruiting_cost=24_750),
HiringEntry(month=10, role="AE #5", department="sales", annual_salary=125_000, recruiting_cost=18_750),
HiringEntry(month=13, role="DevOps Engineer", department="engineering", annual_salary=150_000, recruiting_cost=22_500),
HiringEntry(month=16, role="AE #6", department="sales", annual_salary=125_000, recruiting_cost=18_750),
]
# Bear: 5% MoM growth, hiring freeze after month 3
bear_hiring = [
HiringEntry(month=2, role="AE #1", department="sales", annual_salary=120_000, recruiting_cost=18_000),
HiringEntry(month=3, role="Senior SWE #1", department="engineering", annual_salary=160_000, recruiting_cost=24_000),
]
return [
ModelConfig(name="BULL (15% MoM, full hiring)", mrr_growth_rate=0.15, hiring_plan=bull_hiring, **common_kwargs),
ModelConfig(name="BASE (10% MoM, planned hiring)", mrr_growth_rate=0.10, hiring_plan=base_hiring, **common_kwargs),
ModelConfig(name="BEAR ( 5% MoM, hiring freeze M3+)", mrr_growth_rate=0.05, hiring_plan=bear_hiring, **common_kwargs),
ModelConfig(name="DISTRESS (0% growth, freeze now)", mrr_growth_rate=0.00, hiring_plan=[], **common_kwargs),
]
# ---------------------------------------------------------------------------
# Entry point
# ---------------------------------------------------------------------------
def main() -> None:
parser = argparse.ArgumentParser(description="Startup Burn Rate & Runway Calculator")
parser.add_argument("--csv", action="store_true", help="Export full monthly data as CSV to stdout")
parser.add_argument("--scenario", choices=["bull", "base", "bear", "distress", "all"], default="all")
args = parser.parse_args()
configs = make_sample_configs()
if args.scenario != "all":
configs = [c for c in configs if args.scenario.upper() in c.name.upper()]
all_results: list[tuple[str, list[MonthResult]]] = []
print("\n" + "="*60)
print(" BURN RATE & RUNWAY CALCULATOR")
print(" Sample Company: Series A SaaS Startup")
print(" Starting cash: $3M | Starting MRR: $125K | 18 employees")
print("="*60)
for cfg in configs:
calc = RunwayCalculator(cfg)
results = calc.run()
all_results.append((cfg.name, results))
print_summary(cfg.name, results, calc)
print_monthly_table(results)
# Comparison summary
print("\n" + "="*60)
print(" SCENARIO COMPARISON")
print("="*60)
print(f" {'Scenario':<40} {'Runway':>8} {'Cash Out':<30} {'Burn Mult':>10}")
print(" " + "-"*88)
for cfg, (name, results) in zip(configs, all_results):
calc = RunwayCalculator(cfg)
cash_out = calc.cash_out_date(results) or "Survives model period"
bm = calc.burn_multiple(results)
final_runway = results[-1].runway_months
runway_str = f"{final_runway:.1f}mo" if final_runway != float("inf") else "∞"
bm_str = f"{bm:.2f}x" if bm != float("inf") else "∞"
print(f" {name:<40} {runway_str:>8} {cash_out:<30} {bm_str:>10}")
print("\n Decision Trigger Reference:")
print(" 9 months runway → Start fundraise process")
print(" 6 months runway → Begin cost reduction planning")
print(" 4 months runway → Execute cuts; explore bridge financing")
print(" 3 months runway → Emergency plan only")
if args.csv:
print("\n\n--- CSV EXPORT ---\n")
sys.stdout.write(export_csv(all_results))
if __name__ == "__main__":
main()
FILE:scripts/fundraising_model.py
#!/usr/bin/env python3
"""
Fundraising Model
==================
Cap table management, dilution modeling, and multi-round scenario planning.
Know exactly what you're giving up before you walk into any negotiation.
Covers:
- Cap table state at each round
- Dilution per shareholder per round
- Option pool shuffle impact
- Multi-round projections (Seed → A → B → C)
- Return scenarios at different exit valuations
Usage:
python fundraising_model.py
python fundraising_model.py --exit 150 # model at $150M exit
python fundraising_model.py --csv
Stdlib only. No dependencies.
"""
import argparse
import csv
import io
import sys
from dataclasses import dataclass, field
from typing import Optional
# ---------------------------------------------------------------------------
# Data structures
# ---------------------------------------------------------------------------
@dataclass
class Shareholder:
"""A shareholder in the cap table."""
name: str
share_class: str # "common", "preferred", "option"
shares: float
invested: float = 0.0 # total cash invested
is_option_pool: bool = False
@dataclass
class RoundConfig:
"""Configuration for a financing round."""
name: str # e.g. "Series A"
pre_money_valuation: float
investment_amount: float
new_option_pool_pct: float = 0.0 # % of POST-money to allocate to new options
option_pool_pre_round: bool = True # True = pool created before round (dilutes founders)
lead_investor_name: str = "New Investor"
share_price_override: Optional[float] = None # if None, computed from valuation
@dataclass
class CapTableEntry:
"""A row in the cap table at a point in time."""
name: str
share_class: str
shares: float
pct_ownership: float
invested: float
is_option_pool: bool = False
@dataclass
class RoundResult:
"""Snapshot of cap table after a round closes."""
round_name: str
pre_money_valuation: float
investment_amount: float
post_money_valuation: float
price_per_share: float
new_shares_issued: float
option_pool_shares_created: float
total_shares: float
cap_table: list[CapTableEntry]
@dataclass
class ExitAnalysis:
"""Proceeds to each shareholder at an exit."""
exit_valuation: float
shareholder: str
shares: float
ownership_pct: float
proceeds_common: float # if all preferred converts to common
invested: float
moic: float # multiple on invested capital (for investors)
# ---------------------------------------------------------------------------
# Core cap table engine
# ---------------------------------------------------------------------------
class CapTable:
"""Manages a cap table through multiple rounds."""
def __init__(self):
self.shareholders: list[Shareholder] = []
self._total_shares: float = 0.0
def add_shareholder(self, sh: Shareholder) -> None:
self.shareholders.append(sh)
self._total_shares += sh.shares
def total_shares(self) -> float:
return sum(s.shares for s in self.shareholders)
def snapshot(self, label: str = "") -> list[CapTableEntry]:
total = self.total_shares()
return [
CapTableEntry(
name=s.name,
share_class=s.share_class,
shares=s.shares,
pct_ownership=s.shares / total if total > 0 else 0,
invested=s.invested,
is_option_pool=s.is_option_pool,
)
for s in self.shareholders
]
def execute_round(self, config: RoundConfig) -> RoundResult:
"""
Execute a financing round:
1. (Optional) Create option pool pre-round (dilutes existing shareholders)
2. Issue new shares to investor at round price
Returns a RoundResult with full cap table snapshot.
"""
current_total = self.total_shares()
# Step 1: Option pool shuffle (if pre-round)
option_pool_shares_created = 0.0
if config.new_option_pool_pct > 0 and config.option_pool_pre_round:
# Target: post-round option pool = new_option_pool_pct of total post-money shares
# Solve: pool_shares / (current_total + pool_shares + new_investor_shares) = target_pct
# This requires iteration because new_investor_shares also depends on pool_shares
# Simplification: create pool based on post-round total (slightly approximated)
target_post_round_pct = config.new_option_pool_pct
post_money = config.pre_money_valuation + config.investment_amount
# Estimate shares per dollar (price per share)
price_per_share = config.pre_money_valuation / current_total
new_investor_shares_estimate = config.investment_amount / price_per_share
# Pool shares needed so that pool / total_post = target_pct
total_post_estimate = current_total + new_investor_shares_estimate
pool_shares_needed = (target_post_round_pct * total_post_estimate) / (1 - target_post_round_pct)
# Check if existing pool is sufficient
existing_pool = next(
(s.shares for s in self.shareholders if s.is_option_pool), 0
)
additional_pool_needed = max(0, pool_shares_needed - existing_pool)
if additional_pool_needed > 0:
option_pool_shares_created = additional_pool_needed
# Add to existing pool or create new
pool_sh = next((s for s in self.shareholders if s.is_option_pool), None)
if pool_sh:
pool_sh.shares += additional_pool_needed
else:
self.shareholders.append(Shareholder(
name="Option Pool",
share_class="option",
shares=additional_pool_needed,
is_option_pool=True,
))
# Step 2: Price per share (after pool creation)
current_total_post_pool = self.total_shares()
if config.share_price_override:
price_per_share = config.share_price_override
else:
price_per_share = config.pre_money_valuation / current_total_post_pool
# Step 3: New shares for investor
new_shares = config.investment_amount / price_per_share
# Step 4: Add investor to cap table
self.shareholders.append(Shareholder(
name=config.lead_investor_name,
share_class="preferred",
shares=new_shares,
invested=config.investment_amount,
))
post_money = config.pre_money_valuation + config.investment_amount
total_post = self.total_shares()
return RoundResult(
round_name=config.name,
pre_money_valuation=config.pre_money_valuation,
investment_amount=config.investment_amount,
post_money_valuation=post_money,
price_per_share=price_per_share,
new_shares_issued=new_shares,
option_pool_shares_created=option_pool_shares_created,
total_shares=total_post,
cap_table=self.snapshot(),
)
def analyze_exit(self, exit_valuation: float) -> list[ExitAnalysis]:
"""
Simple exit analysis: all preferred converts to common, proceeds split pro-rata.
(Does not model liquidation preferences — see fundraising_playbook.md for that.)
"""
total = self.total_shares()
price_per_share = exit_valuation / total
results = []
for s in self.shareholders:
if s.is_option_pool:
continue # unissued options don't receive proceeds
proceeds = s.shares * price_per_share
moic = proceeds / s.invested if s.invested > 0 else 0.0
results.append(ExitAnalysis(
exit_valuation=exit_valuation,
shareholder=s.name,
shares=s.shares,
ownership_pct=s.shares / total,
proceeds_common=proceeds,
invested=s.invested,
moic=moic,
))
return sorted(results, key=lambda x: x.proceeds_common, reverse=True)
# ---------------------------------------------------------------------------
# Reporting
# ---------------------------------------------------------------------------
def fmt(value: float, prefix: str = "$") -> str:
if value == float("inf"):
return "∞"
if abs(value) >= 1_000_000:
return f"{prefix}{value/1_000_000:.2f}M"
if abs(value) >= 1_000:
return f"{prefix}{value/1_000:.0f}K"
return f"{prefix}{value:.2f}"
def print_round_result(result: RoundResult, prev_cap_table: Optional[list[CapTableEntry]] = None) -> None:
print(f"\n{'='*70}")
print(f" {result.round_name.upper()}")
print(f"{'='*70}")
print(f" Pre-money valuation: {fmt(result.pre_money_valuation)}")
print(f" Investment: {fmt(result.investment_amount)}")
print(f" Post-money valuation: {fmt(result.post_money_valuation)}")
print(f" Price per share: {fmt(result.price_per_share, '$')}")
print(f" New shares issued: {result.new_shares_issued:,.0f}")
if result.option_pool_shares_created > 0:
print(f" Option pool created: {result.option_pool_shares_created:,.0f} shares")
print(f" ⚠️ Pool created pre-round: dilutes existing shareholders, not new investor")
print(f" Total shares post: {result.total_shares:,.0f}")
print(f"\n {'Shareholder':<22} {'Shares':>12} {'Ownership':>10} {'Invested':>10} {'Δ Ownership':>12}")
print(" " + "-"*68)
prev_map = {e.name: e.pct_ownership for e in prev_cap_table} if prev_cap_table else {}
for entry in result.cap_table:
delta = ""
if entry.name in prev_map:
change = (entry.pct_ownership - prev_map[entry.name]) * 100
delta = f"{change:+.1f}pp"
elif not entry.is_option_pool:
delta = "new"
invested_str = fmt(entry.invested) if entry.invested > 0 else "-"
print(
f" {entry.name:<22} {entry.shares:>12,.0f} "
f"{entry.pct_ownership*100:>9.2f}% {invested_str:>10} {delta:>12}"
)
def print_exit_analysis(results: list[ExitAnalysis], exit_valuation: float) -> None:
print(f"\n{'='*70}")
print(f" EXIT ANALYSIS @ {fmt(exit_valuation)} (all preferred converts to common)")
print(f"{'='*70}")
print(f"\n {'Shareholder':<22} {'Ownership':>10} {'Proceeds':>12} {'Invested':>10} {'MOIC':>8}")
print(" " + "-"*65)
for r in results:
moic_str = f"{r.moic:.1f}x" if r.moic > 0 else "n/a"
invested_str = fmt(r.invested) if r.invested > 0 else "-"
print(
f" {r.shareholder:<22} {r.ownership_pct*100:>9.2f}% "
f"{fmt(r.proceeds_common):>12} {invested_str:>10} {moic_str:>8}"
)
print(f"\n Note: Does not model liquidation preferences.")
print(f" Participating preferred reduces founder proceeds in most real exits.")
print(f" See references/fundraising_playbook.md for full liquidation waterfall.")
def print_dilution_summary(rounds: list[RoundResult]) -> None:
print(f"\n{'='*70}")
print(f" DILUTION SUMMARY — FOUNDER PERSPECTIVE")
print(f"{'='*70}")
# Find all founders (common shareholders who aren't investors or option pool)
founder_names = []
for entry in rounds[0].cap_table:
if entry.share_class == "common" and not entry.is_option_pool:
founder_names.append(entry.name)
if not founder_names:
print(" No common shareholders found in initial cap table.")
return
header = f" {'Round':<16}" + "".join(f" {n:<16}" for n in founder_names) + f" {'Total Inv':>12}"
print(header)
print(" " + "-" * (16 + 18 * len(founder_names) + 14))
for result in rounds:
cap_map = {e.name: e for e in result.cap_table}
total_invested = sum(e.invested for e in result.cap_table if not e.is_option_pool)
row = f" {result.round_name:<16}"
for name in founder_names:
pct = cap_map[name].pct_ownership * 100 if name in cap_map else 0
row += f" {pct:>6.2f}% "
row += f" {fmt(total_invested):>12}"
print(row)
def export_csv_rounds(rounds: list[RoundResult]) -> str:
buf = io.StringIO()
writer = csv.writer(buf)
writer.writerow(["Round", "Shareholder", "Share Class", "Shares", "Ownership Pct",
"Invested", "Pre Money", "Post Money", "Price Per Share"])
for r in rounds:
for entry in r.cap_table:
writer.writerow([
r.round_name, entry.name, entry.share_class,
round(entry.shares, 0), round(entry.pct_ownership * 100, 4),
round(entry.invested, 2), round(r.pre_money_valuation, 0),
round(r.post_money_valuation, 0), round(r.price_per_share, 4),
])
return buf.getvalue()
# ---------------------------------------------------------------------------
# Sample data: typical two-founder Series A/B/C startup
# ---------------------------------------------------------------------------
def build_sample_model() -> tuple[CapTable, list[RoundResult]]:
"""
Sample company:
- 2 founders, started with 10M shares each
- 1M shares for early advisor
- Raises Pre-seed → Seed → Series A → Series B → Series C
"""
cap = CapTable()
SHARES_PER_FOUNDER = 4_000_000
SHARES_ADVISOR = 200_000
# Founding state
cap.add_shareholder(Shareholder("Founder A (CEO)", "common", SHARES_PER_FOUNDER))
cap.add_shareholder(Shareholder("Founder B (CTO)", "common", SHARES_PER_FOUNDER))
cap.add_shareholder(Shareholder("Advisor", "common", SHARES_ADVISOR))
rounds: list[RoundResult] = []
prev_cap = cap.snapshot()
# Round 1: Pre-seed — $500K at $4.5M pre, 10% option pool created
r1 = cap.execute_round(RoundConfig(
name="Pre-seed",
pre_money_valuation=4_500_000,
investment_amount=500_000,
new_option_pool_pct=0.10,
option_pool_pre_round=True,
lead_investor_name="Angel Syndicate",
))
rounds.append(r1)
prev_r1 = r1.cap_table[:]
# Round 2: Seed — $2M at $9M pre, expand option pool to 12%
r2 = cap.execute_round(RoundConfig(
name="Seed",
pre_money_valuation=9_000_000,
investment_amount=2_000_000,
new_option_pool_pct=0.12,
option_pool_pre_round=True,
lead_investor_name="Seed Fund",
))
rounds.append(r2)
# Round 3: Series A — $12M at $38M pre, refresh option pool to 15%
r3 = cap.execute_round(RoundConfig(
name="Series A",
pre_money_valuation=38_000_000,
investment_amount=12_000_000,
new_option_pool_pct=0.15,
option_pool_pre_round=True,
lead_investor_name="Series A Fund",
))
rounds.append(r3)
# Round 4: Series B — $25M at $95M pre, refresh pool to 12%
r4 = cap.execute_round(RoundConfig(
name="Series B",
pre_money_valuation=95_000_000,
investment_amount=25_000_000,
new_option_pool_pct=0.12,
option_pool_pre_round=True,
lead_investor_name="Series B Fund",
))
rounds.append(r4)
# Round 5: Series C — $40M at $185M pre, refresh pool to 10%
r5 = cap.execute_round(RoundConfig(
name="Series C",
pre_money_valuation=185_000_000,
investment_amount=40_000_000,
new_option_pool_pct=0.10,
option_pool_pre_round=True,
lead_investor_name="Series C Fund",
))
rounds.append(r5)
return cap, rounds
# ---------------------------------------------------------------------------
# Entry point
# ---------------------------------------------------------------------------
def main() -> None:
parser = argparse.ArgumentParser(description="Fundraising Model — Cap Table & Dilution")
parser.add_argument("--exit", type=float, default=250.0,
help="Exit valuation in $M for return analysis (default: 250)")
parser.add_argument("--csv", action="store_true", help="Export round data as CSV to stdout")
args = parser.parse_args()
exit_valuation = args.exit * 1_000_000
print("\n" + "="*70)
print(" FUNDRAISING MODEL — CAP TABLE & DILUTION ANALYSIS")
print(" Sample Company: Two-founder SaaS startup")
print(" Pre-seed → Seed → Series A → Series B → Series C")
print("="*70)
cap, rounds = build_sample_model()
# Print each round
prev = None
for r in rounds:
print_round_result(r, prev)
prev = r.cap_table
# Dilution summary table
print_dilution_summary(rounds)
# Exit analysis at specified valuation
exit_results = cap.analyze_exit(exit_valuation)
print_exit_analysis(exit_results, exit_valuation)
# Also print at 2x and 5x for sensitivity
print("\n Exit Sensitivity — Founder A Proceeds:")
print(f" {'Exit Valuation':<20} {'Founder A %':>12} {'Founder A $':>14} {'MOIC':>8}")
print(" " + "-"*56)
for mult in [0.5, 1.0, 1.5, 2.0, 3.0, 5.0]:
val = rounds[-1].post_money_valuation * mult
ex = cap.analyze_exit(val)
founder_a = next((r for r in ex if r.shareholder == "Founder A (CEO)"), None)
if founder_a:
print(f" {fmt(val):<20} {founder_a.ownership_pct*100:>11.2f}% "
f"{fmt(founder_a.proceeds_common):>14} {'n/a':>8}")
print("\n Key Takeaways:")
final = rounds[-1].cap_table
total = sum(e.shares for e in final)
founder_a_final = next((e for e in final if e.name == "Founder A (CEO)"), None)
if founder_a_final:
print(f" Founder A final ownership: {founder_a_final.pct_ownership*100:.2f}%")
total_raised = sum(e.invested for e in final)
print(f" Total capital raised: {fmt(total_raised)}")
print(f" Total shares outstanding: {total:,.0f}")
print(f" Final post-money: {fmt(rounds[-1].post_money_valuation)}")
print("\n Run with --exit <$M> to model proceeds at different exit valuations.")
print(" Example: python fundraising_model.py --exit 500")
if args.csv:
print("\n\n--- CSV EXPORT ---\n")
sys.stdout.write(export_csv_rounds(rounds))
if __name__ == "__main__":
main()
FILE:scripts/unit_economics_analyzer.py
#!/usr/bin/env python3
"""
Unit Economics Analyzer
========================
Per-cohort LTV, per-channel CAC, payback periods, and LTV:CAC ratios.
Never blended averages — those hide what's actually happening.
Usage:
python unit_economics_analyzer.py
python unit_economics_analyzer.py --csv
Stdlib only. No dependencies.
"""
import argparse
import csv
import io
import sys
from dataclasses import dataclass, field
from typing import Optional
# ---------------------------------------------------------------------------
# Data structures
# ---------------------------------------------------------------------------
@dataclass
class CohortData:
"""
Revenue data for a group of customers acquired in the same period.
Revenue is tracked monthly: revenue[0] = month 1, revenue[1] = month 2, etc.
"""
label: str # e.g. "Q1 2024"
acquisition_period: str # human-readable label
customers_acquired: int
total_cac_spend: float # total S&M spend to acquire this cohort
monthly_revenue: list[float] # revenue per month from this cohort
gross_margin_pct: float = 0.70 # blended gross margin for this cohort
@dataclass
class ChannelData:
"""Acquisition cost and customer data for a single channel."""
channel: str
spend: float
customers_acquired: int
avg_arpa: float # average revenue per account (monthly)
gross_margin_pct: float = 0.70
avg_monthly_churn: float = 0.02 # monthly churn rate for customers from this channel
@dataclass
class UnitEconomicsResult:
"""Computed unit economics for a cohort or channel."""
label: str
customers: int
cac: float
arpa: float # average revenue per account per month
gross_margin_pct: float
monthly_churn: float
ltv: float
ltv_cac_ratio: float
payback_months: float
# Cohort-specific
m1_revenue: Optional[float] = None
m6_revenue: Optional[float] = None
m12_revenue: Optional[float] = None
m24_revenue: Optional[float] = None
m12_ltv: Optional[float] = None # realized LTV through month 12
retention_m6: Optional[float] = None # % of M1 revenue retained at M6
retention_m12: Optional[float] = None
# ---------------------------------------------------------------------------
# Calculators
# ---------------------------------------------------------------------------
def calc_ltv(arpa: float, gross_margin_pct: float, monthly_churn: float) -> float:
"""
LTV = (ARPA × Gross Margin) / Monthly Churn Rate
Assumes constant churn (simplified; cohort method is more accurate).
"""
if monthly_churn <= 0:
return float("inf")
return (arpa * gross_margin_pct) / monthly_churn
def calc_payback(cac: float, arpa: float, gross_margin_pct: float) -> float:
"""
CAC Payback (months) = CAC / (ARPA × Gross Margin)
"""
denominator = arpa * gross_margin_pct
if denominator <= 0:
return float("inf")
return cac / denominator
def analyze_cohort(cohort: CohortData) -> UnitEconomicsResult:
"""Compute full unit economics for a cohort."""
n = cohort.customers_acquired
if n == 0:
raise ValueError(f"Cohort {cohort.label}: customers_acquired cannot be 0")
cac = cohort.total_cac_spend / n
# ARPA from month 1 revenue
m1_rev = cohort.monthly_revenue[0] if cohort.monthly_revenue else 0
arpa = m1_rev / n if n > 0 else 0
# Observed monthly churn from cohort data
# Use revenue decline from M1 to M12 to estimate churn
months_available = len(cohort.monthly_revenue)
if months_available >= 12:
m12_rev = cohort.monthly_revenue[11]
# Revenue retention over 12 months: (M12/M1)^(1/11) per month on average
# Implied monthly retention rate
if m1_rev > 0 and m12_rev > 0:
monthly_retention = (m12_rev / m1_rev) ** (1 / 11)
monthly_churn = 1 - monthly_retention
else:
monthly_churn = 0.02 # default
elif months_available >= 6:
m6_rev = cohort.monthly_revenue[5]
if m1_rev > 0 and m6_rev > 0:
monthly_retention = (m6_rev / m1_rev) ** (1 / 5)
monthly_churn = 1 - monthly_retention
else:
monthly_churn = 0.02
else:
monthly_churn = 0.02 # default if < 6 months data
# Clamp to reasonable range
monthly_churn = max(0.001, min(monthly_churn, 0.30))
ltv = calc_ltv(arpa, cohort.gross_margin_pct, monthly_churn)
payback = calc_payback(cac, arpa, cohort.gross_margin_pct)
ltv_cac = ltv / cac if cac > 0 else float("inf")
# Snapshot revenues
def rev_at(month_idx: int) -> Optional[float]:
if months_available > month_idx:
return cohort.monthly_revenue[month_idx]
return None
m6 = rev_at(5)
m12 = rev_at(11)
m24 = rev_at(23)
# Realized LTV through observed months (actual gross profit)
m12_ltv = sum(cohort.monthly_revenue[:12]) * cohort.gross_margin_pct if months_available >= 12 else None
# Retention rates
ret_m6 = (m6 / m1_rev) if (m6 is not None and m1_rev > 0) else None
ret_m12 = (m12 / m1_rev) if (m12 is not None and m1_rev > 0) else None
return UnitEconomicsResult(
label=cohort.label,
customers=n,
cac=cac,
arpa=arpa,
gross_margin_pct=cohort.gross_margin_pct,
monthly_churn=monthly_churn,
ltv=ltv,
ltv_cac_ratio=ltv_cac,
payback_months=payback,
m1_revenue=m1_rev,
m6_revenue=m6,
m12_revenue=m12,
m24_revenue=m24,
m12_ltv=m12_ltv,
retention_m6=ret_m6,
retention_m12=ret_m12,
)
def analyze_channel(ch: ChannelData) -> UnitEconomicsResult:
"""Compute unit economics for an acquisition channel."""
if ch.customers_acquired == 0:
raise ValueError(f"Channel {ch.channel}: customers_acquired cannot be 0")
cac = ch.spend / ch.customers_acquired
ltv = calc_ltv(ch.avg_arpa, ch.gross_margin_pct, ch.avg_monthly_churn)
payback = calc_payback(cac, ch.avg_arpa, ch.gross_margin_pct)
ltv_cac = ltv / cac if cac > 0 else float("inf")
return UnitEconomicsResult(
label=ch.channel,
customers=ch.customers_acquired,
cac=cac,
arpa=ch.avg_arpa,
gross_margin_pct=ch.gross_margin_pct,
monthly_churn=ch.avg_monthly_churn,
ltv=ltv,
ltv_cac_ratio=ltv_cac,
payback_months=payback,
)
# ---------------------------------------------------------------------------
# Blended metrics (for comparison)
# ---------------------------------------------------------------------------
def blended_cac(channels: list[ChannelData]) -> float:
total_spend = sum(c.spend for c in channels)
total_customers = sum(c.customers_acquired for c in channels)
return total_spend / total_customers if total_customers > 0 else 0
def blended_ltv(channels: list[ChannelData]) -> float:
"""Weighted average LTV by customers acquired."""
total_customers = sum(c.customers_acquired for c in channels)
if total_customers == 0:
return 0
weighted = sum(
calc_ltv(c.avg_arpa, c.gross_margin_pct, c.avg_monthly_churn) * c.customers_acquired
for c in channels
)
return weighted / total_customers
# ---------------------------------------------------------------------------
# Reporting
# ---------------------------------------------------------------------------
def fmt(value: float, prefix: str = "$", decimals: int = 0) -> str:
if value == float("inf"):
return "∞"
if abs(value) >= 1_000_000:
return f"{prefix}{value/1_000_000:.2f}M"
if abs(value) >= 1_000:
return f"{prefix}{value/1_000:.1f}K"
return f"{prefix}{value:.{decimals}f}"
def pct(value: Optional[float]) -> str:
if value is None:
return "n/a"
return f"{value*100:.1f}%"
def rating(ltv_cac: float, payback: float) -> str:
if ltv_cac == float("inf"):
return "∞"
if ltv_cac >= 5 and payback <= 12:
return "🟢 Excellent"
if ltv_cac >= 3 and payback <= 18:
return "🟡 Good"
if ltv_cac >= 2 and payback <= 24:
return "🟠 Marginal"
return "🔴 Poor"
def print_cohort_analysis(results: list[UnitEconomicsResult]) -> None:
print("\n" + "="*80)
print(" COHORT ANALYSIS")
print("="*80)
print(f" {'Cohort':<12} {'Cust':>5} {'CAC':>8} {'ARPA/mo':>9} {'Churn/mo':>10} "
f"{'LTV':>10} {'LTV:CAC':>8} {'Payback':>9} {'Ret@M12':>8}")
print(" " + "-"*88)
for r in results:
payback_str = f"{r.payback_months:.1f}mo" if r.payback_months != float("inf") else "∞"
ltv_str = fmt(r.ltv) if r.ltv != float("inf") else "∞"
ltv_cac_str = f"{r.ltv_cac_ratio:.1f}x" if r.ltv_cac_ratio != float("inf") else "∞"
print(
f" {r.label:<12} {r.customers:>5} {fmt(r.cac):>8} {fmt(r.arpa):>9} "
f"{pct(r.monthly_churn):>10} {ltv_str:>10} {ltv_cac_str:>8} "
f"{payback_str:>9} {pct(r.retention_m12):>8}"
)
# Trend analysis
print("\n Cohort Trend (is the business getting better or worse?):")
if len(results) >= 3:
ltv_cac_values = [r.ltv_cac_ratio for r in results if r.ltv_cac_ratio != float("inf")]
cac_values = [r.cac for r in results]
churn_values = [r.monthly_churn for r in results]
if len(ltv_cac_values) >= 2:
ltv_cac_trend = "↑ Improving" if ltv_cac_values[-1] > ltv_cac_values[0] else "↓ Deteriorating"
else:
ltv_cac_trend = "n/a"
cac_trend = "↓ Decreasing (good)" if cac_values[-1] < cac_values[0] else "↑ Increasing"
churn_trend = "↓ Improving" if churn_values[-1] < churn_values[0] else "↑ Worsening"
print(f" LTV:CAC: {ltv_cac_trend}")
print(f" CAC: {cac_trend}")
print(f" Churn rate: {churn_trend}")
def print_channel_analysis(results: list[UnitEconomicsResult], channels: list[ChannelData]) -> None:
print("\n" + "="*80)
print(" CHANNEL ANALYSIS (Per-Channel vs Blended)")
print("="*80)
print(f" {'Channel':<22} {'Spend':>9} {'Cust':>5} {'CAC':>8} {'LTV':>10} {'LTV:CAC':>8} {'Payback':>9} {'Rating'}")
print(" " + "-"*90)
for r, ch in zip(results, channels):
payback_str = f"{r.payback_months:.1f}mo" if r.payback_months != float("inf") else "∞"
ltv_str = fmt(r.ltv) if r.ltv != float("inf") else "∞"
ltv_cac_str = f"{r.ltv_cac_ratio:.1f}x" if r.ltv_cac_ratio != float("inf") else "∞"
print(
f" {r.label:<22} {fmt(ch.spend):>9} {r.customers:>5} {fmt(r.cac):>8} "
f"{ltv_str:>10} {ltv_cac_str:>8} {payback_str:>9} {rating(r.ltv_cac_ratio, r.payback_months)}"
)
# Blended comparison
b_cac = blended_cac(channels)
b_ltv = blended_ltv(channels)
b_ltv_cac = b_ltv / b_cac if b_cac > 0 else 0
total_spend = sum(c.spend for c in channels)
total_customers = sum(c.customers_acquired for c in channels)
avg_payback = sum(
calc_payback(b_cac, c.avg_arpa, c.gross_margin_pct) * c.customers_acquired
for c in channels
) / total_customers
print(" " + "-"*90)
print(
f" {'BLENDED (dangerous)':<22} {fmt(total_spend):>9} {total_customers:>5} "
f"{fmt(b_cac):>8} {fmt(b_ltv):>10} {b_ltv_cac:.1f}x{'':<7} "
f"{avg_payback:.1f}mo{'':<4} {rating(b_ltv_cac, avg_payback)}"
)
print("\n ⚠️ Blended numbers hide channel-level problems. Manage channels individually.")
# Budget reallocation
print("\n Recommended Budget Reallocation:")
sorted_results = sorted(zip(results, channels), key=lambda x: x[0].ltv_cac_ratio, reverse=True)
for r, ch in sorted_results:
if r.ltv_cac_ratio >= 3:
action = "✅ Scale"
elif r.ltv_cac_ratio >= 2:
action = "🔄 Optimize"
else:
action = "❌ Cut / pause"
print(f" {ch.channel:<22} LTV:CAC = {r.ltv_cac_ratio:.1f}x → {action}")
def export_csv_results(cohort_results: list[UnitEconomicsResult], channel_results: list[UnitEconomicsResult]) -> str:
buf = io.StringIO()
writer = csv.writer(buf)
writer.writerow(["Type", "Label", "Customers", "CAC", "ARPA_Monthly", "Gross_Margin_Pct",
"Monthly_Churn", "LTV", "LTV_CAC_Ratio", "Payback_Months",
"Retention_M6", "Retention_M12"])
for r in cohort_results:
writer.writerow(["cohort", r.label, r.customers, round(r.cac, 2), round(r.arpa, 2),
r.gross_margin_pct, round(r.monthly_churn, 4),
round(r.ltv, 2) if r.ltv != float("inf") else "inf",
round(r.ltv_cac_ratio, 2) if r.ltv_cac_ratio != float("inf") else "inf",
round(r.payback_months, 2) if r.payback_months != float("inf") else "inf",
round(r.retention_m6, 3) if r.retention_m6 else "",
round(r.retention_m12, 3) if r.retention_m12 else ""])
for r in channel_results:
writer.writerow(["channel", r.label, r.customers, round(r.cac, 2), round(r.arpa, 2),
r.gross_margin_pct, round(r.monthly_churn, 4),
round(r.ltv, 2) if r.ltv != float("inf") else "inf",
round(r.ltv_cac_ratio, 2) if r.ltv_cac_ratio != float("inf") else "inf",
round(r.payback_months, 2) if r.payback_months != float("inf") else "inf",
"", ""])
return buf.getvalue()
# ---------------------------------------------------------------------------
# Sample data
# ---------------------------------------------------------------------------
def make_sample_cohorts() -> list[CohortData]:
"""
Series A SaaS company, 8 quarters of cohort data.
Shows a business improving on all dimensions over time.
"""
return [
CohortData(
label="Q1 2023", acquisition_period="Jan-Mar 2023",
customers_acquired=12, total_cac_spend=54_000,
gross_margin_pct=0.68,
monthly_revenue=[
10_200, 9_600, 9_100, 8_700, 8_300, 8_000, # M1-M6
7_800, 7_600, 7_400, 7_200, 7_000, 6_800, # M7-M12
6_700, 6_600, 6_500, 6_400, 6_300, 6_200, # M13-M18
6_100, 6_000, 5_900, 5_800, 5_700, 5_600, # M19-M24
],
),
CohortData(
label="Q2 2023", acquisition_period="Apr-Jun 2023",
customers_acquired=15, total_cac_spend=60_000,
gross_margin_pct=0.69,
monthly_revenue=[
13_500, 12_900, 12_500, 12_100, 11_800, 11_500,
11_300, 11_100, 10_900, 10_700, 10_500, 10_300,
10_200, 10_100, 10_000, 9_900, 9_800, 9_700,
],
),
CohortData(
label="Q3 2023", acquisition_period="Jul-Sep 2023",
customers_acquired=18, total_cac_spend=63_000,
gross_margin_pct=0.70,
monthly_revenue=[
16_200, 15_800, 15_400, 15_100, 14_800, 14_600,
14_400, 14_200, 14_000, 13_900, 13_800, 13_700,
13_600, 13_500, 13_400, 13_300,
],
),
CohortData(
label="Q4 2023", acquisition_period="Oct-Dec 2023",
customers_acquired=22, total_cac_spend=70_400,
gross_margin_pct=0.71,
monthly_revenue=[
20_900, 20_500, 20_200, 19_900, 19_700, 19_500,
19_300, 19_100, 19_000, 18_900, 18_800, 18_700,
],
),
CohortData(
label="Q1 2024", acquisition_period="Jan-Mar 2024",
customers_acquired=28, total_cac_spend=81_200,
gross_margin_pct=0.72,
monthly_revenue=[
27_200, 26_900, 26_600, 26_400, 26_200, 26_000,
25_800, 25_700, 25_600, 25_500,
],
),
CohortData(
label="Q2 2024", acquisition_period="Apr-Jun 2024",
customers_acquired=34, total_cac_spend=91_800,
gross_margin_pct=0.72,
monthly_revenue=[
33_300, 33_000, 32_800, 32_600, 32_400, 32_200,
],
),
CohortData(
label="Q3 2024", acquisition_period="Jul-Sep 2024",
customers_acquired=40, total_cac_spend=100_000,
gross_margin_pct=0.73,
monthly_revenue=[
39_600, 39_400, 39_200,
],
),
CohortData(
label="Q4 2024", acquisition_period="Oct-Dec 2024",
customers_acquired=47, total_cac_spend=112_800,
gross_margin_pct=0.73,
monthly_revenue=[
47_000,
],
),
]
def make_sample_channels() -> list[ChannelData]:
"""
Q4 2024 channel breakdown. Blended looks fine; per-channel reveals problems.
"""
return [
ChannelData("Organic / SEO", spend=9_500, customers_acquired=14, avg_arpa=950, gross_margin_pct=0.73, avg_monthly_churn=0.015),
ChannelData("Paid Search (SEM)", spend=48_000, customers_acquired=18, avg_arpa=980, gross_margin_pct=0.73, avg_monthly_churn=0.020),
ChannelData("Paid Social", spend=32_000, customers_acquired=8, avg_arpa=900, gross_margin_pct=0.72, avg_monthly_churn=0.025),
ChannelData("Content / Inbound", spend=11_000, customers_acquired=6, avg_arpa=1100, gross_margin_pct=0.74, avg_monthly_churn=0.012),
ChannelData("Outbound SDR", spend=22_000, customers_acquired=4, avg_arpa=1200, gross_margin_pct=0.73, avg_monthly_churn=0.022),
ChannelData("Events / Webinars", spend=18_500, customers_acquired=3, avg_arpa=1050, gross_margin_pct=0.72, avg_monthly_churn=0.028),
ChannelData("Partner / Referral", spend=7_800, customers_acquired=7, avg_arpa=1000, gross_margin_pct=0.73, avg_monthly_churn=0.013),
]
# ---------------------------------------------------------------------------
# Entry point
# ---------------------------------------------------------------------------
def main() -> None:
parser = argparse.ArgumentParser(description="Unit Economics Analyzer")
parser.add_argument("--csv", action="store_true", help="Export results as CSV to stdout")
args = parser.parse_args()
cohorts = make_sample_cohorts()
channels = make_sample_channels()
print("\n" + "="*80)
print(" UNIT ECONOMICS ANALYZER")
print(" Sample Company: Series A SaaS | Q4 2024 Snapshot")
print(" Gross Margin: ~72% | Monthly Churn: derived from cohort data")
print("="*80)
cohort_results = [analyze_cohort(c) for c in cohorts]
channel_results = [analyze_channel(c) for c in channels]
print_cohort_analysis(cohort_results)
print_channel_analysis(channel_results, channels)
# Health summary
print("\n" + "="*80)
print(" HEALTH SUMMARY")
print("="*80)
latest = cohort_results[-1]
prev = cohort_results[-4] if len(cohort_results) >= 4 else cohort_results[0]
print(f"\n Latest Cohort ({latest.label}):")
print(f" CAC: {fmt(latest.cac)}")
ltv_str = fmt(latest.ltv) if latest.ltv != float("inf") else "∞"
ltv_cac_str = f"{latest.ltv_cac_ratio:.1f}x" if latest.ltv_cac_ratio != float("inf") else "∞"
payback_str = f"{latest.payback_months:.1f} months" if latest.payback_months != float("inf") else "∞"
print(f" LTV: {ltv_str}")
print(f" LTV:CAC: {ltv_cac_str} (target: > 3x)")
print(f" CAC Payback: {payback_str} (target: < 18mo)")
print(f" Rating: {rating(latest.ltv_cac_ratio, latest.payback_months)}")
# Trend vs 4 quarters ago
print(f"\n Trend vs {prev.label}:")
cac_delta = (latest.cac - prev.cac) / prev.cac * 100
ltv_delta_str = "n/a"
if latest.ltv != float("inf") and prev.ltv != float("inf"):
ltv_delta = (latest.ltv - prev.ltv) / prev.ltv * 100
ltv_delta_str = f"{ltv_delta:+.1f}%"
cac_str = "↓ Better" if cac_delta < 0 else "↑ Worse"
print(f" CAC: {cac_delta:+.1f}% ({cac_str})")
print(f" LTV: {ltv_delta_str}")
print("\n Benchmark Reference:")
print(" LTV:CAC > 5x → Scale aggressively")
print(" LTV:CAC 3-5x → Healthy; grow at current pace")
print(" LTV:CAC 2-3x → Marginal; optimize before scaling")
print(" LTV:CAC < 2x → Acquiring unprofitably; stop and fix")
print(" Payback < 12mo → Outstanding capital efficiency")
print(" Payback 12-18mo → Good for B2B SaaS")
print(" Payback > 24mo → Requires long-dated capital to scale")
if args.csv:
print("\n\n--- CSV EXPORT ---\n")
sys.stdout.write(export_csv_results(cohort_results, channel_results))
if __name__ == "__main__":
main()
Thiết kế chiến lược observability kết hợp metrics, logs, traces, gồm SLI/SLO, golden signals và tối ưu cảnh báo.
---
name: "observability-designer"
description: "Design production-ready observability strategies combining metrics, logs, and traces. Includes SLI/SLO design, golden-signals monitoring, alert optimization. Use when adding observability to a new service, refactoring alerting that is too noisy, or designing an SLO program before scaling production load."
---
# Observability Designer (POWERFUL)
**Category:** Engineering
**Tier:** POWERFUL
**Description:** Design comprehensive observability strategies for production systems including SLI/SLO frameworks, alerting optimization, and dashboard generation.
## Overview
Observability Designer enables you to create production-ready observability strategies that provide deep insights into system behavior, performance, and reliability. This skill combines the three pillars of observability (metrics, logs, traces) with proven frameworks like SLI/SLO design, golden signals monitoring, and alert optimization to create comprehensive observability solutions.
## Core Competencies
### SLI/SLO/SLA Framework Design
- **Service Level Indicators (SLI):** Define measurable signals that indicate service health
- **Service Level Objectives (SLO):** Set reliability targets based on user experience
- **Service Level Agreements (SLA):** Establish customer-facing commitments with consequences
- **Error Budget Management:** Calculate and track error budget consumption
- **Burn Rate Alerting:** Multi-window burn rate alerts for proactive SLO protection
### Three Pillars of Observability
#### Metrics
- **Golden Signals:** Latency, traffic, errors, and saturation monitoring
- **RED Method:** Rate, Errors, and Duration for request-driven services
- **USE Method:** Utilization, Saturation, and Errors for resource monitoring
- **Business Metrics:** Revenue, user engagement, and feature adoption tracking
- **Infrastructure Metrics:** CPU, memory, disk, network, and custom resource metrics
#### Logs
- **Structured Logging:** JSON-based log formats with consistent fields
- **Log Aggregation:** Centralized log collection and indexing strategies
- **Log Levels:** Appropriate use of DEBUG, INFO, WARN, ERROR, FATAL levels
- **Correlation IDs:** Request tracing through distributed systems
- **Log Sampling:** Volume management for high-throughput systems
#### Traces
- **Distributed Tracing:** End-to-end request flow visualization
- **Span Design:** Meaningful span boundaries and metadata
- **Trace Sampling:** Intelligent sampling strategies for performance and cost
- **Service Maps:** Automatic dependency discovery through traces
- **Root Cause Analysis:** Trace-driven debugging workflows
### Dashboard Design Principles
#### Information Architecture
- **Hierarchy:** Overview → Service → Component → Instance drill-down paths
- **Golden Ratio:** 80% operational metrics, 20% exploratory metrics
- **Cognitive Load:** Maximum 7±2 panels per dashboard screen
- **User Journey:** Role-based dashboard personas (SRE, Developer, Executive)
#### Visualization Best Practices
- **Chart Selection:** Time series for trends, heatmaps for distributions, gauges for status
- **Color Theory:** Red for critical, amber for warning, green for healthy states
- **Reference Lines:** SLO targets, capacity thresholds, and historical baselines
- **Time Ranges:** Default to meaningful windows (4h for incidents, 7d for trends)
#### Panel Design
- **Metric Queries:** Efficient Prometheus/InfluxDB queries with proper aggregation
- **Alerting Integration:** Visual alert state indicators on relevant panels
- **Interactive Elements:** Template variables, drill-down links, and annotation overlays
- **Performance:** Sub-second render times through query optimization
### Alert Design and Optimization
#### Alert Classification
- **Severity Levels:**
- **Critical:** Service down, SLO burn rate high
- **Warning:** Approaching thresholds, non-user-facing issues
- **Info:** Deployment notifications, capacity planning alerts
- **Actionability:** Every alert must have a clear response action
- **Alert Routing:** Escalation policies based on severity and team ownership
#### Alert Fatigue Prevention
- **Signal vs Noise:** High precision (few false positives) over high recall
- **Hysteresis:** Different thresholds for firing and resolving alerts
- **Suppression:** Dependent alert suppression during known outages
- **Grouping:** Related alerts grouped into single notifications
#### Alert Rule Design
- **Threshold Selection:** Statistical methods for threshold determination
- **Window Functions:** Appropriate averaging windows and percentile calculations
- **Alert Lifecycle:** Clear firing conditions and automatic resolution criteria
- **Testing:** Alert rule validation against historical data
### Runbook Generation and Incident Response
#### Runbook Structure
- **Alert Context:** What the alert means and why it fired
- **Impact Assessment:** User-facing vs internal impact evaluation
- **Investigation Steps:** Ordered troubleshooting procedures with time estimates
- **Resolution Actions:** Common fixes and escalation procedures
- **Post-Incident:** Follow-up tasks and prevention measures
#### Incident Detection Patterns
- **Anomaly Detection:** Statistical methods for detecting unusual patterns
- **Composite Alerts:** Multi-signal alerts for complex failure modes
- **Predictive Alerts:** Capacity and trend-based forward-looking alerts
- **Canary Monitoring:** Early detection through progressive deployment monitoring
### Golden Signals Framework
#### Latency Monitoring
- **Request Latency:** P50, P95, P99 response time tracking
- **Queue Latency:** Time spent waiting in processing queues
- **Network Latency:** Inter-service communication delays
- **Database Latency:** Query execution and connection pool metrics
#### Traffic Monitoring
- **Request Rate:** Requests per second with burst detection
- **Bandwidth Usage:** Network throughput and capacity utilization
- **User Sessions:** Active user tracking and session duration
- **Feature Usage:** API endpoint and feature adoption metrics
#### Error Monitoring
- **Error Rate:** 4xx and 5xx HTTP response code tracking
- **Error Budget:** SLO-based error rate targets and consumption
- **Error Distribution:** Error type classification and trending
- **Silent Failures:** Detection of processing failures without HTTP errors
#### Saturation Monitoring
- **Resource Utilization:** CPU, memory, disk, and network usage
- **Queue Depth:** Processing queue length and wait times
- **Connection Pools:** Database and service connection saturation
- **Rate Limiting:** API throttling and quota exhaustion tracking
### Distributed Tracing Strategies
#### Trace Architecture
- **Sampling Strategy:** Head-based, tail-based, and adaptive sampling
- **Trace Propagation:** Context propagation across service boundaries
- **Span Correlation:** Parent-child relationship modeling
- **Trace Storage:** Retention policies and storage optimization
#### Service Instrumentation
- **Auto-Instrumentation:** Framework-based automatic trace generation
- **Manual Instrumentation:** Custom span creation for business logic
- **Baggage Handling:** Cross-cutting concern propagation
- **Performance Impact:** Instrumentation overhead measurement and optimization
### Log Aggregation Patterns
#### Collection Architecture
- **Agent Deployment:** Log shipping agent strategies (push vs pull)
- **Log Routing:** Topic-based routing and filtering
- **Parsing Strategies:** Structured vs unstructured log handling
- **Schema Evolution:** Log format versioning and migration
#### Storage and Indexing
- **Index Design:** Optimized field indexing for common query patterns
- **Retention Policies:** Time and volume-based log retention
- **Compression:** Log data compression and archival strategies
- **Search Performance:** Query optimization and result caching
### Cost Optimization for Observability
#### Data Management
- **Metric Retention:** Tiered retention based on metric importance
- **Log Sampling:** Intelligent sampling to reduce ingestion costs
- **Trace Sampling:** Cost-effective trace collection strategies
- **Data Archival:** Cold storage for historical observability data
#### Resource Optimization
- **Query Efficiency:** Optimized metric and log queries
- **Storage Costs:** Appropriate storage tiers for different data types
- **Ingestion Rate Limiting:** Controlled data ingestion to manage costs
- **Cardinality Management:** High-cardinality metric detection and mitigation
## Scripts Overview
This skill includes three powerful Python scripts for comprehensive observability design:
### 1. SLO Designer (`slo_designer.py`)
Generates complete SLI/SLO frameworks based on service characteristics:
- **Input:** Service description JSON (type, criticality, dependencies)
- **Output:** SLI definitions, SLO targets, error budgets, burn rate alerts, SLA recommendations
- **Features:** Multi-window burn rate calculations, error budget policies, alert rule generation
### 2. Alert Optimizer (`alert_optimizer.py`)
Analyzes and optimizes existing alert configurations:
- **Input:** Alert configuration JSON with rules, thresholds, and routing
- **Output:** Optimization report and improved alert configuration
- **Features:** Noise detection, coverage gaps, duplicate identification, threshold optimization
### 3. Dashboard Generator (`dashboard_generator.py`)
Creates comprehensive dashboard specifications:
- **Input:** Service/system description JSON
- **Output:** Grafana-compatible dashboard JSON and documentation
- **Features:** Golden signals coverage, RED/USE methods, drill-down paths, role-based views
## Integration Patterns
### Monitoring Stack Integration
- **Prometheus:** Metric collection and alerting rule generation
- **Grafana:** Dashboard creation and visualization configuration
- **Elasticsearch/Kibana:** Log analysis and dashboard integration
- **Jaeger/Zipkin:** Distributed tracing configuration and analysis
### CI/CD Integration
- **Pipeline Monitoring:** Build, test, and deployment observability
- **Deployment Correlation:** Release impact tracking and rollback triggers
- **Feature Flag Monitoring:** A/B test and feature rollout observability
- **Performance Regression:** Automated performance monitoring in pipelines
### Incident Management Integration
- **PagerDuty/VictorOps:** Alert routing and escalation policies
- **Slack/Teams:** Notification and collaboration integration
- **JIRA/ServiceNow:** Incident tracking and resolution workflows
- **Post-Mortem:** Automated incident analysis and improvement tracking
## Advanced Patterns
### Multi-Cloud Observability
- **Cross-Cloud Metrics:** Unified metrics across AWS, GCP, Azure
- **Network Observability:** Inter-cloud connectivity monitoring
- **Cost Attribution:** Cloud resource cost tracking and optimization
- **Compliance Monitoring:** Security and compliance posture tracking
### Microservices Observability
- **Service Mesh Integration:** Istio/Linkerd observability configuration
- **API Gateway Monitoring:** Request routing and rate limiting observability
- **Container Orchestration:** Kubernetes cluster and workload monitoring
- **Service Discovery:** Dynamic service monitoring and health checks
### Machine Learning Observability
- **Model Performance:** Accuracy, drift, and bias monitoring
- **Feature Store Monitoring:** Feature quality and freshness tracking
- **Pipeline Observability:** ML pipeline execution and performance monitoring
- **A/B Test Analysis:** Statistical significance and business impact measurement
## Best Practices
### Organizational Alignment
- **SLO Setting:** Collaborative target setting between product and engineering
- **Alert Ownership:** Clear escalation paths and team responsibilities
- **Dashboard Governance:** Centralized dashboard management and standards
- **Training Programs:** Team education on observability tools and practices
### Technical Excellence
- **Infrastructure as Code:** Observability configuration version control
- **Testing Strategy:** Alert rule testing and dashboard validation
- **Performance Monitoring:** Observability system performance tracking
- **Security Considerations:** Access control and data privacy in observability
### Continuous Improvement
- **Metrics Review:** Regular SLI/SLO effectiveness assessment
- **Alert Tuning:** Ongoing alert threshold and routing optimization
- **Dashboard Evolution:** User feedback-driven dashboard improvements
- **Tool Evaluation:** Regular assessment of observability tool effectiveness
## Success Metrics
### Operational Metrics
- **Mean Time to Detection (MTTD):** How quickly issues are identified
- **Mean Time to Resolution (MTTR):** Time from detection to resolution
- **Alert Precision:** Percentage of actionable alerts
- **SLO Achievement:** Percentage of SLO targets met consistently
### Business Metrics
- **System Reliability:** Overall uptime and user experience quality
- **Engineering Velocity:** Development team productivity and deployment frequency
- **Cost Efficiency:** Observability cost as percentage of infrastructure spend
- **Customer Satisfaction:** User-reported reliability and performance satisfaction
This comprehensive observability design skill enables organizations to build robust, scalable monitoring and alerting systems that provide actionable insights while maintaining cost efficiency and operational excellence.
FILE:assets/sample_alerts.json
{
"alerts": [
{
"alert": "HighLatency",
"expr": "histogram_quantile(0.95, rate(http_request_duration_seconds_bucket{service=\"payment-service\"}[5m])) > 0.5",
"for": "5m",
"labels": {
"severity": "warning",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "High request latency detected",
"description": "95th percentile latency is {{ $value }}s for payment-service",
"runbook_url": "https://runbooks.company.com/high-latency"
},
"historical_data": {
"fires_per_day": 2.5,
"false_positive_rate": 0.15,
"average_duration_minutes": 12
}
},
{
"alert": "ServiceDown",
"expr": "up{service=\"payment-service\"} == 0",
"labels": {
"severity": "critical",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "Payment service is down",
"description": "Payment service has been down for more than 1 minute",
"runbook_url": "https://runbooks.company.com/service-down"
},
"historical_data": {
"fires_per_day": 0.1,
"false_positive_rate": 0.05,
"average_duration_minutes": 3
}
},
{
"alert": "HighErrorRate",
"expr": "sum(rate(http_requests_total{service=\"payment-service\",code=~\"5..\"}[5m])) / sum(rate(http_requests_total{service=\"payment-service\"}[5m])) > 0.01",
"for": "2m",
"labels": {
"severity": "warning",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "High error rate detected",
"description": "Error rate is {{ $value | humanizePercentage }} for payment-service",
"runbook_url": "https://runbooks.company.com/high-error-rate"
},
"historical_data": {
"fires_per_day": 1.8,
"false_positive_rate": 0.25,
"average_duration_minutes": 8
}
},
{
"alert": "HighCPUUsage",
"expr": "rate(process_cpu_seconds_total{service=\"payment-service\"}[5m]) * 100 > 80",
"labels": {
"severity": "warning",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "High CPU usage",
"description": "CPU usage is {{ $value }}% for payment-service"
},
"historical_data": {
"fires_per_day": 15.2,
"false_positive_rate": 0.8,
"average_duration_minutes": 45
}
},
{
"alert": "HighMemoryUsage",
"expr": "process_resident_memory_bytes{service=\"payment-service\"} / process_virtual_memory_max_bytes{service=\"payment-service\"} * 100 > 85",
"labels": {
"severity": "info",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "High memory usage",
"description": "Memory usage is {{ $value }}% for payment-service"
},
"historical_data": {
"fires_per_day": 8.5,
"false_positive_rate": 0.6,
"average_duration_minutes": 30
}
},
{
"alert": "DatabaseConnectionPoolExhaustion",
"expr": "db_connections_active{service=\"payment-service\"} / db_connections_max{service=\"payment-service\"} > 0.9",
"for": "1m",
"labels": {
"severity": "critical",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "Database connection pool near exhaustion",
"description": "Connection pool utilization is {{ $value | humanizePercentage }}",
"runbook_url": "https://runbooks.company.com/db-connections"
},
"historical_data": {
"fires_per_day": 0.3,
"false_positive_rate": 0.1,
"average_duration_minutes": 5
}
},
{
"alert": "LowTraffic",
"expr": "sum(rate(http_requests_total{service=\"payment-service\"}[5m])) < 10",
"for": "10m",
"labels": {
"severity": "warning",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "Unusually low traffic",
"description": "Request rate is {{ $value }} RPS, which is unusually low"
},
"historical_data": {
"fires_per_day": 12.0,
"false_positive_rate": 0.9,
"average_duration_minutes": 120
}
},
{
"alert": "HighLatencyDuplicate",
"expr": "histogram_quantile(0.95, rate(http_request_duration_seconds_bucket{service=\"payment-service\"}[5m])) > 0.5",
"for": "5m",
"labels": {
"severity": "warning",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "High request latency detected (duplicate)",
"description": "95th percentile latency is {{ $value }}s for payment-service"
},
"historical_data": {
"fires_per_day": 2.5,
"false_positive_rate": 0.15,
"average_duration_minutes": 12
}
},
{
"alert": "VeryLowErrorRate",
"expr": "sum(rate(http_requests_total{service=\"payment-service\",code=~\"5..\"}[5m])) / sum(rate(http_requests_total{service=\"payment-service\"}[5m])) > 0.001",
"labels": {
"severity": "info",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "Error rate above 0.1%",
"description": "Error rate is {{ $value | humanizePercentage }}"
},
"historical_data": {
"fires_per_day": 25.0,
"false_positive_rate": 0.95,
"average_duration_minutes": 5
}
},
{
"alert": "DiskUsageHigh",
"expr": "disk_usage_percent{service=\"payment-service\"} > 85",
"labels": {
"severity": "warning",
"service": "payment-service",
"team": "payments"
},
"annotations": {
"summary": "Disk usage high",
"description": "Disk usage is {{ $value }}%"
},
"historical_data": {
"fires_per_day": 3.2,
"false_positive_rate": 0.4,
"average_duration_minutes": 240
}
}
],
"services": [
{
"name": "payment-service",
"type": "api",
"criticality": "critical",
"team": "payments"
},
{
"name": "user-service",
"type": "api",
"criticality": "high",
"team": "identity"
},
{
"name": "notification-service",
"type": "api",
"criticality": "medium",
"team": "communications"
}
],
"alert_routing": {
"routes": [
{
"match": {
"severity": "critical"
},
"receiver": "pager-critical",
"group_wait": "10s",
"group_interval": "1m",
"repeat_interval": "5m"
},
{
"match": {
"severity": "warning"
},
"receiver": "slack-warnings",
"group_wait": "30s",
"group_interval": "5m",
"repeat_interval": "1h"
},
{
"match": {
"severity": "info"
},
"receiver": "email-info",
"group_wait": "2m",
"group_interval": "10m",
"repeat_interval": "24h"
}
]
},
"receivers": [
{
"name": "pager-critical",
"pagerduty_configs": [
{
"routing_key": "pager-key-critical",
"description": "Critical alert: {{ range .Alerts }}{{ .Annotations.summary }}{{ end }}"
}
]
},
{
"name": "slack-warnings",
"slack_configs": [
{
"api_url": "https://hooks.slack.com/services/warnings",
"channel": "#alerts-warnings",
"title": "Warning Alert",
"text": "{{ range .Alerts }}{{ .Annotations.description }}{{ end }}"
}
]
},
{
"name": "email-info",
"email_configs": [
{
"to": "team-notifications@company.com",
"subject": "Info Alert: {{ .GroupLabels.alertname }}",
"body": "{{ range .Alerts }}{{ .Annotations.description }}{{ end }}"
}
]
}
]
}
FILE:assets/sample_service_api.json
{
"name": "payment-service",
"type": "api",
"criticality": "critical",
"user_facing": true,
"description": "Handles payment processing and transaction management",
"team": "payments",
"environment": "production",
"dependencies": [
{
"name": "user-service",
"type": "api",
"criticality": "high"
},
{
"name": "payment-gateway",
"type": "external",
"criticality": "critical"
},
{
"name": "fraud-detection",
"type": "ml",
"criticality": "high"
}
],
"endpoints": [
{
"path": "/api/v1/payments",
"method": "POST",
"sla_latency_ms": 500,
"expected_tps": 100
},
{
"path": "/api/v1/payments/{id}",
"method": "GET",
"sla_latency_ms": 200,
"expected_tps": 500
},
{
"path": "/api/v1/payments/{id}/refund",
"method": "POST",
"sla_latency_ms": 1000,
"expected_tps": 10
}
],
"business_metrics": {
"revenue_per_hour": {
"metric": "sum(payment_amount * rate(payments_successful_total[1h]))",
"target": 50000,
"unit": "USD"
},
"conversion_rate": {
"metric": "sum(rate(payments_successful_total[5m])) / sum(rate(payment_attempts_total[5m]))",
"target": 0.95,
"unit": "percentage"
}
},
"infrastructure": {
"container_orchestrator": "kubernetes",
"replicas": 6,
"cpu_limit": "2000m",
"memory_limit": "4Gi",
"database": {
"type": "postgresql",
"connection_pool_size": 20
},
"cache": {
"type": "redis",
"cluster_size": 3
}
},
"compliance_requirements": [
"PCI-DSS",
"SOX",
"GDPR"
],
"tags": [
"payment",
"transaction",
"critical-path",
"revenue-generating"
]
}
FILE:assets/sample_service_web.json
{
"name": "customer-portal",
"type": "web",
"criticality": "high",
"user_facing": true,
"description": "Customer-facing web application for account management and billing",
"team": "frontend",
"environment": "production",
"dependencies": [
{
"name": "user-service",
"type": "api",
"criticality": "high"
},
{
"name": "billing-service",
"type": "api",
"criticality": "high"
},
{
"name": "notification-service",
"type": "api",
"criticality": "medium"
},
{
"name": "cdn",
"type": "external",
"criticality": "medium"
}
],
"pages": [
{
"path": "/dashboard",
"sla_load_time_ms": 2000,
"expected_concurrent_users": 1000
},
{
"path": "/billing",
"sla_load_time_ms": 3000,
"expected_concurrent_users": 200
},
{
"path": "/settings",
"sla_load_time_ms": 1500,
"expected_concurrent_users": 100
}
],
"business_metrics": {
"daily_active_users": {
"metric": "count(user_sessions_started_total[1d])",
"target": 10000,
"unit": "users"
},
"session_duration": {
"metric": "avg(user_session_duration_seconds)",
"target": 300,
"unit": "seconds"
},
"bounce_rate": {
"metric": "sum(rate(page_views_bounced_total[1h])) / sum(rate(page_views_total[1h]))",
"target": 0.3,
"unit": "percentage"
}
},
"infrastructure": {
"container_orchestrator": "kubernetes",
"replicas": 4,
"cpu_limit": "1000m",
"memory_limit": "2Gi",
"storage": {
"type": "nfs",
"size": "50Gi"
},
"ingress": {
"type": "nginx",
"ssl_termination": true,
"rate_limiting": {
"requests_per_second": 100,
"burst": 200
}
}
},
"monitoring": {
"synthetic_checks": [
{
"name": "login_flow",
"url": "/auth/login",
"frequency": "1m",
"locations": ["us-east", "eu-west", "ap-south"]
},
{
"name": "checkout_flow",
"url": "/billing/checkout",
"frequency": "5m",
"locations": ["us-east", "eu-west"]
}
],
"rum": {
"enabled": true,
"sampling_rate": 0.1
}
},
"compliance_requirements": [
"GDPR",
"CCPA"
],
"tags": [
"frontend",
"customer-facing",
"billing",
"high-traffic"
]
}
FILE:expected_outputs/sample_dashboard.json
{
"metadata": {
"title": "customer-portal - SRE Dashboard",
"service": {
"name": "customer-portal",
"type": "web",
"criticality": "high",
"user_facing": true,
"description": "Customer-facing web application for account management and billing",
"team": "frontend",
"environment": "production",
"dependencies": [
{
"name": "user-service",
"type": "api",
"criticality": "high"
},
{
"name": "billing-service",
"type": "api",
"criticality": "high"
},
{
"name": "notification-service",
"type": "api",
"criticality": "medium"
},
{
"name": "cdn",
"type": "external",
"criticality": "medium"
}
],
"pages": [
{
"path": "/dashboard",
"sla_load_time_ms": 2000,
"expected_concurrent_users": 1000
},
{
"path": "/billing",
"sla_load_time_ms": 3000,
"expected_concurrent_users": 200
},
{
"path": "/settings",
"sla_load_time_ms": 1500,
"expected_concurrent_users": 100
}
],
"business_metrics": {
"daily_active_users": {
"metric": "count(user_sessions_started_total[1d])",
"target": 10000,
"unit": "users"
},
"session_duration": {
"metric": "avg(user_session_duration_seconds)",
"target": 300,
"unit": "seconds"
},
"bounce_rate": {
"metric": "sum(rate(page_views_bounced_total[1h])) / sum(rate(page_views_total[1h]))",
"target": 0.3,
"unit": "percentage"
}
},
"infrastructure": {
"container_orchestrator": "kubernetes",
"replicas": 4,
"cpu_limit": "1000m",
"memory_limit": "2Gi",
"storage": {
"type": "nfs",
"size": "50Gi"
},
"ingress": {
"type": "nginx",
"ssl_termination": true,
"rate_limiting": {
"requests_per_second": 100,
"burst": 200
}
}
},
"monitoring": {
"synthetic_checks": [
{
"name": "login_flow",
"url": "/auth/login",
"frequency": "1m",
"locations": [
"us-east",
"eu-west",
"ap-south"
]
},
{
"name": "checkout_flow",
"url": "/billing/checkout",
"frequency": "5m",
"locations": [
"us-east",
"eu-west"
]
}
],
"rum": {
"enabled": true,
"sampling_rate": 0.1
}
},
"compliance_requirements": [
"GDPR",
"CCPA"
],
"tags": [
"frontend",
"customer-facing",
"billing",
"high-traffic"
]
},
"target_role": "sre",
"generated_at": "2026-02-16T14:02:03.421248Z",
"version": "1.0"
},
"configuration": {
"time_ranges": [
"1h",
"6h",
"1d",
"7d"
],
"default_time_range": "6h",
"refresh_interval": "30s",
"timezone": "UTC",
"theme": "dark"
},
"layout": {
"grid_settings": {
"width": 24,
"height_unit": "px",
"cell_height": 30
},
"sections": [
{
"title": "Service Overview",
"collapsed": false,
"y_position": 0,
"panels": [
"service_status",
"slo_summary",
"error_budget"
]
},
{
"title": "Golden Signals",
"collapsed": false,
"y_position": 8,
"panels": [
"latency",
"traffic",
"errors",
"saturation"
]
},
{
"title": "Resource Utilization",
"collapsed": false,
"y_position": 16,
"panels": [
"cpu_usage",
"memory_usage",
"network_io",
"disk_io"
]
},
{
"title": "Dependencies & Downstream",
"collapsed": true,
"y_position": 24,
"panels": [
"dependency_status",
"downstream_latency",
"circuit_breakers"
]
}
]
},
"panels": [
{
"id": "service_status",
"title": "Service Status",
"type": "stat",
"grid_pos": {
"x": 0,
"y": 0,
"w": 6,
"h": 4
},
"targets": [
{
"expr": "up{service=\"customer-portal\"}",
"legendFormat": "Status"
}
],
"field_config": {
"overrides": [
{
"matcher": {
"id": "byName",
"options": "Status"
},
"properties": [
{
"id": "color",
"value": {
"mode": "thresholds"
}
},
{
"id": "thresholds",
"value": {
"steps": [
{
"color": "red",
"value": 0
},
{
"color": "green",
"value": 1
}
]
}
},
{
"id": "mappings",
"value": [
{
"options": {
"0": {
"text": "DOWN"
}
},
"type": "value"
},
{
"options": {
"1": {
"text": "UP"
}
},
"type": "value"
}
]
}
]
}
]
},
"options": {
"orientation": "horizontal",
"textMode": "value_and_name"
}
},
{
"id": "slo_summary",
"title": "SLO Achievement (30d)",
"type": "stat",
"grid_pos": {
"x": 6,
"y": 0,
"w": 9,
"h": 4
},
"targets": [
{
"expr": "(1 - (increase(http_requests_total{service=\"customer-portal\",code=~\"5..\"}[30d]) / increase(http_requests_total{service=\"customer-portal\"}[30d]))) * 100",
"legendFormat": "Availability"
},
{
"expr": "histogram_quantile(0.95, increase(http_request_duration_seconds_bucket{service=\"customer-portal\"}[30d])) * 1000",
"legendFormat": "P95 Latency (ms)"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "thresholds"
},
"thresholds": {
"steps": [
{
"color": "red",
"value": 0
},
{
"color": "yellow",
"value": 99.0
},
{
"color": "green",
"value": 99.9
}
]
}
}
},
"options": {
"orientation": "horizontal",
"textMode": "value_and_name"
}
},
{
"id": "error_budget",
"title": "Error Budget Remaining",
"type": "gauge",
"grid_pos": {
"x": 15,
"y": 0,
"w": 9,
"h": 4
},
"targets": [
{
"expr": "(1 - (increase(http_requests_total{service=\"customer-portal\",code=~\"5..\"}[30d]) / increase(http_requests_total{service=\"customer-portal\"}[30d])) - 0.999) / 0.001 * 100",
"legendFormat": "Error Budget %"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "thresholds"
},
"min": 0,
"max": 100,
"thresholds": {
"steps": [
{
"color": "red",
"value": 0
},
{
"color": "yellow",
"value": 25
},
{
"color": "green",
"value": 50
}
]
},
"unit": "percent"
}
},
"options": {
"showThresholdLabels": true,
"showThresholdMarkers": true
}
},
{
"id": "latency",
"title": "Request Latency",
"type": "timeseries",
"grid_pos": {
"x": 0,
"y": 8,
"w": 12,
"h": 6
},
"targets": [
{
"expr": "histogram_quantile(0.50, rate(http_request_duration_seconds_bucket{service=\"customer-portal\"}[5m])) * 1000",
"legendFormat": "P50 Latency"
},
{
"expr": "histogram_quantile(0.95, rate(http_request_duration_seconds_bucket{service=\"customer-portal\"}[5m])) * 1000",
"legendFormat": "P95 Latency"
},
{
"expr": "histogram_quantile(0.99, rate(http_request_duration_seconds_bucket{service=\"customer-portal\"}[5m])) * 1000",
"legendFormat": "P99 Latency"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"unit": "ms",
"custom": {
"drawStyle": "line",
"lineInterpolation": "linear",
"lineWidth": 1,
"fillOpacity": 10
}
}
},
"options": {
"tooltip": {
"mode": "multi",
"sort": "desc"
},
"legend": {
"displayMode": "table",
"placement": "bottom"
}
}
},
{
"id": "traffic",
"title": "Request Rate",
"type": "timeseries",
"grid_pos": {
"x": 12,
"y": 8,
"w": 12,
"h": 6
},
"targets": [
{
"expr": "sum(rate(http_requests_total{service=\"customer-portal\"}[5m]))",
"legendFormat": "Total RPS"
},
{
"expr": "sum(rate(http_requests_total{service=\"customer-portal\",code=~\"2..\"}[5m]))",
"legendFormat": "2xx RPS"
},
{
"expr": "sum(rate(http_requests_total{service=\"customer-portal\",code=~\"4..\"}[5m]))",
"legendFormat": "4xx RPS"
},
{
"expr": "sum(rate(http_requests_total{service=\"customer-portal\",code=~\"5..\"}[5m]))",
"legendFormat": "5xx RPS"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"unit": "reqps",
"custom": {
"drawStyle": "line",
"lineInterpolation": "linear",
"lineWidth": 1,
"fillOpacity": 0
}
}
},
"options": {
"tooltip": {
"mode": "multi",
"sort": "desc"
},
"legend": {
"displayMode": "table",
"placement": "bottom"
}
}
},
{
"id": "errors",
"title": "Error Rate",
"type": "timeseries",
"grid_pos": {
"x": 0,
"y": 14,
"w": 12,
"h": 6
},
"targets": [
{
"expr": "sum(rate(http_requests_total{service=\"customer-portal\",code=~\"5..\"}[5m])) / sum(rate(http_requests_total{service=\"customer-portal\"}[5m])) * 100",
"legendFormat": "5xx Error Rate"
},
{
"expr": "sum(rate(http_requests_total{service=\"customer-portal\",code=~\"4..\"}[5m])) / sum(rate(http_requests_total{service=\"customer-portal\"}[5m])) * 100",
"legendFormat": "4xx Error Rate"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"unit": "percent",
"custom": {
"drawStyle": "line",
"lineInterpolation": "linear",
"lineWidth": 2,
"fillOpacity": 20
}
},
"overrides": [
{
"matcher": {
"id": "byName",
"options": "5xx Error Rate"
},
"properties": [
{
"id": "color",
"value": {
"fixedColor": "red"
}
}
]
}
]
},
"options": {
"tooltip": {
"mode": "multi",
"sort": "desc"
},
"legend": {
"displayMode": "table",
"placement": "bottom"
}
}
},
{
"id": "saturation",
"title": "Saturation Metrics",
"type": "timeseries",
"grid_pos": {
"x": 12,
"y": 14,
"w": 12,
"h": 6
},
"targets": [
{
"expr": "rate(process_cpu_seconds_total{service=\"customer-portal\"}[5m]) * 100",
"legendFormat": "CPU Usage %"
},
{
"expr": "process_resident_memory_bytes{service=\"customer-portal\"} / process_virtual_memory_max_bytes{service=\"customer-portal\"} * 100",
"legendFormat": "Memory Usage %"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"unit": "percent",
"max": 100,
"custom": {
"drawStyle": "line",
"lineInterpolation": "linear",
"lineWidth": 1,
"fillOpacity": 10
}
}
},
"options": {
"tooltip": {
"mode": "multi",
"sort": "desc"
},
"legend": {
"displayMode": "table",
"placement": "bottom"
}
}
},
{
"id": "cpu_usage",
"title": "CPU Usage",
"type": "gauge",
"grid_pos": {
"x": 0,
"y": 20,
"w": 6,
"h": 4
},
"targets": [
{
"expr": "rate(process_cpu_seconds_total{service=\"customer-portal\"}[5m]) * 100",
"legendFormat": "CPU %"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "thresholds"
},
"unit": "percent",
"min": 0,
"max": 100,
"thresholds": {
"steps": [
{
"color": "green",
"value": 0
},
{
"color": "yellow",
"value": 70
},
{
"color": "red",
"value": 90
}
]
}
}
},
"options": {
"showThresholdLabels": true,
"showThresholdMarkers": true
}
},
{
"id": "memory_usage",
"title": "Memory Usage",
"type": "gauge",
"grid_pos": {
"x": 6,
"y": 20,
"w": 6,
"h": 4
},
"targets": [
{
"expr": "process_resident_memory_bytes{service=\"customer-portal\"} / 1024 / 1024",
"legendFormat": "Memory MB"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "thresholds"
},
"unit": "decbytes",
"thresholds": {
"steps": [
{
"color": "green",
"value": 0
},
{
"color": "yellow",
"value": 512000000
},
{
"color": "red",
"value": 1024000000
}
]
}
}
}
},
{
"id": "network_io",
"title": "Network I/O",
"type": "timeseries",
"grid_pos": {
"x": 12,
"y": 20,
"w": 6,
"h": 4
},
"targets": [
{
"expr": "rate(process_network_receive_bytes_total{service=\"customer-portal\"}[5m])",
"legendFormat": "RX Bytes/s"
},
{
"expr": "rate(process_network_transmit_bytes_total{service=\"customer-portal\"}[5m])",
"legendFormat": "TX Bytes/s"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"unit": "binBps"
}
}
},
{
"id": "disk_io",
"title": "Disk I/O",
"type": "timeseries",
"grid_pos": {
"x": 18,
"y": 20,
"w": 6,
"h": 4
},
"targets": [
{
"expr": "rate(process_disk_read_bytes_total{service=\"customer-portal\"}[5m])",
"legendFormat": "Read Bytes/s"
},
{
"expr": "rate(process_disk_write_bytes_total{service=\"customer-portal\"}[5m])",
"legendFormat": "Write Bytes/s"
}
],
"field_config": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"unit": "binBps"
}
}
}
],
"variables": [
{
"name": "environment",
"type": "query",
"query": "label_values(environment)",
"current": {
"text": "production",
"value": "production"
},
"includeAll": false,
"multi": false,
"refresh": "on_dashboard_load"
},
{
"name": "instance",
"type": "query",
"query": "label_values(up{service=\"customer-portal\"}, instance)",
"current": {
"text": "All",
"value": "$__all"
},
"includeAll": true,
"multi": true,
"refresh": "on_time_range_change"
},
{
"name": "handler",
"type": "query",
"query": "label_values(http_requests_total{service=\"customer-portal\"}, handler)",
"current": {
"text": "All",
"value": "$__all"
},
"includeAll": true,
"multi": true,
"refresh": "on_time_range_change"
}
],
"alerts_integration": {
"alert_annotations": true,
"alert_rules_query": "ALERTS{service=\"customer-portal\"}",
"alert_panels": [
{
"title": "Active Alerts",
"type": "table",
"query": "ALERTS{service=\"customer-portal\",alertstate=\"firing\"}",
"columns": [
"alertname",
"severity",
"instance",
"description"
]
}
]
},
"drill_down_paths": {
"service_overview": {
"from": "service_status",
"to": "detailed_health_dashboard",
"url": "/d/service-health/customer-portal-health",
"params": [
"var-service",
"var-environment"
]
},
"error_investigation": {
"from": "errors",
"to": "error_details_dashboard",
"url": "/d/errors/customer-portal-errors",
"params": [
"var-service",
"var-time_range"
]
},
"latency_analysis": {
"from": "latency",
"to": "trace_analysis_dashboard",
"url": "/d/traces/customer-portal-traces",
"params": [
"var-service",
"var-handler"
]
},
"capacity_planning": {
"from": "saturation",
"to": "capacity_dashboard",
"url": "/d/capacity/customer-portal-capacity",
"params": [
"var-service",
"var-time_range"
]
}
}
}
FILE:expected_outputs/sample_slo_framework.json
{
"metadata": {
"service": {
"name": "payment-service",
"type": "api",
"criticality": "critical",
"user_facing": true,
"description": "Handles payment processing and transaction management",
"team": "payments",
"environment": "production",
"dependencies": [
{
"name": "user-service",
"type": "api",
"criticality": "high"
},
{
"name": "payment-gateway",
"type": "external",
"criticality": "critical"
},
{
"name": "fraud-detection",
"type": "ml",
"criticality": "high"
}
],
"endpoints": [
{
"path": "/api/v1/payments",
"method": "POST",
"sla_latency_ms": 500,
"expected_tps": 100
},
{
"path": "/api/v1/payments/{id}",
"method": "GET",
"sla_latency_ms": 200,
"expected_tps": 500
},
{
"path": "/api/v1/payments/{id}/refund",
"method": "POST",
"sla_latency_ms": 1000,
"expected_tps": 10
}
],
"business_metrics": {
"revenue_per_hour": {
"metric": "sum(payment_amount * rate(payments_successful_total[1h]))",
"target": 50000,
"unit": "USD"
},
"conversion_rate": {
"metric": "sum(rate(payments_successful_total[5m])) / sum(rate(payment_attempts_total[5m]))",
"target": 0.95,
"unit": "percentage"
}
},
"infrastructure": {
"container_orchestrator": "kubernetes",
"replicas": 6,
"cpu_limit": "2000m",
"memory_limit": "4Gi",
"database": {
"type": "postgresql",
"connection_pool_size": 20
},
"cache": {
"type": "redis",
"cluster_size": 3
}
},
"compliance_requirements": [
"PCI-DSS",
"SOX",
"GDPR"
],
"tags": [
"payment",
"transaction",
"critical-path",
"revenue-generating"
]
},
"generated_at": "2026-02-16T14:01:57.572080Z",
"framework_version": "1.0"
},
"slis": [
{
"name": "Availability",
"description": "Percentage of successful requests",
"type": "ratio",
"good_events": "sum(rate(http_requests_total{service=\"payment-service\",code!~\"5..\"}))",
"total_events": "sum(rate(http_requests_total{service=\"payment-service\"}))",
"unit": "percentage"
},
{
"name": "Request Latency P95",
"description": "95th percentile of request latency",
"type": "threshold",
"query": "histogram_quantile(0.95, rate(http_request_duration_seconds_bucket{service=\"payment-service\"}[5m]))",
"unit": "seconds"
},
{
"name": "Error Rate",
"description": "Rate of 5xx errors",
"type": "ratio",
"good_events": "sum(rate(http_requests_total{service=\"payment-service\",code!~\"5..\"}))",
"total_events": "sum(rate(http_requests_total{service=\"payment-service\"}))",
"unit": "percentage"
},
{
"name": "Request Throughput",
"description": "Requests per second",
"type": "gauge",
"query": "sum(rate(http_requests_total{service=\"payment-service\"}[5m]))",
"unit": "requests/sec"
},
{
"name": "User Journey Success Rate",
"description": "Percentage of successful complete user journeys",
"type": "ratio",
"good_events": "sum(rate(user_journey_total{service=\"payment-service\",status=\"success\"}[5m]))",
"total_events": "sum(rate(user_journey_total{service=\"payment-service\"}[5m]))",
"unit": "percentage"
},
{
"name": "Feature Availability",
"description": "Percentage of time key features are available",
"type": "ratio",
"good_events": "sum(rate(feature_checks_total{service=\"payment-service\",status=\"available\"}[5m]))",
"total_events": "sum(rate(feature_checks_total{service=\"payment-service\"}[5m]))",
"unit": "percentage"
}
],
"slos": [
{
"name": "Availability SLO",
"description": "Service level objective for percentage of successful requests",
"sli_name": "Availability",
"target_value": 0.9999,
"target_display": "99.99%",
"operator": ">=",
"time_windows": [
"1h",
"1d",
"7d",
"30d"
],
"measurement_window": "30d",
"service": "payment-service",
"criticality": "critical"
},
{
"name": "Request Latency P95 SLO",
"description": "Service level objective for 95th percentile of request latency",
"sli_name": "Request Latency P95",
"target_value": 100,
"target_display": "0.1s",
"operator": "<=",
"time_windows": [
"1h",
"1d",
"7d",
"30d"
],
"measurement_window": "30d",
"service": "payment-service",
"criticality": "critical"
},
{
"name": "Error Rate SLO",
"description": "Service level objective for rate of 5xx errors",
"sli_name": "Error Rate",
"target_value": 0.001,
"target_display": "0.1%",
"operator": "<=",
"time_windows": [
"1h",
"1d",
"7d",
"30d"
],
"measurement_window": "30d",
"service": "payment-service",
"criticality": "critical"
},
{
"name": "User Journey Success Rate SLO",
"description": "Service level objective for percentage of successful complete user journeys",
"sli_name": "User Journey Success Rate",
"target_value": 0.9999,
"target_display": "99.99%",
"operator": ">=",
"time_windows": [
"1h",
"1d",
"7d",
"30d"
],
"measurement_window": "30d",
"service": "payment-service",
"criticality": "critical"
},
{
"name": "Feature Availability SLO",
"description": "Service level objective for percentage of time key features are available",
"sli_name": "Feature Availability",
"target_value": 0.9999,
"target_display": "99.99%",
"operator": ">=",
"time_windows": [
"1h",
"1d",
"7d",
"30d"
],
"measurement_window": "30d",
"service": "payment-service",
"criticality": "critical"
}
],
"error_budgets": [
{
"slo_name": "Availability SLO",
"error_budget_rate": 9.999999999998899e-05,
"error_budget_percentage": "0.010%",
"budgets_by_window": {
"1h": "0.4 seconds",
"1d": "8.6 seconds",
"7d": "1.0 minutes",
"30d": "4.3 minutes"
},
"burn_rate_alerts": [
{
"name": "Availability Burn Rate 2% Alert",
"description": "Alert when Availability is consuming error budget at 14.4x rate",
"severity": "critical",
"short_window": "5m",
"long_window": "1h",
"burn_rate_threshold": 14.4,
"budget_consumed": "2%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 14.4) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 14.4)",
"annotations": {
"summary": "High burn rate detected for Availability",
"description": "Error budget consumption rate is 14.4x normal, will exhaust 2% of monthly budget"
}
},
{
"name": "Availability Burn Rate 5% Alert",
"description": "Alert when Availability is consuming error budget at 6x rate",
"severity": "warning",
"short_window": "30m",
"long_window": "6h",
"burn_rate_threshold": 6,
"budget_consumed": "5%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 6) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 6)",
"annotations": {
"summary": "High burn rate detected for Availability",
"description": "Error budget consumption rate is 6x normal, will exhaust 5% of monthly budget"
}
},
{
"name": "Availability Burn Rate 10% Alert",
"description": "Alert when Availability is consuming error budget at 3x rate",
"severity": "info",
"short_window": "2h",
"long_window": "1d",
"burn_rate_threshold": 3,
"budget_consumed": "10%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 3) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 3)",
"annotations": {
"summary": "High burn rate detected for Availability",
"description": "Error budget consumption rate is 3x normal, will exhaust 10% of monthly budget"
}
},
{
"name": "Availability Burn Rate 10% Alert",
"description": "Alert when Availability is consuming error budget at 1x rate",
"severity": "info",
"short_window": "6h",
"long_window": "3d",
"burn_rate_threshold": 1,
"budget_consumed": "10%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 1) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 1)",
"annotations": {
"summary": "High burn rate detected for Availability",
"description": "Error budget consumption rate is 1x normal, will exhaust 10% of monthly budget"
}
}
]
},
{
"slo_name": "User Journey Success Rate SLO",
"error_budget_rate": 9.999999999998899e-05,
"error_budget_percentage": "0.010%",
"budgets_by_window": {
"1h": "0.4 seconds",
"1d": "8.6 seconds",
"7d": "1.0 minutes",
"30d": "4.3 minutes"
},
"burn_rate_alerts": [
{
"name": "User Journey Success Rate Burn Rate 2% Alert",
"description": "Alert when User Journey Success Rate is consuming error budget at 14.4x rate",
"severity": "critical",
"short_window": "5m",
"long_window": "1h",
"burn_rate_threshold": 14.4,
"budget_consumed": "2%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 14.4) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 14.4)",
"annotations": {
"summary": "High burn rate detected for User Journey Success Rate",
"description": "Error budget consumption rate is 14.4x normal, will exhaust 2% of monthly budget"
}
},
{
"name": "User Journey Success Rate Burn Rate 5% Alert",
"description": "Alert when User Journey Success Rate is consuming error budget at 6x rate",
"severity": "warning",
"short_window": "30m",
"long_window": "6h",
"burn_rate_threshold": 6,
"budget_consumed": "5%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 6) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 6)",
"annotations": {
"summary": "High burn rate detected for User Journey Success Rate",
"description": "Error budget consumption rate is 6x normal, will exhaust 5% of monthly budget"
}
},
{
"name": "User Journey Success Rate Burn Rate 10% Alert",
"description": "Alert when User Journey Success Rate is consuming error budget at 3x rate",
"severity": "info",
"short_window": "2h",
"long_window": "1d",
"burn_rate_threshold": 3,
"budget_consumed": "10%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 3) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 3)",
"annotations": {
"summary": "High burn rate detected for User Journey Success Rate",
"description": "Error budget consumption rate is 3x normal, will exhaust 10% of monthly budget"
}
},
{
"name": "User Journey Success Rate Burn Rate 10% Alert",
"description": "Alert when User Journey Success Rate is consuming error budget at 1x rate",
"severity": "info",
"short_window": "6h",
"long_window": "3d",
"burn_rate_threshold": 1,
"budget_consumed": "10%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 1) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 1)",
"annotations": {
"summary": "High burn rate detected for User Journey Success Rate",
"description": "Error budget consumption rate is 1x normal, will exhaust 10% of monthly budget"
}
}
]
},
{
"slo_name": "Feature Availability SLO",
"error_budget_rate": 9.999999999998899e-05,
"error_budget_percentage": "0.010%",
"budgets_by_window": {
"1h": "0.4 seconds",
"1d": "8.6 seconds",
"7d": "1.0 minutes",
"30d": "4.3 minutes"
},
"burn_rate_alerts": [
{
"name": "Feature Availability Burn Rate 2% Alert",
"description": "Alert when Feature Availability is consuming error budget at 14.4x rate",
"severity": "critical",
"short_window": "5m",
"long_window": "1h",
"burn_rate_threshold": 14.4,
"budget_consumed": "2%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 14.4) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 14.4)",
"annotations": {
"summary": "High burn rate detected for Feature Availability",
"description": "Error budget consumption rate is 14.4x normal, will exhaust 2% of monthly budget"
}
},
{
"name": "Feature Availability Burn Rate 5% Alert",
"description": "Alert when Feature Availability is consuming error budget at 6x rate",
"severity": "warning",
"short_window": "30m",
"long_window": "6h",
"burn_rate_threshold": 6,
"budget_consumed": "5%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 6) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 6)",
"annotations": {
"summary": "High burn rate detected for Feature Availability",
"description": "Error budget consumption rate is 6x normal, will exhaust 5% of monthly budget"
}
},
{
"name": "Feature Availability Burn Rate 10% Alert",
"description": "Alert when Feature Availability is consuming error budget at 3x rate",
"severity": "info",
"short_window": "2h",
"long_window": "1d",
"burn_rate_threshold": 3,
"budget_consumed": "10%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 3) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 3)",
"annotations": {
"summary": "High burn rate detected for Feature Availability",
"description": "Error budget consumption rate is 3x normal, will exhaust 10% of monthly budget"
}
},
{
"name": "Feature Availability Burn Rate 10% Alert",
"description": "Alert when Feature Availability is consuming error budget at 1x rate",
"severity": "info",
"short_window": "6h",
"long_window": "3d",
"burn_rate_threshold": 1,
"budget_consumed": "10%",
"condition": "((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_short > 1) and ((1 - (sum(rate(http_requests_total{service='payment-service',code!~'5..'})) / sum(rate(http_requests_total{service='payment-service'}))))_long > 1)",
"annotations": {
"summary": "High burn rate detected for Feature Availability",
"description": "Error budget consumption rate is 1x normal, will exhaust 10% of monthly budget"
}
}
]
}
],
"sla_recommendations": {
"applicable": true,
"service": "payment-service",
"commitments": [
{
"metric": "Availability",
"target": 0.9989,
"target_display": "99.89%",
"measurement_window": "monthly",
"measurement_method": "Uptime monitoring with 1-minute granularity"
},
{
"metric": "Feature Availability",
"target": 0.9989,
"target_display": "99.89%",
"measurement_window": "monthly",
"measurement_method": "Uptime monitoring with 1-minute granularity"
}
],
"penalties": [
{
"breach_threshold": "< 99.99%",
"credit_percentage": 10
},
{
"breach_threshold": "< 99.9%",
"credit_percentage": 25
},
{
"breach_threshold": "< 99%",
"credit_percentage": 50
}
],
"measurement_methodology": "External synthetic monitoring from multiple geographic locations",
"exclusions": [
"Planned maintenance windows (with 72h advance notice)",
"Customer-side network or infrastructure issues",
"Force majeure events",
"Third-party service dependencies beyond our control"
]
},
"monitoring_recommendations": {
"metrics": {
"collection": "Prometheus with service discovery",
"retention": "90 days for raw metrics, 1 year for aggregated",
"alerting": "Prometheus Alertmanager with multi-window burn rate alerts"
},
"logging": {
"format": "Structured JSON logs with correlation IDs",
"aggregation": "ELK stack or equivalent with proper indexing",
"retention": "30 days for debug logs, 90 days for error logs"
},
"tracing": {
"sampling": "Adaptive sampling with 1% base rate",
"storage": "Jaeger or Zipkin with 7-day retention",
"integration": "OpenTelemetry instrumentation"
}
},
"implementation_guide": {
"prerequisites": [
"Service instrumented with metrics collection (Prometheus format)",
"Structured logging with correlation IDs",
"Monitoring infrastructure (Prometheus, Grafana, Alertmanager)",
"Incident response processes and escalation policies"
],
"implementation_steps": [
{
"step": 1,
"title": "Instrument Service",
"description": "Add metrics collection for all defined SLIs",
"estimated_effort": "1-2 days"
},
{
"step": 2,
"title": "Configure Recording Rules",
"description": "Set up Prometheus recording rules for SLI calculations",
"estimated_effort": "4-8 hours"
},
{
"step": 3,
"title": "Implement Burn Rate Alerts",
"description": "Configure multi-window burn rate alerting rules",
"estimated_effort": "1 day"
},
{
"step": 4,
"title": "Create SLO Dashboard",
"description": "Build Grafana dashboard for SLO tracking and error budget monitoring",
"estimated_effort": "4-6 hours"
},
{
"step": 5,
"title": "Test and Validate",
"description": "Test alerting and validate SLI measurements against expectations",
"estimated_effort": "1-2 days"
},
{
"step": 6,
"title": "Documentation and Training",
"description": "Document runbooks and train team on SLO monitoring",
"estimated_effort": "1 day"
}
],
"validation_checklist": [
"All SLIs produce expected metric values",
"Burn rate alerts fire correctly during simulated outages",
"Error budget calculations match manual verification",
"Dashboard displays accurate SLO achievement rates",
"Alert routing reaches correct escalation paths",
"Runbooks are complete and tested"
]
}
}
FILE:README.md
# Observability Designer
A comprehensive toolkit for designing production-ready observability strategies including SLI/SLO frameworks, alert optimization, and dashboard generation.
## Overview
The Observability Designer skill provides three powerful Python scripts that help you create, optimize, and maintain observability systems:
- **SLO Designer**: Generate complete SLI/SLO frameworks with error budgets and burn rate alerts
- **Alert Optimizer**: Analyze and optimize existing alert configurations to reduce noise and improve effectiveness
- **Dashboard Generator**: Create comprehensive dashboard specifications with role-based layouts and drill-down paths
## Quick Start
### Prerequisites
- Python 3.7+
- No external dependencies required (uses Python standard library only)
### Basic Usage
```bash
# Generate SLO framework for a service
python3 scripts/slo_designer.py --service-type api --criticality critical --user-facing true --service-name payment-service
# Optimize existing alerts
python3 scripts/alert_optimizer.py --input assets/sample_alerts.json --analyze-only
# Generate a dashboard specification
python3 scripts/dashboard_generator.py --service-type web --name "Customer Portal" --role sre
```
## Scripts Documentation
### SLO Designer (`slo_designer.py`)
Generates comprehensive SLO frameworks based on service characteristics.
#### Features
- **Automatic SLI Selection**: Recommends appropriate SLIs based on service type
- **Target Setting**: Suggests SLO targets based on service criticality
- **Error Budget Calculation**: Computes error budgets and burn rate thresholds
- **Multi-Window Burn Rate Alerts**: Generates 4-window burn rate alerting rules
- **SLA Recommendations**: Provides customer-facing SLA guidance
#### Usage Examples
```bash
# From service definition file
python3 scripts/slo_designer.py --input assets/sample_service_api.json --output slo_framework.json
# From command line parameters
python3 scripts/slo_designer.py \
--service-type api \
--criticality critical \
--user-facing true \
--service-name payment-service \
--output payment_slos.json
# Generate and display summary only
python3 scripts/slo_designer.py --input assets/sample_service_web.json --summary-only
```
#### Service Definition Format
```json
{
"name": "payment-service",
"type": "api",
"criticality": "critical",
"user_facing": true,
"description": "Handles payment processing",
"team": "payments",
"environment": "production",
"dependencies": [
{
"name": "user-service",
"type": "api",
"criticality": "high"
}
]
}
```
#### Supported Service Types
- **api**: REST APIs, GraphQL services
- **web**: Web applications, SPAs
- **database**: Database services, data stores
- **queue**: Message queues, event streams
- **batch**: Batch processing jobs
- **ml**: Machine learning services
#### Criticality Levels
- **critical**: 99.99% availability, <100ms P95 latency, <0.1% error rate
- **high**: 99.9% availability, <200ms P95 latency, <0.5% error rate
- **medium**: 99.5% availability, <500ms P95 latency, <1% error rate
- **low**: 99% availability, <1s P95 latency, <2% error rate
### Alert Optimizer (`alert_optimizer.py`)
Analyzes existing alert configurations and provides optimization recommendations.
#### Features
- **Noise Detection**: Identifies alerts with high false positive rates
- **Coverage Analysis**: Finds gaps in monitoring coverage
- **Duplicate Detection**: Locates redundant or overlapping alerts
- **Threshold Analysis**: Reviews alert thresholds for appropriateness
- **Fatigue Assessment**: Evaluates alert volume and routing
#### Usage Examples
```bash
# Analyze existing alerts
python3 scripts/alert_optimizer.py --input assets/sample_alerts.json --analyze-only
# Generate optimized configuration
python3 scripts/alert_optimizer.py \
--input assets/sample_alerts.json \
--output optimized_alerts.json
# Generate HTML report
python3 scripts/alert_optimizer.py \
--input assets/sample_alerts.json \
--report alert_analysis.html \
--format html
```
#### Alert Configuration Format
```json
{
"alerts": [
{
"alert": "HighLatency",
"expr": "histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m])) > 0.5",
"for": "5m",
"labels": {
"severity": "warning",
"service": "payment-service"
},
"annotations": {
"summary": "High request latency detected",
"runbook_url": "https://runbooks.company.com/high-latency"
},
"historical_data": {
"fires_per_day": 2.5,
"false_positive_rate": 0.15
}
}
],
"services": [
{
"name": "payment-service",
"criticality": "critical"
}
]
}
```
#### Analysis Categories
- **Golden Signals**: Latency, traffic, errors, saturation
- **Resource Utilization**: CPU, memory, disk, network
- **Business Metrics**: Revenue, conversion, user engagement
- **Security**: Auth failures, suspicious activity
- **Availability**: Uptime, health checks
### Dashboard Generator (`dashboard_generator.py`)
Creates comprehensive dashboard specifications with role-based optimization.
#### Features
- **Role-Based Layouts**: Optimized for SRE, Developer, Executive, and Ops personas
- **Golden Signals Coverage**: Automatic inclusion of key monitoring metrics
- **Service-Type Specific Panels**: Tailored panels based on service characteristics
- **Interactive Elements**: Template variables, drill-down paths, time range controls
- **Grafana Compatibility**: Generates Grafana-compatible JSON
#### Usage Examples
```bash
# From service definition
python3 scripts/dashboard_generator.py \
--input assets/sample_service_web.json \
--output dashboard.json
# With specific role optimization
python3 scripts/dashboard_generator.py \
--service-type api \
--name "Payment Service" \
--role developer \
--output payment_dev_dashboard.json
# Generate Grafana-compatible JSON
python3 scripts/dashboard_generator.py \
--input assets/sample_service_api.json \
--output dashboard.json \
--format grafana
# With documentation
python3 scripts/dashboard_generator.py \
--service-type web \
--name "Customer Portal" \
--output portal_dashboard.json \
--doc-output portal_docs.md
```
#### Target Roles
- **sre**: Focus on availability, latency, errors, resource utilization
- **developer**: Emphasize latency, errors, throughput, business metrics
- **executive**: Highlight availability, business metrics, user experience
- **ops**: Priority on resource utilization, capacity, alerts, deployments
#### Panel Types
- **Stat**: Single value displays with thresholds
- **Gauge**: Resource utilization and capacity metrics
- **Timeseries**: Trend analysis and historical data
- **Table**: Top N lists and detailed breakdowns
- **Heatmap**: Distribution and correlation analysis
## Sample Data
The `assets/` directory contains sample configurations for testing:
- `sample_service_api.json`: Critical API service definition
- `sample_service_web.json`: High-priority web application definition
- `sample_alerts.json`: Alert configuration with optimization opportunities
The `expected_outputs/` directory shows example outputs from each script:
- `sample_slo_framework.json`: Complete SLO framework for API service
- `optimized_alerts.json`: Optimized alert configuration
- `sample_dashboard.json`: SRE dashboard specification
## Best Practices
### SLO Design
- Start with 1-2 SLOs per service and iterate
- Choose SLIs that directly impact user experience
- Set targets based on user needs, not technical capabilities
- Use error budgets to balance reliability and velocity
### Alert Optimization
- Every alert must be actionable
- Alert on symptoms, not causes
- Use multi-window burn rate alerts for SLO protection
- Implement proper escalation and routing policies
### Dashboard Design
- Follow the F-pattern for visual hierarchy
- Use consistent color semantics across dashboards
- Include drill-down paths for effective troubleshooting
- Optimize for the target role's specific needs
## Integration Patterns
### CI/CD Integration
```bash
# Generate SLOs during service onboarding
python3 scripts/slo_designer.py --input service-config.json --output slos.json
# Validate alert configurations in pipeline
python3 scripts/alert_optimizer.py --input alerts.json --analyze-only --report validation.html
# Auto-generate dashboards for new services
python3 scripts/dashboard_generator.py --input service-config.json --format grafana --output dashboard.json
```
### Monitoring Stack Integration
- **Prometheus**: Generated alert rules and recording rules
- **Grafana**: Dashboard JSON for direct import
- **Alertmanager**: Routing and escalation policies
- **PagerDuty**: Escalation configuration
### GitOps Workflow
1. Store service definitions in version control
2. Generate observability configurations in CI/CD
3. Deploy configurations via GitOps
4. Monitor effectiveness and iterate
## Advanced Usage
### Custom SLO Targets
Override default targets by including them in service definitions:
```json
{
"name": "special-service",
"type": "api",
"criticality": "high",
"custom_slos": {
"availability_target": 0.9995,
"latency_p95_target_ms": 150,
"error_rate_target": 0.002
}
}
```
### Alert Rule Templates
Use template variables for reusable alert rules:
```yaml
# Generated Prometheus alert rule
- alert: {{ service_name }}_HighLatency
expr: histogram_quantile(0.95, rate(http_request_duration_seconds_bucket{service="{{ service_name }}"}[5m])) > {{ latency_threshold }}
for: 5m
labels:
severity: warning
service: "{{ service_name }}"
```
### Dashboard Variants
Generate multiple dashboard variants for different use cases:
```bash
# SRE operational dashboard
python3 scripts/dashboard_generator.py --input service.json --role sre --output sre-dashboard.json
# Developer debugging dashboard
python3 scripts/dashboard_generator.py --input service.json --role developer --output dev-dashboard.json
# Executive business dashboard
python3 scripts/dashboard_generator.py --input service.json --role executive --output exec-dashboard.json
```
## Troubleshooting
### Common Issues
#### Script Execution Errors
- Ensure Python 3.7+ is installed
- Check file paths and permissions
- Validate JSON syntax in input files
#### Invalid Service Definitions
- Required fields: `name`, `type`, `criticality`
- Valid service types: `api`, `web`, `database`, `queue`, `batch`, `ml`
- Valid criticality levels: `critical`, `high`, `medium`, `low`
#### Missing Historical Data
- Alert historical data is optional but improves analysis
- Include `fires_per_day` and `false_positive_rate` when available
- Use monitoring system APIs to populate historical metrics
### Debug Mode
Enable verbose logging by setting environment variable:
```bash
export DEBUG=1
python3 scripts/slo_designer.py --input service.json
```
## Contributing
### Development Setup
```bash
# Clone the repository
git clone <repository-url>
cd engineering/observability-designer
# Run tests
python3 -m pytest tests/
# Lint code
python3 -m flake8 scripts/
```
### Adding New Features
1. Follow existing code patterns and error handling
2. Include comprehensive docstrings and type hints
3. Add test cases for new functionality
4. Update documentation and examples
## Support
For questions, issues, or feature requests:
- Check existing documentation and examples
- Review the reference materials in `references/`
- Open an issue with detailed reproduction steps
- Include sample configurations when reporting bugs
---
*This skill is part of the Claude Skills marketplace. For more information about observability best practices, see the reference documentation in the `references/` directory.*
FILE:references/alert_design_patterns.md
# Alert Design Patterns: A Guide to Effective Alerting
## Introduction
Well-designed alerts are the difference between a reliable system and 3 AM pages about non-issues. This guide provides patterns and anti-patterns for creating alerts that provide value without causing fatigue.
## Fundamental Principles
### The Golden Rules of Alerting
1. **Every alert should be actionable** - If you can't do something about it, don't alert
2. **Every alert should require human intelligence** - If a script can handle it, automate the response
3. **Every alert should be novel** - Don't alert on known, ongoing issues
4. **Every alert should represent a user-visible impact** - Internal metrics matter only if users are affected
### Alert Classification
#### Critical Alerts
- Service is completely down
- Data loss is occurring
- Security breach detected
- SLO burn rate indicates imminent SLO violation
#### Warning Alerts
- Service degradation affecting some users
- Approaching resource limits
- Dependent service issues
- Elevated error rates within SLO
#### Info Alerts
- Deployment notifications
- Capacity planning triggers
- Configuration changes
- Maintenance windows
## Alert Design Patterns
### Pattern 1: Symptoms, Not Causes
**Good**: Alert on user-visible symptoms
```yaml
- alert: HighLatency
expr: histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m])) > 0.5
for: 5m
annotations:
summary: "API latency is high"
description: "95th percentile latency is {{ $value }}s, above 500ms threshold"
```
**Bad**: Alert on internal metrics that may not affect users
```yaml
- alert: HighCPU
expr: cpu_usage > 80
# This might not affect users at all!
```
### Pattern 2: Multi-Window Alerting
Reduce false positives by requiring sustained problems:
```yaml
- alert: ServiceDown
expr: (
avg_over_time(up[2m]) == 0 # Short window: immediate detection
and
avg_over_time(up[10m]) < 0.8 # Long window: avoid flapping
)
for: 1m
```
### Pattern 3: Burn Rate Alerting
Alert based on error budget consumption rate:
```yaml
# Fast burn: 2% of monthly budget in 1 hour
- alert: ErrorBudgetFastBurn
expr: (
error_rate_5m > (14.4 * error_budget_slo)
and
error_rate_1h > (14.4 * error_budget_slo)
)
for: 2m
labels:
severity: critical
# Slow burn: 10% of monthly budget in 3 days
- alert: ErrorBudgetSlowBurn
expr: (
error_rate_6h > (1.0 * error_budget_slo)
and
error_rate_3d > (1.0 * error_budget_slo)
)
for: 15m
labels:
severity: warning
```
### Pattern 4: Hysteresis
Use different thresholds for firing and resolving to prevent flapping:
```yaml
- alert: HighErrorRate
expr: error_rate > 0.05 # Fire at 5%
for: 5m
# Resolution happens automatically when error_rate < 0.03 (3%)
# This prevents flapping around the 5% threshold
```
### Pattern 5: Composite Alerts
Alert when multiple conditions indicate a problem:
```yaml
- alert: ServiceDegraded
expr: (
(latency_p95 > latency_threshold)
or
(error_rate > error_threshold)
or
(availability < availability_threshold)
) and (
request_rate > min_request_rate # Only alert if we have traffic
)
```
### Pattern 6: Contextual Alerting
Include relevant context in alerts:
```yaml
- alert: DatabaseConnections
expr: db_connections_active / db_connections_max > 0.8
for: 5m
annotations:
summary: "Database connection pool nearly exhausted"
description: "{{ $labels.database }} has {{ $value | humanizePercentage }} connection utilization"
runbook_url: "https://runbooks.company.com/database-connections"
impact: "New requests may be rejected, causing 500 errors"
suggested_action: "Check for connection leaks or increase pool size"
```
## Alert Routing and Escalation
### Routing by Impact and Urgency
#### Critical Path Services
```yaml
route:
group_by: ['service']
routes:
- match:
service: 'payment-api'
severity: 'critical'
receiver: 'payment-team-pager'
continue: true
- match:
service: 'payment-api'
severity: 'warning'
receiver: 'payment-team-slack'
```
#### Time-Based Routing
```yaml
route:
routes:
- match:
severity: 'critical'
receiver: 'oncall-pager'
- match:
severity: 'warning'
time: 'business_hours' # 9 AM - 5 PM
receiver: 'team-slack'
- match:
severity: 'warning'
time: 'after_hours'
receiver: 'team-email' # Lower urgency outside business hours
```
### Escalation Patterns
#### Linear Escalation
```yaml
receivers:
- name: 'primary-oncall'
pagerduty_configs:
- escalation_policy: 'P1-Escalation'
# 0 min: Primary on-call
# 5 min: Secondary on-call
# 15 min: Engineering manager
# 30 min: Director of engineering
```
#### Severity-Based Escalation
```yaml
# Critical: Immediate escalation
- match:
severity: 'critical'
receiver: 'critical-escalation'
# Warning: Team-first escalation
- match:
severity: 'warning'
receiver: 'team-escalation'
```
## Alert Fatigue Prevention
### Grouping and Suppression
#### Time-Based Grouping
```yaml
route:
group_wait: 30s # Wait 30s to group similar alerts
group_interval: 2m # Send grouped alerts every 2 minutes
repeat_interval: 1h # Re-send unresolved alerts every hour
```
#### Dependent Service Suppression
```yaml
- alert: ServiceDown
expr: up == 0
- alert: HighLatency
expr: latency_p95 > 1
# This alert is suppressed when ServiceDown is firing
inhibit_rules:
- source_match:
alertname: 'ServiceDown'
target_match:
alertname: 'HighLatency'
equal: ['service']
```
### Alert Throttling
```yaml
# Limit to 1 alert per 10 minutes for noisy conditions
- alert: HighMemoryUsage
expr: memory_usage_percent > 85
for: 10m # Longer 'for' duration reduces noise
annotations:
summary: "Memory usage has been high for 10+ minutes"
```
### Smart Defaults
```yaml
# Use business logic to set intelligent thresholds
- alert: LowTraffic
expr: request_rate < (
avg_over_time(request_rate[7d]) * 0.1 # 10% of weekly average
)
# Only alert during business hours when low traffic is unusual
for: 30m
```
## Runbook Integration
### Runbook Structure Template
```markdown
# Alert: {{ $labels.alertname }}
## Immediate Actions
1. Check service status dashboard
2. Verify if users are affected
3. Look at recent deployments/changes
## Investigation Steps
1. Check logs for errors in the last 30 minutes
2. Verify dependent services are healthy
3. Check resource utilization (CPU, memory, disk)
4. Review recent alerts for patterns
## Resolution Actions
- If deployment-related: Consider rollback
- If resource-related: Scale up or optimize queries
- If dependency-related: Engage appropriate team
## Escalation
- Primary: @team-oncall
- Secondary: @engineering-manager
- Emergency: @site-reliability-team
```
### Runbook Integration in Alerts
```yaml
annotations:
runbook_url: "https://runbooks.company.com/alerts/{{ $labels.alertname }}"
quick_debug: |
1. curl -s https://{{ $labels.instance }}/health
2. kubectl logs {{ $labels.pod }} --tail=50
3. Check dashboard: https://grafana.company.com/d/service-{{ $labels.service }}
```
## Testing and Validation
### Alert Testing Strategies
#### Chaos Engineering Integration
```python
# Test that alerts fire during controlled failures
def test_alert_during_cpu_spike():
with chaos.cpu_spike(target='payment-api', duration='2m'):
assert wait_for_alert('HighCPU', timeout=180)
def test_alert_during_network_partition():
with chaos.network_partition(target='database'):
assert wait_for_alert('DatabaseUnreachable', timeout=60)
```
#### Historical Alert Analysis
```prometheus
# Query to find alerts that fired without incidents
count by (alertname) (
ALERTS{alertstate="firing"}[30d]
) unless on (alertname) (
count by (alertname) (
incident_created{source="alert"}[30d]
)
)
```
### Alert Quality Metrics
#### Alert Precision
```
Precision = True Positives / (True Positives + False Positives)
```
Track alerts that resulted in actual incidents vs false alarms.
#### Time to Resolution
```prometheus
# Average time from alert firing to resolution
avg_over_time(
(alert_resolved_timestamp - alert_fired_timestamp)[30d]
) by (alertname)
```
#### Alert Fatigue Indicators
```prometheus
# Alerts per day by team
sum by (team) (
increase(alerts_fired_total[1d])
)
# Percentage of alerts acknowledged within 15 minutes
sum(alerts_acked_within_15m) / sum(alerts_fired) * 100
```
## Advanced Patterns
### Machine Learning-Enhanced Alerting
#### Anomaly Detection
```yaml
- alert: AnomalousTraffic
expr: |
abs(request_rate - predict_linear(request_rate[1h], 300)) /
stddev_over_time(request_rate[1h]) > 3
for: 10m
annotations:
summary: "Traffic pattern is anomalous"
description: "Current traffic deviates from predicted pattern by >3 standard deviations"
```
#### Dynamic Thresholds
```yaml
- alert: DynamicHighLatency
expr: |
latency_p95 > (
quantile_over_time(0.95, latency_p95[7d]) + # Historical 95th percentile
2 * stddev_over_time(latency_p95[7d]) # Plus 2 standard deviations
)
```
### Business Hours Awareness
```yaml
# Different thresholds for business vs off hours
- alert: HighLatencyBusinessHours
expr: latency_p95 > 0.2 # Stricter during business hours
for: 2m
# Active 9 AM - 5 PM weekdays
- alert: HighLatencyOffHours
expr: latency_p95 > 0.5 # More lenient after hours
for: 5m
# Active nights and weekends
```
### Progressive Alerting
```yaml
# Escalating alert severity based on duration
- alert: ServiceLatencyElevated
expr: latency_p95 > 0.5
for: 5m
labels:
severity: info
- alert: ServiceLatencyHigh
expr: latency_p95 > 0.5
for: 15m # Same condition, longer duration
labels:
severity: warning
- alert: ServiceLatencyCritical
expr: latency_p95 > 0.5
for: 30m # Same condition, even longer duration
labels:
severity: critical
```
## Anti-Patterns to Avoid
### Anti-Pattern 1: Alerting on Everything
**Problem**: Too many alerts create noise and fatigue
**Solution**: Be selective; only alert on user-impacting issues
### Anti-Pattern 2: Vague Alert Messages
**Problem**: "Service X is down" - which instance? what's the impact?
**Solution**: Include specific details and context
### Anti-Pattern 3: Alerts Without Runbooks
**Problem**: Alerts that don't explain what to do
**Solution**: Every alert must have an associated runbook
### Anti-Pattern 4: Static Thresholds
**Problem**: 80% CPU might be normal during peak hours
**Solution**: Use contextual, adaptive thresholds
### Anti-Pattern 5: Ignoring Alert Quality
**Problem**: Accepting high false positive rates
**Solution**: Regularly review and tune alert precision
## Implementation Checklist
### Pre-Implementation
- [ ] Define alert severity levels and escalation policies
- [ ] Create runbook templates
- [ ] Set up alert routing configuration
- [ ] Define SLOs that alerts will protect
### Alert Development
- [ ] Each alert has clear success criteria
- [ ] Alert conditions tested against historical data
- [ ] Runbook created and accessible
- [ ] Severity and routing configured
- [ ] Context and suggested actions included
### Post-Implementation
- [ ] Monitor alert precision and recall
- [ ] Regular review of alert fatigue metrics
- [ ] Quarterly alert effectiveness review
- [ ] Team training on alert response procedures
### Quality Assurance
- [ ] Test alerts fire during controlled failures
- [ ] Verify alerts resolve when conditions improve
- [ ] Confirm runbooks are accurate and helpful
- [ ] Validate escalation paths work correctly
Remember: Great alerts are invisible when things work and invaluable when things break. Focus on quality over quantity, and always optimize for the human who will respond to the alert at 3 AM.
FILE:references/dashboard_best_practices.md
# Dashboard Best Practices: Design for Insight and Action
## Introduction
A well-designed dashboard is like a good story - it guides you through the data with purpose and clarity. This guide provides practical patterns for creating dashboards that inform decisions and enable quick troubleshooting.
## Design Principles
### The Hierarchy of Information
#### Primary Information (Top Third)
- Service health status
- SLO achievement
- Critical alerts
- Business KPIs
#### Secondary Information (Middle Third)
- Golden signals (latency, traffic, errors, saturation)
- Resource utilization
- Throughput and performance metrics
#### Tertiary Information (Bottom Third)
- Detailed breakdowns
- Historical trends
- Dependency status
- Debug information
### Visual Design Principles
#### Rule of 7±2
- Maximum 7±2 panels per screen
- Group related information together
- Use sections to organize complexity
#### Color Psychology
- **Red**: Critical issues, danger, immediate attention needed
- **Yellow/Orange**: Warnings, caution, degraded state
- **Green**: Healthy, normal operation, success
- **Blue**: Information, neutral metrics, capacity
- **Gray**: Disabled, unknown, or baseline states
#### Chart Selection Guide
- **Line charts**: Time series, trends, comparisons over time
- **Bar charts**: Categorical comparisons, top N lists
- **Gauges**: Single value with defined good/bad ranges
- **Stat panels**: Key metrics, percentages, counts
- **Heatmaps**: Distribution data, correlation analysis
- **Tables**: Detailed breakdowns, multi-dimensional data
## Dashboard Archetypes
### The Overview Dashboard
**Purpose**: High-level health check and business metrics
**Audience**: Executives, managers, cross-team stakeholders
**Update Frequency**: 5-15 minutes
```yaml
sections:
- title: "Business Health"
panels:
- service_availability_summary
- revenue_per_hour
- active_users
- conversion_rate
- title: "System Health"
panels:
- critical_alerts_count
- slo_achievement_summary
- error_budget_remaining
- deployment_status
```
### The SRE Operational Dashboard
**Purpose**: Real-time monitoring and incident response
**Audience**: SRE, on-call engineers
**Update Frequency**: 15-30 seconds
```yaml
sections:
- title: "Service Status"
panels:
- service_up_status
- active_incidents
- recent_deployments
- title: "Golden Signals"
panels:
- latency_percentiles
- request_rate
- error_rate
- resource_saturation
- title: "Infrastructure"
panels:
- cpu_memory_utilization
- network_io
- disk_space
```
### The Developer Debug Dashboard
**Purpose**: Deep-dive troubleshooting and performance analysis
**Audience**: Development teams
**Update Frequency**: 30 seconds - 2 minutes
```yaml
sections:
- title: "Application Performance"
panels:
- endpoint_latency_breakdown
- database_query_performance
- cache_hit_rates
- queue_depths
- title: "Errors and Logs"
panels:
- error_rate_by_endpoint
- log_volume_by_level
- exception_types
- slow_queries
```
## Layout Patterns
### The F-Pattern Layout
Based on eye-tracking studies, users scan in an F-pattern:
```
[Critical Status] [SLO Summary ] [Error Budget ]
[Latency ] [Traffic ] [Errors ]
[Saturation ] [Resource Use ] [Detailed View]
[Historical ] [Dependencies ] [Debug Info ]
```
### The Z-Pattern Layout
For executive dashboards, follow the Z-pattern:
```
[Business KPIs ] → [System Status]
↓ ↓
[Trend Analysis ] ← [Key Metrics ]
```
### Responsive Design
#### Desktop (1920x1080)
- 24-column grid
- Panels can be 6, 8, 12, or 24 units wide
- 4-6 rows visible without scrolling
#### Laptop (1366x768)
- Stack wider panels vertically
- Reduce panel heights
- Prioritize most critical information
#### Mobile (768px width)
- Single column layout
- Simplified panels
- Touch-friendly controls
## Effective Panel Design
### Stat Panels
```yaml
# Good: Clear value with context
- title: "API Availability"
type: stat
targets:
- expr: avg(up{service="api"}) * 100
field_config:
unit: percent
thresholds:
steps:
- color: red
value: 0
- color: yellow
value: 99
- color: green
value: 99.9
options:
color_mode: background
text_mode: value_and_name
```
### Time Series Panels
```yaml
# Good: Multiple related metrics with clear legend
- title: "Request Latency"
type: timeseries
targets:
- expr: histogram_quantile(0.50, rate(http_duration_bucket[5m]))
legend: "P50"
- expr: histogram_quantile(0.95, rate(http_duration_bucket[5m]))
legend: "P95"
- expr: histogram_quantile(0.99, rate(http_duration_bucket[5m]))
legend: "P99"
field_config:
unit: ms
custom:
draw_style: line
fill_opacity: 10
options:
legend:
display_mode: table
placement: bottom
values: [min, max, mean, last]
```
### Table Panels
```yaml
# Good: Top N with relevant columns
- title: "Slowest Endpoints"
type: table
targets:
- expr: topk(10, histogram_quantile(0.95, sum by (handler)(rate(http_duration_bucket[5m]))))
format: table
instant: true
transformations:
- id: organize
options:
exclude_by_name:
Time: true
rename_by_name:
Value: "P95 Latency (ms)"
handler: "Endpoint"
```
## Color and Visualization Best Practices
### Threshold Configuration
```yaml
# Traffic light system with meaningful boundaries
thresholds:
steps:
- color: green # Good performance
value: null # Default
- color: yellow # Degraded performance
value: 95 # 95th percentile of historical normal
- color: orange # Poor performance
value: 99 # 99th percentile of historical normal
- color: red # Critical performance
value: 99.9 # Worst case scenario
```
### Color Blind Friendly Palettes
```yaml
# Use patterns and shapes in addition to color
field_config:
overrides:
- matcher:
id: byName
options: "Critical"
properties:
- id: color
value:
mode: fixed
fixed_color: "#d73027" # Red-orange for protanopia
- id: custom.draw_style
value: "points" # Different shape
```
### Consistent Color Semantics
- **Success/Health**: Green (#28a745)
- **Warning/Degraded**: Yellow (#ffc107)
- **Error/Critical**: Red (#dc3545)
- **Information**: Blue (#007bff)
- **Neutral**: Gray (#6c757d)
## Time Range Strategy
### Default Time Ranges by Dashboard Type
#### Real-time Operational
- **Default**: Last 15 minutes
- **Quick options**: 5m, 15m, 1h, 4h
- **Auto-refresh**: 15-30 seconds
#### Troubleshooting
- **Default**: Last 1 hour
- **Quick options**: 15m, 1h, 4h, 12h, 1d
- **Auto-refresh**: 1 minute
#### Business Review
- **Default**: Last 24 hours
- **Quick options**: 1d, 7d, 30d, 90d
- **Auto-refresh**: 5 minutes
#### Capacity Planning
- **Default**: Last 7 days
- **Quick options**: 7d, 30d, 90d, 1y
- **Auto-refresh**: 15 minutes
### Time Range Annotations
```yaml
# Add context for time-based events
annotations:
- name: "Deployments"
datasource: "Prometheus"
expr: "deployment_timestamp"
title_format: "Deploy {{ version }}"
text_format: "Deployed version {{ version }} to {{ environment }}"
- name: "Incidents"
datasource: "Incident API"
query: "incidents.json?service={{ service }}"
color: "red"
```
## Interactive Features
### Template Variables
```yaml
# Service selector
- name: service
type: query
query: label_values(up, service)
current:
text: All
value: $__all
include_all: true
multi: true
# Environment selector
- name: environment
type: query
query: label_values(up{service="$service"}, environment)
current:
text: production
value: production
```
### Drill-Down Links
```yaml
# Panel-level drill-downs
- title: "Error Rate"
type: timeseries
# ... other config ...
options:
data_links:
- title: "View Error Logs"
url: "/d/logs-dashboard?var-service=__field.labels.service&from=__from&to=__to"
- title: "Error Traces"
url: "/d/traces-dashboard?var-service=__field.labels.service"
```
### Dynamic Panel Titles
```yaml
- title: "service - Request Rate" # Uses template variable
type: timeseries
# Title updates automatically when service variable changes
```
## Performance Optimization
### Query Optimization
#### Use Recording Rules
```yaml
# Instead of complex queries in dashboards
groups:
- name: http_requests
rules:
- record: http_request_rate_5m
expr: sum(rate(http_requests_total[5m])) by (service, method, handler)
- record: http_request_latency_p95_5m
expr: histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (service, le))
```
#### Limit Data Points
```yaml
# Good: Reasonable resolution for dashboard
- expr: http_request_rate_5m[1h]
interval: 15s # One point every 15 seconds
# Bad: Too many points for visualization
- expr: http_request_rate_1s[1h] # 3600 points!
```
### Dashboard Performance
#### Panel Limits
- **Maximum panels per dashboard**: 20-30
- **Maximum queries per panel**: 10
- **Maximum time series per panel**: 50
#### Caching Strategy
```yaml
# Use appropriate cache headers
cache_timeout: 30 # Cache for 30 seconds on fast-changing panels
cache_timeout: 300 # Cache for 5 minutes on slow-changing panels
```
## Accessibility
### Screen Reader Support
```yaml
# Provide text alternatives for visual elements
- title: "Service Health Status"
type: stat
options:
text_mode: value_and_name # Includes both value and description
field_config:
mappings:
- options:
"1":
text: "Healthy"
color: "green"
"0":
text: "Unhealthy"
color: "red"
```
### Keyboard Navigation
- Ensure all interactive elements are keyboard accessible
- Provide logical tab order
- Include skip links for complex dashboards
### High Contrast Mode
```yaml
# Test dashboards work in high contrast mode
theme: high_contrast
colors:
- "#000000" # Pure black
- "#ffffff" # Pure white
- "#ffff00" # Pure yellow
- "#ff0000" # Pure red
```
## Testing and Validation
### Dashboard Testing Checklist
#### Functional Testing
- [ ] All panels load without errors
- [ ] Template variables filter correctly
- [ ] Time range changes update all panels
- [ ] Drill-down links work as expected
- [ ] Auto-refresh functions properly
#### Visual Testing
- [ ] Dashboard renders correctly on different screen sizes
- [ ] Colors are distinguishable and meaningful
- [ ] Text is readable at normal zoom levels
- [ ] Legends and labels are clear
#### Performance Testing
- [ ] Dashboard loads in < 5 seconds
- [ ] No queries timeout under normal load
- [ ] Auto-refresh doesn't cause browser lag
- [ ] Memory usage remains reasonable
#### Usability Testing
- [ ] New team members can understand the dashboard
- [ ] Action items are clear during incidents
- [ ] Key information is quickly discoverable
- [ ] Dashboard supports common troubleshooting workflows
## Maintenance and Governance
### Dashboard Lifecycle
#### Creation
1. Define dashboard purpose and audience
2. Identify key metrics and success criteria
3. Design layout following established patterns
4. Implement with consistent styling
5. Test with real data and user scenarios
#### Maintenance
- **Weekly**: Check for broken panels or queries
- **Monthly**: Review dashboard usage analytics
- **Quarterly**: Gather user feedback and iterate
- **Annually**: Major review and potential redesign
#### Retirement
- Archive dashboards that are no longer used
- Migrate users to replacement dashboards
- Document lessons learned
### Dashboard Standards
```yaml
# Organization dashboard standards
standards:
naming_convention: "[Team] [Service] - [Purpose]"
tags: [team, service_type, environment, purpose]
refresh_intervals: [15s, 30s, 1m, 5m, 15m]
time_ranges: [5m, 15m, 1h, 4h, 1d, 7d, 30d]
color_scheme: "company_standard"
max_panels_per_dashboard: 25
```
## Advanced Patterns
### Composite Dashboards
```yaml
# Dashboard that includes panels from other dashboards
- title: "Service Overview"
type: dashlist
targets:
- "service-health"
- "service-performance"
- "service-business-metrics"
options:
show_headings: true
max_items: 10
```
### Dynamic Dashboard Generation
```python
# Generate dashboards from service definitions
def generate_service_dashboard(service_config):
panels = []
# Always include golden signals
panels.extend(generate_golden_signals_panels(service_config))
# Add service-specific panels
if service_config.type == 'database':
panels.extend(generate_database_panels(service_config))
elif service_config.type == 'queue':
panels.extend(generate_queue_panels(service_config))
return {
'title': f"{service_config.name} - Operational Dashboard",
'panels': panels,
'variables': generate_variables(service_config)
}
```
### A/B Testing for Dashboards
```yaml
# Test different dashboard designs with different teams
experiment:
name: "dashboard_layout_test"
variants:
- name: "traditional_layout"
weight: 50
config: "dashboard_v1.json"
- name: "f_pattern_layout"
weight: 50
config: "dashboard_v2.json"
success_metrics:
- "time_to_insight"
- "user_satisfaction"
- "troubleshooting_efficiency"
```
Remember: A dashboard should tell a story about your system's health and guide users toward the right actions. Focus on clarity over complexity, and always optimize for the person who will use it during a stressful incident.
FILE:references/slo_cookbook.md
# SLO Cookbook: A Practical Guide to Service Level Objectives
## Introduction
Service Level Objectives (SLOs) are a key tool for managing service reliability. This cookbook provides practical guidance for implementing SLOs that actually improve system reliability rather than just creating meaningless metrics.
## Fundamentals
### The SLI/SLO/SLA Hierarchy
- **SLI (Service Level Indicator)**: A quantifiable measure of service quality
- **SLO (Service Level Objective)**: A target range of values for an SLI
- **SLA (Service Level Agreement)**: A business agreement with consequences for missing SLO targets
### Golden Rule of SLOs
**Start simple, iterate based on learning.** Your first SLOs won't be perfect, and that's okay.
## Choosing Good SLIs
### The Four Golden Signals
1. **Latency**: How long requests take to complete
2. **Traffic**: How many requests are coming in
3. **Errors**: How many requests are failing
4. **Saturation**: How "full" your service is
### SLI Selection Criteria
A good SLI should be:
- **Measurable**: You can collect data for it
- **Meaningful**: It reflects user experience
- **Controllable**: You can take action to improve it
- **Proportional**: Changes in the SLI reflect changes in user happiness
### Service Type Specific SLIs
#### HTTP APIs
- **Request latency**: P95 or P99 response time
- **Availability**: Proportion of successful requests (non-5xx)
- **Throughput**: Requests per second capacity
```prometheus
# Availability SLI
sum(rate(http_requests_total{code!~"5.."}[5m])) / sum(rate(http_requests_total[5m]))
# Latency SLI
histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m]))
```
#### Batch Jobs
- **Freshness**: Age of the last successful run
- **Correctness**: Proportion of jobs completing successfully
- **Throughput**: Items processed per unit time
#### Data Pipelines
- **Data freshness**: Time since last successful update
- **Data quality**: Proportion of records passing validation
- **Processing latency**: Time from ingestion to availability
### Anti-Patterns in SLI Selection
❌ **Don't use**: CPU usage, memory usage, disk space as primary SLIs
- These are symptoms, not user-facing impacts
❌ **Don't use**: Counts instead of rates or proportions
- "Number of errors" vs "Error rate"
❌ **Don't use**: Internal metrics that users don't care about
- Queue depth, cache hit rate (unless they directly impact user experience)
## Setting SLO Targets
### The Art of Target Setting
Setting SLO targets is balancing act between:
- **User happiness**: Targets should reflect acceptable user experience
- **Business value**: Tighter SLOs cost more to maintain
- **Current performance**: Targets should be achievable but aspirational
### Target Setting Strategies
#### Historical Performance Method
1. Collect 4-6 weeks of historical data
2. Calculate the worst user-visible performance in that period
3. Set your SLO slightly better than the worst acceptable performance
#### User Journey Mapping
1. Map critical user journeys
2. Identify acceptable performance for each step
3. Work backwards to component SLOs
#### Error Budget Approach
1. Decide how much unreliability you can afford
2. Set SLO targets based on acceptable error budget consumption
3. Example: 99.9% availability = 43.8 minutes downtime per month
### SLO Target Examples by Service Criticality
#### Critical Services (Revenue Impact)
- **Availability**: 99.95% - 99.99%
- **Latency (P95)**: 100-200ms
- **Error Rate**: < 0.1%
#### High Priority Services
- **Availability**: 99.9% - 99.95%
- **Latency (P95)**: 200-500ms
- **Error Rate**: < 0.5%
#### Standard Services
- **Availability**: 99.5% - 99.9%
- **Latency (P95)**: 500ms - 1s
- **Error Rate**: < 1%
## Error Budget Management
### What is an Error Budget?
Your error budget is the maximum amount of unreliability you can accumulate while still meeting your SLO. It's calculated as:
```
Error Budget = (1 - SLO) × Time Window
```
For a 99.9% availability SLO over 30 days:
```
Error Budget = (1 - 0.999) × 30 days = 0.001 × 30 days = 43.8 minutes
```
### Error Budget Policies
Define what happens when you consume your error budget:
#### Conservative Policy (High-Risk Services)
- **> 50% consumed**: Freeze non-critical feature releases
- **> 75% consumed**: Focus entirely on reliability improvements
- **> 90% consumed**: Consider emergency measures (traffic shaping, etc.)
#### Balanced Policy (Standard Services)
- **> 75% consumed**: Increase focus on reliability work
- **> 90% consumed**: Pause feature work, focus on reliability
#### Aggressive Policy (Early Stage Services)
- **> 90% consumed**: Review but continue normal operations
- **100% consumed**: Evaluate SLO appropriateness
### Burn Rate Alerting
Multi-window burn rate alerts help you catch SLO violations before they become critical:
```yaml
# Fast burn: 2% budget consumed in 1 hour
- alert: FastBurnSLOViolation
expr: (
(1 - (sum(rate(http_requests_total{code!~"5.."}[5m])) / sum(rate(http_requests_total[5m])))) > (14.4 * 0.001)
and
(1 - (sum(rate(http_requests_total{code!~"5.."}[1h])) / sum(rate(http_requests_total[1h])))) > (14.4 * 0.001)
)
for: 2m
# Slow burn: 10% budget consumed in 3 days
- alert: SlowBurnSLOViolation
expr: (
(1 - (sum(rate(http_requests_total{code!~"5.."}[6h])) / sum(rate(http_requests_total[6h])))) > (1.0 * 0.001)
and
(1 - (sum(rate(http_requests_total{code!~"5.."}[3d])) / sum(rate(http_requests_total[3d])))) > (1.0 * 0.001)
)
for: 15m
```
## Implementation Patterns
### The SLO Implementation Ladder
#### Level 1: Basic SLOs
- Choose 1-2 SLIs that matter most to users
- Set aspirational but achievable targets
- Implement basic alerting when SLOs are missed
#### Level 2: Operational SLOs
- Add burn rate alerting
- Create error budget dashboards
- Establish error budget policies
- Regular SLO review meetings
#### Level 3: Advanced SLOs
- Multi-window burn rate alerts
- Automated error budget policy enforcement
- SLO-driven incident prioritization
- Integration with CI/CD for deployment decisions
### SLO Measurement Architecture
#### Push vs Pull Metrics
- **Pull** (Prometheus): Good for infrastructure metrics, real-time alerting
- **Push** (StatsD): Good for application metrics, business events
#### Measurement Points
- **Server-side**: More reliable, easier to implement
- **Client-side**: Better reflects user experience
- **Synthetic**: Consistent, predictable, may not reflect real user experience
### SLO Dashboard Design
Essential elements for SLO dashboards:
1. **Current SLO Achievement**: Large, prominent display
2. **Error Budget Remaining**: Visual indicator (gauge, progress bar)
3. **Burn Rate**: Time series showing error budget consumption rate
4. **Historical Trends**: 4-week view of SLO achievement
5. **Alerts**: Current and recent SLO-related alerts
## Advanced Topics
### Dependency SLOs
For services with dependencies:
```
SLO_service ≤ min(SLO_inherent, ∏SLO_dependencies)
```
If your service depends on 3 other services each with 99.9% SLO:
```
Maximum_SLO = 0.999³ = 0.997 = 99.7%
```
### User Journey SLOs
Track end-to-end user experiences:
```prometheus
# Registration success rate
sum(rate(user_registration_success_total[5m])) / sum(rate(user_registration_attempts_total[5m]))
# Purchase completion latency
histogram_quantile(0.95, rate(purchase_completion_duration_seconds_bucket[5m]))
```
### SLOs for Batch Systems
Special considerations for non-request/response systems:
#### Freshness SLO
```prometheus
# Data should be no more than 4 hours old
(time() - last_successful_update_timestamp) < (4 * 3600)
```
#### Throughput SLO
```prometheus
# Should process at least 1000 items per hour
rate(items_processed_total[1h]) >= 1000
```
#### Quality SLO
```prometheus
# At least 99.5% of records should pass validation
sum(rate(records_valid_total[5m])) / sum(rate(records_processed_total[5m])) >= 0.995
```
## Common Mistakes and How to Avoid Them
### Mistake 1: Too Many SLOs
**Problem**: Drowning in metrics, losing focus
**Solution**: Start with 1-2 SLOs per service, add more only when needed
### Mistake 2: Internal Metrics as SLIs
**Problem**: Optimizing for metrics that don't impact users
**Solution**: Always ask "If this metric changes, do users notice?"
### Mistake 3: Perfectionist SLOs
**Problem**: 99.99% SLO when 99.9% would be fine
**Solution**: Higher SLOs cost exponentially more; pick the minimum acceptable level
### Mistake 4: Ignoring Error Budgets
**Problem**: Treating any SLO miss as an emergency
**Solution**: Error budgets exist to be spent; use them to balance feature velocity and reliability
### Mistake 5: Static SLOs
**Problem**: Setting SLOs once and never updating them
**Solution**: Review SLOs quarterly; adjust based on user feedback and business changes
## SLO Review Process
### Monthly SLO Review Agenda
1. **SLO Achievement Review**: Did we meet our SLOs?
2. **Error Budget Analysis**: How did we spend our error budget?
3. **Incident Correlation**: Which incidents impacted our SLOs?
4. **SLI Quality Assessment**: Are our SLIs still meaningful?
5. **Target Adjustment**: Should we change any targets?
### Quarterly SLO Health Check
1. **User Impact Validation**: Survey users about acceptable performance
2. **Business Alignment**: Do SLOs still reflect business priorities?
3. **Measurement Quality**: Are we measuring the right things?
4. **Cost/Benefit Analysis**: Are tighter SLOs worth the investment?
## Tooling and Automation
### Essential Tools
1. **Metrics Collection**: Prometheus, InfluxDB, CloudWatch
2. **Alerting**: Alertmanager, PagerDuty, OpsGenie
3. **Dashboards**: Grafana, DataDog, New Relic
4. **SLO Platforms**: Sloth, Pyrra, Service Level Blue
### Automation Opportunities
- **Burn rate alert generation** from SLO definitions
- **Dashboard creation** from SLO specifications
- **Error budget calculation** and tracking
- **Release blocking** based on error budget consumption
## Getting Started Checklist
- [ ] Identify your service's critical user journeys
- [ ] Choose 1-2 SLIs that best reflect user experience
- [ ] Collect 4-6 weeks of baseline data
- [ ] Set initial SLO targets based on historical performance
- [ ] Implement basic SLO monitoring and alerting
- [ ] Create an SLO dashboard
- [ ] Define error budget policies
- [ ] Schedule monthly SLO reviews
- [ ] Plan for quarterly SLO health checks
Remember: SLOs are a journey, not a destination. Start simple, learn from experience, and iterate toward better reliability management.
FILE:scripts/alert_optimizer.py
#!/usr/bin/env python3
"""
Alert Optimizer - Analyze and optimize alert configurations
This script analyzes existing alert configurations and identifies optimization opportunities:
- Noisy alerts with high false positive rates
- Missing coverage gaps in monitoring
- Duplicate or redundant alerts
- Poor threshold settings and alert fatigue risks
- Missing runbooks and documentation
- Routing and escalation policy improvements
Usage:
python alert_optimizer.py --input alert_config.json --output optimized_config.json
python alert_optimizer.py --input alerts.json --analyze-only --report report.html
"""
import json
import argparse
import sys
import re
import math
from typing import Dict, List, Any, Tuple, Set
from datetime import datetime, timedelta
from collections import defaultdict, Counter
class AlertOptimizer:
"""Analyze and optimize alert configurations."""
# Alert severity priority mapping
SEVERITY_PRIORITY = {
'critical': 1,
'high': 2,
'warning': 3,
'info': 4
}
# Common noisy alert patterns
NOISY_PATTERNS = [
r'disk.*usage.*>.*[89]\d%', # Disk usage > 80% often noisy
r'memory.*>.*[89]\d%', # Memory > 80% often noisy
r'cpu.*>.*[789]\d%', # CPU > 70% can be noisy
r'response.*time.*>.*\d+ms', # Low latency thresholds
r'error.*rate.*>.*0\.[01]%' # Very low error rate thresholds
]
# Essential monitoring categories
COVERAGE_CATEGORIES = [
'availability',
'latency',
'error_rate',
'resource_utilization',
'security',
'business_metrics'
]
# Golden signals that should always be monitored
GOLDEN_SIGNALS = [
'latency',
'traffic',
'errors',
'saturation'
]
def __init__(self):
"""Initialize the Alert Optimizer."""
self.alert_config = {}
self.optimization_results = {}
self.alert_analysis = {}
def load_alert_config(self, file_path: str) -> Dict[str, Any]:
"""Load alert configuration from JSON file."""
try:
with open(file_path, 'r') as f:
return json.load(f)
except FileNotFoundError:
raise ValueError(f"Alert configuration file not found: {file_path}")
except json.JSONDecodeError as e:
raise ValueError(f"Invalid JSON in alert configuration: {e}")
def analyze_alert_noise(self, alerts: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Identify potentially noisy alerts."""
noisy_alerts = []
for alert in alerts:
noise_score = 0
noise_reasons = []
alert_rule = alert.get('expr', alert.get('condition', ''))
alert_name = alert.get('alert', alert.get('name', 'Unknown'))
# Check for common noisy patterns
for pattern in self.NOISY_PATTERNS:
if re.search(pattern, alert_rule, re.IGNORECASE):
noise_score += 3
noise_reasons.append(f"Matches noisy pattern: {pattern}")
# Check for very frequent evaluation intervals
evaluation_interval = alert.get('for', '0s')
if self._parse_duration(evaluation_interval) < 60: # Less than 1 minute
noise_score += 2
noise_reasons.append("Very short evaluation interval")
# Check for lack of 'for' clause
if not alert.get('for') or alert.get('for') == '0s':
noise_score += 2
noise_reasons.append("No 'for' clause - may cause alert flapping")
# Check for overly sensitive thresholds
if self._has_sensitive_threshold(alert_rule):
noise_score += 2
noise_reasons.append("Potentially sensitive threshold")
# Check historical firing rate if available
historical_data = alert.get('historical_data', {})
if historical_data:
firing_rate = historical_data.get('fires_per_day', 0)
if firing_rate > 10: # More than 10 fires per day
noise_score += 3
noise_reasons.append(f"High firing rate: {firing_rate} times/day")
false_positive_rate = historical_data.get('false_positive_rate', 0)
if false_positive_rate > 0.3: # > 30% false positives
noise_score += 4
noise_reasons.append(f"High false positive rate: {false_positive_rate*100:.1f}%")
if noise_score >= 3: # Threshold for considering an alert noisy
noisy_alert = {
'alert_name': alert_name,
'noise_score': noise_score,
'reasons': noise_reasons,
'current_rule': alert_rule,
'recommendations': self._generate_noise_reduction_recommendations(alert, noise_reasons)
}
noisy_alerts.append(noisy_alert)
return sorted(noisy_alerts, key=lambda x: x['noise_score'], reverse=True)
def _parse_duration(self, duration_str: str) -> int:
"""Parse duration string to seconds."""
if not duration_str or duration_str == '0s':
return 0
duration_map = {'s': 1, 'm': 60, 'h': 3600, 'd': 86400}
match = re.match(r'(\d+)([smhd])', duration_str)
if match:
value, unit = match.groups()
return int(value) * duration_map.get(unit, 1)
return 0
def _has_sensitive_threshold(self, rule: str) -> bool:
"""Check if alert rule has potentially sensitive thresholds."""
# Look for very low error rates or very tight latency thresholds
sensitive_patterns = [
r'error.*rate.*>.*0\.0[01]', # Error rate > 0.01% or 0.001%
r'latency.*>.*[12]\d\d?ms', # Latency > 100-299ms
r'response.*time.*>.*0\.[12]', # Response time > 0.1-0.2s
r'cpu.*>.*[456]\d%' # CPU > 40-69% (too sensitive for most cases)
]
for pattern in sensitive_patterns:
if re.search(pattern, rule, re.IGNORECASE):
return True
return False
def _generate_noise_reduction_recommendations(self, alert: Dict[str, Any],
reasons: List[str]) -> List[str]:
"""Generate recommendations to reduce alert noise."""
recommendations = []
if "No 'for' clause" in str(reasons):
recommendations.append("Add 'for: 5m' clause to prevent flapping")
if "Very short evaluation interval" in str(reasons):
recommendations.append("Increase evaluation interval to at least 1 minute")
if "sensitive threshold" in str(reasons):
recommendations.append("Review and increase threshold based on historical data")
if "High firing rate" in str(reasons):
recommendations.append("Analyze historical firing patterns and adjust thresholds")
if "High false positive rate" in str(reasons):
recommendations.append("Implement more specific conditions to reduce false positives")
if "noisy pattern" in str(reasons):
recommendations.append("Consider using percentile-based thresholds instead of absolute values")
return recommendations
def identify_coverage_gaps(self, alerts: List[Dict[str, Any]],
services: List[Dict[str, Any]] = None) -> Dict[str, Any]:
"""Identify gaps in monitoring coverage."""
coverage_analysis = {
'missing_categories': [],
'missing_golden_signals': [],
'service_coverage_gaps': [],
'critical_gaps': [],
'recommendations': []
}
# Analyze coverage by category
covered_categories = set()
alert_categories = []
for alert in alerts:
alert_rule = alert.get('expr', alert.get('condition', ''))
alert_name = alert.get('alert', alert.get('name', ''))
category = self._classify_alert_category(alert_rule, alert_name)
if category:
covered_categories.add(category)
alert_categories.append(category)
# Check for missing essential categories
missing_categories = set(self.COVERAGE_CATEGORIES) - covered_categories
coverage_analysis['missing_categories'] = list(missing_categories)
# Check for missing golden signals
covered_signals = set()
for alert in alerts:
alert_rule = alert.get('expr', alert.get('condition', ''))
signal = self._identify_golden_signal(alert_rule)
if signal:
covered_signals.add(signal)
missing_signals = set(self.GOLDEN_SIGNALS) - covered_signals
coverage_analysis['missing_golden_signals'] = list(missing_signals)
# Analyze service-specific coverage if service list provided
if services:
service_coverage = self._analyze_service_coverage(alerts, services)
coverage_analysis['service_coverage_gaps'] = service_coverage
# Identify critical gaps
critical_gaps = []
if 'availability' in missing_categories:
critical_gaps.append("Missing availability monitoring")
if 'error_rate' in missing_categories:
critical_gaps.append("Missing error rate monitoring")
if 'errors' in missing_signals:
critical_gaps.append("Missing error signal monitoring")
coverage_analysis['critical_gaps'] = critical_gaps
# Generate recommendations
recommendations = self._generate_coverage_recommendations(coverage_analysis)
coverage_analysis['recommendations'] = recommendations
return coverage_analysis
def _classify_alert_category(self, rule: str, alert_name: str) -> str:
"""Classify alert into monitoring category."""
rule_lower = rule.lower()
name_lower = alert_name.lower()
if any(keyword in rule_lower or keyword in name_lower
for keyword in ['up', 'down', 'available', 'reachable']):
return 'availability'
if any(keyword in rule_lower or keyword in name_lower
for keyword in ['latency', 'response_time', 'duration']):
return 'latency'
if any(keyword in rule_lower or keyword in name_lower
for keyword in ['error', 'fail', '5xx', '4xx']):
return 'error_rate'
if any(keyword in rule_lower or keyword in name_lower
for keyword in ['cpu', 'memory', 'disk', 'network', 'utilization']):
return 'resource_utilization'
if any(keyword in rule_lower or keyword in name_lower
for keyword in ['security', 'auth', 'login', 'breach']):
return 'security'
if any(keyword in rule_lower or keyword in name_lower
for keyword in ['revenue', 'conversion', 'user', 'business']):
return 'business_metrics'
return 'other'
def _identify_golden_signal(self, rule: str) -> str:
"""Identify which golden signal an alert covers."""
rule_lower = rule.lower()
if any(keyword in rule_lower for keyword in ['latency', 'response_time', 'duration']):
return 'latency'
if any(keyword in rule_lower for keyword in ['rate', 'rps', 'qps', 'throughput']):
return 'traffic'
if any(keyword in rule_lower for keyword in ['error', 'fail', '5xx']):
return 'errors'
if any(keyword in rule_lower for keyword in ['cpu', 'memory', 'disk', 'utilization']):
return 'saturation'
return None
def _analyze_service_coverage(self, alerts: List[Dict[str, Any]],
services: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Analyze monitoring coverage per service."""
service_coverage = []
for service in services:
service_name = service.get('name', '')
service_alerts = [alert for alert in alerts
if service_name in alert.get('expr', '') or
service_name in alert.get('labels', {}).get('service', '')]
covered_signals = set()
for alert in service_alerts:
signal = self._identify_golden_signal(alert.get('expr', ''))
if signal:
covered_signals.add(signal)
missing_signals = set(self.GOLDEN_SIGNALS) - covered_signals
if missing_signals or len(service_alerts) < 3: # Less than 3 alerts per service
coverage_gap = {
'service': service_name,
'alert_count': len(service_alerts),
'covered_signals': list(covered_signals),
'missing_signals': list(missing_signals),
'criticality': service.get('criticality', 'medium'),
'recommendations': []
}
if len(service_alerts) == 0:
coverage_gap['recommendations'].append("Add basic availability monitoring")
if 'errors' in missing_signals:
coverage_gap['recommendations'].append("Add error rate monitoring")
if 'latency' in missing_signals:
coverage_gap['recommendations'].append("Add latency monitoring")
service_coverage.append(coverage_gap)
return service_coverage
def _generate_coverage_recommendations(self, coverage_analysis: Dict[str, Any]) -> List[str]:
"""Generate recommendations to improve monitoring coverage."""
recommendations = []
for missing_category in coverage_analysis['missing_categories']:
if missing_category == 'availability':
recommendations.append("Add service availability/uptime monitoring")
elif missing_category == 'latency':
recommendations.append("Add response time and latency monitoring")
elif missing_category == 'error_rate':
recommendations.append("Add error rate and HTTP status code monitoring")
elif missing_category == 'resource_utilization':
recommendations.append("Add CPU, memory, and disk utilization monitoring")
elif missing_category == 'security':
recommendations.append("Add security monitoring (auth failures, suspicious activity)")
elif missing_category == 'business_metrics':
recommendations.append("Add business KPI monitoring")
for missing_signal in coverage_analysis['missing_golden_signals']:
recommendations.append(f"Implement {missing_signal} monitoring (Golden Signal)")
if coverage_analysis['critical_gaps']:
recommendations.append("Address critical monitoring gaps as highest priority")
return recommendations
def find_duplicate_alerts(self, alerts: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Identify duplicate or redundant alerts."""
duplicates = []
alert_signatures = defaultdict(list)
# Group alerts by signature
for i, alert in enumerate(alerts):
signature = self._generate_alert_signature(alert)
alert_signatures[signature].append((i, alert))
# Find exact duplicates
for signature, alert_group in alert_signatures.items():
if len(alert_group) > 1:
duplicate_group = {
'type': 'exact_duplicate',
'signature': signature,
'alerts': [{'index': i, 'name': alert.get('alert', alert.get('name', f'Alert_{i}'))}
for i, alert in alert_group],
'recommendation': 'Remove duplicate alerts, keep the most comprehensive one'
}
duplicates.append(duplicate_group)
# Find semantic duplicates (similar but not identical)
semantic_duplicates = self._find_semantic_duplicates(alerts)
duplicates.extend(semantic_duplicates)
return duplicates
def _generate_alert_signature(self, alert: Dict[str, Any]) -> str:
"""Generate a signature for alert comparison."""
expr = alert.get('expr', alert.get('condition', ''))
labels = alert.get('labels', {})
# Normalize the expression by removing whitespace and standardizing
normalized_expr = re.sub(r'\s+', ' ', expr).strip()
# Create signature from expression and key labels
key_labels = {k: v for k, v in labels.items()
if k in ['service', 'severity', 'team']}
return f"{normalized_expr}::{json.dumps(key_labels, sort_keys=True)}"
def _find_semantic_duplicates(self, alerts: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Find semantically similar alerts."""
semantic_duplicates = []
# Group alerts by service and metric type
service_groups = defaultdict(list)
for i, alert in enumerate(alerts):
service = self._extract_service_from_alert(alert)
metric_type = self._extract_metric_type_from_alert(alert)
key = f"{service}::{metric_type}"
service_groups[key].append((i, alert))
# Look for similar alerts within each group
for key, alert_group in service_groups.items():
if len(alert_group) > 1:
similar_alerts = self._identify_similar_alerts(alert_group)
if similar_alerts:
semantic_duplicates.extend(similar_alerts)
return semantic_duplicates
def _extract_service_from_alert(self, alert: Dict[str, Any]) -> str:
"""Extract service name from alert."""
labels = alert.get('labels', {})
if 'service' in labels:
return labels['service']
expr = alert.get('expr', alert.get('condition', ''))
# Try to extract service from metric labels
service_match = re.search(r'service="([^"]+)"', expr)
if service_match:
return service_match.group(1)
return 'unknown'
def _extract_metric_type_from_alert(self, alert: Dict[str, Any]) -> str:
"""Extract metric type from alert."""
expr = alert.get('expr', alert.get('condition', ''))
# Common metric patterns
if 'up' in expr.lower():
return 'availability'
elif any(keyword in expr.lower() for keyword in ['latency', 'duration', 'response_time']):
return 'latency'
elif any(keyword in expr.lower() for keyword in ['error', 'fail', '5xx']):
return 'error_rate'
elif any(keyword in expr.lower() for keyword in ['cpu', 'memory', 'disk']):
return 'resource'
return 'other'
def _identify_similar_alerts(self, alert_group: List[Tuple[int, Dict[str, Any]]]) -> List[Dict[str, Any]]:
"""Identify similar alerts within a group."""
similar_groups = []
# Simple similarity check based on threshold values and conditions
threshold_groups = defaultdict(list)
for index, alert in alert_group:
expr = alert.get('expr', alert.get('condition', ''))
threshold = self._extract_threshold_from_expression(expr)
severity = alert.get('labels', {}).get('severity', 'unknown')
similarity_key = f"{threshold}::{severity}"
threshold_groups[similarity_key].append((index, alert))
# If multiple alerts have very similar thresholds, they might be redundant
for similarity_key, similar_alerts in threshold_groups.items():
if len(similar_alerts) > 1:
similar_group = {
'type': 'semantic_duplicate',
'similarity_key': similarity_key,
'alerts': [{'index': i, 'name': alert.get('alert', alert.get('name', f'Alert_{i}'))}
for i, alert in similar_alerts],
'recommendation': 'Review for potential consolidation - similar thresholds and conditions'
}
similar_groups.append(similar_group)
return similar_groups
def _extract_threshold_from_expression(self, expr: str) -> str:
"""Extract threshold value from alert expression."""
# Look for common threshold patterns
threshold_patterns = [
r'>[\s]*([0-9.]+)',
r'<[\s]*([0-9.]+)',
r'>=[\s]*([0-9.]+)',
r'<=[\s]*([0-9.]+)',
r'==[\s]*([0-9.]+)'
]
for pattern in threshold_patterns:
match = re.search(pattern, expr)
if match:
return match.group(1)
return 'unknown'
def analyze_thresholds(self, alerts: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Analyze alert thresholds for optimization opportunities."""
threshold_analysis = []
for alert in alerts:
alert_name = alert.get('alert', alert.get('name', 'Unknown'))
expr = alert.get('expr', alert.get('condition', ''))
analysis = {
'alert_name': alert_name,
'current_expression': expr,
'threshold_issues': [],
'recommendations': []
}
# Check for hard-coded thresholds
if re.search(r'[><=]\s*[0-9.]+', expr):
analysis['threshold_issues'].append('Hard-coded threshold value')
analysis['recommendations'].append('Consider parameterizing thresholds')
# Check for percentage-based thresholds that might be too strict
percentage_match = re.search(r'([><=])\s*0?\.\d+', expr)
if percentage_match:
operator = percentage_match.group(1)
if operator in ['>', '>='] and 'error' in expr.lower():
analysis['threshold_issues'].append('Very low error rate threshold')
analysis['recommendations'].append('Consider increasing error rate threshold based on SLO')
# Check for missing hysteresis
if '>' in expr and 'for:' not in str(alert):
analysis['threshold_issues'].append('No hysteresis (for clause)')
analysis['recommendations'].append('Add "for" clause to prevent alert flapping')
# Check for resource utilization thresholds
if any(resource in expr.lower() for resource in ['cpu', 'memory', 'disk']):
threshold_value = self._extract_threshold_from_expression(expr)
if threshold_value and threshold_value.replace('.', '').isdigit():
threshold_num = float(threshold_value)
if threshold_num < 0.7: # Less than 70%
analysis['threshold_issues'].append('Low resource utilization threshold')
analysis['recommendations'].append('Consider increasing threshold to reduce noise')
# Add historical data analysis if available
historical_data = alert.get('historical_data', {})
if historical_data:
false_positive_rate = historical_data.get('false_positive_rate', 0)
if false_positive_rate > 0.2:
analysis['threshold_issues'].append(f'High false positive rate: {false_positive_rate*100:.1f}%')
analysis['recommendations'].append('Analyze historical data and adjust threshold')
if analysis['threshold_issues']:
threshold_analysis.append(analysis)
return threshold_analysis
def assess_alert_fatigue_risk(self, alerts: List[Dict[str, Any]]) -> Dict[str, Any]:
"""Assess risk of alert fatigue."""
fatigue_assessment = {
'total_alerts': len(alerts),
'risk_level': 'low',
'risk_factors': [],
'metrics': {},
'recommendations': []
}
# Count alerts by severity
severity_counts = Counter()
for alert in alerts:
severity = alert.get('labels', {}).get('severity', 'unknown')
severity_counts[severity] += 1
fatigue_assessment['metrics']['severity_distribution'] = dict(severity_counts)
# Calculate risk factors
critical_count = severity_counts.get('critical', 0)
warning_count = severity_counts.get('warning', 0) + severity_counts.get('high', 0)
total_high_priority = critical_count + warning_count
# Too many high-priority alerts
if total_high_priority > 50:
fatigue_assessment['risk_factors'].append('High number of critical/warning alerts')
fatigue_assessment['recommendations'].append('Review and reduce number of high-priority alerts')
# Poor critical to warning ratio
if critical_count > 0 and warning_count > 0:
critical_ratio = critical_count / (critical_count + warning_count)
if critical_ratio > 0.3: # More than 30% critical
fatigue_assessment['risk_factors'].append('High ratio of critical alerts')
fatigue_assessment['recommendations'].append('Review critical alert criteria - not everything should be critical')
# Estimate daily alert volume
daily_estimate = self._estimate_daily_alert_volume(alerts)
fatigue_assessment['metrics']['estimated_daily_alerts'] = daily_estimate
if daily_estimate > 100:
fatigue_assessment['risk_factors'].append('High estimated daily alert volume')
fatigue_assessment['recommendations'].append('Implement alert grouping and suppression rules')
# Check for missing runbooks
alerts_without_runbooks = [alert for alert in alerts
if not alert.get('annotations', {}).get('runbook_url')]
runbook_ratio = len(alerts_without_runbooks) / len(alerts) if alerts else 0
if runbook_ratio > 0.5:
fatigue_assessment['risk_factors'].append('Many alerts lack runbooks')
fatigue_assessment['recommendations'].append('Create runbooks for alerts to improve response efficiency')
# Determine overall risk level
risk_score = len(fatigue_assessment['risk_factors'])
if risk_score >= 3:
fatigue_assessment['risk_level'] = 'high'
elif risk_score >= 1:
fatigue_assessment['risk_level'] = 'medium'
return fatigue_assessment
def _estimate_daily_alert_volume(self, alerts: List[Dict[str, Any]]) -> int:
"""Estimate daily alert volume."""
total_estimated = 0
for alert in alerts:
# Use historical data if available
historical_data = alert.get('historical_data', {})
if historical_data and 'fires_per_day' in historical_data:
total_estimated += historical_data['fires_per_day']
continue
# Otherwise estimate based on alert characteristics
expr = alert.get('expr', alert.get('condition', ''))
severity = alert.get('labels', {}).get('severity', 'warning')
# Base estimate by severity
base_estimates = {
'critical': 0.1, # Critical should rarely fire
'high': 0.5,
'warning': 2,
'info': 5
}
estimate = base_estimates.get(severity, 1)
# Adjust based on alert type
if 'error_rate' in expr.lower():
estimate *= 1.5 # Error rate alerts tend to be more frequent
elif 'availability' in expr.lower() or 'up' in expr.lower():
estimate *= 0.5 # Availability alerts should be rare
total_estimated += estimate
return int(total_estimated)
def generate_optimized_config(self, alerts: List[Dict[str, Any]],
analysis_results: Dict[str, Any]) -> Dict[str, Any]:
"""Generate optimized alert configuration."""
optimized_alerts = []
for i, alert in enumerate(alerts):
optimized_alert = alert.copy()
alert_name = alert.get('alert', alert.get('name', f'Alert_{i}'))
# Apply noise reduction optimizations
noisy_alerts = analysis_results.get('noisy_alerts', [])
for noisy_alert in noisy_alerts:
if noisy_alert['alert_name'] == alert_name:
optimized_alert = self._apply_noise_reduction(optimized_alert, noisy_alert)
break
# Apply threshold optimizations
threshold_issues = analysis_results.get('threshold_analysis', [])
for threshold_issue in threshold_issues:
if threshold_issue['alert_name'] == alert_name:
optimized_alert = self._apply_threshold_optimization(optimized_alert, threshold_issue)
break
# Ensure proper alert metadata
optimized_alert = self._ensure_alert_metadata(optimized_alert)
optimized_alerts.append(optimized_alert)
# Remove duplicates based on analysis
if 'duplicate_alerts' in analysis_results:
optimized_alerts = self._remove_duplicate_alerts(optimized_alerts,
analysis_results['duplicate_alerts'])
# Add missing alerts for coverage gaps
if 'coverage_gaps' in analysis_results:
new_alerts = self._generate_missing_alerts(analysis_results['coverage_gaps'])
optimized_alerts.extend(new_alerts)
optimized_config = {
'alerts': optimized_alerts,
'optimization_metadata': {
'optimized_at': datetime.utcnow().isoformat() + 'Z',
'original_count': len(alerts),
'optimized_count': len(optimized_alerts),
'changes_applied': analysis_results.get('optimizations_applied', [])
}
}
return optimized_config
def _apply_noise_reduction(self, alert: Dict[str, Any],
noise_analysis: Dict[str, Any]) -> Dict[str, Any]:
"""Apply noise reduction optimizations to an alert."""
optimized_alert = alert.copy()
for recommendation in noise_analysis['recommendations']:
if 'for:' in recommendation and not alert.get('for'):
optimized_alert['for'] = '5m'
elif 'threshold' in recommendation.lower():
# This would require more sophisticated threshold adjustment
# For now, add annotation for manual review
if 'annotations' not in optimized_alert:
optimized_alert['annotations'] = {}
optimized_alert['annotations']['optimization_note'] = 'Review threshold - potentially too sensitive'
return optimized_alert
def _apply_threshold_optimization(self, alert: Dict[str, Any],
threshold_analysis: Dict[str, Any]) -> Dict[str, Any]:
"""Apply threshold optimizations to an alert."""
optimized_alert = alert.copy()
# Add 'for' clause if missing
if 'No hysteresis' in str(threshold_analysis['threshold_issues']):
if not alert.get('for'):
optimized_alert['for'] = '5m'
# Add optimization annotations
if threshold_analysis['recommendations']:
if 'annotations' not in optimized_alert:
optimized_alert['annotations'] = {}
optimized_alert['annotations']['threshold_recommendations'] = '; '.join(threshold_analysis['recommendations'])
return optimized_alert
def _ensure_alert_metadata(self, alert: Dict[str, Any]) -> Dict[str, Any]:
"""Ensure alert has proper metadata."""
optimized_alert = alert.copy()
# Ensure annotations exist
if 'annotations' not in optimized_alert:
optimized_alert['annotations'] = {}
# Add summary if missing
if 'summary' not in optimized_alert['annotations']:
alert_name = alert.get('alert', alert.get('name', 'Alert'))
optimized_alert['annotations']['summary'] = f"Alert: {alert_name}"
# Add description if missing
if 'description' not in optimized_alert['annotations']:
optimized_alert['annotations']['description'] = 'This alert requires a description. Please update with specific details about the condition and impact.'
# Ensure proper labels
if 'labels' not in optimized_alert:
optimized_alert['labels'] = {}
if 'severity' not in optimized_alert['labels']:
optimized_alert['labels']['severity'] = 'warning'
return optimized_alert
def _remove_duplicate_alerts(self, alerts: List[Dict[str, Any]],
duplicates: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Remove duplicate alerts from the list."""
indices_to_remove = set()
for duplicate_group in duplicates:
if duplicate_group['type'] == 'exact_duplicate':
# Keep the first alert, remove the rest
alert_indices = [alert_info['index'] for alert_info in duplicate_group['alerts']]
indices_to_remove.update(alert_indices[1:]) # Remove all but first
return [alert for i, alert in enumerate(alerts) if i not in indices_to_remove]
def _generate_missing_alerts(self, coverage_gaps: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Generate alerts for missing coverage."""
new_alerts = []
for missing_signal in coverage_gaps.get('missing_golden_signals', []):
if missing_signal == 'latency':
new_alert = {
'alert': 'HighLatency',
'expr': 'histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m])) > 0.5',
'for': '5m',
'labels': {
'severity': 'warning'
},
'annotations': {
'summary': 'High request latency detected',
'description': 'The 95th percentile latency is above 500ms for 5 minutes.',
'generated': 'true'
}
}
new_alerts.append(new_alert)
elif missing_signal == 'errors':
new_alert = {
'alert': 'HighErrorRate',
'expr': 'sum(rate(http_requests_total{code=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) > 0.01',
'for': '5m',
'labels': {
'severity': 'warning'
},
'annotations': {
'summary': 'High error rate detected',
'description': 'Error rate is above 1% for 5 minutes.',
'generated': 'true'
}
}
new_alerts.append(new_alert)
return new_alerts
def analyze_configuration(self, alert_config: Dict[str, Any]) -> Dict[str, Any]:
"""Perform comprehensive analysis of alert configuration."""
alerts = alert_config.get('alerts', alert_config.get('rules', []))
services = alert_config.get('services', [])
analysis_results = {
'summary': {
'total_alerts': len(alerts),
'analysis_timestamp': datetime.utcnow().isoformat() + 'Z'
},
'noisy_alerts': self.analyze_alert_noise(alerts),
'coverage_gaps': self.identify_coverage_gaps(alerts, services),
'duplicate_alerts': self.find_duplicate_alerts(alerts),
'threshold_analysis': self.analyze_thresholds(alerts),
'alert_fatigue_assessment': self.assess_alert_fatigue_risk(alerts)
}
# Generate overall recommendations
analysis_results['overall_recommendations'] = self._generate_overall_recommendations(analysis_results)
return analysis_results
def _generate_overall_recommendations(self, analysis_results: Dict[str, Any]) -> List[str]:
"""Generate overall recommendations based on complete analysis."""
recommendations = []
# High-priority recommendations
if analysis_results['alert_fatigue_assessment']['risk_level'] == 'high':
recommendations.append("HIGH PRIORITY: Address alert fatigue risk by reducing alert volume")
if len(analysis_results['coverage_gaps']['critical_gaps']) > 0:
recommendations.append("HIGH PRIORITY: Address critical monitoring gaps")
# Medium-priority recommendations
if len(analysis_results['noisy_alerts']) > 0:
recommendations.append(f"Optimize {len(analysis_results['noisy_alerts'])} noisy alerts to reduce false positives")
if len(analysis_results['duplicate_alerts']) > 0:
recommendations.append(f"Remove or consolidate {len(analysis_results['duplicate_alerts'])} duplicate alert groups")
# General recommendations
recommendations.append("Implement proper alert routing and escalation policies")
recommendations.append("Create runbooks for all production alerts")
recommendations.append("Set up alert effectiveness monitoring and regular reviews")
return recommendations
def export_analysis(self, analysis_results: Dict[str, Any], output_file: str,
format_type: str = 'json'):
"""Export analysis results."""
if format_type.lower() == 'json':
with open(output_file, 'w') as f:
json.dump(analysis_results, f, indent=2)
elif format_type.lower() == 'html':
self._export_html_report(analysis_results, output_file)
else:
raise ValueError(f"Unsupported format: {format_type}")
def _export_html_report(self, analysis_results: Dict[str, Any], output_file: str):
"""Export analysis as HTML report."""
html_content = self._generate_html_report(analysis_results)
with open(output_file, 'w') as f:
f.write(html_content)
def _generate_html_report(self, analysis_results: Dict[str, Any]) -> str:
"""Generate HTML report of analysis results."""
html = f"""
<!DOCTYPE html>
<html>
<head>
<title>Alert Configuration Analysis Report</title>
<style>
body {{ font-family: Arial, sans-serif; margin: 20px; }}
.header {{ background: #f4f4f4; padding: 20px; border-radius: 5px; }}
.section {{ margin: 20px 0; padding: 15px; border: 1px solid #ddd; border-radius: 5px; }}
.critical {{ border-left: 5px solid #ff0000; }}
.warning {{ border-left: 5px solid #ff9900; }}
.info {{ border-left: 5px solid #0066cc; }}
.success {{ border-left: 5px solid #00aa00; }}
ul {{ margin: 10px 0; }}
li {{ margin: 5px 0; }}
</style>
</head>
<body>
<div class="header">
<h1>Alert Configuration Analysis Report</h1>
<p>Generated: {analysis_results['summary']['analysis_timestamp']}</p>
<p>Total Alerts Analyzed: {analysis_results['summary']['total_alerts']}</p>
</div>
<div class="section critical">
<h2>Overall Recommendations</h2>
<ul>
{''.join(f'<li>{rec}</li>' for rec in analysis_results['overall_recommendations'])}
</ul>
</div>
<div class="section warning">
<h2>Alert Fatigue Assessment</h2>
<p><strong>Risk Level:</strong> {analysis_results['alert_fatigue_assessment']['risk_level'].upper()}</p>
<p><strong>Risk Factors:</strong></p>
<ul>
{''.join(f'<li>{factor}</li>' for factor in analysis_results['alert_fatigue_assessment']['risk_factors'])}
</ul>
</div>
<div class="section info">
<h2>Noisy Alerts ({len(analysis_results['noisy_alerts'])})</h2>
{''.join(f'<div><strong>{alert["alert_name"]}</strong> (Score: {alert["noise_score"]})<ul>{"".join(f"<li>{reason}</li>" for reason in alert["reasons"])}</ul></div>'
for alert in analysis_results['noisy_alerts'][:5])}
</div>
<div class="section info">
<h2>Coverage Gaps</h2>
<p><strong>Missing Categories:</strong> {', '.join(analysis_results['coverage_gaps']['missing_categories']) or 'None'}</p>
<p><strong>Missing Golden Signals:</strong> {', '.join(analysis_results['coverage_gaps']['missing_golden_signals']) or 'None'}</p>
<p><strong>Critical Gaps:</strong> {len(analysis_results['coverage_gaps']['critical_gaps'])}</p>
</div>
</body>
</html>
"""
return html
def print_summary(self, analysis_results: Dict[str, Any]):
"""Print human-readable summary of analysis."""
print(f"\n{'='*60}")
print(f"ALERT CONFIGURATION ANALYSIS SUMMARY")
print(f"{'='*60}")
summary = analysis_results['summary']
print(f"\nOverall Statistics:")
print(f" Total Alerts: {summary['total_alerts']}")
print(f" Analysis Date: {summary['analysis_timestamp']}")
# Alert fatigue assessment
fatigue = analysis_results['alert_fatigue_assessment']
print(f"\nAlert Fatigue Risk: {fatigue['risk_level'].upper()}")
if fatigue['risk_factors']:
print(f" Risk Factors:")
for factor in fatigue['risk_factors']:
print(f" • {factor}")
# Noisy alerts
noisy = analysis_results['noisy_alerts']
print(f"\nNoisy Alerts: {len(noisy)}")
if noisy:
print(f" Top 3 Noisiest:")
for alert in noisy[:3]:
print(f" • {alert['alert_name']} (Score: {alert['noise_score']})")
# Coverage gaps
gaps = analysis_results['coverage_gaps']
print(f"\nMonitoring Coverage:")
print(f" Missing Categories: {len(gaps['missing_categories'])}")
print(f" Missing Golden Signals: {len(gaps['missing_golden_signals'])}")
print(f" Critical Gaps: {len(gaps['critical_gaps'])}")
# Duplicates
duplicates = analysis_results['duplicate_alerts']
print(f"\nDuplicate Alerts: {len(duplicates)} groups")
# Overall recommendations
recommendations = analysis_results['overall_recommendations']
print(f"\nTop Recommendations:")
for i, rec in enumerate(recommendations[:5], 1):
print(f" {i}. {rec}")
print(f"\n{'='*60}\n")
def main():
"""Main function for CLI usage."""
parser = argparse.ArgumentParser(
description='Analyze and optimize alert configurations',
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
# Analyze alert configuration
python alert_optimizer.py --input alerts.json --analyze-only
# Generate optimized configuration
python alert_optimizer.py --input alerts.json --output optimized_alerts.json
# Generate HTML report
python alert_optimizer.py --input alerts.json --report report.html --format html
"""
)
parser.add_argument('--input', '-i', required=True,
help='Input alert configuration JSON file')
parser.add_argument('--output', '-o',
help='Output optimized configuration JSON file')
parser.add_argument('--report', '-r',
help='Generate analysis report file')
parser.add_argument('--format', choices=['json', 'html'], default='json',
help='Report format (json or html)')
parser.add_argument('--analyze-only', action='store_true',
help='Only perform analysis, do not generate optimized config')
args = parser.parse_args()
optimizer = AlertOptimizer()
try:
# Load alert configuration
alert_config = optimizer.load_alert_config(args.input)
# Perform analysis
analysis_results = optimizer.analyze_configuration(alert_config)
# Generate optimized configuration if requested
if not args.analyze_only:
optimized_config = optimizer.generate_optimized_config(
alert_config.get('alerts', alert_config.get('rules', [])),
analysis_results
)
output_file = args.output or 'optimized_alerts.json'
optimizer.export_analysis(optimized_config, output_file, 'json')
print(f"Optimized configuration saved to: {output_file}")
# Generate report if requested
if args.report:
optimizer.export_analysis(analysis_results, args.report, args.format)
print(f"Analysis report saved to: {args.report}")
# Always show summary
optimizer.print_summary(analysis_results)
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
sys.exit(1)
if __name__ == '__main__':
main()
FILE:scripts/dashboard_generator.py
#!/usr/bin/env python3
"""
Dashboard Generator - Generate comprehensive dashboard specifications
This script generates dashboard specifications based on service/system descriptions:
- Panel layout optimized for different screen sizes and roles
- Metric queries (Prometheus-style) for comprehensive monitoring
- Visualization types appropriate for different metric types
- Drill-down paths for effective troubleshooting workflows
- Golden signals coverage (latency, traffic, errors, saturation)
- RED/USE method implementation
- Business metrics integration
Usage:
python dashboard_generator.py --input service_definition.json --output dashboard_spec.json
python dashboard_generator.py --service-type api --name "Payment Service" --output payment_dashboard.json
"""
import json
import argparse
import sys
import math
from typing import Dict, List, Any, Tuple
from datetime import datetime, timedelta
class DashboardGenerator:
"""Generate comprehensive dashboard specifications."""
# Dashboard layout templates by role
ROLE_LAYOUTS = {
'sre': {
'primary_focus': ['availability', 'latency', 'errors', 'resource_utilization'],
'secondary_focus': ['throughput', 'capacity', 'dependencies'],
'time_ranges': ['1h', '6h', '1d', '7d'],
'default_refresh': '30s'
},
'developer': {
'primary_focus': ['latency', 'errors', 'throughput', 'business_metrics'],
'secondary_focus': ['resource_utilization', 'dependencies'],
'time_ranges': ['15m', '1h', '6h', '1d'],
'default_refresh': '1m'
},
'executive': {
'primary_focus': ['availability', 'business_metrics', 'user_experience'],
'secondary_focus': ['cost', 'capacity_trends'],
'time_ranges': ['1d', '7d', '30d'],
'default_refresh': '5m'
},
'ops': {
'primary_focus': ['resource_utilization', 'capacity', 'alerts', 'deployments'],
'secondary_focus': ['throughput', 'latency'],
'time_ranges': ['5m', '30m', '2h', '1d'],
'default_refresh': '15s'
}
}
# Service type specific metric configurations
SERVICE_METRICS = {
'api': {
'golden_signals': ['latency', 'traffic', 'errors', 'saturation'],
'key_metrics': [
'http_requests_total',
'http_request_duration_seconds',
'http_request_size_bytes',
'http_response_size_bytes'
],
'resource_metrics': ['cpu_usage', 'memory_usage', 'goroutines']
},
'web': {
'golden_signals': ['latency', 'traffic', 'errors', 'saturation'],
'key_metrics': [
'http_requests_total',
'http_request_duration_seconds',
'page_load_time',
'user_sessions'
],
'resource_metrics': ['cpu_usage', 'memory_usage', 'connections']
},
'database': {
'golden_signals': ['latency', 'traffic', 'errors', 'saturation'],
'key_metrics': [
'db_connections_active',
'db_query_duration_seconds',
'db_queries_total',
'db_slow_queries_total'
],
'resource_metrics': ['cpu_usage', 'memory_usage', 'disk_io', 'connections']
},
'queue': {
'golden_signals': ['latency', 'traffic', 'errors', 'saturation'],
'key_metrics': [
'queue_depth',
'message_processing_duration',
'messages_published_total',
'messages_consumed_total'
],
'resource_metrics': ['cpu_usage', 'memory_usage', 'disk_usage']
}
}
# Visualization type recommendations
VISUALIZATION_TYPES = {
'latency': 'line_chart',
'throughput': 'line_chart',
'error_rate': 'line_chart',
'success_rate': 'stat',
'resource_utilization': 'gauge',
'queue_depth': 'bar_chart',
'status': 'stat',
'distribution': 'heatmap',
'alerts': 'table',
'logs': 'logs_panel'
}
def __init__(self):
"""Initialize the Dashboard Generator."""
self.service_config = {}
self.dashboard_spec = {}
def load_service_definition(self, file_path: str) -> Dict[str, Any]:
"""Load service definition from JSON file."""
try:
with open(file_path, 'r') as f:
return json.load(f)
except FileNotFoundError:
raise ValueError(f"Service definition file not found: {file_path}")
except json.JSONDecodeError as e:
raise ValueError(f"Invalid JSON in service definition: {e}")
def create_service_definition(self, service_type: str, name: str,
criticality: str = 'medium') -> Dict[str, Any]:
"""Create a service definition from parameters."""
return {
'name': name,
'type': service_type,
'criticality': criticality,
'description': f'{name} - A {criticality} criticality {service_type} service',
'team': 'platform',
'environment': 'production',
'dependencies': [],
'tags': []
}
def generate_dashboard_specification(self, service_def: Dict[str, Any],
target_role: str = 'sre') -> Dict[str, Any]:
"""Generate comprehensive dashboard specification."""
service_name = service_def.get('name', 'Service')
service_type = service_def.get('type', 'api')
# Get role-specific configuration
role_config = self.ROLE_LAYOUTS.get(target_role, self.ROLE_LAYOUTS['sre'])
dashboard_spec = {
'metadata': {
'title': f"{service_name} - {target_role.upper()} Dashboard",
'service': service_def,
'target_role': target_role,
'generated_at': datetime.utcnow().isoformat() + 'Z',
'version': '1.0'
},
'configuration': {
'time_ranges': role_config['time_ranges'],
'default_time_range': role_config['time_ranges'][1], # Second option as default
'refresh_interval': role_config['default_refresh'],
'timezone': 'UTC',
'theme': 'dark'
},
'layout': self._generate_dashboard_layout(service_def, role_config),
'panels': self._generate_panels(service_def, role_config),
'variables': self._generate_template_variables(service_def),
'alerts_integration': self._generate_alerts_integration(service_def),
'drill_down_paths': self._generate_drill_down_paths(service_def)
}
return dashboard_spec
def _generate_dashboard_layout(self, service_def: Dict[str, Any],
role_config: Dict[str, Any]) -> Dict[str, Any]:
"""Generate dashboard layout configuration."""
return {
'grid_settings': {
'width': 24, # Grafana-style 24-column grid
'height_unit': 'px',
'cell_height': 30
},
'sections': [
{
'title': 'Service Overview',
'collapsed': False,
'y_position': 0,
'panels': ['service_status', 'slo_summary', 'error_budget']
},
{
'title': 'Golden Signals',
'collapsed': False,
'y_position': 8,
'panels': ['latency', 'traffic', 'errors', 'saturation']
},
{
'title': 'Resource Utilization',
'collapsed': False,
'y_position': 16,
'panels': ['cpu_usage', 'memory_usage', 'network_io', 'disk_io']
},
{
'title': 'Dependencies & Downstream',
'collapsed': True,
'y_position': 24,
'panels': ['dependency_status', 'downstream_latency', 'circuit_breakers']
}
]
}
def _generate_panels(self, service_def: Dict[str, Any],
role_config: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Generate dashboard panels based on service and role."""
service_name = service_def.get('name', 'service')
service_type = service_def.get('type', 'api')
panels = []
# Service Overview Panels
panels.extend(self._create_overview_panels(service_def))
# Golden Signals Panels
panels.extend(self._create_golden_signals_panels(service_def))
# Resource Utilization Panels
panels.extend(self._create_resource_panels(service_def))
# Service-specific panels
if service_type == 'api':
panels.extend(self._create_api_specific_panels(service_def))
elif service_type == 'database':
panels.extend(self._create_database_specific_panels(service_def))
elif service_type == 'queue':
panels.extend(self._create_queue_specific_panels(service_def))
# Role-specific additional panels
if 'business_metrics' in role_config['primary_focus']:
panels.extend(self._create_business_metrics_panels(service_def))
if 'capacity' in role_config['primary_focus']:
panels.extend(self._create_capacity_panels(service_def))
return panels
def _create_overview_panels(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Create service overview panels."""
service_name = service_def.get('name', 'service')
return [
{
'id': 'service_status',
'title': 'Service Status',
'type': 'stat',
'grid_pos': {'x': 0, 'y': 0, 'w': 6, 'h': 4},
'targets': [
{
'expr': f'up{{service="{service_name}"}}',
'legendFormat': 'Status'
}
],
'field_config': {
'overrides': [
{
'matcher': {'id': 'byName', 'options': 'Status'},
'properties': [
{'id': 'color', 'value': {'mode': 'thresholds'}},
{'id': 'thresholds', 'value': {
'steps': [
{'color': 'red', 'value': 0},
{'color': 'green', 'value': 1}
]
}},
{'id': 'mappings', 'value': [
{'options': {'0': {'text': 'DOWN'}}, 'type': 'value'},
{'options': {'1': {'text': 'UP'}}, 'type': 'value'}
]}
]
}
]
},
'options': {
'orientation': 'horizontal',
'textMode': 'value_and_name'
}
},
{
'id': 'slo_summary',
'title': 'SLO Achievement (30d)',
'type': 'stat',
'grid_pos': {'x': 6, 'y': 0, 'w': 9, 'h': 4},
'targets': [
{
'expr': f'(1 - (increase(http_requests_total{{service="{service_name}",code=~"5.."}}[30d]) / increase(http_requests_total{{service="{service_name}"}}[30d]))) * 100',
'legendFormat': 'Availability'
},
{
'expr': f'histogram_quantile(0.95, increase(http_request_duration_seconds_bucket{{service="{service_name}"}}[30d])) * 1000',
'legendFormat': 'P95 Latency (ms)'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'thresholds'},
'thresholds': {
'steps': [
{'color': 'red', 'value': 0},
{'color': 'yellow', 'value': 99.0},
{'color': 'green', 'value': 99.9}
]
}
}
},
'options': {
'orientation': 'horizontal',
'textMode': 'value_and_name'
}
},
{
'id': 'error_budget',
'title': 'Error Budget Remaining',
'type': 'gauge',
'grid_pos': {'x': 15, 'y': 0, 'w': 9, 'h': 4},
'targets': [
{
'expr': f'(1 - (increase(http_requests_total{{service="{service_name}",code=~"5.."}}[30d]) / increase(http_requests_total{{service="{service_name}"}}[30d])) - 0.999) / 0.001 * 100',
'legendFormat': 'Error Budget %'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'thresholds'},
'min': 0,
'max': 100,
'thresholds': {
'steps': [
{'color': 'red', 'value': 0},
{'color': 'yellow', 'value': 25},
{'color': 'green', 'value': 50}
]
},
'unit': 'percent'
}
},
'options': {
'showThresholdLabels': True,
'showThresholdMarkers': True
}
}
]
def _create_golden_signals_panels(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Create golden signals monitoring panels."""
service_name = service_def.get('name', 'service')
return [
{
'id': 'latency',
'title': 'Request Latency',
'type': 'timeseries',
'grid_pos': {'x': 0, 'y': 8, 'w': 12, 'h': 6},
'targets': [
{
'expr': f'histogram_quantile(0.50, rate(http_request_duration_seconds_bucket{{service="{service_name}"}}[5m])) * 1000',
'legendFormat': 'P50 Latency'
},
{
'expr': f'histogram_quantile(0.95, rate(http_request_duration_seconds_bucket{{service="{service_name}"}}[5m])) * 1000',
'legendFormat': 'P95 Latency'
},
{
'expr': f'histogram_quantile(0.99, rate(http_request_duration_seconds_bucket{{service="{service_name}"}}[5m])) * 1000',
'legendFormat': 'P99 Latency'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'palette-classic'},
'unit': 'ms',
'custom': {
'drawStyle': 'line',
'lineInterpolation': 'linear',
'lineWidth': 1,
'fillOpacity': 10
}
}
},
'options': {
'tooltip': {'mode': 'multi', 'sort': 'desc'},
'legend': {'displayMode': 'table', 'placement': 'bottom'}
}
},
{
'id': 'traffic',
'title': 'Request Rate',
'type': 'timeseries',
'grid_pos': {'x': 12, 'y': 8, 'w': 12, 'h': 6},
'targets': [
{
'expr': f'sum(rate(http_requests_total{{service="{service_name}"}}[5m]))',
'legendFormat': 'Total RPS'
},
{
'expr': f'sum(rate(http_requests_total{{service="{service_name}",code=~"2.."}}[5m]))',
'legendFormat': '2xx RPS'
},
{
'expr': f'sum(rate(http_requests_total{{service="{service_name}",code=~"4.."}}[5m]))',
'legendFormat': '4xx RPS'
},
{
'expr': f'sum(rate(http_requests_total{{service="{service_name}",code=~"5.."}}[5m]))',
'legendFormat': '5xx RPS'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'palette-classic'},
'unit': 'reqps',
'custom': {
'drawStyle': 'line',
'lineInterpolation': 'linear',
'lineWidth': 1,
'fillOpacity': 0
}
}
},
'options': {
'tooltip': {'mode': 'multi', 'sort': 'desc'},
'legend': {'displayMode': 'table', 'placement': 'bottom'}
}
},
{
'id': 'errors',
'title': 'Error Rate',
'type': 'timeseries',
'grid_pos': {'x': 0, 'y': 14, 'w': 12, 'h': 6},
'targets': [
{
'expr': f'sum(rate(http_requests_total{{service="{service_name}",code=~"5.."}}[5m])) / sum(rate(http_requests_total{{service="{service_name}"}}[5m])) * 100',
'legendFormat': '5xx Error Rate'
},
{
'expr': f'sum(rate(http_requests_total{{service="{service_name}",code=~"4.."}}[5m])) / sum(rate(http_requests_total{{service="{service_name}"}}[5m])) * 100',
'legendFormat': '4xx Error Rate'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'palette-classic'},
'unit': 'percent',
'custom': {
'drawStyle': 'line',
'lineInterpolation': 'linear',
'lineWidth': 2,
'fillOpacity': 20
}
},
'overrides': [
{
'matcher': {'id': 'byName', 'options': '5xx Error Rate'},
'properties': [{'id': 'color', 'value': {'fixedColor': 'red'}}]
}
]
},
'options': {
'tooltip': {'mode': 'multi', 'sort': 'desc'},
'legend': {'displayMode': 'table', 'placement': 'bottom'}
}
},
{
'id': 'saturation',
'title': 'Saturation Metrics',
'type': 'timeseries',
'grid_pos': {'x': 12, 'y': 14, 'w': 12, 'h': 6},
'targets': [
{
'expr': f'rate(process_cpu_seconds_total{{service="{service_name}"}}[5m]) * 100',
'legendFormat': 'CPU Usage %'
},
{
'expr': f'process_resident_memory_bytes{{service="{service_name}"}} / process_virtual_memory_max_bytes{{service="{service_name}"}} * 100',
'legendFormat': 'Memory Usage %'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'palette-classic'},
'unit': 'percent',
'max': 100,
'custom': {
'drawStyle': 'line',
'lineInterpolation': 'linear',
'lineWidth': 1,
'fillOpacity': 10
}
}
},
'options': {
'tooltip': {'mode': 'multi', 'sort': 'desc'},
'legend': {'displayMode': 'table', 'placement': 'bottom'}
}
}
]
def _create_resource_panels(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Create resource utilization panels."""
service_name = service_def.get('name', 'service')
return [
{
'id': 'cpu_usage',
'title': 'CPU Usage',
'type': 'gauge',
'grid_pos': {'x': 0, 'y': 20, 'w': 6, 'h': 4},
'targets': [
{
'expr': f'rate(process_cpu_seconds_total{{service="{service_name}"}}[5m]) * 100',
'legendFormat': 'CPU %'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'thresholds'},
'unit': 'percent',
'min': 0,
'max': 100,
'thresholds': {
'steps': [
{'color': 'green', 'value': 0},
{'color': 'yellow', 'value': 70},
{'color': 'red', 'value': 90}
]
}
}
},
'options': {
'showThresholdLabels': True,
'showThresholdMarkers': True
}
},
{
'id': 'memory_usage',
'title': 'Memory Usage',
'type': 'gauge',
'grid_pos': {'x': 6, 'y': 20, 'w': 6, 'h': 4},
'targets': [
{
'expr': f'process_resident_memory_bytes{{service="{service_name}"}} / 1024 / 1024',
'legendFormat': 'Memory MB'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'thresholds'},
'unit': 'decbytes',
'thresholds': {
'steps': [
{'color': 'green', 'value': 0},
{'color': 'yellow', 'value': 512000000}, # 512MB
{'color': 'red', 'value': 1024000000} # 1GB
]
}
}
}
},
{
'id': 'network_io',
'title': 'Network I/O',
'type': 'timeseries',
'grid_pos': {'x': 12, 'y': 20, 'w': 6, 'h': 4},
'targets': [
{
'expr': f'rate(process_network_receive_bytes_total{{service="{service_name}"}}[5m])',
'legendFormat': 'RX Bytes/s'
},
{
'expr': f'rate(process_network_transmit_bytes_total{{service="{service_name}"}}[5m])',
'legendFormat': 'TX Bytes/s'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'palette-classic'},
'unit': 'binBps'
}
}
},
{
'id': 'disk_io',
'title': 'Disk I/O',
'type': 'timeseries',
'grid_pos': {'x': 18, 'y': 20, 'w': 6, 'h': 4},
'targets': [
{
'expr': f'rate(process_disk_read_bytes_total{{service="{service_name}"}}[5m])',
'legendFormat': 'Read Bytes/s'
},
{
'expr': f'rate(process_disk_write_bytes_total{{service="{service_name}"}}[5m])',
'legendFormat': 'Write Bytes/s'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'palette-classic'},
'unit': 'binBps'
}
}
}
]
def _create_api_specific_panels(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Create API-specific panels."""
service_name = service_def.get('name', 'service')
return [
{
'id': 'endpoint_latency',
'title': 'Top Slowest Endpoints',
'type': 'table',
'grid_pos': {'x': 0, 'y': 24, 'w': 12, 'h': 6},
'targets': [
{
'expr': f'topk(10, histogram_quantile(0.95, sum by (handler) (rate(http_request_duration_seconds_bucket{{service="{service_name}"}}[5m])))) * 1000',
'legendFormat': '{{handler}}',
'format': 'table',
'instant': True
}
],
'transformations': [
{
'id': 'organize',
'options': {
'excludeByName': {'Time': True},
'renameByName': {'Value': 'P95 Latency (ms)'}
}
}
],
'field_config': {
'overrides': [
{
'matcher': {'id': 'byName', 'options': 'P95 Latency (ms)'},
'properties': [
{'id': 'color', 'value': {'mode': 'thresholds'}},
{'id': 'thresholds', 'value': {
'steps': [
{'color': 'green', 'value': 0},
{'color': 'yellow', 'value': 100},
{'color': 'red', 'value': 500}
]
}}
]
}
]
}
},
{
'id': 'request_size_distribution',
'title': 'Request Size Distribution',
'type': 'heatmap',
'grid_pos': {'x': 12, 'y': 24, 'w': 12, 'h': 6},
'targets': [
{
'expr': f'sum by (le) (rate(http_request_size_bytes_bucket{{service="{service_name}"}}[5m]))',
'legendFormat': '{{le}}'
}
],
'options': {
'calculate': True,
'yAxis': {'unit': 'bytes'},
'color': {'scheme': 'Spectral'}
}
}
]
def _create_database_specific_panels(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Create database-specific panels."""
service_name = service_def.get('name', 'service')
return [
{
'id': 'db_connections',
'title': 'Database Connections',
'type': 'timeseries',
'grid_pos': {'x': 0, 'y': 24, 'w': 8, 'h': 6},
'targets': [
{
'expr': f'db_connections_active{{service="{service_name}"}}',
'legendFormat': 'Active Connections'
},
{
'expr': f'db_connections_idle{{service="{service_name}"}}',
'legendFormat': 'Idle Connections'
},
{
'expr': f'db_connections_max{{service="{service_name}"}}',
'legendFormat': 'Max Connections'
}
]
},
{
'id': 'query_performance',
'title': 'Query Performance',
'type': 'timeseries',
'grid_pos': {'x': 8, 'y': 24, 'w': 8, 'h': 6},
'targets': [
{
'expr': f'rate(db_queries_total{{service="{service_name}"}}[5m])',
'legendFormat': 'Queries/sec'
},
{
'expr': f'rate(db_slow_queries_total{{service="{service_name}"}}[5m])',
'legendFormat': 'Slow Queries/sec'
}
]
},
{
'id': 'db_locks',
'title': 'Database Locks',
'type': 'stat',
'grid_pos': {'x': 16, 'y': 24, 'w': 8, 'h': 6},
'targets': [
{
'expr': f'db_locks_waiting{{service="{service_name}"}}',
'legendFormat': 'Waiting Locks'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'thresholds'},
'thresholds': {
'steps': [
{'color': 'green', 'value': 0},
{'color': 'yellow', 'value': 1},
{'color': 'red', 'value': 5}
]
}
}
}
}
]
def _create_queue_specific_panels(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Create queue-specific panels."""
service_name = service_def.get('name', 'service')
return [
{
'id': 'queue_depth',
'title': 'Queue Depth',
'type': 'timeseries',
'grid_pos': {'x': 0, 'y': 24, 'w': 12, 'h': 6},
'targets': [
{
'expr': f'queue_depth{{service="{service_name}"}}',
'legendFormat': 'Messages in Queue'
}
]
},
{
'id': 'message_throughput',
'title': 'Message Throughput',
'type': 'timeseries',
'grid_pos': {'x': 12, 'y': 24, 'w': 12, 'h': 6},
'targets': [
{
'expr': f'rate(messages_published_total{{service="{service_name}"}}[5m])',
'legendFormat': 'Published/sec'
},
{
'expr': f'rate(messages_consumed_total{{service="{service_name}"}}[5m])',
'legendFormat': 'Consumed/sec'
}
]
}
]
def _create_business_metrics_panels(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Create business metrics panels."""
service_name = service_def.get('name', 'service')
return [
{
'id': 'business_kpis',
'title': 'Business KPIs',
'type': 'stat',
'grid_pos': {'x': 0, 'y': 30, 'w': 24, 'h': 4},
'targets': [
{
'expr': f'rate(business_transactions_total{{service="{service_name}"}}[1h])',
'legendFormat': 'Transactions/hour'
},
{
'expr': f'avg(business_transaction_value{{service="{service_name}"}}) * rate(business_transactions_total{{service="{service_name}"}}[1h])',
'legendFormat': 'Revenue/hour'
},
{
'expr': f'rate(user_registrations_total{{service="{service_name}"}}[1h])',
'legendFormat': 'New Users/hour'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'palette-classic'},
'custom': {
'displayMode': 'basic'
}
}
},
'options': {
'orientation': 'horizontal',
'textMode': 'value_and_name'
}
}
]
def _create_capacity_panels(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Create capacity planning panels."""
service_name = service_def.get('name', 'service')
return [
{
'id': 'capacity_trends',
'title': 'Capacity Trends (7d)',
'type': 'timeseries',
'grid_pos': {'x': 0, 'y': 34, 'w': 24, 'h': 6},
'targets': [
{
'expr': f'predict_linear(avg_over_time(rate(http_requests_total{{service="{service_name}"}}[5m])[7d:1h]), 7*24*3600)',
'legendFormat': 'Predicted Traffic (7d)'
},
{
'expr': f'predict_linear(avg_over_time(process_resident_memory_bytes{{service="{service_name}"}}[7d:1h]), 7*24*3600)',
'legendFormat': 'Predicted Memory Usage (7d)'
}
],
'field_config': {
'defaults': {
'color': {'mode': 'palette-classic'},
'custom': {
'drawStyle': 'line',
'lineStyle': {'dash': [10, 10]}
}
}
}
}
]
def _generate_template_variables(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Generate template variables for dynamic dashboard filtering."""
service_name = service_def.get('name', 'service')
return [
{
'name': 'environment',
'type': 'query',
'query': 'label_values(environment)',
'current': {'text': 'production', 'value': 'production'},
'includeAll': False,
'multi': False,
'refresh': 'on_dashboard_load'
},
{
'name': 'instance',
'type': 'query',
'query': f'label_values(up{{service="{service_name}"}}, instance)',
'current': {'text': 'All', 'value': '$__all'},
'includeAll': True,
'multi': True,
'refresh': 'on_time_range_change'
},
{
'name': 'handler',
'type': 'query',
'query': f'label_values(http_requests_total{{service="{service_name}"}}, handler)',
'current': {'text': 'All', 'value': '$__all'},
'includeAll': True,
'multi': True,
'refresh': 'on_time_range_change'
}
]
def _generate_alerts_integration(self, service_def: Dict[str, Any]) -> Dict[str, Any]:
"""Generate alerts integration configuration."""
service_name = service_def.get('name', 'service')
return {
'alert_annotations': True,
'alert_rules_query': f'ALERTS{{service="{service_name}"}}',
'alert_panels': [
{
'title': 'Active Alerts',
'type': 'table',
'query': f'ALERTS{{service="{service_name}",alertstate="firing"}}',
'columns': ['alertname', 'severity', 'instance', 'description']
}
]
}
def _generate_drill_down_paths(self, service_def: Dict[str, Any]) -> Dict[str, Any]:
"""Generate drill-down navigation paths."""
service_name = service_def.get('name', 'service')
return {
'service_overview': {
'from': 'service_status',
'to': 'detailed_health_dashboard',
'url': f'/d/service-health/{service_name}-health',
'params': ['var-service', 'var-environment']
},
'error_investigation': {
'from': 'errors',
'to': 'error_details_dashboard',
'url': f'/d/errors/{service_name}-errors',
'params': ['var-service', 'var-time_range']
},
'latency_analysis': {
'from': 'latency',
'to': 'trace_analysis_dashboard',
'url': f'/d/traces/{service_name}-traces',
'params': ['var-service', 'var-handler']
},
'capacity_planning': {
'from': 'saturation',
'to': 'capacity_dashboard',
'url': f'/d/capacity/{service_name}-capacity',
'params': ['var-service', 'var-time_range']
}
}
def generate_grafana_json(self, dashboard_spec: Dict[str, Any]) -> Dict[str, Any]:
"""Convert dashboard specification to Grafana JSON format."""
metadata = dashboard_spec['metadata']
config = dashboard_spec['configuration']
grafana_json = {
'dashboard': {
'id': None,
'title': metadata['title'],
'tags': [metadata['service']['type'], metadata['target_role'], 'generated'],
'timezone': config['timezone'],
'refresh': config['refresh_interval'],
'time': {
'from': 'now-1h',
'to': 'now'
},
'templating': {
'list': dashboard_spec['variables']
},
'panels': self._convert_panels_to_grafana_format(dashboard_spec['panels']),
'version': 1,
'schemaVersion': 30
},
'overwrite': True
}
return grafana_json
def _convert_panels_to_grafana_format(self, panels: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Convert panel specifications to Grafana format."""
grafana_panels = []
for panel in panels:
grafana_panel = {
'id': hash(panel['id']) % 1000, # Generate numeric ID
'title': panel['title'],
'type': panel['type'],
'gridPos': panel['grid_pos'],
'targets': panel['targets'],
'fieldConfig': panel.get('field_config', {}),
'options': panel.get('options', {}),
'transformations': panel.get('transformations', [])
}
grafana_panels.append(grafana_panel)
return grafana_panels
def generate_documentation(self, dashboard_spec: Dict[str, Any]) -> str:
"""Generate documentation for the dashboard."""
metadata = dashboard_spec['metadata']
service = metadata['service']
doc_content = f"""# {metadata['title']} Documentation
## Overview
This dashboard provides comprehensive monitoring for {service['name']}, a {service['type']} service with {service['criticality']} criticality.
**Target Audience:** {metadata['target_role'].upper()} teams
**Generated:** {metadata['generated_at']}
## Dashboard Sections
### Service Overview
- **Service Status**: Real-time availability status
- **SLO Achievement**: 30-day SLO compliance metrics
- **Error Budget**: Remaining error budget visualization
### Golden Signals Monitoring
- **Latency**: P50, P95, P99 response times
- **Traffic**: Request rate by status code
- **Errors**: Error rates for 4xx and 5xx responses
- **Saturation**: CPU and memory utilization
### Resource Utilization
- **CPU Usage**: Process CPU consumption
- **Memory Usage**: Memory utilization tracking
- **Network I/O**: Network throughput metrics
- **Disk I/O**: Disk read/write operations
## Key Metrics
### SLIs Tracked
"""
# Add service-type specific metrics
service_type = service.get('type', 'api')
if service_type in self.SERVICE_METRICS:
metrics = self.SERVICE_METRICS[service_type]['key_metrics']
for metric in metrics:
doc_content += f"- `{metric}`: Core service metric\n"
doc_content += f"""
## Alert Integration
- Active alerts are displayed in context with relevant panels
- Alert annotations show on time series charts
- Click-through to alert management system available
## Drill-Down Paths
"""
drill_downs = dashboard_spec.get('drill_down_paths', {})
for path_name, path_config in drill_downs.items():
doc_content += f"- **{path_name}**: From {path_config['from']} → {path_config['to']}\n"
doc_content += f"""
## Usage Guidelines
### Time Ranges
Use appropriate time ranges for different investigation types:
- **Real-time monitoring**: 15m - 1h
- **Recent incident investigation**: 1h - 6h
- **Trend analysis**: 1d - 7d
- **Capacity planning**: 7d - 30d
### Variables
- **environment**: Filter by deployment environment
- **instance**: Focus on specific service instances
- **handler**: Filter by API endpoint or handler
### Performance Optimization
- Use longer time ranges for capacity planning
- Refresh intervals are optimized per role:
- SRE: 30s for operational awareness
- Developer: 1m for troubleshooting
- Executive: 5m for high-level monitoring
## Maintenance
- Dashboard panels automatically adapt to service changes
- Template variables refresh based on actual metric labels
- Review and update business metrics quarterly
"""
return doc_content
def export_specification(self, dashboard_spec: Dict[str, Any], output_file: str,
format_type: str = 'json'):
"""Export dashboard specification."""
if format_type.lower() == 'json':
with open(output_file, 'w') as f:
json.dump(dashboard_spec, f, indent=2)
elif format_type.lower() == 'grafana':
grafana_json = self.generate_grafana_json(dashboard_spec)
with open(output_file, 'w') as f:
json.dump(grafana_json, f, indent=2)
else:
raise ValueError(f"Unsupported format: {format_type}")
def print_summary(self, dashboard_spec: Dict[str, Any]):
"""Print human-readable summary of dashboard specification."""
metadata = dashboard_spec['metadata']
service = metadata['service']
config = dashboard_spec['configuration']
panels = dashboard_spec['panels']
print(f"\n{'='*60}")
print(f"DASHBOARD SPECIFICATION SUMMARY")
print(f"{'='*60}")
print(f"\nDashboard Details:")
print(f" Title: {metadata['title']}")
print(f" Target Role: {metadata['target_role'].upper()}")
print(f" Service: {service['name']} ({service['type']})")
print(f" Criticality: {service['criticality']}")
print(f" Generated: {metadata['generated_at']}")
print(f"\nConfiguration:")
print(f" Default Time Range: {config['default_time_range']}")
print(f" Refresh Interval: {config['refresh_interval']}")
print(f" Available Time Ranges: {', '.join(config['time_ranges'])}")
print(f"\nPanels ({len(panels)}):")
panel_types = {}
for panel in panels:
panel_type = panel['type']
panel_types[panel_type] = panel_types.get(panel_type, 0) + 1
for panel_type, count in panel_types.items():
print(f" {panel_type}: {count}")
variables = dashboard_spec.get('variables', [])
print(f"\nTemplate Variables ({len(variables)}):")
for var in variables:
print(f" {var['name']} ({var['type']})")
drill_downs = dashboard_spec.get('drill_down_paths', {})
print(f"\nDrill-down Paths: {len(drill_downs)}")
print(f"\nKey Features:")
print(f" • Golden Signals monitoring")
print(f" • Resource utilization tracking")
print(f" • Alert integration")
print(f" • Role-optimized layout")
print(f" • Service-type specific panels")
print(f"\n{'='*60}\n")
def main():
"""Main function for CLI usage."""
parser = argparse.ArgumentParser(
description='Generate comprehensive dashboard specifications',
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
# Generate from service definition file
python dashboard_generator.py --input service.json --output dashboard.json
# Generate from command line parameters
python dashboard_generator.py --service-type api --name "Payment Service" --output payment_dashboard.json
# Generate Grafana-compatible JSON
python dashboard_generator.py --input service.json --output dashboard.json --format grafana
# Generate with specific role focus
python dashboard_generator.py --service-type web --name "Frontend" --role developer --output frontend_dev.json
"""
)
parser.add_argument('--input', '-i',
help='Input service definition JSON file')
parser.add_argument('--output', '-o',
help='Output dashboard specification file')
parser.add_argument('--service-type',
choices=['api', 'web', 'database', 'queue', 'batch', 'ml'],
help='Service type')
parser.add_argument('--name',
help='Service name')
parser.add_argument('--criticality',
choices=['critical', 'high', 'medium', 'low'],
default='medium',
help='Service criticality level')
parser.add_argument('--role',
choices=['sre', 'developer', 'executive', 'ops'],
default='sre',
help='Target role for dashboard optimization')
parser.add_argument('--format',
choices=['json', 'grafana'],
default='json',
help='Output format (json specification or grafana compatible)')
parser.add_argument('--doc-output',
help='Generate documentation file')
parser.add_argument('--summary-only', action='store_true',
help='Only display summary, do not save files')
args = parser.parse_args()
if not args.input and not (args.service_type and args.name):
parser.error("Must provide either --input file or --service-type and --name")
generator = DashboardGenerator()
try:
# Load or create service definition
if args.input:
service_def = generator.load_service_definition(args.input)
else:
service_def = generator.create_service_definition(
args.service_type, args.name, args.criticality
)
# Generate dashboard specification
dashboard_spec = generator.generate_dashboard_specification(service_def, args.role)
# Output results
if not args.summary_only:
output_file = args.output or f"{service_def['name'].replace(' ', '_').lower()}_dashboard.json"
generator.export_specification(dashboard_spec, output_file, args.format)
print(f"Dashboard specification saved to: {output_file}")
# Generate documentation if requested
if args.doc_output:
documentation = generator.generate_documentation(dashboard_spec)
with open(args.doc_output, 'w') as f:
f.write(documentation)
print(f"Documentation saved to: {args.doc_output}")
# Always show summary
generator.print_summary(dashboard_spec)
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
sys.exit(1)
if __name__ == '__main__':
main()
FILE:scripts/slo_designer.py
#!/usr/bin/env python3
"""
SLO Designer - Generate comprehensive SLI/SLO frameworks for services
This script analyzes service descriptions and generates complete SLO frameworks including:
- SLI definitions based on service characteristics
- SLO targets based on criticality and user impact
- Error budget calculations and policies
- Multi-window burn rate alerts
- SLA recommendations for customer-facing services
Usage:
python slo_designer.py --input service_definition.json --output slo_framework.json
python slo_designer.py --service-type api --criticality high --user-facing true
"""
import json
import argparse
import sys
import math
from typing import Dict, List, Any, Tuple
from datetime import datetime, timedelta
class SLODesigner:
"""Design and generate SLO frameworks for services."""
# SLO target recommendations based on service criticality
SLO_TARGETS = {
'critical': {
'availability': 0.9999, # 99.99% - 4.38 minutes downtime/month
'latency_p95': 100, # 95th percentile latency in ms
'latency_p99': 500, # 99th percentile latency in ms
'error_rate': 0.001 # 0.1% error rate
},
'high': {
'availability': 0.999, # 99.9% - 43.8 minutes downtime/month
'latency_p95': 200, # 95th percentile latency in ms
'latency_p99': 1000, # 99th percentile latency in ms
'error_rate': 0.005 # 0.5% error rate
},
'medium': {
'availability': 0.995, # 99.5% - 3.65 hours downtime/month
'latency_p95': 500, # 95th percentile latency in ms
'latency_p99': 2000, # 99th percentile latency in ms
'error_rate': 0.01 # 1% error rate
},
'low': {
'availability': 0.99, # 99% - 7.3 hours downtime/month
'latency_p95': 1000, # 95th percentile latency in ms
'latency_p99': 5000, # 99th percentile latency in ms
'error_rate': 0.02 # 2% error rate
}
}
# Burn rate windows for multi-window alerting
BURN_RATE_WINDOWS = [
{'short': '5m', 'long': '1h', 'burn_rate': 14.4, 'budget_consumed': '2%'},
{'short': '30m', 'long': '6h', 'burn_rate': 6, 'budget_consumed': '5%'},
{'short': '2h', 'long': '1d', 'burn_rate': 3, 'budget_consumed': '10%'},
{'short': '6h', 'long': '3d', 'burn_rate': 1, 'budget_consumed': '10%'}
]
# Service type specific SLI recommendations
SERVICE_TYPE_SLIS = {
'api': ['availability', 'latency', 'error_rate', 'throughput'],
'web': ['availability', 'latency', 'error_rate', 'page_load_time'],
'database': ['availability', 'query_latency', 'connection_success_rate', 'replication_lag'],
'queue': ['availability', 'message_processing_time', 'queue_depth', 'message_loss_rate'],
'batch': ['job_success_rate', 'job_duration', 'data_freshness', 'resource_utilization'],
'ml': ['model_accuracy', 'prediction_latency', 'training_success_rate', 'feature_freshness']
}
def __init__(self):
"""Initialize the SLO Designer."""
self.service_config = {}
self.slo_framework = {}
def load_service_definition(self, file_path: str) -> Dict[str, Any]:
"""Load service definition from JSON file."""
try:
with open(file_path, 'r') as f:
return json.load(f)
except FileNotFoundError:
raise ValueError(f"Service definition file not found: {file_path}")
except json.JSONDecodeError as e:
raise ValueError(f"Invalid JSON in service definition: {e}")
def create_service_definition(self, service_type: str, criticality: str,
user_facing: bool, name: str = None) -> Dict[str, Any]:
"""Create a service definition from parameters."""
return {
'name': name or f'{service_type}_service',
'type': service_type,
'criticality': criticality,
'user_facing': user_facing,
'description': f'A {criticality} criticality {service_type} service',
'dependencies': [],
'team': 'platform',
'environment': 'production'
}
def generate_slis(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Generate Service Level Indicators based on service characteristics."""
service_type = service_def.get('type', 'api')
base_slis = self.SERVICE_TYPE_SLIS.get(service_type, ['availability', 'latency', 'error_rate'])
slis = []
for sli_name in base_slis:
sli = self._create_sli_definition(sli_name, service_def)
if sli:
slis.append(sli)
# Add user-facing specific SLIs
if service_def.get('user_facing', False):
user_slis = self._generate_user_facing_slis(service_def)
slis.extend(user_slis)
return slis
def _create_sli_definition(self, sli_name: str, service_def: Dict[str, Any]) -> Dict[str, Any]:
"""Create detailed SLI definition."""
service_name = service_def.get('name', 'service')
sli_definitions = {
'availability': {
'name': 'Availability',
'description': 'Percentage of successful requests',
'type': 'ratio',
'good_events': f'sum(rate(http_requests_total{{service="{service_name}",code!~"5.."}}))',
'total_events': f'sum(rate(http_requests_total{{service="{service_name}"}}))',
'unit': 'percentage'
},
'latency': {
'name': 'Request Latency P95',
'description': '95th percentile of request latency',
'type': 'threshold',
'query': f'histogram_quantile(0.95, rate(http_request_duration_seconds_bucket{{service="{service_name}"}}[5m]))',
'unit': 'seconds'
},
'error_rate': {
'name': 'Error Rate',
'description': 'Rate of 5xx errors',
'type': 'ratio',
'good_events': f'sum(rate(http_requests_total{{service="{service_name}",code!~"5.."}}))',
'total_events': f'sum(rate(http_requests_total{{service="{service_name}"}}))',
'unit': 'percentage'
},
'throughput': {
'name': 'Request Throughput',
'description': 'Requests per second',
'type': 'gauge',
'query': f'sum(rate(http_requests_total{{service="{service_name}"}}[5m]))',
'unit': 'requests/sec'
},
'page_load_time': {
'name': 'Page Load Time P95',
'description': '95th percentile of page load time',
'type': 'threshold',
'query': f'histogram_quantile(0.95, rate(page_load_duration_seconds_bucket{{service="{service_name}"}}[5m]))',
'unit': 'seconds'
},
'query_latency': {
'name': 'Database Query Latency P95',
'description': '95th percentile of database query latency',
'type': 'threshold',
'query': f'histogram_quantile(0.95, rate(db_query_duration_seconds_bucket{{service="{service_name}"}}[5m]))',
'unit': 'seconds'
},
'connection_success_rate': {
'name': 'Database Connection Success Rate',
'description': 'Percentage of successful database connections',
'type': 'ratio',
'good_events': f'sum(rate(db_connections_total{{service="{service_name}",status="success"}}[5m]))',
'total_events': f'sum(rate(db_connections_total{{service="{service_name}"}}[5m]))',
'unit': 'percentage'
}
}
return sli_definitions.get(sli_name)
def _generate_user_facing_slis(self, service_def: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Generate additional SLIs for user-facing services."""
service_name = service_def.get('name', 'service')
return [
{
'name': 'User Journey Success Rate',
'description': 'Percentage of successful complete user journeys',
'type': 'ratio',
'good_events': f'sum(rate(user_journey_total{{service="{service_name}",status="success"}}[5m]))',
'total_events': f'sum(rate(user_journey_total{{service="{service_name}"}}[5m]))',
'unit': 'percentage'
},
{
'name': 'Feature Availability',
'description': 'Percentage of time key features are available',
'type': 'ratio',
'good_events': f'sum(rate(feature_checks_total{{service="{service_name}",status="available"}}[5m]))',
'total_events': f'sum(rate(feature_checks_total{{service="{service_name}"}}[5m]))',
'unit': 'percentage'
}
]
def generate_slos(self, service_def: Dict[str, Any], slis: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Generate Service Level Objectives based on service criticality."""
criticality = service_def.get('criticality', 'medium')
targets = self.SLO_TARGETS.get(criticality, self.SLO_TARGETS['medium'])
slos = []
for sli in slis:
slo = self._create_slo_from_sli(sli, targets, service_def)
if slo:
slos.append(slo)
return slos
def _create_slo_from_sli(self, sli: Dict[str, Any], targets: Dict[str, float],
service_def: Dict[str, Any]) -> Dict[str, Any]:
"""Create SLO definition from SLI."""
sli_name = sli['name'].lower().replace(' ', '_')
# Map SLI names to target keys
target_mapping = {
'availability': 'availability',
'request_latency_p95': 'latency_p95',
'error_rate': 'error_rate',
'user_journey_success_rate': 'availability',
'feature_availability': 'availability',
'page_load_time_p95': 'latency_p95',
'database_query_latency_p95': 'latency_p95',
'database_connection_success_rate': 'availability'
}
target_key = target_mapping.get(sli_name)
if not target_key:
return None
target_value = targets.get(target_key)
if target_value is None:
return None
# Determine comparison operator and format target
if 'latency' in sli_name or 'duration' in sli_name:
operator = '<='
target_display = f"{target_value}ms" if target_value < 10 else f"{target_value/1000}s"
elif 'rate' in sli_name and 'error' in sli_name:
operator = '<='
target_display = f"{target_value * 100}%"
target_value = target_value # Keep as decimal
else:
operator = '>='
target_display = f"{target_value * 100}%"
# Calculate time windows
time_windows = ['1h', '1d', '7d', '30d']
slo = {
'name': f"{sli['name']} SLO",
'description': f"Service level objective for {sli['description'].lower()}",
'sli_name': sli['name'],
'target_value': target_value,
'target_display': target_display,
'operator': operator,
'time_windows': time_windows,
'measurement_window': '30d',
'service': service_def.get('name', 'service'),
'criticality': service_def.get('criticality', 'medium')
}
return slo
def calculate_error_budgets(self, slos: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Calculate error budgets for SLOs."""
error_budgets = []
for slo in slos:
if slo['operator'] == '>=': # Availability-type SLOs
target = slo['target_value']
error_budget_rate = 1 - target
# Calculate budget for different time windows
time_windows = {
'1h': 3600,
'1d': 86400,
'7d': 604800,
'30d': 2592000
}
budgets = {}
for window, seconds in time_windows.items():
budget_seconds = seconds * error_budget_rate
if budget_seconds < 60:
budgets[window] = f"{budget_seconds:.1f} seconds"
elif budget_seconds < 3600:
budgets[window] = f"{budget_seconds/60:.1f} minutes"
else:
budgets[window] = f"{budget_seconds/3600:.1f} hours"
error_budget = {
'slo_name': slo['name'],
'error_budget_rate': error_budget_rate,
'error_budget_percentage': f"{error_budget_rate * 100:.3f}%",
'budgets_by_window': budgets,
'burn_rate_alerts': self._generate_burn_rate_alerts(slo, error_budget_rate)
}
error_budgets.append(error_budget)
return error_budgets
def _generate_burn_rate_alerts(self, slo: Dict[str, Any], error_budget_rate: float) -> List[Dict[str, Any]]:
"""Generate multi-window burn rate alerts."""
alerts = []
service_name = slo['service']
sli_query = self._get_sli_query_for_burn_rate(slo)
for window_config in self.BURN_RATE_WINDOWS:
alert = {
'name': f"{slo['sli_name']} Burn Rate {window_config['budget_consumed']} Alert",
'description': f"Alert when {slo['sli_name']} is consuming error budget at {window_config['burn_rate']}x rate",
'severity': self._determine_alert_severity(float(window_config['budget_consumed'].rstrip('%'))),
'short_window': window_config['short'],
'long_window': window_config['long'],
'burn_rate_threshold': window_config['burn_rate'],
'budget_consumed': window_config['budget_consumed'],
'condition': f"({sli_query}_short > {window_config['burn_rate']}) and ({sli_query}_long > {window_config['burn_rate']})",
'annotations': {
'summary': f"High burn rate detected for {slo['sli_name']}",
'description': f"Error budget consumption rate is {window_config['burn_rate']}x normal, will exhaust {window_config['budget_consumed']} of monthly budget"
}
}
alerts.append(alert)
return alerts
def _get_sli_query_for_burn_rate(self, slo: Dict[str, Any]) -> str:
"""Generate SLI query fragment for burn rate calculation."""
service_name = slo['service']
sli_name = slo['sli_name'].lower().replace(' ', '_')
if 'availability' in sli_name or 'success' in sli_name:
return f"(1 - (sum(rate(http_requests_total{{service='{service_name}',code!~'5..'}})) / sum(rate(http_requests_total{{service='{service_name}'}}))))"
elif 'error' in sli_name:
return f"(sum(rate(http_requests_total{{service='{service_name}',code=~'5..'}})) / sum(rate(http_requests_total{{service='{service_name}'}})))"
else:
return f"sli_burn_rate_{sli_name}"
def _determine_alert_severity(self, budget_consumed_percent: float) -> str:
"""Determine alert severity based on budget consumption rate."""
if budget_consumed_percent <= 2:
return 'critical'
elif budget_consumed_percent <= 5:
return 'warning'
else:
return 'info'
def generate_sla_recommendations(self, service_def: Dict[str, Any],
slos: List[Dict[str, Any]]) -> Dict[str, Any]:
"""Generate SLA recommendations for customer-facing services."""
if not service_def.get('user_facing', False):
return {
'applicable': False,
'reason': 'SLA not recommended for non-user-facing services'
}
criticality = service_def.get('criticality', 'medium')
# SLA targets should be more conservative than SLO targets
sla_buffer = 0.001 # 0.1% buffer below SLO
sla_recommendations = {
'applicable': True,
'service': service_def.get('name'),
'commitments': [],
'penalties': self._generate_penalty_structure(criticality),
'measurement_methodology': 'External synthetic monitoring from multiple geographic locations',
'exclusions': [
'Planned maintenance windows (with 72h advance notice)',
'Customer-side network or infrastructure issues',
'Force majeure events',
'Third-party service dependencies beyond our control'
]
}
for slo in slos:
if slo['operator'] == '>=' and 'availability' in slo['sli_name'].lower():
sla_target = max(0.9, slo['target_value'] - sla_buffer)
commitment = {
'metric': slo['sli_name'],
'target': sla_target,
'target_display': f"{sla_target * 100:.2f}%",
'measurement_window': 'monthly',
'measurement_method': 'Uptime monitoring with 1-minute granularity'
}
sla_recommendations['commitments'].append(commitment)
return sla_recommendations
def _generate_penalty_structure(self, criticality: str) -> List[Dict[str, Any]]:
"""Generate penalty structure based on service criticality."""
penalty_structures = {
'critical': [
{'breach_threshold': '< 99.99%', 'credit_percentage': 10},
{'breach_threshold': '< 99.9%', 'credit_percentage': 25},
{'breach_threshold': '< 99%', 'credit_percentage': 50}
],
'high': [
{'breach_threshold': '< 99.9%', 'credit_percentage': 10},
{'breach_threshold': '< 99.5%', 'credit_percentage': 25}
],
'medium': [
{'breach_threshold': '< 99.5%', 'credit_percentage': 10}
],
'low': []
}
return penalty_structures.get(criticality, [])
def generate_framework(self, service_def: Dict[str, Any]) -> Dict[str, Any]:
"""Generate complete SLO framework."""
# Generate SLIs
slis = self.generate_slis(service_def)
# Generate SLOs
slos = self.generate_slos(service_def, slis)
# Calculate error budgets
error_budgets = self.calculate_error_budgets(slos)
# Generate SLA recommendations
sla_recommendations = self.generate_sla_recommendations(service_def, slos)
# Create comprehensive framework
framework = {
'metadata': {
'service': service_def,
'generated_at': datetime.utcnow().isoformat() + 'Z',
'framework_version': '1.0'
},
'slis': slis,
'slos': slos,
'error_budgets': error_budgets,
'sla_recommendations': sla_recommendations,
'monitoring_recommendations': self._generate_monitoring_recommendations(service_def),
'implementation_guide': self._generate_implementation_guide(service_def, slis, slos)
}
return framework
def _generate_monitoring_recommendations(self, service_def: Dict[str, Any]) -> Dict[str, Any]:
"""Generate monitoring tool recommendations."""
service_type = service_def.get('type', 'api')
recommendations = {
'metrics': {
'collection': 'Prometheus with service discovery',
'retention': '90 days for raw metrics, 1 year for aggregated',
'alerting': 'Prometheus Alertmanager with multi-window burn rate alerts'
},
'logging': {
'format': 'Structured JSON logs with correlation IDs',
'aggregation': 'ELK stack or equivalent with proper indexing',
'retention': '30 days for debug logs, 90 days for error logs'
},
'tracing': {
'sampling': 'Adaptive sampling with 1% base rate',
'storage': 'Jaeger or Zipkin with 7-day retention',
'integration': 'OpenTelemetry instrumentation'
}
}
if service_type == 'web':
recommendations['synthetic_monitoring'] = {
'frequency': 'Every 1 minute from 3+ geographic locations',
'checks': 'Full user journey simulation',
'tools': 'Pingdom, DataDog Synthetics, or equivalent'
}
return recommendations
def _generate_implementation_guide(self, service_def: Dict[str, Any],
slis: List[Dict[str, Any]],
slos: List[Dict[str, Any]]) -> Dict[str, Any]:
"""Generate implementation guide for the SLO framework."""
return {
'prerequisites': [
'Service instrumented with metrics collection (Prometheus format)',
'Structured logging with correlation IDs',
'Monitoring infrastructure (Prometheus, Grafana, Alertmanager)',
'Incident response processes and escalation policies'
],
'implementation_steps': [
{
'step': 1,
'title': 'Instrument Service',
'description': 'Add metrics collection for all defined SLIs',
'estimated_effort': '1-2 days'
},
{
'step': 2,
'title': 'Configure Recording Rules',
'description': 'Set up Prometheus recording rules for SLI calculations',
'estimated_effort': '4-8 hours'
},
{
'step': 3,
'title': 'Implement Burn Rate Alerts',
'description': 'Configure multi-window burn rate alerting rules',
'estimated_effort': '1 day'
},
{
'step': 4,
'title': 'Create SLO Dashboard',
'description': 'Build Grafana dashboard for SLO tracking and error budget monitoring',
'estimated_effort': '4-6 hours'
},
{
'step': 5,
'title': 'Test and Validate',
'description': 'Test alerting and validate SLI measurements against expectations',
'estimated_effort': '1-2 days'
},
{
'step': 6,
'title': 'Documentation and Training',
'description': 'Document runbooks and train team on SLO monitoring',
'estimated_effort': '1 day'
}
],
'validation_checklist': [
'All SLIs produce expected metric values',
'Burn rate alerts fire correctly during simulated outages',
'Error budget calculations match manual verification',
'Dashboard displays accurate SLO achievement rates',
'Alert routing reaches correct escalation paths',
'Runbooks are complete and tested'
]
}
def export_json(self, framework: Dict[str, Any], output_file: str):
"""Export framework as JSON."""
with open(output_file, 'w') as f:
json.dump(framework, f, indent=2)
def print_summary(self, framework: Dict[str, Any]):
"""Print human-readable summary of the SLO framework."""
service = framework['metadata']['service']
slis = framework['slis']
slos = framework['slos']
error_budgets = framework['error_budgets']
print(f"\n{'='*60}")
print(f"SLO FRAMEWORK SUMMARY FOR {service['name'].upper()}")
print(f"{'='*60}")
print(f"\nService Details:")
print(f" Type: {service['type']}")
print(f" Criticality: {service['criticality']}")
print(f" User Facing: {'Yes' if service.get('user_facing') else 'No'}")
print(f" Team: {service.get('team', 'Unknown')}")
print(f"\nService Level Indicators ({len(slis)}):")
for i, sli in enumerate(slis, 1):
print(f" {i}. {sli['name']}")
print(f" Description: {sli['description']}")
print(f" Type: {sli['type']}")
print()
print(f"Service Level Objectives ({len(slos)}):")
for i, slo in enumerate(slos, 1):
print(f" {i}. {slo['name']}")
print(f" Target: {slo['target_display']}")
print(f" Measurement Window: {slo['measurement_window']}")
print()
print(f"Error Budget Summary:")
for budget in error_budgets:
print(f" {budget['slo_name']}:")
print(f" Monthly Budget: {budget['error_budget_percentage']}")
print(f" Burn Rate Alerts: {len(budget['burn_rate_alerts'])}")
print()
sla = framework['sla_recommendations']
if sla['applicable']:
print(f"SLA Recommendations:")
print(f" Commitments: {len(sla['commitments'])}")
print(f" Penalty Tiers: {len(sla['penalties'])}")
else:
print(f"SLA Recommendations: {sla['reason']}")
print(f"\nImplementation Timeline: 1-2 weeks")
print(f"Framework generated at: {framework['metadata']['generated_at']}")
print(f"{'='*60}\n")
def main():
"""Main function for CLI usage."""
parser = argparse.ArgumentParser(
description='Generate comprehensive SLO frameworks for services',
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
# Generate from service definition file
python slo_designer.py --input service.json --output framework.json
# Generate from command line parameters
python slo_designer.py --service-type api --criticality high --user-facing true --output framework.json
# Generate and display summary only
python slo_designer.py --service-type web --criticality critical --user-facing true --summary-only
"""
)
parser.add_argument('--input', '-i',
help='Input service definition JSON file')
parser.add_argument('--output', '-o',
help='Output framework JSON file')
parser.add_argument('--service-type',
choices=['api', 'web', 'database', 'queue', 'batch', 'ml'],
help='Service type')
parser.add_argument('--criticality',
choices=['critical', 'high', 'medium', 'low'],
help='Service criticality level')
parser.add_argument('--user-facing',
choices=['true', 'false'],
help='Whether service is user-facing')
parser.add_argument('--service-name',
help='Service name')
parser.add_argument('--summary-only', action='store_true',
help='Only display summary, do not save JSON')
args = parser.parse_args()
if not args.input and not (args.service_type and args.criticality and args.user_facing):
parser.error("Must provide either --input file or --service-type, --criticality, and --user-facing")
designer = SLODesigner()
try:
# Load or create service definition
if args.input:
service_def = designer.load_service_definition(args.input)
else:
user_facing = args.user_facing.lower() == 'true'
service_def = designer.create_service_definition(
args.service_type, args.criticality, user_facing, args.service_name
)
# Generate framework
framework = designer.generate_framework(service_def)
# Output results
if not args.summary_only:
output_file = args.output or f"{service_def['name']}_slo_framework.json"
designer.export_json(framework, output_file)
print(f"SLO framework saved to: {output_file}")
# Always show summary
designer.print_summary(framework)
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
sys.exit(1)
if __name__ == '__main__':
main()Lập kế hoạch di chuyển không gián đoạn, kiểm tra tương thích và chiến lược rollback cho di chuyển hệ thống, cơ sở dữ liệu và hạ tầng.
---
name: "migration-architect"
description: "Zero-downtime migration planning, compatibility validation, and rollback strategy generation. Tools for system, database, and infrastructure migrations with minimal business impact. Use when planning a database migration, infrastructure cutover, system replacement, or any high-risk transition that needs explicit rollback paths."
---
# Migration Architect
**Tier:** POWERFUL
**Category:** Engineering - Migration Strategy
**Purpose:** Zero-downtime migration planning, compatibility validation, and rollback strategy generation
## Overview
The Migration Architect skill provides comprehensive tools and methodologies for planning, executing, and validating complex system migrations with minimal business impact. This skill combines proven migration patterns with automated planning tools to ensure successful transitions between systems, databases, and infrastructure.
## Core Capabilities
### 1. Migration Strategy Planning
- **Phased Migration Planning:** Break complex migrations into manageable phases with clear validation gates
- **Risk Assessment:** Identify potential failure points and mitigation strategies before execution
- **Timeline Estimation:** Generate realistic timelines based on migration complexity and resource constraints
- **Stakeholder Communication:** Create communication templates and progress dashboards
### 2. Compatibility Analysis
- **Schema Evolution:** Analyze database schema changes for backward compatibility issues
- **API Versioning:** Detect breaking changes in REST/GraphQL APIs and microservice interfaces
- **Data Type Validation:** Identify data format mismatches and conversion requirements
- **Constraint Analysis:** Validate referential integrity and business rule changes
### 3. Rollback Strategy Generation
- **Automated Rollback Plans:** Generate comprehensive rollback procedures for each migration phase
- **Data Recovery Scripts:** Create point-in-time data restoration procedures
- **Service Rollback:** Plan service version rollbacks with traffic management
- **Validation Checkpoints:** Define success criteria and rollback triggers
## Migration Patterns
### Database Migrations
#### Schema Evolution Patterns
1. **Expand-Contract Pattern**
- **Expand:** Add new columns/tables alongside existing schema
- **Dual Write:** Application writes to both old and new schema
- **Migration:** Backfill historical data to new schema
- **Contract:** Remove old columns/tables after validation
2. **Parallel Schema Pattern**
- Run new schema in parallel with existing schema
- Use feature flags to route traffic between schemas
- Validate data consistency between parallel systems
- Cutover when confidence is high
3. **Event Sourcing Migration**
- Capture all changes as events during migration window
- Apply events to new schema for consistency
- Enable replay capability for rollback scenarios
#### Data Migration Strategies
1. **Bulk Data Migration**
- **Snapshot Approach:** Full data copy during maintenance window
- **Incremental Sync:** Continuous data synchronization with change tracking
- **Stream Processing:** Real-time data transformation pipelines
2. **Dual-Write Pattern**
- Write to both source and target systems during migration
- Implement compensation patterns for write failures
- Use distributed transactions where consistency is critical
3. **Change Data Capture (CDC)**
- Stream database changes to target system
- Maintain eventual consistency during migration
- Enable zero-downtime migrations for large datasets
### Service Migrations
#### Strangler Fig Pattern
1. **Intercept Requests:** Route traffic through proxy/gateway
2. **Gradually Replace:** Implement new service functionality incrementally
3. **Legacy Retirement:** Remove old service components as new ones prove stable
4. **Monitoring:** Track performance and error rates throughout transition
```mermaid
graph TD
A[Client Requests] --> B[API Gateway]
B --> C{Route Decision}
C -->|Legacy Path| D[Legacy Service]
C -->|New Path| E[New Service]
D --> F[Legacy Database]
E --> G[New Database]
```
#### Parallel Run Pattern
1. **Dual Execution:** Run both old and new services simultaneously
2. **Shadow Traffic:** Route production traffic to both systems
3. **Result Comparison:** Compare outputs to validate correctness
4. **Gradual Cutover:** Shift traffic percentage based on confidence
#### Canary Deployment Pattern
1. **Limited Rollout:** Deploy new service to small percentage of users
2. **Monitoring:** Track key metrics (latency, errors, business KPIs)
3. **Gradual Increase:** Increase traffic percentage as confidence grows
4. **Full Rollout:** Complete migration once validation passes
### Infrastructure Migrations
#### Cloud-to-Cloud Migration
1. **Assessment Phase**
- Inventory existing resources and dependencies
- Map services to target cloud equivalents
- Identify vendor-specific features requiring refactoring
2. **Pilot Migration**
- Migrate non-critical workloads first
- Validate performance and cost models
- Refine migration procedures
3. **Production Migration**
- Use infrastructure as code for consistency
- Implement cross-cloud networking during transition
- Maintain disaster recovery capabilities
#### On-Premises to Cloud Migration
1. **Lift and Shift**
- Minimal changes to existing applications
- Quick migration with optimization later
- Use cloud migration tools and services
2. **Re-architecture**
- Redesign applications for cloud-native patterns
- Adopt microservices, containers, and serverless
- Implement cloud security and scaling practices
3. **Hybrid Approach**
- Keep sensitive data on-premises
- Migrate compute workloads to cloud
- Implement secure connectivity between environments
## Feature Flags for Migrations
### Progressive Feature Rollout
```python
# Example feature flag implementation
class MigrationFeatureFlag:
def __init__(self, flag_name, rollout_percentage=0):
self.flag_name = flag_name
self.rollout_percentage = rollout_percentage
def is_enabled_for_user(self, user_id):
hash_value = hash(f"{self.flag_name}:{user_id}")
return (hash_value % 100) < self.rollout_percentage
def gradual_rollout(self, target_percentage, step_size=10):
while self.rollout_percentage < target_percentage:
self.rollout_percentage = min(
self.rollout_percentage + step_size,
target_percentage
)
yield self.rollout_percentage
```
### Circuit Breaker Pattern
Implement automatic fallback to legacy systems when new systems show degraded performance:
```python
class MigrationCircuitBreaker:
def __init__(self, failure_threshold=5, timeout=60):
self.failure_count = 0
self.failure_threshold = failure_threshold
self.timeout = timeout
self.last_failure_time = None
self.state = 'CLOSED' # CLOSED, OPEN, HALF_OPEN
def call_new_service(self, request):
if self.state == 'OPEN':
if self.should_attempt_reset():
self.state = 'HALF_OPEN'
else:
return self.fallback_to_legacy(request)
try:
response = self.new_service.process(request)
self.on_success()
return response
except Exception as e:
self.on_failure()
return self.fallback_to_legacy(request)
```
## Data Validation and Reconciliation
### Validation Strategies
1. **Row Count Validation**
- Compare record counts between source and target
- Account for soft deletes and filtered records
- Implement threshold-based alerting
2. **Checksums and Hashing**
- Generate checksums for critical data subsets
- Compare hash values to detect data drift
- Use sampling for large datasets
3. **Business Logic Validation**
- Run critical business queries on both systems
- Compare aggregate results (sums, counts, averages)
- Validate derived data and calculations
### Reconciliation Patterns
1. **Delta Detection**
```sql
-- Example delta query for reconciliation
SELECT 'missing_in_target' as issue_type, source_id
FROM source_table s
WHERE NOT EXISTS (
SELECT 1 FROM target_table t
WHERE t.id = s.id
)
UNION ALL
SELECT 'extra_in_target' as issue_type, target_id
FROM target_table t
WHERE NOT EXISTS (
SELECT 1 FROM source_table s
WHERE s.id = t.id
);
```
2. **Automated Correction**
- Implement data repair scripts for common issues
- Use idempotent operations for safe re-execution
- Log all correction actions for audit trails
## Rollback Strategies
### Database Rollback
1. **Schema Rollback**
- Maintain schema version control
- Use backward-compatible migrations when possible
- Keep rollback scripts for each migration step
2. **Data Rollback**
- Point-in-time recovery using database backups
- Transaction log replay for precise rollback points
- Maintain data snapshots at migration checkpoints
### Service Rollback
1. **Blue-Green Deployment**
- Keep previous service version running during migration
- Switch traffic back to blue environment if issues arise
- Maintain parallel infrastructure during migration window
2. **Rolling Rollback**
- Gradually shift traffic back to previous version
- Monitor system health during rollback process
- Implement automated rollback triggers
### Infrastructure Rollback
1. **Infrastructure as Code**
- Version control all infrastructure definitions
- Maintain rollback terraform/CloudFormation templates
- Test rollback procedures in staging environments
2. **Data Persistence**
- Preserve data in original location during migration
- Implement data sync back to original systems
- Maintain backup strategies across both environments
## Risk Assessment Framework
### Risk Categories
1. **Technical Risks**
- Data loss or corruption
- Service downtime or degraded performance
- Integration failures with dependent systems
- Scalability issues under production load
2. **Business Risks**
- Revenue impact from service disruption
- Customer experience degradation
- Compliance and regulatory concerns
- Brand reputation impact
3. **Operational Risks**
- Team knowledge gaps
- Insufficient testing coverage
- Inadequate monitoring and alerting
- Communication breakdowns
### Risk Mitigation Strategies
1. **Technical Mitigations**
- Comprehensive testing (unit, integration, load, chaos)
- Gradual rollout with automated rollback triggers
- Data validation and reconciliation processes
- Performance monitoring and alerting
2. **Business Mitigations**
- Stakeholder communication plans
- Business continuity procedures
- Customer notification strategies
- Revenue protection measures
3. **Operational Mitigations**
- Team training and documentation
- Runbook creation and testing
- On-call rotation planning
- Post-migration review processes
## Migration Runbooks
### Pre-Migration Checklist
- [ ] Migration plan reviewed and approved
- [ ] Rollback procedures tested and validated
- [ ] Monitoring and alerting configured
- [ ] Team roles and responsibilities defined
- [ ] Stakeholder communication plan activated
- [ ] Backup and recovery procedures verified
- [ ] Test environment validation complete
- [ ] Performance benchmarks established
- [ ] Security review completed
- [ ] Compliance requirements verified
### During Migration
- [ ] Execute migration phases in planned order
- [ ] Monitor key performance indicators continuously
- [ ] Validate data consistency at each checkpoint
- [ ] Communicate progress to stakeholders
- [ ] Document any deviations from plan
- [ ] Execute rollback if success criteria not met
- [ ] Coordinate with dependent teams
- [ ] Maintain detailed execution logs
### Post-Migration
- [ ] Validate all success criteria met
- [ ] Perform comprehensive system health checks
- [ ] Execute data reconciliation procedures
- [ ] Monitor system performance over 72 hours
- [ ] Update documentation and runbooks
- [ ] Decommission legacy systems (if applicable)
- [ ] Conduct post-migration retrospective
- [ ] Archive migration artifacts
- [ ] Update disaster recovery procedures
## Communication Templates
### Executive Summary Template
```
Migration Status: [IN_PROGRESS | COMPLETED | ROLLED_BACK]
Start Time: [YYYY-MM-DD HH:MM UTC]
Current Phase: [X of Y]
Overall Progress: [X%]
Key Metrics:
- System Availability: [X.XX%]
- Data Migration Progress: [X.XX%]
- Performance Impact: [+/-X%]
- Issues Encountered: [X]
Next Steps:
1. [Action item 1]
2. [Action item 2]
Risk Assessment: [LOW | MEDIUM | HIGH]
Rollback Status: [AVAILABLE | NOT_AVAILABLE]
```
### Technical Team Update Template
```
Phase: [Phase Name] - [Status]
Duration: [Started] - [Expected End]
Completed Tasks:
✓ [Task 1]
✓ [Task 2]
In Progress:
🔄 [Task 3] - [X% complete]
Upcoming:
⏳ [Task 4] - [Expected start time]
Issues:
⚠️ [Issue description] - [Severity] - [ETA resolution]
Metrics:
- Migration Rate: [X records/minute]
- Error Rate: [X.XX%]
- System Load: [CPU/Memory/Disk]
```
## Success Metrics
### Technical Metrics
- **Migration Completion Rate:** Percentage of data/services successfully migrated
- **Downtime Duration:** Total system unavailability during migration
- **Data Consistency Score:** Percentage of data validation checks passing
- **Performance Delta:** Performance change compared to baseline
- **Error Rate:** Percentage of failed operations during migration
### Business Metrics
- **Customer Impact Score:** Measure of customer experience degradation
- **Revenue Protection:** Percentage of revenue maintained during migration
- **Time to Value:** Duration from migration start to business value realization
- **Stakeholder Satisfaction:** Post-migration stakeholder feedback scores
### Operational Metrics
- **Plan Adherence:** Percentage of migration executed according to plan
- **Issue Resolution Time:** Average time to resolve migration issues
- **Team Efficiency:** Resource utilization and productivity metrics
- **Knowledge Transfer Score:** Team readiness for post-migration operations
## Tools and Technologies
### Migration Planning Tools
- **migration_planner.py:** Automated migration plan generation
- **compatibility_checker.py:** Schema and API compatibility analysis
- **rollback_generator.py:** Comprehensive rollback procedure generation
### Validation Tools
- Database comparison utilities (schema and data)
- API contract testing frameworks
- Performance benchmarking tools
- Data quality validation pipelines
### Monitoring and Alerting
- Real-time migration progress dashboards
- Automated rollback trigger systems
- Business metric monitoring
- Stakeholder notification systems
## Best Practices
### Planning Phase
1. **Start with Risk Assessment:** Identify all potential failure modes before planning
2. **Design for Rollback:** Every migration step should have a tested rollback procedure
3. **Validate in Staging:** Execute full migration process in production-like environment
4. **Plan for Gradual Rollout:** Use feature flags and traffic routing for controlled migration
### Execution Phase
1. **Monitor Continuously:** Track both technical and business metrics throughout
2. **Communicate Proactively:** Keep all stakeholders informed of progress and issues
3. **Document Everything:** Maintain detailed logs for post-migration analysis
4. **Stay Flexible:** Be prepared to adjust timeline based on real-world performance
### Validation Phase
1. **Automate Validation:** Use automated tools for data consistency and performance checks
2. **Business Logic Testing:** Validate critical business processes end-to-end
3. **Load Testing:** Verify system performance under expected production load
4. **Security Validation:** Ensure security controls function properly in new environment
## Integration with Development Lifecycle
### CI/CD Integration
```yaml
# Example migration pipeline stage
migration_validation:
stage: test
script:
- python scripts/compatibility_checker.py --before=old_schema.json --after=new_schema.json
- python scripts/migration_planner.py --config=migration_config.json --validate
artifacts:
reports:
- compatibility_report.json
- migration_plan.json
```
### Infrastructure as Code
```terraform
# Example Terraform for blue-green infrastructure
resource "aws_instance" "blue_environment" {
count = var.migration_phase == "preparation" ? var.instance_count : 0
# Blue environment configuration
}
resource "aws_instance" "green_environment" {
count = var.migration_phase == "execution" ? var.instance_count : 0
# Green environment configuration
}
```
This Migration Architect skill provides a comprehensive framework for planning, executing, and validating complex system migrations while minimizing business impact and technical risk. The combination of automated tools, proven patterns, and detailed procedures enables organizations to confidently undertake even the most complex migration projects.
FILE:assets/database_schema_after.json
{
"schema_version": "2.0",
"database": "user_management_v2",
"tables": {
"users": {
"columns": {
"id": {
"type": "bigint",
"nullable": false,
"primary_key": true,
"auto_increment": true
},
"username": {
"type": "varchar",
"length": 50,
"nullable": false,
"unique": true
},
"email": {
"type": "varchar",
"length": 320,
"nullable": false,
"unique": true
},
"password_hash": {
"type": "varchar",
"length": 255,
"nullable": false
},
"first_name": {
"type": "varchar",
"length": 100,
"nullable": true
},
"last_name": {
"type": "varchar",
"length": 100,
"nullable": true
},
"created_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP"
},
"updated_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP"
},
"is_active": {
"type": "boolean",
"nullable": false,
"default": true
},
"phone": {
"type": "varchar",
"length": 20,
"nullable": true
},
"email_verified_at": {
"type": "timestamp",
"nullable": true,
"comment": "When email was verified"
},
"phone_verified_at": {
"type": "timestamp",
"nullable": true,
"comment": "When phone was verified"
},
"two_factor_enabled": {
"type": "boolean",
"nullable": false,
"default": false
},
"last_login_at": {
"type": "timestamp",
"nullable": true
}
},
"constraints": {
"primary_key": ["id"],
"unique": [
"username",
"email"
],
"foreign_key": [],
"check": [
"email LIKE '%@%'",
"LENGTH(password_hash) >= 60",
"phone IS NULL OR LENGTH(phone) >= 10"
]
},
"indexes": [
{
"name": "idx_users_email",
"columns": ["email"],
"unique": true
},
{
"name": "idx_users_username",
"columns": ["username"],
"unique": true
},
{
"name": "idx_users_created_at",
"columns": ["created_at"]
},
{
"name": "idx_users_email_verified",
"columns": ["email_verified_at"]
},
{
"name": "idx_users_last_login",
"columns": ["last_login_at"]
}
]
},
"user_profiles": {
"columns": {
"id": {
"type": "bigint",
"nullable": false,
"primary_key": true,
"auto_increment": true
},
"user_id": {
"type": "bigint",
"nullable": false
},
"bio": {
"type": "text",
"nullable": true
},
"avatar_url": {
"type": "varchar",
"length": 500,
"nullable": true
},
"birth_date": {
"type": "date",
"nullable": true
},
"location": {
"type": "varchar",
"length": 100,
"nullable": true
},
"website": {
"type": "varchar",
"length": 255,
"nullable": true
},
"privacy_level": {
"type": "varchar",
"length": 20,
"nullable": false,
"default": "public"
},
"timezone": {
"type": "varchar",
"length": 50,
"nullable": true,
"default": "UTC"
},
"language": {
"type": "varchar",
"length": 10,
"nullable": false,
"default": "en"
},
"created_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP"
},
"updated_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP"
}
},
"constraints": {
"primary_key": ["id"],
"unique": [],
"foreign_key": [
{
"columns": ["user_id"],
"references": "users(id)",
"on_delete": "CASCADE"
}
],
"check": [
"privacy_level IN ('public', 'private', 'friends_only')",
"bio IS NULL OR LENGTH(bio) <= 2000",
"language IN ('en', 'es', 'fr', 'de', 'it', 'pt', 'ru', 'ja', 'ko', 'zh')"
]
},
"indexes": [
{
"name": "idx_user_profiles_user_id",
"columns": ["user_id"],
"unique": true
},
{
"name": "idx_user_profiles_privacy",
"columns": ["privacy_level"]
},
{
"name": "idx_user_profiles_language",
"columns": ["language"]
}
]
},
"user_sessions": {
"columns": {
"id": {
"type": "varchar",
"length": 128,
"nullable": false,
"primary_key": true
},
"user_id": {
"type": "bigint",
"nullable": false
},
"ip_address": {
"type": "varchar",
"length": 45,
"nullable": true
},
"user_agent": {
"type": "text",
"nullable": true
},
"expires_at": {
"type": "timestamp",
"nullable": false
},
"created_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP"
},
"last_activity": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP"
},
"session_type": {
"type": "varchar",
"length": 20,
"nullable": false,
"default": "web"
},
"is_mobile": {
"type": "boolean",
"nullable": false,
"default": false
}
},
"constraints": {
"primary_key": ["id"],
"unique": [],
"foreign_key": [
{
"columns": ["user_id"],
"references": "users(id)",
"on_delete": "CASCADE"
}
],
"check": [
"session_type IN ('web', 'mobile', 'api', 'admin')"
]
},
"indexes": [
{
"name": "idx_user_sessions_user_id",
"columns": ["user_id"]
},
{
"name": "idx_user_sessions_expires",
"columns": ["expires_at"]
},
{
"name": "idx_user_sessions_type",
"columns": ["session_type"]
}
]
},
"user_preferences": {
"columns": {
"id": {
"type": "bigint",
"nullable": false,
"primary_key": true,
"auto_increment": true
},
"user_id": {
"type": "bigint",
"nullable": false
},
"preference_key": {
"type": "varchar",
"length": 100,
"nullable": false
},
"preference_value": {
"type": "json",
"nullable": true
},
"created_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP"
},
"updated_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP"
}
},
"constraints": {
"primary_key": ["id"],
"unique": [
["user_id", "preference_key"]
],
"foreign_key": [
{
"columns": ["user_id"],
"references": "users(id)",
"on_delete": "CASCADE"
}
],
"check": []
},
"indexes": [
{
"name": "idx_user_preferences_user_key",
"columns": ["user_id", "preference_key"],
"unique": true
}
]
}
},
"views": {
"active_users": {
"definition": "SELECT u.id, u.username, u.email, u.first_name, u.last_name, u.email_verified_at, u.last_login_at FROM users u WHERE u.is_active = true",
"columns": ["id", "username", "email", "first_name", "last_name", "email_verified_at", "last_login_at"]
},
"verified_users": {
"definition": "SELECT u.id, u.username, u.email FROM users u WHERE u.is_active = true AND u.email_verified_at IS NOT NULL",
"columns": ["id", "username", "email"]
}
},
"procedures": [
{
"name": "cleanup_expired_sessions",
"parameters": [],
"definition": "DELETE FROM user_sessions WHERE expires_at < NOW()"
},
{
"name": "get_user_with_profile",
"parameters": ["user_id BIGINT"],
"definition": "SELECT u.*, p.bio, p.avatar_url, p.privacy_level FROM users u LEFT JOIN user_profiles p ON u.id = p.user_id WHERE u.id = user_id"
}
]
}
FILE:assets/database_schema_before.json
{
"schema_version": "1.0",
"database": "user_management",
"tables": {
"users": {
"columns": {
"id": {
"type": "bigint",
"nullable": false,
"primary_key": true,
"auto_increment": true
},
"username": {
"type": "varchar",
"length": 50,
"nullable": false,
"unique": true
},
"email": {
"type": "varchar",
"length": 255,
"nullable": false,
"unique": true
},
"password_hash": {
"type": "varchar",
"length": 255,
"nullable": false
},
"first_name": {
"type": "varchar",
"length": 100,
"nullable": true
},
"last_name": {
"type": "varchar",
"length": 100,
"nullable": true
},
"created_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP"
},
"updated_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP"
},
"is_active": {
"type": "boolean",
"nullable": false,
"default": true
},
"phone": {
"type": "varchar",
"length": 20,
"nullable": true
}
},
"constraints": {
"primary_key": ["id"],
"unique": [
"username",
"email"
],
"foreign_key": [],
"check": [
"email LIKE '%@%'",
"LENGTH(password_hash) >= 60"
]
},
"indexes": [
{
"name": "idx_users_email",
"columns": ["email"],
"unique": true
},
{
"name": "idx_users_username",
"columns": ["username"],
"unique": true
},
{
"name": "idx_users_created_at",
"columns": ["created_at"]
}
]
},
"user_profiles": {
"columns": {
"id": {
"type": "bigint",
"nullable": false,
"primary_key": true,
"auto_increment": true
},
"user_id": {
"type": "bigint",
"nullable": false
},
"bio": {
"type": "varchar",
"length": 255,
"nullable": true
},
"avatar_url": {
"type": "varchar",
"length": 500,
"nullable": true
},
"birth_date": {
"type": "date",
"nullable": true
},
"location": {
"type": "varchar",
"length": 100,
"nullable": true
},
"website": {
"type": "varchar",
"length": 255,
"nullable": true
},
"privacy_level": {
"type": "varchar",
"length": 20,
"nullable": false,
"default": "public"
},
"created_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP"
},
"updated_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP"
}
},
"constraints": {
"primary_key": ["id"],
"unique": [],
"foreign_key": [
{
"columns": ["user_id"],
"references": "users(id)",
"on_delete": "CASCADE"
}
],
"check": [
"privacy_level IN ('public', 'private', 'friends_only')"
]
},
"indexes": [
{
"name": "idx_user_profiles_user_id",
"columns": ["user_id"],
"unique": true
},
{
"name": "idx_user_profiles_privacy",
"columns": ["privacy_level"]
}
]
},
"user_sessions": {
"columns": {
"id": {
"type": "varchar",
"length": 128,
"nullable": false,
"primary_key": true
},
"user_id": {
"type": "bigint",
"nullable": false
},
"ip_address": {
"type": "varchar",
"length": 45,
"nullable": true
},
"user_agent": {
"type": "varchar",
"length": 500,
"nullable": true
},
"expires_at": {
"type": "timestamp",
"nullable": false
},
"created_at": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP"
},
"last_activity": {
"type": "timestamp",
"nullable": false,
"default": "CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP"
}
},
"constraints": {
"primary_key": ["id"],
"unique": [],
"foreign_key": [
{
"columns": ["user_id"],
"references": "users(id)",
"on_delete": "CASCADE"
}
],
"check": []
},
"indexes": [
{
"name": "idx_user_sessions_user_id",
"columns": ["user_id"]
},
{
"name": "idx_user_sessions_expires",
"columns": ["expires_at"]
}
]
}
},
"views": {
"active_users": {
"definition": "SELECT u.id, u.username, u.email, u.first_name, u.last_name FROM users u WHERE u.is_active = true",
"columns": ["id", "username", "email", "first_name", "last_name"]
}
},
"procedures": [
{
"name": "cleanup_expired_sessions",
"parameters": [],
"definition": "DELETE FROM user_sessions WHERE expires_at < NOW()"
}
]
}
FILE:assets/sample_database_migration.json
{
"type": "database",
"pattern": "schema_change",
"source": "PostgreSQL 13 Production Database",
"target": "PostgreSQL 15 Cloud Database",
"description": "Migrate user management system from on-premises PostgreSQL to cloud with schema updates",
"constraints": {
"max_downtime_minutes": 30,
"data_volume_gb": 2500,
"dependencies": [
"user_service_api",
"authentication_service",
"notification_service",
"analytics_pipeline",
"backup_service"
],
"compliance_requirements": [
"GDPR",
"SOX"
],
"special_requirements": [
"zero_data_loss",
"referential_integrity",
"performance_baseline_maintained"
]
},
"tables_to_migrate": [
{
"name": "users",
"row_count": 1500000,
"size_mb": 450,
"critical": true
},
{
"name": "user_profiles",
"row_count": 1500000,
"size_mb": 890,
"critical": true
},
{
"name": "user_sessions",
"row_count": 25000000,
"size_mb": 1200,
"critical": false
},
{
"name": "audit_logs",
"row_count": 50000000,
"size_mb": 2800,
"critical": false
}
],
"schema_changes": [
{
"table": "users",
"changes": [
{
"type": "add_column",
"column": "email_verified_at",
"data_type": "timestamp",
"nullable": true
},
{
"type": "add_column",
"column": "phone_verified_at",
"data_type": "timestamp",
"nullable": true
}
]
},
{
"table": "user_profiles",
"changes": [
{
"type": "modify_column",
"column": "bio",
"old_type": "varchar(255)",
"new_type": "text"
},
{
"type": "add_constraint",
"constraint_type": "check",
"constraint_name": "bio_length_check",
"definition": "LENGTH(bio) <= 2000"
}
]
}
],
"performance_requirements": {
"max_query_response_time_ms": 100,
"concurrent_connections": 500,
"transactions_per_second": 1000
},
"business_continuity": {
"critical_business_hours": {
"start": "08:00",
"end": "18:00",
"timezone": "UTC"
},
"preferred_migration_window": {
"start": "02:00",
"end": "06:00",
"timezone": "UTC"
}
}
}
FILE:assets/sample_service_migration.json
{
"type": "service",
"pattern": "strangler_fig",
"source": "Legacy User Service (Java Spring Boot 2.x)",
"target": "New User Service (Node.js + TypeScript)",
"description": "Migrate legacy user management service to modern microservices architecture",
"constraints": {
"max_downtime_minutes": 0,
"data_volume_gb": 50,
"dependencies": [
"payment_service",
"order_service",
"notification_service",
"analytics_service",
"mobile_app_v1",
"mobile_app_v2",
"web_frontend",
"admin_dashboard"
],
"compliance_requirements": [
"PCI_DSS",
"GDPR"
],
"special_requirements": [
"api_backward_compatibility",
"session_continuity",
"rate_limit_preservation"
]
},
"service_details": {
"legacy_service": {
"endpoints": [
"GET /api/v1/users/{id}",
"POST /api/v1/users",
"PUT /api/v1/users/{id}",
"DELETE /api/v1/users/{id}",
"GET /api/v1/users/{id}/profile",
"PUT /api/v1/users/{id}/profile",
"POST /api/v1/users/{id}/verify-email",
"POST /api/v1/users/login",
"POST /api/v1/users/logout"
],
"current_load": {
"requests_per_second": 850,
"peak_requests_per_second": 2000,
"average_response_time_ms": 120,
"p95_response_time_ms": 300
},
"infrastructure": {
"instances": 4,
"cpu_cores_per_instance": 4,
"memory_gb_per_instance": 8,
"load_balancer": "AWS ELB Classic"
}
},
"new_service": {
"endpoints": [
"GET /api/v2/users/{id}",
"POST /api/v2/users",
"PUT /api/v2/users/{id}",
"DELETE /api/v2/users/{id}",
"GET /api/v2/users/{id}/profile",
"PUT /api/v2/users/{id}/profile",
"POST /api/v2/users/{id}/verify-email",
"POST /api/v2/users/{id}/verify-phone",
"POST /api/v2/auth/login",
"POST /api/v2/auth/logout",
"POST /api/v2/auth/refresh"
],
"target_performance": {
"requests_per_second": 1500,
"peak_requests_per_second": 3000,
"average_response_time_ms": 80,
"p95_response_time_ms": 200
},
"infrastructure": {
"container_platform": "Kubernetes",
"initial_replicas": 3,
"max_replicas": 10,
"cpu_request_millicores": 500,
"cpu_limit_millicores": 1000,
"memory_request_mb": 512,
"memory_limit_mb": 1024,
"load_balancer": "AWS ALB"
}
}
},
"migration_phases": [
{
"phase": "preparation",
"description": "Deploy new service and configure routing",
"estimated_duration_hours": 8
},
{
"phase": "intercept",
"description": "Configure API gateway to route to new service",
"estimated_duration_hours": 2
},
{
"phase": "gradual_migration",
"description": "Gradually increase traffic to new service",
"estimated_duration_hours": 48
},
{
"phase": "validation",
"description": "Validate new service performance and functionality",
"estimated_duration_hours": 24
},
{
"phase": "decommission",
"description": "Remove legacy service after validation",
"estimated_duration_hours": 4
}
],
"feature_flags": [
{
"name": "enable_new_user_service",
"description": "Route user service requests to new implementation",
"initial_percentage": 5,
"rollout_schedule": [
{"percentage": 5, "duration_hours": 24},
{"percentage": 25, "duration_hours": 24},
{"percentage": 50, "duration_hours": 24},
{"percentage": 100, "duration_hours": 0}
]
},
{
"name": "enable_new_auth_endpoints",
"description": "Enable new authentication endpoints",
"initial_percentage": 0,
"rollout_schedule": [
{"percentage": 10, "duration_hours": 12},
{"percentage": 50, "duration_hours": 12},
{"percentage": 100, "duration_hours": 0}
]
}
],
"monitoring": {
"critical_metrics": [
"request_rate",
"error_rate",
"response_time_p95",
"response_time_p99",
"cpu_utilization",
"memory_utilization",
"database_connection_pool"
],
"alert_thresholds": {
"error_rate": 0.05,
"response_time_p95": 250,
"cpu_utilization": 0.80,
"memory_utilization": 0.85
}
},
"rollback_triggers": [
{
"metric": "error_rate",
"threshold": 0.10,
"duration_minutes": 5,
"action": "automatic_rollback"
},
{
"metric": "response_time_p95",
"threshold": 500,
"duration_minutes": 10,
"action": "alert_team"
},
{
"metric": "cpu_utilization",
"threshold": 0.95,
"duration_minutes": 5,
"action": "scale_up"
}
]
}
FILE:expected_outputs/rollback_runbook.json
{
"runbook_id": "rb_921c0bca",
"migration_id": "23a52ed1507f",
"created_at": "2026-02-16T13:47:31.108500",
"rollback_phases": [
{
"phase_name": "rollback_cleanup",
"description": "Rollback changes made during cleanup phase",
"urgency_level": "medium",
"estimated_duration_minutes": 570,
"prerequisites": [
"Incident commander assigned and briefed",
"All team members notified of rollback initiation",
"Monitoring systems confirmed operational",
"Backup systems verified and accessible"
],
"steps": [
{
"step_id": "rb_validate_0_final",
"name": "Validate rollback completion",
"description": "Comprehensive validation that cleanup rollback completed successfully",
"script_type": "manual",
"script_content": "Execute validation checklist for this phase",
"estimated_duration_minutes": 10,
"dependencies": [],
"validation_commands": [
"SELECT COUNT(*) FROM {table_name};",
"SELECT COUNT(*) FROM information_schema.tables WHERE table_name = '{table_name}';",
"SELECT COUNT(*) FROM information_schema.columns WHERE table_name = '{table_name}' AND column_name = '{column_name}';",
"SELECT COUNT(DISTINCT {primary_key}) FROM {table_name};",
"SELECT MAX({timestamp_column}) FROM {table_name};"
],
"success_criteria": [
"cleanup fully rolled back",
"All validation checks pass"
],
"failure_escalation": "Investigate cleanup rollback failures",
"rollback_order": 99
}
],
"validation_checkpoints": [
"cleanup rollback steps completed",
"System health checks passing",
"No critical errors in logs",
"Key metrics within acceptable ranges",
"Validation command passed: SELECT COUNT(*) FROM {table_name};...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH..."
],
"communication_requirements": [
"Notify incident commander of phase start/completion",
"Update rollback status dashboard",
"Log all actions and decisions"
],
"risk_level": "medium"
},
{
"phase_name": "rollback_contract",
"description": "Rollback changes made during contract phase",
"urgency_level": "medium",
"estimated_duration_minutes": 570,
"prerequisites": [
"Incident commander assigned and briefed",
"All team members notified of rollback initiation",
"Monitoring systems confirmed operational",
"Backup systems verified and accessible",
"Previous rollback phase completed successfully"
],
"steps": [
{
"step_id": "rb_validate_1_final",
"name": "Validate rollback completion",
"description": "Comprehensive validation that contract rollback completed successfully",
"script_type": "manual",
"script_content": "Execute validation checklist for this phase",
"estimated_duration_minutes": 10,
"dependencies": [],
"validation_commands": [
"SELECT COUNT(*) FROM {table_name};",
"SELECT COUNT(*) FROM information_schema.tables WHERE table_name = '{table_name}';",
"SELECT COUNT(*) FROM information_schema.columns WHERE table_name = '{table_name}' AND column_name = '{column_name}';",
"SELECT COUNT(DISTINCT {primary_key}) FROM {table_name};",
"SELECT MAX({timestamp_column}) FROM {table_name};"
],
"success_criteria": [
"contract fully rolled back",
"All validation checks pass"
],
"failure_escalation": "Investigate contract rollback failures",
"rollback_order": 99
}
],
"validation_checkpoints": [
"contract rollback steps completed",
"System health checks passing",
"No critical errors in logs",
"Key metrics within acceptable ranges",
"Validation command passed: SELECT COUNT(*) FROM {table_name};...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH..."
],
"communication_requirements": [
"Notify incident commander of phase start/completion",
"Update rollback status dashboard",
"Log all actions and decisions"
],
"risk_level": "medium"
},
{
"phase_name": "rollback_migrate",
"description": "Rollback changes made during migrate phase",
"urgency_level": "medium",
"estimated_duration_minutes": 570,
"prerequisites": [
"Incident commander assigned and briefed",
"All team members notified of rollback initiation",
"Monitoring systems confirmed operational",
"Backup systems verified and accessible",
"Previous rollback phase completed successfully"
],
"steps": [
{
"step_id": "rb_validate_2_final",
"name": "Validate rollback completion",
"description": "Comprehensive validation that migrate rollback completed successfully",
"script_type": "manual",
"script_content": "Execute validation checklist for this phase",
"estimated_duration_minutes": 10,
"dependencies": [],
"validation_commands": [
"SELECT COUNT(*) FROM {table_name};",
"SELECT COUNT(*) FROM information_schema.tables WHERE table_name = '{table_name}';",
"SELECT COUNT(*) FROM information_schema.columns WHERE table_name = '{table_name}' AND column_name = '{column_name}';",
"SELECT COUNT(DISTINCT {primary_key}) FROM {table_name};",
"SELECT MAX({timestamp_column}) FROM {table_name};"
],
"success_criteria": [
"migrate fully rolled back",
"All validation checks pass"
],
"failure_escalation": "Investigate migrate rollback failures",
"rollback_order": 99
}
],
"validation_checkpoints": [
"migrate rollback steps completed",
"System health checks passing",
"No critical errors in logs",
"Key metrics within acceptable ranges",
"Validation command passed: SELECT COUNT(*) FROM {table_name};...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH..."
],
"communication_requirements": [
"Notify incident commander of phase start/completion",
"Update rollback status dashboard",
"Log all actions and decisions"
],
"risk_level": "medium"
},
{
"phase_name": "rollback_expand",
"description": "Rollback changes made during expand phase",
"urgency_level": "medium",
"estimated_duration_minutes": 570,
"prerequisites": [
"Incident commander assigned and briefed",
"All team members notified of rollback initiation",
"Monitoring systems confirmed operational",
"Backup systems verified and accessible",
"Previous rollback phase completed successfully"
],
"steps": [
{
"step_id": "rb_validate_3_final",
"name": "Validate rollback completion",
"description": "Comprehensive validation that expand rollback completed successfully",
"script_type": "manual",
"script_content": "Execute validation checklist for this phase",
"estimated_duration_minutes": 10,
"dependencies": [],
"validation_commands": [
"SELECT COUNT(*) FROM {table_name};",
"SELECT COUNT(*) FROM information_schema.tables WHERE table_name = '{table_name}';",
"SELECT COUNT(*) FROM information_schema.columns WHERE table_name = '{table_name}' AND column_name = '{column_name}';",
"SELECT COUNT(DISTINCT {primary_key}) FROM {table_name};",
"SELECT MAX({timestamp_column}) FROM {table_name};"
],
"success_criteria": [
"expand fully rolled back",
"All validation checks pass"
],
"failure_escalation": "Investigate expand rollback failures",
"rollback_order": 99
}
],
"validation_checkpoints": [
"expand rollback steps completed",
"System health checks passing",
"No critical errors in logs",
"Key metrics within acceptable ranges",
"Validation command passed: SELECT COUNT(*) FROM {table_name};...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH..."
],
"communication_requirements": [
"Notify incident commander of phase start/completion",
"Update rollback status dashboard",
"Log all actions and decisions"
],
"risk_level": "medium"
},
{
"phase_name": "rollback_preparation",
"description": "Rollback changes made during preparation phase",
"urgency_level": "medium",
"estimated_duration_minutes": 570,
"prerequisites": [
"Incident commander assigned and briefed",
"All team members notified of rollback initiation",
"Monitoring systems confirmed operational",
"Backup systems verified and accessible",
"Previous rollback phase completed successfully"
],
"steps": [
{
"step_id": "rb_schema_4_01",
"name": "Drop migration artifacts",
"description": "Remove temporary migration tables and procedures",
"script_type": "sql",
"script_content": "-- Drop migration artifacts\nDROP TABLE IF EXISTS migration_log;\nDROP PROCEDURE IF EXISTS migrate_data();",
"estimated_duration_minutes": 5,
"dependencies": [],
"validation_commands": [
"SELECT COUNT(*) FROM information_schema.tables WHERE table_name LIKE '%migration%';"
],
"success_criteria": [
"No migration artifacts remain"
],
"failure_escalation": "Manual cleanup required",
"rollback_order": 1
},
{
"step_id": "rb_validate_4_final",
"name": "Validate rollback completion",
"description": "Comprehensive validation that preparation rollback completed successfully",
"script_type": "manual",
"script_content": "Execute validation checklist for this phase",
"estimated_duration_minutes": 10,
"dependencies": [
"rb_schema_4_01"
],
"validation_commands": [
"SELECT COUNT(*) FROM {table_name};",
"SELECT COUNT(*) FROM information_schema.tables WHERE table_name = '{table_name}';",
"SELECT COUNT(*) FROM information_schema.columns WHERE table_name = '{table_name}' AND column_name = '{column_name}';",
"SELECT COUNT(DISTINCT {primary_key}) FROM {table_name};",
"SELECT MAX({timestamp_column}) FROM {table_name};"
],
"success_criteria": [
"preparation fully rolled back",
"All validation checks pass"
],
"failure_escalation": "Investigate preparation rollback failures",
"rollback_order": 99
}
],
"validation_checkpoints": [
"preparation rollback steps completed",
"System health checks passing",
"No critical errors in logs",
"Key metrics within acceptable ranges",
"Validation command passed: SELECT COUNT(*) FROM {table_name};...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...",
"Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH..."
],
"communication_requirements": [
"Notify incident commander of phase start/completion",
"Update rollback status dashboard",
"Log all actions and decisions"
],
"risk_level": "medium"
}
],
"trigger_conditions": [
{
"trigger_id": "error_rate_spike",
"name": "Error Rate Spike",
"condition": "error_rate > baseline * 5 for 5 minutes",
"metric_threshold": {
"metric": "error_rate",
"operator": "greater_than",
"value": "baseline_error_rate * 5",
"duration_minutes": 5
},
"evaluation_window_minutes": 5,
"auto_execute": true,
"escalation_contacts": [
"on_call_engineer",
"migration_lead"
]
},
{
"trigger_id": "response_time_degradation",
"name": "Response Time Degradation",
"condition": "p95_response_time > baseline * 3 for 10 minutes",
"metric_threshold": {
"metric": "p95_response_time",
"operator": "greater_than",
"value": "baseline_p95 * 3",
"duration_minutes": 10
},
"evaluation_window_minutes": 10,
"auto_execute": false,
"escalation_contacts": [
"performance_team",
"migration_lead"
]
},
{
"trigger_id": "availability_drop",
"name": "Service Availability Drop",
"condition": "availability < 95% for 2 minutes",
"metric_threshold": {
"metric": "availability",
"operator": "less_than",
"value": 0.95,
"duration_minutes": 2
},
"evaluation_window_minutes": 2,
"auto_execute": true,
"escalation_contacts": [
"sre_team",
"incident_commander"
]
},
{
"trigger_id": "data_integrity_failure",
"name": "Data Integrity Check Failure",
"condition": "data_validation_failures > 0",
"metric_threshold": {
"metric": "data_validation_failures",
"operator": "greater_than",
"value": 0,
"duration_minutes": 1
},
"evaluation_window_minutes": 1,
"auto_execute": true,
"escalation_contacts": [
"dba_team",
"data_team"
]
},
{
"trigger_id": "migration_progress_stalled",
"name": "Migration Progress Stalled",
"condition": "migration_progress unchanged for 30 minutes",
"metric_threshold": {
"metric": "migration_progress_rate",
"operator": "equals",
"value": 0,
"duration_minutes": 30
},
"evaluation_window_minutes": 30,
"auto_execute": false,
"escalation_contacts": [
"migration_team",
"dba_team"
]
}
],
"data_recovery_plan": {
"recovery_method": "point_in_time",
"backup_location": "/backups/pre_migration_{migration_id}_{timestamp}.sql",
"recovery_scripts": [
"pg_restore -d production -c /backups/pre_migration_backup.sql",
"SELECT pg_create_restore_point('rollback_point');",
"VACUUM ANALYZE; -- Refresh statistics after restore"
],
"data_validation_queries": [
"SELECT COUNT(*) FROM critical_business_table;",
"SELECT MAX(created_at) FROM audit_log;",
"SELECT COUNT(DISTINCT user_id) FROM user_sessions;",
"SELECT SUM(amount) FROM financial_transactions WHERE date = CURRENT_DATE;"
],
"estimated_recovery_time_minutes": 45,
"recovery_dependencies": [
"database_instance_running",
"backup_file_accessible"
]
},
"communication_templates": [
{
"template_type": "rollback_start",
"audience": "technical",
"subject": "ROLLBACK INITIATED: {migration_name}",
"body": "Team,\n\nWe have initiated rollback for migration: {migration_name}\nRollback ID: {rollback_id}\nStart Time: {start_time}\nEstimated Duration: {estimated_duration}\n\nReason: {rollback_reason}\n\nCurrent Status: Rolling back phase {current_phase}\n\nNext Updates: Every 15 minutes or upon phase completion\n\nActions Required:\n- Monitor system health dashboards\n- Stand by for escalation if needed\n- Do not make manual changes during rollback\n\nIncident Commander: {incident_commander}\n",
"urgency": "medium",
"delivery_methods": [
"email",
"slack"
]
},
{
"template_type": "rollback_start",
"audience": "business",
"subject": "System Rollback In Progress - {system_name}",
"body": "Business Stakeholders,\n\nWe are currently performing a planned rollback of the {system_name} migration due to {rollback_reason}.\n\nImpact: {business_impact}\nExpected Resolution: {estimated_completion_time}\nAffected Services: {affected_services}\n\nWe will provide updates every 30 minutes.\n\nContact: {business_contact}\n",
"urgency": "medium",
"delivery_methods": [
"email"
]
},
{
"template_type": "rollback_start",
"audience": "executive",
"subject": "EXEC ALERT: Critical System Rollback - {system_name}",
"body": "Executive Team,\n\nA critical rollback is in progress for {system_name}.\n\nSummary:\n- Rollback Reason: {rollback_reason}\n- Business Impact: {business_impact}\n- Expected Resolution: {estimated_completion_time}\n- Customer Impact: {customer_impact}\n\nWe are following established procedures and will update hourly.\n\nEscalation: {escalation_contact}\n",
"urgency": "high",
"delivery_methods": [
"email"
]
},
{
"template_type": "rollback_complete",
"audience": "technical",
"subject": "ROLLBACK COMPLETED: {migration_name}",
"body": "Team,\n\nRollback has been successfully completed for migration: {migration_name}\n\nSummary:\n- Start Time: {start_time}\n- End Time: {end_time}\n- Duration: {actual_duration}\n- Phases Completed: {completed_phases}\n\nValidation Results:\n{validation_results}\n\nSystem Status: {system_status}\n\nNext Steps:\n- Continue monitoring for 24 hours\n- Post-rollback review scheduled for {review_date}\n- Root cause analysis to begin\n\nAll clear to resume normal operations.\n\nIncident Commander: {incident_commander}\n",
"urgency": "medium",
"delivery_methods": [
"email",
"slack"
]
},
{
"template_type": "emergency_escalation",
"audience": "executive",
"subject": "CRITICAL: Rollback Emergency - {migration_name}",
"body": "CRITICAL SITUATION - IMMEDIATE ATTENTION REQUIRED\n\nMigration: {migration_name}\nIssue: Rollback procedure has encountered critical failures\n\nCurrent Status: {current_status}\nFailed Components: {failed_components}\nBusiness Impact: {business_impact}\nCustomer Impact: {customer_impact}\n\nImmediate Actions:\n1. Emergency response team activated\n2. {emergency_action_1}\n3. {emergency_action_2}\n\nWar Room: {war_room_location}\nBridge Line: {conference_bridge}\n\nNext Update: {next_update_time}\n\nIncident Commander: {incident_commander}\nExecutive On-Call: {executive_on_call}\n",
"urgency": "emergency",
"delivery_methods": [
"email",
"sms",
"phone_call"
]
}
],
"escalation_matrix": {
"level_1": {
"trigger": "Single component failure",
"response_time_minutes": 5,
"contacts": [
"on_call_engineer",
"migration_lead"
],
"actions": [
"Investigate issue",
"Attempt automated remediation",
"Monitor closely"
]
},
"level_2": {
"trigger": "Multiple component failures or single critical failure",
"response_time_minutes": 2,
"contacts": [
"senior_engineer",
"team_lead",
"devops_lead"
],
"actions": [
"Initiate rollback",
"Establish war room",
"Notify stakeholders"
]
},
"level_3": {
"trigger": "System-wide failure or data corruption",
"response_time_minutes": 1,
"contacts": [
"engineering_manager",
"cto",
"incident_commander"
],
"actions": [
"Emergency rollback",
"All hands on deck",
"Executive notification"
]
},
"emergency": {
"trigger": "Business-critical failure with customer impact",
"response_time_minutes": 0,
"contacts": [
"ceo",
"cto",
"head_of_operations"
],
"actions": [
"Emergency procedures",
"Customer communication",
"Media preparation if needed"
]
}
},
"validation_checklist": [
"Verify system is responding to health checks",
"Confirm error rates are within normal parameters",
"Validate response times meet SLA requirements",
"Check all critical business processes are functioning",
"Verify monitoring and alerting systems are operational",
"Confirm no data corruption has occurred",
"Validate security controls are functioning properly",
"Check backup systems are working correctly",
"Verify integration points with downstream systems",
"Confirm user authentication and authorization working",
"Validate database schema matches expected state",
"Confirm referential integrity constraints",
"Check database performance metrics",
"Verify data consistency across related tables",
"Validate indexes and statistics are optimal",
"Confirm transaction logs are clean",
"Check database connections and connection pooling"
],
"post_rollback_procedures": [
"Monitor system stability for 24-48 hours post-rollback",
"Conduct thorough post-rollback testing of all critical paths",
"Review and analyze rollback metrics and timing",
"Document lessons learned and rollback procedure improvements",
"Schedule post-mortem meeting with all stakeholders",
"Update rollback procedures based on actual experience",
"Communicate rollback completion to all stakeholders",
"Archive rollback logs and artifacts for future reference",
"Review and update monitoring thresholds if needed",
"Plan for next migration attempt with improved procedures",
"Conduct security review to ensure no vulnerabilities introduced",
"Update disaster recovery procedures if affected by rollback",
"Review capacity planning based on rollback resource usage",
"Update documentation with rollback experience and timings"
],
"emergency_contacts": [
{
"role": "Incident Commander",
"name": "TBD - Assigned during migration",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "incident.commander@company.com",
"backup_contact": "backup.commander@company.com"
},
{
"role": "Technical Lead",
"name": "TBD - Migration technical owner",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "tech.lead@company.com",
"backup_contact": "senior.engineer@company.com"
},
{
"role": "Business Owner",
"name": "TBD - Business stakeholder",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "business.owner@company.com",
"backup_contact": "product.manager@company.com"
},
{
"role": "On-Call Engineer",
"name": "Current on-call rotation",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "oncall@company.com",
"backup_contact": "backup.oncall@company.com"
},
{
"role": "Executive Escalation",
"name": "CTO/VP Engineering",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "cto@company.com",
"backup_contact": "vp.engineering@company.com"
}
]
}
FILE:expected_outputs/rollback_runbook.txt
================================================================================
ROLLBACK RUNBOOK: rb_921c0bca
================================================================================
Migration ID: 23a52ed1507f
Created: 2026-02-16T13:47:31.108500
EMERGENCY CONTACTS
----------------------------------------
Incident Commander: TBD - Assigned during migration
Phone: +1-XXX-XXX-XXXX
Email: incident.commander@company.com
Backup: backup.commander@company.com
Technical Lead: TBD - Migration technical owner
Phone: +1-XXX-XXX-XXXX
Email: tech.lead@company.com
Backup: senior.engineer@company.com
Business Owner: TBD - Business stakeholder
Phone: +1-XXX-XXX-XXXX
Email: business.owner@company.com
Backup: product.manager@company.com
On-Call Engineer: Current on-call rotation
Phone: +1-XXX-XXX-XXXX
Email: oncall@company.com
Backup: backup.oncall@company.com
Executive Escalation: CTO/VP Engineering
Phone: +1-XXX-XXX-XXXX
Email: cto@company.com
Backup: vp.engineering@company.com
ESCALATION MATRIX
----------------------------------------
LEVEL_1:
Trigger: Single component failure
Response Time: 5 minutes
Contacts: on_call_engineer, migration_lead
Actions: Investigate issue, Attempt automated remediation, Monitor closely
LEVEL_2:
Trigger: Multiple component failures or single critical failure
Response Time: 2 minutes
Contacts: senior_engineer, team_lead, devops_lead
Actions: Initiate rollback, Establish war room, Notify stakeholders
LEVEL_3:
Trigger: System-wide failure or data corruption
Response Time: 1 minutes
Contacts: engineering_manager, cto, incident_commander
Actions: Emergency rollback, All hands on deck, Executive notification
EMERGENCY:
Trigger: Business-critical failure with customer impact
Response Time: 0 minutes
Contacts: ceo, cto, head_of_operations
Actions: Emergency procedures, Customer communication, Media preparation if needed
AUTOMATIC ROLLBACK TRIGGERS
----------------------------------------
• Error Rate Spike
Condition: error_rate > baseline * 5 for 5 minutes
Auto-Execute: Yes
Evaluation Window: 5 minutes
Contacts: on_call_engineer, migration_lead
• Response Time Degradation
Condition: p95_response_time > baseline * 3 for 10 minutes
Auto-Execute: No
Evaluation Window: 10 minutes
Contacts: performance_team, migration_lead
• Service Availability Drop
Condition: availability < 95% for 2 minutes
Auto-Execute: Yes
Evaluation Window: 2 minutes
Contacts: sre_team, incident_commander
• Data Integrity Check Failure
Condition: data_validation_failures > 0
Auto-Execute: Yes
Evaluation Window: 1 minutes
Contacts: dba_team, data_team
• Migration Progress Stalled
Condition: migration_progress unchanged for 30 minutes
Auto-Execute: No
Evaluation Window: 30 minutes
Contacts: migration_team, dba_team
ROLLBACK PHASES
----------------------------------------
1. ROLLBACK_CLEANUP
Description: Rollback changes made during cleanup phase
Urgency: MEDIUM
Duration: 570 minutes
Risk Level: MEDIUM
Prerequisites:
✓ Incident commander assigned and briefed
✓ All team members notified of rollback initiation
✓ Monitoring systems confirmed operational
✓ Backup systems verified and accessible
Steps:
99. Validate rollback completion
Duration: 10 min
Type: manual
Success Criteria: cleanup fully rolled back, All validation checks pass
Validation Checkpoints:
☐ cleanup rollback steps completed
☐ System health checks passing
☐ No critical errors in logs
☐ Key metrics within acceptable ranges
☐ Validation command passed: SELECT COUNT(*) FROM {table_name};...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH...
2. ROLLBACK_CONTRACT
Description: Rollback changes made during contract phase
Urgency: MEDIUM
Duration: 570 minutes
Risk Level: MEDIUM
Prerequisites:
✓ Incident commander assigned and briefed
✓ All team members notified of rollback initiation
✓ Monitoring systems confirmed operational
✓ Backup systems verified and accessible
✓ Previous rollback phase completed successfully
Steps:
99. Validate rollback completion
Duration: 10 min
Type: manual
Success Criteria: contract fully rolled back, All validation checks pass
Validation Checkpoints:
☐ contract rollback steps completed
☐ System health checks passing
☐ No critical errors in logs
☐ Key metrics within acceptable ranges
☐ Validation command passed: SELECT COUNT(*) FROM {table_name};...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH...
3. ROLLBACK_MIGRATE
Description: Rollback changes made during migrate phase
Urgency: MEDIUM
Duration: 570 minutes
Risk Level: MEDIUM
Prerequisites:
✓ Incident commander assigned and briefed
✓ All team members notified of rollback initiation
✓ Monitoring systems confirmed operational
✓ Backup systems verified and accessible
✓ Previous rollback phase completed successfully
Steps:
99. Validate rollback completion
Duration: 10 min
Type: manual
Success Criteria: migrate fully rolled back, All validation checks pass
Validation Checkpoints:
☐ migrate rollback steps completed
☐ System health checks passing
☐ No critical errors in logs
☐ Key metrics within acceptable ranges
☐ Validation command passed: SELECT COUNT(*) FROM {table_name};...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH...
4. ROLLBACK_EXPAND
Description: Rollback changes made during expand phase
Urgency: MEDIUM
Duration: 570 minutes
Risk Level: MEDIUM
Prerequisites:
✓ Incident commander assigned and briefed
✓ All team members notified of rollback initiation
✓ Monitoring systems confirmed operational
✓ Backup systems verified and accessible
✓ Previous rollback phase completed successfully
Steps:
99. Validate rollback completion
Duration: 10 min
Type: manual
Success Criteria: expand fully rolled back, All validation checks pass
Validation Checkpoints:
☐ expand rollback steps completed
☐ System health checks passing
☐ No critical errors in logs
☐ Key metrics within acceptable ranges
☐ Validation command passed: SELECT COUNT(*) FROM {table_name};...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH...
5. ROLLBACK_PREPARATION
Description: Rollback changes made during preparation phase
Urgency: MEDIUM
Duration: 570 minutes
Risk Level: MEDIUM
Prerequisites:
✓ Incident commander assigned and briefed
✓ All team members notified of rollback initiation
✓ Monitoring systems confirmed operational
✓ Backup systems verified and accessible
✓ Previous rollback phase completed successfully
Steps:
1. Drop migration artifacts
Duration: 5 min
Type: sql
Script:
-- Drop migration artifacts
DROP TABLE IF EXISTS migration_log;
DROP PROCEDURE IF EXISTS migrate_data();
Success Criteria: No migration artifacts remain
99. Validate rollback completion
Duration: 10 min
Type: manual
Success Criteria: preparation fully rolled back, All validation checks pass
Validation Checkpoints:
☐ preparation rollback steps completed
☐ System health checks passing
☐ No critical errors in logs
☐ Key metrics within acceptable ranges
☐ Validation command passed: SELECT COUNT(*) FROM {table_name};...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.tables WHE...
☐ Validation command passed: SELECT COUNT(*) FROM information_schema.columns WH...
DATA RECOVERY PLAN
----------------------------------------
Recovery Method: point_in_time
Backup Location: /backups/pre_migration_{migration_id}_{timestamp}.sql
Estimated Recovery Time: 45 minutes
Recovery Scripts:
• pg_restore -d production -c /backups/pre_migration_backup.sql
• SELECT pg_create_restore_point('rollback_point');
• VACUUM ANALYZE; -- Refresh statistics after restore
Validation Queries:
• SELECT COUNT(*) FROM critical_business_table;
• SELECT MAX(created_at) FROM audit_log;
• SELECT COUNT(DISTINCT user_id) FROM user_sessions;
• SELECT SUM(amount) FROM financial_transactions WHERE date = CURRENT_DATE;
POST-ROLLBACK VALIDATION CHECKLIST
----------------------------------------
1. ☐ Verify system is responding to health checks
2. ☐ Confirm error rates are within normal parameters
3. ☐ Validate response times meet SLA requirements
4. ☐ Check all critical business processes are functioning
5. ☐ Verify monitoring and alerting systems are operational
6. ☐ Confirm no data corruption has occurred
7. ☐ Validate security controls are functioning properly
8. ☐ Check backup systems are working correctly
9. ☐ Verify integration points with downstream systems
10. ☐ Confirm user authentication and authorization working
11. ☐ Validate database schema matches expected state
12. ☐ Confirm referential integrity constraints
13. ☐ Check database performance metrics
14. ☐ Verify data consistency across related tables
15. ☐ Validate indexes and statistics are optimal
16. ☐ Confirm transaction logs are clean
17. ☐ Check database connections and connection pooling
POST-ROLLBACK PROCEDURES
----------------------------------------
1. Monitor system stability for 24-48 hours post-rollback
2. Conduct thorough post-rollback testing of all critical paths
3. Review and analyze rollback metrics and timing
4. Document lessons learned and rollback procedure improvements
5. Schedule post-mortem meeting with all stakeholders
6. Update rollback procedures based on actual experience
7. Communicate rollback completion to all stakeholders
8. Archive rollback logs and artifacts for future reference
9. Review and update monitoring thresholds if needed
10. Plan for next migration attempt with improved procedures
11. Conduct security review to ensure no vulnerabilities introduced
12. Update disaster recovery procedures if affected by rollback
13. Review capacity planning based on rollback resource usage
14. Update documentation with rollback experience and timings
FILE:expected_outputs/sample_database_migration_plan.json
{
"migration_id": "23a52ed1507f",
"source_system": "PostgreSQL 13 Production Database",
"target_system": "PostgreSQL 15 Cloud Database",
"migration_type": "database",
"complexity": "critical",
"estimated_duration_hours": 95,
"phases": [
{
"name": "preparation",
"description": "Prepare systems and teams for migration",
"duration_hours": 19,
"dependencies": [],
"validation_criteria": [
"All backups completed successfully",
"Monitoring systems operational",
"Team members briefed and ready",
"Rollback procedures tested"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Backup source system",
"Set up monitoring and alerting",
"Prepare rollback procedures",
"Communicate migration timeline",
"Validate prerequisites"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
},
{
"name": "expand",
"description": "Execute expand phase",
"duration_hours": 19,
"dependencies": [
"preparation"
],
"validation_criteria": [
"Expand phase completed successfully"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Complete expand activities"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
},
{
"name": "migrate",
"description": "Execute migrate phase",
"duration_hours": 19,
"dependencies": [
"expand"
],
"validation_criteria": [
"Migrate phase completed successfully"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Complete migrate activities"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
},
{
"name": "contract",
"description": "Execute contract phase",
"duration_hours": 19,
"dependencies": [
"migrate"
],
"validation_criteria": [
"Contract phase completed successfully"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Complete contract activities"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
},
{
"name": "cleanup",
"description": "Execute cleanup phase",
"duration_hours": 19,
"dependencies": [
"contract"
],
"validation_criteria": [
"Cleanup phase completed successfully"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Complete cleanup activities"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
}
],
"risks": [
{
"category": "technical",
"description": "Data corruption during migration",
"probability": "low",
"impact": "critical",
"severity": "high",
"mitigation": "Implement comprehensive backup and validation procedures",
"owner": "DBA Team"
},
{
"category": "technical",
"description": "Extended downtime due to migration complexity",
"probability": "medium",
"impact": "high",
"severity": "high",
"mitigation": "Use blue-green deployment and phased migration approach",
"owner": "DevOps Team"
},
{
"category": "business",
"description": "Business process disruption",
"probability": "medium",
"impact": "high",
"severity": "high",
"mitigation": "Communicate timeline and provide alternate workflows",
"owner": "Business Owner"
},
{
"category": "operational",
"description": "Insufficient rollback testing",
"probability": "high",
"impact": "critical",
"severity": "critical",
"mitigation": "Execute full rollback procedures in staging environment",
"owner": "QA Team"
},
{
"category": "business",
"description": "Zero-downtime requirement increases complexity",
"probability": "high",
"impact": "medium",
"severity": "high",
"mitigation": "Implement blue-green deployment or rolling update strategy",
"owner": "DevOps Team"
},
{
"category": "compliance",
"description": "Regulatory compliance requirements",
"probability": "medium",
"impact": "high",
"severity": "high",
"mitigation": "Ensure all compliance checks are integrated into migration process",
"owner": "Compliance Team"
}
],
"success_criteria": [
"All data successfully migrated with 100% integrity",
"System performance meets or exceeds baseline",
"All business processes functioning normally",
"No critical security vulnerabilities introduced",
"Stakeholder acceptance criteria met",
"Documentation and runbooks updated"
],
"rollback_plan": {
"rollback_phases": [
{
"phase": "cleanup",
"rollback_actions": [
"Revert cleanup changes",
"Restore pre-cleanup state",
"Validate cleanup rollback success"
],
"validation_criteria": [
"System restored to pre-cleanup state",
"All cleanup changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 285
},
{
"phase": "contract",
"rollback_actions": [
"Revert contract changes",
"Restore pre-contract state",
"Validate contract rollback success"
],
"validation_criteria": [
"System restored to pre-contract state",
"All contract changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 285
},
{
"phase": "migrate",
"rollback_actions": [
"Revert migrate changes",
"Restore pre-migrate state",
"Validate migrate rollback success"
],
"validation_criteria": [
"System restored to pre-migrate state",
"All migrate changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 285
},
{
"phase": "expand",
"rollback_actions": [
"Revert expand changes",
"Restore pre-expand state",
"Validate expand rollback success"
],
"validation_criteria": [
"System restored to pre-expand state",
"All expand changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 285
},
{
"phase": "preparation",
"rollback_actions": [
"Revert preparation changes",
"Restore pre-preparation state",
"Validate preparation rollback success"
],
"validation_criteria": [
"System restored to pre-preparation state",
"All preparation changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 285
}
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Migration timeline exceeded by > 50%",
"Business-critical functionality unavailable",
"Security breach detected",
"Stakeholder decision to abort"
],
"rollback_decision_matrix": {
"low_severity": "Continue with monitoring",
"medium_severity": "Assess and decide within 15 minutes",
"high_severity": "Immediate rollback initiation",
"critical_severity": "Emergency rollback - all hands"
},
"rollback_contacts": [
"Migration Lead",
"Technical Lead",
"Business Owner",
"On-call Engineer"
]
},
"stakeholders": [
"Business Owner",
"Technical Lead",
"DevOps Team",
"QA Team",
"Security Team",
"End Users"
],
"created_at": "2026-02-16T13:47:23.704502"
}
FILE:expected_outputs/sample_database_migration_plan.txt
================================================================================
MIGRATION PLAN: 23a52ed1507f
================================================================================
Source System: PostgreSQL 13 Production Database
Target System: PostgreSQL 15 Cloud Database
Migration Type: DATABASE
Complexity Level: CRITICAL
Estimated Duration: 95 hours (4.0 days)
Created: 2026-02-16T13:47:23.704502
MIGRATION PHASES
----------------------------------------
1. PREPARATION (19h)
Description: Prepare systems and teams for migration
Risk Level: MEDIUM
Tasks:
• Backup source system
• Set up monitoring and alerting
• Prepare rollback procedures
• Communicate migration timeline
• Validate prerequisites
Success Criteria:
✓ All backups completed successfully
✓ Monitoring systems operational
✓ Team members briefed and ready
✓ Rollback procedures tested
2. EXPAND (19h)
Description: Execute expand phase
Risk Level: MEDIUM
Dependencies: preparation
Tasks:
• Complete expand activities
Success Criteria:
✓ Expand phase completed successfully
3. MIGRATE (19h)
Description: Execute migrate phase
Risk Level: MEDIUM
Dependencies: expand
Tasks:
• Complete migrate activities
Success Criteria:
✓ Migrate phase completed successfully
4. CONTRACT (19h)
Description: Execute contract phase
Risk Level: MEDIUM
Dependencies: migrate
Tasks:
• Complete contract activities
Success Criteria:
✓ Contract phase completed successfully
5. CLEANUP (19h)
Description: Execute cleanup phase
Risk Level: MEDIUM
Dependencies: contract
Tasks:
• Complete cleanup activities
Success Criteria:
✓ Cleanup phase completed successfully
RISK ASSESSMENT
----------------------------------------
CRITICAL SEVERITY RISKS:
• Insufficient rollback testing
Category: operational
Probability: high | Impact: critical
Mitigation: Execute full rollback procedures in staging environment
Owner: QA Team
HIGH SEVERITY RISKS:
• Data corruption during migration
Category: technical
Probability: low | Impact: critical
Mitigation: Implement comprehensive backup and validation procedures
Owner: DBA Team
• Extended downtime due to migration complexity
Category: technical
Probability: medium | Impact: high
Mitigation: Use blue-green deployment and phased migration approach
Owner: DevOps Team
• Business process disruption
Category: business
Probability: medium | Impact: high
Mitigation: Communicate timeline and provide alternate workflows
Owner: Business Owner
• Zero-downtime requirement increases complexity
Category: business
Probability: high | Impact: medium
Mitigation: Implement blue-green deployment or rolling update strategy
Owner: DevOps Team
• Regulatory compliance requirements
Category: compliance
Probability: medium | Impact: high
Mitigation: Ensure all compliance checks are integrated into migration process
Owner: Compliance Team
ROLLBACK STRATEGY
----------------------------------------
Rollback Triggers:
• Critical system failure
• Data corruption detected
• Migration timeline exceeded by > 50%
• Business-critical functionality unavailable
• Security breach detected
• Stakeholder decision to abort
Rollback Phases:
CLEANUP:
- Revert cleanup changes
- Restore pre-cleanup state
- Validate cleanup rollback success
Estimated Time: 285 minutes
CONTRACT:
- Revert contract changes
- Restore pre-contract state
- Validate contract rollback success
Estimated Time: 285 minutes
MIGRATE:
- Revert migrate changes
- Restore pre-migrate state
- Validate migrate rollback success
Estimated Time: 285 minutes
EXPAND:
- Revert expand changes
- Restore pre-expand state
- Validate expand rollback success
Estimated Time: 285 minutes
PREPARATION:
- Revert preparation changes
- Restore pre-preparation state
- Validate preparation rollback success
Estimated Time: 285 minutes
SUCCESS CRITERIA
----------------------------------------
✓ All data successfully migrated with 100% integrity
✓ System performance meets or exceeds baseline
✓ All business processes functioning normally
✓ No critical security vulnerabilities introduced
✓ Stakeholder acceptance criteria met
✓ Documentation and runbooks updated
STAKEHOLDERS
----------------------------------------
• Business Owner
• Technical Lead
• DevOps Team
• QA Team
• Security Team
• End Users
FILE:expected_outputs/sample_service_migration_plan.json
{
"migration_id": "21031930da18",
"source_system": "Legacy User Service (Java Spring Boot 2.x)",
"target_system": "New User Service (Node.js + TypeScript)",
"migration_type": "service",
"complexity": "critical",
"estimated_duration_hours": 500,
"phases": [
{
"name": "intercept",
"description": "Execute intercept phase",
"duration_hours": 100,
"dependencies": [],
"validation_criteria": [
"Intercept phase completed successfully"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Complete intercept activities"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
},
{
"name": "implement",
"description": "Execute implement phase",
"duration_hours": 100,
"dependencies": [
"intercept"
],
"validation_criteria": [
"Implement phase completed successfully"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Complete implement activities"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
},
{
"name": "redirect",
"description": "Execute redirect phase",
"duration_hours": 100,
"dependencies": [
"implement"
],
"validation_criteria": [
"Redirect phase completed successfully"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Complete redirect activities"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
},
{
"name": "validate",
"description": "Execute validate phase",
"duration_hours": 100,
"dependencies": [
"redirect"
],
"validation_criteria": [
"Validate phase completed successfully"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Complete validate activities"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
},
{
"name": "retire",
"description": "Execute retire phase",
"duration_hours": 100,
"dependencies": [
"validate"
],
"validation_criteria": [
"Retire phase completed successfully"
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
],
"tasks": [
"Complete retire activities"
],
"risk_level": "medium",
"resources_required": [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
}
],
"risks": [
{
"category": "technical",
"description": "Service compatibility issues",
"probability": "medium",
"impact": "high",
"severity": "high",
"mitigation": "Implement comprehensive integration testing",
"owner": "Development Team"
},
{
"category": "technical",
"description": "Performance degradation",
"probability": "medium",
"impact": "medium",
"severity": "medium",
"mitigation": "Conduct load testing and performance benchmarking",
"owner": "DevOps Team"
},
{
"category": "business",
"description": "Feature parity gaps",
"probability": "high",
"impact": "high",
"severity": "high",
"mitigation": "Document feature mapping and acceptance criteria",
"owner": "Product Owner"
},
{
"category": "operational",
"description": "Monitoring gap during transition",
"probability": "medium",
"impact": "medium",
"severity": "medium",
"mitigation": "Set up dual monitoring and alerting systems",
"owner": "SRE Team"
},
{
"category": "business",
"description": "Zero-downtime requirement increases complexity",
"probability": "high",
"impact": "medium",
"severity": "high",
"mitigation": "Implement blue-green deployment or rolling update strategy",
"owner": "DevOps Team"
},
{
"category": "compliance",
"description": "Regulatory compliance requirements",
"probability": "medium",
"impact": "high",
"severity": "high",
"mitigation": "Ensure all compliance checks are integrated into migration process",
"owner": "Compliance Team"
}
],
"success_criteria": [
"All data successfully migrated with 100% integrity",
"System performance meets or exceeds baseline",
"All business processes functioning normally",
"No critical security vulnerabilities introduced",
"Stakeholder acceptance criteria met",
"Documentation and runbooks updated"
],
"rollback_plan": {
"rollback_phases": [
{
"phase": "retire",
"rollback_actions": [
"Revert retire changes",
"Restore pre-retire state",
"Validate retire rollback success"
],
"validation_criteria": [
"System restored to pre-retire state",
"All retire changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 1500
},
{
"phase": "validate",
"rollback_actions": [
"Revert validate changes",
"Restore pre-validate state",
"Validate validate rollback success"
],
"validation_criteria": [
"System restored to pre-validate state",
"All validate changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 1500
},
{
"phase": "redirect",
"rollback_actions": [
"Revert redirect changes",
"Restore pre-redirect state",
"Validate redirect rollback success"
],
"validation_criteria": [
"System restored to pre-redirect state",
"All redirect changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 1500
},
{
"phase": "implement",
"rollback_actions": [
"Revert implement changes",
"Restore pre-implement state",
"Validate implement rollback success"
],
"validation_criteria": [
"System restored to pre-implement state",
"All implement changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 1500
},
{
"phase": "intercept",
"rollback_actions": [
"Revert intercept changes",
"Restore pre-intercept state",
"Validate intercept rollback success"
],
"validation_criteria": [
"System restored to pre-intercept state",
"All intercept changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": 1500
}
],
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Migration timeline exceeded by > 50%",
"Business-critical functionality unavailable",
"Security breach detected",
"Stakeholder decision to abort"
],
"rollback_decision_matrix": {
"low_severity": "Continue with monitoring",
"medium_severity": "Assess and decide within 15 minutes",
"high_severity": "Immediate rollback initiation",
"critical_severity": "Emergency rollback - all hands"
},
"rollback_contacts": [
"Migration Lead",
"Technical Lead",
"Business Owner",
"On-call Engineer"
]
},
"stakeholders": [
"Business Owner",
"Technical Lead",
"DevOps Team",
"QA Team",
"Security Team",
"End Users"
],
"created_at": "2026-02-16T13:47:34.565896"
}
FILE:expected_outputs/sample_service_migration_plan.txt
================================================================================
MIGRATION PLAN: 21031930da18
================================================================================
Source System: Legacy User Service (Java Spring Boot 2.x)
Target System: New User Service (Node.js + TypeScript)
Migration Type: SERVICE
Complexity Level: CRITICAL
Estimated Duration: 500 hours (20.8 days)
Created: 2026-02-16T13:47:34.565896
MIGRATION PHASES
----------------------------------------
1. INTERCEPT (100h)
Description: Execute intercept phase
Risk Level: MEDIUM
Tasks:
• Complete intercept activities
Success Criteria:
✓ Intercept phase completed successfully
2. IMPLEMENT (100h)
Description: Execute implement phase
Risk Level: MEDIUM
Dependencies: intercept
Tasks:
• Complete implement activities
Success Criteria:
✓ Implement phase completed successfully
3. REDIRECT (100h)
Description: Execute redirect phase
Risk Level: MEDIUM
Dependencies: implement
Tasks:
• Complete redirect activities
Success Criteria:
✓ Redirect phase completed successfully
4. VALIDATE (100h)
Description: Execute validate phase
Risk Level: MEDIUM
Dependencies: redirect
Tasks:
• Complete validate activities
Success Criteria:
✓ Validate phase completed successfully
5. RETIRE (100h)
Description: Execute retire phase
Risk Level: MEDIUM
Dependencies: validate
Tasks:
• Complete retire activities
Success Criteria:
✓ Retire phase completed successfully
RISK ASSESSMENT
----------------------------------------
HIGH SEVERITY RISKS:
• Service compatibility issues
Category: technical
Probability: medium | Impact: high
Mitigation: Implement comprehensive integration testing
Owner: Development Team
• Feature parity gaps
Category: business
Probability: high | Impact: high
Mitigation: Document feature mapping and acceptance criteria
Owner: Product Owner
• Zero-downtime requirement increases complexity
Category: business
Probability: high | Impact: medium
Mitigation: Implement blue-green deployment or rolling update strategy
Owner: DevOps Team
• Regulatory compliance requirements
Category: compliance
Probability: medium | Impact: high
Mitigation: Ensure all compliance checks are integrated into migration process
Owner: Compliance Team
MEDIUM SEVERITY RISKS:
• Performance degradation
Category: technical
Probability: medium | Impact: medium
Mitigation: Conduct load testing and performance benchmarking
Owner: DevOps Team
• Monitoring gap during transition
Category: operational
Probability: medium | Impact: medium
Mitigation: Set up dual monitoring and alerting systems
Owner: SRE Team
ROLLBACK STRATEGY
----------------------------------------
Rollback Triggers:
• Critical system failure
• Data corruption detected
• Migration timeline exceeded by > 50%
• Business-critical functionality unavailable
• Security breach detected
• Stakeholder decision to abort
Rollback Phases:
RETIRE:
- Revert retire changes
- Restore pre-retire state
- Validate retire rollback success
Estimated Time: 1500 minutes
VALIDATE:
- Revert validate changes
- Restore pre-validate state
- Validate validate rollback success
Estimated Time: 1500 minutes
REDIRECT:
- Revert redirect changes
- Restore pre-redirect state
- Validate redirect rollback success
Estimated Time: 1500 minutes
IMPLEMENT:
- Revert implement changes
- Restore pre-implement state
- Validate implement rollback success
Estimated Time: 1500 minutes
INTERCEPT:
- Revert intercept changes
- Restore pre-intercept state
- Validate intercept rollback success
Estimated Time: 1500 minutes
SUCCESS CRITERIA
----------------------------------------
✓ All data successfully migrated with 100% integrity
✓ System performance meets or exceeds baseline
✓ All business processes functioning normally
✓ No critical security vulnerabilities introduced
✓ Stakeholder acceptance criteria met
✓ Documentation and runbooks updated
STAKEHOLDERS
----------------------------------------
• Business Owner
• Technical Lead
• DevOps Team
• QA Team
• Security Team
• End Users
FILE:expected_outputs/schema_compatibility_report.json
{
"schema_before": "{\n \"schema_version\": \"1.0\",\n \"database\": \"user_management\",\n \"tables\": {\n \"users\": {\n \"columns\": {\n \"id\": {\n \"type\": \"bigint\",\n \"nullable\": false,\n \"primary_key\": true,\n \"auto_increment\": true\n },\n \"username\": {\n \"type\": \"varchar\",\n \"length\": 50,\n \"nullable\": false,\n \"unique\": true\n },\n \"email\": {\n \"type\": \"varchar\",\n \"length\": 255,\n \"nullable\": false,\n...",
"schema_after": "{\n \"schema_version\": \"2.0\",\n \"database\": \"user_management_v2\",\n \"tables\": {\n \"users\": {\n \"columns\": {\n \"id\": {\n \"type\": \"bigint\",\n \"nullable\": false,\n \"primary_key\": true,\n \"auto_increment\": true\n },\n \"username\": {\n \"type\": \"varchar\",\n \"length\": 50,\n \"nullable\": false,\n \"unique\": true\n },\n \"email\": {\n \"type\": \"varchar\",\n \"length\": 320,\n \"nullable\": fals...",
"analysis_date": "2026-02-16T13:47:27.050459",
"overall_compatibility": "potentially_incompatible",
"breaking_changes_count": 0,
"potentially_breaking_count": 4,
"non_breaking_changes_count": 0,
"additive_changes_count": 0,
"issues": [
{
"type": "check_added",
"severity": "potentially_breaking",
"description": "New check constraint 'phone IS NULL OR LENGTH(phone) >= 10' added to table 'users'",
"field_path": "tables.users.constraints.check",
"old_value": null,
"new_value": "phone IS NULL OR LENGTH(phone) >= 10",
"impact": "New check constraint may reject existing data",
"suggested_migration": "Validate existing data complies with new constraint",
"affected_operations": [
"INSERT",
"UPDATE"
]
},
{
"type": "check_added",
"severity": "potentially_breaking",
"description": "New check constraint 'bio IS NULL OR LENGTH(bio) <= 2000' added to table 'user_profiles'",
"field_path": "tables.user_profiles.constraints.check",
"old_value": null,
"new_value": "bio IS NULL OR LENGTH(bio) <= 2000",
"impact": "New check constraint may reject existing data",
"suggested_migration": "Validate existing data complies with new constraint",
"affected_operations": [
"INSERT",
"UPDATE"
]
},
{
"type": "check_added",
"severity": "potentially_breaking",
"description": "New check constraint 'language IN ('en', 'es', 'fr', 'de', 'it', 'pt', 'ru', 'ja', 'ko', 'zh')' added to table 'user_profiles'",
"field_path": "tables.user_profiles.constraints.check",
"old_value": null,
"new_value": "language IN ('en', 'es', 'fr', 'de', 'it', 'pt', 'ru', 'ja', 'ko', 'zh')",
"impact": "New check constraint may reject existing data",
"suggested_migration": "Validate existing data complies with new constraint",
"affected_operations": [
"INSERT",
"UPDATE"
]
},
{
"type": "check_added",
"severity": "potentially_breaking",
"description": "New check constraint 'session_type IN ('web', 'mobile', 'api', 'admin')' added to table 'user_sessions'",
"field_path": "tables.user_sessions.constraints.check",
"old_value": null,
"new_value": "session_type IN ('web', 'mobile', 'api', 'admin')",
"impact": "New check constraint may reject existing data",
"suggested_migration": "Validate existing data complies with new constraint",
"affected_operations": [
"INSERT",
"UPDATE"
]
}
],
"migration_scripts": [
{
"script_type": "sql",
"description": "Create new table user_preferences",
"script_content": "CREATE TABLE user_preferences (\n id bigint NOT NULL,\n user_id bigint NOT NULL,\n preference_key varchar NOT NULL,\n preference_value json,\n created_at timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP,\n updated_at timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP\n);",
"rollback_script": "DROP TABLE IF EXISTS user_preferences;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.tables WHERE table_name = 'user_preferences';"
},
{
"script_type": "sql",
"description": "Add column email_verified_at to table users",
"script_content": "ALTER TABLE users ADD COLUMN email_verified_at timestamp;",
"rollback_script": "ALTER TABLE users DROP COLUMN email_verified_at;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.columns WHERE table_name = 'users' AND column_name = 'email_verified_at';"
},
{
"script_type": "sql",
"description": "Add column phone_verified_at to table users",
"script_content": "ALTER TABLE users ADD COLUMN phone_verified_at timestamp;",
"rollback_script": "ALTER TABLE users DROP COLUMN phone_verified_at;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.columns WHERE table_name = 'users' AND column_name = 'phone_verified_at';"
},
{
"script_type": "sql",
"description": "Add column two_factor_enabled to table users",
"script_content": "ALTER TABLE users ADD COLUMN two_factor_enabled boolean NOT NULL DEFAULT False;",
"rollback_script": "ALTER TABLE users DROP COLUMN two_factor_enabled;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.columns WHERE table_name = 'users' AND column_name = 'two_factor_enabled';"
},
{
"script_type": "sql",
"description": "Add column last_login_at to table users",
"script_content": "ALTER TABLE users ADD COLUMN last_login_at timestamp;",
"rollback_script": "ALTER TABLE users DROP COLUMN last_login_at;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.columns WHERE table_name = 'users' AND column_name = 'last_login_at';"
},
{
"script_type": "sql",
"description": "Add check constraint to users",
"script_content": "ALTER TABLE users ADD CONSTRAINT check_users CHECK (phone IS NULL OR LENGTH(phone) >= 10);",
"rollback_script": "ALTER TABLE users DROP CONSTRAINT check_users;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.table_constraints WHERE table_name = 'users' AND constraint_type = 'CHECK';"
},
{
"script_type": "sql",
"description": "Add column timezone to table user_profiles",
"script_content": "ALTER TABLE user_profiles ADD COLUMN timezone varchar DEFAULT UTC;",
"rollback_script": "ALTER TABLE user_profiles DROP COLUMN timezone;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.columns WHERE table_name = 'user_profiles' AND column_name = 'timezone';"
},
{
"script_type": "sql",
"description": "Add column language to table user_profiles",
"script_content": "ALTER TABLE user_profiles ADD COLUMN language varchar NOT NULL DEFAULT en;",
"rollback_script": "ALTER TABLE user_profiles DROP COLUMN language;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.columns WHERE table_name = 'user_profiles' AND column_name = 'language';"
},
{
"script_type": "sql",
"description": "Add check constraint to user_profiles",
"script_content": "ALTER TABLE user_profiles ADD CONSTRAINT check_user_profiles CHECK (bio IS NULL OR LENGTH(bio) <= 2000);",
"rollback_script": "ALTER TABLE user_profiles DROP CONSTRAINT check_user_profiles;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.table_constraints WHERE table_name = 'user_profiles' AND constraint_type = 'CHECK';"
},
{
"script_type": "sql",
"description": "Add check constraint to user_profiles",
"script_content": "ALTER TABLE user_profiles ADD CONSTRAINT check_user_profiles CHECK (language IN ('en', 'es', 'fr', 'de', 'it', 'pt', 'ru', 'ja', 'ko', 'zh'));",
"rollback_script": "ALTER TABLE user_profiles DROP CONSTRAINT check_user_profiles;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.table_constraints WHERE table_name = 'user_profiles' AND constraint_type = 'CHECK';"
},
{
"script_type": "sql",
"description": "Add column session_type to table user_sessions",
"script_content": "ALTER TABLE user_sessions ADD COLUMN session_type varchar NOT NULL DEFAULT web;",
"rollback_script": "ALTER TABLE user_sessions DROP COLUMN session_type;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.columns WHERE table_name = 'user_sessions' AND column_name = 'session_type';"
},
{
"script_type": "sql",
"description": "Add column is_mobile to table user_sessions",
"script_content": "ALTER TABLE user_sessions ADD COLUMN is_mobile boolean NOT NULL DEFAULT False;",
"rollback_script": "ALTER TABLE user_sessions DROP COLUMN is_mobile;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.columns WHERE table_name = 'user_sessions' AND column_name = 'is_mobile';"
},
{
"script_type": "sql",
"description": "Add check constraint to user_sessions",
"script_content": "ALTER TABLE user_sessions ADD CONSTRAINT check_user_sessions CHECK (session_type IN ('web', 'mobile', 'api', 'admin'));",
"rollback_script": "ALTER TABLE user_sessions DROP CONSTRAINT check_user_sessions;",
"dependencies": [],
"validation_query": "SELECT COUNT(*) FROM information_schema.table_constraints WHERE table_name = 'user_sessions' AND constraint_type = 'CHECK';"
}
],
"risk_assessment": {
"overall_risk": "medium",
"deployment_risk": "safe_independent_deployment",
"rollback_complexity": "low",
"testing_requirements": [
"integration_testing",
"regression_testing",
"data_migration_testing"
]
},
"recommendations": [
"Conduct thorough testing with realistic data volumes",
"Implement monitoring for migration success metrics",
"Test all migration scripts in staging environment",
"Implement migration progress monitoring",
"Create detailed communication plan for stakeholders",
"Implement feature flags for gradual rollout"
]
}
FILE:expected_outputs/schema_compatibility_report.txt
================================================================================
COMPATIBILITY ANALYSIS REPORT
================================================================================
Analysis Date: 2026-02-16T13:47:27.050459
Overall Compatibility: POTENTIALLY_INCOMPATIBLE
SUMMARY
----------------------------------------
Breaking Changes: 0
Potentially Breaking: 4
Non-Breaking Changes: 0
Additive Changes: 0
Total Issues Found: 4
RISK ASSESSMENT
----------------------------------------
Overall Risk: medium
Deployment Risk: safe_independent_deployment
Rollback Complexity: low
Testing Requirements: ['integration_testing', 'regression_testing', 'data_migration_testing']
POTENTIALLY BREAKING ISSUES
----------------------------------------
• New check constraint 'phone IS NULL OR LENGTH(phone) >= 10' added to table 'users'
Field: tables.users.constraints.check
Impact: New check constraint may reject existing data
Migration: Validate existing data complies with new constraint
Affected Operations: INSERT, UPDATE
• New check constraint 'bio IS NULL OR LENGTH(bio) <= 2000' added to table 'user_profiles'
Field: tables.user_profiles.constraints.check
Impact: New check constraint may reject existing data
Migration: Validate existing data complies with new constraint
Affected Operations: INSERT, UPDATE
• New check constraint 'language IN ('en', 'es', 'fr', 'de', 'it', 'pt', 'ru', 'ja', 'ko', 'zh')' added to table 'user_profiles'
Field: tables.user_profiles.constraints.check
Impact: New check constraint may reject existing data
Migration: Validate existing data complies with new constraint
Affected Operations: INSERT, UPDATE
• New check constraint 'session_type IN ('web', 'mobile', 'api', 'admin')' added to table 'user_sessions'
Field: tables.user_sessions.constraints.check
Impact: New check constraint may reject existing data
Migration: Validate existing data complies with new constraint
Affected Operations: INSERT, UPDATE
SUGGESTED MIGRATION SCRIPTS
----------------------------------------
1. Create new table user_preferences
Type: sql
Script:
CREATE TABLE user_preferences (
id bigint NOT NULL,
user_id bigint NOT NULL,
preference_key varchar NOT NULL,
preference_value json,
created_at timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP,
updated_at timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP
);
2. Add column email_verified_at to table users
Type: sql
Script:
ALTER TABLE users ADD COLUMN email_verified_at timestamp;
3. Add column phone_verified_at to table users
Type: sql
Script:
ALTER TABLE users ADD COLUMN phone_verified_at timestamp;
4. Add column two_factor_enabled to table users
Type: sql
Script:
ALTER TABLE users ADD COLUMN two_factor_enabled boolean NOT NULL DEFAULT False;
5. Add column last_login_at to table users
Type: sql
Script:
ALTER TABLE users ADD COLUMN last_login_at timestamp;
6. Add check constraint to users
Type: sql
Script:
ALTER TABLE users ADD CONSTRAINT check_users CHECK (phone IS NULL OR LENGTH(phone) >= 10);
7. Add column timezone to table user_profiles
Type: sql
Script:
ALTER TABLE user_profiles ADD COLUMN timezone varchar DEFAULT UTC;
8. Add column language to table user_profiles
Type: sql
Script:
ALTER TABLE user_profiles ADD COLUMN language varchar NOT NULL DEFAULT en;
9. Add check constraint to user_profiles
Type: sql
Script:
ALTER TABLE user_profiles ADD CONSTRAINT check_user_profiles CHECK (bio IS NULL OR LENGTH(bio) <= 2000);
10. Add check constraint to user_profiles
Type: sql
Script:
ALTER TABLE user_profiles ADD CONSTRAINT check_user_profiles CHECK (language IN ('en', 'es', 'fr', 'de', 'it', 'pt', 'ru', 'ja', 'ko', 'zh'));
11. Add column session_type to table user_sessions
Type: sql
Script:
ALTER TABLE user_sessions ADD COLUMN session_type varchar NOT NULL DEFAULT web;
12. Add column is_mobile to table user_sessions
Type: sql
Script:
ALTER TABLE user_sessions ADD COLUMN is_mobile boolean NOT NULL DEFAULT False;
13. Add check constraint to user_sessions
Type: sql
Script:
ALTER TABLE user_sessions ADD CONSTRAINT check_user_sessions CHECK (session_type IN ('web', 'mobile', 'api', 'admin'));
RECOMMENDATIONS
----------------------------------------
1. Conduct thorough testing with realistic data volumes
2. Implement monitoring for migration success metrics
3. Test all migration scripts in staging environment
4. Implement migration progress monitoring
5. Create detailed communication plan for stakeholders
6. Implement feature flags for gradual rollout
FILE:README.md
# Migration Architect
**Tier:** POWERFUL
**Category:** Engineering - Migration Strategy
**Purpose:** Zero-downtime migration planning, compatibility validation, and rollback strategy generation
## Overview
The Migration Architect skill provides comprehensive tools and methodologies for planning, executing, and validating complex system migrations with minimal business impact. This skill combines proven migration patterns with automated planning tools to ensure successful transitions between systems, databases, and infrastructure.
## Components
### Core Scripts
1. **migration_planner.py** - Automated migration plan generation
2. **compatibility_checker.py** - Schema and API compatibility analysis
3. **rollback_generator.py** - Comprehensive rollback procedure generation
### Reference Documentation
- **migration_patterns_catalog.md** - Detailed catalog of proven migration patterns
- **zero_downtime_techniques.md** - Comprehensive zero-downtime migration techniques
- **data_reconciliation_strategies.md** - Advanced data consistency and reconciliation strategies
### Sample Assets
- **sample_database_migration.json** - Example database migration specification
- **sample_service_migration.json** - Example service migration specification
- **database_schema_before.json** - Sample "before" database schema
- **database_schema_after.json** - Sample "after" database schema
## Quick Start
### 1. Generate a Migration Plan
```bash
python3 scripts/migration_planner.py \
--input assets/sample_database_migration.json \
--output migration_plan.json \
--format both
```
**Input:** Migration specification with source, target, constraints, and requirements
**Output:** Detailed phased migration plan with risk assessment, timeline, and validation gates
### 2. Check Compatibility
```bash
python3 scripts/compatibility_checker.py \
--before assets/database_schema_before.json \
--after assets/database_schema_after.json \
--type database \
--output compatibility_report.json \
--format both
```
**Input:** Before and after schema definitions
**Output:** Compatibility report with breaking changes, migration scripts, and recommendations
### 3. Generate Rollback Procedures
```bash
python3 scripts/rollback_generator.py \
--input migration_plan.json \
--output rollback_runbook.json \
--format both
```
**Input:** Migration plan from step 1
**Output:** Comprehensive rollback runbook with procedures, triggers, and communication templates
## Script Details
### Migration Planner (`migration_planner.py`)
Generates comprehensive migration plans with:
- **Phased approach** with dependencies and validation gates
- **Risk assessment** with mitigation strategies
- **Timeline estimation** based on complexity and constraints
- **Rollback triggers** and success criteria
- **Stakeholder communication** templates
**Usage:**
```bash
python3 scripts/migration_planner.py [OPTIONS]
Options:
--input, -i Input migration specification file (JSON) [required]
--output, -o Output file for migration plan (JSON)
--format, -f Output format: json, text, both (default: both)
--validate Validate migration specification only
```
**Input Format:**
```json
{
"type": "database|service|infrastructure",
"pattern": "schema_change|strangler_fig|blue_green",
"source": "Source system description",
"target": "Target system description",
"constraints": {
"max_downtime_minutes": 30,
"data_volume_gb": 2500,
"dependencies": ["service1", "service2"],
"compliance_requirements": ["GDPR", "SOX"]
}
}
```
### Compatibility Checker (`compatibility_checker.py`)
Analyzes compatibility between schema versions:
- **Breaking change detection** (removed fields, type changes, constraint additions)
- **Data migration requirements** identification
- **Suggested migration scripts** generation
- **Risk assessment** for each change
**Usage:**
```bash
python3 scripts/compatibility_checker.py [OPTIONS]
Options:
--before Before schema file (JSON) [required]
--after After schema file (JSON) [required]
--type Schema type: database, api (default: database)
--output, -o Output file for compatibility report (JSON)
--format, -f Output format: json, text, both (default: both)
```
**Exit Codes:**
- `0`: No compatibility issues
- `1`: Potentially breaking changes found
- `2`: Breaking changes found
### Rollback Generator (`rollback_generator.py`)
Creates comprehensive rollback procedures:
- **Phase-by-phase rollback** steps
- **Automated trigger conditions** for rollback
- **Data recovery procedures**
- **Communication templates** for different audiences
- **Validation checklists** for rollback success
**Usage:**
```bash
python3 scripts/rollback_generator.py [OPTIONS]
Options:
--input, -i Input migration plan file (JSON) [required]
--output, -o Output file for rollback runbook (JSON)
--format, -f Output format: json, text, both (default: both)
```
## Migration Patterns Supported
### Database Migrations
- **Expand-Contract Pattern** - Zero-downtime schema evolution
- **Parallel Schema Pattern** - Side-by-side schema migration
- **Event Sourcing Migration** - Event-driven data migration
### Service Migrations
- **Strangler Fig Pattern** - Gradual legacy system replacement
- **Parallel Run Pattern** - Risk mitigation through dual execution
- **Blue-Green Deployment** - Zero-downtime service updates
### Infrastructure Migrations
- **Lift and Shift** - Quick cloud migration with minimal changes
- **Hybrid Cloud Migration** - Gradual cloud adoption
- **Multi-Cloud Migration** - Distribution across multiple providers
## Sample Workflow
### 1. Database Schema Migration
```bash
# Generate migration plan
python3 scripts/migration_planner.py \
--input assets/sample_database_migration.json \
--output db_migration_plan.json
# Check schema compatibility
python3 scripts/compatibility_checker.py \
--before assets/database_schema_before.json \
--after assets/database_schema_after.json \
--type database \
--output schema_compatibility.json
# Generate rollback procedures
python3 scripts/rollback_generator.py \
--input db_migration_plan.json \
--output db_rollback_runbook.json
```
### 2. Service Migration
```bash
# Generate service migration plan
python3 scripts/migration_planner.py \
--input assets/sample_service_migration.json \
--output service_migration_plan.json
# Generate rollback procedures
python3 scripts/rollback_generator.py \
--input service_migration_plan.json \
--output service_rollback_runbook.json
```
## Output Examples
### Migration Plan Structure
```json
{
"migration_id": "abc123def456",
"source_system": "Legacy User Service",
"target_system": "New User Service",
"migration_type": "service",
"complexity": "medium",
"estimated_duration_hours": 72,
"phases": [
{
"name": "preparation",
"description": "Prepare systems and teams for migration",
"duration_hours": 8,
"validation_criteria": ["All backups completed successfully"],
"rollback_triggers": ["Critical system failure"],
"risk_level": "medium"
}
],
"risks": [
{
"category": "technical",
"description": "Service compatibility issues",
"severity": "high",
"mitigation": "Comprehensive integration testing"
}
]
}
```
### Compatibility Report Structure
```json
{
"overall_compatibility": "potentially_incompatible",
"breaking_changes_count": 2,
"potentially_breaking_count": 3,
"issues": [
{
"type": "required_column_added",
"severity": "breaking",
"description": "Required column 'email_verified_at' added",
"suggested_migration": "Add default value initially"
}
],
"migration_scripts": [
{
"script_type": "sql",
"description": "Add email verification columns",
"script_content": "ALTER TABLE users ADD COLUMN email_verified_at TIMESTAMP;",
"rollback_script": "ALTER TABLE users DROP COLUMN email_verified_at;"
}
]
}
```
## Best Practices
### Planning Phase
1. **Start with risk assessment** - Identify failure modes before planning
2. **Design for rollback** - Every step should have a tested rollback procedure
3. **Validate in staging** - Execute full migration in production-like environment
4. **Plan gradual rollout** - Use feature flags and traffic routing
### Execution Phase
1. **Monitor continuously** - Track technical and business metrics
2. **Communicate proactively** - Keep stakeholders informed
3. **Document everything** - Maintain detailed logs for analysis
4. **Stay flexible** - Be prepared to adjust based on real-world performance
### Validation Phase
1. **Automate validation** - Use automated consistency and performance checks
2. **Test business logic** - Validate critical business processes end-to-end
3. **Load test** - Verify performance under expected production load
4. **Security validation** - Ensure security controls function properly
## Integration
### CI/CD Pipeline Integration
```yaml
# Example GitHub Actions workflow
name: Migration Validation
on: [push, pull_request]
jobs:
validate-migration:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v2
- name: Validate Migration Plan
run: |
python3 scripts/migration_planner.py \
--input migration_spec.json \
--validate
- name: Check Compatibility
run: |
python3 scripts/compatibility_checker.py \
--before schema_before.json \
--after schema_after.json \
--type database
```
### Monitoring Integration
The tools generate metrics and alerts that can be integrated with:
- **Prometheus** - For metrics collection
- **Grafana** - For visualization and dashboards
- **PagerDuty** - For incident management
- **Slack** - For team notifications
## Advanced Features
### Machine Learning Integration
- Anomaly detection for data consistency issues
- Predictive analysis for migration success probability
- Automated pattern recognition for migration optimization
### Performance Optimization
- Parallel processing for large-scale migrations
- Incremental reconciliation strategies
- Statistical sampling for validation
### Compliance Support
- GDPR compliance tracking
- SOX audit trail generation
- HIPAA security validation
## Troubleshooting
### Common Issues
**"Migration plan validation failed"**
- Check JSON syntax in migration specification
- Ensure all required fields are present
- Validate constraint values are realistic
**"Compatibility checker reports false positives"**
- Review excluded fields configuration
- Check data type mapping compatibility
- Adjust tolerance settings for numerical comparisons
**"Rollback procedures seem incomplete"**
- Ensure migration plan includes all phases
- Verify database backup locations are specified
- Check that all dependencies are documented
### Getting Help
1. **Review documentation** - Check reference docs for patterns and techniques
2. **Examine sample files** - Use provided assets as templates
3. **Check expected outputs** - Compare your results with sample outputs
4. **Validate inputs** - Ensure input files match expected format
## Contributing
To extend or modify the Migration Architect skill:
1. **Add new patterns** - Extend pattern templates in migration_planner.py
2. **Enhance compatibility checks** - Add new validation rules in compatibility_checker.py
3. **Improve rollback procedures** - Add specialized rollback steps in rollback_generator.py
4. **Update documentation** - Keep reference docs current with new patterns
## License
This skill is part of the claude-skills repository and follows the same license terms.
FILE:references/data_reconciliation_strategies.md
# Data Reconciliation Strategies
## Overview
Data reconciliation is the process of ensuring data consistency and integrity across systems during and after migrations. This document provides comprehensive strategies, tools, and implementation patterns for detecting, measuring, and correcting data discrepancies in migration scenarios.
## Core Principles
### 1. Eventually Consistent
Accept that perfect real-time consistency may not be achievable during migrations, but ensure eventual consistency through reconciliation processes.
### 2. Idempotent Operations
All reconciliation operations must be safe to run multiple times without causing additional issues.
### 3. Audit Trail
Maintain detailed logs of all reconciliation actions for compliance and debugging.
### 4. Non-Destructive
Reconciliation should prefer addition over deletion, and always maintain backups before corrections.
## Types of Data Inconsistencies
### 1. Missing Records
Records that exist in source but not in target system.
### 2. Extra Records
Records that exist in target but not in source system.
### 3. Field Mismatches
Records exist in both systems but with different field values.
### 4. Referential Integrity Violations
Foreign key relationships that are broken during migration.
### 5. Temporal Inconsistencies
Data with incorrect timestamps or ordering.
### 6. Schema Drift
Structural differences between source and target schemas.
## Detection Strategies
### 1. Row Count Validation
#### Simple Count Comparison
```sql
-- Compare total row counts
SELECT
'source' as system,
COUNT(*) as row_count
FROM source_table
UNION ALL
SELECT
'target' as system,
COUNT(*) as row_count
FROM target_table;
```
#### Filtered Count Comparison
```sql
-- Compare counts with business logic filters
WITH source_counts AS (
SELECT
status,
created_date::date as date,
COUNT(*) as count
FROM source_orders
WHERE created_date >= '2024-01-01'
GROUP BY status, created_date::date
),
target_counts AS (
SELECT
status,
created_date::date as date,
COUNT(*) as count
FROM target_orders
WHERE created_date >= '2024-01-01'
GROUP BY status, created_date::date
)
SELECT
COALESCE(s.status, t.status) as status,
COALESCE(s.date, t.date) as date,
COALESCE(s.count, 0) as source_count,
COALESCE(t.count, 0) as target_count,
COALESCE(s.count, 0) - COALESCE(t.count, 0) as difference
FROM source_counts s
FULL OUTER JOIN target_counts t
ON s.status = t.status AND s.date = t.date
WHERE COALESCE(s.count, 0) != COALESCE(t.count, 0);
```
### 2. Checksum-Based Validation
#### Record-Level Checksums
```python
import hashlib
import json
class RecordChecksum:
def __init__(self, exclude_fields=None):
self.exclude_fields = exclude_fields or ['updated_at', 'version']
def calculate_checksum(self, record):
"""Calculate MD5 checksum for a database record"""
# Remove excluded fields and sort for consistency
filtered_record = {
k: v for k, v in record.items()
if k not in self.exclude_fields
}
# Convert to sorted JSON string for consistent hashing
normalized = json.dumps(filtered_record, sort_keys=True, default=str)
return hashlib.md5(normalized.encode('utf-8')).hexdigest()
def compare_records(self, source_record, target_record):
"""Compare two records using checksums"""
source_checksum = self.calculate_checksum(source_record)
target_checksum = self.calculate_checksum(target_record)
return {
'match': source_checksum == target_checksum,
'source_checksum': source_checksum,
'target_checksum': target_checksum
}
# Usage example
checksum_calculator = RecordChecksum(exclude_fields=['updated_at', 'migration_flag'])
source_records = fetch_records_from_source()
target_records = fetch_records_from_target()
mismatches = []
for source_id, source_record in source_records.items():
if source_id in target_records:
comparison = checksum_calculator.compare_records(
source_record, target_records[source_id]
)
if not comparison['match']:
mismatches.append({
'record_id': source_id,
'source_checksum': comparison['source_checksum'],
'target_checksum': comparison['target_checksum']
})
```
#### Aggregate Checksums
```sql
-- Calculate aggregate checksums for data validation
WITH source_aggregates AS (
SELECT
DATE_TRUNC('day', created_at) as day,
status,
COUNT(*) as record_count,
SUM(amount) as total_amount,
MD5(STRING_AGG(CAST(id AS VARCHAR) || ':' || CAST(amount AS VARCHAR), '|' ORDER BY id)) as checksum
FROM source_transactions
GROUP BY DATE_TRUNC('day', created_at), status
),
target_aggregates AS (
SELECT
DATE_TRUNC('day', created_at) as day,
status,
COUNT(*) as record_count,
SUM(amount) as total_amount,
MD5(STRING_AGG(CAST(id AS VARCHAR) || ':' || CAST(amount AS VARCHAR), '|' ORDER BY id)) as checksum
FROM target_transactions
GROUP BY DATE_TRUNC('day', created_at), status
)
SELECT
COALESCE(s.day, t.day) as day,
COALESCE(s.status, t.status) as status,
COALESCE(s.record_count, 0) as source_count,
COALESCE(t.record_count, 0) as target_count,
COALESCE(s.total_amount, 0) as source_amount,
COALESCE(t.total_amount, 0) as target_amount,
s.checksum as source_checksum,
t.checksum as target_checksum,
CASE WHEN s.checksum = t.checksum THEN 'MATCH' ELSE 'MISMATCH' END as status
FROM source_aggregates s
FULL OUTER JOIN target_aggregates t
ON s.day = t.day AND s.status = t.status
WHERE s.checksum != t.checksum OR s.checksum IS NULL OR t.checksum IS NULL;
```
### 3. Delta Detection
#### Change Data Capture (CDC) Based
```python
class CDCReconciler:
def __init__(self, kafka_client, database_client):
self.kafka = kafka_client
self.db = database_client
self.processed_changes = set()
def process_cdc_stream(self, topic_name):
"""Process CDC events and track changes for reconciliation"""
consumer = self.kafka.consumer(topic_name)
for message in consumer:
change_event = json.loads(message.value)
change_id = f"{change_event['table']}:{change_event['key']}:{change_event['timestamp']}"
if change_id in self.processed_changes:
continue # Skip duplicate events
try:
self.apply_change(change_event)
self.processed_changes.add(change_id)
# Commit offset only after successful processing
consumer.commit()
except Exception as e:
# Log failure and continue - will be caught by reconciliation
self.log_processing_failure(change_id, str(e))
def apply_change(self, change_event):
"""Apply CDC change to target system"""
table = change_event['table']
operation = change_event['operation']
key = change_event['key']
data = change_event.get('data', {})
if operation == 'INSERT':
self.db.insert(table, data)
elif operation == 'UPDATE':
self.db.update(table, key, data)
elif operation == 'DELETE':
self.db.delete(table, key)
def reconcile_missed_changes(self, start_timestamp, end_timestamp):
"""Find and apply changes that may have been missed"""
# Query source database for changes in time window
source_changes = self.db.get_changes_in_window(
start_timestamp, end_timestamp
)
missed_changes = []
for change in source_changes:
change_id = f"{change['table']}:{change['key']}:{change['timestamp']}"
if change_id not in self.processed_changes:
missed_changes.append(change)
# Apply missed changes
for change in missed_changes:
try:
self.apply_change(change)
print(f"Applied missed change: {change['table']}:{change['key']}")
except Exception as e:
print(f"Failed to apply missed change: {e}")
```
### 4. Business Logic Validation
#### Critical Business Rules Validation
```python
class BusinessLogicValidator:
def __init__(self, source_db, target_db):
self.source_db = source_db
self.target_db = target_db
def validate_financial_consistency(self):
"""Validate critical financial calculations"""
validation_rules = [
{
'name': 'daily_transaction_totals',
'source_query': """
SELECT DATE(created_at) as date, SUM(amount) as total
FROM source_transactions
WHERE created_at >= CURRENT_DATE - INTERVAL '30 days'
GROUP BY DATE(created_at)
""",
'target_query': """
SELECT DATE(created_at) as date, SUM(amount) as total
FROM target_transactions
WHERE created_at >= CURRENT_DATE - INTERVAL '30 days'
GROUP BY DATE(created_at)
""",
'tolerance': 0.01 # Allow $0.01 difference for rounding
},
{
'name': 'customer_balance_totals',
'source_query': """
SELECT customer_id, SUM(balance) as total_balance
FROM source_accounts
GROUP BY customer_id
HAVING SUM(balance) > 0
""",
'target_query': """
SELECT customer_id, SUM(balance) as total_balance
FROM target_accounts
GROUP BY customer_id
HAVING SUM(balance) > 0
""",
'tolerance': 0.01
}
]
validation_results = []
for rule in validation_rules:
source_data = self.source_db.execute_query(rule['source_query'])
target_data = self.target_db.execute_query(rule['target_query'])
differences = self.compare_financial_data(
source_data, target_data, rule['tolerance']
)
validation_results.append({
'rule_name': rule['name'],
'differences_found': len(differences),
'differences': differences[:10], # First 10 differences
'status': 'PASS' if len(differences) == 0 else 'FAIL'
})
return validation_results
def compare_financial_data(self, source_data, target_data, tolerance):
"""Compare financial data with tolerance for rounding differences"""
source_dict = {
tuple(row[:-1]): row[-1] for row in source_data
} # Last column is the amount
target_dict = {
tuple(row[:-1]): row[-1] for row in target_data
}
differences = []
# Check for missing records and value differences
for key, source_value in source_dict.items():
if key not in target_dict:
differences.append({
'key': key,
'source_value': source_value,
'target_value': None,
'difference_type': 'MISSING_IN_TARGET'
})
else:
target_value = target_dict[key]
if abs(float(source_value) - float(target_value)) > tolerance:
differences.append({
'key': key,
'source_value': source_value,
'target_value': target_value,
'difference': float(source_value) - float(target_value),
'difference_type': 'VALUE_MISMATCH'
})
# Check for extra records in target
for key, target_value in target_dict.items():
if key not in source_dict:
differences.append({
'key': key,
'source_value': None,
'target_value': target_value,
'difference_type': 'EXTRA_IN_TARGET'
})
return differences
```
## Correction Strategies
### 1. Automated Correction
#### Missing Record Insertion
```python
class AutoCorrector:
def __init__(self, source_db, target_db, dry_run=True):
self.source_db = source_db
self.target_db = target_db
self.dry_run = dry_run
self.correction_log = []
def correct_missing_records(self, table_name, key_field):
"""Add missing records from source to target"""
# Find records in source but not in target
missing_query = f"""
SELECT s.*
FROM source_{table_name} s
LEFT JOIN target_{table_name} t ON s.{key_field} = t.{key_field}
WHERE t.{key_field} IS NULL
"""
missing_records = self.source_db.execute_query(missing_query)
for record in missing_records:
correction = {
'table': table_name,
'operation': 'INSERT',
'key': record[key_field],
'data': record,
'timestamp': datetime.utcnow()
}
if not self.dry_run:
try:
self.target_db.insert(table_name, record)
correction['status'] = 'SUCCESS'
except Exception as e:
correction['status'] = 'FAILED'
correction['error'] = str(e)
else:
correction['status'] = 'DRY_RUN'
self.correction_log.append(correction)
return len(missing_records)
def correct_field_mismatches(self, table_name, key_field, fields_to_correct):
"""Correct field value mismatches"""
mismatch_query = f"""
SELECT s.{key_field}, {', '.join([f's.{f} as source_{f}, t.{f} as target_{f}' for f in fields_to_correct])}
FROM source_{table_name} s
JOIN target_{table_name} t ON s.{key_field} = t.{key_field}
WHERE {' OR '.join([f's.{f} != t.{f}' for f in fields_to_correct])}
"""
mismatched_records = self.source_db.execute_query(mismatch_query)
for record in mismatched_records:
key_value = record[key_field]
updates = {}
for field in fields_to_correct:
source_value = record[f'source_{field}']
target_value = record[f'target_{field}']
if source_value != target_value:
updates[field] = source_value
if updates:
correction = {
'table': table_name,
'operation': 'UPDATE',
'key': key_value,
'updates': updates,
'timestamp': datetime.utcnow()
}
if not self.dry_run:
try:
self.target_db.update(table_name, {key_field: key_value}, updates)
correction['status'] = 'SUCCESS'
except Exception as e:
correction['status'] = 'FAILED'
correction['error'] = str(e)
else:
correction['status'] = 'DRY_RUN'
self.correction_log.append(correction)
return len(mismatched_records)
```
### 2. Manual Review Process
#### Correction Workflow
```python
class ManualReviewSystem:
def __init__(self, database_client):
self.db = database_client
self.review_queue = []
def queue_for_review(self, discrepancy):
"""Add discrepancy to manual review queue"""
review_item = {
'id': str(uuid.uuid4()),
'discrepancy_type': discrepancy['type'],
'table': discrepancy['table'],
'record_key': discrepancy['key'],
'source_data': discrepancy.get('source_data'),
'target_data': discrepancy.get('target_data'),
'description': discrepancy['description'],
'severity': discrepancy.get('severity', 'medium'),
'status': 'PENDING',
'created_at': datetime.utcnow(),
'reviewed_by': None,
'reviewed_at': None,
'resolution': None
}
self.review_queue.append(review_item)
# Persist to review database
self.db.insert('manual_review_queue', review_item)
return review_item['id']
def process_review(self, review_id, reviewer, action, notes=None):
"""Process manual review decision"""
review_item = self.get_review_item(review_id)
if not review_item:
raise ValueError(f"Review item {review_id} not found")
review_item.update({
'status': 'REVIEWED',
'reviewed_by': reviewer,
'reviewed_at': datetime.utcnow(),
'resolution': {
'action': action, # 'APPLY_SOURCE', 'KEEP_TARGET', 'CUSTOM_FIX'
'notes': notes
}
})
# Apply the resolution
if action == 'APPLY_SOURCE':
self.apply_source_data(review_item)
elif action == 'KEEP_TARGET':
pass # No action needed
elif action == 'CUSTOM_FIX':
# Custom fix would be applied separately
pass
# Update review record
self.db.update('manual_review_queue',
{'id': review_id},
review_item)
return review_item
def generate_review_report(self):
"""Generate summary report of manual reviews"""
reviews = self.db.query("""
SELECT
discrepancy_type,
severity,
status,
COUNT(*) as count,
MIN(created_at) as oldest_review,
MAX(created_at) as newest_review
FROM manual_review_queue
GROUP BY discrepancy_type, severity, status
ORDER BY severity DESC, discrepancy_type
""")
return reviews
```
### 3. Reconciliation Scheduling
#### Automated Reconciliation Jobs
```python
import schedule
import time
from datetime import datetime, timedelta
class ReconciliationScheduler:
def __init__(self, reconciler):
self.reconciler = reconciler
self.job_history = []
def setup_schedules(self):
"""Set up automated reconciliation schedules"""
# Quick reconciliation every 15 minutes during migration
schedule.every(15).minutes.do(self.quick_reconciliation)
# Comprehensive reconciliation every 4 hours
schedule.every(4).hours.do(self.comprehensive_reconciliation)
# Deep validation daily
schedule.every().day.at("02:00").do(self.deep_validation)
# Weekly business logic validation
schedule.every().sunday.at("03:00").do(self.business_logic_validation)
def quick_reconciliation(self):
"""Quick count-based reconciliation"""
job_start = datetime.utcnow()
try:
# Check critical tables only
critical_tables = [
'transactions', 'orders', 'customers', 'accounts'
]
results = []
for table in critical_tables:
count_diff = self.reconciler.check_row_counts(table)
if abs(count_diff) > 0:
results.append({
'table': table,
'count_difference': count_diff,
'severity': 'high' if abs(count_diff) > 100 else 'medium'
})
job_result = {
'job_type': 'quick_reconciliation',
'start_time': job_start,
'end_time': datetime.utcnow(),
'status': 'completed',
'issues_found': len(results),
'details': results
}
# Alert if significant issues found
if any(r['severity'] == 'high' for r in results):
self.send_alert(job_result)
except Exception as e:
job_result = {
'job_type': 'quick_reconciliation',
'start_time': job_start,
'end_time': datetime.utcnow(),
'status': 'failed',
'error': str(e)
}
self.job_history.append(job_result)
def comprehensive_reconciliation(self):
"""Comprehensive checksum-based reconciliation"""
job_start = datetime.utcnow()
try:
tables_to_check = self.get_migration_tables()
issues = []
for table in tables_to_check:
# Sample-based checksum validation
sample_issues = self.reconciler.validate_sample_checksums(
table, sample_size=1000
)
issues.extend(sample_issues)
# Auto-correct simple issues
auto_corrections = 0
for issue in issues:
if issue['auto_correctable']:
self.reconciler.auto_correct_issue(issue)
auto_corrections += 1
else:
# Queue for manual review
self.reconciler.queue_for_manual_review(issue)
job_result = {
'job_type': 'comprehensive_reconciliation',
'start_time': job_start,
'end_time': datetime.utcnow(),
'status': 'completed',
'total_issues': len(issues),
'auto_corrections': auto_corrections,
'manual_reviews_queued': len(issues) - auto_corrections
}
except Exception as e:
job_result = {
'job_type': 'comprehensive_reconciliation',
'start_time': job_start,
'end_time': datetime.utcnow(),
'status': 'failed',
'error': str(e)
}
self.job_history.append(job_result)
def run_scheduler(self):
"""Run the reconciliation scheduler"""
print("Starting reconciliation scheduler...")
while True:
schedule.run_pending()
time.sleep(60) # Check every minute
```
## Monitoring and Reporting
### 1. Reconciliation Metrics
```python
class ReconciliationMetrics:
def __init__(self, prometheus_client):
self.prometheus = prometheus_client
# Define metrics
self.inconsistencies_found = Counter(
'reconciliation_inconsistencies_total',
'Number of inconsistencies found',
['table', 'type', 'severity']
)
self.reconciliation_duration = Histogram(
'reconciliation_duration_seconds',
'Time spent on reconciliation jobs',
['job_type']
)
self.auto_corrections = Counter(
'reconciliation_auto_corrections_total',
'Number of automatically corrected inconsistencies',
['table', 'correction_type']
)
self.data_drift_gauge = Gauge(
'data_drift_percentage',
'Percentage of records with inconsistencies',
['table']
)
def record_inconsistency(self, table, inconsistency_type, severity):
"""Record a found inconsistency"""
self.inconsistencies_found.labels(
table=table,
type=inconsistency_type,
severity=severity
).inc()
def record_auto_correction(self, table, correction_type):
"""Record an automatic correction"""
self.auto_corrections.labels(
table=table,
correction_type=correction_type
).inc()
def update_data_drift(self, table, drift_percentage):
"""Update data drift gauge"""
self.data_drift_gauge.labels(table=table).set(drift_percentage)
def record_job_duration(self, job_type, duration_seconds):
"""Record reconciliation job duration"""
self.reconciliation_duration.labels(job_type=job_type).observe(duration_seconds)
```
### 2. Alerting Rules
```yaml
# Prometheus alerting rules for data reconciliation
groups:
- name: data_reconciliation
rules:
- alert: HighDataInconsistency
expr: reconciliation_inconsistencies_total > 100
for: 5m
labels:
severity: critical
annotations:
summary: "High number of data inconsistencies detected"
description: "{{ $value }} inconsistencies found in the last 5 minutes"
- alert: DataDriftHigh
expr: data_drift_percentage > 5
for: 10m
labels:
severity: warning
annotations:
summary: "Data drift percentage is high"
description: "{{ $labels.table }} has {{ $value }}% data drift"
- alert: ReconciliationJobFailed
expr: up{job="reconciliation"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: "Reconciliation job is down"
description: "The data reconciliation service is not responding"
- alert: AutoCorrectionRateHigh
expr: rate(reconciliation_auto_corrections_total[10m]) > 10
for: 5m
labels:
severity: warning
annotations:
summary: "High rate of automatic corrections"
description: "Auto-correction rate is {{ $value }} per second"
```
### 3. Dashboard and Reporting
```python
class ReconciliationDashboard:
def __init__(self, database_client, metrics_client):
self.db = database_client
self.metrics = metrics_client
def generate_daily_report(self, date=None):
"""Generate daily reconciliation report"""
if not date:
date = datetime.utcnow().date()
# Query reconciliation results for the day
daily_stats = self.db.query("""
SELECT
table_name,
inconsistency_type,
COUNT(*) as count,
AVG(CASE WHEN resolution = 'AUTO_CORRECTED' THEN 1 ELSE 0 END) as auto_correction_rate
FROM reconciliation_log
WHERE DATE(created_at) = %s
GROUP BY table_name, inconsistency_type
""", (date,))
# Generate summary
summary = {
'date': date.isoformat(),
'total_inconsistencies': sum(row['count'] for row in daily_stats),
'auto_correction_rate': sum(row['auto_correction_rate'] * row['count'] for row in daily_stats) / max(sum(row['count'] for row in daily_stats), 1),
'tables_affected': len(set(row['table_name'] for row in daily_stats)),
'details_by_table': {}
}
# Group by table
for row in daily_stats:
table = row['table_name']
if table not in summary['details_by_table']:
summary['details_by_table'][table] = []
summary['details_by_table'][table].append({
'inconsistency_type': row['inconsistency_type'],
'count': row['count'],
'auto_correction_rate': row['auto_correction_rate']
})
return summary
def generate_trend_analysis(self, days=7):
"""Generate trend analysis for reconciliation metrics"""
end_date = datetime.utcnow().date()
start_date = end_date - timedelta(days=days)
trends = self.db.query("""
SELECT
DATE(created_at) as date,
table_name,
COUNT(*) as inconsistencies,
AVG(CASE WHEN resolution = 'AUTO_CORRECTED' THEN 1 ELSE 0 END) as auto_correction_rate
FROM reconciliation_log
WHERE DATE(created_at) BETWEEN %s AND %s
GROUP BY DATE(created_at), table_name
ORDER BY date, table_name
""", (start_date, end_date))
# Calculate trends
trend_analysis = {
'period': f"{start_date} to {end_date}",
'trends': {},
'overall_trend': 'stable'
}
for table in set(row['table_name'] for row in trends):
table_data = [row for row in trends if row['table_name'] == table]
if len(table_data) >= 2:
first_count = table_data[0]['inconsistencies']
last_count = table_data[-1]['inconsistencies']
if last_count > first_count * 1.2:
trend = 'increasing'
elif last_count < first_count * 0.8:
trend = 'decreasing'
else:
trend = 'stable'
trend_analysis['trends'][table] = {
'direction': trend,
'first_day_count': first_count,
'last_day_count': last_count,
'change_percentage': ((last_count - first_count) / max(first_count, 1)) * 100
}
return trend_analysis
```
## Advanced Reconciliation Techniques
### 1. Machine Learning-Based Anomaly Detection
```python
from sklearn.isolation import IsolationForest
from sklearn.preprocessing import StandardScaler
import numpy as np
class MLAnomalyDetector:
def __init__(self):
self.models = {}
self.scalers = {}
def train_anomaly_detector(self, table_name, training_data):
"""Train anomaly detection model for a specific table"""
# Prepare features (convert records to numerical features)
features = self.extract_features(training_data)
# Scale features
scaler = StandardScaler()
scaled_features = scaler.fit_transform(features)
# Train isolation forest
model = IsolationForest(contamination=0.05, random_state=42)
model.fit(scaled_features)
# Store model and scaler
self.models[table_name] = model
self.scalers[table_name] = scaler
def detect_anomalies(self, table_name, data):
"""Detect anomalous records that may indicate reconciliation issues"""
if table_name not in self.models:
raise ValueError(f"No trained model for table {table_name}")
# Extract features
features = self.extract_features(data)
# Scale features
scaled_features = self.scalers[table_name].transform(features)
# Predict anomalies
anomaly_scores = self.models[table_name].decision_function(scaled_features)
anomaly_predictions = self.models[table_name].predict(scaled_features)
# Return anomalous records with scores
anomalies = []
for i, (record, score, is_anomaly) in enumerate(zip(data, anomaly_scores, anomaly_predictions)):
if is_anomaly == -1: # Isolation forest returns -1 for anomalies
anomalies.append({
'record_index': i,
'record': record,
'anomaly_score': score,
'severity': 'high' if score < -0.5 else 'medium'
})
return anomalies
def extract_features(self, data):
"""Extract numerical features from database records"""
features = []
for record in data:
record_features = []
for key, value in record.items():
if isinstance(value, (int, float)):
record_features.append(value)
elif isinstance(value, str):
# Convert string to hash-based feature
record_features.append(hash(value) % 10000)
elif isinstance(value, datetime):
# Convert datetime to timestamp
record_features.append(value.timestamp())
else:
# Default value for other types
record_features.append(0)
features.append(record_features)
return np.array(features)
```
### 2. Probabilistic Reconciliation
```python
import random
from typing import List, Dict, Tuple
class ProbabilisticReconciler:
def __init__(self, confidence_threshold=0.95):
self.confidence_threshold = confidence_threshold
def statistical_sampling_validation(self, table_name: str, population_size: int) -> Dict:
"""Use statistical sampling to validate large datasets"""
# Calculate sample size for 95% confidence, 5% margin of error
confidence_level = 0.95
margin_of_error = 0.05
z_score = 1.96 # for 95% confidence
p = 0.5 # assume 50% error rate for maximum sample size
sample_size = (z_score ** 2 * p * (1 - p)) / (margin_of_error ** 2)
if population_size < 10000:
# Finite population correction
sample_size = sample_size / (1 + (sample_size - 1) / population_size)
sample_size = min(int(sample_size), population_size)
# Generate random sample
sample_ids = self.generate_random_sample(table_name, sample_size)
# Validate sample
sample_results = self.validate_sample_records(table_name, sample_ids)
# Calculate population estimates
error_rate = sample_results['errors'] / sample_size
estimated_errors = int(population_size * error_rate)
# Calculate confidence interval
standard_error = (error_rate * (1 - error_rate) / sample_size) ** 0.5
margin_of_error_actual = z_score * standard_error
confidence_interval = (
max(0, error_rate - margin_of_error_actual),
min(1, error_rate + margin_of_error_actual)
)
return {
'table_name': table_name,
'population_size': population_size,
'sample_size': sample_size,
'sample_error_rate': error_rate,
'estimated_total_errors': estimated_errors,
'confidence_interval': confidence_interval,
'confidence_level': confidence_level,
'recommendation': self.generate_recommendation(error_rate, confidence_interval)
}
def generate_random_sample(self, table_name: str, sample_size: int) -> List[int]:
"""Generate random sample of record IDs"""
# Get total record count and ID range
id_range = self.db.query(f"SELECT MIN(id), MAX(id) FROM {table_name}")[0]
min_id, max_id = id_range
# Generate random IDs
sample_ids = []
attempts = 0
max_attempts = sample_size * 10 # Avoid infinite loop
while len(sample_ids) < sample_size and attempts < max_attempts:
candidate_id = random.randint(min_id, max_id)
# Check if ID exists
exists = self.db.query(f"SELECT 1 FROM {table_name} WHERE id = %s", (candidate_id,))
if exists and candidate_id not in sample_ids:
sample_ids.append(candidate_id)
attempts += 1
return sample_ids
def validate_sample_records(self, table_name: str, sample_ids: List[int]) -> Dict:
"""Validate a sample of records"""
validation_results = {
'total_checked': len(sample_ids),
'errors': 0,
'error_details': []
}
for record_id in sample_ids:
# Get record from both source and target
source_record = self.source_db.get_record(table_name, record_id)
target_record = self.target_db.get_record(table_name, record_id)
if not target_record:
validation_results['errors'] += 1
validation_results['error_details'].append({
'id': record_id,
'error_type': 'MISSING_IN_TARGET'
})
elif not self.records_match(source_record, target_record):
validation_results['errors'] += 1
validation_results['error_details'].append({
'id': record_id,
'error_type': 'DATA_MISMATCH',
'differences': self.find_differences(source_record, target_record)
})
return validation_results
def generate_recommendation(self, error_rate: float, confidence_interval: Tuple[float, float]) -> str:
"""Generate recommendation based on error rate and confidence"""
if confidence_interval[1] < 0.01: # Less than 1% error rate with confidence
return "Data quality is excellent. Continue with normal reconciliation schedule."
elif confidence_interval[1] < 0.05: # Less than 5% error rate with confidence
return "Data quality is acceptable. Monitor closely and investigate sample errors."
elif confidence_interval[0] > 0.1: # More than 10% error rate with confidence
return "Data quality is poor. Immediate comprehensive reconciliation required."
else:
return "Data quality is uncertain. Increase sample size for better estimates."
```
## Performance Optimization
### 1. Parallel Processing
```python
import asyncio
import multiprocessing as mp
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
class ParallelReconciler:
def __init__(self, max_workers=None):
self.max_workers = max_workers or mp.cpu_count()
async def parallel_table_reconciliation(self, tables: List[str]):
"""Reconcile multiple tables in parallel"""
async with asyncio.Semaphore(self.max_workers):
tasks = [
self.reconcile_table_async(table)
for table in tables
]
results = await asyncio.gather(*tasks, return_exceptions=True)
# Process results
summary = {
'total_tables': len(tables),
'successful': 0,
'failed': 0,
'results': {}
}
for table, result in zip(tables, results):
if isinstance(result, Exception):
summary['failed'] += 1
summary['results'][table] = {
'status': 'failed',
'error': str(result)
}
else:
summary['successful'] += 1
summary['results'][table] = result
return summary
def parallel_chunk_processing(self, table_name: str, chunk_size: int = 10000):
"""Process table reconciliation in parallel chunks"""
# Get total record count
total_records = self.db.get_record_count(table_name)
num_chunks = (total_records + chunk_size - 1) // chunk_size
# Create chunk specifications
chunks = []
for i in range(num_chunks):
start_id = i * chunk_size
end_id = min((i + 1) * chunk_size - 1, total_records - 1)
chunks.append({
'table': table_name,
'start_id': start_id,
'end_id': end_id,
'chunk_number': i + 1
})
# Process chunks in parallel
with ProcessPoolExecutor(max_workers=self.max_workers) as executor:
chunk_results = list(executor.map(self.process_chunk, chunks))
# Aggregate results
total_inconsistencies = sum(r['inconsistencies'] for r in chunk_results)
total_corrections = sum(r['corrections'] for r in chunk_results)
return {
'table': table_name,
'total_records': total_records,
'chunks_processed': len(chunks),
'total_inconsistencies': total_inconsistencies,
'total_corrections': total_corrections,
'chunk_details': chunk_results
}
def process_chunk(self, chunk_spec: Dict) -> Dict:
"""Process a single chunk of records"""
# This runs in a separate process
table = chunk_spec['table']
start_id = chunk_spec['start_id']
end_id = chunk_spec['end_id']
# Initialize database connections for this process
local_source_db = SourceDatabase()
local_target_db = TargetDatabase()
# Get records in chunk
source_records = local_source_db.get_records_range(table, start_id, end_id)
target_records = local_target_db.get_records_range(table, start_id, end_id)
# Reconcile chunk
inconsistencies = 0
corrections = 0
for source_record in source_records:
target_record = target_records.get(source_record['id'])
if not target_record:
inconsistencies += 1
# Auto-correct if possible
try:
local_target_db.insert(table, source_record)
corrections += 1
except Exception:
pass # Log error in production
elif not self.records_match(source_record, target_record):
inconsistencies += 1
# Auto-correct field mismatches
try:
updates = self.calculate_updates(source_record, target_record)
local_target_db.update(table, source_record['id'], updates)
corrections += 1
except Exception:
pass # Log error in production
return {
'chunk_number': chunk_spec['chunk_number'],
'start_id': start_id,
'end_id': end_id,
'records_processed': len(source_records),
'inconsistencies': inconsistencies,
'corrections': corrections
}
```
### 2. Incremental Reconciliation
```python
class IncrementalReconciler:
def __init__(self, source_db, target_db):
self.source_db = source_db
self.target_db = target_db
self.last_reconciliation_times = {}
def incremental_reconciliation(self, table_name: str):
"""Reconcile only records changed since last reconciliation"""
last_reconciled = self.get_last_reconciliation_time(table_name)
# Get records modified since last reconciliation
modified_source = self.source_db.get_records_modified_since(
table_name, last_reconciled
)
modified_target = self.target_db.get_records_modified_since(
table_name, last_reconciled
)
# Create lookup dictionaries
source_dict = {r['id']: r for r in modified_source}
target_dict = {r['id']: r for r in modified_target}
# Find all record IDs to check
all_ids = set(source_dict.keys()) | set(target_dict.keys())
inconsistencies = []
for record_id in all_ids:
source_record = source_dict.get(record_id)
target_record = target_dict.get(record_id)
if source_record and not target_record:
inconsistencies.append({
'type': 'missing_in_target',
'table': table_name,
'id': record_id,
'source_record': source_record
})
elif not source_record and target_record:
inconsistencies.append({
'type': 'extra_in_target',
'table': table_name,
'id': record_id,
'target_record': target_record
})
elif source_record and target_record:
if not self.records_match(source_record, target_record):
inconsistencies.append({
'type': 'data_mismatch',
'table': table_name,
'id': record_id,
'source_record': source_record,
'target_record': target_record,
'differences': self.find_differences(source_record, target_record)
})
# Update last reconciliation time
self.update_last_reconciliation_time(table_name, datetime.utcnow())
return {
'table': table_name,
'reconciliation_time': datetime.utcnow(),
'records_checked': len(all_ids),
'inconsistencies_found': len(inconsistencies),
'inconsistencies': inconsistencies
}
def get_last_reconciliation_time(self, table_name: str) -> datetime:
"""Get the last reconciliation timestamp for a table"""
result = self.source_db.query("""
SELECT last_reconciled_at
FROM reconciliation_metadata
WHERE table_name = %s
""", (table_name,))
if result:
return result[0]['last_reconciled_at']
else:
# First time reconciliation - start from beginning of migration
return self.get_migration_start_time()
def update_last_reconciliation_time(self, table_name: str, timestamp: datetime):
"""Update the last reconciliation timestamp"""
self.source_db.execute("""
INSERT INTO reconciliation_metadata (table_name, last_reconciled_at)
VALUES (%s, %s)
ON CONFLICT (table_name)
DO UPDATE SET last_reconciled_at = %s
""", (table_name, timestamp, timestamp))
```
This comprehensive guide provides the framework and tools necessary for implementing robust data reconciliation strategies during migrations, ensuring data integrity and consistency while minimizing business disruption.
FILE:references/migration_patterns_catalog.md
# Migration Patterns Catalog
## Overview
This catalog provides detailed descriptions of proven migration patterns, their use cases, implementation guidelines, and best practices. Each pattern includes code examples, diagrams, and lessons learned from real-world implementations.
## Database Migration Patterns
### 1. Expand-Contract Pattern
**Use Case:** Schema evolution with zero downtime
**Complexity:** Medium
**Risk Level:** Low-Medium
#### Description
The Expand-Contract pattern allows for schema changes without downtime by following a three-phase approach:
1. **Expand:** Add new schema elements alongside existing ones
2. **Migrate:** Dual-write to both old and new schema during transition
3. **Contract:** Remove old schema elements after validation
#### Implementation Steps
```sql
-- Phase 1: Expand
ALTER TABLE users ADD COLUMN email_new VARCHAR(255);
CREATE INDEX CONCURRENTLY idx_users_email_new ON users(email_new);
-- Phase 2: Migrate (Application Code)
-- Write to both columns during transition period
INSERT INTO users (name, email, email_new) VALUES (?, ?, ?);
-- Backfill existing data
UPDATE users SET email_new = email WHERE email_new IS NULL;
-- Phase 3: Contract (after validation)
ALTER TABLE users DROP COLUMN email;
ALTER TABLE users RENAME COLUMN email_new TO email;
```
#### Pros and Cons
**Pros:**
- Zero downtime deployments
- Safe rollback at any point
- Gradual transition with validation
**Cons:**
- Increased storage during transition
- More complex application logic
- Extended migration timeline
### 2. Parallel Schema Pattern
**Use Case:** Major database restructuring
**Complexity:** High
**Risk Level:** Medium
#### Description
Run new and old schemas in parallel, using feature flags to gradually route traffic to the new schema while maintaining the ability to rollback quickly.
#### Implementation Example
```python
class DatabaseRouter:
def __init__(self, feature_flag_service):
self.feature_flags = feature_flag_service
self.old_db = OldDatabaseConnection()
self.new_db = NewDatabaseConnection()
def route_query(self, user_id, query_type):
if self.feature_flags.is_enabled("new_schema", user_id):
return self.new_db.execute(query_type)
else:
return self.old_db.execute(query_type)
def dual_write(self, data):
# Write to both databases for consistency
success_old = self.old_db.write(data)
success_new = self.new_db.write(transform_data(data))
if not (success_old and success_new):
# Handle partial failures
self.handle_dual_write_failure(data, success_old, success_new)
```
#### Best Practices
- Implement data consistency checks between schemas
- Use circuit breakers for automatic failover
- Monitor performance impact of dual writes
- Plan for data reconciliation processes
### 3. Event Sourcing Migration
**Use Case:** Migrating systems with complex business logic
**Complexity:** High
**Risk Level:** Medium-High
#### Description
Capture all changes as events during migration, enabling replay and reconciliation capabilities.
#### Event Store Schema
```sql
CREATE TABLE migration_events (
event_id UUID PRIMARY KEY,
aggregate_id UUID NOT NULL,
event_type VARCHAR(100) NOT NULL,
event_data JSONB NOT NULL,
event_version INTEGER NOT NULL,
occurred_at TIMESTAMP WITH TIME ZONE DEFAULT NOW(),
processed_at TIMESTAMP WITH TIME ZONE
);
```
#### Migration Event Handler
```python
class MigrationEventHandler:
def __init__(self, old_store, new_store):
self.old_store = old_store
self.new_store = new_store
self.event_log = []
def handle_update(self, entity_id, old_data, new_data):
# Log the change as an event
event = MigrationEvent(
entity_id=entity_id,
event_type="entity_migrated",
old_data=old_data,
new_data=new_data,
timestamp=datetime.now()
)
self.event_log.append(event)
# Apply to new store
success = self.new_store.update(entity_id, new_data)
if not success:
# Mark for retry
event.status = "failed"
self.schedule_retry(event)
return success
def replay_events(self, from_timestamp=None):
"""Replay events for reconciliation"""
events = self.get_events_since(from_timestamp)
for event in events:
self.apply_event(event)
```
## Service Migration Patterns
### 1. Strangler Fig Pattern
**Use Case:** Legacy system replacement
**Complexity:** Medium-High
**Risk Level:** Medium
#### Description
Gradually replace legacy functionality by intercepting calls and routing them to new services, eventually "strangling" the legacy system.
#### Implementation Architecture
```yaml
# API Gateway Configuration
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
name: user-service-migration
spec:
http:
- match:
- headers:
migration-flag:
exact: "new"
route:
- destination:
host: user-service-v2
- route:
- destination:
host: user-service-v1
```
#### Strangler Proxy Implementation
```python
class StranglerProxy:
def __init__(self):
self.legacy_service = LegacyUserService()
self.new_service = NewUserService()
self.feature_flags = FeatureFlagService()
def handle_request(self, request):
route = self.determine_route(request)
if route == "new":
return self.handle_with_new_service(request)
elif route == "both":
return self.handle_with_both_services(request)
else:
return self.handle_with_legacy_service(request)
def determine_route(self, request):
user_id = request.get('user_id')
if self.feature_flags.is_enabled("new_user_service", user_id):
if self.feature_flags.is_enabled("dual_write", user_id):
return "both"
else:
return "new"
else:
return "legacy"
```
### 2. Parallel Run Pattern
**Use Case:** Risk mitigation for critical services
**Complexity:** Medium
**Risk Level:** Low-Medium
#### Description
Run both old and new services simultaneously, comparing outputs to validate correctness before switching traffic.
#### Implementation
```python
class ParallelRunManager:
def __init__(self):
self.primary_service = PrimaryService()
self.candidate_service = CandidateService()
self.comparator = ResponseComparator()
self.metrics = MetricsCollector()
async def parallel_execute(self, request):
# Execute both services concurrently
primary_task = asyncio.create_task(
self.primary_service.process(request)
)
candidate_task = asyncio.create_task(
self.candidate_service.process(request)
)
# Always wait for primary
primary_result = await primary_task
try:
# Wait for candidate with timeout
candidate_result = await asyncio.wait_for(
candidate_task, timeout=5.0
)
# Compare results
comparison = self.comparator.compare(
primary_result, candidate_result
)
# Record metrics
self.metrics.record_comparison(comparison)
except asyncio.TimeoutError:
self.metrics.record_timeout("candidate")
except Exception as e:
self.metrics.record_error("candidate", str(e))
# Always return primary result
return primary_result
```
### 3. Blue-Green Deployment Pattern
**Use Case:** Zero-downtime service updates
**Complexity:** Low-Medium
**Risk Level:** Low
#### Description
Maintain two identical production environments (blue and green), switching traffic between them for deployments.
#### Kubernetes Implementation
```yaml
# Blue Deployment
apiVersion: apps/v1
kind: Deployment
metadata:
name: app-blue
labels:
version: blue
spec:
replicas: 3
selector:
matchLabels:
app: myapp
version: blue
template:
metadata:
labels:
app: myapp
version: blue
spec:
containers:
- name: app
image: myapp:v1.0.0
---
# Green Deployment
apiVersion: apps/v1
kind: Deployment
metadata:
name: app-green
labels:
version: green
spec:
replicas: 3
selector:
matchLabels:
app: myapp
version: green
template:
metadata:
labels:
app: myapp
version: green
spec:
containers:
- name: app
image: myapp:v2.0.0
---
# Service (switches between blue and green)
apiVersion: v1
kind: Service
metadata:
name: app-service
spec:
selector:
app: myapp
version: blue # Change to green for deployment
ports:
- port: 80
targetPort: 8080
```
## Infrastructure Migration Patterns
### 1. Lift and Shift Pattern
**Use Case:** Quick cloud migration with minimal changes
**Complexity:** Low-Medium
**Risk Level:** Low
#### Description
Migrate applications to cloud infrastructure with minimal or no code changes, focusing on infrastructure compatibility.
#### Migration Checklist
```yaml
Pre-Migration Assessment:
- inventory_current_infrastructure:
- servers_and_specifications
- network_configuration
- storage_requirements
- security_configurations
- identify_dependencies:
- database_connections
- external_service_integrations
- file_system_dependencies
- assess_compatibility:
- operating_system_versions
- runtime_dependencies
- license_requirements
Migration Execution:
- provision_target_infrastructure:
- compute_instances
- storage_volumes
- network_configuration
- security_groups
- migrate_data:
- database_backup_restore
- file_system_replication
- configuration_files
- update_configurations:
- connection_strings
- environment_variables
- dns_records
- validate_functionality:
- application_health_checks
- end_to_end_testing
- performance_validation
```
### 2. Hybrid Cloud Migration
**Use Case:** Gradual cloud adoption with on-premises integration
**Complexity:** High
**Risk Level:** Medium-High
#### Description
Maintain some components on-premises while migrating others to cloud, requiring secure connectivity and data synchronization.
#### Network Architecture
```hcl
# Terraform configuration for hybrid connectivity
resource "aws_vpc" "main" {
cidr_block = "10.0.0.0/16"
enable_dns_hostnames = true
enable_dns_support = true
}
resource "aws_vpn_gateway" "main" {
vpc_id = aws_vpc.main.id
tags = {
Name = "hybrid-vpn-gateway"
}
}
resource "aws_customer_gateway" "main" {
bgp_asn = 65000
ip_address = var.on_premises_public_ip
type = "ipsec.1"
tags = {
Name = "on-premises-gateway"
}
}
resource "aws_vpn_connection" "main" {
vpn_gateway_id = aws_vpn_gateway.main.id
customer_gateway_id = aws_customer_gateway.main.id
type = "ipsec.1"
static_routes_only = true
}
```
#### Data Synchronization Pattern
```python
class HybridDataSync:
def __init__(self):
self.on_prem_db = OnPremiseDatabase()
self.cloud_db = CloudDatabase()
self.sync_log = SyncLogManager()
async def bidirectional_sync(self):
"""Synchronize data between on-premises and cloud"""
# Get last sync timestamp
last_sync = self.sync_log.get_last_sync_time()
# Sync on-prem changes to cloud
on_prem_changes = self.on_prem_db.get_changes_since(last_sync)
for change in on_prem_changes:
await self.apply_change_to_cloud(change)
# Sync cloud changes to on-prem
cloud_changes = self.cloud_db.get_changes_since(last_sync)
for change in cloud_changes:
await self.apply_change_to_on_prem(change)
# Handle conflicts
conflicts = self.detect_conflicts(on_prem_changes, cloud_changes)
for conflict in conflicts:
await self.resolve_conflict(conflict)
# Update sync timestamp
self.sync_log.record_sync_completion()
async def apply_change_to_cloud(self, change):
"""Apply on-premises change to cloud database"""
try:
if change.operation == "INSERT":
await self.cloud_db.insert(change.table, change.data)
elif change.operation == "UPDATE":
await self.cloud_db.update(change.table, change.key, change.data)
elif change.operation == "DELETE":
await self.cloud_db.delete(change.table, change.key)
self.sync_log.record_success(change.id, "cloud")
except Exception as e:
self.sync_log.record_failure(change.id, "cloud", str(e))
raise
```
### 3. Multi-Cloud Migration
**Use Case:** Avoiding vendor lock-in or regulatory requirements
**Complexity:** Very High
**Risk Level:** High
#### Description
Distribute workloads across multiple cloud providers for resilience, compliance, or cost optimization.
#### Service Mesh Configuration
```yaml
# Istio configuration for multi-cloud service mesh
apiVersion: networking.istio.io/v1beta1
kind: ServiceEntry
metadata:
name: aws-service
spec:
hosts:
- aws-service.company.com
ports:
- number: 443
name: https
protocol: HTTPS
location: MESH_EXTERNAL
resolution: DNS
---
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
name: multi-cloud-routing
spec:
hosts:
- user-service
http:
- match:
- headers:
region:
exact: "us-east"
route:
- destination:
host: aws-service.company.com
weight: 100
- match:
- headers:
region:
exact: "eu-west"
route:
- destination:
host: gcp-service.company.com
weight: 100
- route: # Default routing
- destination:
host: user-service
subset: local
weight: 80
- destination:
host: aws-service.company.com
weight: 20
```
## Feature Flag Patterns
### 1. Progressive Rollout Pattern
**Use Case:** Gradual feature deployment with risk mitigation
**Implementation:**
```python
class ProgressiveRollout:
def __init__(self, feature_name):
self.feature_name = feature_name
self.rollout_percentage = 0
self.user_buckets = {}
def is_enabled_for_user(self, user_id):
# Consistent user bucketing
user_hash = hashlib.md5(f"{self.feature_name}:{user_id}".encode()).hexdigest()
bucket = int(user_hash, 16) % 100
return bucket < self.rollout_percentage
def increase_rollout(self, target_percentage, step_size=10):
"""Gradually increase rollout percentage"""
while self.rollout_percentage < target_percentage:
self.rollout_percentage = min(
self.rollout_percentage + step_size,
target_percentage
)
# Monitor metrics before next increase
yield self.rollout_percentage
time.sleep(300) # Wait 5 minutes between increases
```
### 2. Circuit Breaker Pattern
**Use Case:** Automatic fallback during migration issues
```python
class MigrationCircuitBreaker:
def __init__(self, failure_threshold=5, timeout=60):
self.failure_count = 0
self.failure_threshold = failure_threshold
self.timeout = timeout
self.last_failure_time = None
self.state = 'CLOSED' # CLOSED, OPEN, HALF_OPEN
def call_new_service(self, request):
if self.state == 'OPEN':
if self.should_attempt_reset():
self.state = 'HALF_OPEN'
else:
return self.fallback_to_legacy(request)
try:
response = self.new_service.process(request)
self.on_success()
return response
except Exception as e:
self.on_failure()
return self.fallback_to_legacy(request)
def on_success(self):
self.failure_count = 0
self.state = 'CLOSED'
def on_failure(self):
self.failure_count += 1
self.last_failure_time = time.time()
if self.failure_count >= self.failure_threshold:
self.state = 'OPEN'
def should_attempt_reset(self):
return (time.time() - self.last_failure_time) >= self.timeout
```
## Migration Anti-Patterns
### 1. Big Bang Migration (Anti-Pattern)
**Why to Avoid:**
- High risk of complete system failure
- Difficult to rollback
- Extended downtime
- All-or-nothing deployment
**Better Alternative:** Use incremental migration patterns like Strangler Fig or Parallel Run.
### 2. No Rollback Plan (Anti-Pattern)
**Why to Avoid:**
- Cannot recover from failures
- Increases business risk
- Panic-driven decisions during issues
**Better Alternative:** Always implement comprehensive rollback procedures before migration.
### 3. Insufficient Testing (Anti-Pattern)
**Why to Avoid:**
- Unknown compatibility issues
- Performance degradation
- Data corruption risks
**Better Alternative:** Implement comprehensive testing at each migration phase.
## Pattern Selection Matrix
| Migration Type | Complexity | Downtime Tolerance | Recommended Pattern |
|---------------|------------|-------------------|-------------------|
| Schema Change | Low | Zero | Expand-Contract |
| Schema Change | High | Zero | Parallel Schema |
| Service Replace | Medium | Zero | Strangler Fig |
| Service Update | Low | Zero | Blue-Green |
| Data Migration | High | Some | Event Sourcing |
| Infrastructure | Low | Some | Lift and Shift |
| Infrastructure | High | Zero | Hybrid Cloud |
## Success Metrics
### Technical Metrics
- Migration completion rate
- System availability during migration
- Performance impact (response time, throughput)
- Error rate changes
- Rollback execution time
### Business Metrics
- Customer impact score
- Revenue protection
- Time to value realization
- Stakeholder satisfaction
### Operational Metrics
- Team efficiency
- Knowledge transfer effectiveness
- Post-migration support requirements
- Documentation completeness
## Lessons Learned
### Common Pitfalls
1. **Underestimating data dependencies** - Always map all data relationships
2. **Insufficient monitoring** - Implement comprehensive observability before migration
3. **Poor communication** - Keep all stakeholders informed throughout the process
4. **Rushed timelines** - Allow adequate time for testing and validation
5. **Ignoring performance impact** - Benchmark before and after migration
### Best Practices
1. **Start with low-risk migrations** - Build confidence and experience
2. **Automate everything possible** - Reduce human error and increase repeatability
3. **Test rollback procedures** - Ensure you can recover from any failure
4. **Monitor continuously** - Use real-time dashboards and alerting
5. **Document everything** - Create comprehensive runbooks and documentation
This catalog serves as a reference for selecting appropriate migration patterns based on specific requirements, risk tolerance, and technical constraints.
FILE:references/zero_downtime_techniques.md
# Zero-Downtime Migration Techniques
## Overview
Zero-downtime migrations are critical for maintaining business continuity and user experience during system changes. This guide provides comprehensive techniques, patterns, and implementation strategies for achieving true zero-downtime migrations across different system components.
## Core Principles
### 1. Backward Compatibility
Every change must be backward compatible until all clients have migrated to the new version.
### 2. Incremental Changes
Break large changes into smaller, independent increments that can be deployed and validated separately.
### 3. Feature Flags
Use feature toggles to control the rollout of new functionality without code deployments.
### 4. Graceful Degradation
Ensure systems continue to function even when some components are unavailable or degraded.
## Database Zero-Downtime Techniques
### Schema Evolution Without Downtime
#### 1. Additive Changes Only
**Principle:** Only add new elements; never remove or modify existing ones directly.
```sql
-- ✅ Good: Additive change
ALTER TABLE users ADD COLUMN middle_name VARCHAR(50);
-- ❌ Bad: Breaking change
ALTER TABLE users DROP COLUMN email;
```
#### 2. Multi-Phase Schema Evolution
**Phase 1: Expand**
```sql
-- Add new column alongside existing one
ALTER TABLE users ADD COLUMN email_address VARCHAR(255);
-- Add index concurrently (PostgreSQL)
CREATE INDEX CONCURRENTLY idx_users_email_address ON users(email_address);
```
**Phase 2: Dual Write (Application Code)**
```python
class UserService:
def create_user(self, name, email):
# Write to both old and new columns
user = User(
name=name,
email=email, # Old column
email_address=email # New column
)
return user.save()
def update_email(self, user_id, new_email):
# Update both columns
user = User.objects.get(id=user_id)
user.email = new_email
user.email_address = new_email
user.save()
return user
```
**Phase 3: Backfill Data**
```sql
-- Backfill existing data (in batches)
UPDATE users
SET email_address = email
WHERE email_address IS NULL
AND id BETWEEN ? AND ?;
```
**Phase 4: Switch Reads**
```python
class UserService:
def get_user_email(self, user_id):
user = User.objects.get(id=user_id)
# Switch to reading from new column
return user.email_address or user.email
```
**Phase 5: Contract**
```sql
-- After validation, remove old column
ALTER TABLE users DROP COLUMN email;
-- Rename new column if needed
ALTER TABLE users RENAME COLUMN email_address TO email;
```
### 3. Online Schema Changes
#### PostgreSQL Techniques
```sql
-- Safe column addition
ALTER TABLE orders ADD COLUMN status_new VARCHAR(20) DEFAULT 'pending';
-- Safe index creation
CREATE INDEX CONCURRENTLY idx_orders_status_new ON orders(status_new);
-- Safe constraint addition (after data validation)
ALTER TABLE orders ADD CONSTRAINT check_status_new
CHECK (status_new IN ('pending', 'processing', 'completed', 'cancelled'));
```
#### MySQL Techniques
```sql
-- Use pt-online-schema-change for large tables
pt-online-schema-change \
--alter "ADD COLUMN status VARCHAR(20) DEFAULT 'pending'" \
--execute \
D=mydb,t=orders
-- Online DDL (MySQL 5.6+)
ALTER TABLE orders
ADD COLUMN priority INT DEFAULT 1,
ALGORITHM=INPLACE,
LOCK=NONE;
```
### 4. Data Migration Strategies
#### Chunked Data Migration
```python
class DataMigrator:
def __init__(self, source_table, target_table, chunk_size=1000):
self.source_table = source_table
self.target_table = target_table
self.chunk_size = chunk_size
def migrate_data(self):
last_id = 0
total_migrated = 0
while True:
# Get next chunk
chunk = self.get_chunk(last_id, self.chunk_size)
if not chunk:
break
# Transform and migrate chunk
for record in chunk:
transformed = self.transform_record(record)
self.insert_or_update(transformed)
last_id = chunk[-1]['id']
total_migrated += len(chunk)
# Brief pause to avoid overwhelming the database
time.sleep(0.1)
self.log_progress(total_migrated)
return total_migrated
def get_chunk(self, last_id, limit):
return db.execute(f"""
SELECT * FROM {self.source_table}
WHERE id > %s
ORDER BY id
LIMIT %s
""", (last_id, limit))
```
#### Change Data Capture (CDC)
```python
class CDCProcessor:
def __init__(self):
self.kafka_consumer = KafkaConsumer('db_changes')
self.target_db = TargetDatabase()
def process_changes(self):
for message in self.kafka_consumer:
change = json.loads(message.value)
if change['operation'] == 'INSERT':
self.handle_insert(change)
elif change['operation'] == 'UPDATE':
self.handle_update(change)
elif change['operation'] == 'DELETE':
self.handle_delete(change)
def handle_insert(self, change):
transformed_data = self.transform_data(change['after'])
self.target_db.insert(change['table'], transformed_data)
def handle_update(self, change):
key = change['key']
transformed_data = self.transform_data(change['after'])
self.target_db.update(change['table'], key, transformed_data)
```
## Application Zero-Downtime Techniques
### 1. Blue-Green Deployments
#### Infrastructure Setup
```yaml
# Blue Environment (Current Production)
apiVersion: apps/v1
kind: Deployment
metadata:
name: app-blue
labels:
version: blue
app: myapp
spec:
replicas: 3
selector:
matchLabels:
app: myapp
version: blue
template:
metadata:
labels:
app: myapp
version: blue
spec:
containers:
- name: app
image: myapp:1.0.0
ports:
- containerPort: 8080
readinessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 15
periodSeconds: 10
---
# Green Environment (New Version)
apiVersion: apps/v1
kind: Deployment
metadata:
name: app-green
labels:
version: green
app: myapp
spec:
replicas: 3
selector:
matchLabels:
app: myapp
version: green
template:
metadata:
labels:
app: myapp
version: green
spec:
containers:
- name: app
image: myapp:2.0.0
ports:
- containerPort: 8080
readinessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
```
#### Service Switching
```yaml
# Service (switches between blue and green)
apiVersion: v1
kind: Service
metadata:
name: app-service
spec:
selector:
app: myapp
version: blue # Switch to 'green' for deployment
ports:
- port: 80
targetPort: 8080
type: LoadBalancer
```
#### Automated Deployment Script
```bash
#!/bin/bash
# Blue-Green Deployment Script
NAMESPACE="production"
APP_NAME="myapp"
NEW_IMAGE="myapp:2.0.0"
# Determine current and target environments
CURRENT_VERSION=$(kubectl get service $APP_NAME-service -o jsonpath='{.spec.selector.version}')
if [ "$CURRENT_VERSION" = "blue" ]; then
TARGET_VERSION="green"
else
TARGET_VERSION="blue"
fi
echo "Current version: $CURRENT_VERSION"
echo "Target version: $TARGET_VERSION"
# Update target environment with new image
kubectl set image deployment/$APP_NAME-$TARGET_VERSION app=$NEW_IMAGE
# Wait for rollout to complete
kubectl rollout status deployment/$APP_NAME-$TARGET_VERSION --timeout=300s
# Run health checks
echo "Running health checks..."
TARGET_IP=$(kubectl get service $APP_NAME-$TARGET_VERSION -o jsonpath='{.status.loadBalancer.ingress[0].ip}')
for i in {1..30}; do
if curl -f http://$TARGET_IP/health; then
echo "Health check passed"
break
fi
if [ $i -eq 30 ]; then
echo "Health check failed after 30 attempts"
exit 1
fi
sleep 2
done
# Switch traffic to new version
kubectl patch service $APP_NAME-service -p '{"spec":{"selector":{"version":"'$TARGET_VERSION'"}}}'
echo "Traffic switched to $TARGET_VERSION"
# Monitor for 5 minutes
echo "Monitoring new version..."
sleep 300
# Check if rollback is needed
ERROR_RATE=$(curl -s "http://monitoring.company.com/api/error_rate?service=$APP_NAME" | jq '.error_rate')
if (( $(echo "$ERROR_RATE > 0.05" | bc -l) )); then
echo "Error rate too high ($ERROR_RATE), rolling back..."
kubectl patch service $APP_NAME-service -p '{"spec":{"selector":{"version":"'$CURRENT_VERSION'"}}}'
exit 1
fi
echo "Deployment successful!"
```
### 2. Canary Deployments
#### Progressive Canary with Istio
```yaml
# Destination Rule
apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata:
name: myapp-destination
spec:
host: myapp
subsets:
- name: v1
labels:
version: v1
- name: v2
labels:
version: v2
---
# Virtual Service for Canary
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
name: myapp-canary
spec:
hosts:
- myapp
http:
- match:
- headers:
canary:
exact: "true"
route:
- destination:
host: myapp
subset: v2
- route:
- destination:
host: myapp
subset: v1
weight: 95
- destination:
host: myapp
subset: v2
weight: 5
```
#### Automated Canary Controller
```python
class CanaryController:
def __init__(self, istio_client, prometheus_client):
self.istio = istio_client
self.prometheus = prometheus_client
self.canary_weight = 5
self.max_weight = 100
self.weight_increment = 5
self.validation_window = 300 # 5 minutes
async def deploy_canary(self, app_name, new_version):
"""Deploy new version using canary strategy"""
# Start with small percentage
await self.update_traffic_split(app_name, self.canary_weight)
while self.canary_weight < self.max_weight:
# Monitor metrics for validation window
await asyncio.sleep(self.validation_window)
# Check canary health
if not await self.is_canary_healthy(app_name, new_version):
await self.rollback_canary(app_name)
raise Exception("Canary deployment failed health checks")
# Increase traffic to canary
self.canary_weight = min(
self.canary_weight + self.weight_increment,
self.max_weight
)
await self.update_traffic_split(app_name, self.canary_weight)
print(f"Canary traffic increased to {self.canary_weight}%")
print("Canary deployment completed successfully")
async def is_canary_healthy(self, app_name, version):
"""Check if canary version is healthy"""
# Check error rate
error_rate = await self.prometheus.query(
f'rate(http_requests_total{{app="{app_name}", version="{version}", status=~"5.."}}'
f'[5m]) / rate(http_requests_total{{app="{app_name}", version="{version}"}}[5m])'
)
if error_rate > 0.05: # 5% error rate threshold
return False
# Check response time
p95_latency = await self.prometheus.query(
f'histogram_quantile(0.95, rate(http_request_duration_seconds_bucket'
f'{{app="{app_name}", version="{version}"}}[5m]))'
)
if p95_latency > 2.0: # 2 second p95 threshold
return False
return True
async def update_traffic_split(self, app_name, canary_weight):
"""Update Istio virtual service with new traffic split"""
stable_weight = 100 - canary_weight
virtual_service = {
"apiVersion": "networking.istio.io/v1beta1",
"kind": "VirtualService",
"metadata": {"name": f"{app_name}-canary"},
"spec": {
"hosts": [app_name],
"http": [{
"route": [
{
"destination": {"host": app_name, "subset": "stable"},
"weight": stable_weight
},
{
"destination": {"host": app_name, "subset": "canary"},
"weight": canary_weight
}
]
}]
}
}
await self.istio.apply_virtual_service(virtual_service)
```
### 3. Rolling Updates
#### Kubernetes Rolling Update Strategy
```yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: rolling-update-app
spec:
replicas: 10
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 2 # Can have 2 extra pods during update
maxUnavailable: 1 # At most 1 pod can be unavailable
selector:
matchLabels:
app: rolling-update-app
template:
metadata:
labels:
app: rolling-update-app
spec:
containers:
- name: app
image: myapp:2.0.0
ports:
- containerPort: 8080
readinessProbe:
httpGet:
path: /ready
port: 8080
initialDelaySeconds: 5
periodSeconds: 2
timeoutSeconds: 1
successThreshold: 1
failureThreshold: 3
livenessProbe:
httpGet:
path: /live
port: 8080
initialDelaySeconds: 10
periodSeconds: 10
```
#### Custom Rolling Update Controller
```python
class RollingUpdateController:
def __init__(self, k8s_client):
self.k8s = k8s_client
self.max_surge = 2
self.max_unavailable = 1
async def rolling_update(self, deployment_name, new_image):
"""Perform rolling update with custom logic"""
deployment = await self.k8s.get_deployment(deployment_name)
total_replicas = deployment.spec.replicas
# Calculate batch size
batch_size = min(self.max_surge, total_replicas // 5) # Update 20% at a time
updated_pods = []
for i in range(0, total_replicas, batch_size):
batch_end = min(i + batch_size, total_replicas)
# Update batch of pods
for pod_index in range(i, batch_end):
old_pod = await self.get_pod_by_index(deployment_name, pod_index)
# Create new pod with new image
new_pod = await self.create_updated_pod(old_pod, new_image)
# Wait for new pod to be ready
await self.wait_for_pod_ready(new_pod.metadata.name)
# Remove old pod
await self.k8s.delete_pod(old_pod.metadata.name)
updated_pods.append(new_pod)
# Brief pause between pod updates
await asyncio.sleep(2)
# Validate batch health before continuing
if not await self.validate_batch_health(updated_pods[-batch_size:]):
# Rollback batch
await self.rollback_batch(updated_pods[-batch_size:])
raise Exception("Rolling update failed validation")
print(f"Updated {batch_end}/{total_replicas} pods")
print("Rolling update completed successfully")
```
## Load Balancer and Traffic Management
### 1. Weighted Routing
#### NGINX Configuration
```nginx
upstream backend {
# Old version - 80% traffic
server old-app-1:8080 weight=4;
server old-app-2:8080 weight=4;
# New version - 20% traffic
server new-app-1:8080 weight=1;
server new-app-2:8080 weight=1;
}
server {
listen 80;
location / {
proxy_pass http://backend;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
# Health check headers
proxy_set_header X-Health-Check-Timeout 5s;
}
}
```
#### HAProxy Configuration
```haproxy
backend app_servers
balance roundrobin
option httpchk GET /health
# Old version servers
server old-app-1 old-app-1:8080 check weight 80
server old-app-2 old-app-2:8080 check weight 80
# New version servers
server new-app-1 new-app-1:8080 check weight 20
server new-app-2 new-app-2:8080 check weight 20
frontend app_frontend
bind *:80
default_backend app_servers
# Custom health check endpoint
acl health_check path_beg /health
http-request return status 200 content-type text/plain string "OK" if health_check
```
### 2. Circuit Breaker Implementation
```python
class CircuitBreaker:
def __init__(self, failure_threshold=5, recovery_timeout=60, expected_exception=Exception):
self.failure_threshold = failure_threshold
self.recovery_timeout = recovery_timeout
self.expected_exception = expected_exception
self.failure_count = 0
self.last_failure_time = None
self.state = 'CLOSED' # CLOSED, OPEN, HALF_OPEN
def call(self, func, *args, **kwargs):
"""Execute function with circuit breaker protection"""
if self.state == 'OPEN':
if self._should_attempt_reset():
self.state = 'HALF_OPEN'
else:
raise CircuitBreakerOpenException("Circuit breaker is OPEN")
try:
result = func(*args, **kwargs)
self._on_success()
return result
except self.expected_exception as e:
self._on_failure()
raise
def _should_attempt_reset(self):
return (
self.last_failure_time and
time.time() - self.last_failure_time >= self.recovery_timeout
)
def _on_success(self):
self.failure_count = 0
self.state = 'CLOSED'
def _on_failure(self):
self.failure_count += 1
self.last_failure_time = time.time()
if self.failure_count >= self.failure_threshold:
self.state = 'OPEN'
# Usage with service migration
@CircuitBreaker(failure_threshold=3, recovery_timeout=30)
def call_new_service(request):
return new_service.process(request)
def handle_request(request):
try:
return call_new_service(request)
except CircuitBreakerOpenException:
# Fallback to old service
return old_service.process(request)
```
## Monitoring and Validation
### 1. Health Check Implementation
```python
class HealthChecker:
def __init__(self):
self.checks = []
def add_check(self, name, check_func, timeout=5):
self.checks.append({
'name': name,
'func': check_func,
'timeout': timeout
})
async def run_checks(self):
"""Run all health checks and return status"""
results = {}
overall_status = 'healthy'
for check in self.checks:
try:
result = await asyncio.wait_for(
check['func'](),
timeout=check['timeout']
)
results[check['name']] = {
'status': 'healthy',
'result': result
}
except asyncio.TimeoutError:
results[check['name']] = {
'status': 'unhealthy',
'error': 'timeout'
}
overall_status = 'unhealthy'
except Exception as e:
results[check['name']] = {
'status': 'unhealthy',
'error': str(e)
}
overall_status = 'unhealthy'
return {
'status': overall_status,
'checks': results,
'timestamp': datetime.utcnow().isoformat()
}
# Example health checks
health_checker = HealthChecker()
async def database_check():
"""Check database connectivity"""
result = await db.execute("SELECT 1")
return result is not None
async def external_api_check():
"""Check external API availability"""
response = await http_client.get("https://api.example.com/health")
return response.status_code == 200
async def memory_check():
"""Check memory usage"""
memory_usage = psutil.virtual_memory().percent
if memory_usage > 90:
raise Exception(f"Memory usage too high: {memory_usage}%")
return f"Memory usage: {memory_usage}%"
health_checker.add_check("database", database_check)
health_checker.add_check("external_api", external_api_check)
health_checker.add_check("memory", memory_check)
```
### 2. Readiness vs Liveness Probes
```yaml
# Kubernetes Pod with proper health checks
apiVersion: v1
kind: Pod
metadata:
name: app-pod
spec:
containers:
- name: app
image: myapp:2.0.0
ports:
- containerPort: 8080
# Readiness probe - determines if pod should receive traffic
readinessProbe:
httpGet:
path: /ready
port: 8080
initialDelaySeconds: 5
periodSeconds: 3
timeoutSeconds: 2
successThreshold: 1
failureThreshold: 3
# Liveness probe - determines if pod should be restarted
livenessProbe:
httpGet:
path: /live
port: 8080
initialDelaySeconds: 30
periodSeconds: 10
timeoutSeconds: 5
successThreshold: 1
failureThreshold: 3
# Startup probe - gives app time to start before other probes
startupProbe:
httpGet:
path: /startup
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
timeoutSeconds: 3
successThreshold: 1
failureThreshold: 30 # Allow up to 150 seconds for startup
```
### 3. Metrics and Alerting
```python
class MigrationMetrics:
def __init__(self, prometheus_client):
self.prometheus = prometheus_client
# Define custom metrics
self.migration_progress = Counter(
'migration_progress_total',
'Total migration operations completed',
['operation', 'status']
)
self.migration_duration = Histogram(
'migration_operation_duration_seconds',
'Time spent on migration operations',
['operation']
)
self.system_health = Gauge(
'system_health_score',
'Overall system health score (0-1)',
['component']
)
self.traffic_split = Gauge(
'traffic_split_percentage',
'Percentage of traffic going to each version',
['version']
)
def record_migration_step(self, operation, status, duration=None):
"""Record completion of a migration step"""
self.migration_progress.labels(operation=operation, status=status).inc()
if duration:
self.migration_duration.labels(operation=operation).observe(duration)
def update_health_score(self, component, score):
"""Update health score for a component"""
self.system_health.labels(component=component).set(score)
def update_traffic_split(self, version_weights):
"""Update traffic split metrics"""
for version, weight in version_weights.items():
self.traffic_split.labels(version=version).set(weight)
# Usage in migration
metrics = MigrationMetrics(prometheus_client)
def perform_migration_step(operation):
start_time = time.time()
try:
# Perform migration operation
result = execute_migration_operation(operation)
# Record success
duration = time.time() - start_time
metrics.record_migration_step(operation, 'success', duration)
return result
except Exception as e:
# Record failure
duration = time.time() - start_time
metrics.record_migration_step(operation, 'failure', duration)
raise
```
## Rollback Strategies
### 1. Immediate Rollback Triggers
```python
class AutoRollbackSystem:
def __init__(self, metrics_client, deployment_client):
self.metrics = metrics_client
self.deployment = deployment_client
self.rollback_triggers = {
'error_rate_spike': {
'threshold': 0.05, # 5% error rate
'window': 300, # 5 minutes
'auto_rollback': True
},
'latency_increase': {
'threshold': 2.0, # 2x baseline latency
'window': 600, # 10 minutes
'auto_rollback': False # Manual confirmation required
},
'availability_drop': {
'threshold': 0.95, # Below 95% availability
'window': 120, # 2 minutes
'auto_rollback': True
}
}
async def monitor_and_rollback(self, deployment_name):
"""Monitor deployment and trigger rollback if needed"""
while True:
for trigger_name, config in self.rollback_triggers.items():
if await self.check_trigger(trigger_name, config):
if config['auto_rollback']:
await self.execute_rollback(deployment_name, trigger_name)
else:
await self.alert_for_manual_rollback(deployment_name, trigger_name)
await asyncio.sleep(30) # Check every 30 seconds
async def check_trigger(self, trigger_name, config):
"""Check if rollback trigger condition is met"""
current_value = await self.metrics.get_current_value(trigger_name)
baseline_value = await self.metrics.get_baseline_value(trigger_name)
if trigger_name == 'error_rate_spike':
return current_value > config['threshold']
elif trigger_name == 'latency_increase':
return current_value > baseline_value * config['threshold']
elif trigger_name == 'availability_drop':
return current_value < config['threshold']
return False
async def execute_rollback(self, deployment_name, reason):
"""Execute automatic rollback"""
print(f"Executing automatic rollback for {deployment_name}. Reason: {reason}")
# Get previous revision
previous_revision = await self.deployment.get_previous_revision(deployment_name)
# Perform rollback
await self.deployment.rollback_to_revision(deployment_name, previous_revision)
# Notify stakeholders
await self.notify_rollback_executed(deployment_name, reason)
```
### 2. Data Rollback Strategies
```sql
-- Point-in-time recovery setup
-- Create restore point before migration
SELECT pg_create_restore_point('pre_migration_' || to_char(now(), 'YYYYMMDD_HH24MISS'));
-- Rollback using point-in-time recovery
-- (This would be executed on a separate recovery instance)
-- recovery.conf:
-- recovery_target_name = 'pre_migration_20240101_120000'
-- recovery_target_action = 'promote'
```
```python
class DataRollbackManager:
def __init__(self, database_client, backup_service):
self.db = database_client
self.backup = backup_service
async def create_rollback_point(self, migration_id):
"""Create a rollback point before migration"""
rollback_point = {
'migration_id': migration_id,
'timestamp': datetime.utcnow(),
'backup_location': None,
'schema_snapshot': None
}
# Create database backup
backup_path = await self.backup.create_backup(
f"pre_migration_{migration_id}_{int(time.time())}"
)
rollback_point['backup_location'] = backup_path
# Capture schema snapshot
schema_snapshot = await self.capture_schema_snapshot()
rollback_point['schema_snapshot'] = schema_snapshot
# Store rollback point metadata
await self.store_rollback_metadata(rollback_point)
return rollback_point
async def execute_rollback(self, migration_id):
"""Execute data rollback to specified point"""
rollback_point = await self.get_rollback_metadata(migration_id)
if not rollback_point:
raise Exception(f"No rollback point found for migration {migration_id}")
# Stop application traffic
await self.stop_application_traffic()
try:
# Restore from backup
await self.backup.restore_from_backup(
rollback_point['backup_location']
)
# Validate data integrity
await self.validate_data_integrity(
rollback_point['schema_snapshot']
)
# Update application configuration
await self.update_application_config(rollback_point)
# Resume application traffic
await self.resume_application_traffic()
print(f"Data rollback completed successfully for migration {migration_id}")
except Exception as e:
# If rollback fails, we have a serious problem
await self.escalate_rollback_failure(migration_id, str(e))
raise
```
## Best Practices Summary
### 1. Pre-Migration Checklist
- [ ] Comprehensive backup strategy in place
- [ ] Rollback procedures tested in staging
- [ ] Monitoring and alerting configured
- [ ] Health checks implemented
- [ ] Feature flags configured
- [ ] Team communication plan established
- [ ] Load balancer configuration prepared
- [ ] Database connection pooling optimized
### 2. During Migration
- [ ] Monitor key metrics continuously
- [ ] Validate each phase before proceeding
- [ ] Maintain detailed logs of all actions
- [ ] Keep stakeholders informed of progress
- [ ] Have rollback trigger ready
- [ ] Monitor user experience metrics
- [ ] Watch for performance degradation
- [ ] Validate data consistency
### 3. Post-Migration
- [ ] Continue monitoring for 24-48 hours
- [ ] Validate all business processes
- [ ] Update documentation
- [ ] Conduct post-migration retrospective
- [ ] Archive migration artifacts
- [ ] Update disaster recovery procedures
- [ ] Plan for legacy system decommissioning
### 4. Common Pitfalls to Avoid
- Don't skip testing rollback procedures
- Don't ignore performance impact
- Don't rush through validation phases
- Don't forget to communicate with stakeholders
- Don't assume health checks are sufficient
- Don't neglect data consistency validation
- Don't underestimate time requirements
- Don't overlook dependency impacts
This comprehensive guide provides the foundation for implementing zero-downtime migrations across various system components while maintaining high availability and data integrity.
FILE:scripts/compatibility_checker.py
#!/usr/bin/env python3
"""
Compatibility Checker - Analyze schema and API compatibility between versions
This tool analyzes schema and API changes between versions and identifies backward
compatibility issues including breaking changes, data type mismatches, missing fields,
constraint violations, and generates migration scripts suggestions.
Author: Migration Architect Skill
Version: 1.0.0
License: MIT
"""
import json
import argparse
import sys
import re
import datetime
from typing import Dict, List, Any, Optional, Tuple, Set
from dataclasses import dataclass, asdict
from enum import Enum
class ChangeType(Enum):
"""Types of changes detected"""
BREAKING = "breaking"
POTENTIALLY_BREAKING = "potentially_breaking"
NON_BREAKING = "non_breaking"
ADDITIVE = "additive"
class CompatibilityLevel(Enum):
"""Compatibility assessment levels"""
FULLY_COMPATIBLE = "fully_compatible"
BACKWARD_COMPATIBLE = "backward_compatible"
POTENTIALLY_INCOMPATIBLE = "potentially_incompatible"
BREAKING_CHANGES = "breaking_changes"
@dataclass
class CompatibilityIssue:
"""Individual compatibility issue"""
type: str
severity: str
description: str
field_path: str
old_value: Any
new_value: Any
impact: str
suggested_migration: str
affected_operations: List[str]
@dataclass
class MigrationScript:
"""Migration script suggestion"""
script_type: str # sql, api, config
description: str
script_content: str
rollback_script: str
dependencies: List[str]
validation_query: str
@dataclass
class CompatibilityReport:
"""Complete compatibility analysis report"""
schema_before: str
schema_after: str
analysis_date: str
overall_compatibility: str
breaking_changes_count: int
potentially_breaking_count: int
non_breaking_changes_count: int
additive_changes_count: int
issues: List[CompatibilityIssue]
migration_scripts: List[MigrationScript]
risk_assessment: Dict[str, Any]
recommendations: List[str]
class SchemaCompatibilityChecker:
"""Main schema compatibility checker class"""
def __init__(self):
self.type_compatibility_matrix = self._build_type_compatibility_matrix()
self.constraint_implications = self._build_constraint_implications()
def _build_type_compatibility_matrix(self) -> Dict[str, Dict[str, str]]:
"""Build data type compatibility matrix"""
return {
# SQL data types compatibility
"varchar": {
"text": "compatible",
"char": "potentially_breaking", # length might be different
"nvarchar": "compatible",
"int": "breaking",
"bigint": "breaking",
"decimal": "breaking",
"datetime": "breaking",
"boolean": "breaking"
},
"int": {
"bigint": "compatible",
"smallint": "potentially_breaking", # range reduction
"decimal": "compatible",
"float": "potentially_breaking", # precision loss
"varchar": "breaking",
"boolean": "breaking"
},
"bigint": {
"int": "potentially_breaking", # range reduction
"decimal": "compatible",
"varchar": "breaking",
"boolean": "breaking"
},
"decimal": {
"float": "potentially_breaking", # precision loss
"int": "potentially_breaking", # precision loss
"bigint": "potentially_breaking", # precision loss
"varchar": "breaking",
"boolean": "breaking"
},
"datetime": {
"timestamp": "compatible",
"date": "potentially_breaking", # time component lost
"varchar": "breaking",
"int": "breaking"
},
"boolean": {
"tinyint": "compatible",
"varchar": "breaking",
"int": "breaking"
},
# JSON/API field types
"string": {
"number": "breaking",
"boolean": "breaking",
"array": "breaking",
"object": "breaking",
"null": "potentially_breaking"
},
"number": {
"string": "breaking",
"boolean": "breaking",
"array": "breaking",
"object": "breaking",
"null": "potentially_breaking"
},
"boolean": {
"string": "breaking",
"number": "breaking",
"array": "breaking",
"object": "breaking",
"null": "potentially_breaking"
},
"array": {
"string": "breaking",
"number": "breaking",
"boolean": "breaking",
"object": "breaking",
"null": "potentially_breaking"
},
"object": {
"string": "breaking",
"number": "breaking",
"boolean": "breaking",
"array": "breaking",
"null": "potentially_breaking"
}
}
def _build_constraint_implications(self) -> Dict[str, Dict[str, str]]:
"""Build constraint change implications"""
return {
"required": {
"added": "breaking", # Previously optional field now required
"removed": "non_breaking" # Previously required field now optional
},
"not_null": {
"added": "breaking", # Previously nullable now NOT NULL
"removed": "non_breaking" # Previously NOT NULL now nullable
},
"unique": {
"added": "potentially_breaking", # May fail if duplicates exist
"removed": "non_breaking" # No longer enforcing uniqueness
},
"primary_key": {
"added": "breaking", # Major structural change
"removed": "breaking", # Major structural change
"modified": "breaking" # Primary key change is always breaking
},
"foreign_key": {
"added": "potentially_breaking", # May fail if referential integrity violated
"removed": "potentially_breaking", # May allow orphaned records
"modified": "breaking" # Reference change is breaking
},
"check": {
"added": "potentially_breaking", # May fail if existing data violates check
"removed": "non_breaking", # No longer enforcing check
"modified": "potentially_breaking" # Different validation rules
},
"index": {
"added": "non_breaking", # Performance improvement
"removed": "non_breaking", # Performance impact only
"modified": "non_breaking" # Performance impact only
}
}
def analyze_database_schema(self, before_schema: Dict[str, Any],
after_schema: Dict[str, Any]) -> CompatibilityReport:
"""Analyze database schema compatibility"""
issues = []
migration_scripts = []
before_tables = before_schema.get("tables", {})
after_tables = after_schema.get("tables", {})
# Check for removed tables
for table_name in before_tables:
if table_name not in after_tables:
issues.append(CompatibilityIssue(
type="table_removed",
severity="breaking",
description=f"Table '{table_name}' has been removed",
field_path=f"tables.{table_name}",
old_value=before_tables[table_name],
new_value=None,
impact="All operations on this table will fail",
suggested_migration=f"CREATE VIEW {table_name} AS SELECT * FROM replacement_table;",
affected_operations=["SELECT", "INSERT", "UPDATE", "DELETE"]
))
# Check for added tables
for table_name in after_tables:
if table_name not in before_tables:
migration_scripts.append(MigrationScript(
script_type="sql",
description=f"Create new table {table_name}",
script_content=self._generate_create_table_sql(table_name, after_tables[table_name]),
rollback_script=f"DROP TABLE IF EXISTS {table_name};",
dependencies=[],
validation_query=f"SELECT COUNT(*) FROM information_schema.tables WHERE table_name = '{table_name}';"
))
# Check for modified tables
for table_name in set(before_tables.keys()) & set(after_tables.keys()):
table_issues, table_scripts = self._analyze_table_changes(
table_name, before_tables[table_name], after_tables[table_name]
)
issues.extend(table_issues)
migration_scripts.extend(table_scripts)
return self._build_compatibility_report(
before_schema, after_schema, issues, migration_scripts
)
def analyze_api_schema(self, before_schema: Dict[str, Any],
after_schema: Dict[str, Any]) -> CompatibilityReport:
"""Analyze REST API schema compatibility"""
issues = []
migration_scripts = []
# Analyze endpoints
before_paths = before_schema.get("paths", {})
after_paths = after_schema.get("paths", {})
# Check for removed endpoints
for path in before_paths:
if path not in after_paths:
for method in before_paths[path]:
issues.append(CompatibilityIssue(
type="endpoint_removed",
severity="breaking",
description=f"Endpoint {method.upper()} {path} has been removed",
field_path=f"paths.{path}.{method}",
old_value=before_paths[path][method],
new_value=None,
impact="Client requests to this endpoint will fail with 404",
suggested_migration=f"Implement redirect to replacement endpoint or maintain backward compatibility stub",
affected_operations=[f"{method.upper()} {path}"]
))
# Check for modified endpoints
for path in set(before_paths.keys()) & set(after_paths.keys()):
path_issues, path_scripts = self._analyze_endpoint_changes(
path, before_paths[path], after_paths[path]
)
issues.extend(path_issues)
migration_scripts.extend(path_scripts)
# Analyze data models
before_components = before_schema.get("components", {}).get("schemas", {})
after_components = after_schema.get("components", {}).get("schemas", {})
for model_name in set(before_components.keys()) & set(after_components.keys()):
model_issues, model_scripts = self._analyze_model_changes(
model_name, before_components[model_name], after_components[model_name]
)
issues.extend(model_issues)
migration_scripts.extend(model_scripts)
return self._build_compatibility_report(
before_schema, after_schema, issues, migration_scripts
)
def _analyze_table_changes(self, table_name: str, before_table: Dict[str, Any],
after_table: Dict[str, Any]) -> Tuple[List[CompatibilityIssue], List[MigrationScript]]:
"""Analyze changes to a specific table"""
issues = []
scripts = []
before_columns = before_table.get("columns", {})
after_columns = after_table.get("columns", {})
# Check for removed columns
for col_name in before_columns:
if col_name not in after_columns:
issues.append(CompatibilityIssue(
type="column_removed",
severity="breaking",
description=f"Column '{col_name}' removed from table '{table_name}'",
field_path=f"tables.{table_name}.columns.{col_name}",
old_value=before_columns[col_name],
new_value=None,
impact="SELECT statements including this column will fail",
suggested_migration=f"ALTER TABLE {table_name} ADD COLUMN {col_name}_deprecated AS computed_value;",
affected_operations=["SELECT", "INSERT", "UPDATE"]
))
# Check for added columns
for col_name in after_columns:
if col_name not in before_columns:
col_def = after_columns[col_name]
is_required = col_def.get("nullable", True) == False and col_def.get("default") is None
if is_required:
issues.append(CompatibilityIssue(
type="required_column_added",
severity="breaking",
description=f"Required column '{col_name}' added to table '{table_name}'",
field_path=f"tables.{table_name}.columns.{col_name}",
old_value=None,
new_value=col_def,
impact="INSERT statements without this column will fail",
suggested_migration=f"Add default value or make column nullable initially",
affected_operations=["INSERT"]
))
scripts.append(MigrationScript(
script_type="sql",
description=f"Add column {col_name} to table {table_name}",
script_content=f"ALTER TABLE {table_name} ADD COLUMN {self._generate_column_definition(col_name, col_def)};",
rollback_script=f"ALTER TABLE {table_name} DROP COLUMN {col_name};",
dependencies=[],
validation_query=f"SELECT COUNT(*) FROM information_schema.columns WHERE table_name = '{table_name}' AND column_name = '{col_name}';"
))
# Check for modified columns
for col_name in set(before_columns.keys()) & set(after_columns.keys()):
col_issues, col_scripts = self._analyze_column_changes(
table_name, col_name, before_columns[col_name], after_columns[col_name]
)
issues.extend(col_issues)
scripts.extend(col_scripts)
# Check constraint changes
before_constraints = before_table.get("constraints", {})
after_constraints = after_table.get("constraints", {})
constraint_issues, constraint_scripts = self._analyze_constraint_changes(
table_name, before_constraints, after_constraints
)
issues.extend(constraint_issues)
scripts.extend(constraint_scripts)
return issues, scripts
def _analyze_column_changes(self, table_name: str, col_name: str,
before_col: Dict[str, Any], after_col: Dict[str, Any]) -> Tuple[List[CompatibilityIssue], List[MigrationScript]]:
"""Analyze changes to a specific column"""
issues = []
scripts = []
# Check data type changes
before_type = before_col.get("type", "").lower()
after_type = after_col.get("type", "").lower()
if before_type != after_type:
compatibility = self.type_compatibility_matrix.get(before_type, {}).get(after_type, "breaking")
if compatibility == "breaking":
issues.append(CompatibilityIssue(
type="incompatible_type_change",
severity="breaking",
description=f"Column '{col_name}' type changed from {before_type} to {after_type}",
field_path=f"tables.{table_name}.columns.{col_name}.type",
old_value=before_type,
new_value=after_type,
impact="Data conversion may fail or lose precision",
suggested_migration=f"Add conversion logic and validate data integrity",
affected_operations=["SELECT", "INSERT", "UPDATE", "WHERE clauses"]
))
scripts.append(MigrationScript(
script_type="sql",
description=f"Convert column {col_name} from {before_type} to {after_type}",
script_content=f"ALTER TABLE {table_name} ALTER COLUMN {col_name} TYPE {after_type} USING {col_name}::{after_type};",
rollback_script=f"ALTER TABLE {table_name} ALTER COLUMN {col_name} TYPE {before_type};",
dependencies=[f"backup_{table_name}"],
validation_query=f"SELECT COUNT(*) FROM {table_name} WHERE {col_name} IS NOT NULL;"
))
elif compatibility == "potentially_breaking":
issues.append(CompatibilityIssue(
type="risky_type_change",
severity="potentially_breaking",
description=f"Column '{col_name}' type changed from {before_type} to {after_type} - may lose data",
field_path=f"tables.{table_name}.columns.{col_name}.type",
old_value=before_type,
new_value=after_type,
impact="Potential data loss or precision reduction",
suggested_migration=f"Validate all existing data can be converted safely",
affected_operations=["Data integrity"]
))
# Check nullability changes
before_nullable = before_col.get("nullable", True)
after_nullable = after_col.get("nullable", True)
if before_nullable != after_nullable:
if before_nullable and not after_nullable: # null -> not null
issues.append(CompatibilityIssue(
type="nullability_restriction",
severity="breaking",
description=f"Column '{col_name}' changed from nullable to NOT NULL",
field_path=f"tables.{table_name}.columns.{col_name}.nullable",
old_value=before_nullable,
new_value=after_nullable,
impact="Existing NULL values will cause constraint violations",
suggested_migration=f"Update NULL values to valid defaults before applying NOT NULL constraint",
affected_operations=["INSERT", "UPDATE"]
))
scripts.append(MigrationScript(
script_type="sql",
description=f"Make column {col_name} NOT NULL",
script_content=f"""
-- Update NULL values first
UPDATE {table_name} SET {col_name} = 'DEFAULT_VALUE' WHERE {col_name} IS NULL;
-- Add NOT NULL constraint
ALTER TABLE {table_name} ALTER COLUMN {col_name} SET NOT NULL;
""",
rollback_script=f"ALTER TABLE {table_name} ALTER COLUMN {col_name} DROP NOT NULL;",
dependencies=[],
validation_query=f"SELECT COUNT(*) FROM {table_name} WHERE {col_name} IS NULL;"
))
# Check length/precision changes
before_length = before_col.get("length")
after_length = after_col.get("length")
if before_length and after_length and before_length != after_length:
if after_length < before_length:
issues.append(CompatibilityIssue(
type="length_reduction",
severity="potentially_breaking",
description=f"Column '{col_name}' length reduced from {before_length} to {after_length}",
field_path=f"tables.{table_name}.columns.{col_name}.length",
old_value=before_length,
new_value=after_length,
impact="Data truncation may occur for values exceeding new length",
suggested_migration=f"Validate no existing data exceeds new length limit",
affected_operations=["INSERT", "UPDATE"]
))
return issues, scripts
def _analyze_constraint_changes(self, table_name: str, before_constraints: Dict[str, Any],
after_constraints: Dict[str, Any]) -> Tuple[List[CompatibilityIssue], List[MigrationScript]]:
"""Analyze constraint changes"""
issues = []
scripts = []
for constraint_type in ["primary_key", "foreign_key", "unique", "check"]:
before_constraint = before_constraints.get(constraint_type, [])
after_constraint = after_constraints.get(constraint_type, [])
# Convert to sets for comparison
before_set = set(str(c) for c in before_constraint) if isinstance(before_constraint, list) else {str(before_constraint)} if before_constraint else set()
after_set = set(str(c) for c in after_constraint) if isinstance(after_constraint, list) else {str(after_constraint)} if after_constraint else set()
# Check for removed constraints
for constraint in before_set - after_set:
implication = self.constraint_implications.get(constraint_type, {}).get("removed", "non_breaking")
issues.append(CompatibilityIssue(
type=f"{constraint_type}_removed",
severity=implication,
description=f"{constraint_type.replace('_', ' ').title()} constraint '{constraint}' removed from table '{table_name}'",
field_path=f"tables.{table_name}.constraints.{constraint_type}",
old_value=constraint,
new_value=None,
impact=f"No longer enforcing {constraint_type} constraint",
suggested_migration=f"Consider application-level validation for removed constraint",
affected_operations=["INSERT", "UPDATE", "DELETE"]
))
# Check for added constraints
for constraint in after_set - before_set:
implication = self.constraint_implications.get(constraint_type, {}).get("added", "potentially_breaking")
issues.append(CompatibilityIssue(
type=f"{constraint_type}_added",
severity=implication,
description=f"New {constraint_type.replace('_', ' ')} constraint '{constraint}' added to table '{table_name}'",
field_path=f"tables.{table_name}.constraints.{constraint_type}",
old_value=None,
new_value=constraint,
impact=f"New {constraint_type} constraint may reject existing data",
suggested_migration=f"Validate existing data complies with new constraint",
affected_operations=["INSERT", "UPDATE"]
))
scripts.append(MigrationScript(
script_type="sql",
description=f"Add {constraint_type} constraint to {table_name}",
script_content=f"ALTER TABLE {table_name} ADD CONSTRAINT {constraint_type}_{table_name} {constraint_type.upper()} ({constraint});",
rollback_script=f"ALTER TABLE {table_name} DROP CONSTRAINT {constraint_type}_{table_name};",
dependencies=[],
validation_query=f"SELECT COUNT(*) FROM information_schema.table_constraints WHERE table_name = '{table_name}' AND constraint_type = '{constraint_type.upper()}';"
))
return issues, scripts
def _analyze_endpoint_changes(self, path: str, before_endpoint: Dict[str, Any],
after_endpoint: Dict[str, Any]) -> Tuple[List[CompatibilityIssue], List[MigrationScript]]:
"""Analyze changes to an API endpoint"""
issues = []
scripts = []
for method in set(before_endpoint.keys()) & set(after_endpoint.keys()):
before_method = before_endpoint[method]
after_method = after_endpoint[method]
# Check parameter changes
before_params = before_method.get("parameters", [])
after_params = after_method.get("parameters", [])
before_param_names = {p["name"] for p in before_params}
after_param_names = {p["name"] for p in after_params}
# Check for removed required parameters
for param_name in before_param_names - after_param_names:
param = next(p for p in before_params if p["name"] == param_name)
if param.get("required", False):
issues.append(CompatibilityIssue(
type="required_parameter_removed",
severity="breaking",
description=f"Required parameter '{param_name}' removed from {method.upper()} {path}",
field_path=f"paths.{path}.{method}.parameters",
old_value=param,
new_value=None,
impact="Client requests with this parameter will fail",
suggested_migration="Implement parameter validation with backward compatibility",
affected_operations=[f"{method.upper()} {path}"]
))
# Check for added required parameters
for param_name in after_param_names - before_param_names:
param = next(p for p in after_params if p["name"] == param_name)
if param.get("required", False):
issues.append(CompatibilityIssue(
type="required_parameter_added",
severity="breaking",
description=f"New required parameter '{param_name}' added to {method.upper()} {path}",
field_path=f"paths.{path}.{method}.parameters",
old_value=None,
new_value=param,
impact="Client requests without this parameter will fail",
suggested_migration="Provide default value or make parameter optional initially",
affected_operations=[f"{method.upper()} {path}"]
))
# Check response schema changes
before_responses = before_method.get("responses", {})
after_responses = after_method.get("responses", {})
for status_code in before_responses:
if status_code in after_responses:
before_schema = before_responses[status_code].get("content", {}).get("application/json", {}).get("schema", {})
after_schema = after_responses[status_code].get("content", {}).get("application/json", {}).get("schema", {})
if before_schema != after_schema:
issues.append(CompatibilityIssue(
type="response_schema_changed",
severity="potentially_breaking",
description=f"Response schema changed for {method.upper()} {path} (status {status_code})",
field_path=f"paths.{path}.{method}.responses.{status_code}",
old_value=before_schema,
new_value=after_schema,
impact="Client response parsing may fail",
suggested_migration="Implement versioned API responses",
affected_operations=[f"{method.upper()} {path}"]
))
return issues, scripts
def _analyze_model_changes(self, model_name: str, before_model: Dict[str, Any],
after_model: Dict[str, Any]) -> Tuple[List[CompatibilityIssue], List[MigrationScript]]:
"""Analyze changes to an API data model"""
issues = []
scripts = []
before_props = before_model.get("properties", {})
after_props = after_model.get("properties", {})
before_required = set(before_model.get("required", []))
after_required = set(after_model.get("required", []))
# Check for removed properties
for prop_name in set(before_props.keys()) - set(after_props.keys()):
issues.append(CompatibilityIssue(
type="property_removed",
severity="breaking",
description=f"Property '{prop_name}' removed from model '{model_name}'",
field_path=f"components.schemas.{model_name}.properties.{prop_name}",
old_value=before_props[prop_name],
new_value=None,
impact="Client code expecting this property will fail",
suggested_migration="Use API versioning to maintain backward compatibility",
affected_operations=["Serialization", "Deserialization"]
))
# Check for newly required properties
for prop_name in after_required - before_required:
issues.append(CompatibilityIssue(
type="property_made_required",
severity="breaking",
description=f"Property '{prop_name}' is now required in model '{model_name}'",
field_path=f"components.schemas.{model_name}.required",
old_value=list(before_required),
new_value=list(after_required),
impact="Client requests without this property will fail validation",
suggested_migration="Provide default values or implement gradual rollout",
affected_operations=["Request validation"]
))
# Check for property type changes
for prop_name in set(before_props.keys()) & set(after_props.keys()):
before_type = before_props[prop_name].get("type")
after_type = after_props[prop_name].get("type")
if before_type != after_type:
compatibility = self.type_compatibility_matrix.get(before_type, {}).get(after_type, "breaking")
issues.append(CompatibilityIssue(
type="property_type_changed",
severity=compatibility,
description=f"Property '{prop_name}' type changed from {before_type} to {after_type} in model '{model_name}'",
field_path=f"components.schemas.{model_name}.properties.{prop_name}.type",
old_value=before_type,
new_value=after_type,
impact="Client serialization/deserialization may fail",
suggested_migration="Implement type coercion or API versioning",
affected_operations=["Serialization", "Deserialization"]
))
return issues, scripts
def _build_compatibility_report(self, before_schema: Dict[str, Any], after_schema: Dict[str, Any],
issues: List[CompatibilityIssue], migration_scripts: List[MigrationScript]) -> CompatibilityReport:
"""Build the final compatibility report"""
# Count issues by severity
breaking_count = sum(1 for issue in issues if issue.severity == "breaking")
potentially_breaking_count = sum(1 for issue in issues if issue.severity == "potentially_breaking")
non_breaking_count = sum(1 for issue in issues if issue.severity == "non_breaking")
additive_count = sum(1 for issue in issues if issue.type == "additive")
# Determine overall compatibility
if breaking_count > 0:
overall_compatibility = "breaking_changes"
elif potentially_breaking_count > 0:
overall_compatibility = "potentially_incompatible"
elif non_breaking_count > 0:
overall_compatibility = "backward_compatible"
else:
overall_compatibility = "fully_compatible"
# Generate risk assessment
risk_assessment = {
"overall_risk": "high" if breaking_count > 0 else "medium" if potentially_breaking_count > 0 else "low",
"deployment_risk": "requires_coordinated_deployment" if breaking_count > 0 else "safe_independent_deployment",
"rollback_complexity": "high" if breaking_count > 3 else "medium" if breaking_count > 0 else "low",
"testing_requirements": ["integration_testing", "regression_testing"] +
(["data_migration_testing"] if any(s.script_type == "sql" for s in migration_scripts) else [])
}
# Generate recommendations
recommendations = []
if breaking_count > 0:
recommendations.append("Implement API versioning to maintain backward compatibility")
recommendations.append("Plan for coordinated deployment with all clients")
recommendations.append("Implement comprehensive rollback procedures")
if potentially_breaking_count > 0:
recommendations.append("Conduct thorough testing with realistic data volumes")
recommendations.append("Implement monitoring for migration success metrics")
if migration_scripts:
recommendations.append("Test all migration scripts in staging environment")
recommendations.append("Implement migration progress monitoring")
recommendations.append("Create detailed communication plan for stakeholders")
recommendations.append("Implement feature flags for gradual rollout")
return CompatibilityReport(
schema_before=json.dumps(before_schema, indent=2)[:500] + "..." if len(json.dumps(before_schema)) > 500 else json.dumps(before_schema, indent=2),
schema_after=json.dumps(after_schema, indent=2)[:500] + "..." if len(json.dumps(after_schema)) > 500 else json.dumps(after_schema, indent=2),
analysis_date=datetime.datetime.now().isoformat(),
overall_compatibility=overall_compatibility,
breaking_changes_count=breaking_count,
potentially_breaking_count=potentially_breaking_count,
non_breaking_changes_count=non_breaking_count,
additive_changes_count=additive_count,
issues=issues,
migration_scripts=migration_scripts,
risk_assessment=risk_assessment,
recommendations=recommendations
)
def _generate_create_table_sql(self, table_name: str, table_def: Dict[str, Any]) -> str:
"""Generate CREATE TABLE SQL statement"""
columns = []
for col_name, col_def in table_def.get("columns", {}).items():
columns.append(self._generate_column_definition(col_name, col_def))
return f"CREATE TABLE {table_name} (\n " + ",\n ".join(columns) + "\n);"
def _generate_column_definition(self, col_name: str, col_def: Dict[str, Any]) -> str:
"""Generate column definition for SQL"""
col_type = col_def.get("type", "VARCHAR(255)")
nullable = "" if col_def.get("nullable", True) else " NOT NULL"
default = f" DEFAULT {col_def.get('default')}" if col_def.get("default") is not None else ""
return f"{col_name} {col_type}{nullable}{default}"
def generate_human_readable_report(self, report: CompatibilityReport) -> str:
"""Generate human-readable compatibility report"""
output = []
output.append("=" * 80)
output.append("COMPATIBILITY ANALYSIS REPORT")
output.append("=" * 80)
output.append(f"Analysis Date: {report.analysis_date}")
output.append(f"Overall Compatibility: {report.overall_compatibility.upper()}")
output.append("")
# Summary
output.append("SUMMARY")
output.append("-" * 40)
output.append(f"Breaking Changes: {report.breaking_changes_count}")
output.append(f"Potentially Breaking: {report.potentially_breaking_count}")
output.append(f"Non-Breaking Changes: {report.non_breaking_changes_count}")
output.append(f"Additive Changes: {report.additive_changes_count}")
output.append(f"Total Issues Found: {len(report.issues)}")
output.append("")
# Risk Assessment
output.append("RISK ASSESSMENT")
output.append("-" * 40)
for key, value in report.risk_assessment.items():
output.append(f"{key.replace('_', ' ').title()}: {value}")
output.append("")
# Issues by Severity
issues_by_severity = {}
for issue in report.issues:
if issue.severity not in issues_by_severity:
issues_by_severity[issue.severity] = []
issues_by_severity[issue.severity].append(issue)
for severity in ["breaking", "potentially_breaking", "non_breaking"]:
if severity in issues_by_severity:
output.append(f"{severity.upper().replace('_', ' ')} ISSUES")
output.append("-" * 40)
for issue in issues_by_severity[severity]:
output.append(f"• {issue.description}")
output.append(f" Field: {issue.field_path}")
output.append(f" Impact: {issue.impact}")
output.append(f" Migration: {issue.suggested_migration}")
if issue.affected_operations:
output.append(f" Affected Operations: {', '.join(issue.affected_operations)}")
output.append("")
# Migration Scripts
if report.migration_scripts:
output.append("SUGGESTED MIGRATION SCRIPTS")
output.append("-" * 40)
for i, script in enumerate(report.migration_scripts, 1):
output.append(f"{i}. {script.description}")
output.append(f" Type: {script.script_type}")
output.append(" Script:")
for line in script.script_content.split('\n'):
output.append(f" {line}")
output.append("")
# Recommendations
output.append("RECOMMENDATIONS")
output.append("-" * 40)
for i, rec in enumerate(report.recommendations, 1):
output.append(f"{i}. {rec}")
output.append("")
return "\n".join(output)
def main():
"""Main function with command line interface"""
parser = argparse.ArgumentParser(description="Analyze schema and API compatibility between versions")
parser.add_argument("--before", required=True, help="Before schema file (JSON)")
parser.add_argument("--after", required=True, help="After schema file (JSON)")
parser.add_argument("--type", choices=["database", "api"], default="database", help="Schema type to analyze")
parser.add_argument("--output", "-o", help="Output file for compatibility report (JSON)")
parser.add_argument("--format", "-f", choices=["json", "text", "both"], default="both", help="Output format")
args = parser.parse_args()
try:
# Load schemas
with open(args.before, 'r') as f:
before_schema = json.load(f)
with open(args.after, 'r') as f:
after_schema = json.load(f)
# Analyze compatibility
checker = SchemaCompatibilityChecker()
if args.type == "database":
report = checker.analyze_database_schema(before_schema, after_schema)
else: # api
report = checker.analyze_api_schema(before_schema, after_schema)
# Output results
if args.format in ["json", "both"]:
report_dict = asdict(report)
if args.output:
with open(args.output, 'w') as f:
json.dump(report_dict, f, indent=2)
print(f"Compatibility report saved to {args.output}")
else:
print(json.dumps(report_dict, indent=2))
if args.format in ["text", "both"]:
human_report = checker.generate_human_readable_report(report)
text_output = args.output.replace('.json', '.txt') if args.output else None
if text_output:
with open(text_output, 'w') as f:
f.write(human_report)
print(f"Human-readable report saved to {text_output}")
else:
print("\n" + "="*80)
print("HUMAN-READABLE COMPATIBILITY REPORT")
print("="*80)
print(human_report)
# Return exit code based on compatibility
if report.breaking_changes_count > 0:
return 2 # Breaking changes found
elif report.potentially_breaking_count > 0:
return 1 # Potentially breaking changes found
else:
return 0 # No compatibility issues
except FileNotFoundError as e:
print(f"Error: File not found: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON: {e}", file=sys.stderr)
return 1
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
return 1
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/migration_planner.py
#!/usr/bin/env python3
"""
Migration Planner - Generate comprehensive migration plans with risk assessment
This tool analyzes migration specifications and generates detailed, phased migration plans
including pre-migration checklists, validation gates, rollback triggers, timeline estimates,
and risk matrices.
Author: Migration Architect Skill
Version: 1.0.0
License: MIT
"""
import json
import argparse
import sys
import datetime
import hashlib
import math
from typing import Dict, List, Any, Optional, Tuple
from dataclasses import dataclass, asdict
from enum import Enum
class MigrationType(Enum):
"""Migration type enumeration"""
DATABASE = "database"
SERVICE = "service"
INFRASTRUCTURE = "infrastructure"
DATA = "data"
API = "api"
class MigrationComplexity(Enum):
"""Migration complexity levels"""
LOW = "low"
MEDIUM = "medium"
HIGH = "high"
CRITICAL = "critical"
class RiskLevel(Enum):
"""Risk assessment levels"""
LOW = "low"
MEDIUM = "medium"
HIGH = "high"
CRITICAL = "critical"
@dataclass
class MigrationConstraint:
"""Migration constraint definition"""
type: str
description: str
impact: str
mitigation: str
@dataclass
class MigrationPhase:
"""Individual migration phase"""
name: str
description: str
duration_hours: int
dependencies: List[str]
validation_criteria: List[str]
rollback_triggers: List[str]
tasks: List[str]
risk_level: str
resources_required: List[str]
@dataclass
class RiskItem:
"""Individual risk assessment item"""
category: str
description: str
probability: str # low, medium, high
impact: str # low, medium, high
severity: str # low, medium, high, critical
mitigation: str
owner: str
@dataclass
class MigrationPlan:
"""Complete migration plan structure"""
migration_id: str
source_system: str
target_system: str
migration_type: str
complexity: str
estimated_duration_hours: int
phases: List[MigrationPhase]
risks: List[RiskItem]
success_criteria: List[str]
rollback_plan: Dict[str, Any]
stakeholders: List[str]
created_at: str
class MigrationPlanner:
"""Main migration planner class"""
def __init__(self):
self.migration_patterns = self._load_migration_patterns()
self.risk_templates = self._load_risk_templates()
def _load_migration_patterns(self) -> Dict[str, Any]:
"""Load predefined migration patterns"""
return {
"database": {
"schema_change": {
"phases": ["preparation", "expand", "migrate", "contract", "cleanup"],
"base_duration": 24,
"complexity_multiplier": {"low": 1.0, "medium": 1.5, "high": 2.5, "critical": 4.0}
},
"data_migration": {
"phases": ["assessment", "setup", "bulk_copy", "delta_sync", "validation", "cutover"],
"base_duration": 48,
"complexity_multiplier": {"low": 1.2, "medium": 2.0, "high": 3.0, "critical": 5.0}
}
},
"service": {
"strangler_fig": {
"phases": ["intercept", "implement", "redirect", "validate", "retire"],
"base_duration": 168, # 1 week
"complexity_multiplier": {"low": 0.8, "medium": 1.0, "high": 1.8, "critical": 3.0}
},
"parallel_run": {
"phases": ["setup", "deploy", "shadow", "compare", "cutover", "cleanup"],
"base_duration": 72,
"complexity_multiplier": {"low": 1.0, "medium": 1.3, "high": 2.0, "critical": 3.5}
}
},
"infrastructure": {
"cloud_migration": {
"phases": ["assessment", "design", "pilot", "migration", "optimization", "decommission"],
"base_duration": 720, # 30 days
"complexity_multiplier": {"low": 0.6, "medium": 1.0, "high": 1.5, "critical": 2.5}
},
"on_prem_to_cloud": {
"phases": ["discovery", "planning", "pilot", "migration", "validation", "cutover"],
"base_duration": 480, # 20 days
"complexity_multiplier": {"low": 0.8, "medium": 1.2, "high": 2.0, "critical": 3.0}
}
}
}
def _load_risk_templates(self) -> Dict[str, List[RiskItem]]:
"""Load risk templates for different migration types"""
return {
"database": [
RiskItem("technical", "Data corruption during migration", "low", "critical", "high",
"Implement comprehensive backup and validation procedures", "DBA Team"),
RiskItem("technical", "Extended downtime due to migration complexity", "medium", "high", "high",
"Use blue-green deployment and phased migration approach", "DevOps Team"),
RiskItem("business", "Business process disruption", "medium", "high", "high",
"Communicate timeline and provide alternate workflows", "Business Owner"),
RiskItem("operational", "Insufficient rollback testing", "high", "critical", "critical",
"Execute full rollback procedures in staging environment", "QA Team")
],
"service": [
RiskItem("technical", "Service compatibility issues", "medium", "high", "high",
"Implement comprehensive integration testing", "Development Team"),
RiskItem("technical", "Performance degradation", "medium", "medium", "medium",
"Conduct load testing and performance benchmarking", "DevOps Team"),
RiskItem("business", "Feature parity gaps", "high", "high", "high",
"Document feature mapping and acceptance criteria", "Product Owner"),
RiskItem("operational", "Monitoring gap during transition", "medium", "medium", "medium",
"Set up dual monitoring and alerting systems", "SRE Team")
],
"infrastructure": [
RiskItem("technical", "Network connectivity issues", "medium", "critical", "high",
"Implement redundant network paths and monitoring", "Network Team"),
RiskItem("technical", "Security configuration drift", "high", "high", "high",
"Automated security scanning and compliance checks", "Security Team"),
RiskItem("business", "Cost overrun during transition", "high", "medium", "medium",
"Implement cost monitoring and budget alerts", "Finance Team"),
RiskItem("operational", "Team knowledge gaps", "high", "medium", "medium",
"Provide training and create detailed documentation", "Platform Team")
]
}
def _calculate_complexity(self, spec: Dict[str, Any]) -> str:
"""Calculate migration complexity based on specification"""
complexity_score = 0
# Data volume complexity
data_volume = spec.get("constraints", {}).get("data_volume_gb", 0)
if data_volume > 10000:
complexity_score += 3
elif data_volume > 1000:
complexity_score += 2
elif data_volume > 100:
complexity_score += 1
# System dependencies
dependencies = len(spec.get("constraints", {}).get("dependencies", []))
if dependencies > 10:
complexity_score += 3
elif dependencies > 5:
complexity_score += 2
elif dependencies > 2:
complexity_score += 1
# Downtime constraints
max_downtime = spec.get("constraints", {}).get("max_downtime_minutes", 480)
if max_downtime < 60:
complexity_score += 3
elif max_downtime < 240:
complexity_score += 2
elif max_downtime < 480:
complexity_score += 1
# Special requirements
special_reqs = spec.get("constraints", {}).get("special_requirements", [])
complexity_score += len(special_reqs)
if complexity_score >= 8:
return "critical"
elif complexity_score >= 5:
return "high"
elif complexity_score >= 3:
return "medium"
else:
return "low"
def _estimate_duration(self, migration_type: str, migration_pattern: str, complexity: str) -> int:
"""Estimate migration duration based on type, pattern, and complexity"""
pattern_info = self.migration_patterns.get(migration_type, {}).get(migration_pattern, {})
base_duration = pattern_info.get("base_duration", 48)
multiplier = pattern_info.get("complexity_multiplier", {}).get(complexity, 1.5)
return int(base_duration * multiplier)
def _generate_phases(self, spec: Dict[str, Any]) -> List[MigrationPhase]:
"""Generate migration phases based on specification"""
migration_type = spec.get("type")
migration_pattern = spec.get("pattern", "")
complexity = self._calculate_complexity(spec)
pattern_info = self.migration_patterns.get(migration_type, {})
if migration_pattern in pattern_info:
phase_names = pattern_info[migration_pattern]["phases"]
else:
# Default phases based on migration type
phase_names = {
"database": ["preparation", "migration", "validation", "cutover"],
"service": ["preparation", "deployment", "testing", "cutover"],
"infrastructure": ["assessment", "preparation", "migration", "validation"]
}.get(migration_type, ["preparation", "execution", "validation", "cleanup"])
phases = []
total_duration = self._estimate_duration(migration_type, migration_pattern, complexity)
phase_duration = total_duration // len(phase_names)
for i, phase_name in enumerate(phase_names):
phase = self._create_phase(phase_name, phase_duration, complexity, i, phase_names)
phases.append(phase)
return phases
def _create_phase(self, phase_name: str, duration: int, complexity: str,
phase_index: int, all_phases: List[str]) -> MigrationPhase:
"""Create a detailed migration phase"""
phase_templates = {
"preparation": {
"description": "Prepare systems and teams for migration",
"tasks": [
"Backup source system",
"Set up monitoring and alerting",
"Prepare rollback procedures",
"Communicate migration timeline",
"Validate prerequisites"
],
"validation_criteria": [
"All backups completed successfully",
"Monitoring systems operational",
"Team members briefed and ready",
"Rollback procedures tested"
],
"risk_level": "medium"
},
"assessment": {
"description": "Assess current state and migration requirements",
"tasks": [
"Inventory existing systems and dependencies",
"Analyze data volumes and complexity",
"Identify integration points",
"Document current architecture",
"Create migration mapping"
],
"validation_criteria": [
"Complete system inventory documented",
"Dependencies mapped and validated",
"Migration scope clearly defined",
"Resource requirements identified"
],
"risk_level": "low"
},
"migration": {
"description": "Execute core migration processes",
"tasks": [
"Begin data/service migration",
"Monitor migration progress",
"Validate data consistency",
"Handle migration errors",
"Update configuration"
],
"validation_criteria": [
"Migration progress within expected parameters",
"Data consistency checks passing",
"Error rates within acceptable limits",
"Performance metrics stable"
],
"risk_level": "high"
},
"validation": {
"description": "Validate migration success and system health",
"tasks": [
"Execute comprehensive testing",
"Validate business processes",
"Check system performance",
"Verify data integrity",
"Confirm security controls"
],
"validation_criteria": [
"All critical tests passing",
"Performance within acceptable range",
"Security controls functioning",
"Business processes operational"
],
"risk_level": "medium"
},
"cutover": {
"description": "Switch production traffic to new system",
"tasks": [
"Update DNS/load balancer configuration",
"Redirect production traffic",
"Monitor system performance",
"Validate end-user experience",
"Confirm business operations"
],
"validation_criteria": [
"Traffic successfully redirected",
"System performance stable",
"User experience satisfactory",
"Business operations normal"
],
"risk_level": "critical"
}
}
template = phase_templates.get(phase_name, {
"description": f"Execute {phase_name} phase",
"tasks": [f"Complete {phase_name} activities"],
"validation_criteria": [f"{phase_name.title()} phase completed successfully"],
"risk_level": "medium"
})
dependencies = []
if phase_index > 0:
dependencies.append(all_phases[phase_index - 1])
rollback_triggers = [
"Critical system failure",
"Data corruption detected",
"Performance degradation > 50%",
"Business process failure"
]
resources_required = [
"Technical team availability",
"System access and permissions",
"Monitoring and alerting systems",
"Communication channels"
]
return MigrationPhase(
name=phase_name,
description=template["description"],
duration_hours=duration,
dependencies=dependencies,
validation_criteria=template["validation_criteria"],
rollback_triggers=rollback_triggers,
tasks=template["tasks"],
risk_level=template["risk_level"],
resources_required=resources_required
)
def _assess_risks(self, spec: Dict[str, Any]) -> List[RiskItem]:
"""Generate risk assessment for migration"""
migration_type = spec.get("type")
base_risks = self.risk_templates.get(migration_type, [])
# Add specification-specific risks
additional_risks = []
constraints = spec.get("constraints", {})
if constraints.get("max_downtime_minutes", 480) < 60:
additional_risks.append(
RiskItem("business", "Zero-downtime requirement increases complexity", "high", "medium", "high",
"Implement blue-green deployment or rolling update strategy", "DevOps Team")
)
if constraints.get("data_volume_gb", 0) > 5000:
additional_risks.append(
RiskItem("technical", "Large data volumes may cause extended migration time", "high", "medium", "medium",
"Implement parallel processing and progress monitoring", "Data Team")
)
compliance_reqs = constraints.get("compliance_requirements", [])
if compliance_reqs:
additional_risks.append(
RiskItem("compliance", "Regulatory compliance requirements", "medium", "high", "high",
"Ensure all compliance checks are integrated into migration process", "Compliance Team")
)
return base_risks + additional_risks
def _generate_rollback_plan(self, phases: List[MigrationPhase]) -> Dict[str, Any]:
"""Generate comprehensive rollback plan"""
rollback_phases = []
for phase in reversed(phases):
rollback_phase = {
"phase": phase.name,
"rollback_actions": [
f"Revert {phase.name} changes",
f"Restore pre-{phase.name} state",
f"Validate {phase.name} rollback success"
],
"validation_criteria": [
f"System restored to pre-{phase.name} state",
f"All {phase.name} changes successfully reverted",
"System functionality confirmed"
],
"estimated_time_minutes": phase.duration_hours * 15 # 25% of original phase time
}
rollback_phases.append(rollback_phase)
return {
"rollback_phases": rollback_phases,
"rollback_triggers": [
"Critical system failure",
"Data corruption detected",
"Migration timeline exceeded by > 50%",
"Business-critical functionality unavailable",
"Security breach detected",
"Stakeholder decision to abort"
],
"rollback_decision_matrix": {
"low_severity": "Continue with monitoring",
"medium_severity": "Assess and decide within 15 minutes",
"high_severity": "Immediate rollback initiation",
"critical_severity": "Emergency rollback - all hands"
},
"rollback_contacts": [
"Migration Lead",
"Technical Lead",
"Business Owner",
"On-call Engineer"
]
}
def generate_plan(self, spec: Dict[str, Any]) -> MigrationPlan:
"""Generate complete migration plan from specification"""
migration_id = hashlib.md5(json.dumps(spec, sort_keys=True).encode()).hexdigest()[:12]
complexity = self._calculate_complexity(spec)
phases = self._generate_phases(spec)
risks = self._assess_risks(spec)
total_duration = sum(phase.duration_hours for phase in phases)
rollback_plan = self._generate_rollback_plan(phases)
success_criteria = [
"All data successfully migrated with 100% integrity",
"System performance meets or exceeds baseline",
"All business processes functioning normally",
"No critical security vulnerabilities introduced",
"Stakeholder acceptance criteria met",
"Documentation and runbooks updated"
]
stakeholders = [
"Business Owner",
"Technical Lead",
"DevOps Team",
"QA Team",
"Security Team",
"End Users"
]
return MigrationPlan(
migration_id=migration_id,
source_system=spec.get("source", "Unknown"),
target_system=spec.get("target", "Unknown"),
migration_type=spec.get("type", "Unknown"),
complexity=complexity,
estimated_duration_hours=total_duration,
phases=phases,
risks=risks,
success_criteria=success_criteria,
rollback_plan=rollback_plan,
stakeholders=stakeholders,
created_at=datetime.datetime.now().isoformat()
)
def generate_human_readable_plan(self, plan: MigrationPlan) -> str:
"""Generate human-readable migration plan"""
output = []
output.append("=" * 80)
output.append(f"MIGRATION PLAN: {plan.migration_id}")
output.append("=" * 80)
output.append(f"Source System: {plan.source_system}")
output.append(f"Target System: {plan.target_system}")
output.append(f"Migration Type: {plan.migration_type.upper()}")
output.append(f"Complexity Level: {plan.complexity.upper()}")
output.append(f"Estimated Duration: {plan.estimated_duration_hours} hours ({plan.estimated_duration_hours/24:.1f} days)")
output.append(f"Created: {plan.created_at}")
output.append("")
# Phases
output.append("MIGRATION PHASES")
output.append("-" * 40)
for i, phase in enumerate(plan.phases, 1):
output.append(f"{i}. {phase.name.upper()} ({phase.duration_hours}h)")
output.append(f" Description: {phase.description}")
output.append(f" Risk Level: {phase.risk_level.upper()}")
if phase.dependencies:
output.append(f" Dependencies: {', '.join(phase.dependencies)}")
output.append(" Tasks:")
for task in phase.tasks:
output.append(f" • {task}")
output.append(" Success Criteria:")
for criteria in phase.validation_criteria:
output.append(f" ✓ {criteria}")
output.append("")
# Risk Assessment
output.append("RISK ASSESSMENT")
output.append("-" * 40)
risk_by_severity = {}
for risk in plan.risks:
if risk.severity not in risk_by_severity:
risk_by_severity[risk.severity] = []
risk_by_severity[risk.severity].append(risk)
for severity in ["critical", "high", "medium", "low"]:
if severity in risk_by_severity:
output.append(f"{severity.upper()} SEVERITY RISKS:")
for risk in risk_by_severity[severity]:
output.append(f" • {risk.description}")
output.append(f" Category: {risk.category}")
output.append(f" Probability: {risk.probability} | Impact: {risk.impact}")
output.append(f" Mitigation: {risk.mitigation}")
output.append(f" Owner: {risk.owner}")
output.append("")
# Rollback Plan
output.append("ROLLBACK STRATEGY")
output.append("-" * 40)
output.append("Rollback Triggers:")
for trigger in plan.rollback_plan["rollback_triggers"]:
output.append(f" • {trigger}")
output.append("")
output.append("Rollback Phases:")
for rb_phase in plan.rollback_plan["rollback_phases"]:
output.append(f" {rb_phase['phase'].upper()}:")
for action in rb_phase["rollback_actions"]:
output.append(f" - {action}")
output.append(f" Estimated Time: {rb_phase['estimated_time_minutes']} minutes")
output.append("")
# Success Criteria
output.append("SUCCESS CRITERIA")
output.append("-" * 40)
for criteria in plan.success_criteria:
output.append(f"✓ {criteria}")
output.append("")
# Stakeholders
output.append("STAKEHOLDERS")
output.append("-" * 40)
for stakeholder in plan.stakeholders:
output.append(f"• {stakeholder}")
output.append("")
return "\n".join(output)
def main():
"""Main function with command line interface"""
parser = argparse.ArgumentParser(description="Generate comprehensive migration plans")
parser.add_argument("--input", "-i", required=True, help="Input migration specification file (JSON)")
parser.add_argument("--output", "-o", help="Output file for migration plan (JSON)")
parser.add_argument("--format", "-f", choices=["json", "text", "both"], default="both",
help="Output format")
parser.add_argument("--validate", action="store_true", help="Validate migration specification only")
args = parser.parse_args()
try:
# Load migration specification
with open(args.input, 'r') as f:
spec = json.load(f)
# Validate required fields
required_fields = ["type", "source", "target"]
for field in required_fields:
if field not in spec:
print(f"Error: Missing required field '{field}' in specification", file=sys.stderr)
return 1
if args.validate:
print("Migration specification is valid")
return 0
# Generate migration plan
planner = MigrationPlanner()
plan = planner.generate_plan(spec)
# Output results
if args.format in ["json", "both"]:
plan_dict = asdict(plan)
if args.output:
with open(args.output, 'w') as f:
json.dump(plan_dict, f, indent=2)
print(f"Migration plan saved to {args.output}")
else:
print(json.dumps(plan_dict, indent=2))
if args.format in ["text", "both"]:
human_plan = planner.generate_human_readable_plan(plan)
text_output = args.output.replace('.json', '.txt') if args.output else None
if text_output:
with open(text_output, 'w') as f:
f.write(human_plan)
print(f"Human-readable plan saved to {text_output}")
else:
print("\n" + "="*80)
print("HUMAN-READABLE MIGRATION PLAN")
print("="*80)
print(human_plan)
except FileNotFoundError:
print(f"Error: Input file '{args.input}' not found", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON in input file: {e}", file=sys.stderr)
return 1
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
return 1
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/rollback_generator.py
#!/usr/bin/env python3
"""
Rollback Generator - Generate comprehensive rollback procedures for migrations
This tool takes a migration plan and generates detailed rollback procedures for each phase,
including data rollback scripts, service rollback steps, validation checks, and communication
templates to ensure safe and reliable migration reversals.
Author: Migration Architect Skill
Version: 1.0.0
License: MIT
"""
import json
import argparse
import sys
import datetime
import hashlib
from typing import Dict, List, Any, Optional, Tuple
from dataclasses import dataclass, asdict
from enum import Enum
class RollbackTrigger(Enum):
"""Types of rollback triggers"""
MANUAL = "manual"
AUTOMATED = "automated"
THRESHOLD_BASED = "threshold_based"
TIME_BASED = "time_based"
class RollbackUrgency(Enum):
"""Rollback urgency levels"""
LOW = "low"
MEDIUM = "medium"
HIGH = "high"
EMERGENCY = "emergency"
@dataclass
class RollbackStep:
"""Individual rollback step"""
step_id: str
name: str
description: str
script_type: str # sql, bash, api, manual
script_content: str
estimated_duration_minutes: int
dependencies: List[str]
validation_commands: List[str]
success_criteria: List[str]
failure_escalation: str
rollback_order: int
@dataclass
class RollbackPhase:
"""Rollback phase containing multiple steps"""
phase_name: str
description: str
urgency_level: str
estimated_duration_minutes: int
prerequisites: List[str]
steps: List[RollbackStep]
validation_checkpoints: List[str]
communication_requirements: List[str]
risk_level: str
@dataclass
class RollbackTriggerCondition:
"""Conditions that trigger automatic rollback"""
trigger_id: str
name: str
condition: str
metric_threshold: Optional[Dict[str, Any]]
evaluation_window_minutes: int
auto_execute: bool
escalation_contacts: List[str]
@dataclass
class DataRecoveryPlan:
"""Data recovery and restoration plan"""
recovery_method: str # backup_restore, point_in_time, event_replay
backup_location: str
recovery_scripts: List[str]
data_validation_queries: List[str]
estimated_recovery_time_minutes: int
recovery_dependencies: List[str]
@dataclass
class CommunicationTemplate:
"""Communication template for rollback scenarios"""
template_type: str # start, progress, completion, escalation
audience: str # technical, business, executive, customers
subject: str
body: str
urgency: str
delivery_methods: List[str]
@dataclass
class RollbackRunbook:
"""Complete rollback runbook"""
runbook_id: str
migration_id: str
created_at: str
rollback_phases: List[RollbackPhase]
trigger_conditions: List[RollbackTriggerCondition]
data_recovery_plan: DataRecoveryPlan
communication_templates: List[CommunicationTemplate]
escalation_matrix: Dict[str, Any]
validation_checklist: List[str]
post_rollback_procedures: List[str]
emergency_contacts: List[Dict[str, str]]
class RollbackGenerator:
"""Main rollback generator class"""
def __init__(self):
self.rollback_templates = self._load_rollback_templates()
self.validation_templates = self._load_validation_templates()
self.communication_templates = self._load_communication_templates()
def _load_rollback_templates(self) -> Dict[str, Any]:
"""Load rollback script templates for different migration types"""
return {
"database": {
"schema_rollback": {
"drop_table": "DROP TABLE IF EXISTS {table_name};",
"drop_column": "ALTER TABLE {table_name} DROP COLUMN IF EXISTS {column_name};",
"restore_column": "ALTER TABLE {table_name} ADD COLUMN {column_definition};",
"revert_type": "ALTER TABLE {table_name} ALTER COLUMN {column_name} TYPE {original_type};",
"drop_constraint": "ALTER TABLE {table_name} DROP CONSTRAINT {constraint_name};",
"add_constraint": "ALTER TABLE {table_name} ADD CONSTRAINT {constraint_name} {constraint_definition};"
},
"data_rollback": {
"restore_backup": "pg_restore -d {database_name} -c {backup_file}",
"point_in_time_recovery": "SELECT pg_create_restore_point('pre_migration_{timestamp}');",
"delete_migrated_data": "DELETE FROM {table_name} WHERE migration_batch_id = '{batch_id}';",
"restore_original_values": "UPDATE {table_name} SET {column_name} = backup_{column_name} WHERE migration_flag = true;"
}
},
"service": {
"deployment_rollback": {
"rollback_blue_green": "kubectl patch service {service_name} -p '{\"spec\":{\"selector\":{\"version\":\"blue\"}}}'",
"rollback_canary": "kubectl scale deployment {service_name}-canary --replicas=0",
"restore_previous_version": "kubectl rollout undo deployment/{service_name} --to-revision={revision_number}",
"update_load_balancer": "aws elbv2 modify-rule --rule-arn {rule_arn} --actions Type=forward,TargetGroupArn={original_target_group}"
},
"configuration_rollback": {
"restore_config_map": "kubectl apply -f {original_config_file}",
"revert_feature_flags": "curl -X PUT {feature_flag_api}/flags/{flag_name} -d '{\"enabled\": false}'",
"restore_environment_vars": "kubectl set env deployment/{deployment_name} {env_var_name}={original_value}"
}
},
"infrastructure": {
"cloud_rollback": {
"revert_terraform": "terraform apply -target={resource_name} {rollback_plan_file}",
"restore_dns": "aws route53 change-resource-record-sets --hosted-zone-id {zone_id} --change-batch file://{rollback_dns_changes}",
"rollback_security_groups": "aws ec2 authorize-security-group-ingress --group-id {group_id} --protocol {protocol} --port {port} --cidr {cidr}",
"restore_iam_policies": "aws iam put-role-policy --role-name {role_name} --policy-name {policy_name} --policy-document file://{original_policy}"
},
"network_rollback": {
"restore_routing": "aws ec2 replace-route --route-table-id {route_table_id} --destination-cidr-block {cidr} --gateway-id {original_gateway}",
"revert_load_balancer": "aws elbv2 modify-load-balancer --load-balancer-arn {lb_arn} --scheme {original_scheme}",
"restore_firewall_rules": "aws ec2 revoke-security-group-ingress --group-id {group_id} --protocol {protocol} --port {port} --source-group {source_group}"
}
}
}
def _load_validation_templates(self) -> Dict[str, List[str]]:
"""Load validation command templates"""
return {
"database": [
"SELECT COUNT(*) FROM {table_name};",
"SELECT COUNT(*) FROM information_schema.tables WHERE table_name = '{table_name}';",
"SELECT COUNT(*) FROM information_schema.columns WHERE table_name = '{table_name}' AND column_name = '{column_name}';",
"SELECT COUNT(DISTINCT {primary_key}) FROM {table_name};",
"SELECT MAX({timestamp_column}) FROM {table_name};"
],
"service": [
"curl -f {health_check_url}",
"kubectl get pods -l app={service_name} --field-selector=status.phase=Running",
"kubectl logs deployment/{service_name} --tail=100 | grep -i error",
"curl -f {service_endpoint}/api/v1/status"
],
"infrastructure": [
"aws ec2 describe-instances --instance-ids {instance_id} --query 'Reservations[*].Instances[*].State.Name'",
"nslookup {domain_name}",
"curl -I {load_balancer_url}",
"aws elbv2 describe-target-health --target-group-arn {target_group_arn}"
]
}
def _load_communication_templates(self) -> Dict[str, Dict[str, str]]:
"""Load communication templates"""
return {
"rollback_start": {
"technical": {
"subject": "ROLLBACK INITIATED: {migration_name}",
"body": """Team,
We have initiated rollback for migration: {migration_name}
Rollback ID: {rollback_id}
Start Time: {start_time}
Estimated Duration: {estimated_duration}
Reason: {rollback_reason}
Current Status: Rolling back phase {current_phase}
Next Updates: Every 15 minutes or upon phase completion
Actions Required:
- Monitor system health dashboards
- Stand by for escalation if needed
- Do not make manual changes during rollback
Incident Commander: {incident_commander}
"""
},
"business": {
"subject": "System Rollback In Progress - {system_name}",
"body": """Business Stakeholders,
We are currently performing a planned rollback of the {system_name} migration due to {rollback_reason}.
Impact: {business_impact}
Expected Resolution: {estimated_completion_time}
Affected Services: {affected_services}
We will provide updates every 30 minutes.
Contact: {business_contact}
"""
},
"executive": {
"subject": "EXEC ALERT: Critical System Rollback - {system_name}",
"body": """Executive Team,
A critical rollback is in progress for {system_name}.
Summary:
- Rollback Reason: {rollback_reason}
- Business Impact: {business_impact}
- Expected Resolution: {estimated_completion_time}
- Customer Impact: {customer_impact}
We are following established procedures and will update hourly.
Escalation: {escalation_contact}
"""
}
},
"rollback_complete": {
"technical": {
"subject": "ROLLBACK COMPLETED: {migration_name}",
"body": """Team,
Rollback has been successfully completed for migration: {migration_name}
Summary:
- Start Time: {start_time}
- End Time: {end_time}
- Duration: {actual_duration}
- Phases Completed: {completed_phases}
Validation Results:
{validation_results}
System Status: {system_status}
Next Steps:
- Continue monitoring for 24 hours
- Post-rollback review scheduled for {review_date}
- Root cause analysis to begin
All clear to resume normal operations.
Incident Commander: {incident_commander}
"""
}
}
}
def generate_rollback_runbook(self, migration_plan: Dict[str, Any]) -> RollbackRunbook:
"""Generate comprehensive rollback runbook from migration plan"""
runbook_id = f"rb_{hashlib.md5(str(migration_plan).encode()).hexdigest()[:8]}"
migration_id = migration_plan.get("migration_id", "unknown")
migration_type = migration_plan.get("migration_type", "unknown")
# Generate rollback phases (reverse order of migration phases)
rollback_phases = self._generate_rollback_phases(migration_plan)
# Generate trigger conditions
trigger_conditions = self._generate_trigger_conditions(migration_plan)
# Generate data recovery plan
data_recovery_plan = self._generate_data_recovery_plan(migration_plan)
# Generate communication templates
communication_templates = self._generate_communication_templates(migration_plan)
# Generate escalation matrix
escalation_matrix = self._generate_escalation_matrix(migration_plan)
# Generate validation checklist
validation_checklist = self._generate_validation_checklist(migration_plan)
# Generate post-rollback procedures
post_rollback_procedures = self._generate_post_rollback_procedures(migration_plan)
# Generate emergency contacts
emergency_contacts = self._generate_emergency_contacts(migration_plan)
return RollbackRunbook(
runbook_id=runbook_id,
migration_id=migration_id,
created_at=datetime.datetime.now().isoformat(),
rollback_phases=rollback_phases,
trigger_conditions=trigger_conditions,
data_recovery_plan=data_recovery_plan,
communication_templates=communication_templates,
escalation_matrix=escalation_matrix,
validation_checklist=validation_checklist,
post_rollback_procedures=post_rollback_procedures,
emergency_contacts=emergency_contacts
)
def _generate_rollback_phases(self, migration_plan: Dict[str, Any]) -> List[RollbackPhase]:
"""Generate rollback phases from migration plan"""
migration_phases = migration_plan.get("phases", [])
migration_type = migration_plan.get("migration_type", "unknown")
rollback_phases = []
# Reverse the order of migration phases for rollback
for i, phase in enumerate(reversed(migration_phases)):
if isinstance(phase, dict):
phase_name = phase.get("name", f"phase_{i}")
phase_duration = phase.get("duration_hours", 2) * 60 # Convert to minutes
phase_risk = phase.get("risk_level", "medium")
else:
phase_name = str(phase)
phase_duration = 120 # Default 2 hours
phase_risk = "medium"
rollback_steps = self._generate_rollback_steps(phase_name, migration_type, i)
rollback_phase = RollbackPhase(
phase_name=f"rollback_{phase_name}",
description=f"Rollback changes made during {phase_name} phase",
urgency_level=self._calculate_urgency(phase_risk),
estimated_duration_minutes=phase_duration // 2, # Rollback typically faster
prerequisites=self._get_rollback_prerequisites(phase_name, i),
steps=rollback_steps,
validation_checkpoints=self._get_validation_checkpoints(phase_name, migration_type),
communication_requirements=self._get_communication_requirements(phase_name, phase_risk),
risk_level=phase_risk
)
rollback_phases.append(rollback_phase)
return rollback_phases
def _generate_rollback_steps(self, phase_name: str, migration_type: str, phase_index: int) -> List[RollbackStep]:
"""Generate specific rollback steps for a phase"""
steps = []
templates = self.rollback_templates.get(migration_type, {})
if migration_type == "database":
if "migration" in phase_name.lower() or "cutover" in phase_name.lower():
# Data rollback steps
steps.extend([
RollbackStep(
step_id=f"rb_data_{phase_index}_01",
name="Stop data migration processes",
description="Halt all ongoing data migration processes",
script_type="sql",
script_content="-- Stop migration processes\nSELECT pg_cancel_backend(pid) FROM pg_stat_activity WHERE query LIKE '%migration%';",
estimated_duration_minutes=5,
dependencies=[],
validation_commands=["SELECT COUNT(*) FROM pg_stat_activity WHERE query LIKE '%migration%';"],
success_criteria=["No active migration processes"],
failure_escalation="Contact DBA immediately",
rollback_order=1
),
RollbackStep(
step_id=f"rb_data_{phase_index}_02",
name="Restore from backup",
description="Restore database from pre-migration backup",
script_type="bash",
script_content=templates.get("data_rollback", {}).get("restore_backup", "pg_restore -d {database_name} -c {backup_file}"),
estimated_duration_minutes=30,
dependencies=[f"rb_data_{phase_index}_01"],
validation_commands=["SELECT COUNT(*) FROM information_schema.tables;"],
success_criteria=["Database restored successfully", "All expected tables present"],
failure_escalation="Escalate to senior DBA and infrastructure team",
rollback_order=2
)
])
if "preparation" in phase_name.lower():
# Schema rollback steps
steps.append(
RollbackStep(
step_id=f"rb_schema_{phase_index}_01",
name="Drop migration artifacts",
description="Remove temporary migration tables and procedures",
script_type="sql",
script_content="-- Drop migration artifacts\nDROP TABLE IF EXISTS migration_log;\nDROP PROCEDURE IF EXISTS migrate_data();",
estimated_duration_minutes=5,
dependencies=[],
validation_commands=["SELECT COUNT(*) FROM information_schema.tables WHERE table_name LIKE '%migration%';"],
success_criteria=["No migration artifacts remain"],
failure_escalation="Manual cleanup required",
rollback_order=1
)
)
elif migration_type == "service":
if "cutover" in phase_name.lower():
# Service rollback steps
steps.extend([
RollbackStep(
step_id=f"rb_service_{phase_index}_01",
name="Redirect traffic back to old service",
description="Update load balancer to route traffic back to previous service version",
script_type="bash",
script_content=templates.get("deployment_rollback", {}).get("update_load_balancer", "aws elbv2 modify-rule --rule-arn {rule_arn} --actions Type=forward,TargetGroupArn={original_target_group}"),
estimated_duration_minutes=2,
dependencies=[],
validation_commands=["curl -f {health_check_url}"],
success_criteria=["Traffic routing to original service", "Health checks passing"],
failure_escalation="Emergency procedure - manual traffic routing",
rollback_order=1
),
RollbackStep(
step_id=f"rb_service_{phase_index}_02",
name="Rollback service deployment",
description="Revert to previous service deployment version",
script_type="bash",
script_content=templates.get("deployment_rollback", {}).get("restore_previous_version", "kubectl rollout undo deployment/{service_name} --to-revision={revision_number}"),
estimated_duration_minutes=10,
dependencies=[f"rb_service_{phase_index}_01"],
validation_commands=["kubectl get pods -l app={service_name} --field-selector=status.phase=Running"],
success_criteria=["Previous version deployed", "All pods running"],
failure_escalation="Manual pod management required",
rollback_order=2
)
])
elif migration_type == "infrastructure":
steps.extend([
RollbackStep(
step_id=f"rb_infra_{phase_index}_01",
name="Revert infrastructure changes",
description="Apply terraform plan to revert infrastructure to previous state",
script_type="bash",
script_content=templates.get("cloud_rollback", {}).get("revert_terraform", "terraform apply -target={resource_name} {rollback_plan_file}"),
estimated_duration_minutes=15,
dependencies=[],
validation_commands=["terraform plan -detailed-exitcode"],
success_criteria=["Infrastructure matches previous state", "No planned changes"],
failure_escalation="Manual infrastructure review required",
rollback_order=1
),
RollbackStep(
step_id=f"rb_infra_{phase_index}_02",
name="Restore DNS configuration",
description="Revert DNS changes to point back to original infrastructure",
script_type="bash",
script_content=templates.get("cloud_rollback", {}).get("restore_dns", "aws route53 change-resource-record-sets --hosted-zone-id {zone_id} --change-batch file://{rollback_dns_changes}"),
estimated_duration_minutes=10,
dependencies=[f"rb_infra_{phase_index}_01"],
validation_commands=["nslookup {domain_name}"],
success_criteria=["DNS resolves to original endpoints"],
failure_escalation="Contact DNS administrator",
rollback_order=2
)
])
# Add generic validation step for all migration types
steps.append(
RollbackStep(
step_id=f"rb_validate_{phase_index}_final",
name="Validate rollback completion",
description=f"Comprehensive validation that {phase_name} rollback completed successfully",
script_type="manual",
script_content="Execute validation checklist for this phase",
estimated_duration_minutes=10,
dependencies=[step.step_id for step in steps],
validation_commands=self.validation_templates.get(migration_type, []),
success_criteria=[f"{phase_name} fully rolled back", "All validation checks pass"],
failure_escalation=f"Investigate {phase_name} rollback failures",
rollback_order=99
)
)
return steps
def _generate_trigger_conditions(self, migration_plan: Dict[str, Any]) -> List[RollbackTriggerCondition]:
"""Generate automatic rollback trigger conditions"""
triggers = []
migration_type = migration_plan.get("migration_type", "unknown")
# Generic triggers for all migration types
triggers.extend([
RollbackTriggerCondition(
trigger_id="error_rate_spike",
name="Error Rate Spike",
condition="error_rate > baseline * 5 for 5 minutes",
metric_threshold={
"metric": "error_rate",
"operator": "greater_than",
"value": "baseline_error_rate * 5",
"duration_minutes": 5
},
evaluation_window_minutes=5,
auto_execute=True,
escalation_contacts=["on_call_engineer", "migration_lead"]
),
RollbackTriggerCondition(
trigger_id="response_time_degradation",
name="Response Time Degradation",
condition="p95_response_time > baseline * 3 for 10 minutes",
metric_threshold={
"metric": "p95_response_time",
"operator": "greater_than",
"value": "baseline_p95 * 3",
"duration_minutes": 10
},
evaluation_window_minutes=10,
auto_execute=False,
escalation_contacts=["performance_team", "migration_lead"]
),
RollbackTriggerCondition(
trigger_id="availability_drop",
name="Service Availability Drop",
condition="availability < 95% for 2 minutes",
metric_threshold={
"metric": "availability",
"operator": "less_than",
"value": 0.95,
"duration_minutes": 2
},
evaluation_window_minutes=2,
auto_execute=True,
escalation_contacts=["sre_team", "incident_commander"]
)
])
# Migration-type specific triggers
if migration_type == "database":
triggers.extend([
RollbackTriggerCondition(
trigger_id="data_integrity_failure",
name="Data Integrity Check Failure",
condition="data_validation_failures > 0",
metric_threshold={
"metric": "data_validation_failures",
"operator": "greater_than",
"value": 0,
"duration_minutes": 1
},
evaluation_window_minutes=1,
auto_execute=True,
escalation_contacts=["dba_team", "data_team"]
),
RollbackTriggerCondition(
trigger_id="migration_progress_stalled",
name="Migration Progress Stalled",
condition="migration_progress unchanged for 30 minutes",
metric_threshold={
"metric": "migration_progress_rate",
"operator": "equals",
"value": 0,
"duration_minutes": 30
},
evaluation_window_minutes=30,
auto_execute=False,
escalation_contacts=["migration_team", "dba_team"]
)
])
elif migration_type == "service":
triggers.extend([
RollbackTriggerCondition(
trigger_id="cpu_utilization_spike",
name="CPU Utilization Spike",
condition="cpu_utilization > 90% for 15 minutes",
metric_threshold={
"metric": "cpu_utilization",
"operator": "greater_than",
"value": 0.90,
"duration_minutes": 15
},
evaluation_window_minutes=15,
auto_execute=False,
escalation_contacts=["devops_team", "infrastructure_team"]
),
RollbackTriggerCondition(
trigger_id="memory_leak_detected",
name="Memory Leak Detected",
condition="memory_usage increasing continuously for 20 minutes",
metric_threshold={
"metric": "memory_growth_rate",
"operator": "greater_than",
"value": "1MB/minute",
"duration_minutes": 20
},
evaluation_window_minutes=20,
auto_execute=True,
escalation_contacts=["development_team", "sre_team"]
)
])
return triggers
def _generate_data_recovery_plan(self, migration_plan: Dict[str, Any]) -> DataRecoveryPlan:
"""Generate data recovery plan"""
migration_type = migration_plan.get("migration_type", "unknown")
if migration_type == "database":
return DataRecoveryPlan(
recovery_method="point_in_time",
backup_location="/backups/pre_migration_{migration_id}_{timestamp}.sql",
recovery_scripts=[
"pg_restore -d production -c /backups/pre_migration_backup.sql",
"SELECT pg_create_restore_point('rollback_point');",
"VACUUM ANALYZE; -- Refresh statistics after restore"
],
data_validation_queries=[
"SELECT COUNT(*) FROM critical_business_table;",
"SELECT MAX(created_at) FROM audit_log;",
"SELECT COUNT(DISTINCT user_id) FROM user_sessions;",
"SELECT SUM(amount) FROM financial_transactions WHERE date = CURRENT_DATE;"
],
estimated_recovery_time_minutes=45,
recovery_dependencies=["database_instance_running", "backup_file_accessible"]
)
else:
return DataRecoveryPlan(
recovery_method="backup_restore",
backup_location="/backups/pre_migration_state",
recovery_scripts=[
"# Restore configuration files from backup",
"cp -r /backups/pre_migration_state/config/* /app/config/",
"# Restart services with previous configuration",
"systemctl restart application_service"
],
data_validation_queries=[
"curl -f http://localhost:8080/health",
"curl -f http://localhost:8080/api/status"
],
estimated_recovery_time_minutes=20,
recovery_dependencies=["service_stopped", "backup_accessible"]
)
def _generate_communication_templates(self, migration_plan: Dict[str, Any]) -> List[CommunicationTemplate]:
"""Generate communication templates for rollback scenarios"""
templates = []
base_templates = self.communication_templates
# Rollback start notifications
for audience in ["technical", "business", "executive"]:
if audience in base_templates["rollback_start"]:
template_data = base_templates["rollback_start"][audience]
templates.append(CommunicationTemplate(
template_type="rollback_start",
audience=audience,
subject=template_data["subject"],
body=template_data["body"],
urgency="high" if audience == "executive" else "medium",
delivery_methods=["email", "slack"] if audience == "technical" else ["email"]
))
# Rollback completion notifications
for audience in ["technical", "business"]:
if audience in base_templates.get("rollback_complete", {}):
template_data = base_templates["rollback_complete"][audience]
templates.append(CommunicationTemplate(
template_type="rollback_complete",
audience=audience,
subject=template_data["subject"],
body=template_data["body"],
urgency="medium",
delivery_methods=["email", "slack"] if audience == "technical" else ["email"]
))
# Emergency escalation template
templates.append(CommunicationTemplate(
template_type="emergency_escalation",
audience="executive",
subject="CRITICAL: Rollback Emergency - {migration_name}",
body="""CRITICAL SITUATION - IMMEDIATE ATTENTION REQUIRED
Migration: {migration_name}
Issue: Rollback procedure has encountered critical failures
Current Status: {current_status}
Failed Components: {failed_components}
Business Impact: {business_impact}
Customer Impact: {customer_impact}
Immediate Actions:
1. Emergency response team activated
2. {emergency_action_1}
3. {emergency_action_2}
War Room: {war_room_location}
Bridge Line: {conference_bridge}
Next Update: {next_update_time}
Incident Commander: {incident_commander}
Executive On-Call: {executive_on_call}
""",
urgency="emergency",
delivery_methods=["email", "sms", "phone_call"]
))
return templates
def _generate_escalation_matrix(self, migration_plan: Dict[str, Any]) -> Dict[str, Any]:
"""Generate escalation matrix for different failure scenarios"""
return {
"level_1": {
"trigger": "Single component failure",
"response_time_minutes": 5,
"contacts": ["on_call_engineer", "migration_lead"],
"actions": ["Investigate issue", "Attempt automated remediation", "Monitor closely"]
},
"level_2": {
"trigger": "Multiple component failures or single critical failure",
"response_time_minutes": 2,
"contacts": ["senior_engineer", "team_lead", "devops_lead"],
"actions": ["Initiate rollback", "Establish war room", "Notify stakeholders"]
},
"level_3": {
"trigger": "System-wide failure or data corruption",
"response_time_minutes": 1,
"contacts": ["engineering_manager", "cto", "incident_commander"],
"actions": ["Emergency rollback", "All hands on deck", "Executive notification"]
},
"emergency": {
"trigger": "Business-critical failure with customer impact",
"response_time_minutes": 0,
"contacts": ["ceo", "cto", "head_of_operations"],
"actions": ["Emergency procedures", "Customer communication", "Media preparation if needed"]
}
}
def _generate_validation_checklist(self, migration_plan: Dict[str, Any]) -> List[str]:
"""Generate comprehensive validation checklist"""
migration_type = migration_plan.get("migration_type", "unknown")
base_checklist = [
"Verify system is responding to health checks",
"Confirm error rates are within normal parameters",
"Validate response times meet SLA requirements",
"Check all critical business processes are functioning",
"Verify monitoring and alerting systems are operational",
"Confirm no data corruption has occurred",
"Validate security controls are functioning properly",
"Check backup systems are working correctly",
"Verify integration points with downstream systems",
"Confirm user authentication and authorization working"
]
if migration_type == "database":
base_checklist.extend([
"Validate database schema matches expected state",
"Confirm referential integrity constraints",
"Check database performance metrics",
"Verify data consistency across related tables",
"Validate indexes and statistics are optimal",
"Confirm transaction logs are clean",
"Check database connections and connection pooling"
])
elif migration_type == "service":
base_checklist.extend([
"Verify service discovery is working correctly",
"Confirm load balancing is distributing traffic properly",
"Check service-to-service communication",
"Validate API endpoints are responding correctly",
"Confirm feature flags are in correct state",
"Check resource utilization (CPU, memory, disk)",
"Verify container orchestration is healthy"
])
elif migration_type == "infrastructure":
base_checklist.extend([
"Verify network connectivity between components",
"Confirm DNS resolution is working correctly",
"Check firewall rules and security groups",
"Validate load balancer configuration",
"Confirm SSL/TLS certificates are valid",
"Check storage systems are accessible",
"Verify backup and disaster recovery systems"
])
return base_checklist
def _generate_post_rollback_procedures(self, migration_plan: Dict[str, Any]) -> List[str]:
"""Generate post-rollback procedures"""
return [
"Monitor system stability for 24-48 hours post-rollback",
"Conduct thorough post-rollback testing of all critical paths",
"Review and analyze rollback metrics and timing",
"Document lessons learned and rollback procedure improvements",
"Schedule post-mortem meeting with all stakeholders",
"Update rollback procedures based on actual experience",
"Communicate rollback completion to all stakeholders",
"Archive rollback logs and artifacts for future reference",
"Review and update monitoring thresholds if needed",
"Plan for next migration attempt with improved procedures",
"Conduct security review to ensure no vulnerabilities introduced",
"Update disaster recovery procedures if affected by rollback",
"Review capacity planning based on rollback resource usage",
"Update documentation with rollback experience and timings"
]
def _generate_emergency_contacts(self, migration_plan: Dict[str, Any]) -> List[Dict[str, str]]:
"""Generate emergency contact list"""
return [
{
"role": "Incident Commander",
"name": "TBD - Assigned during migration",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "incident.commander@company.com",
"backup_contact": "backup.commander@company.com"
},
{
"role": "Technical Lead",
"name": "TBD - Migration technical owner",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "tech.lead@company.com",
"backup_contact": "senior.engineer@company.com"
},
{
"role": "Business Owner",
"name": "TBD - Business stakeholder",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "business.owner@company.com",
"backup_contact": "product.manager@company.com"
},
{
"role": "On-Call Engineer",
"name": "Current on-call rotation",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "oncall@company.com",
"backup_contact": "backup.oncall@company.com"
},
{
"role": "Executive Escalation",
"name": "CTO/VP Engineering",
"primary_phone": "+1-XXX-XXX-XXXX",
"email": "cto@company.com",
"backup_contact": "vp.engineering@company.com"
}
]
def _calculate_urgency(self, risk_level: str) -> str:
"""Calculate rollback urgency based on risk level"""
risk_to_urgency = {
"low": "low",
"medium": "medium",
"high": "high",
"critical": "emergency"
}
return risk_to_urgency.get(risk_level, "medium")
def _get_rollback_prerequisites(self, phase_name: str, phase_index: int) -> List[str]:
"""Get prerequisites for rollback phase"""
prerequisites = [
"Incident commander assigned and briefed",
"All team members notified of rollback initiation",
"Monitoring systems confirmed operational",
"Backup systems verified and accessible"
]
if phase_index > 0:
prerequisites.append("Previous rollback phase completed successfully")
if "cutover" in phase_name.lower():
prerequisites.extend([
"Traffic redirection capabilities confirmed",
"Load balancer configuration backed up",
"DNS changes prepared for quick execution"
])
if "data" in phase_name.lower() or "migration" in phase_name.lower():
prerequisites.extend([
"Database backup verified and accessible",
"Data validation queries prepared",
"Database administrator on standby"
])
return prerequisites
def _get_validation_checkpoints(self, phase_name: str, migration_type: str) -> List[str]:
"""Get validation checkpoints for rollback phase"""
checkpoints = [
f"{phase_name} rollback steps completed",
"System health checks passing",
"No critical errors in logs",
"Key metrics within acceptable ranges"
]
validation_commands = self.validation_templates.get(migration_type, [])
checkpoints.extend([f"Validation command passed: {cmd[:50]}..." for cmd in validation_commands[:3]])
return checkpoints
def _get_communication_requirements(self, phase_name: str, risk_level: str) -> List[str]:
"""Get communication requirements for rollback phase"""
base_requirements = [
"Notify incident commander of phase start/completion",
"Update rollback status dashboard",
"Log all actions and decisions"
]
if risk_level in ["high", "critical"]:
base_requirements.extend([
"Notify all stakeholders of phase progress",
"Update executive team if rollback extends beyond expected time",
"Prepare customer communication if needed"
])
if "cutover" in phase_name.lower():
base_requirements.append("Immediate notification when traffic is redirected")
return base_requirements
def generate_human_readable_runbook(self, runbook: RollbackRunbook) -> str:
"""Generate human-readable rollback runbook"""
output = []
output.append("=" * 80)
output.append(f"ROLLBACK RUNBOOK: {runbook.runbook_id}")
output.append("=" * 80)
output.append(f"Migration ID: {runbook.migration_id}")
output.append(f"Created: {runbook.created_at}")
output.append("")
# Emergency Contacts
output.append("EMERGENCY CONTACTS")
output.append("-" * 40)
for contact in runbook.emergency_contacts:
output.append(f"{contact['role']}: {contact['name']}")
output.append(f" Phone: {contact['primary_phone']}")
output.append(f" Email: {contact['email']}")
output.append(f" Backup: {contact['backup_contact']}")
output.append("")
# Escalation Matrix
output.append("ESCALATION MATRIX")
output.append("-" * 40)
for level, details in runbook.escalation_matrix.items():
output.append(f"{level.upper()}:")
output.append(f" Trigger: {details['trigger']}")
output.append(f" Response Time: {details['response_time_minutes']} minutes")
output.append(f" Contacts: {', '.join(details['contacts'])}")
output.append(f" Actions: {', '.join(details['actions'])}")
output.append("")
# Rollback Trigger Conditions
output.append("AUTOMATIC ROLLBACK TRIGGERS")
output.append("-" * 40)
for trigger in runbook.trigger_conditions:
output.append(f"• {trigger.name}")
output.append(f" Condition: {trigger.condition}")
output.append(f" Auto-Execute: {'Yes' if trigger.auto_execute else 'No'}")
output.append(f" Evaluation Window: {trigger.evaluation_window_minutes} minutes")
output.append(f" Contacts: {', '.join(trigger.escalation_contacts)}")
output.append("")
# Rollback Phases
output.append("ROLLBACK PHASES")
output.append("-" * 40)
for i, phase in enumerate(runbook.rollback_phases, 1):
output.append(f"{i}. {phase.phase_name.upper()}")
output.append(f" Description: {phase.description}")
output.append(f" Urgency: {phase.urgency_level.upper()}")
output.append(f" Duration: {phase.estimated_duration_minutes} minutes")
output.append(f" Risk Level: {phase.risk_level.upper()}")
if phase.prerequisites:
output.append(" Prerequisites:")
for prereq in phase.prerequisites:
output.append(f" ✓ {prereq}")
output.append(" Steps:")
for step in sorted(phase.steps, key=lambda x: x.rollback_order):
output.append(f" {step.rollback_order}. {step.name}")
output.append(f" Duration: {step.estimated_duration_minutes} min")
output.append(f" Type: {step.script_type}")
if step.script_content and step.script_type != "manual":
output.append(" Script:")
for line in step.script_content.split('\n')[:3]: # Show first 3 lines
output.append(f" {line}")
if len(step.script_content.split('\n')) > 3:
output.append(" ...")
output.append(f" Success Criteria: {', '.join(step.success_criteria)}")
output.append("")
if phase.validation_checkpoints:
output.append(" Validation Checkpoints:")
for checkpoint in phase.validation_checkpoints:
output.append(f" ☐ {checkpoint}")
output.append("")
# Data Recovery Plan
output.append("DATA RECOVERY PLAN")
output.append("-" * 40)
drp = runbook.data_recovery_plan
output.append(f"Recovery Method: {drp.recovery_method}")
output.append(f"Backup Location: {drp.backup_location}")
output.append(f"Estimated Recovery Time: {drp.estimated_recovery_time_minutes} minutes")
output.append("Recovery Scripts:")
for script in drp.recovery_scripts:
output.append(f" • {script}")
output.append("Validation Queries:")
for query in drp.data_validation_queries:
output.append(f" • {query}")
output.append("")
# Validation Checklist
output.append("POST-ROLLBACK VALIDATION CHECKLIST")
output.append("-" * 40)
for i, item in enumerate(runbook.validation_checklist, 1):
output.append(f"{i:2d}. ☐ {item}")
output.append("")
# Post-Rollback Procedures
output.append("POST-ROLLBACK PROCEDURES")
output.append("-" * 40)
for i, procedure in enumerate(runbook.post_rollback_procedures, 1):
output.append(f"{i:2d}. {procedure}")
output.append("")
return "\n".join(output)
def main():
"""Main function with command line interface"""
parser = argparse.ArgumentParser(description="Generate comprehensive rollback runbooks from migration plans")
parser.add_argument("--input", "-i", required=True, help="Input migration plan file (JSON)")
parser.add_argument("--output", "-o", help="Output file for rollback runbook (JSON)")
parser.add_argument("--format", "-f", choices=["json", "text", "both"], default="both", help="Output format")
args = parser.parse_args()
try:
# Load migration plan
with open(args.input, 'r') as f:
migration_plan = json.load(f)
# Validate required fields
if "migration_id" not in migration_plan and "source" not in migration_plan:
print("Error: Migration plan must contain migration_id or source field", file=sys.stderr)
return 1
# Generate rollback runbook
generator = RollbackGenerator()
runbook = generator.generate_rollback_runbook(migration_plan)
# Output results
if args.format in ["json", "both"]:
runbook_dict = asdict(runbook)
if args.output:
with open(args.output, 'w') as f:
json.dump(runbook_dict, f, indent=2)
print(f"Rollback runbook saved to {args.output}")
else:
print(json.dumps(runbook_dict, indent=2))
if args.format in ["text", "both"]:
human_runbook = generator.generate_human_readable_runbook(runbook)
text_output = args.output.replace('.json', '.txt') if args.output else None
if text_output:
with open(text_output, 'w') as f:
f.write(human_runbook)
print(f"Human-readable runbook saved to {text_output}")
else:
print("\n" + "="*80)
print("HUMAN-READABLE ROLLBACK RUNBOOK")
print("="*80)
print(human_runbook)
except FileNotFoundError:
print(f"Error: Input file '{args.input}' not found", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON in input file: {e}", file=sys.stderr)
return 1
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
return 1
return 0
if __name__ == "__main__":
sys.exit(main())Tuân thủ quy định ban hành và kiểm soát tài liệu của Elmich khi soạn, đặt mã, đặt tên, trình duyệt, ban hành, lưu trữ chính sách và quy trình.
--- name: elmich-document-control description: Tuân thủ Quy định ban hành và kiểm soát tài liệu và Quy trình hệ thống nội bộ của Công ty cổ phần Elmich (QĐ.HCNS.02/ELM, hiệu lực 05/10/2026). Dùng khi soạn, rà soát, sửa đổi, đặt mã, đặt tên, trình duyệt, ban hành hoặc lưu trữ chính sách, quy chế, quy định, quy trình, SOP, hướng dẫn, biểu mẫu của Elmich; khi cần mã hiệu, phiên bản, trang kiểm soát, header, thẩm quyền phê duyệt, SLA ban hành, cấu trúc SharePoint. --- # Kiểm soát tài liệu hệ thống – Elmich Skill này giúp mọi tài liệu quản trị nội bộ của Công ty cổ phần Elmich được soạn, đặt mã, phê duyệt, ban hành và lưu trữ đúng Quy định ban hành và kiểm soát tài liệu và Quy trình hệ thống nội bộ (mã hiệu QĐ.HCNS.02/ELM, ban hành lần 01, hiệu lực 05/10/2026; 6 chương, 28 điều). Nguồn: bản Quyết định số 0310/2026/QĐ-ELM do Tổng Giám đốc ký. Lưu ý: trang Quyết định ghi mã "QĐ.NS.02/ELM" còn bìa và header ghi "QĐ.HCNS.02/ELM"; skill dùng mã trên bìa và header, và cần báo cho HCNS thống nhất lại. Khi áp dụng: nếu người dùng yêu cầu soạn tài liệu, làm theo các mục dưới đây. Nếu yêu cầu rà soát, đối chiếu từng mục và trả về bảng Đạt, Chưa đạt, Cần bổ sung kèm số Điều. Không tự bịa mã lĩnh vực, số thứ tự tài liệu, ngày hiệu lực hoặc tên người phê duyệt; điền chỗ trống và nói rõ ai cấp. ## 1. Phân loại và cấp tài liệu (Điều 6 – 8) Nhóm: văn bản điều hành (nghị quyết, quyết định, thông báo, công văn); tài liệu quản trị hệ thống; tài liệu pháp lý (hợp đồng, thỏa thuận, NDA, MOU, hồ sơ pháp nhân); tài liệu bên ngoài (luật, nghị định, thông tư, tiêu chuẩn, bản vẽ, thông số, yêu cầu khách hàng). Loại tài liệu hệ thống, mã và cấp quản trị: - Chính sách (CS), Quy chế (QC): cấp 1. Xác lập định hướng, nguyên tắc, cơ chế tổ chức, thẩm quyền, phối hợp. - Quy định hoặc Nội quy (QD), Tiêu chuẩn (TC), Định mức (ĐM): cấp 2. Yêu cầu bắt buộc, giới hạn, tiêu chí, mức chuẩn. - Quy trình (QT): cấp 3. Chuỗi hoạt động đầu đến cuối, phân định trách nhiệm, SLA, điểm kiểm soát. - SOP (SOP), Hướng dẫn công việc (HD), Workflow (WF), Checklist (CL), Sổ tay hoặc Cẩm nang (ST): cấp 4. Chuẩn hóa chi tiết thực hiện, số hóa nghiệp vụ. - Biểu mẫu chuẩn (BM), Báo cáo chuẩn (BC): cấp 5. Thu thập, ghi nhận, cung cấp thông tin quản trị. Quy tắc: cấp tài liệu thể hiện mức quản trị của nội dung, không mặc nhiên tương ứng cấp chức danh phê duyệt. Tài liệu cấp dưới không được trái hoặc vượt nguyên tắc, thẩm quyền, hạn mức của cấp trên. Sổ tay chỉ tổng hợp, hướng dẫn, tra cứu, không tạo quy định trái hoặc thay thế tài liệu nguồn. Biểu mẫu, checklist, báo cáo chuẩn ở trạng thái mẫu thuộc hệ thống tài liệu; sau khi điền, xác nhận hoặc phát hành thì thành hồ sơ hoặc bản ghi. Khi tài liệu chuyên ngành quy định chặt hơn thì áp dụng quy định chặt hơn. ## 2. Đặt tên và mã hóa (Điều 9) - Tên phản ánh đúng đối tượng hoặc kết quả quản trị; không dùng chuỗi hành động thay tên quy trình. Với quy trình ưu tiên cấu trúc "Quy trình + đối tượng hoặc kết quả quản trị", ví dụ Quy trình lập kế hoạch kinh doanh năm, Quy trình xử lý khiếu nại khách hàng. - Văn bản chính: XX.YY.ZZ/ELM, trong đó XX là loại tài liệu, YY là mã lĩnh vực hoặc đơn vị phát hành, ZZ là số thứ tự tài liệu. Ví dụ QĐ.HCNS.02/ELM. - Văn bản phái sinh: XXnn.[mã văn bản chính], nn là thứ tự phái sinh. Ví dụ BM01.QĐ.HCNS.02/ELM. - Mỗi tài liệu một mã duy nhất; không dùng lại mã của tài liệu đã hủy. Đổi cơ cấu nhưng phạm vi quản trị không đổi thì ưu tiên giữ nguyên mã. - Danh mục mã lĩnh vực và đơn vị do Đơn vị quản trị hệ thống duy trì; không tự tạo mã mới. Nếu chưa có mã, ghi "Chờ HCNS cấp mã" thay vì tự đặt. ## 3. Phiên bản, hiệu lực, lịch sử thay đổi (Điều 10, 18) - V1.0 ban hành lần đầu. V1.1, V1.2, V1.3 là sửa đổi nhỏ (không đổi cơ bản phạm vi, thẩm quyền, trách nhiệm, luồng xử lý, điểm kiểm soát trọng yếu). V2.0, V3.0... là sửa đổi lớn hoặc ban hành lại sau tối đa 03 lần sửa đổi nhỏ. Thay đổi lớn phải ban hành phiên bản mới, không phụ thuộc số lần sửa. - Sửa đổi lớn gồm thay đổi phạm vi, bước trọng yếu, Chủ sở hữu hoặc trách nhiệm chính, cấp phê duyệt, hạn mức, SLA trọng yếu, cơ chế kiểm soát, quyền hoặc nghĩa vụ, tác động tài chính hoặc phân quyền hệ thống; phải làm lại tham vấn, thẩm định, phê duyệt, phát hành. - Trạng thái hiệu lực chỉ có hai: Có hiệu lực, Hết hiệu lực. Phiên bản mới có hiệu lực thì phiên bản cũ hết hiệu lực, được lưu và nhận diện rõ để tránh dùng nhầm, và thu hồi bản kiểm soát đang lưu hành. - Mỗi lần sửa phải ghi tối thiểu: phiên bản, ngày thay đổi, nội dung thay đổi chính, người phê duyệt. Cập nhật lịch sử và phiên bản trước khi áp dụng. ## 4. Thể thức trình bày (Điều 11) Áp dụng cho tài liệu thuộc hệ thống quản trị, theo mẫu và loại tài liệu tương ứng: - Khổ A4, mặc định dọc; được dùng ngang cho bảng hoặc lưu đồ rộng. - Font Arial. Nội dung 10 – 11 pt; bảng 8 – 10 pt; tiêu đề 12 – 14 pt. - Lề trên và trái 20 – 25 mm; dưới và phải 15 – 20 mm; thống nhất trong cùng tài liệu. - Header theo mẫu từng loại: logo Elmich, dòng "Tài liệu quản lý chất lượng", tên tài liệu, và bốn ô Mã hiệu, Ngày hiệu lực, Lần BH/SĐ (ví dụ 01/00), Trang (x/tổng). - Trang kiểm soát cho tài liệu cần kiểm soát soạn thảo, soát xét, phê duyệt, lịch sử hoặc phân phối: gồm Bảng phân phối tài liệu, Lịch sử sửa đổi (lần sửa đổi, ngày hiệu lực, nội dung, ghi chú), và khối Soạn thảo, Soát xét, Phê duyệt (họ tên, chức danh, ngày ký). - Đánh số Chương, Điều, Khoản, Điểm; quy trình có thể dùng B01, B02... cho bước thực hiện. - Bảng và lưu đồ trình bày rõ, lặp tiêu đề cột khi qua trang, hạn chế chia một hàng qua hai trang. - File phát hành ưu tiên PDF hoặc định dạng chỉ đọc; biểu mẫu theo định dạng phù hợp để dùng. Tiếng Việt là ngôn ngữ chính, thuật ngữ nước ngoài khi cần thiết. ## 5. Cấu trúc tối thiểu của Quy trình (Điều 12) Quy trình phải có đủ 12 nội dung: mục đích; phạm vi, đối tượng áp dụng; thuật ngữ và tài liệu liên quan; nguyên tắc thực hiện (điều kiện, giới hạn, phân quyền); điểm bắt đầu và kết thúc; lưu đồ (trình tự, trách nhiệm, bàn giao, kiểm tra hoặc phê duyệt, nhánh chính); diễn giải bước; điểm kiểm soát (phê duyệt, hạn mức, ngoại lệ trọng yếu); chỉ số đầu ra; biểu mẫu, hồ sơ, hệ thống; tổ chức thực hiện (Chủ sở hữu, giám sát, cập nhật); hiệu lực và tài liệu thay thế. Bảng diễn giải bước tối thiểu gồm cột: Bước, Trách nhiệm, Nội dung hoặc hành động, Thời gian hoặc SLA, Đầu ra. Lưu đồ và bảng diễn giải phải thống nhất mã bước và nội dung. Không bắt buộc nhiều chỉ số; ưu tiên ít chỉ số phản ánh trực tiếp hiệu quả và chất lượng đầu ra. ## 6. Thẩm quyền phê duyệt (Điều 14) - HĐQT hoặc Chủ tịch: tài liệu thuộc thẩm quyền theo Điều lệ, quy chế quản trị hoặc phân quyền của Công ty. - Tổng Giám đốc: tài liệu áp dụng toàn Công ty, liên đơn vị hoặc có ảnh hưởng trọng yếu đến cơ cấu, phân quyền, P&L, khách hàng, pháp lý, chất lượng, dữ liệu, an toàn. - Giám đốc Khối hoặc Trưởng đơn vị: SOP, biểu mẫu, hướng dẫn công việc thuộc quy trình hoặc quy định đã duyệt, với điều kiện không trái tài liệu cấp trên, không tạo nghĩa vụ cho đơn vị khác, không vượt ngân sách hoặc hạn mức. - Không hạ cấp phê duyệt đối với nội dung thuộc thẩm quyền cấp cao hơn. ## 7. Tham vấn, thẩm định, soát xét trước phê duyệt (Điều 15) Tài liệu ảnh hưởng đơn vị nào phải lấy ý kiến đơn vị đó. Nội dung chuyên môn trọng yếu phải có chức năng liên quan thẩm định: - Pháp lý (quy định pháp luật, hợp đồng, quyền nghĩa vụ với bên thứ ba, dữ liệu cá nhân): Pháp chế hoặc chức năng được giao. - Tài chính, Kế toán (thu chi, ngân sách, giá thành, công nợ, thuế, cơ chế thanh toán): Tài chính – Kế toán. - Nhân sự (cơ cấu, chức danh, định biên, tuyển dụng, lương thưởng, đánh giá, kỷ luật): Nhân sự. - CNTT và dữ liệu (phần mềm, tài khoản, phân quyền, tích hợp, workflow điện tử, bảo mật, sao lưu): CNTT hoặc đơn vị quản trị dữ liệu. - Chất lượng, Kỹ thuật (tiêu chuẩn sản phẩm, nguyên vật liệu, kiểm nghiệm): Chất lượng, Kỹ thuật, Nhà máy theo phạm vi. - HSE (an toàn lao động, PCCC, môi trường, máy móc): HSE hoặc chức năng chuyên trách. - Kinh doanh, Thương mại (giá bán, chiết khấu, khuyến mại, điều kiện bán hàng): Kinh doanh và Tài chính. - Marketing, Thương hiệu, Content Ads (nhận diện, truyền thông, nội dung công bố ra ngoài, hình ảnh thương hiệu): Marketing, Thương hiệu, Content Ads. - Kế hoạch, Cung ứng, Logistics (dự báo, mua hàng, sản xuất, tồn kho, vận chuyển, S&OP): Kế hoạch, Cung ứng, Logistics theo phạm vi. - Workflow và tự động hóa: Chủ sở hữu quy trình cùng CNTT và Đơn vị quản trị hệ thống. Đơn vị quản trị hệ thống soát xét phân loại, mã, cấu trúc, tính thống nhất, trùng lặp, tính đầy đủ trước khi trình duyệt. Tham vấn, thẩm định, soát xét và phê duyệt là các vai trò độc lập; góp ý hoặc xác nhận không đồng nghĩa với quyền phê duyệt. Phê duyệt và phát hành là hai việc độc lập: người có thẩm quyền duyệt nội dung, rồi Đơn vị quản trị hệ thống kiểm soát và phát hành bản chính thức. ## 7b. Vai trò (Điều 13) Chủ sở hữu tài liệu: chịu trách nhiệm cuối cùng về nội dung, tính đúng đắn, khả thi, hiệu quả, đề xuất sửa đổi. Đơn vị quản trị hệ thống tài liệu: phân loại, mã, phiên bản, thể thức, danh mục, hiệu lực, kho chính thức. Đơn vị chuyên môn: góp ý, thẩm định. Người có thẩm quyền: phê duyệt. Hành chính hoặc Văn thư: số văn bản, bản ký gốc, đóng dấu, hồ sơ phát hành. CNTT: kỹ thuật hệ thống, quyền truy cập, sao lưu, workflow theo tài liệu đã duyệt. ## 8. Quy trình 6 bước và SLA (Điều 16) - B01 Đề xuất (Chủ sở hữu hoặc đơn vị đề xuất): 01 ngày làm việc; đầu ra đề xuất xây dựng hoặc sửa đổi. - B02 Soạn thảo (Chủ sở hữu hoặc đơn vị soạn thảo): 03 – 05 ngày làm việc; đầu ra dự thảo. - B03 Tham vấn, thẩm định (Chủ sở hữu, chức năng liên quan, đơn vị quản trị hệ thống): 02 – 03 ngày làm việc; các đơn vị ký xác nhận đồng ý theo BM01. - B04 Phê duyệt (Chủ sở hữu và người phê duyệt): 01 – 02 ngày làm việc. - B05 Phát hành (Đơn vị quản trị hệ thống): chậm nhất 01 ngày làm việc sau phê duyệt; chốt mã, phiên bản, ngày hiệu lực, gửi email ban hành, cập nhật danh mục theo BM02, lưu kho chính thức, chuyển bản cũ sang hết hiệu lực. - B06 Truyền thông, áp dụng (Chủ sở hữu và đơn vị liên quan): trong 01 – 03 ngày làm việc sau phát hành; thông báo, đào tạo, cấu hình workflow hoặc hệ thống theo kế hoạch đã duyệt, theo dõi áp dụng. Sau đào tạo với quy trình, quy định mới, nhân sự ký cam kết theo BM03. Khẩn cấp: cấp có thẩm quyền có thể cho rút gọn tham vấn, thẩm định, nhưng tài liệu vẫn phải được phê duyệt, nhận diện, phát hành và hoàn thiện hồ sơ kiểm soát sau đó. ## 9. Bản chính thức và kiểm soát sử dụng (Điều 17) - Chỉ tài liệu đã phê duyệt, có mã, phiên bản, ngày hiệu lực và công bố trên kho tài liệu chính thức mới có giá trị áp dụng. - Phát hành mặc định bằng thông báo qua email hoặc nền tảng nội bộ kèm đường dẫn đến bản hiện hành; không dùng file đính kèm làm nguồn áp dụng chính thức. - Không tự lưu hành file riêng ngoài kho kiểm soát. Bản tải xuống hoặc bản in là bản không kiểm soát, trừ khi được đăng ký và nhận diện là BẢN KIỂM SOÁT. Email, tin nhắn, bản sao chỉ có giá trị thông báo, tham khảo. - Quyền xem, tải, in, sao chép, chỉnh sửa, chia sẻ theo phạm vi sử dụng và mức độ bảo mật. ## 10. Lưu trữ trên SharePoint (Điều 20) SharePoint là kho điện tử chính thức. Cấu trúc 4 tầng: Tầng 1 là khu vực (CEO, PUBLIC, TOÀN QUỐC, MIỀN BẮC, MIỀN NAM, NHÀ MÁY); Tầng 2 là phòng ban hoặc chức năng (riêng PUBLIC theo nhóm nội dung dùng chung); Tầng 3 là nghiệp vụ, cấp phân quyền chính; Tầng 4 là Năm, Tháng, Quý, Kỳ (chỉ với hồ sơ, dữ liệu có kỳ; tài liệu chuẩn quản lý theo phiên bản và ngày hiệu lực, không chia theo tháng). Tổ chức dữ liệu: Khu vực, Chức năng, Nghiệp vụ, Thời gian (ví dụ MIỀN NAM, HCNS, TUYỂN DỤNG, 2026, 09). Dữ liệu nhạy cảm (lương thưởng, dữ liệu cá nhân, kỷ luật, đánh giá cán bộ, pháp lý, thông tin mật) phải tách vùng lưu trữ và phân quyền riêng, không mặc nhiên kế thừa quyền chung của phòng ban. CEO không là nơi lưu dữ liệu nguồn của các đơn vị. PUBLIC là khu vực công bố và dùng chung, nhưng không có nghĩa mọi nội dung trong PUBLIC mở cho toàn bộ CBNV. TOÀN QUỐC lấy dữ liệu tự động từ Miền Bắc, Miền Nam, Nhà máy; không nhập hoặc sao chép lại khi đã có nguồn chuẩn. Phân quyền theo 3 yếu tố: Chức năng hoặc nghiệp vụ, Phạm vi quản lý, Mức quyền (Xem; Cập nhật; Quản trị). Quy tắc: người cùng phòng ban không mặc nhiên xem toàn bộ dữ liệu phòng ban; ưu tiên phân quyền theo nhóm người dùng ở Library, Folder lớn hoặc Tầng 3, hạn chế phân quyền lẻ từng file; quyền kỹ thuật của CNTT không đồng nghĩa quyền khai thác nội dung nghiệp vụ; tên nhóm quyền theo cấu trúc [Chức năng]_[Nghiệp vụ]_[Phạm vi]_[Mức quyền]; đổi nhân sự bằng thêm hoặc bớt thành viên khỏi nhóm quyền. Power Query và Power BI: luồng chuẩn MIỀN BẮC + MIỀN NAM + NHÀ MÁY, qua Power Query hoặc Power BI, đến TOÀN QUỐC; chỉ kết nối vùng DATA đã xác định, không quét toàn bộ thư mục; các nguồn cùng nghiệp vụ thống nhất cấu trúc file, tên bảng, tên cột, kiểu dữ liệu, mã đơn vị, kỳ dữ liệu; tài khoản kết nối chỉ có quyền đọc đúng nguồn. Không tự ý đổi cấu trúc thư mục, tên folder, tên file chuẩn, cấu trúc bảng hoặc quyền truy cập nếu có thể ảnh hưởng Power Query, Power BI, workflow, báo cáo; mọi thay đổi có ảnh hưởng phải được Chủ sở hữu dữ liệu thống nhất với CNTT và Đơn vị quản trị hệ thống trước khi thực hiện. ## 11. Tài liệu bên ngoài, bản cứng, tiêu hủy (Điều 19, 21, 22) - Tài liệu bên ngoài dùng làm căn cứ phải được nhận diện, theo dõi tối thiểu: tên và số hiệu, nguồn ban hành, phiên bản và ngày hiệu lực, nơi lưu hoặc link nguồn, phạm vi áp dụng. Khi thay đổi, Chủ sở hữu đánh giá tác động và cập nhật tài liệu, quy trình, hệ thống liên quan. Tài liệu kỹ thuật, bản vẽ, tiêu chuẩn khách hàng, tài liệu hạn chế phải phân quyền đúng đối tượng. - Bản cứng lưu khi pháp luật, hợp đồng, kiểm toán, thẩm quyền ký hoặc nhu cầu chứng minh bản gốc yêu cầu (hồ sơ pháp nhân, giấy phép, hồ sơ HĐQT, BĐH, quyết định quan trọng, hợp đồng, thỏa thuận có chữ ký gốc). Hành chính hoặc Văn thư lưu bản ký gốc; bản cứng và bản điện tử liên kết được theo mã hoặc tên tài liệu; sắp xếp theo Đơn vị, Nhóm hồ sơ, Năm hoặc kỳ, Số văn bản hoặc thời gian; mỗi bìa hoặc tập có danh mục tài liệu ở đầu tập. Gáy bìa còng: nền trắng, logo, tên công ty, phòng ban, tên hồ sơ, số thứ tự hoặc ngày, chữ in hoa đậm, chữ dọc, font Arial. - Tiêu hủy: đơn vị sở hữu rà soát hồ sơ hết thời hạn lưu và lập danh mục đề nghị tiêu hủy; không tiêu hủy hồ sơ liên quan tranh chấp, kiểm toán, thanh tra, điều tra, yêu cầu pháp lý hoặc lưu giữ đặc biệt; phải được phê duyệt theo thẩm quyền; phương thức đảm bảo không thể khôi phục; lập biên bản và cập nhật danh mục hồ sơ. ## 12. Rà soát, ngoại lệ, cải tiến (Điều 23, 24) - Chính sách, Quy chế, Quy định, Quy trình rà soát tối thiểu 12 tháng một lần hoặc khi có thay đổi trọng yếu. SOP, Hướng dẫn, Checklist, Biểu mẫu rà soát khi tài liệu nguồn, nghiệp vụ hoặc hệ thống liên quan thay đổi. Ngoài chu kỳ, rà soát khi đổi pháp luật, cơ cấu, phân quyền, quy trình, hệ thống, sản phẩm, khách hàng hoặc có rủi ro, sai lệch trọng yếu. Rà soát không mặc nhiên dẫn đến sửa đổi; nếu vẫn phù hợp, Chủ sở hữu ghi nhận kết quả và tiếp tục áp dụng. - Ngoại lệ so với tài liệu hiện hành phải được người có thẩm quyền phê duyệt, xác định rõ lý do, phạm vi, thời hạn, rủi ro và biện pháp kiểm soát thay thế. Ngoại lệ lặp lại hoặc kéo dài phải xem xét sửa đổi tài liệu hoặc xử lý nguyên nhân gốc. - Cải tiến ưu tiên loại bỏ việc không tạo giá trị, giảm bước phê duyệt hoặc bàn giao không cần thiết, rút ngắn thời gian xử lý, chuẩn hóa dữ liệu và tự động hóa phù hợp. Workflow hoặc hệ thống không được thiết lập trái với tài liệu đã được phê duyệt. ## 13. Biểu mẫu kèm theo (Điều 26) và chuyển đổi (Điều 27) - BM01.QĐ.NS.02/ELM Phiếu xác nhận thông qua tài liệu (HCNS lưu, theo thời hiệu của tài liệu). BM02 Danh mục lưu trữ văn bản tài liệu (HCNS, vĩnh viễn; cột: danh mục tài liệu, loại, link, đơn vị soạn thảo, người phê duyệt, mã hiệu, ngày ban hành, lần ban hành, lần sửa đổi, cập nhật hiện trạng, ghi chú). BM03 Phiếu cam kết thực hiện quy trình quy định (HCNS, theo thời hiệu của tài liệu). Mã biểu mẫu trong bản gốc ghi BMxx.QĐ.NS.02/ELM; khi trích dẫn, dùng đúng như bản đang lưu hành. - Tài liệu hiện hữu được rà soát và phân loại: Tiếp tục áp dụng; Cần sửa đổi; Cần hợp nhất; Cần thay thế; Cần ban hành mới; Hết hiệu lực. Không mặc nhiên coi tài liệu hiện có là phù hợp chỉ vì đã từng ban hành. ## 14. Nguyên tắc nền (Điều 5) Một nội dung một nguồn chính thức; một quy trình một Chủ sở hữu (không đồng chủ trì); tuân thủ thứ bậc; quy trình phải đầy đủ đầu vào, đầu ra, bước, trách nhiệm, thời hạn, điểm kiểm soát, hồ sơ; tách biệt phê duyệt và phát hành; quy trình trước, hệ thống sau (workflow, phần mềm chỉ cấu hình chính thức sau khi quy trình, phân quyền, điều kiện phê duyệt đã được duyệt); bảo đảm truy xuất, truy vết (Chủ sở hữu, người phê duyệt, phiên bản, ngày hiệu lực, lịch sử, nơi lưu); kho chính thức là nguồn áp dụng; kiểm tra phiên bản còn hiệu lực trước khi dùng; tài liệu phù hợp thực tế vận hành; kiểm soát quyền truy cập; rà soát định kỳ. ## Cách trả lời - Soạn tài liệu mới: đề xuất loại, mã (hoặc "chờ cấp mã"), cấp, người phê duyệt theo Điều 14, các đơn vị cần thẩm định theo Điều 15, rồi soạn theo thể thức Điều 11 và cấu trúc Điều 12 nếu là quy trình. Kèm trang kiểm soát (phân phối, lịch sử, soạn thảo, soát xét, phê duyệt) để trống chữ ký. - Rà soát tài liệu có sẵn: trả bảng đối chiếu theo các mục 2, 3, 4, 5, 6, 7 với kết luận từng dòng và đề xuất sửa; không tự sửa nội dung thuộc thẩm quyền người phê duyệt. - Không đưa ra cam kết về ngày hiệu lực, số quyết định, chữ ký; đó là việc của Đơn vị quản trị hệ thống và người có thẩm quyền. - Quy định có thể được cập nhật; nếu người dùng cho biết bản mới, ưu tiên bản mới.
Lãnh đạo nhân sự: chiến lược tuyển dụng, thiết kế lương thưởng, cơ cấu tổ chức, văn hóa và giữ chân nhân tài.
---
name: "chro-advisor"
description: "People leadership for scaling companies. Hiring strategy, compensation design, org structure, culture, and retention. Use when building hiring plans, designing comp frameworks, restructuring teams, managing performance, building culture, or when user mentions CHRO, HR, people strategy, talent, headcount, compensation, org design, retention, or performance management."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: c-level
domain: chro-leadership
updated: 2026-03-05
python-tools: hiring_plan_modeler.py, comp_benchmarker.py
frameworks: people-strategy, comp-frameworks, org-design
---
# CHRO Advisor
People strategy and operational HR frameworks for business-aligned hiring, compensation, org design, and culture that scales.
## Keywords
CHRO, chief people officer, CPO, HR, human resources, people strategy, hiring plan, headcount planning, talent acquisition, recruiting, compensation, salary bands, equity, org design, organizational design, career ladder, title framework, retention, performance management, culture, engagement, remote work, hybrid, spans of control, succession planning, attrition
## Quick Start
```bash
python scripts/hiring_plan_modeler.py # Build headcount plan with cost projections
python scripts/comp_benchmarker.py # Benchmark salaries and model total comp
```
## Core Responsibilities
### 1. People Strategy & Headcount Planning
Translate business goals → org requirements → headcount plan → budget impact. Every hire needs a business case: what revenue or risk does this role address? See `references/people_strategy.md` for hiring at each growth stage.
### 2. Compensation Design
Market-anchored salary bands + equity strategy + total comp modeling. See `references/comp_frameworks.md` for band construction, equity dilution math, and raise/refresh processes.
### 3. Org Design
Right structure for the stage. Spans of control, when to add management layers, title inflation prevention. See `references/org_design.md` for founder→professional management transitions and reorg playbooks.
### 4. Retention & Performance
Retention starts at hire. Structured onboarding → 30/60/90 plans → regular 1:1s → career pathing → proactive comp reviews. See `references/people_strategy.md` for what actually moves the needle.
**Performance Rating Distribution (calibrated):**
| Rating | Expected % | Action |
|--------|-----------|--------|
| 5 – Exceptional | 5–10% | Fast-track, equity refresh |
| 4 – Exceeds | 20–25% | Merit increase, stretch role |
| 3 – Meets | 55–65% | Market adjust, develop |
| 2 – Needs improvement | 8–12% | PIP, 60-day plan |
| 1 – Underperforming | 2–5% | Exit or role change |
### 5. Culture & Engagement
Culture is behavior, not values on a wall. Measure eNPS quarterly. Act on results within 30 days or don't ask.
## Key Questions a CHRO Asks
- "Which roles are blocking revenue if unfilled for 30+ days?"
- "What's our regrettable attrition rate? Who left that we wish hadn't?"
- "Are managers our retention asset or our attrition cause?"
- "Can a new hire explain their career path in 12 months?"
- "Where are we paying below P50? Who's a flight risk because of it?"
- "What's the cost of this hire vs. the cost of not hiring?"
## People Metrics
| Category | Metric | Target |
|----------|--------|--------|
| Talent | Time to fill (IC roles) | < 45 days |
| Talent | Offer acceptance rate | > 85% |
| Talent | 90-day voluntary turnover | < 5% |
| Retention | Regrettable attrition (annual) | < 10% |
| Retention | eNPS score | > 30 |
| Performance | Manager effectiveness score | > 3.8/5 |
| Comp | % employees within band | > 90% |
| Comp | Compa-ratio (avg) | 0.95–1.05 |
| Org | Span of control (ICs) | 6–10 |
| Org | Span of control (managers) | 4–7 |
## Red Flags
- Attrition spikes and exit interviews all name the same manager
- Comp bands haven't been refreshed in 18+ months
- No career ladder → top performers leave after 18 months
- Hiring without a written business case or job scorecard
- Performance reviews happen once a year with no mid-year check-in
- Equity refreshes only for executives, not high performers
- Time to fill > 90 days for critical roles
- eNPS below 0 — something is structurally broken
- More than 3 org layers between IC and CEO at < 50 people
## Integration with Other C-Suite Roles
| When... | CHRO works with... | To... |
|---------|-------------------|-------|
| Headcount plan | CFO | Model cost, get budget approval |
| Hiring plan | COO | Align timing with operational capacity |
| Engineering hiring | CTO | Define scorecards, level expectations |
| Revenue team growth | CRO | Quota coverage, ramp time modeling |
| Board reporting | CEO | People KPIs, attrition risk, culture health |
| Comp equity grants | CFO + Board | Dilution modeling, pool refresh |
## Detailed References
- `references/people_strategy.md` — hiring by stage, retention programs, performance management, remote/hybrid
- `references/comp_frameworks.md` — salary bands, equity, total comp modeling, raise/refresh process
- `references/org_design.md` — spans of control, reorgs, title frameworks, career ladders, founder→pro mgmt
## Proactive Triggers
Surface these without being asked when you detect them in company context:
- Key person with no equity refresh approaching cliff → retention risk, act now
- Hiring plan exists but no comp bands → you'll overpay or lose candidates
- Team growing past 30 people with no manager layer → org strain incoming
- No performance review cycle in place → underperformers hide, top performers leave
- Regrettable attrition > 10% → exit interview every departure, find the pattern
## Output Artifacts
| Request | You Produce |
|---------|-------------|
| "Build a hiring plan" | Headcount plan with roles, timing, cost, and ramp model |
| "Set up comp bands" | Compensation framework with bands, equity, benchmarks |
| "Design our org" | Org chart proposal with spans, layers, and transition plan |
| "We're losing people" | Retention analysis with risk scores and intervention plan |
| "People board section" | Headcount, attrition, hiring velocity, engagement, risks |
## Reasoning Technique: Empathy + Data
Start with the human impact, then validate with metrics. Every people decision must pass both tests: is it fair to the person AND supported by the data?
## Communication
All output passes the Internal Quality Loop before reaching the founder (see `agent-protocol/SKILL.md`).
- Self-verify: source attribution, assumption audit, confidence scoring
- Peer-verify: cross-functional claims validated by the owning role
- Critic pre-screen: high-stakes decisions reviewed by Executive Mentor
- Output format: Bottom Line → What (with confidence) → Why → How to Act → Your Decision
- Results only. Every finding tagged: 🟢 verified, 🟡 medium, 🔴 assumed.
## Context Integration
- **Always** read `company-context.md` before responding (if it exists)
- **During board meetings:** Use only your own analysis in Phase 2 (no cross-pollination)
- **Invocation:** You can request input from other roles: `[INVOKE:role|question]`
FILE:references/comp_frameworks.md
# Compensation Frameworks Reference
Salary bands, equity design, total comp modeling, comp philosophy, and raise/refresh processes.
---
## Comp Philosophy — The Foundation
Before building bands, define your philosophy. Ambiguity in comp philosophy = pay equity lawsuits and trust erosion.
**The five decisions:**
### 1. What market percentile do you target?
- **P25 (below market):** Only viable with exceptional mission, equity, or growth opportunity. Flight risk is high after 18 months.
- **P50 (market median):** Standard for most Series A–B companies. Competitive without premium.
- **P75 (above market):** Premium talent strategy. Used by high-margin or talent-intensive businesses. Netflix model.
- **P90+:** Top-of-market for specific functions (ML at AI companies, senior engineers at FAANG feeders).
**Common hybrid:** P50 base + above-market equity = total comp at P65–75.
### 2. What's in your total comp package?
Define each component explicitly:
- **Base salary** — cash, market-benchmarked
- **Variable / bonus** — % of base, tied to what criteria
- **Equity** — options vs. RSUs, vesting schedule, refresh cadence
- **Benefits** — health, retirement, PTO policy
- **Learning & development budget**
- **Remote/location allowances**
### 3. Are bands public internally?
Recommended: Yes. Pay transparency reduces equity complaints, builds trust, and forces you to maintain clean bands.
### 4. How often do you refresh bands?
Minimum: annually. High-growth markets: every 6 months (engineering specifically in hot markets).
### 5. How do you handle individual negotiation?
Options:
- **Fixed bands, no negotiation** (Buffer model) — simple, fair, loses some candidates
- **Band range with manager discretion** — most common, requires calibration guardrails
- **Individual negotiation within band** — flexible, creates pay equity drift over time
---
## Salary Bands: Construction
### Step 1: Define levels
Standard IC levels (adapt to company):
| Level | Title example | Scope |
|-------|--------------|-------|
| L1 | Junior / Associate | Execution with guidance |
| L2 | Mid-level | Independent execution |
| L3 | Senior | Leads workstreams, mentors L1-L2 |
| L4 | Staff / Principal | Cross-team technical leadership |
| L5 | Distinguished / Fellow | Company-wide technical direction |
Management track:
| Level | Title | Scope |
|-------|-------|-------|
| M1 | Manager | Team of 4–8 ICs |
| M2 | Senior Manager | Manager of managers or larger team |
| M3 | Director | Function or large org |
| M4 | VP | Business unit, company-wide |
| M5 | SVP / C-Suite | Executive |
### Step 2: Gather market data
**Data sources (by quality):**
1. **Radford / Aon** — Gold standard. Expensive ($10K+/year). Worth it at Series B+.
2. **Levels.fyi** — Excellent for engineering. Free. Self-reported but large sample.
3. **Glassdoor Salary** — Broad coverage. Less precise for startups.
4. **Pave / Carta Total Comp** — VC-backed companies. Good peer benchmarking.
5. **LinkedIn Salary** — Free tier. Reasonable signal for G&A roles.
6. **Offer letter data** — What candidates are bringing from other companies. Real-time signal.
**What to pull:** P25, P50, P75, P90 for each role × level × geography.
### Step 3: Set band structure
**Band width (range within a level):**
- IC bands: 80–120% of midpoint (i.e., ±20% from center)
- Manager bands: 85–115% of midpoint
- Wider bands allow room for differentiation within level; narrower bands reduce pay equity drift
**Band overlap between levels:**
- 10–20% overlap is normal (top of L2 overlaps with bottom of L3)
- > 30% overlap: your levels are too close together
- No overlap: new hires jump too much between levels (compression risk)
**Example engineering band structure (US, Series B company, P50 target):**
| Level | Band Min | Midpoint | Band Max |
|-------|----------|----------|----------|
| L1 Software Engineer | $90K | $105K | $125K |
| L2 Software Engineer | $115K | $135K | $160K |
| L3 Senior SWE | $150K | $175K | $205K |
| L4 Staff SWE | $195K | $225K $260K |
| M1 Eng Manager | $175K | $205K | $235K |
| M2 Sr Eng Manager | $215K | $250K | $285K |
| M3 Director, Eng | $255K | $300K | $345K |
*Adjust by 15–25% for non-SF/NYC markets. Adjust -40% to -60% for European markets.*
### Step 4: Place employees in bands
**Compa-ratio** = Employee salary / Band midpoint
| Compa-ratio | Interpretation |
|------------|---------------|
| < 0.85 | Below range — immediate risk |
| 0.85–0.95 | Developing in role |
| 0.95–1.05 | Fully performing (target zone) |
| 1.05–1.15 | Senior/expert in role |
| > 1.15 | Above range — flag for review |
**Audit report:** Run quarterly. Flag anyone below 0.85 (flight risk) or above 1.15 (overpaid for level, or needs promotion).
---
## Equity Frameworks for Startups
### Option Basics
**ISO vs NSO:**
- ISO (Incentive Stock Options): For employees. Favorable tax treatment if held 1+ year post-exercise.
- NSO (Non-Qualified Stock Options): For advisors, contractors, sometimes employees. Taxed as ordinary income on exercise.
**Strike price:** Set to 409A valuation at grant. Lower is better for employees. Early employees win on strike price.
**Vesting schedule standards:**
- 4-year vest, 1-year cliff: Standard
- 4-year vest, 6-month cliff: Startup market adapting to faster pace
- 1-year cliff means: nothing until 12 months; monthly or quarterly after
**Post-termination exercise window (PTEW):**
- Standard: 90 days. Often too short for employees who can't afford exercise.
- Better: 1–5 years or until IPO. Use as a talent differentiator.
- Companies extending PTEW: Stripe, Airbnb (pre-IPO), Square, most employee-friendly startups.
### Equity Grant Ranges by Stage and Level
*Expressed as % of fully diluted shares at grant. Ranges vary significantly by market, stage, and funding.*
**Seed stage:**
| Role | Equity % |
|------|----------|
| Co-founder | 20–40% |
| First engineering hire | 0.5–1.5% |
| First non-technical exec hire | 0.25–0.75% |
| IC (L2-L3) | 0.1–0.4% |
| IC (L3-L4) | 0.2–0.6% |
**Series A:**
| Role | Equity % |
|------|----------|
| VP / Head of function | 0.3–0.75% |
| Director | 0.1–0.3% |
| Senior IC (L3) | 0.05–0.15% |
| Mid IC (L2) | 0.02–0.08% |
| Junior IC (L1) | 0.01–0.05% |
**Series B:**
| Role | Equity % |
|------|----------|
| VP / Head of function | 0.1–0.3% |
| Director | 0.05–0.15% |
| Senior IC (L3) | 0.02–0.07% |
| Mid IC (L2) | 0.01–0.03% |
*At Series B+, equity is increasingly expressed in dollar value (grant value = X shares × current 409A). Use Carta or Pulley to model dilution.*
### Equity Refresh Program
**Why it matters:** Employees hired at Series A with 4-year vesting will be fully vested by Series B. No unvested equity = no retention hook.
**When to refresh:**
- After every significant funding round
- Annually for high performers (top 20%)
- After promotion (role-commensurate top-up)
- Counter-offer situations (use carefully — signals you underpaid initially)
**Refresh models:**
1. **Anniversary grant:** Annual cliff-free refresh for all employees above a performance threshold
2. **Evergreen model:** Continuous vesting maintained — refresh annually so employee always has 2–3 years remaining
3. **Event-based:** Refresh tied to milestones (promotion, funding, annual review cycle)
**Dilution awareness:** Every refresh dilutes existing shareholders. Model pool usage quarterly. Replenish option pool before it drops below 10–12% of fully diluted shares.
---
## Total Comp Modeling
### Components of Total Comp
```
Total Compensation = Base Salary
+ Annual Bonus (target %)
+ Equity Value (annualized grant / vesting period)
+ Benefits (employer-paid premiums, retirement match)
+ Allowances (home office, internet, L&D, commuter)
```
### Annualizing Equity Value
For comparison to cash compensation:
```
Annual equity value = (Grant shares × Current 409A price) / Vesting years
```
Example: 10,000 options at $2 strike, current 409A = $8, 4-year vest
- Grant value at current 409A = 10,000 × $8 = $80,000
- Annual value = $80,000 / 4 = $20,000/year
- If base is $150K, total comp is ~$170K/year
*Note: For recruiting purposes, you can use last preferred share price (VC price) to show upside — but be transparent about the difference between 409A and preferred.*
### Benefits Valuation
Frequently undervalued in offers. Quantify explicitly:
| Benefit | Typical employer cost |
|---------|----------------------|
| Health insurance (employee) | $4K–8K/year |
| Health insurance (family) | $15K–25K/year |
| 401K match (4% of salary) | $5K–10K/year |
| L&D budget ($2K/year) | $2K/year |
| Home office stipend ($500) | $500/year |
A $140K offer with family health coverage + 4% 401K match is worth $165K+ total.
---
## Raise and Refresh Process
### Annual Compensation Review Cycle
**Recommended cadence:**
- October/November: Market data refresh, band updates
- November/December: Manager merit recommendations
- December/January: Calibration and approvals
- January/February: Effective date for new salaries + equity grants
**Budget allocation:**
- **Merit budget** (performance-based raises): 3–5% of total payroll typically
- **Market adjustment budget** (fixing below-band salaries): Separate from merit. Non-negotiable to avoid attrition.
- **Promotion budget:** Separate. Promotions should not come from merit pool.
### Merit Increase Guidelines
| Performance Rating | Merit Increase Range |
|-------------------|---------------------|
| 5 – Exceptional | 8–15% |
| 4 – Exceeds | 5–8% |
| 3 – Meets | 2–4% |
| 2 – Needs improvement | 0–1% |
| 1 – Underperforming | 0% (PIP active) |
*Adjust based on compa-ratio. A high performer at P90 of their band gets a smaller increase than a high performer at P50.*
### Compa-Ratio Adjustment Matrix
| Performance \ Compa-Ratio | < 0.90 | 0.90–1.00 | 1.00–1.10 | > 1.10 |
|---------------------------|--------|-----------|-----------|--------|
| Exceptional (5) | 12–15% | 8–12% | 5–8% | 3–5% |
| Exceeds (4) | 8–12% | 5–8% | 3–5% | 1–3% |
| Meets (3) | 5–8% | 3–5% | 2–3% | 0–2% |
| Needs impr (2) | 0–2% | 0–1% | 0% | 0% |
### Promotion vs. Merit — Keep These Separate
**Common mistake:** Using merit budget to fund promotions. This forces a choice between rewarding performance and recognizing level change.
**Promotion increase guidelines:**
- One level (e.g., L2 → L3): 10–20% increase, new equity grant
- Two levels (rare): 20–35% increase, new equity grant at new level
- Manager track (IC → M1): 15–25% increase, new equity grant
**Promotion criteria process:**
1. Manager nominates with written business case
2. Calibration committee reviews cross-functionally
3. HR validates against band (no off-band exceptions without CHRO sign-off)
4. Employee informed before annual review — never surprised at review meeting
### Off-Cycle Adjustments
When to do them:
- Counter-offer situations (see below)
- Competitive intelligence reveals underpay for a specific role
- New market data shows a role significantly under-benchmarked
- Internal equity audit reveals unexplained gaps
**Counter-offer policy:**
Three options:
1. **Match** — Risk: signals you underpay; sets precedent
2. **Partial match** — "We can do X, which is the top of your band" — cleaner
3. **Decline** — Accept the attrition, improve the band for the next hire
**Rule:** If you're regularly in counter-offer conversations, your bands are stale. Fix the bands.
---
## Pay Equity Audit
Run annually. Non-negotiable at Series B+.
**What to audit:**
- Pay gap by gender within each level and function
- Pay gap by ethnicity within each level and function
- Compa-ratio distribution across demographics
- Time-to-promotion by demographic group
**Methodology:**
1. Pull all employee data: level, function, salary, tenure, performance ratings, gender, ethnicity
2. Run regression controlling for level, tenure, and performance
3. Unexplained gap after controls = the problem to fix
4. Flag and remediate within the same review cycle
**Legal exposure:** In many jurisdictions, documented pay gaps without remediation plans are litigation risk. The audit creates a record of intent; remediation closes the risk.
**Remediation budget:** Set aside 0.5–1% of payroll annually for equity adjustments. If you're doing it right, this shrinks over time.
FILE:references/org_design.md
# Org Design Reference
Spans of control, layering decisions, reorgs, title frameworks, career ladders, and the founder→professional management transition.
---
## Core Org Design Principles
1. **Structure follows strategy.** Reorg after strategy shifts, not before.
2. **Optimize for the bottleneck.** Where does work get slow? Design around that.
3. **Minimize coordination cost.** Conway's Law: your org structure becomes your product architecture. Design intentionally.
4. **Bias toward flatness until it breaks.** Adding layers adds cost and slows decisions.
5. **Reorgs have transition costs.** Relationships reset. Count the cost before you restructure.
---
## Spans of Control
Span of control = number of direct reports a manager has.
### Benchmarks
| Role Type | Optimal Span | Min | Max |
|-----------|-------------|-----|-----|
| IC manager (predictable work) | 7–10 | 5 | 12 |
| IC manager (complex/creative work) | 5–7 | 4 | 8 |
| Manager of managers | 4–6 | 3 | 7 |
| VP / Director | 4–7 | 3 | 8 |
| C-Suite | 5–9 | 4 | 10 |
**Too narrow (< 4 ICs):** Over-management, high cost per output, manager becomes a bottleneck
**Too wide (> 12 ICs):** Under-management, degraded 1:1 quality, feedback loops collapse
### Factors that allow wider spans
- Highly autonomous, senior team (L3+ ICs)
- Predictable, well-defined work (support, ops)
- Strong tooling and process (reduces manager overhead)
- Experienced manager
### Factors that require narrower spans
- High-complexity, undefined problems (research, early product)
- Junior or newly promoted team members
- High interdependence between reports (coordination overhead)
- Manager is also an IC contributor (player-coach)
---
## When to Add Management Layers
**The wrong reason to add layers:** "We need to give good people somewhere to grow."
**The right reason:** "This manager has too many direct reports to do the job well."
### Layer triggers by growth stage
**0 → 15 people:** No layers. Everyone reports to founders.
**15 → 30 people:** First managers emerge. Usually technical leads or function leads. Should still be player-coaches.
**30 → 60 people:** Second layer forms. Engineering splits into squads. Sales gets a frontline manager. Each function has a head.
**60 → 150 people:** Director layer becomes necessary in large functions. Engineering VP + Engineering Directors + Team Managers.
**150+ people:** VP layer fully staffed. Senior Director / Director split. Clear IC → M → Senior M → Director → VP paths.
### The Rule of 7
When any manager has 7 or more direct reports and:
- 1:1s are skipped regularly
- Feedback quality drops
- Manager can't answer "how is each person doing?" without checking notes
→ Time to split or hire a manager.
### Management overhead cost
Every manager layer costs 10–15% in decision speed (communication hops).
Every management role without a team = pure overhead.
**Litmus test for each management role:**
- Does this person have at least 4 ICs under them?
- Would removing this role improve decision speed?
- Is this a management job or a "we ran out of IC levels" job?
---
## Functional vs. Product Org Structures
### Functional Structure (by discipline)
```
CEO
├── VP Engineering
│ ├── Backend Team
│ ├── Frontend Team
│ └── DevOps
├── VP Product
│ ├── PM (Feature A)
│ └── PM (Feature B)
└── VP Design
└── UX Designers
```
**Best for:** Early stage, < 100 people, single product
**Advantage:** Deep expertise development, clear career paths per discipline
**Disadvantage:** Cross-functional coordination is heavy; features require synchronization across silos
### Product/Pod Structure (by product area)
```
CEO
├── Product Area A (autonomous team)
│ ├── EM
│ ├── PM
│ └── Designer
├── Product Area B (autonomous team)
│ ├── EM
│ ├── PM
│ └── Designer
└── Platform (shared services)
└── Platform EM + team
```
**Best for:** Multiple products or large user segments, 50+ in product/eng
**Advantage:** Speed and autonomy; less cross-team coordination for most features
**Disadvantage:** Duplication risk; harder to maintain technical coherence; harder career paths
### When to shift from Functional → Product org
- You have 2+ distinct product lines that rarely share features
- Cross-functional feature delivery takes > 3 sprints of coordination overhead
- Teams are > 8 engineers and still waiting on shared resources
### Hybrid / Matrix (avoid unless necessary)
Matrix reporting (e.g., engineer reports to EM + PM) creates accountability confusion. Avoid at < 500 people.
---
## Title Frameworks
### The Problem with Title Inflation
Early startups over-title to compete with cash. "VP of Engineering" with 2 reports. "Head of Marketing" with no team.
**Consequences:**
- Can't add leadership above inflated titles without awkward conversations
- Candidates from mature companies expect scope commensurate with titles
- Internal equity breaks when the same title means different things
### Preventing Title Inflation
**Rule 1:** VP titles require managing managers (not just ICs).
**Rule 2:** Director titles require managing multiple ICs or a large function.
**Rule 3:** No more than one "Head of X" per function.
**Rule 4:** Document scope expectations per title before making offers.
### Engineering Title Ladder (example)
| Title | Level | Scope | Reports |
|-------|-------|-------|---------|
| Software Engineer I | L1 | Executes defined tasks | — |
| Software Engineer II | L2 | Independent delivery | — |
| Senior Software Engineer | L3 | Leads features, mentors | — |
| Staff Software Engineer | L4 | Cross-team technical leadership | — |
| Principal Software Engineer | L5 | Company-wide technical direction | — |
| Distinguished Engineer | L6 | External recognition, defining practice | — |
| Engineering Manager | M1 | Team of 4–8 engineers | 4–8 ICs |
| Senior Engineering Manager | M2 | Larger team or manager of managers | 2–4 managers |
| Director of Engineering | M3 | Functional area | Multiple managers |
| VP of Engineering | M4 | Engineering org | Directors |
| CTO | M5 | Technical organization + strategy | VPs |
**IC vs. Management track:** Explicitly separate. Senior ICs should not need to move to management for career advancement. Staff/Principal/Distinguished track provides this.
### Go-to-Market Title Ladder (example)
| Title | Level | Focus |
|-------|-------|-------|
| SDR / BDR | S1 | Outbound prospecting |
| Account Executive I | S2 | SMB closing |
| Account Executive II | S3 | Mid-market closing |
| Senior Account Executive | S4 | Enterprise closing |
| Principal / Strategic AE | S5 | Named accounts, complex deals |
| Sales Manager | M1 | 6–8 reps |
| Director of Sales | M2 | Multiple teams or segments |
| VP of Sales | M3 | Full sales org |
| CRO | M4 | Revenue org (sales + CS + marketing) |
---
## Career Ladders
A career ladder is a documented set of expectations per level. Not aspirational — behavioral. "What does a P3 engineer do that a P2 doesn't?"
### Why career ladders matter for HR
1. **Retention:** Employees can see where they're going
2. **Consistency:** Managers use the same criteria for promotions
3. **Compensation:** Bands anchor to levels; levels require definitions
4. **Equity:** Removes "who's the manager's favorite" from promotion decisions
### Career Ladder Structure
For each level, define 4 dimensions:
**1. Scope** — How big is the problem space? Team / cross-team / org-wide / company-wide?
**2. Impact** — How does work connect to outcomes? (Task → Feature → Product → Business)
**3. Craft** — Technical/functional skill expectations
**4. Influence** — How does this person improve others? (Self → peers → team → org)
**Example: Senior Software Engineer (L3) vs. Staff Software Engineer (L4)**
| Dimension | L3 (Senior SWE) | L4 (Staff SWE) |
|-----------|----------------|----------------|
| Scope | Owns features or services | Owns technical domains across teams |
| Impact | Ships features that improve user outcomes | Shapes technical direction for a product area |
| Craft | Writes high-quality code, good design skills | Sets coding standards, contributes to architecture |
| Influence | Mentors L1–L2, code reviews | Mentors L3+, identifies org-wide technical gaps |
### How to build a career ladder from scratch
1. **Interview your best performers** — "What do you do that your junior peers don't?" Collect behaviors, not aspirations.
2. **Draft 3 levels** — Don't start with 6. Start with junior, mid, senior. Add staff/principal only when you have enough people to warrant it.
3. **Manager calibration** — Every manager rates 5 current employees against the draft. Gaps surface immediately.
4. **Publish and iterate** — Don't wait for perfection. A 70% ladder shipped is better than a 100% ladder in a drawer.
---
## Reorg Playbook
### When reorgs are necessary
- Strategy pivot requires different team structure (e.g., single product → multi-product)
- Acquisition or team merger
- Function is genuinely too slow due to coordination overhead
- Leadership departure creates structural opportunity
### When reorgs are a mistake
- "We need to shake things up" (disruption for its own sake)
- Avoiding a specific personnel decision (use the right tool)
- Solving a cultural problem with a structural change
- Reacting to one team's complaint without systemic evidence
### Reorg Process (4–8 weeks)
**Week 1–2: Diagnose**
- Map current org: every role, reporting line, team output
- Identify where work is slow, duplicated, or falling through cracks
- Interview 5–10 people across teams: "What takes longer than it should? What decisions are hard to make?"
**Week 3–4: Design options**
- Draft 2–3 structural alternatives
- For each: estimated coordination costs, manager span impact, open roles created
- Validate with CEO + 1–2 trusted operators. Don't crowdsource the design.
**Week 5–6: Decide and prepare**
- Select option; finalize all reporting changes
- Prepare communications for every affected person (individual conversations before all-hands)
- Write the "why" — employees need to understand the business reason, not just the result
**Week 7–8: Communicate and implement**
- Individual conversations with all manager+ changes (first)
- Team-level conversations with managers (second)
- All-hands with full context (third)
- Updated org chart published within 24 hours of announcement
### Communication sequence (non-negotiable)
1. Affected individuals first (private, before anything else)
2. Affected managers second (to prepare for team conversations)
3. Full team/company third (all-hands or company note)
4. External (clients, board) only if materially impacted
**Never:** Email blast first. No individual conversations. Discovered on the org chart.
---
## Founder → Professional Management Transition
The most common scaling failure point in startups.
### Stage 1: Founder-Led (0–30 people)
Founders make all decisions, know everyone personally, set culture through behavior. Works because trust and context are built directly.
**What breaks:**
- Decisions bottleneck at founders
- New hires don't get enough context (founders can't be everywhere)
- Culture transmitted through osmosis, not documentation
### Stage 2: First Managers (30–80 people)
Founders can no longer manage all ICs. First manager layer typically = promoted high performers.
**The "brilliant IC → struggling manager" trap:**
- Individual contributor skills ≠ management skills
- Promoted ICs often continue doing IC work while ignoring management work
- No one holds them accountable to management output (1:1 quality, team health, performance feedback)
**What to do:**
- Explicit manager training before promotion (not after)
- Management KPIs separate from IC KPIs
- Peer community for new managers (monthly cohort session)
- HR check-ins on manager health at 30/60/90 days
### Stage 3: Professional Management (80–200 people)
External hires at Director/VP level bring professional management skills but lack company context.
**Common failure modes:**
- Hired "too senior" — VP who's used to 200-person teams in a 50-person function
- Culture clash — Big-company manager who adds process that kills startup speed
- Authority vacuum — External VP doesn't earn trust; team ignores them; founder continues to bypass hierarchy
**Mitigation:**
- Hiring bar: Has this person scaled from this stage to 2x this stage before? Not managed a team at 2x — built a team to 2x.
- Explicit onboarding on "how we make decisions here"
- 90-day milestones focused on relationship-building before any structural changes
- Founders explicitly hand off ownership and reinforce new manager's authority publicly
### Stage 4: Founder Transition from Operator to Executive
The hardest personal transition. Founder moves from doing to enabling.
**Signs you haven't made the transition:**
- You're still in every technical decision
- Teams come to you instead of their manager for approvals
- You know more about the team's work than the manager does
- Managers feel they need to check in before acting
**What the transition requires:**
- Explicit authority delegation in writing (not just verbal)
- Willingness to let managers make decisions you'd make differently
- Redirecting team members to their manager consistently
- Measuring managers on outcomes, not just process adherence
- Letting managers hire and fire without founder override (except final call on VPs)
FILE:references/people_strategy.md
# People Strategy Reference
Hiring, retention, performance, and remote/hybrid frameworks for each growth stage.
---
## Hiring Strategy by Growth Stage
### Pre-Seed / Seed (1–15 people)
**Who you're hiring:** Generalists who can do multiple jobs. Specialists are a luxury you can't afford unless the specialty is your core product.
**The test:** Could this person be the 5th employee at a startup and thrive? If they need a defined role, clear process, and a manager — not yet.
**Sourcing at this stage:**
- Founder networks first (highest signal, lowest cost)
- Angel List / Wellfound — self-selected for startup risk tolerance
- Referrals from existing employees (offer a referral bonus from day 1)
- GitHub / Dribbble / published work for technical roles
- Avoid: Big job boards, recruiters (unless technical retained search for C-suite)
**Interview process (keep it lean):**
1. 30-min intro call (culture/motivation fit, comp alignment)
2. Take-home or live work sample (2–4 hours max, paid for senior roles)
3. 60-min deep-dive with founders
4. Reference checks (3 calls, not emails — you want the real story)
**Offer timeline:** Decision within 48 hours. Top candidates have multiple offers.
**What to get right:**
- Written job scorecard (outcomes expected in 30/60/90 days) — not a job description
- Equity range disclosed in first conversation
- No exploding offers. Pressure tactics lose good people.
---
### Series A (15–50 people)
**The hiring shift:** You need some specialists now. First management layer emerges. First "culture carries" — people who reinforce what you want to become.
**Critical hires at this stage (in priority order):**
1. VP/Head of Engineering (if founder isn't technical)
2. Head of Product
3. First dedicated recruiter (when you're hiring > 10/year)
4. First Finance/Operations hire
5. Head of Sales (when product-market fit is real)
**Building the recruiting function:**
- First recruiter should be a generalist with hustle, not a specialist
- Set up an ATS (Ashby, Greenhouse, or Lever) before you need it — not after
- Create interview scorecards for every role
- Track: time to fill, offer acceptance rate, source quality
**Common mistakes at Series A:**
- Promoting top ICs to management without management training
- Hiring "brand name" executives who've never operated lean
- Over-indexing on experience, under-indexing on trajectory
- No onboarding process → 90-day regrettable turnover
**Job scorecards (required for every role):**
```
Role: [Title]
Reports to: [Manager]
Start date: [Target]
Why this role now: [Business case in 1-2 sentences]
Outcomes (90 days):
- [Concrete deliverable 1]
- [Concrete deliverable 2]
- [Concrete deliverable 3]
Outcomes (12 months):
- [Strategic impact 1]
- [Strategic impact 2]
Competencies (top 3 only):
- [What, why it matters for THIS role]
- [What, why it matters for THIS role]
- [What, why it matters for THIS role]
Comp range: [Base] + [Equity] + [Benefits summary]
```
---
### Series B (50–150 people)
**The scaling inflection point.** Tribal knowledge breaks. Process matters now. Culture requires deliberate investment.
**What changes:**
- Recruiters become specialists (technical, GTM, exec)
- Manager training becomes non-negotiable
- Performance management needs structure (not just "we'll know it when we see it")
- Onboarding needs to scale without founders in every session
- Comp bands become essential — people are comparing notes
**Hiring velocity benchmarks (Series B):**
| Function | Avg time to fill | Avg interviews | Benchmark offer acceptance |
|----------|-----------------|----------------|---------------------------|
| Engineering IC | 35–45 days | 4–5 rounds | 80–85% |
| Engineering Manager | 45–60 days | 5–6 rounds | 75–80% |
| Sales IC | 25–35 days | 3–4 rounds | 85–90% |
| Sales Manager | 40–55 days | 4–5 rounds | 80–85% |
| G&A (Finance, HR, Ops) | 30–45 days | 3–4 rounds | 85–90% |
**Internal mobility:** By 50 people, start tracking internal promotion rates. Target: 20–30% of manager+ roles filled internally. If it's < 10%, your career development is failing.
---
### Series C+ (150+ people)
**Professional management era.** Founders can't know everyone. Systems and culture carry what personal relationships used to.
**HR function maturity required:**
- Dedicated HRBPs per business unit (1:75–100 employees)
- L&D budget (1–2% of salary budget minimum)
- Succession planning for all VP+ roles
- Structured calibration process for performance reviews
- Total rewards strategy reviewed annually with board
---
## Retention Programs That Actually Work
### What drives retention (in order of impact)
1. **Manager quality** — Gallup: 70% of team engagement variance is explained by the manager. Fix managers first.
2. **Growth trajectory** — People leave when they can't see their next role. Career ladders are retention tools.
3. **Compensation competitiveness** — Being at P25 on salary is a slow leak. Audit annually.
4. **Mission/product belief** — Especially for senior ICs. They want to work on something that matters.
5. **Team quality** — "I stay because of the people I work with." True at every level.
6. **Flexibility** — Location, hours, autonomy. Low cost, high impact.
### What doesn't work (but companies do anyway)
- Pizza parties and ping pong tables
- "Perks" that substitute for salary
- Annual reviews with no action on feedback
- Forced fun events
- Vague "culture improvement" initiatives without specific behavior changes
### The 30-60-90 Onboarding Framework
Structured onboarding cuts 90-day turnover by 50%+.
**Days 1–30: Learn**
- Complete admin setup (day 1, before lunch)
- Meet all key stakeholders (scheduled by their manager, not on the new hire)
- Understand: business model, current priorities, team processes, how success is measured
- No deliverables expected. Learning is the job.
- Weekly 1:1 with manager: "What's confusing? What do you need?"
**Days 31–60: Contribute**
- First real project (scoped to be completable)
- Present findings or work to the team
- Identify one process that could be improved (observation only — don't fix yet)
- 30-day check-in: formal feedback from manager
**Days 61–90: Lead**
- Own a deliverable end-to-end
- Offer one specific improvement recommendation with data
- 90-day review: mutual assessment — manager on new hire, new hire on onboarding
- Set 6-month goals
### Stay Interviews (underused, high ROI)
Run with every employee once per year. Not their manager — HR or skip-level.
**Questions that surface real risk:**
- "What's keeping you here?"
- "What would make you consider leaving?"
- "What's one thing your manager could do differently?"
- "Is your role what you expected when you joined?"
- "What career path do you want? Are we helping you get there?"
- "Are you fairly compensated? Do you know how you'd get a raise?"
**Act on answers within 30 days or don't ask.** Unanswered feedback is worse than no feedback.
### Exit Interviews — What to Actually Learn
Skip the happiness survey. Ask these:
- "When did you first think about leaving?"
- "Was there a specific event that triggered your decision?"
- "What could we have done to retain you?"
- "Where are you going and why?" (What does the other offer have that we don't?)
- "Would you recommend us as an employer? Why or why not?"
Track exit themes by manager. If one manager's exits cite "micromanagement" three times — that's data.
---
## Performance Management
### The System That Works
**Continuous > annual.** Annual reviews with no mid-year touchpoints are theater.
**Structure:**
- **Weekly 1:1s** (30 min): blockers, priorities, relationship
- **Monthly check-ins** (1 hr): progress against goals, feedback exchange
- **Quarterly reviews** (formal): written self-assessment + manager assessment + goal revision
- **Annual calibration** (rating + comp): cross-manager calibration session, then individual conversations
### Calibration Sessions
**Purpose:** Prevent manager bias. Ensure "exceeds expectations" means the same thing across teams.
**Process:**
1. Managers submit preliminary ratings independently
2. HR facilitates 2-hr calibration with all managers in a function
3. Managers must justify outliers (top and bottom)
4. Ratings adjusted for consistency
5. Managers deliver final ratings with rationale
**Distribution guidance (enforce with calibration):**
- Exceptional (5): < 10% — if everyone's exceptional, no one is
- Exceeds (4): 20–25%
- Meets (3): 55–65%
- Needs improvement (2): 8–12%
- Underperforming (1): 2–5%
### Managing Underperformers
**The most avoided management task. And the most damaging when avoided.**
High performers notice when underperformers are tolerated. They leave.
**The 4-step framework:**
**Step 1: Diagnose before acting** (Week 1–2)
- Is this a skill gap (can't do it) or a will gap (won't do it)?
- Skill gap → training, clearer expectations, different role
- Will gap → direct feedback, clear consequences, then PIP
**Step 2: Direct feedback conversation** (Week 2–3)
- Specific: "Your last 3 sprint deliveries were 40% incomplete"
- Not: "You're not meeting expectations"
- Document. Send written summary after every feedback conversation.
**Step 3: Performance Improvement Plan (PIP)**
Required when: two rounds of direct feedback haven't produced change.
PIP structure:
```
Name: [Employee]
Manager: [Name]
Date: [Start]
Review date: [30/60 days out]
Current performance issues:
- [Specific, observable behavior with examples and dates]
- [Metric not met: target X, actual Y for Z weeks]
Required improvements:
- [Specific, measurable outcome 1] by [date]
- [Specific, measurable outcome 2] by [date]
Support provided:
- [Training, coaching, additional resources]
Consequences if not met: [Role change / separation]
Check-in schedule: [Weekly with manager + HR]
```
**Step 4: Exit or role change**
- If PIP milestones not met: proceed to separation
- Don't extend PIPs indefinitely — it's unfair to the employee and the team
- Offer a graceful exit where possible: "This role isn't the right fit. Here's a package and a reference."
**What not to do:**
- "Quiet manage out" without clear feedback (legally risky, unfair)
- PIP as a formality before termination (if you know you're firing them, just do it)
- Tolerating underperformance "because we're understaffed" (it makes understaffing worse)
---
## Remote / Hybrid Strategy
### The question isn't "remote or not" — it's "what kind of collaboration does our work require?"
**Work type taxonomy:**
| Work type | Remote-compatible? | Hybrid compatible? |
|-----------|-------------------|-------------------|
| Deep individual work (coding, writing, analysis) | Yes | Yes |
| Async collaboration (code review, doc review) | Yes | Yes |
| Synchronous problem-solving (debugging, design) | Yes (video) | Yes |
| Relationship-building (onboarding, new team) | Harder | Yes |
| Executive alignment, strategy | Harder | Yes — quarterly in-person |
| Sales (enterprise, relationship-based) | No | Depends on market |
### Making Hybrid Work (Not Just a Policy)
**The failure mode:** "Hybrid" = go to office on Tuesday/Thursday, but no one coordinates, all meetings are still Zoom anyway.
**What actually works:**
1. **Anchor days with purpose** — Office days should have things that require the office: workshops, team rituals, whiteboarding sessions. Not just "presence."
2. **Async-first culture, not async-only** — Document decisions. Write things down. Use Loom for walkthroughs. Reduce "quick sync" meetings.
3. **Equal experience for remote participants** — If some are in the room and some are on video, the remote folks are second-class. Either everyone's remote or set up rooms properly.
4. **Manager standards for remote teams:**
- 1:1s are non-negotiable (video, not async)
- Over-communicate on priorities (people can't absorb hallway context)
- Write down decisions (remote employees miss casual office decisions)
- Recognize work publicly (Slack shoutouts, all-hands wins)
### Remote Compensation Philosophy (pick one, be explicit)
**Option A: Location-based pay**
Pay based on where the employee lives. Lower cost in lower-cost markets. Harder to hire in high-cost cities.
**Option B: Role-based (location-neutral)**
One band for each role regardless of location. Simpler, more equitable. Higher overall payroll cost.
**Option C: Zone-based**
Define 2–3 geographic zones (e.g., Tier 1 cities, Tier 2 cities, international). Set bands per zone. Common at mid-stage startups.
**The wrong answer:** No stated policy, and every offer is negotiated individually. Creates pay equity problems fast.
FILE:scripts/comp_benchmarker.py
#!/usr/bin/env python3
"""
Compensation Benchmarker
========================
Salary benchmarking and total comp modeling for startup teams.
Analyzes pay equity, compa-ratios, and total comp vs. market.
Usage:
python comp_benchmarker.py # Run with built-in sample data
python comp_benchmarker.py --config roster.json # Load from JSON
python comp_benchmarker.py --help
Output: Band compliance report, compa-ratio distribution, pay equity flags,
equity value analysis, and total comp vs. market.
"""
import argparse
import json
import csv
import io
import sys
from dataclasses import dataclass, field, asdict
from typing import Optional
from datetime import date
import math
# ---------------------------------------------------------------------------
# Data structures
# ---------------------------------------------------------------------------
@dataclass
class BandDefinition:
"""Salary band for a role level."""
level: str # L1, L2, L3, L4, M1, M2, M3, VP
function: str # Engineering, Sales, Product, G&A, Marketing, CS
band_min: int # Annual USD
band_mid: int # P50 anchor
band_max: int # Band ceiling
market_p25: int # Market 25th percentile
market_p50: int # Market median (should align with band_mid for P50 strategy)
market_p75: int # Market 75th percentile
location_zone: str # Tier1 (SF/NYC), Tier2 (Austin/Denver), Tier3 (Remote/other), EU
@dataclass
class Employee:
"""One employee record."""
id: str
name: str
role: str
level: str
function: str
location_zone: str
base_salary: int
bonus_target_pct: float # % of base
equity_shares: int # Total unvested options/RSUs
equity_strike: float # Strike price (0 for RSUs)
equity_current_409a: float # Current 409A share price
equity_vest_years_remaining: float # How many years of vesting remain
benefits_annual: int # Employer-paid benefits cost
gender: str # M/F/NB/Undisclosed (for equity audit)
ethnicity: str # For equity audit — can be "Undisclosed"
tenure_years: float
performance_rating: int # 1–5
last_raise_months_ago: int
last_equity_refresh_months_ago: Optional[int] = None
@dataclass
class CompRoster:
company: str
as_of_date: str # ISO date
funding_stage: str # Seed, Series A, Series B, etc.
comp_philosophy_target: str # P50, P65, P75 — your target percentile
preferred_stock_price: float # Last round price (for offer modeling)
employees: list[Employee] = field(default_factory=list)
bands: list[BandDefinition] = field(default_factory=list)
# ---------------------------------------------------------------------------
# Band lookup
# ---------------------------------------------------------------------------
def find_band(roster: CompRoster, level: str, function: str, zone: str) -> Optional[BandDefinition]:
"""Find best-matching band. Falls back to any matching level+function if zone not found."""
matches = [b for b in roster.bands if b.level == level and b.function == function and b.location_zone == zone]
if matches:
return matches[0]
# Fallback: same level+function, any zone
matches = [b for b in roster.bands if b.level == level and b.function == function]
if matches:
return matches[0]
# Fallback: same level, any function
matches = [b for b in roster.bands if b.level == level]
if matches:
return matches[0]
return None
# ---------------------------------------------------------------------------
# Compensation analysis
# ---------------------------------------------------------------------------
def compa_ratio(salary: int, band_mid: int) -> float:
return salary / band_mid if band_mid > 0 else 0.0
def band_position(salary: int, band_min: int, band_max: int) -> float:
"""Position in band: 0.0 = at min, 1.0 = at max."""
if band_max == band_min:
return 0.5
return (salary - band_min) / (band_max - band_min)
def annualized_equity_value(emp: Employee) -> int:
"""Current 409A value of unvested equity, annualized."""
if emp.equity_vest_years_remaining <= 0:
return 0
if emp.equity_current_409a > emp.equity_strike:
intrinsic = (emp.equity_current_409a - emp.equity_strike) * emp.equity_shares
else:
# Options underwater — still show at current FMV for RSUs or future value for options
intrinsic = emp.equity_current_409a * emp.equity_shares if emp.equity_strike == 0 else 0
return int(intrinsic / emp.equity_vest_years_remaining)
def total_comp(emp: Employee) -> int:
bonus = int(emp.base_salary * emp.bonus_target_pct)
equity = annualized_equity_value(emp)
return emp.base_salary + bonus + equity + emp.benefits_annual
def analyze_employee(emp: Employee, roster: CompRoster) -> dict:
band = find_band(roster, emp.level, emp.function, emp.location_zone)
result = {
"id": emp.id,
"name": emp.name,
"role": emp.role,
"level": emp.level,
"function": emp.function,
"zone": emp.location_zone,
"base": emp.base_salary,
"bonus_target": int(emp.base_salary * emp.bonus_target_pct),
"equity_annual": annualized_equity_value(emp),
"benefits": emp.benefits_annual,
"total_comp": total_comp(emp),
"performance": emp.performance_rating,
"tenure_years": emp.tenure_years,
"last_raise_months": emp.last_raise_months_ago,
"band": band,
"compa_ratio": None,
"band_position": None,
"vs_market_p50": None,
"flags": [],
}
if band:
cr = compa_ratio(emp.base_salary, band.band_mid)
bp = band_position(emp.base_salary, band.band_min, band.band_max)
result["compa_ratio"] = round(cr, 3)
result["band_position"] = round(bp, 3)
result["vs_market_p50"] = round((emp.base_salary - band.market_p50) / band.market_p50 * 100, 1)
# Flags
if emp.base_salary < band.band_min:
result["flags"].append(("CRITICAL", "Base below band minimum — immediate attrition risk"))
elif cr < 0.88:
result["flags"].append(("HIGH", f"Compa-ratio {cr:.2f} — significantly below midpoint"))
elif cr < 0.93:
result["flags"].append(("MEDIUM", f"Compa-ratio {cr:.2f} — below target zone (0.95–1.05)"))
if emp.base_salary > band.band_max:
result["flags"].append(("HIGH", "Base above band maximum — review for promotion or band update"))
if emp.performance_rating >= 4 and cr < 0.95:
result["flags"].append(("HIGH", f"High performer (rating {emp.performance_rating}) underpaid — flight risk"))
if emp.last_raise_months_ago > 18:
result["flags"].append(("MEDIUM", f"No raise in {emp.last_raise_months_ago} months — review due"))
if emp.equity_vest_years_remaining < 1.0 and (emp.last_equity_refresh_months_ago is None or emp.last_equity_refresh_months_ago > 24):
result["flags"].append(("HIGH", "Equity nearly fully vested with no refresh — retention hook gone"))
else:
result["flags"].append(("INFO", "No band found for this level/function/zone"))
return result
# ---------------------------------------------------------------------------
# Aggregate analysis
# ---------------------------------------------------------------------------
def pay_equity_audit(analyses: list[dict], employees: list[Employee]) -> dict:
"""Simple pay equity analysis by gender and ethnicity."""
emp_by_id = {e.id: e for e in employees}
def group_stats(group_key_fn):
groups: dict[str, list[float]] = {}
for a in analyses:
if a["compa_ratio"] is None:
continue
emp = emp_by_id.get(a["id"])
if not emp:
continue
key = group_key_fn(emp)
if key not in groups:
groups[key] = []
groups[key].append(a["compa_ratio"])
return {k: {"n": len(v), "avg_cr": round(sum(v)/len(v), 3), "min_cr": round(min(v), 3), "max_cr": round(max(v), 3)}
for k, v in groups.items() if v}
gender_stats = group_stats(lambda e: e.gender)
ethnicity_stats = group_stats(lambda e: e.ethnicity)
# Compute gap vs. the largest group
def compute_gap(stats: dict) -> dict[str, float]:
if not stats:
return {}
largest = max(stats.items(), key=lambda x: x[1]["n"])
ref_cr = largest[1]["avg_cr"]
return {k: round((v["avg_cr"] - ref_cr) / ref_cr * 100, 1) for k, v in stats.items()}
gender_gaps = compute_gap(gender_stats)
ethnicity_gaps = compute_gap(ethnicity_stats)
return {
"gender": gender_stats,
"gender_gaps_pct": gender_gaps,
"ethnicity": ethnicity_stats,
"ethnicity_gaps_pct": ethnicity_gaps,
}
def compa_ratio_distribution(analyses: list[dict]) -> dict:
crs = [a["compa_ratio"] for a in analyses if a["compa_ratio"] is not None]
if not crs:
return {}
buckets = {
"< 0.85 (below band)": 0,
"0.85–0.94 (developing)": 0,
"0.95–1.05 (target zone)": 0,
"1.06–1.15 (senior in role)": 0,
"> 1.15 (above band)": 0,
}
for cr in crs:
if cr < 0.85:
buckets["< 0.85 (below band)"] += 1
elif cr < 0.95:
buckets["0.85–0.94 (developing)"] += 1
elif cr <= 1.05:
buckets["0.95–1.05 (target zone)"] += 1
elif cr <= 1.15:
buckets["1.06–1.15 (senior in role)"] += 1
else:
buckets["> 1.15 (above band)"] += 1
avg = sum(crs) / len(crs)
return {"distribution": buckets, "avg_compa_ratio": round(avg, 3), "n": len(crs)}
# ---------------------------------------------------------------------------
# Report output
# ---------------------------------------------------------------------------
def fmt(n) -> str:
return f",.0f"
def bar(value: float, width: int = 20) -> str:
filled = min(width, max(0, int(value * width)))
return "█" * filled + "░" * (width - filled)
def print_report(roster: CompRoster):
WIDTH = 76
SEP = "=" * WIDTH
sep = "-" * WIDTH
analyses = [analyze_employee(e, roster) for e in roster.employees]
cr_dist = compa_ratio_distribution(analyses)
equity_audit = pay_equity_audit(analyses, roster.employees)
print(SEP)
print(f" COMPENSATION BENCHMARKING REPORT — {roster.company}")
print(f" As of: {roster.as_of_date} | Stage: {roster.funding_stage} | Target: {roster.comp_philosophy_target}")
print(SEP)
# Summary stats
total_emps = len(roster.employees)
flagged = sum(1 for a in analyses if any(s in ["CRITICAL", "HIGH"] for s, _ in a["flags"]))
total_payroll = sum(e.base_salary for e in roster.employees)
avg_total_comp = sum(a["total_comp"] for a in analyses) // total_emps if total_emps else 0
print(f"\n[ SUMMARY ]")
print(sep)
print(f" Employees analyzed: {total_emps}")
print(f" Flagged (critical/high): {flagged}")
print(f" Total base payroll: {fmt(total_payroll)}/year")
print(f" Avg total comp: {fmt(avg_total_comp)}/year")
if cr_dist:
print(f" Avg compa-ratio: {cr_dist['avg_compa_ratio']:.3f}")
# Compa-ratio distribution
if cr_dist:
print(f"\n[ COMPA-RATIO DISTRIBUTION ]")
print(sep)
total_n = cr_dist["n"]
for label, count in cr_dist["distribution"].items():
pct = count / total_n if total_n else 0
bar_str = bar(pct, 25)
print(f" {label:<30} {bar_str} {count:3d} ({pct*100:4.0f}%)")
# Pay equity audit
print(f"\n[ PAY EQUITY AUDIT ]")
print(sep)
print(f" By Gender:")
for group, stats in equity_audit["gender"].items():
gap = equity_audit["gender_gaps_pct"].get(group, 0.0)
gap_str = f" gap: {gap:+.1f}%" if gap != 0 else " (reference group)"
flag = " ⚠" if abs(gap) > 5 else ""
print(f" {group:<15} n={stats['n']} avg_CR={stats['avg_cr']:.3f}{gap_str}{flag}")
print(f"\n By Ethnicity:")
for group, stats in equity_audit["ethnicity"].items():
gap = equity_audit["ethnicity_gaps_pct"].get(group, 0.0)
gap_str = f" gap: {gap:+.1f}%" if gap != 0 else " (reference group)"
flag = " ⚠" if abs(gap) > 5 else ""
print(f" {group:<20} n={stats['n']} avg_CR={stats['avg_cr']:.3f}{gap_str}{flag}")
print(f"\n ⚠ = gap > 5%. Investigate with regression controlling for level, tenure, and performance.")
# Employee detail with flags
print(f"\n[ EMPLOYEE DETAIL ]")
print(sep)
# Group by function
functions = sorted(set(e.function for e in roster.employees))
for fn in functions:
fn_analyses = [a for a in analyses if a["function"] == fn]
if not fn_analyses:
continue
print(f"\n ── {fn} ──")
print(f" {'Name':<22} {'Role':<28} {'Lvl':<5} {'Base':>10} {'TotalComp':>11} {'CR':>6} {'Perf':>5} Flags")
print(f" {'-'*22} {'-'*28} {'-'*5} {'-'*10} {'-'*11} {'-'*6} {'-'*5} {'-'*20}")
for a in sorted(fn_analyses, key=lambda x: -x["base"]):
cr_str = f"{a['compa_ratio']:.2f}" if a["compa_ratio"] else "N/A"
flag_summary = ", ".join(s for s, _ in a["flags"] if s in ("CRITICAL", "HIGH", "MEDIUM"))
flag_str = flag_summary if flag_summary else "OK"
print(f" {a['name']:<22} {a['role']:<28} {a['level']:<5} "
f"{fmt(a['base']):>10} {fmt(a['total_comp']):>11} {cr_str:>6} {a['performance']:>5} {flag_str}")
# Print flag detail for critical/high
for severity, msg in a["flags"]:
if severity in ("CRITICAL", "HIGH"):
print(f" {'':>22} ↳ [{severity}] {msg}")
# Action items
critical = [(a["name"], msg) for a in analyses for sev, msg in a["flags"] if sev == "CRITICAL"]
high = [(a["name"], msg) for a in analyses for sev, msg in a["flags"] if sev == "HIGH"]
medium = [(a["name"], msg) for a in analyses for sev, msg in a["flags"] if sev == "MEDIUM"]
print(f"\n[ ACTION ITEMS ]")
print(sep)
if critical:
print(f"\n CRITICAL — Address this review cycle:")
for name, msg in critical:
print(f" • {name}: {msg}")
if high:
print(f"\n HIGH — Address within 30 days:")
for name, msg in high[:10]:
print(f" • {name}: {msg}")
if len(high) > 10:
print(f" ... and {len(high)-10} more")
if medium:
print(f"\n MEDIUM — Address in next comp cycle:")
for name, msg in medium[:8]:
print(f" • {name}: {msg}")
if len(medium) > 8:
print(f" ... and {len(medium)-8} more")
if not critical and not high and not medium:
print(f"\n No critical or high-severity issues. Compensation appears well-managed.")
# Remediation cost estimate
below_min = [a for a in analyses if a["band"] and a["base"] < a["band"].band_min]
below_mid = [a for a in analyses if a["compa_ratio"] and a["compa_ratio"] < 0.90]
if below_min or below_mid:
print(f"\n[ REMEDIATION COST ESTIMATE ]")
print(sep)
if below_min:
cost_to_min = sum(a["band"].band_min - a["base"] for a in below_min)
print(f" Cost to bring below-minimum to band min: {fmt(cost_to_min)}/year ({len(below_min)} employees)")
if below_mid:
cost_to_90 = sum(int(a["band"].band_mid * 0.90) - a["base"] for a in below_mid if a["base"] < int(a["band"].band_mid * 0.90))
cost_to_90 = max(0, cost_to_90)
print(f" Cost to bring CR < 0.90 to CR = 0.90: {fmt(cost_to_90)}/year ({len(below_mid)} employees)")
total_payroll_impact = sum(e.base_salary for e in roster.employees)
total_remediation = (below_min and cost_to_min or 0)
print(f"\n Total payroll before remediation: {fmt(total_payroll_impact)}/year")
print(f" Remediation as % of payroll: {total_remediation/total_payroll_impact*100:.1f}%")
print(f"\n{SEP}\n")
def export_csv(roster: CompRoster) -> str:
analyses = [analyze_employee(e, roster) for e in roster.employees]
output = io.StringIO()
writer = csv.writer(output)
writer.writerow(["ID", "Name", "Role", "Level", "Function", "Zone",
"Base", "Bonus Target", "Equity Annual", "Benefits", "Total Comp",
"Compa Ratio", "Band Position", "vs Market P50 %",
"Performance", "Tenure Years", "Last Raise (mo)",
"Gender", "Ethnicity", "Critical Flags", "High Flags"])
for a, e in zip(analyses, roster.employees):
critical_flags = "; ".join(msg for sev, msg in a["flags"] if sev == "CRITICAL")
high_flags = "; ".join(msg for sev, msg in a["flags"] if sev == "HIGH")
writer.writerow([a["id"], a["name"], a["role"], a["level"], a["function"], a["zone"],
a["base"], a["bonus_target"], a["equity_annual"], a["benefits"], a["total_comp"],
a["compa_ratio"], a["band_position"], a["vs_market_p50"],
a["performance"], a["tenure_years"], a["last_raise_months"],
e.gender, e.ethnicity, critical_flags, high_flags])
return output.getvalue()
# ---------------------------------------------------------------------------
# Sample data
# ---------------------------------------------------------------------------
def build_sample_roster() -> CompRoster:
roster = CompRoster(
company="AcmeTech (Series A)",
as_of_date=date.today().isoformat(),
funding_stage="Series A",
comp_philosophy_target="P50",
preferred_stock_price=8.50,
)
# Bands (Engineering, P50 target, Tier1 = SF/NYC)
roster.bands = [
BandDefinition("L2", "Engineering", 115_000, 132_000, 155_000, 110_000, 132_000, 155_000, "Tier1"),
BandDefinition("L3", "Engineering", 148_000, 170_000, 198_000, 145_000, 170_000, 198_000, "Tier1"),
BandDefinition("L4", "Engineering", 185_000, 215_000, 248_000, 182_000, 215_000, 250_000, "Tier1"),
BandDefinition("M1", "Engineering", 170_000, 195_000, 225_000, 168_000, 195_000, 225_000, "Tier1"),
BandDefinition("L2", "Engineering", 95_000, 108_000, 125_000, 92_000, 108_000, 126_000, "Tier2"),
BandDefinition("L3", "Engineering", 122_000, 140_000, 162_000, 120_000, 140_000, 162_000, "Tier2"),
BandDefinition("L2", "Sales", 80_000, 92_000, 108_000, 78_000, 92_000, 108_000, "Tier1"),
BandDefinition("L3", "Sales", 95_000, 110_000, 128_000, 93_000, 110_000, 128_000, "Tier1"),
BandDefinition("M1", "Sales", 130_000, 150_000, 172_000, 128_000, 150_000, 172_000, "Tier1"),
BandDefinition("L2", "Product", 125_000, 145_000, 168_000, 123_000, 145_000, 168_000, "Tier1"),
BandDefinition("L3", "Product", 155_000, 178_000, 205_000, 153_000, 178_000, 205_000, "Tier1"),
BandDefinition("L2", "G&A", 85_000, 98_000, 115_000, 83_000, 98_000, 115_000, "Tier1"),
BandDefinition("L3", "G&A", 110_000, 128_000, 148_000, 108_000, 128_000, 148_000, "Tier1"),
]
roster.employees = [
# Engineering — mix of scenarios
Employee("E001", "Aarav Shah", "Senior SWE (Backend)", "L3", "Engineering", "Tier1",
base_salary=168_000, bonus_target_pct=0.0, equity_shares=40_000,
equity_strike=1.50, equity_current_409a=6.80, equity_vest_years_remaining=2.5,
benefits_annual=18_000, gender="M", ethnicity="Asian",
tenure_years=2.5, performance_rating=4, last_raise_months_ago=14,
last_equity_refresh_months_ago=None),
Employee("E002", "Yuki Tanaka", "Senior SWE (Frontend)", "L3", "Engineering", "Tier1",
base_salary=152_000, bonus_target_pct=0.0, equity_shares=30_000,
equity_strike=2.20, equity_current_409a=6.80, equity_vest_years_remaining=0.5,
benefits_annual=18_000, gender="F", ethnicity="Asian",
tenure_years=3.8, performance_rating=5, last_raise_months_ago=11,
last_equity_refresh_months_ago=30),
# Note: Yuki is high performer, near-vested, no recent refresh — flag expected
Employee("E003", "Marcus Johnson", "SWE II (Backend)", "L2", "Engineering", "Tier1",
base_salary=110_000, bonus_target_pct=0.0, equity_shares=15_000,
equity_strike=2.50, equity_current_409a=6.80, equity_vest_years_remaining=3.0,
benefits_annual=15_000, gender="M", ethnicity="Black",
tenure_years=1.2, performance_rating=3, last_raise_months_ago=12,
last_equity_refresh_months_ago=None),
# Note: Below band midpoint, recently hired — developing flag
Employee("E004", "Priya Nair", "Staff SWE", "L4", "Engineering", "Tier1",
base_salary=222_000, bonus_target_pct=0.0, equity_shares=60_000,
equity_strike=0.80, equity_current_409a=6.80, equity_vest_years_remaining=2.0,
benefits_annual=18_000, gender="F", ethnicity="Asian",
tenure_years=4.2, performance_rating=5, last_raise_months_ago=8,
last_equity_refresh_months_ago=8),
Employee("E005", "Tom Rivera", "SWE II (Platform)", "L2", "Engineering", "Tier2",
base_salary=88_000, bonus_target_pct=0.0, equity_shares=12_000,
equity_strike=3.00, equity_current_409a=6.80, equity_vest_years_remaining=2.5,
benefits_annual=14_000, gender="M", ethnicity="Hispanic",
tenure_years=1.8, performance_rating=4, last_raise_months_ago=22,
last_equity_refresh_months_ago=None),
# Note: No raise in 22 months, high performer — flag expected
Employee("E006", "Sarah Kim", "Eng Manager", "M1", "Engineering", "Tier1",
base_salary=192_000, bonus_target_pct=0.10, equity_shares=35_000,
equity_strike=1.20, equity_current_409a=6.80, equity_vest_years_remaining=1.8,
benefits_annual=18_000, gender="F", ethnicity="Asian",
tenure_years=2.8, performance_rating=4, last_raise_months_ago=9,
last_equity_refresh_months_ago=9),
# Sales
Employee("S001", "David Chen", "Account Executive (MM)", "L3", "Sales", "Tier1",
base_salary=105_000, bonus_target_pct=0.50, equity_shares=8_000,
equity_strike=3.50, equity_current_409a=6.80, equity_vest_years_remaining=2.0,
benefits_annual=15_000, gender="M", ethnicity="Asian",
tenure_years=1.5, performance_rating=3, last_raise_months_ago=15,
last_equity_refresh_months_ago=None),
Employee("S002", "Amara Osei", "AE (Mid-Market)", "L3", "Sales", "Tier1",
base_salary=98_000, bonus_target_pct=0.50, equity_shares=6_000,
equity_strike=3.50, equity_current_409a=6.80, equity_vest_years_remaining=2.5,
benefits_annual=15_000, gender="F", ethnicity="Black",
tenure_years=1.0, performance_rating=4, last_raise_months_ago=12,
last_equity_refresh_months_ago=None),
# Note: High performer, significantly below midpoint — flag expected
Employee("S003", "Jordan Blake", "Sales Manager", "M1", "Sales", "Tier1",
base_salary=155_000, bonus_target_pct=0.20, equity_shares=20_000,
equity_strike=2.00, equity_current_409a=6.80, equity_vest_years_remaining=1.5,
benefits_annual=16_000, gender="NB", ethnicity="White",
tenure_years=2.2, performance_rating=3, last_raise_months_ago=10,
last_equity_refresh_months_ago=10),
# Product
Employee("P001", "Nina Patel", "Senior PM", "L3", "Product", "Tier1",
base_salary=176_000, bonus_target_pct=0.10, equity_shares=22_000,
equity_strike=1.80, equity_current_409a=6.80, equity_vest_years_remaining=2.0,
benefits_annual=17_000, gender="F", ethnicity="Asian",
tenure_years=2.0, performance_rating=4, last_raise_months_ago=12,
last_equity_refresh_months_ago=12),
# G&A
Employee("G001", "Chris Mueller", "Finance Manager", "L3", "G&A", "Tier1",
base_salary=125_000, bonus_target_pct=0.10, equity_shares=10_000,
equity_strike=2.80, equity_current_409a=6.80, equity_vest_years_remaining=3.0,
benefits_annual=16_000, gender="M", ethnicity="White",
tenure_years=1.5, performance_rating=3, last_raise_months_ago=15,
last_equity_refresh_months_ago=None),
Employee("G002", "Fatima Al-Hassan", "HR Operations", "L2", "G&A", "Tier1",
base_salary=82_000, bonus_target_pct=0.08, equity_shares=5_000,
equity_strike=4.00, equity_current_409a=6.80, equity_vest_years_remaining=3.5,
benefits_annual=14_000, gender="F", ethnicity="Middle Eastern",
tenure_years=0.8, performance_rating=3, last_raise_months_ago=8,
last_equity_refresh_months_ago=None),
# Note: Below band minimum — critical flag expected
]
return roster
# ---------------------------------------------------------------------------
# CLI
# ---------------------------------------------------------------------------
def load_roster_from_json(path: str) -> CompRoster:
with open(path) as f:
data = json.load(f)
employees = [Employee(**e) for e in data.pop("employees", [])]
bands = [BandDefinition(**b) for b in data.pop("bands", [])]
roster = CompRoster(**data)
roster.employees = employees
roster.bands = bands
return roster
def main():
parser = argparse.ArgumentParser(
description="Compensation Benchmarker — salary analysis and pay equity audit",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
python comp_benchmarker.py # Run sample roster
python comp_benchmarker.py --config roster.json # Load from JSON
python comp_benchmarker.py --export-csv # Output CSV
python comp_benchmarker.py --export-json # Output JSON template
"""
)
parser.add_argument("--config", help="Path to JSON roster file")
parser.add_argument("--export-csv", action="store_true", help="Export analysis as CSV")
parser.add_argument("--export-json", action="store_true", help="Export sample roster as JSON template")
args = parser.parse_args()
if args.config:
roster = load_roster_from_json(args.config)
else:
roster = build_sample_roster()
if args.export_json:
data = asdict(roster)
print(json.dumps(data, indent=2))
return
if args.export_csv:
print(export_csv(roster))
return
print_report(roster)
if __name__ == "__main__":
main()
FILE:scripts/hiring_plan_modeler.py
#!/usr/bin/env python3
"""
Hiring Plan Modeler
===================
Builds hiring plans from business goals with cost projections.
Outputs quarterly headcount plan, cost model, and risk assessment.
Usage:
python hiring_plan_modeler.py # Run with built-in sample data
python hiring_plan_modeler.py --config plan.json # Load from JSON config
python hiring_plan_modeler.py --help
"""
import argparse
import json
import sys
from dataclasses import dataclass, field, asdict
from datetime import datetime, date
from typing import Optional
import csv
import io
# ---------------------------------------------------------------------------
# Data structures
# ---------------------------------------------------------------------------
@dataclass
class HireTarget:
"""One planned hire."""
role: str
level: str # L1, L2, L3, L4, M1, M2, M3, VP, C-Suite
function: str # Engineering, Sales, Product, G&A, Marketing, CS
quarter: str # Q1-2025, Q2-2025, etc.
base_salary: int # Annual, USD
bonus_pct: float # % of base (e.g., 0.10 for 10%)
equity_annual_usd: int # Annualized equity value at current 409A
benefits_annual: int # Employer-paid benefits
recruiter_fee_pct: float= 0.20 # Agency fee if used (0 for internal recruiter)
ramp_months: int = 3 # Months to full productivity
priority: str = "High" # High / Medium / Low
business_case: str = ""
open_to_internal: bool = False
@dataclass
class HiringPlan:
company: str
plan_period: str # e.g., "2025 Annual"
current_headcount: int
target_revenue: int # Annual target revenue ($)
current_revenue: int # Current ARR ($)
hires: list[HireTarget] = field(default_factory=list)
# Cost overheads beyond comp
overhead_rate: float = 0.25 # Workspace, software, onboarding overhead as % of base
internal_recruiter_cost: int = 0 # If you have an internal recruiter, annual cost
# ---------------------------------------------------------------------------
# Computation
# ---------------------------------------------------------------------------
def quarter_to_sortkey(q: str) -> tuple[int, int]:
"""Parse 'Q2-2025' → (2025, 2)"""
parts = q.upper().split("-")
if len(parts) == 2:
q_num = int(parts[0].replace("Q", ""))
year = int(parts[1])
return (year, q_num)
return (9999, 9)
def get_quarters(hires: list[HireTarget]) -> list[str]:
"""Return sorted unique quarters from hire list."""
quarters = sorted(set(h.quarter for h in hires), key=quarter_to_sortkey)
return quarters
def compute_hire_costs(hire: HireTarget) -> dict:
"""Compute total first-year cost for one hire."""
total_comp = hire.base_salary + int(hire.base_salary * hire.bonus_pct) + hire.equity_annual_usd + hire.benefits_annual
recruiter_fee = int(hire.base_salary * hire.recruiter_fee_pct)
overhead = int(hire.base_salary * 0.25) # workspace, tools, onboarding
ramp_productivity_cost = int(hire.base_salary * (hire.ramp_months / 12)) # cost during ramp
return {
"base_salary": hire.base_salary,
"target_bonus": int(hire.base_salary * hire.bonus_pct),
"equity_annual": hire.equity_annual_usd,
"benefits": hire.benefits_annual,
"total_comp": total_comp,
"recruiter_fee": recruiter_fee,
"overhead": overhead,
"ramp_cost": ramp_productivity_cost,
"first_year_total": total_comp + recruiter_fee + overhead,
"fully_loaded_first_year": total_comp + recruiter_fee + overhead + ramp_productivity_cost,
}
def summarize_by_quarter(plan: HiringPlan) -> dict[str, dict]:
"""Aggregate headcount and costs per quarter."""
quarters = get_quarters(plan.hires)
summary = {}
running_headcount = plan.current_headcount
for q in quarters:
q_hires = [h for h in plan.hires if h.quarter == q]
q_costs = [compute_hire_costs(h) for h in q_hires]
total_comp = sum(c["total_comp"] for c in q_costs)
total_first_year = sum(c["first_year_total"] for c in q_costs)
recruiter_fees = sum(c["recruiter_fee"] for c in q_costs)
running_headcount += len(q_hires)
summary[q] = {
"new_hires": len(q_hires),
"headcount_eop": running_headcount,
"total_annual_comp_added": total_comp,
"total_first_year_cost": total_first_year,
"recruiter_fees": recruiter_fees,
"hires": q_hires,
"costs": q_costs,
}
return summary
def summarize_by_function(plan: HiringPlan) -> dict[str, dict]:
"""Aggregate headcount and costs per function."""
functions: dict[str, dict] = {}
for hire in plan.hires:
fn = hire.function
if fn not in functions:
functions[fn] = {"count": 0, "total_comp": 0, "total_first_year": 0, "roles": []}
costs = compute_hire_costs(hire)
functions[fn]["count"] += 1
functions[fn]["total_comp"] += costs["total_comp"]
functions[fn]["total_first_year"] += costs["first_year_total"]
functions[fn]["roles"].append(hire.role)
return functions
def compute_totals(plan: HiringPlan) -> dict:
all_costs = [compute_hire_costs(h) for h in plan.hires]
total_hires = len(plan.hires)
total_comp = sum(c["total_comp"] for c in all_costs)
total_first_year = sum(c["first_year_total"] for c in all_costs)
total_fully_loaded = sum(c["fully_loaded_first_year"] for c in all_costs)
total_recruiter = sum(c["recruiter_fee"] for c in all_costs)
final_headcount = plan.current_headcount + total_hires
revenue_per_employee = plan.target_revenue / final_headcount if final_headcount > 0 else 0
revenue_per_employee_current = plan.current_revenue / plan.current_headcount if plan.current_headcount > 0 else 0
return {
"total_hires": total_hires,
"final_headcount": final_headcount,
"headcount_growth_pct": ((final_headcount - plan.current_headcount) / plan.current_headcount * 100) if plan.current_headcount > 0 else 0,
"total_annual_comp_added": total_comp,
"total_first_year_cost": total_first_year,
"total_fully_loaded_first_year": total_fully_loaded,
"total_recruiter_fees": total_recruiter,
"revenue_per_employee_target": revenue_per_employee,
"revenue_per_employee_current": revenue_per_employee_current,
"avg_comp_per_hire": total_comp // total_hires if total_hires > 0 else 0,
}
# ---------------------------------------------------------------------------
# Risk assessment
# ---------------------------------------------------------------------------
def assess_risks(plan: HiringPlan, totals: dict) -> list[dict]:
risks = []
# Headcount growth too fast
growth_pct = totals["headcount_growth_pct"]
if growth_pct > 80:
risks.append({
"severity": "HIGH",
"category": "Execution",
"finding": f"Headcount growing {growth_pct:.0f}% this period. "
"Culture and processes rarely scale this fast without breakage.",
"recommendation": "Stagger Q3/Q4 hires. Validate Q1/Q2 cohort is onboarded before next wave."
})
elif growth_pct > 50:
risks.append({
"severity": "MEDIUM",
"category": "Execution",
"finding": f"Headcount growing {growth_pct:.0f}% — significant scaling challenge.",
"recommendation": "Ensure onboarding infrastructure scales. Assign buddy/mentor to each hire."
})
# High concentration in one quarter
quarters = get_quarters(plan.hires)
q_counts = {q: sum(1 for h in plan.hires if h.quarter == q) for q in quarters}
max_q = max(q_counts.values()) if q_counts else 0
if max_q > len(plan.hires) * 0.5 and max_q > 4:
heavy_q = [q for q, c in q_counts.items() if c == max_q][0]
risks.append({
"severity": "MEDIUM",
"category": "Hiring Execution",
"finding": f"More than 50% of hires planned in {heavy_q} ({max_q} hires). "
"Recruiting capacity and onboarding bandwidth may be insufficient.",
"recommendation": "Spread hires across quarters. Hiring pipeline needs to start 60–90 days before target start date."
})
# Revenue per employee declining
if totals["revenue_per_employee_target"] < totals["revenue_per_employee_current"] * 0.7:
risks.append({
"severity": "HIGH",
"category": "Financial",
"finding": f"Revenue per employee declining from ,.0f to "
f",.0f — a {((totals['revenue_per_employee_target']/totals['revenue_per_employee_current'])-1)*100:.0f}% drop.",
"recommendation": "Validate that revenue model supports this headcount. Is target revenue achievable with this team?"
})
# Low priority hires consuming budget
low_priority_hires = [h for h in plan.hires if h.priority == "Low"]
if low_priority_hires:
lp_cost = sum(compute_hire_costs(h)["first_year_total"] for h in low_priority_hires)
risks.append({
"severity": "MEDIUM",
"category": "Prioritization",
"finding": f"{len(low_priority_hires)} 'Low' priority hires consuming ,.0f in first-year costs.",
"recommendation": "Consider deferring Low priority hires to preserve runway. Cut these first if budget tightens."
})
# Hires without business cases
no_case = [h for h in plan.hires if not h.business_case]
if no_case:
risks.append({
"severity": "MEDIUM",
"category": "Governance",
"finding": f"{len(no_case)} hires have no documented business case: {', '.join(h.role for h in no_case[:5])}{'...' if len(no_case) > 5 else ''}",
"recommendation": "Every hire over $80K should have a written business case. What revenue or risk does this role address?"
})
# High recruiter fee exposure
if totals["total_recruiter_fees"] > 100_000:
risks.append({
"severity": "LOW",
"category": "Cost",
"finding": f",.0f in recruiter fees. "
"Consider whether internal recruiter investment would be cheaper at this hiring volume.",
"recommendation": f"Internal recruiter at $120–150K fully loaded pays off at 3–4 hires/year vs. agency fees."
})
# No risks — that's itself a flag
if not risks:
risks.append({
"severity": "INFO",
"category": "General",
"finding": "No major risks flagged. Plan appears well-structured.",
"recommendation": "Validate assumptions: time-to-fill estimates, revenue model, and Q1 hiring pipeline status."
})
return risks
# ---------------------------------------------------------------------------
# Formatting / Output
# ---------------------------------------------------------------------------
def fmt(n: int) -> str:
return f",.0f"
def pct(n: float) -> str:
return f"{n:.1f}%"
def print_report(plan: HiringPlan):
WIDTH = 72
SEP = "=" * WIDTH
sep = "-" * WIDTH
print(SEP)
print(f" HIRING PLAN: {plan.company}")
print(f" Period: {plan.plan_period} | Generated: {date.today().isoformat()}")
print(SEP)
totals = compute_totals(plan)
q_summary = summarize_by_quarter(plan)
fn_summary = summarize_by_function(plan)
risks = assess_risks(plan, totals)
# Executive summary
print("\n[ EXECUTIVE SUMMARY ]")
print(sep)
print(f" Current headcount: {plan.current_headcount:>5}")
print(f" Planned hires: {totals['total_hires']:>5}")
print(f" Final headcount: {totals['final_headcount']:>5} (+{totals['headcount_growth_pct']:.0f}%)")
print(f" Current ARR: {fmt(plan.current_revenue):>12}")
print(f" Target revenue: {fmt(plan.target_revenue):>12}")
print(f" Revenue/employee now: {fmt(int(totals['revenue_per_employee_current'])):>12}")
print(f" Revenue/employee target: {fmt(int(totals['revenue_per_employee_target'])):>12}")
print()
print(f" Total annual comp added: {fmt(totals['total_annual_comp_added']):>12}")
print(f" Total first-year cost: {fmt(totals['total_first_year_cost']):>12}")
print(f" Fully loaded (w/ ramp): {fmt(totals['total_fully_loaded_first_year']):>12}")
print(f" Recruiter fees: {fmt(totals['total_recruiter_fees']):>12}")
print(f" Avg comp per hire: {fmt(totals['avg_comp_per_hire']):>12}")
# Quarterly breakdown
print(f"\n[ QUARTERLY HEADCOUNT PLAN ]")
print(sep)
print(f" {'Quarter':<10} {'New Hires':>10} {'HC (EOP)':>10} {'Comp Added':>14} {'1yr Cost':>14} {'Recruiter $':>12}")
print(f" {'-'*10} {'-'*10} {'-'*10} {'-'*14} {'-'*14} {'-'*12}")
for q, data in q_summary.items():
print(f" {q:<10} {data['new_hires']:>10} {data['headcount_eop']:>10} "
f"{fmt(data['total_annual_comp_added']):>14} "
f"{fmt(data['total_first_year_cost']):>14} "
f"{fmt(data['recruiter_fees']):>12}")
# By function
print(f"\n[ HEADCOUNT BY FUNCTION ]")
print(sep)
print(f" {'Function':<18} {'Hires':>7} {'Annual Comp':>14} {'1yr Cost':>14}")
print(f" {'-'*18} {'-'*7} {'-'*14} {'-'*14}")
for fn, data in sorted(fn_summary.items(), key=lambda x: -x[1]["count"]):
print(f" {fn:<18} {data['count']:>7} {fmt(data['total_comp']):>14} {fmt(data['total_first_year']):>14}")
# Hire detail
print(f"\n[ HIRE DETAIL ]")
print(sep)
print(f" {'Role':<30} {'Fn':<14} {'Lvl':<6} {'Q':<8} {'Base':>10} {'Total Comp':>12} {'Priority':<8}")
print(f" {'-'*30} {'-'*14} {'-'*6} {'-'*8} {'-'*10} {'-'*12} {'-'*8}")
for h in sorted(plan.hires, key=lambda x: quarter_to_sortkey(x.quarter)):
costs = compute_hire_costs(h)
print(f" {h.role:<30} {h.function:<14} {h.level:<6} {h.quarter:<8} "
f"{fmt(h.base_salary):>10} {fmt(costs['total_comp']):>12} {h.priority:<8}")
if h.business_case:
bc = h.business_case[:60] + "..." if len(h.business_case) > 60 else h.business_case
print(f" {'':>30} ↳ {bc}")
# Risk assessment
print(f"\n[ RISK ASSESSMENT ]")
print(sep)
sev_order = {"HIGH": 0, "MEDIUM": 1, "LOW": 2, "INFO": 3}
for risk in sorted(risks, key=lambda r: sev_order.get(r["severity"], 99)):
sev = risk["severity"]
marker = {"HIGH": "⚠ HIGH", "MEDIUM": "◆ MED ", "LOW": "◇ LOW ", "INFO": "ℹ INFO"}[sev]
print(f"\n [{marker}] {risk['category']}")
# Wrap finding
finding = risk["finding"]
words = finding.split()
line = " Finding: "
for w in words:
if len(line) + len(w) + 1 > WIDTH - 2:
print(line)
line = " " + w + " "
else:
line += w + " "
if line.strip():
print(line)
reco = risk["recommendation"]
words = reco.split()
line = " Action: "
for w in words:
if len(line) + len(w) + 1 > WIDTH - 2:
print(line)
line = " " + w + " "
else:
line += w + " "
if line.strip():
print(line)
print(f"\n{SEP}\n")
def export_csv(plan: HiringPlan) -> str:
"""Return CSV of hire detail."""
output = io.StringIO()
writer = csv.writer(output)
writer.writerow(["Role", "Function", "Level", "Quarter", "Priority",
"Base Salary", "Bonus Target", "Equity Annual", "Benefits",
"Total Comp", "Recruiter Fee", "Overhead", "First Year Total",
"Ramp Months", "Open to Internal", "Business Case"])
for h in plan.hires:
c = compute_hire_costs(h)
writer.writerow([h.role, h.function, h.level, h.quarter, h.priority,
h.base_salary, c["target_bonus"], h.equity_annual_usd, h.benefits_annual,
c["total_comp"], c["recruiter_fee"], c["overhead"], c["first_year_total"],
h.ramp_months, h.open_to_internal, h.business_case])
return output.getvalue()
# ---------------------------------------------------------------------------
# Sample data
# ---------------------------------------------------------------------------
def build_sample_plan() -> HiringPlan:
"""Sample Series A → B hiring plan."""
plan = HiringPlan(
company="AcmeTech (Series A)",
plan_period="2025 Annual",
current_headcount=32,
current_revenue=3_500_000,
target_revenue=8_000_000,
overhead_rate=0.25,
internal_recruiter_cost=140_000,
)
plan.hires = [
# Q1 — Foundation hires
HireTarget(
role="Staff Software Engineer (Backend)",
level="L4", function="Engineering", quarter="Q1-2025",
base_salary=185_000, bonus_pct=0.0, equity_annual_usd=25_000,
benefits_annual=18_000, recruiter_fee_pct=0.0, ramp_months=2,
priority="High", open_to_internal=True,
business_case="Core API team is bottleneck for 3 roadmap items. Staff-level needed to lead architecture."
),
HireTarget(
role="Account Executive (Mid-Market)",
level="L3", function="Sales", quarter="Q1-2025",
base_salary=95_000, bonus_pct=0.50, equity_annual_usd=10_000,
benefits_annual=15_000, recruiter_fee_pct=0.18, ramp_months=4,
priority="High",
business_case="Pipeline coverage at 1.8x quota. Need 2.5x by Q2. AE adds $600K ARR/year at ramp."
),
HireTarget(
role="Product Designer (Senior)",
level="L3", function="Product", quarter="Q1-2025",
base_salary=145_000, bonus_pct=0.0, equity_annual_usd=18_000,
benefits_annual=18_000, recruiter_fee_pct=0.0, ramp_months=2,
priority="High",
business_case="Single designer for 4 squads. UX debt slowing enterprise deals requiring onboarding improvements."
),
# Q2 — Growth hires
HireTarget(
role="Engineering Manager (Frontend)",
level="M1", function="Engineering", quarter="Q2-2025",
base_salary=175_000, bonus_pct=0.10, equity_annual_usd=22_000,
benefits_annual=18_000, recruiter_fee_pct=0.20, ramp_months=3,
priority="High",
business_case="Frontend team at 7 ICs with no dedicated EM. Performance review debt is high; manager needed."
),
HireTarget(
role="Account Executive (Mid-Market)",
level="L2", function="Sales", quarter="Q2-2025",
base_salary=85_000, bonus_pct=0.50, equity_annual_usd=8_000,
benefits_annual=15_000, recruiter_fee_pct=0.18, ramp_months=4,
priority="High",
business_case="Second AE to reach 2.5x pipeline coverage target."
),
HireTarget(
role="Customer Success Manager",
level="L2", function="Customer Success", quarter="Q2-2025",
base_salary=90_000, bonus_pct=0.15, equity_annual_usd=8_000,
benefits_annual=15_000, recruiter_fee_pct=0.0, ramp_months=2,
priority="Medium",
business_case="CSM:account ratio at 1:60, industry standard 1:30. NRR has dipped 4pts in 2 quarters."
),
HireTarget(
role="Data Engineer",
level="L2", function="Engineering", quarter="Q2-2025",
base_salary=155_000, bonus_pct=0.0, equity_annual_usd=18_000,
benefits_annual=18_000, recruiter_fee_pct=0.0, ramp_months=3,
priority="Medium",
business_case="Analytics infrastructure blocking product analytics, customer dashboards, and board metrics."
),
# Q3 — Scale hires
HireTarget(
role="Senior Software Engineer (Backend)",
level="L3", function="Engineering", quarter="Q3-2025",
base_salary=165_000, bonus_pct=0.0, equity_annual_usd=20_000,
benefits_annual=18_000, recruiter_fee_pct=0.0, ramp_months=2,
priority="High",
business_case="Backend team needs capacity to deliver Q3 roadmap without delaying Q4 items."
),
HireTarget(
role="Head of Marketing",
level="M3", function="Marketing", quarter="Q3-2025",
base_salary=180_000, bonus_pct=0.15, equity_annual_usd=30_000,
benefits_annual=18_000, recruiter_fee_pct=0.20, ramp_months=3,
priority="High",
business_case="No marketing function. 100% of pipeline is outbound. Need inbound by Q1-2026 for Series B."
),
HireTarget(
role="People Operations Manager",
level="M1", function="G&A", quarter="Q3-2025",
base_salary=120_000, bonus_pct=0.10, equity_annual_usd=12_000,
benefits_annual=16_000, recruiter_fee_pct=0.0, ramp_months=2,
priority="Medium",
business_case="Founders spending 8hrs/week on HR ops at 40 employees. Unscalable. First dedicated HR hire."
),
# Q4 — Stretch hires (conditional on revenue milestone)
HireTarget(
role="Senior Software Engineer (Frontend)",
level="L3", function="Engineering", quarter="Q4-2025",
base_salary=160_000, bonus_pct=0.0, equity_annual_usd=18_000,
benefits_annual=18_000, recruiter_fee_pct=0.0, ramp_months=2,
priority="Medium",
business_case="Conditional on Q3 ARR exceeding $5.5M. Frontend team capacity planning for 2026 roadmap."
),
HireTarget(
role="Account Executive (Enterprise)",
level="L4", function="Sales", quarter="Q4-2025",
base_salary=120_000, bonus_pct=0.60, equity_annual_usd=15_000,
benefits_annual=15_000, recruiter_fee_pct=0.20, ramp_months=6,
priority="Low",
business_case="Enterprise motion exploratory. Requires ICP validation in Q2-Q3 before committing."
),
HireTarget(
role="DevOps / Platform Engineer",
level="L3", function="Engineering", quarter="Q4-2025",
base_salary=150_000, bonus_pct=0.0, equity_annual_usd=18_000,
benefits_annual=18_000, recruiter_fee_pct=0.0, ramp_months=3,
priority="Low",
business_case="Platform reliability becoming bottleneck. Conditional on uptime SLA breaches continuing in Q3."
),
]
return plan
# ---------------------------------------------------------------------------
# CLI
# ---------------------------------------------------------------------------
def load_plan_from_json(path: str) -> HiringPlan:
with open(path) as f:
data = json.load(f)
hires = [HireTarget(**h) for h in data.pop("hires", [])]
plan = HiringPlan(**data)
plan.hires = hires
return plan
def main():
parser = argparse.ArgumentParser(
description="Hiring Plan Modeler — build headcount plans with cost projections",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
python hiring_plan_modeler.py # Run sample plan
python hiring_plan_modeler.py --config plan.json # Load from JSON
python hiring_plan_modeler.py --export-csv # Output CSV of hires
python hiring_plan_modeler.py --export-json # Output plan as JSON template
"""
)
parser.add_argument("--config", help="Path to JSON plan file")
parser.add_argument("--export-csv", action="store_true", help="Export hire detail as CSV")
parser.add_argument("--export-json", action="store_true", help="Export sample plan as JSON template")
args = parser.parse_args()
if args.config:
plan = load_plan_from_json(args.config)
else:
plan = build_sample_plan()
if args.export_json:
data = asdict(plan)
print(json.dumps(data, indent=2))
return
if args.export_csv:
print(export_csv(plan))
return
print_report(plan)
if __name__ == "__main__":
main()
Tư vấn pháp lý cho startup: rà soát hợp đồng (MSA, SaaS, NDA, DPA), chiến lược sở hữu trí tuệ, term sheet và bản đồ quy định.
---
name: "general-counsel-advisor"
description: "General Counsel advisory for startups: contract review (MSA, SaaS, NDA, DPA, employment), IP strategy, term sheet decoding, and regulatory landscape mapping. Use when reviewing any contract or term sheet, deciding when to engage outside counsel, defining IP strategy, evaluating regulatory exposure (HIPAA, GDPR, FDA, fintech), or when user mentions general counsel, GC, legal review, contract risk, term sheet, IP assignment, or regulatory exposure. NOT a substitute for licensed counsel — surfaces questions to bring to qualified attorneys."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: c-level
domain: general-counsel-leadership
updated: 2026-05-12
python-tools: contract_risk_scanner.py, term_sheet_analyzer.py
frameworks: contract-review, ip-strategy, term-sheet-decoding, regulatory-mapping
---
# General Counsel Advisor
Strategic legal frameworks for startup General Counsels and founders without one. Contract risk, IP strategy, term sheet decoding, regulatory landscape.
This is **not legal advice**. It surfaces the right questions to bring to qualified outside counsel and catches the obvious traps before they reach a signature. Treat every output as a starting point for a conversation with a licensed attorney, not as a substitute for one.
## Keywords
general counsel, GC, legal review, contract review, MSA, SaaS agreement, NDA, DPA, employment agreement, contractor agreement, IP assignment, invention assignment, open source license, OSS compliance, term sheet, liquidation preference, anti-dilution, option pool, vesting, acceleration, drag-along, pro-rata, board composition, regulatory, HIPAA, GDPR, CCPA, FDA, MDR, fintech, BSA/AML, money transmitter, AI Act, indemnity, liability cap, force majeure, auto-renewal, choice of law, venue, non-compete, non-solicit
## Quick Start
```bash
# Scan a contract for risky clauses (uses bundled sample if no path given)
python scripts/contract_risk_scanner.py
python scripts/contract_risk_scanner.py path/to/contract.txt
# Analyze a term sheet for founder-friendliness
python scripts/term_sheet_analyzer.py
python scripts/term_sheet_analyzer.py path/to/term_sheet.json
```
## Key Questions (ask these first)
- **Who owns the IP being created or shared?** (Founders forget that contractors don't auto-assign IP without a written clause.)
- **What's the liability cap, and what's carved out?** (Standard: 12 months of fees, with carve-outs for IP infringement, data breach, willful misconduct.)
- **Is there a DPA in place if any personal data flows?** (GDPR, CCPA, state laws — non-negotiable if EU/CA data is touched.)
- **What's the termination right, notice period, and auto-renewal trap?** (5-year auto-renew with 60-day notice is a common founder mistake.)
- **Does this contract or product launch trigger a new regulatory regime?** (Healthcare → HIPAA. Fintech → BSA/AML. Medical device → FDA/MDR.)
- **For term sheets: liquidation preference, pre-money option pool, anti-dilution flavor?** (Three places where 5% of founder economics can quietly disappear.)
## Core Responsibilities
### 1. Contract Review
Standard contracts a startup signs in its first 5 years:
- **Vendor MSA** — Master Service Agreement (cloud, tooling, services)
- **Customer SaaS Agreement** — your standard customer paper + customer redlines
- **NDA** — mutual + one-way, with carve-outs for residuals + independent development
- **DPA** — Data Processing Agreement (required when personal data flows)
- **Employment Agreement** — offer letter, IP assignment, non-compete (where enforceable), arbitration
- **Contractor / 1099 Agreement** — IP assignment is critical; misclassification risk
- **Equity Agreements** — option grants, RSU agreements, advisor grants (FAST template, YC SAFE for advisors)
**Run** `contract_risk_scanner.py` on the text. It flags the 12 most common founder-killer clauses.
### 2. IP Strategy
- **Invention assignment** — every employee and contractor signs one. No exceptions.
- **Open source license compliance** — track every OSS dependency's license; AGPL and GPL trigger copyleft obligations.
- **Trade secrets** — define what's protected and how (clean room dev, access controls, NDAs).
- **Patents** — file provisional within 12 months of disclosure; PCT for international.
- **Trademarks** — register the word mark first, design mark second; clear before launch.
- **Copyright** — automatic on creation, but register for statutory damages eligibility.
See `references/ip_and_regulatory.md`.
### 3. Term Sheet Decoding
When a term sheet arrives, the difference between a founder-friendly and founder-hostile sheet often hides in three clauses:
- **Liquidation preference** — 1x non-participating is standard; 1x participating or 2x is hostile
- **Pre-money vs post-money option pool** — pre-money pool dilutes founders; post-money dilutes everyone proportionally
- **Anti-dilution** — broad-based weighted average is standard; full ratchet is hostile
**Run** `term_sheet_analyzer.py` to get a 0-100 founder-friendliness score with flags.
### 4. Regulatory Landscape
When to engage outside counsel **before** committing:
| Trigger | Regime | First Step |
|---|---|---|
| Healthcare data | HIPAA, HITECH, state breach laws | Specialist health-tech counsel |
| Cardholder data | PCI DSS (industry standard, not law, but contractually required) | QSA + counsel |
| Money movement | BSA/AML, state money-transmitter (50-state patchwork) | Fintech specialist |
| Medical device claims | FDA 510(k) / De Novo / PMA, MDR (EU), ISO 13485 | Medical-device specialist |
| EU residents' personal data | GDPR + EU AI Act if AI is deployed | EU privacy counsel |
| California residents | CCPA / CPRA | Privacy generalist |
| Securities (tokens, equity crowdfunding) | SEC rules (Reg D, Reg A+, Reg CF) | Securities counsel |
| Defense / aerospace customers | ITAR, EAR, DFARS, CMMC | Export-control counsel |
| AI in EU | EU AI Act (risk-tiered) | EU privacy + product counsel |
| AI for hiring (NYC, CO, IL) | Local bias-audit laws | Employment counsel |
See `references/ip_and_regulatory.md` for sequencing.
## Workflows
### Workflow 1: Contract Review
1. Save the contract as plain text
2. Run `contract_risk_scanner.py path/to/contract.txt`
3. For each HIGH risk finding, draft a counter-proposal
4. Bring the redline + counter-proposals to outside counsel
5. Log the decision via `/cs:decide`
### Workflow 2: Term Sheet Response
1. Save the term sheet as a JSON file matching the schema in `term_sheet_analyzer.py --help`
2. Run `python scripts/term_sheet_analyzer.py path/to/term_sheet.json`
3. Review the founder-friendliness score and per-clause flags
4. Negotiate the worst 3 clauses (don't try to win all 20)
5. Always have a securities/venture attorney review before signing
6. Log via `/cs:decide` with `/cs:freeze 30` to prevent regret-driven re-opening
### Workflow 3: IP Hygiene Audit
1. Confirm every employee and contractor (past 12 months) signed invention assignment
2. Run an OSS license inventory (`pip-licenses`, `license-checker` for npm)
3. Map AGPL/GPL dependencies and confirm compliance (or remove)
4. File provisional patents on novel inventions (12-month deadline from disclosure)
5. Register word-mark trademarks for the product name
### Workflow 4: Regulatory Trigger Assessment
1. List planned product features for the next 12 months
2. Map each feature to the trigger table in this document
3. For any HIPAA / FDA / fintech trigger, engage a specialist counsel **before** building
4. Document the regulatory roadmap and budget alongside the product roadmap
5. Pair with `cs-ciso-advisor` for ISO 27001 / SOC 2 sequencing
## Output Standard (when invoked via `/cs:gc-review`)
```
**Bottom Line:** [sign / negotiate / do not sign]
**The Risks:** [3 highest-severity issues]
**Counter-Proposals:** [specific language]
**Outside Counsel Action Items:** [what to bring to the attorney]
**Your Decision:** [the call only the founder can make]
```
## Adjacent Skills
- `../ciso-advisor/` — Compliance overlap (SOC 2, ISO 27001, HIPAA technical safeguards)
- `../cfo-advisor/` — Term sheet → dilution math
- `../ma-playbook/` — Acquisition agreements, integration playbooks
- `../../../ra-qm-team/` — ISO 13485, MDR, FDA 510(k), GDPR execution
- `../../c-level-agents/skills/gc-review/SKILL.md` — `/cs:gc-review` slash command
## References
- [contracts_playbook.md](references/contracts_playbook.md) — Standard contracts, clause checklist, common founder traps
- [ip_and_regulatory.md](references/ip_and_regulatory.md) — IP protection + regulatory landscape mapping
- [term_sheet_decoder.md](references/term_sheet_decoder.md) — Term sheet glossary + founder-friendly defaults + pushback strategies
---
**Version:** 1.0.0
**Status:** Production Ready
**Disclaimer:** Not legal advice. Always engage qualified counsel for binding decisions.
FILE:references/contracts_playbook.md
# Contracts Playbook — Standard Startup Agreements
Reference for the 7 contracts every startup signs in its first 5 years and the clause traps to avoid in each. **Not legal advice.** Bring redlines to qualified counsel.
## 1. Master Service Agreement (MSA) — Vendor Side (you signing theirs)
**What it is:** The umbrella contract for an ongoing relationship with a vendor (cloud, tooling, services, agencies). Usually paired with one or more SOWs / Order Forms.
**Top 5 redlines to push:**
1. **Auto-renewal:** Cut notice period to 30 days max. Reject 60/90/180 day notice.
2. **Liability cap:** Insist on 12 months of fees. Reject "fees in the preceding 3 months" (too narrow).
3. **Mutual indemnification:** Reject one-sided. Mirror the scope on both sides.
4. **IP ownership of deliverables:** All work product belongs to you. Vendor retains rights to pre-existing tools / methodologies, granted back to you for use.
5. **Data: DPA + return-or-destroy on termination.** Specifically: vendor cannot use your data to train AI models.
**Bonus catch:** Watch for "Vendor may modify these terms upon notice" — this means the contract you signed isn't the contract you have.
## 2. Customer SaaS Agreement (your paper)
**Standard structure:**
1. License grant (subscription, scope, term)
2. Acceptable use policy (what customer can/can't do)
3. Fees & payment (annual prepay vs. monthly, late fee, currency)
4. Service Level Agreement (uptime %, credits, exclusions)
5. Confidentiality (mutual, residuals carve-out)
6. Data Protection (DPA exhibit, subprocessor list, security commitments)
7. Warranties (limited, disclaim implied)
8. Indemnification (mutual, IP-infringement focused)
9. Limitation of liability (12 months fees, carve-outs for IP/data breach/willful)
10. Term & termination (term, termination for cause, termination for convenience)
**Founder traps when accepting customer redlines:**
- "Most-favored-nation" pricing (means you can never give anyone else a better deal).
- Uncapped liability for data breach with no minimum threshold.
- Customer right to perpetual license-back of "improvements" to your product.
- Customer "ownership" of any custom configuration (often hiding IP creep).
- Source-code escrow with auto-release triggers tied to customer convenience.
## 3. Non-Disclosure Agreement (NDA)
**One-way (you receiving):** Acceptable to sign without redlines for short evaluations.
**Mutual NDA (both directions):** The default for ongoing discussions.
**Critical carve-outs (always include):**
- **Residuals:** Information retained in unaided memory after end of engagement is not confidential.
- **Independent development:** If you build something similar without using their info, it's yours.
- **Public domain:** Information already public is not confidential.
- **Rightfully received:** Information received from a third party without confidentiality obligation.
- **Required by law:** Information disclosed under subpoena (with notice).
**Founder trap:** NDAs that prevent you from "engaging in similar business" — that's a non-compete in disguise. Strip it out.
## 4. Data Processing Agreement (DPA)
**Required when:** Personal data of EU residents flows (GDPR Article 28), or California residents (CCPA / CPRA), or HIPAA-covered data, or biometrics in IL/TX/WA (BIPA).
**Standard structure (GDPR-aligned):**
- Scope of processing (what data, what purpose)
- Controller / Processor designation
- Subprocessor list + flow-down obligations
- Data subject rights (access, deletion, portability)
- Security measures (encryption, access controls, training)
- Breach notification timelines (within 72 hours for GDPR)
- Audit rights (annual, reasonable)
- International transfer mechanism (SCCs, adequacy decision, BCRs)
- Return-or-destroy on termination
**Templates:** Use IAPP, EU Commission SCCs, or vendor-friendly DPA (e.g., Vanta's, Stripe's).
**Founder trap:** Missing DPA when EU/CA data flows = contract may be unenforceable AND regulatory fine exposure.
## 5. Employment Agreement / Offer Letter
**Must-have provisions:**
- **At-will employment** (US most states; not enforceable in MT for example)
- **Compensation:** salary, bonus structure, equity (option grant separately documented)
- **Invention assignment:** all IP created during employment using company resources belongs to company
- **Confidentiality:** ongoing duty, surviving termination
- **Non-solicit:** 12 months post-termination, employees + customers (carve out general advertising)
- **Non-compete:** state-dependent (CA, ND, OK, DC: void; many other states: enforceable if reasonable)
- **Arbitration:** mutual, AAA or JAMS rules, employer pays fees
**Founder traps:**
- Forgetting to require employees to sign **before** starting work (otherwise IP assignment is weak).
- Not including a "previously created inventions" exhibit (lets founders document pre-existing IP brought into the company).
- Skipping background checks for senior hires.
## 6. Contractor / 1099 Agreement
**Critical differences from employment:**
- **IP assignment is NOT automatic.** Without a written clause, the contractor owns what they create (under US law, "work for hire" applies only to specific categories of work).
- **Misclassification risk:** If a contractor functions like an employee (controlled hours, exclusive engagement, supplied equipment), tax authorities can reclassify, triggering back taxes + penalties.
- **No benefits, no withholding, contractor handles their own taxes.**
**Must-have provisions:**
- **Explicit work-for-hire OR written IP assignment** ("Contractor hereby assigns all right, title, and interest...").
- **Independent contractor status:** contractor controls means and methods.
- **Termination:** 30-day notice, immediate for cause.
- **Indemnification:** contractor indemnifies you for misclassification claims if they misrepresent status.
**Tooling:** Use Deel, Remote, or Velocity Global for international contractors to handle classification correctly.
## 7. Equity Agreements (Option Grants, Advisor Grants)
**Employee option grant:**
- **Strike price:** must be ≥ fair market value (FMV) at grant date (409A valuation, refreshed annually).
- **Vesting:** standard 4 years, 1 year cliff, monthly thereafter.
- **Exercise window post-termination:** 90 days standard; 7-10 years is founder-friendly.
- **ISO vs NSO:** ISOs have tax advantages (long-term capital gains if held) but limits ($100K vest/year) and US-citizen-only.
**Advisor grant (FAST template by Founder Institute):**
- 0.1% - 1% equity vested over 1-2 years, depending on level and stage.
- 2-year vesting, no cliff (advisors are tested through engagement, not retention).
- Single trigger acceleration on change of control (rare; double trigger more common).
**Founder trap:**
- Issuing options before completing the 409A valuation — strike price might be challenged by IRS.
- Verbal promises about acceleration — must be in writing.
- Forgetting to issue option grants to early employees within 90 days of hire (loses ISO eligibility).
## Quick Triage Heuristics
When you have 5 minutes to look at a contract:
1. **Find the liability cap.** No cap or > 24 months of fees = red flag.
2. **Find the indemnity clauses.** One-sided = red flag.
3. **Find the IP clause.** Vague or "as agreed" = red flag.
4. **Find the term + termination.** Auto-renewal with > 30 day notice = red flag.
5. **Find the choice of law/venue.** Exclusive in counterparty home jurisdiction = red flag.
Run `scripts/contract_risk_scanner.py` for the automated version.
---
**Final reminder:** This is a triage playbook. Every contract over $100K or longer than 1 year deserves outside counsel review. Every contract that touches personal data deserves a privacy attorney. Every term sheet deserves a securities / venture attorney. Period.
FILE:references/ip_and_regulatory.md
# IP Strategy & Regulatory Landscape
The two areas where startups most often discover legal exposure after it's too late to fix cheaply: IP ownership and regulatory triggers. **Not legal advice.**
## Part 1: IP Strategy
### IP Inventory — The Four Categories
| Type | What it protects | How you get it | How you lose it |
|---|---|---|---|
| **Patents** | Inventions (novel, non-obvious, useful) | File application | Public disclosure > 12 months before filing |
| **Copyright** | Original works of authorship (code, content, designs) | Automatic on fixation | Almost never; can be assigned away |
| **Trademark** | Brand identifiers (names, logos, slogans) | Use in commerce + registration | Not policing infringement; becoming generic |
| **Trade secret** | Confidential business information | Reasonable measures to keep secret | Public disclosure; failure to maintain confidentiality |
### Invention Assignment — The Single Most Important IP Practice
**Rule:** Every person who touches the company's product or systems must sign an invention assignment agreement **before** they start work.
This includes:
- Co-founders (often forgotten — usually fixed via founder restricted-stock purchase agreements)
- Employees (in employment agreement)
- Contractors (in contractor agreement; NOT automatic in US law)
- Interns (often forgotten — use a short standalone IP agreement)
- Advisors (in advisor agreement, scope limited to inventions related to company)
**Why it matters:** Without written assignment, the creator retains ownership. A contractor who built a critical service for 6 months and never signed an assignment can come back years later and demand a license — or assert that competitors can also use what they built.
**The "previously created inventions" exhibit:** Every IP assignment should include an exhibit where the signer lists pre-existing inventions they want to exclude. This protects everyone — the signer's prior work isn't accidentally assigned, and the company has documentation of what came in.
### Open Source License Compliance
**Permissive licenses** (MIT, Apache 2.0, BSD 2/3): Use freely, attribute, no copyleft.
**Weak copyleft** (LGPL, MPL): Can use in proprietary product; modifications to the OSS itself must be released. Distribution model matters.
**Strong copyleft** (GPL v2, GPL v3, AGPL): Distribution / SaaS use of a strong-copyleft component can require releasing your derivative work under the same license. **AGPL is the most aggressive** — it applies even when you only run the software on a server (SaaS / network use).
**Practice:**
1. Maintain an OSS inventory: `pip-licenses`, `license-checker` (npm), `cargo-license`, `go-licenses`.
2. Identify any GPL / AGPL / SSPL dependencies.
3. For each: either (a) comply with the license, (b) replace with a permissively-licensed alternative, or (c) document the carve-out (some companies build internally with GPL but only ship the binary externally — verify with counsel).
4. Run the inventory before any due diligence (acquisition, financing).
### Patents — When to File
**File when:**
- You have a genuinely novel technical invention (algorithm, hardware design, materials, biotech process).
- You face well-funded competitors who could copy without consequence.
- You're in a patent-dense industry (semiconductors, pharma, networking, medical devices).
- Filing strengthens fundraising / acquisition optics (limited weight for software-only startups).
**Don't bother when:**
- Your "invention" is a UX flow or business method (these are extremely hard to patent post-Alice Corp).
- You're in early stage with limited capital and no competitors close enough to copy.
- Defensive only and joining a patent pool (LOT Network, OIN) might be cheaper.
**Process:**
1. **Provisional patent** ($300-500 USPTO fee + $3K-5K attorney). 12 months to file non-provisional.
2. **Non-provisional / utility patent** ($1K USPTO fee + $10K-15K attorney + prosecution costs).
3. **PCT application** for international filings ($5K-10K).
4. **National phase entries** in each country you care about ($5K-15K per country).
Budget $25K-50K total for one well-prosecuted patent family with international coverage.
### Trade Secrets
**Reasonable measures required for legal protection:**
- NDA / confidentiality clauses with everyone who has access.
- Access controls (need-to-know basis, not company-wide).
- Marking documents "Confidential."
- Departure procedures (return of materials, exit interview, deactivation).
- Training employees on what's a trade secret.
**Without these measures, the information may not qualify for trade secret protection if disclosed — even by a thief.**
**Common trade secrets:**
- Customer lists with usage / pricing data
- Algorithms not disclosed in published patents
- Manufacturing processes
- Sales playbooks and pricing models
- Internal financial projections
- Source code (unless OSS)
### Trademark Strategy
**Search before launch:**
- USPTO TESS search (free, but limited; doesn't catch common-law marks).
- Professional search via attorney ($500-2K) catches common-law marks and similar-mark conflicts.
- International searches via WIPO Global Brand Database.
**Register early:**
- US: Intent-to-use application (1B) lets you reserve a mark before launch.
- International: Madrid Protocol filing extends to 100+ countries.
- Word marks first (the brand name itself), design marks second (logos).
**Policing:**
- Set up Google Alerts and USPTO TMNG for your mark.
- Send cease-and-desist letters promptly; failure to police can weaken the mark.
---
## Part 2: Regulatory Landscape — When to Engage Counsel
The startups that survive their first regulatory encounter engage specialist counsel **before** building, not after. The ones that don't usually pivot, retreat, or pay heavy fines.
### Trigger Matrix
| Trigger | Regulatory Regime | Specialist Needed | Earliest Action |
|---|---|---|---|
| Healthcare data (patient records, claims, PHI) | HIPAA, HITECH, state breach laws | Health-tech attorney | Business Associate Agreement, OCR-aligned risk assessment |
| Cardholder data | PCI DSS (industry standard; contractually required) | QSA + counsel | Scope reduction, tokenization, certified processor |
| Money movement (transmitting funds, custody, crypto) | BSA/AML, state money-transmitter (50-state patchwork) | Fintech attorney | Stripe Treasury / Banking as a Service to avoid MT registration |
| Lending | Truth in Lending Act, state usury laws, ECOA | Fintech / consumer-finance attorney | Bank partnership, state licensing analysis |
| Medical device claims | FDA 510(k), De Novo, PMA; EU MDR; ISO 13485 | Medical-device regulatory specialist | Pre-submission meeting with FDA |
| EU residents' personal data | GDPR + ePrivacy + EU AI Act if AI | EU privacy attorney | DPA, SCCs for international transfer, DPIA |
| California residents | CCPA / CPRA | Privacy generalist | Privacy notice, opt-out mechanisms, vendor management |
| Children's data (under 13 US, under 16 in some EU states) | COPPA, GDPR-K | Privacy attorney | Parental consent, no-track defaults |
| Securities (tokens, equity crowdfunding, advisory boards) | SEC rules (Reg D, Reg A+, Reg CF, Howey test) | Securities attorney | Token sale legal opinion, Form D filing |
| Defense / aerospace customers | ITAR, EAR, DFARS, CMMC | Export-control attorney | Export classification, registered with State Dept |
| AI in EU | EU AI Act (risk-tiered: prohibited / high-risk / limited / minimal) | EU privacy + product attorney | Risk assessment, conformity assessment for high-risk |
| AI for hiring | NYC Local Law 144, CO SB 21-169, IL HB 53 | Employment attorney | Bias audit, candidate notice |
| Telehealth / online prescribing | State medical board rules, DEA registration for controlled substances | Telehealth specialist | State-by-state physician licensing strategy |
| Insurance (sale, underwriting, brokerage) | State insurance commissioners | Insurance attorney | State licensing, agency agreement |
### Sequencing: SOC 2 → ISO 27001 → Industry-Specific
For most B2B SaaS, the security/compliance sequence is:
1. **SOC 2 Type 1** (point-in-time audit) — ~$15K-25K, 3-6 months prep
2. **SOC 2 Type 2** (continuous, ~6-12 month audit window) — ~$25K-50K
3. **ISO 27001** if expanding internationally — ~$30K-60K, builds on SOC 2 controls
4. **ISO 42001** if AI is core to product — first AI management system standard
5. **Industry overlays:** HIPAA technical safeguards, FedRAMP (federal customers), PCI DSS (cardholder data)
**Sequencing logic:** SOC 2 unlocks the majority of enterprise sales. ISO 27001 unlocks European and Asia-Pacific. Industry overlays are required for specific verticals.
### When to Get a General Counsel Hire
| Stage | GC need |
|---|---|
| Pre-seed / seed | None. Use outside counsel ad-hoc + Clerky/Stripe Atlas templates |
| Series A | Fractional GC (~$10-20K/month) OR senior associate at firm |
| Series B | Full-time GC if regulated industry, customer contracts are heavy, or fundraising is constant |
| Series C+ | Full-time GC + Deputy/Associate GC if international |
**Signs you need a GC hire:**
- You're spending > $200K/year on outside counsel
- You're signing > 1 enterprise contract per week with customer redlines
- You're in a regulated industry (healthcare, fintech, defense)
- You're preparing for IPO or going-public transaction
- You're acquiring companies
### Cross-Border Considerations
**Hiring international employees:**
- Use Deel / Remote / Velocity Global for first 1-5 contractors per country.
- Establish an entity (subsidiary or EOR-to-entity transition) at 5-10+ employees.
- Tax residency, permanent establishment risk, and equity grants vary significantly.
**International data flows:**
- EU → US: SCCs + Transfer Impact Assessment (TIA); DPF if certified.
- China → outbound: PIPL approval + standard contract + security assessment.
- UK → outside: UK SCCs (similar to EU).
- Schrems / DPF status changes regularly — monitor with privacy counsel.
**International IP:**
- Patent: PCT application within 12 months of first national filing.
- Trademark: Madrid Protocol for multi-country filings.
- Copyright: Berne Convention covers most countries automatically.
---
## Closing: The General Counsel's Three Rules
1. **Get it in writing.** Verbal agreements and "we'll figure it out later" produce 80% of post-engagement disputes.
2. **Identify the regulatory trigger before you build.** It's 10x cheaper to design around a regulation than to retrofit.
3. **Always have outside counsel review anything binding.** This document is triage; real legal review is mandatory.
FILE:references/term_sheet_decoder.md
# Term Sheet Decoder
Glossary + founder-friendly defaults + pushback strategies for every clause in a standard venture term sheet. **Not legal advice.** Always engage venture / securities counsel before responding.
## The Three Clauses That Matter Most
In any term sheet review, focus disproportionately on these three. They drive ~80% of the founder economics impact.
### 1. Liquidation Preference
**What it is:** Investors get their investment back (the "preference") before founders see anything in an exit.
**The dimensions:**
- **Multiple:** 1x (standard) means $1 back per $1 invested. 2x means $2 back. Higher = more hostile.
- **Participating vs Non-participating:**
- **Non-participating (founder-friendly):** Investor chooses preference OR convert to common at exit. Most exits hit the conversion threshold, so preference is effectively just downside protection.
- **Participating ("double-dip"):** Investor gets preference back AND a pro-rata share of remaining proceeds as if converted. Significantly increases investor take in mid-range exits.
- **Cap:** Caps the total return at, say, 2x or 3x of investment for participating preferences. Limits the double-dip.
**Standard (Series A/B):** 1x non-participating.
**Hostile flavors:**
- 1x participating uncapped (significant founder dilution at exit)
- 2x preference (only acceptable in distressed rounds)
- Multi-stack preferences (Series A + Series B both get their preferences before any common)
**Pushback:** "Our standard is 1x non-participating. Participating preferences create misalignment with management at exit."
### 2. Option Pool — Pre-Money vs Post-Money
**The "option pool shuffle":** Investors typically require an unallocated option pool (10-20% of post-money) to be created **before** the new investment. If this comes out of pre-money, founders are diluted; if post-money, all shareholders dilute proportionally.
**Example math (Series A):**
| Scenario | Pre-Money | Pool Size | Effective Pre-Money for Founders |
|---|---|---|---|
| $30M pre, 10% pool pre-money | $30M | 10% of post | ~$26M (10% comes from founders) |
| $30M pre, 10% pool post-money | $30M | 10% of post | $30M (pool spread across all) |
**Standard:** 10-15% pool, often pre-money at Series A. Founder-friendly: smaller pool or post-money.
**Pushback:** "We've modeled our hiring plan and 8% supports the next 18 months. Let's right-size to actual need, not standard percentage." Or: "Pool top-up should come out of post-money so the new investor shares the dilution."
### 3. Anti-Dilution
**What it is:** Protection for investors against future down rounds. If a later round prices below the current, the current investor's price is adjusted retroactively.
**Flavors (least to most hostile):**
- **None:** Rare; only in seed SAFEs sometimes.
- **Broad-based weighted average (standard):** Adjusts using all shares (common, options, warrants). Modest founder dilution in a down round.
- **Narrow-based weighted average:** Uses only preferred. More dilutive than broad-based.
- **Full ratchet (hostile):** Investor's price resets entirely to the new round's price. Massively dilutive to founders.
**Standard:** Broad-based weighted average.
**Pushback:** "Full ratchet is non-starter at this stage. Narrow-based is unusual. We need broad-based weighted average — this is the NVCA standard."
---
## The Full Glossary
### Board Composition
**Standard at Series A:** 2 founders / 1 investor / 1 independent (or 1 founder / 1 investor / 1 independent for solo founders).
**At Series B:** Often 2 / 2 / 1 (balanced with independent tie-breaker).
**At Series C+:** Often investors get majority (signals control transition).
**Founder protection:** Always insist on the independent seat. Independent directors prevent deadlock and provide a neutral voice.
**Pushback on investor-majority boards at A:** "Investor control of the board at Series A is premature. Let's keep founder control with an independent tie-breaker until Series B."
### Vesting (for founders)
**Founder vesting in a financing:** Investors often require founder shares to be subject to vesting (re-vesting if you already exercised). Standard: 4 years, 1-year cliff. Often the cliff is waived if you've been at the company > 1 year.
**Acceleration:**
- **Single trigger:** All unvested shares vest immediately upon change of control. Founder-friendly but rare; investors resist.
- **Double trigger (standard):** Acceleration requires (a) change of control AND (b) involuntary termination of the founder within X months. Industry standard at Series A+.
**Pushback:** "Double-trigger acceleration is industry standard. Without it, founders are exposed to acquirer post-acquisition staffing decisions."
### Pro-Rata Rights
**What it is:** The right (but not obligation) to participate in future rounds proportionally to maintain ownership.
**Standard:** Lead investor + major investors (typically those above some ownership threshold) get pro-rata. Smaller checks often don't.
**Founder impact:** Granting pro-rata is generally fine — it shows investor conviction and aligns long-term. The cost is small dilution in future rounds.
**Pushback:** Only push back if there's a long tail of small investors each demanding pro-rata; cap to "major investors" defined by ownership %.
### Drag-Along
**What it is:** If a majority approves a sale, all shareholders must agree (including minority holders, including founders who later become minority).
**Founder-friendly version:** Drag-along requires founder consent OR a minimum sale price threshold (e.g., > 3x liquidation preference).
**Hostile version:** Drag-along with no founder consent and no price floor. Investors can force a sale at any price over founder objection.
**Pushback:** "Drag-along is standard, but we need founder consent OR a price floor."
### Protective Provisions
**What it is:** Investor consent rights for certain corporate decisions.
**Standard (NVCA model):**
- Issuing new senior or pari-passu preferred stock
- Authorizing new shares above existing pool
- Liquidating, merging, or selling the company
- Amending the charter or bylaws
- Increasing the board size
- Paying dividends
- Major debt
**Aggressive (push back):**
- Approving the annual budget
- Hiring or firing executives
- Setting compensation above thresholds
- Approving individual contracts above thresholds
- Capital expenditures above thresholds
**Pushback:** "We're aligned on the NVCA standard list. Operating decisions like budget and hiring are management's responsibility — protective provisions are for fundamental corporate changes."
### Information Rights
**Standard:** Quarterly unaudited financials, annual audited financials, annual budget.
**Aggressive (push back):** Monthly financials, board observer rights, weekly KPI dashboards, inspection rights at will.
**Pushback:** "Standard quarterly + annual is enough. Monthly creates significant CFO overhead at our stage. We'll commit to ad-hoc updates on material events."
### Dividends
**Standard:** None (default).
**Acceptable:** Non-cumulative dividends "when and if declared by the board" — almost never paid in practice.
**Hostile:** Cumulative dividends accrue every year regardless of declaration and must be paid in cash at exit. This is a creeping liquidation preference.
**Pushback:** "Cumulative dividends create a hidden liquidation preference that accrues over time. Non-cumulative when-declared, or none, is standard."
### Right of First Refusal (ROFR) / Co-Sale
**What it is:** If founders try to sell shares to a third party, investors have the right to buy first (ROFR) or to sell alongside (co-sale).
**Founder-friendly:** Standard ROFR + co-sale for all preferred; founders can still do secondary up to small thresholds without triggering.
**Hostile:** No secondary at all without unanimous investor consent.
**Pushback:** "We need to allow modest founder secondary (e.g., up to $1M aggregate) without investor consent — this is needed for founder financial planning."
### Founder Liquidity
**What it is:** Built-in secondary at later rounds (Series B/C) where founders sell some shares.
**Standard:** Becoming more common; 10-20% of round size as founder secondary.
**Pushback:** Raise this in Series B+ discussions; not typically negotiated at Series A.
### Most Favored Nation (MFN)
**What it is:** If you give a later investor better terms, the MFN-holder gets the same terms retroactively.
**Common in:** Seed SAFEs and convertible notes; rare in priced rounds.
**Founder trap:** MFN provisions can prevent you from offering competitive terms to new lead investors later. Be specific about what's covered (just SAFE terms? all terms?).
### No-Shop / Exclusivity
**What it is:** During due diligence, you can't shop the round to other investors.
**Standard:** 30-45 days. Founder-friendly. Investor-aligned because it shows commitment.
**Pushback only if:** > 60 days, or if it extends post-execution of definitive docs.
---
## Founder-Friendly Defaults (Cheat Sheet)
| Clause | Founder-Friendly Default |
|---|---|
| Liquidation preference | 1x non-participating |
| Anti-dilution | Broad-based weighted average |
| Option pool | 8-12%, post-money |
| Board (Series A) | 2F / 1I / 1Indep |
| Vesting (founder re-vest) | 4yr / 1yr cliff, often with credit for time served |
| Acceleration | Double-trigger |
| Pro-rata | For lead + major investors |
| Drag-along | Requires founder consent or price floor |
| Protective provisions | NVCA standard list only |
| Information rights | Quarterly + annual + budget |
| Dividends | None or non-cumulative when-declared |
| ROFR / co-sale | Standard, with carve-out for modest founder secondary |
| MFN (in notes/SAFEs) | Avoid if possible; if not, narrow scope |
| No-shop | 30-45 days |
---
## Negotiation Strategy
**Pick your battles:** A term sheet has 25-40 clauses. Winning every one is impossible and signals you don't understand priorities.
**Focus on the top 3 mistakes (in order):**
1. Liquidation preference flavor (participating vs non-participating)
2. Option pool pre-money vs post-money + size
3. Board control and protective provisions
These are the clauses where you can save 5-10% of founder economics or retain operating control. Everything else is secondary.
**The "founder-friendly NVCA" framing:** Many investors signal their posture by deviating from the NVCA model (the industry standard documents published by the National Venture Capital Association). Pushing back to "let's use the NVCA standard" is rarely rejected and resolves most issues.
**Walking away:** If a lead insists on:
- 1x participating uncapped preference
- Full ratchet anti-dilution
- Investor-majority board at Series A
- Cumulative dividends
These are not standard. A founder-friendly lead doesn't insist on these. Either walk or get specific written justification (sometimes a distressed cap-table situation justifies one of them, but never all).
---
## After Signing
Once the term sheet is signed:
1. **No-shop is active.** Don't talk to other investors except to officially decline.
2. **Definitive documents (SPA, IRA, Voting Agreement, ROFR Agreement) take 4-6 weeks.** Don't lose energy here; main fight was the term sheet.
3. **Closing conditions:** legal opinion, secretary's certificate, charter filing, capitalization confirmation.
4. **Wire timing:** Investors often wire 1-3 days after charter filing. Plan accordingly.
Run `scripts/term_sheet_analyzer.py` on the structured JSON of the term sheet for an automated scoring + flag analysis.
---
**Final reminder:** This document is a decoder, not a negotiation manual. Real term sheet response always involves your venture / securities counsel + your lead investor's diligence + your board (if any). Use this as a primer before those conversations.
FILE:scripts/contract_risk_scanner.py
#!/usr/bin/env python3
"""contract_risk_scanner.py — Scan a contract for founder-killer clauses.
Stdlib-only. Outputs human-readable or JSON. Detects 12 common risk patterns:
1. Unilateral termination favoring the counterparty
2. Auto-renewal with long notice (60+ days)
3. Uncapped liability or exclusion of standard caps
4. Broad indemnification flowing one direction
5. Non-mutual confidentiality
6. Missing or vague IP ownership clauses
7. Aggressive non-compete / non-solicit
8. Choice of law/venue in counterparty's home jurisdiction (one-sided)
9. Force majeure favoring only the counterparty
10. Missing DPA reference when personal data flows
11. Most-favored-nation pricing clauses
12. Audit rights without reciprocity
NOT legal advice. Use this to triage; bring findings to qualified counsel.
Usage:
python contract_risk_scanner.py # uses embedded sample
python contract_risk_scanner.py path/to/contract.txt
python contract_risk_scanner.py contract.txt --output json
python contract_risk_scanner.py --help
"""
import argparse
import json
import re
import sys
from dataclasses import dataclass, asdict
from typing import List
SAMPLE_CONTRACT = """\
MASTER SERVICES AGREEMENT
This Agreement shall automatically renew for successive one (1) year terms
unless either party provides ninety (90) days written notice of non-renewal.
LIMITATION OF LIABILITY. In no event shall Provider's aggregate liability
arising out of this Agreement exceed the fees paid by Customer in the
twelve (12) months preceding the claim. Notwithstanding the foregoing,
Customer's indemnification obligations under Section 8 shall be uncapped.
INDEMNIFICATION. Customer shall defend, indemnify and hold harmless
Provider, its affiliates, officers, directors and employees from and against
any and all claims, damages, losses and expenses arising out of or relating
to Customer's use of the Services.
INTELLECTUAL PROPERTY. The parties agree that intellectual property created
during the engagement shall belong to the party who develops it.
NON-COMPETE. For a period of three (3) years following termination, Customer
shall not engage with any competitor of Provider in any capacity, in any
geography.
GOVERNING LAW. This Agreement shall be governed by the laws of Delaware,
and any disputes shall be resolved exclusively in the state and federal
courts located in Wilmington, Delaware.
FORCE MAJEURE. Provider shall not be liable for any failure to perform due
to causes beyond its reasonable control.
"""
@dataclass
class Finding:
rule_id: str
severity: str # CRITICAL | HIGH | MEDIUM | LOW
title: str
excerpt: str
why_it_matters: str
suggested_redline: str
RULES = [
{
"id": "AUTO_RENEW_LONG_NOTICE",
"severity": "HIGH",
"title": "Auto-renewal with long notice period",
"pattern": re.compile(
r"automatically renew.{0,200}?(\d+|sixty|ninety|one hundred|180)\s*(\(\d+\))?\s*day",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"Auto-renewal with >30 day notice is a classic vendor trap: founders forget the "
"deadline and get locked into another full term. Especially painful on multi-year contracts."
),
"redline": (
"Counter: '...unless either party provides thirty (30) days written notice of non-renewal' "
"OR remove auto-renewal entirely and require affirmative re-signature."
),
},
{
"id": "UNCAPPED_CUSTOMER_INDEMNITY",
"severity": "CRITICAL",
"title": "Customer indemnity carved out from liability cap (uncapped)",
"pattern": re.compile(
r"(customer'?s|your)\s+indemnification.{0,200}?(uncapped|shall be uncapped|excluded from)",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"Uncapped customer indemnity means a single bad claim can exceed all fees ever paid. "
"Standard practice: mutual indemnity, both sides capped at fees, with narrow carve-outs "
"(IP infringement, data breach, gross negligence)."
),
"redline": (
"Counter: cap customer indemnity at 12 months of fees, mutual indemnity, carve-outs only "
"for willful misconduct and breach of confidentiality."
),
},
{
"id": "ONE_SIDED_INDEMNITY",
"severity": "HIGH",
"title": "Indemnification flows in one direction only",
"pattern": re.compile(
r"(customer|client)\s+shall\s+(defend|indemnify).{0,500}?(provider|company|vendor)",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"One-sided indemnity means you take on risk for the counterparty's actions without reciprocity. "
"A balanced contract has mutual indemnification with mirrored carve-outs."
),
"redline": (
"Counter: 'Each party shall defend, indemnify and hold harmless the other party...' with "
"mirrored scope and equal caps."
),
},
{
"id": "VAGUE_IP",
"severity": "CRITICAL",
"title": "Vague IP ownership clause",
"pattern": re.compile(
r"intellectual property.{0,200}?(belong to the party who develops it|jointly owned|to be determined|as agreed)",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"Vague IP language is the #1 source of post-engagement disputes. Joint ownership often means "
"neither party can license freely without the other's consent. 'As agreed' is unenforceable."
),
"redline": (
"Counter: 'All work product, deliverables, and derivative works created under this Agreement "
"shall be the sole and exclusive property of Customer. Provider hereby assigns all right, title "
"and interest...' Or explicitly carve out Provider's pre-existing IP and tools with a license back."
),
},
{
"id": "AGGRESSIVE_NONCOMPETE",
"severity": "HIGH",
"title": "Aggressive non-compete (long duration or broad geography)",
"pattern": re.compile(
r"non.compete.{0,300}?(two|three|four|five|2|3|4|5)\s*\(?\d*\)?\s*year",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"Non-competes >12 months or with unbounded geography are often unenforceable (especially in "
"California, and increasingly federally) but create chilling effects. They also signal the "
"counterparty's overall negotiation posture."
),
"redline": (
"Counter: maximum 12 months, specific competitor list (not 'any competitor'), specific "
"geography. For California-resident counterparties, remove entirely (California labor code "
"voids most non-competes)."
),
},
{
"id": "ONE_SIDED_VENUE",
"severity": "MEDIUM",
"title": "Choice of law/venue exclusively in counterparty jurisdiction",
"pattern": re.compile(
r"(exclusively in|exclusive jurisdiction).{0,300}?(courts? located in|state and federal courts of)",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"Exclusive venue in counterparty's jurisdiction means you bear travel cost and out-of-state "
"counsel cost for any dispute. For startups this can effectively prevent enforcement."
),
"redline": (
"Counter: neutral venue (Delaware is common), or 'venue in the jurisdiction of the defendant' "
"(forces plaintiff to travel), or arbitration in a neutral location with AAA/JAMS rules."
),
},
{
"id": "ONE_SIDED_FORCE_MAJEURE",
"severity": "MEDIUM",
"title": "Force majeure clause favors one party",
"pattern": re.compile(
r"(provider|company|vendor)\s+shall not be liable.{0,200}?(force majeure|causes beyond)",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"If only the vendor gets force-majeure protection, you pay full price during a pandemic / "
"outage / supply chain disruption but receive nothing. Mutual force majeure is standard."
),
"redline": (
"Counter: 'Neither party shall be liable...' with explicit list of qualifying events "
"(pandemic, war, natural disaster, government action) and a termination right after 30 days."
),
},
{
"id": "MISSING_DPA",
"severity": "HIGH",
"title": "Personal data appears to flow but no DPA referenced",
"pattern": re.compile(
r"(personal data|personally identifiable|user data|customer data|PII)(?!.{0,500}(DPA|data processing agreement|GDPR))",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"If personal data of EU residents (or California residents) flows, a DPA is legally required. "
"Missing DPA = GDPR Article 28 violation, potential 4%-of-revenue fine, contract unenforceable "
"with EU customers."
),
"redline": (
"Counter: 'The parties shall execute a Data Processing Agreement substantially in the form "
"of Exhibit X prior to any processing of Personal Data.' Use IAPP or Vendor-friendly DPA template."
),
},
{
"id": "MOST_FAVORED_NATION",
"severity": "MEDIUM",
"title": "Most-favored-nation (MFN) pricing clause",
"pattern": re.compile(
r"(most.favored.nation|MFN|best price|lowest price).{0,200}?(offered to|charged to)",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"MFN clauses prevent you from offering volume discounts or strategic pricing to anyone else. "
"If you sign with one customer, every future customer can demand the same price."
),
"redline": (
"Counter: remove the MFN entirely. If kept, narrow to 'similarly situated customers, same "
"tier and volume, excluding strategic / launch / migration discounts.'"
),
},
{
"id": "ONE_SIDED_AUDIT",
"severity": "MEDIUM",
"title": "Audit rights without reciprocity",
"pattern": re.compile(
r"(customer|client).{0,100}?right to audit",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"One-sided audit rights mean the counterparty can demand records on demand, often at your "
"expense. Reciprocity is standard for B2B agreements."
),
"redline": (
"Counter: mutual audit rights, max once per year, at requesting party's expense, with "
"30-day notice, during business hours, narrowed to specific compliance categories."
),
},
{
"id": "BROAD_NON_SOLICIT",
"severity": "MEDIUM",
"title": "Broad non-solicit (employees AND customers, long duration)",
"pattern": re.compile(
r"non.solicit.{0,300}?(employees? and customers?|customers? and employees?)",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"Combined employee + customer non-solicits, especially with long duration, can severely "
"limit hiring and business development. Many states limit enforceability."
),
"redline": (
"Counter: split into employee-only (12 months max) and customer-only (12 months max) clauses, "
"with carve-outs for general advertising / open job postings and for customers who initiate "
"contact independently."
),
},
{
"id": "PERPETUAL_LICENSE_BACK",
"severity": "HIGH",
"title": "Perpetual license-back to counterparty of your data or work",
"pattern": re.compile(
r"perpetual.{0,100}?(license|right).{0,300}?(customer data|user data|work product|deliverables)",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"A perpetual license-back lets the counterparty use your data or deliverables forever, even "
"after termination. This is acceptable for usage analytics, NOT for customer data or core IP."
),
"redline": (
"Counter: time-limited license (for the term of the agreement only), specific purpose "
"(service delivery only, not training AI models, not sharing with third parties), and "
"post-termination return-or-destroy obligation."
),
},
]
def scan(text: str) -> List[Finding]:
findings: List[Finding] = []
for rule in RULES:
for match in rule["pattern"].finditer(text):
excerpt = match.group(0).strip()
# truncate long excerpts
if len(excerpt) > 300:
excerpt = excerpt[:297] + "..."
findings.append(Finding(
rule_id=rule["id"],
severity=rule["severity"],
title=rule["title"],
excerpt=excerpt,
why_it_matters=rule["why_it_matters"],
suggested_redline=rule["redline"],
))
# rank by severity then rule order
severity_order = {"CRITICAL": 0, "HIGH": 1, "MEDIUM": 2, "LOW": 3}
findings.sort(key=lambda f: (severity_order.get(f.severity, 9), f.rule_id))
return findings
def render_text(findings: List[Finding], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("CONTRACT RISK SCAN")
lines.append(f"Source: {source}")
lines.append(f"Findings: {len(findings)}")
lines.append("=" * 72)
lines.append("")
if not findings:
lines.append("No risk patterns matched. (Absence of findings does not mean the contract is safe;")
lines.append("it means the 12 common patterns this scanner checks did not trigger.)")
lines.append("")
lines.append("Always engage qualified counsel before signing.")
return "\n".join(lines)
severity_counts = {}
for f in findings:
severity_counts[f.severity] = severity_counts.get(f.severity, 0) + 1
severity_summary = " ".join(
f"{sev}: {severity_counts.get(sev, 0)}"
for sev in ("CRITICAL", "HIGH", "MEDIUM", "LOW")
if severity_counts.get(sev, 0) > 0
)
lines.append(f"Severity: {severity_summary}")
lines.append("")
for i, f in enumerate(findings, 1):
lines.append(f"[{i}] {f.severity} — {f.title}")
lines.append(f" Rule: {f.rule_id}")
lines.append(f" Excerpt: \"{f.excerpt}\"")
lines.append("")
lines.append(f" Why it matters:")
for line in _wrap(f.why_it_matters, 4):
lines.append(line)
lines.append("")
lines.append(f" Suggested redline:")
for line in _wrap(f.suggested_redline, 4):
lines.append(line)
lines.append("")
lines.append("-" * 72)
lines.append("")
lines.append("REMINDER: This scanner triages obvious traps. Always bring redlines to qualified counsel.")
return "\n".join(lines)
def _wrap(text: str, indent: int, width: int = 68) -> List[str]:
import textwrap
return textwrap.wrap(text, width=width, initial_indent=" " * indent, subsequent_indent=" " * indent) or [" " * indent + text]
def main() -> int:
parser = argparse.ArgumentParser(
description="Scan a contract for the 12 most common founder-killer clauses.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to contract text file (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
text = f.read()
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
else:
text = SAMPLE_CONTRACT
source = "<embedded sample MSA>"
findings = scan(text)
if args.output == "json":
payload = {
"source": source,
"findings_count": len(findings),
"findings": [asdict(f) for f in findings],
}
print(json.dumps(payload, indent=2))
else:
print(render_text(findings, source))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/term_sheet_analyzer.py
#!/usr/bin/env python3
"""term_sheet_analyzer.py — Score a term sheet on founder-friendliness.
Stdlib-only. Computes a 0-100 score across 12 dimensions and flags
hostile clauses. Outputs human-readable or JSON.
NOT legal advice — surfaces questions for venture / securities counsel.
Input schema (JSON):
{
"round": "Series A",
"pre_money": 30000000,
"raise_amount": 8000000,
"liquidation_preference": {
"multiple": 1.0,
"participating": false,
"cap": null
},
"anti_dilution": "broad_based_weighted_average", // | "narrow_based_weighted_average" | "full_ratchet" | "none"
"option_pool": {
"size_pct": 12.0,
"pre_money": true
},
"board_composition": {
"investor_seats": 1,
"founder_seats": 2,
"independent_seats": 1
},
"vesting": {
"standard_years": 4,
"cliff_months": 12,
"single_trigger_acceleration": false,
"double_trigger_acceleration": true
},
"pro_rata": true,
"drag_along": {
"exists": true,
"founder_consent_required": true
},
"protective_provisions": "standard", // | "standard" | "aggressive"
"information_rights": "standard", // | "standard" | "aggressive"
"dividends": "none" // | "none" | "non_cumulative_when_declared" | "cumulative"
}
Usage:
python term_sheet_analyzer.py # uses embedded sample
python term_sheet_analyzer.py path/to/term_sheet.json
python term_sheet_analyzer.py term_sheet.json --output json
python term_sheet_analyzer.py --help
"""
import argparse
import json
import sys
from typing import Any, Dict, List, Tuple
SAMPLE = {
"round": "Series A",
"pre_money": 30_000_000,
"raise_amount": 8_000_000,
"liquidation_preference": {"multiple": 1.0, "participating": False, "cap": None},
"anti_dilution": "broad_based_weighted_average",
"option_pool": {"size_pct": 12.0, "pre_money": True},
"board_composition": {"investor_seats": 1, "founder_seats": 2, "independent_seats": 1},
"vesting": {
"standard_years": 4,
"cliff_months": 12,
"single_trigger_acceleration": False,
"double_trigger_acceleration": True,
},
"pro_rata": True,
"drag_along": {"exists": True, "founder_consent_required": True},
"protective_provisions": "standard",
"information_rights": "standard",
"dividends": "none",
}
def score(ts: Dict[str, Any]) -> Tuple[int, List[Dict[str, Any]]]:
"""Returns (total_score_0_to_100, list_of_findings).
Each dimension is scored 0-100, then averaged. Findings list contains
per-clause analysis with severity.
"""
findings: List[Dict[str, Any]] = []
scores: List[int] = []
# --- 1. Liquidation Preference (high signal) ---
lp = ts.get("liquidation_preference", {})
lp_mult = lp.get("multiple", 1.0)
lp_part = lp.get("participating", False)
lp_cap = lp.get("cap")
if lp_mult == 1.0 and not lp_part:
lp_score = 100
findings.append(_ok("liquidation_preference", "1x non-participating — founder-friendly standard."))
elif lp_mult == 1.0 and lp_part and lp_cap and lp_cap <= 3:
lp_score = 55
findings.append(_warn("liquidation_preference",
f"1x participating with {lp_cap}x cap. Investor double-dips up to cap. "
"Push for non-participating; if accepted, accept cap < 3x."))
elif lp_mult == 1.0 and lp_part and not lp_cap:
lp_score = 25
findings.append(_crit("liquidation_preference",
"1x PARTICIPATING UNCAPPED. Investor gets their money back AND a pro-rata share of remaining proceeds, "
"forever. Hostile. Push to non-participating or at minimum cap at 2x."))
elif lp_mult > 1.0:
lp_score = 10
findings.append(_crit("liquidation_preference",
f"{lp_mult}x preference. Investor gets {lp_mult}x their money back before founders see a dollar. "
"Hostile; only acceptable in distressed rounds."))
else:
lp_score = 80
findings.append(_ok("liquidation_preference", f"{lp_mult}x configuration acceptable."))
scores.append(lp_score)
# --- 2. Anti-Dilution ---
ad = ts.get("anti_dilution", "broad_based_weighted_average")
if ad == "broad_based_weighted_average":
ad_score = 100
findings.append(_ok("anti_dilution", "Broad-based weighted average — founder-friendly standard."))
elif ad == "narrow_based_weighted_average":
ad_score = 70
findings.append(_warn("anti_dilution",
"Narrow-based weighted average. More dilutive to founders than broad-based in a down round. "
"Push to broad-based."))
elif ad == "full_ratchet":
ad_score = 10
findings.append(_crit("anti_dilution",
"FULL RATCHET. In a down round, investor's price is reset to the new round price entirely, "
"massively diluting founders. Hostile; reject."))
elif ad == "none":
ad_score = 100
findings.append(_ok("anti_dilution", "No anti-dilution provision. Unusual but founder-friendly."))
else:
ad_score = 50
findings.append(_warn("anti_dilution", f"Unrecognized anti-dilution type: {ad}. Verify with counsel."))
scores.append(ad_score)
# --- 3. Option Pool (pre-money vs post-money) ---
op = ts.get("option_pool", {})
op_pre = op.get("pre_money", True)
op_size = op.get("size_pct", 10.0)
if not op_pre:
op_score = 100
findings.append(_ok("option_pool",
f"Pool of {op_size}% sits post-money — dilutes all shareholders proportionally."))
elif op_pre and op_size <= 10.0:
op_score = 70
findings.append(_warn("option_pool",
f"Pool of {op_size}% pre-money — comes out of founders' shares. Reasonable size, but consider "
"negotiating post-money or sharing the pool top-up across the round."))
elif op_pre and op_size > 10.0:
op_score = 30
findings.append(_crit("option_pool",
f"Pool of {op_size}% PRE-MONEY. This is the 'option pool shuffle' — typically reduces pre-money "
f"by ~{op_size}%, diluting founders silently. Negotiate hard: justify the size with a hiring plan "
"or push for post-money."))
else:
op_score = 60
findings.append(_warn("option_pool", "Option pool structure unclear; verify."))
scores.append(op_score)
# --- 4. Board Composition ---
bc = ts.get("board_composition", {})
inv = bc.get("investor_seats", 0)
fnd = bc.get("founder_seats", 0)
ind = bc.get("independent_seats", 0)
total = inv + fnd + ind
if total == 0:
bc_score = 50
findings.append(_warn("board_composition", "Board composition unspecified."))
elif fnd > inv and ind >= 1:
bc_score = 100
findings.append(_ok("board_composition",
f"{fnd} founder / {inv} investor / {ind} independent — founder-friendly; founders retain control "
"with independent tie-breaker."))
elif fnd == inv and ind >= 1:
bc_score = 75
findings.append(_ok("board_composition",
f"{fnd} founder / {inv} investor / {ind} independent — balanced, independent is critical."))
elif inv > fnd:
bc_score = 30
findings.append(_crit("board_composition",
f"{fnd} founder / {inv} investor / {ind} independent — investors control the board at Series A. "
"This is unusually early; investor control typically arrives at Series B or later."))
else:
bc_score = 50
findings.append(_warn("board_composition", f"Composition: {fnd}F/{inv}I/{ind}Ind — verify with counsel."))
scores.append(bc_score)
# --- 5. Vesting & Acceleration ---
vest = ts.get("vesting", {})
years = vest.get("standard_years", 4)
cliff = vest.get("cliff_months", 12)
single = vest.get("single_trigger_acceleration", False)
double = vest.get("double_trigger_acceleration", False)
if years == 4 and cliff == 12 and double and not single:
vest_score = 100
findings.append(_ok("vesting",
"4yr/1yr cliff with double-trigger acceleration — founder-friendly standard. "
"Single-trigger is rare and not recommended by counsel."))
elif years == 4 and cliff == 12 and not double:
vest_score = 60
findings.append(_warn("vesting",
"4yr/1yr cliff WITHOUT acceleration. Push for double-trigger (change of control + termination "
"without cause) to protect founder upside in acquisition scenarios."))
elif years > 4:
vest_score = 20
findings.append(_crit("vesting",
f"{years}-year vesting. Non-standard; reject. 4 years is industry norm."))
else:
vest_score = 70
findings.append(_warn("vesting", f"{years}yr/{cliff}mo cliff — verify acceleration with counsel."))
scores.append(vest_score)
# --- 6. Pro-Rata Rights ---
if ts.get("pro_rata", True):
pr_score = 100
findings.append(_ok("pro_rata", "Pro-rata rights — standard for the lead and major investors."))
else:
pr_score = 60
findings.append(_warn("pro_rata",
"No pro-rata rights. Unusual; if investor is offering this, ask why (signals weak conviction "
"or competitive pressure). Pro-rata is generally fine for founders to grant."))
scores.append(pr_score)
# --- 7. Drag-Along ---
drag = ts.get("drag_along", {})
if drag.get("exists") and drag.get("founder_consent_required"):
drag_score = 100
findings.append(_ok("drag_along",
"Drag-along exists but requires founder consent — balanced."))
elif drag.get("exists") and not drag.get("founder_consent_required"):
drag_score = 40
findings.append(_crit("drag_along",
"Drag-along WITHOUT founder consent. Investors can force a sale over founder objection. "
"Push for founder consent OR a minimum price threshold (e.g., 3x preference) to trigger drag."))
else:
drag_score = 80
findings.append(_ok("drag_along", "No drag-along — neutral; common at early stages."))
scores.append(drag_score)
# --- 8. Protective Provisions ---
pp = ts.get("protective_provisions", "standard")
if pp == "standard":
pp_score = 100
findings.append(_ok("protective_provisions",
"Standard protective provisions (NVCA model) — acceptable."))
elif pp == "aggressive":
pp_score = 40
findings.append(_crit("protective_provisions",
"Aggressive protective provisions can require investor consent for routine operating "
"decisions (hiring execs, budget changes, vendor contracts). Push back to NVCA standard."))
else:
pp_score = 70
findings.append(_warn("protective_provisions", f"Verify scope with counsel: {pp}"))
scores.append(pp_score)
# --- 9. Information Rights ---
ir = ts.get("information_rights", "standard")
if ir == "standard":
ir_score = 100
findings.append(_ok("information_rights",
"Standard information rights (quarterly financials, annual audited, budget) — acceptable."))
elif ir == "aggressive":
ir_score = 60
findings.append(_warn("information_rights",
"Aggressive information rights (monthly financials, board observer rights, inspection rights). "
"Reasonable for lead at Series B+; at Series A, push to quarterly."))
else:
ir_score = 75
findings.append(_warn("information_rights", f"Verify: {ir}"))
scores.append(ir_score)
# --- 10. Dividends ---
div = ts.get("dividends", "none")
if div == "none":
div_score = 100
findings.append(_ok("dividends", "No dividend obligation — founder-friendly standard."))
elif div == "non_cumulative_when_declared":
div_score = 80
findings.append(_ok("dividends",
"Non-cumulative when-declared dividends — acceptable; rare to actually be paid."))
elif div == "cumulative":
div_score = 30
findings.append(_crit("dividends",
"CUMULATIVE dividends accrue every year regardless of declaration and must be paid at exit. "
"Hostile; push to non-cumulative or none."))
else:
div_score = 60
findings.append(_warn("dividends", f"Verify dividend type: {div}"))
scores.append(div_score)
# --- 11. Valuation Sanity ---
pre = ts.get("pre_money", 0)
raise_amt = ts.get("raise_amount", 0)
if pre and raise_amt:
post = pre + raise_amt
dilution = (raise_amt / post) * 100
if dilution > 30:
val_score = 40
findings.append(_crit("valuation",
f"Round dilutes {dilution:.1f}% (raise , on , pre = , post). "
"Over 30% in a single round is heavy; standard is 15-25%."))
elif dilution > 25:
val_score = 70
findings.append(_warn("valuation",
f"Round dilutes {dilution:.1f}%. Acceptable but on the high end. Standard 15-25%."))
else:
val_score = 100
findings.append(_ok("valuation",
f"Round dilutes {dilution:.1f}% — within standard 15-25% range."))
scores.append(val_score)
# --- 12. Holistic posture ---
crit_count = sum(1 for f in findings if f["severity"] == "CRITICAL")
if crit_count >= 3:
findings.append(_crit("holistic",
f"{crit_count} CRITICAL flags. This is a hostile term sheet. Either renegotiate the worst clauses "
"or walk. Do not sign as-is."))
elif crit_count >= 1:
findings.append(_warn("holistic",
f"{crit_count} CRITICAL flag(s). Address before signing; the rest is negotiable but not "
"disqualifying."))
else:
findings.append(_ok("holistic", "No critical flags. Standard founder-friendly term sheet."))
total_score = round(sum(scores) / len(scores)) if scores else 0
return total_score, findings
def _ok(clause: str, msg: str) -> Dict[str, Any]:
return {"clause": clause, "severity": "OK", "message": msg}
def _warn(clause: str, msg: str) -> Dict[str, Any]:
return {"clause": clause, "severity": "WARN", "message": msg}
def _crit(clause: str, msg: str) -> Dict[str, Any]:
return {"clause": clause, "severity": "CRITICAL", "message": msg}
def render_text(score_val: int, findings: List[Dict[str, Any]], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("TERM SHEET ANALYSIS")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
grade = (
"🟢 FOUNDER-FRIENDLY" if score_val >= 85 else
"🟡 NEGOTIATE" if score_val >= 65 else
"🔴 HOSTILE"
)
lines.append(f"Founder-friendliness score: {score_val}/100 {grade}")
lines.append("")
lines.append("-" * 72)
for f in findings:
sev = f["severity"]
marker = {"OK": "✅", "WARN": "⚠️ ", "CRITICAL": "🚨"}.get(sev, "•")
lines.append(f"{marker} [{sev:>8}] {f['clause']}")
lines.append(f" {f['message']}")
lines.append("")
lines.append("-" * 72)
lines.append("REMINDER: This tool is not legal advice. Always engage venture / securities counsel.")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="Score a term sheet on founder-friendliness across 12 dimensions.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to term sheet JSON file (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
ts = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
ts = SAMPLE
source = "<embedded sample Series A term sheet>"
score_val, findings = score(ts)
if args.output == "json":
print(json.dumps({
"source": source,
"score": score_val,
"grade": "FOUNDER_FRIENDLY" if score_val >= 85 else "NEGOTIATE" if score_val >= 65 else "HOSTILE",
"findings": findings,
}, indent=2))
else:
print(render_text(score_val, findings, source))
return 0
if __name__ == "__main__":
sys.exit(main())
Hoạch định chiến lược ra mắt sản phẩm hoặc phát hành tính năng, gồm Product Hunt, beta, early access, waitlist, kế hoạch GTM và checklist ra mắt.
---
name: "launch-strategy"
description: "When the user wants to plan a product launch, feature announcement, or release strategy. Also use when the user mentions 'launch,' 'Product Hunt,' 'feature release,' 'announcement,' 'go-to-market,' 'beta launch,' 'early access,' 'waitlist,' 'product update,' 'GTM plan,' 'launch checklist,' or 'launch momentum.' This skill covers phased launches, channel strategy, and ongoing launch momentum."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: marketing
updated: 2026-03-06
---
# Launch Strategy
You are an expert in SaaS product launches and feature announcements. Your goal is to help users plan launches that build momentum, capture attention, and convert interest into users.
## Before Starting
**Check for product marketing context first:**
If `.claude/product-marketing-context.md` exists, read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
---
## Core Philosophy
→ See references/launch-frameworks-and-checklists.md for details
## Task-Specific Questions
1. What are you launching? (New product, major feature, minor update)
2. What's your current audience size and engagement?
3. What owned channels do you have? (Email list size, blog traffic, community)
4. What's your timeline for launch?
5. Have you launched before? What worked/didn't work?
6. Are you considering Product Hunt? What's your preparation status?
---
## Proactive Triggers
Proactively offer launch planning when:
1. **Feature ship date mentioned** — When an engineering delivery date is discussed, immediately ask about the launch plan; shipping without a marketing plan is a missed opportunity.
2. **Waitlist or early access mentioned** — Offer to design the full phased launch funnel from alpha through full GA, not just the landing page.
3. **Product Hunt consideration** — Any mention of Product Hunt should trigger the full PH strategy section including pre-launch relationship building timeline.
4. **Post-launch silence** — If a user launched recently but hasn't followed up with momentum content, proactively suggest the post-launch marketing actions (comparison pages, roundup email, interactive demo).
5. **Pricing change planned** — Pricing updates are a launch opportunity; offer to build an announcement campaign treating it as a product update.
---
## Output Artifacts
| Artifact | Format | Description |
|----------|--------|-------------|
| Launch Plan | Markdown doc | Phase-by-phase plan with owners, dates, channels, and success metrics |
| ORB Channel Map | Table | Owned/Rented/Borrowed channel strategy with tactics per channel |
| Launch Day Checklist | Checklist | Complete day-of execution checklist with time-boxed actions |
| Product Hunt Brief | Markdown doc | Listing copy, asset specs, pre-launch timeline, engagement playbook |
| Post-Launch Momentum Plan | Bulleted list | 30-day post-launch actions to sustain and compound the launch |
---
## Communication
Launch plans should be concrete, time-bound, and channel-specific — no vague "post on social media" recommendations. Every output should specify who does what and when. Reference `marketing-context` to ensure the launch narrative matches ICP language and positioning before drafting any copy. Quality bar: a launch plan is only complete when it covers all three ORB channel types and includes both launch-day and post-launch actions.
---
## Related Skills
- **email-sequence** — USE for building the launch announcement and post-launch onboarding email sequences; NOT as a substitute for the full channel strategy.
- **social-content** — USE for drafting the specific social posts and threads for launch day; NOT for channel selection strategy.
- **paid-ads** — USE when the launch plan includes a paid amplification component; NOT for organic launch-only strategies.
- **content-strategy** — USE when the launch requires a sustained content program (blog posts, case studies) in the weeks after; NOT for single-day launch execution.
- **pricing-strategy** — USE when the launch involves a pricing change or new tier introduction; NOT for feature-only launches.
- **marketing-context** — USE as foundation to align launch messaging with ICP and brand voice; always load first.
FILE:references/launch-frameworks-and-checklists.md
# launch-strategy reference
## Core Philosophy
The best companies don't just launch once—they launch again and again. Every new feature, improvement, and update is an opportunity to capture attention and engage your audience.
A strong launch isn't about a single moment. It's about:
- Getting your product into users' hands early
- Learning from real feedback
- Making a splash at every stage
- Building momentum that compounds over time
---
## The ORB Framework
Structure your launch marketing across three channel types. Everything should ultimately lead back to owned channels.
### Owned Channels
You own the channel (though not the audience). Direct access without algorithms or platform rules.
**Examples:**
- Email list
- Blog
- Podcast
- Branded community (Slack, Discord)
- Website/product
**Why they matter:**
- Get more effective over time
- No algorithm changes or pay-to-play
- Direct relationship with audience
- Compound value from content
**Start with 1-2 based on audience:**
- Industry lacks quality content → Start a blog
- People want direct updates → Focus on email
- Engagement matters → Build a community
**Example - Superhuman:**
Built demand through an invite-only waitlist and one-on-one onboarding sessions. Every new user got a 30-minute live demo. This created exclusivity, FOMO, and word-of-mouth—all through owned relationships. Years later, their original onboarding materials still drive engagement.
### Rented Channels
Platforms that provide visibility but you don't control. Algorithms shift, rules change, pay-to-play increases.
**Examples:**
- Social media (Twitter/X, LinkedIn, Instagram)
- App stores and marketplaces
- YouTube
- Reddit
**How to use correctly:**
- Pick 1-2 platforms where your audience is active
- Use them to drive traffic to owned channels
- Don't rely on them as your only strategy
**Example - Notion:**
Hacked virality through Twitter, YouTube, and Reddit where productivity enthusiasts were active. Encouraged community to share templates and workflows. But they funneled all visibility into owned assets—every viral post led to signups, then targeted email onboarding.
**Platform-specific tactics:**
- Twitter/X: Threads that spark conversation → link to newsletter
- LinkedIn: High-value posts → lead to gated content or email signup
- Marketplaces (Shopify, Slack): Optimize listing → drive to site for more
Rented channels give speed, not stability. Capture momentum by bringing users into your owned ecosystem.
### Borrowed Channels
Tap into someone else's audience to shortcut the hardest part—getting noticed.
**Examples:**
- Guest content (blog posts, podcast interviews, newsletter features)
- Collaborations (webinars, co-marketing, social takeovers)
- Speaking engagements (conferences, panels, virtual summits)
- Influencer partnerships
**Be proactive, not passive:**
1. List industry leaders your audience follows
2. Pitch win-win collaborations
3. Use tools like SparkToro or Listen Notes to find audience overlap
4. Set up affiliate/referral incentives
**Example - TRMNL:**
Sent a free e-ink display to YouTuber Snazzy Labs—not a paid sponsorship, just hoping he'd like it. He created an in-depth review that racked up 500K+ views and drove $500K+ in sales. They also set up an affiliate program for ongoing promotion.
Borrowed channels give instant credibility, but only work if you convert borrowed attention into owned relationships.
---
## Five-Phase Launch Approach
Launching isn't a one-day event. It's a phased process that builds momentum.
### Phase 1: Internal Launch
Gather initial feedback and iron out major issues before going public.
**Actions:**
- Recruit early users one-on-one to test for free
- Collect feedback on usability gaps and missing features
- Ensure prototype is functional enough to demo (doesn't need to be production-ready)
**Goal:** Validate core functionality with friendly users.
### Phase 2: Alpha Launch
Put the product in front of external users in a controlled way.
**Actions:**
- Create landing page with early access signup form
- Announce the product exists
- Invite users individually to start testing
- MVP should be working in production (even if still evolving)
**Goal:** First external validation and initial waitlist building.
### Phase 3: Beta Launch
Scale up early access while generating external buzz.
**Actions:**
- Work through early access list (some free, some paid)
- Start marketing with teasers about problems you solve
- Recruit friends, investors, and influencers to test and share
**Consider adding:**
- Coming soon landing page or waitlist
- "Beta" sticker in dashboard navigation
- Email invites to early access list
- Early access toggle in settings for experimental features
**Goal:** Build buzz and refine product with broader feedback.
### Phase 4: Early Access Launch
Shift from small-scale testing to controlled expansion.
**Actions:**
- Leak product details: screenshots, feature GIFs, demos
- Gather quantitative usage data and qualitative feedback
- Run user research with engaged users (incentivize with credits)
- Optionally run product/market fit survey to refine messaging
**Expansion options:**
- Option A: Throttle invites in batches (5-10% at a time)
- Option B: Invite all users at once under "early access" framing
**Goal:** Validate at scale and prepare for full launch.
### Phase 5: Full Launch
Open the floodgates.
**Actions:**
- Open self-serve signups
- Start charging (if not already)
- Announce general availability across all channels
**Launch touchpoints:**
- Customer emails
- In-app popups and product tours
- Website banner linking to launch assets
- "New" sticker in dashboard navigation
- Blog post announcement
- Social posts across platforms
- Product Hunt, BetaList, Hacker News, etc.
**Goal:** Maximum visibility and conversion to paying users.
---
## Product Hunt Launch Strategy
Product Hunt can be powerful for reaching early adopters, but it's not magic—it requires preparation.
### Pros
- Exposure to tech-savvy early adopter audience
- Credibility bump (especially if Product of the Day)
- Potential PR coverage and backlinks
### Cons
- Very competitive to rank well
- Short-lived traffic spikes
- Requires significant pre-launch planning
### How to Launch Successfully
**Before launch day:**
1. Build relationships with influential supporters, content hubs, and communities
2. Optimize your listing: compelling tagline, polished visuals, short demo video
3. Study successful launches to identify what worked
4. Engage in relevant communities—provide value before pitching
5. Prepare your team for all-day engagement
**On launch day:**
1. Treat it as an all-day event
2. Respond to every comment in real-time
3. Answer questions and spark discussions
4. Encourage your existing audience to engage
5. Direct traffic back to your site to capture signups
**After launch day:**
1. Follow up with everyone who engaged
2. Convert Product Hunt traffic into owned relationships (email signups)
3. Continue momentum with post-launch content
### Case Studies
**SavvyCal** (Scheduling tool):
- Optimized landing page and onboarding before launch
- Built relationships with productivity/SaaS influencers in advance
- Responded to every comment on launch day
- Result: #2 Product of the Month
**Reform** (Form builder):
- Studied successful launches and applied insights
- Crafted clear tagline, polished visuals, demo video
- Engaged in communities before launch (provided value first)
- Treated launch as all-day engagement event
- Directed traffic to capture signups
- Result: #1 Product of the Day
---
## Post-Launch Product Marketing
Your launch isn't over when the announcement goes live. Now comes adoption and retention work.
### Immediate Post-Launch Actions
**Educate new users:**
Set up automated onboarding email sequence introducing key features and use cases.
**Reinforce the launch:**
Include announcement in your weekly/biweekly/monthly roundup email to catch people who missed it.
**Differentiate against competitors:**
Publish comparison pages highlighting why you're the obvious choice.
**Update web pages:**
Add dedicated sections about the new feature/product across your site.
**Offer hands-on preview:**
Create no-code interactive demo (using tools like Navattic) so visitors can explore before signing up.
### Keep Momentum Going
It's easier to build on existing momentum than start from scratch. Every touchpoint reinforces the launch.
---
## Ongoing Launch Strategy
Don't rely on a single launch event. Regular updates and feature rollouts sustain engagement.
### How to Prioritize What to Announce
Use this matrix to decide how much marketing each update deserves:
**Major updates** (new features, product overhauls):
- Full campaign across multiple channels
- Blog post, email campaign, in-app messages, social media
- Maximize exposure
**Medium updates** (new integrations, UI enhancements):
- Targeted announcement
- Email to relevant segments, in-app banner
- Don't need full fanfare
**Minor updates** (bug fixes, small tweaks):
- Changelog and release notes
- Signal that product is improving
- Don't dominate marketing
### Announcement Tactics
**Space out releases:**
Instead of shipping everything at once, stagger announcements to maintain momentum.
**Reuse high-performing tactics:**
If a previous announcement resonated, apply those insights to future updates.
**Keep engaging:**
Continue using email, social, and in-app messaging to highlight improvements.
**Signal active development:**
Even small changelog updates remind customers your product is evolving. This builds retention and word-of-mouth—customers feel confident you'll be around.
---
## Launch Checklist
### Pre-Launch
- [ ] Landing page with clear value proposition
- [ ] Email capture / waitlist signup
- [ ] Early access list built
- [ ] Owned channels established (email, blog, community)
- [ ] Rented channel presence (social profiles optimized)
- [ ] Borrowed channel opportunities identified (podcasts, influencers)
- [ ] Product Hunt listing prepared (if using)
- [ ] Launch assets created (screenshots, demo video, GIFs)
- [ ] Onboarding flow ready
- [ ] Analytics/tracking in place
### Launch Day
- [ ] Announcement email to list
- [ ] Blog post published
- [ ] Social posts scheduled and posted
- [ ] Product Hunt listing live (if using)
- [ ] In-app announcement for existing users
- [ ] Website banner/notification active
- [ ] Team ready to engage and respond
- [ ] Monitor for issues and feedback
### Post-Launch
- [ ] Onboarding email sequence active
- [ ] Follow-up with engaged prospects
- [ ] Roundup email includes announcement
- [ ] Comparison pages published
- [ ] Interactive demo created
- [ ] Gather and act on feedback
- [ ] Plan next launch moment
---
FILE:scripts/launch_readiness_scorer.py
#!/usr/bin/env python3
"""
launch_readiness_scorer.py — Product Launch Readiness Scorer
100% stdlib, no pip installs required.
Usage:
python3 launch_readiness_scorer.py # demo mode
python3 launch_readiness_scorer.py --checklist checklist.json
python3 launch_readiness_scorer.py --checklist checklist.json --json
python3 launch_readiness_scorer.py --export-template > my_checklist.json
checklist.json format:
{
"product": [
{"item": "Beta tested with 10+ users", "status": "done"},
{"item": "Documentation ready", "status": "partial"},
{"item": "Support team trained", "status": "not_started"}
],
"marketing": [...],
"technical": [...]
}
Valid status values: "done" | "partial" | "not_started"
"""
import argparse
import json
import sys
from datetime import datetime, timezone
# ---------------------------------------------------------------------------
# Default checklist template
# ---------------------------------------------------------------------------
DEFAULT_CHECKLIST = {
"product": [
{"item": "Beta tested with real users (≥10)", "status": "done", "weight": 3},
{"item": "Core user journey validated end-to-end", "status": "done", "weight": 3},
{"item": "Known P0/P1 bugs resolved", "status": "partial", "weight": 3},
{"item": "User-facing documentation complete", "status": "partial", "weight": 2},
{"item": "In-app onboarding / empty states ready", "status": "done", "weight": 2},
{"item": "Support team trained on common Q&A", "status": "not_started", "weight": 2},
{"item": "Pricing finalised and live", "status": "done", "weight": 2},
{"item": "Accessibility basics checked (WCAG AA)", "status": "not_started", "weight": 1},
{"item": "Localisation / i18n ready (if applicable)", "status": "done", "weight": 1},
{"item": "Feedback collection mechanism in place", "status": "partial", "weight": 1},
],
"marketing": [
{"item": "Landing page live and conversion-optimised", "status": "done", "weight": 3},
{"item": "Email announcement list ready (≥100)", "status": "done", "weight": 3},
{"item": "Press / media kit prepared", "status": "partial", "weight": 2},
{"item": "Social media assets created", "status": "done", "weight": 2},
{"item": "Product Hunt / launch platform submission", "status": "not_started", "weight": 2},
{"item": "SEO meta tags and OG images set", "status": "done", "weight": 2},
{"item": "Influencer / community outreach planned", "status": "partial", "weight": 2},
{"item": "Launch-day email sequence scheduled", "status": "not_started", "weight": 2},
{"item": "Paid ads creative prepared (if applicable)", "status": "not_started", "weight": 1},
{"item": "Referral / viral loop mechanism designed", "status": "not_started", "weight": 1},
],
"technical": [
{"item": "Production monitoring & alerting active", "status": "done", "weight": 3},
{"item": "Load / performance tested at 5× expected", "status": "partial", "weight": 3},
{"item": "Rollback plan documented and rehearsed", "status": "not_started", "weight": 3},
{"item": "Database backups verified and automated", "status": "done", "weight": 2},
{"item": "CDN / caching configured", "status": "done", "weight": 2},
{"item": "Error tracking (Sentry/similar) live", "status": "done", "weight": 2},
{"item": "SSL / HTTPS confirmed on all endpoints", "status": "done", "weight": 2},
{"item": "Analytics events firing correctly", "status": "partial", "weight": 2},
{"item": "Rate limiting / DDoS protection in place", "status": "partial", "weight": 2},
{"item": "Feature flags configured for safe rollout", "status": "not_started", "weight": 1},
],
}
CATEGORY_META = {
"product": {"emoji": "🛠 ", "label": "Product Readiness"},
"marketing": {"emoji": "📣 ", "label": "Marketing Readiness"},
"technical": {"emoji": "⚙️ ", "label": "Technical Readiness"},
}
STATUS_WEIGHTS = {
"done": 1.0,
"partial": 0.5,
"not_started": 0.0,
}
BLOCKERS_THRESHOLD = 0.0 # not_started items with weight ≥3 are blockers
# ---------------------------------------------------------------------------
# Core scoring
# ---------------------------------------------------------------------------
def score_category(items: list) -> dict:
"""Score a single category 0-100 using weighted item scores."""
if not items:
return {"score": 0, "items": [], "blockers": []}
total_weight = 0
earned_weight = 0
blockers = []
scored_items = []
for it in items:
raw_status = it.get("status", "not_started").strip().lower()
status = raw_status if raw_status in STATUS_WEIGHTS else "not_started"
weight = it.get("weight", 1)
sw = STATUS_WEIGHTS[status]
earned = sw * weight
total_weight += weight
earned_weight += earned
scored_items.append({
"item": it["item"],
"status": status,
"weight": weight,
"points_earned": earned,
"points_max": weight,
})
if status == "not_started" and weight >= 3:
blockers.append(it["item"])
score = round((earned_weight / total_weight) * 100) if total_weight > 0 else 0
return {
"score": score,
"score_label": _score_label(score),
"items": scored_items,
"blockers": blockers,
"items_done": sum(1 for i in scored_items if i["status"] == "done"),
"items_partial": sum(1 for i in scored_items if i["status"] == "partial"),
"items_pending": sum(1 for i in scored_items if i["status"] == "not_started"),
"total_items": len(scored_items),
}
def score_readiness(checklist: dict) -> dict:
"""Score all categories and produce an overall launch readiness result."""
categories = {}
all_scores = []
all_blockers = []
for cat, items in checklist.items():
result = score_category(items)
categories[cat] = result
all_scores.append(result["score"])
all_blockers.extend(result["blockers"])
overall = round(sum(all_scores) / len(all_scores)) if all_scores else 0
return {
"overall": {
"score": overall,
"score_label": _score_label(overall),
"launch_decision": _launch_decision(overall, all_blockers),
"blockers": all_blockers,
"generated_at": datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ"),
},
"categories": {
cat: {**CATEGORY_META.get(cat, {"emoji": "📋", "label": cat.title()}),
**res}
for cat, res in categories.items()
},
"action_plan": _action_plan(categories),
}
def _launch_decision(score: int, blockers: list) -> str:
if blockers:
return f"⛔ NOT READY — {len(blockers)} blocker(s) must be resolved before launch."
if score >= 80:
return "✅ LAUNCH READY — all categories are in good shape."
if score >= 60:
return "🟡 CONDITIONAL — address partial items but launch is defensible."
if score >= 40:
return "🟠 CAUTION — significant gaps; soft launch / waitlist recommended."
return "🔴 NOT READY — major preparation required across multiple areas."
def _action_plan(categories: dict) -> list:
"""Build a prioritised action list: blockers first, then by score ascending."""
actions = []
for cat, res in categories.items():
label = CATEGORY_META.get(cat, {}).get("label", cat.title())
for bl in res.get("blockers", []):
actions.append({
"priority": "🚨 BLOCKER",
"category": label,
"action": bl,
})
for cat, res in sorted(categories.items(), key=lambda x: x[1]["score"]):
label = CATEGORY_META.get(cat, {}).get("label", cat.title())
for it in res.get("items", []):
if it["status"] == "partial":
actions.append({
"priority": "⚠️ PARTIAL",
"category": label,
"action": f"Complete: {it['item']}",
})
return actions[:15] # top 15 actions
def _score_label(s: int) -> str:
if s >= 90: return "Excellent"
if s >= 75: return "Good"
if s >= 60: return "Fair"
if s >= 40: return "Poor"
return "Critical"
# ---------------------------------------------------------------------------
# Pretty-print
# ---------------------------------------------------------------------------
def pretty_print(result: dict) -> None:
ov = result["overall"]
print("\n" + "=" * 65)
print(" 🚀 LAUNCH READINESS SCORER")
print("=" * 65)
print(f"\n Overall Score : {ov['score']}/100 ({ov['score_label']})")
print(f" Launch Decision : {ov['launch_decision']}")
if ov["blockers"]:
print(f"\n 🚨 BLOCKERS ({len(ov['blockers'])}):")
for b in ov["blockers"]:
print(f" • {b}")
print(f"\n{'─'*65}")
print(f" {'CATEGORY':<30} {'SCORE':>6} {'DONE':>5} {'PARTIAL':>7} {'PENDING':>7}")
print(f"{'─'*65}")
for cat, res in result["categories"].items():
bar = "█" * (res["score"] // 10) + "░" * (10 - res["score"] // 10)
print(f" {res['emoji']} {res['label']:<27} {res['score']:>5}/100 "
f"{res['items_done']:>5} {res['items_partial']:>7} {res['items_pending']:>7} {bar}")
print(f"\n{'─'*65}")
print(f" 🗂 CATEGORY DETAILS\n")
for cat, res in result["categories"].items():
print(f" {res['emoji']} {res['label']} — {res['score']}/100 ({res['score_label']})")
for it in res["items"]:
icon = {"done": "✅", "partial": "🔶", "not_started": "⬜"}.get(it["status"], "⬜")
print(f" {icon} [{it['status']:<11}] (w={it['weight']}) {it['item']}")
print()
ap = result["action_plan"]
if ap:
print(f" 📋 ACTION PLAN (top {len(ap)} items)\n")
for i, a in enumerate(ap, 1):
print(f" {i:>2}. {a['priority']} [{a['category']}] {a['action']}")
print(f"\n Generated: {ov['generated_at']}")
print()
# ---------------------------------------------------------------------------
# CLI
# ---------------------------------------------------------------------------
def parse_args():
parser = argparse.ArgumentParser(
description="Score product launch readiness across categories (stdlib only).",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("--checklist", type=str, default=None,
help="Path to JSON checklist file")
parser.add_argument("--json", action="store_true",
help="Output results as JSON")
parser.add_argument("--export-template", action="store_true",
help="Print the default checklist template as JSON and exit")
return parser.parse_args()
def main():
args = parse_args()
if args.export_template:
print(json.dumps(DEFAULT_CHECKLIST, indent=2))
return
if args.checklist:
with open(args.checklist) as f:
checklist = json.load(f)
else:
print("🔬 DEMO MODE — using embedded sample checklist\n")
checklist = DEFAULT_CHECKLIST
result = score_readiness(checklist)
if args.json:
print(json.dumps(result, indent=2))
else:
pretty_print(result)
if __name__ == "__main__":
main()
Triển khai và duy trì hệ thống quản lý chất lượng ISO 13485 cho thiết bị y tế: thiết kế QMS, kiểm soát tài liệu, đánh giá nội bộ, CAPA và hỗ trợ chứng nhận.
---
name: "quality-manager-qms-iso13485"
description: ISO 13485 Quality Management System implementation and maintenance for medical device organizations. Provides QMS design, documentation control, internal auditing, CAPA management, and certification support. Use when working with medical device quality systems, preparing for ISO 13485 audits, managing regulatory compliance documentation, setting up corrective actions, or building audit preparation programs. Useful for quality management, audit preparation, regulatory compliance, medical device documentation, and corrective action workflows.
triggers:
- ISO 13485
- QMS implementation
- quality management system
- document control
- internal audit
- management review
- quality manual
- CAPA process
- process validation
- design control
- supplier qualification
- quality records
---
# Quality Manager - QMS ISO 13485 Specialist
ISO 13485:2016 Quality Management System implementation, maintenance, and certification support for medical device organizations.
---
## Table of Contents
- [QMS Implementation Workflow](#qms-implementation-workflow)
- [Document Control Workflow](#document-control-workflow)
- [Internal Audit Workflow](#internal-audit-workflow)
- [Process Validation Workflow](#process-validation-workflow)
- [Supplier Qualification Workflow](#supplier-qualification-workflow)
- [QMS Process Reference](#qms-process-reference)
- [Decision Frameworks](#decision-frameworks)
- [Tools and References](#tools-and-references)
---
## QMS Implementation Workflow
Implement ISO 13485:2016 compliant quality management system from gap analysis through certification.
### Workflow: Initial QMS Implementation
1. Conduct gap analysis against ISO 13485:2016 requirements
2. Document current state vs. required state for each clause
3. Prioritize gaps by:
- Regulatory criticality
- Risk to product safety
- Resource requirements
4. Develop implementation roadmap with milestones
5. Establish Quality Manual per Clause 4.2.2:
- QMS scope with justified exclusions
- Process interactions
- Procedure references
6. Create required documented procedures — see [Mandatory Documented Procedures](#quick-reference-mandatory-documented-procedures) for the full list
7. Deploy processes with training
8. **Validation:** Gap analysis complete; Quality Manual approved; all required procedures documented and trained
> Use the Gap Analysis Matrix template in [qms-process-templates.md](references/qms-process-templates.md) to document clause-by-clause current state, gaps, priority, and actions.
### QMS Structure
| Level | Document Type | Example |
|-------|---------------|---------|
| 1 | Quality Manual | QM-001 |
| 2 | Procedures | SOP-02-001 |
| 3 | Work Instructions | WI-06-012 |
| 4 | Records | Training records |
---
## Document Control Workflow
Establish and maintain document control per ISO 13485 Clause 4.2.3.
### Workflow: Document Creation and Approval
1. Identify need for new document or revision
2. Assign document number per numbering convention:
- Format: `[TYPE]-[AREA]-[SEQUENCE]-[REV]`
- Example: `SOP-02-001-01`
3. Draft document using approved template
4. Route for review to subject matter experts
5. Collect and address review comments
6. Obtain required approvals based on document type
7. Update Document Master List
8. **Validation:** Document numbered correctly; all reviewers signed; Master List updated
### Document Numbering Convention
| Prefix | Document Type | Approval Authority |
|--------|---------------|-------------------|
| QM | Quality Manual | Management Rep + CEO |
| POL | Policy | Department Head + QA |
| SOP | Procedure | Process Owner + QA |
| WI | Work Instruction | Supervisor + QA |
| TF | Template/Form | Process Owner |
| SPEC | Specification | Engineering + QA |
### Area Codes
| Code | Area | Examples |
|------|------|----------|
| 01 | Quality Management | Quality Manual, policy |
| 02 | Document Control | This procedure |
| 03 | Training | Competency procedures |
| 04 | Design | Design control |
| 05 | Purchasing | Supplier management |
| 06 | Production | Manufacturing |
| 07 | Quality Control | Inspection, testing |
| 08 | CAPA | Corrective actions |
### Document Change Control
| Change Type | Approval Level | Examples |
|-------------|----------------|----------|
| Administrative | Document Control | Typos, formatting |
| Minor | Process Owner + QA | Clarifications |
| Major | Full review cycle | Process changes |
| Emergency | Expedited + retrospective | Safety issues |
### Document Review Schedule
| Document Type | Review Period | Trigger for Unscheduled Review |
|---------------|---------------|-------------------------------|
| Quality Manual | Annual | Organizational change |
| Procedures | Annual | Audit finding, regulation change |
| Work Instructions | 2 years | Process change |
| Forms | 2 years | User feedback |
---
## Internal Audit Workflow
Plan and execute internal audits per ISO 13485 Clause 8.2.4.
### Workflow: Annual Audit Program
1. Identify processes and areas requiring audit coverage
2. Assess risk factors for audit frequency:
- Previous audit findings
- Regulatory changes
- Process changes
- Complaint trends
3. Assign qualified auditors (independent of area audited)
4. Develop annual audit schedule
5. Obtain management approval
6. Communicate schedule to process owners
7. Track completion and reschedule as needed
8. **Validation:** All processes covered; auditors qualified and independent; schedule approved
> Use the Audit Program Template in [qms-process-templates.md](references/qms-process-templates.md) to schedule audits by clause and quarter across processes such as Document Control (4.2.3/4.2.4), Management Review (5.6), Design Control (7.3), Production (7.5), and CAPA (8.5.2/8.5.3).
### Workflow: Individual Audit Execution
1. Prepare audit plan with scope, criteria, and schedule
2. Notify auditee minimum 1 week prior
3. Review procedures and previous audit results
4. Prepare audit checklist
5. Conduct opening meeting
6. Collect evidence through:
- Document review
- Record sampling
- Process observation
- Personnel interviews
7. Classify findings:
- Major NC: Absence or breakdown of system
- Minor NC: Single lapse or deviation
- Observation: Risk of future NC
8. Conduct closing meeting
9. Issue audit report within 5 business days
10. **Validation:** All checklist items addressed; findings supported by evidence; report distributed
### Auditor Qualification Requirements
| Criterion | Requirement |
|-----------|-------------|
| Training | ISO 13485 awareness + auditor training |
| Experience | Minimum 1 audit as observer |
| Independence | Not auditing own work area |
| Competence | Understanding of audited process |
### Finding Classification Guide
| Classification | Criteria | Response Time |
|----------------|----------|---------------|
| Major NC | System absence, total breakdown, regulatory violation | 30 days for CAPA |
| Minor NC | Single instance, partial compliance | 60 days for CAPA |
| Observation | Potential risk, improvement opportunity | Track in next audit |
---
## Process Validation Workflow
Validate special processes per ISO 13485 Clause 7.5.6.
### Workflow: Process Validation Protocol
1. Identify processes requiring validation:
- Output cannot be verified by inspection
- Deficiencies appear only in use
- Sterilization, welding, sealing, software
2. Form validation team with subject matter experts
3. Write validation protocol including:
- Process description and parameters
- Equipment and materials
- Acceptance criteria
- Statistical approach
4. Execute IQ: verify equipment installed correctly and document specifications
5. Execute OQ: test parameter ranges and verify process control
6. Execute PQ: run production conditions and verify output meets requirements
7. Write validation report with conclusions
8. **Validation:** IQ/OQ/PQ complete; acceptance criteria met; validation report approved
### Validation Documentation Requirements
| Phase | Content | Evidence |
|-------|---------|----------|
| Protocol | Objectives, methods, criteria | Approved protocol |
| IQ | Equipment verification | Installation records |
| OQ | Parameter verification | Test results |
| PQ | Performance verification | Production data |
| Report | Summary, conclusions | Approval signatures |
### Revalidation Triggers
| Trigger | Action Required |
|---------|-----------------|
| Equipment change | Assess impact, revalidate affected phases |
| Parameter change | OQ and PQ minimum |
| Material change | Assess impact, PQ minimum |
| Process failure | Full revalidation |
| Periodic | Per validation schedule (typically 3 years) |
### Special Process Examples
| Process | Validation Standard | Critical Parameters |
|---------|--------------------|--------------------|
| EO Sterilization | ISO 11135 | Temperature, humidity, EO concentration, time |
| Steam Sterilization | ISO 17665 | Temperature, pressure, time |
| Radiation Sterilization | ISO 11137 | Dose, dose uniformity |
| Sealing | Internal | Temperature, pressure, dwell time |
| Welding | ISO 11607 | Heat, pressure, speed |
---
## Supplier Qualification Workflow
Evaluate and approve suppliers per ISO 13485 Clause 7.4.
### Workflow: New Supplier Qualification
1. Identify supplier category:
- Category A: Critical (affects safety/performance)
- Category B: Major (affects quality)
- Category C: Minor (indirect impact)
2. Request supplier information:
- Quality certifications
- Product specifications
- Quality history
3. Evaluate supplier based on:
- Quality system (ISO certification)
- Technical capability
- Quality history
- Financial stability
4. For Category A suppliers:
- Conduct on-site audit
- Require quality agreement
5. Calculate qualification score
6. Make approval decision:
- >80: Approved
- 60-80: Conditional approval
- <60: Not approved
7. Add to Approved Supplier List
8. **Validation:** Evaluation criteria scored; qualification records complete; supplier categorized
### Supplier Evaluation Criteria
| Criterion | Weight | Scoring |
|-----------|--------|---------|
| Quality System | 30% | ISO 13485=30, ISO 9001=20, Documented=10, None=0 |
| Quality History | 25% | Reject rate: <1%=25, 1-3%=15, >3%=0 |
| Delivery | 20% | On-time: >95%=20, 90-95%=10, <90%=0 |
| Technical Capability | 15% | Exceeds=15, Meets=10, Marginal=5 |
| Financial Stability | 10% | Strong=10, Adequate=5, Questionable=0 |
### Supplier Category Requirements
| Category | Qualification | Monitoring | Agreement |
|----------|---------------|------------|-----------|
| A - Critical | On-site audit | Annual review | Quality agreement |
| B - Major | Questionnaire | Semi-annual review | Quality requirements |
| C - Minor | Assessment | Issue-based | Standard terms |
### Supplier Performance Metrics
| Metric | Target | Calculation |
|--------|--------|-------------|
| Accept Rate | >98% | (Accepted lots / Total lots) × 100 |
| On-Time Delivery | >95% | (On-time / Total orders) × 100 |
| Response Time | <5 days | Average days to resolve issues |
| Documentation | 100% | (Complete CoCs / Required CoCs) × 100 |
---
## QMS Process Reference
For detailed requirements and audit questions for each ISO 13485:2016 clause, see [iso13485-clause-requirements.md](references/iso13485-clause-requirements.md).
### Management Review Required Inputs (Clause 5.6.2)
| Input | Source | Prepared By |
|-------|--------|-------------|
| Audit results | Internal and external audits | QA Manager |
| Customer feedback | Complaints, surveys | Customer Quality |
| Process performance | Process metrics | Process Owners |
| Product conformity | Inspection data, NCs | QC Manager |
| CAPA status | CAPA system | CAPA Officer |
| Previous actions | Prior review records | QMR |
| Changes affecting QMS | Regulatory, organizational | RA Manager |
| Recommendations | All sources | All Managers |
### Record Retention Requirements
| Record Type | Minimum Retention | Regulatory Basis |
|-------------|-------------------|------------------|
| Device Master Record | Life of device + 2 years | 21 CFR 820.181 |
| Device History Record | Life of device + 2 years | 21 CFR 820.184 |
| Design History File | Life of device + 2 years | 21 CFR 820.30 |
| Complaint Records | Life of device + 2 years | 21 CFR 820.198 |
| Training Records | Employment + 3 years | Best practice |
| Audit Records | 7 years | Best practice |
| CAPA Records | 7 years | Best practice |
| Calibration Records | Equipment life + 2 years | Best practice |
---
## Decision Frameworks
### Exclusion Justification (Clause 4.2.2)
| Clause | Permissible Exclusion | Justification Required |
|--------|----------------------|------------------------|
| 6.4.2 | Contamination control | Product not affected by contamination |
| 7.3 | Design and development | Organization does not design products |
| 7.5.2 | Product cleanliness | No cleanliness requirements |
| 7.5.3 | Installation | No installation activities |
| 7.5.4 | Servicing | No servicing activities |
| 7.5.5 | Sterile products | No sterile products |
### Nonconformity Disposition Decision Tree
```
Nonconforming Product Identified
│
▼
Can it be reworked?
│
Yes──┴──No
│ │
▼ ▼
Is rework Can it be used
procedure as is?
available? │
│ Yes──┴──No
Yes─┴─No │ │
│ │ ▼ ▼
▼ ▼ Concession Scrap or
Rework Create approval return to
per SOP rework needed? supplier
procedure │
Yes─┴─No
│ │
▼ ▼
Customer Use as is
approval with MRB
approval
```
### CAPA Initiation Criteria
| Source | Automatic CAPA | Evaluate for CAPA |
|--------|----------------|-------------------|
| Customer complaint | Safety-related | All others |
| External audit | Major NC | Minor NC |
| Internal audit | Major NC | Repeat minor NC |
| Product NC | Field failure | Trend exceeds threshold |
| Process deviation | Safety impact | Repeated deviations |
---
## Tools and References
### Scripts
| Tool | Purpose | Usage |
|------|---------|-------|
| [qms_audit_checklist.py](scripts/qms_audit_checklist.py) | Generate audit checklists by clause or process | `python qms_audit_checklist.py --help` |
**Audit Checklist Generator Features:**
- Generate clause-specific checklists (e.g., `--clause 7.3`)
- Generate process-based checklists (e.g., `--process design-control`)
- Full system audit checklist (`--audit-type system`)
- Text or JSON output formats
- Interactive mode for guided selection
### References
| Document | Content |
|----------|---------|
| [iso13485-clause-requirements.md](references/iso13485-clause-requirements.md) | Detailed requirements for each ISO 13485:2016 clause with audit questions |
| [qms-process-templates.md](references/qms-process-templates.md) | Ready-to-use templates for gap analysis, audit program, document control, CAPA, supplier, training |
### Quick Reference: Mandatory Documented Procedures
| Procedure | Clause | Key Elements |
|-----------|--------|--------------|
| Document Control | 4.2.3 | Approval, distribution, obsolete control |
| Record Control | 4.2.4 | Identification, retention, disposal |
| Internal Audit | 8.2.4 | Program, auditor qualification, reporting |
| NC Product Control | 8.3 | Identification, segregation, disposition |
| Corrective Action | 8.5.2 | Root cause, implementation, verification |
| Preventive Action | 8.5.3 | Risk identification, implementation |
---
## Related Skills
| Skill | Integration Point |
|-------|-------------------|
| [quality-manager-qmr](../quality-manager-qmr/) | Management review, quality policy |
| [capa-officer](../capa-officer/) | CAPA system management |
| [qms-audit-expert](../qms-audit-expert/) | Advanced audit techniques |
| [quality-documentation-manager](../quality-documentation-manager/) | DHF, DMR, DHR management |
| [risk-management-specialist](../risk-management-specialist/) | ISO 14971 integration |
FILE:references/iso13485-clause-requirements.md
# ISO 13485:2016 Clause Requirements
Detailed requirements for each ISO 13485:2016 clause with implementation guidance and audit criteria.
---
## Table of Contents
- [Clause 4: Quality Management System](#clause-4-quality-management-system)
- [Clause 5: Management Responsibility](#clause-5-management-responsibility)
- [Clause 6: Resource Management](#clause-6-resource-management)
- [Clause 7: Product Realization](#clause-7-product-realization)
- [Clause 8: Measurement, Analysis and Improvement](#clause-8-measurement-analysis-and-improvement)
---
## Clause 4: Quality Management System
### 4.1 General Requirements
| Requirement | Implementation | Evidence |
|-------------|----------------|----------|
| Determine processes needed | Process map showing QMS processes | Documented process map |
| Determine sequence and interaction | Process interaction diagram | Cross-reference matrix |
| Determine criteria for operation | Process metrics and acceptance criteria | Documented criteria per process |
| Ensure resources available | Resource allocation per process | Training records, equipment logs |
| Monitor, measure, analyze | Process monitoring procedures | Trend data, performance reports |
| Implement actions for results | Improvement projects, CAPAs | Action records with verification |
| Document processes | Procedures, work instructions | Controlled document list |
**Audit Questions:**
- How are QMS processes identified and documented?
- What criteria determine if processes are operating effectively?
- How is outsourced process control demonstrated?
### 4.2 Documentation Requirements
#### 4.2.1 General
| Document Type | Requirement | Retention |
|---------------|-------------|-----------|
| Quality Policy | Documented statement of commitment | Life of QMS |
| Quality Objectives | Measurable objectives at relevant functions | Life of QMS |
| Quality Manual | QMS scope and processes | Current version |
| Documented Procedures | Required by standard | Life of QMS + 2 years |
| Records | Evidence of conformity | As defined per record type |
#### 4.2.2 Quality Manual
**Required Content:**
1. Scope of QMS including justification for exclusions
2. Documented procedures or reference to them
3. Description of process interactions
**Quality Manual Template Structure:**
```
QUALITY MANUAL
1. Company Overview
1.1 Company Description
1.2 Scope of QMS
1.3 Exclusions and Justification
2. Quality Policy
3. Quality Objectives
4. QMS Structure
4.1 Process Map
4.2 Process Interactions
4.3 Organizational Chart
5. Procedure References
5.1 Document Control
5.2 Record Control
5.3 Management Review
5.4 Internal Audit
5.5 Nonconformity Control
5.6 CAPA
6. Appendices
6.1 Glossary
6.2 Regulatory Cross-Reference
```
#### 4.2.3 Control of Documents
| Control Element | Requirement | Method |
|-----------------|-------------|--------|
| Approval | Adequate prior to issue | Signature/electronic approval |
| Review and update | Re-approval after changes | Periodic review process |
| Identification of changes | Change history visible | Revision log in document |
| Revision status | Current revision identifiable | Document master list |
| Legibility | Readable and identifiable | Format standards |
| External documents | Identified and controlled | Incoming document log |
| Obsolete documents | Prevented from unintended use | Archive system |
**Document Numbering Convention:**
```
[TYPE]-[AREA]-[SEQUENCE]-[REV]
TYPE:
QM = Quality Manual
SOP = Standard Operating Procedure
WI = Work Instruction
TF = Template/Form
POL = Policy
AREA:
01 = Quality Management
02 = Document Control
03 = Training
04 = Design
05 = Purchasing
06 = Production
07 = Quality Control
08 = CAPA
Example: SOP-02-001-03 = Document Control SOP, Revision 03
```
#### 4.2.4 Control of Records
| Record Category | Minimum Retention | Basis |
|-----------------|-------------------|-------|
| Device Master Record | Life of device + 2 years | 21 CFR 820.181 |
| Device History Record | Life of device + 2 years | 21 CFR 820.184 |
| Design History File | Life of device + 2 years | 21 CFR 820.30 |
| Training Records | Employment + 3 years | Best practice |
| Audit Records | 7 years | Best practice |
| Complaint Records | Life of device + 2 years | 21 CFR 820.198 |
| CAPA Records | 7 years | Best practice |
| Calibration Records | Equipment life + 2 years | Best practice |
| Supplier Records | Relationship + 3 years | Best practice |
---
## Clause 5: Management Responsibility
### 5.1 Management Commitment
| Commitment Area | Evidence Required |
|-----------------|-------------------|
| Communicate importance of requirements | Meeting minutes, communications |
| Establish quality policy | Documented policy, communication records |
| Ensure quality objectives established | Objective documentation |
| Conduct management reviews | Management review records |
| Ensure resources available | Budget records, staffing records |
### 5.2 Customer Focus
| Requirement | Implementation | Verification |
|-------------|----------------|--------------|
| Customer requirements determined | Requirements review process | Contract review records |
| Requirements met | Process controls | Inspection and test data |
| Regulatory requirements met | Regulatory register | Compliance assessments |
| Customer satisfaction enhanced | Feedback collection | Satisfaction data, complaints |
### 5.3 Quality Policy
**Policy Requirements:**
- Appropriate to organization purpose
- Commitment to compliance and effectiveness
- Framework for quality objectives
- Communicated and understood
- Reviewed for continuing suitability
**Sample Quality Policy Elements:**
```
[Company Name] Quality Policy
We are committed to:
- Designing and manufacturing safe, effective medical devices
- Meeting customer and regulatory requirements
- Maintaining an effective Quality Management System
- Continuously improving our processes and products
- Providing resources for QMS effectiveness
Signed: [Executive]
Date: [Date]
Review Date: [Annual]
```
### 5.4 Planning
#### 5.4.1 Quality Objectives
| Objective Criteria | Requirement |
|-------------------|-------------|
| Measurable | Quantifiable targets |
| Consistent with policy | Aligned to policy statements |
| Relevant functions | Cascaded to departments |
| Includes compliance | Regulatory and customer requirements |
| Includes product conformity | Product-related targets |
**Objective Template:**
```
QUALITY OBJECTIVE [Year]
Objective: [Statement]
Metric: [How measured]
Target: [Specific value]
Baseline: [Current performance]
Owner: [Responsible person]
Due Date: [Target date]
Reporting: [Frequency]
```
#### 5.4.2 Quality Management System Planning
**Planning Requirements:**
- QMS meets general requirements (4.1)
- QMS meets quality objectives (5.4.1)
- Integrity maintained during changes
### 5.5 Responsibility, Authority and Communication
#### 5.5.1 Responsibility and Authority
| Role | Responsibilities | Authority |
|------|-----------------|-----------|
| Top Management | QMS commitment, resources, policy | Budget, staffing, strategic decisions |
| Quality Manager | QMS implementation, reporting | Document approval, CAPA approval |
| Department Managers | Process ownership, resources | Process changes, training |
| Process Owners | Process performance, improvements | Procedure changes within scope |
#### 5.5.2 Management Representative
| QMR Responsibility | Activities |
|-------------------|------------|
| QMS establishment | Process definition, documentation |
| QMS implementation | Training, deployment, monitoring |
| QMS maintenance | Audits, reviews, improvements |
| Reporting to top management | Performance reports, recommendations |
| Awareness promotion | Training, communications |
#### 5.5.3 Internal Communication
| Communication Type | Method | Frequency |
|-------------------|--------|-----------|
| Policy and objectives | Posting, training | Annual and on change |
| QMS performance | Dashboards, reports | Monthly |
| Changes affecting quality | Email, meetings | As needed |
| Audit results | Reports, presentations | Per audit |
### 5.6 Management Review
#### 5.6.1 General
| Requirement | Specification |
|-------------|---------------|
| Frequency | Planned intervals (typically quarterly/semi-annually) |
| Purpose | Assess QMS suitability, adequacy, effectiveness |
| Records | Documented meeting records |
#### 5.6.2 Review Input
| Input | Source | Responsible |
|-------|--------|-------------|
| Audit results | Internal/external audits | QA Manager |
| Customer feedback | Complaints, surveys | Customer Quality |
| Process performance | Metrics, yields | Process Owners |
| Product conformity | Inspection data | QC Manager |
| CAPA status | CAPA system | CAPA Officer |
| Previous actions | Prior review records | QMR |
| Changes affecting QMS | Regulatory, organizational | RA, HR |
| Recommendations | All sources | All Managers |
#### 5.6.3 Review Output
| Output | Documentation |
|--------|---------------|
| QMS improvement decisions | Action items with owners |
| Process improvements | Project charters |
| Resource needs | Resource allocation plans |
| Product improvements | Design change requests |
---
## Clause 6: Resource Management
### 6.1 Provision of Resources
**Resource Categories:**
- Human resources (competent personnel)
- Infrastructure (facilities, equipment, software)
- Work environment (environmental conditions)
### 6.2 Human Resources
| Requirement | Implementation | Evidence |
|-------------|----------------|----------|
| Competence determined | Job descriptions, competency matrix | Role definitions |
| Training provided | Training programs | Training records |
| Effectiveness evaluated | Assessments, observations | Competency verification |
| Awareness ensured | Orientation, ongoing training | Acknowledgments |
| Records maintained | Training database | Training files |
**Competency Matrix Template:**
```
COMPETENCY MATRIX
Role: [Job Title]
Department: [Department]
Required Competencies:
| Competency | Requirement Level | Method | Verification |
|------------|------------------|--------|--------------|
| [Skill 1] | Expert/Proficient/Basic | Training/OJT | Assessment |
| [Skill 2] | Expert/Proficient/Basic | Training/OJT | Assessment |
Training Requirements:
| Training | Initial | Refresher | Record |
|----------|---------|-----------|--------|
| ISO 13485 Awareness | Yes | Annual | TR-001 |
| Document Control | Yes | On Change | TR-002 |
```
### 6.3 Infrastructure
| Infrastructure Type | Control Requirements |
|--------------------|---------------------|
| Buildings and workspace | Cleaning, maintenance schedules |
| Process equipment | Maintenance, calibration |
| Supporting services | Utilities, IT systems |
| Information systems | Backup, security, validation |
### 6.4 Work Environment and Contamination Control
| Environment Factor | Control Method | Monitoring |
|-------------------|----------------|------------|
| Temperature | HVAC control | Continuous logging |
| Humidity | HVAC control | Continuous logging |
| Cleanliness | Cleaning procedures | Particle counts |
| Lighting | Lux levels | Periodic verification |
| ESD protection | Grounding, ionization | Periodic testing |
---
## Clause 7: Product Realization
### 7.1 Planning of Product Realization
| Planning Element | Content |
|-----------------|---------|
| Quality objectives for product | Product-specific quality targets |
| Processes and documentation | Process flow, required documents |
| Verification and validation | Test methods, acceptance criteria |
| Records | Required quality records |
| Risk management | Per ISO 14971 |
### 7.2 Customer-Related Processes
#### 7.2.1 Determination of Requirements
| Requirement Type | Source |
|-----------------|--------|
| Customer-specified | Contract, purchase order |
| Not stated but necessary | Intended use analysis |
| Regulatory | Applicable standards, regulations |
| Organization-defined | Internal specifications |
#### 7.2.2 Review of Requirements
| Review Element | Verification |
|----------------|--------------|
| Requirements defined | Complete specification |
| Differences resolved | Documented resolution |
| Ability to meet | Feasibility assessment |
| Risk management | Initial risk assessment |
#### 7.2.3 Communication
| Communication Type | Method |
|-------------------|--------|
| Product information | Catalogs, IFU |
| Inquiries and orders | Sales process |
| Feedback and complaints | Customer feedback system |
| Advisory notices | Field safety notices |
### 7.3 Design and Development
| Stage | Clause | Requirements |
|-------|--------|--------------|
| Planning | 7.3.2 | Stages, reviews, responsibilities |
| Inputs | 7.3.3 | Functional, performance, regulatory |
| Outputs | 7.3.4 | Meet inputs, acceptance criteria |
| Review | 7.3.5 | Evaluate ability to meet requirements |
| Verification | 7.3.6 | Outputs meet inputs |
| Validation | 7.3.7 | Product meets intended use |
| Transfer | 7.3.8 | Verified before production |
| Changes | 7.3.9 | Controlled, reviewed, verified |
### 7.4 Purchasing
#### 7.4.1 Purchasing Process
| Control Element | Implementation |
|-----------------|----------------|
| Supplier evaluation | Qualification procedure |
| Selection criteria | Quality, delivery, cost |
| Monitoring | Performance metrics |
| Re-evaluation | Periodic review |
**Supplier Classification:**
```
Category A: Critical - Affects product safety/performance
- Full qualification audit
- Annual performance review
- Quality agreement required
Category B: Major - Affects product quality
- Qualification questionnaire
- Periodic performance review
- Quality requirements communicated
Category C: Minor - Indirect impact
- Initial assessment
- Issue-based review
- Standard terms
```
#### 7.4.2 Purchasing Information
| Information Required | Purpose |
|---------------------|---------|
| Product specifications | Clear requirements |
| QMS requirements | Supplier system expectations |
| Personnel competence | Where applicable |
| Approval requirements | Where applicable |
#### 7.4.3 Verification of Purchased Product
| Verification Method | Application |
|--------------------|-------------|
| Incoming inspection | Standard verification |
| Source inspection | Critical items |
| Certificate of Conformance | Documented evidence |
| Certificate of Analysis | Material verification |
### 7.5 Production and Service Provision
#### 7.5.1 Control of Production and Service Provision
| Control Element | Implementation |
|-----------------|----------------|
| Product information | Specifications, drawings |
| Work instructions | Where necessary |
| Suitable equipment | Qualified equipment |
| Monitoring devices | Calibrated instruments |
| Implementation of monitoring | Inspections, tests |
| Defined processes | Process parameters |
| Labeling and packaging | Per requirements |
#### 7.5.2 Cleanliness of Product
| Cleanliness Control | Method |
|--------------------|--------|
| Product cleaning | Validated procedures |
| Contamination prevention | Controlled environment |
| Process aids | Qualified, controlled |
#### 7.5.3 Installation Activities
| Requirement | Implementation |
|-------------|----------------|
| Installation requirements | Documented instructions |
| Acceptance criteria | Defined criteria |
| Records | Installation records |
#### 7.5.4 Servicing Activities
| Requirement | Implementation |
|-------------|----------------|
| Documented requirements | Service procedures |
| Reference materials | Service manuals |
| Measurement equipment | Calibrated |
| Records | Service records |
#### 7.5.5 Particular Requirements for Sterile Medical Devices
| Process | Control |
|---------|---------|
| Sterilization validation | Per ISO 11135/11137/17665 |
| Parameter control | Monitoring records |
| Sterile barrier | Validated packaging |
#### 7.5.6 Validation of Processes
| Validation Required When | Evidence |
|-------------------------|----------|
| Output cannot be verified | Validation protocol and report |
| Deficiencies appear only in use | Process capability data |
| Special processes | Qualified operators |
**Process Validation Elements:**
- Equipment qualification (IQ/OQ/PQ)
- Process parameters
- Monitoring methods
- Operator qualification
- Revalidation criteria
#### 7.5.7 Particular Requirements for Validation
| Requirement | Implementation |
|-------------|----------------|
| Documented procedures | Validation SOPs |
| Defined methods | Statistical methods |
| Acceptance criteria | Predefined criteria |
| Software validation | Where applicable |
| Revalidation | Change-triggered |
#### 7.5.8 Identification
| Identification Type | Method |
|--------------------|--------|
| Product | Labels, markings |
| Documentation | Document numbers |
| Unique Device Identification | UDI per regulation |
#### 7.5.9 Traceability
| Traceability Element | Record |
|---------------------|--------|
| Components | Lot/batch numbers |
| Materials | Certificates |
| Work environment | Environmental records |
| Measurement equipment | Calibration records |
| Personnel | Training records |
| Distribution | Shipping records |
#### 7.5.10 Customer Property
| Control | Implementation |
|---------|----------------|
| Identification | Marking, segregation |
| Verification | Incoming inspection |
| Protection | Storage conditions |
| Safeguarding | Security measures |
| Reporting | Loss/damage notification |
#### 7.5.11 Preservation of Product
| Preservation Element | Control |
|---------------------|---------|
| Identification | Labels, markings |
| Handling | Procedures |
| Packaging | Specifications |
| Storage | Conditions, FIFO |
| Protection | Environmental controls |
### 7.6 Control of Monitoring and Measuring Equipment
| Control Element | Implementation |
|-----------------|----------------|
| Calibration | At specified intervals |
| Adjustment | As needed |
| Identification | Calibration status |
| Safeguarding | Protection from damage |
| Software validation | Where applicable |
| Records | Calibration records |
---
## Clause 8: Measurement, Analysis and Improvement
### 8.1 General
**Monitoring and Measurement Requirements:**
- Demonstrate product conformity
- Ensure QMS conformity
- Maintain QMS effectiveness
### 8.2 Monitoring and Measurement
#### 8.2.1 Feedback
| Feedback Source | Collection Method |
|-----------------|-------------------|
| Customer complaints | Complaint system |
| Customer surveys | Periodic surveys |
| Field feedback | Service reports |
| Regulatory feedback | Inspection findings |
#### 8.2.2 Complaint Handling
| Process Step | Requirements |
|--------------|--------------|
| Receipt | Timely logging |
| Investigation | Root cause analysis |
| Corrective action | If warranted |
| Regulatory reporting | If required |
| Trend analysis | Aggregate review |
#### 8.2.3 Reporting to Regulatory Authorities
| Report Type | Trigger | Timeline |
|-------------|---------|----------|
| MDR (Medical Device Report) | Death/serious injury | 30 days (5 if awareness) |
| FSCA (Field Safety Corrective Action) | Safety issue | Without delay |
| Periodic Safety Update | Per regulation | Per schedule |
#### 8.2.4 Internal Audit
| Audit Element | Requirement |
|---------------|-------------|
| Planned program | Risk-based schedule |
| Criteria and scope | Defined per audit |
| Auditor selection | Independent, competent |
| Procedure | Documented process |
| Records | Audit reports, findings |
| Follow-up | CAPA, verification |
**Audit Program Template:**
```
ANNUAL INTERNAL AUDIT PROGRAM
Year: [Year]
| Audit # | Area/Process | Scope | Auditor | Planned Date | Status |
|---------|--------------|-------|---------|--------------|--------|
| IA-01 | Document Control | 4.2.3, 4.2.4 | [Name] | Q1 | |
| IA-02 | Design Control | 7.3 | [Name] | Q2 | |
| IA-03 | Production | 7.5 | [Name] | Q2 | |
| IA-04 | Purchasing | 7.4 | [Name] | Q3 | |
| IA-05 | CAPA | 8.5.2, 8.5.3 | [Name] | Q3 | |
| IA-06 | Management Review | 5.6 | [Name] | Q4 | |
Risk Considerations:
- Previous audit findings
- Regulatory changes
- Process changes
- Complaint trends
```
#### 8.2.5 Monitoring and Measurement of Processes
| Monitoring Type | Method |
|-----------------|--------|
| Process metrics | KPIs, trend analysis |
| Process audits | Internal audits |
| Process reviews | Management review |
#### 8.2.6 Monitoring and Measurement of Product
| Stage | Verification |
|-------|--------------|
| Incoming | Incoming inspection |
| In-process | In-process inspection |
| Final | Final inspection and test |
| Release | Authorized release |
### 8.3 Control of Nonconforming Product
| Control Element | Requirement |
|-----------------|-------------|
| Identification | Clear marking |
| Segregation | Physical separation |
| Documentation | NC record |
| Disposition | Use as is/rework/scrap/return |
| Concession | If accepted |
| Reinspection | After rework |
| Investigation | For detected after delivery |
**Nonconformity Disposition Options:**
```
1. Use As Is (Concession)
- Does not affect safety/performance
- Customer approval if applicable
- Documented justification
2. Rework
- Per approved procedure
- Reinspection required
- Records maintained
3. Scrap/Reject
- Physical destruction or marking
- Prevented from reentry
- Documented disposal
4. Return to Supplier
- Communication with supplier
- Replacement or credit
- Root cause if systemic
```
### 8.4 Analysis of Data
| Data Source | Analysis |
|-------------|----------|
| Feedback | Complaint trends, satisfaction |
| Nonconformity | Defect Pareto, trends |
| Process performance | Capability, trends |
| Supplier | Performance trends |
| Audit | Finding trends |
### 8.5 Improvement
#### 8.5.1 General
**Improvement Sources:**
- Quality policy
- Quality objectives
- Audit results
- Data analysis
- Corrective actions
- Preventive actions
- Management review
#### 8.5.2 Corrective Action
| Process Step | Requirement |
|--------------|-------------|
| Review nonconformity | Including complaints |
| Determine cause | Root cause analysis |
| Evaluate action need | Based on risk |
| Determine action | Proportionate to risk |
| Implement action | Execute plan |
| Document results | Records |
| Review effectiveness | Verification |
#### 8.5.3 Preventive Action
| Process Step | Requirement |
|--------------|-------------|
| Determine potential NC | Risk analysis, trends |
| Evaluate action need | Prevention opportunity |
| Determine action | Proportionate to risk |
| Implement action | Execute plan |
| Document results | Records |
| Review effectiveness | Verification |
FILE:references/qms-process-templates.md
# QMS Process Templates
Ready-to-use templates for ISO 13485 QMS processes including document control, internal audit, CAPA, and supplier management.
---
## Table of Contents
- [Document Control Templates](#document-control-templates)
- [Internal Audit Templates](#internal-audit-templates)
- [CAPA Templates](#capa-templates)
- [Supplier Management Templates](#supplier-management-templates)
- [Training Templates](#training-templates)
- [Nonconformity Templates](#nonconformity-templates)
---
## Document Control Templates
### Document Master List
```
DOCUMENT MASTER LIST
Organization: [Company Name]
Last Updated: [Date]
Maintained By: Document Control
| Doc # | Title | Rev | Effective Date | Status | Owner | Next Review |
|-------|-------|-----|----------------|--------|-------|-------------|
| QM-001 | Quality Manual | 03 | 2024-01-15 | Effective | QMR | 2025-01-15 |
| SOP-01-001 | Document Control | 04 | 2024-03-01 | Effective | QA Mgr | 2025-03-01 |
| SOP-01-002 | Record Control | 02 | 2024-02-01 | Effective | QA Mgr | 2025-02-01 |
| | | | | | | |
Status Values: Draft, Under Review, Effective, Obsolete
```
### Document Change Request
```
DOCUMENT CHANGE REQUEST
DCR Number: DCR-[YYYY]-[NNN]
Date Submitted: [Date]
Submitted By: [Name]
DOCUMENT INFORMATION
Document Number: [Number]
Document Title: [Title]
Current Revision: [Rev]
CHANGE REQUEST
Change Type: [ ] Administrative [ ] Minor [ ] Major [ ] Emergency
Requested Change: [Description of change]
Reason for Change:
[ ] Regulatory requirement
[ ] Process improvement
[ ] Nonconformity/CAPA
[ ] Organizational change
[ ] Error correction
[ ] Other: [Specify]
Justification: [Detailed justification]
IMPACT ASSESSMENT
Training Required: [ ] Yes [ ] No
If yes, who: [Roles/departments]
Other Documents Affected: [List]
Regulatory Filing Impact: [ ] Yes [ ] No
If yes, details: [Explain]
APPROVALS
Requested By: _________________ Date: _______
Document Owner: _________________ Date: _______
QA Approval: _________________ Date: _______
COMPLETION
New Revision: [Rev]
Effective Date: [Date]
Training Completed: [ ] Yes [ ] N/A
Distribution Completed: [ ] Yes
```
### Document Review Record
```
DOCUMENT REVIEW RECORD
Document Number: [Number]
Document Title: [Title]
Current Revision: [Rev]
Review Due Date: [Date]
Review Completed: [Date]
REVIEWERS
| Reviewer | Role | Review Date | Comments | Signature |
|----------|------|-------------|----------|-----------|
| [Name] | [Role] | [Date] | [Comments] | |
| [Name] | [Role] | [Date] | [Comments] | |
REVIEW OUTCOME
[ ] No changes required - document remains current
[ ] Minor changes required - see attached DCR
[ ] Major revision required - see attached DCR
[ ] Document obsolete - initiate retirement
NEXT REVIEW
Next Review Date: [Date]
APPROVAL
Review Completed By: _________________ Date: _______
Approved By: _________________ Date: _______
```
---
## Internal Audit Templates
### Annual Audit Schedule
```
INTERNAL AUDIT SCHEDULE
Year: [Year]
Prepared By: [Name]
Approved By: [Name]
Date: [Date]
AUDIT SCHEDULE
| Audit # | Process/Area | ISO Clauses | Lead Auditor | Q1 | Q2 | Q3 | Q4 |
|---------|--------------|-------------|--------------|----|----|----|----|
| IA-001 | Document Control | 4.2.3, 4.2.4 | [Name] | X | | | |
| IA-002 | Management Review | 5.6 | [Name] | | X | | |
| IA-003 | Training | 6.2 | [Name] | | X | | |
| IA-004 | Design Control | 7.3 | [Name] | | | X | |
| IA-005 | Purchasing | 7.4 | [Name] | | | X | |
| IA-006 | Production | 7.5 | [Name] | | | | X |
| IA-007 | CAPA | 8.5.2, 8.5.3 | [Name] | | | | X |
RISK FACTORS CONSIDERED
[ ] Previous audit findings
[ ] Regulatory changes
[ ] Process changes
[ ] Complaint trends
[ ] Management concerns
SCHEDULE REVISION LOG
| Rev | Date | Change | Approved By |
|-----|------|--------|-------------|
| 00 | [Date] | Initial release | [Name] |
```
### Audit Plan
```
INTERNAL AUDIT PLAN
Audit Number: IA-[YYYY]-[NNN]
Audit Date(s): [Date(s)]
Audit Type: [ ] Process [ ] System [ ] Product
SCOPE
Process/Area: [Name]
ISO 13485 Clauses: [List]
Regulatory Requirements: [If applicable]
Locations: [Locations]
AUDIT TEAM
Lead Auditor: [Name]
Auditor(s): [Names]
Observer(s): [If any]
AUDITEE CONTACTS
Process Owner: [Name]
Other Contacts: [Names]
AUDIT CRITERIA
- ISO 13485:2016
- [Organization procedures]
- [Regulatory requirements]
AUDIT SCHEDULE
| Time | Activity | Participants |
|------|----------|--------------|
| 09:00 | Opening meeting | All |
| 09:30 | Document review | Auditor, Doc Control |
| 10:30 | Process observation | Auditor, Operators |
| 12:00 | Lunch | |
| 13:00 | Record review | Auditor, QA |
| 14:30 | Interviews | Selected personnel |
| 15:30 | Auditor caucus | Audit team |
| 16:00 | Closing meeting | All |
PREPARATION CHECKLIST
[ ] Previous audit reports reviewed
[ ] Procedures reviewed
[ ] Checklist prepared
[ ] Auditees notified
[ ] Resources arranged
```
### Audit Checklist Template
```
INTERNAL AUDIT CHECKLIST
Audit Number: IA-[YYYY]-[NNN]
Process: [Process Name]
Auditor: [Name]
Date: [Date]
INSTRUCTIONS
C = Conforming, NC = Nonconforming, OBS = Observation, N/A = Not Applicable
CHECKLIST
| # | Requirement | Reference | Evidence Reviewed | Finding | Notes |
|---|-------------|-----------|-------------------|---------|-------|
| 1 | Is the procedure current and approved? | 4.2.3 | [Evidence] | C/NC/OBS | |
| 2 | Are personnel trained on the procedure? | 6.2 | [Evidence] | C/NC/OBS | |
| 3 | Are records maintained as required? | 4.2.4 | [Evidence] | C/NC/OBS | |
| 4 | Is the process performed as documented? | 4.1 | [Evidence] | C/NC/OBS | |
| 5 | Are monitoring activities performed? | 8.2.5 | [Evidence] | C/NC/OBS | |
INTERVIEWS CONDUCTED
| Person | Role | Topics Discussed |
|--------|------|------------------|
| [Name] | [Role] | [Topics] |
DOCUMENTS REVIEWED
| Document # | Title | Rev | Findings |
|------------|-------|-----|----------|
| [Number] | [Title] | [Rev] | [Findings] |
RECORDS SAMPLED
| Record Type | Sample Size | Sample IDs | Findings |
|-------------|-------------|------------|----------|
| [Type] | [N] | [IDs] | [Findings] |
AUDITOR SIGNATURE: _________________ Date: _______
```
### Audit Report
```
INTERNAL AUDIT REPORT
Audit Number: IA-[YYYY]-[NNN]
Report Date: [Date]
Report Status: [ ] Draft [ ] Final
AUDIT SUMMARY
Audit Date(s): [Date(s)]
Process/Area: [Name]
ISO Clauses Covered: [List]
Lead Auditor: [Name]
Audit Team: [Names]
AUDIT SCOPE
[Description of scope]
AUDIT OBJECTIVES
[List objectives]
EXECUTIVE SUMMARY
[Brief summary of audit results]
FINDINGS SUMMARY
| Type | Count |
|------|-------|
| Major Nonconformity | [N] |
| Minor Nonconformity | [N] |
| Observation | [N] |
| Opportunity for Improvement | [N] |
DETAILED FINDINGS
FINDING 1
Number: IA-[YYYY]-[NNN]-F01
Classification: [ ] Major NC [ ] Minor NC [ ] Observation [ ] OFI
Requirement: [Clause/requirement reference]
Statement: [Objective description of finding]
Evidence: [Evidence supporting finding]
Auditee Response Due: [Date]
[Repeat for each finding]
POSITIVE OBSERVATIONS
[List areas of good practice observed]
CONCLUSION
[Overall conclusion on process effectiveness]
REPORT DISTRIBUTION
| Name | Role | Date |
|------|------|------|
| [Name] | Process Owner | [Date] |
| [Name] | QA Manager | [Date] |
| [Name] | Management Rep | [Date] |
APPROVALS
Lead Auditor: _________________ Date: _______
QA Manager: _________________ Date: _______
```
---
## CAPA Templates
### CAPA Request Form
```
CORRECTIVE AND PREVENTIVE ACTION REQUEST
CAPA Number: CAPA-[YYYY]-[NNN]
Date Opened: [Date]
Initiated By: [Name]
CAPA TYPE
[ ] Corrective Action (response to existing nonconformity)
[ ] Preventive Action (prevent potential nonconformity)
SOURCE
[ ] Customer complaint: Reference #_______
[ ] Internal audit: Audit #_______
[ ] External audit: Audit #_______
[ ] Nonconformity: NC #_______
[ ] Process deviation
[ ] Management review action
[ ] Trend analysis
[ ] Risk assessment
[ ] Other: _______
CLASSIFICATION
Severity: [ ] Critical [ ] Major [ ] Minor
Regulatory Reportable: [ ] Yes [ ] No
PROBLEM DESCRIPTION
[Detailed description of the problem or potential problem]
IMMEDIATE CONTAINMENT (if applicable)
Actions Taken: [Description]
Date: [Date]
Responsible: [Name]
ASSIGNMENT
Process Owner: [Name]
CAPA Owner: [Name]
Due Date for Root Cause: [Date]
Target Closure Date: [Date]
APPROVAL TO PROCEED
Approved By: _________________ Date: _______
```
### Root Cause Analysis Record
```
ROOT CAUSE ANALYSIS
CAPA Number: CAPA-[YYYY]-[NNN]
Analysis Date: [Date]
Analyst: [Name]
PROBLEM STATEMENT
[Clear, specific statement of the problem]
INVESTIGATION TEAM
| Name | Role | Contribution |
|------|------|--------------|
| [Name] | [Role] | [Area of expertise] |
INVESTIGATION METHOD
[ ] 5 Why Analysis
[ ] Fishbone Diagram
[ ] Fault Tree Analysis
[ ] Human Factors Analysis
[ ] Other: _______
INVESTIGATION DETAILS
5 WHY ANALYSIS
Why 1: [First why]
Answer: [Answer]
Why 2: [Second why based on answer]
Answer: [Answer]
Why 3: [Third why based on answer]
Answer: [Answer]
Why 4: [Fourth why based on answer]
Answer: [Answer]
Why 5: [Fifth why based on answer]
Answer: [Answer]
ROOT CAUSE STATEMENT
[Clear statement of identified root cause]
ROOT CAUSE CATEGORY
[ ] Process/Procedure
[ ] Training/Competency
[ ] Equipment/Material
[ ] Design
[ ] Human Error
[ ] Communication
[ ] Management System
[ ] External Factor
CONTRIBUTING FACTORS
[List any contributing factors]
EVIDENCE SUPPORTING ROOT CAUSE
[List evidence]
APPROVAL
Analysis By: _________________ Date: _______
Reviewed By: _________________ Date: _______
```
### CAPA Action Plan
```
CAPA ACTION PLAN
CAPA Number: CAPA-[YYYY]-[NNN]
Root Cause: [Brief statement]
Plan Date: [Date]
Plan Owner: [Name]
CORRECTIVE/PREVENTIVE ACTIONS
Action 1:
Description: [Detailed action description]
Responsible: [Name]
Due Date: [Date]
Resources Required: [Resources]
Success Criteria: [How completion verified]
Action 2:
Description: [Detailed action description]
Responsible: [Name]
Due Date: [Date]
Resources Required: [Resources]
Success Criteria: [How completion verified]
[Continue for additional actions]
RELATED CHANGES
Documents Affected: [List]
Training Required: [Description]
Process Changes: [Description]
Equipment Changes: [Description]
RISK ASSESSMENT
Residual Risk After Implementation: [ ] High [ ] Medium [ ] Low
Justification: [Explanation]
APPROVAL
Plan Developed By: _________________ Date: _______
Approved By: _________________ Date: _______
```
### CAPA Effectiveness Verification
```
CAPA EFFECTIVENESS VERIFICATION
CAPA Number: CAPA-[YYYY]-[NNN]
Verification Date: [Date]
Verified By: [Name]
ACTIONS COMPLETED
| Action | Completion Date | Evidence |
|--------|-----------------|----------|
| [Action 1] | [Date] | [Reference] |
| [Action 2] | [Date] | [Reference] |
EFFECTIVENESS CRITERIA
[Criteria established during action planning]
VERIFICATION METHOD
[ ] Data analysis (trends, metrics)
[ ] Process audit
[ ] Record review
[ ] Product inspection
[ ] Customer feedback review
[ ] Other: _______
VERIFICATION PERIOD
From: [Date] To: [Date]
VERIFICATION RESULTS
[Detailed results of verification activities]
DATA/EVIDENCE REVIEWED
| Data Type | Period | Result |
|-----------|--------|--------|
| [Type] | [Period] | [Result] |
EFFECTIVENESS CONCLUSION
[ ] Effective - Root cause eliminated, problem resolved
[ ] Partially Effective - Improvement noted, additional action needed
[ ] Not Effective - Problem persists, reopen CAPA
If not effective, describe additional actions:
[Description]
CAPA CLOSURE
[ ] Approved for closure
[ ] Not approved - additional action required
Verified By: _________________ Date: _______
Approved By: _________________ Date: _______
```
---
## Supplier Management Templates
### Approved Supplier List
```
APPROVED SUPPLIER LIST
Organization: [Company Name]
Last Updated: [Date]
Maintained By: [Name]
| Supplier | Supplier # | Category | Products/Services | Status | Qualification Date | Next Review |
|----------|-----------|----------|-------------------|--------|-------------------|-------------|
| [Name] | SUP-001 | A | [Products] | Approved | [Date] | [Date] |
| [Name] | SUP-002 | B | [Products] | Conditional | [Date] | [Date] |
Category:
A = Critical (affects safety/performance)
B = Major (affects quality)
C = Minor (indirect impact)
Status:
Approved = Full use authorized
Conditional = Limited use, monitoring
Probation = Performance issues, enhanced monitoring
Disqualified = Use not authorized
Revision History:
| Rev | Date | Change | Approved By |
|-----|------|--------|-------------|
| 01 | [Date] | Initial release | [Name] |
```
### Supplier Evaluation Form
```
SUPPLIER EVALUATION
Supplier Name: [Name]
Supplier Number: [Number]
Evaluation Date: [Date]
Evaluated By: [Name]
Evaluation Type: [ ] Initial [ ] Periodic [ ] For Cause
SUPPLIER INFORMATION
Address: [Address]
Contact: [Name, Title]
Phone: [Phone]
Email: [Email]
Products/Services: [Description]
PROPOSED CATEGORY
[ ] A - Critical (affects safety/performance)
[ ] B - Major (affects quality)
[ ] C - Minor (indirect impact)
EVALUATION CRITERIA
1. QUALITY MANAGEMENT SYSTEM (30 points max)
[ ] ISO 13485 Certified (30 pts)
[ ] ISO 9001 Certified (20 pts)
[ ] Documented QMS (10 pts)
[ ] No formal QMS (0 pts)
Score: ___/30
2. QUALITY HISTORY (25 points max)
Reject Rate: ___% (0-1% = 25 pts, 1-3% = 15 pts, >3% = 0 pts)
Score: ___/25
3. DELIVERY PERFORMANCE (20 points max)
On-Time Delivery: ___% (>95% = 20 pts, 90-95% = 10 pts, <90% = 0 pts)
Score: ___/20
4. TECHNICAL CAPABILITY (15 points max)
[ ] Exceeds requirements (15 pts)
[ ] Meets requirements (10 pts)
[ ] Marginally meets (5 pts)
Score: ___/15
5. FINANCIAL STABILITY (10 points max)
[ ] Strong (10 pts)
[ ] Adequate (5 pts)
[ ] Questionable (0 pts)
Score: ___/10
TOTAL SCORE: ___/100
QUALIFICATION DECISION
>80 = Approved
60-80 = Conditional (monitoring required)
<60 = Not Approved
Decision: [ ] Approved [ ] Conditional [ ] Not Approved
APPROVAL
Evaluated By: _________________ Date: _______
QA Approval: _________________ Date: _______
```
### Supplier Performance Scorecard
```
SUPPLIER PERFORMANCE SCORECARD
Supplier: [Name]
Supplier #: [Number]
Period: [Q1/Q2/Q3/Q4] [Year]
Prepared By: [Name]
PERFORMANCE METRICS
1. QUALITY (40% weight)
Total Lots Received: [N]
Lots Rejected: [N]
Accept Rate: ___% Target: >98%
Score: ___/40
2. DELIVERY (30% weight)
Total Orders: [N]
On-Time Deliveries: [N]
On-Time Rate: ___% Target: >95%
Score: ___/30
3. RESPONSIVENESS (15% weight)
Issues Reported: [N]
Resolved <5 days: [N]
Response Rate: ___% Target: >90%
Score: ___/15
4. DOCUMENTATION (15% weight)
CoC Required: [N]
CoC Complete: [N]
Documentation Rate: ___% Target: 100%
Score: ___/15
TOTAL SCORE: ___/100
PERFORMANCE TREND
| Period | Quality | Delivery | Response | Docs | Total |
|--------|---------|----------|----------|------|-------|
| Q1 | | | | | |
| Q2 | | | | | |
| Q3 | | | | | |
| Q4 | | | | | |
ISSUES/CONCERNS
[List any quality or delivery issues during period]
ACTIONS REQUIRED
[ ] None - Performance acceptable
[ ] Enhanced monitoring
[ ] Supplier corrective action request
[ ] Supplier audit
[ ] Consider alternative supplier
NEXT REVIEW: [Date]
Prepared By: _________________ Date: _______
Reviewed By: _________________ Date: _______
```
---
## Training Templates
### Training Record
```
EMPLOYEE TRAINING RECORD
Employee Name: [Name]
Employee ID: [ID]
Department: [Department]
Job Title: [Title]
Date of Hire: [Date]
REQUIRED TRAINING
| Training | Requirement | Initial Date | Last Date | Next Due | Status |
|----------|-------------|--------------|-----------|----------|--------|
| ISO 13485 Awareness | Initial + Annual | [Date] | [Date] | [Date] | Current |
| Document Control | Initial + On Change | [Date] | [Date] | [Date] | Current |
| CAPA Procedure | Initial + On Change | [Date] | [Date] | [Date] | Due |
| Job-Specific | Per competency matrix | [Date] | [Date] | [Date] | Current |
TRAINING HISTORY
| Date | Training | Method | Duration | Trainer | Assessment | Result |
|------|----------|--------|----------|---------|------------|--------|
| [Date] | [Title] | Classroom | 2 hrs | [Name] | Written test | Pass |
| [Date] | [Title] | OJT | 4 hrs | [Name] | Observation | Pass |
COMPETENCY VERIFICATION
| Competency | Method | Date | Verified By | Result |
|------------|--------|------|-------------|--------|
| [Skill] | Observation | [Date] | [Name] | Qualified |
| [Skill] | Test | [Date] | [Name] | Qualified |
Employee Signature: _________________ Date: _______
Supervisor Signature: _________________ Date: _______
```
### Training Attendance Record
```
TRAINING ATTENDANCE RECORD
Training Title: [Title]
Training Date: [Date]
Trainer: [Name]
Location: [Location]
Duration: [Hours]
TRAINING CONTENT
[Brief description of content covered]
ATTENDEES
| Name | Employee ID | Department | Signature | Assessment Result |
|------|-------------|------------|-----------|-------------------|
| [Name] | [ID] | [Dept] | | Pass/Fail |
| [Name] | [ID] | [Dept] | | Pass/Fail |
ASSESSMENT METHOD
[ ] Written test (attach copy)
[ ] Practical demonstration
[ ] Verbal Q&A
[ ] Observation
[ ] N/A
TRAINING MATERIALS
[ ] Presentation: [Reference]
[ ] Procedure: [Reference]
[ ] Other: [Reference]
Trainer Signature: _________________ Date: _______
Training Coordinator: _________________ Date: _______
```
---
## Nonconformity Templates
### Nonconformity Report
```
NONCONFORMITY REPORT
NC Number: NC-[YYYY]-[NNN]
Date Identified: [Date]
Identified By: [Name]
NONCONFORMITY TYPE
[ ] Product [ ] Process [ ] Document [ ] System
NONCONFORMITY SOURCE
[ ] Incoming inspection
[ ] In-process inspection
[ ] Final inspection
[ ] Customer complaint
[ ] Internal audit
[ ] External audit
[ ] Other: _______
PRODUCT IDENTIFICATION (if applicable)
Product Name: [Name]
Part Number: [Number]
Lot/Batch: [Number]
Quantity Affected: [N]
NONCONFORMITY DESCRIPTION
[Detailed, objective description of the nonconformity]
REQUIREMENT
[Reference to requirement that was not met]
CONTAINMENT ACTION
Action Taken: [Description]
Quantity Contained: [N]
Location: [Location]
Date: [Date]
By: [Name]
DISPOSITION
[ ] Use As Is - Justification: _______
[ ] Rework - Per procedure: _______
[ ] Scrap - Method: _______
[ ] Return to Supplier - RMA #: _______
[ ] Other: _______
Disposition By: [Name]
Disposition Date: [Date]
CAPA REQUIRED?
[ ] Yes - CAPA #: _______
[ ] No - Justification: _______
CLOSURE
All actions complete: [ ] Yes
NC Closed By: _________________ Date: _______
QA Approval: _________________ Date: _______
```
### Material Review Board Record
```
MATERIAL REVIEW BOARD (MRB) RECORD
MRB Number: MRB-[YYYY]-[NNN]
Date: [Date]
NC Reference: NC-[YYYY]-[NNN]
NONCONFORMING MATERIAL
Product: [Name]
Part Number: [Number]
Lot/Batch: [Number]
Quantity: [N]
NONCONFORMITY DESCRIPTION
[Description from NC report]
MRB PARTICIPANTS
| Name | Role | Signature |
|------|------|-----------|
| [Name] | QA Representative | |
| [Name] | Engineering | |
| [Name] | Production | |
| [Name] | Other | |
DISPOSITION OPTIONS CONSIDERED
1. Use As Is
Technical Justification: [Justification]
Risk Assessment: [Assessment]
2. Rework
Procedure: [Reference]
Feasibility: [Assessment]
3. Scrap
Cost Impact: [Amount]
MRB DECISION
[ ] Use As Is - Customer notification required: [ ] Yes [ ] No
[ ] Rework per: [Procedure reference]
[ ] Scrap
[ ] Return to Supplier
RATIONALE
[Detailed rationale for decision]
APPROVALS
| Role | Name | Signature | Date |
|------|------|-----------|------|
| QA | [Name] | | [Date] |
| Engineering | [Name] | | [Date] |
| Production | [Name] | | [Date] |
FOLLOW-UP ACTIONS
[ ] CAPA initiated: CAPA-_______
[ ] Customer notified: Date: _______
[ ] Supplier notified: Date: _______
[ ] Other: _______
```
FILE:scripts/qms_audit_checklist.py
#!/usr/bin/env python3
"""
QMS Internal Audit Checklist Generator
Generates audit checklists for ISO 13485:2016 clauses and QMS processes.
Supports process audits, system audits, and clause-specific audits.
Usage:
python qms_audit_checklist.py --clause 7.3
python qms_audit_checklist.py --process design-control
python qms_audit_checklist.py --audit-type system --output json
python qms_audit_checklist.py --interactive
"""
import argparse
import json
import sys
from datetime import datetime
from typing import Optional
# ISO 13485:2016 Clause Structure with Audit Questions
ISO13485_CLAUSES = {
"4.1": {
"title": "General Requirements",
"questions": [
"Are QMS processes identified and documented?",
"Is the sequence and interaction of processes defined?",
"Are criteria and methods for process operation determined?",
"Are resources and information available for process operation?",
"Are processes monitored, measured, and analyzed?",
"Are actions taken to achieve planned results?",
"Is outsourced process control documented?",
"Are changes to processes managed?"
]
},
"4.2.1": {
"title": "Documentation Requirements - General",
"questions": [
"Is a quality policy documented?",
"Are quality objectives documented?",
"Is a quality manual maintained?",
"Are required documented procedures established?",
"Are documents needed for process planning and operation maintained?",
"Are required records maintained?",
"Is a medical device file established for each device type?"
]
},
"4.2.2": {
"title": "Quality Manual",
"questions": [
"Does the quality manual include QMS scope?",
"Are exclusions justified?",
"Are documented procedures included or referenced?",
"Is the interaction between processes described?",
"Is the quality manual controlled?"
]
},
"4.2.3": {
"title": "Control of Documents",
"questions": [
"Are documents approved before issue?",
"Are documents reviewed and updated as necessary?",
"Are changes and revision status identified?",
"Are current versions available at points of use?",
"Are documents legible and identifiable?",
"Are external documents identified and controlled?",
"Is unintended use of obsolete documents prevented?",
"Is there a document change control process?"
]
},
"4.2.4": {
"title": "Control of Records",
"questions": [
"Is there a procedure for record control?",
"Are records legible and identifiable?",
"Are records retrievable?",
"Are retention times defined?",
"Is protection from damage ensured?",
"Are confidential records protected?",
"Is record disposal controlled?"
]
},
"5.1": {
"title": "Management Commitment",
"questions": [
"Is there evidence of management commitment to QMS?",
"Is the importance of regulatory requirements communicated?",
"Is a quality policy established?",
"Are quality objectives established?",
"Are management reviews conducted?",
"Are resources provided for QMS?"
]
},
"5.2": {
"title": "Customer Focus",
"questions": [
"Are customer requirements determined?",
"Are applicable regulatory requirements determined?",
"Are customer and regulatory requirements met?",
"Is customer satisfaction enhanced?"
]
},
"5.3": {
"title": "Quality Policy",
"questions": [
"Is the quality policy appropriate to the organization?",
"Does it include commitment to compliance?",
"Does it include commitment to effectiveness?",
"Does it provide framework for quality objectives?",
"Is it communicated and understood?",
"Is it reviewed for continuing suitability?"
]
},
"5.4.1": {
"title": "Quality Objectives",
"questions": [
"Are quality objectives measurable?",
"Are they consistent with quality policy?",
"Are they established at relevant functions?",
"Do they include product requirements?",
"Do they include compliance requirements?"
]
},
"5.4.2": {
"title": "QMS Planning",
"questions": [
"Is QMS planning carried out to meet requirements?",
"Is QMS planning done to meet quality objectives?",
"Is QMS integrity maintained during changes?"
]
},
"5.5.1": {
"title": "Responsibility and Authority",
"questions": [
"Are responsibilities and authorities defined?",
"Are they documented?",
"Are they communicated?",
"Are interrelationships defined?"
]
},
"5.5.2": {
"title": "Management Representative",
"questions": [
"Is a management representative appointed?",
"Is authority to ensure QMS processes established?",
"Is authority to report to top management defined?",
"Is authority to promote awareness of requirements defined?"
]
},
"5.5.3": {
"title": "Internal Communication",
"questions": [
"Are communication processes established?",
"Is QMS effectiveness communicated?",
"Is information communicated appropriately?"
]
},
"5.6": {
"title": "Management Review",
"questions": [
"Are management reviews planned?",
"Are all required inputs reviewed?",
"Are outputs documented?",
"Are action items followed up?",
"Are records maintained?"
]
},
"6.1": {
"title": "Provision of Resources",
"questions": [
"Are resources determined?",
"Are resources provided for QMS?",
"Are resources provided for customer satisfaction?",
"Are resources provided for regulatory compliance?"
]
},
"6.2": {
"title": "Human Resources",
"questions": [
"Is competence defined for personnel?",
"Is training provided to achieve competence?",
"Is training effectiveness evaluated?",
"Is awareness of job relevance ensured?",
"Are training records maintained?"
]
},
"6.3": {
"title": "Infrastructure",
"questions": [
"Is necessary infrastructure determined?",
"Are buildings and workspace adequate?",
"Is process equipment adequate?",
"Are supporting services adequate?",
"Are maintenance requirements documented?"
]
},
"6.4": {
"title": "Work Environment",
"questions": [
"Is work environment determined?",
"Are environmental requirements documented?",
"Is contamination control adequate?",
"Are personnel health and cleanliness controlled?",
"Are environmental conditions monitored?"
]
},
"7.1": {
"title": "Planning of Product Realization",
"questions": [
"Are quality objectives for product defined?",
"Are processes needed determined?",
"Is verification and validation defined?",
"Are records requirements defined?",
"Is risk management applied?"
]
},
"7.2": {
"title": "Customer-Related Processes",
"questions": [
"Are customer requirements determined?",
"Are regulatory requirements determined?",
"Are requirements reviewed before commitment?",
"Are differences resolved before acceptance?",
"Is communication with customers effective?"
]
},
"7.3.1": {
"title": "Design and Development Planning",
"questions": [
"Are design stages determined?",
"Are review activities defined?",
"Are verification activities defined?",
"Are validation activities defined?",
"Are responsibilities assigned?",
"Are interfaces managed?"
]
},
"7.3.2": {
"title": "Design and Development Inputs",
"questions": [
"Are functional requirements defined?",
"Are performance requirements defined?",
"Are safety requirements defined?",
"Are regulatory requirements identified?",
"Are previous design inputs considered?",
"Are risk management outputs included?"
]
},
"7.3.3": {
"title": "Design and Development Outputs",
"questions": [
"Do outputs meet input requirements?",
"Is purchasing information provided?",
"Are acceptance criteria defined?",
"Are essential characteristics specified?",
"Are outputs approved before release?"
]
},
"7.3.4": {
"title": "Design and Development Review",
"questions": [
"Are design reviews conducted at suitable stages?",
"Is ability to meet requirements evaluated?",
"Are problems identified?",
"Are follow-up actions recorded?",
"Are appropriate functions represented?"
]
},
"7.3.5": {
"title": "Design and Development Verification",
"questions": [
"Is verification performed per plan?",
"Do outputs meet inputs?",
"Are verification records maintained?",
"Are verification methods appropriate?"
]
},
"7.3.6": {
"title": "Design and Development Validation",
"questions": [
"Is validation performed per plan?",
"Is product evaluated for intended use?",
"Is clinical evaluation included?",
"Are validation records maintained?",
"Is validation completed before product delivery?"
]
},
"7.3.7": {
"title": "Design and Development Transfer",
"questions": [
"Are outputs verified before transfer?",
"Is manufacturing capability verified?",
"Are transfer activities documented?"
]
},
"7.3.8": {
"title": "Control of Design and Development Changes",
"questions": [
"Are design changes identified?",
"Are changes reviewed?",
"Are changes verified?",
"Are changes validated as appropriate?",
"Is impact on product assessed?",
"Are changes approved before implementation?"
]
},
"7.4.1": {
"title": "Purchasing Process",
"questions": [
"Are suppliers evaluated and selected?",
"Are evaluation criteria established?",
"Is supplier performance monitored?",
"Are re-evaluation criteria defined?",
"Is purchased product verified?"
]
},
"7.4.2": {
"title": "Purchasing Information",
"questions": [
"Is purchasing information adequate?",
"Are product requirements specified?",
"Are QMS requirements specified?",
"Are personnel requirements specified?"
]
},
"7.4.3": {
"title": "Verification of Purchased Product",
"questions": [
"Is incoming inspection adequate?",
"Are verification activities defined?",
"Are verification records maintained?",
"Is source verification defined if applicable?"
]
},
"7.5.1": {
"title": "Control of Production and Service Provision",
"questions": [
"Is product information available?",
"Are work instructions available?",
"Is suitable equipment used?",
"Are monitoring devices available?",
"Is monitoring implemented?",
"Are release activities defined?",
"Are labeling requirements met?"
]
},
"7.5.2": {
"title": "Cleanliness of Product",
"questions": [
"Are cleanliness requirements documented?",
"Is contamination controlled?",
"Are process agents controlled?"
]
},
"7.5.3": {
"title": "Installation Activities",
"questions": [
"Are installation requirements documented?",
"Are acceptance criteria defined?",
"Are installation records maintained?"
]
},
"7.5.4": {
"title": "Servicing Activities",
"questions": [
"Are servicing procedures documented?",
"Are reference materials controlled?",
"Are service records maintained?",
"Is feedback analyzed?"
]
},
"7.5.5": {
"title": "Sterile Medical Devices",
"questions": [
"Is sterilization validated?",
"Are process parameters controlled?",
"Is sterile barrier validated?",
"Are sterilization records maintained?"
]
},
"7.5.6": {
"title": "Validation of Processes",
"questions": [
"Are special processes identified?",
"Are validation procedures documented?",
"Is equipment qualified?",
"Are personnel qualified?",
"Are validation records maintained?",
"Are revalidation criteria defined?"
]
},
"7.5.7": {
"title": "Particular Requirements for Validation",
"questions": [
"Are validation methods defined?",
"Are acceptance criteria established?",
"Is software validation appropriate?",
"Are validation records maintained?"
]
},
"7.5.8": {
"title": "Identification",
"questions": [
"Is product identified throughout realization?",
"Is documentation identified?",
"Is UDI implemented as required?"
]
},
"7.5.9": {
"title": "Traceability",
"questions": [
"Are traceability procedures documented?",
"Are components traceable?",
"Is work environment recorded?",
"Is distribution recorded?",
"Is traceability extent defined?"
]
},
"7.5.10": {
"title": "Customer Property",
"questions": [
"Is customer property identified?",
"Is it verified on receipt?",
"Is it protected and safeguarded?",
"Is loss or damage reported?"
]
},
"7.5.11": {
"title": "Preservation of Product",
"questions": [
"Is product identified?",
"Is handling controlled?",
"Is packaging controlled?",
"Is storage controlled?",
"Is protection adequate?"
]
},
"7.6": {
"title": "Control of Monitoring and Measuring Equipment",
"questions": [
"Is equipment calibrated?",
"Is calibration traceable?",
"Is calibration status identified?",
"Is equipment protected from damage?",
"Is software validated?",
"Are records maintained?"
]
},
"8.1": {
"title": "Measurement, Analysis and Improvement - General",
"questions": [
"Are monitoring activities planned?",
"Are analysis activities planned?",
"Are improvement activities planned?"
]
},
"8.2.1": {
"title": "Feedback",
"questions": [
"Is feedback collected?",
"Is feedback analyzed?",
"Is feedback used for improvement?",
"Is regulatory feedback included?"
]
},
"8.2.2": {
"title": "Complaint Handling",
"questions": [
"Is there a complaint procedure?",
"Are complaints investigated?",
"Are regulatory reports made if required?",
"Is trend analysis performed?",
"Are CAPAs initiated when warranted?"
]
},
"8.2.3": {
"title": "Reporting to Regulatory Authorities",
"questions": [
"Are reporting requirements identified?",
"Are reports submitted timely?",
"Are records maintained?"
]
},
"8.2.4": {
"title": "Internal Audit",
"questions": [
"Is an audit program established?",
"Are audit criteria defined?",
"Are auditors independent?",
"Are auditors competent?",
"Are audit records maintained?",
"Are findings followed up?"
]
},
"8.2.5": {
"title": "Monitoring and Measurement of Processes",
"questions": [
"Are processes monitored?",
"Are suitable methods used?",
"Is process capability demonstrated?",
"Are corrections made when needed?"
]
},
"8.2.6": {
"title": "Monitoring and Measurement of Product",
"questions": [
"Is product inspected?",
"Are acceptance criteria met?",
"Is release authorized?",
"Is traceability to inspection recorded?",
"Are records maintained?"
]
},
"8.3": {
"title": "Control of Nonconforming Product",
"questions": [
"Is nonconforming product identified?",
"Is it documented?",
"Is it evaluated?",
"Is it segregated?",
"Is disposition determined?",
"Is rework verified?",
"Is concession controlled?",
"Is post-delivery NC investigated?"
]
},
"8.4": {
"title": "Analysis of Data",
"questions": [
"Is data collected?",
"Is feedback analyzed?",
"Is conformity data analyzed?",
"Is process data analyzed?",
"Is supplier data analyzed?",
"Are audit results analyzed?"
]
},
"8.5.1": {
"title": "Improvement - General",
"questions": [
"Is continual improvement pursued?",
"Are policy, objectives, audits, data, actions, and reviews used?"
]
},
"8.5.2": {
"title": "Corrective Action",
"questions": [
"Is there a CA procedure?",
"Are NCs reviewed (including complaints)?",
"Is root cause determined?",
"Is action needed evaluated?",
"Is action determined and implemented?",
"Are results documented?",
"Is effectiveness verified?"
]
},
"8.5.3": {
"title": "Preventive Action",
"questions": [
"Is there a PA procedure?",
"Are potential NCs identified?",
"Is action needed evaluated?",
"Is action determined and implemented?",
"Are results documented?",
"Is effectiveness verified?"
]
}
}
# Process-to-Clause Mapping
PROCESS_MAPPING = {
"document-control": ["4.2.1", "4.2.2", "4.2.3", "4.2.4"],
"management-review": ["5.6"],
"internal-audit": ["8.2.4"],
"training": ["6.2"],
"design-control": ["7.3.1", "7.3.2", "7.3.3", "7.3.4", "7.3.5", "7.3.6", "7.3.7", "7.3.8"],
"purchasing": ["7.4.1", "7.4.2", "7.4.3"],
"production": ["7.5.1", "7.5.2", "7.5.6", "7.5.7", "7.5.8", "7.5.9", "7.5.11"],
"capa": ["8.5.2", "8.5.3"],
"nonconformity": ["8.3"],
"calibration": ["7.6"],
"complaint-handling": ["8.2.1", "8.2.2", "8.2.3"],
"risk-management": ["7.1"],
"infrastructure": ["6.3", "6.4"],
"customer-requirements": ["5.2", "7.2"]
}
def get_clause_checklist(clause: str) -> dict:
"""Get audit checklist for a specific clause."""
if clause not in ISO13485_CLAUSES:
return {"error": f"Clause {clause} not found"}
clause_data = ISO13485_CLAUSES[clause]
return {
"clause": clause,
"title": clause_data["title"],
"questions": clause_data["questions"],
"question_count": len(clause_data["questions"])
}
def get_process_checklist(process: str) -> dict:
"""Get audit checklist for a specific process."""
if process not in PROCESS_MAPPING:
available = ", ".join(sorted(PROCESS_MAPPING.keys()))
return {"error": f"Process '{process}' not found. Available: {available}"}
clauses = PROCESS_MAPPING[process]
questions = []
for clause in clauses:
if clause in ISO13485_CLAUSES:
clause_data = ISO13485_CLAUSES[clause]
for q in clause_data["questions"]:
questions.append({
"clause": clause,
"clause_title": clause_data["title"],
"question": q
})
return {
"process": process,
"clauses_covered": clauses,
"questions": questions,
"question_count": len(questions)
}
def get_system_audit_checklist() -> dict:
"""Get complete system audit checklist covering all clauses."""
all_questions = []
for clause, data in sorted(ISO13485_CLAUSES.items()):
for q in data["questions"]:
all_questions.append({
"clause": clause,
"clause_title": data["title"],
"question": q
})
return {
"audit_type": "system",
"clauses_covered": list(ISO13485_CLAUSES.keys()),
"questions": all_questions,
"question_count": len(all_questions)
}
def format_checklist_text(checklist: dict) -> str:
"""Format checklist for text output."""
lines = []
if "error" in checklist:
return f"Error: {checklist['error']}"
lines.append("=" * 70)
lines.append("ISO 13485:2016 INTERNAL AUDIT CHECKLIST")
lines.append(f"Generated: {datetime.now().strftime('%Y-%m-%d %H:%M')}")
lines.append("=" * 70)
if "clause" in checklist:
lines.append(f"\nClause: {checklist['clause']} - {checklist['title']}")
lines.append("-" * 50)
for i, q in enumerate(checklist["questions"], 1):
lines.append(f"\n{i}. {q}")
lines.append(" [ ] C [ ] NC [ ] OBS [ ] N/A")
lines.append(" Evidence: _________________________________")
lines.append(" Notes: ____________________________________")
elif "process" in checklist:
lines.append(f"\nProcess: {checklist['process'].replace('-', ' ').title()}")
lines.append(f"Clauses Covered: {', '.join(checklist['clauses_covered'])}")
lines.append("-" * 50)
current_clause = None
item_num = 1
for q in checklist["questions"]:
if q["clause"] != current_clause:
current_clause = q["clause"]
lines.append(f"\n--- {q['clause']} {q['clause_title']} ---")
lines.append(f"\n{item_num}. {q['question']}")
lines.append(" [ ] C [ ] NC [ ] OBS [ ] N/A")
lines.append(" Evidence: _________________________________")
lines.append(" Notes: ____________________________________")
item_num += 1
elif "audit_type" in checklist:
lines.append(f"\nAudit Type: Full System Audit")
lines.append(f"Total Clauses: {len(checklist['clauses_covered'])}")
lines.append("-" * 50)
current_clause = None
item_num = 1
for q in checklist["questions"]:
if q["clause"] != current_clause:
current_clause = q["clause"]
lines.append(f"\n{'=' * 40}")
lines.append(f"CLAUSE {q['clause']}: {q['clause_title']}")
lines.append("=" * 40)
lines.append(f"\n{item_num}. {q['question']}")
lines.append(" [ ] C [ ] NC [ ] OBS [ ] N/A")
lines.append(" Evidence: _________________________________")
item_num += 1
lines.append("\n" + "=" * 70)
lines.append(f"Total Questions: {checklist['question_count']}")
lines.append("")
lines.append("Legend: C=Conforming, NC=Nonconforming, OBS=Observation, N/A=Not Applicable")
lines.append("=" * 70)
return "\n".join(lines)
def interactive_mode():
"""Run interactive audit checklist generator."""
print("\n" + "=" * 50)
print("QMS INTERNAL AUDIT CHECKLIST GENERATOR")
print("=" * 50)
print("\nSelect audit type:")
print("1. Clause-specific audit")
print("2. Process audit")
print("3. Full system audit")
print("4. List available processes")
print("5. List all clauses")
print("6. Exit")
choice = input("\nEnter choice (1-6): ").strip()
if choice == "1":
print("\nAvailable clause sections:")
print(" 4.x - Quality Management System")
print(" 5.x - Management Responsibility")
print(" 6.x - Resource Management")
print(" 7.x - Product Realization")
print(" 8.x - Measurement, Analysis, Improvement")
clause = input("\nEnter clause number (e.g., 7.3.1): ").strip()
checklist = get_clause_checklist(clause)
print(format_checklist_text(checklist))
elif choice == "2":
processes = sorted(PROCESS_MAPPING.keys())
print("\nAvailable processes:")
for i, p in enumerate(processes, 1):
clauses = PROCESS_MAPPING[p]
print(f" {i}. {p} (clauses: {', '.join(clauses)})")
process = input("\nEnter process name: ").strip().lower()
checklist = get_process_checklist(process)
print(format_checklist_text(checklist))
elif choice == "3":
print("\nGenerating full system audit checklist...")
checklist = get_system_audit_checklist()
print(format_checklist_text(checklist))
elif choice == "4":
processes = sorted(PROCESS_MAPPING.keys())
print("\nAvailable QMS Processes:")
print("-" * 50)
for p in processes:
clauses = PROCESS_MAPPING[p]
print(f" {p}")
print(f" Clauses: {', '.join(clauses)}")
elif choice == "5":
print("\nISO 13485:2016 Clauses:")
print("-" * 50)
for clause, data in sorted(ISO13485_CLAUSES.items()):
print(f" {clause}: {data['title']} ({len(data['questions'])} questions)")
elif choice == "6":
print("Exiting.")
return
else:
print("Invalid choice.")
def main():
parser = argparse.ArgumentParser(
description="Generate ISO 13485:2016 internal audit checklists",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
python qms_audit_checklist.py --clause 7.3
python qms_audit_checklist.py --process design-control
python qms_audit_checklist.py --audit-type system --output json
python qms_audit_checklist.py --list-processes
python qms_audit_checklist.py --list-clauses
python qms_audit_checklist.py --interactive
"""
)
parser.add_argument(
"--clause",
help="Generate checklist for specific clause (e.g., 7.3.1, 8.5.2)"
)
parser.add_argument(
"--process",
help="Generate checklist for process (e.g., design-control, capa)"
)
parser.add_argument(
"--audit-type",
choices=["clause", "process", "system"],
help="Audit type for checklist generation"
)
parser.add_argument(
"--output",
choices=["text", "json"],
default="text",
help="Output format (default: text)"
)
parser.add_argument(
"--list-processes",
action="store_true",
help="List available QMS processes"
)
parser.add_argument(
"--list-clauses",
action="store_true",
help="List all ISO 13485 clauses"
)
parser.add_argument(
"--interactive",
action="store_true",
help="Run in interactive mode"
)
args = parser.parse_args()
if args.interactive:
interactive_mode()
return
if args.list_processes:
processes = sorted(PROCESS_MAPPING.keys())
if args.output == "json":
result = {p: PROCESS_MAPPING[p] for p in processes}
print(json.dumps(result, indent=2))
else:
print("\nAvailable QMS Processes:")
print("-" * 50)
for p in processes:
clauses = PROCESS_MAPPING[p]
print(f" {p}: {', '.join(clauses)}")
return
if args.list_clauses:
if args.output == "json":
result = {c: {"title": d["title"], "question_count": len(d["questions"])}
for c, d in sorted(ISO13485_CLAUSES.items())}
print(json.dumps(result, indent=2))
else:
print("\nISO 13485:2016 Clauses:")
print("-" * 50)
for clause, data in sorted(ISO13485_CLAUSES.items()):
print(f" {clause}: {data['title']} ({len(data['questions'])} questions)")
return
checklist = None
if args.clause:
checklist = get_clause_checklist(args.clause)
elif args.process:
checklist = get_process_checklist(args.process)
elif args.audit_type == "system":
checklist = get_system_audit_checklist()
else:
parser.print_help()
return
if checklist:
if args.output == "json":
print(json.dumps(checklist, indent=2))
else:
print(format_checklist_text(checklist))
if __name__ == "__main__":
main()
Bộ 10 skill sản phẩm: PM toolkit (RICE), PO agile, chiến lược OKR, nghiên cứu UX, design system UI, phân tích đối thủ, landing page, SaaS scaffolder.
--- name: "product-skills" description: "10 product agent skills and plugins for Claude Code, Codex, Gemini CLI, Cursor, OpenClaw. PM toolkit (RICE), agile PO, product strategist (OKR), UX researcher, UI design system, competitive teardown, landing page generator, SaaS scaffolder, research summarizer. Python tools (stdlib-only)." version: 2.9.0 author: Alireza Rezvani license: MIT tags: - product - product-management - ux - ui - saas - agile agents: - claude-code - codex-cli - openclaw --- # Product Team Skills 8 production-ready product skills covering product management, UX/UI design, and SaaS development. ## Quick Start ### Claude Code ``` /read product-team/product-manager-toolkit/SKILL.md ``` ### Codex CLI ```bash npx agent-skills-cli add alirezarezvani/claude-skills/product-team ``` ## Skills Overview | Skill | Folder | Focus | |-------|--------|-------| | Product Manager Toolkit | `product-manager-toolkit/` | RICE prioritization, customer discovery, PRDs | | Agile Product Owner | `agile-product-owner/` | User stories, sprint planning, backlog | | Product Strategist | `product-strategist/` | OKR cascades, market analysis, vision | | UX Researcher Designer | `ux-researcher-designer/` | Personas, journey maps, usability testing | | UI Design System | `ui-design-system/` | Design tokens, component docs, responsive | | Competitive Teardown | `competitive-teardown/` | Systematic competitor analysis | | Landing Page Generator | `landing-page-generator/` | Conversion-optimized pages | | SaaS Scaffolder | `saas-scaffolder/` | Production SaaS boilerplate | ## Python Tools 9 scripts, all stdlib-only: ```bash python3 product-manager-toolkit/scripts/rice_prioritizer.py --help python3 product-strategist/scripts/okr_cascade_generator.py --help ``` ## Rules - Load only the specific skill SKILL.md you need - Use Python tools for scoring and analysis, not manual judgment
Tạo skill agent mới với cấu trúc đúng chuẩn, tiết lộ thông tin dần dần và tài nguyên đi kèm.
---
name: write-a-skill
description: Create new agent skills with proper structure, progressive disclosure, and bundled resources. Use when user wants to create, write, build, or author a new skill.
license: MIT
metadata:
derived_from: "https://github.com/mattpocock/skills/tree/main/skills/productivity/write-a-skill"
original_author: "Matt Pocock (@mattpocock)"
original_license: MIT
voice: "Matt Pocock — direct, concrete, imperative, example-driven"
version: 1.0.0
---
# Writing Skills
> Derived from [Matt Pocock's write-a-skill](https://github.com/mattpocock/skills/tree/main/skills/productivity/write-a-skill) (MIT). Matt's voice and 3-phase workflow preserved verbatim. Additions: validation tools + references + cs-* wrapper (see *Tooling + Companions* below).
## Process
1. **Gather requirements** - ask user about:
- What task/domain does the skill cover?
- What specific use cases should it handle?
- Does it need executable scripts or just instructions?
- Any reference materials to include?
2. **Draft the skill** - create:
- SKILL.md with concise instructions
- Additional reference files if content exceeds 500 lines
- Utility scripts if deterministic operations needed
3. **Review with user** - present draft and ask:
- Does this cover your use cases?
- Anything missing or unclear?
- Should any section be more/less detailed?
## Skill Structure
```
skill-name/
├── SKILL.md # Main instructions (required)
├── REFERENCE.md # Detailed docs (if needed)
├── EXAMPLES.md # Usage examples (if needed)
└── scripts/ # Utility scripts (if needed)
└── helper.js
```
## SKILL.md Template
```md
---
name: skill-name
description: Brief description of capability. Use when [specific triggers].
---
# Skill Name
## Quick start
[Minimal working example]
## Workflows
[Step-by-step processes with checklists for complex tasks]
## Advanced features
[Link to separate files: See [REFERENCE.md](REFERENCE.md)]
```
## Description Requirements
The description is **the only thing your agent sees** when deciding which skill to load. It's surfaced in the system prompt alongside all other installed skills. Your agent reads these descriptions and picks the relevant skill based on the user's request.
**Goal**: Give your agent just enough info to know:
1. What capability this skill provides
2. When/why to trigger it (specific keywords, contexts, file types)
**Format**:
- Max 1024 chars
- Write in third person
- First sentence: what it does
- Second sentence: "Use when [specific triggers]"
**Good example**:
```
Extract text and tables from PDF files, fill forms, merge documents. Use when working with PDF files or when user mentions PDFs, forms, or document extraction.
```
**Bad example**:
```
Helps with documents.
```
The bad example gives your agent no way to distinguish this from other document skills.
## When to Add Scripts
Add utility scripts when:
- Operation is deterministic (validation, formatting)
- Same code would be generated repeatedly
- Errors need explicit handling
Scripts save tokens and improve reliability vs generated code.
## When to Split Files
Split into separate files when:
- SKILL.md exceeds 100 lines
- Content has distinct domains (finance vs sales schemas)
- Advanced features are rarely needed
## Review Checklist
After drafting, verify:
- [ ] Description includes triggers ("Use when...")
- [ ] SKILL.md under 100 lines
- [ ] No time-sensitive info
- [ ] Consistent terminology
- [ ] Concrete examples included
- [ ] References one level deep
## Tooling + Companions
Validation tools + cs-* wrapper sit alongside this skill. Run all 6 review-checklist items programmatically:
```
python scripts/skill_review_checklist_runner.py path/to/skill-folder
```
See [references/companion_tooling.md](references/companion_tooling.md) for the tool catalogue, cs-skill-author persona agent, and `/cs:write-a-skill` slash command.
---
**Version:** 1.0.0
**Derived:** Matt Pocock (MIT) + this repo's wrapper
FILE:references/companion_tooling.md
# Companion Tooling
Validation tools + cs-* wrapper layered on top of Matt's write-a-skill. Use these when authoring a new skill in this repo.
## Validation Tools (stdlib Python)
| Tool | Purpose | Run before |
|---|---|---|
| `scripts/skill_description_validator.py` | Validates description: ≤1024 chars, third person, "Use when" trigger, action verb in first sentence | First draft of SKILL.md |
| `scripts/skill_structure_validator.py` | Validates folder structure: SKILL.md present, ≤100 lines, references one level deep, no circular refs | Pre-commit |
| `scripts/skill_review_checklist_runner.py` | Runs all 6 review-checklist items from Matt's write-a-skill against a skill folder | Final check before PR |
All three tools:
- Stdlib-only (no external dependencies)
- Run with embedded sample if no path provided
- Output text or JSON (`--output json`)
- Exit code: 0 if PASS, 1 if FAIL/WARN
## cs-skill-author Persona Agent
Lives at `../agents/cs-skill-author.md`. Voice: forcing-question interrogator. Surfaces Matt's skill-authoring workflow as an interrogation before any new skill commit.
**Opening question:** "What capability does this skill provide, and what's the trigger phrase that distinguishes it from existing skills?"
**Six forcing questions** (matches the review checklist):
1. What's the description? Is it ≤1024 chars + third person + has "Use when ..."?
2. Is SKILL.md under 100 lines? If not, where will the split land (REFERENCE.md / EXAMPLES.md / references/)?
3. Are there time-sensitive claims (dates, "as of YYYY")?
4. Is terminology consistent — same word for the same concept throughout?
5. Concrete examples — at least 1 code block, ideally good/bad contrast?
6. References one level deep, no circular refs?
## `/cs:write-a-skill` Slash Command
Lives at `../commands/cs-write-a-skill.md`. Three-step flow:
1. Run `cs-skill-author` interrogation (6 questions)
2. Draft skill files per Matt's structure pattern
3. Run all 3 validation tools; show verdict; fix until PASS
Use when: starting a new skill in this repo from scratch.
## Why Wrap Matt's Original
Matt's write-a-skill is a tight, principled, ~93-line skill — perfect as-is for individual authoring sessions. The wrapper layers add three things this repo benefits from at scale:
1. **Programmatic enforcement** of Matt's review checklist (the validation tools) — prevents human review-checklist drift across 100+ skills.
2. **Forcing-question interrogation** (the cs-skill-author persona) — adapts Matt's "review with user" phase to the cs-* persona pattern used elsewhere in this repo.
3. **Citation-backed references** — Matt links to his own materials; the wrapper adds 5+ authoritative external sources per reference (Anthropic skill docs + community precedent + research) for newcomers learning the pattern.
This is the [hybrid voice approach](../SKILL.md): Matt's words for the principles, our additions for the tooling.
## Attribution
Original: [matt-pocock/skills/skills/productivity/write-a-skill](https://github.com/mattpocock/skills/tree/main/skills/productivity/write-a-skill) (MIT).
---
**Source authorities (non-exhaustive):**
- **Matt Pocock — write-a-skill** (https://github.com/mattpocock/skills/, MIT, 2024) — the upstream source
- **Anthropic — Skills documentation** (https://docs.claude.com/en/docs/agents/skills) — official guidance on skill structure
- **Anthropic Engineering Blog — Skills patterns** (continuously updated) — patterns for skill authoring
- **Karpathy, A. — "Software 3.0" + LLM coding pitfalls** (X.com posts 2024-2025) — discipline reference applied throughout this repo's karpathy-coder skill
- **Pareto principle applied to documentation** — concise = trustworthy; 80% of value in 20% of words
- **Hyrum's Law** as applied to skill descriptions — once a description shape is observed, downstream agents depend on it
- **Conway's Law as applied to skill libraries** — skill organization mirrors team responsibilities; progressive disclosure mirrors information needs across team boundaries
FILE:references/description_design_patterns.md
# Description Design Patterns for Skills
This reference answers exactly one decision: **how do we write a skill description that an agent actually picks correctly when faced with a long skill list?**
Pair with `scripts/skill_description_validator.py` for automated enforcement.
## Matt Pocock's Foundational Rule
> "The description is **the only thing your agent sees** when deciding which skill to load."
>
> — Matt Pocock, write-a-skill
Implication: the description is not marketing copy. It's a routing signal for the agent. Every word competes with every other skill's description for activation attention.
## The Four Format Rules (per Matt)
1. **Max 1024 chars** — beyond this, agents lose the early sentences when condensing context
2. **Third person** — first-person ("I help with...") confuses agent self-identification; second-person ("You can...") confuses pronoun reference
3. **First sentence: what it does** — front-load the verb + object
4. **Second sentence: "Use when [specific triggers]"** — agent's most reliable activation cue
## Good vs Bad Examples (Matt's pattern, expanded)
**Good** (Matt's PDF example):
```
Extract text and tables from PDF files, fill forms, merge documents. Use when working with PDF files or when user mentions PDFs, forms, or document extraction.
```
**Why good:**
- Front-loaded verbs: Extract, fill, merge
- Concrete objects: text, tables, PDF files, forms
- Explicit trigger: "Use when working with PDF files"
- Specific keywords for matching: "PDFs", "forms", "document extraction"
**Bad** (Matt's):
```
Helps with documents.
```
**Why bad:**
- "Helps" is content-free
- "Documents" is generic — every doc skill has this
- No trigger
- No keyword variety
**Bad in different way** (over-specified):
```
This skill performs comprehensive PDF document processing including but not limited to extraction, manipulation, format conversion, content analysis, metadata management, and security operations on PDF files, with support for various PDF versions and embedded media types.
```
**Why bad:** verbose, no triggers, agent can't extract the key keywords from the wall of text.
## The Trigger Sentence Pattern
The "Use when" sentence is the highest-leverage part of the description. Patterns that work:
**Keyword triggers** (when user types specific words):
```
Use when user mentions PDFs, forms, or document extraction.
```
**File-type triggers** (when agent sees specific files):
```
Use when working with `.tsx` files or React component tests.
```
**Context triggers** (when agent is in a specific state):
```
Use when the user requests a code review of a pull request.
```
**Workflow triggers** (when agent is mid-workflow):
```
Use after running tests and before committing changes.
```
## Vocabulary Selection
The description's words must overlap with words users + agents naturally use for the task.
| Bad keyword | Better keyword | Why |
|---|---|---|
| "documents" | "PDF files" / "Word docs" | More specific = less collision |
| "improve" | "refactor" / "fix" / "optimize" | Specific verb = clearer routing |
| "various" | (delete; just list them) | Hedge language = no info |
| "modern" | (cite the actual tool/version) | Trend words age badly |
| "comprehensive" | (delete; just list capabilities) | Adjective inflation |
## Length Optimization
Below 1024 chars, shorter is usually better. Target: 100-300 chars for most skills.
Where complexity demands more chars, prioritize:
1. The verb-object pair (what it does) — never compress
2. The trigger phrase — never compress
3. Keyword variety (different ways users describe it) — expand here if space allows
4. Anti-keyword (what it does NOT do) — only if there's a frequently-confused sibling skill
## Anti-Patterns to Avoid
1. **First-person voice** — "I extract PDFs" — confuses agent self-reference
2. **Marketing language** — "fast, powerful, intuitive" — agent doesn't care, ignores adjectives
3. **Trigger-less descriptions** — every skill needs "Use when X"
4. **Multi-purpose dumping** — if your skill does 10 unrelated things, it's probably 10 skills
5. **Pronouns and hedges** — "you can also use this if you want to" — drop entirely
6. **Recursive descriptions** — "Use this skill when you need this skill" — adds nothing
7. **Implementation details** — "Built on Python + stdlib" — agent doesn't care; matters for README, not description
## Pre-Commit Discipline
Run before every skill PR:
```bash
python scripts/skill_description_validator.py path/to/SKILL.md
```
If validator returns FAIL, fix before merging. If WARN, justify and document the trade-off.
## When This Reference Doesn't Help
- **Naming the skill itself** — different concern; see naming-conventions guidance per-repo
- **Skill discovery in marketplaces** — different audience (humans browsing), different rules
- **System-prompt design for the agent that loads skills** — upstream concern
---
**Source authorities (non-exhaustive):**
- **Matt Pocock — write-a-skill** (https://github.com/mattpocock/skills/, MIT) — the 4 format rules + good/bad example pattern
- **Anthropic — Building agents with skills** (https://docs.claude.com/en/docs/agents/skills) — official format guidance
- **Anthropic Engineering — Effective system prompts** (continuously updated blog) — same principles applied to system-prompt design
- **Claude Code documentation — Skill registry** — how Claude's skill-loader uses descriptions
- **Karpathy, A. — public commentary on LLM prompt design** — emphasis on specificity + lack of ambiguity
- **Garrett, J.J. — "The Elements of User Experience"** (2002) + information architecture principles — labels must match user mental models
- **Nielsen Norman Group — Microcontent guidelines** — applies to skill descriptions: front-load value, hard-cap length, scannable structure
- **Search-engine + SEO patterns adapted for agent routing** — keyword density, intent matching, semantic field coverage
FILE:references/progressive_disclosure_principles.md
# Progressive Disclosure for Skill Files
This reference answers exactly one decision: **when should a SKILL.md be split into reference files, and how do we keep the disclosure ladder shallow + scannable?**
Pair with `scripts/skill_structure_validator.py` for automated enforcement of the 100-line ceiling + one-level-deep rule.
## What "Progressive Disclosure" Means in Skill Files
Progressive disclosure = present the minimum needed to act, with paths to deeper detail when needed. For agent skills:
- **SKILL.md** = the description + minimum workflow the agent needs to invoke the skill
- **REFERENCE.md / EXAMPLES.md / references/*.md** = deep detail invoked only when the SKILL.md workflow points there
- **scripts/** = deterministic operations (no LLM token cost; no inconsistency risk)
The goal: agent reads SKILL.md and either has enough to act, or has a clear link to the specific reference file that resolves its question. No deeper than that.
## Matt Pocock's Original Rule (the 100-Line Ceiling)
> "Split into separate files when:
> - SKILL.md exceeds 100 lines
> - Content has distinct domains (finance vs sales schemas)
> - Advanced features are rarely needed"
>
> — Matt Pocock, write-a-skill
The 100-line ceiling is empirical: agents reading >100 lines of SKILL.md tend to over-condition on tangential detail; below 100 lines, the agent reads the entire skill and routes correctly to references or scripts when needed.
## When the Ceiling Is Right vs Wrong
| Situation | 100-line ceiling appropriate? |
|---|---|
| Single-action skill (e.g., format-json) | Yes — fits comfortably under 50 lines |
| Mid-complexity skill with 2-3 workflows | Yes — 70-100 lines |
| Skill with 4+ workflows + extensive examples | No — split workflows into separate reference files |
| Domain-spanning skill (multi-framework like compliance-os) | No — split per-framework into separate references |
| Skill that wraps another (derived/extension) | Special case — wrapper additions push past 100; treat as warning, not failure |
## The One-Level-Deep Rule
> "References one level deep" — Matt Pocock review checklist
Why: agent loading a reference file should resolve its question without further indirection. If `REFERENCE.md` says "see `references/foo.md` for more on bar," then bar's content is the leaf — it shouldn't say "see references/foo/bar/baz.md."
Operational consequence: keep `references/` flat. No nested subfolders.
## Anti-Patterns to Avoid
1. **SKILL.md as a complete manual** — 300-line SKILL.md with every workflow inline. Agent over-conditions; token cost on every invocation.
2. **Reference soup** — 20 reference files at one level. Hard to scan; agent can't tell which to load.
3. **Circular references** — `A.md` → `B.md` → `A.md`. Agent loops or fails.
4. **No examples in SKILL.md** — "see EXAMPLES.md for usage." Forces agent to load another file to do anything. Provide a *minimum* example in SKILL.md.
5. **Versioned references** — `references/v1/` and `references/v2/`. Maintenance burden; pick one.
6. **Auto-generated table-of-contents** — agents don't need this; humans rarely browse `references/`.
## How to Apply Progressive Disclosure Concretely
1. Draft SKILL.md with the workflow you want the agent to use 80% of the time
2. Count lines. If > 100, identify the next-largest section. Move it to `references/<topic>.md`.
3. Replace the moved section with a 1-2-line pointer: "See [references/topic.md](references/topic.md) for X."
4. Repeat until SKILL.md ≤ 100 lines.
5. Validate: `python scripts/skill_structure_validator.py path/to/skill-folder/`
## When 100 Is Too Restrictive
For skills that wrap or extend other skills (like this `write-a-skill` itself, which preserves Matt's full original content + adds wrapper sections), the 100-line ceiling becomes an artifact of attribution rather than over-conditioning. Two options:
- Accept the line-count WARN as documentation of intentional preservation
- Move attribution/wrapper notes to `README.md` (which lives outside the SKILL.md ceiling)
This `write-a-skill` skill demonstrates option 1.
## When This Reference Doesn't Help
- **Choosing what to put in scripts/ vs references/** — see Matt's "When to Add Scripts" guidance in main SKILL.md.
- **Information architecture for documentation sites** — see DocOps + DITA references.
- **Token-budget optimization beyond skill files** — different scope (system-prompt design, context engineering).
---
**Source authorities (non-exhaustive):**
- **Matt Pocock — write-a-skill** (https://github.com/mattpocock/skills/, MIT) — the 100-line ceiling + one-level-deep rule originator
- **Anthropic — Building agents with skills** (https://docs.claude.com/en/docs/agents/skills) — official skill structure documentation
- **Anthropic Engineering Blog — Prompt design + context engineering** — concise context = lower hallucination + better routing
- **Don Norman — "The Design of Everyday Things"** (1988) + progressive disclosure HCI principle — origin of the term
- **Information Foraging Theory** — Pirolli & Card (1995) — humans + agents search info using cost/benefit tradeoffs analogous to foraging
- **John Maeda — "The Laws of Simplicity"** (2006) — reduction principle applied to UX, directly applicable to skill files
- **Lean Documentation movement** — DocOps + DITA practitioners on minimum-viable-documentation patterns
- **Pareto principle (80/20 rule)** applied to skill workflows — most agent invocations use the same 20% of skill content
FILE:references/quality_gates_for_skills.md
# Quality Gates for Skill Libraries
This reference answers exactly one decision: **what checks must pass before a new skill enters the library, and why?**
Pair with `scripts/skill_review_checklist_runner.py` for the automated gate.
## The Six Mandatory Gates (per Matt Pocock's checklist)
| # | Check | Why it matters |
|---|---|---|
| 1 | Description includes triggers ("Use when ...") | Without trigger, agent guesses when to activate — high false-positive rate |
| 2 | SKILL.md under 100 lines | Over-conditioning; agent reads tangential detail and misroutes |
| 3 | No time-sensitive info | Dates/versions/year refs rot; agent receives stale guidance |
| 4 | Consistent terminology | Synonym drift confuses the agent + downstream users |
| 5 | Concrete examples included | Without an example, agent constructs from scratch and hallucinates |
| 6 | References one level deep | Deep nesting = agent gives up resolving the reference chain |
## Why Programmatic, Not Manual
Manual review of these 6 items:
- Drifts across reviewers (different humans interpret "concrete example" differently)
- Slows PR cadence (every reviewer re-reads every skill against every check)
- Misses regressions (a skill once compliant can drift across updates)
Programmatic gate (the `skill_review_checklist_runner.py` tool):
- Same verdict regardless of reviewer
- Runs in CI in seconds
- Catches regressions automatically
- Documents the explicit criteria — no implicit reviewer judgment
## Beyond Matt's Six: Additional Quality Dimensions
Matt's 6 are the floor. For a mature skill library, add:
### Citation density (this repo's standard)
Every reference file in `references/` should cite ≥ 5 authoritative sources. Why: skills inspired by public material need traceable provenance. Tool: grep-based count of bibliography entries.
### Tool determinism (karpathy-coder discipline)
Every script in `scripts/` should:
- Be stdlib-only (no external dependencies)
- Have embedded sample input
- Support `--output {text,json}`
- Be deterministic (no randomness, no LLM calls)
Tool: `engineering/karpathy-coder/skills/karpathy-coder/scripts/complexity_checker.py`
### Cross-skill compatibility
For skills that reference other skills (via `Adjacent Skills` sections), every cross-reference must resolve to an existing skill. Tool: link-integrity grep across skill folders.
### Attribution discipline (this repo's standard)
Skills derived from external sources (MIT-licensed or public-domain) must:
- Name the original author
- Link to the original source
- State the license
- Note what's preserved vs added
Tool: presence-of-attribution grep in plugin.json + README.md.
## Quality Gate Sequencing
Apply gates in this order during PR:
```
1. Description validator (fast; catches most issues early)
2. Structure validator (fast; folder layout + line counts)
3. Review checklist runner (combined; all 6 of Matt's items)
4. Karpathy complexity check (code quality; only if scripts/ exists)
5. Karpathy assumption linter (code quality; only if scripts/ exists)
6. Link integrity scan (cross-skill references)
7. Citation density check (references/ bibliography)
```
If any gate fails, PR is blocked. WARN status (1 check fails out of 6) requires reviewer justification in PR description.
## CI Integration Pattern
```yaml
# .github/workflows/skill-quality-gate.yml (illustrative)
on: [pull_request]
jobs:
skill-quality:
steps:
- uses: actions/checkout@v4
- name: Run review checklist
run: |
for skill in $(find . -name "SKILL.md" -type f); do
python engineering/write-a-skill/skills/write-a-skill/scripts/skill_review_checklist_runner.py "$(dirname $skill)"
done
- name: Run karpathy gate
run: python engineering/karpathy-coder/skills/karpathy-coder/scripts/complexity_checker.py .
```
## Common Failure Modes (and Fixes)
| Failure | Common cause | Fix |
|---|---|---|
| Description >1024 chars | Trying to describe every feature | Cut to verbs + objects + triggers; move details to SKILL.md |
| SKILL.md >100 lines | Inline workflows that belong in references | Move workflows to `references/<workflow>.md`; replace with 1-line pointers |
| Missing "Use when" | Description written as marketing copy | Rewrite second sentence to start with "Use when ..." |
| Time-sensitive info | "As of October 2024 ..." | Remove date; describe pattern that doesn't depend on date |
| No examples | Abstract guidance only | Add at least 1 code block showing minimum invocation |
| Deep references | Subfolder structure under references/ | Flatten to one level |
## Quality Gate Anti-Patterns
1. **Disabling gates "just for this skill"** — once disabled, never re-enabled. If a gate genuinely doesn't apply, document the exception in skill metadata.
2. **Reviewer override without rationale** — if a reviewer bypasses a check, they own future regressions. Require justification.
3. **Manual review for what tools can check** — wastes reviewer attention on mechanical items. Reserve manual review for judgment calls (is the workflow correct? Does the skill cover the stated use case?).
4. **Gate proliferation** — adding new gates faster than they're enforced creates fatigue. Cap at ~10 gates total; merge similar ones.
## Binding vs Advisory for Legacy Skills
Matt's 6-item checklist is **binding for new skills** (any skill authored after v2.6.0 must PASS all 6 before merge). For **legacy skills** authored before this discipline was established, the same rules apply as **advisory** signals to triage, not blockers.
The reason: this repo has 298 SKILL.md files written under different conventions over time. Auditing them against the v2.6.0 checklist surfaces real tech debt, but retro-fitting all 298 in one sweep would require ~50-100 hours of careful editing. Forcing the gate as blocking would either delay all PRs or require disabling the gate.
The pragmatic split:
| Skill cohort | Gate status | Action on failure |
|---|---|---|
| **New skills (post-v2.6.0)** | **Blocking** — must PASS all 6 | Fix before PR merge |
| **Legacy skills (pre-v2.6.0)** | **Advisory** — WARN/FAIL surfaced but non-blocking | Track in audit report; fix opportunistically |
How to tell which cohort a skill belongs to:
- New: matches the `engineering/<skill>/skills/<skill>/` wrapper pattern with `attribution` in plugin.json, OR was added in a PR tagged for v2.6.0+
- Legacy: pre-existing structure without the wrapper pattern, or pre-v2.6.0 git history
Re-running `scripts/audit_skills.py` periodically captures the legacy backlog drift. The numerator (PASS count) is the metric to grow over time, not "force every skill to PASS by Friday."
## Common Cohort-Specific Issues
**Legacy SKILL.md > 100 lines (88% of repo):** the dominant violation. Most legacy skills predate the 100-line ceiling. Splitting them into `references/` is invasive. The advisory frame: a 200-line legacy SKILL.md isn't urgent unless the skill is actively being edited.
**Legacy missing "Use when" trigger (26% of repo after v2.6.1 validator fix):** highest-leverage fix because it's a 1-line edit per skill. Even legacy skills should adopt this in the next time they're touched.
**Legacy placeholder descriptions (e.g., "Migration Architect" as the only description text):** these are real bugs, not just lint failures. Fix on sight. v2.6.1 fixed 10 of these in the engineering POWERFUL tier.
## When This Reference Doesn't Help
- **Performance optimization of skills** — different concern; benchmark agent token usage, not skill files
- **Skill discovery + organization in marketplaces** — different audience (humans), different rules
- **A/B testing skills** — different mode; quality gates are preconditions, not A/B subjects
---
**Source authorities (non-exhaustive):**
- **Matt Pocock — write-a-skill** (https://github.com/mattpocock/skills/, MIT) — the 6-item review checklist
- **Karpathy, A. — public commentary on LLM coding pitfalls** (X.com, 2024-2025) — discipline framework adopted as `engineering/karpathy-coder/`
- **Anthropic — Building agents with skills** (https://docs.claude.com/en/docs/agents/skills) — official skill quality guidance
- **Continuous Integration / Continuous Deployment patterns** — Humble & Farley (Continuous Delivery, 2010) — gate sequencing principles
- **The Phoenix Project** (Kim et al., 2013) + Three Ways of DevOps — quality gates as constraint management
- **Hyrum's Law** as applied to skill libraries — once a skill's behavior is observed, downstream depends on it; quality gates prevent drift
- **Software craftsmanship + the Boy Scout Rule** — leave each skill cleaner than you found it; gates enforce the floor
FILE:scripts/skill_description_validator.py
#!/usr/bin/env python3
"""skill_description_validator.py — Validate a skill's description against Matt Pocock's rules.
Stdlib-only. Parses YAML frontmatter of a SKILL.md and checks the `description`
field against the criteria from Matt Pocock's write-a-skill:
1. Description present (non-empty after `description:` key)
2. Length <= 1024 characters
3. Written in third person (no first-person pronouns I/me/my; no second-person you)
4. Has explicit trigger phrase: "Use when ..." (or similar trigger pattern)
5. First sentence describes what the skill does (heuristic: at least one verb)
Outputs pass/fail per check + overall verdict.
Deterministic logic. No LLM calls. Stdlib only.
Usage:
python skill_description_validator.py # uses embedded sample
python skill_description_validator.py path/to/SKILL.md
python skill_description_validator.py path/to/SKILL.md --output json
"""
import argparse
import json
import re
import sys
from typing import Any, Dict, List, Optional
# Embedded sample: a SKILL.md description that PASSES all checks
SAMPLE_DESCRIPTION = (
"Extract text and tables from PDF files, fill forms, merge documents. "
"Use when working with PDF files or when user mentions PDFs, forms, or document extraction."
)
# Embedded sample: SKILL.md content (just the frontmatter + body shell)
SAMPLE_SKILL_MD = f"""---
name: pdf-tools
description: {SAMPLE_DESCRIPTION}
---
# PDF Tools
## Quick start
...
"""
# First-person pronouns + second-person pronouns to flag
FIRST_PERSON = {"i", "me", "my", "myself", "we", "us", "our", "ours", "ourselves"}
SECOND_PERSON = {"you", "your", "yours", "yourself"}
# Trigger phrases that count as explicit "use when" triggers
# Per Matt Pocock's rule: descriptions need an explicit trigger so agents know when to invoke.
# Natural English variants are all accepted: "Use when/before/during/after/for/while ..." etc.
TRIGGER_PATTERNS = [
re.compile(r"\buse\s+when\b", re.IGNORECASE),
re.compile(r"\buse\s+for\b", re.IGNORECASE),
re.compile(r"\buse\s+before\b", re.IGNORECASE),
re.compile(r"\buse\s+during\b", re.IGNORECASE),
re.compile(r"\buse\s+after\b", re.IGNORECASE),
re.compile(r"\buse\s+while\b", re.IGNORECASE),
re.compile(r"\binvoke\s+when\b", re.IGNORECASE),
re.compile(r"\binvoke\s+before\b", re.IGNORECASE),
re.compile(r"\binvoke\s+after\b", re.IGNORECASE),
re.compile(r"\btrigger\s+when\b", re.IGNORECASE),
re.compile(r"\bapply\s+when\b", re.IGNORECASE),
re.compile(r"\brun\s+when\b", re.IGNORECASE),
re.compile(r"\brun\s+before\b", re.IGNORECASE),
]
def extract_frontmatter(text: str) -> Dict[str, str]:
"""Extract YAML frontmatter as a flat dict. Stdlib-only — minimal YAML parser
sufficient for SKILL.md frontmatter (key: value pairs, no nesting)."""
if not text.startswith("---"):
return {}
end = text.find("\n---", 3)
if end == -1:
return {}
block = text[3:end].strip()
out: Dict[str, str] = {}
current_key: Optional[str] = None
buffer: List[str] = []
for line in block.splitlines():
if ":" in line and not line.startswith(" ") and not line.startswith("\t"):
# Flush previous
if current_key:
out[current_key] = " ".join(buffer).strip()
buffer = []
key, _, val = line.partition(":")
current_key = key.strip()
val = val.strip()
if val and val != ">":
buffer.append(val)
elif current_key and line.strip():
buffer.append(line.strip())
if current_key:
out[current_key] = " ".join(buffer).strip()
return out
def check_present(desc: str) -> Dict[str, Any]:
return {
"rule": "description_present",
"pass": bool(desc and desc.strip()),
"detail": f"Length: {len(desc)} chars" if desc else "Missing or empty description field",
}
def check_length(desc: str, max_chars: int = 1024) -> Dict[str, Any]:
n = len(desc)
return {
"rule": "description_length",
"pass": n <= max_chars,
"detail": f"{n} chars (limit {max_chars})",
}
def check_third_person(desc: str) -> Dict[str, Any]:
words = re.findall(r"\b[a-zA-Z]+\b", desc.lower())
flagged_first = [w for w in words if w in FIRST_PERSON]
flagged_second = [w for w in words if w in SECOND_PERSON]
flagged = flagged_first + flagged_second
return {
"rule": "third_person",
"pass": len(flagged) == 0,
"detail": f"Found pronouns: {sorted(set(flagged))}" if flagged else "No 1st/2nd-person pronouns",
}
def check_trigger(desc: str) -> Dict[str, Any]:
for pattern in TRIGGER_PATTERNS:
if pattern.search(desc):
return {
"rule": "explicit_trigger",
"pass": True,
"detail": f"Found trigger phrase matching: {pattern.pattern}",
}
return {
"rule": "explicit_trigger",
"pass": False,
"detail": 'No explicit trigger ("Use when..." or similar). Agent will struggle to know when to invoke.',
}
# Action verb vocabulary used to detect "first sentence describes what the skill does"
# This is content data, not an assumption — these are the verbs we look for in skill descriptions.
ACTION_VERB_VOCABULARY = (
"extract", "fill", "merge", "create", "build", "generate", "analyze", "analyse",
"validate", "check", "run", "format", "parse", "render", "review", "audit", "scan",
"compute", "score", "track", "report", "transform", "convert", "deploy", "test",
"monitor", "log", "search", "find", "fetch", "store", "send", "read", "write",
"refresh", "remove", "process", "manage", "apply", "implement", "interrogate",
"orchestrate", "classify",
)
ACTION_VERB_RE = re.compile(
r"\b(" + "|".join(ACTION_VERB_VOCABULARY) + r")s?\b",
re.IGNORECASE,
)
def check_first_sentence_has_verb(desc: str) -> Dict[str, Any]:
# Heuristic: split on first period; first sentence should have an action verb
parts = re.split(r"\.\s+", desc, maxsplit=1)
first = parts[0] if parts else desc
verbs = ACTION_VERB_RE.findall(first)
return {
"rule": "first_sentence_has_action_verb",
"pass": len(verbs) >= 1,
"detail": f"Verb(s) found in first sentence: {verbs}" if verbs else "No action verb detected in first sentence",
}
def analyze(skill_md_text: str) -> Dict[str, Any]:
fm = extract_frontmatter(skill_md_text)
desc = fm.get("description", "")
checks = [
check_present(desc),
check_length(desc),
check_third_person(desc),
check_trigger(desc),
check_first_sentence_has_verb(desc),
]
passed = sum(1 for c in checks if c["pass"])
overall = "PASS" if passed == len(checks) else ("WARN" if passed >= 3 else "FAIL")
return {
"description": desc,
"checks": checks,
"passed": passed,
"total": len(checks),
"overall": overall,
}
def render_text(r: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("SKILL DESCRIPTION VALIDATOR")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append(f"Description ({len(r['description'])} chars):")
lines.append(f" {r['description'][:200]}{'...' if len(r['description']) > 200 else ''}")
lines.append("")
lines.append("-" * 72)
lines.append(f"Checks: {r['passed']} / {r['total']} passed")
lines.append("")
for c in r["checks"]:
marker = "PASS" if c["pass"] else "FAIL"
lines.append(f" [{marker}] {c['rule']:30s} {c['detail']}")
lines.append("")
lines.append("-" * 72)
lines.append(f"Verdict: {r['overall']}")
lines.append("")
lines.append("Rules (per Matt Pocock's write-a-skill):")
lines.append(" - Max 1024 chars")
lines.append(" - Third person (no I/we/you)")
lines.append(" - First sentence: what it does (action verb)")
lines.append(" - Second sentence: 'Use when [specific triggers]'")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="Validate a SKILL.md description per Matt Pocock's rules.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to SKILL.md (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
text = f.read()
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
else:
text = SAMPLE_SKILL_MD
source = "<embedded sample: pdf-tools description (PASS expected)>"
result = analyze(text)
if args.output == "json":
print(json.dumps({"source": source, **result}, indent=2))
else:
print(render_text(result, source))
return 0 if result["overall"] == "PASS" else 1
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/skill_review_checklist_runner.py
#!/usr/bin/env python3
"""skill_review_checklist_runner.py — Run Matt Pocock's 6-item review checklist programmatically.
Stdlib-only. Combines the description-validator + structure-validator into a single
report that mirrors Matt Pocock's review checklist from write-a-skill:
1. [ ] Description includes triggers ("Use when...")
2. [ ] SKILL.md under 100 lines
3. [ ] No time-sensitive info (heuristic: no year mentions / "as of" claims / version-specific dates)
4. [ ] Consistent terminology (heuristic: no obvious synonym pairs in same doc — light check)
5. [ ] Concrete examples included (>=1 code block)
6. [ ] References one level deep
This is the canonical pre-commit check for any new skill in this repo.
Deterministic logic. No LLM calls. Stdlib only.
Usage:
python skill_review_checklist_runner.py # uses embedded sample (this skill's own folder)
python skill_review_checklist_runner.py path/to/skill-folder/
python skill_review_checklist_runner.py path/to/skill-folder/ --output json
"""
import argparse
import json
import os
import re
import sys
from typing import Any, Dict, List
# Phrases that suggest time-sensitive content
TIME_SENSITIVE_PATTERNS = [
re.compile(r"\bas\s+of\s+\d{4}\b", re.IGNORECASE),
re.compile(r"\bin\s+(20\d{2})\b", re.IGNORECASE),
re.compile(r"\b(?:january|february|march|april|may|june|july|august|september|october|november|december)\s+\d{4}\b", re.IGNORECASE),
re.compile(r"\b(?:released|launched|published|updated)\s+(?:on|in)\b", re.IGNORECASE),
]
def find_skill_md(folder: str) -> str:
candidate = os.path.join(folder, "SKILL.md")
return candidate if os.path.isfile(candidate) else ""
def extract_frontmatter_description(text: str) -> str:
"""Extract description from YAML frontmatter (single key)."""
if not text.startswith("---"):
return ""
end = text.find("\n---", 3)
if end == -1:
return ""
block = text[3:end]
# Match "description: ..." potentially spanning multiple lines (>- folded)
match = re.search(r"^description:\s*(.*)$(?:\n[ ]+(.*))*", block, re.MULTILINE)
if not match:
return ""
val = match.group(1).strip()
if val == ">" or val == "|":
# Folded scalar — collect indented continuation lines
lines_iter = iter(block.splitlines())
for line in lines_iter:
if line.strip().startswith("description:"):
break
collected = []
for line in lines_iter:
if line.startswith(" ") or line.startswith("\t"):
collected.append(line.strip())
else:
break
val = " ".join(collected)
return val
# Trigger phrases that count as explicit "use when ..." triggers in a description.
# Per Matt Pocock's rule: explicit trigger phrase. Natural English variants all accepted.
TRIGGER_PATTERNS = [
re.compile(r"\buse\s+when\b", re.IGNORECASE),
re.compile(r"\buse\s+for\b", re.IGNORECASE),
re.compile(r"\buse\s+before\b", re.IGNORECASE),
re.compile(r"\buse\s+during\b", re.IGNORECASE),
re.compile(r"\buse\s+after\b", re.IGNORECASE),
re.compile(r"\buse\s+while\b", re.IGNORECASE),
re.compile(r"\binvoke\s+when\b", re.IGNORECASE),
re.compile(r"\binvoke\s+before\b", re.IGNORECASE),
re.compile(r"\binvoke\s+after\b", re.IGNORECASE),
re.compile(r"\btrigger\s+when\b", re.IGNORECASE),
re.compile(r"\bapply\s+when\b", re.IGNORECASE),
re.compile(r"\brun\s+when\b", re.IGNORECASE),
re.compile(r"\brun\s+before\b", re.IGNORECASE),
]
def check_description_has_trigger(text: str) -> Dict[str, Any]:
desc = extract_frontmatter_description(text)
has_trigger = any(p.search(desc) for p in TRIGGER_PATTERNS)
return {
"rule": "1. Description includes triggers",
"pass": has_trigger,
"detail": ("Found explicit trigger phrase" if has_trigger
else "Missing explicit trigger phrase (Use when/before/after/for ...)"),
}
def check_skill_md_length(filepath: str, max_lines: int = 100) -> Dict[str, Any]:
with open(filepath, "r", encoding="utf-8") as f:
lines = sum(1 for _ in f)
return {
"rule": f"2. SKILL.md under {max_lines} lines",
"pass": lines <= max_lines,
"detail": f"{lines} lines",
}
def check_no_time_sensitive(text: str) -> Dict[str, Any]:
flagged = []
for pattern in TIME_SENSITIVE_PATTERNS:
for m in pattern.finditer(text):
flagged.append(m.group(0))
# Limit
flagged = list(dict.fromkeys(flagged))[:5]
return {
"rule": "3. No time-sensitive info",
"pass": len(flagged) == 0,
"detail": ("No date/year/version-bound claims detected" if not flagged
else f"Flagged phrases: {flagged}"),
}
def check_consistent_terminology(text: str) -> Dict[str, Any]:
"""Light check for common synonym mismatches in the same doc."""
synonyms = [
("agent", "bot"),
("skill", "tool"),
("user", "developer"),
]
findings = []
text_lower = text.lower()
for a, b in synonyms:
if re.search(rf"\b{re.escape(a)}\b", text_lower) and re.search(rf"\b{re.escape(b)}\b", text_lower):
findings.append(f"Both '{a}' and '{b}' used")
return {
"rule": "4. Consistent terminology",
"pass": len(findings) == 0,
"detail": ("No obvious synonym pairs detected" if not findings
else "; ".join(findings)),
}
def check_concrete_examples(text: str) -> Dict[str, Any]:
code_blocks = re.findall(r"```", text)
has_examples = len(code_blocks) >= 2 # opening + closing = 1 block
return {
"rule": "5. Concrete examples included",
"pass": has_examples,
"detail": f"{len(code_blocks) // 2} code block(s) found",
}
def _find_nested_md(refs_subdir: str) -> List[str]:
"""Return .md files nested deeper than refs_subdir."""
nested: List[str] = []
if not os.path.isdir(refs_subdir):
return nested
for root, _, files in os.walk(refs_subdir):
if root == refs_subdir:
continue
nested.extend(os.path.join(root, f) for f in files if f.endswith(".md"))
return nested
def check_references_one_level_deep(folder: str) -> Dict[str, Any]:
deeper = _find_nested_md(os.path.join(folder, "references"))
return {
"rule": "6. References one level deep",
"pass": len(deeper) == 0,
"detail": ("All references at one level" if not deeper
else f"Found nested ref files: {deeper}"),
}
def analyze(folder: str) -> Dict[str, Any]:
skill_md = find_skill_md(folder)
if not skill_md:
detail = f"SKILL.md not found at {folder}"
missing_check = {"rule": "skill_md_present", "pass": False, "detail": detail}
return {
"folder": folder,
"checks": [missing_check],
"passed": 0,
"total": 1,
"overall": "FAIL",
}
with open(skill_md, "r", encoding="utf-8") as f:
text = f.read()
checks = [
check_description_has_trigger(text),
check_skill_md_length(skill_md, max_lines=100),
check_no_time_sensitive(text),
check_consistent_terminology(text),
check_concrete_examples(text),
check_references_one_level_deep(folder),
]
passed = sum(1 for c in checks if c["pass"])
total = len(checks)
overall = "PASS" if passed == total else ("WARN" if passed >= total - 1 else "FAIL")
return {
"folder": folder,
"skill_md": skill_md,
"checks": checks,
"passed": passed,
"total": total,
"overall": overall,
}
def render_text(r: Dict[str, Any]) -> str:
lines = []
lines.append("=" * 72)
lines.append("SKILL REVIEW CHECKLIST RUNNER (per Matt Pocock's write-a-skill)")
lines.append(f"Folder: {r['folder']}")
lines.append("=" * 72)
lines.append("")
lines.append(f"Checks: {r['passed']} / {r['total']} passed")
lines.append("")
for c in r["checks"]:
marker = "[x]" if c["pass"] else "[ ]"
lines.append(f" {marker} {c['rule']}")
lines.append(f" {c['detail']}")
lines.append("")
lines.append("-" * 72)
lines.append(f"Verdict: {r['overall']}")
lines.append("")
lines.append("Reference: Matt Pocock's 6-item review checklist from write-a-skill (MIT).")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="Run Matt Pocock's 6-item review checklist on a skill folder.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to skill folder (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
folder = args.path
else:
folder = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
if not os.path.isdir(folder):
print(f"error: not a directory: {folder}", file=sys.stderr)
return 1
result = analyze(folder)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_text(result))
return 0 if result["overall"] == "PASS" else 1
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/skill_structure_validator.py
#!/usr/bin/env python3
"""skill_structure_validator.py — Validate a skill folder structure against Matt Pocock's pattern.
Stdlib-only. Walks a skill folder and checks:
1. SKILL.md present at folder root
2. SKILL.md <= 100 lines (Matt's ceiling; configurable via --max-lines)
3. If SKILL.md > limit, separate reference files exist (REFERENCE.md, EXAMPLES.md, or references/*.md)
4. Reference files are one level deep (no nested references in subfolders)
5. No circular cross-references between markdown files (file A links to B which links back to A)
6. Scripts present in scripts/ subfolder when SKILL.md mentions executable operations
Deterministic logic. No LLM calls. Stdlib only.
Usage:
python skill_structure_validator.py # uses embedded sample (current write-a-skill folder)
python skill_structure_validator.py path/to/skill-folder/
python skill_structure_validator.py path/to/skill-folder/ --output json
python skill_structure_validator.py path/to/skill-folder/ --max-lines 100
"""
import argparse
import json
import os
import re
import sys
from typing import Any, Dict, List, Set, Tuple
# Default max-lines threshold from Matt Pocock's write-a-skill review checklist
DEFAULT_MAX_LINES = 100
# Reference filename patterns Matt's pattern recognizes
REFERENCE_FILE_PATTERNS = ["REFERENCE.md", "EXAMPLES.md", "references", "examples"]
# Script folder names
SCRIPT_FOLDERS = ["scripts"]
def find_skill_md(folder: str) -> str:
"""Find SKILL.md at folder root; return its path or empty string."""
candidate = os.path.join(folder, "SKILL.md")
if os.path.isfile(candidate):
return candidate
return ""
def count_lines(filepath: str) -> int:
with open(filepath, "r", encoding="utf-8") as f:
return sum(1 for _ in f)
def _list_md_in_subdir(subdir: str) -> List[str]:
"""List .md files directly inside a subdirectory (not recursive)."""
out: List[str] = []
if not os.path.isdir(subdir):
return out
for name in sorted(os.listdir(subdir)):
full = os.path.join(subdir, name)
if os.path.isfile(full) and name.endswith(".md"):
out.append(full)
return out
def find_reference_files(folder: str) -> List[str]:
"""Find reference files at folder root + one-level-deep references/ subfolder."""
refs: List[str] = []
for name in os.listdir(folder):
full = os.path.join(folder, name)
if os.path.isfile(full) and name.endswith(".md") and name != "SKILL.md":
refs.append(full)
elif os.path.isdir(full) and name in ("references", "examples"):
refs.extend(_list_md_in_subdir(full))
return refs
def find_deeper_references(folder: str) -> List[str]:
"""Find markdown files nested deeper than one level (violation of one-level-deep rule)."""
deeper: List[str] = []
refs_subdir = os.path.join(folder, "references")
if not os.path.isdir(refs_subdir):
return deeper
for root, _, files in os.walk(refs_subdir):
if root == refs_subdir:
continue
for f in files:
if f.endswith(".md"):
deeper.append(os.path.join(root, f))
return deeper
def has_scripts_folder(folder: str) -> bool:
return os.path.isdir(os.path.join(folder, "scripts"))
def extract_md_links(text: str) -> List[str]:
"""Extract local markdown links: [...](path.md), excluding URLs."""
pattern = re.compile(r"\[[^\]]+\]\(([^)]+\.md(?:#[^)]*)?)\)")
links = []
for m in pattern.finditer(text):
target = m.group(1).split("#", 1)[0]
if not target.startswith("http"):
links.append(target)
return links
def _collect_links_for_file(filepath: str, files: List[str]) -> Set[str]:
"""Read filepath, return set of links that resolve to other files in `files`."""
out: Set[str] = set()
try:
with open(filepath, "r", encoding="utf-8") as fh:
text = fh.read()
except (IOError, OSError):
return out
for link in extract_md_links(text):
target = os.path.normpath(os.path.join(os.path.dirname(filepath), link))
if target in files:
out.add(target)
return out
def detect_circular_refs(folder: str, files: List[str]) -> List[Tuple[str, str]]:
"""Detect circular references: file A -> file B -> file A.
Returns list of (file_a, file_b) tuples."""
graph: Dict[str, Set[str]] = {f: _collect_links_for_file(f, files) for f in files}
seen_pairs: Set[Tuple[str, str]] = set()
circular: List[Tuple[str, str]] = []
for a, neighbors in graph.items():
for b in neighbors:
if a not in graph.get(b, set()):
continue
pair = tuple(sorted([a, b]))
if pair in seen_pairs:
continue
seen_pairs.add(pair)
circular.append((a, b))
return circular
def analyze(folder: str, max_lines: int) -> Dict[str, Any]:
folder = folder.rstrip("/")
findings: List[Dict[str, Any]] = []
skill_md = find_skill_md(folder)
if not skill_md:
findings.append({
"rule": "skill_md_present",
"pass": False,
"detail": f"SKILL.md not found at {folder}",
})
return {"folder": folder, "checks": findings, "passed": 0, "total": 1, "overall": "FAIL"}
findings.append({
"rule": "skill_md_present",
"pass": True,
"detail": skill_md,
})
lines = count_lines(skill_md)
skill_md_under_ceiling = lines <= max_lines
findings.append({
"rule": "skill_md_line_count",
"pass": skill_md_under_ceiling,
"detail": f"{lines} lines (limit {max_lines})",
})
refs = find_reference_files(folder)
if not skill_md_under_ceiling:
# When SKILL.md exceeds ceiling, reference files SHOULD exist
findings.append({
"rule": "reference_files_when_split_needed",
"pass": len(refs) > 0,
"detail": f"Found {len(refs)} reference file(s)" if refs
else "SKILL.md exceeds ceiling but no reference files present",
})
else:
findings.append({
"rule": "reference_files_when_split_needed",
"pass": True,
"detail": "SKILL.md under ceiling; reference split not required",
})
deeper = find_deeper_references(folder)
findings.append({
"rule": "references_one_level_deep",
"pass": len(deeper) == 0,
"detail": f"Found nested ref files (violations): {deeper}" if deeper
else "All references are one level deep (or at root)",
})
all_md = [skill_md] + refs
circular = detect_circular_refs(folder, all_md)
findings.append({
"rule": "no_circular_references",
"pass": len(circular) == 0,
"detail": f"Circular refs detected: {circular}" if circular
else "No circular references between markdown files",
})
has_scripts = has_scripts_folder(folder)
findings.append({
"rule": "scripts_folder_present",
"pass": True,
"detail": "scripts/ folder exists" if has_scripts
else "No scripts/ folder (optional per Matt's pattern)",
})
passed = sum(1 for c in findings if c["pass"])
overall = "PASS" if passed == len(findings) else ("WARN" if passed >= len(findings) - 1 else "FAIL")
return {
"folder": folder,
"max_lines_threshold": max_lines,
"skill_md": skill_md,
"skill_md_lines": lines,
"reference_files": refs,
"checks": findings,
"passed": passed,
"total": len(findings),
"overall": overall,
}
def render_text(r: Dict[str, Any]) -> str:
lines = []
lines.append("=" * 72)
lines.append("SKILL STRUCTURE VALIDATOR")
lines.append(f"Folder: {r['folder']}")
lines.append(f"Max-lines threshold: {r['max_lines_threshold']}")
lines.append("=" * 72)
lines.append("")
lines.append(f"SKILL.md: {r.get('skill_md', '<missing>')} ({r.get('skill_md_lines', 0)} lines)")
lines.append(f"Reference files: {len(r.get('reference_files', []))}")
lines.append("")
lines.append("-" * 72)
lines.append(f"Checks: {r['passed']} / {r['total']} passed")
lines.append("")
for c in r["checks"]:
marker = "PASS" if c["pass"] else "FAIL"
lines.append(f" [{marker}] {c['rule']:35s} {c['detail']}")
lines.append("")
lines.append("-" * 72)
lines.append(f"Verdict: {r['overall']}")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="Validate skill folder structure per Matt Pocock's write-a-skill pattern.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
max_lines_help = f"SKILL.md line ceiling (default: {DEFAULT_MAX_LINES} per Matt's rule)"
parser.add_argument("path", nargs="?", help="Path to skill folder (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
parser.add_argument("--max-lines", type=int, default=DEFAULT_MAX_LINES, help=max_lines_help)
args = parser.parse_args()
if args.path:
folder = args.path
else:
# Embedded sample: validate this skill's own folder
folder = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
if not os.path.isdir(folder):
print(f"error: not a directory: {folder}", file=sys.stderr)
return 1
result = analyze(folder, args.max_lines)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_text(result))
return 0 if result["overall"] == "PASS" else 1
if __name__ == "__main__":
sys.exit(main())
Tạo bản tóm tắt chiến lược một trang từ buổi office hours, bước đầu của quy trình sprint chiến lược.
---
name: "brief"
description: "/cs:brief <topic> — Generate a one-page strategy brief from an office-hours intake. First step in the strategic sprint pipeline."
---
# /cs:brief — One-Page Strategy Brief
**Command:** `/cs:brief <topic>` or `/cs:brief <office-hours-output>`
Turns intake (raw question or office-hours output) into a one-page strategy brief that the boardroom can deliberate on. This is **Step 1** of the strategic sprint pipeline.
## Pipeline Position
```
/cs:office-hours → /cs:brief → /cs:boardroom → /cs:decide → /cs:execute → /cs:post-mortem
↑ you are here
```
## Inputs
- A topic string, **or**
- An office-hours brief (preferred — more rigor)
- `~/.claude/company-context.md` (loaded automatically)
## Output
A single Markdown file under `~/.claude/briefs/YYYY-MM-DD-<slug>.md` with this structure:
```markdown
# Strategy Brief: <topic>
**Date:** YYYY-MM-DD
**Author:** cs-chief-of-staff
**Status:** DRAFT | UNDER REVIEW | APPROVED | RETIRED
## Context
[1-2 paragraphs: where the company sits today on this topic — pulled from company-context.md]
## Question
[The one sentence question the boardroom must answer]
## Options
1. **Option A:** <name> — <one-sentence summary>
2. **Option B:** <name> — <one-sentence summary>
3. **Option C:** <name> — <one-sentence summary>
(Minimum 2 options. "Do nothing" is always an option.)
## Assumptions
- <assumption 1 — explicit>
- <assumption 2>
- <assumption 3>
## Constraints
- Time: <by when must this decide>
- Money: <budget envelope>
- People: <who can / can't be reallocated>
- Reversibility: <one-way door | two-way door>
## Affected Roles
[Which cs-* advisors should weigh in. Used to route to /cs:boardroom panel composition.]
- [ ] cs-ceo-advisor
- [ ] cs-cfo-advisor
- [ ] cs-cto-advisor
- [ ] cs-cmo-advisor
- [ ] cs-cro-advisor
- [ ] cs-cpo-advisor
- [ ] cs-coo-advisor
- [ ] cs-chro-advisor
- [ ] cs-ciso-advisor
- [ ] cs-chief-of-staff
## Success Criteria
[Measurable outcomes that define success — set BEFORE the decision]
- <metric 1, threshold, timeframe>
- <metric 2, threshold, timeframe>
## Kill Criteria
[What signal would tell you in 90 days that this was the wrong call]
- <metric, threshold, action if missed>
```
## Workflow
1. Load company-context.md via context-engine
2. If input is office-hours output, parse the 6 answers
3. If input is a raw topic, prompt the founder for the missing pieces
4. Draft 2-3 options (never just one — every brief needs a counterfactual)
5. Make assumptions and constraints explicit
6. Identify affected roles → drives panel composition for `/cs:boardroom`
7. Write success + kill criteria BEFORE the decision (this is the rigor moment)
8. Save to `~/.claude/briefs/`
## Why This Step Exists
The biggest decision-making failure is debating implementation before agreeing on the question. The brief locks the question, options, and success criteria so the boardroom can deliberate without scope creep.
This is also the **artifact handoff** — the next command consumes this file, not your memory.
## Routing
- `/cs:boardroom <brief>` — multi-role deliberation
- `/cs:cross-eval <brief>` — multi-model sanity check before boardroom (for high-stakes)
- `/cs:freeze <brief>` — cooldown lock for irreversible decisions
## Related
- Agent: [`cs-chief-of-staff`](../../agents/cs-chief-of-staff.md)
- Skills: [`context-engine`](../../../skills/context-engine/SKILL.md), [`board-meeting`](../../../skills/board-meeting/SKILL.md)
---
**Version:** 1.0.0
Kiểm tra và tối ưu nội dung theo E-E-A-T để được các LLM như ChatGPT, Perplexity, Claude trích dẫn, theo dõi bằng sổ ghi cục bộ.
---
name: "cs-aeo"
description: "/cs:aeo — Answer Engine Optimization workflow. Audit content for E-E-A-T + structure signals that drive LLM citation (ChatGPT, Perplexity, Claude, Gemini, Mistral). Optimize content in 3 modes (conservative/balanced/aggressive). Track which LLMs cite which pages via local ledger. Industry-aware thresholds (8 industries with YMYL calibration). Distinct from SEO — refuses to optimize one at expense of the other."
---
# /cs:aeo — Answer Engine Optimization
**Command:** `/cs:aeo [action] [args]`
The `cs-aeo` command is the **entry point for AEO workflows**: audit → optimize → publish → track citations.
## Distinct From `/cs:seo-audit`
These share a foundation (E-E-A-T) but optimize for different conversion events:
- **`/cs:seo-audit`** — optimizes for ranking + click-through in Google/Bing search results
- **`/cs:aeo`** (this command) — optimizes for being cited as authoritative source by LLMs
They can run on the same content. The cs-aeo agent will surface this and recommend running both for high-leverage pages.
## When To Run
- Auditing existing content for AI-search readiness (E-E-A-T + structure signals)
- Optimizing a page for LLM citation before publishing
- Tracking which LLMs cite which pages over time (citation ledger)
- Researching whether AEO investment is worth it for a given content piece
- Benchmarking against competitor citation rates
## When NOT To Run
- Pure click-through SEO without AI-citation intent → use `/cs:seo-audit`
- Brand-voice content with no factual claims (citations require facts)
- Time-sensitive news (LLM training lag means citation comes months later)
- Topics where LLMs already have strong training (e.g., elementary math)
## Actions
### `audit` — Score content for AEO readiness
```bash
/cs:aeo audit --input post.md --industry saas
/cs:aeo audit --url https://example.com/blog/post --industry healthcare
/cs:aeo audit --sample
```
Returns composite 0-100 with per-dimension breakdown (E-E-A-T + Structure) and top 5 fixes in priority order.
### `optimize` — Generate AEO-improved variant
```bash
/cs:aeo optimize --input post.md --mode balanced --output post-aeo.md
/cs:aeo optimize --input post.md --mode aggressive --industry finance
```
Three modes:
- `conservative` — touch <10% of words (schema + corrections footer only)
- `balanced` — touch <30% (citation markers + heading restructure + schema + footer)
- `aggressive` — full restructure + fact-first lede + maximum citation density
### `track` — Log a citation you observed in an LLM response
```bash
/cs:aeo track --url https://example.com/post --llm perplexity --query "what is AEO" --date 2026-05-17
```
Maintains a local ledger at `~/.aeo-data/citations.json`. No telemetry.
### `report` — Aggregate citation report for a URL
```bash
/cs:aeo report --url https://example.com/post
```
Returns total citations, LLM coverage, velocity, top queries, verdict (EARLY / EMERGING / STRONG).
### `export` — Emit citation ledger as CSV
```bash
/cs:aeo export --output citations.csv
```
For reporting to clients / stakeholders.
## Minimal Intake (3 Questions)
| Q | Asks | When |
|---|---|---|
| Q1 | What action — audit / optimize / track / report? | Always |
| Q2 | Industry (saas / healthcare / finance / legal / ecommerce / b2b / media / education) | Always (calibrates thresholds) |
| Q3 | For `optimize`: mode (conservative / balanced / aggressive)? | Only when action=optimize |
Most invocations exit intake after Q2.
## Workflow
```bash
# Phase 1: Audit
python3 marketing-skill/skills/aeo/scripts/aeo_audit.py --input <file> --industry <industry>
# → composite score 0-100 + top fixes
# Phase 2: Optimize (if audit < industry threshold)
python3 marketing-skill/skills/aeo/scripts/aeo_optimizer.py \
--input <file> --mode <mode> --industry <industry> --output <file>-aeo.md
# → optimized variant + changelog
# Phase 3: Publish (manual step — review the optimized variant, then deploy)
# Phase 4: Track (over 4-12 weeks)
python3 marketing-skill/skills/aeo/scripts/citation_tracker.py \
--action add --url <url> --llm <llm> --query <query> --date <YYYY-MM-DD>
# → ledger updated
# Phase 5: Report (monthly)
python3 marketing-skill/skills/aeo/scripts/citation_tracker.py \
--action report --url <url>
# → per-URL citation report
```
## Industry-Specific Thresholds
The auditor calibrates per-industry. YMYL ("Your Money or Your Life") topics use stricter thresholds:
| Industry | Min Composite | Why |
|---|---|---|
| Healthcare | 85 | Direct health implications |
| Finance | 85 | Real financial decisions |
| Legal | 85 | Legal jeopardy if misapplied |
| Education | 75 | Learning outcomes |
| SaaS, B2B, Media | 70 | Business decisions, moderate stakes |
| E-commerce | 65 | Product reviews, lower individual risk |
Content for YMYL topics scoring below threshold is unlikely to be cited regardless of other signals — the cs-aeo agent will flag this and refuse aggressive optimization until the foundational dimensions improve.
## Anti-Patterns Rejected
- LLM-generated AEO content with no human review (RAG retrieval deprioritizes generic LLM output)
- Fabricated credentials in author bylines (LLMs cross-reference via LinkedIn/Wikipedia)
- Schema spam (false structured-data markup gets filtered)
- Authority laundering (linking out doesn't confer authority)
- Per-LLM optimization tunnel-vision (73% cross-LLM citation correlation — optimize for shared signals)
- Optimizing AEO at expense of SEO (and vice versa) — they complement, don't substitute
## Trigger Phrases
- "AEO audit"
- "optimize for ChatGPT / Perplexity / Claude / Gemini"
- "get cited by [LLM]"
- "LLM citation strategy"
- "answer engine optimization"
- "E-E-A-T audit"
- "content for AI search"
- "track AI citations"
- "schema for AI"
## Related
- Agent: [`cs-aeo`](../agents/cs-aeo.md)
- Skill: [`aeo`](../skills/aeo/SKILL.md)
- Companion: `/cs:seo-audit` (SEO + AEO often run together)
- Source: ported from [`alirezarezvani/aeo-box`](https://github.com/alirezarezvani/aeo-box)
---
**Version:** 2.7.3
**License:** MIT
Thêm, gỡ bỏ và kiểm tra feature flag: kế hoạch rollout, kill switch, phát hiện flag cũ và các câu hỏi về triển khai tiến dần.
---
name: feature-flags-architect
description: Use when adding, retiring, or auditing feature flags. Triggers on "add a flag", "ship behind a flag", "rollout plan", "kill switch", "stale flags", "flag debt", "LaunchDarkly", "GrowthBook", "Statsig", "Unleash", "Flipt", or any progressive-delivery question. Ships flag debt scanner, rollout planner, and kill-switch auditor (all stdlib Python), 4 references on flag taxonomy + provider trade-offs + rollout strategies + lifecycle, plus a /flag-cleanup slash command.
context: fork
version: 2.9.0
author: claude-code-skills
license: MIT
tags: [feature-flags, progressive-delivery, rollout, kill-switch, launchdarkly, growthbook, statsig, unleash, flipt, release-engineering]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# Feature Flags Architect
End-to-end discipline for feature flags: classify them, ship them, ramp them, and retire them. Most teams treat flags as throwaway `if`-statements; this skill treats them as a controlled lifecycle with measurable debt.
## When to use
- Adding a new flag and need a rollout plan
- Auditing a codebase for stale or orphaned flags
- Choosing a flag provider (LaunchDarkly vs GrowthBook vs Statsig vs Unleash vs Flipt vs build-your-own)
- Designing a kill-switch path for a risky launch
- Cleaning up flag debt before a release freeze
- Reviewing whether a feature should ship behind a flag at all
## Core principle: flags are a lifecycle, not an `if`
```
request → design → ship → ramp → cleanup → archive
```
Flags that skip cleanup become debt: dead branches, stale defaults, untested code paths, unbounded blast radius. The three scripts in this skill enforce the lifecycle.
## Quick start
```bash
# 1. Audit the repo for flag debt
python scripts/flag_debt_scanner.py --repo . --max-age-days 90
# 2. Plan a progressive rollout for a new flag
python scripts/rollout_planner.py --population 100000 --target-percent 100 --duration-days 14 --strategy ring
# 3. Verify every flag has a documented kill switch
python scripts/kill_switch_audit.py --repo . --flag-doc docs/feature-flags.md
```
## The 4 flag types (taxonomy)
Different flag types have different lifespans and ownership. Misclassifying creates debt.
| Type | Purpose | Typical lifespan | Owner | Cleanup trigger |
|---|---|---|---|---|
| **Release** | Hide unfinished features in production | days–weeks | Eng | 100% rollout reached |
| **Experiment** | A/B test variants | weeks | Product/Marketing | Test concluded; winner picked |
| **Operational** | Circuit breakers, perf toggles, kill switches | months–years | Eng/SRE | Replaced by autoscaling/feature retirement |
| **Permission** | Entitlements per user/account/plan | years (permanent) | Product | Plan/role removed |
Only Release and Experiment flags should be on a debt-scanner watchlist. Operational and Permission flags are by design long-lived. See `references/flag_taxonomy.md` for decision tree.
## The 3 Python tools
All three are stdlib-only. Run with `--help`.
### `flag_debt_scanner.py`
Finds flags older than `--max-age-days` with low usage, suggesting candidates for cleanup.
```bash
python scripts/flag_debt_scanner.py --repo . --max-age-days 90 --format text
python scripts/flag_debt_scanner.py --repo . --max-age-days 60 --format json > debt.json
```
**Detection heuristic:**
1. Walk `--repo` for code references matching common flag-call patterns:
- `flag("...")`, `isFlagEnabled("...")`, `featureFlag("...")`, `getFlag("...")`
- `client.variation("...", ...)`, `unleash.isEnabled("...")`, `growthbook.feature("...")`
2. For each unique flag identifier, find the oldest commit that introduced it (`git log --diff-filter=A -S <name>`).
3. Flag as DEBT if introduced > `--max-age-days` ago AND used in ≤`--min-uses` places.
Outputs flag name, age in days, file references, suggested action. JSON mode is CI-friendly.
### `rollout_planner.py`
Generates a phased rollout schedule from population size, target percent, duration, and strategy.
```bash
python scripts/rollout_planner.py --population 100000 --target-percent 100 --duration-days 14 --strategy ring
python scripts/rollout_planner.py --population 50000 --target-percent 25 --duration-days 7 --strategy linear
python scripts/rollout_planner.py --population 1000000 --target-percent 100 --duration-days 30 --strategy log
```
**Strategies:**
- `ring`: 1% → 5% → 25% → 50% → 100%, evenly spaced. Default for risky launches.
- `linear`: constant rate per day. Default for medium-risk.
- `log`: rapid early, slow tail. Default for low-risk launches with confidence.
- `cohort`: by named cohort (internal → beta → free → paid → all).
Outputs a markdown table with date, percent, expected user count, abort criteria, and verification step per phase.
### `kill_switch_audit.py`
Cross-references code-discovered flags against documentation to verify each has a kill switch path written down.
```bash
python scripts/kill_switch_audit.py --repo . --flag-doc docs/feature-flags.md
python scripts/kill_switch_audit.py --repo . --flag-doc runbooks/flags.md --format json
```
**What it checks:**
1. Every code-discovered flag has an entry in `--flag-doc`
2. Each entry declares: owner, type, kill-switch trigger, monitoring dashboard
3. Reports flags missing documentation (FAIL) or missing fields (WARN)
Use as a pre-merge gate before any new flag ships.
## Provider chooser (5 + DIY)
| Provider | Best for | Pricing model | Lock-in risk | OSS option |
|---|---|---|---|---|
| **LaunchDarkly** | Enterprise, complex targeting, audit/compliance | Per-MAU, expensive | High | No |
| **GrowthBook** | Mid-market, A/B testing focused, OSS-friendly | Per-MAU + OSS | Low | Yes (self-host) |
| **Statsig** | Growth/product teams, advanced experimentation | Free tier + per-MAU | Medium | No |
| **Unleash** | OSS-first, self-hosted, dev-friendly | OSS + Enterprise | Low | Yes |
| **Flipt** | Lightweight, k8s-native, simple needs | OSS-only | None | Yes |
| **DIY** | <100 flags, no targeting, full control | None | None | N/A |
Decision rules:
- <50 flags + no targeting → DIY with config file or env vars
- Need analytics + experimentation → Statsig or GrowthBook
- Compliance/SOC2 audit logs required → LaunchDarkly
- Self-hosting required (data residency / air-gapped) → Unleash or Flipt
- See `references/provider_comparison.md` for detail.
## Workflows
### Workflow 1: Ship a new feature behind a flag
```
1. Classify: which of the 4 flag types?
→ Release (most common for engineering work)
2. Run rollout_planner.py to design the ramp
3. Add flag entry to docs/feature-flags.md BEFORE writing code:
- name, owner, type, kill-switch trigger, dashboard URL
4. Write the code with the flag
5. Run kill_switch_audit.py — must pass before merge
6. Deploy at 0%; verify kill switch works
7. Execute rollout schedule; abort if abort criteria met
8. At 100% for 7+ days: remove flag, delete dead branch, archive doc entry
```
### Workflow 2: Quarterly flag cleanup
```
1. Run flag_debt_scanner.py --repo . --max-age-days 90 > debt.md
2. For each flagged item:
a. Confirm it reached 100% (or was killed)
b. Find the issue/PR that introduced it; verify owner agrees to remove
c. Delete dead branches; remove flag config
d. Run kill_switch_audit.py — should now show one fewer flag
3. Update CHANGELOG: "Removed N stale flags"
```
### Workflow 3: Choose a provider
```
1. Estimate flag count (current + 12-month projection)
2. Required features:
- Targeting rules (user, account, geo, %)?
- A/B testing + stats?
- Audit log / SOC2?
- Self-hosting / data residency?
3. Pricing budget (MAU * cost-per-MAU)
4. See provider_comparison.md decision tree
5. Build a 30-day proof-of-concept before signing
```
### Workflow 4: Design a kill switch
```
1. Identify the failure modes:
- Latency spike (which threshold?)
- Error rate spike (which threshold?)
- Business metric regression (which threshold?)
2. Wire each to an abort:
- Manual: dashboard link + on-call playbook
- Automated: alert threshold flips flag back to 0%
3. Test the kill switch in staging BEFORE production rollout
4. Document in flag-doc; pass kill_switch_audit.py
```
## References
- `references/flag_taxonomy.md` — 4 types, decision tree, ownership, lifespan
- `references/provider_comparison.md` — LaunchDarkly / GrowthBook / Statsig / Unleash / Flipt / DIY trade-offs
- `references/rollout_strategies.md` — ring / linear / log / cohort / geo, abort criteria, monitoring
- `references/flag_lifecycle.md` — request → design → ship → ramp → cleanup → archive
## Slash command
`/flag-cleanup` — Run the full cleanup workflow on the current repo: scan for debt, generate a removal plan, audit kill switches.
## Asset templates
- `assets/flag_request_template.md` — fill-in form for new flag requests (name, owner, type, kill switch, rollout plan)
## Anti-patterns
- **Permanent flag with `if (FLAG_FOO)` 50 places** — should be a Permission flag with a runtime config, not a Release flag
- **Flag with no owner** — when the original engineer leaves, no one cleans it up
- **No kill switch documented** — when the feature breaks, no one knows how to disable it
- **A/B test that ran 6 months** — pick a winner; running indefinitely is debt
- **Flags as feature toggles for cosmetic changes** — ship via deploy, not flag
## Verifiable success
A team using this skill should achieve:
- 100% of new flags pass `kill_switch_audit.py` at merge time
- `flag_debt_scanner.py --max-age-days 90` returns ≤5 stale flags repo-wide
- Every flag has a documented owner, type, and kill switch
- Mean time to retire a Release flag: <60 days from 100% rollout
FILE:assets/flag_request_template.md
# Feature flag request
Fill in every section before opening a PR that adds the flag.
## Basics
- **Name:** `<kebab-case-flag-name>` (e.g., `new-checkout-flow`)
- **Owner:** `<your-handle@team>`
- **Type:** [ ] Release [ ] Experiment [ ] Operational [ ] Permission
- **Created:** `<YYYY-MM-DD>`
- **Expected cleanup:** `<YYYY-MM-DD or "permanent">`
## Justification
> Why a flag and not a direct deploy?
(Examples: risky launch, A/B test, kill-switch needed, gradual rollout, compliance requirement)
## Rollout plan
> Generated by `rollout_planner.py`. Paste output below.
```
<paste output here>
```
## Kill switch
- **Trigger:** `<concrete signal that flips the flag back to 0%>`
- **Threshold:** `<numeric threshold>`
- **Method:** [ ] Manual via dashboard URL [ ] Automated via alert webhook
- **Runbook:** `<link to on-call runbook>`
## Monitoring
- **Dashboard:** `<URL>`
- **Key metrics to watch:**
- `<metric 1>` baseline: `<value>`, abort threshold: `<value>`
- `<metric 2>` baseline: `<value>`, abort threshold: `<value>`
## Code locations
- **Decision point:** `<file:line>` (single point of conditional)
- **Provider used:** `<LaunchDarkly | GrowthBook | Statsig | Unleash | Flipt | DIY>`
- **SDK:** `<sdk version / config file path>`
## Tests
- [ ] Test for ON branch
- [ ] Test for OFF branch
- [ ] Kill-switch test in staging (verify flag flip works)
## Cleanup criteria
> When can this flag be removed?
(Example: at 100% rollout for ≥7 days with no incidents)
## Pre-merge checklist
- [ ] `kill_switch_audit.py` passes
- [ ] flag-doc entry added with all required fields
- [ ] PR description links to this template
- [ ] Owner has write access to the provider dashboard
- [ ] Abort criteria are concrete numbers, not vague
FILE:references/flag_lifecycle.md
# Flag lifecycle
Every flag passes through 6 phases. Skipping any phase creates debt.
```
request → design → ship → ramp → cleanup → archive
```
## Phase 1: Request
Triggered by an engineer or PM identifying a need.
**Required:**
- Flag name (kebab-case, descriptive: `new-checkout-flow` not `flag1`)
- Owner (named individual; not a team)
- Type (Release / Experiment / Operational / Permission)
- Justification (why a flag, not direct deploy?)
- Expected lifespan (days for Release, weeks for Experiment)
**Tool:** `assets/flag_request_template.md`
**Reject the request if:**
- It's a cosmetic change with no risk → ship via deploy
- It has no clear cleanup criteria → not a flag, refactor instead
- It duplicates an existing flag → reuse
## Phase 2: Design
Before writing code. Document decisions.
**Required artifacts:**
- Entry in `docs/feature-flags.md` (or your flag registry) with: name, owner, type, kill switch, dashboard URL
- Rollout plan generated by `rollout_planner.py`
- Kill-switch trigger and runbook
- Abort criteria with concrete thresholds
**Code location:**
- Single point of decision (not 5 `if (flag)` scattered)
- Use a strategy/feature-toggle pattern at module boundary
```python
# Good: one decision at module entry
if flags.is_enabled("new-checkout"):
return new_checkout(request)
return legacy_checkout(request)
# Bad: flag check scattered through the function
def checkout(request):
if flags.is_enabled("new-checkout"):
validate_v2(request)
else:
validate_v1(request)
if flags.is_enabled("new-checkout"):
format_v2(request)
else:
format_v1(request)
# ... many more
```
## Phase 3: Ship
Deploy with flag at **0% in production**, **100% in dev/staging**.
**Verification before merge:**
- [ ] `kill_switch_audit.py` passes
- [ ] Both branches (on/off) covered by tests
- [ ] Provider dashboard shows the flag at 0%
- [ ] Kill switch tested in staging (flip to ON, observe; flip to OFF, observe)
- [ ] Monitoring dashboard linked from flag-doc entry
**Common shipping mistakes:**
- Default-to-true in production (skip the safety wheels)
- Test only the new path; assume the old path still works
- Forget to update the flag-doc
## Phase 4: Ramp
Execute the rollout plan from `rollout_planner.py`. Hold each phase per `rollout_strategies.md`.
**Decision points:**
- After each phase: check abort criteria → hold | rollback | advance
- Communicate progress in team channel
- Update flag-doc with current percent and any abort events
## Phase 5: Cleanup
Once at 100% (or experiment concluded with a winner picked), remove the flag.
**Cleanup checklist:**
- [ ] Flag at 100% for ≥7 days (Release flags) OR test concluded (Experiment)
- [ ] Owner confirms no rollback risk
- [ ] Code change: delete the conditional, keep the new branch, delete the old branch
- [ ] Delete the flag in the provider dashboard
- [ ] Mark the flag-doc entry as ARCHIVED with date and PR link
- [ ] Add to CHANGELOG: "Removed feature flag: <name>"
**Common cleanup mistakes:**
- Removing the flag from code but forgetting the provider config (orphaned)
- Removing both branches (keep the new one)
- Not updating flag-doc (audit trail lost)
- Not running tests after removal (latent break)
## Phase 6: Archive
Move the flag-doc entry to an archive section. Keep the audit trail.
```markdown
## Archived
### new-checkout-flow [removed 2026-04-12, PR #1234]
- Owner: jane@team
- Type: Release
- Lifespan: 38 days from request to removal
- Outcome: Shipped at 100%; no incidents
```
## Lifecycle automation
| Phase | Tool / process |
|---|---|
| Request | `flag_request_template.md` filled in PR description |
| Design | `rollout_planner.py` output committed to PR |
| Ship | `kill_switch_audit.py` as pre-merge CI gate |
| Ramp | Provider dashboard execution; abort wired to alerts |
| Cleanup | Quarterly run of `flag_debt_scanner.py` |
| Archive | Manual (engineer cleanup PR) |
## SLAs by phase
| Phase | Max duration | Trigger if exceeded |
|---|---|---|
| Request → Design | 7 days | Owner ping |
| Design → Ship | 30 days | Owner ping; close request if stale |
| Ship → Ramp start | 7 days | Owner ping |
| Ramp → 100% (Release) | 30 days | Pause, review |
| 100% → Cleanup | 30 days | `flag_debt_scanner.py` flags it |
| Cleanup → Archive | 7 days | PR review reminder |
## Worked example
**Day 0:** Engineer files request: `new-search-relevance` Release flag, owner @bob, expected 21-day rollout.
**Day 2:** Design done. flag-doc entry created. `rollout_planner.py` output: ring strategy, 5 rings over 14 days. Kill-switch: any drop in CTR > 5%, set flag to 0% via provider API.
**Day 4:** Code shipped, flag at 0%. `kill_switch_audit.py` green. Smoke test passes.
**Day 5:** Ring 1 — 1% rollout. CTR within bounds. Hold 48h.
**Day 7:** Ring 2 — 5%. p99 latency +5% (within bounds). Hold 48h.
**Day 9:** Ring 3 — 25%. CTR +2% — winning. Hold 48h.
**Day 11:** Ring 4 — 50%. CTR +2.5%. Hold 48h.
**Day 13:** Ring 5 — 100%. Hold 7 days for stability.
**Day 20:** Cleanup PR opens — remove conditional, delete old branch.
**Day 21:** PR merged. Flag deleted in provider. flag-doc entry archived.
**Total elapsed: 21 days.** This is the target.
## When the lifecycle breaks
| Symptom | Diagnosis | Fix |
|---|---|---|
| Flag at 100% in code 6+ months | Cleanup phase skipped | Run `flag_debt_scanner.py` quarterly |
| Flag has no owner | Owner left; not reassigned | Assign to team's tech-debt owner; cleanup or transfer in 30 days |
| Two flags doing the same thing | Request phase missed dedup check | Consolidate; archive duplicate |
| Flag-doc entry missing | Design phase skipped | `kill_switch_audit.py` must be a CI gate |
| Flag flipped without rollout plan | Ramp phase skipped | Treat as incident; review cause |
FILE:references/flag_taxonomy.md
# Flag taxonomy — the 4 types
Misclassifying a flag is the root cause of flag debt. Pick one type at the moment you create the flag.
## Decision tree
```
Is the flag intended to be permanent (entitlement, plan tier, role-based access)?
├── YES → Permission flag
└── NO → Will it eventually be removed?
├── Will it be removed when feature is fully shipped?
│ └── Yes → Release flag
├── Will it be removed when an A/B test concludes?
│ └── Yes → Experiment flag
└── Will it remain as a circuit breaker / safety toggle?
└── Yes → Operational flag
```
## 1. Release flag
**Purpose:** Hide an unfinished or risky feature in production while it's being built or rolled out.
| Property | Value |
|---|---|
| Lifespan | Days to weeks (≤90 days target) |
| Default | OFF in prod, ON in dev/staging |
| Owner | Engineer who created it |
| Cleanup trigger | Reached 100% rollout AND stable for 7+ days |
| Debt risk | High — easy to forget |
| Storage | Provider (LD/GrowthBook) or config file |
**Examples:**
- `new-checkout-flow` — gating a UI rewrite
- `payment-v2-engine` — gating backend rewrite during cutover
- `enable-search-relevance-v3` — A/B test of new ranking
**Anti-pattern:** Release flag still at 100% in code 6+ months later. The branch the flag protects is dead code; remove it.
## 2. Experiment flag
**Purpose:** Run an A/B test or multivariate experiment.
| Property | Value |
|---|---|
| Lifespan | 2-8 weeks (until significance) |
| Default | OFF; control group |
| Owner | Product or Marketing |
| Cleanup trigger | Test concluded; winner shipped |
| Debt risk | Medium |
| Storage | Provider with experimentation features |
**Examples:**
- `homepage-headline-v2` — testing new copy
- `pricing-page-monthly-vs-annual-default` — testing default toggle
- `onboarding-checklist-vs-tour` — testing onboarding pattern
**Anti-pattern:** Experiment running for 6 months because no one decided to call it. Either declare a winner or kill the test.
## 3. Operational flag
**Purpose:** Circuit breakers, kill switches, performance toggles. Designed to be flipped during incidents.
| Property | Value |
|---|---|
| Lifespan | Months to years (long-lived by design) |
| Default | ON (active path) |
| Owner | SRE / on-call team |
| Cleanup trigger | Replaced by autoscaling, retired feature |
| Debt risk | Low — they're meant to persist |
| Storage | Provider with low-latency global edge |
**Examples:**
- `enable-rate-limit-v2` — kill switch if v2 misbehaves
- `disable-recommendations-engine` — emergency cutoff
- `use-fallback-search` — degraded mode toggle
**Anti-pattern:** Operational flag that no one knows how to use during an incident. Document the trigger and runbook.
## 4. Permission flag
**Purpose:** Entitlements per user/account/plan/role. Permanent by design.
| Property | Value |
|---|---|
| Lifespan | Indefinite (plan/role lifetime) |
| Default | OFF; granted by entitlement system |
| Owner | Product (plan/role definitions) |
| Cleanup trigger | Plan or role retired |
| Debt risk | Very low |
| Storage | User/account database, NOT a flag provider |
**Examples:**
- `feature.advanced-analytics` — enterprise-only
- `feature.export-csv` — paid plans only
- `role.admin-dashboard` — admin-only UI
**Anti-pattern:** Permission flags stored in a flag provider with per-user targeting rules. Move them to your entitlements system; they're not feature flags.
## Classification matrix
When you can't decide, ask:
| Question | If YES | If NO |
|---|---|---|
| Will this be at 100% in <90 days? | Release | next ↓ |
| Will this run an A/B test? | Experiment | next ↓ |
| Is this a kill switch / safety toggle? | Operational | next ↓ |
| Is this a plan/role entitlement? | Permission | reconsider |
If none fit: you don't need a flag. Either ship the feature directly via deploy, or use a different mechanism (config, env var, role).
## Ownership rules
- Every flag must have a named owner at creation
- When the owner leaves, the flag is reassigned within 30 days or removed
- Release flags lapse to the team's tech-debt owner if not reassigned
## Lifespan SLAs
| Type | Max acceptable lifespan | Cleanup automation |
|---|---|---|
| Release | 90 days | `flag_debt_scanner.py` |
| Experiment | 60 days | Provider auto-stop on significance |
| Operational | none | Annual review |
| Permission | none | Tied to plan/role retirement |
FILE:references/provider_comparison.md
# Provider comparison
Five mainstream providers + DIY. Pick based on flag count, targeting needs, compliance, and self-hosting requirements.
## At-a-glance matrix
| Provider | Flag count sweet spot | Targeting | A/B testing | Audit log | Self-host | OSS | Pricing model |
|---|---|---|---|---|---|---|---|
| **LaunchDarkly** | 100+ | Best-in-class | Yes (Galaxy) | Full SOC2 audit trail | Edge SDK only | No | Per-MAU, expensive |
| **GrowthBook** | 20-500 | Good | Yes (built-in) | Yes | Yes (Docker/k8s) | Yes (MIT) | Free OSS + Cloud per-MAU |
| **Statsig** | 50-500 | Good | Best-in-class | Yes (paid) | No | No | Free tier (1M events), then per-MAU |
| **Unleash** | 10-200 | Good | Limited | Yes (Enterprise) | Yes (Docker/k8s) | Yes (Apache 2) | Free OSS + Hosted/Enterprise |
| **Flipt** | 5-100 | Basic | No | Limited | Yes (Docker/k8s) | Yes (MIT) | OSS only |
| **DIY** | <50 | None to basic | None | Whatever you build | Always | N/A | None |
## When to choose each
### LaunchDarkly
Choose if:
- Enterprise team with 100+ flags across many services
- Compliance requires SOC2 / ISO 27001 / FedRAMP audit logs
- Need fine-grained targeting (cohorts, custom attributes, percentages by attribute)
- Need experimentation + targeting + audit in one platform
- Budget for enterprise tooling ($20-100k/year typical)
Avoid if:
- Small team / <50 flags (overkill)
- Strict data residency (no on-prem; relays only)
- Low budget
### GrowthBook
Choose if:
- Mid-market team that wants OSS option for self-hosting
- Need built-in A/B testing with proper stats (frequentist + Bayesian)
- Want SQL-based experimentation (define metrics from your warehouse)
- Self-host on k8s or run their hosted Cloud
Avoid if:
- Need real-time targeting at edge (use LD or Statsig)
- Need enterprise audit features (Cloud only)
### Statsig
Choose if:
- Growth/product team for whom experimentation is the core use
- Need advanced stats (CUPED, sequential testing)
- Want generous free tier (good for early-stage)
- Want best-in-class metric library and platform-side experimentation logic
Avoid if:
- Strict data residency / self-host requirement (no on-prem option)
- Don't need experimentation, just toggles (overkill)
### Unleash
Choose if:
- OSS-first culture; want to self-host
- Dev-friendly with good SDKs and a clean API
- Don't need full A/B testing platform
- Need Open Source license for compliance (Apache 2)
Avoid if:
- Need experimentation + stats out of the box
- Need enterprise-grade audit (Enterprise tier only)
### Flipt
Choose if:
- Lightweight needs, <100 flags
- k8s-native (Flipt is operator-friendly)
- Want pure OSS, no commercial component
- Don't need A/B testing
Avoid if:
- Need targeting beyond simple boolean rules
- Need experimentation
- Need analytics or audit features
### DIY (env vars / config file)
Choose if:
- <50 flags total
- No targeting beyond `enabled: true/false`
- No A/B testing needs
- Want zero external dependencies
- Strict cost control
Implementation:
```yaml
# config/flags.yaml
flags:
new-checkout: { enabled: true, owner: jane@team }
payment-v2: { enabled: false, owner: bob@team, kill_switch: PagerDuty alert "payment-v2 SEV1" }
```
Or env-var based:
```bash
FLAG_NEW_CHECKOUT=true
FLAG_PAYMENT_V2=false
```
Avoid if:
- Flag count growing past 50
- Need percentage rollouts (you'll re-implement provider logic poorly)
- Need audit log (compliance)
- Multiple teams / multiple deploy cadences
## Cost rule of thumb
| Team stage | Typical monthly cost |
|---|---|
| Pre-seed / solo | $0 (DIY or OSS) |
| Seed (Series A) | $0-200 (Statsig free tier, Unleash OSS) |
| Series B-C | $500-3,000 (GrowthBook Cloud, Unleash Pro) |
| Series D+ / Enterprise | $5,000-20,000+ (LaunchDarkly, Statsig Pro, Unleash Enterprise) |
## Migration paths
Easy migrations:
- DIY → Unleash / Flipt (similar simple model)
- Unleash ↔ GrowthBook (similar feature surface)
Hard migrations:
- LaunchDarkly → anywhere (proprietary targeting language)
- Statsig → anywhere (proprietary experimentation logic)
**Lock-in mitigation:** Wrap your provider behind an interface in code:
```ts
interface FlagProvider {
isEnabled(name: string, context?: UserContext): boolean;
getValue<T>(name: string, defaultValue: T, context?: UserContext): T;
}
```
Swap providers by writing a new adapter, not by rewriting every call site.
## Build-vs-buy threshold
Buy a provider when:
- Flag count > 50
- Multiple teams need to manage flags independently
- Targeting needs include percentages, cohorts, or custom attributes
- Compliance requires audit log
- Need real-time updates without redeploy
Build (DIY) when:
- All of the above are NO
## Selection checklist
Before signing a contract:
- [ ] Estimate flag count over 12 months
- [ ] List required targeting dimensions (user/account/geo/%/custom)
- [ ] Confirm SDK availability for every language in your stack
- [ ] Check edge latency (p99 < 50ms for prod)
- [ ] Verify failure mode if provider is unreachable (default-to-safe)
- [ ] Confirm SOC2 / data residency if needed
- [ ] Run a 30-day proof-of-concept; measure actual cost at projected MAU
FILE:references/rollout_strategies.md
# Rollout strategies
Pick a strategy by risk, not by preference. Higher-risk launches get slower, more granular ramps.
## The 4 strategies
### 1. Ring (canary) — risky launches
`1% → 5% → 25% → 50% → 100%`
| Property | Value |
|---|---|
| Use when | Touches payments, auth, data integrity, performance-sensitive paths |
| Duration | 14-30 days typical |
| Hold time per ring | 24-72 hours minimum (long enough to detect anomalies) |
| Abort cost | Low (only 1-25% affected) |
| Verification | Full metrics suite at each ring |
**Phases:**
1. **0% (deploy)** — code ships dark; verify it deploys without flag turned on
2. **1%** — internal users + low-traffic cohort; full metric verification
3. **5%** — broader smoke test; watch for tail-of-distribution issues
4. **25%** — significant load; performance and infra checks
5. **50%** — half-and-half; perfect for A/B comparison
6. **100%** — fully on; hold 7 days before removing flag
**Abort triggers per ring:**
- Error rate > baseline + 1pp
- p99 latency > baseline × 1.2
- Business metric regression (conversion, retention) > baseline × 0.95
### 2. Linear — medium risk
Constant percent-per-day until target.
| Property | Value |
|---|---|
| Use when | Standard feature launches without high-risk paths |
| Duration | 7-14 days |
| Step size | (target / duration_days) per day |
| Abort cost | Medium |
| Verification | Daily metric check |
Example: 100% over 10 days = 10% per day.
### 3. Log (front-loaded) — low risk
Fast early ramp, slow tail. Reaches majority of population in first 1/3 of duration.
| Property | Value |
|---|---|
| Use when | Low-risk launch with high confidence; UI tweaks; copy changes |
| Duration | 3-7 days |
| Curve | `pct(t) = target × log(1+t) / log(1+T)` |
| Abort cost | Higher (most users on early) |
| Verification | Light — metric check at start and end |
### 4. Cohort — entitlement-aware
Named segments rolled in order: `internal → beta → free → paid → all`
| Property | Value |
|---|---|
| Use when | Feature has different value/risk per cohort; beta access; paying-tier first |
| Duration | Variable (gate by cohort size, not days) |
| Step size | Whole cohort at a time |
| Abort cost | Cohort-bounded |
| Verification | Per-cohort metrics |
**Order rules:**
1. Internal first — your own team finds bugs cheaply
2. Beta opt-in users — they expect rough edges
3. Free tier — broader signal at lower commercial risk
4. Paid plans — most valuable users last (or first for premium features)
5. All — flag fully on; remove flag
## Geo-staged variant
For internationally-distributed products, layer geo on top of any strategy:
```
Phase A: 100% in NZ/AU (low-traffic, English, off-business-hours US)
Phase B: 100% in EU (test data residency / GDPR paths)
Phase C: 100% in US (high traffic; full validation)
```
Useful for catching i18n, timezone, and regional infrastructure issues before peak load.
## Abort criteria
Hard-coded thresholds that auto-flip the flag back to 0% (or trigger paging):
| Signal | Threshold | Severity |
|---|---|---|
| Error rate (5xx) | > baseline + 1 percentage point | SEV1 |
| Error rate (4xx) | > baseline + 5 percentage points | SEV2 |
| p99 latency | > baseline × 1.2 | SEV2 |
| p999 latency | > baseline × 1.5 | SEV1 |
| Conversion rate | < baseline × 0.95 | SEV2 |
| Retention (D1/D7/D30) | < baseline × 0.95 | SEV2 |
| Database CPU | > 80% | SEV1 |
| Saturation alarm | any | SEV1 |
**Automate:** wire each threshold to a webhook that sets the flag to 0% via provider API.
## Verification per phase
At each phase, confirm:
1. **Health metrics** are within abort thresholds
2. **Business metrics** match or exceed control
3. **Logs** show no new error patterns
4. **User reports** (support tickets) show no spike for the affected feature
5. **Ops on-call** acknowledges no anomalies
If any signal is off, hold the phase. Don't advance on schedule alone.
## Hold-time rules
- **Off-hours hold time** doesn't count toward bake-in (e.g., a phase started Friday 6pm in PST is held until Monday 9am)
- **Weekend rollouts** require explicit owner approval and on-call coverage
- **Holiday rollouts** require VP-level approval
## Common mistakes
| Mistake | Fix |
|---|---|
| Skipping rings to "just get it done" | Don't. Aborts cost less than incidents. |
| 100% on Friday afternoon | Wait until Monday morning. |
| Rolling forward when metrics regress slightly | Stop. Investigate. The next ring exposes 5× more users. |
| No verification step defined per ring | Define it before starting. |
| Manual abort only (no automated kill switch) | Wire a threshold-based auto-abort. |
| Holding "for a few hours" then forgetting | Set a calendar event with the next phase + abort criteria. |
## Tools
- `scripts/rollout_planner.py` — generates a markdown plan
- Provider dashboards — for execution and real-time abort
- Metrics dashboard linked from `flag-doc` entry
- On-call runbook with kill-switch trigger words
FILE:scripts/flag_debt_scanner.py
#!/usr/bin/env python3
"""Scan a repo for stale feature flags (Karpathy goal-driven cleanup).
Detects flag identifiers from common code patterns, dates each one by its
introducing commit, and flags items older than --max-age-days that appear in
fewer than --min-uses places as cleanup candidates.
"""
import argparse
import json
import os
import re
import subprocess
import sys
from collections import defaultdict
from datetime import datetime, timezone
FLAG_PATTERNS = [
re.compile(r'\b(?:isFlagEnabled|isEnabled|featureFlag|getFlag|flag|useFlag|useExperiment)\(\s*["\']([\w.\-:]+)["\']'),
re.compile(r'\b(?:client|ld|unleash|growthbook|statsig)\.(?:variation|isEnabled|feature|getValue|getExperiment)\(\s*["\']([\w.\-:]+)["\']'),
]
CODE_EXTS = {".py", ".js", ".ts", ".tsx", ".jsx", ".go", ".rb", ".java", ".kt", ".cs", ".rs", ".php"}
SKIP_DIRS = {".git", "node_modules", ".venv", "venv", "dist", "build", "__pycache__", ".next"}
def _walk_code_files(repo):
for root, dirs, files in os.walk(repo):
dirs[:] = [d for d in dirs if d not in SKIP_DIRS]
for f in files:
if os.path.splitext(f)[1] in CODE_EXTS:
yield os.path.join(root, f)
def _scan_file(path):
try:
with open(path, "r", encoding="utf-8", errors="replace") as f:
text = f.read()
except OSError:
return []
found = set()
for pat in FLAG_PATTERNS:
for m in pat.finditer(text):
found.add(m.group(1))
return list(found)
def _first_commit_date(repo, flag_name):
try:
out = subprocess.run(
["git", "-C", repo, "log", "--diff-filter=A", "--format=%cI", "-S", flag_name],
capture_output=True, text=True, timeout=10, check=False,
)
except (subprocess.SubprocessError, OSError):
return None
lines = [ln for ln in out.stdout.strip().split("\n") if ln]
if not lines:
return None
try:
return datetime.fromisoformat(lines[-1])
except ValueError:
return None
def _age_days(when):
if when is None:
return None
now = datetime.now(timezone.utc)
return (now - when).days
def collect_flags(repo):
flags_to_paths = defaultdict(list)
for path in _walk_code_files(repo):
for name in _scan_file(path):
flags_to_paths[name].append(os.path.relpath(path, repo))
return flags_to_paths
def assess(repo, flags_to_paths, max_age_days, min_uses):
rows = []
for name in sorted(flags_to_paths.keys()):
paths = flags_to_paths[name]
when = _first_commit_date(repo, name)
age = _age_days(when)
is_debt = (
age is not None
and age > max_age_days
and len(paths) <= min_uses
)
rows.append({
"flag": name,
"uses": len(paths),
"age_days": age,
"first_seen": when.date().isoformat() if when else None,
"files": paths[:5],
"is_debt": is_debt,
})
return rows
def render_text(rows, max_age_days):
debt = [r for r in rows if r["is_debt"]]
print(f"Flag Debt Scanner — {len(rows)} flags found, {len(debt)} stale (>{max_age_days}d, ≤2 uses)")
print("")
if not debt:
print("No debt detected. Nice.")
return
print(f"{'flag':40} {'age':>6} {'uses':>4} files")
print("-" * 80)
for r in debt:
files = ", ".join(r["files"][:2]) + ("…" if len(r["files"]) > 2 else "")
age = f"{r['age_days']}d" if r["age_days"] is not None else "?"
print(f"{r['flag']:40} {age:>6} {r['uses']:>4} {files}")
print("")
print("Suggested action: confirm reached 100% (or killed); delete dead branch; remove flag.")
def main():
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--repo", default=".", help="Path to repo root (default: .)")
ap.add_argument("--max-age-days", type=int, default=90, help="Flags older than this are debt candidates (default: 90)")
ap.add_argument("--min-uses", type=int, default=2, help="Flags with ≤ this many uses are debt candidates (default: 2)")
ap.add_argument("--format", choices=["text", "json"], default="text")
args = ap.parse_args()
repo = os.path.abspath(args.repo)
if not os.path.isdir(os.path.join(repo, ".git")):
print(f"WARN: {repo} is not a git repo; age detection disabled", file=sys.stderr)
flags = collect_flags(repo)
rows = assess(repo, flags, args.max_age_days, args.min_uses)
if args.format == "json":
print(json.dumps(rows, indent=2, default=str))
else:
render_text(rows, args.max_age_days)
return 1 if any(r["is_debt"] for r in rows) else 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/kill_switch_audit.py
#!/usr/bin/env python3
"""Verify every feature flag in code has a documented kill switch.
Cross-references flag identifiers found in source code against a markdown
flag registry. Each documented flag must declare: owner, type, kill switch,
dashboard. Reports undocumented flags (FAIL) and incompletely-documented
flags (WARN). Use as a pre-merge gate.
"""
import argparse
import json
import os
import re
import sys
FLAG_PATTERNS = [
re.compile(r'\b(?:isFlagEnabled|isEnabled|featureFlag|getFlag|flag|useFlag|useExperiment)\(\s*["\']([\w.\-:]+)["\']'),
re.compile(r'\b(?:client|ld|unleash|growthbook|statsig)\.(?:variation|isEnabled|feature|getValue|getExperiment)\(\s*["\']([\w.\-:]+)["\']'),
]
REQUIRED_FIELDS = ("owner", "type", "kill switch", "dashboard")
CODE_EXTS = {".py", ".js", ".ts", ".tsx", ".jsx", ".go", ".rb", ".java", ".kt", ".cs", ".rs", ".php"}
SKIP_DIRS = {".git", "node_modules", ".venv", "venv", "dist", "build", "__pycache__", ".next"}
def _walk_code_files(repo):
for root, dirs, files in os.walk(repo):
dirs[:] = [d for d in dirs if d not in SKIP_DIRS]
for f in files:
if os.path.splitext(f)[1] in CODE_EXTS:
yield os.path.join(root, f)
def discover_code_flags(repo):
found = set()
for path in _walk_code_files(repo):
try:
with open(path, "r", encoding="utf-8", errors="replace") as f:
text = f.read()
except OSError:
continue
for pat in FLAG_PATTERNS:
for m in pat.finditer(text):
found.add(m.group(1))
return found
def _split_sections(text):
"""Split flag-doc into per-flag sections by H2 (## flag-name) or H3."""
sections = {}
current = None
buf = []
for line in text.splitlines():
m = re.match(r"^#{2,3}\s+([\w.\-:]+)\s*$", line)
if m:
if current is not None:
sections[current] = "\n".join(buf)
current = m.group(1)
buf = []
else:
buf.append(line)
if current is not None:
sections[current] = "\n".join(buf)
return sections
def _missing_fields(section_text):
lower = section_text.lower()
return [f for f in REQUIRED_FIELDS if f not in lower]
def audit(repo, flag_doc_path):
if not os.path.isfile(flag_doc_path):
return {"error": f"flag-doc not found: {flag_doc_path}"}
with open(flag_doc_path, "r", encoding="utf-8") as f:
doc_text = f.read()
sections = _split_sections(doc_text)
documented = set(sections.keys())
code_flags = discover_code_flags(repo)
undocumented = sorted(code_flags - documented)
orphaned_docs = sorted(documented - code_flags)
incomplete = []
for name in sorted(code_flags & documented):
missing = _missing_fields(sections[name])
if missing:
incomplete.append({"flag": name, "missing": missing})
return {
"code_flags": sorted(code_flags),
"documented_flags": sorted(documented),
"undocumented": undocumented,
"incomplete": incomplete,
"orphaned_in_doc": orphaned_docs,
}
def render_text(result):
if "error" in result:
print(f"ERROR: {result['error']}")
return
code, doc = result["code_flags"], result["documented_flags"]
print(f"Kill Switch Audit — {len(code)} flags in code, {len(doc)} documented")
print("")
if result["undocumented"]:
print(f"FAIL: {len(result['undocumented'])} undocumented flag(s):")
for f in result["undocumented"]:
print(f" - {f}")
print("")
if result["incomplete"]:
print(f"WARN: {len(result['incomplete'])} flag(s) with incomplete documentation:")
for item in result["incomplete"]:
print(f" - {item['flag']}: missing {', '.join(item['missing'])}")
print("")
if result["orphaned_in_doc"]:
print(f"INFO: {len(result['orphaned_in_doc'])} doc entry(s) for flags not in code:")
for f in result["orphaned_in_doc"]:
print(f" - {f}")
print("")
if not (result["undocumented"] or result["incomplete"]):
print("PASS: every code flag is fully documented.")
def main():
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--repo", default=".", help="Path to repo root (default: .)")
ap.add_argument("--flag-doc", required=True, help="Path to markdown flag registry (e.g., docs/feature-flags.md)")
ap.add_argument("--format", choices=["text", "json"], default="text")
args = ap.parse_args()
result = audit(os.path.abspath(args.repo), args.flag_doc)
if args.format == "json":
print(json.dumps(result, indent=2))
else:
render_text(result)
if "error" in result:
return 2
if result["undocumented"] or result["incomplete"]:
return 1
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/rollout_planner.py
#!/usr/bin/env python3
"""Generate a phased rollout schedule for a feature flag.
Strategies:
ring 1% → 5% → 25% → 50% → 100% — risky launches
linear constant percent-per-day — medium risk
log fast early, slow tail — low risk
cohort named cohorts (internal → beta → free → paid → all) — entitlement-aware
"""
import argparse
import json
import math
import sys
from datetime import datetime, timedelta
DEFAULT_RING_STOPS = [1, 5, 25, 50, 100]
DEFAULT_COHORTS = ["internal", "beta", "free", "paid", "all"]
def _ring(target):
return [s for s in DEFAULT_RING_STOPS if s <= target] + ([target] if target not in DEFAULT_RING_STOPS else [])
def _linear(target, days):
if days < 1:
return [target]
step = target / days
return [round((i + 1) * step, 2) for i in range(days)]
def _log_curve(target, days):
if days < 1:
return [target]
out = []
for i in range(days):
frac = math.log1p(i + 1) / math.log1p(days)
out.append(round(target * frac, 2))
return out
def _dedupe_sorted(values):
seen = set()
out = []
for v in values:
if v not in seen:
seen.add(v)
out.append(v)
return out
def build_schedule(strategy, target, duration_days, population, start_date):
if strategy == "ring":
percents = _ring(target)
elif strategy == "linear":
percents = _linear(target, duration_days)
elif strategy == "log":
percents = _log_curve(target, duration_days)
elif strategy == "cohort":
per_step = target / len(DEFAULT_COHORTS)
percents = [round(per_step * (i + 1), 2) for i in range(len(DEFAULT_COHORTS))]
else:
raise ValueError(f"unknown strategy: {strategy}")
percents = _dedupe_sorted(percents)
n = len(percents)
interval = max(1, duration_days // max(n - 1, 1))
rows = []
for i, pct in enumerate(percents):
date = start_date + timedelta(days=i * interval)
users = int(population * pct / 100)
cohort = DEFAULT_COHORTS[min(i, len(DEFAULT_COHORTS) - 1)] if strategy == "cohort" else None
rows.append({
"phase": i + 1,
"date": date.date().isoformat(),
"percent": pct,
"users": users,
"cohort": cohort,
"abort_if": "error_rate > baseline + 1pp OR p99_latency > baseline * 1.2",
"verify": "compare metrics dashboard against control",
})
return rows
def render_markdown(rows, strategy, target, duration_days, population):
print(f"# Rollout plan — strategy={strategy}, target={target}%, duration={duration_days}d, population={population:,}")
print("")
headers = ["Phase", "Date", "Percent", "Users", "Cohort", "Abort criteria", "Verify"]
print("| " + " | ".join(headers) + " |")
print("|" + "|".join(["---"] * len(headers)) + "|")
for r in rows:
cohort = r["cohort"] or "—"
print(f"| {r['phase']} | {r['date']} | {r['percent']}% | {r['users']:,} | {cohort} | {r['abort_if']} | {r['verify']} |")
def main():
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--population", type=int, required=True, help="Total user population")
ap.add_argument("--target-percent", type=float, default=100, help="Final rollout percent (default: 100)")
ap.add_argument("--duration-days", type=int, default=14, help="Total rollout duration (default: 14)")
ap.add_argument("--strategy", choices=["ring", "linear", "log", "cohort"], default="ring")
ap.add_argument("--start-date", default=None, help="ISO date YYYY-MM-DD (default: today)")
ap.add_argument("--format", choices=["markdown", "json"], default="markdown")
args = ap.parse_args()
if not 0 < args.target_percent <= 100:
print("ERROR: --target-percent must be in (0, 100]", file=sys.stderr)
return 2
if args.population < 1:
print("ERROR: --population must be >= 1", file=sys.stderr)
return 2
start = datetime.fromisoformat(args.start_date) if args.start_date else datetime.utcnow()
rows = build_schedule(args.strategy, args.target_percent, args.duration_days, args.population, start)
if args.format == "json":
print(json.dumps(rows, indent=2, default=str))
else:
render_markdown(rows, args.strategy, args.target_percent, args.duration_days, args.population)
return 0
if __name__ == "__main__":
sys.exit(main())
Xây và duy trì một câu chuyện công ty nhất quán cho nhân viên, nhà đầu tư, khách hàng, ứng viên và đối tác, phát hiện mâu thuẫn.
--- name: "internal-narrative" description: "Build and maintain one coherent company story across all audiences — employees, investors, customers, candidates, and partners. Detects narrative contradictions and ensures the same truth is framed for each audience's needs. Use when preparing investor updates, all-hands presentations, board communications, recruiting narratives, crisis communications, or when user mentions company narrative, messaging consistency, storytelling, all-hands, investor update, or crisis communication." license: MIT metadata: version: 1.0.0 author: Alireza Rezvani category: c-level domain: narrative-strategy updated: 2026-03-05 frameworks: narrative-frameworks, all-hands-template --- # Internal Narrative Builder One company. Many audiences. Same truth — different lenses. Narrative inconsistency is trust erosion. This skill builds and maintains coherent communication across every stakeholder group. ## Keywords narrative, company story, internal communication, investor update, all-hands, board communication, crisis communication, messaging, storytelling, narrative consistency, audience translation, founder narrative, employee communication, candidate narrative, partner communication ## Core Principle **The same fact lands differently depending on who hears it and what they need.** "We're shifting resources from Product A to Product B" means: - To employees: "Is my job safe? Why are we abandoning what I built?" - To investors: "Smart capital allocation — they're doubling down on the winner" - To customers of Product A: "Are they abandoning us?" - To candidates: "Exciting new focus — are they decisive?" Same fact. Four different narratives needed. The skill is maintaining truth while serving each audience's actual question. --- ## Framework ### Step 1: Build the Core Narrative One paragraph that every other communication derives from. This is the source of truth. **Core narrative template:** > [Company name] exists to [mission — present tense, specific]. We're building [what you're building] because [the problem you're solving]. Our approach is [your unique way of doing this]. We're at [honest description of current state] and heading toward [where you're going in concrete terms]. **Good core narrative (example):** > Acme Health exists to reduce preventable falls in elderly care using smartphone-based mobility analysis. We're building an AI diagnostic tool for care teams because current fall risk assessments are subjective, infrequent, and often wrong. Our approach — using the phone's camera during a 10-second walking test — means no new hardware, no specialist required. We have 80 care facilities in DACH paying us €800K ARR, and we're heading to €3M ARR by demonstrating clinical value at scale before our Series B. **Bad core narrative:** > Acme Health is an innovative AI company revolutionizing elderly care through cutting-edge technology that empowers care providers and improves patient outcomes across the continuum of care. The good version is usable. The bad version says nothing. --- ### Step 2: Audience Translation Matrix Take the core narrative and translate it for each audience. Same truth, different frame. | Fact | Employees need to hear | Investors need to hear | Customers need to hear | Candidates need to hear | |------|----------------------|----------------------|----------------------|------------------------| | We have 80 customers | "We've proven the model — your work matters" | "Product-market fit signal, capital efficient" | "80 care facilities trust us" | "Traction you'd be joining" | | We pivoted from hardware | "We were honest enough to change course" | "Capital-efficient pivot to better unit economics" | "We found a faster, simpler way to serve you" | "We make decisions based on evidence, not ego" | | We missed Q2 revenue | "Here's why, here's the plan, here's what you can do" | "Revenue mix shifted — trailing indicator improving" | [Usually don't tell customers revenue misses] | [Usually not shared externally] | | We're hiring fast | "The team is growing — your network matters" | "Headcount plan aligned to growth" | [Not relevant unless it affects service] | "This is a rocket ship moment" | **Rules:** - Never contradict yourself across audiences. Different framing ≠ different facts. - "We told investors growth, told employees efficiency" is a contradiction. Audit for this. - Investors and employees see each other. Board members talk to your team. Candidates google you. --- ### Step 3: Contradiction Detection Before any major communication, run the contradiction check: **Question 1:** What did we tell investors last month about [topic]? **Question 2:** What did we tell employees about the same topic? **Question 3:** Are these consistent? If not — which version is true? **Common contradictions:** - "Efficient growth" to investors + "we're hiring aggressively" to candidates - "Strong pipeline" to investors + "sales is struggling" at all-hands - "Customer-first" in culture + recent decisions that clearly prioritized revenue over customer need **When you catch a contradiction:** Fix the less accurate version, then communicate the correction explicitly. "Last month I said X. After more reflection, X is not quite right. Here's the clearer version." Correcting yourself before someone else catches it builds more trust than getting caught. --- ### Step 4: Audience-Specific Communication Cadence | Audience | Format | Frequency | Owner | |----------|--------|-----------|-------| | Employees | All-hands | Monthly | CEO | | Employees | Team updates | Weekly | Team leads | | Investors | Written update | Monthly | CEO + CFO | | Board | Board meeting + memo | Quarterly | CEO | | Customers | Product updates | Per release | CPO / CS | | Candidates | Careers page + interview narrative | Ongoing | CHRO + Founders | | Partners | Quarterly business review | Quarterly | BD Lead | --- ### Step 5: All-Hands Structure and Cadence See `templates/all-hands-template.md` for the full template. **Principles:** - Lead with honest state of the company. No spin. - Connect company performance to individual work: "Here's how what you built contributed to this outcome." - Give people a reason to be proud of their choice to work here. - Leave time for real Q&A — not curated questions. **All-hands failure modes:** - CEO speaks for 55 of 60 minutes; Q&A is "any quick questions?" - All good news, all the time — employees know when you're not being honest - Metrics without context: "ARR grew 15%" without explaining if that's good, bad, or expected - Questions deflected: "That's a great point, we should follow up on that" → never followed up --- ### Step 6: Crisis Communication When the narrative breaks — someone leaves publicly, a product fails, a security breach, a press article. **The 4-hour rule:** If something is public or about to be, communicate internally within 4 hours. Employees should never learn about company news from Twitter. **Crisis communication sequence:** **Hour 0–4 (internal first):** 1. CEO or relevant leader sends an internal message 2. Acknowledge what happened factually 3. State what you know and what you don't know yet 4. Tell people what you're doing about it 5. Tell people what they should do if they're asked about it **Hour 4–24 (external if needed):** 1. External statement (press, social) only if the event is public 2. Consistent with the internal message — same facts, audience-appropriate framing 3. Legal review if any claims or liability involved **What not to do in a crisis:** - Silence: letting rumors fill the vacuum - Spin: people can detect it and it destroys trust - "No comment": says "we have something to hide" - Blaming: even if someone else caused the problem, your audience only cares what you're doing about it **Template for crisis internal communication:** > "Here's what happened: [factual description]. Here's what we know right now: [known facts]. Here's what we don't know yet: [honest uncertainty]. Here's what we're doing: [specific actions]. Here's what you should do if you're asked about this: [specific guidance]. I'll update you by [specific time] with more information." --- ## Narrative Consistency Checklist Run before any major external communication: - [ ] Is this consistent with what we told investors last month? - [ ] Is this consistent with what we told employees at the last all-hands? - [ ] Does this contradict anything on our website, careers page, or press releases? - [ ] If an employee read this external communication, would they recognize the company being described? - [ ] If an investor read our internal all-hands deck, would they find anything inconsistent? - [ ] Have we been accurate about our current state, or are we projecting an aspiration? --- ## Key Questions for Narrative - "Could a new employee explain to a friend why our company exists? What would they say?" - "What do we tell investors about our strategy? What do we tell employees? Are these the same?" - "If a journalist asked our team members to describe the company independently, what would they say?" - "When did we last update our 'why we exist' story? Is it still true?" - "What's the hardest question we'd get from each audience? Do we have an honest answer?" ## Red Flags - Different departments describe the company mission differently - Investor narrative emphasizes growth; employee narrative emphasizes stability (or vice versa) - All-hands presentations are mostly slides, mostly one-way - Q&A questions are screened or deflected - Bad news reaches employees through Slack rumors before leadership communication - Careers page describes a culture that employees don't recognize ## Integration with Other C-Suite Roles | When... | Work with... | To... | |---------|-------------|-------| | Investor update prep | CFO | Align financial narrative with company narrative | | Reorg or leadership change | CHRO + CEO | Sequence: employees first, then external | | Product pivot | CPO | Align customer communication with investor story | | Culture change | Culture Architect | Ensure internal story is consistent with external employer brand | | M&A or partnership | CEO + COO | Control information flow, prevent narrative leaks | | Crisis | All C-suite | Single voice, consistent story, internal first | ## Detailed References - `references/narrative-frameworks.md` — Storytelling structures, founder narrative, bad news delivery, all-hands templates - `templates/all-hands-template.md` — All-hands presentation template FILE:references/narrative-frameworks.md # Narrative Frameworks Reference frameworks for building compelling, consistent business narratives. --- ## 1. Storytelling Structure for Business ### The SCR Framework (Situation, Complication, Resolution) Barbara Minto's Pyramid Principle adapted for business narrative. Works for any audience. **Situation:** The established facts everyone agrees on. **Complication:** What changed, what problem arose, what makes the situation untenable. **Resolution:** What you're doing about it, and why this solution works. **Example — Investor update:** > **Situation:** We entered Q2 with €650K ARR and a target of €800K. Our DACH pipeline was strong at 3x coverage. > > **Complication:** Two large deals (€90K combined ARR) that were expected to close in May pushed to Q3 due to procurement delays on the customer side. We ended Q2 at €710K — below target but within the range we'd flag as manageable. > > **Resolution:** Both deals are now signed with June start dates. We're entering Q3 at €800K ARR. We've added a new procurement risk flag to our pipeline methodology to catch this pattern earlier. **Why it works:** It respects the audience's intelligence, acknowledges the problem directly, and frames your response before they can object. --- ### The Problem-Solution-Evidence Structure Best for pitches, product announcements, and strategy communications. 1. **The world as it is:** What's the current reality? 2. **What's broken about it:** Why is the status quo painful or inefficient? 3. **What we're doing:** Your specific solution 4. **Why it works:** Evidence, mechanism, or proof 5. **What happens next:** Call to action or forward look **Example — All-hands strategy communication:** > The world as it is: We have 80 customers in DACH. Churn is 8% annually. That means we're losing 6–7 customers a year just to stay flat. > > What's broken: Our onboarding takes 6 weeks. By week 4, customers haven't seen value yet and they're questioning the decision. We've traced 60% of churn to customers who never completed onboarding. > > What we're doing: We're redesigning onboarding to show the first meaningful mobility report within 48 hours of account activation. > > Why it will work: We ran this with 5 pilot customers in Q2. Time-to-first-value dropped from 4 weeks to 2 days. 4 of 5 expanded their contract within 60 days. > > What happens next: Engineering ships the new onboarding flow by August 15. CS is retrained by August 22. We'll run the new flow with all new customers from September 1 and report back at the October all-hands. --- ## 2. The Founder's Narrative The founder's personal story is one of the most underutilized assets in a startup. Used well, it anchors the company's mission and creates genuine connection. ### The Founder Story Structure **Origin:** What led you to this problem? (Ideally personal — you experienced it, someone you loved experienced it, you couldn't stop thinking about why nobody was solving it) **Insight:** What did you see that others didn't? (Your unique perspective or unfair advantage) **Decision:** The moment you committed. (Specific, not aspirational — "I left my job on March 14" not "I decided to pursue my passion") **What you've learned:** 2–3 honest observations that shaped your approach. (Including what you got wrong) **Where you're going:** Connection from your personal why to where the company is heading. **Example (condensed):** > My mother had a fall in 2018 that broke her hip. She spent 3 months in rehabilitation. The terrifying part: nobody saw it coming. Her doctor had assessed her fall risk 4 months earlier — using a paper questionnaire. I spent two years talking to geriatricians trying to understand why this assessment was still done by hand, on paper, in 2018. The answer: nobody had made it easy enough for a non-specialist to do it digitally. That's what we're building. **Why it matters:** Investors, candidates, and customers all respond to a founder who started from a real problem rather than a market opportunity. The narrative makes you memorable and makes the mission credible. --- ## 3. How to Deliver Bad News Across Audiences ### Universal principles 1. **Internal first.** Always. Every time. No exceptions. 2. **Direct, not hedged.** "We missed our Q2 target by 12%" beats "Q2 performance came in below our expectations." 3. **Own it before explaining it.** Context comes after acknowledgment, not before. 4. **State what you're doing.** Bad news without a response plan creates panic. 5. **Give a timeline.** "We'll know more by [date]" is better than open-ended uncertainty. ### Delivering bad news to employees **Format:** Synchronous (all-hands or team meeting), followed by written summary. **What to say:** - What happened (factual, no spin) - What it means for the company - What it means for them specifically (will roles change? Will comp change?) - What you're doing about it - When you'll have more information **What not to say:** - "I can't share the details" (share everything you legally can) - "This is actually good news because..." (if it's bad news, don't reframe it before acknowledging it) - "We saw this coming" (if you did, why didn't you tell them?) **Example — Missed fundraise:** > "I have to share news that's disappointing. We went out to raise a Series A in Q1, and we didn't close the round. We had term sheets that fell through when the market conditions shifted in April. We're not in crisis — we have 12 months of runway — but we need to recalibrate. Here's what that means concretely: we're pausing 3 open headcount. Everyone currently on the team keeps their role. We're going back to market in Q4 with stronger metrics. I'll share our updated financial model with everyone by Friday and answer every question you have." --- ### Delivering bad news to investors **Format:** Written update (monthly update format) + proactive call if material. **What to say:** - Headline the bad news in the first paragraph (don't bury it) - Context: what changed and what didn't change - What you're doing about it - What you need from them (if anything) **What investors hate:** - Finding out from someone other than you - Bad news wrapped in so much context they have to work to find it - "We're watching it closely" without specific action - Consistent over-optimism followed by consistent misses **Example — Investor update paragraph:** > "Revenue miss: We ended Q2 at €710K ARR vs. a target of €800K. Two deals totaling €90K pushed to Q3 due to customer procurement delays (not product or relationship issues — both have since signed). We've adjusted our sales process to flag procurement risk earlier. Q3 is starting at €800K with those deals live." --- ### Delivering bad news to customers **Scope:** Only share bad news that affects them. Don't share internal struggles that aren't relevant to their experience. **Format:** Proactive communication from their account owner or a senior leader. **What customers need:** - What happened (that affects them) - What you're doing about it - What they should do (if anything) - Who to contact **What customers don't need:** - Your internal financial struggles - Drama about team changes - More detail than affects their use of your product **Example — Service disruption:** > "Yesterday evening we experienced a 90-minute service outage that affected your access to [feature]. We've identified the root cause (a failed database migration) and deployed a fix. Your data is intact and complete. We've implemented additional monitoring to prevent this from recurring. I'd like to schedule a brief call to answer any questions you have." --- ## 4. Narrative Consistency Checklist Use before any significant external communication. ### Pre-communication audit **Factual consistency:** - [ ] Is the ARR/revenue figure consistent with what we've shared with investors? - [ ] Is the team size consistent with what's on LinkedIn and our careers page? - [ ] Are our stated priorities consistent with our published roadmap? - [ ] Is our "stage" description consistent across all channels? (We can't be "early stage" to investors and "established leader" to customers) **Message consistency:** - [ ] Does this message conflict with anything said in the last 90 days? - [ ] If an employee read this external message, would they recognize the company? - [ ] If an investor read our internal all-hands, would they find anything that contradicts what we've told them? **Audience appropriateness:** - [ ] Have we answered the key question for this specific audience? - [ ] Have we avoided sharing information this audience doesn't need and shouldn't have? - [ ] Have we framed the message for what this audience cares about — not what we want them to care about? --- ## 5. All-Hands Presentation Templates See `templates/all-hands-template.md` for the complete slide-by-slide template. ### Monthly all-hands (30–45 min) **Structure:** 1. State of the company (10 min) — honest, metric-driven 2. Progress on quarterly rocks (5 min) — on track / off track / done 3. Team spotlight (5 min) — one team's work, why it matters 4. What's coming next 30 days (5 min) — what to expect 5. Q&A (10–15 min) — real questions, real answers ### Quarterly all-hands (60–90 min) **Structure:** 1. Last quarter results vs. targets (15 min) 2. What we learned (10 min) — honest reflection on what didn't work 3. Next quarter priorities (15 min) — company rocks, why these three 4. Strategy update (10 min) — anything changing? Why? 5. Team recognition (10 min) — specific, values-linked examples 6. Q&A (15–20 min) ### Annual all-hands (2–4 hours, often a full day) **Structure:** 1. Year in review: what we achieved (30 min) 2. What we learned — what we'd do differently (20 min) 3. State of the company: financial health, competitive position (20 min) 4. 3-year vision update (30 min) 5. Next year's strategy and priorities (30 min) 6. Department presentations: what each team is building (60 min) 7. Celebrations and recognition (20 min) 8. Q&A + social (open-ended) ### The "no-BS questions" technique At any all-hands, reserve the last 5 minutes for: "What question are you afraid to ask publicly? Submit anonymously via [link]." Read 3–5 of the hardest ones out loud and answer them honestly. This builds more trust than 45 minutes of polished presentation. FILE:templates/all-hands-template.md # All-Hands Presentation Template **Monthly format (30–45 min) | Adjust timing for quarterly/annual** --- ## Slide 1: State of the Company **Headline:** One honest sentence about where we are right now. > "We're ahead on revenue, behind on hiring, and Q3 is looking strong." **3 key metrics (vs. target):** | Metric | Target | Actual | Status | |--------|--------|--------|--------| | ARR / Revenue | | | 🟢/🟡/🔴 | | [Key growth metric] | | | | | [Key health metric] | | | | **One sentence on momentum:** Are we accelerating, steady, or facing headwinds? Be honest. --- ## Slide 2: Progress on Quarterly Rocks For each company-level rock: | Rock | Owner | Status | |------|-------|--------| | [Rock 1 description] | [Name] | ✅ Done / 🟡 On track / 🔴 At risk | | [Rock 2 description] | [Name] | | | [Rock 3 description] | [Name] | | For any 🔴 at-risk rock: one sentence on what changed and what we're doing about it. --- ## Slide 3: What We're Proud Of **One specific win from the last 30 days.** Not the metric — the story behind it. > "CS team saved the Müller Group account after a critical feature gap was flagged 72 hours before their renewal. They pulled together engineering, product, and sales in 24 hours and presented a roadmap commitment that converted a churned account into an expansion. That's what customer obsession looks like." Tie to a company value. Name the people involved. --- ## Slide 4: What We Learned / What Didn't Work **One honest thing that didn't go as planned.** > "Our Q2 product launch was delayed 3 weeks because we underestimated the testing scope for the new export feature. We shipped it, customers are using it, but we learned that our pre-launch testing checklist needs to include third-party integration validation. We've added that to the template." If it's small: 2 sentences. If it's big: more time here, less elsewhere. **Why this slide exists:** A company that only celebrates wins teaches people to hide problems. This slide teaches people that honesty is valued. --- ## Slide 5: What's Coming Next 30 Days **3 things to know about:** 1. [Upcoming release / launch / event] — [What it is and why it matters] 2. [Hiring update or org change] — [Honest current state] 3. [External event / partnership / market development] — [What we're watching] **What NOT to include:** Vague aspirations. Only things people can actually act on or prepare for. --- ## Slide 6: Q&A **Format:** Live questions preferred. Anonymous submission option always available. **CEO rules for Q&A:** - Answer the question asked, not the one you wish they'd asked - "I don't know, but I'll find out and share by [date]" > vague answer - "I can't share that yet because [reason], but I will when I can" > "no comment" - If the same question has been asked three times across all-hands, it's a communication gap — fix it **Closing line:** > "Thanks for your time. If you have a question you didn't get to ask — Slack me directly, or use the anonymous form. I read every one." --- ## Presenter Notes **Before every all-hands:** - [ ] Review last all-hands deck — anything promised that wasn't delivered? - [ ] Check: is there anything employees should have heard from us before this meeting? - [ ] Have 3–5 real Q&A answers prepared for the hardest questions you'd expect **During Q&A:** - [ ] Don't deflect hard questions — they remember - [ ] Don't over-explain — short answers signal confidence - [ ] Don't let one person dominate — "let's take that to a 1:1" is a valid response **After every all-hands:** - [ ] Send a written summary within 24 hours (key metrics, decisions, answers to top questions) - [ ] Follow up on any commitments made during Q&A within the stated timeframe
Chấm điểm sức khỏe sprint và phân tích velocity cho đội agile.
---
name: sprint-health
description: Sprint health scoring and velocity analysis for agile teams. Usage: /sprint-health <analyze|velocity> [options]
---
# /sprint-health
Score sprint health across delivery, quality, and team metrics with velocity trend analysis.
## Usage
```
/sprint-health analyze <sprint_data.json> Full sprint health score
/sprint-health velocity <sprint_data.json> Velocity trend analysis
```
## Input Format
```json
{
"sprint_name": "Sprint 24",
"committed_points": 34,
"completed_points": 29,
"stories": {"total": 12, "completed": 10, "carried_over": 2},
"blockers": [{"description": "API dependency", "days_blocked": 3}],
"ceremonies": {"planning": true, "daily": true, "review": true, "retro": true}
}
```
## Examples
```
/sprint-health analyze sprint-24.json
/sprint-health velocity last-6-sprints.json
/sprint-health analyze sprint-24.json --format json
```
## Scripts
- `project-management/scrum-master/scripts/sprint_health_scorer.py` — Sprint health scorer (`<data_file> [--format text|json]`)
- `project-management/scrum-master/scripts/velocity_analyzer.py` — Velocity analyzer (`<data_file> [--format text|json]`)
## Skill Reference
> `project-management/scrum-master/SKILL.md`
Tạo danh sách đọc bổ sung từ giáo trình môn học bằng tìm kiếm học thuật Consensus, phù hợp trình độ và đối tượng khóa học.
---
name: syllabus
description: "Generates a curated supplementary reading list from any course syllabus using Consensus academic search. Grill-me intake (syllabus input format + course audience + year range) plus a grouping forcing-options checkpoint before any search runs — so the reading list matches the course's level and recency need. Parses the syllabus to extract topics and learning outcomes, searches Consensus for recent peer-reviewed papers per topic, and produces a professionally formatted .docx with clickable Consensus links, plain-language summaries calibrated to audience level, and Bloom-higher-order discussion questions tied to course learning goals. Triggers whenever a user uploads a syllabus, course outline, or curriculum document and wants supplementary readings. Also triggers on: 'syllabus reading list', 'find papers for my course', 'create a reading list from this syllabus', 'recent research for my class', 'supplementary readings', 'find journal articles for these topics', 'what recent papers cover this material', 'any new research on these course topics', 'update my syllabus with recent papers'. Even casual mentions when a syllabus is attached should trigger this skill."
license: MIT
metadata:
source_spec: "megaprompts/10-syllabus-megaprompt.md"
build_pattern: "Path B (direct conversion)"
research_pack_convention: "Agent Integrity Rules verbatim per PR #657 audit; bundled-JS-DOCX-generator variant"
version: 1.0.0
---
# Syllabus — Course Supplementary Reading List
> **Portability:** Requires a Consensus MCP connection, Node.js with `docx` package, and file reading capability for the syllabus. Works in Claude Code CLI natively. In Claude.ai with Consensus MCP + Code Execution + file upload, the workflow is supported.
For an instructor or student with a course syllabus, produce a professional supplementary reading list as `.docx` containing recent peer-reviewed papers per course section.
## Architectural Pattern: Bundled Script
This skill uses a **bundled JavaScript helper script** for DOCX generation rather than inlining the 300+ lines of layout code:
- DOCX generation logic is reusable + complex
- Better separation of concerns: skill = orchestration + intelligence; script = mechanical document assembly
- Token-efficient: skill doesn't re-derive layout each run
- Easier to maintain and version
The bundled script is at `scripts/generate_reading_list.js`. The skill orchestrates the pipeline + invokes the script with JSON input.
## Agent Integrity Rules (Research-Pack Convention)
Locked verbatim per PR #657 audit.
- **Only use what Consensus returns.** Every paper title, author, journal, year, URL must come from this session's tool calls. Training-knowledge papers labeled `[Not from Consensus — model knowledge]` and excluded.
- **Confirm before moving on.** A search isn't complete until response received and inspected.
- **Track three counts.** Queries sent / papers received / papers cited. Surface in audit summary.
- **Surface gaps, don't fill them.** Section with one paper + note about limited results > section padded with fabrications.
## Phase 0: Grill-Me Intake (3 forcing questions)
### Q1 (root) — Syllabus input
> **Provide the syllabus — pick one:**
>
> 1. File path (PDF, DOCX, text) — I'll read it
> 2. Pasted content — paste below
> 3. Image of a printed syllabus — attach the image
>
> *Why I'm asking:* Each format needs a different reader (PDF / DOCX parser / vision). Picking upfront prevents wasted attempts.
Forcing choice. Refuse to start without a syllabus.
### Q2 (depends on Q1) — Course audience
> **Course audience — pick one:**
>
> 1. Undergraduate (intro level)
> 2. Undergraduate (advanced / upper division)
> 3. Graduate (Masters / early PhD)
> 4. Graduate (doctoral / advanced)
> 5. Professional / continuing education
> 6. Mixed
>
> *Why I'm asking:* Audience dictates summary jargon level and discussion-question complexity. Undergrad summaries define every term; grad summaries assume technical fluency. Discussion questions for undergrads test analysis; for grads test critique and extension.
See [`references/audience_calibration.md`](references/audience_calibration.md) for the canon.
### Q3 (depends on Q1) — Year range
> **Year range for papers — pick one:**
>
> 1. Last 1 year (most recent only)
> 2. Last 2 years (default — recent + a year of context)
> 3. Last 5 years (broader, includes foundational recent work)
>
> *Why I'm asking:* Reading lists go stale fast. 1-year filters keep things fresh; 5-year filters surface foundational recent work that's already standard. Drives the year_min parameter on every Consensus search.
Forcing choice with default (last 2 years).
**Stop condition:** 3 questions max before Phase 1. The post-Phase-2 group-and-confirm checkpoint is its own grill-me moment.
## Phase 1: Parse the Syllabus
Per Q1 input format:
- **PDF**: use PDF reader; extract text
- **DOCX**: use pandoc or DOCX parser; extract text
- **Text/pasted**: read directly
- **Image**: use vision; extract text
From extracted text:
1. Course title + instructor + term
2. Topic list (lecture titles, week-by-week breakdown, etc.)
3. Learning outcomes (if explicit; if missing, infer 3-5 from description)
Mark inferred learning outcomes as `[inferred]` in the DOCX.
## Phase 2: Group Topics + Confirm with User
### Group via topic_grouper.py
Use `scripts/topic_grouper.py` to cluster related topics into 6-12 sections. Heuristic: closely-related topics merge; cross-cutting topics get their own section.
### Group-and-Confirm Checkpoint (Forcing Options)
After grouping, present:
> **Proposed sections: [list with item counts]. Pick one:**
>
> 1. "Looks good — proceed with these sections"
> 2. "Merge sections [X] and [Y]"
> 3. "Split section [X] into two"
> 4. "Add a section for [topic]"
> 5. "Remove section [X]"
>
> *Why I'm asking:* Grouping drives search allocation. Wrong grouping wastes the search budget on bad clusters. This is the **last cheap moment** to correct course before searches consume Consensus calls.
**Refuse to start Phase 3 without explicit user choice.**
## Phase 3: Search Consensus per Section
Sequential, 1 q/sec. 1-2 queries per section.
### Applied-Domain Weaving (Critical)
Don't just search the topic — **search the topic + applied domain**:
| ❌ Generic | ✅ Applied-domain |
|---|---|
| "enzyme kinetics" | "enzyme kinetics food processing applications" |
| "machine learning" | "machine learning clinical decision support" |
| "thermodynamics" | "thermodynamics renewable energy systems" |
| "social network analysis" | "social network analysis public health interventions" |
Boosts paper relevance dramatically. See [`references/applied_domain_weaving.md`](references/applied_domain_weaving.md) for the canon.
### Per-Section Pattern
```
For each section:
1. Construct query: "{topic-keywords} {applied-domain-angle}" + year_min from Q3
2. Submit to Consensus (sequential, 1 q/sec gap enforced by citation_tracker)
3. Receive results
4. (If thin) submit one fallback query without applied-domain angle
5. Select 1-3 papers per section (15-25 total across all sections)
```
### Selection Priorities
1. **Relevance** — paper directly addresses the section topic
2. **Reviews / meta-analyses** — synthesize the field
3. **Citation count** — established work
4. **Applied-domain connection** — tied to the course's domain (e.g., engineering vs theory)
## Phase 4: Write Summaries + Discussion Questions
### Summary writing
Per paper:
- Plain language (calibrated to audience from Q2)
- 2-3 sentences
- Define jargon if undergraduate audience; assume fluency if graduate
### Quality bars
| ✅ Good summary | ❌ Bad summary |
|---|---|
| "This review maps how different diets — Mediterranean, Nordic, vegetarian — reshape the types of fat molecules circulating in your blood, with implications for heart disease risk." | "This paper reviews lipidomic profiles across dietary interventions and their cardiometabolic implications." |
### Discussion question writing
Per paper:
- Bloom **higher-order** (apply / analyze / evaluate)
- Tied to a specific course learning outcome
- Promotes discussion, not just recall
| ✅ Good question | ❌ Bad question |
|---|---|
| "If dietary fat quality can reshape your lipoprotein lipidome, what does this suggest about the biochemical basis for dietary guidelines recommending unsaturated over saturated fats?" | "What did the authors find?" (Just recall) |
Use `scripts/discussion_question_validator.py` to flag recall-only questions.
## Phase 5: Generate .docx via Bundled Script
```bash
node ../scripts/generate_reading_list.js \
--input /tmp/syllabus_data.json \
--output /path/to/reading_list_<course>_<date>.docx
```
The script accepts JSON with this schema:
```json
{
"courseTitle": "string",
"courseSubtitle": "string",
"generatedDate": "string",
"yearRange": "string",
"introText": "string",
"learningOutcomes": ["string", ...],
"sections": [
{
"heading": "string",
"papers": [
{
"title": "string",
"authors": "string",
"journal": "string",
"year": number,
"url": "string",
"summary": "string",
"question": "string"
}
]
}
],
"auditLog": {
"totalQueriesSent": number,
"totalPapersReceived": number,
"totalPapersCited": number,
"toolConstraints": "string",
"searchDetails": [
{
"section": "string",
"query": "string",
"papersReturned": number,
"papersSelected": number,
"status": "string"
}
],
"failures": []
}
}
```
The script handles:
- `docx` package require with multi-location fallback
- Title page, intro with Consensus link, learning outcomes box, numbered papers per section
- `ExternalHyperlink` with full Consensus URLs (never truncated)
- `LevelFormat.BULLET` for lists (not unicode bullets)
- Footer with generation metadata
- Input validation (missing fields → graceful error)
See [`references/bundled_script_pattern.md`](references/bundled_script_pattern.md) for why bundled vs inline.
## Phase 6: Deliver
- File path
- Audit summary in chat: "Saved {file}. {N} sections × {M} papers / {K} cited. Plan tier: {tier}."
- Validate: `python scripts/office/validate.py <docx>`
## Tooling
| Script | Role |
|---|---|
| `scripts/citation_tracker.py` | Consensus three-count audit + 1s sequential discipline at `~/.syllabus_sessions/<session>.json` |
| `scripts/topic_grouper.py` | Heuristic 6-12 section grouping from extracted topics |
| `scripts/discussion_question_validator.py` | Bloom higher-order quality check; flags recall-only questions |
| `scripts/generate_reading_list.js` | **Bundled Node.js DOCX generator** — JSON input → .docx output |
## References
- [`references/applied_domain_weaving.md`](references/applied_domain_weaving.md) — search-quality canon (7+ sources)
- [`references/audience_calibration.md`](references/audience_calibration.md) — undergrad vs grad summary jargon (7+ sources)
- [`references/bundled_script_pattern.md`](references/bundled_script_pattern.md) — why bundle vs inline (7+ sources)
## Error Handling
| Failure | Behavior |
|---|---|
| Consensus rate-limit hit | Wait 3s, retry once, log |
| Search returns 0 for a section | Note section as "limited results — consider manual supplementation" |
| 3 consecutive failures | Stop, alert user, share collected so far |
| `docx` package not installed | Script attempts `npm install`; if still failing, fail with clear message |
| DOCX validation fails | Unpack XML, log issue, ask user to retry |
| Syllabus format unsupported | List supported formats, ask user to convert |
| Learning outcomes can't be extracted | Infer 3-5 from course description; mark as inferred in document |
## Anti-Patterns To Reject
- Parallelizing Consensus calls (rate limit)
- Searching topics without applied-domain angle (poor relevance)
- Padding sections with fabricated entries when Consensus returns thin
- Generic discussion questions ("What did the authors find?")
- Jargon-heavy summaries unsuitable for the course's audience level
- Skipping the group-and-confirm step (wastes searches)
- Truncating Consensus URLs in hyperlinks
- Inlining 300 lines of docx-generation JavaScript in the skill body (use bundled script)
---
**Version:** 1.0.0
**Source spec:** [`megaprompts/10-syllabus-megaprompt.md`](../../../../megaprompts/10-syllabus-megaprompt.md)
**Build pattern:** Path B (direct conversion). Bundled-JS-DOCX-generator variant.
FILE:references/applied_domain_weaving.md
# Applied-Domain Weaving — The Search-Quality Multiplier
This reference answers exactly one decision: **why does the syllabus skill always weave the applied domain into Consensus queries, and what makes a generic search produce thin results?**
## The Core Insight
A query like `"enzyme kinetics"` returns **review papers and theoretical treatments** — useful for a biochemistry course but unhelpful for a *food science* course where students need to know how enzyme kinetics applies to bread fermentation, cheese ripening, and meat tenderization.
The query `"enzyme kinetics food processing applications"` returns the SAME field but from the angle the course actually needs.
> **Applied-domain weaving = search the topic + the course's applied domain.**
This is the single highest-leverage technique in the skill. Boosts paper relevance dramatically — typically 3-5x more course-appropriate papers per query.
## Concrete Examples by Discipline
### Engineering / Applied Sciences
| Topic | Generic search | Applied-domain search |
|---|---|---|
| Thermodynamics | "thermodynamics" | "thermodynamics renewable energy systems" |
| Fluid mechanics | "fluid mechanics" | "fluid mechanics biomedical device design" |
| Control systems | "PID control" | "PID control HVAC building automation" |
| Materials science | "polymer composites" | "polymer composites aerospace structural" |
### Health Sciences
| Topic | Generic search | Applied-domain search |
|---|---|---|
| Pharmacology | "drug interactions" | "drug interactions pediatric oncology" |
| Public health | "social determinants" | "social determinants rural health disparities" |
| Nutrition | "lipid metabolism" | "lipid metabolism Mediterranean diet" |
| Immunology | "innate immunity" | "innate immunity vaccine development" |
### Computer Science / Data Science
| Topic | Generic search | Applied-domain search |
|---|---|---|
| Machine learning | "neural networks" | "neural networks medical imaging diagnosis" |
| Distributed systems | "consensus algorithms" | "consensus algorithms blockchain finance" |
| Database systems | "query optimization" | "query optimization warehouse analytics" |
| HCI | "user interface design" | "user interface design accessibility" |
### Business / Social Sciences
| Topic | Generic search | Applied-domain search |
|---|---|---|
| Game theory | "Nash equilibrium" | "Nash equilibrium auction design" |
| Behavioral econ | "loss aversion" | "loss aversion retirement savings" |
| Org psychology | "team dynamics" | "team dynamics remote engineering" |
| Marketing | "consumer behavior" | "consumer behavior subscription services" |
### Physical Sciences
| Topic | Generic search | Applied-domain search |
|---|---|---|
| Quantum mechanics | "entanglement" | "entanglement quantum computing applications" |
| Astrophysics | "stellar evolution" | "stellar evolution exoplanet habitability" |
| Geology | "plate tectonics" | "plate tectonics earthquake hazard" |
## Why This Works
The applied-domain term:
1. **Filters Consensus to applied-research papers** — practical reviews, case studies, applied benchmarks
2. **Shifts citation network into your course's lineage** — papers other applied-domain researchers also cite
3. **Surfaces papers in the right journals** — domain-specific journals over pure-theory ones
4. **Gives papers students can connect to** — abstract theory → "I see how this matters"
## How to Identify the Applied Domain
The applied domain comes from one or more of:
1. **Course title** — "Food Science 301" → "food processing applications"
2. **Department / college** — Engineering → "engineering applications"
3. **Course description** — explicit "applied to X" / "for Y industry"
4. **Learning outcomes** — operational outcomes signal applied focus
If the syllabus is genuinely theoretical (e.g., a pure-math course), use **methodological angle** instead:
- Theoretical CS → "theoretical CS algorithm complexity"
- Pure math → "pure math applications" (or skip — pure-theory queries are fine here)
## When to Skip Applied-Domain Weaving
- **Pure theory courses** — no applied angle. Search topic only.
- **Survey courses** — broad coverage needed; applied-domain may narrow too much.
- **Topic genuinely doesn't have a natural applied domain** — e.g., "intro to research methods" — skip and search the topic + "review" or "introduction".
If applied-domain search returns < 3 papers, **fall back to generic search** for that section. Don't pad with fabrications.
## Operational Pattern
In Phase 3 of the skill:
```
For each section in [proposed sections]:
1. Construct primary query: "{topic} {applied-domain-keyword}" + year_min
2. Submit to Consensus (sequential, 1 q/sec gap)
3. If results >= 3: select papers, move on
4. If results < 3: submit fallback "{topic}" + year_min
5. Select 1-3 papers from combined results
```
## Anti-Patterns
### "Just search the topic"
Most common mistake. Produces theoretically rigorous but unhelpful papers for an applied course. Students can't connect them to course goals. Engagement drops.
### "Search the applied domain alone"
Without the topic anchor, query is too broad. "Food processing" returns 10,000+ papers across all subfields. Topic + applied-domain is the sweet spot.
### "Use multiple applied domains in one query"
"Enzyme kinetics food processing biomedical industrial applications" overconstrains. Each query targets ONE applied domain. If a section spans multiple domains, run separate queries.
### "Weave domain into queries even for pure-theory courses"
Pure-theory courses don't have applied domains. Forcing one in produces awkward queries that miss the actual theoretical literature.
### "Skip applied-domain weaving to save query budget"
The applied-domain weaving doesn't add queries — it modifies them. Same query budget, dramatically better relevance.
## Operational Checklist
- [ ] Course's applied domain identified (from title / department / description / learning outcomes)
- [ ] Each Phase 3 query: `{topic} + {applied-domain}` format
- [ ] Fallback to generic search if applied-domain returns < 3 papers
- [ ] Pure-theory courses: skip applied-domain weaving (use generic)
- [ ] Multi-domain sections: separate query per domain (don't stack in one query)
## Citations (7 sources)
1. **Bloom, B. S. (ed.), *Taxonomy of Educational Objectives* (1956).** Source for the application-tier of learning that justifies the applied-domain framing. Higher-tier learning (apply / analyze / evaluate) requires applied examples; pure-theory readings only support recall + comprehension.
2. **Mayer, R. E., *Multimedia Learning* (Cambridge, 2nd ed. 2009).** Empirical research on how applied examples accelerate learning vs abstract presentation. Source for the engagement-drop signal that pure-theory readings produce in applied courses.
3. **Fink, L. D., *Creating Significant Learning Experiences* (Jossey-Bass, 2003).** Source for the "integration" learning category — the discipline of connecting course content to students' applied contexts. Applied-domain weaving operationalizes this.
4. **Donald, J. G., *Learning to Think: Disciplinary Perspectives* (Jossey-Bass, 2002).** Empirical study of disciplinary thinking patterns. Justifies the per-discipline query-pattern table — engineering thinks differently from biology thinks differently from CS.
5. **Lave, J. & Wenger, E., *Situated Learning* (Cambridge, 1991).** Source for "situated cognition" — knowledge is best learned in the context of its application. Applied-domain weaving brings the readings into the situated context.
6. **Chickering, A. W. & Gamson, Z. F., "Seven Principles for Good Practice in Undergraduate Education" — *AAHE Bulletin*, 1987.** Principle #5 ("Emphasize Time on Task") + Principle #7 ("Respect Diverse Talents") favor applied-domain readings over pure-theory abstracts that don't connect to student backgrounds.
7. **Boyer, E. L., *Scholarship Reconsidered* (Carnegie Foundation, 1990).** Source for the "Scholarship of Application" framing. Applied-domain papers represent this scholarship category; weaving them into reading lists honors that scholarship.
FILE:references/audience_calibration.md
# Audience Calibration — Undergrad vs Grad Summary Jargon + Question Complexity
This reference answers exactly one decision: **how does the syllabus skill calibrate summary jargon and discussion question complexity to the course's audience (Q2)?**
## The Core Rule
The same paper needs **different summaries** for different audiences:
- **Undergrad-intro**: define every technical term; assume zero prior knowledge
- **Undergrad-advanced**: assume foundational vocabulary; explain field-specific terms
- **Grad-Masters**: assume technical fluency; brief context for novel concepts
- **Grad-doctoral**: assume technical + methodological fluency; brief mention only of established context
Same paper, different summaries. Generic summaries miss the engagement target.
## Audience Buckets (Q2)
| Bucket | Vocabulary assumption | Method assumption | Discussion question complexity |
|---|---|---|---|
| Undergraduate (intro) | Zero specialized | Zero | Recall + comprehension + simple application |
| Undergraduate (advanced) | Foundational vocab | Common methods | Application + analysis |
| Graduate (Masters / early PhD) | Technical fluency | Common research methods | Analysis + evaluation |
| Graduate (doctoral / advanced) | Technical + methodological fluency | Methods specifics | Evaluation + critique + synthesis |
| Professional / continuing ed | Field-specific assumed | Methods context-dependent | Application to practice |
| Mixed | Lowest bucket present | Same | Same |
## Summary Calibration
### Undergrad-intro
Every technical term defined. Plain language. Connects to common experience.
| ❌ Too jargon | ✅ Calibrated |
|---|---|
| "This RCT compared lipidomic profiles across dietary interventions to assess cardiometabolic risk modulation." | "This randomized study compared what happens to fat molecules in the blood when people eat different diets — Mediterranean, Nordic, vegetarian — and looked at how those changes might affect heart disease risk." |
| "The phylogenetic analysis identified convergent evolution of toxin-resistant Na+ channels across reptilian lineages." | "Researchers compared sodium-channel genes across snake species and found that snakes from very different evolutionary branches independently developed similar resistance to toxic prey." |
### Undergrad-advanced
Foundational vocabulary assumed. Explain field-specific terms briefly.
| ❌ Too dumbed-down | ✅ Calibrated |
|---|---|
| "This randomized study compared what happens to fat molecules in the blood..." | "This RCT (n=240) tracked lipidomic shifts across three dietary patterns — Mediterranean, Nordic, vegetarian — over 12 weeks. Cardiometabolic markers improved most in the Mediterranean arm." |
| "Researchers compared sodium-channel genes..." | "Phylogenetic analysis across 47 reptilian lineages identifies convergent evolution of Na+ channel modifications conferring resistance to neurotoxic prey." |
### Grad (Masters or doctoral)
Technical fluency assumed. Brief context for novel concepts. Method specifics if relevant.
| ❌ Too verbose | ✅ Calibrated |
|---|---|
| "This RCT (n=240) tracked lipidomic shifts across three dietary patterns over 12 weeks. Cardiometabolic markers improved most in Mediterranean." | "RCT (n=240, 12-week, parallel-arm) comparing Mediterranean / Nordic / vegetarian. Mediterranean → 14% lower LDL-particle count, 22% lower oxidized LDL; differences plausibly mediated by MUFA:SFA ratio." |
| "Phylogenetic analysis across 47 reptilian lineages identifies convergent evolution..." | "Bayesian phylogenetic analysis (47 lineages, BEAST 2.7) supports independent emergence of Na+ channel S6-domain modifications in 6 lineages; convergence rate inconsistent with neutral drift (PP > 0.95)." |
### Professional / continuing ed
Field-specific terms assumed. Emphasize practice implications.
| ❌ Too academic | ✅ Calibrated |
|---|---|
| "RCT (n=240, 12-week)... LDL-particle count down 14%..." | "12-week RCT shows Mediterranean diet improves LDL-particle metrics 14-22% vs comparators. Practice implication: nutritional counseling for cardiovascular-risk patients should emphasize MUFA-rich foods specifically, not just 'low-fat'." |
## Discussion Question Calibration
Use Bloom's revised taxonomy (Anderson & Krathwohl 2001):
| Level | Action verbs | Question pattern |
|---|---|---|
| Remember | identify, list, recall | "What is X?" "Name the components" |
| Understand | explain, summarize, classify | "Why does X happen?" "How would you describe Y?" |
| Apply | use, apply, demonstrate | "How could this method be applied to...?" "What would happen if we used X for Y?" |
| Analyze | compare, contrast, examine | "What patterns connect X and Y?" "Why do X and Y produce different results?" |
| Evaluate | judge, critique, defend | "Is this study's conclusion warranted by its methods?" "Which approach better serves goal Z, and why?" |
| Create | design, propose, construct | "Design a study that would test the limits of X." "Propose a novel application of Y to Z." |
### Calibration by audience
| Audience | Question levels | Avoid |
|---|---|---|
| Undergrad-intro | Remember + Understand + simple Apply | Pure recall ("what did authors find?") |
| Undergrad-advanced | Understand + Apply + simple Analyze | Sophisticated Evaluate / Create |
| Grad-Masters | Apply + Analyze + Evaluate | Pure recall (insulting) |
| Grad-doctoral | Analyze + Evaluate + Create | Anything below Apply |
### Examples per audience
#### Undergrad-intro
| ❌ Recall only | ✅ Calibrated |
|---|---|
| "What did the authors find?" | "If you wanted to lower your heart disease risk through diet, what does this study suggest you should change?" (Apply) |
#### Grad-doctoral
| ❌ Below level | ✅ Calibrated |
|---|---|
| "What did this RCT show?" | "How would you redesign this RCT to test whether MUFA:SFA ratio specifically (vs total fat composition) drives the lipidomic shift?" (Create) |
## Discussion Question Validator
`scripts/discussion_question_validator.py` flags:
- **Recall-only questions** (any audience): "what did authors find?", "summarize", "describe"
- **Below-audience questions**: undergrad-intro questions in grad course → flag
- **Above-audience questions**: doctoral-level questions in undergrad-intro → flag
Validator suggests upgrades by replacing verbs with audience-appropriate Bloom verbs.
## Tying Discussion Questions to Learning Outcomes
Beyond audience calibration, each question should **explicitly tie to a learning outcome**:
| Without LO tie | With LO tie |
|---|---|
| "How could this approach be applied to...?" | "Course outcome 3 says students should be able to design enzymatic processes. How would the kinetics described in this paper inform a process design for cheese ripening?" |
The LO tie:
- Reinforces course goals
- Shows students why the reading matters
- Creates assessable discussion behaviors
If learning outcomes were inferred (`[inferred]`), still tie discussion questions to them — flag both as inferred.
## Anti-Patterns
### "Same summary for all audiences"
The biggest engagement killer. Undergrad summaries that read like graduate abstracts produce blank stares; graduate summaries that read like K-12 explainers feel patronizing.
### "Add jargon to look academic in undergrad summaries"
Engagement signal: students underline / highlight content. Jargon-heavy summaries get less highlighting in undergrad classes. Plain-language summaries get more.
### "Generic discussion questions"
"What did the authors find?" works for any audience — and serves none. The discussion question is the engagement hook; generic questions waste it.
### "All discussion questions at the highest Bloom level"
In a grad-doctoral course, even one Create-level question per paper is taxing. Mix Analyze, Evaluate, Create. Don't make every reading require students to design a follow-up study.
## Operational Checklist
- [ ] Q2 audience parsed → calibration bucket selected
- [ ] All summaries calibrated to bucket
- [ ] All discussion questions calibrated to bucket's Bloom range
- [ ] Each discussion question tied to a learning outcome (explicit or inferred)
- [ ] Validator (`discussion_question_validator.py`) run on all questions
- [ ] Recall-only questions rejected
- [ ] Below-audience or above-audience questions reworked
## Citations (7 sources)
1. **Bloom, B. S. (1956); Anderson, L. W. & Krathwohl, D. R. (2001), *A Taxonomy for Learning, Teaching, and Assessing*.** The revised Bloom's taxonomy. Source for the 6-level question hierarchy + action verb lexicon.
2. **Marzano, R. J. & Kendall, J. S., *The New Taxonomy of Educational Objectives* (Corwin, 2007).** Modern alternative to Bloom; emphasizes meta-cognitive and self-system levels. Source for the validator's "below-level vs above-level" distinction.
3. **Hattie, J., *Visible Learning* (Routledge, 2008/2023 update).** Meta-meta-analysis of educational interventions. Effect size 0.6+ for "teacher clarity" justifies the audience-calibrated summary discipline (clarity is audience-relative).
4. **Bain, K., *What the Best College Teachers Do* (Harvard, 2004).** Source for the "tied to learning outcome" discipline. Bain's research found great teachers connect every reading explicitly to course-level goals; generic readings produce engagement drop.
5. **Walvoord, B. E. & Anderson, V. J., *Effective Grading* (Jossey-Bass, 2nd ed. 2010).** Source for the "discussion question is assessable behavior" framing. Each discussion question = an opportunity to assess whether learning outcomes are being met.
6. **Brookfield, S. D. & Preskill, S., *Discussion as a Way of Teaching* (Jossey-Bass, 2nd ed. 2005).** Source for the engagement-vs-jargon trade-off in summary writing. Brookfield's research: students engage with content they can paraphrase; jargon-heavy summaries reduce paraphrase capability.
7. **Bjork, R. A. & Bjork, E. L., "Making Things Hard on Yourself, but in a Good Way" — *Psychology and the Real World* (FABBS Foundation, 2011).** Source for the "desirable difficulty" framing. Discussion questions should be challenging at the audience's edge, not below it (insulting) or above it (defeating).
FILE:references/bundled_script_pattern.md
# Bundled Script Pattern — Why JS for DOCX Generation, Not Inline
This reference answers exactly one decision: **why does the syllabus skill ship a bundled `generate_reading_list.js` script rather than inlining the DOCX generation logic in SKILL.md?**
## The Core Trade
DOCX generation requires ~300 lines of `docx`-package boilerplate (table layouts, hyperlink patterns, list formatting, page setup, etc.). This logic is:
1. **Reusable** across runs — every reading list uses the same DOCX layout
2. **Mechanical** — no LLM judgment required; just JSON-in / DOCX-out
3. **Long-lived** — the layout doesn't change between runs
Inlining 300 lines of mechanical layout code in SKILL.md means:
- The skill prompt is much longer (token cost on every invocation)
- Layout changes require editing the skill prompt (high-risk)
- The skill body has to re-derive the same logic each run
Bundling the logic in `scripts/generate_reading_list.js` means:
- The skill body is ~200 lines lighter (token-efficient)
- Layout changes are isolated to one file
- The skill orchestrates; the script executes mechanically
## When to Bundle (vs Inline)
### Bundle when:
- ✅ The logic is mechanical (no LLM judgment)
- ✅ The logic is reusable across runs (same layout / same algorithm)
- ✅ The logic is non-trivial (>50 lines)
- ✅ The logic is in a non-Python language (JS, Go, Rust, etc.)
- ✅ The logic has external dependencies (`docx` package, `requests`, etc.)
### Inline (in SKILL.md) when:
- The logic requires LLM judgment per run (e.g., paper-summary writing)
- The logic is short (<20 lines) and run-specific
- The logic is in-context-only (uses session-specific tool calls)
- The logic varies significantly per invocation
## The Pattern Used Here
`scripts/generate_reading_list.js`:
1. **Accepts JSON input + output path as CLI args**
```bash
node generate_reading_list.js --input data.json --output result.docx
```
2. **Has a documented JSON schema** (in SKILL.md so the orchestrator knows what to produce)
3. **Handles `docx` require with multi-location fallback** (works whether `docx` is installed locally, globally, or in a parent dir)
4. **Validates input** (missing fields → graceful error, not silent failure)
5. **Produces a clean professional DOCX** with:
- Title page
- Introduction (with Consensus link)
- Learning outcomes box
- Numbered papers per section
- Footer with metadata
6. **Uses canonical `docx` patterns**:
- `ExternalHyperlink` with full URLs
- `LevelFormat.BULLET` for lists
- Dual-width tables (`columnWidths` + cell `width`)
## Skill Orchestrator's Role
The skill body (SKILL.md):
1. Walks Phase 0 intake
2. Parses syllabus + extracts topics
3. Walks group-and-confirm checkpoint
4. Runs Consensus searches (LLM judgment per query)
5. Writes summaries + discussion questions (LLM judgment per paper)
6. **Constructs the JSON payload** matching the bundled script's schema
7. **Invokes the script** with the JSON
8. Validates output + delivers
The skill body is responsible for **what goes in the document**. The script is responsible for **how it's laid out**.
## Why Node.js Specifically
The `docx` library is a JavaScript library (npm package). Could the skill use a Python `docx` library (`python-docx`)? Yes, but:
- The repo's other research-pack DOCX-generating skills (litreview, grants, dossier) all use Node.js + `docx`
- Consistency: one DOCX library across the research pack
- The `docx` JS library is more actively maintained + has richer features
- `python-docx` doesn't support all the features the skill needs (advanced hyperlinks, table styling)
## File Structure
```
research/syllabus/skills/syllabus/scripts/
├── citation_tracker.py ← stdlib Python (orchestration helper)
├── topic_grouper.py ← stdlib Python (orchestration helper)
├── discussion_question_validator.py ← stdlib Python (orchestration helper)
└── generate_reading_list.js ← BUNDLED Node.js (mechanical DOCX assembly)
```
The Python scripts are stateless helpers (per-run). The JS script is the bundled mechanical assembler (called once per run).
## Anti-Patterns
### "Inline the JS into a Python script via subprocess"
Adds an unnecessary layer. The skill should call `node` directly.
### "Convert JS logic to Python to keep all scripts in one language"
Loses access to the better-maintained `docx` JS library. Worse: would diverge from sibling skills (litreview, grants, dossier all use `docx` JS).
### "Keep the JS script but inline the JSON schema in the script"
The JSON schema needs to be IN SKILL.md so the orchestrator knows what to construct. Documenting it in the script alone hides it from the orchestrator's prompt context.
### "Inline 300 lines of docx code in SKILL.md"
The original anti-pattern. Bloats the prompt, makes layout changes risky, makes the skill body harder to read.
### "Import the script from another skill"
Cross-skill dependencies break the per-skill self-contained discipline (per CLAUDE.md anti-patterns). Even though it would save duplication, the bundled script lives within syllabus's own folder.
## Operational Checklist
- [ ] `scripts/generate_reading_list.js` exists in syllabus's scripts/ folder
- [ ] Script accepts `--input <json>` + `--output <docx>` CLI args
- [ ] Script handles `docx` require with multi-location fallback
- [ ] Script validates input (missing fields → graceful error)
- [ ] JSON schema documented in SKILL.md (not just in the script)
- [ ] Skill orchestrator constructs JSON matching the schema
- [ ] Skill orchestrator invokes the script via `node` (not `python`)
- [ ] DOCX output validated post-generation
## Citations (7 sources)
1. **Karpathy-coder discipline + write-a-skill conventions** (this repo's `engineering/write-a-skill/`). Source for the "stdlib-only Python tools, bundled non-Python scripts allowed for mechanical jobs" pattern.
2. **CLAUDE.md anti-pattern: "Don't add features beyond what the task requires."** The bundled script honors this — it does ONE thing (DOCX layout) and does it mechanically.
3. **`docx` Node.js package — github.com/dolanmiu/docx (MIT).** Authoritative source for the API patterns the bundled script uses. Active maintenance, comprehensive feature set.
4. **CommonJS / Node.js module resolution algorithm.** Source for the "multi-location fallback" pattern in the require statement. Ensures the script works in development (local node_modules) and production (global install).
5. **Twelve-Factor App principles — III. Config: store config in the environment.** Source for the CLI-args-not-config pattern. Script accepts input/output as args, not via env vars or config files.
6. **Brian Kernighan & P. J. Plauger, *Software Tools* (1976).** Source for the "do one thing well + compose" pattern. The bundled script does exactly one thing (mechanical DOCX assembly); the skill body composes it with the rest of the pipeline.
7. **Doug McIlroy / Unix philosophy.** Source for the broader pattern: "Write programs that do one thing and do it well. Write programs to work together. Write programs to handle text streams, because that is a universal interface." JSON-in / DOCX-out is the modern equivalent.
FILE:scripts/citation_tracker.py
#!/usr/bin/env python3
"""citation_tracker.py — Syllabus three-count audit + 1s sequential discipline.
Stdlib-only. Mirrors litreview's citation_tracker (research-pack convention)
adapted for syllabus's per-section search budget.
Tracked counts:
- searches_total
- searches_per_section
- papers_received
- papers_cited
Per-section detail recorded for DOCX audit log.
Enforces 1s sequential gap.
Usage:
python citation_tracker.py --action start --session syllabus-bio101-20260515 --course "Intro Biology"
python citation_tracker.py --action record_search --session ... --section "Cell Biology" --query "..."
python citation_tracker.py --action record_received --session ... --section "Cell Biology" --count 3
python citation_tracker.py --action record_cited --session ... --section "Cell Biology" --url "..."
python citation_tracker.py --action status --session ...
"""
import argparse
import json
import sys
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Dict, List, Optional
SESSIONS_DIR = Path.home() / ".syllabus_sessions"
MIN_GAP_SECONDS = 1.0
def session_path(name: str) -> Path:
return SESSIONS_DIR / f"{name}.json"
def load_session(name: str) -> Dict[str, Any]:
p = session_path(name)
if not p.exists():
raise FileNotFoundError(f"Session not found: {name}")
return json.loads(p.read_text(encoding="utf-8"))
def save_session(name: str, data: Dict[str, Any]) -> None:
SESSIONS_DIR.mkdir(parents=True, exist_ok=True)
session_path(name).write_text(json.dumps(data, indent=2), encoding="utf-8")
def now_iso() -> str:
return datetime.now(timezone.utc).isoformat()
def now_ts() -> float:
return datetime.now(timezone.utc).timestamp()
def action_start(name: str, course: Optional[str], audience: Optional[str], year_range: Optional[str]) -> Dict[str, Any]:
if session_path(name).exists():
raise FileExistsError(f"Session already exists: {name}")
data: Dict[str, Any] = {
"session": name,
"course": course or "",
"audience": audience or "",
"year_range": year_range or "",
"consensus_tier": None,
"started_at": now_iso(),
"ended_at": None,
"searches": [],
"received_log": [],
"cited": [],
"counts": {
"searches_total": 0,
"papers_received_total": 0,
"papers_cited_total": 0,
},
"by_section": {},
}
save_session(name, data)
return data
def action_record_search(name: str, section: str, query: str, tier: Optional[str]) -> Dict[str, Any]:
data = load_session(name)
if data["searches"]:
last_ts = data["searches"][-1].get("ts", 0)
gap = now_ts() - last_ts
if gap < MIN_GAP_SECONDS:
raise RuntimeError(
f"Sequential discipline violated: {gap:.2f}s gap (need >= {MIN_GAP_SECONDS}s). "
f"Wait {MIN_GAP_SECONDS - gap:.2f}s more."
)
if tier and not data["consensus_tier"]:
data["consensus_tier"] = tier
data["searches"].append({"section": section, "query": query, "tier": tier, "at": now_iso(), "ts": now_ts()})
data["counts"]["searches_total"] += 1
if section not in data["by_section"]:
data["by_section"][section] = {"searches": 0, "received": 0, "cited": 0}
data["by_section"][section]["searches"] += 1
save_session(name, data)
return data
def action_record_received(name: str, section: str, count: int) -> Dict[str, Any]:
data = load_session(name)
data["received_log"].append({"section": section, "count": count, "at": now_iso()})
data["counts"]["papers_received_total"] += count
if section not in data["by_section"]:
data["by_section"][section] = {"searches": 0, "received": 0, "cited": 0}
data["by_section"][section]["received"] += count
save_session(name, data)
return data
def action_record_cited(name: str, section: str, url: str, title: Optional[str]) -> Dict[str, Any]:
data = load_session(name)
if any(c["url"] == url for c in data["cited"]):
return data
data["cited"].append({"section": section, "url": url, "title": title, "at": now_iso()})
data["counts"]["papers_cited_total"] += 1
if section not in data["by_section"]:
data["by_section"][section] = {"searches": 0, "received": 0, "cited": 0}
data["by_section"][section]["cited"] += 1
save_session(name, data)
return data
def action_status(name: str) -> Dict[str, Any]:
return load_session(name)
def action_close(name: str) -> Dict[str, Any]:
data = load_session(name)
if data.get("ended_at") is None:
data["ended_at"] = now_iso()
save_session(name, data)
return data
def render_status_human(data: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"Session: {data['session']}")
out.append(f"Course: {data.get('course', '(unset)')}")
out.append(f"Audience: {data.get('audience', '(unset)')}")
out.append(f"Year range: {data.get('year_range', '(unset)')}")
out.append(f"Consensus tier: {data.get('consensus_tier') or '(not detected)'}")
out.append(f"Started: {data['started_at']}")
out.append(f"Ended: {data.get('ended_at') or '(active)'}")
out.append("")
c = data["counts"]
out.append(f"Total searches: {c['searches_total']}")
out.append(f"Total received: {c['papers_received_total']}")
out.append(f"Total cited: {c['papers_cited_total']}")
out.append("")
if data["by_section"]:
out.append("Per-section breakdown:")
for section, stats in data["by_section"].items():
out.append(f" {section:<40s} {stats['searches']} searches → {stats['received']} received → {stats['cited']} cited")
out.append("")
out.append("Audit block (paste in DOCX audit-log section):")
out.append(
f" Total queries: {c['searches_total']}. Papers received: {c['papers_received_total']}. "
f"Papers cited: {c['papers_cited_total']}. "
f"Plan tier: {data.get('consensus_tier') or 'undetected'}."
)
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--action", required=True, choices=["start", "record_search", "record_received", "record_cited", "status", "list", "close"])
parser.add_argument("--session")
parser.add_argument("--course")
parser.add_argument("--audience")
parser.add_argument("--year-range")
parser.add_argument("--section")
parser.add_argument("--query")
parser.add_argument("--tier")
parser.add_argument("--count", type=int)
parser.add_argument("--url")
parser.add_argument("--title")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
try:
if args.action == "start":
result = action_start(args.session, args.course, args.audience, args.year_range)
elif args.action == "record_search":
result = action_record_search(args.session, args.section, args.query, args.tier)
elif args.action == "record_received":
result = action_record_received(args.session, args.section, args.count)
elif args.action == "record_cited":
result = action_record_cited(args.session, args.section, args.url, args.title)
elif args.action == "status":
result = action_status(args.session)
elif args.action == "close":
result = action_close(args.session)
else:
SESSIONS_DIR.mkdir(parents=True, exist_ok=True)
result = [{"session": p.stem, "data": json.loads(p.read_text(encoding="utf-8"))} for p in sorted(SESSIONS_DIR.glob("*.json"))]
except (FileNotFoundError, FileExistsError, RuntimeError) as e:
print(f"error: {e}", file=sys.stderr); return 2
if args.output == "json":
print(json.dumps(result, indent=2, default=str))
else:
if args.action == "list":
print(json.dumps(result, indent=2, default=str))
else:
print(render_status_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/discussion_question_validator.py
#!/usr/bin/env python3
"""discussion_question_validator.py — Bloom higher-order quality check.
Stdlib-only. Validates each discussion question against Bloom's revised
taxonomy (Anderson & Krathwohl 2001). Flags:
- Recall-only questions (any audience): "what did authors find?", "summarize", etc.
- Below-audience questions (e.g., grad-doctoral course with undergrad-intro questions)
- Above-audience questions (e.g., undergrad-intro course with doctoral-level questions)
Suggests upgrades by replacing low-tier verbs with audience-appropriate Bloom verbs.
NO LLM CALLS. Pure regex + verb classification.
Usage:
python discussion_question_validator.py --questions-file /tmp/questions.json --audience grad_masters
python discussion_question_validator.py --question "What did the authors find?" --audience undergrad_intro
python discussion_question_validator.py --sample
"""
import argparse
import json
import re
import sys
from typing import Any, Dict, List, Optional
VALID_AUDIENCES = ["undergrad_intro", "undergrad_advanced", "grad_masters", "grad_doctoral", "professional", "mixed"]
# Bloom's revised taxonomy verb classification
BLOOM_VERBS = {
"remember": ["identify", "list", "recall", "name", "define", "label", "match", "recognize", "state", "what is", "what are", "what did", "describe what"],
"understand": ["explain", "summarize", "classify", "compare", "contrast", "describe how", "interpret", "paraphrase", "translate"],
"apply": ["use", "apply", "demonstrate", "implement", "execute", "carry out", "how could you use", "how would you apply", "how could this be applied", "what would happen if"],
"analyze": ["compare", "contrast", "examine", "differentiate", "organize", "what patterns", "why do", "what connections", "deconstruct"],
"evaluate": ["judge", "critique", "defend", "justify", "argue", "is this", "should we", "which is better", "do you agree", "evaluate the"],
"create": ["design", "propose", "construct", "develop", "formulate", "create a", "design a", "what would you propose", "how would you redesign"],
}
# Audience → minimum acceptable Bloom level
AUDIENCE_MIN_BLOOM = {
"undergrad_intro": 1, # Remember+ acceptable, but apply+ preferred
"undergrad_advanced": 2, # Understand+
"grad_masters": 3, # Apply+
"grad_doctoral": 4, # Analyze+
"professional": 3, # Apply+ (practice-oriented)
"mixed": 2, # Understand+ (lowest bucket present)
}
BLOOM_LEVEL_ORDER = ["remember", "understand", "apply", "analyze", "evaluate", "create"]
def classify_question(question: str) -> Dict[str, Any]:
"""Classify question by Bloom level."""
q_lower = question.lower()
detected_levels: List[str] = []
matched_phrases: Dict[str, List[str]] = {}
for level, verbs in BLOOM_VERBS.items():
for verb in verbs:
if re.search(rf"\b{re.escape(verb)}\b", q_lower):
if level not in detected_levels:
detected_levels.append(level)
matched_phrases.setdefault(level, []).append(verb)
if not detected_levels:
# Default heuristic: if starts with "what/why/how", probably understand or apply
if q_lower.strip().startswith(("what", "why", "how")):
detected_levels = ["understand"]
matched_phrases["understand"] = ["(inferred from interrogative)"]
else:
detected_levels = ["unknown"]
# Highest Bloom level detected
highest_level = "unknown"
highest_idx = -1
for level in detected_levels:
if level in BLOOM_LEVEL_ORDER:
idx = BLOOM_LEVEL_ORDER.index(level)
if idx > highest_idx:
highest_idx = idx
highest_level = level
return {
"question": question,
"detected_levels": detected_levels,
"highest_level": highest_level,
"highest_level_index": highest_idx,
"matched_phrases": matched_phrases,
}
def validate_against_audience(question: str, audience: str) -> Dict[str, Any]:
if audience not in VALID_AUDIENCES:
raise ValueError(f"Invalid audience '{audience}'. Pick from: {VALID_AUDIENCES}")
classification = classify_question(question)
min_required_idx = AUDIENCE_MIN_BLOOM[audience] - 1 # convert level to 0-indexed
detected_idx = classification["highest_level_index"]
if detected_idx == -1:
verdict = "WARN"
message = f"Could not detect Bloom level. Manual review recommended."
elif detected_idx < min_required_idx:
verdict = "FAIL"
required_level = BLOOM_LEVEL_ORDER[min_required_idx]
message = (
f"Question level '{classification['highest_level']}' is BELOW required minimum "
f"'{required_level}' for {audience}. Rework with verbs from higher Bloom levels."
)
elif detected_idx > min_required_idx + 2:
verdict = "WARN"
target_level = BLOOM_LEVEL_ORDER[min_required_idx]
message = (
f"Question level '{classification['highest_level']}' may be ABOVE typical "
f"{audience} level. Consider whether students can engage at {target_level} level."
)
else:
verdict = "PASS"
message = f"Question level '{classification['highest_level']}' appropriate for {audience}."
suggested_upgrades: List[str] = []
if verdict == "FAIL":
target_level = BLOOM_LEVEL_ORDER[min_required_idx]
suggested_upgrades = [
f"Replace verb with: {', '.join(BLOOM_VERBS[target_level][:5])}",
f"Pattern: '{BLOOM_VERBS[target_level][0]} [the {target_level} concept]...'",
]
return {
"verdict": verdict,
"audience": audience,
"min_required_level": BLOOM_LEVEL_ORDER[min_required_idx] if min_required_idx >= 0 else "unknown",
"classification": classification,
"message": message,
"suggested_upgrades": suggested_upgrades,
}
SAMPLE_QUESTIONS = [
{"question": "What did the authors find?", "audience": "undergrad_intro"},
{"question": "What did the authors find?", "audience": "grad_doctoral"},
{"question": "How could you apply this method to clinical decision support for sepsis?", "audience": "grad_masters"},
{"question": "Design a follow-up study that would test whether MUFA:SFA ratio specifically drives the lipidomic shift.", "audience": "grad_doctoral"},
{"question": "Why does the Mediterranean diet improve lipoprotein profiles?", "audience": "undergrad_intro"},
]
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--question", help="Single question to validate")
parser.add_argument("--questions-file", help="JSON file with [{question, audience}, ...] entries")
parser.add_argument("--audience", choices=VALID_AUDIENCES, help="Course audience for the question(s)")
parser.add_argument("--sample", action="store_true")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
results: List[Dict[str, Any]] = []
try:
if args.sample:
for sq in SAMPLE_QUESTIONS:
results.append(validate_against_audience(sq["question"], sq["audience"]))
elif args.question and args.audience:
results.append(validate_against_audience(args.question, args.audience))
elif args.questions_file:
from pathlib import Path
p = Path(args.questions_file)
if not p.exists():
print(f"error: {args.questions_file} not found", file=sys.stderr); return 2
data = json.loads(p.read_text(encoding="utf-8"))
for item in data:
results.append(validate_against_audience(item["question"], item["audience"]))
else:
parser.print_help(); return 0
except ValueError as e:
print(f"error: {e}", file=sys.stderr); return 2
if args.output == "json":
print(json.dumps(results, indent=2))
else:
for r in results:
marker = {"PASS": "[ok]", "WARN": "[warn]", "FAIL": "[FAIL]"}[r["verdict"]]
print(f"{marker} ({r['audience']:<20s}) {r['classification']['question'][:80]}")
print(f" Highest Bloom level: {r['classification']['highest_level']}; required: {r['min_required_level']}")
print(f" → {r['message']}")
if r["suggested_upgrades"]:
print(f" Suggested upgrades:")
for s in r["suggested_upgrades"]:
print(f" - {s}")
print()
fail_count = sum(1 for r in results if r["verdict"] == "FAIL")
return 1 if fail_count > 0 else 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/generate_reading_list.js
#!/usr/bin/env node
/**
* generate_reading_list.js — Bundled DOCX generator for syllabus skill.
*
* Accepts JSON input + output path as CLI args. Produces a clean professional
* .docx reading list with title page, learning outcomes, sections of papers
* (each with hyperlinked title + audience-calibrated summary + Bloom-tied
* discussion question), and footer.
*
* Path-B build: this is the bundled mechanical layout logic. The skill
* orchestrator constructs JSON; this script assembles the DOCX. ~300 lines.
*
* Handles `docx` package require with multi-location fallback (works whether
* `docx` is installed locally, globally, or in a parent dir).
*
* JSON schema (documented in SKILL.md):
* { courseTitle, courseSubtitle, generatedDate, yearRange, introText,
* learningOutcomes: [], sections: [{ heading, papers: [...] }],
* auditLog: { totalQueriesSent, totalPapersReceived, totalPapersCited,
* toolConstraints, searchDetails: [], failures: [] } }
*
* Usage:
* node generate_reading_list.js --input data.json --output result.docx
*/
'use strict';
const fs = require('fs');
const path = require('path');
// Multi-location require for docx package
function loadDocx() {
const candidates = [
'docx', // Local node_modules
path.join(process.cwd(), 'node_modules', 'docx'), // Explicit local
'/usr/lib/node_modules/docx', // Global Linux
'/usr/local/lib/node_modules/docx', // Global macOS / brew
path.join(process.env.HOME || '', '.npm-global', 'lib', 'node_modules', 'docx'),
];
for (const candidate of candidates) {
try {
return require(candidate);
} catch (e) {
// try next
}
}
console.error('error: cannot find `docx` npm package. Install with: npm install docx');
process.exit(2);
}
const docx = loadDocx();
const {
Document, Paragraph, TextRun, Packer, AlignmentType, HeadingLevel,
ExternalHyperlink, Table, TableRow, TableCell, WidthType, ShadingType,
LevelFormat, Footer, Header, PageNumber, PageBreak, BorderStyle,
} = docx;
// ----------------------------------------------------------------------------
// CLI args
// ----------------------------------------------------------------------------
function parseArgs() {
const args = process.argv.slice(2);
const opts = {};
for (let i = 0; i < args.length; i++) {
if (args[i] === '--input') opts.input = args[++i];
else if (args[i] === '--output') opts.output = args[++i];
else if (args[i] === '--help' || args[i] === '-h') {
console.log('Usage: node generate_reading_list.js --input <data.json> --output <result.docx>');
process.exit(0);
}
}
if (!opts.input || !opts.output) {
console.error('error: both --input and --output are required');
console.error('Usage: node generate_reading_list.js --input <data.json> --output <result.docx>');
process.exit(2);
}
return opts;
}
// ----------------------------------------------------------------------------
// Input validation
// ----------------------------------------------------------------------------
function validateInput(data) {
const required = ['courseTitle', 'sections'];
for (const field of required) {
if (!data[field]) {
console.error(`error: missing required field 'field' in input JSON`);
process.exit(2);
}
}
if (!Array.isArray(data.sections) || data.sections.length === 0) {
console.error('error: sections must be a non-empty array');
process.exit(2);
}
for (const section of data.sections) {
if (!section.heading || !Array.isArray(section.papers)) {
console.error('error: each section must have heading + papers array');
process.exit(2);
}
for (const paper of section.papers) {
if (!paper.title || !paper.url) {
console.error('error: each paper must have title + url');
process.exit(2);
}
}
}
}
// ----------------------------------------------------------------------------
// DOCX building blocks
// ----------------------------------------------------------------------------
const NAVY = '1A3A5C';
const LIGHT_BLUE = 'E8F0F8';
const ACCENT_BLUE = '2E5C8A';
const GRAY = '808080';
const DARK_GRAY = '404040';
function buildTitlePage(data) {
return [
new Paragraph({
children: [new TextRun({ text: data.courseTitle, bold: true, size: 48, color: NAVY })],
alignment: AlignmentType.CENTER,
spacing: { before: 2400, after: 200 },
}),
new Paragraph({
children: [new TextRun({ text: 'Supplementary Reading List', bold: false, size: 28, color: ACCENT_BLUE })],
alignment: AlignmentType.CENTER,
spacing: { after: 200 },
}),
data.courseSubtitle ? new Paragraph({
children: [new TextRun({ text: data.courseSubtitle, italics: true, size: 22, color: DARK_GRAY })],
alignment: AlignmentType.CENTER,
spacing: { after: 800 },
}) : null,
new Paragraph({
children: [new TextRun({ text: `Generated: data.generatedDate || new Date().toISOString().split('T')[0]`, size: 18, color: GRAY })],
alignment: AlignmentType.CENTER,
spacing: { after: 100 },
}),
new Paragraph({
children: [new TextRun({ text: `Year range: data.yearRange || 'last 2 years'`, size: 18, color: GRAY })],
alignment: AlignmentType.CENTER,
spacing: { after: 200 },
}),
new Paragraph({ children: [new PageBreak()] }),
].filter(Boolean);
}
function buildIntroSection(data) {
const introText = data.introText || 'This supplementary reading list collects recent peer-reviewed research relevant to each section of the course. Each entry includes a plain-language summary calibrated to the course audience and a discussion question tied to the course learning outcomes.';
return [
new Paragraph({
heading: HeadingLevel.HEADING_1,
children: [new TextRun({ text: 'Introduction', color: NAVY, bold: true, size: 32 })],
spacing: { after: 200 },
}),
new Paragraph({
children: [new TextRun({ text: introText, size: 22 })],
spacing: { after: 200 },
}),
new Paragraph({
children: [
new TextRun({ text: 'Papers sourced via ', size: 20 }),
new ExternalHyperlink({
link: 'https://consensus.app',
children: [new TextRun({ text: 'Consensus', style: 'Hyperlink', size: 20 })],
}),
new TextRun({ text: ' academic search. URLs in this document link directly to Consensus paper records.', size: 20 }),
],
spacing: { after: 400 },
}),
];
}
function buildLearningOutcomesBox(outcomes) {
if (!outcomes || outcomes.length === 0) return [];
const cells = [
new TableRow({
children: [
new TableCell({
width: { size: 9000, type: WidthType.DXA },
shading: { type: ShadingType.CLEAR, color: 'auto', fill: LIGHT_BLUE },
children: [
new Paragraph({
children: [new TextRun({ text: 'Course Learning Outcomes', bold: true, size: 24, color: NAVY })],
spacing: { after: 100 },
}),
...outcomes.map(outcome => new Paragraph({
children: [new TextRun({ text: '• ' + outcome, size: 20 })],
spacing: { after: 60 },
})),
],
}),
],
}),
];
return [
new Table({
columnWidths: [9000],
rows: cells,
}),
new Paragraph({ children: [new TextRun({ text: '', size: 4 })], spacing: { after: 400 } }),
];
}
function buildSection(section, sectionIndex) {
const elements = [
new Paragraph({
heading: HeadingLevel.HEADING_1,
children: [new TextRun({ text: `sectionIndex. section.heading`, color: NAVY, bold: true, size: 28 })],
spacing: { before: 400, after: 200 },
}),
];
for (let i = 0; i < section.papers.length; i++) {
const paper = section.papers[i];
const paperNum = `sectionIndex.i + 1`;
// Title (hyperlinked)
elements.push(new Paragraph({
children: [
new TextRun({ text: `paperNum. `, bold: true, size: 22 }),
new ExternalHyperlink({
link: paper.url,
children: [new TextRun({ text: paper.title, style: 'Hyperlink', size: 22, bold: true })],
}),
],
spacing: { after: 60 },
}));
// Author / journal / year (italic gray)
const meta = `paper.authors || ''''''`;
if (meta.trim()) {
elements.push(new Paragraph({
children: [new TextRun({ text: meta, italics: true, size: 18, color: GRAY })],
spacing: { after: 60 },
}));
}
// Summary
if (paper.summary) {
elements.push(new Paragraph({
children: [
new TextRun({ text: 'Summary: ', bold: true, size: 20 }),
new TextRun({ text: paper.summary, size: 20 }),
],
spacing: { after: 60 },
}));
}
// Discussion question (blue accent)
if (paper.question) {
elements.push(new Paragraph({
children: [
new TextRun({ text: 'Discussion: ', bold: true, size: 20, color: ACCENT_BLUE }),
new TextRun({ text: paper.question, size: 20 }),
],
spacing: { after: 200 },
}));
}
}
return elements;
}
function buildAuditLogSection(audit) {
if (!audit) return [];
const elements = [
new Paragraph({ children: [new PageBreak()] }),
new Paragraph({
heading: HeadingLevel.HEADING_1,
children: [new TextRun({ text: 'Audit Log', color: NAVY, bold: true, size: 28 })],
spacing: { after: 200 },
}),
new Paragraph({
children: [
new TextRun({ text: `Total queries sent: `, bold: true, size: 20 }),
new TextRun({ text: `audit.totalQueriesSent || 0`, size: 20 }),
],
spacing: { after: 60 },
}),
new Paragraph({
children: [
new TextRun({ text: `Total papers received: `, bold: true, size: 20 }),
new TextRun({ text: `audit.totalPapersReceived || 0`, size: 20 }),
],
spacing: { after: 60 },
}),
new Paragraph({
children: [
new TextRun({ text: `Total papers cited in this list: `, bold: true, size: 20 }),
new TextRun({ text: `audit.totalPapersCited || 0`, size: 20 }),
],
spacing: { after: 200 },
}),
];
if (audit.toolConstraints) {
elements.push(new Paragraph({
children: [
new TextRun({ text: 'Tool constraints: ', bold: true, size: 20 }),
new TextRun({ text: audit.toolConstraints, size: 20 }),
],
spacing: { after: 200 },
}));
}
if (Array.isArray(audit.searchDetails) && audit.searchDetails.length > 0) {
elements.push(new Paragraph({
children: [new TextRun({ text: 'Per-search detail:', bold: true, size: 22, color: NAVY })],
spacing: { after: 100 },
}));
for (const sd of audit.searchDetails) {
elements.push(new Paragraph({
children: [
new TextRun({ text: `• sd.section || 'Unassigned': `, bold: true, size: 18 }),
new TextRun({ text: `"sd.query" → sd.papersReturned || 0 returned, sd.papersSelected || 0 selected (sd.status || 'OK')`, size: 18 }),
],
spacing: { after: 40 },
}));
}
}
if (Array.isArray(audit.failures) && audit.failures.length > 0) {
elements.push(new Paragraph({
children: [new TextRun({ text: 'Failures:', bold: true, size: 22, color: 'AA0000' })],
spacing: { before: 200, after: 100 },
}));
for (const f of audit.failures) {
elements.push(new Paragraph({
children: [new TextRun({ text: `• f`, size: 18, color: '880000' })],
spacing: { after: 40 },
}));
}
}
return elements;
}
function buildFooter(data) {
return new Footer({
children: [
new Paragraph({
children: [new TextRun({ text: `data.courseTitle — Supplementary Reading List`, size: 16, color: GRAY })],
alignment: AlignmentType.CENTER,
}),
],
});
}
// ----------------------------------------------------------------------------
// Main
// ----------------------------------------------------------------------------
function main() {
const opts = parseArgs();
let data;
try {
data = JSON.parse(fs.readFileSync(opts.input, 'utf-8'));
} catch (e) {
console.error(`error: cannot read input JSON opts.input: e.message`);
process.exit(2);
}
validateInput(data);
const sections = data.sections.map((s, i) => buildSection(s, i + 1)).flat();
const doc = new Document({
creator: 'syllabus skill',
title: `data.courseTitle — Supplementary Reading List`,
description: 'Generated by syllabus skill via bundled generate_reading_list.js',
sections: [
{
properties: {
page: {
margin: { top: 1440, right: 1440, bottom: 1440, left: 1440 }, // 1 inch
size: { width: 12240, height: 15840 }, // US Letter
},
},
footers: { default: buildFooter(data) },
children: [
...buildTitlePage(data),
...buildIntroSection(data),
...buildLearningOutcomesBox(data.learningOutcomes),
...sections,
...buildAuditLogSection(data.auditLog),
],
},
],
});
Packer.toBuffer(doc).then(buffer => {
fs.writeFileSync(opts.output, buffer);
console.log(`Generated: opts.output (buffer.length bytes, data.sections.length sections, data.sections.reduce((sum, s) => sum + s.papers.length, 0) papers)`);
}).catch(e => {
console.error(`error: DOCX packing failed: e.message`);
process.exit(2);
});
}
main();
FILE:scripts/topic_grouper.py
#!/usr/bin/env python3
"""topic_grouper.py — Heuristic 6-12 section grouping from extracted syllabus topics.
Stdlib-only. Given a list of extracted course topics, produce a proposed
grouping into 6-12 sections by detecting shared keywords.
The output feeds the Phase 2 group-and-confirm checkpoint where the user
can override (proceed / merge / split / add / remove).
Algorithm:
1. Tokenize each topic into significant words (stop-words removed)
2. Build word → topics inverted index
3. Greedy clustering: topics sharing 2+ significant words → same section
4. Cap at 12 sections (over-cap → merge smallest); ensure minimum 6 (under → split largest)
5. Each section gets a heading derived from its dominant shared keyword
NO LLM CALLS. Pure tokenization + clustering.
Usage:
python topic_grouper.py --topics "Cell biology, DNA replication, Protein synthesis, ..."
python topic_grouper.py --topics-file /tmp/topics.json
python topic_grouper.py --sample
"""
import argparse
import json
import re
import sys
from collections import Counter, defaultdict
from typing import Any, Dict, List, Set
STOP_WORDS = {
"the", "a", "an", "and", "or", "but", "if", "of", "in", "on", "at", "to",
"for", "with", "by", "from", "is", "are", "was", "were", "be", "been",
"this", "that", "these", "those", "introduction", "overview", "basics",
"fundamentals", "principles", "concepts", "topics", "review", "advanced",
"intermediate", "i", "ii", "iii", "iv", "v", "1", "2", "3", "4", "5",
"6", "7", "8", "9", "10", "11", "12", "week", "lecture", "chapter", "unit",
"module", "lesson", "section",
}
MIN_SECTIONS = 6
MAX_SECTIONS = 12
SHARED_WORD_THRESHOLD = 2
def tokenize(topic: str) -> Set[str]:
"""Extract significant words from a topic string."""
words = re.findall(r"\b[a-z]{3,}\b", topic.lower())
return {w for w in words if w not in STOP_WORDS}
def cluster_topics(topics: List[str]) -> List[Dict[str, Any]]:
"""Cluster topics by shared significant words."""
topic_tokens = [(i, t, tokenize(t)) for i, t in enumerate(topics)]
clusters: List[List[int]] = [] # list of topic-index lists
assigned: Set[int] = set()
for i, _, tokens_i in topic_tokens:
if i in assigned:
continue
# Start a new cluster with topic i
cluster = [i]
assigned.add(i)
# Try to add other topics that share >= SHARED_WORD_THRESHOLD tokens
for j, _, tokens_j in topic_tokens:
if j in assigned or j == i:
continue
shared = tokens_i & tokens_j
if len(shared) >= SHARED_WORD_THRESHOLD:
cluster.append(j)
assigned.add(j)
clusters.append(cluster)
return _normalize_to_size(clusters, topic_tokens)
def _normalize_to_size(clusters: List[List[int]], topic_tokens: List[tuple]) -> List[Dict[str, Any]]:
"""Ensure 6-12 sections by merging smallest or splitting largest."""
# Merge smallest if over MAX_SECTIONS
while len(clusters) > MAX_SECTIONS:
clusters.sort(key=len)
smallest = clusters.pop(0)
# Merge into next-smallest
if clusters:
clusters[0].extend(smallest)
else:
clusters.append(smallest)
# Split largest if under MIN_SECTIONS (and largest has >= 4 items)
while len(clusters) < MIN_SECTIONS and clusters:
clusters.sort(key=len, reverse=True)
largest = clusters.pop(0)
if len(largest) >= 4:
mid = len(largest) // 2
clusters.extend([largest[:mid], largest[mid:]])
else:
clusters.insert(0, largest)
break # Can't split further
# Generate section heading per cluster (most-common shared word)
sections: List[Dict[str, Any]] = []
topic_lookup = {i: (t, tokens) for i, t, tokens in topic_tokens}
for cluster_indices in clusters:
all_tokens: Counter = Counter()
cluster_topics: List[str] = []
for idx in cluster_indices:
topic, tokens = topic_lookup[idx]
cluster_topics.append(topic)
all_tokens.update(tokens)
# Heading = top 1-3 most common tokens, capitalized
top_words = [w for w, _ in all_tokens.most_common(2)]
heading = " + ".join(w.capitalize() for w in top_words) if top_words else f"Section {len(sections) + 1}"
sections.append({
"heading": heading,
"topic_count": len(cluster_indices),
"topics": cluster_topics,
})
return sections
SAMPLE_TOPICS = [
"Cell Biology Fundamentals",
"DNA Replication",
"Protein Synthesis",
"Cell Division and Mitosis",
"Mendelian Genetics",
"Population Genetics",
"Evolution and Natural Selection",
"Speciation",
"Ecology Basics",
"Ecosystem Dynamics",
"Energy Flow in Ecosystems",
"Conservation Biology",
"Plant Anatomy",
"Plant Physiology",
"Animal Anatomy Overview",
"Animal Behavior",
"Microbiology Introduction",
"Bacterial Genetics",
"Viruses and Pathogens",
]
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--topics", help="Comma-separated topic list")
parser.add_argument("--topics-file", help="Path to JSON file with topics array")
parser.add_argument("--sample", action="store_true")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
topics = SAMPLE_TOPICS
elif args.topics:
topics = [t.strip() for t in args.topics.split(",") if t.strip()]
elif args.topics_file:
from pathlib import Path
p = Path(args.topics_file)
if not p.exists():
print(f"error: {args.topics_file} not found", file=sys.stderr); return 2
topics = json.loads(p.read_text(encoding="utf-8"))
else:
parser.print_help(); return 0
sections = cluster_topics(topics)
result = {
"input_topic_count": len(topics),
"section_count": len(sections),
"sections": sections,
}
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(f"Input topics: {len(topics)}")
print(f"Output sections: {len(sections)} (target: {MIN_SECTIONS}-{MAX_SECTIONS})")
print()
print("Proposed sections (present this at Phase 2 checkpoint):")
for i, s in enumerate(sections, 1):
print(f"")
print(f" Section {i}: {s['heading']} ({s['topic_count']} topics)")
for t in s["topics"]:
print(f" - {t}")
print()
print("Group-and-confirm checkpoint forcing options:")
print(" 1. Looks good — proceed with these sections")
print(" 2. Merge sections [X] and [Y]")
print(" 3. Split section [X] into two")
print(" 4. Add a section for [topic]")
print(" 5. Remove section [X]")
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
Xác định KPI, thiết lập dashboard, thiết kế thí nghiệm và diễn giải kết quả kiểm thử cho sản phẩm.
--- name: cs-product-analyst description: Product analytics agent for KPI definition, dashboard setup, experiment design, and test result interpretation. skills: - product-team/product-analytics - product-team/experiment-designer domain: product model: sonnet tools: [Read, Write, Bash, Grep, Glob] --- # Product Analyst Agent ## Skill Links - `../../product-team/product-analytics/SKILL.md` - `../../product-team/experiment-designer/SKILL.md` ## Primary Workflows 1. Metric framework and KPI definition 2. Dashboard design and cohort/retention analysis 3. Experiment design with hypothesis + sample sizing 4. Result interpretation and decision recommendations ## Tooling - `../../product-team/product-analytics/scripts/metrics_calculator.py` - `../../product-team/experiment-designer/scripts/sample_size_calculator.py` ## Usage Notes - Define decision metrics before analysis to avoid post-hoc bias. - Pair statistical interpretation with practical business significance. - Use guardrail metrics to prevent local optimization mistakes.
Theo dõi đối thủ có hệ thống, phục vụ định vị marketing, battlecard bán hàng và quyết định lộ trình sản phẩm.
---
name: "competitive-intel"
description: "Systematic competitor tracking that feeds CMO positioning, CRO battlecards, and CPO roadmap decisions. Use when analyzing competitors, building sales battlecards, tracking market moves, positioning against alternatives, or when user mentions competitive intelligence, competitive analysis, competitor research, battlecards, win/loss, or market positioning."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: c-level
domain: competitive-strategy
updated: 2026-03-05
frameworks: ci-playbook, battlecard-template
---
# Competitive Intelligence
Systematic competitor tracking. Not obsession — intelligence that drives real decisions.
## Keywords
competitive intelligence, competitor analysis, battlecard, win/loss analysis, competitive positioning, competitive tracking, market intelligence, competitor research, SWOT, competitive map, feature gap analysis, competitive strategy
## Quick Start
```
/ci:landscape — Map your competitive space (direct, indirect, future)
/ci:battlecard [name] — Build a sales battlecard for a specific competitor
/ci:winloss — Analyze recent wins and losses by reason
/ci:update [name] — Track what a competitor did recently
/ci:map — Build competitive positioning map
```
## Framework: 5-Layer Intelligence System
### Layer 1: Competitor Identification
**Direct competitors:** Same ICP, same problem, comparable solution, similar price point.
**Indirect competitors:** Same budget, different solution (including "do nothing" and "build in-house").
**Future competitors:** Well-funded startups in adjacent space; large incumbents with stated roadmap overlap.
**The 2x2 Threat Matrix:**
| | Same ICP | Different ICP |
|---|---|---|
| **Same problem** | Direct threat | Adjacent (watch) |
| **Different problem** | Displacement risk | Ignore for now |
Update this quarterly. Who's moved quadrants?
### Layer 2: Tracking Dimensions
Track these 8 dimensions per competitor:
| Dimension | Sources | Cadence |
|-----------|---------|---------|
| **Product moves** | Changelog, G2/Capterra reviews, Twitter/LinkedIn | Monthly |
| **Pricing changes** | Pricing page, sales call intel, customer feedback | Triggered |
| **Funding** | Crunchbase, TechCrunch, LinkedIn | Triggered |
| **Hiring signals** | LinkedIn job postings, Indeed | Monthly |
| **Partnerships** | Press releases, co-marketing | Triggered |
| **Customer wins** | Case studies, review sites, LinkedIn | Monthly |
| **Customer losses** | Win/loss interviews, churned accounts | Ongoing |
| **Messaging shifts** | Homepage, ads (Facebook/Google Ad Library) | Quarterly |
### Layer 3: Analysis Frameworks
**SWOT per Competitor:**
- Strengths: What do they do well? Where do they win?
- Weaknesses: Where do they lose? What do customers complain about?
- Opportunities: What could they do that would threaten you?
- Threats: What's their existential risk?
**Competitive Positioning Map (2 axis):**
Choose axes that matter for your buyers:
- Common: Price vs Feature Depth; Enterprise-ready vs SMB-ready; Easy to implement vs Configurable
- Pick axes that show YOUR differentiation clearly
**Feature Gap Analysis:**
| Feature | You | Competitor A | Competitor B | Gap status |
|---------|-----|-------------|-------------|------------|
| [Feature] | ✅ | ✅ | ❌ | Your advantage |
| [Feature] | ❌ | ✅ | ✅ | Gap — roadmap? |
| [Feature] | ✅ | ❌ | ❌ | Moat |
| [Feature] | ❌ | ❌ | ✅ | Competitor B only |
### Layer 4: Output Formats
**For Sales (CRO):** Battlecards — one page per competitor, designed for pre-call prep.
See `templates/battlecard-template.md`
**For Marketing (CMO):** Positioning update — message shifts, new differentiators, claims to stop or start making.
**For Product (CPO):** Feature gap summary — what customers ask for that we don't have, what competitors ship, what to reprioritize.
**For CEO/Board:** Monthly competitive summary — 1-page: who moved, what it means, recommended responses.
### Layer 5: Intelligence Cadence
**Monthly (scheduled):**
- Review all tier-1 competitors (direct threats, top 3)
- Update battlecards with new intel
- Publish 1-page summary to leadership
**Triggered (event-based):**
- Competitor raises funding → assess implications within 48 hours
- Competitor launches major feature → product + sales response within 1 week
- Competitor poaches key customer → win/loss interview within 2 weeks
- Competitor changes pricing → analyze and respond within 1 week
**Quarterly:**
- Full competitive landscape review
- Update positioning map
- Refresh ICP competitive threat assessment
- Add/remove companies from tracking list
---
## Win/Loss Analysis
This is the highest-signal competitive data you have. Most companies do it too rarely.
**When to interview:**
- Every lost deal >$50K ACV
- Every churn >6 months tenure
- Every competitive win (learn why — it may not be what you think)
**Who conducts it:**
- NOT the AE who worked the deal (too close, prospect won't be candid)
- Customer success, product team, or external researcher
**Question structure:**
1. "Walk me through your evaluation process"
2. "Who else were you considering?"
3. "What were the top 3 criteria in your decision?"
4. "Where did [our product] fall short?"
5. "What was the deciding factor?"
6. "What would have changed your decision?"
**Aggregate findings monthly:**
- Win reasons (rank by frequency)
- Loss reasons (rank by frequency)
- Competitor win rates (by competitor, by segment)
- Patterns over time
---
## The Balance: Intelligence Without Obsession
**Signs you're over-tracking competitors:**
- Roadmap decisions are primarily driven by "they just shipped X"
- Team morale drops when competitors fundraise
- You're shipping features you don't believe in to match their checklist
- Pricing discussions always start with "well, they charge X"
**Signs you're under-tracking:**
- Your AEs get blindsided on calls
- Prospects know more about competitors than your team does
- You missed a major product launch until customers told you
- Your positioning hasn't changed in 12+ months despite market moves
**The right posture:**
- Know competitors well enough to win against them
- Don't let them set your agenda
- Your roadmap is led by customer problems, informed by competitive gaps
---
## Distributing Intelligence
| Audience | Format | Cadence | Owner |
|----------|--------|---------|-------|
| AEs + SDRs | Updated battlecards in CRM | Monthly + triggered | CRO |
| Product | Feature gap analysis | Quarterly | CPO |
| Marketing | Positioning brief | Quarterly | CMO |
| Leadership | 1-page competitive summary | Monthly | CEO/COO |
| Board | Competitive landscape slide | Quarterly | CEO |
**One source of truth:** All competitive intel lives in one place (Notion, Confluence, Salesforce). Avoid Slack-only distribution — it disappears.
---
## Red Flags in Competitive Intelligence
| Signal | What it means |
|--------|---------------|
| Competitor's win rate >50% in your core segment | Fundamental positioning problem, not sales problem |
| Same objection from 5+ deals: "competitor has X" | Feature gap that's real, not just optics |
| Competitor hired 10 engineers in your domain | Major product investment incoming |
| Competitor raised >$20M and targets your ICP | 12-month runway for them to compete hard |
| Prospects evaluate you to justify competitor decision | You're the "check box" — fix perception or segment |
## Integration with C-Suite Roles
| Intelligence Type | Feeds To | Output Format |
|------------------|----------|---------------|
| Product moves | CPO | Roadmap input, feature gap analysis |
| Pricing changes | CRO, CFO | Pricing response recommendations |
| Funding rounds | CEO, CFO | Strategic positioning update |
| Hiring signals | CHRO, CTO | Talent market intelligence |
| Customer wins/losses | CRO, CMO | Battlecard updates, positioning shifts |
| Marketing campaigns | CMO | Counter-positioning, channel intelligence |
## References
- `references/ci-playbook.md` — OSINT sources, win/loss framework, positioning map construction
- `templates/battlecard-template.md` — sales battlecard template
FILE:references/ci-playbook.md
# Competitive Intelligence Playbook
## OSINT Sources for Competitor Tracking
### Free, Reliable Sources
**Company & Product:**
- **Their website** — pricing page (archive.org for history), product changelog, careers page
- **G2 / Capterra / Trustpilot** — customer reviews; filter by recency; read 1-star reviews carefully
- **LinkedIn** — job postings signal roadmap; company page for headcount trend; employees for leaks
- **GitHub** — open source activity; what they're building; engineering team size; tech stack
- **Crunchbase / PitchBook** (free tier) — funding history, investors, team changes
- **BuiltWith** — tech stack they use; signals about infrastructure maturity
**Messaging & Positioning:**
- **Facebook Ad Library** — see their current ad copy and creative; what messages they're testing
- **Google Keyword Planner** — which keywords they're bidding on
- **SEMrush / Ahrefs** (free trial or limited) — their organic keywords, backlink profile
- **Wayback Machine** — homepage evolution over time; when positioning shifted
- **Their blog** — content strategy reveals priorities and ICP assumptions
**News & Events:**
- **TechCrunch, VentureBeat** — funding announcements, major launches
- **Twitter/X / LinkedIn** — CEO + founders; direct signals about strategy
- **Podcast appearances** — founders talk more openly on podcasts than press releases
- **Job descriptions** — "Senior Engineer - Payments" means they're building payments
### Paid (Worth It for Tier-1 Competitors)
- **G2 Buyer Intent** — which prospects are researching your competitor right now
- **Bombora** — intent data for account-level research signals
- **PitchBook** — funding, investors, valuation estimates
- **Klue / Crayon / Kompyte** — dedicated CI platforms that aggregate automatically
### Primary Research (Best Signal)
- **Win/loss interviews** — the single highest-signal source (see below)
- **Talk to churned customers** — why did they switch? To whom?
- **Talk to their customers** — LinkedIn outreach; honest conversations
- **Industry events** — competitor presentations reveal roadmap; talk to attendees
- **Former employees** — LinkedIn; respectful outreach; no NDA violations
---
## Competitive Battlecard Format
A battlecard is a 1-page (or single screen) document for sales reps to reference before and during calls.
**Design principles:**
- Written for a rep with 2 minutes to prep, not a product manager
- Action-oriented: tells reps what to SAY, not just what to know
- Updated monthly at minimum; never more than 90 days old
### Battlecard Structure
```
COMPETITOR: [Name]
Last updated: [Date] | Owner: [Name]
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
THE 30-SECOND SUMMARY
[One paragraph. Who they are, who they sell to, why they win.]
THEIR STRENGTHS (know these — don't dismiss them)
• [Strength 1] — what customers actually love about them
• [Strength 2]
• [Strength 3]
THEIR REAL WEAKNESSES (from win/loss data, not assumptions)
• [Weakness 1] — source: [customer quote / win/loss theme]
• [Weakness 2]
• [Weakness 3]
OUR DIFFERENTIATED ADVANTAGES
• [Advantage 1] — proof point: [metric/customer/case study]
• [Advantage 2] — proof point:
• [Advantage 3] — proof point:
COMMON OBJECTIONS + RESPONSES
"They have [feature] and you don't."
→ [Response. Acknowledge, reframe, redirect.]
"They're cheaper."
→ [Response with ROI angle or TCO comparison.]
"They're more established / bigger."
→ [Response. Size isn't always advantage; use to your benefit.]
TRAP-SETTING QUESTIONS (ask these early to shift the eval criteria)
• "How important is [your differentiator] to your team?"
• "Have you looked at [pain point they create]?"
• "What happens to your workflow when [their known limitation occurs]?"
WHEN WE WIN
• [Segment or scenario where we almost always beat them]
• [Use case where we're clearly stronger]
WHEN WE LOSE (be honest)
• [Scenario where they're genuinely better — don't fight these battles]
• [Segment where they have structural advantages]
DO NOT SAY
• Don't claim [X] — it's not true and they'll call it out
• Don't say [Y] — prospect will already know it and it sounds desperate
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
```
---
## Win/Loss Analysis Framework
### Why Most Companies Do This Wrong
- They survey instead of interview (surveys get polite answers)
- The AE conducts it (too emotionally invested; prospect won't be candid)
- They do it 6 months after the decision (memory fades)
- They look for confirmation of what they believe
### The Right Process
**Timing:** Within 30 days of deal closed/lost/churned.
**Interviewer:** Customer success, product, or external researcher. Never the AE.
**Duration:** 30 minutes (budget 45).
**Incentive:** $100 gift card gets you 80% acceptance. Worth it.
**Interview Guide:**
*Opening:*
"I'm [name] from [company]. I'm not in sales — I'm trying to understand what drove your decision so we can improve. There's nothing you can say that will change the outcome. I just want honest feedback."
*Core questions:*
1. "Can you walk me through your evaluation process from the beginning?"
2. "Who were the key stakeholders involved in the decision?"
3. "What were the 3 most important criteria you were evaluating against?"
4. "Which vendors did you seriously consider?"
5. "Where did [company] fall short of your expectations?" (For losses)
OR "What tipped the decision in [company]'s favor?" (For wins)
6. "Was price a factor? How significant?"
7. "What would have had to be different for you to choose [us / the other option]?"
8. "Any advice for our team on how we handled the process?"
**Data aggregation:**
- Tag every response: [criterion], [competitor mentioned], [product gap], [sales process], [price], [trust/credibility]
- Monthly rollup: top 5 win reasons, top 5 loss reasons, competitor win rate
- Share with: CEO, CRO, CPO, CMO — not just sales
---
## Competitive Positioning Map Construction
A positioning map shows where you sit relative to competitors on 2 dimensions that BUYERS care about.
### Step 1: Choose Your Axes
- Pick dimensions that actually drive purchase decisions in your segment
- At least one axis should be where you win
- Avoid generic axes ("feature-rich vs. simple" tells you nothing)
**Good axis pairs:**
- Implementation time (days vs. months) × Customization depth
- Price point × Enterprise readiness
- Automation level × Human-in-the-loop control
- Time-to-value × Total cost of ownership
**Bad axes:**
- Quality (too vague)
- "Innovation" (unmeasurable)
- Any axis where all competitors cluster in the same spot
### Step 2: Place Competitors Objectively
- Use customer quotes and win/loss data to justify placement
- Don't place competitors where you WANT them — where they ACTUALLY are
- If you're unsure, ask 5 customers to place them
### Step 3: Find and Name Your White Space
- Where is there a position no competitor holds?
- Is that white space there because it's valuable (opportunity) or worthless (avoid)?
- Can you credibly occupy it?
### Step 4: Test Your Positioning
- Show the map to 5 prospects: "Does this match your perception?"
- Show it to 5 lost prospects: "Where would you place [the winner] and us?"
- Adjust until map matches buyer reality, not internal perception
---
## Intelligence Sharing Across Roles
### What Each Role Needs and When
**CRO (Sales):**
- Needs: Battlecards, win rates by competitor, competitor objections + responses
- Cadence: Updated battlecards monthly; triggered updates on major competitor moves
- Format: 1-pager per competitor in CRM, linked from deal record
**CMO (Marketing):**
- Needs: Messaging shifts, new claims, ad spend signals, keyword battles
- Cadence: Quarterly positioning review, triggered on major launches
- Format: Positioning brief with recommended response to messaging shifts
**CPO (Product):**
- Needs: Feature gap analysis, competitor roadmap signals (job postings, changelog), what we lose to
- Cadence: Monthly feature gap update, triggered on major launches
- Format: Feature comparison matrix + gap prioritization recommendation
**CTO (Engineering):**
- Needs: Tech stack signals, infrastructure approaches, scale they've achieved
- Cadence: Quarterly
- Format: Technical comparison notes, relevant for architectural decisions
**CEO:**
- Needs: Summary of threat landscape, recommended responses, board-level narrative
- Cadence: Monthly 1-pager + quarterly deep dive
- Format: 1-page brief: who moved, what it means, what we do
### The Single Source of Truth Rule
All competitive intel in one place. Suggest:
- Notion database per competitor: profile, battlecard, changelog, win/loss notes
- Slack channel: `#competitive-intel` for real-time triggered alerts
- Monthly digest email to leadership
If it lives only in Slack, it disappears. If it lives only in a wiki that nobody reads, it doesn't matter. Combine both.
---
## How to Track Without Obsessing
**Set up the system, then let it run:**
- Google Alerts for competitor names + CEO names
- LinkedIn Saved Searches for their job postings
- Klue/Crayon if budget allows (automated aggregation)
- Monthly 60-minute competitive review meeting (not 4 hours)
**What to do when competitor makes a big move:**
1. Read the announcement objectively
2. Talk to 3 customers: "Did you see this? What do you think?"
3. Assess: does this change any buying criteria in your deals?
4. If yes: update battlecard and positioning within 1 week
5. If no: log it, move on
**The test:** After reviewing a competitor move, do you feel urgency to ship something? If yes, you're reacting. The right feeling is "noted — let's see if customers care."
FILE:templates/battlecard-template.md
# Sales Battlecard Template
**COMPETITOR:** [Name]
**Last updated:** [YYYY-MM-DD] | **Owner:** [Name]
**Win rate vs this competitor:** [X]% | **Deals tracked:** [N]
---
## 30-Second Summary
[Who they are. Who they target. Why they win. What they're known for. 3-4 sentences max.]
---
## Their Strengths
*Know these. Don't dismiss them. Prospects have already heard their pitch.*
- **[Strength]:** [What customers genuinely love; source if available]
- **[Strength]:** [Specific capability or trait]
- **[Strength]:** [Brand, market position, or ecosystem advantage]
---
## Their Real Weaknesses
*From win/loss data only — not wishful thinking.*
- **[Weakness]:** "[Customer quote]" — seen in [N] deals
- **[Weakness]:** [Documented limitation with evidence]
- **[Weakness]:** [Implementation, support, or pricing issue]
---
## Our Differentiated Advantages
*Must be real and provable. Each needs a proof point.*
- **[Advantage]:** [Proof: metric / customer quote / case study]
- **[Advantage]:** [Proof]
- **[Advantage]:** [Proof]
---
## Common Objections + Responses
**"They have [feature X] and you don't."**
> [Acknowledge. Reframe to your strength. Redirect to outcome.
> "You're right that they have X. What we've found is that customers who care most about X tend to also care about [Y], where we're significantly stronger. Can I show you [specific example]?"]
**"They're cheaper."**
> [Don't fight on price. Reframe to TCO or ROI.
> "They are lower in initial cost. Most customers find the total cost over 12 months is actually comparable when you factor in [implementation time / support costs / integrations]. Want to walk through that?"]
**"They've been around longer / they're more established."**
> [Reframe tenure as potential liability or irrelevance.
> "Their longevity means they have a lot of technical debt and a big customer base that pulls their roadmap in every direction. Our customers tell us that's exactly why they chose us — we move faster and we're laser-focused on [their specific use case]."]
**"[Competitor] is already used by [big customer they respect]."**
> [Name-drop your wins in their segment.
> "We work with [comparable logo]. Want me to connect you with their [role] to ask how they made the decision?"]
---
## Trap-Setting Questions
*Ask early in discovery to establish criteria that favor you.*
- "How important is [your key differentiator] to your workflow?"
- "What happens when [their known limitation] occurs? Has that been an issue before?"
- "How long does your team typically take to onboard a new tool?"
- "Who manages the integration work — do you have dedicated engineering resources for that?"
- "What does your current vendor do when you need support?"
---
## When We Win
- [Scenario or segment where we consistently beat them]
- [Use case that plays to our strengths]
- [Buyer profile that prefers us]
## When We Lose (Be Honest)
- [Scenario where they genuinely win — don't fight here]
- [Segment where their strengths matter more than ours]
---
## Do NOT Say
- ❌ Don't claim [X] — it's not accurate and they'll check
- ❌ Don't attack [Y] — it backfires and makes us look insecure
- ❌ Don't say "we're better" without specifics — be concrete
---
## Recent Intel
*Last 90 days only. Older than 90 days: archive.*
- [Date]: [What happened — funding, product launch, pricing change, key hire]
- [Date]: [Customer feedback from win/loss interview]
- [Date]: [Any notable market move]
---
*Battlecards are only useful if current. If this is >90 days old, flag to [owner] for update.*
Điểm vào mặc định cho mọi yêu cầu nghiên cứu: phân loại câu hỏi rồi chuyển cho skill chuyên biệt như xu hướng, tài trợ NIH, tài liệu học thuật, sáng chế.
---
name: research
description: Default entry point for any research request — a hybrid router that classifies the question deterministically and either delegates to a specialist research skill (pulse for trends/sentiment, grants for NIH funding, litreview for academic literature, syllabus for course reading, patent for prior-art + IP landscape, dossier for entity research) or runs its own plan-decompose-multi-source-search-synthesize-cite fallback workflow when no specialist matches. Always surfaces the routing decision so users can override. Triggers — "research [topic]", "look into [topic]", "what do we know about [topic]", "investigate [topic]", "find me information on [topic]", "do some research on [topic]", "I need to understand [topic]", or any research request that doesn't obviously match a more-specific specialist skill. Output is a markdown briefing (default) or .docx document (on request) with full citations and an audit log.
---
# Research — Hybrid Router + Fallback
**The runtime orchestrator for the research domain.** Architecture C: deterministic classification → specialist delegation OR own plan-decompose-search-synthesize-cite workflow.
## Portability
Requires `WebSearch` + `WebFetch` for the fallback workflow; specialist skills (`pulse`, `grants`, `litreview`, `syllabus`, `patent`, `dossier`) must be present for delegation to work. Node.js with `docx` package required if Q2 = document mode. Works in Claude Code CLI natively. In Claude.ai with web tools + Code Execution, the workflow is supported.
## Distinct From `engineering/autoresearch-agent`
These two skills share the word "research" but serve **completely different use cases**:
- **`research/research/`** (this skill) — research-query router + fallback workflow ("Research X")
- **`engineering/autoresearch-agent/`** — Karpathy's autonomous file-optimization experiment loop ("Make this code faster")
No overlap. They coexist.
## Hybrid Architecture (C)
Every invocation produces one of three outcomes:
1. **Delegation** — Classified as specialist-domain. Routes there. User sees the specialist's output.
2. **Fallback execution** — Classified as general research. Runs own plan → search → synthesize workflow.
3. **Clarification request** — Classification ambiguous. Asks one forcing question to disambiguate, then routes.
The skill **never silently runs its fallback** when a specialist would have done better. **Routing transparency** is what makes the hybrid architecture trustworthy.
## Specialist Registry
| Specialist | Routing signals | Domain |
|---|---|---|
| `pulse` | reddit / hn / x / buzz / sentiment / trending / "what's people saying" / "pulse on" / "take the pulse" / "current conversation" | Multi-source recency research |
| `grants` | NIH / grant / R01 / K-award / RePORTER / NOSI / "grants for" / FDA / "study section" / "principal investigator" | NIH grant-funding intelligence |
| `litreview` | literature review / PICO / SPIDER / systematic review / "review papers on" / meta-analysis | Academic literature orientation |
| `syllabus` | syllabus / course outline / curriculum / "reading list" / "for my class" / "for my students" | Course supplementary reading |
| `patent` | prior art / FTO / freedom to operate / patent / "patent landscape" / invention / novelty search / "ip landscape" | Patent prior-art + landscape |
| `dossier` | "dossier on" / "due diligence" / "background check" / "prep me for" / "competitor research" / "investor diligence" / "interview prep" / "background on" | Decision-grade entity research |
## Agent Integrity Rules
This skill obeys the research-pack convention:
- **Execution discipline (fallback only)**: Sequential searches. 1 q/sec rate limit. Confirm response received before next call.
- **Source discipline**: Cite only sources returned by this session's tool calls. Training knowledge labeled `[Background — not from search]` and excluded from counts.
- **Three-count tracking (fallback only)**: Queries sent / sources received / sources cited.
- **Retry policy**: On failure → wait 3s → retry once → log. After 3 consecutive failures: stop, alert user.
- **Plan-tier detection**: If delegated to Consensus-using specialist, that specialist handles detection. In fallback mode, surface any rate-limit signals.
- **Routing discipline**: Never delegate silently. Always state the decision + accept override.
## Phase 1: Grill-Me Intake (2–4 Questions)
Intake is intentionally minimal — the goal is to route fast, not to interrogate. One question per turn.
### Q1 (always) — Research question
> **What's the research question? State it in 1–2 sentences. Specific is better than broad — "AI for healthcare" gets you a vague survey; "How are health systems integrating LLM-based clinical decision support in 2026?" gets you a useful answer.**
>
> *Why I'm asking:* Specificity dictates classification accuracy and search precision. A vague question routes to fallback; a specific question often matches a specialist cleanly.
**Refuse mush.** If user says "research AI", push back once: "What about AI specifically — adoption, safety, capability, funding, regulation, comparison? Pick an angle."
### Q2 (always) — Output preference
> **What output do you want? Pick one:**
> 1. Quick chat briefing (5-min read, markdown in chat)
> 2. Standalone document (.docx with citations, shareable)
>
> *Why I'm asking:* Document mode triggers deeper search budgets and full audit logs. Chat mode optimizes for fast delivery.
Forcing choice.
### Q3 (asked only if classification ambiguous — ≤1 signal) — Domain disambiguation
> **Quick clarification — pick the closest match:**
> 1. Academic literature (papers, peer-reviewed)
> 2. Industry / trends (what's the buzz, news, sentiment)
> 3. Specific entity (a company, person, organization)
> 4. Technology / patents (prior art, IP landscape)
> 5. Grant funding (NIH, foundations)
> 6. Course material (syllabus or curriculum)
> 7. None of the above — run general research
>
> *Why I'm asking:* I couldn't classify confidently from your question alone. This routes you to the right specialist or confirms general-research fallback.
**Skip if Q1 + Q2 produced clear specialist match (≥2 signals).**
### Q4 (asked only if Q3 was needed AND user picked "none of the above") — General-research scope
> **For general research, what's your time horizon — quick scan (5 searches) or thorough (15 searches)?**
>
> *Why I'm asking:* General research has no specialist budget; you pick it. Quick is good for "what's the lay of the land". Thorough is for "I'll make a decision based on this".
Skip if a specialist took over.
**Stop condition:** After Q4 (or earlier if dependency skips applied), commit and start Phase 2. **Most invocations exit intake after Q1 + Q2.**
## Phase 2: Deterministic Classification
This is **deterministic, not LLM-reasoned** — for speed, debuggability, and consistency.
```python
SIGNALS = {
pulse: ["reddit", "hn", "hacker news", "x.com", "twitter", "buzz",
"sentiment", "trending", "what are people saying",
"what's happening", "the conversation around",
"pulse on", "take the pulse", "current conversation"],
grants: ["nih", "grant", "grants for", "r01", "r21", "k-award", "reporter",
"nosi", "funding", "fda", "study section", "principal investigator"],
litreview:["literature review", "lit review", "litreview", "pico", "spider",
"systematic review", "review papers on", "research papers on",
"papers about", "meta-analysis"],
syllabus: ["syllabus", "course outline", "curriculum", "reading list",
"for my class", "for my students", "course material"],
patent: ["prior art", "fto", "freedom to operate", "patent",
"patent landscape", "invention", "novelty search",
"patent search", "ip landscape"],
dossier: ["dossier on", "due diligence", "background check",
"prep me for", "competitor research", "investor diligence",
"interview prep", "research my competitor", "background on"]
}
# Signals are case-insensitive literal phrases (multi-word substring match).
# Bracketed placeholders (e.g., "research [company]") are intentionally NOT
# signals — they over-trigger on generic "research X" queries that should
# fall back to general research, not auto-route to dossier. Specific phrases
# pair the verb with the noun ("dossier on", "background on") and route reliably.
For each specialist S:
score[S] = count of SIGNALS[S] phrases matched in question (case-insensitive substring)
if max(score) >= 2:
route_to = argmax(score) # high confidence
elif max(score) == 1 and only one specialist has score 1:
route_to = that specialist # weak match, single specialist
else:
route_to = "fallback" # ambiguous or no match — ask Q3
```
**Implementation:** `scripts/classifier.py --question "..."` returns the routing decision + matched signals + per-specialist scores. Use it; don't re-implement.
## Phase 3a: Specialist Delegation (≥2 signals OR single weak match)
When delegating:
1. Pass the user's question **verbatim** plus the output preference (Q2)
2. **Let the specialist run its own grill-me intake** — do NOT pre-answer specialist questions
3. Return specialist output as the user-visible result
4. Tag the result with `[Delegated to: research → {specialist}]` in the chat output so the user knows what skill produced it
5. Tag the audit log via `scripts/routing_transparency_logger.py --action record_delegation`
## Phase 3b: Own Fallback Workflow
If routing produced no specialist match, run the 8-step fallback.
### Step 1: Decompose
Break the research question into 3–5 sub-questions. Use the framework: what / why / how / who / what's next. Show the decomposition to the user before searching. Use `scripts/fallback_decomposer.py --question "..."` for a deterministic starting point.
### Step 2: Source Selection
For each sub-question, choose source(s) deterministically:
- **Recency-sensitive** → WebSearch + WebFetch + (optionally Reddit/HN if signal)
- **Technical specs / docs** → WebSearch + WebFetch
- **Academic** → Consensus MCP if connected; otherwise WebSearch with `scholar.google.com` site filter
- **Data / numbers** → WebSearch for sources; then WebFetch for primary documents
- **Person / company entity-level** → consider routing to `dossier` (offer override)
### Step 3: Search
Sequential per sub-question. 1 q/sec etiquette. Per source: 2–4 queries, broad-to-narrow.
### Step 4: Read + Extract
For each result that looks high-signal: WebFetch and extract the relevant section. Note the source URL.
### Step 5: Synthesize
Per sub-question: 2–4 paragraphs answering it with inline citations. Surface disagreement when sources disagree.
### Step 6: Cross-Cutting Patterns
After per-sub-question synthesis: 1–2 paragraphs of patterns across sub-questions — consensus, controversy, gaps.
### Step 7: Output
Markdown brief by default (Q2 choice). DOCX if user picked document mode.
### Step 8: Audit Log
Three-count summary (sent / received / cited) + per-source list with reliability tier (primary / secondary / tertiary).
## Routing Transparency Protocol (Mandatory)
After classification, the skill **always**:
1. **States the decision** in one sentence: "Routing to `litreview` because you mentioned PICO and meta-analysis (2 signals)."
2. **Offers override**: "If you want general research instead OR a different specialist, say so now. Otherwise proceeding in 5 seconds."
3. **Waits 1 turn** for confirmation (or auto-proceeds after 5s in interactive contexts).
4. **If user overrides** → accept, re-route, log the override via `routing_transparency_logger.py --action record_override`.
**Never delegates silently.** This is the trust-building property that makes the hybrid pattern work.
## Output Format
### Markdown brief (Q2 = quick chat briefing)
```markdown
# [Research Question] — Briefing
*Generated: [DATE] | Routed: [delegated specialist | fallback]*
## TL;DR
[2-3 sentences]
## Findings
### [Sub-question 1]
[2-4 paragraphs with inline citations]
### [Sub-question 2]
...
## Cross-Cutting Patterns
[1-2 paragraphs]
## Sources
[Numbered list with hyperlinks, reliability tier per source]
## Audit
[Three counts + per-source tier + failures]
```
### DOCX (Q2 = standalone document)
Use the standard research-pack DOCX patterns: Arial 12pt, navy headings, blue table headers, hyperlinked sources, mandatory audit log section. Reference the `docx` skill for setup.
## Audit Log Requirement (Fallback Mode)
```
Queries sent: N
Sources received: M
Sources cited: K
Failures: F (3-consecutive-failures triggered: yes/no)
Per-source tier: [URL — primary | secondary | tertiary]
Routing decision: fallback (no specialist matched)
Sub-questions: [list]
```
All routing decisions + overrides also logged to `~/.research_sessions/<session>.json` via `routing_transparency_logger.py`.
## Failure Modes
| Failure | Behavior |
|---|---|
| Classification ambiguous (≤1 signal) | Ask Q3 (domain disambiguation). |
| Specialist delegation fails | Note in chat. Offer to retry or fall back to general research. |
| User overrides routing | Accept. Re-route to chosen specialist or fallback. Log the override. |
| Fallback search returns thin results | Surface explicitly. Suggest the question may be too niche or too new. Do not fabricate. |
| 3 consecutive tool failures in fallback | Stop, alert user, share what was collected. |
| Question is non-research (e.g., "write me code") | Decline politely. Suggest the user invoke an appropriate skill. |
| Sub-question can't be answered | Note in synthesis as "limited public signal on this"; don't omit silently. |
| Output format mismatch | Honor Q2 preference; if format unavailable, fall back to markdown with note. |
| Specialist skill missing from environment | Skip it in classification scoring; route to fallback or next-best specialist. |
## Anti-Patterns Rejected
- LLM-reasoned classification (must be deterministic keyword + intent matching)
- Silent delegation (always surface routing decision)
- Refusing to route to a specialist when ≥2 signals match
- Routing to a specialist when classification is genuinely ambiguous (≤1 signal across all)
- Pre-answering the specialist's grill-me intake (let it run its own)
- Running fallback when a specialist would clearly do better
- Fabricating sources in fallback when search is thin
- Skipping audit log in fallback mode
- Treating "dossier on [company]" as fallback when `dossier` is the right specialist (the verb-noun-paired phrase, not the generic "research X" form, is what routes)
- Treating "what are people saying about X" as fallback when `pulse` is the right specialist
- Auto-routing generic "research [topic]" queries to a specialist when the user hasn't paired the verb with a specialist-specific noun (e.g., "research Microsoft" alone is ambiguous — could be dossier or general; ask Q3 instead of guessing)
## Tooling
### Python (stdlib only)
- **`scripts/classifier.py`** — Deterministic SIGNALS matching → routing decision + per-specialist score + matched phrases. `--question "..." --output json`.
- **`scripts/routing_transparency_logger.py`** — JSON-backed audit log at `~/.research_sessions/<session>.json`. Records every routing decision, override, and delegation handoff.
- **`scripts/fallback_decomposer.py`** — Heuristic question → 3–5 sub-questions using what / why / how / who / what's next framework.
### Reference Docs (each cites 7+ authoritative sources)
- **`references/hybrid_router_architecture.md`** — router-vs-run trade-offs + routing transparency principle
- **`references/deterministic_classification_canon.md`** — why keyword > LLM-reasoned for routing
- **`references/fallback_workflow_canon.md`** — plan-decompose-search-synthesize methodology
## Dependencies
- **`WebSearch`** + **`WebFetch`** — Required for fallback workflow
- **Specialist skills** — Required for delegation: `pulse`, `grants`, `litreview`, `syllabus`, `patent`, `dossier`. If a specialist is missing, the router skips it in classification and routes to fallback instead.
- **Node.js `docx` library** — Required if user picks document output (Q2 = standalone)
- **Consensus MCP** — Optional; used in fallback if academic sub-questions surface
## Trigger Phrases
- "research [topic]"
- "look into [topic]"
- "what do we know about [topic]"
- "investigate [topic]"
- "find me information on [topic]"
- "do some research on [topic]"
- "I need to understand [topic]"
- Any research request that doesn't obviously match a more-specific specialist
---
**Version:** 1.0.0
**Source spec:** [`megaprompts/13-research-megaprompt.md`](../../../../megaprompts/13-research-megaprompt.md)
**Build pattern:** Path B (direct conversion)
FILE:references/deterministic_classification_canon.md
# Deterministic Classification — Why Keyword Beats LLM-Reasoned For Routing
This reference answers one decision: **should the routing classifier use deterministic keyword matching or LLM reasoning over the query?** The answer is **deterministic keyword matching** for query-routing purposes, with LLM reasoning reserved for cases where keyword matching has genuinely exhausted the signal space.
## The Trade-Off Spectrum
| Approach | Latency | Cost | Determinism | Debuggability | Coverage of fuzzy intent |
|---|---|---|---|---|---|
| **Keyword + intent signals** (this skill) | <1ms | $0 | 100% | High (signals named explicitly) | Low |
| **Embedding similarity to specialist descriptions** | ~10-100ms | Cents/100K queries | High (deterministic given embeddings) | Medium (need to inspect cosine scores) | Medium |
| **LLM reasoning over query + specialist list** | ~500ms-2s | ~$0.001-0.01/query | Low (same query → varied outputs) | Low (prompt-dependent) | High |
The trade-off: as you move down the table, coverage of fuzzy intent improves, but latency, cost, and unpredictability all worsen. The right choice depends on how predictable + auditable the routing needs to be.
## For Query Routing, Determinism Wins
Routing is **fundamentally a control-flow decision**: it determines which subsystem runs next. Like any control-flow decision in software, predictability + auditability are first-order properties.
Compare to other deterministic control-flow systems:
- **Compilers** use deterministic lexer + parser, not LLMs.
- **Routers** (network sense) use deterministic CIDR matching, not LLMs.
- **CI/CD systems** use deterministic file-pattern triggers, not LLMs.
- **Linters + formatters** use deterministic AST-walking, not LLMs.
These are all systems where users need to predict + debug behavior. LLM-reasoned routing in any of them would be a regression. Same applies to skill routing.
## The Bracketed-Placeholder Anti-Pattern
A common mistake when building keyword classifiers: using bracketed placeholders as signals.
**Wrong:**
```python
SIGNALS = {
dossier: ["dossier on [company]", "background check on [person]", "research [entity]"]
}
```
**Why wrong:** the "research [entity]" pattern collapses to "research" as a substring match, which matches every research request ever. The signal over-triggers + breaks the classifier.
**Right:**
```python
SIGNALS = {
dossier: ["dossier on", "background check", "background on", "competitor research"]
}
```
**Why right:** verb-noun pairs ("dossier on", "background on", "competitor research") are specific to dossier intent. Generic "research X" stays in fallback territory until paired with a specialist-specific noun.
This is the post-PR-#657-audit lesson encoded as a hard rule.
## What Counts As A "Signal"
A signal is a **case-insensitive literal phrase (multi-word substring)** that, when present in the user's question, indicates a specialist domain. Good signals are:
- **Specific enough** that they don't appear in unrelated queries (good: "literature review", bad: "research")
- **Common enough** that users actually say them (good: "due diligence", bad: "actuarial diligence assessment framework")
- **Diverse enough** to cover surface variations (good: "lit review" + "literature review" + "litreview"; bad: only one form)
- **Verb-noun-paired** when the noun alone is ambiguous (good: "dossier on" + "background on"; bad: just "company name")
## Confidence Thresholds
The skill commits to a specialist at **≥2 signals** for two reasons:
1. **2 signals reliably indicate intent.** "PICO + meta-analysis" doesn't show up in unrelated queries.
2. **1 signal isn't strong enough.** "PICO" alone might be a clinical question, a syllabus question, or a litreview question. The second signal distinguishes.
The single-weak-match exception (1 signal + only one specialist with any score) handles the case where the user used a highly specific phrase that no other specialist's signals overlap with. "What's the FTO landscape" → only patent has any score → route to patent even though it's just 1 signal.
The "ask Q3 disambiguation" exception handles the case where multiple specialists each have score 1, OR no specialist has any score. Both indicate genuine ambiguity that the classifier can't resolve.
## What Goes Wrong With LLM-Reasoned Classification
### Non-determinism
Same query, different responses across invocations. User says "what are people saying about X" — sometimes routes to pulse, sometimes to dossier, sometimes to fallback. User can't develop intuition for the system.
### Cost
500ms-2s per classification × hundreds of routing decisions/day adds up. Deterministic classifier is sub-millisecond + free.
### Debuggability
When LLM routes "weirdly," there's no signal to inspect. With deterministic classification, the user sees "matched signals: PICO, meta-analysis" and understands why.
### Prompt drift
LLM classifier behavior changes when the underlying model version changes. Deterministic classifier behavior is locked to the signals list. Auditable + reproducible.
## What Goes Wrong With Pure Keyword Classification
### Fuzzy intent
User says "I want to understand what the academic community thinks about CRISPR safety." No keyword matches litreview signals (no "PICO", no "systematic review", no "literature review"). Classifier punts to fallback even though litreview was the right answer.
**Mitigation:** Q3 disambiguation handles this. User picks "academic literature" → routes to litreview. The architecture's clarification path covers the fuzzy-intent case.
### Surface-form proliferation
Users say "lit review", "literature review", "litreview", "review the literature on", "review papers on", "look at the papers about", "what does the research say about" — that's 7 surface forms for the same intent. Signals list grows.
**Mitigation:** Cover the top-N surface forms (3-5 per specialist). Let Q3 handle the long tail.
### Polysemy
"Patent" could mean a legal patent (route to patent specialist) OR a medical term ("the symptoms are patent" = obvious). Keyword matching can't distinguish.
**Mitigation:** Multi-signal requirement reduces false positives. "Patent + prior art" is unambiguously patent intent.
## The Right Hybrid: Deterministic First, Clarify When Stuck
The architecture combines:
1. **Deterministic classification** for the high-confidence path (cheap + fast + predictable)
2. **Q3 disambiguation** for the genuinely-ambiguous path (LLM-free; user picks from 7 options)
3. **Fallback workflow** for the no-specialist path
This is strictly better than pure-LLM classification (cheaper, faster, more predictable) and strictly better than pure-keyword classification (handles fuzzy intent via Q3).
## Operational Discipline
When adding a new signal to the SIGNALS map:
- [ ] Verify the signal doesn't appear in queries that should route elsewhere (false positive check)
- [ ] Verify the signal does appear in queries that should route to this specialist (false negative check)
- [ ] Check for case-insensitivity (the matcher is case-insensitive, but be explicit)
- [ ] Avoid bracketed placeholders
- [ ] Use verb-noun pairs when the noun alone is ambiguous
- [ ] Document why this signal was added (which queries it covers)
When removing a signal:
- [ ] Check what queries previously routed via this signal
- [ ] Confirm they still route correctly (via another signal OR via Q3)
- [ ] Update the documentation
## Tooling
`scripts/classifier.py` implements the deterministic SIGNALS-matching algorithm. Use it; don't re-implement. It returns:
- `route_to`: specialist name OR "fallback"
- `confidence`: "high (N signals)" OR "weak (1 signal, single specialist)" OR "ambiguous"
- `matched_signals`: dict of specialist → list of matched phrases
- `scores`: dict of specialist → integer score
The CLI: `classifier.py --question "..." --output json`.
## Citations (7 sources)
1. **Aho, Sethi, Ullman — "Compilers: Principles, Techniques, and Tools" (Dragon Book, 1986).** Source for the deterministic lexer + parser as the canonical control-flow classifier in software. Compilers don't use LLMs for tokenization; routing shouldn't either.
2. **Cisco IOS — Access Control List (ACL) implementation guides.** Source for the deterministic CIDR-matching pattern in network routing. Predictability + auditability are first-order requirements; same applies to skill routing.
3. **Google Search Engineering blog — Query Classification (2020+).** Source for the production-grade query-classification pattern. Google uses deterministic signal matching as the first layer + LLM reasoning only for residual queries that signals miss. Same architecture as this skill (Q3 as the LLM-equivalent escape hatch).
4. **Mikolov et al. — "Distributed Representations of Words and Phrases" (Word2Vec, 2013).** Source for the embedding-similarity baseline. Embeddings are an intermediate point between keywords + LLM reasoning; this skill chooses keywords for cost + determinism reasons but acknowledges embedding-similarity as a valid alternative.
5. **Karpathy, Andrej — "Software 2.0" (blog post, 2017).** Source for the framing that not everything should be ML. Deterministic systems (compilers, routers, type checkers) remain superior for control-flow decisions even in the LLM era. https://karpathy.github.io/2017/11/11/software-2-0/
6. **Anthropic — Tool Use + Function Calling documentation.** Source for the production pattern of LLM-routes-to-deterministic-tool: the LLM decides intent at the top level, then deterministic tools handle the actual work. Same shape as this skill (intake → deterministic classifier → specialist tool). https://docs.anthropic.com/
7. **NIST — "Information Retrieval Evaluation" (TREC reports).** Source for the canonical evaluation methodology for classifiers: precision + recall measured against held-out queries. Keyword classifiers reliably outperform LLM-reasoned classifiers on precision for domain-specific routing tasks. https://trec.nist.gov/
FILE:references/fallback_workflow_canon.md
# Fallback Workflow Canon — Plan / Decompose / Search / Synthesize / Cite
This reference answers one decision: **when no specialist matches, what workflow does the orchestrator run instead?** The answer is an **8-step plan-decompose-multi-source-search-synthesize-cite** workflow grounded in the canonical research-pack conventions.
## The Eight Steps
The fallback workflow is documented in `SKILL.md`. This reference explains the **why** behind each step + the failure modes per step + the tooling that supports it.
### Step 1: Decompose
Break the research question into 3–5 sub-questions. Use the framework: **what / why / how / who / what's next**.
**Why decompose?** A 1-sentence research question rarely has a 1-source answer. Decomposition forces the orchestrator to enumerate the actual claim shape before searching, which makes search precise + makes synthesis structured.
**Failure mode:** decomposing into too many sub-questions (>5) wastes search budget on diminishing returns. Cap at 5.
**Tooling:** `scripts/fallback_decomposer.py` returns a deterministic starting point. Override + refine before searching.
### Step 2: Source Selection
For each sub-question, pick the right source class. Use the deterministic mapping in SKILL.md:
- Recency-sensitive → WebSearch + WebFetch (+ optional Reddit/HN signal)
- Technical specs → WebSearch + WebFetch
- Academic → Consensus MCP if available; else WebSearch + scholar.google.com filter
- Data / numbers → WebSearch for primary documents
- Entity-level → consider routing back to `dossier`
**Failure mode:** using a wrong-class source (e.g., WebSearch for academic when Consensus would have produced higher-quality results). The mapping is deterministic for a reason.
### Step 3: Search
Sequential per sub-question. **1 q/sec rate limit** (research-pack convention). Per source: 2–4 queries, broad-to-narrow.
**Why broad-to-narrow?** Broad queries map the landscape; narrow queries find the high-signal sources within it. Going narrow-only often misses the orienting overview.
**Failure mode:** parallel search bursts that trigger rate-limiting or get blocked. Sequential is the discipline.
### Step 4: Read + Extract
For each high-signal result: WebFetch the full content + extract the relevant section + note the URL.
**Why extract, not summarize?** Direct quotes + section references make citations verifiable. Summaries hide the source structure.
**Failure mode:** synthesizing from search snippets without WebFetch. Snippets are not sources.
### Step 5: Synthesize Per Sub-Question
For each sub-question: 2–4 paragraphs with inline citations. Surface disagreement when sources disagree.
**Why per-sub-question?** Sub-question structure carries through to the output. Reader can navigate to the part they care about.
**Failure mode:** synthesizing across sub-questions in one mega-paragraph. Loses the navigability + makes disagreements harder to surface.
### Step 6: Cross-Cutting Patterns
After per-sub-question synthesis: 1–2 paragraphs of patterns across all sub-questions — consensus, controversy, gaps.
**Why a separate section?** Pattern-level claims (e.g., "all sources agree on X but disagree on Y") are valuable for the reader's understanding but don't belong inside any single sub-question's synthesis.
**Failure mode:** skipping this step because "the sub-questions cover it". They don't — the cross-cutting view is its own contribution.
### Step 7: Output
Markdown brief by default. DOCX if Q2 = document mode. Honor user preference.
**Why honor preference?** Document mode triggers deeper search budgets + full audit logs. Brief mode is optimized for fast delivery. Different goals → different output shapes.
**Failure mode:** producing DOCX when user wanted brief (overkill) or producing brief when user wanted DOCX (loses citations).
### Step 8: Audit Log
Three-count summary (queries sent / sources received / sources cited) + per-source list with reliability tier.
**Why audit?** Research-pack convention. Lets the reader verify the orchestrator didn't fabricate sources or hide failures.
**Failure mode:** skipping the audit. Audit is what makes the fallback output trustworthy.
## The Three-Count Convention
The research-pack convention requires tracking three integers throughout the fallback workflow:
- **Sent**: queries actually issued (WebSearch + WebFetch + Consensus calls)
- **Received**: results returned from those calls (after filtering)
- **Cited**: sources actually cited in the final output
The relationship `sent >= received >= cited` is always true. When it isn't, something went wrong.
**Why three counts?** They make the orchestrator's search productivity visible. If sent=15, received=3, cited=1, the question was too niche or the search strategy was off. If sent=5, received=20, cited=15, the orchestrator found a rich vein. The reader can interpret the result quality based on the counts.
## Source Discipline
The orchestrator cites **only sources returned by this session's tool calls**. Training knowledge is labeled `[Background — not from search]` and excluded from the three-count.
**Why?** Citations must be verifiable. A "cited" source that wasn't actually retrieved is a fabrication, regardless of how well it matches the orchestrator's training data.
**Failure mode:** inferring a citation from background knowledge + presenting it as if retrieved. This is the highest-severity research-pack violation.
## Retry + Failure Policy
- **On single failure**: wait 3s → retry once → log.
- **After 3 consecutive failures**: stop, alert user, share what was collected.
**Why 3s + single retry?** Most failures are transient (rate limit, network blip). 3s + retry catches them. After 3 in a row, something structural is wrong (API outage, blocked endpoint, query-format issue); halt + escalate.
**Failure mode:** infinite retry loops that consume the session budget. The 3-consecutive-failure stop is the safety valve.
## Reliability Tier Classification
Per source, classify as:
- **Primary** — original source (peer-reviewed paper, government document, company filing, original announcement)
- **Secondary** — derivative reporting (news article summarizing a paper, blog post analyzing a filing)
- **Tertiary** — aggregator or wiki (Wikipedia, news aggregator, opinion piece)
**Why surface tiers?** Reader needs to know which claims rest on primary evidence vs derivative reporting. A consensus claim backed by 5 secondary sources is weaker than the same claim backed by 1 primary source.
**Failure mode:** misclassifying tier to make the audit look better. Honest tiering > polished audit.
## Disagreement Surfacing
When two sources disagree on a sub-question's answer:
- **Name both positions** in the synthesis
- **Cite both sources**
- **State which seems stronger** + why (primary vs secondary, recency, methodology)
- **Don't pick a winner without reasoning**
**Why?** Hiding disagreement misleads the reader. Surfacing it lets them apply their own judgment.
**Failure mode:** averaging two disagreeing sources into a mushy middle that neither source actually supports. This is the synthesis equivalent of fabrication.
## When To Stop Searching (Fallback Mode)
The fallback workflow is **not infinite**. Q4 sets the budget (5 searches for quick scan, 15 for thorough). Stop when:
- Budget exhausted
- All sub-questions have ≥1 high-signal source
- 3-consecutive-failure threshold hit
- User says "stop" or "that's enough"
- Diminishing returns (last 3 searches produced no new high-signal sources)
**Why budget the search?** Open-ended search is the failure mode that turns "research X" into a 30-minute exploration. Budget forces commitment + delivery.
## What Goes Wrong With Fallback
### Fabricated sources
The orchestrator infers a citation from background knowledge. Highest-severity violation. **Prevention:** strict source discipline + three-count tracking makes this auditable.
### Thin results presented as comprehensive
Search returned 2 sources. Orchestrator presents conclusions as if backed by 10. **Prevention:** surface the audit counts. Reader sees `cited: 2` + adjusts confidence.
### Skipping cross-cutting patterns
Per-sub-question synthesis without cross-cutting view. Reader misses the pattern-level insight. **Prevention:** Step 6 is mandatory.
### Skipping audit
Output without the audit section. **Prevention:** Audit is part of the output format, not optional.
### Wrong output format
User asked for brief, got DOCX. Or vice versa. **Prevention:** Q2 captures preference + Step 7 honors it.
### Synthesis without decomposition
Orchestrator searches first, organizes later. Output is unstructured. **Prevention:** Step 1 (decompose) before Step 3 (search) is non-negotiable.
## When To Choose Fallback Over Specialist
The classifier handles this deterministically. But conceptually, fallback is right when:
- No specialist's signal vocabulary fits the question
- User explicitly picked "none of the above" in Q3
- User overrode the routing decision to fallback
- A specialist failed + user opted to retry as fallback
Fallback is **wrong** when:
- A specialist clearly matched (≥2 signals) but the orchestrator ran fallback anyway
- The question is structurally a specialist's domain but used non-canonical phrasing (this is the Q3 case — disambiguate, then route)
## Operational Checklist (Per Fallback Run)
- [ ] Q1 specific enough to decompose (push back if vague)
- [ ] Decomposition produced 3-5 sub-questions
- [ ] Source class chosen per sub-question
- [ ] Sequential 1 q/sec search discipline
- [ ] WebFetch on every cited result
- [ ] Per-sub-question synthesis with citations
- [ ] Cross-cutting patterns section
- [ ] Output format honors Q2
- [ ] Three-count tracked
- [ ] Reliability tier per source
- [ ] Audit log included
- [ ] No fabricated citations
## Citations (7 sources)
1. **Cooper, Hedges, Valentine — "The Handbook of Research Synthesis and Meta-Analysis" (2009, 3rd ed.).** Source for the canonical research-synthesis workflow: question → decomposition → systematic search → extraction → synthesis → reporting. The fallback workflow is a lightweight adaptation of this for AI-orchestrated general research.
2. **Cochrane Collaboration — Handbook for Systematic Reviews of Interventions (current ed.).** Source for the rigor of source classification (primary vs secondary vs tertiary), explicit search protocols, and audit requirements. The three-count + per-source-tier conventions trace to Cochrane practice.
3. **PRISMA 2020 Statement — Page et al., BMJ 2021.** Source for the canonical reporting checklist for research synthesis: searches conducted + sources screened + sources included + sources excluded with reasons. The audit log in fallback mode parallels PRISMA's flow diagram.
4. **Karpathy, Andrej — "On chunking and search in LLMs" (talks 2024-2025).** Source for the principle that decomposition before retrieval beats single-shot retrieval. Sub-questions drive precise queries; whole-question retrieval is too broad. https://karpathy.ai/
5. **Anthropic — Multi-Agent Research System (2024-2025).** Source for the orchestrator-runs-fallback-with-audit pattern. Anthropic's research orchestrator includes explicit audit + source-tier surfacing as trust mechanisms. https://www.anthropic.com/research
6. **Tufte, Edward — "The Visual Display of Quantitative Information" (1983).** Source (by analogy) for the principle of surfacing data integrity to the reader rather than hiding methodology. The three-count + audit log are the textual analogue of Tufte's data-ink ratio: report what you did so the reader can interpret what you found.
7. **NIST — Special Publication 800-53 (Audit Logging guidance).** Source for the operational discipline of immutable, structured audit logs. The `routing_transparency_logger.py` JSON-backed log + the fallback audit section both implement this discipline at different scales.
FILE:references/hybrid_router_architecture.md
# Hybrid Router + Fallback Architecture — When To Delegate, When To Run
This reference answers one decision: **should a research request be delegated to a specialist OR run directly by the orchestrator?** The answer is "either — depending on classification confidence," and the trustability property is **routing transparency**.
## The Core Trade-Off
A purely router-based architecture forces the user to know which specialist applies. A purely monolithic skill produces mediocre output for cases where a specialist would have done better.
The **hybrid** answer: route when confidence is high, run a fallback when it isn't, always surface the decision so the user can correct.
| Architecture | Strength | Weakness |
|---|---|---|
| **Pure router** | Always lands in the right specialist when it knows which one. | Brittle: every miss is a failure (no graceful degradation). |
| **Pure monolith** | Always answers. | Generic answers when a specialist would have done better. |
| **Hybrid (this skill)** | Specialist quality when matched; fallback when not. | Adds a classification step — but it's deterministic + fast. |
## Why Routing Transparency Is Mandatory
The hybrid is **only trustworthy if the user can see the routing decision and override it**. Otherwise the user can't tell when the orchestrator silently downgraded their request to a generic fallback (when a specialist would have done better) or upgraded it to a specialist (when fallback was what they actually wanted).
This is the same property that makes well-designed CI/CD systems trustworthy: the system tells you what stage it's in and lets you intervene. Silent routing is a black box; transparent routing is operable.
## The Three Outcomes (Forcing Frame)
Every invocation produces exactly one of:
1. **Delegation** (classified as specialist-domain, ≥2 signals OR single weak match): hand off to specialist verbatim, return their output, log the delegation.
2. **Fallback execution** (no specialist matched OR Q3 user picked "none of the above"): run the 8-step plan-decompose-search-synthesize-cite workflow.
3. **Clarification request** (classification ambiguous — ≤1 signal across all specialists): ask Q3 (domain disambiguation), then route based on the answer.
Frame this way to refuse the trap of "router silently runs its fallback because the user didn't explicitly ask for a specialist." That's the failure mode the architecture exists to prevent.
## What Makes A Good Routing Decision
A routing decision is good when:
1. **It uses signal-based deterministic logic** (keyword matching, not LLM reasoning over the query)
2. **It commits at high confidence** (≥2 signals for a specialist)
3. **It refuses to commit at low confidence** (1 signal across multiple specialists, or 0 across all → fallback or clarification)
4. **It surfaces the decision** to the user with the matched signals named
5. **It accepts override** without penalty
Bad routing decisions: LLM-only "vibes" classification, silent delegation, refusal to delegate at high-confidence matches, eager delegation at ambiguous matches.
## Forcing-Function Trade-Offs
The orchestrator's job is to make the routing decision **fast** and **visible**, not to do the research itself when a specialist exists. This forces three design constraints:
- **Minimal intake** — 2-4 questions max. Goal is to route, not to interrogate. Specialist handles its own grill-me.
- **Deterministic classifier** — no LLM round-trip. Signal matching is sub-millisecond.
- **Pass-through delegation** — don't pre-answer specialist questions. Their intake is intentional.
When these constraints are violated, the orchestrator slowly becomes a competitor to the specialists rather than their router.
## Sequencing: What Runs When
```
T+0 User invokes /cs:research with their question
T+0 Q1 (research question) — always asked
T+0 Q2 (output preference) — always asked
T+0 Classifier runs (deterministic, sub-millisecond)
T+0 IF score >= 2 OR single specialist with score 1:
Routing transparency: "Routing to X because Y"
Wait 1 turn for override (or 5s timeout)
Delegate verbatim + return specialist output
ELSE:
Q3 (domain disambiguation) — only when ambiguous
IF Q3 picks specialist: delegate
IF Q3 picks "none of the above": Q4 → fallback
T+~5s Specialist output OR fallback workflow complete
```
This sequencing is what keeps the orchestrator fast on the happy path (specialist matched cleanly) while still degrading gracefully (Q3 + Q4 + fallback for the edge cases).
## What Goes Wrong With Each Component
### Silent delegation (no routing transparency)
User asks "what's the buzz about Anthropic," skill silently routes to `pulse`. User never sees the routing. If they wanted general research instead, they have to notice the output came from pulse, then re-invoke. This burns trust + a session.
**Fix:** Routing transparency is mandatory. State decision + accept override.
### LLM-reasoned classification
Skill uses Claude to "decide" which specialist matches. Adds latency, costs tokens, is non-deterministic across invocations (same query → different route). User can't predict what will route where.
**Fix:** Deterministic keyword matching. Predictability is the value.
### Over-eager specialist routing
Skill routes "research Microsoft" to `dossier` based on the word "research". But the user might want general research about Microsoft, not a competitor dossier. The single weak signal isn't strong enough.
**Fix:** Generic "research [topic]" doesn't route. Specific phrases like "dossier on Microsoft" or "background on Microsoft" do.
### Specialist intake pre-answering
Orchestrator collects Q1 + Q2 + Q3 + Q4 + Q5 (passing all into the specialist). Specialist's own grill-me is now redundant; user has to confirm answers twice.
**Fix:** Pass Q1 + Q2 only. Let specialist run its own intake.
### Fallback when specialist would have done better
User asks "review papers on GLP-1 receptor agonists" but skill runs fallback because the classifier missed "review papers" → "literature review" stemming. User gets generic web-search summary instead of structured litreview output.
**Fix:** Signals list must include all reasonable surface forms ("review papers on", "literature review", "lit review", "litreview", etc.).
## When Hybrid Is The Right Architecture
The hybrid pattern is most valuable when:
- Specialists exist + cover non-trivial portion of likely requests
- Specialists have different intake/output shapes (forcing user to know which to use is a tax)
- Generic fallback exists + is acceptable (better than rejecting the request)
- Routing can be made deterministic (predictable classification > LLM "vibes")
When these aren't true, simpler architectures win:
- No specialists yet? Build the monolith.
- One dominant specialist? Just expose it.
- Routing requires deep reasoning over intent? Use LLM classification (accept the cost).
- Fallback would mislead users? Reject instead of falling back.
## Operational Checklist
Before deploying a hybrid router skill:
- [ ] Specialist registry documented with explicit routing signals per specialist
- [ ] Classifier is deterministic (no LLM in the loop)
- [ ] Confidence threshold defined (≥2 signals for commit)
- [ ] Single-weak-match policy defined (1 signal + only one specialist → route)
- [ ] Ambiguity policy defined (≤1 across all → Q3 disambiguation)
- [ ] Routing transparency is mandatory (decision + override surface)
- [ ] Override path tested
- [ ] Fallback workflow specified end-to-end
- [ ] Audit log captures routing decisions + overrides for later review
- [ ] Anti-patterns documented (LLM classification, silent delegation, etc.)
## Citations (8 sources)
1. **Karpathy, Andrej — "LLM OS" talk (2024).** Source for the orchestrator pattern: a smart top-level dispatcher routing to specialized capabilities is more effective than a single monolithic LLM call. Frames the router-with-fallback as a kernel-vs-syscalls analogy. https://karpathy.ai/
2. **Anthropic — Multi-Agent Research System (2024-2025).** Source for the hybrid router-vs-run trade-off in agentic systems. Anthropic's research orchestrator surfaces routing decisions explicitly + accepts user overrides. Practical implementation of the pattern this skill formalizes. https://www.anthropic.com/research
3. **Schaubroeck et al. — "Bounded Confidence in Multi-Agent Systems" (2018).** Source for the academic framing of why bounded-confidence routing (commit only above threshold) outperforms always-route-or-always-defer architectures. Confidence thresholds prevent both over-eager + under-eager commitment.
4. **Google Search Engineering — Query Classification (industry posts).** Source for the deterministic-keyword-matching pattern in production query routers. Google's query classifier uses signal-based deterministic routing for predictability + debuggability, with LLM-reasoned routing only for the residual that signals miss.
5. **Robert Frost, "The Road Not Taken" (1916).** Cited tongue-in-cheek for the routing decision as a one-way door: once delegated, the user sees the specialist's output, not what fallback would have produced. Routing transparency is what gives the user the option to take the other road.
6. **Kubernetes API server — admission controller chain.** Source for the chain-of-responsibility pattern: each handler classifies + either acts or passes to next. Routing transparency in Kubernetes is the auditable admission decision log. Same property in this skill via `routing_transparency_logger.py`.
7. **Tom Preston-Werner — Semantic Versioning specification.** Source for the principle of explicit, predictable contracts over implicit behavior. SemVer's predictability is what made it adoptable; the same property applies to this skill's deterministic routing.
8. **Jeff Hodges — "Notes on Distributed Systems for Young Bloods" (2013).** Source for the principle that explicit + visible system state is what makes operators trust + intervene. Routing transparency is the operator-trust property for skill orchestration. https://www.somethingsimilar.com/2013/01/14/notes-on-distributed-systems-for-young-bloods/
FILE:scripts/classifier.py
#!/usr/bin/env python3
"""
classifier.py — Deterministic SIGNALS-based routing classifier for the research orchestrator.
Given a research question, returns the routing decision (specialist name or "fallback"),
matched signals per specialist, and confidence reasoning.
The SIGNALS map is the post-PR-#657-audit canonical version: verb-noun-paired phrases
that route reliably, with NO bracketed placeholders (those over-trigger on generic
"research [topic]" queries that should fall back instead).
Usage:
python classifier.py --question "What's the literature on PICO for sepsis?"
python classifier.py --question "..." --output json
python classifier.py --sample
"""
import argparse
import json
import sys
SIGNALS = {
"pulse": [
"reddit", "hn", "hacker news", "x.com", "twitter", "buzz",
"sentiment", "trending", "what are people saying",
"what's happening", "the conversation around",
"pulse on", "take the pulse", "current conversation",
],
"grants": [
"nih", "grant", "grants for", "r01", "r21", "k-award", "reporter",
"nosi", "funding", "fda", "study section", "principal investigator",
],
"litreview": [
"literature review", "lit review", "litreview", "pico", "spider",
"systematic review", "review papers on", "research papers on",
"papers about", "meta-analysis",
],
"syllabus": [
"syllabus", "course outline", "curriculum", "reading list",
"for my class", "for my students", "course material",
],
"patent": [
"prior art", "fto", "freedom to operate", "patent",
"patent landscape", "invention", "novelty search",
"patent search", "ip landscape",
],
"dossier": [
"dossier on", "due diligence", "background check",
"prep me for", "competitor research", "investor diligence",
"interview prep", "research my competitor", "background on",
],
}
def classify(question: str) -> dict:
"""
Apply the deterministic routing algorithm:
- score[S] = count of SIGNALS[S] substrings matched (case-insensitive)
- if max(score) >= 2: route to argmax
- elif max(score) == 1 AND only one specialist scored 1: route to that one
- else: route to "fallback"
"""
q = question.lower()
scores = {}
matched = {}
for specialist, phrases in SIGNALS.items():
hits = [p for p in phrases if p in q]
scores[specialist] = len(hits)
if hits:
matched[specialist] = hits
max_score = max(scores.values()) if scores else 0
top = [s for s, sc in scores.items() if sc == max_score and sc > 0]
if max_score >= 2:
route_to = top[0] if len(top) == 1 else _pick_highest_priority(top, scores)
confidence = f"high ({max_score} signals)"
elif max_score == 1:
single_scorers = [s for s, sc in scores.items() if sc == 1]
if len(single_scorers) == 1:
route_to = single_scorers[0]
confidence = "weak (1 signal, single specialist)"
else:
route_to = "fallback"
confidence = "ambiguous (multiple specialists with 1 signal)"
else:
route_to = "fallback"
confidence = "no signals matched"
return {
"route_to": route_to,
"confidence": confidence,
"scores": scores,
"matched_signals": matched,
"question": question,
}
def _pick_highest_priority(candidates: list, scores: dict) -> str:
"""When max(score) is tied across specialists, prefer the one with the
most specific signals (longest matched phrase across SIGNALS map). This is
a tie-breaker; in practice ties at ≥2 are rare."""
return sorted(candidates)[0]
def render_human(result: dict) -> str:
lines = [
f"Question: {result['question']}",
f"Route to: {result['route_to']}",
f"Confidence: {result['confidence']}",
"",
"Per-specialist scores:",
]
for s, sc in sorted(result["scores"].items(), key=lambda kv: -kv[1]):
lines.append(f" {s}: {sc}")
if result["matched_signals"]:
lines.append("")
lines.append("Matched signals:")
for s, phrases in result["matched_signals"].items():
lines.append(f" {s}: {', '.join(repr(p) for p in phrases)}")
if result["route_to"] != "fallback":
lines.append("")
lines.append(
f"Routing transparency: 'Routing to `{result['route_to']}` because "
f"of {result['confidence']}. Override or proceed in 5s.'"
)
else:
lines.append("")
lines.append("Routing transparency: 'No specialist matched. Running fallback.'")
return "\n".join(lines)
def main():
p = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
p.add_argument("--question", help="The research question to classify.")
p.add_argument("--output", choices=["human", "json"], default="human")
p.add_argument("--sample", action="store_true", help="Run with built-in sample question.")
args = p.parse_args()
if args.sample:
args.question = "Can you do a systematic review of PICO frameworks for sepsis treatment? I need a meta-analysis."
if not args.question:
p.error("either --question or --sample is required")
result = classify(args.question)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
if __name__ == "__main__":
main()
FILE:scripts/fallback_decomposer.py
#!/usr/bin/env python3
"""
fallback_decomposer.py — Heuristic question decomposer for the fallback workflow.
Given a research question, returns 3-5 sub-questions using the
what / why / how / who / what's next framework. Deterministic + stdlib only.
The output is a starting point; the orchestrator + user should refine before
search budget is committed.
Usage:
python fallback_decomposer.py --question "How are health systems integrating LLM-based clinical decision support in 2026?"
python fallback_decomposer.py --question "..." --output json
python fallback_decomposer.py --sample
"""
import argparse
import json
import re
FRAMEWORK = [
("what", "What is {topic} — definition, scope, and current state?"),
("why", "Why does {topic} matter now — the forces driving attention or change?"),
("how", "How is {topic} being implemented or applied — methods, players, examples?"),
("who", "Who are the key actors in {topic} — leaders, critics, regulators, adopters?"),
("whats_next", "What's next for {topic} — near-term trajectory, open questions, watchpoints?"),
]
def _extract_topic(question: str) -> str:
"""Strip leading 'research', interrogatives, framing verbs to surface the topic noun phrase."""
q = question.strip().rstrip("?").strip()
q = re.sub(
r"^(can you |could you |please |i need to |i want to |help me )",
"", q, flags=re.IGNORECASE,
).strip()
q = re.sub(
r"^(research |look into |investigate |find me information on |"
r"find information on |do some research on |what do we know about |"
r"what is |what's |how are |how is |how do |why is |why are |"
r"who is |who are |when |where |tell me about )",
"", q, flags=re.IGNORECASE,
).strip()
q = re.sub(r"\s+", " ", q)
return q or question.strip().rstrip("?")
def decompose(question: str, n: int = 5) -> dict:
"""Build 3-5 sub-questions from the framework. n is capped at 5 and floored at 3."""
n = max(3, min(5, n))
topic = _extract_topic(question)
selected = FRAMEWORK[:n]
sub_questions = [
{"label": label, "question": template.format(topic=topic)}
for label, template in selected
]
return {
"question": question,
"extracted_topic": topic,
"sub_question_count": len(sub_questions),
"framework": "what/why/how/who/what's next",
"sub_questions": sub_questions,
"note": ("Starting point only. Refine sub-questions with the user "
"before committing search budget. Drop any that don't fit; "
"rewrite ones that do."),
}
def render_human(result: dict) -> str:
lines = [
f"Question: {result['question']}",
f"Extracted topic: {result['extracted_topic']}",
f"Framework: {result['framework']}",
f"Sub-questions ({result['sub_question_count']}):",
]
for i, sq in enumerate(result["sub_questions"], 1):
lines.append(f" {i}. [{sq['label']}] {sq['question']}")
lines.append("")
lines.append(f"Note: {result['note']}")
return "\n".join(lines)
def main():
p = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
p.add_argument("--question", help="The research question to decompose.")
p.add_argument("--n", type=int, default=5, help="Number of sub-questions (3-5; default 5).")
p.add_argument("--output", choices=["human", "json"], default="human")
p.add_argument("--sample", action="store_true", help="Run with built-in sample question.")
args = p.parse_args()
if args.sample:
args.question = "How are health systems integrating LLM-based clinical decision support in 2026?"
if not args.question:
p.error("either --question or --sample is required")
result = decompose(args.question, n=args.n)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
if __name__ == "__main__":
main()
FILE:scripts/routing_transparency_logger.py
#!/usr/bin/env python3
"""
routing_transparency_logger.py — JSON-backed audit log for the research orchestrator.
Records every routing decision, override, and delegation handoff to a
per-session JSON file at ~/.research_sessions/<session>.json. Stdlib only.
Schema:
{
"session": "<name>",
"created_at": "<iso8601>",
"events": [
{"at": "<iso8601>", "type": "decision", "question": "...", "route_to": "...", "confidence": "...", "matched": {...}},
{"at": "<iso8601>", "type": "override", "from": "...", "to": "...", "reason": "..."},
{"at": "<iso8601>", "type": "delegation", "target": "...", "signals": "..."}
]
}
Usage:
python routing_transparency_logger.py --action record_decision --session demo --question "..." --route-to litreview --confidence "high (2 signals)"
python routing_transparency_logger.py --action record_override --session demo --from litreview --to fallback --reason "wanted general scope"
python routing_transparency_logger.py --action record_delegation --session demo --target litreview --signals "pico,meta-analysis"
python routing_transparency_logger.py --action read --session demo
python routing_transparency_logger.py --sample
"""
import argparse
import json
import os
import sys
from datetime import datetime, timezone
from pathlib import Path
def _now() -> str:
return datetime.now(timezone.utc).isoformat()
def _session_path(session: str) -> Path:
base = Path.home() / ".research_sessions"
base.mkdir(parents=True, exist_ok=True)
safe = "".join(c if c.isalnum() or c in ("-", "_") else "_" for c in session)
return base / f"{safe}.json"
def _load(session: str) -> dict:
path = _session_path(session)
if not path.exists():
return {"session": session, "created_at": _now(), "events": []}
return json.loads(path.read_text(encoding="utf-8"))
def _save(session: str, data: dict) -> Path:
path = _session_path(session)
path.write_text(json.dumps(data, indent=2), encoding="utf-8")
return path
def record_decision(session: str, question: str, route_to: str, confidence: str,
matched: dict | None = None) -> dict:
data = _load(session)
event = {
"at": _now(),
"type": "decision",
"question": question,
"route_to": route_to,
"confidence": confidence,
"matched": matched or {},
}
data["events"].append(event)
_save(session, data)
return event
def record_override(session: str, from_target: str, to_target: str, reason: str) -> dict:
data = _load(session)
event = {
"at": _now(),
"type": "override",
"from": from_target,
"to": to_target,
"reason": reason,
}
data["events"].append(event)
_save(session, data)
return event
def record_delegation(session: str, target: str, signals: str) -> dict:
data = _load(session)
event = {
"at": _now(),
"type": "delegation",
"target": target,
"signals": signals,
}
data["events"].append(event)
_save(session, data)
return event
def read(session: str) -> dict:
return _load(session)
def render_human(result: dict) -> str:
if "events" in result:
lines = [
f"Session: {result['session']}",
f"Created: {result['created_at']}",
f"Events ({len(result['events'])}):",
]
for e in result["events"]:
t = e.get("type")
if t == "decision":
lines.append(f" [{e['at']}] decision → {e['route_to']} ({e['confidence']})")
elif t == "override":
lines.append(f" [{e['at']}] override {e['from']} → {e['to']} ({e['reason']})")
elif t == "delegation":
lines.append(f" [{e['at']}] delegation → {e['target']} (signals: {e['signals']})")
else:
lines.append(f" [{e['at']}] {t}: {e}")
return "\n".join(lines)
return json.dumps(result, indent=2)
def main():
p = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
p.add_argument("--action",
choices=["record_decision", "record_override", "record_delegation", "read"],
help="What to do.")
p.add_argument("--session", help="Session name (used as filename stem).")
p.add_argument("--question", help="(record_decision) The classified question.")
p.add_argument("--route-to", dest="route_to", help="(record_decision) Routing target.")
p.add_argument("--confidence", help="(record_decision) Confidence string.")
p.add_argument("--matched", help="(record_decision) Matched signals (JSON).")
p.add_argument("--from", dest="from_target", help="(record_override) Previous target.")
p.add_argument("--to", dest="to_target", help="(record_override) New target.")
p.add_argument("--reason", help="(record_override) Why user overrode.")
p.add_argument("--target", help="(record_delegation) Specialist target.")
p.add_argument("--signals", help="(record_delegation) Signals that matched.")
p.add_argument("--output", choices=["human", "json"], default="human")
p.add_argument("--sample", action="store_true", help="Run a built-in 4-event sample sequence.")
args = p.parse_args()
if args.sample:
session = "sample"
path = _session_path(session)
if path.exists():
path.unlink()
record_decision(session,
"Can you review the literature on PICO for sepsis?",
"litreview",
"high (2 signals)",
{"litreview": ["pico", "literature"]})
record_delegation(session, "litreview", "pico,literature")
record_decision(session,
"What's the buzz about Anthropic on HN?",
"pulse",
"high (2 signals)",
{"pulse": ["hn", "buzz"]})
record_override(session, "pulse", "fallback", "wanted general scope")
result = read(session)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return
if not args.action:
p.error("--action is required (unless --sample)")
if not args.session:
p.error("--session is required")
if args.action == "record_decision":
if not (args.question and args.route_to and args.confidence):
p.error("record_decision requires --question, --route-to, --confidence")
matched = json.loads(args.matched) if args.matched else None
out = record_decision(args.session, args.question, args.route_to, args.confidence, matched)
elif args.action == "record_override":
if not (args.from_target and args.to_target and args.reason):
p.error("record_override requires --from, --to, --reason")
out = record_override(args.session, args.from_target, args.to_target, args.reason)
elif args.action == "record_delegation":
if not (args.target and args.signals):
p.error("record_delegation requires --target, --signals")
out = record_delegation(args.session, args.target, args.signals)
elif args.action == "read":
out = read(args.session)
else:
p.error(f"unknown action {args.action}")
if args.output == "json":
print(json.dumps(out, indent=2))
else:
print(render_human(out))
if __name__ == "__main__":
main()